Back to timeline

Research · December 2014

Deep Speech

Baidu showed a system in which one network turns sound into text with no phonemes, no pronunciation dictionary and no hidden Markov models.

Why it matters

A pipeline of four separately tuned parts was replaced by a single learned function.

The classical scheme required an acoustic model, a pronunciation dictionary, a language model and a decoder, each tuned separately. Deep Speech trains end to end from spectrogram to characters. The system proved more robust to noise than commercial products. Whisper used the same approach eight years later, adding scale of data.

Event record

Event date
December 2014
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0227

Preprint of 17 December 2014.

Sources

Related events

Records that link to this one