Deep Speech
Baidu showed a system in which one network turns sound into text with no phonemes, no pronunciation dictionary and no hidden Markov models.
Why it matters
A pipeline of four separately tuned parts was replaced by a single learned function.
The classical scheme required an acoustic model, a pronunciation dictionary, a language model and a decoder, each tuned separately. Deep Speech trains end to end from spectrogram to characters. The system proved more robust to noise than commercial products. Whisper used the same approach eight years later, adding scale of data.