Deep networks in speech recognition
Mohamed, Dahl and Hinton replaced the Gaussian mixtures in the acoustic model with a deep network and beat the best published results on TIMIT.
Why it matters
The first field where deep learning beat the established method, and it happened in speech rather than vision, three years before AlexNet.
Gaussian mixtures had been the standard of acoustic modelling since the 1980s. The network gave a better representation of the same signal and slotted into the existing hidden Markov scheme, so industry could migrate gradually. By 2011 Microsoft, Google and IBM had moved production systems onto it. Speech moved faster than vision precisely because it already had a shared ruler.