Deep networks in production recognition
A Microsoft group replaced the acoustic model with a deep network on a real transcription task and cut error from 27.4 to 18.5 percent.
Why it matters
A third of the errors disappeared from one substitution, and that was enough to move the whole industry to deep networks within two years.
A one-third reduction was unheard of in a field where progress had been measured in percentage points for years. The network dropped in where the Gaussian mixtures had been, inside the existing hidden Markov scheme, so migration needed no rewrite. Google, IBM and Baidu did the same during 2012 and 2013. It is the first case of deep learning winning in production rather than in a paper.