Deep networks on graphics processors
Raina, Madhavan and Ng showed that a model with a hundred million parameters trains on graphics processors in a day instead of several weeks.
Why it matters
The computational obstacle that kept deep networks small disappeared, and that opened the road to 2012.
The speedup was roughly seventyfold against a central processor. The authors state plainly that model scale is from now on limited by the hardware budget rather than by the algorithm. This is the link between CUDA in 2007 and AlexNet in 2012: first it was shown to be possible, then it was shown to win.