Batch normalisation
Ioffe and Szegedy proposed normalising each layer's inputs by the statistics of the current batch, and training sped up several times over.
Why it matters
Deep networks began training predictably: careful initialisation and a tiny learning rate stopped being necessary.
The distribution of each layer's inputs shifts during training because the weights of earlier layers change, so every layer chases a moving target. Normalisation fixes the scale. The effect was strong enough that the method entered practically every architecture. Why exactly it works is still debated; the authors' original hypothesis was later disputed.