The Adam optimizer
Kingma and Ba described a method that picks a learning rate for each parameter separately from estimates of the first and second moments of the gradient.
Why it matters
Tuning training stopped being a craft: Adam worked acceptably almost everywhere without a parameter search.
Before Adam the learning rate was a leading reason a network failed to train, and choosing it took experience. Adam adapts the step for each weight and corrects the bias in its estimates early in training. It is not the best method in every case but it is a reliable default. It is among the most cited papers in machine learning precisely for being practical.