A Stochastic Approximation Method
Robbins and Monro showed how to find the root of a function from noisy measurements by taking small steps of decreasing size.
Why it matters
The step conditions derived here are still the ones a learning-rate schedule is chosen against.
The paper is framed as mathematical statistics and does not mention machine learning. It sets the requirement on the step sequence: the sum diverges, the sum of squares converges. Convergence proofs for gradient descent on noisy estimates rest on this result.