Temporal-difference learning
Sutton showed that a prediction can be corrected against the next prediction rather than the final outcome, and proved it works.
Why it matters
An agent stopped waiting for the end of an episode: it learns from every step instead of from the finish.
The paper formalises the TD(lambda) family, where lambda governs how far back a correction propagates. Sutton proves convergence for the linear case and shows TD beats Monte Carlo on problems with Markov structure. This is the core TD-Gammon was built on four years later, and deep reinforcement learning thirty years later.