Q-learning
Watkins described an algorithm that learns the value of an action in a state directly from experience, needing no model of the environment.
Why it matters
An agent learned optimal behaviour without knowing the rules of its world and without a teacher, only from the consequences of its own actions.
The key property is being off-policy: the agent may act however it likes while learning optimal behaviour. Watkins gave a sketch of a convergence proof, completed with Dayan in 1992. The algorithm's simplicity made it the standard entry point to reinforcement learning, and in 2013 the deep version, DQN, produced the first results on Atari games.