Back to timeline

Research · May 1989

Q-learning

Watkins described an algorithm that learns the value of an action in a state directly from experience, needing no model of the environment.

Why it matters

An agent learned optimal behaviour without knowing the rules of its world and without a teacher, only from the consequences of its own actions.

The key property is being off-policy: the agent may act however it likes while learning optimal behaviour. Watkins gave a sketch of a convergence proof, completed with Dayan in 1992. The algorithm's simplicity made it the standard entry point to reinforcement learning, and in 2013 the deep version, DQN, produced the first results on Atari games.

Event record

Event date
May 1989
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0160

The thesis was submitted to the University of Cambridge in May 1989.

Sources

Related events

Records that link to this one