Back to timeline

Research · August 1988

Temporal-difference learning

Sutton showed that a prediction can be corrected against the next prediction rather than the final outcome, and proved it works.

Why it matters

An agent stopped waiting for the end of an episode: it learns from every step instead of from the finish.

The paper formalises the TD(lambda) family, where lambda governs how far back a correction propagates. Sutton proves convergence for the linear case and shows TD beats Monte Carlo on problems with Markov structure. This is the core TD-Gammon was built on four years later, and deep reinforcement learning thirty years later.

Event record

Event date
August 1988
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0157

Machine Learning volume 3, August 1988.

Sources

Related events

Records that link to this one