Back to timeline

Benchmark · 1992

TD-Gammon

Tesauro's program learned to play backgammon at the level of the best humans, purely by playing itself, with no game database and no hints.

Why it matters

A network combined with temporal-difference learning reached expert level, the first convincing evidence that the pair works.

The network evaluated a position and temporal-difference learning corrected that estimate against the next move. After a million self-play games it matched the world elite. Some of its opening moves contradicted established theory, and players later revised the theory in its favour. This is the same recipe that produced AlphaGo twenty-four years later.

Event record

Event date
1992
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0165

Machine Learning volume 8, 1992.

Sources

Related events

Records that link to this one