TD-Gammon
Tesauro's program learned to play backgammon at the level of the best humans, purely by playing itself, with no game database and no hints.
Why it matters
A network combined with temporal-difference learning reached expert level, the first convincing evidence that the pair works.
The network evaluated a position and temporal-difference learning corrected that estimate against the next move. After a million self-play games it matched the world elite. Some of its opening moves contradicted established theory, and players later revised the theory in its favour. This is the same recipe that produced AlphaGo twenty-four years later.