Matchboxes that learn to play
In November 1963 Donald Michie described in The Computer Journal MENACE, a machine that learned noughts and crosses by trial and error: one matchbox for each of 287 essentially distinct positions, with coloured beads in it for the possible moves. After a draw each open box got one extra bead, after a win three, and after a defeat one was taken away. Michie had described the machine and its first tournament in 1961.
Why it matters
Reinforcement learning became a device that could be built on a table and counted: a move that led to a win comes up more often, one that led to a loss less often. The 1963 paper moved the same scheme onto a Pegasus 2 computer and turned it into an experiment with three parameters, in which reinforcement weakens from the end of the game back towards its start.
The Pegasus 2 simulation, written with D. J. M. Martin of Ferranti, played both sides at about a game a second; when both learned, both were playing near-expert games after a few hundred. The model's parameters are A, B and a decay factor D; the next paper of the series was to describe how learning depends on them. Michie explains his machine in the terms of experimental psychology and cites no work but his own paper of 1961. What the record does not claim: the length of the first tournament - the figure of 220 games in common retellings does not occur in the paper; the page names no affiliation for the author. The research report behind the candidate said that scientific papers call the machine "Machine Educable"; the paper itself says "Matchbox Educable". The publisher's page sits behind a Cloudflare check, so the text was read in the copy on Rodney Brooks's MIT pages, and the month and pages come from Crossref.