Reinforcement Learning: An Introduction
Sutton and Barto gathered scattered methods into one discipline with shared notation, problems and theory.
Why it matters
A direction that had existed for twenty years as a set of tricks became a subject with a textbook, and so became teachable.
The book introduces a common language: Markov decision process, policy, value function, the trade-off between exploration and exploitation. It also separates honestly what is proved from what only works empirically. The authors put the full text online. The 2018 second edition added deep methods; this is the book the generation that built AlphaGo learned from.