Actor and critic
Barto, Sutton and Anderson split learning into two elements: one chooses an action, the other judges how much better a state is than expected.
Why it matters
Reinforcement learning acquired an architecture that turns a rare reward into a signal available at every step.
The task is balancing a pole on a cart, where the failure signal arrives only when it falls. The critic builds its own value function over states and gives the actor an evaluation at every step. It is the direct ancestor of temporal-difference learning, which Sutton formalised in 1988, and of actor-critic schemes in modern algorithms.