Back to timeline

Research · 1983

Actor and critic

Barto, Sutton and Anderson split learning into two elements: one chooses an action, the other judges how much better a state is than expected.

Why it matters

Reinforcement learning acquired an architecture that turns a rare reward into a signal available at every step.

The task is balancing a pole on a cart, where the failure signal arrives only when it falls. The critic builds its own value function over states and gives the actor an evaluation at every step. It is the direct ancestor of temporal-difference learning, which Sutton formalised in 1988, and of actor-critic schemes in modern algorithms.

Event record

Event date
1983
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0143

IEEE Transactions on Systems, Man, and Cybernetics volume SMC-13, 1983.

Sources

Related events

Records that link to this one