Applied Optimal Control
Bryson and Ho showed how to get the gradient of a cost with respect to the control at every stage of a multistage system: Lagrange multipliers, which they also call influence functions, are computed backward from the final stage.
Why it matters
The computational core of back-propagation existed in control theory seventeen years before the 1986 paper.
The book solves optimal control for a multistage system in which each stage depends on the previous one, and notes that such systems gained special significance because digital computers solve continuous problems in stages. Working out how each change of control propagates forward "would be tedious", so the multipliers are chosen to carry the effect backward instead; the book calls the resulting equations adjoint to the perturbation equations. It is structurally the same problem as training a multilayer network. The connection between the two fields went unnoticed until Werbos carried the device across directly.