A Theory of Adaptive Pattern Classifiers
Amari described training a classifier by stochastic gradient descent and proved convergence even for non-separable distributions.
Why it matters
Gradient descent on noisy estimates acquired a theory just as the field was turning away from neural networks.
The paper also treats the multilayer case with hidden units trained by gradient. It appeared two years before Perceptrons and had almost no effect on the English-language discussion of the time. Amari later developed information geometry and the natural gradient.