Schlesinger: unsupervised learning as repeated learning
In Kibernetika in 1968 (No. 2, pp. 81-88) M. I. Schlesinger described a self-learning algorithm for pattern recognition: each iteration first recognises the patterns, computing the a posteriori probabilities of the classes, and then solves the learning problem with those probabilities. The paper proves that the likelihood strictly increases, and that the limiting parameter values are maximum likelihood estimates.
Why it matters
Learning without a teacher, that is without class labels, was reduced to alternating two known procedures, recognition and learning, with a proven rise in likelihood. Nine years later statistics would call this scheme the EM algorithm; Schlesinger himself, in a book of 2002, places his paper before the work of Dempster, Laird and Rubin of 1977. This is an editorial assessment resting on his own account.
The paper was read in a scan of the original (received by the editors on 20 April 1967) and in the full English translation (Cybernetics, vol. 4, no. 2, pp. 66-71). The translation distinguishes learning, where each pattern is given with its class, from self-organization, where it is not. The paper cites Glushkov (1962), Rosenblatt, Robbins's empirical Bayes approach and Cooper and Cooper (1964). On EM: the bibliographical notes of Schlesinger and Hlaváč's book of 2002, p. 274 (read only through Google Books search snippets), say the task was solved by 'Schlesinger in the year 1968' and later by others (Dempster et al., 1977), and that the procedure is known today as EM. This is the author's own account; the record has no independent source on priority. What the record does not claim: where Schlesinger worked in 1968, since neither the original nor the translation prints an affiliation; that Dempster, Laird and Rubin knew the paper.