Automatic recognition of spoken digits
In November 1952 K. H. Davis, R. Biddulph and S. Balashek of Bell Labs described a circuit that automatically recognises ten telephone-quality digits spoken at normal speech rates by a single individual, with an accuracy varying between 97 and 99 percent. The spectrum was split into two bands, one below and one above 900 cycles per second, axis-crossing counts were made in each, and the result was compared against ten built-in standards, one per digit, with the best match selected.
Why it matters
For the first time a machine recognised spoken words to a measured accuracy rather than in principle. In the same breath the paper named the limit that would hold for two decades: the circuit had to be adjusted to one particular talker, and without that adjustment it did not perform equally well across a series of voices.
The full text is paywalled and the record rests on the paper's abstract, where both the accuracy and the method are stated outright. The abstract says that after some preliminary analysis of the speech of any individual the circuit can be adjusted to deliver a similar accuracy on that person, and that in its present configuration it is not capable of performing equally well on a series of talkers without recourse to such adjustment. What the record does not claim: that the two measured bands gave the first and second formants, that the standards were resistive matching circuits, or that the apparatus was called Audrey. None of those three appears in the abstract; the research report behind this batch asserted all three and none entered. The abstract also names no count of talkers tested, only that it was one individual.