Shoebox does arithmetic by voice
In 1961 the IBM engineer William C. Dersch introduced Shoebox at a press conference: a machine the size of a shoebox that recognised 16 words, the digits zero through nine and six commands among them plus, minus, total and subtotal. Dersch said "Seven plus three plus six plus nine plus five. Subtotal", the machine typed out each digit and command and returned 30. The circuit needed 31 transistors, fewer than two per word, where earlier machines, by IBM's account, needed as many as 200 transistors per word.
Why it matters
Speech recognition left the laboratory for the first time as a thing you could carry and show: not a rack of circuits but a box driving an ordinary adding machine. And for the first time it was built deliberately unlike human hearing - Dersch used features the ear does not discern, because he thought copying people was the wrong way to build a machine.
Both sources are IBM's own, so every figure here is a maker's claim about its own product and there is no independent measurement; confidence is recorded as medium for that reason. The method was this: the machine sorted sounds into 'machine vowels', voiced sounds originating in the throat such as a, o and m, and 'machine consonants', frictional sounds made by air passing tongue and teeth such as f, v and th, with the unvoiced sounds further graded weak or strong to separate similar words like one and nine. Because it keyed on patterns that occur only in the human voice, ambient room noise barely troubled it. What the record does not claim: any date finer than the year, and that Shoebox was the world's first speech-recognition system, which IBM's headline says but which the Bell Labs circuit of 1952, with its measured accuracy, precedes by nine years. The research report behind this batch gave the fricatives as f, s and th; the page says f, v and th.