Whisper
OpenAI released a model trained on 680 thousand hours of multilingual web audio that transcribes 99 languages and translates into English.
Why it matters
Speech recognition became free, multilingual and robust to noise, and with it the line begun in 1971 was substantially closed.
The bet was on diversity of data rather than cleanliness: web recordings with imperfect subtitles instead of studio corpora. That is why the model holds up on accents, noise and background music, exactly where systems trained on Switchboard broke. Open weights meant transcription no longer required a cloud service or a subscription. Fifty-one years separate this from the ARPA programme. How many languages. The paper accepted at ICML 2023 counts 117,000 hours of non-English audio covering 96 other languages, that is 97 with English; the string 99 does not occur in it. The model card in OpenAI’s repository says 98 other languages. The 99 in the summary comes from OpenAI’s own release materials; the three counts differ in what they count, and none of them is an independent measurement.