wav2vec 2.0
The model learned to represent speech from unlabelled recordings, then reached good accuracy with only ten minutes of transcribed audio.
Why it matters
The main obstacle for most of the world's languages disappeared: recognition stopped requiring thousands of hours of manual transcription.
Classical systems needed hundreds or thousands of hours of transcribed speech, which is why they existed for about a dozen languages. Here the model first learns to predict masked parts of its own audio without labels, and labelled data is needed only for brief fine-tuning. The practical consequence is that recognition became reachable for languages nobody ever collected corpora for.