Back to timeline

Research · June 2020

wav2vec 2.0

The model learned to represent speech from unlabelled recordings, then reached good accuracy with only ten minutes of transcribed audio.

Why it matters

The main obstacle for most of the world's languages disappeared: recognition stopped requiring thousands of hours of manual transcription.

Classical systems needed hundreds or thousands of hours of transcribed speech, which is why they existed for about a dozen languages. Here the model first learns to predict masked parts of its own audio without labels, and labelled data is needed only for brief fine-tuning. The practical consequence is that recognition became reachable for languages nobody ever collected corpora for.

Event record

Event date
June 2020
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0240

Preprint of 20 June 2020; presented at NeurIPS 2020.

Sources

Related events

Records that link to this one