LibriSpeech: a thousand hours of audiobooks
In April 2015 Vassil Panayotov, Daniel Povey and colleagues at Johns Hopkins University published LibriSpeech, about 1,000 hours of read English at 16 kHz aligned with its text, taken from public-domain LibriVox audiobooks.
Why it matters
An open corpus of this size for training and testing speech recognition appeared for the first time, with Kaldi scripts and language models. Its "clean" and "other" test sets became the common yardstick of recognition systems for years.
The training subsets are train-clean-100 (100.6 hours), train-clean-360 (363.6) and train-other-500 (496.7); speakers were split into "clean" and "other" by the error rate of a model trained on WSJ. Models trained on LibriSpeech had lower error on the Wall Street Journal test sets than models trained on WSJ itself. The record gives no WER figures from the paper's table: the extracted table text is broken.