AudioSet: a vocabulary of sounds and two million clips
In March 2017 Jort Gemmeke and colleagues at Google presented AudioSet, a hierarchical ontology of sound events and 2,084,320 human-labelled 10-second clips from YouTube videos; labellers checked whether a given class of sound was present in a segment.
Why it matters
Recognising sounds other than speech, such as instruments, animals and household noise, gained a shared large dataset, as vision had gained ImageNet. Later music generation models built their evaluation sets from its clips.
The dataset page on 8 March 2017: an ontology of 632 classes, 527 classes labelled, 2.1 million videos, 5.8 thousand hours of audio. The paper's abstract says 635 classes; the two figures come from different documents of the same month. The full paper was not read: the PDF on Google's site downloads instead of opening.