CIFAR-10: tiny images with labels
On 8 April 2009 Alex Krizhevsky of the University of Toronto issued a technical report describing CIFAR-10, a labelled part of the 80 million tiny images collection at 32 by 32 pixels: 10 classes of 6,000 images, with 5,000 from each class in the training part. The same report has CIFAR-100, 100 classes of 600.
Why it matters
The tiny images collected from the web had no reliable labels and so were unusable for object recognition. CIFAR-10 gave a small but honestly labelled set on which methods could be compared quickly; dropout in 2012 and diffusion models in 2020 were among those tested on it.
The labels were set by students paid a fixed sum per hour, so that there was no incentive to rush; every label was checked by the authors themselves. The set is named after the Canadian Institute for Advanced Research (CIFAR), which funded the project. The totals of 60,000, 50,000 for training and 10,000 for testing come from the dataset page, not the report; the page names Krizhevsky, Vinod Nair and Geoffrey Hinton as its creators. The 80 million image collection itself was gathered by groups at MIT and NYU.