Eighty million images at 32 × 32
In November 2008 Antonio Torralba, Rob Fergus and William Freeman of MIT described in IEEE PAMI a set of 79,302,017 images of 32 × 32 pixels. Over eight months seven search engines returned 97,245,098 pictures for queries on 75,846 WordNet nouns; after cleaning, images remained for 75,062 words. Each image’s only label is the query that found it.
Why it matters
The work tested whether the size of the data could stand in for the complexity of the model: the simplest nearest-neighbour search over tens of millions of tiny pictures recognised scenes and, for people, came near dedicated detectors of the Viola-Jones kind. The label noise remained: CIFAR-10 in 2009 would hand-label a small part of the set, and ImageNet would set clean labels and full resolution against it.
IEEE Transactions on Pattern Analysis and Machine Intelligence 30(11), pp. 1958-1970. The search engines: Altavista, Ask, Flickr, Cydral, Google, Picsearch and Webshots; at most 3000 images a word; about 1% of words returned nothing. The whole set occupies 760 GB on one disk. The resolution was chosen from psychophysics: in colour at 32 × 32 people assign a scene to one of 15 categories correctly more than 80% of the time, while in greyscale they need about 64 × 64. The share of correct labels drops after the hundredth image for a query and settles near 44%. MIT’s technical report of 23 April 2007 described the same work with different figures: 10⁸ images and 70,399 nouns. The copy of the paper read is the authors’ manuscript without volume numbers; the record takes its figures from it.