A hundred and one classes instead of five
Fei-Fei Li, Rob Fergus and Pietro Perona assembled a set of 101 object categories plus a background class, and presented a Bayesian method that learns from a handful of examples. Object class recognition had until then been tested mostly on one or two categories, on four by Weber and colleagues and on six by Fergus and colleagues; the authors call their dataset fifteen times larger than anything before it.
Why it matters
Object recognition acquired a proving ground on which one method's advantage over another was still visible while the choice of favourite classes no longer decided anything. The field spent the following years on this set, and its exhaustion is what opened the way first to PASCAL VOC and then to ImageNet.
The category names were generated by flipping through the Webster Collegiate Dictionary and taking words that came with a drawing; images were gathered with Google Image Search and two graduate students unconnected with the experiment weeded out what did not belong. The background class was collected on the keyword "things". The counts differ between the paper and the dataset page, and the record names both. The paper's figure caption says each category contains between 45 and 400 images. Caltech's own record of the released dataset says about 40 to 800 images per category, most categories about 50, each image roughly 300 by 200 pixels, collected in September 2003. The results the paper reports: with one training example the Bayesian method averages 71 percent and the best results are over 80 percent; at three examples both Bayesian variants are near 75 percent while maximum likelihood is still at chance, catching up only at fifteen. The record does not claim "9,144 images" and does not claim "16.0 percent accuracy at 15 training images". Neither figure is in the paper; what it measures is a class against background clutter, not a choice of one category out of 101.