Back to timeline

Benchmark · 2004

A hundred and one classes instead of five

Fei-Fei Li, Rob Fergus and Pietro Perona assembled a set of 101 object categories plus a background class, and presented a Bayesian method that learns from a handful of examples. Object class recognition had until then been tested mostly on one or two categories, on four by Weber and colleagues and on six by Fergus and colleagues; the authors call their dataset fifteen times larger than anything before it.

Why it matters

Object recognition acquired a proving ground on which one method's advantage over another was still visible while the choice of favourite classes no longer decided anything. The field spent the following years on this set, and its exhaustion is what opened the way first to PASCAL VOC and then to ImageNet.

The category names were generated by flipping through the Webster Collegiate Dictionary and taking words that came with a drawing; images were gathered with Google Image Search and two graduate students unconnected with the experiment weeded out what did not belong. The background class was collected on the keyword "things". The counts differ between the paper and the dataset page, and the record names both. The paper's figure caption says each category contains between 45 and 400 images. Caltech's own record of the released dataset says about 40 to 800 images per category, most categories about 50, each image roughly 300 by 200 pixels, collected in September 2003. The results the paper reports: with one training example the Bayesian method averages 71 percent and the best results are over 80 percent; at three examples both Bayesian variants are near 75 percent while maximum likelihood is still at chance, catching up only at fifteen. The record does not claim "9,144 images" and does not claim "16.0 percent accuracy at 15 training images". Neither figure is in the paper; what it measures is a class against background clutter, not a choice of one category out of 101.

Event record

Event date
2004
Timeline date
Event date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0497

Presented at the CVPR 2004 Workshop on Generative-Model Based Vision in Washington, DC. Neither the paper nor the Crossref record names a month, so the record keeps the year. The images themselves were collected in September 2003, per Caltech's own record.

Sources

Related events

Records that link to this one