Back to timeline

Research · February 2003

Words and pictures: a joint model

In February 2003 the Journal of Machine Learning Research printed "Matching Words and Pictures" by Barnard, Duygulu, Forsyth, de Freitas, Blei and Jordan. Several models of the joint distribution of words and image regions predicted captions for whole images and named single regions; the data were images from 160 Corel CDs of 100 each.

Why it matters

Auto-annotation and region naming were measured on the same data for a whole family of models: hierarchical aspect, translation, and a multimodal extension of LDA. Images became documents with words, open to modelling by the same means as text.

Samples of 80 of the 160 Corel CDs were split 75% for training and 25% held out, and the remaining CDs gave a harder "novel" set. Region features were quantized into 500 clusters; the vocabulary was around 155 words depending on the test set. The authors explain a typical gain of 0.160 in the normalised score as reaching about 80% of the words in about 40 guesses instead of 70 for the empirical distribution. The translation model is taken from the machine translation of Brown et al. The record does not claim "16,000 images": the text has 160 CDs of 100 images and never states the sum.

Event record

Event date
February 2003
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0800

The month of the issue by JMLR ("Published 2/03"); the paper was submitted in April 2002.

Sources

Related events