Back to timeline

Research · April 29, 2002

Object recognition as translation

In the ECCV 2002 volume, published on 29 April 2002, Pinar Duygulu, Kobus Barnard, Nando de Freitas and David Forsyth of Berkeley and UBC cast recognition as translation: image regions are the words of one language, keywords those of another. On 4,500 Corel images with 371 words, EM learned a lexicon between 500 region types and the words.

Why it matters

Images and text were trained together as an aligned bilingual corpus, by the same method statistical translation used to learn lexicons from parallel texts. The machine learned to name parts of an image from captions of whole images alone, with no region labels: the problem later taken up by captioning models and CLIP.

Images were segmented by Normalized Cuts into 5-10 regions, each described by 33 features and quantized by k-means into 500 "blobs"; each image had 4-5 keywords. The model is model 2 of Brown et al. Only 80 of the 371 words could be predicted; on 500 held-out images the best words reached a recall of 0.83 (sky) and a precision of 1.00 (petals). Words the features cannot tell apart (train and locomotive) were clustered. Barnard's page names the ECVision best paper award in cognitive computer vision. The author copy has no text layer; the figures were read from rendered pages.

Event record

Event date
April 29, 2002
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0799

The publication day of LNCS volume 2353 by Springer; the dates of the Copenhagen conference itself are not in the sources read.

Sources

Related events

Records that link to this one