Object recognition as translation
In the ECCV 2002 volume, published on 29 April 2002, Pinar Duygulu, Kobus Barnard, Nando de Freitas and David Forsyth of Berkeley and UBC cast recognition as translation: image regions are the words of one language, keywords those of another. On 4,500 Corel images with 371 words, EM learned a lexicon between 500 region types and the words.
Why it matters
Images and text were trained together as an aligned bilingual corpus, by the same method statistical translation used to learn lexicons from parallel texts. The machine learned to name parts of an image from captions of whole images alone, with no region labels: the problem later taken up by captioning models and CLIP.
Images were segmented by Normalized Cuts into 5-10 regions, each described by 33 features and quantized by k-means into 500 "blobs"; each image had 4-5 keywords. The model is model 2 of Brown et al. Only 80 of the 371 words could be predicted; on 500 held-out images the best words reached a recall of 0.83 (sky) and a precision of 1.00 (petals). Words the features cannot tell apart (train and locomotive) were clustered. Barnard's page names the ECVision best paper award in cognitive computer vision. The author copy has no text layer; the figures were read from rendered pages.