Back to timeline

Research · November 17, 2014

Show and Tell: a caption for an image

On 17 November 2014 Oriol Vinyals and colleagues at Google described a network that composes a sentence about an image: a convolutional network encodes the picture and an LSTM writes the description. On Pascal the BLEU score was 59 against a previous best of 25 and human performance around 69.

Why it matters

Describing an image became a single model trained end to end on image-sentence pairs, just as translation had become sentence to sentence. The line of models joining vision and language starts here.

By the first version of the preprint, BLEU on Flickr30k rose from 55 to 66 and on SBU from 19 to 27. The MSCOCO training split is 82,783 images. The record does not claim the 27.7 on COCO: it is not in the first version.

Event record

Event date
November 17, 2014
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0768

The day the first version of the preprint was submitted; the figures were read in it.

Sources

Related events

Records that link to this one