Show and Tell: a caption for an image
On 17 November 2014 Oriol Vinyals and colleagues at Google described a network that composes a sentence about an image: a convolutional network encodes the picture and an LSTM writes the description. On Pascal the BLEU score was 59 against a previous best of 25 and human performance around 69.
Why it matters
Describing an image became a single model trained end to end on image-sentence pairs, just as translation had become sentence to sentence. The line of models joining vision and language starts here.
By the first version of the preprint, BLEU on Flickr30k rose from 55 to 66 and on SBU from 19 to 27. The MSCOCO training split is 82,783 images. The record does not claim the 27.7 on COCO: it is not in the first version.