MS COCO: objects in context
On 1 May 2014 Tsung-Yi Lin and colleagues from Cornell, Caltech, Brown, UC Irvine and Microsoft Research released Microsoft COCO, 328,000 images of everyday scenes with 2.5 million labelled instances of 91 object categories, each with its own segmentation mask.
Why it matters
The set deliberately took not iconic photos of one large object but scenes with many objects in their natural surroundings, and labelled them with masks rather than boxes. Image captioning, as in Show and Tell, was among what was trained and tested on COCO.
The annotation was done by workers on Amazon Mechanical Turk through purpose-built interfaces; by the first preprint version, over 32,000 worker hours, 11,972 of them on one stage and 6,715 on another. The figure of 70,000 hours is not in the first version.