Back to timeline

Research · May 3, 2015

VQA: questions put to an image

On 3 May 2015 Stanislaw Antol, Devi Parikh and colleagues proposed the task of open-ended answers to questions about images, with a dataset for it. The first version held 123,285 MS COCO images, 215,150 questions and 430,920 answers, and 10,000 abstract scenes with 30,000 questions.

Why it matters

Understanding an image became measurable: a few-word answer can be checked automatically against human ones, while a caption is hard to score. The task demands recognition, counting and common sense together.

At submission the dataset was incomplete: answers had been collected for only 50,000 real images. There are two forms, open-ended and multiple choice. A random answer scores under 0.5%, answering "yes" to everything 16.29%, and 56.43% on yes/no questions; every baseline is worse than humans. The record does not claim the widely quoted 0.25 million images, 0.76 million questions and 10 million answers, nor ten answers per question: they are not in the first version.

Event record

Event date
May 3, 2015
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0769

The day the first version of the preprint was submitted; the figures were read in it.

Sources

Related events