Back to timeline

Research · April 29, 2022

Flamingo: vision and language from a few examples

On 29 April 2022 Jean-Baptiste Alayrac and colleagues at DeepMind described Flamingo, models that join a frozen vision encoder to a frozen language model and take text, images and video interleaved. The largest has 80 billion parameters and beat fine-tuned models on 6 of 16 tasks with only 32 examples.

Why it matters

A single model began to solve new image tasks from a few examples in the prompt, as GPT-3 did with text, instead of being fine-tuned for each dataset. The line of large language models that see runs from it.

Between the encoder and the language model sit a Perceiver Resampler and new gated cross-attention layers, which alone are trained. The data include M3W, text and images from the HTML of about 43 million web pages in their order. The tasks include open-ended ones: visual question answering, captioning and dialogue.

Event record

Event date
April 29, 2022
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
ID
evt-0774

The day the first version of the preprint was submitted; the PDF header says 28-04-2022. The figures were read in the first version.

Sources

Related events

Records that link to this one