Flamingo: vision and language from a few examples
On 29 April 2022 Jean-Baptiste Alayrac and colleagues at DeepMind described Flamingo, models that join a frozen vision encoder to a frozen language model and take text, images and video interleaved. The largest has 80 billion parameters and beat fine-tuned models on 6 of 16 tasks with only 32 examples.
Why it matters
A single model began to solve new image tasks from a few examples in the prompt, as GPT-3 did with text, instead of being fine-tuned for each dataset. The line of large language models that see runs from it.
Between the encoder and the language model sit a Perceiver Resampler and new gated cross-attention layers, which alone are trained. The data include M3W, text and images from the HTML of about 43 million web pages in their order. The tasks include open-ended ones: visual question answering, captioning and dialogue.