CLIP
CLIP trained joint image-text representations.
Why it matters
It linked image and text representations without a task-specific labeled dataset.
The web-derived training data can carry biases and noise into the model.
Research · January 5, 2021
CLIP trained joint image-text representations.
It linked image and text representations without a task-specific labeled dataset.
The web-derived training data can carry biases and noise into the model.
OpenAI · Published January 5, 2021
arXiv · Published February 26, 2021
CLIP uses a Transformer text encoder, documented in the technical paper included alongside its announcement.
CLIPBoth connect text and images: CLIP matches them, whereas Muse Image generates images. No direct model dependency is claimed.
The text condition is supplied through a representation trained to match text and images.
CLIP did the selecting: a pair was kept when the cosine between the text and image embeddings was at least 0.3.
LAION-400-MILLION OPEN DATASETFlamingo's vision encoder is pretrained contrastively "a la CLIP", with the two-term contrastive loss from the CLIP paper.
Flamingo: a Visual Language Model for Few-Shot Learning (arXiv:2204.14198v1)LLaVA's vision encoder is the pretrained CLIP ViT-L/14, joined to the language model by a single linear projection.
Visual Instruction Tuning (arXiv:2304.08485v1)