LLaVA: an open assistant that sees
On 17 April 2023 Haotian Liu and colleagues at the University of Wisconsin-Madison, Microsoft Research and Columbia University described LLaVA: a CLIP vision encoder joined to the LLaMA language model by a single linear projection and fine-tuned on 158,000 instructions generated by text-only GPT-4.
Why it matters
Conversation about images was trained without human-written instructions and released with its data and code. A month after GPT-4 was announced, an open model reproduced part of its behaviour with pictures.
The instructions are 58,000 conversations, 23,000 detailed descriptions and 77,000 reasoning tasks; alignment before that used 595,000 pairs from CC3M. On a synthetic set LLaVA scores 85.1% relative to GPT-4, and 92.53% on Science QA combined with GPT-4. The first version names LLaMA as the language model; Vicuna appears only among related work.