Back to timeline

Research · April 17, 2023

LLaVA: an open assistant that sees

On 17 April 2023 Haotian Liu and colleagues at the University of Wisconsin-Madison, Microsoft Research and Columbia University described LLaVA: a CLIP vision encoder joined to the LLaMA language model by a single linear projection and fine-tuned on 158,000 instructions generated by text-only GPT-4.

Why it matters

Conversation about images was trained without human-written instructions and released with its data and code. A month after GPT-4 was announced, an open model reproduced part of its behaviour with pictures.

The instructions are 58,000 conversations, 23,000 detailed descriptions and 77,000 reasoning tasks; alignment before that used 595,000 pairs from CC3M. On a synthetic set LLaVA scores 85.1% relative to GPT-4, and 92.53% on Science QA combined with GPT-4. The first version names LLaMA as the language model; Vicuna appears only among related work.

Event record

Event date
April 17, 2023
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0777

The day the first version of the preprint was submitted; the figures were read in it.

Sources

Related events