Back to timeline

Availability · August 20, 2021

LAION-400M: image-caption pairs in the open

On 20 August 2021 the non-profit network LAION published 400 million English image-text pairs extracted from Common Crawl. CLIP did the selecting: a pair was kept when the cosine similarity between text and image was at least 0.3.

Why it matters

Until then a set of that size existed only inside OpenAI - the same one CLIP and DALL-E were trained on, which nobody outside OpenAI had seen. After this day anyone could take comparable material; all later open image generation rests on it.

There are 413 million unique samples. What was distributed is not images but metadata: 32 parquet files of roughly a gigabyte each, about 50 GB in all, with the fields URL, TEXT, LICENSE, NSFW, similarity, WIDTH and HEIGHT. Each user downloads the images themselves. The 0.3 threshold was chosen by hand - the release says it was determined through human evaluations and seemed a good heuristic for whether a caption matches an image. What the release says about content filtering is worth quoting exactly: "Our filtering protocol only removed NSFW images detected as illegal" - only what a detector judged illegal was dropped, and the rest, tagged, stayed in the set. Two years later the limits of that filtering became the subject of an audit of its own.

Event record

Event date
August 20, 2021
Timeline date
Event date
Verification
Sources gathered automatically · September 22, 2026
Lines
ID
evt-0520

The date in the byline of the release post on LAION's own site.

Sources

Related events

Records that link to this one