LAION-400M: image-caption pairs in the open
On 20 August 2021 the non-profit network LAION published 400 million English image-text pairs extracted from Common Crawl. CLIP did the selecting: a pair was kept when the cosine similarity between text and image was at least 0.3.
Why it matters
Until then a set of that size existed only inside OpenAI - the same one CLIP and DALL-E were trained on, which nobody outside OpenAI had seen. After this day anyone could take comparable material; all later open image generation rests on it.
There are 413 million unique samples. What was distributed is not images but metadata: 32 parquet files of roughly a gigabyte each, about 50 GB in all, with the fields URL, TEXT, LICENSE, NSFW, similarity, WIDTH and HEIGHT. Each user downloads the images themselves. The 0.3 threshold was chosen by hand - the release says it was determined through human evaluations and seemed a good heuristic for whether a caption matches an image. What the release says about content filtering is worth quoting exactly: "Our filtering protocol only removed NSFW images detected as illegal" - only what a detector judged illegal was dropped, and the rest, tagged, stayed in the set. Two years later the limits of that filtering became the subject of an audit of its own.