Back to timeline

Availability · March 31, 2022

LAION-5B: 5.85 billion pairs, a third of them not in English

On 31 March 2022 LAION published 5.85 billion image-text pairs, fourteen times more than LAION-400M. Of these 2.3 billion are English, another 2.2 billion are in more than a hundred other languages, and about a billion carry text in no assignable language.

Why it matters

Until then open sets at this scale existed only in English, and there was nothing to train a multimodal model on outside it. Here half the set is not English for the first time. Within six months a model trained on a subset of it shipped with weights anyone could run at home.

The set was distributed with CLIP ViT-L/14 embeddings, kNN indices, and per-entry probability scores for watermarking and for unsafe content. The licence has to be stated precisely, because this is the easiest thing here to get wrong. What goes out under Creative Commons CC-BY 4.0 is the metadata - parquet files of links and captions. The images themselves stay under their own copyright; LAION neither distributes them nor holds rights in them. That distinction is what every later argument about this set turns on. The record gives no size in gigabytes for the metadata: the release carries no such figure. The finer split into subsets - 2.32, 2.26 and 1.27 billion - comes not from here but from a later outside report.

Event record

Event date
March 31, 2022
Timeline date
Event date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0521

The date in the byline of the release post on LAION's site. The peer-reviewed paper came later, at NeurIPS 2022; the day recorded here is the day the set became available.

Sources

Related events

Records that link to this one