Back to timeline

Research · December 6, 2022

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

ROOTS: a corpus whose sources were mostly chosen by people

On 6 December 2022 the BigScience consortium described ROOTS - a multilingual corpus of 1.6TB spanning 59 languages, 46 natural and 13 programming. 62 per cent of its size in bytes comes from a catalogue of 252 sources selected and documented by participants rather than by an automatic crawl.

Why it matters

Until then a large corpus meant whatever could be pulled out of the web, and its composition was described afterwards. Here the order is reversed: language communities first named what was worth taking, and only then was it collected. For the first time provenance was not a report on the work but the method of doing it.

The catalogue of 252 sources was assembled with at least 21 for every language category considered. The remaining 38 per cent is text from the pre-processed web crawl OSCAR - the same automatic route, but as the smaller part. The 46 natural languages belong to nine language families across three macroareas. Among the 13 programming languages, Java, PHP and C++ account for more than half the documents. The open 176-billion-parameter BLOOM model was trained on this corpus. The record gives no token count: the paper carries none, and the often-quoted 341 billion does not come from here.

Event record

Event date
December 6, 2022
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0522

The publication date in the NeurIPS 2022 proceedings, per the page's own metadata. The arXiv preprint came later, on 7 March 2023.

Sources

Related events

Records that link to this one