ROOTS: a corpus whose sources were mostly chosen by people
On 6 December 2022 the BigScience consortium described ROOTS - a multilingual corpus of 1.6TB spanning 59 languages, 46 natural and 13 programming. 62 per cent of its size in bytes comes from a catalogue of 252 sources selected and documented by participants rather than by an automatic crawl.
Why it matters
Until then a large corpus meant whatever could be pulled out of the web, and its composition was described afterwards. Here the order is reversed: language communities first named what was worth taking, and only then was it collected. For the first time provenance was not a report on the work but the method of doing it.
The catalogue of 252 sources was assembled with at least 21 for every language category considered. The remaining 38 per cent is text from the pre-processed web crawl OSCAR - the same automatic route, but as the smaller part. The 46 natural languages belong to nine language families across three macroareas. Among the 13 programming languages, Java, PHP and C++ account for more than half the documents. The open 176-billion-parameter BLOOM model was trained on this corpus. The record gives no token count: the paper carries none, and the often-quoted 341 billion does not come from here.