C4: the web put through six rules
On 23 October 2019 Google released C4 - about 750 GB of English text made from a single monthly slice of Common Crawl, the one for April 2019. The corpus was distributed through TensorFlow Datasets alongside the paper on the T5 models.
Why it matters
Until then open text corpora were either small or handed over as a raw dump. Here a large web corpus went out with an exact list of the rules by which things had been thrown away - and it was those rules that outsiders were able to check eighteen months later.
Six cleaning rules: keep only lines ending in terminal punctuation; drop any page containing any word from the "List of Dirty, Naughty, Obscene or Otherwise Bad Words"; drop lines containing the word Javascript; drop pages containing "lorem ipsum"; drop pages containing a curly bracket, since it occurs in code and not in natural text; keep one copy of any three-sentence span that occurs more than once. Then a language filter: a page was kept if it was classified as English with a probability of at least 0.99. One monthly Common Crawl pass yields around 20TB of text; after these rules about 750 GB remained. The record deliberately gives no token count, no licence name and no Hugging Face distribution: version 1 of the paper carries none of them. The figure of 156 billion tokens comes from a later outside description of the corpus, not from here.