Back to timeline

Availability · October 23, 2019

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

C4: the web put through six rules

On 23 October 2019 Google released C4 - about 750 GB of English text made from a single monthly slice of Common Crawl, the one for April 2019. The corpus was distributed through TensorFlow Datasets alongside the paper on the T5 models.

Why it matters

Until then open text corpora were either small or handed over as a raw dump. Here a large web corpus went out with an exact list of the rules by which things had been thrown away - and it was those rules that outsiders were able to check eighteen months later.

Six cleaning rules: keep only lines ending in terminal punctuation; drop any page containing any word from the "List of Dirty, Naughty, Obscene or Otherwise Bad Words"; drop lines containing the word Javascript; drop pages containing "lorem ipsum"; drop pages containing a curly bracket, since it occurs in code and not in natural text; keep one copy of any three-sentence span that occurs more than once. Then a language filter: a page was kept if it was classified as English with a probability of at least 0.99. One monthly Common Crawl pass yields around 20TB of text; after these rules about 750 GB remained. The record deliberately gives no token count, no licence name and no Hugging Face distribution: version 1 of the paper carries none of them. The figure of 156 billion tokens comes from a later outside description of the corpus, not from here.

Event record

Event date
October 23, 2019
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0518

The day of the first version of the preprint, per the arXiv submission history. Every number here was read in version 1 rather than on the abstract page, which shows a 2023 version.

Sources

Related events

Records that link to this one