What is actually inside the corpus T5 was trained on
On 18 April 2021 the first description of the C4 corpus appeared - 156 billion tokens, 365 million documents - the corpus the T5 models were trained on. A blocklist filter removed African American English at a rate of 42 per cent against 6.2 per cent for White-aligned English, and verbatim matches turned up in all five evaluation sets checked.
Why it matters
Until then large text corpora were handed over with almost no description: size, collection method, and that was all. This work showed that filters which look merely technical decide whose language reaches the model, and that a privately held test set can be sitting in the pretraining data. From then on "what is this made of" became a question with numbers rather than a rhetorical one.
Three versions of the corpus: uncleaned (1.1 billion documents, 2.3 TB), without the blocklist (395 million), and C4.EN itself - 365 million documents, 305 GB, 156 billion tokens. The most frequent source is patents.google.com, where a notable share of the text is machine-translated or OCR'd. The central measurement is about the filter. The blocklist of "bad" words was meant to remove hateful and obscene language. It was measured to remove documents in African American English at a rate of 42 per cent and in Hispanic-aligned English at 32 per cent, against 6.2 and 7.2 per cent for White-aligned and other English. Within C4.EN itself, 97.8 per cent of documents are assigned White-aligned English and 0.07 per cent African American English. Contamination has to be stated precisely, because it is the easiest thing here to overstate. The authors checked five sets and summarise themselves this way: all five carry some level of contamination, though for most it is small. The rates: about 1 per cent of target summaries in CNN/Daily Mail, 4.6 and 3.3 per cent of LAMA T-REx and Google-RE examples verbatim in the corpus, 2.4 per cent of BoolQ test questions, and 14.4 per cent of CoLA test instances by input. SQuAD, MNLI, SST-2 and WNLI were not checked in this work and are never mentioned in it. A searchable index of the corpus was released with the paper so that others could repeat the check. The first version of the preprint carries a shorter title than the version published at EMNLP.