The Pile: 825 gibibytes from twenty-two sources
On 31 December 2020 the independent group EleutherAI released The Pile - an English corpus of 825.18 gibibytes assembled from twenty-two separate sets. The largest of them, the web crawl Pile-CC, supplies 227.12 GiB, under a fifth of the whole.
Why it matters
Until then an open corpus meant the web cleaned of everything that did not look like the web. Here it is the reverse: arXiv and PubMed Central papers, code from GitHub, court opinions, patents, subtitles, books. For the first time material to train a GPT-3-class model existed outside a large lab, and at the same time what that material consisted of was named item by item.
Twenty-two sets with their sizes in Table 1: Pile-CC 227.12 GiB, PubMed Central 90.27, Books3 100.96, OpenWebText2 62.77, arXiv 56.21, GitHub 95.16, FreeLaw 51.15, Stack Exchange 32.20, and so down to 0.88 GiB of Enron emails. On Books3 the paper is direct: these are books taken from a copy of the contents of the Bibliotik private tracker, made available by Shawn Presser. It gives no count of books, so none appears here. The often-quoted "108 GB" is the same 100.96 GiB in decimal units. A section of its own, 7.1, is about legality. The authors write that the community discusses the legality of training on copyright material but barely discusses the legality of processing and distributing it - and then argue that their own use is fair: non-commercial, transformative, and needing the full text because language has long-range dependencies.