Back to timeline

Availability · December 31, 2020

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

The Pile: 825 gibibytes from twenty-two sources

On 31 December 2020 the independent group EleutherAI released The Pile - an English corpus of 825.18 gibibytes assembled from twenty-two separate sets. The largest of them, the web crawl Pile-CC, supplies 227.12 GiB, under a fifth of the whole.

Why it matters

Until then an open corpus meant the web cleaned of everything that did not look like the web. Here it is the reverse: arXiv and PubMed Central papers, code from GitHub, court opinions, patents, subtitles, books. For the first time material to train a GPT-3-class model existed outside a large lab, and at the same time what that material consisted of was named item by item.

Twenty-two sets with their sizes in Table 1: Pile-CC 227.12 GiB, PubMed Central 90.27, Books3 100.96, OpenWebText2 62.77, arXiv 56.21, GitHub 95.16, FreeLaw 51.15, Stack Exchange 32.20, and so down to 0.88 GiB of Enron emails. On Books3 the paper is direct: these are books taken from a copy of the contents of the Bibliotik private tracker, made available by Shawn Presser. It gives no count of books, so none appears here. The often-quoted "108 GB" is the same 100.96 GiB in decimal units. A section of its own, 7.1, is about legality. The authors write that the community discusses the legality of training on copyright material but barely discusses the legality of processing and distributing it - and then argue that their own use is fair: non-commercial, transformative, and needing the full text because language has long-range dependencies.

Event record

Event date
December 31, 2020
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 22, 2026
Lines
ID
evt-0519

The day of the preprint's only version, per the arXiv submission history.

Sources

Related events

Records that link to this one