A crawl of the web becomes available to anyone
On 23 January 2012 the Common Crawl Foundation announced that its web crawl archive - five billion pages - was hosted on Amazon's Public Data Sets. Amazon pays for the storage; a user pays only for their own compute.
Why it matters
Until then a full slice of the web belonged to whoever crawled it, which meant a search company. From that day it belonged to anyone with a cloud account. Nearly every large text corpus of the following decade, and nearly every argument about what models are made of, starts in this archive.
The foundation was founded in 2007 by Gil Elbaz; data has been collected since 2008. As of 2026 a crawl is published roughly once a month and holds more than two billion pages at a time, with the whole archive past ten pebibytes. What this date did was not to create the corpus but to remove the cost of reaching it. Before the Public Data Sets hosting, getting five billion pages meant downloading them. The record gives no size in terabytes, no count of archive files and no processing tool: the foundation's announcement carries none of those numbers. The share in training data also needs care. In Table 2.2 of the GPT-3 paper, filtered Common Crawl supplies 410 billion of the 499 billion tokens in the assembled dataset, but its weight during training is 60 per cent, and the authors state that the weights are deliberately not proportional to size.