Back to timeline

Availability · January 23, 2012

A crawl of the web becomes available to anyone

On 23 January 2012 the Common Crawl Foundation announced that its web crawl archive - five billion pages - was hosted on Amazon's Public Data Sets. Amazon pays for the storage; a user pays only for their own compute.

Why it matters

Until then a full slice of the web belonged to whoever crawled it, which meant a search company. From that day it belonged to anyone with a cloud account. Nearly every large text corpus of the following decade, and nearly every argument about what models are made of, starts in this archive.

The foundation was founded in 2007 by Gil Elbaz; data has been collected since 2008. As of 2026 a crawl is published roughly once a month and holds more than two billion pages at a time, with the whole archive past ten pebibytes. What this date did was not to create the corpus but to remove the cost of reaching it. Before the Public Data Sets hosting, getting five billion pages meant downloading them. The record gives no size in terabytes, no count of archive files and no processing tool: the foundation's announcement carries none of those numbers. The share in training data also needs care. In Table 2.2 of the GPT-3 paper, filtered Common Crawl supplies 410 billion of the 499 billion tokens in the assembled dataset, but its weight during training is 60 per cent, and the authors state that the weights are deliberately not proportional to size.

Event record

Event date
January 23, 2012
Timeline date
Event date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0516

The day of the announcement on the foundation's own site. The crawl itself has run since 2008; this day is about access to it, not about its start.

Sources

Related events

Records that link to this one