DBpedia: a Wikipedia you can query
In November 2007 a group from Berlin, Leipzig and Pennsylvania presented DBpedia at the ISWC conference: 103 million RDF triples extracted from the infoboxes of the English Wikipedia. The dataset described more than 1.95 million things and answered queries over an open SPARQL endpoint.
Why it matters
Until then Wikipedia could only be read, or searched as full text. DBpedia turned it into a base you could ask a question of - every German city over a million people - and get an answer rather than a list of pages. More consequential still were the 180,000 outward links: from this point other people's datasets had a shared address system to attach themselves to.
The authors were Soeren Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak and Zachary Ives, of the Freie Universitaet Berlin, the University of Leipzig and the University of Pennsylvania. The paper sits in Lecture Notes in Computer Science 4825, pages 722-735. Its figures: 103 million RDF triples; more than 1.95 million "things", among them at least 80,000 persons, 70,000 places, 35,000 music albums and 12,000 films; 657,000 links to images, 1,600,000 links to external web pages and 180,000 external RDF links. Together with the datasets DBpedia is interlinked with, that amounts to about 2 billion triples. What this record does not claim. The two billion is not the size of DBpedia but of the whole linked cloud it belongs to. Nor does the record claim the extracted data is clean: the paper names outright the properties of Wikipedia it inherits - contradictory data, inconsistent taxonomical conventions, errors and spam.