YAGO: Wikipedia sewn onto WordNet
On 8 May 2007 three Max Planck researchers published YAGO, an ontology of roughly a million entities and five million facts assembled automatically: the individuals and things come from Wikipedia's categories and infoboxes, the top of the hierarchy from WordNet. A sampled human check put accuracy at about 95 per cent.
Why it matters
Until then a large fact base was either written by hand and stayed small, or extracted automatically and stayed dirty. YAGO showed a third way: take someone else's clean taxonomy as the top, fill the bottom from an encyclopedia, and millions of facts hold an accuracy you can state as a number. It was the number, not the promise, that made automatic extraction usable.
The authors were Fabian M. Suchanek, Gjergji Kasneci and Gerhard Weikum, all of the Max Planck Institute for Informatics in Saarbruecken. The paper appeared in the proceedings of the 16th International World Wide Web Conference in Banff. Its figures: roughly one million entities and five million facts, plus approximately forty million context facts. Accuracy of the facts is about 95 per cent by empirical evaluation; a separate 97 per cent is given for the WordNet-to-Wikipedia unification itself. The authors do not take a model off the shelf, and the text says why: OWL Full can express properties of relations but is undecidable; OWL Lite and OWL DL cannot express relations between facts; RDFS can, but its semantics are too primitive. So YAGO is a "slight extension of RDFS". Relations of more than two arguments are handled by naming a primary pair and hanging the rest off the fact identifier. What this record does not claim. The entity count is not 900,000: the paper says "1 million" three times, and 900,000 appears nowhere in it. The 95 per cent concerns the correctness of facts, not coverage: how much of what is in Wikipedia never reached the base is not measured.