WordNet: a dictionary arranged by meanings
In December 1990 George Miller's Princeton group published WordNet, a database of English vocabulary built around meanings rather than the alphabet. Its unit is the synset, a set of words expressing one concept; the nouns are split into 25 independent hierarchies.
Why it matters
Until then a computer that met a word got from a dictionary only its spelling and a gloss written for a person. WordNet gave a machine a distance between concepts: one can ask what two words have in common and get an answer out of the structure itself. Word-sense disambiguation was built on this, and fifteen years later so was the business of attaching facts scraped from the web to something with a shape.
The authors were George A. Miller, Richard Beckwith, Christiane Fellbaum, Derek Gross and Katherine J. Miller. The paper opens the special issue of the International Journal of Lexicography 3(4), pages 235-244. The work began, as the text says, in 1985. The lexicon is divided into four parts of speech: nouns, verbs, adjectives, adverbs. Function words are excluded on purpose, the text assuming they are stored separately as part of the syntactic component. The nouns are not gathered under one tree: they are partitioned into twenty-five independent hierarchies, each with its own root, because a shared top would be semantically empty. Nouns are linked by hyponymy and meronymy, verbs by troponymy. What this record does not claim. It gives no size for the database as of 1990. The copy of the five papers Princeton still distributes is marked "Revised August 1993", and its figures - 95,600 word forms and 70,100 meanings - are the state of 1993, not of 1990; filing them under this date would be an error. Nor does the record claim twenty-six unique beginners: the table lists twenty-five, and the text says so twice in words.