The document as a vector of terms
In November 1975 Gerard Salton, A. Wong and C. S. Yang of Cornell described documents and queries as vectors of term weights and showed that retrieval works better where documents lie further apart in that space. On three collections, 424 documents in aerodynamics, 450 in medicine and 425 Time articles, they used this measure to choose an indexing vocabulary.
Why it matters
Retrieval became a matter of geometry: the similarity of a document and a query is a function of the angle between vectors, and the quality of a vocabulary is the density of the space.
The authors describe similarity as the inner product of the vectors 'or alternatively an inverse function of the angle' between them; the word cosine does not occur. Replacing term frequency with inverse document frequency gave, by earlier work of Salton and Yang, about 14 per cent average improvement in precision. The Time collection has 7,569 terms. Table IV: automatic phrases against the standard run, +32, +39 and +17 per cent on the three collections; phrases with a thesaurus, +33, +50 and +18. The name vector space model is in the title; the SMART system the experiments ran on is not described by this record.