Back to timeline

Research · January 1972

Weighting a term by its rarity in the collection

In January 1972 Karen Spärck Jones of Cambridge proposed weighting query terms by how many documents in the collection contain them, so that a match on a rare term counts for more than one on a common term. On three test collections, Cranfield, INSPEC and Keen, the weighting gave a substantial gain over simple counting of matches.

Why it matters

Term specificity could be measured statistically, by use rather than by meaning, and put into search without a thesaurus or a vocabulary. The weight, later called inverse document frequency, sits inside Salton's vector space model and BM25.

The weight of a term occurring in n of N documents is f(N) - f(n) + 1, where f(n) = m such that 2^(m-1) < n <= 2^m, a stepwise base-2 logarithm. In the 200-document Cranfield collection a term occurring 90 times gets weight 2 and one occurring three times gets 7. INSPEC has 541 documents and 1,341 terms, Keen 797 documents. Simply dropping frequent terms made results worse, as they are needed for recall. The gain was tested by the difference in area under the precision-recall curves; the file read carries no figures for it, being the 2004 reprint in Journal of Documentation 60(5), whose Table I and figures are empty. The words 'inverse document frequency' and 'IDF' do not occur in the paper.

Event record

Event date
January 1972
Timeline date
Event date
Verification
Sources gathered automatically · September 24, 2026
Lines
ID
evt-0689

Journal of Documentation 28(1), January 1972 by Crossref; Salton, Wong and Yang cite the same issue as March 1972 in 1975.

Sources

Records that link to this one