'A Plan for Spam'
In August 2002 Paul Graham published the essay 'A Plan for Spam': a filter combining, by Bayes' rule, the spam probabilities of a message's fifteen most telling words let through fewer than 5 spams in 1,000 on his own mail, with no false positives.
Why it matters
The essay set out statistical filtering as an open recipe, two corpora, word counts and one formula, and set it against the usual approach of writing rules that recognise individual properties of spam.
The corpora held about 4,000 spam and 4,000 non-spam messages each, and tokens come from the whole text including headers and HTML. Words seen in only one corpus get probabilities of 0.01 and 0.99; a message is spam above 0.9. The figures are the author's count on his own mail, not an independent measurement. The page carries a later note pointing to an improved algorithm in 'Better Bayesian Filtering'; the record concerns only the algorithm of August 2002. It does not claim where or when such filters came into mass use: no source on that was read.