Back to timeline

Research · 1998

A Bayesian junk-mail filter

In 1998 Mehran Sahami of Stanford and Susan Dumais, David Heckerman and Eric Horvitz of Microsoft Research showed that a naive Bayesian classifier trained on a user's mail filters junk. On 1,789 messages, with simple mail-specific features, it reached 100 per cent precision and 98.3 per cent recall on junk.

Why it matters

A junk filter could be learned from one's own mail rather than written as rules by hand, and its threshold deliberately shifted so that legitimate mail was almost never lost. The authors argued the technology was ready for deployment.

Corpus: 1,789 messages, 1,578 junk and 211 legitimate, split in time into 1,538 for training and 251 for testing. A message counts as junk only if its probability exceeds 99.9 per cent. Table 1, junk precision and recall: words alone 97.1 and 94.3, words and about 35 phrases 97.6 and 94.3, with 20 non-textual mail-specific features 100.0 and 98.3. In a real-use test trained on 2,593 messages from a year, a week of 222 incoming messages, 45 of them junk, lost 80 per cent of the junk and flagged 3 legitimate messages, one a forwarded piece of junk. Splitting junk into pornographic and other made results worse.

Event record

Event date
1998
Timeline date
Event date
Verification
Sources gathered automatically · September 24, 2026
Lines
ID
evt-0695

AAAI Technical Report WS-98-05, 1998; the document read names neither the workshop's month nor its venue.

Sources

Records that link to this one