Brill's tagger: rules the machine learns
At the third ANLP conference in Trento (31 March to 3 April 1992) Eric Brill of the University of Pennsylvania presented a part-of-speech tagger that finds its own rules for correcting its errors. The 71 rules it found brought the error rate on a held-out 5% of the Brown Corpus down to 5.1%.
Why it matters
Brill showed that tagging by rules the machine learns from a corpus reaches the accuracy of stochastic taggers while remaining a list of a few dozen rules a person can read. The version of the same tagger from his 1994 paper assigned the parts of speech in the CoNLL-2000 data.
The initial tagger gives each word its most frequent tag in the training 90% of the Brown Corpus (about 1.1 million words) and gets about 7.9% of words wrong. Then, on a separate 5% of the corpus, the system tries patch templates (for instance, change tag a to b if the previous tag is z) and each time takes the one that removes the most errors. Of the 71 rules, 66 reduced errors on the test set, 3 changed nothing and 2 made things worse; three different splits of the corpus each gave 5%. Brill's own implementation of Church's algorithm on the same samples gave about 4.5%. What the record does not claim. The word transformation does not occur in the 1992 paper, which calls its rules patches; the name transformation-based learning came later. The paper says the tagger performs on par with stochastic ones, although its own comparison with Church slightly favours Church.