Conditional random fields
At ICML 2001 John Lafferty (Carnegie Mellon), Andrew McCallum and Fernando Pereira (WhizBang! Labs, Pereira also the University of Pennsylvania) presented conditional random fields: a model of the probability of a whole label sequence given an observation sequence, normalised over the sequence rather than at each state. On part-of-speech tagging of the Penn treebank the per-word error was 5.55 per cent against 5.69 for a hidden Markov model and 6.37 for a maximum entropy Markov model; with spelling features, 4.27 against 4.81.
Why it matters
A sequence could be labelled using many overlapping features of the input, which a generative HMM cannot represent, and without the label bias of discriminative models normalised state by state, which the paper says can be biased towards states with few successor states. It is the step from the generative model of the 1989 tutorial to discriminative sequence labelling. An editorial assessment.
Figure 4 gives per-word error and error on out-of-vocabulary words: HMM 5.69 and 45.99 per cent, MEMM 6.37 and 54.61, CRF 5.55 and 48.05, MEMM with spelling features 4.81 and 26.99, CRF with them 4.27 and 23.76. Without spelling features CRFs beat the HMM overall but are worse on unseen words. Read in full from the authors' postprint in the University of Pennsylvania repository, which says the definitive version is in the ICML 2001 proceedings, pages 282-289. What the record does not claim: the venue and closing date of the conference, which no source read names, or 28 June as the day of the talk.