CoNLL-2000: eleven systems on one split
On 13 and 14 September 2000 in Lisbon, eleven research systems cut the same text into groups of words, over the same 211,727 training and 47,377 test tokens, and were scored by one shared script. The best reached an F of 93.48 against 77.07 for a simple baseline.
Why it matters
Until then a result on this task meant a number the authors had computed on their own split by their own rules, and two numbers from two papers could not be set side by side. After this day they could be: split, script and baseline sat in one place, and the difference between systems became a difference between systems rather than between ways of counting.
The data was taken from the parsed Wall Street Journal in Penn Treebank II: sections 15 to 18 for training (211,727 tokens, 106,978 chunks), section 20 for testing (47,377 tokens). Part-of-speech tags came from the Brill tagger rather than from the treebank itself, deliberately, so that the scores would be realistic for text that has no parse. Distributed with the data was conlleval, a Perl script for computing precision, recall and F. The eleven systems fell into four families: rule-based, memory-based, statistical, and combinations. All eleven beat the baseline, which simply assigned the most frequent chunk tag for each part-of-speech tag (F 77.07). Six of the eleven landed between 91.5 and 92.5. The top two were clearly above: support vector learning by Kudoh and Matsumoto (F 93.48) and weighted probability distribution voting by van Halteren (F 93.32). This record does not claim that the shared-task format began here. The organisers' own site lists CoNLL-97, CoNLL-98 and CoNLL-99, so this is the fourth conference in a row, and shared tasks did not start with it. The test-token count and conlleval itself are not named in the paper; both come from the task page, which survives only in an archived copy. The date does not come from the paper either: it reads "Lisbon, Portugal, 2000", and the day and month sit on an archived copy of the conference home page. That is why confidence here is medium, although the three sources do not disagree with one another.