WMT-06: the formula and the judges over the same outputs
On 8 and 9 June 2006, fourteen teams from eleven institutions translated between English, French, German and Spanish on the Europarl corpus. For the first time in this competition the outputs were scored both by BLEU and by people - about 180 hours of the participants' own labour - and the two scorings disagreed.
Why it matters
Until then the automatic metric and the human judgement lived in separate papers, and the question of whether they measure the same thing had no annual answer. Here it got one: the rule-based system was scored by the formula below every statistical system, including ones that people rated markedly lower. It became visible that the choice of metric chooses the winner.
Fourteen teams from eleven institutions: commercial companies, industrial research labs and individual graduate students. Six translation directions, the Europarl training corpus (688,000 to 751,000 sentences per pair), a 2,000-sentence in-domain test set, and a newly added out-of-domain test set of editorials. The manual evaluation was done by the participants themselves because there was no funding for judges: about 180 hours of labour, 300 to 400 judgements per judgement type, per system, per language pair. Judges saw five outputs for one sentence at a time and rated adequacy and fluency separately. The organisers named the limit of their own procedure straight away: with the number of judgements collected, only half of the systems can be told apart. The central comparison, in their own words: the organisers confirm that the rule-based Systran system is "not adequately appreciated" by BLEU - on in-domain data it scores below every statistical system, including ones with much worse human scores. On out-of-domain data the effect is weaker, and for English-French Systran has both the best BLEU and the best manual scores there. The spread among the judges was measured too: average fluency judgements per judge ranged from 2.33 to 3.67 and average adequacy judgements from 2.56 to 4.13, so the scores had to be normalised. This record does not claim that this was the first WMT. The volume's introduction says plainly, "This is the second time that this workshop has been held", the first being in 2005 inside an ACL workshop on parallel texts. What was new here is something else: manual evaluation alongside the automatic kind.