Back to timeline

Research · July 2002

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

BLEU: a formula scores translation quality

In July 2002 four IBM researchers showed a way to score machine translation without a person: count matches of runs of one to four words against human references, with a penalty for an answer that is too short. Across five systems the score correlated with human judges at 0.99.

Why it matters

Until then checking a translation meant a panel of people and weeks or months of waiting, so a developer could not see what yesterday's change had done. A formula that runs in seconds and produces the same ordering as the judges turned evaluation from an event into an operation, and by the same move made the metric itself the thing being optimised.

The baseline form of the metric, in the paper's own words: N = 4 with uniform weights wn = 1/N, that is, the geometric mean of modified precision over runs of one to four words. "Modified" means a matching word counts no more times than it occurs in a reference; otherwise "the" repeated seven times would score 7/7. The brevity penalty is computed over the whole corpus rather than sentence by sentence: BP is 1 when the candidate is longer than the reference and e^(1-r/c) when it is shorter. The check was run with two groups of ten judges: a monolingual group of native English speakers and a bilingual group of native Chinese speakers living in the United States. None of the judges was a professional translator. They rated five systems on a subset of a 500-sentence corpus, 250 pairs in all, on a scale from 1 to 5. The correlation coefficient against the monolingual group was 0.99 and against the bilingual group 0.96. What this record does not claim: BLEU does not measure whether a translation is correct, it measures closeness to a set of human versions. The authors themselves call the metric an understudy to the judges rather than a replacement. What that means in practice was measured four years later at WMT-06, and the measurement did not favour the formula. The Anthology record carries the field "Award: NAACL 2018 Test-of-Time" - awarded sixteen years after publication.

Event record

Event date
July 2002
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0503

July 2002, the 40th Annual Meeting of the ACL in Philadelphia. The sources give no day: the paper's header reads "Philadelphia, July 2002" and the Anthology record reads "Month: July, Year: 2002".

Sources

Related events

Records that link to this one