GLUE: nine tasks, one score, closed tests
On 20 April 2018 researchers at New York University gathered nine existing language understanding tasks under one leaderboard with privately held test data and a single averaged score. The best baseline in the paper itself reached 60.3.
Why it matters
Until then a system did one task, and comparison meant comparing specialists. One score over nine unlike tasks made generality measurable rather than asserted: a model trained for a single task now lost visibly. That number became the thing pretrained models aimed at for the next two years.
The nine tasks: grammatical acceptability (CoLA), sentiment (SST-2), paraphrase (MRPC, QQP, STS-B) and textual inference (MNLI, QNLI, RTE, WNLI). The headline score is a macro-average over tasks rather than weighted by data size, so that small tasks do not vanish behind large ones. Alongside them sits a separate hand-built diagnostic set, so that it is visible which linguistic phenomena a model actually handles. The stated feature that did not exist before: the leaderboard rests primarily on private test data, and a system must be run over all nine tasks and its results uploaded, rather than its own numbers reported. This is the closed test server often credited to SQuAD. The best result in the paper itself is a macro-average of 60.3 (BiLSTM with ELMo, trained per task). The authors state the conclusion immediately: multi-task learning gave no substantial gain over training a separate model per task, so the room for generality remains open. What this publication does not contain, despite a common claim: a human baseline. A full-text search of the first version returns no occurrence of 87.1. The human estimate for GLUE was published a year later and in a different paper - in the SuperGLUE paper, as a conservative 87.3 per cent. The date of this record also deserves naming separately: the first version of the preprint was submitted on 20 April 2018; the second did not arrive until 18 September, and 2 May 2018 does not appear anywhere in the submission history.