Back to timeline

Benchmark · May 2, 2019

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

SuperGLUE: a set built for years lasted twenty months

On 2 May 2019 the authors of GLUE declared their own set exhausted and released a harder one, keeping two of the nine tasks. A human estimate shipped with the release: 89.6 against 69.7 for the best baseline. On 6 January 2021 DeBERTa sat on top with 90.3.

Why it matters

Until then nobody had measured the shelf life of a test set, because there were no two sets from the same people under the same procedure. Here it is measured: twenty months from release to defeat, over a single change of model generation. From then on a benchmark had to be planned as something that spoils rather than as a scale.

The first version of the preprint - the one this record is dated to - carries six tasks in its results table: CB, COPA, MultiRC, RTE, WiC and WSC. Its own prose says "seven", and its list of changes says "retains only two of the nine GLUE tasks and replaces the remainder with four new ones", which is six; the source disagrees with itself, which is why confidence here is medium. BoolQ and ReCoRD, and with them an eighth task, appear in the version published later. The numbers in the first version: human estimate 89.6 by macro-average, BERT 66.6, BERT++ 69.7. In words the authors summarise the gap as about 20 points. The largest shortfall is on the Winograd schemas, 32.5 points; the smallest are on CB, RTE and WiC at 11.8, 10.9 and 9.6. The published version carries eight tasks and restates the human estimate as 89.8. The date of defeat is taken not from the leaderboard but from a primary source that cites it: the DeBERTa abstract states that a single model passed the human macro-average for the first time at 89.9 against 89.8, and that the ensemble "sits atop the SuperGLUE leaderboard as of January 6, 2021" with 90.3 against 89.8. This is also where the human estimate for GLUE was first published, which GLUE itself did not carry: a conservative 87.3 per cent for a non-expert. So the human level for a 2018 benchmark arrived a year later and in somebody else's paper. The claim that T5 with Meena scored 90.0 could not be verified, and this record does not carry it.

Event record

Event date
May 2, 2019
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0508

The day of the first version of the preprint, per the arXiv submission history. It is this version, not the later eight-task one, that fixes what this record claims.

Sources

Related events