Back to timeline

Benchmark · June 16, 2016

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

SQuAD: a hundred thousand questions on Wikipedia passages

On 16 June 2016 Stanford published a set of 107,785 questions written by crowdworkers against 536 Wikipedia articles, where the answer is a span of the passage itself. The authors' own model reached 51.0 F1; a human on the test set reached 86.8.

Why it matters

Until then reading-comprehension sets were either small or template-generated, and could not be trained on. A hundred thousand labelled questions with span answers made reading comprehension a task you learn rather than a task you demonstrate, and it set the format in which language models were measured for the next four years.

The passages were selected this way: from the ten thousand highest-ranked English Wikipedia articles by internal PageRank, 536 were sampled at random, yielding 23,215 paragraphs longer than 500 characters. Crowdworkers were paid nine dollars an hour. Scoring used two measures: exact string match (EM) and token-level overlap (F1). The number that shows how hard it was at release: the authors' own logistic regression reached 40.0 EM and 51.0 F1, against 20 per cent for a sliding window. Within four months neural models reached 70.3 F1. On human performance there is a genuine disagreement here. The paper, section 6.2: "77.0 per cent for the exact match metric, and 86.8 per cent for F1", taking the second answer to each question as the human prediction. The SQuAD leaderboard carried a different pair, 82.304 EM and 91.221 F1. Both are real, they were computed differently, and the authors named the reason themselves in the SQuAD 2.0 paper: in 2016 a single human was evaluated, so human accuracy was likely underestimated. The leaderboard also fixes the date the machine crossed the human mark. As of 15 January 2018 the top rows were r-net+ from Microsoft Research Asia (82.650 EM, 3 January) and SLQA+ from Alibaba (82.440 EM, 5 January) against the human 82.304. But on F1 neither passed the human: 88.493 and 88.607 against 91.221. So "the machine beat the human on SQuAD" holds for exactly one of the two measures. This record does not claim that SQuAD introduced a hidden test server: the paper has no such thing, only an ordinary held-out tenth. A closed leaderboard as a stated feature arrived with GLUE two years later.

Event record

Event date
June 16, 2016
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0505

The day of the first version of the preprint, per the arXiv submission history. The peer-reviewed publication followed at EMNLP the same year.

Sources

Related events

Records that link to this one