SQuAD 2.0: the question it is right to refuse
On 11 June 2018, 53,775 questions the passage does not answer were added to SQuAD, written by crowdworkers to look like ones it does. The best model fell from 85.8 to 66.3 F1 while a human held at 89.5.
Why it matters
Until then a high reading-comprehension score meant an ability to find the most similar span in a passage, and that ability had already passed the human. A set where a third of the questions have no answer made abstention measurable for the first time and showed with a number that models do not have it: the gap to the human jumped from 5.4 to 23.2 points on one change of condition.
The 53,775 new questions were written by crowdworkers looking at a passage and inventing a question it does not answer, but for which a plausible-looking span exists. To score well a system must not only find answers but hold back when the paragraph does not support one. The numbers from Table 3 of the paper, on the test split: the best of the three models tried, DocQA with ELMo, reaches 63.4 EM and 66.3 F1; the human reaches 86.9 EM and 89.5 F1; the gap is 23.2 points of F1. The same architecture on SQuAD 1.1 reached 85.8 F1, only 5.4 points behind the human. One separate figure says the most about where things stood: a system that always abstains scores 48.9 F1 on the test set, and the existing models are closer to it than to the human. Human performance was computed differently here than in 2016: on average 4.8 answers were collected per question and the majority taken. The authors explain plainly why the older number is lower: the first paper evaluated a single human, so human accuracy there was likely underestimated. That resolves the disagreement between 77.0/86.8 in the 2016 paper and 82.304/91.221 on the leaderboard. One more measurement, rarely quoted: roughly half of all wrong answers - by machines and by humans alike - matched exactly the plausible span the crowdworker had left as a distractor.