Back to timeline

Benchmark · September 7, 2020

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

MMLU: 57 subjects the model did not study

On 7 September 2020 a test of 15,908 four-choice questions across 57 disciplines appeared - from elementary mathematics to professional medicine and law - to be taken with no task-specific training. GPT-3 at 175 billion parameters scored 43.9 per cent against 25 per cent for random guessing.

Why it matters

Until then a language model was measured on tasks it had been fine-tuned for, and the score said how well the fine-tuning had gone. A set drawn from professional examinations the model had not seen and was not prepared for moved the question from linguistic skill to what the model knows at all, and it set the scale the field used for the next five years.

The 15,908 questions were collected by graduate and undergraduate students from freely available sources: practice questions for the GRE, for the United States Medical Licensing Examination, for undergraduate courses, for readers of Oxford University Press books. The test split holds 14,080 questions, at least a hundred in each subject. Subjects come with difficulty levels: elementary, high school, college, professional. The authors explain the number 57 with an aside: it is also the number of Atari games in the standard reinforcement-learning suite. Human performance has to be stated carefully here, because the first version of the preprint contains none at all. It appeared in the third version, dated January 2021, and there it is not a measurement but an estimate with the authors' own caveat: they take 95th-percentile human test-taker accuracy for the exams the test is built from and, where such information is unavailable, make - their words - an educated guess; that is where "approximately 89.8 per cent" comes from. The measured number in the same paragraph is different and rarely quoted: unspecialised people from Amazon Mechanical Turk score 34.5 per cent. This record does not claim that 89.8 per cent is a recorded human result. It is an estimate assembled from other people's percentiles and from guesses, and it became the bar that was for years announced as surpassed.

Event record

Event date
September 7, 2020
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0509

The day of the first version of the preprint, per the arXiv submission history. The paragraph about human performance did not appear until the third version, on 12 January 2021.

Sources

Related events

Records that link to this one