BIG-bench: 204 tasks assembled by 442 authors
On 9 June 2022, 442 authors from 132 institutions published 204 tasks chosen deliberately to lie beyond what models can do. Running them across sizes from millions to hundreds of billions of parameters showed that quality sometimes jumps at a particular scale - and that sometimes the jump is produced by a brittle metric.
Why it matters
Until then a benchmark was built by one group, and what that group found hard set the limit of what could be measured. An open call answered by a hundred and thirty-two institutions removed that limit and produced material on which a jump in ability with scale can be decomposed rather than proclaimed. One of the causes turned out to be an artefact of measurement.
204 tasks: from logical paradoxes and cryptography to problems in biology, physics, software development, childhood development and social bias. The models tested were OpenAI's GPT models, Google-internal dense transformers and sparse models, across sizes from millions to hundreds of billions of parameters. Separately, a team of human expert raters performed all the tasks to provide a baseline. The findings, as the authors state them: performance and calibration both improve with scale but remain poor in absolute terms, particularly against the raters; different model classes behave remarkably similarly, with a benefit from sparsity; tasks that improve gradually and predictably usually rest on knowledge or memorisation, while tasks with a breakthrough at a critical scale involve multiple steps or components or brittle metrics; social bias typically increases with scale where context is ambiguous, and prompting can reduce it. Those last two points are why this record is here. The authors named the method of measurement among the causes of "emergence", which puts part of the conclusions drawn from their own set in doubt. What this record does not claim: the first version of the paper contains neither "57 languages" nor BIG-bench Hard. Non-English tasks there are tagged and counted differently: non-English 16, multilingual 12, low-resource language 10, translation 10. BIG-bench Hard is a separate, later work. The author count depends on the version too: 442 in the first, 450 in the current one.