HumanEval: code checked by tests
On 7 July 2021 Mark Chen and colleagues at OpenAI released HumanEval, 164 hand-written Python problems with unit tests. Codex solved 28.8% at the first attempt, GPT-3 0% and GPT-J 11.4%; with 100 samples per problem, 70.2%.
Why it matters
Code began to be judged not by how much its text resembles a reference but by whether it passes the tests, the pass@k measure. The same preprint showed that sampling the model repeatedly yields working solutions where a single attempt does not.
Codex's training data were collected in May 2020 from 54 million public GitHub repositories: 179 GB of unique Python files under 1 MB, 159 GB after filtering. The HumanEval problems were written by hand so as not to be among those data. A model fine-tuned on separately collected problems, Codex-S, solved 37.7% at the first attempt.