Back to timeline

Benchmark · July 7, 2021

HumanEval: code checked by tests

On 7 July 2021 Mark Chen and colleagues at OpenAI released HumanEval, 164 hand-written Python problems with unit tests. Codex solved 28.8% at the first attempt, GPT-3 0% and GPT-J 11.4%; with 100 samples per problem, 70.2%.

Why it matters

Code began to be judged not by how much its text resembles a reference but by whether it passes the tests, the pass@k measure. The same preprint showed that sampling the model repeatedly yields working solutions where a single attempt does not.

Codex's training data were collected in May 2020 from 54 million public GitHub repositories: 179 GB of unique Python files under 1 MB, 159 GB after filtering. The HumanEval problems were written by hand so as not to be among those data. A model fine-tuned on separately collected problems, Codex-S, solved 37.7% at the first attempt.

Event record

Event date
July 7, 2021
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0788

The first version of the preprint.

Sources

Related events

Records that link to this one