Back to timeline

Benchmark · October 10, 2023

SWE-bench: real tasks from GitHub

On 10 October 2023 Carlos Jimenez, John Yang and colleagues at Princeton released SWE-bench, 2,294 tasks from real issues and pull requests of twelve popular Python repositories. The model gets a codebase and an issue description and must change the code so the tests pass. Claude 2 solved 4.8%, GPT-4 1.7%.

Why it matters

Evaluation of code moved from single functions to work across a whole repository: finding the right files, changing several places at once and working with very long context. It is on this set that the agents recorded in the atlas were measured in 2024.

The 4.8% and 1.7% are with an oracle retriever that hands the model the files of the reference fix; with BM25 retrieval Claude 2 solved 1.96%. For budget reasons GPT-4 was evaluated on a random quarter of the tasks only. The tasks were selected from 90,000 pull requests. The authors also fine-tuned SWE-Llama at 7 and 13 billion parameters on CodeLlama. The preprint calls HumanEval the current standard for evaluating code generation.

Event record

Event date
October 10, 2023
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0792

The first version of the preprint.

Sources

Related events