SWE-bench: real tasks from GitHub
On 10 October 2023 Carlos Jimenez, John Yang and colleagues at Princeton released SWE-bench, 2,294 tasks from real issues and pull requests of twelve popular Python repositories. The model gets a codebase and an issue description and must change the code so the tests pass. Claude 2 solved 4.8%, GPT-4 1.7%.
Why it matters
Evaluation of code moved from single functions to work across a whole repository: finding the right files, changing several places at once and working with very long context. It is on this set that the agents recorded in the atlas were measured in 2024.
The 4.8% and 1.7% are with an oracle retriever that hands the model the files of the reference fix; with BM25 retrieval Claude 2 solved 1.96%. For budget reasons GPT-4 was evaluated on a random quarter of the tasks only. The tasks were selected from 90,000 pull requests. The authors also fine-tuned SWE-Llama at 7 and 13 billion parameters on CodeLlama. The preprint calls HumanEval the current standard for evaluating code generation.