The interface as a lever for an agent
On 6 May 2024 Princeton showed SWE-agent: not a new model but a command environment rebuilt for a language model. GPT-4 Turbo inside it solved 12.47% of the 2,294 SWE-bench issues against 3.8% for the best non-interactive retrieval approach until then.
Why it matters
Until then the gain on such tasks was sought in the model. This work measured something else: the same model with a rebuilt interface does more than three times better, so how the environment is presented turned out to be a lever comparable to the weights.
What is in the interface. A file is shown through a window of at most 100 lines with scroll and go-to-line commands; editing replaces a named span; a linter is built into the edit command and shows the model a syntax error before the change is applied; search returns a summarised result rather than the whole output. The ablation on a 300-issue subset prices each decision: editing with the linter 18.0 against 15.0 without; a 100-line window against 14.3 for a 30-line one; summarised search 18.0 against 12.0 for iterative; a plain shell instead of the editor 10.3. What the record does not claim. The 12.47% is GPT-4 Turbo; Claude 3 Opus in the same environment gives 10.46%. The claim that the linter drove bad edits "practically to zero" is not in the source: what the source gives is three points on a subset. The 3.8% is the previous best result of a retrieval approach, not a range.