Claude Opus 4 and Sonnet 4
Anthropic released a generation tuned for long tasks: the model could hold one piece of code work for hours, alternating deliberation and tool calls.
Why it matters
The measure of a model became not how well it answers but how long it can work before it has to be corrected.
Until then what limited agents was not intelligence but accumulating error: after a few dozen steps a system lost the thread. Here the models could alternate reasoning with search and execution while holding the goal, and that is what they were measured on. It settled the shift that began with Claude Code: the main use of a frontier model became programming, not as help with a line but as independent work on a task.