Back to timeline

Availability · August 27, 2026

optstop: when to stop measuring a model

On 27 August 2026 the UK AI Security Institute released optstop, an open-source Python package that stops an evaluation where the estimate is already precise enough and keeps going where uncertainty remains. In the institute's tests it saved between 57 and 97 per cent of planned trials under each of nine conditions, without changing the score estimates.

Why it matters

The budget of an evaluation stopped being a number that has to be guessed in advance. Until then the sample size was fixed before the run and could not respond to where certainty was actually missing: part of the budget went on measurements that were already precise, and the hard cases ran out of it. A tool from a government evaluator, and an open one, makes that decision measurable and checkable.

The package plugs into the same institute's Inspect evaluation framework. It tracks the credible interval at two levels — repetitions within a task and tasks within a grouping — and at the grouping level combines them with a hierarchical Bayesian model that accounts for unequal sample sizes. It stops on either of two rules: precision, when the interval is narrow enough, or stabilisation, when the interval has stopped changing. If neither fires, the grouping runs to its full planned budget, so a genuinely noisy estimate is not cut short. A separate safeguard demands more data when estimated success falls below one per cent, because a rare success can be the whole point of the test. Every stopping decision is recorded along with its rationale. The testing covered three kinds of scoring — binary, ordinal and continuous — at three levels of expected performance, so nine settings in all, on public benchmarks including MATH, GPQA Diamond and WritingBench. The institute frames the problem this way: fixed-budget evaluations systematically underestimate frontier agentic capability, and some frontier evaluations now need hundreds of millions of tokens, so every wasted step is expensive. Unlike adaptive item selection, the package needs no pre-calibrated bank of tasks with known difficulty, and no task is dropped from consideration. What the record does not claim. That the saving will be this large in any evaluation: the institute notes that a leaner initial design leaves fewer trials to save. That anyone outside the institute has used the tool: nothing of the kind was read as of the day of this record.

Event record

Event date
August 27, 2026
Timeline date
Event date
Verification
Sources gathered automatically · September 28, 2026
Lines
ID
evt-0902

The day the tool and the institute's post appeared. The paper describing it was posted to arXiv on 14 August 2026.

Sources

Related events