AISI resumes evaluations behind new controls
On 1 October 2026 the UK AI Security Institute said that after the first phase of strengthening its security it could resume most of the evaluation activity paused after its August incident: internet access for agentic cyber evaluations is switched off, the sandboxes' outbound network is blocked by two independent layers, and a synchronous monitor built on a language model can stop a suspicious action before it happens.
Why it matters
A public worked example of what safe evaluation costs: the institute names the conflict between realism of testing and containment of the model outright and writes that the burden falls hardest on smaller and less well-resourced evaluators. It also admits the limit of the monitor: chain-of-thought monitoring is fragile, because models increasingly act without words in the chain or shape it to mislead a monitor.
What was changed (the institute's own report, not independently checked). Internet access for future agentic cyber evaluations is off until safe controls exist to bring it back; outbound networking is blocked within the sandboxes and, separately, by cloud controls on the virtual-machine host; both layers are checked by a phased process that hunts for misconfigurations. The monitor reviews the agent's messages, tool calls and, where available, chain of thought, can block an action and pass it to a person; it was validated on past logs and synthetic transcripts and against models' attempts to evade it; for models whose chain of thought is not provided there is a second monitor working from the sequence of actions alone, expected to be less effective. Evaluation design: tasks that run without internet, clearer boundaries in prompts, automatic checks before a run that the monitor is on and the internet off. New internal governance: technical and security review of research, following NCSC guidance. Security is tested with AI too: static analysis, dynamic analysis, and controlled escape attempts (a flag outside the environment, a monitor that halts the exercise, full isolation from production systems). Planned: a new sandbox service with automatic quarantine and a single security platform. The institute writes that these measures reduce risk and do not eliminate it, and that measures sufficient for today's models may not hold for the next generation. What the record does not claim: that anyone but the institute has tested the controls; that 'most' means all evaluations (the highest-risk ones are not named); the date, which is taken from the blog index.