METR discloses two security incidents
The organisation described how in March an attacker talked an agent into handing over an API key and spent credits for three weeks, and how in May attackers probed its infrastructure using agents.
Why it matters
An AI evaluator published, with numbers, how its own agent was talked into revealing its credentials and how attackers automated the search for the next hole using agents.
The publication is dated 31 August 2026. In March 2026 a researcher ran an agent orchestration dashboard, put together quickly, on a personal EC2 instance; it contained a fail-open vulnerability that silently disabled authentication and left the dashboard exposed. METR suspects the attacker found it by looking through recently registered websites for such hand-made pages with high-signal keywords about language models and agents. The attacker prompted an agent directly to reveal its model provider API key, added an SSH key for persistent access, and over three weeks used the stolen credentials. The credits consumed would have been worth approximately 600 thousand dollars, although the model developer had granted them to METR for free. In early May 2026 attackers systematically probed the organisation's publicly accessible infrastructure with heavy use of agents to automate vulnerability discovery; at the same time a read-only SQL exploit existed in the public transcript viewer, which an independent researcher found and responsibly disclosed. There is no evidence that the attackers discovered it or accessed any non-public data.