Anthropic's first Risk Report
On 24 February 2026 Anthropic published "Risk Report: February 2026", a separate document and the first under version 3 of its scaling policy: an assessment of catastrophic risk from all its models, internal ones included, in redacted form.
Why it matters
A developer published a written risk assessment of its whole activity rather than of one model, with a promise to repeat it. An editorial assessment: the assessments in it are the company's own, not an independent finding.
What the document names. It is not a restatement of the policy but a separate 104-page text that the policy has the company publish every 3-6 months. It covers all of the company's models, including those run only internally. Four risk categories and the author's overall assessments: sabotage of safety research by AI, "very low but not negligible"; automated research and development, "very low"; non-novel chemical and biological weapons, "very low but not negligible"; novel chemical and biological weapons, "low, but with substantial uncertainty". The analysis centres on Claude Opus 4.6. Some content is redacted for intellectual property and public safety. A standalone sabotage risk report for Opus 4.6 is nearly identical to section 2. Later revisions. The address anthropic.com/feb-2026-risk-report now leads to a revised copy: its change log lists edits of 26 May 2026 (after a pilot external review by METR) and 8 July 2026; the record rests on the original file of 24 February. Is it the same text as the policy (evt-0370)? No, it is a separate document. What the record does not claim. No assessment was checked against the company's underlying data; no independent assessment within February was found. The outside report said the document "confirmed the safety" of the system; the document does not say that, it rates the risk as low and names uncertainty.