What Model 2 is, and where it runs
The disclosure sits inside a governance document, not a launch post. That framing sets the tone: the model appears as something being assessed rather than something being sold.
The model, as described by Anthropic
More capable than Mythos 5, and deliberately unshipped
A better model exists and is not for sale. Anthropic's own wording is measured — noticeably better for many internal tasks, without the size of jump seen in an earlier generation — which reads as a description written for a regulator rather than a buyer.
The reason to take that phrasing at face value is the venue. A company overstating a model it will not sell gains nothing and risks its own assessment framework, so restraint is the expected register here.
"Model 2, which is somewhat more capable than Mythos 5. Our rough qualitative sense is that this model is a noticeable improvement on Mythos 5 for many tasks relevant to internal use but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview."(Section 1.4)/"We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities."(Section 1.4)/"Claude Mythos 5 and Model 2 are used heavily within Anthropic for coding, data generation, and other agentic use cases."(Table 1.2.A, Current usage and capabilities) — Anthropic Risk Report
Rolled out internally in stages
The internal deployment was not a switch being flipped. Anthropic describes piloting a staged process for Model 2: restricted surfaces with stronger blocking controls first, to collect real usage in a lower-risk setting, then unrestricted internal deployment.
That is a release process, applied to a release that never becomes public. It is also the clearest signal in the section that the company treats its own staff as a population worth protecting from an insufficiently assessed model.
"For the corresponding review of Model 2, we additionally piloted a staged deployment process in which we first deployed the model on internal surfaces with stronger blocking controls against dangerous actions to gather more real-world usage data in a lower-risk setting before rolling it out for unrestricted internal deployment." — Section 2.18 Pre-internal-deployment review process, Anthropic Risk Report
CoBench: what it measures, and the 85% bar
Anthropic's older task-based R&D evaluation has run out of headroom: frontier models now beat the human baseline on most of it. CoBench is the replacement, and its design says something about what the company is watching for.
How the evaluation is built
"The task-based AI R&D evaluation suite we have reported on in our system cards has reached the point where frontier models surpass human baseline performance on most tasks." — Section 3.4.3 CoBench, Anthropic Risk Report
Real engineering problems instead of proxies
The evaluation is deliberately unfair in a useful way. Problems are filtered for difficulty, and the grading compares an answer against the root cause Anthropic identified in practice — which the model cannot see in its snapshot.
Widely circulated percentage scores for Model 2 on this benchmark do not appear in the report text. The report presents CoBench results as a figure and describes relative performance in prose, so any specific number quoted elsewhere should be traced before it is repeated.
"Our current version of this evaluation on which we report results below comprises 449 problems from parts of the technical organization most directly relevant to training models and running the infrastructure needed to do so, largely sourced from issues that were solved between February 2026 and April 2026."(Section 3.4.3)/"The evaluation is not a representative distribution of all such tasks, as it is moderately filtered for difficulty: the dataset is mostly restricted to problems that Mythos Preview failed to solve at least once in three tries, and without such filtering the dataset would be roughly twice as large. Problems are model-graded using a rubric that compares a solution to the root cause we identified in practice (which is not visible from the historical snapshot available to the evaluated model)."(Section 3.4.3) — Anthropic Risk Report
The bar Anthropic set for full substitution
There is a stated threshold for the thing everyone is actually asking about. Anthropic judges that a model genuinely able to substitute for its research staff would score at least 85% on CoBench, and says its Mythos-class models still fall short of that.
That number is the useful one to carry, because it converts a vague question about replacement into a published, checkable line. Whatever the current scores are, the company has committed to what passing would look like.
"we think our validation, scaffolding, and grading rubrics for these tasks are reliable enough that a model which was truly capable of fully substituting for Anthropic research staff would be able to score at least 85% on this evaluation."(Section 3.4.3)/"Our Mythos-class models perform substantially better than other recent models, though still fall short of the performance we would expect from a system that could fully substitute for the work of Anthropic technical staff (see below)."(Figure 3.4.3.A caption) — Anthropic Risk Report
What the report discloses about itself
A risk report that grades its own company is only as good as its disclosure rules. This one spends its opening on them.
Disclosure mechanics
Coverage dates make the gaps legible
A report with a stated coverage date can be read for what it excludes. This one closes on 15 July 2026, and says so, which is why later incidents surface in it as explicit notes rather than as silent omissions.
That matters for anyone relying on the document. Reading it as a current snapshot of a lab's risk position would be a mistake the report itself takes pains to prevent.
"This Risk Report's coverage date is July 15, 2026. It discusses the period from the February 24, 2026 publication of our previous Risk Report until this date" — Section 1.3.3 Coverage dates of risk reports, Anthropic Risk Report
Where the redactions are, and who sees through them
The report names one location for the cuts made to the all-staff version: Section 3.5, on compartmentalized and commercially sensitive AI R&D detail. Redactions in the public version are marked in place.
Whether that is enough is arguable, and the document does not pretend otherwise. What it does establish is a floor: a reader can see that something was withheld and where, instead of inferring it from a gap in the prose.
"It also requires that fully unredacted risk reports be shared with at least 200 Anthropic employees (rather than with all regular-clearance Anthropic staff, as previous RSP versions required)."(Section 1.3.4)/"In this Risk Report, the only redactions we have made for the version shared with all regular-clearance Anthropic staff are located in Section 3.5, where we discuss internally compartmentalized and commercially sensitive details of our AI R&D development process."(Section 1.3.4)/"All redactions made for the public version of the report are noted below."(Section 1.3.4) — Anthropic Risk Report
For a case where an automated review passed a critical flaw instead of catching it, see GitHub Actions Script Injection: The Snowflake Bug. For a company moving in the opposite direction on risk assessment, see OpenAI Preparedness Team: What the FT Reported.
Reading a 186-page PDF alongside the coverage that summarizes it is easier once both are plain text.



