sakutto
Generative AI

Anthropic Model 2: No Plans to Release It

AnthropicAI SafetyFrontier Models
Anthropic Model 2: No Plans to Release It

What Model 2 is, and where it runs

The disclosure sits inside a governance document, not a launch post. That framing sets the tone: the model appears as something being assessed rather than something being sold.

The model, as described by Anthropic

Capability
Somewhat more capable than Mythos 5; no jump of the size seen from Claude Opus 4.6 to Mythos Preview
Availability
No current plans for external release
Assessment
Has not run the full predeployment suite, so Anthropic holds lower confidence in its own capability estimates
Internal use
Used heavily for coding, data generation and other agentic work

More capable than Mythos 5, and deliberately unshipped

A better model exists and is not for sale. Anthropic's own wording is measured — noticeably better for many internal tasks, without the size of jump seen in an earlier generation — which reads as a description written for a regulator rather than a buyer.

The reason to take that phrasing at face value is the venue. A company overstating a model it will not sell gains nothing and risks its own assessment framework, so restraint is the expected register here.

View official source →
"Model 2, which is somewhat more capable than Mythos 5. Our rough qualitative sense is that this model is a noticeable improvement on Mythos 5 for many tasks relevant to internal use but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview."(Section 1.4)/"We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities."(Section 1.4)/"Claude Mythos 5 and Model 2 are used heavily within Anthropic for coding, data generation, and other agentic use cases."(Table 1.2.A, Current usage and capabilities) — Anthropic Risk Report

Rolled out internally in stages

The internal deployment was not a switch being flipped. Anthropic describes piloting a staged process for Model 2: restricted surfaces with stronger blocking controls first, to collect real usage in a lower-risk setting, then unrestricted internal deployment.

That is a release process, applied to a release that never becomes public. It is also the clearest signal in the section that the company treats its own staff as a population worth protecting from an insufficiently assessed model.

View official source →
"For the corresponding review of Model 2, we additionally piloted a staged deployment process in which we first deployed the model on internal surfaces with stronger blocking controls against dangerous actions to gather more real-world usage data in a lower-risk setting before rolling it out for unrestricted internal deployment." — Section 2.18 Pre-internal-deployment review process, Anthropic Risk Report

CoBench: what it measures, and the 85% bar

Anthropic's older task-based R&D evaluation has run out of headroom: frontier models now beat the human baseline on most of it. CoBench is the replacement, and its design says something about what the company is watching for.

How the evaluation is built

Input
A snapshot of Anthropic's codebase, logs, internal messaging and docs at a past timestamp
Task
Diagnose the root cause of an issue Anthropic engineers went on to solve
Size
449 problems, largely sourced from issues solved between February and April 2026
Filtering
Mostly restricted to problems Mythos Preview failed at least once in three tries
View official source →
"The task-based AI R&D evaluation suite we have reported on in our system cards has reached the point where frontier models surpass human baseline performance on most tasks." — Section 3.4.3 CoBench, Anthropic Risk Report

Real engineering problems instead of proxies

The evaluation is deliberately unfair in a useful way. Problems are filtered for difficulty, and the grading compares an answer against the root cause Anthropic identified in practice — which the model cannot see in its snapshot.

Widely circulated percentage scores for Model 2 on this benchmark do not appear in the report text. The report presents CoBench results as a figure and describes relative performance in prose, so any specific number quoted elsewhere should be traced before it is repeated.

View official source →
"Our current version of this evaluation on which we report results below comprises 449 problems from parts of the technical organization most directly relevant to training models and running the infrastructure needed to do so, largely sourced from issues that were solved between February 2026 and April 2026."(Section 3.4.3)/"The evaluation is not a representative distribution of all such tasks, as it is moderately filtered for difficulty: the dataset is mostly restricted to problems that Mythos Preview failed to solve at least once in three tries, and without such filtering the dataset would be roughly twice as large. Problems are model-graded using a rubric that compares a solution to the root cause we identified in practice (which is not visible from the historical snapshot available to the evaluated model)."(Section 3.4.3) — Anthropic Risk Report

The bar Anthropic set for full substitution

There is a stated threshold for the thing everyone is actually asking about. Anthropic judges that a model genuinely able to substitute for its research staff would score at least 85% on CoBench, and says its Mythos-class models still fall short of that.

That number is the useful one to carry, because it converts a vague question about replacement into a published, checkable line. Whatever the current scores are, the company has committed to what passing would look like.

View official source →
"we think our validation, scaffolding, and grading rubrics for these tasks are reliable enough that a model which was truly capable of fully substituting for Anthropic research staff would be able to score at least 85% on this evaluation."(Section 3.4.3)/"Our Mythos-class models perform substantially better than other recent models, though still fall short of the performance we would expect from a system that could fully substitute for the work of Anthropic technical staff (see below)."(Figure 3.4.3.A caption) — Anthropic Risk Report

What the report discloses about itself

A risk report that grades its own company is only as good as its disclosure rules. This one spends its opening on them.

Disclosure mechanics

Scope
Coverage date 15 July 2026, covering the period since the previous report of 24 February 2026
Internal
Minimally redacted copies to all staff; fully unredacted copies to at least 200 employees
Public
Redactions in the public version are noted in the document itself

Coverage dates make the gaps legible

A report with a stated coverage date can be read for what it excludes. This one closes on 15 July 2026, and says so, which is why later incidents surface in it as explicit notes rather than as silent omissions.

That matters for anyone relying on the document. Reading it as a current snapshot of a lab's risk position would be a mistake the report itself takes pains to prevent.

View official source →
"This Risk Report's coverage date is July 15, 2026. It discusses the period from the February 24, 2026 publication of our previous Risk Report until this date" — Section 1.3.3 Coverage dates of risk reports, Anthropic Risk Report

Where the redactions are, and who sees through them

The report names one location for the cuts made to the all-staff version: Section 3.5, on compartmentalized and commercially sensitive AI R&D detail. Redactions in the public version are marked in place.

Whether that is enough is arguable, and the document does not pretend otherwise. What it does establish is a floor: a reader can see that something was withheld and where, instead of inferring it from a gap in the prose.

View official source →
"It also requires that fully unredacted risk reports be shared with at least 200 Anthropic employees (rather than with all regular-clearance Anthropic staff, as previous RSP versions required)."(Section 1.3.4)/"In this Risk Report, the only redactions we have made for the version shared with all regular-clearance Anthropic staff are located in Section 3.5, where we discuss internally compartmentalized and commercially sensitive details of our AI R&D development process."(Section 1.3.4)/"All redactions made for the public version of the report are noted below."(Section 1.3.4) — Anthropic Risk Report

For a case where an automated review passed a critical flaw instead of catching it, see GitHub Actions Script Injection: The Snowflake Bug. For a company moving in the opposite direction on risk assessment, see OpenAI Preparedness Team: What the FT Reported.

Reading a 186-page PDF alongside the coverage that summarizes it is easier once both are plain text.

Free ToolURL to Markdown ConverterConvert any public web page URL to Markdown. Preserves headings, tables, lists, and links — perfect for LLM and RAG preprocessing, research notes, and archiving web articles.Try it now →

FAQ

Q. Can I use Model 2?
No, and Anthropic gives no timeline. The report states there are no current plans to release it externally, and adds that it has not been through the company's full predeployment assessment suite, which leaves Anthropic itself with lower confidence about what the model can do.
Anthropic — Redacted Risk Report, August 2026
We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities. Anthropic — Redacted Risk Report, August 2026
Q. What does CoBench actually test?
Debugging against Anthropic's own history. The model is handed a snapshot of the codebase, logs, internal messages and docs from a past date, then asked to find the root cause of an issue Anthropic engineers went on to solve. It replaces proxy tasks with work the company actually did.
Anthropic — Redacted Risk Report, August 2026
CoBench is an internal evaluation measuring how well a model, placed at a historical point in Anthropic's infrastructure (that is, given a snapshot of our codebase, logs, internal messaging, and docs at a past timestamp), can diagnose the root causes of issues that Anthropic engineers actually solved. Anthropic — Redacted Risk Report, August 2026
Q. How much of the public report is withheld?
The report says redactions are marked, and that the version circulated to regular-clearance staff is cut in one place only: Section 3.5, covering compartmentalized and commercially sensitive detail about AI R&D. The policy version behind the report also requires that fully unredacted copies reach at least 200 employees.
Anthropic — Redacted Risk Report, August 2026
In this Risk Report, the only redactions we have made for the version shared with all regular-clearance Anthropic staff are located in Section 3.5, where we discuss internally compartmentalized and commercially sensitive details of our AI R&D development process. Anthropic — Redacted Risk Report, August 2026

Related Tools

Related Tool Categories

Articles