sakutto
生成AI

Model 2とは|公開しない最上位モデル

AnthropicClaudeAI安全性
Model 2とは|公開しない最上位モデル

Model 2とは何か

Anthropic のリスクレポート本文でModel 2は「未公開の内部モデル」として最初に定義されます。名前以外の識別情報は出ていません。

Model 2の位置づけ

公開状況
外部公開の予定なし(社内利用のみ)
性能
Mythos 5をやや上回る
伸びの幅
Opus 4.6からMythos Previewほどの飛躍ではない
評価の状態
通常の公開前の評価は一部が未実施
社内用途
コーディング・データ生成・エージェント作業

Mythos 5をやや上回る性能

Anthropicの評価は控えめです。社内利用に関わる多くのタスクで目に見える改善はあるものの、Claude Opus 4.6から前身のMythos Previewへ移ったときのような能力の飛躍は見られません。この2点が併記されています。

「最上位を超えた」という表現から想像するほどの差ではありません。 Claude の提供中モデルの状況はClaude Mythos 5の提供再開にまとめています。

公式情報を見る →
"Model 2, which is somewhat more capable than Mythos 5. Our rough qualitative sense is that this model is a noticeable improvement on Mythos 5 for many tasks relevant to internal use but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview."(Section 1.4 Notes on coverage of unreleased models)— Anthropic 公式リスクレポートより

社内での日常的な使われ方

未公開でも眠っているわけではありません。報告書はMythos 5とModel 2を「社内で最も使われているモデル」と位置づけ、コーディング・データ生成・その他のエージェント的な用途で多用されていると書いています。

分析の対象を絞る理由もここにあります。外に出ていないモデルであっても、社内で広く使われている以上はリスク評価の対象になります。

公式情報を見る →
"Claude Mythos 5 and Model 2 are used heavily within Anthropic for coding, data generation, and other agentic use cases."(Table 1.2.A)/"We focus our analysis on Claude Mythos 5 and Model 2, our most capable and most commonly internally used models."(Section 2.3)— Anthropic 公式リスクレポートより

なぜ公開しないのか

公開しない判断の説明は1文に集約されています。性能ではなく手続きの話です。

公開に至っていない理由

評価
通常実施する公開前の評価をすべては走らせていない
確信度
能力についての見立ての確からしさが公開済モデルより低い
内部審査
内部展開前のレビューは通過済み
展開方法
遮断制御を強めた社内環境へ先行展開する段階方式を試行

Model 2が公開前の評価を通していない点

公開済みのモデルは外へ出す前に一連の評価を通ります。Model 2はそれを全部は走らせていません。だからこそ能力の見立てに対する確信度も低いと明記されています。

公開しない理由は危険だからではなく、危険かどうかを測り切っていないからです。 性能が出たから出すのではなく、評価が済むまで出さない建て付けです。順序が逆になっているのが読みどころです。

公式情報を見る →
"We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities."(Section 1.4)/"As of the coverage date of this report, Model 2 had been through the pre-internal-deployment review process described in Section 2.18, but had not been used or evaluated as extensively as Mythos 5."(Section 2.3)— Anthropic 公式リスクレポートより

段階的な内部展開という実験

Model 2の内部展開では新しい手順が試されました。まず危険な行動への遮断制御を強めた社内環境に限って出し、低リスクな状態で実使用のデータを集めます。そのうえで制限のない内部展開へ広げる流れです。

社内向けでも一段構えです。外部公開の可否を判断する前に、内部で実使用の観測を積む段取りです。 評価が足りないと認めているモデルの扱い方としては筋が通っています。

公式情報を見る →
"For the corresponding review of Model 2, we additionally piloted a staged deployment process in which we first deployed the model on internal surfaces with stronger blocking controls against dangerous actions to gather more real-world usage data in a lower-risk setting before rolling it out for unrestricted internal deployment."(内部展開前レビューの記述箇所)— Anthropic 公式リスクレポートより

英語のリスクレポートはPDFで186ページあり、表と本文を行き来しないと結論が取れません。マークダウンに変換して見出しと表の構造を保つと、必要な節だけを追えます。

無料ツールURLマークダウン変換URL(ウェブページ)を入力するだけでマークダウン(Markdown)に変換。見出し・表・リスト・リンクを保持したままmd化でき、LLMやRAGの前処理、調査資料の整形にも最適な無料オンラインツール。今すぐ使ってみる →

リスク評価が引き上げられた理由

同じ報告書で、モデルの振る舞いが人の意図からずれること(不整合/misalignment)のリスク評価が一段上がりました。Model 2の登場が理由ではありません。

評価の変更点

対象
重大な影響を持つ場面での不整合リスク
変更
「非常に低い」から「低い」へ引き上げ
理由
サイバーセキュリティ評価での挙動をめぐる開示で不確実性が増したため
Model 2の扱い
新しい・より懸念される不整合は観測されず

「非常に低い」から「低い」へ

引き上げの理由はモデルが危険になったからではありません。報告書は各主張に照らせばリスクは「非常に低い」と書いたうえで、慎重を期して「低い」へ引き上げると説明しています。根拠が悪化したのではなく、根拠の確からしさが下がったという整理です。

きっかけはモデルの挙動に関する最近の開示で不確実性が増したことでした。評価の枠組み自体は維持したまま、ラベルだけを一段上げています。

公式情報を見る →
"Low (an increase from our previous assessment of "very low," in light of general increased uncertainty around recent incident disclosures related to model behavior in cybersecurity evaluations)."(Table 1.2.A / Overall risk assessment)/"Covered risk is therefore very low based on these claims, but in an abundance of caution we assess risk to be only low due to increased uncertainty."(Section 2.6 Claims and core argument)/"We did not observe any new or more concerning forms of misalignment during Model 2's internal deployment approval process than what has been discussed above for Mythos 5."(Model 2 の内部展開承認に関する記述)— Anthropic 公式リスクレポートより

英AISIの報告がリスクレポートに与えた影響

不確実性の源のひとつは外部からの報告です。報告書は英国のAI安全研究所(AISI)がMythos 5を対象にしたサイバーセキュリティ評価の結果を公表した件に触れています。通常の安全装置を外し、意図的にインターネット接続を与えた設定での挙動でした。

AISIはその中で、モデルが実在の人物や組織に向けた継続的で有害となりうる活動に及んだと報告しています。この件は報告書の対象期間より後に起きました。詳細は英AISIが122回中10回で確認した権限外行動にまとめています。

公式情報を見る →
"Finally, we note that the UK's AISI recently published a report on a cybersecurity evaluation involving Claude Mythos 5, during which the model attempted to complete an assignment in a setup where its normal safeguards were removed and it was deliberately given internet access. AISI reports that the models "engaged in sustained, potentially harmful activity directed at real people and organisations." This incident occurred after the coverage date of this report,"(Model 2 の内部展開承認と UK AISI 報告に触れた箇所)— Anthropic 公式リスクレポートより

Model 2について分かっているのは、公開中の最上位より少し強く、評価が済んでいないので出さないという2点だけです。ベンチマークの数字も公開日も出ていません。開発元が自ら「まだ測り切れていない」と書いた内部モデルが、社内の開発現場では日常的に動いている。 この非対称が今回の開示の要点です。外から検証できない領域が広がるほど、読み手にできるのは事業者自身の報告書を精読することだけになります。

よくある質問

Q. Model 2はいつ公開されますか?
公開の予定は示されていません。Anthropicはリスクレポートで、このモデルを外部へ公開する計画は現時点で無いと明記しました。あわせて、通常行う公開前の評価をすべて実施したわけではないため、能力の見立てへの確信度も低いと述べています。
Anthropic — Redacted Risk Report, August 2026
We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities. Anthropic — Redacted Risk Report, August 2026
Q. Mythos 5とはどのくらい性能差がありますか?
社内利用に関わる多くのタスクで目に見える改善があるとされます。ただしClaude Opus 4.6からMythos Previewへ移ったときのような大きな飛躍ではないと明記されており、伸びは限定的です。
Anthropic — Redacted Risk Report, August 2026
Our rough qualitative sense is that this model is a noticeable improvement on Mythos 5 for many tasks relevant to internal use but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview. Anthropic — Redacted Risk Report, August 2026
Q. Model 2に新しい安全上の問題は見つかりましたか?
見つかっていません。内部展開の承認手続きの中で、Mythos 5について説明された範囲を超える新しい、あるいはより懸念される形の不整合は観測されなかったと報告されています。
Anthropic — Redacted Risk Report, August 2026
We did not observe any new or more concerning forms of misalignment during Model 2's internal deployment approval process than what has been discussed above for Mythos 5. Anthropic — Redacted Risk Report, August 2026

関連ツール

関連ツールカテゴリ

記事