sakutto
生成AI

Cerebras CS-4が推論速度をGPU比で最大30倍に伸ばす

CerebrasAI半導体AI推論
Cerebras CS-4が推論速度をGPU比で最大30倍に伸ばす

Cerebras CS-4とは何か

Cerebras CS-4 は Cerebras Systems が2026年8月18日に発表した第4世代のシステムです。同社はこれを業界最速のAIアクセラレータであり、最先端AIを動かすための新しい土台だと位置づけています。単体のチップではありません。ラック1本ぶんの製品で、演算・電力・冷却・I/O をまとめて設計し直した点が前世代との最大の違いです。

CS-4 の基本仕様(ラック全体・3ウェハー構成)

発表日
2026年8月18日(Cerebras Systems・NASDAQ: CBRS)
構成
Wafer Scale Engine 3 Turbo(WSE-3T)×3枚のラックスケール機
AI演算
750 PFLOPS(1秒あたりの演算回数の単位。CS-3は1ウェハーで125 PFLOPS)
メモリ帯域
129.6 ペタバイト毎秒(CS-3は21.6)
システムI/O
7.2 テラビット毎秒(CS-3は1.2)
I/Oレイテンシ
最短2マイクロ秒(CS-3は5マイクロ秒)
出荷
最初の出荷は発表と同じ四半期に開始
"Today, we are introducing the fourth generation of our Cerebras System: CS-4. The fastest AI accelerator in the industry, and a new foundation for frontier AI. Built from three new Wafer Scale Engine 3 Turbo processors, CS-4 pairs a more powerful processor with a completely redesigned new rack and system."(公式ブログ冒頭)— Cerebras 公式ブログより。なお上表の CS-3 側の値(AI演算125 PFLOPS/メモリ帯域21.6 PByte/s/システムI/O 1.2 Tbit/s/I/Oレイテンシ5マイクロ秒)は、プレスリリース「CS-3 vs. CS-4: Key Specifications」表の CS-3 列に各セルとして掲載されている
公式情報を見る →

ウェハースケール3枚を束ねた初のNexus機

ウェハースケールとは、切り分けた小さなチップではなくウェハー1枚をまるごと1個のプロセッサとして使う方式です。CS-4 はそのウェハーを3枚束ねた、新しいラック基盤 Cerebras Nexus Platform Architecture の第1号機にあたります。Nexus はラックを演算・電力・I/O の3要素へ分け、それぞれを独立に更新できるようにした設計思想を指します。部品点数は前世代比で50%減り、演算・電力・I/O がそれぞれ完結したアセンブリになったことで、製造と設置が速くなりました。 一部だけを新しくして出荷できるため、次の世代を待たずに改良が顧客へ届く構造でもあります。

ラック1本としての数字は上表のとおりです。ラック全体の演算ファブリックの総帯域は160.5ペタバイト毎秒に上がり、ウェハー間のレイテンシは最短2マイクロ秒まで下がりました。この低遅延があるのでかなり大きなクラスタまで組めると同社は言います。

"The CS-4 is a rack-scale solution built from three new Wafer Scale Engines and revolutionary rack and system designs. The CS-4 is the first member of the next-generation Cerebras Nexus rack-scale platform architecture."/"The rack scale CS-4 is built from three of the newly released Wafer Scale Engine 3 Turbo (WSE-3T). The CS-4 delivers 750 PFLOPs of AI compute, 7.2 terabits per second of I/O, and 129.6 petabytes per second of memory bandwidth. Total compute fabric bandwidth jumps to 160.5 petabytes per second, and wafer to wafer latency drops as low as two microseconds, enabling the creation of very large clusters and the support of models with over 50 trillion parameters."(冒頭/システム諸元段落)— Cerebras プレスリリースより
公式情報を見る →
"With 50 percent fewer components and self-contained assemblies for compute, power, and I/O, CS-4 supports faster manufacturing and rapid deployment. Instead of treating the rack as a tightly coupled collection of parts, Nexus turns it into a platform of purpose-built modules."/"That platform approach also creates a cleaner path for upgrades, allowing innovation in one part of the system to reach customers without waiting for every other part to be redesigned."(The Nexus rack-scale platform 節)— Cerebras 公式ブログより
公式情報を見る →

WSE-3 Turboは4兆トランジスタの新プロセッサ

WSE-3 Turbo(WSE-3T)は CS-4 に載る新しいプロセッサで、前世代の WSE-3 と同じく史上最大のAIプロセッサです。46,225平方ミリのシリコンに4兆個のトランジスタと90万個のAI最適化コアを載せ、44GBのSRAM(プロセッサに直結した高速メモリ)をウェハー上に直接持ちます。規模そのものは据え置きで、伸びたのは性能側です。 ウェハーあたりのAI演算は250 PFLOPSへ、メモリ帯域は43.2ペタバイト毎秒へ、いずれも倍になりました。コアどうしを結ぶ内部配線(ファブリック)の帯域は53.5ペタバイト毎秒、チップ外との I/O は2.4テラビット毎秒で、こちらも倍増しています。

速度と処理量を決めるのはメモリ帯域だと同社は説明しています。その帯域が倍になったのだから改善も段違いになるという理屈です。

"Like the WSE-3, the WSE-3T is the largest AI processor ever built, containing four trillion transistors and 900,000 AI-optimized cores across 46,225 square millimeters of silicon, with 44GB of SRAM integrated directly on the wafer."/"The WSE-3T doubles AI compute to 250 PFLOPS per wafer and doubles memory bandwidth to 43.2 petabytes per second. And since memory bandwidth is the determining factor in driving speed and throughput, this translates to a step-change improvement in both. The on-chip fabric bandwidth and off-chip I/O both double to 53.5 petabytes per second and 2.4 terabits per second respectively."(A New Processor 節)— Cerebras プレスリリースより
公式情報を見る →

「GPU比30倍」はどの条件で出た数字か

30倍が出たのは GPT-OSS-120B というモデルで同一のプロンプトを与えた直接比較です。ユーザーあたり毎秒4,400トークン超という値がその根拠にあたります。条件と注記は公式の2文書に書かれていて、自分が動かす処理でそのまま出る数字ではありません。

4,400トークン/秒はGPT-OSS-120Bでの値

比較に使われたモデルは GPT-OSS-120B です。同一のプロンプトを与えた直接比較で、CS-4 はユーザーあたり毎秒4,400トークン超を出し、これがGPUソリューション比で最大30倍にあたるとされています。注記もはっきり書かれていて、実際のスループットはモデル構造・コンテキスト長・精度・サービング構成によって変わります。 30倍という数字はこの1モデル・この構成での上限値だと読むのが正確です。

同社CTOの Sean Lie はこう述べています。30倍という差は応答が速く感じられるというだけの話ではない。同じ実時間のなかで1桁以上多い推論・検証・ツール利用をエージェントに許すことだ、と。

"In a head-to-head comparison on GPT-OSS-120B, when given identical prompts, the CS-4 delivers [1] more than 4,400 tokens second per user (TPS/user), up to 30 times faster than GPU solutions."/"Being 30 times faster doesn’t just make a response feel fast. It gives an agentic system room for more than an order of magnitude as much reasoning, verification, or tool use in the same wall-clock time,"/"Actual throughput varies by model architecture, context length, precision, and serving configuration."(A New Leader in AI Inference Speed and Capacity 節/脚注[1])— Cerebras プレスリリースより
公式情報を見る →

出どころは社内ベンチと第三者計測の併記

数字の根拠についても公式ブログは出所を明記しています。速度比較のグラフの出所は第三者ベンチマーク機関の Artificial Analysis と社内ベンチマーク(2026年8月)。大規模モデルでの毎秒1,000トークン超という値のほうは社内ベンチマークからの外挿だと注記されています。第三者だけで独立に再現された数字ではありません。

対応できるモデル規模の書き方は2つの公式文書で違います。ブログは10兆パラメータを超えるモデルで毎秒1,000トークン超と書き、プレスリリースは50兆パラメータを超えるモデルに対応できると書いています。前者は速度、後者はクラスタとして扱える上限の話で、同じ土俵の数字ではありません。CS-3 と比べると速度は最大2倍、電力あたりのスループットは最大10倍というのが公式の数字です。

"Source: Artificial analysis and internal benchmarking (August 2026)"/"With this low-latency communication, CS-4 can deliver more than 1,000 tokens per second on models exceeding 10 trillion parameters."/"Low-latency wafer-to-wafer communication preserves interactive decode performance as model size grows. Source: extrapolation from internal benchmarking (August 2026)."/"The CS-4 solution generates tokens up to 30 times faster than production GPU systems while delivering up to 10 times more throughput per watt than CS-3."/"CS-4 expands the ultrafast inference frontier with up to 10x more token capacity and up to 2x faster performance than CS-3. Source: internal benchmarking and projections (August 2026)."(Engineered for speed at scale 節)— Cerebras 公式ブログより
公式情報を見る →

公式ブログとプレスリリースは同じ製品を説明しながら、載せている数字と条件が微妙に食い違います。両方の本文を貼り付けて文字単位で差分を取ると、どこが言い換えでどこが別の主張なのかが一目で分かります。

無料ツールテキスト差分比較2つのテキストの差分をハイライト表示。変更点を素早く発見できます。今すぐ使ってみる →

速さはラック全体の作り直しから来ている

CS-4 の性能向上はプロセッサ1点の改良で説明できるものではありません。同社の説明も「演算・電力・冷却・I/O が同時に前へ進まなければ次の飛躍は来ない」という前提から始まっています。

"CS-4 is designed around a simple idea: the next leap in AI infrastructure cannot come from improving one component in isolation. Compute, power, cooling, and I/O have to move forward together."— Cerebras 公式ブログより
公式情報を見る →

Backpack構造で設置は数日から数時間へ

Nexus の中心にあるのが背面に垂直装着する Wafer-Scale Backpack です。電力変換・直接液冷・高速I/O・制御電子回路を、ウェハーを取り巻く立体的なパッケージへ折りたたんだ自己完結型のアセンブリになっています。演算を電源から切り離した結果、製造が単純になり、設置にかかる時間が数日から数時間へ短縮されたとしています。前世代比で部品点数は50%減、自動化製造の比率は60%増です。

電力の届け方も変わりました。従来のGPUボードではプロセッサからおよそ50ミリメートル離れていた電力変換を約0.5ミリメートルまで、つまり100倍の近さへ寄せています。基板レベルの電力損失がほぼ消えるため、WSE-3T へ2倍の電力を供給でき、動作周波数を上げてトークン生成を速められるという流れです。公式文書は製造プロセスの微細化に一言も触れておらず、説明の力点はもっぱら電力を届ける経路そのものに置かれています。

"By decoupling compute from the power supplies, the Wafer-Scale Backpack simplifies manufacturing and reduces deployment time from days to hours. Compared with the prior-generation system, the Wafer-Scale Backpack has 50% fewer components and uses 60% more automated manufacturing."/"For example, by moving power conversion 100x closer to the processors – from roughly 50 millimeters away from the processor as on conventional GPU boards to approximately 0.5 millimeters – CS-4 nearly eliminates board-level power loss. This delivers twice as much power to the WSE-3T, enabling higher operating frequencies and faster token generation."(New System and Rack Design 節)— Cerebras プレスリリースより
公式情報を見る →

prefillを他社基盤へ渡す分離推論に対応した

もうひとつの設計上の選択が分離推論(disaggregated inference)への対応です。推論はプロンプトを読み込む prefill と、応答を1トークンずつ吐き出す decode の2段階に分かれます。CS-4 はこの2つを別々の基盤に割り当てる構成を標準で想定しました。専用の prefill エンジンがプロンプトを処理してモデルの状態を作り、それを CS-4 へ渡して超低遅延の decode を担わせるという分業です。

組み合わせ先として名前が挙がっているのは AMD Helios(ヘリオス)と AWS Trainium(トレイニウム)です。GPUや ASIC(用途を絞って作った専用チップ)で効率よく prefill を回し、decode だけ CS-4 に任せます。自社製品だけで固めず異種混成を前提にした設計です。接続はイーサネット上の標準規格 RoCE v2 RDMA と、スイッチを介さずウェハー同士を直結する Direct Wafer Links の2モードが用意されています。

超高速な推論を製品体験へ落とし込む動きはモデル提供側でも進んでいます。OpenAI が限定プレビューで出したGPT-5.6 Sol Ultrafastの解説は、同じ「速さそのものを売る」方向の例です。推論専用ハードを自社で持つ流れはOpenAIとBroadcomの推論チップ Jalapeño(ハラペーニョ)の記事で扱っています。

"First, a purpose-built prefill engine processes the incoming prompt and prepares the model state. That state is then transferred to CS-4, where the system performs ultra-low-latency decoding and generates the response."/"For operators, this architecture combines industry-leading Cerebras decode performance with the flexibility to pair CS-4 with complementary prefill platforms, including AMD Helios and AWS Trainium."/"Standards-based RoCE v2 RDMA over Ethernet provides a familiar way to connect CS-4 with existing infrastructure and with an ecosystem of heterogeneous systems. Direct Wafer Links provide for switch-free connections within and across racks."(Native support for disaggregated inference / A new modular I/O subsystem 節)— Cerebras 公式ブログより
公式情報を見る →

まとめ

Cerebras CS-4 は WSE-3 Turbo を3枚束ねたラックスケール機で、ラック全体で750 PFLOPS を出します。目を引く「GPU比最大30倍」は GPT-OSS-120B の直接比較で毎秒4,400トークン超という条件の値です。モデルや構成が変われば数字も変わると同社自身が注記しています。公式が速さの理由として挙げているのは、電力変換をプロセッサへ100倍近づけて供給電力を倍にしたラック設計の作り直しです。最初の出荷は発表と同じ四半期に始まる予定ですが、価格は公表されていません。 大規模モデルでの数値には社内ベンチマークからの外挿が含まれるため、導入を検討するなら自分のモデルと構成での実測を前提にしたほうが安全です。

よくある質問

Q. Cerebras CS-4はいつから使えますか?
最初の出荷は発表と同じ四半期に始まるとされています。具体的な提供開始日や価格は公式発表に含まれていません。詳細仕様は同社が公開しているCS-4データシートで確認できます。
Cerebras — Cerebras Unveils CS-4(Availability)
First CS-4 shipments begin this quarter. Cerebras — Cerebras Unveils CS-4(Availability)
Q. 「GPU比30倍」は何と比べた数字ですか?
GPT-OSS-120Bを使い同一プロンプトを与えた直接比較で、CS-4がユーザーあたり毎秒4,400トークン超を出したという条件での値です。Cerebrasは注記で、実際のスループットはモデル構造・コンテキスト長・精度・サービング構成によって変わるとしています。
Cerebras — Cerebras Unveils CS-4(A New Leader in AI Inference Speed and Capacity/脚注[1])
In a head-to-head comparison on GPT-OSS-120B, when given identical prompts, the CS-4 delivers [1] more than 4,400 tokens second per user (TPS/user), up to 30 times faster than GPU solutions. / Actual throughput varies by model architecture, context length, precision, and serving configuration. Cerebras — Cerebras Unveils CS-4(A New Leader in AI Inference Speed and Capacity/脚注[1])
Q. WSE-3 Turboは前世代のWSE-3と何が違いますか?
トランジスタ数やコア数といった規模は同じで、性能側が倍になっています。ウェハーあたりのAI演算は250 PFLOPSへ、メモリ帯域は43.2ペタバイト毎秒へそれぞれ倍増しました。I/Oレイテンシも5マイクロ秒から最短2マイクロ秒へ縮んでいます。
Cerebras — Cerebras Unveils CS-4(A New Processor)
The WSE-3T doubles AI compute to 250 PFLOPS per wafer and doubles memory bandwidth to 43.2 petabytes per second. / I/O latency shrinks from five microseconds to as low as two microseconds. Cerebras — Cerebras Unveils CS-4(A New Processor)

関連ツール

関連ツールカテゴリ

記事