sakutto
Generative AI

GPT-5.6 Sol Goes Up to 14× Faster: Inside the New Ultrafast Mode

OpenAIGPT-5.6GPT-5.6 Sol
GPT-5.6 Sol Goes Up to 14× Faster: Inside the New Ultrafast Mode

What Ultrafast is

Ultrafast at a glance

Speed
Up to 14× Standard processing / up to 750 tokens per second
Foundation
Cerebras
Availability
OpenAI API, limited preview

Ultrafast is a new service tier announced by OpenAI on August 13, 2026.

A new tier at up to 14× Standard

Ultrafast runs GPT-5.6 Sol at up to 14× the speed of Standard processing. As the word "tier" suggests, this is not a swap to a different model but a new speed grade for running the same top-end model. It launches first in the OpenAI API. For where GPT-5.6 Sol itself sits, see the GPT-5.6 Sol explainer.

View official source →
"Today, we're sharing an early look at Ultrafast, a new service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing, launching first in the OpenAI API." — from the OpenAI official announcement

Removing the speed-or-intelligence choice

This is the crux of the release. Until now, if you wanted real-time responses, the only option was to pick a smaller or more specialized model. OpenAI describes Ultrafast as progress toward more useful work per second. If you can take speed without giving up intelligence, AI can move into work you had written off because of the wait.

View official source →
"Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second."/"When speed no longer requires giving up intelligence, AI can move into the most time-sensitive parts of a business and new kinds of work become possible."(both from the opening explanatory paragraphs) — from the OpenAI official announcement

Where the speed comes from, and where it fits

Use cases named by OpenAI

Incident response
Read logs and recent changes to identify a cause and prepare a fix
Finance and fraud
Assess markets and transactions while conditions are still moving
Voice and support
Return multi-step answers without interrupting the conversation
Commerce
Check inventory and clear checkout issues before the shopper leaves
Live research
Turn an overnight run into an interactive working session

Here is the source of the speed and the use cases the announcement names.

750 tokens per second, powered by Cerebras

The foundation is Cerebras, and it is stated to generate up to 750 output tokens per second. Cerebras is OpenAI's partner for ultra-low-latency inference (the process by which an AI produces an answer). A token is a small unit that text is split into, so 750 tokens per second means long passages arrive with essentially no waiting.

View official source →
"Powered by Cerebras, Ultrafast generates up to 750 output tokens per second, bringing our most intelligent model to products and workflows where every second matters."(opening paragraph)/"Ultrafast marks the next step in our partnership with Cerebras to bring ultra-low-latency inference to OpenAI's platform."(Powered by Cerebras section) — from the OpenAI official announcement

Where seconds change the outcome

The announcement names five areas. Incident response, financial research and security, customer support and voice, commerce, and live research and experimentation. All of them are work where the situation changes before an answer arrives, so speed directly changes the result. For incident response, that means reading logs and recent code changes to narrow down a cause while the outage is still unfolding.

View official source →
"Incident response and reliability: When a critical system fails, analyze application logs, recent code changes, and engineer reports to identify the likely cause and help prepare a fix while the outage is still unfolding."/"Financial research and security: Analyze market signals, assess transactions, and identify suspicious activity while conditions are still changing."/"Customer support and voice: Resolve complex customer issues in real time without interrupting the conversation, even when finding the answer requires multiple steps or systems."/"Commerce: Answer product questions, check inventory, personalize recommendations, and resolve checkout issues while the shopper is still deciding, before hesitation becomes an abandoned cart."/"Live research and experimentation: Turn research that previously took an overnight run into an interactive working session, letting teams test an idea, examine the results, adjust their approach, and run another experiment without breaking their flow."(all five bullets from the use-case list) — from the OpenAI official announcement

Who can use it, and the takeaway

Availability

Now
Limited preview, select customers only
Later
Access expands as capacity grows

Finally, whether you can actually use it today.

Still a limited preview

Ultrafast is a limited preview, open only to a selected group of customers. It is not something anyone can call through the API. OpenAI says access will expand as capacity grows, and the official page offers a sign-up to be notified when it does. If you are considering it, working out in advance which of your own processes change value by the second will let you move when a slot opens.

View official source →
"GPT-5.6 Sol on Ultrafast mode is available in a limited preview today to a select group of customers. We'll expand access as capacity grows."(Availability section)/"If your business requires frontier intelligence at the highest speed, you can sign up to get notified when access expands."(opening explanatory paragraph) — from the OpenAI official announcement

How OpenAI uses it internally

It is already in use inside the company. Research team members used to launch a batch of experiments overnight and review results the next morning. With Ultrafast, OpenAI says that loop tightens enough to support multiple iterations during the workday. Verification that took a night becomes interactive work.

View official source →
"A common workflow in research is for our team members to launch a batch of experiments over night, and review the results in the morning. With Ultrafast, we see this loop tightening to support multiple iterations during the workday instead." — from the OpenAI official announcement

When you want to read an announcement page on your own machine as markdown with its headings and tables intact, the following tool can help.

Free ToolURL to Markdown ConverterConvert any public web page URL to Markdown. Preserves headings, tables, lists, and links — perfect for LLM and RAG preprocessing, research notes, and archiving web articles.Try it now →

FAQ

Q. What is Ultrafast?
It is a new service tier OpenAI announced on August 13, 2026. It runs GPT-5.6 Sol up to 14× faster than Standard processing, and it is launching first in the OpenAI API.
OpenAI Official Announcement
Today, we're sharing an early look at Ultrafast, a new service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing, launching first in the OpenAI API. OpenAI Official Announcement
Q. How fast is Ultrafast?
Powered by Cerebras, it is stated to generate up to 750 output tokens per second. That is positioned as up to 14× Standard processing.
OpenAI Official Announcement
Powered by Cerebras, Ultrafast generates up to 750 output tokens per second, bringing our most intelligent model to products and workflows where every second matters. OpenAI Official Announcement
Q. Can anyone use Ultrafast?
Not at this point. It is offered as a limited preview to a select group of customers, and OpenAI says access will expand as capacity grows.
OpenAI Official Announcement
GPT-5.6 Sol on Ultrafast mode is available in a limited preview today to a select group of customers. We'll expand access as capacity grows. OpenAI Official Announcement

Related Tools

Related Tool Categories

Articles