What Ultrafast is
Ultrafast at a glance
Ultrafast is a new service tier announced by OpenAI on August 13, 2026.
A new tier at up to 14× Standard
Ultrafast runs GPT-5.6 Sol at up to 14× the speed of Standard processing. As the word "tier" suggests, this is not a swap to a different model but a new speed grade for running the same top-end model. It launches first in the OpenAI API. For where GPT-5.6 Sol itself sits, see the GPT-5.6 Sol explainer.
"Today, we're sharing an early look at Ultrafast, a new service tier that runs GPT-5.6 Sol up to 14× faster than Standard processing, launching first in the OpenAI API." — from the OpenAI official announcement
Removing the speed-or-intelligence choice
This is the crux of the release. Until now, if you wanted real-time responses, the only option was to pick a smaller or more specialized model. OpenAI describes Ultrafast as progress toward more useful work per second. If you can take speed without giving up intelligence, AI can move into work you had written off because of the wait.
"Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second."/"When speed no longer requires giving up intelligence, AI can move into the most time-sensitive parts of a business and new kinds of work become possible."(both from the opening explanatory paragraphs) — from the OpenAI official announcement
Where the speed comes from, and where it fits
Use cases named by OpenAI
Here is the source of the speed and the use cases the announcement names.
750 tokens per second, powered by Cerebras
The foundation is Cerebras, and it is stated to generate up to 750 output tokens per second. Cerebras is OpenAI's partner for ultra-low-latency inference (the process by which an AI produces an answer). A token is a small unit that text is split into, so 750 tokens per second means long passages arrive with essentially no waiting.
"Powered by Cerebras, Ultrafast generates up to 750 output tokens per second, bringing our most intelligent model to products and workflows where every second matters."(opening paragraph)/"Ultrafast marks the next step in our partnership with Cerebras to bring ultra-low-latency inference to OpenAI's platform."(Powered by Cerebras section) — from the OpenAI official announcement
Where seconds change the outcome
The announcement names five areas. Incident response, financial research and security, customer support and voice, commerce, and live research and experimentation. All of them are work where the situation changes before an answer arrives, so speed directly changes the result. For incident response, that means reading logs and recent code changes to narrow down a cause while the outage is still unfolding.
"Incident response and reliability: When a critical system fails, analyze application logs, recent code changes, and engineer reports to identify the likely cause and help prepare a fix while the outage is still unfolding."/"Financial research and security: Analyze market signals, assess transactions, and identify suspicious activity while conditions are still changing."/"Customer support and voice: Resolve complex customer issues in real time without interrupting the conversation, even when finding the answer requires multiple steps or systems."/"Commerce: Answer product questions, check inventory, personalize recommendations, and resolve checkout issues while the shopper is still deciding, before hesitation becomes an abandoned cart."/"Live research and experimentation: Turn research that previously took an overnight run into an interactive working session, letting teams test an idea, examine the results, adjust their approach, and run another experiment without breaking their flow."(all five bullets from the use-case list) — from the OpenAI official announcement
Who can use it, and the takeaway
Availability
Finally, whether you can actually use it today.
Still a limited preview
Ultrafast is a limited preview, open only to a selected group of customers. It is not something anyone can call through the API. OpenAI says access will expand as capacity grows, and the official page offers a sign-up to be notified when it does. If you are considering it, working out in advance which of your own processes change value by the second will let you move when a slot opens.
"GPT-5.6 Sol on Ultrafast mode is available in a limited preview today to a select group of customers. We'll expand access as capacity grows."(Availability section)/"If your business requires frontier intelligence at the highest speed, you can sign up to get notified when access expands."(opening explanatory paragraph) — from the OpenAI official announcement
How OpenAI uses it internally
It is already in use inside the company. Research team members used to launch a batch of experiments overnight and review results the next morning. With Ultrafast, OpenAI says that loop tightens enough to support multiple iterations during the workday. Verification that took a night becomes interactive work.
"A common workflow in research is for our team members to launch a batch of experiments over night, and review the results in the morning. With Ultrafast, we see this loop tightening to support multiple iterations during the workday instead." — from the OpenAI official announcement
When you want to read an announcement page on your own machine as markdown with its headings and tables intact, the following tool can help.



