Ultrafast Mode Pushes GPT-5.6 Sol to 14x Speed

OpenAI just launched a new mode called Ultrafast, and it’s built to make the company’s most powerful model, GPT-5.6 Sol, run dramatically faster. OpenAI says Ultrafast can work at 14x the speed of standard processing, cranking out up to 750 output tokens per second. Tokens are the individual pieces of text a language model generates as it responds, so more tokens per second means answers that land almost in real time.

This is significant because speed has usually come with a tradeoff. “Until now, getting real-time speed typically meant choosing a smaller or more specialized model,” OpenAI wrote in a blog post on Thursday. “Ultrafast points to progress in a new direction: more useful work per second.” In plain terms, you no longer have to downgrade to a lighter model just to get quick responses.

What Ultrafast actually does

  1. Massive speed jump. OpenAI claims 14x faster processing than standard, hitting up to 750 output tokens per second. That’s the headline number, and it’s a big one.
  2. Runs on the full model. This isn’t a stripped-down variant. Ultrafast accelerates GPT-5.6 Sol, OpenAI’s latest and most capable model, so you keep the power while gaining the speed.
  3. Powered by Cerebras. The mode runs on OpenAI’s partnership with chipmaker Cerebras, which specializes in hardware built for exactly this kind of high-throughput AI work.
  4. Aimed at real-time work. The whole pitch is doing more useful work per second, which matters most when a delay of a few seconds actually costs you something.

Where it’s meant to be used

OpenAI suggests Ultrafast fits a range of corporate workflows where response time is critical. The company points to a few areas in particular:

  • Incident response, where every second of delay can matter
  • Customer service and support, where fast replies keep conversations flowing
  • Financial market analysis, where speed can be the whole edge
  • E-commerce, where quick answers shape whether a sale happens

What stands out here is that these are all live, high-stakes settings. This isn’t about drafting a blog post faster. It’s about AI keeping pace with situations that move quickly on their own.

How it stacks up

OpenAI isn’t alone in chasing speed. Competitors like Anthropic have launched accelerated versions of their models too. Claude has a fast mode, though it doesn’t deliver the kind of speed OpenAI is claiming here. If those 14x numbers hold up in real use, OpenAI is setting a high bar for the rest of the field.

A quick caveat worth keeping in mind: these speed figures come from OpenAI’s own blog post. Independent testing tends to tell the fuller story, so it’s worth watching how Ultrafast performs once more people get their hands on it.

Who can get it

Here’s where the excitement meets reality. Ultrafast is being released in preview, and right now that preview only reaches a small group of customers. OpenAI says it plans to expand access as “capacity grows,” which is a familiar pattern for hardware-bound features. When your speed depends on a chip partner like Cerebras, availability tends to scale with how much of that hardware you can bring online.

So most users won’t be flipping this on today. But the direction is clear.

Why it matters

Speed has quietly become one of the next big battlegrounds in AI. Raw model intelligence still matters, but for a lot of business use cases, a smart answer that arrives too late is worthless. By pushing full-strength GPT-5.6 Sol to real-time speeds, OpenAI is signaling that you shouldn’t have to choose between capable and fast anymore.

The Cerebras partnership is the piece to watch. It ties OpenAI’s speed story directly to specialized silicon, and that hardware layer is increasingly where these performance gains are won. As capacity grows and access widens, expect speed claims like this to become a standard part of how AI models compete.

For the full details, check OpenAI’s official blog.

Scroll to Top