OpenAI introduces ‘Ultrafast,’ a new mode that makes GPT-5.6 Sol work at 14x the speed

Published: August 13, 2026 — TechCrunch

If you’ve ever found yourself wishing that ChatGPT was a little bit quicker on the uptake, OpenAI seems to be answering your prayers.

The AI lab has rolled out a new mode called Ultrafast, which it says is designed to seriously accelerate the pace at which its latest and most powerful model, GPT-5.6 Sol, accomplishes its work.

The company says that Ultrafast can work at 14x the speed of standard processing, delivering up to 750 output tokens — such tokens represent the distinct pieces of text generated by an LLM when it interacts with a human — per second.

“Until now, getting real-time speed typically meant choosing a smaller or more specialized model,” the company said in the blog post on Thursday. “Ultrafast points to progress in a new direction: more useful work per second.”

OpenAI’s competitors, like Anthropic, have similarly launched accelerated versions of their models. Claude has fast mode, although it doesn’t deliver the kind of speed that OpenAI is offering here.

OpenAI suggests that this high-octane version of GPT-5.6 Sol can be deployed across a number of different corporate workflows, most notably incident response, customer service and support, financial market analysis, and e-commerce, among other relevant areas.

Ultrafast, which is currently being released in preview, is being powered by OpenAI’s partnership with chipmaker Cerebras. Currently, that preview is only being made available to a small group of customers, although OpenAI says that it will expand access to the feature as “capacity grows.”

Cerebras partnership details (from Cerebras blog, Aug 13, 2026)

Cerebras and OpenAI shared an early look at Ultrafast Mode, a new service tier launching first in the OpenAI API and powered by Cerebras. Cerebras powers GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second without quality compromise.

Compared with output speeds reported by Artificial Analysis, GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode.

On Humanity’s Last Exam (2,500 questions), Cerebras reported GPT-5.6 Sol on Ultrafast answered the full set in 11 hours and 11 minutes versus Claude Fable 5 needing 78 hours and 27 minutes — nearly 7× faster at comparable accuracy.

On GDP-Val, Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation.

Ultrafast is powered by Cerebras’ Wafer-Scale Engine architecture with 44 GB of SRAM on each wafer-sized chip, keeping weights on-chip to avoid GPU memory-bandwidth bottlenecks.

GPT-5.6 Sol on Ultrafast mode is available in a limited preview to a select group of customers; access will expand as capacity grows.

Source: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai