Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

OpenAI previews Ultrafast tier running GPT-5.6 Sol up to 14x faster on Cerebras chips

OpenAI unveiled Ultrafast on August 13, a preview tier it says runs GPT-5.6 Sol up to 14 times faster on Cerebras' wafer-scale hardware.

D
Aug 13, 2026 · 1 min read

OpenAI on August 13, 2026, previewed Ultrafast, a new service tier that it says runs its GPT-5.6 Sol model up to 14 times faster than standard processing. Ultrafast generates as many as 750 output tokens per second on Cerebras’ Wafer-Scale Engine hardware.

The tier targets latency-sensitive work — voice interfaces, customer support, developer agents, financial research and security-incident response — where response speed, not raw model size, decides whether an AI system feels usable. OpenAI is offering Ultrafast as a limited preview to select customers, with a waitlist for wider access.

The speed comes from hardware, not a new model. Cerebras says its wafer-scale chip keeps model weights on-chip in 44GB of SRAM, removing the memory-bandwidth bottleneck that throttles GPU-based inference. Running GPT-5.6 Sol on that silicon, OpenAI says Ultrafast delivers a 5.6x end-to-end speedup on the GDP-Val benchmark with no measured quality loss versus the standard configuration.

OpenAI also framed the tier against rivals, claiming it runs about 5 times faster than Anthropic’s Claude Opus 4.8 Fast mode and 11 times faster than Claude Fable 5. Those comparisons come from OpenAI and Cerebras and have not been independently verified; the benchmark figures describe the companies’ own tests rather than third-party measurement.

Speed at this level matters most for agents that chain many model calls, where each saved second compounds. Whether Ultrafast holds its quality claims outside the GDP-Val test, and how OpenAI prices it beyond the preview, will determine if wafer-scale inference moves from demo to default.

More news