The rule used to be simple: want speed, take a smaller model. OpenAI just broke it. Ultrafast Mode, announced August 13, runs GPT-5.6 Sol — the actual flagship, not a distilled sibling — at up to 750 output tokens per second, 14x its standard 53 tok/s. The speed comes from Cerebras wafer-scale hardware, the first product to ship from their $10 billion partnership.
What it actually is
An API service tier, not a new model. Same weights, same intelligence, 14x the output. One data point: all 2,500 Humanity’s Last Exam questions answered in 11 hours instead of 78, at comparable accuracy. OpenAI’s pitch is “more useful work per second” — for agent loops, latency is the real bottleneck, and this attacks it directly.
How to get in
Limited API preview for a small batch of customers. No pricing, no GA date, no public model ID yet — access expands as Cerebras capacity grows. Target workloads: incident response, real-time customer support, financial-market analysis, e-commerce agents. Anywhere a 10-step agent chain used to mean a coffee break, it now means seconds.
You Might Also Like
- Openai gpt 5 6 Same Smarts Half the Price sol Hits 750 Tokens s on Cerebras
- Gpt 5 6 sol on Cerebras 750 Tokens s Openais Frontier Model Gets 15x Faster in July
- Mercury 2 Just hit 1000 Tokens per Second and its not Even Using Transformers
- Openai Trusted Access for Cyber Opens gpt 5 5 to Offensive Security Work for Verified Defenders Only
- Openai gpt 5 6 sol Terra Luna a Three Tier Lineup Only 20 Orgs can Touch

Leave a comment