Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
OpenAI has unveiled Ultrafast mode, a new API service tier that accelerates GPT-5.6 Sol inference to up to 14 times faster speeds, achieving as many as 750 output tokens per second. This performance leap is powered by Cerebras’ specialized AI hardware, marking a significant step in making advanced generative models more responsive for real-time applications. The service targets developers and enterprises needing low-latency AI interactions, such as live coding assistants, real-time translation, or interactive agents. By decoupling speed from model size, OpenAI aims to broaden access to high-performance AI without compromising on capability.
What this means for you
Developers building real-time AI applications should evaluate Ultrafast mode for latency-sensitive use cases, while monitoring cost and availability as the service scales.
Put this to work
Try the tools in this story inside the AI Tool Lab, browse roles in the AI Job Portal, test your skills with Aptitude Tests, or read our in-depth guides.
Keep reading
More AI news
Anthropic CEO says it’s time to pump the brakes on AI
Anthropic CEO Dario Amodei has proposed a strategic shift in AI development, advocating for a measured pace rather than an unchecked race for intelligence. By inviting independent third-party assessors like METR to audit their systems, the company aims to establish a new standard for transparency and safety. This move signals a significant departure from the 'growth-at-all-costs' mentality, suggesting that the industry must prioritize verifiable safety protocols before deploying more powerful models to the public.
Read the briefAnthropic CEO outlines plan to ‘pace the frontier’
Anthropic CEO Dario Amodei has proposed a strategic shift aimed at decelerating the breakneck speed of artificial intelligence development. By advocating for a more deliberate 'pacing' of innovation, the company suggests that industry leaders should prioritize safety evaluations and societal impact assessments over mere technical capability benchmarks. This shift marks a significant departure from the competitive race between major AI labs, highlighting growing concerns that rapid deployment without adequate guardrails could pose existential or systemic risks to global infrastructure.
Read the briefY Combinator’s Garry Tan wants U.S. open-weight AI labs to ‘distill’ frontier models, too
Y Combinator leader Garry Tan is advocating for a shift in how leading artificial intelligence developers share their technology. He contends that since foundational models are built upon vast repositories of human-generated information, the resulting capabilities should be accessible as a public benefit. By encouraging labs to create smaller, distilled versions of frontier models, Tan hopes to democratize access to high-performance AI, moving away from closed-off ecosystems toward a more equitable distribution of innovative tools.
Read the brief