AI Infrastructure

OpenAI previews Ultrafast mode for GPT-5.6 Sol with Cerebras-powered inference

OpenAI says a limited preview of Ultrafast mode can run GPT-5.6 Sol up to 14 times faster than Standard processing and generate up to 750 output tokens per second.

Published Updated
OpenAIGPT-5.6 SolCerebras

OpenAI has previewed Ultrafast, a new service tier that runs GPT-5.6 Sol at much higher speed for applications where response time is part of the product experience. The August 13 announcement says Ultrafast can run the model up to 14 times faster than Standard processing and generate up to 750 output tokens per second. The tier is launching first in the OpenAI API as a limited preview for selected customers, with Cerebras providing the ultra-low-latency inference behind the service.

The product is notable because it tries to reduce a trade-off that has shaped many AI applications. Developers building voice assistants, live support tools, monitoring systems or interactive research products have often chosen smaller models when they needed immediate responses, even when a more capable model would have produced better work. OpenAI’s pitch is that GPT-5.6 Sol on Ultrafast can bring frontier intelligence into time-sensitive workflows without forcing teams to give up too much capability for speed.

OpenAI lists several early use cases. In incident response, a faster frontier model could read logs, recent code changes and engineer reports while an outage is still unfolding, then help identify likely causes and prepare a fix for human review. In financial research and security, it could analyze changing signals, transactions or suspicious behavior before the context goes stale. In customer support and voice, it could resolve complex issues in real time, even when the answer requires multiple tool calls or several pieces of business context.

The company is also testing the tier internally. OpenAI says developers have used GPT-5.6 Sol on Ultrafast to support incident response by reading traces, synthesizing conversations and narrowing the next checks in a fraction of the usual time. Research teams are using it to search knowledge sources, query data and organize information more quickly, turning some workflows that previously depended on overnight batches into multiple iterations during the workday. The key point is not only faster text generation, but a tighter loop between observing, testing and deciding.

Customer feedback in the announcement suggests OpenAI is targeting production environments rather than demonstration tasks. Jane Street described the speed increase from Cerebras as enabling more focused developer workflows. Podium said Ultrafast changed the call experience for complex work in its voice stack. Basis emphasized synchronous user experiences that previously ran into a wall where fast products lacked sufficient intelligence. Rogo described complex financial research becoming closer to a real-time interaction.

The partnership with Cerebras gives the release a broader infrastructure angle. As model providers compete on capability, inference speed and cost are becoming equally important for commercial adoption. A model that is powerful but slow may be acceptable for an overnight research job, but not for a checkout assistant, live analyst, customer call or engineering incident. By offering a speed class for its most intelligent model, OpenAI is trying to make latency a configurable product dimension rather than a fixed limitation.

The limited preview means the claims still need validation outside the initial customer group. Throughput measured in tokens per second does not automatically translate into complete workflow speed, because tool calls, retrieval systems, network latency and review steps can still dominate the user experience. Capacity will also matter as access expands. Even so, the announcement signals where high-end AI services are heading: not only toward better answers, but toward answers that arrive quickly enough to change what users can do while a decision, conversation or operational event is still in motion. If the tier scales, speed may become a purchasing criterion alongside accuracy, safety controls and price per task. It may also pressure competitors to publish clearer latency guarantees for frontier-class models.