toolcall.
ModelsAug 13, 2026, 20:35 UTC

OpenAI previews Ultrafast mode for GPT-5.6 Sol

The Cerebras-powered API tier runs OpenAI’s top model up to 14x faster than Standard processing for time-sensitive business workflows.

OpenAI is previewing Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14x faster than Standard processing. The company says the Cerebras-powered setup can generate up to 750 output tokens per second while still using its most capable Sol model.

The point is latency, not a new model family. OpenAI argues that teams usually had to choose smaller or more specialized models when they needed real-time speed. Ultrafast is meant to bring frontier-level reasoning into workflows where waiting several seconds changes the product experience or business outcome.

OpenAI names incident response, financial research, security analysis, customer support, voice, commerce and live experimentation as early use cases. Internally, it says engineers have tested Ultrafast for reading logs, traces and team conversations during live incidents, and researchers have used it to tighten experiment-review loops that previously took overnight batches.

Access is limited to a select customer preview for now, with expansion planned as capacity grows. For developers, the useful signal is that inference speed is becoming a platform feature in its own right: OpenAI is not only competing on model quality and price, but on whether its strongest models can respond quickly enough for interactive production systems.

Sources

Mentioned

ai-infrastructuredeveloper-toolsinferenceopenai