HeadlinesBriefing favicon HeadlinesBriefing.com

Previewing Ultrafast: GPT-5.6 Sol at 14X Speed

OpenAI Blog •
×

Today, we’re sharing an early look at Ultrafast, a new service tier that runs GPT‑5.6 Sol up to 14× faster than Standard processing, launching first in the Open AI API. Powered by Cerebras, Ultrafast generates up to 750 output tokens per second, bringing our most intelligent model to products and workflows where every second matters. With GPT‑5.6, we’re pushing the frontier on what our models can do and making them more efficient across every layer of our stack. Those improvements have made advanced intelligence more affordable and more useful to more people.

Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second. When speed no longer requires giving up intelligence, AI can move into the most time‑sensitive parts of a business. We’ve already seen encouraging scenarios for Ultrafast: incident response, financial research and security, customer support and voice, commerce, and live research and experimentation.

During the preview period, we’re working with an initial group of customers to understand where this speed makes the biggest difference. GPT‑5.6 Sol on Ultrafast mode is being tested across coding, commerce, financial research, support, and other interactive applications. Inside Open AI, developers are using it for incident response and research, reducing the delay between observing a signal and taking action.

Ultrafast marks the next step in our partnership with Cerebras to bring ultra‑low‑latency inference to Open AI’s platform. Now, with GPT‑5.6 Sol on Ultrafast mode, Cerebras is delivering up to 750 output tokens per second. Availability is limited preview today, with expansion as capacity grows.