HeadlinesBriefing favicon HeadlinesBriefing.com

GPT-5.6 Sol Ultrafast: Fastest AI Inference

Hacker News •
×

Cerebras and OpenAI are sharing an early look at Ultrafast Mode, a new service tier launching first in the OpenAI API and powered by Cerebras. Ultrafast powers GPT-5.6 Sol, delivering up to 750 output tokens per second without quality compromise, accelerating time-sensitive, mission-critical work.

Compared with output speeds reported by Artificial Analysis, GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode. On Humanity's Last Exam, GPT-5.6 Sol answered all 2,500 questions in 11 hours and 11 minutes; Claude Fable 5 needed 78 hours and 27 minutes—nearly 7× slower.

Ultrafast is enabled by Cerebras’ Wafer-Scale Engine architecture, which packs 44 GB of SRAM on each wafer-sized chip, eliminating data movement bottlenecks. It delivers a 5.6x end-to-end speedup on economically valuable tasks without quality degradation.

GPT-5.6 Sol on Ultrafast mode is available in limited preview today to select customers, with access expanding as capacity grows.