OpenAI's Cerebras-powered "Ultrafast" service tier runs GPT-5.6 Sol at up to 750 output tokens per second — the same flagship model, just far faster.
Last verified · 2026-08-13 · by Moe Ameen
GPT-5.6 Sol Ultrafast is not a new model — it is a new speed tier. On August 13, 2026, OpenAI and Cerebras shared an early look at "Ultrafast," a service tier that runs OpenAI's flagship GPT-5.6 Sol at up to 750 output tokens per second, launching first in the OpenAI API. OpenAI positions it as up to roughly 14 times faster than Sol's Standard processing, and both companies stress the point that matters most: it is the same GPT-5.6 Sol intelligence, just returned far faster. There is no quality trade for the speed — on the GDP-Val benchmark the companies reported a 5.6x end-to-end speedup with no measured quality degradation.
The speed comes from the hardware, not a smaller model. Cerebras runs Sol on its Wafer-Scale Engine, a chip that keeps model weights on-chip in 44 GB of SRAM, sidestepping the memory-bandwidth bottleneck that caps frontier-model inference on conventional GPUs. In the companies' own comparisons, Ultrafast runs about 5x faster than Claude Opus 4.8 on Fast mode and about 11x faster than Claude Fable 5; on the Humanity's Last Exam benchmark, it worked through 2,500 questions in 11 hours and 11 minutes versus roughly 78 hours for Fable 5. The use cases OpenAI cites — legal briefs, financial models, engineering reports, production-outage diagnosis, cybersecurity response, and long agent chains — are all cases where waiting on the model was the bottleneck.
Two things keep this honest for a creator. First, access is narrow: Ultrafast is in limited preview, available initially to a select group of customers with access expanding over time, so most people reading about it cannot use it yet. Second, and more important, Ultrafast changes only how fast Sol answers — it does not change what Sol produces. GPT-5.6 Sol is a text-and-code model that reads images and returns text and code (see the base tier, [GPT-5.6 Sol](/ai-tools/gpt-5-6-sol)). It draws no images, renders no video, synthesizes no audio, and publishes nothing. Do not confuse it with [GPT-5.6 Sol Ultra](/ai-tools/gpt-5-6-sol-ultra), which is a separate subagent/max-effort mode wired into Codex — "Ultrafast" is about throughput, "Ultra" is about depth. For content, that means the thinking half of your workflow just got near-instant, and the producing-and-publishing half did not move at all.
The interesting thing about Ultrafast is where it moves the bottleneck. For most creators the slow step was never the model thinking — it was everything after: turning a draft into a video, a carousel, and a week of on-brand posts, then getting them scheduled and out the door. Ultrafast makes the reasoning near-instant, which only sharpens the contrast: you can now generate a month of scripts and angles in one sitting and still have nothing published, because a faster language model produces more text, not more finished content. That downstream half is precisely what [Kompozy](/) exists to run, and Ultrafast makes it a strong upstream partner — the faster you can draft, the more raw material Kompozy has to turn into shipped posts.
Concretely: batch-draft your angles and scripts against Ultrafast (or any fast tier — Kompozy does not care which model you drafted on), then bring that pile into Kompozy as sources. Kompozy rewrites each in your real voice through the [Persona Brief](/glossary/persona-brief) and generates the formats no language model outputs at any speed: [Persona Shorts](/glossary/persona-shorts) and HeyGen avatar video, Carousel Posts and Persona Tweets rendered pixel-exact through [HyperFrames](/glossary/hyperframes), Photo Posts, Quote Graphics, Blog Articles, and Email Newsletters. Then it schedules and publishes the whole batch across Instagram, TikTok, YouTube, LinkedIn, X, Facebook, Pinterest, and Threads, plus Mailchimp and your blog, on Autopilot with a per-post review pipeline. Worth knowing: Kompozy already runs its own copy generation on managed OpenAI and Claude models, so you get frontier-class writing inside the engine without an API key or per-token billing — Ultrafast is optional upstream horsepower, not a dependency.
It is a speed tier, not a new model. OpenAI and Cerebras previewed "Ultrafast" mode on August 13, 2026 — a service tier launching first in the OpenAI API that runs the existing GPT-5.6 Sol at up to 750 output tokens per second, roughly 14x faster than Sol Standard, with the same intelligence and no reported quality loss.
"Ultrafast" is about throughput — the same Sol model returned much faster, powered by Cerebras hardware. "Ultra" is about depth — a subagent and max-reasoning-effort mode on Sol that is being wired into Codex for agentic coding. Different axes: one makes Sol quicker, the other makes it work harder on a single complex task.
Not broadly. As of the August 13, 2026 preview it is in limited availability to a select group of customers, with access expanding over time. It launches first in the OpenAI API, so it is developer-facing rather than a click-to-use consumer feature.
No. Ultrafast only changes how fast GPT-5.6 Sol answers — it still outputs text and code and reads images. It renders no video, images, or audio and publishes nothing. Faster drafting is upstream of content; to turn drafts into finished, on-brand, scheduled posts you pair it with a content engine like Kompozy.
It removes latency from the reasoning step, so you can draft and plan at higher volume in less time. But it does not shorten the producing-and-publishing half — generating persona video, carousels, images, blogs, and newsletters and scheduling them across platforms. That is the part a content engine like Kompozy runs on top of whichever model you drafted on.