// AI NEWS · MODEL RELEASE

OpenAI and Cerebras Preview "Ultrafast" Mode for GPT-5.6 Sol, Running the Flagship at Up to 750 Tokens per Second

A new OpenAI API service tier powered by Cerebras runs the same GPT-5.6 Sol model up to about 14 times faster than Standard, with no reported quality loss — in limited preview.

2026-08-13 · by Moe Ameen

What happened

On August 13, 2026, OpenAI and Cerebras shared an early look at "Ultrafast," a new OpenAI API service tier that runs the flagship GPT-5.6 Sol model at up to 750 output tokens per second — which OpenAI describes as up to roughly 14 times faster than Sol's Standard processing. The key claim is that speed comes with no intelligence trade: it is the same GPT-5.6 Sol, and on the GDP-Val benchmark the companies reported a 5.6x end-to-end speedup with no measured quality degradation.

The acceleration is a hardware story. Cerebras serves Sol on its Wafer-Scale Engine, whose wafer-sized chips hold model weights on-chip in 44 GB of SRAM, avoiding the memory-bandwidth bottleneck that limits frontier-model inference speed on conventional GPUs. In the companies' own comparisons, Ultrafast runs about 5x faster than Claude Opus 4.8 on Fast mode and about 11x faster than Claude Fable 5; on the Humanity's Last Exam benchmark it completed 2,500 questions in 11 hours and 11 minutes, against roughly 78 hours for Fable 5. OpenAI frames the target work as latency-bound and high-stakes: legal briefs, financial models, engineering reports, production-outage diagnosis, cybersecurity response, and long agent chains.

The important caveats are access and scope. Ultrafast launches first in the OpenAI API and is in limited preview — available initially to a select group of customers, with access expanding over time — so most people cannot use it yet. And it is a speed tier, not a capability change: GPT-5.6 Sol still outputs text and code and reads images. It renders no images, video, or audio and publishes nothing. As always with a preview, treat the specific throughput and benchmark figures as a snapshot of the announcement.

Why it matters for creators

  • The bottleneck in an AI workflow is shifting from "waiting on the model" to everything after it — near-instant reasoning does not shorten producing and publishing.
  • Frontier intelligence at interactive speed makes high-volume drafting practical: a month of hooks, scripts, or angles in one sitting instead of across days.
  • It is preview-gated and API-first, so it is a developer capability today, not a click-to-use feature — a gap between "in the news" and "in your hands" worth explaining to an audience.
  • Speed is now a competitive axis between frontier models, not just quality — but for content, faster tokens still are not a finished, scheduled post.
  • Specialized inference hardware (Cerebras' wafer-scale chips) underpinning a frontier model is itself a timely, high-search topic your audience is asking about this week.

How to act on this with Kompozy

The fastest thing you can do with this news is not use Ultrafast — it is publish a clear take on it before everyone else. "OpenAI just made GPT-5.6 Sol 14x faster on Cerebras" is exactly the timely, high-intent topic your audience is searching this week, and most creators will still be waiting for preview access while the interest peaks. Drop your angle into [Kompozy](/) and it fans one point of view into a blog explainer, a carousel that breaks down what a speed tier actually is, short captioned clips, a quote graphic on the headline number, and platform-native posts — all in your voice through the [Persona Brief](/glossary/persona-brief) — then schedules and publishes them across Instagram, TikTok, YouTube, LinkedIn, X, Facebook, Pinterest, and Threads plus a blog and newsletter. Being early and useful on a story like this is how one take becomes a week of content.

There is also a workflow read worth keeping. If a faster model tempts you to wire your pipeline to whichever endpoint is quickest this month, resist it — the fastest tier changes constantly, and token speed is not speed to a published post. Kompozy runs generation on managed Claude and OpenAI models with the model layer abstracted away, so you never re-wire when a new fast tier lands; you set a Persona Brief and approve outputs while the engine renders persona video, carousels, images, blogs, and newsletters and ships them on a schedule. Let Ultrafast (or any fast model) draft upstream if you have access; let Kompozy own the producing-and-publishing half that no inference speedup touches.

Quick takeaways

  • OpenAI and Cerebras previewed Ultrafast mode for GPT-5.6 Sol on August 13, 2026 — up to 750 output tokens/sec, ~14x faster than Standard, same model.
  • The speed is from Cerebras' Wafer-Scale Engine (44 GB on-chip SRAM); the companies reported no quality loss and a 5.6x GDP-Val speedup.
  • It is API-first and in limited preview to select customers, so broad access is not available yet.
  • A faster model produces more text, not finished content — Kompozy turns drafts into 18 formats published across nine platforms, whatever model runs upstream.

Frequently asked questions

What is GPT-5.6 Sol Ultrafast mode?

It is a new OpenAI API service tier, previewed with Cerebras on August 13, 2026, that runs the existing GPT-5.6 Sol model at up to 750 output tokens per second — roughly 14x faster than Sol Standard. The intelligence is unchanged; only the speed is different, powered by Cerebras' wafer-scale hardware.

Can I use Ultrafast mode right now?

Not broadly. At the August 13, 2026 preview it is in limited availability to a select group of customers, with access expanding over time, and it launches first in the OpenAI API — so it is developer-facing rather than a consumer feature you can click on today.

Does Ultrafast change what GPT-5.6 Sol can produce?

No. Ultrafast only makes Sol answer faster. It still outputs text and code and reads images — it renders no video, images, or audio and publishes nothing. Faster drafting sits upstream of content; turning drafts into finished, scheduled posts is the job of a content engine like Kompozy.

How is this different from GPT-5.6 Sol Ultra?

"Ultrafast" is a speed tier — the same Sol, returned faster, on Cerebras hardware. "Ultra" is a separate mode that uses subagents and a max-reasoning-effort setting for depth on hard tasks, and is being wired into Codex for agentic coding. One is about throughput, the other about how hard the model works on a single problem.

Related news

← All AI news · Get started →