Inception's diffusion LLM — a text model that refines tokens in parallel to hit over 1,100 tokens per second at a low cost per token.
Last verified · 2026-09-08 · by Moe Ameen
Mercury 2.5 is a diffusion large language model (a dLLM) from Inception, the company founded by Stanford professor Stefano Ermon to commercialize diffusion-based text generation. It was announced in September 2026, following a preview build in late August. Where a standard autoregressive model writes one token at a time from left to right, a diffusion LLM starts from a rough draft of the whole output and refines tokens in parallel — the mechanism behind its headline speed. Inception calls it the most capable dLLM on the market and reports 1,107 tokens per second on widely available NVIDIA GPUs.
Inception positions Mercury 2.5 in the cost-optimized frontier tier, comparable to models like GPT-5.6 Luna, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5, with roughly a 40% increase in intelligence over Mercury 2 while holding the same speed and cost profile. It carries a 260K-token context window and supports tunable reasoning, parallel tool calls, and schema-aligned JSON, so it works for agents and structured extraction, not just chat.
You reach it through Inception's own API, Baseten, and OpenRouter, over an OpenAI-compatible endpoint — an app already wired to OpenAI can switch with minimal changes. There is a hosted chat at chat.inceptionlabs.ai and free tokens for API testing at launch. Standard pricing is $0.20 per million input tokens and $0.75 per million output, with a steep launch promotion on top.
The boundary worth planning around: Mercury is a text engine, not a content workflow. It drafts words — hooks, scripts, captions, articles, JSON — extremely fast, and stops there. It does not govern brand voice, generate images or video, fan one idea into a multi-format set, or publish anywhere.
Mercury's real gift to a creator is throughput: at over 1,100 tokens per second you can draft fifty hook variants, ten script angles, or a week of caption options in the time a slower model writes one. But raw speed produces raw text, and a pile of drafts is not a content calendar. That is the exact hand-off into [Kompozy](/): use Mercury to generate the volume, then use Kompozy to turn the winners into finished, published content. Paste a Mercury-drafted script into Kompozy and it becomes a [Persona Short](/glossary/persona-shorts) — a HeyGen talking-head avatar reading your words, with auto-captions burned in — while a Mercury-drafted post becomes a brand-exact [Carousel](/glossary/hyperframes), a Quote Graphic, and Photo Posts, each with per-platform copy refined through the [Persona Brief](/glossary/persona-brief).
The pairing works because the two tools optimize opposite ends. Mercury has no brand-voice governance, no image or video generation, and no publishing; Kompozy adds all three and fans one idea across nine platforms plus blog and email with scheduling, autopilot, and a review step. So the workflow is: generate wide and cheap in Mercury, then let Kompozy govern the voice, generate the media, reframe per platform, and ship it. You go from a fast draft to scheduled persona video, carousels, and posts without the manual formatting and posting in between. For another fast, cheap model on the same publishing rails, see [DiffusionGemma](/ai-tools/diffusiongemma).
Mercury 2.5 is a diffusion large language model from Inception, announced in September 2026. Instead of generating one token at a time, it refines tokens in parallel, reaching 1,107 tokens per second on common GPUs. It is built for fast text generation, low-latency agents, and structured output at a low cost per token.
A standard (autoregressive) LLM writes text one token at a time, left to right. A diffusion LLM like Mercury starts from a rough draft of the whole output and refines all the tokens in parallel over several passes. That parallelism is why Mercury reaches much higher tokens-per-second throughput and lower latency than comparable autoregressive models.
It excels at drafting text at volume: hooks, captions, video scripts, article and newsletter drafts, and structured JSON for pipelines. The speed makes it ideal for generating many variants of an idea quickly. It does not make images, video, or branded posts — pair it with a content engine like Kompozy for the media, branding, and publishing.
On its own Mercury produces text and stops. In Kompozy you bring a Mercury-drafted script or post in and it becomes finished content: a Persona Short avatar video, a HyperFrames carousel, or Photo Posts, each refined into your brand voice via the Persona Brief and scheduled across nine platforms plus blog and email — the steps Mercury does not do.
Standard pricing is $0.20 per million input tokens and $0.75 per million output tokens, with a promotional launch discount down to $0.04 per million input and $0.15 per million output. Launch promotions expire, so budget against the standard rates and confirm current pricing on Inception's site.