// AI TOOLS · MERCURY 2.5

Mercury 2.5

Inception's diffusion LLM — a text model that refines tokens in parallel to hit over 1,100 tokens per second at a low cost per token.

Last verified · 2026-09-08 · by Moe Ameen

What Mercury 2.5 is

Mercury 2.5 is a diffusion large language model (a dLLM) from Inception, the company founded by Stanford professor Stefano Ermon to commercialize diffusion-based text generation. It was announced in September 2026, following a preview build in late August. Where a standard autoregressive model writes one token at a time from left to right, a diffusion LLM starts from a rough draft of the whole output and refines tokens in parallel — the mechanism behind its headline speed. Inception calls it the most capable dLLM on the market and reports 1,107 tokens per second on widely available NVIDIA GPUs.

Inception positions Mercury 2.5 in the cost-optimized frontier tier, comparable to models like GPT-5.6 Luna, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5, with roughly a 40% increase in intelligence over Mercury 2 while holding the same speed and cost profile. It carries a 260K-token context window and supports tunable reasoning, parallel tool calls, and schema-aligned JSON, so it works for agents and structured extraction, not just chat.

You reach it through Inception's own API, Baseten, and OpenRouter, over an OpenAI-compatible endpoint — an app already wired to OpenAI can switch with minimal changes. There is a hosted chat at chat.inceptionlabs.ai and free tokens for API testing at launch. Standard pricing is $0.20 per million input tokens and $0.75 per million output, with a steep launch promotion on top.

The boundary worth planning around: Mercury is a text engine, not a content workflow. It drafts words — hooks, scripts, captions, articles, JSON — extremely fast, and stops there. It does not govern brand voice, generate images or video, fan one idea into a multi-format set, or publish anywhere.

What you can make with it

  • High-volume batches of hooks, captions, and post copy drafted in seconds, not minutes
  • Video and short-form scripts you can regenerate in bulk to pick the strongest
  • Long-form article and newsletter drafts, using the 260K-token context for source material
  • Structured JSON output for content pipelines — schema-aligned fields, tags, and metadata
  • Fast agent and tool-calling backends where response latency is the product
  • Rapid A/B variant sets: dozens of alternative angles on one idea to test

How Kompozy turns Mercury 2.5 output into content

Mercury's real gift to a creator is throughput: at over 1,100 tokens per second you can draft fifty hook variants, ten script angles, or a week of caption options in the time a slower model writes one. But raw speed produces raw text, and a pile of drafts is not a content calendar. That is the exact hand-off into [Kompozy](/): use Mercury to generate the volume, then use Kompozy to turn the winners into finished, published content. Paste a Mercury-drafted script into Kompozy and it becomes a [Persona Short](/glossary/persona-shorts) — a HeyGen talking-head avatar reading your words, with auto-captions burned in — while a Mercury-drafted post becomes a brand-exact [Carousel](/glossary/hyperframes), a Quote Graphic, and Photo Posts, each with per-platform copy refined through the [Persona Brief](/glossary/persona-brief).

The pairing works because the two tools optimize opposite ends. Mercury has no brand-voice governance, no image or video generation, and no publishing; Kompozy adds all three and fans one idea across nine platforms plus blog and email with scheduling, autopilot, and a review step. So the workflow is: generate wide and cheap in Mercury, then let Kompozy govern the voice, generate the media, reframe per platform, and ship it. You go from a fast draft to scheduled persona video, carousels, and posts without the manual formatting and posting in between. For another fast, cheap model on the same publishing rails, see [DiffusionGemma](/ai-tools/diffusiongemma).

  1. In Mercury, batch-draft the raw material: fifty hooks, ten script angles, or a week of caption options in one fast pass.
  2. Pick the strongest drafts and bring them into Kompozy.
  3. Turn a script into a Persona Short — a HeyGen avatar video with auto-captions — or a post into a HyperFrames carousel and Photo Posts.
  4. Let the Persona Brief refine each piece into your brand voice with per-platform copy, and generate the images or video Mercury cannot.
  5. Schedule and publish the set across nine platforms plus blog and email on autopilot, with a review step.

Frequently asked questions

What is Mercury 2.5?

Mercury 2.5 is a diffusion large language model from Inception, announced in September 2026. Instead of generating one token at a time, it refines tokens in parallel, reaching 1,107 tokens per second on common GPUs. It is built for fast text generation, low-latency agents, and structured output at a low cost per token.

How is a diffusion LLM different from a normal LLM?

A standard (autoregressive) LLM writes text one token at a time, left to right. A diffusion LLM like Mercury starts from a rough draft of the whole output and refines all the tokens in parallel over several passes. That parallelism is why Mercury reaches much higher tokens-per-second throughput and lower latency than comparable autoregressive models.

What can I use Mercury 2.5 to make for content?

It excels at drafting text at volume: hooks, captions, video scripts, article and newsletter drafts, and structured JSON for pipelines. The speed makes it ideal for generating many variants of an idea quickly. It does not make images, video, or branded posts — pair it with a content engine like Kompozy for the media, branding, and publishing.

How do I turn Mercury 2.5 drafts into published posts?

On its own Mercury produces text and stops. In Kompozy you bring a Mercury-drafted script or post in and it becomes finished content: a Persona Short avatar video, a HyperFrames carousel, or Photo Posts, each refined into your brand voice via the Persona Brief and scheduled across nine platforms plus blog and email — the steps Mercury does not do.

How much does Mercury 2.5 cost?

Standard pricing is $0.20 per million input tokens and $0.75 per million output tokens, with a promotional launch discount down to $0.04 per million input and $0.15 per million output. Launch promotions expire, so budget against the standard rates and confirm current pricing on Inception's site.

Related tools

  • DiffusionGemmaGoogle DeepMind's experimental open-weight diffusion language model — it generates text by refining a whole block of tokens in parallel instead of one at a time, hitting over 1,000 tokens per second on a single H100. Technical report published July 31, 2026.
  • OpenRouterA unified, OpenAI-compatible API that routes one endpoint to 400+ large language models from dozens of providers — with automatic fallback, cost and speed routing, and a single shared credit balance.
  • GPT-5.6 SolOpenAI's flagship GPT-5.6 tier — a frontier reasoning, writing, and tool-orchestration model that reads reference images and drives multi-tool creative pipelines, but generates no media itself.
  • Gemini 3.8 FlashGoogle's fast, low-cost workhorse model that 'works harder' on coding, agents, and analysis — launched September 2, 2026 alongside a defense-focused Cyber variant.
  • DeepSeek-V4-FlashDeepSeek's fast, low-cost frontier language model — a 284B-parameter mixture-of-experts LLM (13B active) with a 1M-token context, open weights under the MIT license, and API pricing near the bottom of the market.

← All AI tools · Get started →