// AI NEWS · PRICING

What Kimi K3 Actually Costs: Moonshot's Frontier Model Lands at About $3/$15 per Million Tokens, Well Above Kimi's Bargain Past

At launch pricing near $3 per million input tokens and $15 per million output, Kimi K3 runs roughly three to four times the previous Kimi flagship — in the same range as Claude Sonnet rather than the cut-rate tiers Kimi built its name on.

KompozyTurn one idea into a week of content — across every platform, published for you.
Get Started →

2026-07-27 · by Moe Ameen

What happened

When Moonshot AI rolled out its Kimi K3 flagship in mid-July 2026, it set API pricing at roughly $3 per million input tokens and $15 per million output, with cached input near $0.30 per million. Those rates apply flat across the model's full 1-million-token (1,048,576-token) context — there is no length-based tiering — and, as of late July 2026, they have not changed since launch. The "update" here is not a price cut or hike; it is that the numbers are now settled and public, and creators are working out what a frontier Kimi model actually costs to run.

The reason it drew attention is the break from Kimi's own history. Moonshot built its reputation on undercutting Western labs: the prior flagship, Kimi K2.6, has been listed around $0.95 per million input tokens and $4 per million output, and the cheaper K2.5 around $0.60/$3.00 (both on a 256K context). Against K2.6, K3 is roughly three times the input price and close to four times the output price. So the model that Moonshot frames as its most capable — and the largest open-weight model to date — is also, by a wide margin, its most expensive to call. Cached input softens that: at about $0.30 per million, a cache hit is ten times cheaper than a fresh read, so workflows that reuse a large fixed context (a long document, a system prompt) pay far less than the headline rate implies.

Placed against the wider market, K3's $3/$15 lands in roughly the same range as Claude Sonnet-class pricing, below the top closed frontier tiers such as Claude Opus (around $5/$25), and well above the cheapest open and Chinese models. That is the shift worth naming: Kimi is no longer competing purely on being the bargain option. It is asking to be judged on capability at a mainstream frontier price. Because Moonshot also published K3's open weights, self-hosting is a theoretical way to avoid per-token API fees — but running a ~2.8-trillion-parameter model takes a multi-GPU cluster and roughly a terabyte-plus of fast memory, so for almost everyone the hosted API (or a third-party host) remains the practical, and metered, path. Confirm the current rates on Moonshot's own pricing pages before budgeting against them, since launch pricing can move.

Why it matters for creators

  • The cheap-frontier-model era has a ceiling. K3 shows a leading open model can still carry mainstream frontier pricing — the assumption that every new Chinese release is a bargain no longer holds automatically.
  • Per-token billing makes content costs variable and hard to predict. A workflow that drafts long scripts, reasons over transcripts, or fans out many variations can run up real output-token spend at $15 per million.
  • Caching is the lever. At roughly $0.30 per million, cached input is 10x cheaper than a fresh read, so reusing a fixed context (a brief, a style guide, a source doc) is where the savings actually live.
  • Price buys reasoning, not finished content. Even at frontier rates, K3 outputs text and code and reads images — it renders no captioned video, branded graphic, or scheduled post, so the token bill is only the first line item of a real content operation.
  • It is a timely explainer. "How much does Kimi K3 cost" is a question your audience is searching this week — a clear, honest cost breakdown is useful content right now.

How to act on this with Kompozy

A pricing page is really a budgeting question, so answer it as one. If your plan is to build a content operation on top of a raw model like K3, the per-token rate is only the visible cost. You are also on the hook for everything the model does not do: a system to turn drafts into captioned video, branded carousels, quote cards, and infographics; a scheduler; and the integrations to publish to each platform. Those are separate tools, separate bills, and separate maintenance — and the token meter keeps running the whole time. The headline "$3/$15" is the cheapest part of the real number.

Kompozy is the flat-fee alternative to that stack. Instead of metering tokens and stitching together a production-and-publishing pipeline, you get a managed subscription on credits — Starter at $99/mo and Pro at $299/mo — that covers generation across 18 formats (persona and avatar video, Clipped Shorts, brand-exact carousels, photo posts, quote graphics, blogs, and newsletters) and publishing to eight social platforms plus blog and email from one queue, with autopilot and a per-post review pipeline. Because Kompozy runs its own managed Claude and OpenAI models under the hood, you never touch a token bill, and a stronger model entering the pool is upside for the engine rather than a new invoice for you. And if you specifically want to route a model like K3, the Founding tier's bring-your-own-key option lets you plug in your own provider key while Kompozy still handles the media, the brand voice via your Persona Brief, and the publishing — so cost stays predictable and the output is finished, not a raw draft you still have to produce around.

Quick takeaways

  • Kimi K3 launched in mid-July 2026 with API pricing near $3 per million input tokens, $15 per million output, and about $0.30 per million on cached input — flat across its 1M-token context, and unchanged since launch as of late July 2026.
  • That is roughly 3x the input and ~4x the output price of the prior Kimi K2.6 flagship (about $0.95/$4.00), so K3 is Moonshot's most capable and most expensive model.
  • K3's $3/$15 sits in Claude Sonnet's pricing range, below top closed tiers like Claude Opus (~$5/$25), and above the cheapest open models — a move away from Kimi's pure bargain positioning.
  • Cached input at ~$0.30/M is 10x cheaper than a fresh read, so reusing a large fixed context is the main way to control cost.
  • Open weights let you self-host to avoid API fees in theory, but serving a ~2.8T-parameter model needs a multi-GPU cluster — the hosted, metered API stays the practical path for most.

Frequently asked questions

How much does Kimi K3 cost?

At its mid-July 2026 launch, Moonshot listed Kimi K3 API pricing around $3 per million input tokens and $15 per million output, with cached input near $0.30 per million, flat across its 1-million-token context. Those rates had not changed as of late July 2026. There is also consumer access through kimi.com and the Kimi apps on a membership basis. Confirm current pricing on Moonshot's own pages before budgeting.

Why is Kimi K3 more expensive than earlier Kimi models?

K3 is Moonshot's largest and most capable model — framed as the biggest open-weight model to date, at around 2.8 trillion parameters. The prior flagship K2.6 has been listed near $0.95/$4.00 per million input/output tokens and K2.5 near $0.60/$3.00, so K3 runs roughly three to four times those rates. The jump reflects the model's scale and a shift away from Kimi's bargain-only positioning toward a mainstream frontier price.

Is Kimi K3 cheaper than Claude or GPT?

K3's roughly $3/$15 per million tokens sits in the same range as Claude Sonnet-class pricing and below top closed tiers such as Claude Opus (around $5/$25), while remaining above the cheapest open models. So it is competitive with mid-tier frontier pricing rather than dramatically undercutting every Western model, which was the reputation earlier Kimi releases earned.

How can I lower the cost of using Kimi K3 for content?

Two levers help. First, prompt caching: cached input is about $0.30 per million versus $3 for a fresh read, so reusing a fixed context pays off. Second, don't pay frontier per-token rates for the whole pipeline — K3 only drafts text and reads images, so you still need media generation and publishing on top. A managed content engine like Kompozy covers generation across 18 formats plus publishing on flat credit plans, and its Founding tier lets you bring your own model key if you want to route K3 directly.

Related news

← All AI news · Get started →