At launch pricing near $3 per million input tokens and $15 per million output, Kimi K3 runs roughly three to four times the previous Kimi flagship — in the same range as Claude Sonnet rather than the cut-rate tiers Kimi built its name on.
2026-07-27 · by Moe Ameen
When Moonshot AI rolled out its Kimi K3 flagship in mid-July 2026, it set API pricing at roughly $3 per million input tokens and $15 per million output, with cached input near $0.30 per million. Those rates apply flat across the model's full 1-million-token (1,048,576-token) context — there is no length-based tiering — and, as of late July 2026, they have not changed since launch. The "update" here is not a price cut or hike; it is that the numbers are now settled and public, and creators are working out what a frontier Kimi model actually costs to run.
The reason it drew attention is the break from Kimi's own history. Moonshot built its reputation on undercutting Western labs: the prior flagship, Kimi K2.6, has been listed around $0.95 per million input tokens and $4 per million output, and the cheaper K2.5 around $0.60/$3.00 (both on a 256K context). Against K2.6, K3 is roughly three times the input price and close to four times the output price. So the model that Moonshot frames as its most capable — and the largest open-weight model to date — is also, by a wide margin, its most expensive to call. Cached input softens that: at about $0.30 per million, a cache hit is ten times cheaper than a fresh read, so workflows that reuse a large fixed context (a long document, a system prompt) pay far less than the headline rate implies.
Placed against the wider market, K3's $3/$15 lands in roughly the same range as Claude Sonnet-class pricing, below the top closed frontier tiers such as Claude Opus (around $5/$25), and well above the cheapest open and Chinese models. That is the shift worth naming: Kimi is no longer competing purely on being the bargain option. It is asking to be judged on capability at a mainstream frontier price. Because Moonshot also published K3's open weights, self-hosting is a theoretical way to avoid per-token API fees — but running a ~2.8-trillion-parameter model takes a multi-GPU cluster and roughly a terabyte-plus of fast memory, so for almost everyone the hosted API (or a third-party host) remains the practical, and metered, path. Confirm the current rates on Moonshot's own pricing pages before budgeting against them, since launch pricing can move.
A pricing page is really a budgeting question, so answer it as one. If your plan is to build a content operation on top of a raw model like K3, the per-token rate is only the visible cost. You are also on the hook for everything the model does not do: a system to turn drafts into captioned video, branded carousels, quote cards, and infographics; a scheduler; and the integrations to publish to each platform. Those are separate tools, separate bills, and separate maintenance — and the token meter keeps running the whole time. The headline "$3/$15" is the cheapest part of the real number.
Kompozy is the flat-fee alternative to that stack. Instead of metering tokens and stitching together a production-and-publishing pipeline, you get a managed subscription on credits — Starter at $99/mo and Pro at $299/mo — that covers generation across 18 formats (persona and avatar video, Clipped Shorts, brand-exact carousels, photo posts, quote graphics, blogs, and newsletters) and publishing to eight social platforms plus blog and email from one queue, with autopilot and a per-post review pipeline. Because Kompozy runs its own managed Claude and OpenAI models under the hood, you never touch a token bill, and a stronger model entering the pool is upside for the engine rather than a new invoice for you. And if you specifically want to route a model like K3, the Founding tier's bring-your-own-key option lets you plug in your own provider key while Kompozy still handles the media, the brand voice via your Persona Brief, and the publishing — so cost stays predictable and the output is finished, not a raw draft you still have to produce around.
At its mid-July 2026 launch, Moonshot listed Kimi K3 API pricing around $3 per million input tokens and $15 per million output, with cached input near $0.30 per million, flat across its 1-million-token context. Those rates had not changed as of late July 2026. There is also consumer access through kimi.com and the Kimi apps on a membership basis. Confirm current pricing on Moonshot's own pages before budgeting.
K3 is Moonshot's largest and most capable model — framed as the biggest open-weight model to date, at around 2.8 trillion parameters. The prior flagship K2.6 has been listed near $0.95/$4.00 per million input/output tokens and K2.5 near $0.60/$3.00, so K3 runs roughly three to four times those rates. The jump reflects the model's scale and a shift away from Kimi's bargain-only positioning toward a mainstream frontier price.
K3's roughly $3/$15 per million tokens sits in the same range as Claude Sonnet-class pricing and below top closed tiers such as Claude Opus (around $5/$25), while remaining above the cheapest open models. So it is competitive with mid-tier frontier pricing rather than dramatically undercutting every Western model, which was the reputation earlier Kimi releases earned.
Two levers help. First, prompt caching: cached input is about $0.30 per million versus $3 for a fresh read, so reusing a fixed context pays off. Second, don't pay frontier per-token rates for the whole pipeline — K3 only drafts text and reads images, so you still need media generation and publishing on top. A managed content engine like Kompozy covers generation across 18 formats plus publishing on flat credit plans, and its Founding tier lets you bring your own model key if you want to route K3 directly.