The weights for the 1M-context flagship went live on Hugging Face a day ahead of schedule, making what Moonshot calls the largest open-weight model publicly available free to download, self-host, and fine-tune — though running it takes roughly a terabyte-plus of fast memory.
2026-07-27 · by Moe Ameen
Moonshot AI published the open weights for Kimi K3 on Hugging Face on July 26, 2026 — a day ahead of the July 27 target the lab had signaled — making the model free to download, inspect, fine-tune, and self-host. K3 had been available through Moonshot's hosted API and the kimi.com consumer platform since its mid-July unveiling; this drop is the model itself, put in anyone's hands. Moonshot frames it as the largest open-weight model released to date.
The published model card lists the architecture in detail: about 2.8 trillion total parameters with roughly 104 billion activated per token, a mixture-of-experts design (896 experts, 16 selected per token) using Moonshot's own Kimi Delta Attention and Attention Residuals, a 1,048,576-token (1M) context window, and native multimodal input through a MoonViT-V2 vision encoder. The weights ship MXFP4-quantized, and the card reports benchmark figures across reasoning, coding, agentic, and vision categories (for example a high GPQA Diamond score and strong coding results). Those numbers are Moonshot's own published card values, and independent leaderboard confirmation was still limited at the authoring date — weigh them accordingly. The release is under Moonshot's permissive Kimi K3 License, in the same open-license lineage as prior Kimi models.
The practical catch is scale. Holding the MXFP4-quantized weights alone takes on the order of 1.4 terabytes of fast memory, before runtime buffers or key-value cache — which in practice means a multi-GPU cluster, not a workstation. So "free to download" is not the same as "free to run": the license and the file are open, the infrastructure to serve a 2.8-trillion-parameter model is not. For most people, hosted API access remains the realistic path; the open weights matter most to teams that want to self-host for data control, audit the model, or fine-tune it on their own data. It is another marker in the broader trend of frontier intelligence getting bigger, cheaper, and more open at once.
The fantasy this release sells a creator is "I'll download K3, fine-tune it on my voice, and have my own branded content model." Two things are true about that. First, it is genuinely hard: you'd assemble training data, fine-tune and evaluate a 2.8-trillion-parameter mixture-of-experts model, and provision roughly a terabyte-plus of GPU memory to serve it — a full ML-infrastructure project. Second, and more important, even if you nailed all of that, you'd have a model that writes in your voice and still renders zero pixels and publishes to nothing. The infrastructure gets you a better drafter, not a content operation.
Kompozy is the shortcut to the actual goal behind the fantasy — "content that sounds like me, everywhere" — without touching a GPU. The brand-consistency that people imagine fine-tuning delivers, Kompozy delivers through the Persona Brief: it governs tone, banned phrases, and audience on every text output, and a face-locked AI Influencer persona keeps your look consistent across avatar video and images — no training run, no weights to host. Then it does the part no fine-tune can: it renders the media a raw model can't (persona and avatar video, Clipped Shorts, carousels, quote cards, infographics), reframes each asset to 9:16, 1:1, and 16:9, and publishes across the eight social platforms plus blog and email on autopilot. And because Kompozy runs its own managed models under the hood, a stronger open model like K3 entering the pool is upside for the engine, not homework for you. If you have the cluster and the reasons, self-host K3 for what it's built for. If your actual goal is on-brand content shipped on a schedule, that's the layer above the model — and it's already built.
Moonshot AI published Kimi K3's open weights on Hugging Face (huggingface.co/moonshotai/Kimi-K3) on July 26, 2026, under a permissive Kimi K3 License, free to download. Downloading is free, but running it is not trivial — holding the MXFP4-quantized weights takes on the order of 1.4 terabytes of fast memory, so self-hosting realistically means a multi-GPU cluster.
Per Moonshot's model card, K3 is a mixture-of-experts model with around 2.8 trillion total parameters and roughly 104 billion activated per token (896 experts, 16 selected), a 1,048,576-token context window, and native multimodal input through a MoonViT-V2 vision encoder, with weights shipped MXFP4-quantized. Its benchmark figures are Moonshot's own published values — treat them as provisional until independent leaderboards confirm them.
For almost all creators, no. Fine-tuning a 2.8-trillion-parameter model and serving it is an ML-infrastructure project, and even done well it produces a better text drafter — not finished media or published posts. The goal behind that idea, on-brand content everywhere, is better served by a content engine like Kompozy, whose Persona Brief governs brand voice, whose face-locked personas keep your look consistent, and which renders video and graphics and publishes across nine platforms with no GPUs to manage.