Thinking Machines Lab's efficient open-weights model — a 276B/12B multimodal MoE that matches the larger Inkling at a quarter of the size.
Last verified · 2026-07-30 · by Moe Ameen
Inkling-Small is the efficient sibling to Inkling, the first model from Thinking Machines Lab — the AI startup founded by former OpenAI CTO Mira Murati. It was released on July 30, 2026 as open weights on Hugging Face, roughly two weeks after the flagship Inkling. It is a natively multimodal mixture-of-experts model with 276 billion total parameters and about 12 billion active for any given token, so it runs cheap and fast for its class. It uses the same encoder-free multimodal architecture as Inkling — audio is read as dMel spectrograms and images as patches — reads text, images, and audio, responds in text, and supports up to a 1M-token context. Thinking Machines says it was trained on NVIDIA GB300 NVL72 systems.
The headline claim is efficiency: Inkling-Small reaches comparable performance to Inkling at a quarter of its size, and on reasoning and agentic-coding benchmarks it actually surpasses the larger model. Thinking Machines got there by post-training an earlier checkpoint with on-policy distillation using Inkling as the teacher, then continuing to scale agentic-coding reinforcement learning for two more weeks. Like Inkling it exposes controllable reasoning effort — dialable from minimal to xhigh — so you trade latency and cost against depth, and it handles tool use and coding. The honest trade-off the lab states plainly: the full-size Inkling still leads on knowledge coverage and factuality, so Small is the pick for throughput and cost, not for maximum recall.
The full weights are downloadable from Hugging Face and shipped with day-0 vLLM support. You can also reach hosted inference and fine-tuning through the company's Tinker platform, where Inkling-Small runs at roughly $1.20 per million output tokens versus about $4.05 for Inkling — and fine-tuning plus chat access is available through Tinker Playground with a limited-time launch discount.
One thing to be clear about: Inkling-Small is a text-output model. It reads images and audio as input but generates no images, video, or audio — it writes, reasons, and codes, but renders no media and publishes nothing. And "open weights" is not the same as fully open source: the checkpoint and its serving support are public, but the full training corpus is not. Like any LLM, its output can be wrong — more so than the larger Inkling on factual recall — so check it before it ships.
Inkling-Small's edge over the full Inkling is cost per draft, not depth — a quarter the size at roughly $1.20 per million output tokens means you can generate ten hook variants, five caption angles, and a fortnight of text-post options for the price of a single pass on a frontier model. That makes it the natural workhorse for volume: fine-tune it on your back catalog through Tinker, run it self-hosted, and it churns out on-voice copy all day. What it will not do is turn a word of that into something you can post. It outputs text only — no images, no video, no burned-in captions, no schedule, no publish step.
Kompozy is the half that makes the volume matter. Pipe an Inkling-Small script or an ingested-recording summary in, and Kompozy generates the media the model can't: HeyGen persona and avatar shorts with captions, face-locked Persona Photos and Persona Tweets, multi-slide Carousels and Quote Graphics rendered pixel-exact through HyperFrames, plus full blog articles and newsletters — each held to your brand voice by the Persona Brief. Then it publishes, fanning every piece across the nine supported platforms (Instagram, Facebook, TikTok, YouTube, LinkedIn, X, Pinterest, Threads) plus Mailchimp email and blog, on a schedule, on autopilot, with a per-post review pipeline. Because Small drafts so cheaply, the practical move is to over-generate — let it write far more options than you need, then use Kompozy to render and ship only the ones that land. The cheap model floods the top of the funnel with words; Kompozy converts the winners into finished, scheduled posts.
Inkling-Small is an efficient open-weights model from Thinking Machines Lab, released on July 30, 2026, about two weeks after the flagship Inkling. It is a natively multimodal mixture-of-experts model — 276B total parameters with roughly 12B active per token — that reads text, images, and audio, responds in text, and supports up to a 1M-token context.
It is a quarter the size (276B vs 975B total) and cheaper to run, yet reaches comparable overall performance and actually surpasses Inkling on reasoning and agentic-coding benchmarks. The trade-off Thinking Machines states is that the larger Inkling still leads on knowledge coverage and factuality, so Small favors throughput and cost over maximum recall.
The weights are freely downloadable from Hugging Face with day-0 vLLM support, and hosted inference and fine-tuning are available on Thinking Machines' Tinker platform (around $1.20 per million output tokens). It is best described as "open weights" rather than fully open source: the checkpoint is public, the full training corpus is not. Your real cost is the infrastructure to run it.
No. It reads images and audio as input but outputs text only — it does not render images, video, or audio. To turn its text into finished posts you pair it with a generation and publishing engine like Kompozy, which produces the media and publishes across platforms.
Because it drafts cheaply, over-generate copy with it, then bring the best scripts or caption sets into Kompozy to render persona video, carousels, quote cards, blogs, or newsletters in your brand voice — and schedule and publish the set across TikTok, Reels, Shorts, X, LinkedIn, and more from one queue.