// OPEN-WEIGHTS LANGUAGE MODEL REVIEW

Inkling-Small Review (2026): Honest Verdict on Thinking Machines Lab's Efficient Open Model

Inkling-Small review (2026): honest verdict on Thinking Machines Lab's efficient 276B/12B open-weights model — specs, benchmarks, pricing, and its real limits.

KompozyTurn one idea into a week of content — across every platform, published for you.
Get Started →
Last verified · 2026-07-30 · by Moe Ameen
The verdict
4.2 / 5

Inkling-Small is an impressive efficiency play: a 276B/12B open-weights multimodal model that matches the full Inkling at a quarter of the size and beats it on reasoning and agentic coding, at roughly a third of the token cost. For drafting and coding at volume it's excellent. The honest caveats: it trades some factual accuracy for speed, and — like any language model — it writes text but renders no media and publishes nothing, so it's a component in a content workflow, not the workflow.

Inkling-Small is the efficient sibling to Inkling, the first model from Thinking Machines Lab — the startup founded by former OpenAI CTO Mira Murati. It was released on July 30, 2026 as open weights on Hugging Face, about two weeks after the flagship Inkling, and I reviewed it as what it is: a fast, cheap, open language model, not a content or publishing platform it never claimed to be.

The short version is that the efficiency claim holds up as a story. It's a 276-billion-parameter mixture-of-experts with only about 12 billion parameters active per token, and Thinking Machines built it by distilling from Inkling on-policy and then scaling agentic-coding reinforcement learning for two more weeks. The result reaches comparable overall performance to Inkling at a quarter of the size and actually surpasses the larger model on reasoning and agentic-coding benchmarks — while costing roughly $1.20 per million output tokens on Tinker versus about $4.05 for Inkling. Day-0 vLLM support and free downloadable weights make it easy to run.

The honest caveats are scope and factuality. Thinking Machines states plainly that the full Inkling still leads on knowledge coverage and factuality, so Small trades some accuracy for its speed and cost — its drafts need closer checking. And like every language model, it reads text, images, and audio but outputs only text: it renders no images or video, burns no captions, designs no carousels, and publishes to nowhere. This review scores it as a model. Where it's a strong fit and where you'd need something else are both below, plus an honest note on where Kompozy sits — not as a rival, but as the layer that turns its text into published content.

What Inkling-Small is

Inkling-Small is an open-weights large language model from Thinking Machines Lab. It's a natively multimodal, encoder-free mixture-of-experts — 276B total parameters, about 12B active per token — that reads text, images (as patches), and audio (as dMel spectrograms) and responds in text, with up to a 1M-token context and controllable reasoning effort from minimal to xhigh. It was trained on NVIDIA GB300 NVL72 systems and post-trained via on-policy distillation using Inkling as the teacher, followed by continued agentic-coding RL. You can download the full weights from Hugging Face with day-0 vLLM support, or reach hosted inference and LoRA fine-tuning through the company's Tinker platform, with chat access via Tinker Playground. What it is not is a content tool: no image or video generation, no captioning, no carousel or graphic rendering, no brand-voice governance, no scheduler, and no publishing. It generates and reasons over text — scripts, summaries, captions, code — and that's where its job ends.

Who Inkling-Small is for

Inkling-Small fits developers and teams who want a cheap, controllable, open model to build on: high-volume drafting pipelines, agentic-coding and tool-use tasks where it beats the larger Inkling, and anyone with a data-control or self-hosting requirement that a hosted API can't meet. For creators specifically, it's a strong drafting engine — feed it a recording or a brief and it writes scripts and caption packs cheaply — but it's a weak fit for anyone who assumed "AI model" meant a shortcut to finished posts, because once the text is written, the work of rendering media, sizing per platform, and publishing is still entirely ahead of them, and Inkling-Small does none of it.

Scoring breakdown

DimensionScoreWhy
Efficiency & cost4.6 / 5Comparable to the full Inkling at a quarter of the size and roughly a third of the token cost — the headline strength.
Reasoning & agentic coding4.4 / 5Surpasses the larger Inkling on reasoning and agentic-coding benchmarks after extended coding RL.
Multimodal input4.0 / 5Reads text, images, and audio natively, so a recording or screenshot becomes structured text.
Context length4.2 / 5Up to a 1M-token context handles long documents and transcripts comfortably.
Openness & self-hosting4.3 / 5Free downloadable weights with day-0 vLLM support give full local control and no lock-in.
Factuality & knowledge coverage3.3 / 5The trade-off: Thinking Machines says the larger Inkling still leads here, so drafts need closer checking.
Availability & ecosystem4.2 / 5Hugging Face weights, Tinker fine-tuning and chat, and broad serving-framework support at launch.
Content & publishing capability1.5 / 5None by design — no media generation, captions, brand voice, or publishing. It stops at text.

Pros and cons

Pros

  • Genuinely efficient: matches the full Inkling at a quarter of the size and cheaper per token.
  • Beats the larger Inkling on reasoning and agentic-coding benchmarks.
  • Open weights on Hugging Face with day-0 vLLM support — full local control, no lock-in.
  • Natively multimodal input: reads audio and images directly into structured text.
  • Up to a 1M-token context plus controllable reasoning effort from minimal to xhigh.
  • Roughly $1.20 per million output tokens on Tinker makes high-volume drafting cheap.
  • Fine-tunable via Tinker, so you can shape it to your voice or domain.

Cons

  • Trades some knowledge accuracy and factuality for efficiency — drafts need closer checking than a frontier model's.
  • Text output only — no image, video, caption, carousel, or media generation of any kind.
  • Publishes to nothing; there is no scheduler or platform integration.
  • Self-hosting a 276B mixture-of-experts is a real infrastructure undertaking.
  • No brand-voice governance, banned-word filtering, or approval workflow without building one.
  • "Open weights" is not fully open source — the training corpus is not public.

Pricing analysis

On price, Inkling-Small is aggressive in the right direction. The weights are free to download from Hugging Face, so your only cost self-hosting is the GPU and infrastructure to serve a 276B mixture-of-experts — non-trivial, but yours to control. On Thinking Machines' Tinker platform, hosted inference runs at roughly $1.20 per million output tokens, versus about $4.05 for the full Inkling: a meaningful discount for work where Small's benchmark performance is comparable or better. Fine-tuning and chat access through Tinker Playground launched with a limited-time discount.

That makes it one of the more economical capable open models to draft with at volume. But it's worth being clear about what the token price does and doesn't buy: it buys text. For a creator, the total cost of turning that text into published content includes whatever you spend on image generation, video, captioning, design, and scheduling on top — tools the model doesn't provide.

For context on the other side of that line, a shipping content engine meters finished output rather than tokens. Kompozy, for example, runs credit-based tiers from $99/mo (5,500 credits) up through a sales-led enterprise plan, covering generation across 18 formats plus publishing. The two aren't priced on the same axis, and that difference — tokens versus finished, distributed posts — is the whole point of the comparison.

Use-case fit

Use caseFitWhy
High-volume text draftingStrongCheap tokens and near-flagship quality make it ideal for generating scripts and captions at scale.
Agentic coding & tool useStrongIt surpasses the larger Inkling on these benchmarks and is cheaper to run for coding agents.
Audio/image-to-text ingestStrongNative multimodal input turns recordings and screenshots into structured text directly.
Self-hosted / data-controlled deploymentsStrongFree open weights let you run it entirely on your own hardware.
High-factuality knowledge workOKCapable, but it trades some accuracy for efficiency — the larger Inkling is the safer pick here.
Making finished social postsWeakIt writes the text but renders no media and publishes nothing — most of the job is left undone.
Multi-platform publishingWeakThere is no scheduler or platform integration; publishing needs a separate engine entirely.
Brand-governed team contentWeakNo Persona Brief, banned-word filtering, or approval pipeline without building one around it.

Alternatives worth considering

  • Inkling — the full-size 975B/41B sibling, stronger on knowledge coverage and factuality when accuracy matters most.
  • Kompozy — a shipping content engine that drafts, generates 18 media formats, and publishes across nine destinations.
  • Apertus — a fully open, multilingual model for teams that want maximum openness and sovereignty.
  • Kimi K3 — Moonshot AI's large, long-context multimodal flagship for a different capability trade-off.
  • Gemma 4 — Google DeepMind's open-weight multimodal family for lighter, widely supported deployments.

How Kompozy compares

The fair way to place Kompozy against Inkling-Small is to admit they're not competing — they're adjacent halves of a pipeline. Inkling-Small is a model: it drafts, reasons, and codes, and it does that cheaply and well. Kompozy is a content engine: it drafts too, but its real work starts where the model's ends — generating the persona and avatar video, carousels, quote graphics, blogs, and newsletters a language model can't render, keeping them on-brand with a Persona Brief and HyperFrames, and publishing to the eight primary social platforms plus blog and email with scheduling and autopilot.

So if you're reading this review to decide whether Inkling-Small is "worth it" for a content operation, the honest answer is that it's worth it as a drafting layer and not much use as the whole solution. Its factuality caveat means you'll check its output regardless, and its text-only scope means you'll still need a rendering and publishing layer on top. The natural setup is to let Inkling-Small do the cheap, high-volume writing it's genuinely good at, then hand the winners to an engine like Kompozy to turn into finished, scheduled posts. One writes the words; the other ships them.

Frequently asked questions

What is Inkling-Small?

Inkling-Small is an efficient open-weights language model from Thinking Machines Lab, released July 30, 2026, about two weeks after the flagship Inkling. It is a natively multimodal mixture-of-experts — 276B total parameters, roughly 12B active per token — that reads text, images, and audio, responds in text, and supports up to a 1M-token context.

Is Inkling-Small better than Inkling?

It depends on the task. Inkling-Small matches the larger Inkling overall at a quarter of the size and surpasses it on reasoning and agentic-coding benchmarks, but Thinking Machines says the full Inkling still leads on knowledge coverage and factuality. Small favors throughput and cost; Inkling favors accuracy.

How much does Inkling-Small cost?

The weights are free to download from Hugging Face, so self-hosting costs only your infrastructure. On Thinking Machines' Tinker platform, hosted inference runs at roughly $1.20 per million output tokens (versus about $4.05 for Inkling), with fine-tuning and chat via Tinker Playground at a limited-time launch discount.

Is Inkling-Small worth it?

As a cheap, capable open model for drafting and coding at volume, yes — it's an excellent efficiency play. As a complete content solution, no: it outputs text only, so you'll still need media generation and a publishing layer, and its factuality trade-off means drafts need checking.

Can Inkling-Small generate images or video?

No. It reads images and audio as input but outputs text only — it does not render images, video, or audio. To turn its text into finished posts you pair it with a generation and publishing engine like Kompozy.

What can Inkling-Small run on?

The full weights are downloadable from Hugging Face with day-0 vLLM support, so you can self-host it, and hosted inference plus LoRA fine-tuning are available on Thinking Machines' Tinker platform. Running a 276B mixture-of-experts locally is a real infrastructure task despite the low active-parameter count.

How do I use Inkling-Small to make social content?

Use it to over-generate cheap drafts — hooks, scripts, caption packs — then bring the best into a content engine like Kompozy to render persona video, carousels, quote cards, blogs, or newsletters in your brand voice and publish them across platforms. The model writes; the engine ships.

Related deep guides

See Inkling-Small vs Kompozy comparison → · Get Started →