// AI TOOLS · INKLING-SMALL

Inkling-Small

Thinking Machines Lab's efficient open-weights model — a 276B/12B multimodal MoE that matches the larger Inkling at a quarter of the size.

KompozyTurn one idea into a week of content — across every platform, published for you.
Get Started →

Last verified · 2026-07-30 · by Moe Ameen

What Inkling-Small is

Inkling-Small is the efficient sibling to Inkling, the first model from Thinking Machines Lab — the AI startup founded by former OpenAI CTO Mira Murati. It was released on July 30, 2026 as open weights on Hugging Face, roughly two weeks after the flagship Inkling. It is a natively multimodal mixture-of-experts model with 276 billion total parameters and about 12 billion active for any given token, so it runs cheap and fast for its class. It uses the same encoder-free multimodal architecture as Inkling — audio is read as dMel spectrograms and images as patches — reads text, images, and audio, responds in text, and supports up to a 1M-token context. Thinking Machines says it was trained on NVIDIA GB300 NVL72 systems.

The headline claim is efficiency: Inkling-Small reaches comparable performance to Inkling at a quarter of its size, and on reasoning and agentic-coding benchmarks it actually surpasses the larger model. Thinking Machines got there by post-training an earlier checkpoint with on-policy distillation using Inkling as the teacher, then continuing to scale agentic-coding reinforcement learning for two more weeks. Like Inkling it exposes controllable reasoning effort — dialable from minimal to xhigh — so you trade latency and cost against depth, and it handles tool use and coding. The honest trade-off the lab states plainly: the full-size Inkling still leads on knowledge coverage and factuality, so Small is the pick for throughput and cost, not for maximum recall.

The full weights are downloadable from Hugging Face and shipped with day-0 vLLM support. You can also reach hosted inference and fine-tuning through the company's Tinker platform, where Inkling-Small runs at roughly $1.20 per million output tokens versus about $4.05 for Inkling — and fine-tuning plus chat access is available through Tinker Playground with a limited-time launch discount.

One thing to be clear about: Inkling-Small is a text-output model. It reads images and audio as input but generates no images, video, or audio — it writes, reasons, and codes, but renders no media and publishes nothing. And "open weights" is not the same as fully open source: the checkpoint and its serving support are public, but the full training corpus is not. Like any LLM, its output can be wrong — more so than the larger Inkling on factual recall — so check it before it ships.

What you can make with it

  • High-volume first-draft hooks, scripts, and caption packs cheaply, thanks to the low active-parameter cost
  • Structured summaries, takeaways, and outlines reasoned directly from a raw audio recording or a screenshot (native multimodal input)
  • Blog drafts, newsletter sections, and text posts from a rough brief or a transcript
  • A self-hosted drafting model a team can run and fine-tune on its own infrastructure via Tinker at a low per-token cost
  • Agentic coding and tool-use tasks where its benchmark strength beats the larger Inkling

How Kompozy turns Inkling-Small output into content

Inkling-Small's edge over the full Inkling is cost per draft, not depth — a quarter the size at roughly $1.20 per million output tokens means you can generate ten hook variants, five caption angles, and a fortnight of text-post options for the price of a single pass on a frontier model. That makes it the natural workhorse for volume: fine-tune it on your back catalog through Tinker, run it self-hosted, and it churns out on-voice copy all day. What it will not do is turn a word of that into something you can post. It outputs text only — no images, no video, no burned-in captions, no schedule, no publish step.

Kompozy is the half that makes the volume matter. Pipe an Inkling-Small script or an ingested-recording summary in, and Kompozy generates the media the model can't: HeyGen persona and avatar shorts with captions, face-locked Persona Photos and Persona Tweets, multi-slide Carousels and Quote Graphics rendered pixel-exact through HyperFrames, plus full blog articles and newsletters — each held to your brand voice by the Persona Brief. Then it publishes, fanning every piece across the nine supported platforms (Instagram, Facebook, TikTok, YouTube, LinkedIn, X, Pinterest, Threads) plus Mailchimp email and blog, on a schedule, on autopilot, with a per-post review pipeline. Because Small drafts so cheaply, the practical move is to over-generate — let it write far more options than you need, then use Kompozy to render and ship only the ones that land. The cheap model floods the top of the funnel with words; Kompozy converts the winners into finished, scheduled posts.

  1. Fine-tune Inkling-Small on your back catalog via Tinker, or run the open weights self-hosted, to draft copy at low per-token cost.
  2. Over-generate: have it produce many hooks, scripts, and caption variants from one brief or recording.
  3. Bring the strongest drafts into Kompozy and pick formats — persona or avatar short, carousel, quote card, blog, newsletter, text posts.
  4. Let Kompozy render each in your voice via the Persona Brief, with branded captions and per-platform reframing.
  5. Schedule and publish the finished set across TikTok, Reels, Shorts, X, LinkedIn, and more from one queue.

Frequently asked questions

What is Inkling-Small?

Inkling-Small is an efficient open-weights model from Thinking Machines Lab, released on July 30, 2026, about two weeks after the flagship Inkling. It is a natively multimodal mixture-of-experts model — 276B total parameters with roughly 12B active per token — that reads text, images, and audio, responds in text, and supports up to a 1M-token context.

How is Inkling-Small different from Inkling?

It is a quarter the size (276B vs 975B total) and cheaper to run, yet reaches comparable overall performance and actually surpasses Inkling on reasoning and agentic-coding benchmarks. The trade-off Thinking Machines states is that the larger Inkling still leads on knowledge coverage and factuality, so Small favors throughput and cost over maximum recall.

Is Inkling-Small free and open source?

The weights are freely downloadable from Hugging Face with day-0 vLLM support, and hosted inference and fine-tuning are available on Thinking Machines' Tinker platform (around $1.20 per million output tokens). It is best described as "open weights" rather than fully open source: the checkpoint is public, the full training corpus is not. Your real cost is the infrastructure to run it.

Can Inkling-Small generate images or video?

No. It reads images and audio as input but outputs text only — it does not render images, video, or audio. To turn its text into finished posts you pair it with a generation and publishing engine like Kompozy, which produces the media and publishes across platforms.

How do I turn Inkling-Small drafts into social posts?

Because it drafts cheaply, over-generate copy with it, then bring the best scripts or caption sets into Kompozy to render persona video, carousels, quote cards, blogs, or newsletters in your brand voice — and schedule and publish the set across TikTok, Reels, Shorts, X, LinkedIn, and more from one queue.

Related tools

  • InklingThinking Machines Lab's first model — a large, natively multimodal open-weights LLM built to be customized, not rented.
  • ApertusA fully open, multilingual foundation model built in Switzerland for sovereign AI.
  • Kimi K3Moonshot AI's new flagship frontier model — a very large, long-context, natively multimodal model that reads images and reasons over million-token inputs, positioned as the largest open-weight model from China.
  • Gemma 4Google DeepMind's open-weight multimodal model family — reads images and audio, generates text, and runs fast and cheap.

← All AI tools · Get started →