// AI NEWS · MODEL RELEASE

Anthropic Launches Claude Haiku 5.5, Its Cheapest and Fastest Small Model, at $0.10 per Million Input Tokens

Released October 7, 2026 as the third model in the Claude 5.5 family, Haiku 5.5 runs about 75% cheaper on average than Haiku 4.5 and is the first Haiku-class model with an adjustable effort setting.

2026-10-07 · by Moe Ameen

What happened

Anthropic released Claude Haiku 5.5 on October 7, 2026, calling it "the cheapest, fastest, and most capable small model we've ever released." It is the third model in the Claude 5.5 family, following Opus 5.5 on September 22 and Sonnet 5.5 on September 28, and is reachable under the identifier `claude-haiku-5-5` on the Claude Platform API and across the major clouds — Amazon Web Services, Google Cloud, and Microsoft Azure.

The headline is price. Haiku 5.5 uses tiered pricing that turns on a 100,000-token prompt boundary: for prompts up to 100k tokens it is $0.10 per million input tokens and $0.50 per million output, and above that it rises to $0.50 input and $2.50 output. Against Haiku 4.5 — which was $1 input and $5 output — Anthropic puts the short-prompt cut at about 90% and the long-prompt cut at about 50%, for an average cost roughly 75% lower. Cache reads start at $0.01 per million tokens. Anthropic also trimmed Sonnet 5.5's cache-read price to $0.10 in the same update.

Haiku 5.5 is the first Haiku-class model with an adjustable effort setting (Low, Medium, High, Xhigh, Max), so you can trade depth for cost per call. Anthropic positions it for high-volume, cost-sensitive work — summaries, classification, extraction, quick lookups — and for speed-sensitive jobs like live customer support, browser use, and acting as a fast subagent paired with Opus 5.5 or Sonnet 5.5 on coding. It is honest about the ceiling: on hard agentic work such as Terminal-Bench 4.0, Haiku 5.5 scores well below Sonnet 5.5, which remains the pick when a task needs more judgment. Treat specific figures as a launch-day snapshot and verify current numbers on Anthropic's pricing page.

Why it matters for creators

  • The cost floor for high-volume content work just dropped. The mechanical steps behind a content pipeline — transcribing and summarizing a source, classifying topics, drafting captions and metadata, routing — are exactly the cheap, fast, repetitive jobs Haiku 5.5 is built for, so running them at a daily cadence gets materially cheaper.
  • Speed compounds across a batch. A small model that returns answers faster is not just nicer in a chat; across hundreds of per-post operations it is the difference between a content run finishing in minutes versus stalling, which is where the launch customers reported 2x-plus latency wins.
  • It is still a text model, so the saving lands on drafting, not production. Haiku 5.5 generates no images, video, or audio and publishes nothing — the parts of a content workflow that were expensive (rendering media, designing, scheduling, distributing) do not get cheaper because a small model did.
  • Model choice is becoming a tuning knob, not a commitment. With an effort setting and a 90%-cheaper short-prompt rate, the smart pattern is routing cheap mechanical steps to Haiku and reserving Sonnet or Opus for the hard ones — orchestration most creators do not want to build by hand.
  • The pricing arms race favors anyone who runs AI at volume. A 75%-average cut at the small-model tier, with a matching Sonnet cache-read trim, keeps pushing the economics of automated, multi-format content in the creator's direction.

How to act on this with Kompozy

The way to read this launch as a creator is not "a new chat model" — it is "the per-operation cost of running an automated content pipeline just fell." A finished post is the sum of many small AI steps: summarize the transcript, pull the angles, classify the topic, draft the caption, write the alt text and metadata. Those are precisely the high-volume, latency-sensitive jobs Haiku 5.5 is priced and tuned for. The catch is that wiring a small model to each of those steps, then deciding when to escalate to Sonnet or Opus for the parts that need judgment, is an orchestration problem — and building it yourself is a project, not an afternoon.

[Kompozy](/) is where this already plays out without you touching an API key. It runs Claude generation under the hood and manages model choice internally, so the mechanical steps ride cheap, fast models while the writing that carries your brand stays Claude-class — governed by your [Persona Brief](/glossary/persona-brief). What you act on today is the output, not the plumbing: drop in one source and Kompozy fans it into a [Persona Short](/glossary/persona-shorts), a brand-exact [HyperFrames](/glossary/hyperframes) carousel, Quote Graphics, a Blog Article, and an Email Newsletter, then [Autopilot](/glossary/autopilot) captions and reframes each to 9:16, 1:1, and 16:9 and schedules the set across eight social platforms plus blog and email behind a per-post review. Cheaper small models make the pipeline cheaper to run; Kompozy is the pipeline.

Quick takeaways

  • Claude Haiku 5.5 launched October 7, 2026 as the third model in the Claude 5.5 family, under the identifier claude-haiku-5-5.
  • Pricing is tiered at a 100k-token prompt boundary: $0.10/$0.50 per million input/output up to 100k, $0.50/$2.50 above — about 90% cheaper than Haiku 4.5 for short prompts, 50% cheaper for long, ~75% on average.
  • It is the first Haiku model with an adjustable effort setting (Low, Medium, High, Xhigh, Max) and is available on the Claude API, AWS, Google Cloud, and Azure.
  • Anthropic positions it for high-volume, cost-sensitive and speed-sensitive work and as a fast subagent; Sonnet 5.5 and Opus 5.5 remain better for hard agentic tasks.
  • The cut lands on drafting, not production — Haiku 5.5 makes no media and publishes nothing. An engine like Kompozy routes cheap models to mechanical steps and then renders and publishes finished, on-brand content.

Frequently asked questions

How much does Claude Haiku 5.5 cost?

It uses tiered pricing at a 100,000-token prompt boundary. For prompts up to 100k tokens it is $0.10 per million input tokens and $0.50 per million output; above 100k it is $0.50 input and $2.50 output, with cache reads from $0.01 per million. Anthropic says that is about 90% cheaper than Haiku 4.5 for short prompts, 50% cheaper for long prompts, and roughly 75% cheaper on average. Verify current figures on Anthropic's pricing page.

When did Claude Haiku 5.5 launch?

Anthropic released it on October 7, 2026, as the third model in the Claude 5.5 family after Opus 5.5 (September 22) and Sonnet 5.5 (September 28). It is available on the Claude Platform API and on AWS, Google Cloud, and Azure.

Is Claude Haiku 5.5 good for making content?

It is a fast, cheap text and reasoning model — excellent for the high-volume mechanical steps behind a content pipeline (summarizing, classifying, drafting captions and metadata) but not a content tool on its own. It generates no images, video, or audio and publishes nothing. To turn its output into finished, published posts you pair it with a content engine like Kompozy that handles media, design, and multi-platform publishing.

How is Haiku 5.5 different from Sonnet 5.5?

Haiku 5.5 is the small, cheap, fast tier; Sonnet 5.5 is the mid tier with more capability on hard tasks. Anthropic says Sonnet 5.5 and Opus 5.5 are the better choices for complex agentic work, while Haiku 5.5 is tuned for cost- and speed-sensitive jobs and as a subagent. The practical pattern is routing mechanical steps to Haiku and escalating to Sonnet or Opus when a task needs judgment.

Related news

← All AI news · Get started →