Claude Haiku 5.5 is the first small model from Anthropic that I would hand real coding work to, and it costs $0.10 per million input tokens. Anthropic released it on October 7, 2026, nine days after Claude Sonnet 5.5, and it beats OpenAI’s GPT-6 Luna on every benchmark where the two share a row, at the same list price.

There is a catch hiding in the pricing table. Above 100,000 tokens of prompt, the price jumps fivefold, and coding agents cross that line all the time. I went through Anthropic’s launch post, the platform docs and the Claude Code configuration docs to work out where Haiku 5.5 fits and where it will quietly cost you more than you planned.

What Is Claude Haiku 5.5?

Claude Haiku 5.5 is Anthropic’s fastest and cheapest current model, built for high-volume work like classification, extraction, routing and subagent tasks. It completes the Claude 5.5 generation that started with Opus 5.5 on September 22.

The specs from the Haiku 5.5 docs page:

  • API ID: claude-haiku-5-5 (also on Amazon Bedrock as anthropic.claude-haiku-5-5, Google Cloud, Microsoft Foundry and Claude Platform on AWS)
  • Context window: 1M tokens
  • Max output: 128K tokens, or 300K on the Batch API with a beta header
  • Thinking: adaptive, steered by effort, default medium
  • Input: text and images, text out
  • Knowledge cutoff: June 2026
  • Retirement: not sooner than October 7, 2027

Two details matter for anyone migrating code. Haiku 5.5 is the first Haiku with an effort setting, with levels from low to max. It also rejects any non-default temperature, top_p or top_k with a 400 error. If your classification pipeline sets temperature: 0 for determinism, that request will fail until you remove the parameter.

Haiku 5.5 has the same 1M context and 128K output as Sonnet 5.5 and Opus 5.5, so the trade you are making is capability per token, not window size.

Claude Haiku 5.5 Benchmarks vs GPT-6 Luna and Sonnet 5.5

Anthropic’s comparison table puts Haiku 5.5 against Haiku 4.5, GPT-6 Luna and Claude Sonnet 5.5. Terminal-Bench 4.0 is the row that tells you most about agentic coding.

Terminal-Bench 4.0 score (vendor-reported)
  • Claude Sonnet 5.5 70.6%
  • Claude Haiku 5.5 39.2%
  • GPT-6 Luna 16.4%
  • Claude Haiku 4.5 0%

Source: Anthropic, Introducing Claude Haiku 5.5 (Oct 2026)

Haiku 4.5 scored 0.0% on that benchmark. Haiku 5.5 scores 39.2%. That is the difference between a model you would never let near a terminal and one that finishes a meaningful share of long shell tasks on its own.

The full table is more interesting than the headline chart:

BenchmarkHaiku 5.5GPT-6 LunaSonnet 5.5
Terminal-Bench 4.039.2%16.4%70.6%
FrontierCode 1.1 (Main)46.4%42.4%52.1%
OSWorld 2.1 (offline subset)72.4%48.9%83.9%
GDPval-AA v2.1 (Elo)162014371840
Humanity’s Last Exam, with tools57.4%not reported64.5%

(Source: Anthropic, Introducing Claude Haiku 5.5, October 7, 2026. All scores vendor-reported; Sonnet 5.5’s FrontierCode score is at xhigh effort.)

Look at FrontierCode next to Terminal-Bench. On FrontierCode, Haiku 5.5 is 5.7 points behind Sonnet 5.5. On Terminal-Bench 4.0 it is 31.4 points behind.

My reading is that Haiku 5.5 writes code nearly as well as Sonnet when the task is bounded, and falls apart much faster when it has to drive a long, stateful session: run the build, read the log, fix, retry. That matches how Anthropic itself pitches the model. Its launch post recommends Sonnet 5.5 and Opus 5.5 for complex agentic coding and Haiku 5.5 as a coding subagent working alongside them.

Against GPT-6 Luna, the gap is large on agentic work. Haiku 5.5 scores 39.2% to Luna’s 16.4% on Terminal-Bench 4.0 and 72.4% to 48.9% on the OSWorld 2.1 offline subset, and SiliconANGLE notes the two models have identical list prices. These are Anthropic’s numbers for OpenAI’s model, so wait for independent runs before treating them as settled.

One sanity check from the press coverage: Decrypt’s reporter asked Haiku 5.5 a simple logic question and got a near-instant wrong answer. One question proves nothing statistically, but it is a fair reminder that the default effort is medium, and speed is the point of this model.

Haiku 5.5 is the strongest small model on Anthropic’s table by a wide margin, and it still trails Sonnet 5.5 badly on long terminal sessions.

Claude Haiku 5.5 Pricing and the 100K Token Cliff

Here is the pricing from the Haiku 5.5 docs page, per million tokens:

ModelInputOutputCache read
Claude Haiku 5.5 (prompt ≤100K)$0.10$0.50$0.01
Claude Haiku 5.5 (prompt >100K)$0.50$2.50$0.05
GPT-6 Luna$0.10$0.50not compared
Claude Haiku 4.5$1.00$5.00$0.10
Claude Sonnet 5.5$2.00$10.00$0.10

(Sources: Claude Haiku 5.5 docs and Anthropic’s launch post, October 2026; GPT-6 Luna list price as reported by SiliconANGLE. Batch API is 50% off on all Claude rows.)

The headline is that Haiku 5.5 is 90% cheaper than Haiku 4.5 below 100K tokens. There are two adjustments you need to make before you believe that for your workload.

The tokenizer counts more tokens. The docs say Haiku 5.5 uses the newer tokenizer from Claude 4.7 onward, so the same text counts as roughly 30% more tokens than on Haiku 4.5. A prompt that cost you 1M tokens on Haiku 4.5 is about 1.3M tokens now, so the effective input price is closer to $0.13 per old million. That is still about 87% cheaper.

The price is set by prompt size, not by the extra tokens. Prompts over 100,000 tokens are billed at $0.50 input and $2.50 output for the whole request, per the docs table. This is the part I would plan around.

Coding subagents accumulate context. Each turn re-sends the conversation, the file reads and the tool results. A subagent that starts at 30K tokens and reads a dozen files can pass 100K by turn ten, and from then on every turn costs five times as much.

Here is my rough math for 1,000 subagent calls, each producing 2,000 output tokens, ignoring caching:

  • 40K-token prompts on Haiku 5.5: 40M input × $0.10 + 2M output × $0.50 = $5
  • 150K-token prompts on Haiku 5.5: 150M × $0.50 + 2M × $2.50 = $80
  • 40K-token prompts on Sonnet 5.5: 40M × $2 + 2M × $10 = $100

The jump from $5 to $80 is not because the prompt is nearly four times longer. It is because the per-token price rose fivefold at the same time. Even at the higher tier, Haiku 5.5 is still four times cheaper per token than Sonnet 5.5, so it stays the cheaper model. The cliff just shrinks your savings from 95% to 75%.

If you are already trying to keep agent context tight for cost reasons, the techniques in my guide to reducing Claude API costs matter more on Haiku 5.5 than on any earlier model, because staying under 100K is now worth a 5x discount.

Anthropic also halved cache reads on Claude Sonnet 5.5, from $0.20 to $0.10 per million tokens, at the same launch. The company says that cuts the cost of most agentic Sonnet 5.5 work by about 20%. That changes the numbers in my GPT-6.1 Sol vs Claude Sonnet 5.5 comparison, where GPT-6.1 Sol’s $0.10 cached input was half Sonnet’s old rate. The two now match. The AI model pricing tracker now lists both changes with sources.

Below 100K tokens, Claude Haiku 5.5 is about 87% cheaper than Haiku 4.5 even after the tokenizer change; above 100K, the discount shrinks to about 35%.

How to Use Claude Haiku 5.5 in Claude Code

This is where Haiku 5.5 will save most developers money without them writing any API code.

According to the Claude Code model configuration docs, the haiku alias resolves to Haiku 5.5 on the Anthropic API and requires Claude Code v2.1.293 or later. On Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS, the same alias still resolves to Haiku 4.5. If you use Claude Code through a cloud provider, pin the full model ID rather than trusting the alias.

The built-in Explore subagent does not use Haiku by default. The subagents docs say Explore runs on the main conversation’s model, so if your session runs Opus 5.5, every codebase search runs on Opus too. To move it to Haiku 5.5, define a project or user subagent named Explore with its own model field:

---
name: explore
description: Fast read-only codebase search
model: haiku
---

The docs say a subagent you define with that name overrides the built-in one and keeps its own model field.

If you want every subagent on Haiku, the docs give a settings.json option:

{
  "env": {
    "CLAUDE_CODE_SUBAGENT_MODEL": "haiku",
    "CLAUDE_CODE_SUBAGENT_MODEL_FORCE": "1"
  }
}

I would not use the forced version. It also overrides subagents you deliberately set to opus or sonnet, like a reviewer that needs the stronger model. Setting model: haiku on the specific subagents that do search, summarisation and log reading is more work and much safer.

A pattern that fits the benchmark data: keep Sonnet 5.5 or Opus 5.5 as the main session, which drives the terminal loop where Haiku’s 39.2% Terminal-Bench score would hurt. Send exploration, file summaries and test-output triage to Haiku subagents, which are bounded tasks where its FrontierCode score holds up.

The cheapest Claude Code change this week is a one-line model: haiku on your Explore subagent, as long as you are on v2.1.293 or later and using the Anthropic API.

Where Claude Haiku 5.5 Falls Short

Haiku 5.5 is a big jump, and there are still places I would not use it.

Long autonomous coding sessions. A 39.2% Terminal-Bench 4.0 score means the model fails most long terminal tasks that Sonnet 5.5 completes. Do not make it your main agent for refactors that span many files and many build cycles.

Security work. Anthropic says Haiku 5.5’s cyber safeguards are stricter than Haiku 4.5’s but looser than Sonnet 5.5’s. They allow more defensive tasks while still blocking penetration testing and attacker-style techniques. Organisations that need more can apply to Anthropic’s Cyber Verification Program.

Sampling-dependent pipelines. Anything that relies on temperature, top_p or top_k needs rework. You steer Haiku 5.5 with effort and prompting now.

Cloud-provider users expecting the alias to upgrade. As noted above, Claude Code’s haiku alias stays on Haiku 4.5 on every non-Anthropic provider.

Claude Haiku 5.5’s weak points are the long, stateful agent loops where Sonnet and Opus earn their price.

Who Should Switch to Claude Haiku 5.5?

If you run Haiku 4.5 in production today, switch after you remove the sampling parameters and re-run your evals. The benchmark jump is large, the price drops by up to 90%, and the 1M window is unchanged. Watch your token counts for the first week, because the tokenizer change will make your dashboards look 30% busier.

If you use GPT-6 Luna for cheap classification or routing, Haiku 5.5 is worth a direct bake-off at the same list price. Anthropic’s table favours Haiku on agentic work, and you can measure on your own data within a day.

If you mainly use Claude Code, the change is the Explore subagent setting above. Anthropic is also adding monthly API credits for Max and Team subscribers, as 9to5Mac reported: $100 for Max 5x, $200 for Max 20x and up to $500 pooled for Team. That is enough to run a lot of Haiku subagent calls for experiments.

One more reason to keep cheap subagents on a short leash: a subagent that can run npm install can pull in a malicious package as easily as a main agent can. I covered how those attacks work, and the settings that block them, in today’s post on AI coding agent supply chain attacks.

Claude Haiku 5.5 should be the default model for every Claude subagent that reads and summarises, and it should stay off the main agent’s chair for long coding sessions.