What you’ll walk away with
By the end of this you’ll have Claude Haiku 5.5 answering real work in the free Claude.ai chat and a clean API call that won’t throw a 400 or surprise-bill you. I spent the first evening after the October 7, 2026 drop wiring a meeting-note compressor and hit every trap the launch posts glossed over. Here’s the path that actually worked.
Haiku 5.5 is Anthropic’s small, fastest model for high-volume work: summaries, extraction, routing, compaction, live support, browser use, and cheap subagents under Sonnet or Opus. Free, Pro, Max, Team, and Enterprise can select it on Claude.ai (web, iOS, Android) as of the launch window.
The two-minute background
Ships as model ID claude-haiku-5-5. 1M context. 128K max output (300K on Batch with the beta header). Knowledge cutoff June 2026. First Haiku with an effort dial (low → max; default medium) and adaptive thinking on by default. Per Anthropic’s launch notes (Oct 7, 2026), list price is $0.10 input / $0.50 output per MTok for prompts up to 100K – about 90% lower than the prior Haiku card under that line – with ~75% average savings after tokenizer mix and real traffic shape.
Benchmarks they published the same day: OSWorld 2.1 at 72.4% (was 15.7% on Haiku 4.5), Terminal-Bench 4.0 at 39.2% (was 0%), GDPval-AA 1620 (was 735). Useful signal. Not a reason to skip the billing traps below.
Method A vs Method B: chat or API?
I ran both the same night.
| Path | Best for | Cost | Gotchas |
|---|---|---|---|
| Method A: Claude.ai | Quick tests, Free/Pro, no code | Included on Free+ | Usage limits; no raw API key |
| Method B: Messages API | Volume, agents, automation | From $0.10/$0.50 per MTok (≤100K prompts) | Param breaks, 100K tier, new tokenizer |
Beginners: start on Method A. Prove the prompt on your real text, then copy the winner into Method B. Order below matches that.
Walkthrough: Haiku 5.5 on Claude.ai (Free works)
Open claude.ai (or iOS/Android). Tap the model name near send. Pick Haiku 5.5. Missing? “More models.”
Effort lives in the same menu. Leave medium on the first pass. Drop to low when the task is dumb-simple and you want less thinking overhead.
- Paste a real short task. Mine: “Turn this messy meeting note into exactly 3 action items, one line each. No preamble.nnNote: Jordan still blocked on the invoice export – needs API retry budget. Priya ships the mobile crash fix Friday if QA signs off. Alex to send the pricing one-pager to finance before Thursday standup.”
- Send. Check speed and whether you got three clean lines.
- Flip effort once. Same note. See if quality moves or you just paid for longer thinking.
- Save the prompt that worked. That string goes straight into the API.
Pro tip: Free and Pro stretch further on Haiku than on Sonnet – lighter hit on the usage limit. Burn classification/summary experiments here while you tune.
That loop is dull on purpose. Ten minutes of real notes beat another polished demo you’ll never ship.
Then the API: one call that won’t 400
Model string: claude-haiku-5-5 (Bedrock: anthropic.claude-haiku-5-5). Adaptive thinking defaults on; effort defaults medium – spelled out in the platform overview.
import anthropic
client = anthropic.Anthropic() # ANTHROPIC_API_KEY in env
msg = client.messages.create(
model="claude-haiku-5-5",
max_tokens=1024,
output_config={"effort": "low"}, # low | medium | high | xhigh | max
messages=[{
"role": "user",
"content": "Turn this messy meeting note into exactly 3 action items, one line each. No preamble.nnNote: Jordan still blocked on the invoice export - needs API retry budget. Priya ships the mobile crash fix Friday if QA signs off. Alex to send the pricing one-pager to finance before Thursday standup."
}],
)
# Never assume first block is text - thinking can land first
text = next(b.text for b in msg.content if b.type == "text")
print(text)
Hard-won rules from the migration guide – skip these and you eat HTTP 400:
- Omit
temperature,top_p,top_k. Non-defaults (or temp + top_p together) fail. - No
budget_tokens. Adaptive + effort replaces it. - End on a user turn. Assistant prefills fail too.
Raise max_tokens if your old 4.5 ceiling was tight. Thinking tokens count against the cap – they can swallow the whole budget and return empty visible text. Select content by type; don’t assume content[0] is the answer.
Edge cases that actually bite
The 100K line is not a soft warning. Prompts ≤100,000 tokens bill at $0.10 / $0.50. Cross it and the entire request – input, output, cache reads – flips to $0.50 / $2.50. Anthropic’s own note: ~90% of prior Haiku traffic sat under the line. Long thread + preserved thinking blocks still shove you over. Count with claude-haiku-5-5, not your old 4.5 totals.
Turns out the same English runs ~30% more tokens on the newer tokenizer (shared with Claude 4.7+). Stack default thinking on top and “cheap” jobs drift. I stopped reusing 4.5 token estimates entirely.
Force a tool with tool_choice any/named and you get the tool call without a thinking block – use auto when you want reasoning first. Handle stop_reason: "refusal" yourself; there’s no server fallback model.
Priority Tier capacity you bought for Haiku 4.5 does not carry over. Plan Haiku 5.5 capacity as its own line item if you rely on reserved throughput.
Is the 100K cliff a cost-control lever or a foot-gun for anyone who pastes full threads? Still unsure. Chart your p95 prompt length once – you’ll know which camp you’re in.
FAQ
Is Claude Haiku 5.5 free?
Yes on Claude.ai for Free, Pro, Max, Team, and Enterprise. API starts at $0.10 / $0.50 per MTok when prompts stay ≤100K (as of the Oct 2026 card).
When should I pick Haiku 5.5 over Sonnet?
Picture 8,000 support notes overnight, or a subagent that only routes and compacts while Sonnet owns multi-step coding. That’s Haiku’s lane: latency-sensitive, narrow, high-volume. Anthropic’s launch numbers show a big jump on computer-use and knowledge-work style evals versus Haiku 4.5 (OSWorld 72.4%, Terminal-Bench 39.2%, GDPval-AA 1620). Sonnet still wins when one wrong step is expensive – keep it on the hard agent loop and farm the repetitive slice downward.
Why did my old Haiku 4.5 code start returning 400?
Sampling knobs, budget_tokens, assistant prefills, stale computer-use tool versions, and replaying edited turns with thinking blocks. Swap the model ID, strip those knobs, select blocks by type, recount tokens. Full checklist: migration guide linked above – no need to re-debug from memory.
Open Claude.ai, switch to Haiku 5.5, run one real note or paragraph from your backlog. Paste the winning prompt into the snippet. That’s the setup.