Skip to content

Prompting Claude Opus 5.5: What Actually Works

Drop "think hard," hand whole tasks with a finish line, and set effort on purpose. Hands-on Opus 5.5 patterns that cut stalls, cache misses, and quiet agent fake-dones.

8 min readBeginner

Two camps showed up the week Claude Opus 5.5 launched. One still wraps every ask in “think carefully / step by step / verify twice.” The other states the full outcome, a checkable finish line, and when the model may interrupt – then lets the effort dial set depth. On this release the second camp finishes more often and spends less.

Opus 5.5 shipped 22 September 2026 (model id claude-opus-5-5). As of those launch posts, Anthropic pegs typical default workloads around 40% lower cost than Opus 5, with output tokens more than 30% faster. Pricing on the API: $4 / MTok input, $20 / MTok output, cache reads $0.20 / MTok. The official prompting guide is blunt: Opus 5 prompts still run, but a few old habits now buy latency for no quality.

This is not a recap. It’s what to type, what to delete, and where runs already stall.

Quick context: what actually changed for prompting

Thinking is always on. You cannot disable it. Depth is the effort setting: low, medium (default), high, xhigh, max. Opus 5 defaulted to high; here the default is medium. In Anthropic’s coding and knowledge-work tests, medium matched or beat Opus 5 at high – often on fewer tokens.

Prompt for outcomes and stop conditions. Control how hard it thinks with effort, not magic phrases.

Lever Opus 5 habit Opus 5.5 move
Thinking Sometimes disabled for speed Always on; use low effort for speed
Default effort high (if unset) medium – set it on purpose
“Think carefully” System-prompt filler Delete; first tokens often arrive sooner
Task framing Step lists and micro-prompts Whole task + checkable “done”
max_tokens Sized for visible answer only Must cover thinking + answer

Context window is 1M tokens; max output is 128K. Fast mode exists as a research preview (as of launch: up to ~2.5× speed at $8 / $40 per MTok in Claude Code and Platform) when every turn is blocking you.

Hands-on: prompts that actually finish

Start in the Claude app or Claude Code on Opus 5.5. Skip the five-part role/format essay for a minute. One message: job, done state, interrupt policy.

Pattern A – knowledge work (planning-deck audit)

Audit this planning deck against the attached spreadsheet.

Done means:
- every chart number matches the sheet (quote mismatches)
- every date that implies a weekday is correct
- names and owners are consistent across slides

Output a table: slide #, issue, evidence quote, severity (block / fix-later).
Mark anything you could not confirm and where you looked.
Stop and ask only if a file is unreadable or two sources conflict with no way to choose.

Why this shape works: you’re defining acceptance tests, not coaching chain-of-thought. Attach the files; don’t retype numbers. Official capability notes put Opus 5.5 stronger on dense visuals and long docs than Opus 5 – this framing uses that.

Pattern B – agentic coding (finish line + interrupt policy)

Refactor the notification service onto the new queue client.
Done means: every producer uses the new client, the legacy queue
helpers are deleted, and integration tests pass green.
Stop and ask me only if a test fails for a reason you can't explain
or before any migration that touches production config.

Same idea as Anthropic’s Claude Code playbook, different task. Early community runs of multi-hour agent jobs needed less babysitting once “done” was checkable and the stop rule was narrow.

Pro tip: Mid-run, don’t restart. Send a follow-up while it works (“Also keep the old topic names as aliases”). Restarts hurt more now because runs go longer.

Effort: the dial that replaced “think harder”

In Claude Code / API, set effort explicitly. On a fresh workload I do this:

  1. Baseline at medium. Measure quality, tokens, wall time on your own eval set.
  2. If quality already holds and cost matters, try low. Anthropic reports low often lands close on several coding evals at much lower cost.
  3. Push high / xhigh / max only where medium misses and you’ve measured a real gain. At the same label, Opus 5.5 tends to think more per turn than Opus 5 – especially at the top end – so an old “xhigh” config can quietly inflate spend.

The catch is API shape. Leave thinking off the request (or send adaptive). thinking: {"type": "disabled"} or a manual budget_tokens returns 400 invalid_request_error. Size max_tokens for thinking plus the reply – thinking still counts when display hides it; long agent turns in Anthropic testing used up to 128K (the model max). Changing top-level effort between requests invalidates the prompt cache; use per-message effort (beta) when one conversation needs mixed levels.

Long runs: stop rules that block fake “done”

Unattended loops are where Opus 5.5 goes quiet in a bad way. It sometimes ends a turn with a tidy progress summary and no tool call. A dumb use treats that text-only end_turn as finished while checklist items stay open. Put named stops in CLAUDE.md (or your system prompt):

When a step doesn't need my input, keep going.
Put status notes in the same message as your next action.
Stop and ask only when you can't continue without me, or before anything
destructive: deleting data, force-pushing, or changing anything outside this repo.

Keep a live checklist in a file (TASKS.md). If a turn ends with open items and no blocker, nudge once or twice – then stop and review so you don’t loop forever. For multi-service audits, fan work to subagents and require evidence before you accept each report.

Design bans and pasted text

“Avoid a generic AI look” mostly swaps one default skin for another. Name the habits you hate (the official guide calls this out):

Build a settings panel for an internal admin tool with placeholder copy.
Do not use a cream or off-white background, italic accent words in headings,
numbered "01 / 02 / 03" section labels, monospace labels, or pill-shaped buttons.

Check what it picked instead; extend the ban list. For user-pasted email or web text, wrap blocks so instructions inside them don’t become yours:

<pasted_content id="ab12">
...pasted text...
</pasted_content id="ab12">

Pair that with a short system note that pasted content may hold untrusted instructions – per the official guide – then measure how cautious the model gets on your real tasks.

Silent traps

  • “Think carefully” left in system prompts. Dead weight. Anthropic’s chat-product test: remove it, first token arrives sooner, quality held.
  • Demanding a dump of internal chain-of-thought in the visible reply. Can hit a reasoning_extraction refusal. Ask for a short choice rationale, or read summarized thinking blocks.
  • max_tokens sized like Opus 5 with thinking off. Replies look randomly cut off because thinking still burns the budget.
  • Agent UI silent between tools. Progress often lives in thinking blocks; default display can omit them. If your use only streams text blocks, turn display on for updates.
  • Status monologue read as completion. Checklist + stop rules beat vibes.
  • Opus 5 effort labels copied without a re-sweep. Same name ≠ same spend.

What holds up after the launch noise

I’m wary of launch tables. Anthropic’s posts claim fewer steps, fewer tokens, and clearer status writing on agentic coding and knowledge-work loads; community threads after week one mostly agree the prose is tighter than Opus 5’s heavier agent voice. Small independent retests of the prompting guide are mixed on secondary claims – treat those as signals.

What holds in my own runs: medium effort + whole-task framing finishes without babysitting more often; stripping think-boilerplate cuts the “still warming up” feel; specific design bans beat vague aesthetic coaching. If the job is already easy at medium, max effort is mostly a latency tax.

Is clearer writing the part people will remember six months out? After Opus 5’s denser agent voice, a lot of threads say Opus 5.5 “writes the way I do.” Subjective – but if your bottleneck is reviewing agent prose, that may matter as much as any bench point.

When not to use this model (or this style)

Don’t burn Opus 5.5 at high effort on bulk classification, short rewrites, or anything a cheaper Sonnet-class model already clears. Lower effort or a smaller model is enough.

Don’t force whole-task autonomy when you need a human checkpoint every step. Pair-programming wants the opposite CLAUDE.md: short plan first, recap at the end, stop often.

Don’t treat Opus 5.5 as a dump-CoT debugging surface. Safeguards around raw reasoning extraction are intentional. Related skills once this clicks: prompt caching for long agent loops, Claude Code subagent budgets, and a structured eval use so effort sweeps aren’t vibes. Separate rabbit holes.

FAQ

Do I need to rewrite all my Opus 5 prompts?

No. Strip think-hard lines, set effort explicitly, re-check max_tokens. That pass is enough for most stacks.

Should I start every task at max effort “just in case”?

Usually not. Picture a 40-file refactor that already passes at medium: max can multiply thinking tokens and wall time for a one-point bump – or none. Reserve xhigh/max for failure modes your eval set shows twice at medium. For chat latency, try low first.

Why did my API call return a 400 about thinking?

Opus 5.5 rejects thinking: disabled and manual budget_tokens. Omit the field (adaptive is default behavior) and control depth with effort. If old code assumed content[0] was always text, fix that – responses can open with thinking blocks even when visible text is empty under default display. Full notes: What’s new in Claude Opus 5.5 and the launch post.

Next action: Take one real task you ran last week on Opus 5. Strip every “think carefully / step by step / double-check” line, add a three-bullet “done means” block, pin effort to medium, run once. Compare tokens, interruptions, and whether you still wanted to babysit. That A/B beats another tip list.