Opus 5.5 Effort Settings: Stop Burning Tokens on High

I’ll admit it. For a long time I treated every AI model like it had one speed: as fast and as hard as possible. Quick text cleanup or a gnarly bug, it didn’t matter. Everything got the same heavy treatment, and my results didn’t get any better for it.

So this post from an AI professional on LinkedIn hit home. The author points out something unusual about Opus 5.5: it got cheaper and stronger at the same time. That almost never happens in AI. Usually a better model costs you more.

The original poster’s warning is that most people will miss it. They’ll skim the release notes and move on, or run high effort on every single task. Simple extraction and hard debugging get the same setting, the bill climbs, and the results barely move.

The expert also admits they got this wrong for years with every tool. What finally drives results was never the biggest model. It was matching the effort to the task. They say they see this every week: the founders who get it early ship faster and spend less. Here’s how to put their one-page Opus 5.5 guide to work, step by step.

Step 1: Know what you’re working with

You can’t tune a tool you don’t understand. So start with the specs the creator pulled together:

  • 1M token context and up to 128K tokens of output
  • Adaptive thinking is always on, and effort is your dial
  • Input takes text and images, output is text
  • Knowledge cutoff is June 2026
  • Batch output goes up to 300K tokens in beta
  • The API model name is “claude-opus-5-5”

The big one here is that adaptive thinking is always on. You don’t flip reasoning on or off. You just decide how hard the model should work, and that’s the effort setting.

Step 2: Learn the price tags

Once you know the dials, look at what each one costs. According to the post:

  • Input: $4 per 1M tokens
  • Output: $20 per 1M tokens
  • Cache reads: $0.20
  • Batch jobs: 50% off
  • Fast mode: up to 2.5x quicker at $8 in and $40 out

Why bother with this step? Because the gap between a cache read and fresh input is huge. Fast mode is handy when speed really matters, but it costs twice as much. When you know these numbers, the cost of a lazy setup becomes obvious.

Step 3: Match the effort to the task

This is the heart of the guide. The author lays out three levels:

  • Low: cheaper
  • Medium: the default
  • High: for hard problems

Then they map them to real work. Use high effort for complex debugging, big migrations, and high-stakes work. Use lower effort for extraction, classification, and rewriting.

Say you’re pulling names and dates out of 500 support tickets. That’s extraction, so go low. Moving a legacy codebase to a new framework? That’s a big migration, so turn it up to high. Same model, very different bills.

Step 4: Point it at the right jobs

Next, know where Opus 5.5 really shines. The post lists its strongest areas as agentic coding, long-running agents, and deep research. It’s also great at computer use and visual understanding.

The original poster also covers what changed compared to Opus 5:

  • Better token efficiency
  • Faster output and clearer communication
  • Stronger prompt-injection resistance
  • Stronger coding and better long-running agents
  • Improved visual understanding

Better token efficiency plus lower prices means each task should cost less than before, as long as you aren’t wasting effort.

Step 5: Build every prompt in this order

This is the part the expert says they’d steal first, and I agree. Build your prompts in this order:

  1. Role
  2. Task
  3. Context
  4. Requirements
  5. Constraints
  6. Output
  7. Verification

The creator calls that last one “the sleeper.” Ask Claude to check its work before it finishes. It’s one extra line, and it catches mistakes before they ever reach you.

Here’s how that might look in a quick example: “You’re a senior data analyst. Summarize churn drivers from the attached survey. Context: we’re a B2B SaaS with 2,000 customers. Requirements: list the top 5 drivers with supporting quotes. Constraints: no guessing beyond the data. Output: a bulleted list. Verification: before you finish, check that every driver is backed by at least one quote.”

Step 6: Lock in the habits

A good prompt only gets you so far. Your daily workflow matters too. The post’s author shares four habits that get better results:

  • ✔ Set clear goals and define what success looks like
  • ✔ Add the relevant context and use tools when needed
  • ✔ Cache repeated contexts and batch high-volume tasks
  • ✔ Verify the output, then raise effort only when needed

That last habit flips the usual approach. Don’t start on high just in case. Start lower, check the output, and only turn it up when the results fall short. Caching and batching tie straight back to Step 2, since that’s where the biggest savings live.

Step 7: Keep the rule simple

After all that detail, the contributor wraps it up in one sentence:

Use Opus 5.5 for difficult, complex, professional work.

I love how clean that is. Not every task needs the most powerful setting. Save the heavy effort for the work that earns it, and let lower settings handle the routine stuff.

Why this matters

It’s tempting to think better results come from a bigger model or a pricier setting. This savvy professional’s point is that the real gains come from being intentional. Know your specs, know your costs, match effort to the task, structure your prompts, and always verify.

The author also suggests sending this to the one person on your team who’s still burning tokens on high. They want to hear where you land on effort settings, too. Check out the full LinkedIn post to see the original one-page breakdown and join the conversation.

Scroll to Top