Claude Opus 5's Hidden Dial Can Cut Your AI Bill by 5X
Claude Opus 5 ships with an adjustable effort setting, five levels from low to max, that changes how many tokens a task burns without changing the per-token price. Set it on purpose instead of leaving it on the default and routine AI work can cost roughly five times less, while contracts, financial reviews, and other high-stakes tasks still get full power.
Your AI is not expensive. Your defaults are.
It is a Tuesday afternoon in Manchester. Priya runs a nine-person marketing studio out of a converted warehouse near Ancoats. A vendor just emailed over a new software contract, twelve pages, the kind of thing that used to sit in her inbox for a week because she is not a lawyer and does not want to pay one 400 pounds to read a renewal clause. She pastes it into the AI tool her team uses for client reports and asks it to flag anything that could bite her. Ninety seconds later she has a ranked list of the three clauses to push back on, in plain English, with the reasoning shown.
Here is the part nobody told her. That same tool, on the same subscription, has been running her routine work, the FAQ replies, the meeting summaries, the first-draft social captions, at the exact same intensity it just used to read a legal contract. She has been paying full price for thinking she never asked for.
Why does the same AI cost five times more some days?
Most people assume an AI tool has one setting: on. You send a prompt, it answers, you get billed for however many tokens that took. What almost nobody outside a developer team knows is that the newest generation of Claude models ships with a dial for exactly this, and most businesses are not touching it.
Anthropic's Claude Opus 5, released on 24 July 2026, introduced what it calls an effort setting. Five levels: low, medium, high, xhigh, max. The setting does not change the price per token. It changes how many tokens the model spends reasoning before it answers. A model left on high, the default, will think longer and write more than the same model told to run on low, even for a task that did not need the extra thinking.
Independent testing on a real business workload backs this up with numbers, not marketing. A production intake pipeline handling ticket triage and contract clause analysis was benchmarked across all five effort levels on the same 500-item evaluation set. Low effort processed the batch for roughly 4.50 dollars per thousand requests. The default, high effort, cost roughly 22.25 dollars for the identical batch, and accuracy on the routine triage work did not measurably change. That is not a rounding error. That is five times the bill for zero extra value on the tasks that did not need it.
What is Claude's effort setting, exactly?
Effort is a request-level control on the Claude API and inside tools built on Claude, including Claude Code and Claude Cowork. It sits alongside the model choice itself. Anthropic's own launch documentation is direct about the purpose: the effort setting lets a business optimise for intelligence when a task demands it, or conserve tokens for faster, cheaper results when it does not. It defaults to high unless someone sets it deliberately.
The five levels, in practice:
- Low. Minimal reasoning, fewer tool calls, tightly scoped output. Built for high-volume, low-stakes work: sorting emails, tagging tickets, drafting a quick reply.
- Medium. A balance of speed and depth. Good for a first-draft report, a standard client update, a routine summary.
- High (the default). More reasoning, better for moderately complex work, but the setting most people never touch, which means most people are paying for it on everything.
- Xhigh. Long-horizon, multi-step work. Anthropic's own guidance recommends this as the starting point for genuinely hard coding and agentic tasks.
- Max. Unconstrained. Reserved for the handful of tasks a quarter where getting it wrong is expensive.
Claude Opus 5 also shipped a second lever, Fast mode, which runs the same model at up to 2.5 times the output speed for exactly double the price. It is a separate decision from effort: effort changes how much the model thinks, Fast mode changes how quickly it delivers what it already decided to say. Most day-to-day business work does not need Fast mode. A live customer chat where someone is sitting there waiting might.
Isn't this just a technical setting for developers?
It looks that way, which is exactly why almost no small business is using it. But the decision underneath it is not technical at all. It is the same decision Priya makes every day without an AI in the loop: how much attention does this task actually deserve?
She does not proofread a one-line Slack message with the same care she gives a client contract. She does not want her AI tool to either, especially when the difference shows up on an invoice at the end of the month. If you run any AI-powered tool for your business, whether you built it yourself, bought it from a vendor, or a Notma sprint set it up for you, the effort question is one sentence long: is this route thinking as hard as the job needs, or as hard as the default happens to be set?
You do not need to write code to act on the answer. In a chat interface, saying "keep this quick" or "take your time and check this carefully" nudges the same lever in plain language. If you have a developer or an agency running your AI workflows, the question to ask them directly is simpler still: which effort level is this route running on, and did anyone choose it on purpose.
How to put an effort dial on your own AI work
You do not need a rebuild. You need an audit, a rule, and a test. Here is the version Priya's studio actually ran, over one afternoon.
- List every task your AI touches. Not the tools, the tasks. Client FAQ replies. Meeting notes. First-draft captions. Contract review before signing. Financial summaries before a stakeholder call. Ten minutes with a notepad is enough.
- Sort each task by what a wrong answer actually costs you. A mistagged support ticket costs a re-tag. A missed liability clause in a vendor contract can cost real money. Be honest about which is which, because most tasks are the first kind.
- Assign an effort level to each bucket. Routine, repeatable, low-stakes work gets low or medium. Anything you would hand to your most careful staff member gets high, xhigh, or max.
- Set it explicitly, per route, not per request. Changing effort mid-conversation resets the model's cache and can cost more than it saves, so the level belongs on the workflow, not typed in fresh each time.
- Test before you trust it. Run the same real task at two effort levels and compare the output side by side. If the cheaper version holds up, keep it there. If it does not, that is exactly the task that earns the higher setting.
| Effort level | What it is built for | A business example |
|---|---|---|
| Low | High-volume, low-stakes, repeatable | Sorting inbound emails, tagging support tickets, drafting a routine reply |
| Medium | Balanced, moderate complexity | A first-draft weekly report, a standard client update |
| High (default) | Moderately complex reasoning | Whatever nobody has reviewed yet, which is most default setups |
| Xhigh | Long-horizon, multi-step, genuinely hard | Rebuilding a pricing model, a multi-document research brief |
| Max | Unconstrained, highest stakes | Reviewing a vendor contract before signing, a quarterly financial deep-dive |

What this looked like on Priya's actual contract
The prompt Priya used, turned up to high effort, was not complicated: flag anything unusually one-sided or that creates open-ended cost, rank the issues by what they could actually cost her, and name the two or three she should push back on first if she only gets one shot to negotiate. No legal jargon required on her end, and no rewriting yet, just where the risk sits before she decides what to fight for. The same approach works on a quarterly report before a stakeholder update: walk through whether costs are growing faster than revenue, flag anywhere the numbers do not match the story, and hand over the three tough questions a skeptical investor would ask.
Both of those are high-effort, high-stakes prompts, and they are worth every extra token. The point is not to run everything cheap. The point is that Priya's studio was running its Tuesday-morning inbox triage at the same intensity as that contract review, every single day, for months, without knowing there was a difference to make.
What you have after this
A week in, you have a task list sorted by stakes instead of guesswork, and a first cheap route running in production. A month in, you have real before-and-after numbers instead of a hunch, because you tested rather than assumed. A quarter in, your AI spend looks like a set of deliberate decisions instead of a bill that arrived and surprised you.
Same tool. Same subscription. A completely different bill.
What is Claude Opus 5's effort parameter?
It is an adjustable setting, low, medium, high, xhigh, or max, that controls how many tokens Claude Opus 5 spends reasoning through a task. It does not change the price per token, only how much thinking, and therefore how many tokens, a given request uses. It defaults to high unless set otherwise.
Does Claude Opus 5 cost more than Opus 4.8?
No. Anthropic priced Opus 5 at 5 dollars per million input tokens and 25 dollars per million output tokens, identical to Opus 4.8, while reporting meaningfully higher performance. The effort setting is what actually changes your bill, since it determines how many tokens a task consumes to reach a given answer.
What is Claude's Fast mode and is it worth paying for?
Fast mode runs Opus 5 at up to 2.5 times its normal output speed for exactly double the standard price. It helps when a person is waiting on a response in real time. It does nothing for background jobs, batch processing, or anything that is not latency-sensitive, so most routine business automation does not need it.
Can a small business use effort settings without hiring a developer?
Yes, in two ways. Inside a chat interface, phrases like keep this quick or take your time and check this carefully nudge the same behavior in plain language. If a vendor or developer runs your AI workflows, ask them directly which effort level each route uses and whether anyone chose it on purpose, since most routes are simply left on the default.
How do I know if my AI tool is running at full effort on everything?
Ask the vendor or the developer who built it. If nobody can answer, the workflow is almost certainly running on the platform default, which is the highest setting most providers ship with. That is the single most common reason a bill looks larger than the actual work being done would suggest.
Find your first high-payback workflow.
See the SprintSources
- Anthropic: Introducing Claude Opus 5
- FoundrySoft: Claude Opus 5 Effort Parameter Guide (real cost numbers, per-effort-level benchmark)
- Hashnode: Claude Opus 5 effort levels and real cost
- Limitless Podcast: Not Using Opus 5 Will Cost You (A Lot), YouTube
- Sohfi Hamid: How Claude Opus 5 Can Help Your Business, YouTube
- YouLasy: Claude Fable 5 vs Opus 5, effort dial by job type, YouTube
Find your first high-payback workflow.
Book a free conversation or start with the fixed-fee Sprint.
Keep reading
Stop Paying to Turn Training Videos Into SOPs. Claude Already Does It.
You do not need to pay for Trupeer, Clueso, or Whale to turn a training video into a searchable SOP. If your business already uses Claude, a free,…
AI Tools & SkillsHubSpot Wanted $80,000 a Year. He Built an AI Agent Instead.
A startup founder turned down an $80,000-a-year HubSpot quote for a ten-person sales team and spent an afternoon building his own CRM instead, then open…
Front-office AIStop Losing Customers to Voicemail. Build an AI Voice Agent That Answers Every Call.
Connect a no-code voice AI platform, such as Bland AI or Vapi, to your existing business phone number, feed it a short written knowledge base of your hours,