Inference & Economics 28 September 2026 8 min read 1,804 words

The default did the saving

Claude Opus 5.5 costs 40% less to run than Opus 5, measured at default settings. The default moved down a notch between the two models, and Anthropic's own migration notes say that at the same setting the newer model thinks more per turn, not less.

The argument

Opus 5.5's 40% cheaper is measured at default settings and the default dropped a notch, so the headline compares two different amounts of thinking rather than two efficiencies of the same amount.

Two sentences in the Claude Opus 5.5 announcement describe the same number and only one of them carries the condition. The opening paragraph says the model "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5." Further down, in the section on cost and speed: "Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads."

Three words in the second sentence matter, and they are not "40% less." They are "at default settings." Anthropic's own migration notes, on a different page, spell out what changed there: "The default effort is medium. A request that omits effort runs at medium; on Claude Opus 5 it ran at high." A few lines later, the same page lists another behaviour difference: "At the same [effort] setting the model tends to think more per turn than Claude Opus 5, most of all at xhigh and max."

Put those together and the headline stops being an efficiency claim. Opus 5.5's 40% is measured at default settings, and the default dropped a notch, so the comparison is between two different amounts of thinking rather than two efficiencies of the same amount. Hold the setting constant and the newer model thinks more, not less. That is not a scandal and nobody hid it; the announcement publishes its whole cost curve and the migration page states the default change in plain words. But the number that travels is "40% cheaper," and the thing that produced it is a configuration.

What the dial actually is

Claude models take a request field, output_config.effort, with five values: low, medium, high, xhigh, max. On Opus 5.5 the default is medium; on almost every other model it is high. Effort is not a token budget. The documentation is explicit: "Effort is a behavioral signal, not a strict token budget," and, on the page about steering thinking, "You don't set a thinking token budget." The only hard limit is max_tokens, which caps one request's total output, thinking and answer together, and which therefore bounds a single call rather than a whole agentic turn made of many calls.

What effort does instead is shift a threshold. Adaptive thinking is always on for Opus 5.5 and cannot be turned off: a request that sends thinking: {"type": "disabled"} returns a 400 error. So the model decides per request whether to think at all and how far, and effort sets how willing it should be. At low it "minimizes thinking. Skips thinking for simple tasks." At max it "thinks the most readily and at the greatest depth, with no constraint on thinking length." Two requests at the same effort level are not two requests of the same length, and there is no parameter that makes them so.

Effort also reaches past thinking. It applies to every output token, including tool calls, and the documentation says lower levels tend to "combine multiple operations into fewer tool calls," "make fewer tool calls," and "proceed directly to action without preamble." That is where the token savings show up in practice, and the release's own testimony reads exactly that way. Lovable reports the model "gathers context once, makes fewer and more complete edits," finishing in "a third to half fewer steps." Kiro reports "about 40% fewer calls and using half the tokens." Box AI reports "a third of the tokens Opus 5 did, and its answers were 40% less verbose without losing accuracy."

So "uses fewer tokens per task," which the announcement names as half of the 40% alongside the lower per-token prices, is a description of how the model behaves under a setting. It is not a property of the weights, and it is not stable across settings. A model that has been asked to think and talk less is a cheaper model. That is a real and defensible engineering result, and it is a different result from a model that does the same work for less.

The table, the default, and the testimony

Once effort is understood as a dial, a model stops having a score.

A footnote under the benchmark table says: "Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort." Terminal-Bench 4.0 is noted separately, reported at xhigh. The table is therefore measured at the top of the dial, and shipped behaviour sits in the middle of it.

The chart captions give both, and the gap is not uniform. On CursorBench 4.0 the table shows 57.8% at max; the caption says "At default effort (medium), Opus 5.5 scores 52.5%." That is 5.3 points bought by turning the dial up. On FrontierCode v1.1 the table shows 54.4% at max; the caption says "At default effort (medium), Opus 5.5 scores 54.6%." That is nothing, or slightly less than nothing. Anthropic publishes a standard error of ±2.6 points for its Terminal-Bench runs and ±3.5 to 5 points for Terminal-Bench-Science, so the honest reading of the FrontierCode pair is not that more thinking hurt. It is that the top of the dial bought no improvement anyone could detect.

One caveat belongs here. Every number above comes from one party: Anthropic ran the evaluations, chose the effort level each result is reported at, and selected the customers quoted. Terminal-Bench makes the point. Its public repository describes an execution harness and a beta core dataset at version 0.1.1 as the leaderboard set, and nothing about the 4.0 dataset this release reports against, so the only account of that instrument I have read is Anthropic's own footnote.

The testimony points the other way again. Deloitte reports that "even at its lowest effort setting, Claude Opus 5.5 caught 72% of known bugs in our code reviews to Opus 5's 56% at high effort," and that on consulting analysis "low thinking effort matched its higher thinking settings on half the output." Rogo reports the lowest setting beating Opus 5 at high effort on its finance benchmark "with about 60% fewer output tokens." Walleye Capital reports that at its lowest setting the model "largely solved our evaluation task."

Three sets of numbers, three different configurations, one model name. The benchmark table is measured at the top, the product ships at the middle, and several customers say the bottom is enough for their work. That is not inconsistency. It is what a five-point dial looks like when someone finally reports all five points, and it means test-time compute is not a scaling axis you can quote a slope for. It is an elasticity, and the elasticity is shaped differently on every task.

The obvious objection

Every serious system has a price and performance knob. Databases have instance sizes, CDNs have cache tiers, compilers have optimisation levels. Quoting your best configuration in the headline table is ordinary vendor behaviour, and Anthropic is doing considerably better than ordinary here by publishing the full accuracy-against-cost scatter with all five effort points plotted and the default labelled. Anyone who deploys a model without measuring it on their own workload has made a beginner's mistake that no amount of disclosure will fix.

All of that is fair, and the charts genuinely are more informative than the bar charts they replace. But three things make this knob different from an instance size.

It is not a resource you provision. Effort is soft guidance over a model's own decision, so two identical requests at medium can cost different amounts, and no setting makes the spend deterministic. Its default moved between versions, which means code that changed nothing except a model string changed its behaviour and its bill. And finding your own point on the curve is not free, because of how the dial interacts with the cache.

That last one is the part worth sitting with. The documentation says the resolved effort value is rendered into the prompt, so changing it between requests invalidates cache breakpoints, and the practical advice is to "pick a thinking configuration and an effort level per conversation and keep them." Meanwhile the announcement notes that cache reads "make up the majority of agentic and coding work costs." So the same page that tells you to "run an effort sweep on your own evals rather than carrying settings over from an earlier model" also tells you that varying effort inside a cached workload throws away the discount that made the workload affordable. Anthropic ships a partial fix, a beta per-message effort change that preserves the cache, gated behind a header. The general shape stands: calibration is a separate expense from production, and it is charged at full price.

What to take from this

Stop quoting a model's benchmark score without its effort setting. "Opus 5.5 scores 57.8% on CursorBench" is not a fact about a model; it is a fact about a model at max, and the model you will actually run scored 52.5%. The same discipline applies in reverse to cost: "40% less" needs "than what, at which setting, on whose workload" attached, and the answer here is Opus 5 at high versus Opus 5.5 at medium on Anthropic's typical workloads.

Then do the thing the vendor's own documentation asks for, and budget for it. Run the sweep on your evals, all five points, and record where your curve flattens, because on one benchmark that happened before medium and on another it had not happened by max. Treat the sweep as a line item, not a Friday afternoon. And read the next release's cost claim for whether it is like-for-like before you read it for the percentage.

There is one more habit worth breaking. The useful unit of comparison between models is no longer a number but a curve, and almost nobody publishes one. When a lab does, as here, the interesting information is rarely the peak. It is the shape: how early the curve flattens, and how differently it flattens per task. A model whose curve is flat from low upward on your work is worth more to you than one with a higher peak, and nothing in a leaderboard will tell you which you have.

Anthropic says, in the same release, that "at these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences," and that in its own use the gap to Fable 5.1 "is narrower than these scores suggest." Take the lab at its word. Discount the table, as invited. What is left standing is the cost axis, and the cost axis now runs through a default that a vendor sets, that changed last week without anyone's code changing, and that you cannot explore without paying for the privilege.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. Introducing Claude Opus 5.5 Anthropic · 2026-09-22
  2. What's new in Claude Opus 5.5 Anthropic · 2026-09-28
  3. Effort Anthropic · 2026-09-28
  4. Steering thinking Anthropic · 2026-09-28
  5. terminal-bench Laude Institute · 2026-09-28

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

effort levelsadaptive thinkinginference costdefaultsprompt caching