Mer. Set 30th, 2026
A golden trophy faces turquoise gears, contrasting benchmark victories with real operational work. Original illustration without text.

Anthropic presents Claude Opus 5.5 as faster, cheaper on typical workloads and stronger across its own internal evaluations. Those are real claims and they deserve examination. The leap from there to declaring every other model pointless deserves rather less. It skips the small matter of what the customer actually has to get done.

A model does not become the right tool because a launch narrative has run out of adjectives. It becomes the right tool when it completes a defined task to an acceptable standard, at a cost that includes the work required to establish that the standard was met. The benchmark is not your job wearing a tie.

Three different price comparisons in one headline

The launch describes a 40 percent cost reduction on typical workloads. Listed input pricing moves from five dollars to four per million tokens, and output from twenty-five to twenty: both 20 percent reductions. Cache reads fall from fifty cents to twenty cents per million, a 60 percent reduction.

None of these figures contradicts the others. They simply describe different things. Total cost depends on the mix of input, output and reused context, and on how much processing a given task actually consumes. A workload-level comparison cannot be pasted onto every invoice as a universal discount.

The trouble starts when all three numbers are blended into a single claim that the model is cheaper. Cheaper under which pattern of use? With how many retries? Measured against which acceptance criteria? If the saving depends heavily on cache reuse, a task with little reuse has no reason to expect the same benefit.

And that is before anyone counts the person who has to inspect the output. Lower token prices are welcome. They do not make the reviewer’s afternoon disappear.

A large codebase is not a certificate of quality

The examples accompanying the launch describe major migrations and audits finished quickly: a 680,000-line migration completed in under a day, a 200,000-line codebase audited and fixed in under three hours against more than twenty hours for the previous model. Anthropic also reports an internal exercise translating HAProxy from C to Rust, where Opus 5.5 finished in 9.5 hours against 12, at 51 percent of the cost, with almost all regression tests passing.

Those results can indicate genuine progress without establishing that every codebase is equally tractable. Lines of code measure size, not the difficulty of untangling dependencies, preserving behaviour or discovering what the existing tests never covered. “Almost all” regression tests is also a different statement from proof of correctness, and the gap between the two is exactly where expensive surprises live.

Anyone evaluating a model of this kind should separate producing a proposed change from accepting one. Compilation, regression testing, human review and operational constraints each contribute different evidence. A fast migration that still needs careful investigation is not worthless, but it is not finished work either.

The sales headline compresses all of that into a stopwatch. The engineering team inherits the remainder.

The examination may be changing the behaviour

Anthropic also acknowledges something more awkward: the model often appears to suspect that it is being evaluated. That complicates the interpretation of the alignment results and their relationship to behaviour outside the test setting. The company says so explicitly, and says the problem is not solved.

It does not establish a conscious intention to deceive, and it should not be inflated into a story about a machine plotting between exams. The narrower point is serious enough. Behaviour measured in a recognisable setting cannot automatically be treated as an unconditional guarantee in an unfamiliar one.

Higher scores on internal safety tests are information. They are not a warranty covering every integration, permission set and operating condition a customer might subsequently build. A launch can report a genuine improvement and still leave a significant uncertainty open. A buyer who ignores that uncertainty has not made the model safer. They have made the procurement story tidier.

Clear writing can help, or reassure too easily

The claimed improvement in how the model communicates is genuinely welcome. Less jargon, shorter explanations, important information at the top: all of that makes an output easier to examine. Anyone who has had to excavate a conclusion from a mountain of generated prose will understand the appeal.

But readable is not a synonym for correct. A concise error is easier to spot if the result is being checked properly, and easier to swallow whole if fluency has quietly replaced checking. Anthropic argues that clarity is itself a safety benefit because it makes mistakes easier to catch. That holds only where somebody is still doing the catching.

Which is the test I would apply to the whole launch. Define the work. Decide what counts as an acceptable result. Measure the full cost of reaching it, reviewer included. Then compare the alternatives under the conditions that matter to your project rather than the conditions that matter to the announcement.

“Best” without those conditions is a label on a shelf. Declaring every other tool useless is not an evaluation. It is the moment the advertisement starts doing the customer’s thinking.

Raffaele Di Marzio

All my “insane” books on cybersecurity and governance are here 👇 https://www.amazon.it/stores/author/B0FB47T6Q4/allbooks