Skip to main content

Opus 5 was meant to stay behind Fable 5. It beats it on 8 evaluations out of 13

Anthropic's official table has landed, and it contradicts the scenario we had defended since 15 July: an Opus 5 close to the flagship, without surpassing it. It surpasses it — decisively on agentic work, by a hair on coding, while staying behind on legal and health. A public audit of our four bets, a row-by-row reading, and a stop on the one measure that really describes our trade: 26% on business workflows.

5/5 Google rating · 130+ projects delivered (combined) · book 30 min — free

Analysis
By Victor
13 min read

Claude Opus 5 has shipped, and the official benchmark table is public. Since 15 July we had been defending a precise scenario, and we had written it down: "an Opus 5 close to Fable 5, without surpassing it" — because an Opus that topped the flagship three weeks after its launch would break the product line. It surpasses it. Across thirteen published evaluations, Opus 5 comes ahead of Fable 5 on eight, including end-to-end agentic work: ARC-AGI-3 from 1.5% to 30.2%, agentic terminal coding doubled. We were wrong, and that is more interesting than being right: this article audits our bets line by line, reads the full table, and stops on the one row nobody is commenting on — the row that directly concerns companies that automate.

In short — Claude Opus 5, released on 24 July 2026, is priced at $5 per million input tokens and $25 per million output tokens — exactly the Opus 4.8 rate, and roughly half Fable 5's cost per task. It scores 43.3% on agentic terminal coding (Frontier-Bench v0.1), 30.2% on ARC-AGI-3, 90.8% on BrowseComp and 26.0% on AutomationBench. It comes ahead of Claude Fable 5 on 8 of the 13 evaluations in Anthropic's official benchmark table, but stays behind GPT-5.6 Sol on DeepSWE v1.1 (68.8% vs 72.7%) and behind Mythos 5 on HealthBench Professional (59.8% vs 66.0%). Against its predecessor Opus 4.8, the largest gain is on autonomous tasks: the ARC-AGI-3 score rises from 1.5% to 30.2%.

Key takeaways

  • Opus 5 comes ahead of Fable 5 on 8 of the 13 published evaluations — while the consensus scenario, ours included, was an Opus "close to Fable, without surpassing it". Against Opus 4.8 the jump is stark: agentic terminal coding from 21.1% to 43.3%, novel problem-solving (ARC-AGI-3) from 1.5% to 30.2%, computer use from 55.7% to 70.6%.
  • Opus 5 does not win everywhere, and that is the honest reading. GPT-5.6 Sol still leads on DeepSWE v1.1 (72.7% vs 68.8%), Fable 5 edges it out on FrontierCode (53.5% vs 53.4%) and on legal work (13.3% vs 11.7%), and Mythos 5 dominates health (66.0% vs 59.8%).
  • The row nobody is commenting on: AutomationBench, 26.0%. That is the business-workflow test — precisely our field. Opus 5 crushes the pack there (17.4% for Fable 5, 18.1% for GPT-5.6 Sol, 17.0% for Opus 4.8), yet completes barely more than a quarter of the tasks.
  • What 26% means for you: the best model in the world, shipped on its own, fails three business workflows out of four. The value is not in the model, it is in the engineering around it — specification, data, guardrails, error recovery.
  • Prompt injection resistance: 0.2% on one attempt, 2.0% over fifteen (Gray Swan IPI benchmark) — against 3.1% / 20.0% for GPT-5.6 Sol and 14.2% / 49.2% for Gemini 3.1 Pro. It is the point Boris Cherny (Anthropic) puts ahead of every other score, and the most decisive one for pointing an agent at inbound documents.
  • The price does not move, the cost per task collapses. $5 / $25 per million tokens — exactly Opus 4.8's rate card. Yet Opus 5 reaches Fable 5's level at half the cost per task on CursorBench 3.2, and beats it on OSWorld 2.0 at a little over a third.
  • Three independent sources converge: maximum effort is not the right default. Dan Shipper (Every), the CursorBench results and CodeRabbit's code-review bench establish it by three different routes — more reasoning improves nothing uniformly, and costs more than double ($3.91 against $8.23 per task). Opus 5 also breaks compatibility with existing automations: it is not a drop-in replacement.
  • Audit of our four bets from 15 July: two won, one half-right, one lost. The lost one matters most — we bet on economic efficiency rather than raw capability. Details in the bet audit.
  • Every figure in this article comes from the official benchmark table published by Anthropic at the launch of Opus 5 (24/07/2026). We have recomputed nothing, extrapolated nothing, and we flag every row where a competitor wins.

Claude Opus 5 official benchmarks, evaluation by evaluation

Below is the full set of evaluations published by Anthropic at the launch of Opus 5, against the three comparison models used in the official table: Fable 5 (the flagship released three weeks earlier), Opus 4.8 (the direct predecessor) and GPT-5.6 Sol (the most-cited competitor). Cells where Opus 5 does not come first are flagged — they are more instructive than the rest.

Official Claude Opus 5 benchmarks compared with Fable 5, Opus 4.8 and GPT-5.6 Sol
Evaluation Opus 5 Fable 5 Opus 4.8 GPT-5.6 Sol
Agentic terminal coding
Frontier-Bench v0.1
43.3%33.7%21.1%34.4%
Knowledge work
GDPval-AA v2 (Elo score)
1,8611,7471,5931,736
Novel problem-solving
ARC-AGI-3
30.2%1.5%7.8%
Agentic search
BrowseComp
90.8%87.4%84.3%90.4%
Multidisciplinary reasoning
Humanity's Last Exam — no tools
56.3%56.5%49.8%
Multidisciplinary reasoning
Humanity's Last Exam — with tools
64.7%63.9%57.9%
Computer use
OSWorld 2.0
70.6%66.1%55.7%62.6%
Agentic coding
DeepSWE v1.1
68.8%69.7%59.0%72.7%
Agentic coding
FrontierCode v1.1, Main
53.4%53.5%46.5%47.5%
Business workflows
AutomationBench
26.0%17.4%17.0%18.1%
Legal
Legal Agent Benchmark, held-out
11.7%13.3%10.4%2.5%
Health
HealthBench Professional
59.8%66.0%
Mythos 5
57.4%60.5%
Biology — hard problems
BioMysteryBench
49.4%46.5%42.4%
Biology — human solved
BioMysteryBench
90.1%89.0%
Mythos 5
88.5%

Source: official Anthropic benchmark table, published at the launch of Claude Opus 5 (24/07/2026). The health and biology rows compare against Mythos 5 where the official table does so. A dash marks an evaluation not reported for that model.

How much does Claude Opus 5 cost?

$5 per million input tokens, $25 per million output tokens — precisely the Opus 4.8 rate. A fast mode, roughly 2.5× quicker, is billed at double. The model is available through the API under the identifier claude-opus-5, becomes the default model on the Claude Max subscription and the most capable one available on Claude Pro. Standard access carries no data-retention requirement (Clubic and BelieveMy, 24/07/2026).

So the sticker price is unchanged. The cost per task, however, collapses — and that is where the news actually is. The CursorBench results published by BridgeMind on 24 July put exact numbers on it, and they say more than any isolated score.

CursorBench results: score, cost and token consumption per task
Model Score Cost / task Tokens / task
Fable 5 Max70.5%$17.32103,525
Opus 5 Max70.0%$8.2361,838
Opus 5 Extra High69.3%$7.3554,239
Fable 5 Extra High68.4%$11.7364,971
GPT-5.6 Sol Max67.2%$5.6928,320
Opus 5 High66.7%$3.9127,932

Source: CursorBench results published by BridgeMind (@bridgemindai), 24/07/2026.

Three readings stand out. One: Opus 5 Max finishes half a point behind Fable 5 Max — 70.0% against 70.5% — at less than half the cost per task ($8.23 against $17.32) and roughly 40% fewer tokens. Two: drop one effort level and Opus 5 Extra High beats Fable 5 Extra High (69.3% against 68.4%) while costing $7.35 against $11.73. Three, the most interesting line for a tight budget: Opus 5 High scores 66.7% for $3.91 per task, half a point behind GPT-5.6 Sol Max which costs $5.69.

This is where the effort parameter stops being a technical detail: the same model family spans $3.91 to $8.23 per task for three to four points of score. On a workflow running thousands of times a month, that slider is the architecture decision. The same mechanism shows on OSWorld 2.0, where Opus 5 beats Fable 5's best score at a little over a third of the cost.

Three new capabilities that matter in production

  • A one-million-token context window, both default and maximum, with performance claimed to hold across the full length — a direct answer to the degradation past 200,000 tokens observed on Opus 4.8.
  • An adjustable effort parameter — trade quality against latency and cost explicitly, request by request, instead of paying the maximum on every call.
  • Automatic fallback routing (beta on the API): disputed requests reroute to another model rather than failing. This was the one credible leak of last week — the one matching the behaviour observed on "Honeycomb" in Cursor on 8 July. It is confirmed.
  • Self-verification: the model tests and validates its own outputs. On a business workflow this shifts part of quality control upstream — without removing it, as the 26% on AutomationBench makes clear.

On safety, Anthropic claims its most aligned model to date: the rate of misaligned behaviour is the lowest measured across its recent models (a score of 2.3 in the alignment audit), and safety classifiers trigger roughly 85% less often than on Fable 5 — a practical point for anyone who abandoned a deployment because of false positives.

Opus 5 vs Fable 5 vs GPT-5.6 Sol: where are the real gaps?

The lazy reading counts the winning cells and concludes "new best model". The useful reading looks at where the gaps are large, and where they are not. Sorted by the size of the jump between Opus 4.8 and Opus 5, the table tells a clear story.

Every jump is on the same side: autonomy. Novel problem-solving (ARC-AGI-3): from 1.5% to 30.2%. Agentic terminal coding (Frontier-Bench): from 21.1% to 43.3%, a doubling. Computer use (OSWorld 2.0): from 55.7% to 70.6%. Business workflows (AutomationBench): from 17.0% to 26.0%. These four evaluations share one property — they measure the ability to chain actions in a real environment without a human taking over at every step.

Every near-tie is on the same side too: "classic" coding and expert domains. On DeepSWE v1.1, GPT-5.6 Sol stays ahead (72.7% vs 68.8%). On FrontierCode v1.1, Fable 5 edges out Opus 5 by a tenth of a point (53.5% vs 53.4%) — not a real gap, but enough to forbid writing that Opus 5 "dominates coding". On legal work, Fable 5 does better (13.3% vs 11.7%); on health, Mythos 5 dominates (66.0% vs 59.8%), and even GPT-5.6 Sol comes ahead of Opus 5 (60.5%).

This relative plateau on pure coding deserves to be stated plainly, because it contradicts the prevailing narrative: across Fable 5, Opus 5 and GPT-5.6 Sol, the spread on classic coding evaluations is about five points. Picking a model will not determine the quality of a custom software development project — architecture, tests and domain knowledge will continue to do that.

Our reading, stated as an opinion

Opus 5 is not "the smartest model", it is the model that runs on its own for longest. The difference is concrete: on a task you can describe in a single prompt, the gap with the previous generation will be modest. On a task requiring ten tool calls, an error recovery and an intermediate decision, it will be dramatic. That is exactly the usage profile of an agent running in production, and not at all that of a writing assistant.

Auditing our four bets — including the one we lost

On 15 July we published four signed, dated bets on Opus 5, promising they could be held against us. The model has shipped: here is the outcome, including where it is unflattering.

Bet 1 — "Opus 5 will ship before the end of July": won

Written on 15 July, when prediction markets gave it roughly even odds and no primary source existed. The model shipped on 23 July. Nothing heroic: it was the most probable inference from Fable 5's commercial calendar.

Bet 2 — "it will be cheaper than Opus 4.8": lost on the letter, won on the substance

Lost in the strict sense: the rate card is identical — $5 / $25 per million tokens, exactly Opus 4.8's pricing. We bet on a lower sticker price; that did not happen.

But the economic intuition was right, and that is what makes this bet worth re-reading. We wrote that "the battle is fought on cost per task, not on benchmark records". That is precisely what Anthropic shipped: at an unchanged sticker price, Opus 5 reaches Fable 5's level at half the cost per task on CursorBench 3.2, and beats it on OSWorld 2.0 at a little over a third. The price cut did happen — not on the label, on the bill.

Bet 3 — "a fourth extension for Fable 5": half right

Anthropic did not extend free access: on 18 July it announced something else — Fable 5 stays in Max and Team Premium, but at reduced limits, while Pro plans lose included access. We had the direction (Anthropic is not cutting Fable 5), not the shape.

Bet 4 — "GPT-6 has nothing to do with this": won

The official table compares Opus 5 with GPT-5.6 Sol. No GPT-6 anywhere, one week after half the AI press invoked it as the driving force.

The bet we lost — "close to Fable 5, without surpassing it"

That was our core thesis, held for a week: Opus 5 would reach the flagship's level without overtaking it, because an Opus topping Fable 5 three weeks after its launch would break the product line. The reasoning rested on a solid fact — Dario Amodei had himself acknowledged that Opus lagged Mythos on benchmarks — and we drew a pricing conclusion from it: "this will not be a feat of engineering, it will be a price cut". The official table contradicts both. Opus 5 comes ahead of Fable 5 on eight of the thirteen published evaluations, and multiplies its novel problem-solving score twentyfold. What we missed: we were reasoning about product-line coherence (where to place Opus relative to Fable), while Anthropic was reasoning about autonomy (how many steps the model holds on its own). On that axis, there was no reason to hold back.

How good is Opus 5 at business workflows? 26% on AutomationBench

This is the row nobody will discuss, and the only one that describes our trade. AutomationBench measures business workflows: chains of enterprise tasks, not laboratory puzzles. Opus 5 scores 26.0% there, against 17.4% for Fable 5, 18.1% for GPT-5.6 Sol and 17.0% for Opus 4.8.

Both halves of that number deserve a sentence.

The good half: the gap to the pack is the widest in the entire table. Where other evaluations separate models by a few points, this one separates them by half. If you automate processes, model choice has just stopped being neutral.

The bad half, and it matters more: 26% is roughly one task in four. The best model in the world, on the evaluation that most resembles a real enterprise process, fails three business workflows out of four. Every commercial promise along the lines of "AI automates your company" runs into this number, published by the vendor itself.

Where we stand

That 26% is the strongest argument against the idea that you just need to "plug in a good model". A workflow that reaches production is not decided by the 26% the model handles alone: it is decided by the remaining 74% — process specification, input data quality, guardrails, error recovery and human checkpoints. That is engineering, not model selection — and it is precisely what separates a pilot that impresses from an automation that runs on Monday morning. It is also the first question to ask any AI agency that approaches you: what do you do about the other 74%?

The figure that really decides a production rollout: prompt injection

This is the point Boris Cherny, creator of Claude Code at Anthropic, puts ahead of every evaluation score — and he is right, because it is the one that decides whether you can point an agent at your real data. In his words, Opus 5 is the hardest model Anthropic has produced to hijack through prompt injection, a result he notes is somewhat buried in the system card.

What are we talking about? Indirect prompt injection is the attack where content your agent reads — an email, a supplier PDF, a web page, a product sheet — carries hidden instructions that hijack its behaviour. It is the number-one risk for any agent processing inbound documents, and the reason many automation projects never leave the pilot stage.

The Gray Swan IPI benchmark measures exactly that: the probability an attack succeeds within k attempts. Lower is better.

Indirect prompt injection robustness — Gray Swan IPI benchmark, probability an attack succeeds within k attempts
Model 1 attempt 10 attempts 15 attempts
Opus 50.2%1.6%2.0%
Mythos 50.3%2.1%2.6%
Fable 50.4%2.3%2.8%
Opus 4.80.5%4.1%5.5%
Sonnet 50.6%4.7%5.9%
Muse Spark2.9%14.3%16.5%
GPT-5.53.0%17.4%20.8%
GPT-5.6 Sol3.1%16.3%20.0%
GPT-5.6 Terra5.4%26.0%30.4%
Gemini 3.6 Flash7.3%32.2%37.3%
GPT-5.6 Luna8.3%38.6%43.9%
Grok 4.513.4%54.2%60.8%
Gemini 3.5 Flash14.1%54.2%60.5%
Gemini 3.1 Pro14.2%45.7%49.2%

Source: Gray Swan IPI benchmark, chart published by Boris Cherny (Anthropic), 24/07/2026. Read it as: the lower the percentage, the more resistant the model.

The gap here is of a different order than on any other evaluation. On a single attempt, Opus 5 drops to 0.2% against 3.1% for GPT-5.6 Sol and 14.2% for Gemini 3.1 Pro. More importantly, it holds up over time: after fifteen attempts the success probability stays at 2.0%, where it climbs to 20% for GPT-5.6 Sol and past 60% for Grok 4.5 and Gemini 3.5 Flash. Cherny adds that layering defences — model alignment, injection probes and automatic mode — drives the attack success rate to roughly zero.

The result drew praise beyond Anthropic: Elon Musk replied "very impressive" to Cherny's post within minutes. The remark is worth noting, because his own company's model, Grok 4.5, sits at the bottom of this ranking — a 13.4% attack success rate on the first attempt, 60.8% after fifteen. When a direct competitor publicly endorses a figure that works against them, the figure is usually solid.

Why we consider this the most important line of the launch

An agent reading your emails, supplier invoices or tender documents is exposed to this risk on every single document. Until now, the honest answer to "can we let an agent act alone on inbound documents?" was: not without tight supervision. Going from 5.5% to 2.0% over fifteen attempts does not make the risk zero — but it changes the nature of the conversation, and moves the line of what can reasonably ship to production. It also makes traceability of automated decisions more necessary, not less: a safer system is still a system to log, not least under Article 50 of the European AI Act.

What early users are saying: "a hard model to love"

It would be dishonest to relay only the enthusiasm. The two most serious independent reviews on launch day are critical — and far more useful than the benchmarks for anyone deciding on a migration.

Claire Vo ran Opus 5 through her own benchmark: seven tasks, seven competing models including GPT-5.6 Sol, Sonnet 5 and Gemini 3.1 Pro, blind scoring. Her verdict is in her title — brilliant, but annoying. Opus 5 earns top marks on at least one use case while proving verbose and "neurotic", to the point of refusing to touch a merge conflict during a real coding session, out of excessive caution.

Dan Shipper (Every) went further: a week of testing across coding, writing, knowledge work and an internal agent. His starting assessment is harsh — the model argues with instructions, stops before the work is finished, and does not play well with existing tooling and automations. Then his team did something counter-intuitive: they deleted everything and started from scratch. Without the elaborate workflows built for earlier generations, Opus 5 became "dramatically better".

Three concrete warnings before migrating

  • It breaks backward compatibility. Wired into automations designed for an earlier model, Opus 5 often stops early or misses part of the instructions. This is not a drop-in replacement.
  • Starting from scratch yields far better results than adapting what exists — an observation from Kieran Klaassen on the same team. This is a model to rebuild around, with a payoff for whoever accepts that cost.
  • Medium or low reasoning effort works better than maximum: the more thinking time it gets, the more the annoying behaviours surface. Shipper is emphatic — do not switch to a smaller model for speed; try Opus 5 on low thinking first.

Shipper closes with a barbed image: Opus 5 has "the personality of the genius" without its top end, leaving it in an uncomfortable middle ground between the elite model and the fast generalist.

The most rigorous bench: CodeRabbit

A third source, and the most methodical: CodeRabbit ran Opus 5 through its code-review bench — around a hundred error patterns drawn from real open-source pull requests, three runs per configuration, compared against their production model mix. They have benched every Opus release since version 4, which gives rare perspective.

Their verdict is nuanced and worth quoting as-is: a precision specialist, to be paired with a recall model. The numbers:

CodeRabbit code-review bench: Opus 5 x-high against the production mix
Measure Opus 5 (max effort) Production baseline
Actionable comment precision39.3%35.2%
Known bugs caught55.2%61.1%
Nitpicks (noise)9223
Full-stream precision28.6%32.8%

Source: CodeRabbit, "Opus 5 model review", 24/07/2026 — ~100 error patterns from real pull requests, averages across three runs.

In other words: Opus 5 says fewer wrong things when it speaks, but it misses more bugs and produces four times as many minor remarks to triage. CodeRabbit draws an architectural conclusion from this, not a ranking: frontier models have split into two camps — recall-oriented ones that find a lot and lean on filtering (GPT-5.6 Sol catches 69.7% of bugs, but only 31.6% of its comments are worth keeping), and precision-oriented ones that say less and are right more often. Opus 5 anchors the second camp.

Two figures matter for a budget. First, Opus 5 reads about 50% more and writes about 65% more than the baseline models for the same job (≈60,500 input and 9,500 output tokens per call, against ≈40,500 and 5,800): part of the real premium comes from volume, not from the rate card. Second, the context window moves to one million tokens by default, with performance claimed to hold throughout — a direct answer to the degradation past 200,000 tokens observed on Opus 4.8.

By category, CodeRabbit finds it solid on configuration errors and code quality, weaker on logic errors, race conditions and API misuse. Practical translation: good as an additional perspective, insufficient as the only safety net on critical code.

What nobody is connecting: three independent signals converge

Three sources that did not coordinate say the same thing about effort settings. Shipper: "low effort works better". CursorBench: Opus 5 on high reasoning scores 66.7% at $3.91 per task, against 70.0% at $8.23 on maximum — three points of score for more than double the cost. CodeRabbit: "more reasoning did not consistently produce a better review", with maximum effort buying precision at the cost of coverage and nothing improving uniformly — they call effort a routing decision, not a quality slider.

The practical conclusion is counter-intuitive and quantifiable: the default setting for a serious deployment is not the maximum, it is the minimum that passes your business tests. And the right use of Opus 5 is not "the single model" but a specialised lane inside a routed ensemble — exactly the interchangeable-intelligence-layer architecture we advocate.

These criticisms match our own reading from the field: a more autonomous, more cautious model is not automatically a more useful one in production. An agent that refuses to act when in doubt beats an agent that breaks something — but every refusal is an exception to handle by hand, and therefore one more line in the 74% discussed above.

What Opus 5 actually changes for a mid-sized company

Three operational consequences, without extrapolation.

1. Autonomous agent projects become defensible — some of them, not all. An agent that must chain steps across your tools benefits directly from the jumps on OSWorld 2.0 and Frontier-Bench. A project shelved six months ago for "too many steps, too many manual recoveries" deserves a second look. A project shelved because the process itself was never written down remains exactly as bad.

2. Model choice becomes an architecture decision, not a preference. The table shows domain-by-domain inversions: legal goes to Fable 5, health to Mythos 5, pure coding to GPT-5.6 Sol on DeepSWE. A serious architecture routes each task to the model that fits, instead of locking in a single vendor — which is what we do in our AI automation projects, where the intelligence layer is explicitly interchangeable. It is also what makes a switch to an open-weight model practical the day it becomes relevant — Kimi K3 being the current case in point.

3. None of this changes your obligations. The European AI Act deadlines stay on the same dates: transparency (Article 50) on 2 August 2026, high-risk Annex III systems on 2 December 2027. A more autonomous model mechanically means more decisions taken without a human in the loop — so more traceability to plan for, not less.

Beyond that, the method does not move: map the tasks, quantify the return on investment, start with the process whose failure costs least, and put nothing into production without a rollback plan. The 26% on AutomationBench is the best reason to stick to it.

FAQ — Claude Opus 5

Is Claude Opus 5 better than GPT-5.6 Sol?

On most evaluations in the official table, yes — clearly on agentic terminal coding (43.3% vs 34.4%), novel problem-solving (30.2% vs 7.8%), computer use (70.6% vs 62.6%) and business workflows (26.0% vs 18.1%). But GPT-5.6 Sol remains ahead on DeepSWE v1.1 (72.7% vs 68.8%) and on HealthBench Professional (60.5% vs 59.8%). The answer therefore depends on the task, not on a general ranking.

How much does Claude Opus 5 cost?

Claude Opus 5 is priced at $5 per million input tokens and $25 per million output tokens — exactly the Opus 4.8 rate card. A fast mode, roughly 2.5× quicker, is billed at double. It is also the default model on the Claude Max subscription and the most capable one available on Claude Pro. So the sticker price did not drop, but the cost per task did: on CursorBench 3.2 Opus 5 reaches Fable 5's level at half the cost per task, and on OSWorld 2.0 it beats it at a little over a third (sources: Clubic and BelieveMy, 24/07/2026).

What is the biggest gain of Opus 5 over Opus 4.8?

In relative terms, novel problem-solving: ARC-AGI-3 goes from 1.5% to 30.2%, twenty times the predecessor's score. In absolute terms and for professional use, agentic terminal coding (from 21.1% to 43.3%) and computer use (from 55.7% to 70.6%) matter more, because they determine whether an agent can work on its own across several steps.

What does AutomationBench measure, and why does 26% matter?

AutomationBench evaluates business workflows, meaning chains of enterprise tasks rather than academic exercises. Opus 5 scores 26.0%, well above Fable 5 (17.4%), GPT-5.6 Sol (18.1%) and Opus 4.8 (17.0%) — but that still means roughly three tasks in four fail. It is the quantified proof that an automation project is not reducible to model choice: most of the work lies in process specification, data, guardrails and error recovery.

Is Claude Opus 5 resistant to prompt injection?

It is the most resistant model in the measured panel. On the Gray Swan IPI benchmark, the probability that an indirect prompt injection attack succeeds is 0.2% on a single attempt and 2.0% over fifteen, against 3.1% and 20.0% for GPT-5.6 Sol, and 14.2% and 49.2% for Gemini 3.1 Pro. Boris Cherny (Anthropic) states that layering defences — model alignment, injection probes, automatic mode — drives the success rate to roughly zero. The risk is therefore not nil, but it becomes manageable for an agent processing inbound documents.

Should you migrate your automations to Opus 5?

Not as a drop-in replacement. Early users report that Opus 5 breaks compatibility with automations designed for previous models: it often stops early or misses part of the instructions. Rebuilding the workflow from scratch yields far better results than adapting what exists. A second useful counter-intuition: medium or low reasoning effort produces better results than maximum effort — and it is also the cheapest mode ($3.91 against $8.23 per task on CursorBench). For single-pass tasks (classifying an email, extracting three fields), the gain stays modest. For an agent chaining several tool calls with error recovery, the margins are significant. In every case, replay your own business test set before switching: a public benchmark does not replace a measurement on your data.

Which evaluations does Opus 5 not lead?

Four rows of the official table: DeepSWE v1.1, led by GPT-5.6 Sol (72.7% vs 68.8%); FrontierCode v1.1, where Fable 5 edges ahead by a tenth of a point (53.5% vs 53.4%); the Legal Agent Benchmark, where Fable 5 does better (13.3% vs 11.7%); and HealthBench Professional, dominated by Mythos 5 (66.0% vs 59.8%). On Humanity's Last Exam without tools, Fable 5 is also ahead by two tenths (56.5% vs 56.3%).

Sources

  • Anthropic — official benchmark table published at the launch of Claude Opus 5, 24/07/2026 (Frontier-Bench v0.1, GDPval-AA v2, ARC-AGI-3, BrowseComp, Humanity's Last Exam, OSWorld 2.0, DeepSWE v1.1, FrontierCode v1.1, AutomationBench, Legal Agent Benchmark, HealthBench Professional, BioMysteryBench).
  • Pricing, availability and capabilities (effort parameter, fallback routing, self-verification, alignment): Clubic, "Anthropic lance Claude Opus 5", 24/07/2026; BelieveMy, "Claude Opus 5", 24/07/2026.
  • Indirect prompt injection robustness: Gray Swan IPI benchmark, chart and commentary published by Boris Cherny (creator of Claude Code, Anthropic), 24/07/2026.
  • Independent review: Claire Vo, "Claude Opus 5 review: this model is brilliant (but annoying)", Lenny's Newsletter, 24/07/2026 — 7-task benchmark across 7 models, blind scoring.
  • Cost per task: CursorBench results published by BridgeMind (@bridgemindai), 24/07/2026.
  • Critical review after a week of testing: Dan Shipper (Every), "Day 0 vibe check" published on X, 24/07/2026, including Kieran Klaassen's observations on rebuilding workflows and reasoning effort.
  • Code-review bench: CodeRabbit, "Opus 5 model review", 24/07/2026 — ~100 error patterns from real open-source pull requests, three runs per configuration.
  • Our signed bets of 15/07/2026 and their interim audit of 21/07/2026, published on this blog before the model shipped.
  • European AI Act — application timeline: Article 50 (transparency) on 02/08/2026, Annex III (high risk) on 02/12/2027.
Victor Gless-Krumhorn

Victor Gless-Krumhorn

Founder & AI Consultant — JAIKIN

AI implementation and automation expert for SMBs and mid-market companies. Works with businesses across France, Germany and Switzerland, from process mapping to production rollout.

1 à 3 tâches automatisables identifiées — diagnostic gratuit

30 minutes avec un expert, un plan d'action écrit sous 24 h.

Free quote within 24h