Hey there! 👋
Welcome back to SavvyMonk, your one-stop for AI and tech news that actually matters. Anthropic launched a new Opus model that it says comes close to matching its own flagship on the hardest coding and reasoning tests, while costing half as much to run.
Let's get into it.
TODAY'S DEEP DIVE
Opus 5 Beats Every Model Except Fable 5 on Anthropic's Benchmarks
On 24 July 2026, Anthropic released Claude Opus 5, positioning it as the strongest model in the Opus line and the new default on Claude Max, with Opus 4.8 stepping down after roughly a year at the top of that tier. The company frames the release around a single trade. Opus 5 gives up a small amount of raw intelligence compared with the pricier Claude Fable 5, and in exchange it costs half as much to run.
The Price And The Pitch
Opus 5 is priced at five dollars per million input tokens and twenty five dollars per million output tokens, unchanged from Opus 4.8. Anthropic's pitch is that the price stayed flat while the model underneath it improved substantially, so the same budget now buys meaningfully more capability than it did with the previous generation. A faster variant, running at roughly two and a half times the default speed, is priced at double the base rate.
On Frontier-Bench v0.1, an internal coding evaluation, Opus 5 more than doubled Opus 4.8's score at a lower cost per task.
On CursorBench 3.2, at its highest effort setting, it landed within half a percentage point of Fable 5's best result while costing half as much per task. Cursor co-founder Sualeh Asif described the model as delivering near Fable 5 intelligence at Opus speed and cost, adding that on CursorBench it sits slightly under Fable 5 with many of the same behaviours.
Where It Beats Everyone Else
Anthropic's own testing puts Opus 5 ahead of every other model, Fable 5 included, on several fronts. On ARC-AGI 3, a test built around novel problems the model has not seen before, Opus 5 scored three times as high as the next best model.
On Zapier's AutomationBench, which measures whether a model can complete a full business task end to end, Opus 5's pass rate ran around one and a half times the next best model at the same cost, and Zapier chief executive Wade Foster said the model took a raw account health workbook and ran a complete churn prevention sequence without a single failed attempt, something no earlier model managed.
On OSWorld 2.0, a computer use benchmark, Opus 5 beat Fable 5's best result while costing a little over a third as much.

These are Anthropic's own internal benchmark results rather than independently audited scores, so they are worth reading as the company's reported figures rather than settled fact. Even so, the pattern across coding, reasoning and business automation tests points the same direction, a smaller gap to the top model at a noticeably lower price.
The One Model It Still Trails
The exception Anthropic is upfront about is Fable 5 itself, which remains ahead on raw coding and knowledge work performance, and Claude Mythos 5, which stays out in front on cybersecurity tasks specifically.
Anthropic says Opus 5 comes close to Mythos 5 at finding software vulnerabilities but remains well behind it at turning those vulnerabilities into working exploits, a gap the company presents as intentional rather than accidental.
What Early Users Are Seeing
Feedback from early access partners lines up with the benchmark story. Devin maker Scott Wu said Opus 5 approaches Fable level performance at half the cost and shows particular strength on difficult debugging work. Lovable co-founder Fabian Hedin said the model came out ahead of every prior model in its family on their internal evaluations, up twenty two per cent over Opus 4.7, while also showing far less variance from one run to the next.
Legal AI firm Eve reported that Opus 5 maintains quality at lower reasoning levels, generating twenty six per cent fewer tokens on average than Opus 4.8 running at maximum reasoning while producing comparable results.
The Safety Picture
Anthropic's automated behavioural audit rated Opus 5 its most aligned model to date, scoring 2.3 on the company's internal misaligned behaviour metric, the lowest of its recent releases, with the company reporting the lowest rates of deceptive behaviour and the least susceptibility to misuse it has measured.
On the cybersecurity side, Opus 5's safety classifiers are set to intervene around eighty five per cent less often than Fable 5's, allowing the model to find vulnerabilities in source code while still blocking binary based scanning, penetration testing and exploit generation. Flagged requests fall back automatically to Opus 4.8 in Claude.ai, Claude Code and Claude Cowork.
The Bottom Line
Opus 5 is not being sold as the smartest model Anthropic makes. It is being sold as the model that gets close enough to the smartest one while costing half as much, which is a different and in some ways more useful pitch for anyone running the model at real scale rather than testing it once. The early access quotes back that framing up more than they oversell it, with several partners specifically noting steadier, more consistent output rather than a single dramatic capability jump.
Whether that trade holds up once millions of people are actually running it in production is the thing worth watching over the next few months.
AI PROMPT OF THE DAY
Category: Vendor Evaluation
"Act as a procurement analyst helping me decide between two AI models for [describe your use case, e.g. coding agent, customer support, data analysis]. One model is priced lower and claims to match the more expensive flagship model on most benchmarks while trailing slightly on a few. Walk me through how to test whether that gap matters for my specific workload, what a fair side by side evaluation would look like using my own tasks rather than published benchmarks, and how to calculate the break-even usage volume where the cheaper model's savings outweigh any performance difference."
ONE LAST THING
Every recent model launch has followed roughly the same script, a chart showing a new model beating the old one, followed by a price that either drops or stays flat. Opus 5 fits that pattern almost exactly, and the more interesting detail sits one layer down.
Anthropic chose to frame its second best model around cost rather than around raw capability, which says something about what the company thinks its customers are actually optimising for once a model moves past the demo stage and into daily use.
Hit reply, I read every response.
See you in the next one.
— Vivek
P.S. Know a developer or founder deciding which Claude model to build on? Forward this to them. They can subscribe at https://savvymonk.beehiiv.com/

