Claude Opus 5 Tops the Benchmarks While Early Testers Call It “Hard to Love”
Anthropic’s new model leads the independent intelligence index and costs half of Fable 5. The reviewers who spent a week with it before launch say it argues with instructions and stops before the work is finished.

Anthropic released Claude Opus 5 on Friday afternoon. By that evening it was top of the Artificial Analysis intelligence leaderboard. The people who had been testing it for a week before launch spent that same evening explaining why they don’t enjoy using it.
That gap is the story. Opus 5 costs $5 per million input tokens and $25 per million output, the same as Opus 4.8. Anthropic says it lands within 0.5% of Fable 5 on CursorBench at half the cost per task. Those numbers hold up. So do the complaints, and they come from the same reviewers who were praising Anthropic’s last model a month ago.
BIG NEWS: Opus 5 is here...and I hate working with it.
— claire vo (@clairevo) July 24, 2026
And yet in a blind taste test, I ranked it above every other model (even Fable and my beloved GPT-5.6)
Claire Vo built ChatPRD and hosts the How I AI podcast. She opened her day-zero review by saying she hates working with it, then said she ranked it above every other model in a blind test, above Fable 5 and above GPT-5.6. Her review is titled “Claude Opus 5 review: this model is brilliant (but annoying)”. It is a seven-model, six-task blind scoring run. Her notes single out the model’s verbosity and what she calls its neurotic personality, including a merge conflict it refused to touch.
What are the testers complaining about?
Dan Shipper and Katie Parrott at Every ran Opus 5 through coding, writing, knowledge work and their internal agent for a week before launch. Their verdict went up as “Vibe Check: Claude Opus 5 Is Brilliant in Flashes, Frustrating in Practice”.
BREAKING: Claude Opus 5 is OUT NOW!
— Dan Shipper (@danshipper) July 24, 2026
And…it's a hard model to love.
We've spent the last week @every testing it across coding, writing, knowledge work, and our internal agent.
It argued with instructions, stopped before the work was finished, and generally didn't play well with our existing skills and plugins like Compound Engineering.
“It’s a poor man’s Fable. It has many of Fable’s personality quirks without Fable’s genius.”
Dan Shipper, Every
The complaints are about behaviour, not capability. It talks back, it stops early, and it ignores scaffolding that worked fine on Opus 4.8. Anthropic’s own prompting guide admits one piece of this: Opus 5’s default responses run longer than previous Opus models. For a release sold on efficiency, that is an odd thing to ship.
So why did it still finish first?
Because on the tests, it does finish first. Artificial Analysis put it at the top of their index and called it the new leader in agentic knowledge work, with 1861 Elo on GDPval-AA v2, more than 100 points clear of both Fable 5 and GPT-5.6 Sol.
Claude Opus 5 is narrowly the most intelligent model on the Artificial Analysis Intelligence Index, offering comparable intelligence to Fable 5 at 26% lower Cost per Task
— Artificial Analysis (@ArtificialAnlys) July 24, 2026
We supported @AnthropicAI to evaluate Claude Opus 5 ahead of release: it sets the highest GDPval-AA v2 and AA-Briefcase scores so far.
Two caveats before you weigh that. Artificial Analysis discloses that Anthropic paid them to evaluate the model ahead of release. And their index runs used Opus 4.8 as a fallback.
Independent results point the same way. ARC Prize recorded 30.2% on ARC-AGI-3 and said the gain comes from stronger logical reasoning. Cursor put it at 66.7 on CursorBench against Fable 5’s 66.5 at default effort. Box CEO Aaron Levie reported gains on his company’s internal eval, including 30% on a life sciences task and 17% on due diligence.
Where does it lose?
Two places, and both sit inside Anthropic’s own numbers.
The first is factual accuracy. Artificial Analysis found accuracy improved 7 points over Opus 4.8, but the model now answers more often when it isn’t sure. That pushed its hallucination rate up 14 points, to 50%. It is the most under-reported finding of the launch.
| Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol | |
|---|---|---|---|---|
| Price per M tokens | $5 / $25 | $10 / $50 | $5 / $25 | — |
| Intelligence Index | 61 | 60 | 56 | 59 |
| Cost per task | $2.03 | $2.75 | $1.80 | $1.04 |
| Agentic coding · FrontierCode | 53.4% | 53.5% | 46.5% | 47.5% |
| Agentic coding · DeepSWE | 68.8% | 69.7% | 59.0% | 72.7% |
| Novel problems · ARC-AGI-3 | 30.2% | — | 1.5% | 7.8% |
The second is agentic coding, where the marketing got ahead of the table. On FrontierCode, Anthropic’s launch chart scored Opus 5 at 53.4% and Fable 5 at 53.5%, then shaded Opus 5 as the winner. Readers on Hacker News caught it within the hour.
Anthropic swapped the image out quietly between 17:02 and 19:46 UTC on launch day and never posted a correction. On the other agentic coding row, DeepSWE, Opus 5 scores 68.8% against Fable 5’s 69.7% and GPT-5.6 Sol’s 72.7%. Opus 5 loses both coding rows in Anthropic’s own table, and one of them was briefly presented as a win.
We pulled every archived capture of Anthropic’s announcement page from launch day and compared which table image each one loads. The 17:02 UTC capture points to an image where Opus 5’s 53.4% carries the winner highlight. The 19:46 UTC capture points to a different image where the highlight sits on Fable 5’s 53.5%. Both files are still live on Anthropic’s CDN, and every other cell in the two versions is identical. Anthropic published no correction notice, so we describe this as a quiet fix rather than an acknowledged one.
Turn the effort dial down
The most useful thing to come out of launch day is that everyone arrived at the same fix independently: give it less.
Shipper’s team found that deleting their existing skills and starting from scratch made results dramatically better, and that medium or low effort beat high. His read was that the more thinking time you give it, the more likely the annoying behaviour shows up.
Anthropic’s own system card has the same finding, as a chart. Opus 5’s FrontierCode score peaks at 53.4% on medium effort, then falls to 48.0% on high and 43.6% on xhigh. Every other model in the chart just climbs.
Anthropic explains why in the same document, and the explanation is the tester complaint in Anthropic’s own words:
“This reflects a tendency for Opus 5 at these effort levels to make more changes than the task requires, e.g. refactoring or making other edits to improve the codebase.”
Claude Opus 5 system card, p.151
Anthropic adds that a short instruction telling the model to stay in scope recovered most of the lost score, and that it published the numbers without that fix applied. So the behaviour is steerable. It just isn’t the default.
Anthropic’s engineers were saying a version of this on launch day too. Adam Wolff, who works on Claude Code, wrote that Opus 5 at medium effort is his go-to for shipping a PR, and that he still reaches for Fable on bigger projects. Thariq on the same team posted that they removed roughly 80% of the Claude Code system prompt for these newer models.
If you’ve built an elaborate prompt library around Opus 4.8, that is the migration cost nobody priced in.
Key background
Anthropic has now shipped three frontier-tier models in about seven weeks. Fable 5 arrived in June as a Mythos-class model, was pulled during a 19-day US export ban, and came back on July 1 behind a safety classifier that routes anything resembling cybersecurity work to Opus 4.8 instead. Users called that nerfed. Then Sonnet 5 landed at a fifth of Fable’s price and took most of the everyday work.
Opus 5 now claims near-Fable performance at half of Fable’s price, without Fable’s classifier in front of it. The question being asked loudest in the Claude community since Friday is what Fable 5 is still for.
One number complicates the cheaper story. Artificial Analysis puts Opus 5’s cost per task at $2.03 at max effort. That is below Fable 5’s $2.75, but above Opus 4.8’s $1.80 and Sonnet 5’s $1.53. Cheaper than the flagship, not cheaper than what you were probably already using.
Introducing Claude Opus 5.
— Claude (@claudeai) July 24, 2026
It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price.
Opus 5 is live now on all paid plans and in the API as claude-opus-5, and Anthropic’s full write-up is on its announcement page. The context window is 1 million tokens, though Pro accounts are being told the 1M version isn’t available to them. If you’re moving over from Opus 4.8, two things changed under you: thinking is on by default, and you can’t switch it off at xhigh or max effort.
Sources & documents
- Anthropic, “Introducing Claude Opus 5,” July 24, 2026 · archived at 17:02 UTC and at 19:46 UTC
- Anthropic · Claude Opus 5 system card (effort-scaling charts, p.151–152)
- Anthropic developer documentation · effort parameter reference
- Artificial Analysis evaluation thread, July 24 · full results and disclosure
- Every, Dan Shipper and Katie Parrott · “Vibe Check: Claude Opus 5” (subscriber content)
- Claire Vo · “Claude Opus 5 review”, Lenny’s Newsletter
- Hacker News launch discussion · item 49038433