AI News

Claude Opus 5 Tops the Benchmarks While Early Testers Call It “Hard to Love”

Anthropic’s new model leads the independent intelligence index and costs half of Fable 5. The reviewers who spent a week with it before launch say it argues with instructions and stops before the work is finished.

Kaustubh Saini
Founder, FavTutor · Writes about AI models, tools and news
Published
Updated · 6 min read
Paper-collage illustration headed Claude Opus 5 with the Claude sunburst mark, beside a figure on a first-place podium rubbing his eyes next to a bar chart and a crumpled sheet of paper

Anthropic released Claude Opus 5 on Friday afternoon. By that evening it was top of the Artificial Analysis intelligence leaderboard. The people who had been testing it for a week before launch spent that same evening explaining why they don’t enjoy using it.

That gap is the story. Opus 5 costs $5 per million input tokens and $25 per million output, the same as Opus 4.8. Anthropic says it lands within 0.5% of Fable 5 on CursorBench at half the cost per task. Those numbers hold up. So do the complaints, and they come from the same reviewers who were praising Anthropic’s last model a month ago.

Claire Vo built ChatPRD and hosts the How I AI podcast. She opened her day-zero review by saying she hates working with it, then said she ranked it above every other model in a blind test, above Fable 5 and above GPT-5.6. Her review is titled “Claude Opus 5 review: this model is brilliant (but annoying)”. It is a seven-model, six-task blind scoring run. Her notes single out the model’s verbosity and what she calls its neurotic personality, including a merge conflict it refused to touch.

What are the testers complaining about?

Dan Shipper and Katie Parrott at Every ran Opus 5 through coding, writing, knowledge work and their internal agent for a week before launch. Their verdict went up as “Vibe Check: Claude Opus 5 Is Brilliant in Flashes, Frustrating in Practice”.

“It’s a poor man’s Fable. It has many of Fable’s personality quirks without Fable’s genius.”

Dan Shipper, Every

The complaints are about behaviour, not capability. It talks back, it stops early, and it ignores scaffolding that worked fine on Opus 4.8. Anthropic’s own prompting guide admits one piece of this: Opus 5’s default responses run longer than previous Opus models. For a release sold on efficiency, that is an odd thing to ship.

So why did it still finish first?

Because on the tests, it does finish first. Artificial Analysis put it at the top of their index and called it the new leader in agentic knowledge work, with 1861 Elo on GDPval-AA v2, more than 100 points clear of both Fable 5 and GPT-5.6 Sol.

Two caveats before you weigh that. Artificial Analysis discloses that Anthropic paid them to evaluate the model ahead of release. And their index runs used Opus 4.8 as a fallback.

Independent results point the same way. ARC Prize recorded 30.2% on ARC-AGI-3 and said the gain comes from stronger logical reasoning. Cursor put it at 66.7 on CursorBench against Fable 5’s 66.5 at default effort. Box CEO Aaron Levie reported gains on his company’s internal eval, including 30% on a life sciences task and 17% on due diligence.

This is the chart the pricing claim rests on. Each line is one model climbing its five effort settings. Opus 5 reaches roughly **70%** at about **$8** per task; Fable 5 needs about **$17** to reach the same place. **Image Credits:** Anthropic
This is the chart the pricing claim rests on. Each line is one model climbing its five effort settings. Opus 5 reaches roughly 70% at about $8 per task; Fable 5 needs about $17 to reach the same place. Image Credits: Anthropic

Where does it lose?

Two places, and both sit inside Anthropic’s own numbers.

The first is factual accuracy. Artificial Analysis found accuracy improved 7 points over Opus 4.8, but the model now answers more often when it isn’t sure. That pushed its hallucination rate up 14 points, to 50%. It is the most under-reported finding of the launch.

Opus 5Fable 5Opus 4.8GPT-5.6 Sol
Price per M tokens$5 / $25$10 / $50$5 / $25
Intelligence Index61605659
Cost per task$2.03$2.75$1.80$1.04
Agentic coding · FrontierCode53.4%53.5%46.5%47.5%
Agentic coding · DeepSWE68.8%69.7%59.0%72.7%
Novel problems · ARC-AGI-330.2%1.5%7.8%
Shaded cell is the leader in that row. Index and cost-per-task figures are Artificial Analysis at max effort; benchmark rows are Anthropic’s own comparison table. Fable 5 was not scored on ARC-AGI-3 by Anthropic.

The second is agentic coding, where the marketing got ahead of the table. On FrontierCode, Anthropic’s launch chart scored Opus 5 at 53.4% and Fable 5 at 53.5%, then shaded Opus 5 as the winner. Readers on Hacker News caught it within the hour.

The same row before and after Anthropic replaced the chart image on launch day. The numbers never changed, only the highlight. **Image Credits:** Anthropic, compiled by FavTutor
The same row before and after Anthropic replaced the chart image on launch day. The numbers never changed, only the highlight. Image Credits: Anthropic, compiled by FavTutor

Anthropic swapped the image out quietly between 17:02 and 19:46 UTC on launch day and never posted a correction. On the other agentic coding row, DeepSWE, Opus 5 scores 68.8% against Fable 5’s 69.7% and GPT-5.6 Sol’s 72.7%. Opus 5 loses both coding rows in Anthropic’s own table, and one of them was briefly presented as a win.

How we checked the chart swap

We pulled every archived capture of Anthropic’s announcement page from launch day and compared which table image each one loads. The 17:02 UTC capture points to an image where Opus 5’s 53.4% carries the winner highlight. The 19:46 UTC capture points to a different image where the highlight sits on Fable 5’s 53.5%. Both files are still live on Anthropic’s CDN, and every other cell in the two versions is identical. Anthropic published no correction notice, so we describe this as a quiet fix rather than an acknowledged one.

Turn the effort dial down

The most useful thing to come out of launch day is that everyone arrived at the same fix independently: give it less.

Shipper’s team found that deleting their existing skills and starting from scratch made results dramatically better, and that medium or low effort beat high. His read was that the more thinking time you give it, the more likely the annoying behaviour shows up.

Anthropic’s own system card has the same finding, as a chart. Opus 5’s FrontierCode score peaks at 53.4% on medium effort, then falls to 48.0% on high and 43.6% on xhigh. Every other model in the chart just climbs.

Opus 5 is the orange line. It peaks at medium effort and then drops, which is the opposite of what the other three models do. **Image Credits:** Anthropic, Claude Opus 5 system card, p.151
Opus 5 is the orange line. It peaks at medium effort and then drops, which is the opposite of what the other three models do. Image Credits: Anthropic, Claude Opus 5 system card, p.151

Anthropic explains why in the same document, and the explanation is the tester complaint in Anthropic’s own words:

“This reflects a tendency for Opus 5 at these effort levels to make more changes than the task requires, e.g. refactoring or making other edits to improve the codebase.”

Claude Opus 5 system card, p.151

Anthropic adds that a short instruction telling the model to stay in scope recovered most of the lost score, and that it published the numbers without that fix applied. So the behaviour is steerable. It just isn’t the default.

Anthropic’s engineers were saying a version of this on launch day too. Adam Wolff, who works on Claude Code, wrote that Opus 5 at medium effort is his go-to for shipping a PR, and that he still reaches for Fable on bigger projects. Thariq on the same team posted that they removed roughly 80% of the Claude Code system prompt for these newer models.

If you’ve built an elaborate prompt library around Opus 4.8, that is the migration cost nobody priced in.

Key background

Anthropic has now shipped three frontier-tier models in about seven weeks. Fable 5 arrived in June as a Mythos-class model, was pulled during a 19-day US export ban, and came back on July 1 behind a safety classifier that routes anything resembling cybersecurity work to Opus 4.8 instead. Users called that nerfed. Then Sonnet 5 landed at a fifth of Fable’s price and took most of the everyday work.

Opus 5 now claims near-Fable performance at half of Fable’s price, without Fable’s classifier in front of it. The question being asked loudest in the Claude community since Friday is what Fable 5 is still for.

One number complicates the cheaper story. Artificial Analysis puts Opus 5’s cost per task at $2.03 at max effort. That is below Fable 5’s $2.75, but above Opus 4.8’s $1.80 and Sonnet 5’s $1.53. Cheaper than the flagship, not cheaper than what you were probably already using.

Opus 5 is live now on all paid plans and in the API as claude-opus-5, and Anthropic’s full write-up is on its announcement page. The context window is 1 million tokens, though Pro accounts are being told the 1M version isn’t available to them. If you’re moving over from Opus 4.8, two things changed under you: thinking is on by default, and you can’t switch it off at xhigh or max effort.

Sources & documents

  1. Anthropic, “Introducing Claude Opus 5,” July 24, 2026 · archived at 17:02 UTC and at 19:46 UTC
  2. Anthropic · Claude Opus 5 system card (effort-scaling charts, p.151–152)
  3. Anthropic developer documentation · effort parameter reference
  4. Artificial Analysis evaluation thread, July 24 · full results and disclosure
  5. Every, Dan Shipper and Katie Parrott · “Vibe Check: Claude Opus 5” (subscriber content)
  6. Claire Vo · “Claude Opus 5 review”, Lenny’s Newsletter
  7. Hacker News launch discussion · item 49038433
Note on the blind test: Claire Vo’s first-place ranking for Opus 5 is her own reported result. The full leaderboard appears only inside her video review and has not been independently reproduced.
TagsClaude Opus 5AnthropicModel LaunchesAI BenchmarksLLM Pricing
Kaustubh Saini, founder of FavTutor
Kaustubh Saini
Founder & Technical Writer · FavTutor

I’m Kaustubh Saini, founder of FavTutor. I love breaking down complex AI concepts, trends, and news, writing about them until an AGI agent takes over my job. When I’m not writing, I’m building AI-powered tools to make learning more accessible and engaging at FavTutor.