TrueSeeker AI · Verified claim report Case 096500d2b4 · 2026-09-26

§ Claim under review · Release

"A Anthropic lançou o Claude Opus 5.5, que rende no nível do Fable 5.1 (modelo mais potente da Anthropic) na maioria das tarefas e custa 40% menos para rodar do que o Opus 5."

Circulating claim, as submitted.

Verdict

Mostly accurate

Confidence

Medium
§

Summary

This post is largely accurate. Anthropic really did release Claude Opus 5.5 on September 22, 2026, and it is fully available now, not a preview. The two main statements in the post are almost word for word what Anthropic itself published: that the model performs at the level of Claude Fable 5.1 on most work and costs 40 percent less to run than Opus 5. The post also deserves credit for naming Anthropic as the source of its benchmark chart and for pointing out two tests where a rival model scored higher. The important missing detail is the 40 percent figure. Anthropic ties that number to default settings on typical workloads, and the actual published price cut is 20 percent on input and output tokens with cache reads down 60 percent. One independent measurement found that if you run the model at its highest effort setting, the cost saving over the previous model largely disappears, and that same independent test found a much narrower coding lead over the competing model than the launch chart shows.

§

The readings

key figures from the evidence
40 %

claimed cost reduction vs Opus 5, per Anthropic's post

§

Why this verdict

Both operative propositions, Fable-level performance on most tasks and 40 percent lower running cost than Opus 5, are near-verbatim renderings of Anthropic's own published sentence, the release is confirmed as generally available on the vendor's official channels as of 2026-09-25, and an independent evaluator's aggregate index corroborates the performance half rather than undercutting it. I considered "Accurate" and rejected it because the cost figure loses its scope in transmission: Anthropic ties it to default settings and typical workloads, and at maximum effort the independently reported per-task saving disappears. I considered "Partially accurate but misleading" and rejected it because the post credits Anthropic as the source of its numbers, discloses the two benchmarks where the model lost, and explicitly warns against hype, so no cited source contradicts the operative propositions and the framing does not reverse their meaning. Confidence is Medium rather than High because the comparative and cost claims rest mainly on vendor-run numbers, and the one independent check available diverges from the vendor table on the headline coding benchmark and on cost at non-default settings.
§

Evidence

The release is real and the two headline propositions are near-verbatim translations of Anthropic's own published sentence. Anthropic's announcement page states that Opus 5.5 is the first model in the new Claude 5.5 family, that it "performs at the level of Claude Fable 5.1 on most work", and that it "costs 40% less to run than Opus 5".

On cost, Anthropic's own page is more specific than the post. It states that "at default settings it will cost 40% less than Opus 5 on typical workloads". The list price change is smaller than 40 percent: input and output tokens are $4 and $20 per million, described by Anthropic as 20 percent below Opus 5, while cache reads fall to $0.20 per million, 60 percent below Opus 5. Anthropic also reports output generated more than 30 percent faster. The 40 percent figure is a blended estimate combining price cuts with token efficiency at default effort, not a list price cut.

An independent evaluator, Artificial Analysis, was measured differently. As reported by several outlets, it placed Opus 5.5 at 58 on its Intelligence Index at maximum effort, ahead of both GPT-6 Astra and Fable 5.1 at 53. That corroborates the performance-parity proposition and arguably exceeds it. On cost, however, the same evaluator's decomposition reportedly found that at maximum effort Opus 5.5 cost about $5.98 per Intelligence Index task against about $5.86 for Opus 5 at maximum effort, meaning the per-task saving disappears at that setting. Opus 5.5 defaults to medium effort while Opus 5 defaulted to high, which is the comparison the 40 percent figure rests on.

On the benchmark table the post reproduces, the figures match Anthropic's published numbers, and the post correctly names Anthropic as the source. The headline coding row carries an asymmetry: Anthropic's table reports Opus 5.5 at 66.4 percent on Terminal-Bench 4.0 run at extra-high effort against GPT-6 Astra at 57.9 percent taken at high effort from OpenAI's own figure. Artificial Analysis, running its own harness, reportedly measured Opus 5.5 at 59.6 percent on the same benchmark, level with Astra rather than ahead. Anthropic itself reports a standard error of plus or minus 2.6 points on that row, and on Terminal-Bench-Science the stated standard error is plus or minus 3.5 to 5.0 points, which is larger than the gap the post highlights on several other rows.

§

Findings

✓ What's accurate 6

  • Anthropic did release Claude Opus 5.5, on September 22, 2026, as the first model in a new Claude 5.5 family. It is generally available, not a preview or a waitlist.
  • Anthropic's own announcement says the model performs at the level of Claude Fable 5.1 on most work, which is what the post says.
  • Anthropic's own announcement says the model costs 40 percent less to run than Opus 5, which is what the post says.
  • The benchmark figures reproduced in the post match Anthropic's published launch table, and the post names Anthropic's table as the source rather than presenting the numbers as independent.
  • The post is correct that the model did not lead everywhere. On Anthropic's own table, GPT-6 Astra leads on AutomationBench at 41.4 percent against 40.0 percent, and on Terminal-Bench-Science at 64.6 percent against 58.7 percent.
  • An independent evaluator's aggregate index supports the performance-parity proposition, placing Opus 5.5 above Fable 5.1 on that index.

≈ What's misleading 5

  • Omitted qualifier: the post states flatly that the model costs 40 percent less to run, and one slide says it simply "got 40 percent cheaper". Anthropic's own wording scopes the figure to default settings on typical workloads. The list price cut is 20 percent on input and output tokens, with cache reads down 60 percent, and the 40 percent figure blends those cuts with the model's token efficiency at its default effort level. A reader who takes it as a flat 40 percent price reduction has the wrong number.
  • Cost compute omission: the saving depends on the effort setting, and the post does not mention effort settings at all. Opus 5.5 defaults to medium effort while Opus 5 defaulted to high. At maximum effort, the independent per-task cost measurement reported for Opus 5.5 is roughly the same as Opus 5, so the advertised saving does not hold across the range of ways someone might actually run the model.
  • Harness mismatch: the slide showing Opus 5.5 leading GPT-6 Astra on agentic coding rests on a row where the two numbers were produced at different effort settings, with Astra's figure taken from OpenAI's own reporting. An independent runner using one harness measured the two as level on that benchmark. The lead is directionally real on the index level, but the size shown is a product of the comparison setup.
  • Marketing as evidence: although the post credits Anthropic's table, it presents vendor-run results as settling where the model leads, including a slide asserting leadership across every benchmark listed. Vendor numbers establish what the vendor claims, not what is independently true, and the one independent check available narrows several of those leads.
  • Omitted qualifier: the computer-use figure of 81.8 percent on OSWorld 2.0 is a partial-credit score. The strict completion score reported in the system card is 48.7 percent. Presenting only the partial figure under the heading of autonomous computer use overstates how often tasks are finished end to end.

? What's uncertain 4

  • Whether Fable 5.1 is correctly described as Anthropic's most powerful model is genuinely ambiguous. Fable sits in a Mythos class that Anthropic positions above the Opus class, and Fable 5.1 was the prior public flagship. But in the same announcement the post draws from, Anthropic calls Opus 5.5 its new leading model, and there is also a restricted-access Mythos 5.1. This is a positioning and definition question rather than a factual error.
  • The Artificial Analysis figures, including the 59.6 percent Terminal-Bench result and the per-task cost decomposition, were read through third-party reports rather than from the evaluator's own pages, so the exact methodology behind them is not confirmed here.
  • Whether the 40 percent saving holds for any particular real workload is untestable from the published material, because "typical workloads" is not defined in the announcement.
  • No independent verification was found for the post's Opus 5 comparison figure on Humanity's Last Exam of 63.6 percent, although nothing contradicts it either.
Distortion flags omitted qualifier cost compute omission harness mismatch marketing as evidence exaggeration
§

Sources

9 of 9 linked to records
[1]

Anthropic, "Introducing Claude Opus 5.5", official announcement and benchmark table

primary vendor official channel
https://www.anthropic.com/claude-opus-5-5 ↗
[2]

Anthropic, Claude Opus product and pricing page

primary vendor official channel
https://www.anthropic.com/claude/opus ↗
[3]

Claude Platform pricing documentation

primary vendor official docs
https://platform.claude.com/docs/en/about-claude/pricing ↗
[4]

AWS, "Claude Opus 5.5 is now available on AWS"

primary cloud provider official channel
https://aws.amazon.com/blogs/machine-learning/claude-opus-5-5-is-now-available-on-aws/ ↗
[5]

Artificial Analysis measurements, as reported by Kingy AI, Latent Space AINews, Emergent and Implicator

secondary independent benchmark runner, read through third parties
https://kingy.ai/blog/claude-opus-5-5-specs-benchmarks-pricing-comparison/ ↗
[6]

TechCrunch, "Anthropic releases Opus 5.5 with lower prices and Fable-level performance"

secondary named-outlet journalism
https://techcrunch.com/2026/09/22/anthropic-releases-opus-5-5-with-lower-prices-and-fable-level-performance/ ↗
[8]

Vellum and Orca Router benchmark breakdowns

secondary specialist commentary
https://www.vellum.ai/blog/claude-opus-5-5-benchmarks-explained ↗
[9]

Anthropic, "Introducing Claude Fable 5.1 and Claude Mythos 5.1"

primary vendor official channel
https://www.anthropic.com/claude-fable-and-mythos-5-1 ↗
How links are chosen. A source is linked only when the address comes from the investigation's own retrieval or from a registry lookup (PubMed, Crossref) that matches the citation's title and year. Author lists shown as registry-verified come from the registry record, not from the report text. Citations that cannot be matched are labeled, never guessed.
This is one case on the record See the full case, browse the archive, and search every checked claim on TrueSeeker AI Open on ai.trueseeker.com →