Anthropic released Claude Opus 5 on July 24, 2026, and the early verdict from independent testers is that the model delivers near top-tier capability at a mid-tier price. Rolled out across Anthropic's consumer app and developer platform, Opus 5 is being positioned as approaching the quality of the company's most powerful internal systems while costing meaningfully less β a framing that lands squarely in the middle of 2026's fierce competition over capability per dollar.
The Headline: Performance Without the Premium
Opus 5 is priced at $5 per million input tokens and $25 per million output tokens β identical to its predecessor Opus 4.8, and roughly half the rate of Anthropic's higher-end Fable 5 tier. Anthropic says it compared the two models across 13 benchmarks and that Opus 5 scored higher on eight of them despite costing about 50% less. That is the core pitch: not a new performance ceiling, but a substantial improvement in the price-to-capability ratio that most enterprise buyers actually optimize for.
The model ships with a 1-million-token context window, up to 128K tokens of output, and a knowledge cutoff of May 2026 β the most current of any Claude model to date. The API identifier is simply `claude-opus-5`, with matching identifiers on major cloud platforms.
Benchmarks That Stand Out
Several results drew attention on launch day, with at least one independently verified.
- ARC-AGI-3 (fluid reasoning): On FranΓ§ois Chollet's benchmark, which drops an agent into interactive environments with no stated rules or goals and forces it to learn by trial and error, Opus 5 scored 30.2%. For context, the best model scored a fraction of a percent when the benchmark launched in March 2026, and Opus 4.8 barely registered at 1.5%. The ARC Prize Foundation independently verified the 30.2% figure the day the model launched.
- Frontier-Bench v0.1 (agentic terminal coding): Opus 5 posted 43.3%, ahead of Fable 5 at 33.7%, Opus 4.8 at 21.1%, and a competing flagship at 34.4%.
- GDPval-AA v2 (knowledge work): A score of 1861, described as the highest published result from any commercial model on this measure of professional analysis, writing, synthesis and planning.
The ARC-AGI-3 leap is the most scrutinized result, precisely because that benchmark is designed to resist pattern-matching and reward something closer to genuine adaptive reasoning. Independent confirmation from the benchmark's own foundation gives it more weight than vendor-reported figures alone.
A New Lever: The Effort Toggle
Beyond raw scores, Opus 5 introduces a per-request effort setting β low, medium or high β that controls how much reasoning the model spends on a given task. Low runs faster and cheaper for routine work; high lets the model think longer on hard problems. It is Anthropic's explicit lever for balancing cost against capability, and it complicates benchmark comparisons, since a model's score can shift with how much reasoning effort it is allowed to expend.
The pricing structure adds further flexibility. Cached input is billed at one-tenth of the base input rate, and asynchronous batch processing carries a 50% discount. A separate fast mode roughly doubles the token rate but runs about 2.5Γ faster β a trade for latency-sensitive applications.
Where It Sits in the Lineup
Anthropic is clear that Opus 5 is not its most capable system for every job. The company still recommends its higher tiers for the most demanding long-horizon autonomous work and for frontier scientific domains. Opus 5's role is different: a strong default for the broad middle of real workloads β coding, analysis, agentic tasks and knowledge work β where the combination of capability and cost matters more than absolute peak performance.
Availability reflects that positioning. Opus 5 becomes the default for the company's top consumer tier and the strongest option on its mid-tier plan, while free users remain on a smaller model. It is available to paying consumers, teams and enterprises, as well as through the API and major cloud marketplaces.
Why It Matters
Opus 5 is the latest data point in a 2026 pattern: the frontier is advancing less through single dramatic capability jumps and more through relentless compression of cost. The industry has been shipping new state-of-the-art systems at a remarkable cadence this year, and the competitive battleground has moved from "who has the smartest model" to "who delivers the most usable intelligence per dollar."
For developers and enterprises, that shift is unambiguously good. Capabilities that would have commanded premium pricing months ago are now available at commodity-adjacent rates, and controls like the effort toggle let teams tune spend against difficulty on a per-request basis. The practical consequence is that ambitious agentic and analytical applications β once gated by inference costs β become financially viable for far more teams. As always, the caveat holds: several launch figures come from the vendor, so production decisions should lean on independently verified results and real-world testing rather than headline numbers alone.
