Claude Opus 5: Frontier AI Agentic Coding at Unchanged Pricing

Anthropic releases Claude Opus 5 with enhanced agentic coding, computer use, 1M context, and same $5/$25 pricing. Best scores in agentic benchmarks.

sábado, 25 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Rendimiento récord en coding y uso de computadora

Anthropic has released Claude Opus 5, its new flagship model in the Opus series, replacing Opus 4.8 with a clear proposition: frontier-level intelligence at an unchanged price. At $5 per million input tokens and $25 per million output tokens, this model promises to match Claude Fable 5's capabilities at half the cost. For businesses seeking advanced intelligent automation, this combination of price and performance opens up real possibilities for large-scale deployment. At Q2BSTUDIO, as a company specialized in custom software, we see this launch as a milestone that redefines what can be achieved with AI agents in production environments.

From a technical perspective, the most relevant change is that thinking is enabled by default. In Opus 4.8 it was optional; now every request includes an internal reasoning and verification process, controlled by the effort parameter. This forces a review of max_token limits, as they encompass both thinking and visible response. Additionally, disabling thinking at high effort levels returns a 400 error. Anthropic recommends removing manual verification instructions because the model already does that internally. These adjustments are crucial for developers integrating AI into workflows, as they directly affect cost and latency. In our cloud AWS/Azure solutions, for example, we optimize these parameters to balance performance and operational cost.

Benchmark results are compelling. On FrontierBench v0.1, Opus 5 achieves 43.3% at max effort, up from 18.7% for Opus 4.8. On SWE-bench Verified it scores 96.0%, nearing perfection in software engineering tasks. But where it truly shines is in agentic evaluations: OSWorld 2.0 rises to 70.57% and Zapier AutomationBench to 26.0%, doubling or tripling predecessors. This means AI agents can now execute complex automation workflows with much higher reliability. Companies using BI/Power BI can integrate these agents to generate dynamic reports or detect anomalies without constant human intervention.

In mathematical reasoning, Opus 5 solved all six IMO 2026 problems without external tools, achieving a perfect 42/42 score equivalent to a gold medal. On ARC-AGI-3, considered a test of general intelligence, it scores 30.16% versus 1.52% for Opus 4.8. This drastic improvement indicates the model does not just memorize patterns but generalizes and applies abstract logic. For development companies like Q2BSTUDIO, this enables building applications that require complex reasoning, such as cybersecurity systems that identify advanced threats or code assistants that understand deep business contexts. AI ceases to be a black box and becomes a reliable collaborator.

One aspect worth attention is cybersecurity. Anthropic did not specifically train Opus 5 on cyberattack tasks, but its enhanced general capability has raised the skill for finding vulnerabilities. On ExploitBench, the model generated 99 complete arbitrary-code-execution exploits, though below Mythos 5's 132. The company has decided to unblock vulnerability finding in source code, while keeping binary exploitation and exploit generation blocked. This is a delicate balance: companies need powerful pentesting tools but without opening the door to malicious use. At Q2BSTUDIO we offer cybersecurity services that can leverage these capabilities in a controlled manner, always under human supervision and with robust security protocols.

Indirect prompt injection is another front where Opus 5 shows significant progress. On the Gray Swan benchmark, attacker success rate dropped from 5.5% on Opus 4.8 to 2.0%. In browser environments, successful attack fell from 31.5% to 3.7% without safeguards, and to 0% with auto mode enabled. This is key for companies integrating agents into web portals or client applications, as it drastically reduces manipulation risks. Trust in AI systems depends on their resistance to these attack vectors, and Opus 5 demonstrates that progress can be made without compromising security.

From a business perspective, the decision to keep Opus 4.8's pricing is strategic. It allows organizations to access frontier intelligence without inflating costs, facilitating the adoption of AI agents in everyday tasks. At Q2BSTUDIO, we combine this type of model with cloud architectures on AWS or Azure to scale applications efficiently. For instance, an agent-based customer service system can handle hundreds of simultaneous queries with predictable operational cost. Moreover, integration with BI tools like Power BI enables agents to extract real-time conclusions from data and update dashboards automatically. All of this is part of a digital transformation strategy where artificial intelligence is the engine and the cloud is the fuel.

The 1-million-token context, both default and maximum, is another differentiating factor. It allows processing extensive documents, complete codebases, or long conversations without losing coherence. For code auditing or contract analysis tasks, this is a qualitative leap. At Q2BSTUDIO, we have seen development teams integrate agents that review entire pull requests or generate technical documentation from complex specifications. The large context window also benefits recommendation systems and semantic search, where considering large volumes of information is necessary to provide accurate answers.

Finally, it is important to note that Anthropic also publishes the model's limitations. Opus 5 shows a slightly higher rate of factual hallucinations compared to Opus 4.8, a reminder that no technology is perfect. To mitigate this, we recommend combining AI with human verification processes and curated knowledge bases. In our custom software implementations, we always add validation layers that ensure the reliability of generated responses. The future of agentic AI is not about replacing people, but about empowering them with increasingly capable and accessible tools. Claude Opus 5 represents a firm step in that direction, and at Q2BSTUDIO we are ready to help companies make that leap with personalized and secure solutions.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.