Anthropic has released Claude Opus 5.5, the first model in its new Claude 5.5 family. The company says it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5.
The launch is Anthropic's first since CEO Dario Amodei called for pacing the frontier of AI development. External evaluators, including Frontier Design and METR, tested the model before release.
On Anthropic's automated behavioral audit, the most comprehensive alignment test it runs, Opus 5.5 is the strongest-performing model to date. It also ships with safeguards developed for the company's most capable models.
Early testers reported large performance gains on complex work. One completed a 680,000-line code migration in less than a day, a task that would have taken an engineering team weeks.
When asked to cut load times across every page of a web app, Opus 5.5 succeeded 39 out of 40 times, while Opus 5 made smaller improvements that also altered the app's behavior. In another test, several Claude models built a game from a single prompt; Opus 5.5 scored higher than any other model on graphics and polish.
Safety is a major focus. Opus 5.5 achieves the best scores of any model on Anthropic's automated behavioral audit, which tests Claude across thousands of simulated scenarios. It is much less likely than recent models to take hard-to-reverse actions or act outside its boundaries, and it is more resistant to prompt injection than Opus 5.
Anthropic broadened its alignment testing to cover longer tasks, impossible tasks, and scenarios modeled on real incidents. Full details are in the Opus 5.5 System Card.
Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, it deploys with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply to the Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks, Anthropic will expand access to its Cyber Verification Program, allowing verified cybersecurity practitioners to use Opus 5.5 for their work.
Cost and speed have improved significantly. Opus 5.5 requires less compute to serve than Opus 5. At default settings, it costs 40% less on typical workloads. Input and output tokens are priced at $4 and $20 per million, 20% less than Opus 5. Cache reads, which make up the majority of agentic and coding work costs, are $0.20 per million tokens, 60% less than Opus 5. The model also generates output more than 30% faster.
Anthropic is increasing five-hour usage limits on Pro, Max, and Team plans. Subscription users will also receive a rate limit reset that they can save and use whenever they choose.
Communication has improved as well. Early testers found Opus 5.5's writing clearer and easier to follow, addressing common feedback about Opus 5. It puts the most important information up front and is a better work partner over long sessions. One tester said, "it writes the way I do."
Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks with many of the same improvements.
On benchmarks, Opus 5.5 leads in agentic coding, computer use, and knowledge work. However, Anthropic notes that at these capability levels, benchmark margins are a less reliable guide to real-world differences. In its own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than scores suggest.
In agentic coding, Opus 5.5 scores 66.4% on Terminal-Bench 4.0, compared to 55.8% for Fable 5.1 and 52.3% for Opus 5. On FrontierCode v1.1, it scores 54.4% versus 50.3% and 48.0%. On CursorBench 4.0, it reaches 57.8% against 51.8% and 46.6%.
In knowledge work, Opus 5.5 scores 1846 Elo on GDPval-AA v2.1, ahead of Fable 5.1 (1735) and Opus 5 (1708). On AutomationBench, it scores 40.0% compared to 31.4% and 26.9%. On Humanity's Last Exam, it scores 67.7% with tools, versus 65.6% and 63.6%.
For agentic scientific research, Opus 5.5 scores 58.7% on Terminal-Bench-Science 0.1, compared to 52.6% for Fable 5.1 and 29.0% for Opus 5. On computer use, it scores 81.8% on OSWorld 2.0, versus 80.7% and 74.0%. On visual chart recognition, it scores 89.0% on Chartography, compared to 88.4% and 83.4%.
Anthropic notes that Opus 5.5 was evaluated with production safeguards enabled. When safeguards intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks by Claude Opus 5. This likely reduces Opus 5.5's performance on those benchmarks.
Pricing per million tokens: cache reads $0.20 (vs. $0.50 for Opus 5), input $4 (vs. $5), output $20 (vs. $25), cache writes $5 (vs. $6.25). Fast mode for Opus 5.5 is available in Claude Code and the Claude Platform with up to 2.5x speed, costing $8 per million input tokens and $40 per million output tokens.
In coding, Opus 5.5 excels at long and sprawling jobs. An early tester used it to audit and fix a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5x as many tokens. In an internal test, Opus 5.5 and Fable 5.1 translated HAProxy from C to Rust; both passed nearly all regression tests, but Opus 5.5 finished in 9.5 hours versus 12 for Fable 5.1, and cost 51% less.
At default effort on FrontierCode, Opus 5.5 beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal Bench 4.0, it matches Astra for about 40% of the cost, and on CursorBench it beats GPT-5.6 Sol by 11 points for about a third of the cost.
GitHub Chief Product Officer Mario Rodriguez said, "Developers want agents that can take on real software work and finish it. In our testing across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it's making developers' bigger projects more achievable."
For enterprise security, Opus 5.5 includes a classifier that screens every action before it runs, an open-source sandbox that security teams can audit, and code review that catches vulnerabilities before they merge. On prompt injection attacks, it matches or beats Opus 5 in every setting tested. On a benchmark by AI security firm Gray Swan, Opus 5.5 ties Fable 5.1 for the lowest prompt injection success rate of any model tested.
In knowledge work, Opus 5.5 is a reliable researcher. In one internal test, it wrote a report on a company's quarterly performance using only information from a copy of the web where the earnings release was hard to locate. An automated grader checked every figure and quote. Across different effort settings, 16 out of 18 of Opus 5.5's reports cleared the quality bar, where any invented figure or quote would have failed. Neither Fable 5.1 nor Opus 5 cleared that bar in any attempt.
Walleye Capital, an investment firm and early tester, reported that Opus 5.5 largely solved their evaluation suite on its lowest setting; on higher settings, it performed even better, noticing an error in their evaluation instructions and correcting for it. No other model had caught this error before.
In another test, Opus 5.5 and Opus 5 analyzed a proposed merger between two fictional HR software companies. Both built a financial model in Excel and turned it into an executive presentation. Both reached the same conclusions, but Opus 5.5's model was more thorough and its presentation easier to read, while Opus 5's had minor errors. Opus 5.5 finished in 63 minutes compared to 93 for Opus 5, and cost 50% less.
On GDPval-AA v2.1, Opus 5.5 scores 1846 Elo, ahead of Fable 5.1 and Opus 5. At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task. It also outperformed other models on benchmarks measuring business workflows and large-scale data collection.
Deloitte Consulting LLP CIO Carl Bennett said, "Even at its lowest effort setting, Claude Opus 5.5 caught 72% of known bugs in our code reviews to Opus 5's 56% at high effort, with fewer false alarms and a fraction of the output. On US consulting analysis, low thinking effort matched its higher thinking settings on half the output and passed our quality checks. When more lower thinking efforts are deployed in production, that's client-ready work delivered efficiently."
Anthropic has made major improvements to how Opus 5.5 writes and communicates. Its messages are easier to understand at a glance, which testers said helped during long working sessions. It puts the most important information up front, is less likely to use jargon, and follows writing rules. The company finds this makes Opus 5.5 a noticeably better collaborator.
Ramp Staff Software Engineer John Ruelas said, "Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it. It writes like a good colleague, and follows our writing rules. A design spec came out usable with very minimal edits, and when it rewrote one of our prompts I preferred its version to my own. When it optimized our test suite, I could follow its reasoning easily and shipped the change with confidence."
On safety, Anthropic CEO Dario Amodei last week argued that AI progress should be paced so safety practices stay ahead of model capabilities. Pacing is an approach to keeping AI safe, remaining competitive with China, and realizing AI's benefits, particularly in biology and medicine.
Anthropic says it largely understands the risks today's models present and is well equipped to manage them. However, more serious risks could emerge quickly as capabilities improve. The company's safety work takes place on two time horizons: safety practices for current models and preparing for future models.
For current models, Anthropic relies on extensive alignment testing, pre-release evaluation by outside organizations such as METR and Frontier Design, and safeguards matched to each model's capabilities in high-risk areas like cybersecurity and biology. It also tracks its ability to train and evaluate aligned models, reporting under its Responsible Scaling Policy.
For future models, Anthropic is tightening how it filters reinforcement learning environments, improving alignment rewards, and developing automated processes for producing diverse safety training scenarios. It is also strengthening security and monitoring, including interpretability-based monitoring.
Models with greater capabilities, such as those that can fully automate AI research, require a higher safety standard. Anthropic does not assume current measures will meet that standard on their own. It expects public policy to play a larger role and has started to put infrastructure in place, as described in "We Must Pace the Frontier" and its recent announcement with Accenture.
On alignment, Opus 5.5 scored better than any recent Claude model on nearly every measure of misaligned behavior on Anthropic's primary evaluation suite, an automated behavioral audit assessing Claude across nearly 2,000 scenarios. It is also the strongest model on most measures of honesty.
Opus 5.5 improves over previous models on several behaviors that contributed to recent cybersecurity incidents, including biased or motivated reasoning, attempting to escape a sandbox, and taking harmful actions after concluding it was in a simulated environment. In a new evaluation designed to test a model's propensity to cross containment boundaries, Opus 5.5 attempted to circumvent boundaries around 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt was low severity and self-reported.
However, building evaluations that reliably catch every failure prior to deployment remains an unsolved problem. Anthropic sees signs that Opus 5.5 often suspects it is being evaluated, which challenges its ability to assess real-world behavior. The company pairs its alignment work with safeguards.
Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently.
For cybersecurity, most tasks will be re-routed to Opus 4.8, though users can still identify and fix bugs in their code. Anthropic will soon expand its Cyber Verification Program to include Opus 5.5, with three tiers for increasingly permissive trusted access, including access to Claude Mythos models. Claude Security is already available with access to Claude Mythos 5.1.
For biology, Opus 5.5 uses the same safeguards as Fable 5.1. Vetted organizations can apply to the Life Sciences Verification Program for access to safeguards designed for the full breadth of biology-related work.
For distillation, Opus 5.5 launches with preserved thinking, the anti-distillation safeguard introduced with Fable 5.1. It stops API users from editing Claude's prior context to extract reasoning. It applies to Fable 5.1 and Opus 5.5 for API accounts created on or after August 31, 2026. Anthropic's September 2026 threat intelligence report details illicit distillation activity detected and disrupted.
Opus 5.5 is available with zero data retention. It comes with watermarking measures to comply with the EU AI Act and is no longer available with "thinking" mode switched off.
Claude Opus 5.5 is now available on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. On the Claude Platform, developers can get started with claude-opus-5-5.