You picked Claude Sonnet 5.5 - but Anthropic may send your request to Sonnet 5
The New Stack

You picked Claude Sonnet 5.5 - but Anthropic may send your request to Sonnet 5

Anthropic’s Claude Sonnet 5.5, released Monday, is the first Sonnet model to launch with cyber safeguards and model fallbacks like those the company developed for its most capable models, putting classifier-driven routing in the tier many developers use for production workloads. Anthropic says Sonnet 5.5 doesn’t advance the frontier of its models’ capabilities, but it rates the model’s cybersecurity skills as comparable to Opus 5’s, and on Terminal-Bench 4.0, an agentic coding benchmark, the cheaper model scores 70.6% against 66.4% for Opus 5.5 at its xhigh effort setting. The release shows how a model can still sit below Anthropic’s most powerful models overall while becoming capable enough in one area to require the same kinds of safeguards.

Exploit gains trigger new safeguards

Anthropic’s system card shows just how much Sonnet 5.5 has improved at offensive security tasks. With its cyber safeguards turned off, the model achieved full arbitrary code execution in 178 of 410 ExploitBench runs. It also completed 46.1% of the challenges on Irregular’s CyScenarioBench, up from just 0.7% for Sonnet 5, and managed 50 control-flow hijacks on a binary exploitation benchmark based on Google’s OSS-Fuzz corpus, compared with three for Sonnet 5. Anthropic still considers Sonnet 5.5 less capable at cybersecurity than Opus 5.5 and Mythos 5.1, but the jump from the previous Sonnet was large enough for the company to apply the same cyber policy it uses with Opus 5 and Opus 5.5.

Three-stage cyber enforcement

Enforcement happens in three stages, beginning with a probe that reads the model’s internal activations, followed by a lightweight classifier running on Sonnet 5.5 itself and a separate trained LLM classifier that weighs the probe’s verdict in deciding whether to block a conversation. Anthropic says the classifiers catch harmful cyber requests at a rate comparable to those on Opus 5, but it chose less aggressive jailbreak protections because Sonnet 5.5 isn’t as capable at cybersecurity as Opus 5 or Fable 5.1. The company also says users should expect more refusals than they saw with Sonnet 5, including on legitimate cybersecurity work.

When Sonnet 5 launched, Anthropic turned on cyber safeguards by default, using protections similar to those on Opus 4.7 and 4.8, models the company said performed substantially better than Sonnet 5 on cyber exploit-development evaluations. The shift grew more pronounced with Opus 5.5 last week, which Anthropic considers comparable to Mythos 5.1 in biology and cybersecurity and which became the first Opus model released with safeguards similar to Fable 5.1’s. Those safeguards determine which model answers a request, since most cybersecurity tasks sent to Opus 5.5 go to Opus 4.8, and requests flagged for biology or frontier-model development go to Opus 5; Anthropic acknowledged that these interventions likely lowered Opus 5.5’s scores on benchmarks it ran with the safeguards switched on.

Routing and fallback

Sonnet 5.5 brings that routing to the cheaper tier, where blocked cyber requests, along with a narrow set of requests tied to frontier LLM development such as kernel work on certain ML accelerators, fall back to Sonnet 5. Blocks for biology, conventional weapons, and anti-distillation classifiers end the request without any fallback model, and Anthropic says these blocks are transparent and do not covertly change the model’s responses.

API fallback is opt-in

How fallback works depends on where developers are using the model. Anthropic’s own apps automatically send blocked cyber requests to Sonnet 5, but API developers must enable that fallback themselves, and other platforms and providers may handle blocked requests differently. That means developers moving from Sonnet 5 to Sonnet 5.5 can’t reliably treat it as a straight model swap. Without fallback enabled, a blocked request stops rather than being passed to Sonnet 5.

Policy details

The cyber policy allows vulnerability discovery in source code, which keeps secure-coding workflows intact, but it blocks vulnerability discovery in compiled binaries. Anthropic’s support documentation also says the checks review everything the model reads, including memory, connector content, web search results, and files, so content nobody typed can trigger a fallback. For agents pulling information from repositories, security advisories, or web pages, content returned by those tools can also trigger the safety system.

Fallback weakens injection defenses

Fallback also introduces another potential weak spot: prompt injection. In Anthropic’s testing of coding environments, 25% of requests sent to Sonnet 5.5 were handled by Sonnet 5 after triggering a cyber block, often because injected instructions to wipe disks or delete files set off the classifier. Of those rerouted requests, 12.01% were successfully compromised. By comparison, Sonnet 5.5 was compromised in just four of the 5,901 requests it handled itself. A separate indirect prompt injection benchmark from AI security company Gray Swan found no performance drop with fallback enabled. Still, Anthropic’s coding tests show that teams using fallback need to account for the security of the older model as well as Sonnet 5.5 itself.

Fable 5 example

Fable 5 showed how disruptive fallback can become in specialized workflows. The day after the model launched in June, a Claude Code user doing defensive threat-intelligence work reported on GitHub that 2,746 of 3,427 main-session messages that day had been generated by Opus 4.8 after fallbacks. The user also said that once the session switched models, it never returned to Fable 5 on its own.

Future plans

Anthropic says it is still tuning Sonnet 5.5’s classifiers to reduce false positives, and it plans to give verified defenders access to the model with fewer restrictions through an expanded Cyber Verification Program. Sonnet 5.5 isn’t available in that program at launch, and the GitHub user who reported the Fable 5 fallbacks said their organization was already enrolled in it.

Read on The New Stack ↗ ← Back to News

Comments

No comments yet. Start the discussion.