Claude Opus 5 Cuts Indirect Prompt Injection Attack Success to 2% in New Benchmark Analysis

Blog WriterCybersecurity News - Original News Source is cybersecuritynews.com

Spread the love

Anthropic’s Claude Opus 5 has recorded the lowest indirect prompt injection attack success rate in Gray Swan’s latest benchmark, according to results provided in its system card.

The model reduced an attacker’s chance of success within 15 attempts to 2.0%. This result places Opus 5 ahead of every other model tested, including earlier Claude releases and competing frontier systems.

The finding highlights growing attention on indirect prompt injection, a security issue that can affect AI agents connected to documents, websites, email, and business tools.

Indirect prompt injection occurs when malicious instructions are hidden in untrusted content that an AI system later reads. A poisoned webpage, document, or message may try to override the user’s task, expose sensitive information, or trigger unsafe actions.

The danger increases when models can use tools, retrieve data, or act across enterprise environments. Strong model behavior alone does not remove the risk. However, it can reduce the chance that hostile instructions successfully change an agent’s decisions.

On the IPI benchmark, Opus 5 improved substantially over Claude Opus 4.8. Its attack success rate at 15 attempts declined from 5.5% to 2.0%. In the single-attempt setting, the rate fell from 0.5% to 0.2%.

The results also surpassed those of Claude Sonnet 5, which recorded 5.9% over 15 attempts, and Claude Mythos 5, which reached 2.6%. Anthropic said Opus 5 was the most robust model evaluated under the benchmark’s test conditions.

Claude Opus 5 Cuts Prompt Injection Success to 2%

Indirect prompt injection attacks from the Gray Swan IPI benchmark (Q1 2026) (source : anthropic)

The gap with non-Claude systems was wider. Muse Spark, the strongest non-Claude model in the test, had a 16.5% success rate within 15 attempts.

That is more than eight times the rate reported for Opus 5. GPT 5.6 Sol, described as the most capable variant, reached 20.0%, close to GPT 5.5’s 20.8%.

The Claude Opus 5 System Card reports that Sol was ten times more likely to be successfully attacked than Opus 5. GPT-5.6 Terra and Luna had higher attack success rates of 30.4% and 43.9%, respectively.

The single-attempt comparison is also notable. A single attack attempt against GPT 5.6 Sol succeeded 3.1% of the time. That figure exceeds the 2.0% success rate achieved against Opus 5, which was reached only after 15 attempts.

The comparison suggests that Opus 5’s protections can resist repeated adversarial prompting more effectively. However, benchmark scores should not be treated as a complete measure of real-world security.

For security teams, the results reinforce the need for layered defenses around AI deployments. Organizations should continue to separate trusted instructions from untrusted data, restrict tool permissions, require confirmation for sensitive actions, and monitor agent activity.

Benchmark leadership is useful, but it does not make prompt injection impossible. Attackers can vary payloads, exploit workflow weaknesses, and target integrations rather than the model alone.

The key operational question is whether an AI system can fail safely when it encounters hostile content. Opus 5’s result indicates meaningful progress, while leaving enterprises responsible for secure architecture, testing, and incident response.

Security leaders should also conduct red-team exercises that simulate malicious files, retrieved web content, and compromised third-party data sources before granting agents broad production access in real workflows.

 Strengthen Your SOC by Accelerating Threat Detection & Rapid Investigations. -> Integrate ANY.RUN With Your SOC Now.