LIVE · cybersecurity feed
Live wire
security

Anthropic’s Opus 5 Is Better at Resisting Prompt Injection

The chart is interesting. On the IPI benchmark, Opus 5 improved over Opus 4.8, reducing the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0%, and from 0.5% to 0.2% on 1 attempt. It also improved on Sonnet 5 (5.9% at k=15) and Mythos 5 (2.6%), making it the most robust model evaluated. Opus 5 also outperformed all non-Claude models on this benchmark. The most robust non-Cl

zeroday.news · 1d ago

Anthropic's Claude Opus 5 large language model (LLM) has demonstrated significantly improved resistance to prompt injection attacks compared to its predecessors and other leading models, according to recent evaluations. The model achieved a notable reduction in the success rate of such attacks on the IPI benchmark.

On the IPI benchmark, Opus 5 lowered the probability of a successful attack within 15 attempts to 2.0%, a substantial improvement from Opus 4.8's 5.5%. For a single attack attempt, the success rate dropped from 0.5% to 0.2%. This performance positions Opus 5 as the most robust model evaluated on this benchmark, surpassing both Claude Sonnet 5, which had a 5.9% success rate at 15 attempts, and Mythos 5, at 2.6%.

Opus 5 also outperformed all non-Claude models in the evaluation. The most robust non-Claude model, Muse Spark, exhibited a 16.5% success rate within 15 attempts, more than eight times higher than Opus 5's rate.

Comparatively, the most capable variant of GPT 5.6, named Sol, showed a 20.0% success rate within 15 attempts, which was similar to its predecessor, GPT 5.5, at 20.8%. This makes GPT 5.6 Sol ten times more susceptible to successful attacks than Claude Opus 5 over 15 attempts. Other GPT 5.6 variants, Terra and Luna, demonstrated even higher vulnerability, with success rates of 30.4% and 43.9% respectively. A single attack attempt against GPT 5.6 Sol succeeded 3.1% of the time, which is higher than the 2.0% success rate Opus 5 experienced after fifteen attempts.

While the complete prevention of prompt injection in all general scenarios is considered unachievable, the advancements seen in models like Claude Opus 5 indicate significant progress in mitigating these attacks in specific contexts.

ShareXLinkedInWhatsAppFacebook

More News

view all →
vulnerability

Coldcard Hardware Wallet Flaw Linked to $70 Million Bitcoin Theft in 41 Minutes

An attacker drained 1,196 Bitcoin addresses in 41 minutes on July 30, taking 1,082.65 BTC worth about $70.2 million at the time. Galaxy Research mapped the sweep and tied it to a firmware flaw in Coldcard, the Bitcoin-only hardware wallet made by Canadian firm Coinkite. A March 2021 firmware integration error routed seed generation to a deterministic software pseudorandom number generator (PRNG

vulnerabilitycritical

Rails patches critical Active Storage flaw with RCE potential

A critical vulnerability in the Active Storage framework can allow an unauthenticated attacker to read arbitrary files from a Rails application, and potentially escalate to remote code execution (RCE). [...]

malware

Russian Hackers Hijack Hotel Wi-Fi to Steal Microsoft 365 Tokens

Microsoft says Russian hackers hijacked hotel Wi-Fi portals to spread malware and steal Microsoft 365 tokens from travelers. Microsoft Threat Intelligence disclosed CaptiveCrunch, a campaign it attributes to Storm-2945, an operational sub-cluster of Midnight Blizzard, the Russian SVR-linked group also known as APT29 and Cozy Bear. Since early May 2026, Storm-2945 has been manipulating DNS […]

CVE-2026-48449critical

Adobe fixed a maximum-severity vulnerability flaw in Campaign Classic

Adobe fixed a maximum severity vulnerability in Campaign Classic that could let attackers run code remotely without user interaction. Adobe has addressed a critical vulnerability, tracked as CVE-2026-48449 (CVSS score of 10.0), in Adobe Campaign Classic, the company’s enterprise marketing automation platform. The flaw is caused by incorrect authorization and could allow attackers to execute […]

security

Balance Theory Raises $19 Million to Help Enterprises Manage Cybersecurity Investments

The funding round was led by SYN Ventures, with participation from existing investors DataTribe and TEDCO. The post Balance Theory Raises $19 Million to Help Enterprises Manage Cybersecurity Investments appeared first on SecurityWeek.

vulnerabilitycritical

Ruby on Rails Patches Critical Vulnerability

The flaw can be exploited by unauthenticated attackers to read arbitrary files and potentially achieve remote code execution (RCE). The post Ruby on Rails Patches Critical Vulnerability appeared first on SecurityWeek.