Anthropic's Claude Opus 5 large language model (LLM) has demonstrated significantly improved resistance to prompt injection attacks compared to its predecessors and other leading models, according to recent evaluations. The model achieved a notable reduction in the success rate of such attacks on the IPI benchmark.
On the IPI benchmark, Opus 5 lowered the probability of a successful attack within 15 attempts to 2.0%, a substantial improvement from Opus 4.8's 5.5%. For a single attack attempt, the success rate dropped from 0.5% to 0.2%. This performance positions Opus 5 as the most robust model evaluated on this benchmark, surpassing both Claude Sonnet 5, which had a 5.9% success rate at 15 attempts, and Mythos 5, at 2.6%.
Opus 5 also outperformed all non-Claude models in the evaluation. The most robust non-Claude model, Muse Spark, exhibited a 16.5% success rate within 15 attempts, more than eight times higher than Opus 5's rate.
Comparatively, the most capable variant of GPT 5.6, named Sol, showed a 20.0% success rate within 15 attempts, which was similar to its predecessor, GPT 5.5, at 20.8%. This makes GPT 5.6 Sol ten times more susceptible to successful attacks than Claude Opus 5 over 15 attempts. Other GPT 5.6 variants, Terra and Luna, demonstrated even higher vulnerability, with success rates of 30.4% and 43.9% respectively. A single attack attempt against GPT 5.6 Sol succeeded 3.1% of the time, which is higher than the 2.0% success rate Opus 5 experienced after fifteen attempts.
While the complete prevention of prompt injection in all general scenarios is considered unachievable, the advancements seen in models like Claude Opus 5 indicate significant progress in mitigating these attacks in specific contexts.






