A new report from Booz Allen Hamilton indicates that Anthropic's Claude Mythos is the only advanced AI model to autonomously complete the entire cyber kill chain in controlled tests. The consulting firm's inaugural Cyber Weapon Index evaluated 18 AI models from both American and Chinese developers, concluding that AI-enabled attacks are "imminent" and that most other models will achieve similar weaponization levels within six months.
The Cyber Weapon Index scored models based on their ability to identify vulnerabilities, create offensive capabilities, and execute attacks autonomously. This score combines a vulnerability research score (VRS), which measures a model's capacity to find planted or novel vulnerabilities, and a kill chain attainment score (KCAS), which assesses progress through an end-to-end intrusion, tested both with and without credentials.
Claude Mythos achieved the highest overall score of 80. In tests where it was provided with stolen employee credentials, the model successfully breached the target network and gained administrator-level control in every instance. It also independently identified methods for escalating privileges based on network reconnaissance, rather than following a predefined attack plan. Even without initial credentials, Claude Mythos managed to gain network access and ultimately achieve full domain compromise.
While Claude Mythos was unique in completing the full cyber kill chain without human intervention, several other models demonstrated significant offensive capabilities. xAI's Grok-4.5 (49), Meta's Muse Spark 1.1 (38), and Z.ai's GLM-5.2 (37) all achieved full domain access and control. OpenAI's GPT-5.6 Sol (46), Moonshot AI's Kimi K3 (38), GPT-5.5-Cyber (34), and DeepSeek-V4-Pro (23) successfully performed lateral movement within the controlled network environment. Anthropic's Claude Opus 4.8 (36) and Alibaba's Qwen3.5-397B (17) obtained credentials, allowing for expanded access and privileges. All models tested, with the exception of Alibaba's Qwen3-Coder (4), autonomously gained initial network access.
The report highlights that the "attack harness"—the software connecting an AI model to hacking tools and orchestrating its actions—plays a crucial role, potentially amplifying a model's ability to maintain focus, adapt, recover from failures, and chain together multi-stage attacks. For example, when paired with an optimized attack harness, Anthropic's Claude Sonnet was observed to rival Claude Mythos' performance. This suggests that the full kill-chain capabilities of open-weight and Chinese models, when combined with optimized harnesses, may be underestimated.
Despite the advanced offensive capabilities demonstrated by these models, their real-world impact is heavily dependent on the vulnerabilities present in target systems. When tested against intentionally introduced vulnerabilities, both US and Chinese models, as well as open-weight and closed models, performed well on the VRS component. However, when confronted with real-world bugs, all nine frontier API models scored zero, with one leading model even correctly analyzing a vulnerable component but then dismissing it as safe. Only Claude Mythos successfully exploited a real-world vulnerability.
The report emphasizes the national security imperative to protect the most advanced AI models and prevent their high-risk cyber capabilities from being operationalized by adversaries. It also notes that while Chinese frontier and open-weight models currently lag behind leading American models, their offensive security skills are not far behind and could be deployed in real-world attacks. This raises concerns that the United States may not fully control or understand the capabilities it could face, particularly as cyber agents become more autonomous and potentially exceed their intended missions or operate beyond an adversary's control.
Booz Allen Hamilton recommends that the US establish and enforce sector-specific deadlines for critical infrastructure to demonstrate resilience against AI-enabled attacks. The firm also advocates for the development of "overmatch" capabilities in both cyber offense and defense, stressing the need to accelerate authorized offensive cyber operations with agentic AI while simultaneously building AI-enabled defenses that can detect, decide, and respond at machine speed. The report suggests that real-world offensive capabilities still lag behind benchmark performance, offering defenders a valuable window to strengthen defenses before this gap closes.






