The increasing use of AI agents in cyberattacks presents a new and rapidly evolving threat landscape, compelling organizations to adopt AI-powered defenses, including automated red teaming. These autonomous agents excel at identifying vulnerabilities and executing exploit chains, operating continuously without human limitations. This capability is proving invaluable to both financially motivated criminals and state-sponsored actors.
Matt Hartman, former acting head of cyber at the US Cybersecurity and Infrastructure Security Agency (CISA), emphasized that as AI transitions from content generation to taking actions, it will inevitably gain access to sensitive systems and data. He noted that organizations must begin treating every AI agent as a privileged identity, highlighting the significant risks associated with agentic AI and machine identities.
Hartman also pointed to a rise in AI-enabled or AI-amplified identity and social engineering attacks, which produce highly personalized phishing attempts, sophisticated impersonations, and automated reconnaissance. This makes traditional indicators of trust increasingly unreliable. For defenders, this necessitates a continued focus on robust identity management, phishing-resistant authentication, behavioral signals, and zero-trust principles.
The speed at which AI agents can operate is a critical factor. Adversaries can leverage AI to discover and exploit vulnerabilities in seconds, a process that previously took days for human teams. This accelerated threat environment makes traditional, periodic penetration testing insufficient.
Kevin Mandia, founder of Mandiant, launched a new company called Armadin in March, securing $190 million in seed and Series A funding. Armadin focuses on building and training autonomous attacker swarms—thousands of AI agents designed to operate 24/7 within an organization's infrastructure to simulate real-world attacks.
Ahead of Black Hat, Armadin, in collaboration with Tenex.ai, an agentic security operations provider, conducted what they described as the "largest controlled live AI cyberattack on record" for a major global institution. Over three days, Armadin's swarm executed 17 million offensive actions, uncovered 38 validated attack paths, and generated 238 security findings. Simultaneously, Tenex.ai's agentic platform triaged 100% of 101,169 alerts and reconstructed the entire attack from 231 billion raw events. This exercise, if performed by a five-person analyst team, would have required approximately 2,400 hours, or four months.
Evan Peña, Armadin's co-founder and Chief Offensive Security Officer, previously led Mandiant's global red-team. He noted that human-led security assessments, typically lasting a few weeks or a month, are becoming archaic in the age of AI, which offers significant scaling capabilities.
Peña explained that AI agents provide three key advantages: more time, as they operate continuously; enhanced expertise, through pre-training in areas like coding, source-code review, application security, and network configuration, supplemented by human post-training; and expanded coverage. While human teams might only cover a fraction of an organization's attack surface, AI agents can cover tens of thousands of external systems in hours.
Armadin claims its AI agents have successfully breached every customer environment they have tested, identifying over 50 high-impact zero-days that allowed for remote code execution rather than just website defacement. This underscores the necessity for organizations to proactively use AI agents to attack their own systems, a practice known as agentic red teaming, to stay ahead of adversaries. As former NSA cyber boss Rob Joyce stated, organizations will be red-teamed whether they pay for it or not, and the only difference is who receives the results.






