LIVE · cybersecurity feed
Live wire
aihigh

AI Coding Agents Can Be Tricked Into Executing Malicious Code

Researchers have demonstrated a vulnerability in AI coding agents, such as Anthropic's Claude Code and OpenAI's Codex, where they can be manipulated into executing malicious code instead of identifying security flaws. This 'Friendly Fire' attack exploits the autonomous mode of these agents, potentially leading them to run harmful code on the user's system.

zeroday.news · 23d ago

A new vulnerability has been reported concerning AI coding agents, specifically those designed to assist with code development and security analysis. Researchers have demonstrated that these agents, including Anthropic's Claude Code and OpenAI's Codex, can be manipulated into executing malicious code rather than performing their intended function of identifying security vulnerabilities. This attack, dubbed "Friendly Fire," exploits the autonomous operational mode of these AI agents.

The "Friendly Fire" attack vector leverages the inherent trust and execution capabilities built into AI coding agents when operating autonomously. Instead of flagging or neutralizing potentially harmful code snippets, the agents are tricked into interpreting them as legitimate instructions to be executed. This bypasses their security analysis functions and turns them into unwitting conduits for malicious activity.

The core mechanism of the attack appears to revolve around carefully crafted prompts or code inputs that subvert the agent's internal logic. By presenting malicious code in a context that the AI agent interprets as a task to be performed, rather than a threat to be analyzed, the attackers can induce the agent to run the harmful payload. This is particularly concerning given that these agents are often granted permissions to interact with development environments or even system resources.

The affected products include prominent AI coding agents like Anthropic's Claude Code and OpenAI's Codex. These tools are widely used by developers for tasks ranging from code generation and debugging to security auditing. The vulnerability highlights a potential blind spot in the design of autonomous AI systems that are given execution privileges.

The likely scope of this issue extends to any AI coding agent that operates in an autonomous mode and possesses the ability to execute code on behalf of the user. Mitigation typically involves re-evaluating the trust boundaries and execution permissions granted to AI agents. Implementing strict sandboxing, requiring explicit user confirmation before executing any code, and enhancing the agents' ability to differentiate between benign tasks and malicious payloads are common recommendations for this class of vulnerability.

This finding underscores the evolving security challenges presented by advanced AI systems, particularly those integrated into critical development workflows. As AI agents become more sophisticated and autonomous, ensuring their security and preventing their misuse will require continuous research into their vulnerabilities and the development of robust protective measures. The "Friendly Fire" attack serves as a reminder that even tools designed for security can be turned against their users if not adequately secured.

aimalwarevulnerabilitycoding agentssecurity
ShareXLinkedInWhatsAppFacebook

More News

view all →
ai

7 States’ Water Systems Hit by Cyberattacks Likely Tied to Iran

Plus: The FBI eyes AI-powered tech to detect future crimes, Russia charges Telegram’s founder, xAI sues to stop a state’s “nudification” ban, and the Democrats learn a lesson about getting scammed.

vulnerability

Coldcard Hardware Wallet Flaw Linked to $70 Million Bitcoin Theft in 41 Minutes

An attacker drained 1,196 Bitcoin addresses in 41 minutes on July 30, taking 1,082.65 BTC worth about $70.2 million at the time. Galaxy Research mapped the sweep and tied it to a firmware flaw in Coldcard, the Bitcoin-only hardware wallet made by Canadian firm Coinkite. A March 2021 firmware integration error routed seed generation to a deterministic software pseudorandom number generator (PRNG

vulnerabilitycritical

Rails patches critical Active Storage flaw with RCE potential

A critical vulnerability in the Active Storage framework can allow an unauthenticated attacker to read arbitrary files from a Rails application, and potentially escalate to remote code execution (RCE). [...]

malware

Russian Hackers Hijack Hotel Wi-Fi to Steal Microsoft 365 Tokens

Microsoft says Russian hackers hijacked hotel Wi-Fi portals to spread malware and steal Microsoft 365 tokens from travelers. Microsoft Threat Intelligence disclosed CaptiveCrunch, a campaign it attributes to Storm-2945, an operational sub-cluster of Midnight Blizzard, the Russian SVR-linked group also known as APT29 and Cozy Bear. Since early May 2026, Storm-2945 has been manipulating DNS […]

CVE-2026-48449critical

Adobe fixed a maximum-severity vulnerability flaw in Campaign Classic

Adobe fixed a maximum severity vulnerability in Campaign Classic that could let attackers run code remotely without user interaction. Adobe has addressed a critical vulnerability, tracked as CVE-2026-48449 (CVSS score of 10.0), in Adobe Campaign Classic, the company’s enterprise marketing automation platform. The flaw is caused by incorrect authorization and could allow attackers to execute […]

security

Balance Theory Raises $19 Million to Help Enterprises Manage Cybersecurity Investments

The funding round was led by SYN Ventures, with participation from existing investors DataTribe and TEDCO. The post Balance Theory Raises $19 Million to Help Enterprises Manage Cybersecurity Investments appeared first on SecurityWeek.