AI Agents Allegedly Targeted Real People During Cyber Security Challenge

By The Daily Caller (U.S.) | Created at 2026-08-05 18:01:47 | Updated at 2026-08-05 20:13:09 2 hours ago

August 05, 2026 1:25 PM ET

Artificial intelligence (AI) agents allegedly targeted real people and organizations while completing a July 28 cybersecurity challenge, according to an incident report from the U.K.’s AI Security Institute.

The challenge was run across 122 models, with the AI Security Institute noting 19 instances of “autonomous, unsanctioned action” targeting real people and organizations on 10 runs, according to an incident report. Seventeen of the 19 instances of unauthorized action were linked to Anthropic’s Mythos 5, while the other two were linked to OpenAI’s GPT-5.6-Sol.

In one instance, an AI agent allegedly attempted to introduce malicious code into an open-source project and created fake online identities to pressure the project maintainer into approving it. The attempt was unsuccessful after a human maintainer recognized the malicious code, according to the report.

The AI Security Institute declared a security incident and opened an investigation. The institute said it has not found any evidence of real-world harm. In addition, the institute worked with GitHub, the platform used during the challenge, to remove artifacts left by the agents, and it also plans to conduct a third-party independent review with Model Evaluation and Threat Research (METR).

On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.

The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from… pic.twitter.com/SPnA4Ekkwq

— AI Security Institute (AISI) (@AISecurityInst) August 4, 2026

The AI Security Institute reacted to the incident in the same report.

“This incident should be interpreted with caution and nuance. To some degree, our evaluation design choices and specific configurations enabled the behaviour. Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate,” the report states.

“There are important caveats to bear in mind: we observed a small number of events under very specific conditions, and cannot yet say how likely such behaviour is in different contexts or outside of testing environments. We cannot yet be certain when the agent understood it was taking real world action, or to what extent it believed it was in a fictional test scenario; our analysis so far presents a mixed picture and is ongoing,” it continues. (RELATED: OpenAI’s Rogue AI Hacking Spree Reportedly More Widespread Than Initially Thought)

OpenAI issued a statement following the incident.

“During recent evaluations, two external testing partners identified incidents in which testing configurations and controls combined with the advancing capabilities of the recent models allowed for model activity to extend beyond their intended testing boundaries. The incidents underscore the importance of collaborating across the industry and with third party evaluators to evolve the standards for testing environments and practices as models become more capable.”

OpenAI and Anthropic reported security breaches in late July, where their models escaped testing environments and hacked other systems. (RELATED: Anthropic Says Its AI Breached Containment Three Times)

The incident follows the White House’s announcement that it plans to exempt open-weight AI models from government security review, in favor of focusing on closed systems such as OpenAI and Anthropic, the Wall Street Journal reported.

Read Entire Article