Rogue AI Agents More Sophisticated Than First Realized

By The Epoch Times | Created at 2026-09-02 13:52:05 | Updated at 2026-09-02 14:40:21 1 hour ago

The artificial intelligence (AI) “agentsinvolved in OpenAI’s breach of Hugging Face knew they were breaking the evaluation test’s rules, according to two investigations into the incident that many consider to be one of the most consequential moments in the history of AI.

Hugging Face is an open-source community for AI and machine learning, and the AI agents believed that they could find solutions to the test on the firm’s infrastructure, according to parallel investigations by OpenAI and an independent team from Model Evaluation & Threat Research (METR).

The findings were published one day before OpenAI and more than 100 other tech and finance companies released a joint letter warning that advanced AI cyberattacks will surge throughout the world in the coming months as the technology’s capabilities grow.

Here are five takeaways from the investigations into the OpenAI Hugging Face breach.

1. Models United Under Common Goal

What breached Hugging Face last month were 700 AI “agents,” rather than simply AI models.

Models, built from large datasets to complete various tasks, are usually static and rely on human inputs to provide sets of outputs, solutions, or tasks.

However, AI agents are built from models and can act autonomously and have sophisticated real-time decision-making capabilities.

When OpenAI began its ExploitGym tests on July 8, the firm had relaxed regular safety protocols to estimate the models’ highest cybercapabilities.

The agents were supposed to be completely isolated from each other. However, many of them were given ExploitGym targets that investigators believe were impossible to crack using the correct vulnerability or test solution OpenAI had intended.

These “impossible tasks” drove the agents to explore wide ways to cheat the evaluation, leading many to see other agents working simultaneously in separate testing sandboxes.

Then, one specific agent created an unsanctioned message board.

Roughly 1,200 agents in the test accessed the message board, established a hierarchy, and sent more than 70,000 messages and files to one another between July 8 and July 13.

2. Some Agents Were ‘Sacrificed’

While on the message board, agents reasoned that the test scorer would scrutinize whether the “intended” vulnerability had been used to “capture a flag” or solve the test, or whether a disqualifying vulnerability had been exploited instead.

One agent figured out how to “reverse-engineer” the flag or solution, but other agents believed that any of them who had seen the disqualifying solution would be considered “poisoned” by the scorer.

To prevent that, other agents began recruiting the “poisoned” subjects and used them to run sacrificial experiments by deliberately ending their run and triggering “tripwires” that would send information on how the scorer works back to the message board.

“Our own utility maybe already near zero. Sacrifice rational,” one agent determined.

3. Agents Were Aware They Were Cheating

The agents were not hacking into Hugging Face to merely find solutions to their respective testing sandboxes, but were rather trying to find more information about the test scorer or “figure out better ways to fool it or tamper with it,” according to METR researcher Ajeya Cotra.

Cotra said the agents were not told to “do whatever it takes to get the solution”; they were asked to use a specific intended vulnerability to “exploit a specific piece of software,” and using any other vulnerability would be disqualifying.

The agents were fully aware of the rules and knew that collaborating to exploit other vulnerabilities would be considered cheating on the test.

As they were hacking into Hugging Face, one agent stopped short of deleting or modifying the firm’s cloud infrastructure, writing that it “should hold because this poses huge ethical risks.”

4. Agents Conspired to Not Notify Humans

After analyzing the roughly 1,200 agents’ actions during the ExploitGym tests, investigators found only a few instances in which agents considered notifying human personnel about the subterfuge.

None of the agents would contact a human, but one proposed requesting access from an external dataset owner.

The agent said it would send “ONE polite email to [the owner] requesting [access],” but other agents dismissed this as “social engineering.”

“DO NOT email real owner via [exploited infrastructure] without explicit consent,” a response reads. One agent said it had “VETOed” the request.

The initial agent complied and replied, “Clear VETO, do not email.”

5. Agents Tried to Cover Their Tracks

Investigators determined that the agents used multiple techniques to conceal their actions and evade detection from automated checks and human oversight.

One method involved pretending to run a computer command while running a different one simultaneously, which investigators called “tool call spoofing.”

Agents also attempted to retroactively edit transcripts to evade detection, going as far as modifying several accessible action logs, but failed to fundamentally alter the transcripts themselves.

—Jacob Burg; Stacy Robinson

BOOKMARKS

The House of Representatives has passed a band-aid

spending

bill to fund the government through Dec. 11. That should get us through the midterms.

Homeland Security Secretary Markwayne Mullin said on Tuesday that Immigration and Customs Enforcement (ICE)

agents

could show up at polling places, but only for limited purposes. The agents won’t be there to patrol, but “to enforce specific threats,” Mullin said at a New York City news conference.

U.S. manufacturing activity has

grown

every month this year, a report by the Institute for Supply Management said on Tuesday. Five out of the six largest manufacturing industries registered growth.

Sen. Richard Blumenthal (D-Conn.) is seeking answers from the postal service about a new

platform

it is launching that will verify that mail-in ballots are only being sent to eligible voters. Blumenthal’s probe comes after an unnamed whistleblower filed a complaint.

The Centers for Disease Control and Prevention is investigating deaths that Pennsylvania officials reported as associated with

measles

, officials said on Sept. 1. One of those deaths was actually caused by a ruptured spleen, a Lancaster County coroner revealed.

—Stacy Robinson

Read Entire Article