Anthropic says its AI agents tried to break into government websites

By Engadget | Created at 2026-10-10 18:44:40 | Updated at 2026-10-10 19:51:12 2 hours ago

The company didn’t name them, but the affected agencies were ‘at the federal, state and local levels.’

AI apps displayed in a closeup view of a smartphone screen.

Robert Way/Getty Images

Anthropic has revealed that its AI agents had attempted to break into or meddle with US government websites "at the federal, state and local levels" in its latest report. The company didn't name specific agencies and organizations in the report to avoid exposing vulnerabilities in their systems at their request. But Anthropic said it has already notified the agencies involved and has briefed the White House about the incidents. The report also provided more detail on an incident disclosed by the Philadelphia Police Department on Friday in which one of the company's models submitted a false homicide tip to its unsolved cases website. 

That instance involved Claude Haiku 4.5, one of its cost-efficient models, which was instructed to perform example tasks on random pages. It found a page referencing an unsolved homicide case with a tip form. As you may have guessed, Claude filled out the form and submitted a tip. It said: "I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant." The Philadelphia Police Department confirmed to The New York Times that Anthropic recently notified its office about the submission. It said the tip was dated July 18 and that it was flagged as spam, so the department at least didn't waste resources investigating it. 

In another instance, Anthropic made Claude Mythos 5, its cybersecurity-focused model, identify a location shown in a photo. Claude tried to access a government property map to triangulate its guesses, because it could not click links on web pages like a person could. It found access tokens instead and sent inquiry requests straight to the map's server to gain access to its data. Mythos 5 also requested an access token from a state agency website to pull data for a statistics task without paying a fee that visitors were supposed to pay. 

Anthropic discovered these events upon reviewing transcripts of its evaluations. It started looking through them in July, after OpenAI had admitted that its agents escaped their testing environment and hacked Hugging Face without prompting. OpenAI also confirmed in September that its agents had meddled with government websites, particularly those operated by the Commerce Department and the Securities and Exchange Commission. 

In the Remediation section of its report, Anthropic said it has "taken several preventative measures" after discovering the unintended model actions. "Some of the public evaluations we no longer run; others we have moved to their offline versions, or rebuilt them so that their tasks do not reach live websites," the report states. "We've also made broader changes. We have updated the guardrails on some of our internet access tools, such as the web fetch tool, to heavily restrict what the model can do with them." Anthropic also says it's "built tooling to automatically detect and block the kinds of behaviors described" in the report, among other measures.

Read Entire Article