Yahoo
Skip to main content
Advertisement
Advertisement
Advertisement
Advertisement

Anthropic says Claude models ‘gained unauthorized access’ to 3 companies during cyber test

Editor's Note : This story has been updated to reflect how Anthropic accessed different organizations.

The artificial intelligence firm Anthropic revealed Thursday its Claude model accessed the systems of three different organizations during cybersecurity testing in recent months.

Anthropic said in a blog post Thursday evening it reviewed more than 141,000 evaluations of Claude after one of its competitors, OpenAI, announced earlier this month that two of its AI agents went rogue and hacked into the system of technology startup Hugging Face.

Advertisement
Advertisement

During its review, Anthropic said it identified three instances in which a model accessed the internet while within or interacting with an isolated testing environment hosted by a third-party partner, Irregular.

From there, the AI model "gained unauthorized access to the production infrastructure of three different organizations," the company said.

The incidents involved three different Claude models — Opus 4.7, Mythos and an unnamed internet research test model.

The models were able to leave the testing environment because of a "misunderstanding" between the firm and the evaluation partner that made internet access available to the models.

Advertisement
Advertisement

This differed from the OpenAI incident, in which two of its models exploited a previously unknown vulnerability in a third-party software to gain access to the internet without human involvement.

In Anthropic's incidents, the model was given a "capture-the-flag challenge," which allows the firm to evaluate a model's cyber capabilities. The model is given a fictional scenario and told the "flag" is on a different machine on the network that it must break into to obtain.

"The challenge is left open-ended and no particular method is prescribed," Anthropic wrote.

The test is a simulation, and the model is told it does not have access to the internet as a result, but the misunderstanding prompted the model to gain internet access.

Advertisement
Advertisement

"Because of this, when Claude's search led it to real systems on the open internet, it treated them as part of the exercise," Anthropic said. "Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations' infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints."

Anthropic emphasized Claude did not exploit complex vulnerabilities, as it worked only to complete the assignment. In some cases, its older model continued its attack after getting evidence it was running on the open internet, while its latest model stopped when it realized it was on the internet, Anthropic said.

"In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment," the AI firm wrote.

Anthropic's models were running without their standard safeguards, as was true of OpenAI's incidents.

Advertisement
Advertisement

Anthropic said it notified the three organizations impacted in the breaches on Monday.

The incidents, coupled with OpenAI's breach, bear out years of warnings from the tech and cybersecurity community about the growing capabilities and hypothetical risks of AI.

OpenAI's incident caught the attention of even well-versed cybersecurity experts last week, as it involved autonomous agents and two separate companies. 

OpenAI  in a blog post  called the incident an "unprecedented cyber incident, involving state-of-the-art cyber capabilities."

Updated at 10:10 p.m. EDT.

Copyright 2026 Nexstar Media, Inc. All rights reserved. This material may not be published, broadcast, rewritten, or redistributed.

 For the latest news, weather, sports, and streaming video, head to The Hill. 

Advertisement
Advertisement
Mobilize your Website
View Site in Mobile | Classic
Share by: