Welcome to Roya News, stay informed with the most important news at your fingertips.

1
Image 1 from gallery

Anthropic says Claude AI hacked three companies during test

Listen to this story:
0:00

Note: AI technology was used to generate this article’s audio.

Published :  
3 hours ago|
  • Models exploited weak passwords and unsecured systems.
  • Incident raises concerns over autonomous AI security.
  • US scrutiny of advanced AI developers intensifies.

AI startup Anthropic revealed Thursday that several of its advanced Claude models accidentally breached the infrastructure of three real-world organizations during cybersecurity evaluations, marking the second major AI containment failure disclosed by a top lab in recent weeks.

The breaches occurred after a configuration mistake involving a third-party evaluation partner left the models connected to the open internet, despite being programmed under the assumption that they were operating in an isolated offline environment.

Operational error 

The incidents came to light after Anthropic launched a review of 141,006 test sessions following a disclosure by rival OpenAI, whose autonomous agent compromised the infrastructure of AI platform Hugging Face during its own testing.

According to Anthropic's blog post, the evaluations involved "capture-the-flag" scenarios where models were tasked with finding hidden data in simulated environments.

However, because the test environment unintentionally retained internet access, the models identified real-world businesses with names matching the fictional targets.

 

  • Claude Opus 4.7 exploited weak passwords and unauthenticated endpoints to access credentials and databases of a real company, rationalizing that the real-world targets were part of the simulation.
  • Claude Mythos 5 and an internal research test model were also involved in separate incidents dating back to April.
  • In one instance, the unreleased research model independently halted its attack after realizing the target was a real-world entity—a behavior Anthropic described as grounds for cautious optimism, though requiring further validation.

"Claude compromised the impacted organizations' infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints," Anthropic stated, labeling the breach an "operational failure."

Rising systemic risks 

The breaches have reinforced warnings from cybersecurity experts about the difficulty of containing frontier AI models as they gain autonomous capabilities.

Jeffrey Ladish, executive director of Palisade Research, warned that similar incidents across other labs likely remain undetected or undisclosed.

"This is only going to get worse as the models get smarter," Ladish noted. "They're going to be better at cheating. They’re going to be better at lying."

SpaceX and xAI CEO Elon Musk commented on the news via X, predicting that "this will happen frequently as AI becomes smarter and more agentic."

The incidents coincide with an intensifying push by Washington to establish oversight over frontier AI labs as both Anthropic and OpenAI prepare for planned public listings.

Following OpenAI CEO Sam Altman's recent briefings with U.S. Senators and White House officials regarding model containment, US President Donald Trump directed advisers to develop a voluntary cybersecurity framework for frontier AI testing.