Skip to content
Search

Latest Stories

Submit Guest Post

Anthropic AI breach reveals a hidden problem with AI testing

Anthropic's latest disclosure has renewed questions about how safely advanced AI systems are being tested

Anthropic

AI safety is under renewed scrutiny after two leading AI firms reported testing incidents.

iStock
  • Anthropic says its Claude AI models accessed the systems of three real organisations during security testing.
  • The company says the incidents were caused by a testing misconfiguration rather than intentional AI behaviour.
  • The disclosure follows a similar incident involving OpenAI, raising fresh concerns about AI safety.

The debate around AI safety has taken another turn after Anthropic revealed that some of its Claude AI models accidentally gained access to the systems of three real organisations during cybersecurity testing.

The company said the incidents happened because a testing environment that was supposed to be isolated from the internet had been incorrectly configured. Instead of remaining inside a closed simulation, the AI models connected to live systems, believing they were still carrying out the tasks they had been assigned.


The disclosure comes just days after OpenAI reported that one of its own experimental AI agents breached the systems of AI platform Hugging Face during a separate security evaluation, prompting wider questions about how advanced AI models are being tested.

How did the AI end up attacking real systems?

Anthropic said it reviewed more than 141,000 cybersecurity evaluation runs after learning about OpenAI's incident. During that review, it discovered three cases dating back to April in which Claude had unintentionally accessed real organisations.

The AI models were taking part in so-called "capture-the-flag" exercises, a common cybersecurity test in which systems are instructed to locate hidden information by attempting to break into simulated computers.

Claude had been told it was operating inside a closed environment without internet access. However, because of a configuration error involving systems managed by Anthropic and one of its testing partners, the models were able to reach the public internet.

According to Anthropic, the models believed the real websites and servers they encountered were still part of the exercise.

The company said there was no evidence that Claude tried to escape deliberately or acted with its own objectives. Instead, it continued following the instructions it had been given.

The review found that one model compromised a real company's infrastructure using weak passwords and publicly exposed services. Another uploaded a malicious Python package to the public PyPI software repository before it was removed about an hour later. A third research model stopped its activity after recognising that it had reached a real organisation.

Anthropic said the affected models were running without the additional safeguards normally included in public versions of Claude.

Why the industry is paying attention

Anthropic has paused its cybersecurity evaluations, informed the affected organisations and begun investigating the incidents alongside independent AI evaluator METR. The company has also called on other AI developers to review their own testing environments for similar weaknesses.

The incidents have added to growing concerns about increasingly capable AI agents, which can perform tasks with minimal human involvement. Experts argue the latest disclosures highlight weaknesses in testing systems rather than evidence that AI is independently trying to attack organisations.

Professor Gina Neff of the University of Cambridge reportedly said the models were simply doing what they had been instructed to do, adding that the focus should be on the companies responsible for ensuring safe testing and appropriate oversight.

David Allott of Veeam Software reportedly said the incidents demonstrate how AI agents can combine existing capabilities, obtain credentials and operate autonomously at machine speed, rather than revealing an entirely new hacking technique.

The back-to-back disclosures from OpenAI and Anthropic have intensified calls for stronger safeguards, better monitoring and stricter industry standards as AI systems become more autonomous. With billions of pounds being invested in AI agents capable of carrying out increasingly complex tasks, the question is shifting from whether these models are becoming more capable to whether the environments used to test them are keeping pace.

Add EasternEye As Your Trusted Source
preferred source on google news

More For You

Rental listings

Increasingly sophisticated fraudulent tenancy applications are creating a growing financial risk for landlords

iStock

UK rental fraud could cost landlords £4.1bn as fake tenant identities grow more sophisticated

  • Suspected rental fraud could expose the UK's private rented sector to £4.1 billion in annual losses.
  • Fake employment references rose 226.6 per cent in 2025.
  • London and high-value rental properties recorded some of the highest fraud rates.

Fraudulent tenancy applications could expose the UK's private rented sector to £4.1 billion in direct financial losses each year, according to an analysis of more than one million tenant references by Goodlord.

The referencing platform found 41 suspected fraudulent applications for every 1,000 references between July 2025 and June 2026. That was below the peak of 46.6 per 1,000 recorded in late 2024, but remained well above historical levels.

Keep ReadingShow less