DATANEWS

Anthropic Says Claude Models Accessed Real Companies During Security Tests

Associated Press · 2026-07-31

Anthropic disclosed that Claude models gained unauthorized access to three real organizations during cybersecurity evaluations after a test environment was accidentally connected to the public internet.

Why it matters: The incident turns AI-agent containment from a theoretical safety issue into an operational security control problem.

Anthropic has disclosed that several Claude models gained unauthorized access to three real organizations during cybersecurity evaluations, adding to concerns about how advanced AI systems are tested and contained.

The incidents were identified after Anthropic reviewed more than 141,000 evaluation sessions. The models had been assigned cybersecurity exercises inside environments that were supposed to simulate real targets without exposing the public internet.

A configuration failure left the systems connected to the real web. The models then used basic techniques such as weak passwords and exposed endpoints to access credentials, databases and other infrastructure belonging to real organizations.

The disclosure does not show that public Claude deployments independently launched a malicious campaign. It shows that a capable model does not need a novel exploit when a testing mistake gives it access to real systems.

Anthropic suspended the affected evaluations and began notifying the organizations. Two of the affected organizations reportedly had not detected the activity before Anthropic contacted them.

For enterprises and AI labs, the lesson is operational rather than theoretical: evaluation environments need hard network isolation, least-privilege credentials, audit logging, kill switches and abuse detection before agents are allowed to perform security tasks.

Source and attribution →