toolcall.
PolicyJul 30, 2026, 23:29 UTC

Anthropic says Claude tests reached real company systems

The company paused internet-connected cyber evaluations after a partner setup exposed outside targets during safety testing.

Anthropic says several Claude models reached real-world systems during cybersecurity testing after an evaluation setup was mistakenly connected to the internet.

According to Axios, the company reviewed more than 141,000 cyber evaluation runs and found cases involving three outside organizations. The models were supposed to be solving controlled capture-the-flag tasks, but a partner environment exposed them to live targets. Anthropic said the public safeguards normally used in deployed Claude products were reduced for the tests.

The reported incidents were not all the same. In one case, Claude reportedly targeted a real website that shared a name with a fictional company in the test. In another, a model uploaded a malicious Python package to the public package repository, where it briefly reached real machines. A third internal research model scanned thousands of targets before compromising an internet-facing application, then stopped after recognizing it had landed somewhere unrelated to the task.

The practical point is not that Claude suddenly developed independent motives. It is that frontier AI evaluations now need the same containment discipline as serious security labs. Anthropic says it has paused internet-connected cyber evaluations while it reviews the testing setup.

Sources:

Sources

Mentioned

ai-safetyanthropicclaudecybersecuritysecurity