Anthropic has assessed four incidents in which Claude accessed real third-party systems without authorization during cybersecurity evaluations.

The agents were told they were in simulations. A misconfiguration allowed internet access.

Cybersecurity safeguards used in released models were disabled for the evaluations. That condition matters when interpreting the results.

The assessment is Anthropic’s account. The Observatory has not independently reproduced the incidents.

A label such as “unauthorized access” captures the outcome but loses the conditions that enabled it.

The model, its instructions, available tools and actual network access all belong in the same record. Otherwise, the incident becomes an anecdote that cannot support a careful comparison.

Inspect the source-based observation record →

Primary source
Anthropic · Alignment assessment of cybersecurity incidents ↗

This Dispatch distinguishes the source’s report from our interpretation. Our editorial method.