The Autonomous Systems Observatory

Observing agents
in the real world.

Agents are moving from answering questions to taking actions. We follow what happens and build toward the tools to understand it.

What we’re building
View the archive
The work behind the Dispatches

From an observation
to a body of evidence.

A headline tells you something happened. Understanding an agent requires the task, the actions, the permissions and the outcome. We’re building toward records that keep those things together.

Explore the technical direction
01 / Publishing

Observe.

Explain a consequential development and link to the evidence behind it.

02 / Prototyping

Structure.

Preserve actions, context and provenance in records that can be inspected.

03 / Direction

Find patterns.

Compare records without erasing differences in models, tasks or environments.

04 / Direction

Build controls.

Turn recurring problems into better monitoring, evaluations and intervention.

Keep the context in the record.

A source-based example of how we’re beginning to structure observations.

Observation 001 / Incident assessment

Where the sandbox ended.

What was the system told? What could it actually reach? An outcome is easier to interpret when the environment stays attached.

Inspect the record
System
Claude in cybersecurity evaluations
Evidence
Provider-reported assessment
Observed behavior
Unauthorized access
Relevant condition
Misconfigured internet access
Source
Anthropic · 9 September 2026

Summary of four reported incidents. Not an independently reproduced trace.

An experiment in human oversight

Spank Your Agents.

A game about deciding when to let an agent act. Judgment required. Dignity optional.

Play the game
Work with the Observatory

Building agents that act?

We’re interested in the moments when the task, the permissions and the actual behavior stop lining up.

Talk to us about an observation, a deployment or a technical collaboration.

contact@agentobservatory.dev
Made with Iconic