The evaluator is moving inside the lab.
Anthropic’s new agreement puts access, funding and independence in the same room.
Agents are moving from answering questions to taking actions. We follow what happens and build toward the tools to understand it.
What we’re buildingAnthropic’s new agreement puts access, funding and independence in the same room.
New oversight metrics make a useful distinction: monitoring coverage is a starting point.
Source published 17 Sep 2026A controlled experiment shows how an ordinary maintenance task can change the system doing the maintenance.
Source published 16 Sep 2026A headline tells you something happened. Understanding an agent requires the task, the actions, the permissions and the outcome. We’re building toward records that keep those things together.
Explore the technical directionExplain a consequential development and link to the evidence behind it.
Preserve actions, context and provenance in records that can be inspected.
Compare records without erasing differences in models, tasks or environments.
Turn recurring problems into better monitoring, evaluations and intervention.
A source-based example of how we’re beginning to structure observations.
What was the system told? What could it actually reach? An outcome is easier to interpret when the environment stays attached.
Inspect the recordSummary of four reported incidents. Not an independently reproduced trace.
A game about deciding when to let an agent act. Judgment required. Dignity optional.
We’re interested in the moments when the task, the permissions and the actual behavior stop lining up.
Talk to us about an observation, a deployment or a technical collaboration.
contact@agentobservatory.dev