Anthropic has published measures of AI’s role in its own development work, including how internal agent actions are monitored.

Its oversight measures include coverage, review latency and escalation rate. These are the company’s own measurements.

Coverage tells you which actions a monitor examines. It does not tell you how often the monitor recognizes a consequential mistake.

Latency matters because review after an action can explain a loss without preventing it.

Escalations need context, too: a quiet monitor might be encountering few problems, or missing them.

The useful record connects the action, the monitor’s decision and the eventual outcome. That connection is what makes oversight assessable, rather than merely countable.

Primary source
Anthropic Institute · Measuring the pace of AI development ↗

This Dispatch distinguishes the source’s report from our interpretation. Our editorial method.