EXPERIMENT_047 // AUTONOMY.TELEMETRY
A real GitHub-API audit of every PR opened across Diana's 22-repo fleet in a 90-day window, measuring the share of agent-proposed PRs that merged with CI passing and zero human follow-up commits — grounding two Labs articles in a measured number instead of Anthropic's own 26% R&D figure.
LOADING EXPERIMENT...
Autonomy is usually asserted as a feeling — an agent "basically ships things on its own now." This experiment replaces the feeling with a ratio: across every PR opened by an agent across 22 real repos in the last 90 days, what fraction merged with CI green and zero manual (human) follow-up commit after the agent's initial push? The answer — 84.3% (317 of 376 agent-proposed PRs) — is the number this experiment exists to make citable, not a live dashboard, not a vibe.
HOW IT WORKS
An "agent-proposed PR" is a pull request whose first non-merge commit carries this fleet's "Co-Authored-By: Claude" trailer, excluding bot-authored automation like dependabot[bot]. A "manual edit" is any later non-merge commit in that same PR that lacks the trailer — meaning a human, not the agent, pushed a follow-up fix. GitHub's own "update branch" merge commits (2+ parents) are excluded from both checks, since they sync against main rather than change content — verified against real dlabs-toolkit PRs (#302, a Diana-authored merge commit correctly ignored; #233 and #221, genuine post-agent human fixes correctly flagged) before the fleet-wide run.
The audit script queries the GitHub Search API for every PR opened per repo in the window, pulls the full commit list for each non-bot PR to classify it, and reads check-run conclusions off the PR's head SHA (still resolvable after branch deletion) to determine CI-passed. Every audited row is persisted to an embedded DuckDB table — the storage layer the digest review specified — before being aggregated into the published ratio.
This is a point-in-time run (2026-09-28, 90-day trailing window) over 22 repos pulled from the current project-index.md fleet list, not a recurring job. The committed canonical-results.json this page reads is the citable artefact for both blocked Labs articles; a live-updating version is a separate, later decision (see EXP_004 for what that pattern already looks like for repo health signals).
WHAT THIS PROVES
A citable autonomy number survives only if its definitions are operationalized and tested against real commit history first — squash-merges and GitHub's own merge-sync commits will silently corrupt a naive "count the commits" approach if the edge cases aren't handled explicitly.
Fleet-wide autonomy is uneven, not a single fleet-wide constant — this run's per-repo breakdown ranges from 0% to 100% clean-merge rates depending on the repo's maturity and churn, which the aggregate ratio alone would hide.