Skip to content

Agent-run measurement

Measure repository readiness with three similar-scope tasks from a fresh environment:

  1. change one public API or deployment-configuration contract;
  2. diagnose and fix one pipeline or storage defect with a regression test;
  3. change one migration or simulation contract and verify its boundary.

Copy template.md for each run. A reviewer confirms command results, regressions, human intervention, and failure classification. Compare guidance or tooling changes only across tasks of similar scope; a faster trivial task is not evidence of improved readiness.

The baseline is complete only when all three records contain reproducible commands and reviewed results. Store secrets, model transcripts, and production data outside the record.