You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Define and collect evidence showing whether the governance plane improves coordination and agent readiness without becoming an expensive central bottleneck or a false-compliance dashboard.
This measurement begins only after PR #7 is merged, #6 applies the administrative baseline, and #8 lands at least two immutable pilot adopters.
Measures
Collect by repository, risk class, and task type:
clean-clone bootstrap success;
fast/full check duration and flake rate;
clarification count before an agent can identify the canonical owner;
human interventions for R0/R1 work;
correctly stopped/proposed R3/R4 attempts;
stale manifest, generated-view, dependency, owner, and GitHub-setting drift;
false-positive and escaped-drift rate;
reusable workflow adoption and immutable-pin freshness;
exceptions opened, expired, renewed, and closed;
rollback/recovery success;
cross-repository compatibility regressions;
ownership conflicts and time to resolve;
contributor friction and time-to-first-green-PR;
public/private leakage test results;
administrator and backup-reviewer coverage.
Guardrails
Do not use repository count, green badges, or manifest presence alone as readiness proof.
Do not reward agents for bypassing approvals or reducing test scope.
Keep operational data minimized; do not publish private repository names, prompts, memories, user data, personal performance profiles, or security incident detail.
Represent unsupported, unavailable, stale, and degraded evidence explicitly.
Separate structural validity, repository verification, administrative application, control effectiveness, and product/protocol conformance.
Acceptance criteria
Baseline is captured before broad rollout.
Metrics have source, calculation, owner, review cadence, retention, and privacy classification.
At least one adversarial exercise tests metadata spoofing, reusable-workflow pin drift, and private-data leakage.
At least one recovery exercise rebuilds generated views and governance evidence from authoritative inputs.
A 30-day pilot review recommends continue, modify, pause, or replace with explicit evidence.
A 90-day review decides whether file-based governance still scales or a projected service catalog is justified.
Any future service remains a projection unless a separate ADR explicitly migrates authority with rollback.
Measurement results inform governance decisions; they do not certify security, privacy, SOC 2, ISO, GDPR, NIST, OpenSSF, SLSA, or product conformance.
Outcome
Define and collect evidence showing whether the governance plane improves coordination and agent readiness without becoming an expensive central bottleneck or a false-compliance dashboard.
This measurement begins only after PR #7 is merged, #6 applies the administrative baseline, and #8 lands at least two immutable pilot adopters.
Measures
Collect by repository, risk class, and task type:
Guardrails
Acceptance criteria
Measurement results inform governance decisions; they do not certify security, privacy, SOC 2, ISO, GDPR, NIST, OpenSSF, SLSA, or product conformance.