How far vibey is from full autonomy¶
vibey's goal is a delivery loop that runs from a specification to a merged, released change
with nobody present, and that stops for a person only where a rule the operator declared says
it must. This page shows how far it is from that, as a measurement. Everything between the
markers below is generated: scripts/autonomy_scorecard.py computes it from primary data
and appends each measurement to the record,
docs/architecture/evidence/autonomy-scorecard.jsonl.
The prose around the markers is written by hand; the numbers are not.
The headline¶
Measured 2026-10-02 11:39Z (the cutoff). Source: docs/architecture/evidence/autonomy-scorecard.jsonl, regenerated by scripts/autonomy_scorecard.py; do not edit inside these markers.
1 of 11 stages autonomous (3 partial, 6 manual, 1 unknown) at the cutoff, 2026-10-02 11:39Z.
Not yet autonomous:
- DESIGN resolves without a person: partial
- BUILD runs unattended: unknown
- Paid engines stay available: manual
- REVIEW resolves without a person: partial
- The pull-request review reaches a verdict: manual
- Approval comes from the delegated approver: manual
- Merges land without the operator: manual
- Promotion and release run without hand steps: manual
- CI stays green without human re-runs: partial
- The queue delivers projects to DONE: manual
The count is of declared stages. It is not a percentage of autonomy, and no stage is worth more than another: a stage either runs without a person by its own declared thresholds or it does not, and the table says which criterion holds it back.
What "full autonomy" means here¶
The six-phase model (CLAUDE.md,
The six-phase model) is designed to talk to a person in DESIGN, REVIEW and the deployment
stages, and to run BUILD unattended. The operator's standing grant ([autonomy] in
.vibey-gh.toml, given 2026-09-30) asks for more: where a step would need a person, replace
it with a declared TOML gate recorded on the ledger, "so we can prove that it is possible to
fully automate software engineering forever". The floor does not move with it. Sub-doctrine 12.d
bounds unattended authority by a gate, never by judgement: no bypass, no self-approval, no
irreversible act without a person, and silence is never consent. Sub-doctrine 12.f lets a
delegated approver give the approval a merge needs, applying a standard the operator fixed
in advance.
So full autonomy, for this scorecard, means every stage below runs without a person by a declared rule, not by someone choosing to look away. The stages are of two kinds:
- The product's loop (DESIGN, BUILD, the engines, REVIEW, the queue reaching DONE),
measured from the local queue, read only. A human gate counts as answered without a person
only when the gate-timeout sweep answered it (a kind the project declared under
[human_gates] timeout_defaults, #1299) or a declared automation did. - The project's own loop (the pull-request review, its trustworthiness, approval,
merge, promotion and release, CI), measured from the forge with
ghand from the review canary's ledger.
The deployment stage set (DEPLOY_DESIGN, DEPLOY_EXECUTE, DEPLOY_REVIEW) is not scored: it is entered only after an explicit opt-in, and an irreversible real-world act without a person is outside what any grant covers (SD-01 §6).
Each stage, its question, its figures, its thresholds and the reason for each threshold are
declared in scripts/autonomy_scorecard.toml.
A stage is as far along as its weakest judged criterion: autonomous when every criterion
meets its threshold, manual when any judged criterion is under its partial threshold,
partial otherwise, and unknown when nothing judged falls short but something could
not be judged. A criterion whose sample is smaller than its
declared minimum is unknown, never a pass.
The scorecard¶
Measured 2026-10-02 11:39Z (the cutoff). Source: docs/architecture/evidence/autonomy-scorecard.jsonl, regenerated by scripts/autonomy_scorecard.py; do not edit inside these markers.
| Stage | Metric | Value | Autonomous at | Criterion | Stage |
|---|---|---|---|---|---|
| DESIGN resolves without a person | DESIGN gates answered without a person | 14 of 23 (61%) (stale, measured 2026-10-02) | ≥ 95%, n ≥ 5 | partial | partial |
| BUILD runs unattended | BUILD jobs finished per finished job or escalation | 1 of 3 (33%) (stale, measured 2026-10-02) | ≥ 90%, n ≥ 10 | unknown | unknown |
| Paid engines stay available | paid engines with a login check inside its time-to-live | 0 of 4 (0%) (stale, measured 2026-10-02) | ≥ 100% | manual | manual |
| REVIEW resolves without a person | REVIEW gate kinds that may resolve without a person | 1 of 2 (50%) | ≥ 100% | partial | partial |
| REVIEW resolves without a person | REVIEW gates answered without a person | unknown | ≥ 95%, n ≥ 3 | unknown | partial |
| The pull-request review reaches a verdict | merged PRs with a review verdict at their head | 27 of 313 (9%) | ≥ 95%, n ≥ 5 | manual | manual |
| The review's verdict is trustworthy | the latest review-canary measurement meets its floor | yes | yes | autonomous | autonomous |
| Approval comes from the delegated approver | merged PRs approved by the delegated approver | 0 of 313 (0%) | ≥ 95%, n ≥ 5 | manual | manual |
| Approval comes from the delegated approver | the review gate is a required check on the integration branch | no | yes | manual | manual |
| Merges land without the operator | merged PRs merged by an account other than the operator's | 5 of 313 (2%) | ≥ 95%, n ≥ 5 | manual | manual |
| Merges land without the operator | merged PRs carrying an approving review (no bypass) | 0 of 313 (0%) | ≥ 95%, n ≥ 5 | manual | manual |
| Promotion and release run without hand steps | promotion runs a person dispatched by hand | 8 | ≤ 0 | manual | manual |
| Promotion and release run without hand steps | promotions merged into the release branch by an account other than the operator's | 0 of 15 (0%) | ≥ 95%, n ≥ 2 | manual | manual |
| Promotion and release run without hand steps | release publishes that succeeded | 5 of 6 (83%) | ≥ 90%, n ≥ 2 | partial | manual |
| CI stays green without human re-runs | CI runs on the integration branch that succeeded | 195 of 288 (68%) | ≥ 90%, n ≥ 10 | partial | partial |
| CI stays green without human re-runs | CI runs on the integration branch that were re-run | 0 of 343 (0%) | ≤ 5%, n ≥ 10 | autonomous | partial |
| The queue delivers projects to DONE | projects that reached DONE | 0 (stale, measured 2026-10-02) | ≥ 1 | manual | manual |
| The queue delivers projects to DONE | hours since the queue's latest event | 50.1 h (stale, measured 2026-10-02) | ≤ 24 h | partial | manual |
Each stage's evidence¶
Every figure names its source, its window or the moment it was read, and how it was
obtained. A stale figure is one this measurement could not re-read: the last good value
is shown with the date it was measured, for at most max_stale_days, and is then unknown.
An unknown figure has no number, and none is invented.
Measured 2026-10-02 11:39Z (the cutoff). Source: docs/architecture/evidence/autonomy-scorecard.jsonl, regenerated by scripts/autonomy_scorecard.py; do not edit inside these markers.
DESIGN resolves without a person¶
Status: partial. Of the DESIGN-phase gates answered in the window, how many were answered by the gate-timeout sweep or a declared automation rather than a person?
- DESIGN gates answered without a person: 14 of 23 (61%) (stale, measured 2026-10-02), stale; autonomous at ≥ 95%, n ≥ 5; criterion partial. Note: not re-read at 2026-10-02: no read-only queue DSN in $VIBEY_AUTONOMY_QUEUE_DSN. Source: queue, window 2026-09-02 to 2026-10-02 11:07Z;
human_gaterows raised in the window by DESIGN jobs and answered; unattended = answered bygate-timeout,automation:%. Why this threshold: DESIGN is designed to talk to a person (CLAUDE.md, the six-phase model); it is autonomous only where every gate it raises is answered by a declared, ledgered rule.
What would close it: Answer the DESIGN interview from declared defaults ([design.interview]) or a declared automation, and record it on the ledger, so no DESIGN gate waits for a person.
Answers to: ADR-0009, ADR-0027, sub-doctrine 12.d.
BUILD runs unattended¶
Status: unknown. Of the BUILD jobs that finished and the BUILD escalations raised in the window, how many were finished jobs rather than gates waiting for a person?
- BUILD jobs finished per finished job or escalation: 1 of 3 (33%) (stale, measured 2026-10-02), stale; autonomous at ≥ 90%, n ≥ 10; criterion unknown. a sample of 3, under the declared minimum of 10. Note: not re-read at 2026-10-02: no read-only queue DSN in $VIBEY_AUTONOMY_QUEUE_DSN. Source: queue, window 2026-09-02 to 2026-10-02 11:07Z; BUILD jobs that reached
succeededin the window, over those plus the gates BUILD jobs raised in it. Why this threshold: BUILD is designed to run unattended; an exhaustion gate is the point where it stops and asks a person.
What would close it: Fewer BUILD escalations (budget, repair or attempts exhausted) per finished job: the escalations are the gates a person answers.
Answers to: ADR-0004, ADR-0005, ADR-0024, ADR-0038.
Paid engines stay available¶
Status: manual. Of the declared paid engines, how many passed a login check within its time-to-live at the cutoff?
- paid engines with a login check inside its time-to-live: 0 of 4 (0%) (stale, measured 2026-10-02), stale; autonomous at ≥ 100%; criterion manual. Note: not re-read at 2026-10-02: no read-only queue DSN in $VIBEY_AUTONOMY_QUEUE_DSN. Source: queue, window 2026-09-02 to 2026-10-02 11:07Z; each paid engine's latest
engine_health.auth_ok_at, within 24 h of the cutoff. Why this threshold: The 2026-09-30 scan found every paid job deferred once a login was a day old; a lapsed login is the dropout.
What would close it: A running worker that keeps every paid engine's login fresh (#1295 rechecks at half its life); an engine whose login lapses drops out of rotation until someone logs in.
Answers to: ADR-0038, ADR-0070, sub-doctrine 8.a.
REVIEW resolves without a person¶
Status: partial. Can each gate REVIEW raises resolve without a person, and did the ones answered in the window?
- REVIEW gate kinds that may resolve without a person: 1 of 2 (50%), declared; autonomous at ≥ 100%; criterion partial. Note: cannot time out:
approval. Source: repo, read 2026-10-02 11:38Z;DEFAULT_ANSWER_KEYSinsrc/vibey/domain/gate_timeout.py, against REVIEW's gate kinds. Why this threshold: A gate kind without a timeout default waits for a person however long (the operator's ruling of 2026-09-30, #1299). - REVIEW gates answered without a person: unknown, unknown; autonomous at ≥ 95%, n ≥ 3; criterion unknown. no read-only queue DSN in $VIBEY_AUTONOMY_QUEUE_DSN. Source: queue, window 2026-09-02 to 2026-10-02 11:39Z; the local queue, read only. Why this threshold: What the declarations allow is not evidence they are used: the answered gates are.
What would close it: A declared, ledgered rule for REVIEW's approval gate, which today carries no default and can never time out; until then REVIEW keeps its person by design.
Answers to: ADR-0009, ADR-0010, sub-doctrine 12.d, sub-doctrine 12.f.
The pull-request review reaches a verdict¶
Status: manual. Of the pull requests merged into the integration branch in the window, how many carry a review verdict (passed or blocked) at their merged head?
- merged PRs with a review verdict at their head: 27 of 313 (9%), measured; autonomous at ≥ 95%, n ≥ 5; criterion manual. Source: forge, window 2026-09-18 to 2026-10-02 11:38Z; the latest
PR review / gatecheck run at each merged head; a verdict is a title matching^PR review: gate \(orfound blocking findings. Why this threshold: A merge without a verdict is a merge a person vouched for, or nobody did.
What would close it: A review that answers every head in time (#1316 bounds the call; #1303 gives it the files), and a merge that waits for the verdict.
Answers to: ADR-0042, sub-doctrine 8.b, sub-doctrine 10.f.
The review's verdict is trustworthy¶
Status: autonomous. Does the latest review-canary measurement meet the declared recall and false-positive floor, for the review's current settings and corpus, within its age limit?
- the latest review-canary measurement meets its floor: yes, measured; autonomous at yes; criterion autonomous. Source: canary, read 2026-10-02 11:38Z;
vibey-gh review-canary status --jsonagainst the committed ledger. Why this threshold: An unattended approver leans on the review; a review whose recall is unmeasured or under the floor is not something to lean on (vibey-gh review-canary status).
What would close it: A weekly review-canary measurement that meets [pr_automation.review_canary]'s floor for the settings the review runs with.
Answers to: ADR-0040, sub-doctrine 10.f, sub-doctrine 12.f.
Approval comes from the delegated approver¶
Status: manual. Of the pull requests merged in the window, how many carry an approving review by the delegated approver, and is the review gate a required check?
- merged PRs approved by the delegated approver: 0 of 313 (0%), measured; autonomous at ≥ 95%, n ≥ 5; criterion manual. Source: forge, window 2026-09-18 to 2026-10-02 11:38Z;
gh pr list --base develop --state merged, merged in the window; an APPROVED review bythevibeyproject. Why this threshold: Sub-doctrine 12.f: approval while nobody watches comes from the operator's declared approver, applying a standard it did not choose. - the review gate is a required check on the integration branch: no, measured; autonomous at yes; criterion manual. Source: forge, window 2026-09-18 to 2026-10-02 11:38Z; active branch rulesets that include
refs/heads/develop, their required status checks (2 rulesets; required: No skip markers, No-loss property suite (10,000 examples), VS Code extension (macos-latest), VS Code extension (ubuntu-latest), gates). Why this threshold: A gate that is not required can be merged past; the approver's standard rests on the gate being green.
What would close it: A workflow that calls the delegated approver once every gate is green, and PR review / gate declared as a required status check on the integration branch.
Answers to: ADR-0046, ADR-0049, sub-doctrine 12.f.
Merges land without the operator¶
Status: manual. Of the pull requests merged in the window, how many were merged by an account other than the operator's, and how many carried the approving review the ruleset requires (no bypass)?
- merged PRs merged by an account other than the operator's: 5 of 313 (2%), measured; autonomous at ≥ 95%, n ≥ 5; criterion manual. Source: forge, window 2026-09-18 to 2026-10-02 11:38Z;
gh pr list --base develop --state merged, merged in the window;mergedByis notadammatthewsteinberger. Why this threshold: A merge by the operator's account is a merge by the operator. - merged PRs carrying an approving review (no bypass): 0 of 313 (0%), measured; autonomous at ≥ 95%, n ≥ 5; criterion manual. Source: forge, window 2026-09-18 to 2026-10-02 11:38Z;
gh pr list --base develop --state merged, merged in the window; an APPROVED review by an account other than the author. Why this threshold: The integration branch requires one approving review; a merge without one went in through the ruleset bypass (12.d: no gate routed around).
What would close it: Merges made by the merge train after an approval, so the ruleset is satisfied rather than bypassed.
Answers to: ADR-0036, ADR-0046, sub-doctrine 12.d.
Promotion and release run without hand steps¶
Status: manual. In the window, how many promotions did a person dispatch by hand, how many promotions merged into the release branch without the operator, and how many release publishes succeeded?
- promotion runs a person dispatched by hand: 8, measured; autonomous at ≤ 0; criterion manual. Source: forge, window 2026-09-18 to 2026-10-02 11:38Z;
promote-to-main.ymlruns created in the window byworkflow_dispatchwhose triggering actor is not a bot (992 other runs in the window started on their own). Why this threshold: A promotion dispatched by hand is a hand step. A share of runs would hide it: the merge train starts hundreds of runs that find nothing to promote, so the count of hand dispatches is the honest figure (at most one a week is partial). - promotions merged into the release branch by an account other than the operator's: 0 of 15 (0%), measured; autonomous at ≥ 95%, n ≥ 2; criterion manual. Source: forge, window 2026-09-18 to 2026-10-02 11:38Z;
gh pr list --base main --state merged, merged in the window;mergedByis notadammatthewsteinberger. Why this threshold: A promotion merged by the operator's account is the operator's act. - release publishes that succeeded: 5 of 6 (83%), measured; autonomous at ≥ 90%, n ≥ 2; criterion partial. Source: forge, window 2026-09-18 to 2026-10-02 11:38Z;
vibey-engine.ymlruns onmain, eventpush, created in the window; cancelled and skipped runs left out. Why this threshold: A failed publish needs a person to repair and re-run it.
What would close it: Promotion started by the merge train or its schedule and merged without the operator's account; a release workflow that publishes on every push to the release branch.
Answers to: ADR-0028, ADR-0069, sub-doctrine 12.e.
CI stays green without human re-runs¶
Status: partial. Of the CI runs on the integration branch in the window, how many succeeded, and how many needed a re-run?
- CI runs on the integration branch that succeeded: 195 of 288 (68%), measured; autonomous at ≥ 90%, n ≥ 10; criterion partial. Source: forge, window 2026-09-18 to 2026-10-02 11:38Z;
ci.ymlruns ondevelop, eventpush, created in the window; cancelled and skipped runs left out. Why this threshold: A red integration branch stops the merge train until someone repairs it. - CI runs on the integration branch that were re-run: 0 of 343 (0%), measured; autonomous at ≤ 5%, n ≥ 10; criterion autonomous. Source: forge, window 2026-09-18 to 2026-10-02 11:38Z;
ci.ymlruns ondevelop, eventpush, created in the window;run_attemptabove 1. Why this threshold: A re-run is a person deciding a failure was a flake.
What would close it: A green integration branch whose failures are real and repaired by a change, not by pressing re-run on a flake.
Answers to: ADR-0023, sub-doctrine 9.c, sub-doctrine 12.e.
The queue delivers projects to DONE¶
Status: manual. Did any project reach DONE in the window, and has the queue recorded any event within a day of the cutoff?
- projects that reached DONE: 0 (stale, measured 2026-10-02), stale; autonomous at ≥ 1; criterion manual. Note: not re-read at 2026-10-02: no read-only queue DSN in $VIBEY_AUTONOMY_QUEUE_DSN. Source: queue, window 2026-09-02 to 2026-10-02 11:07Z;
PhaseTransitionedevents intodonein the window. Why this threshold: DONE is the only completion the six-phase model records; nothing else is delivery. - hours since the queue's latest event: 50.1 h (stale, measured 2026-10-02), stale; autonomous at ≤ 24 h; criterion partial. Note: not re-read at 2026-10-02: no read-only queue DSN in $VIBEY_AUTONOMY_QUEUE_DSN. Source: queue, window 2026-09-02 to 2026-10-02 11:07Z; the latest
event.produced_atbefore the cutoff. Why this threshold: A queue with no event in a day has no worker draining it.
What would close it: A worker and the bridge installed as services (vibey supervisor install) and draining the queue, with projects reaching DONE.
Answers to: ADR-0002, ADR-0044, sub-doctrine 9.c.
Every figure¶
Measured 2026-10-02 11:39Z (the cutoff). Source: docs/architecture/evidence/autonomy-scorecard.jsonl, regenerated by scripts/autonomy_scorecard.py; do not edit inside these markers.
| Figure | Value | Status | Source | Reason |
|---|---|---|---|---|
canary.age_days |
0.5 days | measured | canary | - |
canary.false_positive_upper_bound |
0.21 | measured | canary | - |
canary.meets_floor |
yes | measured | canary | - |
canary.recall_lower_bound |
0.52 | measured | canary | - |
forge.approval.approver_share |
0 of 313 (0%) | measured | forge | - |
forge.approval.gate_required |
no | measured | forge | - |
forge.ci.rerun_share |
0 of 343 (0%) | measured | forge | - |
forge.ci.success_share |
195 of 288 (68%) | measured | forge | - |
forge.merge.not_operator_share |
5 of 313 (2%) | measured | forge | - |
forge.merge.reviewed_share |
0 of 313 (0%) | measured | forge | - |
forge.merge.reviews_per_merge |
0.57 | measured | forge | - |
forge.promotion.hand_dispatches |
8 | measured | forge | - |
forge.promotion.main_not_operator_share |
0 of 15 (0%) | measured | forge | - |
forge.release.success_share |
5 of 6 (83%) | measured | forge | - |
forge.review.verdict_share |
27 of 313 (9%) | measured | forge | - |
queue.build.unattended_share |
1 of 3 (33%) (stale, measured 2026-10-02) | stale | queue | not re-read at 2026-10-02: no read-only queue DSN in $VIBEY_AUTONOMY_QUEUE_DSN |
queue.design.unattended_share |
14 of 23 (61%) (stale, measured 2026-10-02) | stale | queue | not re-read at 2026-10-02: no read-only queue DSN in $VIBEY_AUTONOMY_QUEUE_DSN |
queue.done_projects |
0 (stale, measured 2026-10-02) | stale | queue | not re-read at 2026-10-02: no read-only queue DSN in $VIBEY_AUTONOMY_QUEUE_DSN |
queue.engines.paid_auth_fresh_share |
0 of 4 (0%) (stale, measured 2026-10-02) | stale | queue | not re-read at 2026-10-02: no read-only queue DSN in $VIBEY_AUTONOMY_QUEUE_DSN |
queue.last_event_hours |
50.1 h (stale, measured 2026-10-02) | stale | queue | not re-read at 2026-10-02: no read-only queue DSN in $VIBEY_AUTONOMY_QUEUE_DSN |
queue.review.unattended_share |
unknown | unknown | queue | no read-only queue DSN in $VIBEY_AUTONOMY_QUEUE_DSN |
repo.review.timeout_kinds_share |
1 of 2 (50%) | declared | repo | cannot time out: approval |
How it is kept current¶
A weekly workflow (.github/workflows/autonomy-scorecard.yml) observes the forge and the
review canary on a GitHub-hosted runner and the local queue on the project's self-hosted
runner, appends one line to the record, re-renders this page, the README's summary, the
documentation's landing page and the research paper's table, and opens a pull request. It
never pushes to develop; the change lands through the merge train like any other. A source
a week's run cannot reach keeps its last value, marked stale with its date, and after two
weeks shows as unknown.
The record is append-only: each line carries its sequence number, the digest of the line
before it and its own, through the family's ledger (vibey_gh.review_canary.CanaryLedger).
python scripts/autonomy_scorecard.py check runs in CI. It fails when a line has been
edited, when the latest verdicts do not follow from its figures, when the stages in the TOML
differ from the ones the latest line was judged by (reevaluate re-judges the same figures
and appends the result), or when any generated block is out of date.
To measure by hand:
VIBEY_AUTONOMY_QUEUE_DSN='postgresql:///vibey' \
uv run python scripts/autonomy_scorecard.py measure
uv run python scripts/autonomy_scorecard.py check
Where this came from¶
The first reading was the autonomy scan of 2026-09-30, recorded in
docs/architecture/evidence/autonomy-2026-10-01.md
and in the research paper's section The autonomy scan of 2026-09-30. It found BUILD the
only phase that carried a project forward with nobody present, every merge into develop
made by the operator's account through the ruleset bypass, and an approver account that had
reviewed nothing. This scorecard re-measures those findings every week instead of repeating
them.
Related decisions and law: ADR-0009 (human gates park jobs), ADR-0040 (evidence-bounded status), ADR-0046 (unattended authority is bounded by a gate), ADR-0047 (the toil is automated), ADR-0049 (the delegated approver), and sub-doctrines 10.f, 12.d, 12.e and 12.f in the Twelve Doctrines.