A portable verification procedure for coding agents: implement → verify → review → repair → re-verify, ending in evidence bound to the exact repository state that produced it, rather than in an agent's assurance that it checked.
The problem it solves is narrow and specific. An agent that reports "done" after implementing is asserting something it has not established, and the usual remedy is a person carrying results to a reviewer and findings back. That person is a message bus. This replaces the bus, not the human judgment at the end of it.
TASK
│
▼
IMPLEMENT
│
▼
┌──── VERIFICATION LADDER ◄───┐
│ │ │
│ finding │
│ │ │
│ REPAIR ────────────┘
│
▼
LOCAL EXECUTION ─────┐
CI EXECUTION ────────┼──► COMPOSE ──► COMPLETE = TRUE ──► HUMAN
AGENT ATTESTATIONS ──┘
- A skill (
skills/verification-ladder/) an agent reads per task: nine rungs from baseline through a fresh-reviewer pass, with defined invalidation, re-entry, escalation and termination. - Evidence machinery (
skills/verification-ladder/bin/verify.py): runs a project's gates, records what ran, and binds each result to astate_idtaken over HEAD plus every deviation from it. Standard library andgitonly. - A composition protocol: local execution, CI execution and agent attestations combine into one completion predicate over one state.
A verification result belongs to one repository state. Change the state and
every earlier result is stale. Repair is a state change, so a repair expires the
PASS it was meant to earn. check and compose enforce this; attestations are
discarded outright when the tree moves, because "I reviewed the diff" cannot
survive a change to the diff.
PASS is a claim about evidence you hold. A gate that could not run is
BLOCKED, never PASS. A green gate over a tree that moved while it ran establishes
nothing about either state. A required gate nobody ran reads MISSING and counts
against completion — absence of evidence is not evidence of success.
Every row carries its kind. container-build PASS execution and
self-review PASS attestation are different claims and print differently; a
required gate holding only an attestation reads attested, never executed and
fails the predicate. This is the property that keeps a composite from laundering
an agent's self-assessment into a fact.
VERIFICATION STATE: 7272888 + sha256:3130b1f9…
* lint PASS execution local, ci run 35486373054
* container-build PASS execution ci run 35486373054
diff PASS attestation agent
--------------------------------------------------------------
BLOCKED REQUIRED GATES 0
STALE EVIDENCE 0
STATE MATCH TRUE
CLEAN CLIMB TRUE
--------------------------------------------------------------
READY FOR HUMAN GATE TRUE
The ladder is the same everywhere. What counts as verified is not: each project
declares that in its own committed verification.toml, which names its required
gates, its local commands and the CI step that establishes each CI-only gate.
Changing what "complete" means is then a reviewed change to that project.
A repository with no policy is BLOCKED, not verified. Silence is not consent; an unconfigured project must never read as a passing one.
See INSTALL.md to install it and onboard a repository. Claude Code and Codex can both use a linked skill tree, so the verifier source remains one canonical checkout rather than two copies.
Early. The procedure and the machinery were developed and dogfooded against a real project (Cerberus), which is now a consumer of this repository rather than the owner of the skill. This repository verifies itself with its own ladder on every change.
The evidence model is being revised. docs/design/evidence-v3.md fixes the shape of
verification.ladder.evidence/3, under which a result is admissible only when
its definition, permitted authority, execution provenance, target state, proof
artifacts and verifier identity all refer to the same thing.
The model is implemented and activatable. A project adopts it with
evidence = "verification.ladder.evidence/3" in its committed policy. Until a
checkout activates, records on disk are evidence/2; beside each one the CLI
writes an evidence/3 shadow, and compare and qualify report how the v3
reading differs, while compose refuses to report completion under the older
contract. After verify.py activate writes .verification/authority.json,
new run, attest, and import-ci records are evidence/3, and compose
applies the v3 admissibility rules as the authoritative predicate. Existing
evidence/2 records remain readable for migration diagnostics, but cannot
satisfy activated v3 requirements.
It has qualified against a real consumer. Cerberus adopted evidence/3 on
0.3.0 and reached READY FOR HUMAN GATE TRUE from its own local gate, a real
GitHub Actions run and attested judgment. Six targeted mutations were each
refused for the reason that names them. The record is in
docs/qualification/cerberus-0.3.0.md.
MIT licensed; see LICENSE.