Core ideas
How verification decides
Evidence, authority, and why a model can never sign off a pass: the six verdicts, what counts as evidence, retries, flaky tests and anti-gaming.
On this page
"Done" is a claim an agent makes about its own work. A verdict is what happens when something else goes and looks. This page is the rulebook: what counts as evidence, who is allowed to conclude what, and what stops the whole thing from being gamed.
The six verdicts
A verification resolves to exactly one of six states.
| Verdict | Meaning |
|---|---|
| PASS | Every obligation holds, at or above the authority floor. |
| MISMATCH | An independent record was consulted and contradicted the claim. This is an accusation. |
| FAIL | Conclusive, not a pass, and no independent contradiction was demonstrated. A refusal to certify. |
| UNKNOWN | No finding. Coverage was partial or integrity did not hold, and nothing the agent does will change that. A human is asked to fix an instrument. |
| NOT_RUN | Nothing was looked at. |
| PENDING | Too early to tell, and there is a deadline by which it will not be. |
The distinction between FAIL and MISMATCH matters: only the oracle seam, which
reads a record the verifying side populates, can mint a MISMATCH. An evaluator
reads the bundle the submitter chose, so an evaluator can never author an accusation.
When several checks are folded into one verdict, the order of precedence is:
MISMATCH > FAIL > UNKNOWN > NOT_RUN > PENDING > PASSPASS is last. One unresolved check is enough to sink an otherwise clean set. An
empty set aggregates to NOT_RUN, not to PASS.
PENDING is resolved at read time, never by writing back — a receipt is signed and
append-only. A pending verdict that runs out of deadline resolves to UNKNOWN,
FAIL, MISMATCH or NOT_RUN. It is a type-level guarantee that it can never expire
into PASS.
Who is allowed to conclude what
Every check records the authority of whatever decided it:
deterministic— a machine re-ran something and observed the outcome.human— a person looked.semantic— a model judged.imported— the result came from somewhere else.
Only deterministic and human can carry a set to PASS. This is a membership
test, not an exclusion, so it fails closed: a new authority added later is untrusted
until it is deliberately added to that list.
A PASS also requires a deterministic evidence-integrity proof. Without one, the
determination is rejected outright.
When model output counts as evidence
It counts. It just cannot be the thing that says yes.
A semantic worker can move a verdict down — a cited conflict becomes FAIL — or
route it to a person. It can never move one up. The routing table has six outcomes,
and every one of them resolves to UNKNOWN except a cited conflict, which resolves to
FAIL:
| Route | Taken when |
|---|---|
unknown_uncited | The judgment cited nothing. |
uncalibrated_model_needs_review | The model version has no calibration on record. |
unknown_model_reported | The model itself said it did not know. |
cited_conflict_exception | The model found a cited contradiction. Resolves to FAIL. |
high_consequence_support_needs_human | The judgment supports the claim, but consequence is high or critical. |
cited_support_needs_review | The judgment supports the claim and a person still reviews it. |
For a semantic judgment to be considered at all, it must be grounded: bound to a
frozen evaluation case, with every citation resolving to an admitted artifact carrying
a sha256: quote hash. A judgment with no typed predicate is not refused — it resolves
NOT_RUN, which aggregates to UNKNOWN.
The practical consequence, stated in the product's own posture:
A semantic worker may never issue an authoritative PASS, and must cite the evidence it relied on or return
unknown_uncited. So the strongest outcome an injected instruction can reach is routing a judgement to a person — never minting a pass, never closing a finding.
That is also the answer to prompt injection. An attacker who fully controls the model's output can, at worst, cause a human review.
What counts as evidence
Evidence is typed on several axes at once, and the combination is what gives it weight:
- type —
artifact,observation,attestation,review,import - admission source —
producer,collector,reviewer,human,imported - independence —
independent,self_attested,derived,unknown - mutability —
immutable,versioned,mutable - freshness —
fresh,stale,expired,unknown - accessibility —
accessible,inaccessible,redacted,unknown
Independence is assessed on three separate questions — is the actor distinct, is the execution boundary distinct, is custody distinct — and the assessment must state its basis. It is resolved from structure, not accepted as a label. A producer that declares itself independent is not independent.
A claim only reaches PASS or FAIL through a decidable observable. At the
current tier that means one thing: evidence.content-digest, where the expression is
exactly (= <symbol> "sha256:<64 hex>"). Anything else binds and answers UNKNOWN.
Notably, evidence.present is explicitly declined — the absence of cited evidence is
not falsehood.
Retries, re-runs and flaky tests
A re-run is a new attempt, not an overwrite. Receipts are signed and append-only.
The lawful way to change a verdict is to issue a new receipt carrying
supersedesReceiptId. Nothing edits history, so a fail followed by a repair and a pass
leaves both on the record.
Retry limits that exist today:
| Where | Limit |
|---|---|
| Semantic judgment per check | 2 attempts, then semantic-retry-limit-exceeded |
| Verification API executor | 3 attempts |
| Connector queue | 3 attempts |
| Webhook delivery | 6 attempts, backing off 1m, 5m, 25m, 2h05, 10h25 |
Flaky tests are not modelled. There is no quarantine state, no flake score and no "re-run until green" path in the verification kernel — by design, because the shape of "retry until it passes" is exactly what the system exists to refuse. A test that passes and fails on the same input produces two receipts that disagree, and that disagreement is the finding. If you want flake handling today, it belongs in your test runner, before the result becomes evidence.
Anti-gaming
Six rails, all enforced in code rather than convention.
An agent cannot verify its own work. A completion is legitimate only if at least one actor who never claimed the task performed the verification. The check is scoped to the current cycle, so a stale approval from an earlier round cannot satisfy it after a reopen, and an unnamed actor cannot count as a "distinct" claimer.
An agent cannot verify code it wrote, even if someone else claimed the task. The verifier is refused if it appears as the committing actor of any link anchoring the work — including superseded links, so a rebase does not erase the original author.
The verifier must not share the worker's execution boundary. Same actor is
refused as worker-self-verification; same mutable execution boundary is refused as
shared-execution-boundary. A worker also cannot accept the contract for its own
attempt.
A producer's account of what changed is not evidence. A change manifest is only
admitted if it was independently observed — a diff computed in a store that is never
mounted into the container that produced the change. A producer-declared manifest
returns UNKNOWN, never FAIL and never PASS, with the reason
change-manifest-not-independently-observed.
The baseline is captured before the work starts. Any predicate comparing before and
after needs an authoritative pre-mutation baseline. Without one it produces UNKNOWN.
The correlation token cannot be minted except by the capture step, so no caller can
fake having taken it.
An obligation that cannot fail is refused. Falsifiability is decided by witness: a witness that fails proves the obligation was falsifiable. Vacuous predicates, unreachable fail branches and undecidable predicate kinds are all rejected. A false refusal costs an argument; a false admission signs a verdict that means nothing.
On top of those, a structural gate walks the transitive import closure of every hard-gate decision path and refuses a build in which any of them can reach the statistical scorer. No learned component sits in a gate decision, and the gate refuses rather than passes if its own scan shrinks.
What this does not claim
Independence today is process-distinct, not party-distinct. A verifier running in a separate boundary on the same machine, under the same operator, is structurally separated from the producer — it is not an adversarial third party. Read a receipt for what it is: tamper-evident proof of what was bound and checked together, by a process that could not also have written the code.
Verifying a signature proves the receipt has not been altered since the key holder signed it. It does not prove the key is the tenant's key. If you obtained the key from the same system that issued the verdict, a forged receipt re-signed under an attacker's key verifies exactly as cleanly as a genuine one. Confirm the fingerprint through a channel this system does not control, then pin it.
Related
- Sample receipt — what a signed verdict actually contains.
- Receipts — verifying one.
- Claims and evidence — the shorter version of this page.