Skip to content
All guides
Evaluation9 minute read

How to choose an agent verification system

Twelve questions to ask before you trust an agent workflow: evidence capture, checks, reviewer boundaries, receipts, and exit cost.

On this page

Choosing an AI agent is only half the decision. The other half is choosing how its work gets checked.

This guide is the checklist we use when a team asks whether an agent workflow is ready for real engineering work. It is organized around the parts that matter after the demo: attempts, checks, evidence, reviewer boundaries, receipts, and improvement.

Bucket 1 — Attempts

1. What exactly is recorded?

The system should capture the task, repository state, branch, changed files, agent identity, model, prompt context, tools, artifacts, and final claim.

If the record is only a chat transcript, review will still depend on memory.

2. Are retries preserved?

Failed attempts should not disappear when the agent tries again. You need the first failure to understand whether the prompt, context, check, or agent behavior needs to change.

3. Can a reviewer reconstruct the work?

A useful attempt record lets a reviewer answer: what was requested, what changed, what the agent claimed, and what proof was required.

Bucket 2 — Checks

4. Can checks be task-specific?

Different work needs different proof. A docs update, API change, UI fix, migration, and release task should not all run the same generic checklist.

5. Does the system capture raw output?

Pass/fail summaries are not enough. Keep command names, exit codes, relevant output, screenshots, traces, CI links, and deployment observations.

6. Does it say when it cannot tell?

Unchecked is a real result. A good system does not turn missing evidence into false confidence.

Bucket 3 — Review Boundaries

7. Is verification separate from the agent?

The agent can claim it is done. Verification should run outside that claim.

Look for structural separation: required checks, reviewer identities, and result transitions that cannot be faked by the same actor that made the change.

8. Can humans attach judgment?

Some work needs human review. The system should let reviewers add notes, decisions, and skipped-proof rationale without losing the mechanical evidence.

9. Can failed work be returned cleanly?

The agent should receive the exact failure or uncertainty, not a vague "please fix." Returned results should preserve the failed evidence and start a new attempt.

Bucket 4 — Receipts And Improvement

10. What does an accepted receipt contain?

At minimum: task, attempt, changed files, checks, artifacts, decision, timestamps, and reviewer context.

11. Can receipts be exported?

If you cannot export receipts and evidence references, the audit trail is rented.

12. How are better setups promoted?

Prompt, model, tool, and context changes should be compared on held-out tasks. A setup should not be promoted because it feels better in one session.

The Honest Summary

The right verification system is the one that makes agent work inspectable enough for your team to trust, return, or reject it.

For Excellent, that means a loop around the agent you already use:

  • Start with a task.
  • Record the attempt.
  • Run checks.
  • Attach evidence.
  • Return a result.
  • Keep a receipt.
  • Compare improvements before promotion.

That is the bar to hold any agent workflow against.

Keep going

Done reading. Ready to check a real task?

Install Excellent, connect your AI agent, and turn the next result into an evidence-backed receipt.

How it works