Skip to content
All guides
Rollout8 minute read

Rolling out agent verification: a 30-day plan

A pragmatic four-week plan for replacing ad hoc agent trust with attempts, checks, evidence, results, and receipts.

On this page

Most agent rollouts start with enthusiasm and then hit the same question: who proves the work is actually good?

This is a 30-day plan for adding verification without stopping the team. Each week has one objective, one measurable outcome, and one thing you deliberately do not need to finish.

Week 0 — Choose the first workflow

Do not start with every repository and every agent. Pick one workflow where proof is already valuable:

  • A frontend task with browser behavior.
  • A backend task with a focused test suite.
  • A release chore with CI and deployment evidence.
  • A documentation or contract update with clear review rules.

Write down the current path from request to merge. Include the agent used, the model provider, the repo, the branch pattern, the required commands, the review owner, and what usually gets missed.

The measurable outcome for week 0 is simple: one task type with a known proof envelope.

Week 1 — Record attempts

Install Excellent and run it beside the agent you already use.

Goals for the week:

  1. A real task starts as an Excellent attempt.
  2. The attempt records the request, repo state, model, agent, prompt context, changed files, and agent claim.
  3. The team can open the result and see what happened without reconstructing it from terminal history.

What you are not doing this week:

  • Replacing your agent.
  • Rewriting your release process.
  • Connecting every evidence source.

The goal is observability. Before you improve anything, make the agent's work legible.

Week 2 — Attach checks and evidence

Now connect the proof sources that decide whether the work should be trusted.

Start with the narrow proof envelope from week 0:

  1. Add the task-specific commands.
  2. Capture exit codes and useful output.
  3. Attach browser screenshots or traces when behavior matters.
  4. Link CI or deployment output when the task depends on it.
  5. Record human review notes when judgment is still required.

By the end of week 2, a result should be able to say verified, returned, or unchecked with the evidence attached.

Week 3 — Return failures cleanly

Verification is useful only if failed work comes back with enough detail for the next attempt.

Tune the return path:

  • Send the exact failed check back to the agent.
  • Preserve the original task and failed attempt.
  • Make retries visible instead of overwriting the first run.
  • Require a fresh result for each material change.
  • Mark uncertainty honestly when proof is missing.

This is where teams usually find the highest leverage. A clear returned result turns a vague code review into a concrete repair loop.

Week 4 — Compare the setup

Once attempts and results are consistent, compare better ways to run the agent.

Pick one candidate change:

  • A different model.
  • A smaller or cheaper model for part of the workflow.
  • A new prompt or instruction file.
  • A tighter context package.
  • A different check cadence.
  • A retry rule that explains failures before rerunning commands.

Test the candidate on held-out tasks. Do not promote it because it feels better in a demo. Promote it only when the evidence shows better quality without hiding cost, human help, or regressions.

What is left after 30 days

After the four weeks, you should have:

  • One workflow verified from request to result.
  • Receipts for accepted agent work.
  • Failed attempts returned with concrete evidence.
  • A known set of checks the team trusts.
  • A first candidate comparison for improving the agent setup.

Then repeat the pattern for the next workflow.

When this plan fails

It fails when teams:

  • Start with every repo and every agent at once.
  • Let the agent decide whether its own work is done.
  • Treat CI as the whole proof envelope when product behavior also matters.
  • Hide failed attempts instead of using them to improve checks.
  • Promote a new prompt or model without held-out comparison.

The point is not ceremony. The point is to make agent work inspectable enough that trust can come from evidence.

Keep going

Done reading. Ready to check a real task?

Install Excellent, connect your AI agent, and turn the next result into an evidence-backed receipt.

How it works