For platform and devex teams
Decide whether an agent setup is actually getting better
You own the tooling everyone else runs. Excellent records what each agent run did, checks it against evidence rather than a model's opinion, and tests a change to the prompt, model or tools on held-out work before it becomes the default.
The problems
Three things platform and devex teams keep telling us
01
Prompt changes, model swaps and retry rules happen ad hoc, so nobody can say whether the setup improved or just changed.
02
A tool that asks a model whether the model's work is good is not an independent check, and you know it.
03
A passing check in one repository does not tell you the same setup holds for the next class of task.
Outcomes
What verified agent work looks like for you
Evidence, not a second opinion
Checks are things that run — git state, tests, commands — and a claim with no observation behind it is reported as unchecked rather than assumed.
Receipts you can inspect outside the tool
A receipt binds the work item, attempt, versions, checks and result into a record that can be read and verified after the session is gone.
Setup changes go through a trial
A candidate instruction, model or tool change runs against held-out tasks, and the comparison — not a vibe — decides whether it is promoted.
Go and look
The pages that answer this, rather than a claim that we did
- How verification decidesThe rules behind verified, returned and unchecked — including what Excellent will not conclude.
- ReceiptsThe receipt schema and how one is checked offline.
- Promotion and rollbackWhat has to be true before a candidate setup replaces the default, and how it is reverted.
- ComparisonsExcellent next to the other tools in this category, with dated prices and sourced claims.
Where it stops
What Excellent does not do for you yet
- Browser, CI and deployment observation are experimental. Git and test evidence are the parts to lean on.
- The source is not public yet, so the checks are inspectable through the receipt rather than the repository.
- Nothing is hosted and nothing is shared between machines — there is no central dashboard to roll out.
Solutions
Someone else does this work at your company?
Run it on one real task
Install Excellent, point it at the agent you already use, and read the first result before you trust the claim.