Compare · category
Excellent vs CI and human review: the checks you have vs the claim nobody checks
CI runs the checks that exist. It has no idea what the agent was asked to do, or what it claimed.
CI and human review
CI and human review, in one line
The default, and the honest baseline: your existing CI pipeline plus a human reading the pull request. No new vendor, no new invoice, no new thing to learn.
- Best for
- Almost everything, for most of software's history. CI is deterministic, it is yours, and a good reviewer catches what no tool does.
- Limits
- CI runs whatever checks exist and knows nothing about the task, the attempts behind the change, or what the agent claimed it did. A reviewer can read a diff; a reviewer cannot read fifty agent-authored diffs a day at the same depth. The gap is not the checks — it is that nothing connects the checks to the claim.
- Stronger than us at
- Cost and trust. You already have it, you already understand it, nobody has to approve a purchase, and a senior engineer's judgement about whether a change is a good idea is something no verifier replaces. Excellent does not review design, and it does not replace the reviewer.
- Price 2026-09-23
- No new licence (Runner minutes plus reviewer hours)
Excellent
Excellent in one line
A verification system for agentic tasks — it records what the agent attempted, runs your checks against the claim, and binds the result to evidence you can open.
- Best for
- Teams that want agent work checked before they trust or promote it.
- Limits
- There is no hosted option. You run it on your machines. Verification evidence stays on the machine that produced it — it does not sync to teammates yet. No SSO, and no Mac app has shipped. It does not write tests, and it does not review your diff.
- Stronger than them at
- Keeping the attempt record, running your own checks as the oracle, and trialling prompt, model and tool changes on held-out work from your repository.
- Price 2026-09-23
- Free to install and run (Consulting from $10,000 / month)
Head to head
The capabilities that decide it
Capability by capability — CI and human review on the left, Excellent on the right. Prices are the vendor's own, read on 2026-09-23.
| Capability | CI and human review | Excellent |
|---|---|---|
Runs deterministic checks | ||
Knows what the task was | ||
Records each agent attempt and the claim it made | ||
What comes back when work fails | A red build and a log to go read | The failed check, the missing artifact or the uncertainty, attached to the attempt |
Keeps up with agent output volume | CI does; a human reviewer does not | The record and the checks scale; the judgement call is still a person's |
Tests prompt, model and tool changes on held-out work | ||
Where it runs | Your CI provider's runners, and your team's attention | On your machines; it reads your existing GitHub Actions or GitLab CI results |
Published price (checked 2026-09-23) | No new licence — runner minutes and reviewer time you already pay for. | Free to install and run. Consulting starts at $10,000 / month. |
Together
They are not mutually exclusive
Excellent reads your existing GitHub Actions or GitLab CI results through a config file in your repository. It is not a pipeline and does not want to be one.
Verdict
When to pick which
Do not rip out CI, and do not stop reviewing. Excellent's job is the layer neither of them covers: what the agent was asked, what it tried, what it claimed, and which check settled it. Checkout's engineers built that layer internally and reported that every organisation was independently building similar enforcement systems — which is a reasonable argument for not building it a second time yourself.
Stay with CI and human review if
Almost everything, for most of software's history. CI is deterministic, it is yours, and a good reviewer catches what no tool does.
Add Excellent if
You want agent work captured as attempts, checked against evidence you can open, and promoted only when the result is defensible to someone outside the team.
More matchups
See Excellent against the rest of the category
Trust the evidence, not the agent's confidence
Install Excellent, run it beside the agent you already use, and inspect what the checks actually said.