Skip to content

Compare · category

Excellent vs CI and human review: the checks you have vs the claim nobody checks

CI runs the checks that exist. It has no idea what the agent was asked to do, or what it claimed.

Jump to comparison

CI and human review

CI and human review, in one line

The default, and the honest baseline: your existing CI pipeline plus a human reading the pull request. No new vendor, no new invoice, no new thing to learn.

Best for
Almost everything, for most of software's history. CI is deterministic, it is yours, and a good reviewer catches what no tool does.
Limits
CI runs whatever checks exist and knows nothing about the task, the attempts behind the change, or what the agent claimed it did. A reviewer can read a diff; a reviewer cannot read fifty agent-authored diffs a day at the same depth. The gap is not the checks — it is that nothing connects the checks to the claim.
Stronger than us at
Cost and trust. You already have it, you already understand it, nobody has to approve a purchase, and a senior engineer's judgement about whether a change is a good idea is something no verifier replaces. Excellent does not review design, and it does not replace the reviewer.
Price 2026-09-23
No new licence (Runner minutes plus reviewer hours)

Excellent

Excellent in one line

A verification system for agentic tasks — it records what the agent attempted, runs your checks against the claim, and binds the result to evidence you can open.

Best for
Teams that want agent work checked before they trust or promote it.
Limits
There is no hosted option. You run it on your machines. Verification evidence stays on the machine that produced it — it does not sync to teammates yet. No SSO, and no Mac app has shipped. It does not write tests, and it does not review your diff.
Stronger than them at
Keeping the attempt record, running your own checks as the oracle, and trialling prompt, model and tool changes on held-out work from your repository.
Price 2026-09-23
Free to install and run (Consulting from $10,000 / month)

Head to head

The capabilities that decide it

Capability by capability — CI and human review on the left, Excellent on the right. Prices are the vendor's own, read on 2026-09-23.

CapabilityCI and human reviewExcellent

Runs deterministic checks

Excellent watches the CI you already have rather than replacing it.

YesYes

Knows what the task was

NoYes

Records each agent attempt and the claim it made

NoYes

What comes back when work fails

A red build and a log to go readThe failed check, the missing artifact or the uncertainty, attached to the attempt

Keeps up with agent output volume

CI does; a human reviewer does notThe record and the checks scale; the judgement call is still a person's

Tests prompt, model and tool changes on held-out work

NoYes

Where it runs

Your CI provider's runners, and your team's attentionOn your machines; it reads your existing GitHub Actions or GitLab CI results

Published price (checked 2026-09-23)

No new licence — runner minutes and reviewer time you already pay for.Free to install and run. Consulting starts at $10,000 / month.

Together

They are not mutually exclusive

Excellent reads your existing GitHub Actions or GitLab CI results through a config file in your repository. It is not a pipeline and does not want to be one.

Verdict

When to pick which

Do not rip out CI, and do not stop reviewing. Excellent's job is the layer neither of them covers: what the agent was asked, what it tried, what it claimed, and which check settled it. Checkout's engineers built that layer internally and reported that every organisation was independently building similar enforcement systems — which is a reasonable argument for not building it a second time yourself.

Stay with CI and human review if

Almost everything, for most of software's history. CI is deterministic, it is yours, and a good reviewer catches what no tool does.

Add Excellent if

You want agent work captured as attempts, checked against evidence you can open, and promoted only when the result is defensible to someone outside the team.

More matchups

See Excellent against the rest of the category

Trust the evidence, not the agent's confidence

Install Excellent, run it beside the agent you already use, and inspect what the checks actually said.

See how it works