Skip to content

For platform and devex teams

Decide whether an agent setup is actually getting better

You own the tooling everyone else runs. Excellent records what each agent run did, checks it against evidence rather than a model's opinion, and tests a change to the prompt, model or tools on held-out work before it becomes the default.

The problems

Three things platform and devex teams keep telling us

01

Prompt changes, model swaps and retry rules happen ad hoc, so nobody can say whether the setup improved or just changed.

02

A tool that asks a model whether the model's work is good is not an independent check, and you know it.

03

A passing check in one repository does not tell you the same setup holds for the next class of task.

Outcomes

What verified agent work looks like for you

Evidence, not a second opinion

Checks are things that run — git state, tests, commands — and a claim with no observation behind it is reported as unchecked rather than assumed.

Receipts you can inspect outside the tool

A receipt binds the work item, attempt, versions, checks and result into a record that can be read and verified after the session is gone.

Setup changes go through a trial

A candidate instruction, model or tool change runs against held-out tasks, and the comparison — not a vibe — decides whether it is promoted.

Go and look

The pages that answer this, rather than a claim that we did

Where it stops

What Excellent does not do for you yet

  • Browser, CI and deployment observation are experimental. Git and test evidence are the parts to lean on.
  • The source is not public yet, so the checks are inspectable through the receipt rather than the repository.
  • Nothing is hosted and nothing is shared between machines — there is no central dashboard to roll out.

Solutions

Someone else does this work at your company?

Run it on one real task

Install Excellent, point it at the agent you already use, and read the first result before you trust the claim.

How it works