Our results.
We are building Excellent with Excellent.
This page shows what changed, how we tested it, what failed, and what we still do not know.
Current release.
Release Not published
Date Not published
Source commit Not published
Build digest Not published
Agent version before Not published
Agent version after Not published
Task distribution version Not published
Held-out set digest Not published
Verification contract version Not published
Receipt ID Not publishedNo signed public relaunch release result has been published yet. The page is wired to the public results manifest and will render the release fields when they exist.
Before and after.
Before Current
Verified result rate Not published Not published
Cost per verified task Not published Not published
Human help per task Not published Not published
p50 elapsed time Not published Not published
p95 elapsed time Not published Not published
Known outcome regressions Not published 0Unknown work stays in the denominator. The primary comparison will show an uncertainty interval when a signed result includes one.
What changed.
Model use Not published
Instructions Not published
Context Not published
Tools Not published
Checking Not published
Retrying Not published
Stopping Not publishedEach real change will link to the experiment that supported it. No change is listed as promoted until a receipt exists.
What was tested.
Tasks used for baseline 0
Tasks used for candidate work Not published
Tasks used for validation Not published
Held-out tasks 0
Hidden from candidate proposer Not published
Mechanical checks Git, tests
Model or human review Not published
Outcome observation window 0 days
Environments tested macOS local
Skipped Release manifest, public receipt, confidence intervalModels and providers.
Excellent is designed to sit beside the coding agent and model you choose. It checks the output against repo evidence, so the comparison can survive a model change.
OpenAI: Supported
Anthropic: Supported
Google: Supported
xAI: Supported
Meta: Supported
Mistral: Supported
DeepSeek: Supported
Groq: SupportedCandidates rejected.
Candidate records Not published
Rejected candidate count 0
Invalid candidate count Not publishedRejected candidates are part of the result. This section should show meaningful failures, not only the successful path.
Known limits.
- No signed public relaunch result has been published yet.
Verify it yourself.
excellent receipt verify release-Not published.json- Download receipt: Not published
- View signed payload: Not published
- View public key: Not published
- Receipt format: Read the receipt format
Release history.
Version Date Verified result Cost/result Status Receipt
current - Not published Not published In progress Not published