I0838 — bench compare and the top regressions report
Add bench compare <run-a> <run-b> and make the human report rank the largest regressions, the largest improvements and the slowest targets by p95.
READY wave 5 · p0 · implementation profile · runs alone
Part of M004 — Performance is executable evidence, and the hot path does no canonical work twice.
Ready. Every dependency is done, so majordomus plan start I0838 will be accepted.
Objective
Add bench compare <run-a> <run-b> and make the human report rank the largest regressions, the largest improvements and the slowest targets by p95.
Why
A wall of numbers hides the one that matters.
Current state
Only --check compares, against the baseline.
Desired state
Any two persisted runs compare with the same table as --check; the default report ends with the ranked tail.
Scope
- lib/bench.sh
- docs/CLI.md
- test/cases/85_bench_compare.sh
Dependencies
Acceptance criteria
- compare of a run with itself reports zero deltas
- Case 85 proves the ranking on two fixture runs
Validation
- bash test/run.sh 85_bench_compare
Evidence required
- regression_refused
Evidence
None recorded. Every token above needs a command or an artifact behind it before this issue can be completed; narrative is refused.
Risk
Percentile arithmetic is implemented once and unit-tested on small datasets.
Timeline
- started
- —
- verified
- —
- completed
- —
Those three fields, the evidence above and the state of the dependencies are all the status is made of. There is no status field to disagree with them.
Canonical record: .ai/repo/project/issues/I0838.yaml. Read it back with majordomus plan show I0838.