Skip to content

I0838 — bench compare and the top regressions report

Add bench compare <run-a> <run-b> and make the human report rank the largest regressions, the largest improvements and the slowest targets by p95.

READY wave 5 · p0 · implementation profile · runs alone

Part of M004 — Performance is executable evidence, and the hot path does no canonical work twice.

Ready. Every dependency is done, so majordomus plan start I0838 will be accepted.

Objective

Add bench compare <run-a> <run-b> and make the human report rank the largest regressions, the largest improvements and the slowest targets by p95.

Why

A wall of numbers hides the one that matters.

Current state

Only --check compares, against the baseline.

Desired state

Any two persisted runs compare with the same table as --check; the default report ends with the ranked tail.

Scope

  • lib/bench.sh
  • docs/CLI.md
  • test/cases/85_bench_compare.sh

Dependencies

Acceptance criteria

  • compare of a run with itself reports zero deltas
  • Case 85 proves the ranking on two fixture runs

Validation

  • bash test/run.sh 85_bench_compare

Evidence required

  • regression_refused

Evidence

None recorded. Every token above needs a command or an artifact behind it before this issue can be completed; narrative is refused.

Risk

Percentile arithmetic is implemented once and unit-tested on small datasets.

Timeline

started
verified
completed

Those three fields, the evidence above and the state of the dependencies are all the status is made of. There is no status field to disagree with them.

Canonical record: .ai/repo/project/issues/I0838.yaml. Read it back with majordomus plan show I0838.