Give EvalRun.summary a typed RunSummary value (unified RunError, lenient legacy parsing) so readers stop reaching into a schemaless dict, and route every cross-run rollup — dashboard, scenario ranking, trend, campaign report — through one aggregate_runs seam. Fixes the divergence where stats averaged pass_rate over completed-only runs while the campaign report counted faults as 0.0. Cross-run rule (ADR-0004): genuine faults count 0.0, user-cancelled runs are excluded from both denominators. |
||
|---|---|---|
| .. | ||
| rules | ||
| __init__.py | ||
| campaign_runner.py | ||
| campaign_scheduler.py | ||
| engine.py | ||
| judgement.py | ||
| metrics.py | ||
| report.py | ||