Case studies / webpack/webpack

The estimate said +45.6%. Verification said +14.1%.

VERIFIED · p = 0.019 · N = 10

The point of this exhibit is the gap between those two numbers. Our autoresearch engine estimated a +45.6% wall-clock win for a composed candidate against webpack/webpack. Ten interleaved live runs on the substrate a real client actually uses — GitHub-hosted CI — verified a statistically significant but far smaller win: +14.1%. The change is real. The estimate was roughly three times too optimistic. Only the measurement is safe to guarantee — and only measurement on the real substrate produced it.

The verified result

N=10, GitHub-hosted, interleaved

Candidatesetup-node-cache + remove-needs, composed
Estimate (autoresearch)+45.6% wall-clock
Verified p50 wall-clock delta+14.1% (baseline p50 2757s → candidate p50 2369s)
Mann-Whitney U p-value0.019 (< 0.05 → gate kept / verified)
p50 bootstrap CI[≈0%, +34.1%]
p95 wall-clock delta+21.0% (no tail regression)
N10 baseline / 10 candidate, interleaved — not provisional
calibration_ratio (verified ÷ estimated)0.31

Read it honestly, both ways. It passed: the candidate genuinely makes webpack's CI faster, and the Mann-Whitney rank test is significant because the candidate distribution sits consistently below baseline. And it's a large calibration miss: +14.1% is under a third of the +45.6% estimate, and the p50 bootstrap CI's lower bound sits right at zero — so while the effect is significant by the rank test, its magnitude carries real uncertainty. A guarantee written off the +45.6% estimate would have been off by roughly 3×; a guarantee written off the +14.1% measured number is safe.

Why the estimate ran hot

The critical-path model estimated the saving from mean/median per-task durations on the dependency graph. Real GitHub-hosted wall-clock is set by matrix-tail and queue/runner-start latency variance that don't shrink proportionally when you cache — the kind of substrate noise a static model can't see and a live measurement catches automatically.

What this means for you

This is why the diagnostic step exists, and why it runs on your CI specifically, not a generic model. An estimate tells you where to look. A measurement tells you what to guarantee.


Book an engagement

← Back to case studies