Evaluation
How to measure a model or a pipeline honestly: harnesses, datasets, statistics, and the traps.
-
Evaluation ยท
Bootstrap CIs for eval deltas: is a 4-point gap on 500 items real?
A paired bootstrap for the accuracy gap between two systems on 500 items, with McNemar's test, coverage checked over 1,000 replays, and the code and per-item data.