Why the Score Lies
Stop trusting public numbers and build a private eval set from your own failures.
$87
this phase
- Open1.1
The Saturated Leaderboard
Free previewMMLU and GPQA now sit near the ceiling, so a five-point spread ranks nothing about your product.
- 1.2
Contamination, Measured
Canary strings, n-gram overlap, and verbatim recall probes on the benchmark you were about to trust.
- 1.3
The Leaderboard Illusion
Private variant testing and silent deprecation, and why an Arena rank is not a procurement decision.
- 1.4
Error Analysis First
Read one hundred real traces and open-code the failures before you write a single test case.
- 1.5
A Failure Taxonomy
Group the open codes into categories a grader can score, with counts so you know what to fix first.
- 1.6
Dev Split, Frozen Gate
Two splits, one you iterate against and one only CI may run, plus items dated after the training cutoff.