August 16, 2026
Three results reversed between n≈50 and n=200, in three days
The finding
Every one of them looked solid and directional at small n. Four pre-registered runs were needed to settle a single benchmark question, because three earlier passes at n≈50 scattered across two different verdicts.
Over three days, three results we had computed at n≈40–50 reversed when we re-ran them larger. Not weakened — reversed.
What reversed
A benchmark comparison. Three separate runs at n≈50 scattered across accuracy ratios of 0.638 to 1.131 and returned two different verdicts about whether we were ahead. The question only settled at n=200, with bootstrap confidence intervals and a split-half stability check, and the answer was narrower than any of the small-n readings suggested.
A model capability claim. At n=40 we concluded a prediction head "cannot tell a working circuit from a broken one" — an effect size of Cohen's d = 1.18 on the signal we thought was carrying it. At n=200 the head's discrimination was real (AUROC 0.639, CI [0.561, 0.713], excluding chance), and the signal we had attributed it to scored 0.454 — below chance. We withdrew the finding publicly in our own audit.
Every one of these was seductive in the same way: the small-n reading was clean, directional, and told a story.
What we changed
Three things, and they are cheap:
Pre-specify the criterion before the run. Ours are written into a document and hashed before execution. That is what made the withdrawal above clean rather than embarrassing — we had said in advance what would count, so changing our minds afterwards was not available.
Report an interval, always. A bootstrap CI is what turned four benchmark runs into one answer. Where our intervals exclude parity we say so; where they do not, we say that too.
Check split-half stability. If the result is not stable across halves of the data, n is too small, and no amount of narrative makes it otherwise.
The line we now repeat
A point estimate without an interval is not a measurement.
We are not claiming this is novel — it is textbook. We are reporting that a team that already believed it still produced three reversals in one week, and that the small-n readings were convincing at the time. If you are running benchmarks at n≈50 on device data, the honest expectation is that some fraction of your current conclusions will not survive.