Why this direction
Almost everything we have found that was worth publishing has this shape. Not "our model beats yours" — that framing hands the opponent the benchmark, and it is the one we have failed at repeatedly. What survives is narrower and more durable: a measurement in common use turns out not to track the thing it is used to decide.
These results need no competitor to be beaten, which is what makes them worth running. They are also uncomfortable to publish, because the same lens applied inward keeps finding our own numbers. We publish those too — a methodology critique from someone who exempts themselves is not a methodology critique.
Findings
8- Aug 2026
Random splits overstate surrogate accuracy on qubit design data by 31×
A gradient-boosted model scores 0.12% under random K-fold and 3.81% under a grouped split that holds out whole geometries — the same model, the same data, a 31× difference. The field read this as needing more data. It is a split-methodology problem.
- Aug 2026
Our rule checker has 218 rules, and 218 is the wrong denominator
A median published quantum device exposes enough information to reach a verdict on 7–8 rules. Saturated with what its own field already knows how to publish, it reaches 28. The binding constraint on checkable hardware claims is reporting convention, not rule coverage.
- Aug 2026
A claim in our own copy went unverified for months because a probe timed out 0.7 seconds early
Our validators check whether a simulator is reachable before comparing against it. The check allowed 20 seconds; the environment took 20.7 to start. It failed closed, silently, and every run since had skipped the comparison while printing a line that read like a normal limitation.
- Aug 2026
Eight numbers beat a neural network on quantum weak-measurement records
A linear model on 8 time-binned means scores 0.820 held-out accuracy; an MLP on the raw 64-sample record scores 0.805. Naive summary statistics fail only for lack of temporal resolution — crude binning recovers everything, and the raw series adds nothing.
- Aug 2026
A gradient-boosted tree beats our 297-million-parameter model, including out of distribution
On RSFQ schema-derivable heads a gradient-boosted tree on tabular features wins by 15–89×, and it wins on the out-of-distribution cell-type split too — which is exactly where the large model was supposed to earn its keep.
- Aug 2026
Three results reversed between n≈50 and n=200, in three days
Every one of them looked solid and directional at small n. Four pre-registered runs were needed to settle a single benchmark question, because three earlier passes at n≈50 scattered across two different verdicts.
- Aug 2026
One conformal band across cell types hides a 23× spread in the band each type needs
Fitted per cell type, the conformal quantile is 0.033 for a JTL and 0.767 for a comparator-DFF — a 23× range a single global number cannot express. Splitting the band cut the cells we wrongly served as functional from 11 to 3.
- Aug 2026
We tried per-point conformal intervals and rejected them, and the rejection is in the product
Per-point intervals worked for capacitance — 22% and 10% narrower, width-error correlation +0.45 and +0.49. For resonator frequency they were wider, covered less (0.893 against a 0.95 guarantee), and width correlated with error at −0.03. The constant band ships.
Still open
What this direction has not answered, and what would settle each.
What must a device paper report for its claims to be independently checkable?
The median published device reaches 7–8 rules of our set. Saturated with what its own field already knows how to publish, it reaches 28 — so the binding constraint is reporting convention, not physics or tooling. Growing the database from 16 to 60 superconducting rules barely moved the number.
Settled if
A minimal reporting set lifts median reach from ~8 toward 28 across a large corpus — and closing it changes verdicts, not just counts.