The prescription — not the data — sets the proton's error bar
Closure testing against a planted truth then says which one to trust — and one sector still fails even that.
Eleven slides: two of orientation, three on the problem, four of evidence, two of verdict.
Talk over the picture: every group welds one recipe into its own code, next to its own model and its own data — so the recipe's effect was never separable. Colibri unwelds it: one program, plug in the model, run all three on the identical fit.
Then the numbers: 52 parameters, 3,092 real measurements, and the answer — 12 to 144 times apart.
Phase A found the methods disagree. Phase B asks whether more data fixes it
PHASE_B_PREREG.md before any compute was spent. None can be adjusted after seeing a result; the file is append-only.| card | what | state | headline |
|---|---|---|---|
| PPDF-30 | stage + hadronic port | ✅ | 4,615 pts vs NNPDF's 4,618 |
| PPDF-31 | χ² benchmark | ✅ | χ²/N = 1.2368 |
| PPDF-32 | Fisher ladder | ✅ corrected | 12 → 9 → 8 → 8 · saturates |
| PPDF-33 | three methods | ⚠️ partial | Hessian ✅ MC ✅ Bayesian ✗ |
| PPDF-34 | closure coverage | ⛔ | not attempted — gate failed |
| PPDF-35 | real global data | ✅ | χ²/N = 1.2380 |
| PPDF-36 | Hessian at its minimum | ✅ corrected | negatives unresolved |
| PPDF-37 | would more data help? | ⛔ | W+charm grids absent |
| PPDF-38 | which data holds it down | ✅ | ν-DIS & jets 4× Drell-Yan |
Seven cards carry results. The headline is a negative one about data, and a positive one about where to look instead.
Frame this as: the honest answer to 'throw more data at it' is no, and Phase B says why and what to do instead.
More data helps, then stops. The last 23 datasets removed nothing
R3 = 8 lands exactly on our pre-registered “cured” boundary. We report that as a boundary result, not a pass — eight directions remain unconstrained by every dataset we have.
If asked why this replaces 14 to 12 to 10 to 10: the old numbers came from forward differencing with an untested step size. Exact central differences give these. The saturation survived; the absolute counts did not.
Neutrino DIS and jets do the work. Drell-Yan is largely redundant
33 Drell-Yan datasets and 15 DIS datasets cost the same 2 directions each. Acquiring more of either would be close to worthless for this problem — which is itself a useful recommendation.
The null control is the point of the design: one of the five variants had a known answer before it ran, and it came back exactly right.
Three problems we told as one story live in three different sectors
sea_c6 — four of the top six are sea/strange.uv_c3, uv_c5, uv_del) — a different sector entirely.These need different fixes. Adding gluon-sensitive data cures degeneracy and would do nothing for the up-valence mixing that breaks the sampler.
MC coverage is 75.0% against a nominal 68.3% — mildly over-conservative, consistent with Phase A. That number is MAP-independent, which is why we trust it more than the Gaussian one.
The Hessian method does fail — but not for the reason we published
Conclusion unchanged, reason corrected. Phase A's own study pages hedged this correctly — “consistent with zero at FD resolution” — while the decks asserted it.
Right panel: the Bayesian arm. ESS 226 at 648 points falls to 3.2 at 3,089. Two measurements, not a fitted law, and they differ in preconditioner as well as data volume.
The parametrisation is not the bottleneck — a validation, not a discovery
Reproduced independently to 0.097% across a doubling of core count — which is also how we discovered χ² is only reproducible to ±0.7% across hardware.
If she pushes on parameter counts, concede immediately: raw-vs-effective is not a fair comparison and we are not resting anything on it.
The Bayesian arm did not converge, and we quote no band from it
| diagnostic | gate | measured | verdict |
|---|---|---|---|
| divergences | < 5% | 0.25% | ✅ pass — the integrator is healthy |
| ESSmin | ≥ 100 | 3.2 | ❌ fail — 31× short |
| split-R̂ | ≤ 1.05 | 2.45 | ❌ fail |
| posterior band | — | none quoted | kill condition applied as written |
| PPDF-34 coverage | — | NOT ATTEMPTED | recorded, not estimated |
Reaching the gate needs ~37,500 samples ≈ 14 h, measured not guessed. That run is in progress. Parallel chains would be 3× faster but are blocked by a 200 vCPU quota, not by capacity.
This is the slide to volunteer, not bury. Four confident-looking numbers turned out to be artefacts this week, none of which crashed; each was caught by a consistency argument.
One 37-point dataset would settle the question we cannot answer
| resource | status |
|---|---|
| yamldb dataset → FK mapping | ✅ defined |
| commondata (the measurements) | ✅ 5 datasets, 37 points |
| FK tables (theory predictions) | ❌ absent — no *WCHARM* file exists |
runcard fisher_R4.yaml | ✅ written, 87 datasets, ready |
| theory 40000000 coverage | 88 of 131 mappings lack FK tables |
| what it needs | pineappl/EKO grid generation, not a download |
ν-DIS is the only strangeness-sensitive class we have and it is tied for most efficient of any class. That makes the grid-generation cost an evidence-backed investment rather than a guess.
Also worth asking: raise the Nebius non-GPU vCPU quota above 200. That takes the Bayesian gate from 15 h to about 4 h, permanently.