partonmap
Research read-out for M. Ubiali · HEP-PBSP, Cambridge · 5 August 2026

The prescription — not the data — sets the proton's error bar

BEFORE MSHT20 Chebyshev, 52 par. Hessian only their own data CT18 Bernstein, 29 par. Hessian only their own data NNPDF4.0 neural net, 763 par. Monte-Carlo only their own data model + method + data all change together — no clean comparison Colibri AFTER Colibri — one program proton model you plug any one in the same data held fixed one switch, three settings Hessian Monte-Carlo Bayesian one model, one dataset, three methods — finally comparable
52
free parameters — the MSHT20 form, a realistic proton, not a toy
3,092
real measurements — full-DIS, 19 datasets, Q₀ = 1.65 GeV
3 → 1
prescriptions, now in one program — so the recipe can finally be isolated
12–144×
how far apart their error bars come out — 28 pre-registered runs

Closure testing against a planted truth then says which one to trust — and one sector still fails even that.

xpartonmap.com · with M. Ubiali, HEP-PBSP Cambridge · full 27-slide read-out and results ledger online
Speaker notes

Eleven slides: two of orientation, three on the problem, four of evidence, two of verdict.

Talk over the picture: every group welds one recipe into its own code, next to its own model and its own data — so the recipe's effect was never separable. Colibri unwelds it: one program, plug in the model, run all three on the identical fit.

Then the numbers: 52 parameters, 3,092 real measurements, and the answer — 12 to 144 times apart.

Phase B · the question

Phase A found the methods disagree. Phase B asks whether more data fixes it

the setup82 datasets, 4,615 points — the NNPDF4.0-like global set. Theory, model and charm treatment held identical to Phase A, so data is the only variable.
the pre-registrationEvery gate was written into PHASE_B_PREREG.md before any compute was spent. None can be adjusted after seeing a result; the file is append-only.
what changedThree previously published numbers were corrected. In every case the conclusion survived and the absolute numbers did not. Those corrections are on the cards, dated.
cardwhatstateheadline
PPDF-30stage + hadronic port4,615 pts vs NNPDF's 4,618
PPDF-31χ² benchmarkχ²/N = 1.2368
PPDF-32Fisher ladder✅ corrected12 → 9 → 8 → 8 · saturates
PPDF-33three methods⚠️ partialHessian ✅ MC ✅ Bayesian ✗
PPDF-34closure coveragenot attempted — gate failed
PPDF-35real global dataχ²/N = 1.2380
PPDF-36Hessian at its minimum✅ correctednegatives unresolved
PPDF-37would more data help?W+charm grids absent
PPDF-38which data holds it downν-DIS & jets 4× Drell-Yan

Seven cards carry results. The headline is a negative one about data, and a positive one about where to look instead.

x1 · 82 datasets · 4,615 points · theory held fixed
Speaker notes

Frame this as: the honest answer to 'throw more data at it' is no, and Phase B says why and what to do instead.

the headline result

More data helps, then stops. The last 23 datasets removed nothing

what is plottedUnconstrained parameter directions against how much data is in the fit. 12 → 9 → 8 → 8.
the findingDrell-Yan/W/Z removes three. Jets and photon remove one. Top-quark data — 23 datasets, the largest block — removes zero.
why trust itFisher information at a planted truth: noise-free, no sampler. Exact central differences, converged (step 10⁻⁴ ≡ 10⁻⁵), zero negative eigenvalues.
More data helps, then stops. The last 23 datasets removed nothing
OUR MEASUREMENTexact Fisher spectrum, four dataset sizes

R3 = 8 lands exactly on our pre-registered “cured” boundary. We report that as a boundary result, not a pass — eight directions remain unconstrained by every dataset we have.

x2 · the degeneracy saturates at 8 of 52
Speaker notes

If asked why this replaces 14 to 12 to 10 to 10: the old numbers came from forward differencing with an untested step size. Exact central differences give these. The saturation survived; the absolute counts did not.

which data actually works

Neutrino DIS and jets do the work. Drell-Yan is largely redundant

the experimentRemove one data class at a time from the full set and recount. Marginal value — what each class uniquely contributes given everything else.
the resultν-DIS and jets: 0.250 flat directions per dataset. Charged-lepton DIS 0.133. Drell-Yan 0.061. Top 0.000.
the built-in checkRemoving top costs zero — matching the saturation result from the other direction. Two independent methods agreeing exactly.
Neutrino DIS and jets do the work. Drell-Yan is largely redundant
OUR MEASUREMENTfive leave-one-out variants, all converged

33 Drell-Yan datasets and 15 DIS datasets cost the same 2 directions each. Acquiring more of either would be close to worthless for this problem — which is itself a useful recommendation.

x3 · ν-DIS and jets are ~4× Drell-Yan per dataset
Speaker notes

The null control is the point of the design: one of the five variants had a known answer before it ran, and it came back exactly right.

the structural finding

Three problems we told as one story live in three different sectors

degeneracyThe flat directions sit in strange / sea.
method disagreementMonte-Carlo and Hessian bands agree on a typical parameter (1.09×) and differ by 31× on sea_c6 — four of the top six are sea/strange.
sampler mixingWorst in up-valence (uv_c3, uv_c5, uv_del) — a different sector entirely.
Three problems we told as one story live in three different sectors
OUR MEASUREMENT39 MC replicas vs the Hessian band at the MAP

These need different fixes. Adding gluon-sensitive data cures degeneracy and would do nothing for the up-valence mixing that breaks the sampler.

x4 · flatness, disagreement and mixing are not the same pathology
Speaker notes

MC coverage is 75.0% against a nominal 68.3% — mildly over-conservative, consistent with Phase A. That number is MAP-independent, which is why we trust it more than the Gaussian one.

a correction we owed

The Hessian method does fail — but not for the reason we published

what we said“10–11 of 52 curvature eigenvalues come out negative.” It was asserted in three places.
what is trueMeasured at a converged minimum, λ′min swings 15× non-monotonically with step size. Those negatives are below the numerical resolution of the method that produced them.
what survivesThe condition number is 1.6–1.8×10¹¹, stable at every step size. A matrix that ill-conditioned cannot be inverted into a covariance whatever the signs are.
The Hessian method does fail — but not for the reason we published
OUR MEASUREMENTleft: real data vs an exact minimum · right: sampler scaling

Conclusion unchanged, reason corrected. Phase A's own study pages hedged this correctly — “consistent with zero at FD resolution” — while the decks asserted it.

x5 · ill-conditioning, not wrong-sign curvature
Speaker notes

Right panel: the Bayesian arm. ESS 226 at 648 points falls to 3.2 at 3,089. Two measurements, not a fitted law, and they differ in preconditioner as well as data volume.

the external anchor

The parametrisation is not the bottleneck — a validation, not a discovery

the numberχ²/N = 1.238 on all 82 datasets, against NNPDF4.0's published 1.17. Pre-registered pass gate 1.35.
what it is forPhase A had no external anchor on real data. This closes the objection that our results are artefacts of too rigid a model.
what it is notMSHT20 already fits global data at this quality. We do not claim 14× fewer parameters as a headline: NNPDF's 763 are regularised network weights and our own effective dof is ~10–14.
The parametrisation is not the bottleneck — a validation, not a discovery
OUR MEASUREMENTvs a published global fit, plus the reproduction check

Reproduced independently to 0.097% across a doubling of core count — which is also how we discovered χ² is only reproducible to ±0.7% across hardware.

x6 · χ²/N = 1.238 ± 0.7% (hardware)
Speaker notes

If she pushes on parameter counts, concede immediately: raw-vs-effective is not a fair comparison and we are not resting anything on it.

what failed, stated plainly

The Bayesian arm did not converge, and we quote no band from it

the gateESS ≥ 100, split-R̂ ≤ 1.05, divergences < 5%. Written before the run.
the resultDivergences 0.25% — the sampler is healthy. But ESSmin = 3.2 and R̂ = 2.45. Mixing fails, not the integrator.
the consequenceNo posterior band anywhere. Coverage (PPDF-34) is recorded NOT ATTEMPTED rather than estimated — a coverage number from an unconverged chain measures sampler failure, not method performance.
diagnosticgatemeasuredverdict
divergences< 5%0.25%✅ pass — the integrator is healthy
ESSmin≥ 1003.2❌ fail — 31× short
split-R̂≤ 1.052.45❌ fail
posterior bandnone quotedkill condition applied as written
PPDF-34 coverageNOT ATTEMPTEDrecorded, not estimated

Reaching the gate needs ~37,500 samples ≈ 14 h, measured not guessed. That run is in progress. Parallel chains would be 3× faster but are blocked by a 200 vCPU quota, not by capacity.

x7 · the kill condition, applied as written
Speaker notes

This is the slide to volunteer, not bury. Four confident-looking numbers turned out to be artefacts this week, none of which crashed; each was caught by a consistency argument.

the ask

One 37-point dataset would settle the question we cannot answer

the open questionIs the residual degeneracy a data problem or structural to the 52-parameter form? Our leave-one-out says data still has purchase, so it is not purely structural.
the decisive testR4 = R3 + W+charm — the canonical direct probe of the strange PDF, absent from the global set. Runcard written; runs in five minutes.
the blockerW+charm FK tables do not exist in theory 40000000 — 88 of its 131 dataset mappings lack them. This needs pineappl/EKO grid generation, not a download.
resourcestatus
yamldb dataset → FK mapping✅ defined
commondata (the measurements)✅ 5 datasets, 37 points
FK tables (theory predictions)absent — no *WCHARM* file exists
runcard fisher_R4.yaml✅ written, 87 datasets, ready
theory 40000000 coverage88 of 131 mappings lack FK tables
what it needspineappl/EKO grid generation, not a download

ν-DIS is the only strangeness-sensitive class we have and it is tied for most efficient of any class. That makes the grid-generation cost an evidence-backed investment rather than a guess.

x8 · the one thing that would move this forward
Speaker notes

Also worth asking: raise the Nebius non-GPU vCPU quota above 200. That takes the Bayesian gate from 15 h to about 4 h, permanently.

PDF ↓PDF + notes ↓
1 / 12