The prescription — not the data — sets the proton's error bar
Closure testing against a planted truth then says which one to trust — and one sector still fails even that.
Eleven slides: two of orientation, three on the problem, four of evidence, two of verdict.
Talk over the picture: every group welds one recipe into its own code, next to its own model and its own data — so the recipe's effect was never separable. Colibri unwelds it: one program, plug in the model, run all three on the identical fit.
Then the numbers: 52 parameters, 3,092 real measurements, and the answer — 12 to 144 times apart.
Proton structure cannot be calculated. It can only be fitted
Each curve also carries a published ± of a few percent — 0.24% of the height of this plot, so it is not drawn. How that ± is arrived at, and how much the choice of method changes it, is the subject of this deck.
Two confusions to head off. First: readers assume f(x) is bounded by 1 and are derailed when x·f reads 10 at x = 10⁻⁴ (f = 100,000 there). Second: on a log axis the small-x tail looks dominant but contributes little.
Bands are deliberately not drawn here — MSHT20's gluon ± is a median 1.4%, invisible on a 5.3-decade axis. Drawing it would suggest proton uncertainties are tiny and settled, which is the opposite of the setup. Uncertainty arrives on the next slide, where it is legible.
If asked about Q²: boosting the proton changes nothing — x is a ratio, so it is boost-invariant. The total momentum stays exactly 1 at every scale.
The groups agree on the proton. They do not agree on the error bar
Comparing published sets can never isolate the recipe: the groups also differ in data, model and code.
This is the slide that motivates the whole project. The recipe's effect was structurally unmeasurable until one code could run all three on the same fit.
Three prescriptions, three distinct failure modes
Does: measures the curvature of χ² at the minimum and inverts it.
Assumes: a positive-definite bowl.
Fails when: a flat direction gives zero or negative curvature — H⁻¹ is then meaningless. Not a wide band: no band at all.
Does: refits many noisy copies of the data; the spread is the uncertainty.
Assumes: frequentist scatter = uncertainty.
Fails when: refits roll to the prior bound — it then measures the box, not the data.
Does: maps the posterior by sampling; the spread is the uncertainty.
Assumes: a prior, and a converged sampler.
Fails when: sampling is under-converged — it reports a spuriously tight band.
Insist on the Hessian wording: it does not give a wide error bar, it gives none. The covariance is not positive-definite; there is nothing to quote.
What I actually ran
| parametrisation | MSHT20 — 52 free parameters, 4 analytic sum rules |
| framework | Colibri — JAX, auto-differentiable likelihood |
| data | full-DIS: 19 datasets, ≈3,092 points after cuts |
| which data | SLAC, BCDMS, NMC fixed-target F₂; HERA NC/CC; CHORUS, NuTeV ν-DIS |
| theory | FK tables, theoryid 40000000, t₀ covariance |
| starting scale | Q₀ = 1.65 GeV |
| runs | 28 experiments, pre-registered criteria, one ledger |
Colibri's published demonstration used a 13-parameter benchmark with no flat directions. In that regime all three prescriptions agree — as they must.
My job was to take it to a realistic parametrisation, where the degeneracy is real. The problem only becomes visible at that data-to-parameter ratio.
The rigid 13-parameter form floors at χ²/N ≈ 5.6 on this data; MSHT20 reaches χ²/N ≈ 1.0, ΔlogZ ≈ +1450. The realistic model is not optional — it is the only one that fits.
Scope: this is a DIS-only study. Every statement here is about methods at this data-to-parameter ratio — not about MSHT20's published global fit.
The scope caveat matters. Do not let it be read as a criticism of MSHT20's released uncertainties.
First: the machinery is correct. Without this, divergence is indistinguishable from a bug

The implementations are right. Everything that follows is therefore physics, not a coding error.
If she challenges the whole result, this is the slide to return to. The methods were validated where they should agree, before being trusted where they do not.
The degeneracy is measured, not asserted

Two unknowns, one measurement that sees only their sum: total = 10.0 ± 0.5. (5,5), (2,8) and (8,2) all fit identically; (7,5) is off by four error bars. Moving along the sum is free; moving across it is not.
The data has no opinion about the split — so the prescription supplies one.
The Fisher information spectrum at the planted truth — how sharply the data pins down each independent direction. Noise-free, so this is geometry, not luck.
12 of 52 directions unconstrained on DIS. Spectrum spans ~11 orders of magnitude. Flattest: s−s̄, then strange sea, then high-x down-valence.
Adding Drell–Yan, W and Z moves it from 14 to 12. Two directions. An independent earlier analysis found 16→12 at condition ≈10¹³ — the same partial-cure conclusion.
Emphasise "measured". The strange sector was identified by the Fisher analysis independently of the coverage result on slide 11 — two routes to the same place.
Same model, same data, only the recipe swapped: error bars 12–144× apart
The Hessian does not give a wide band — it gives none. On real data its curvature matrix has condition ≈ 6×10¹¹: inverting it is numerically meaningless, so there is nothing to quote. (An earlier version also cited “10 of 52 negative eigenvalues”; that count sits at the finite-difference resolution limit. The condition number is the argument and is unaffected.)
Be ready for "which dataset?" — the 12–144× is the full-DIS fit; the five panels are the 648-point subset, where the ratios are smaller. Both are in the ledger.
Note also 11/52 negative eigenvalues appears in an earlier run (PPDF-14). The 10/52 quoted here is the real-data blind result.
I built the test that could have destroyed the headline, and ran it

My first result claimed Bayesian bands collapsing to ~2% of the PDF value. That was a sampling artefact at ESS ≈ 24.
Coverage climbed 6.5% → 49% → 51%, then went flat. Naive run: ESS ≈ 8, 262/400 divergences. Converged: ESS median 226, 41/3000 divergences, R̂ < 1.05.
Saturation is the proof: ESS 51 → 226 moved coverage by under 2 points.
A converged real-data run on a different machine with a different initialisation (ESS ≈ 181) reproduced the earlier band exactly.
The dramatic magnitude did not survive. The divergence did. Both are in the ledger, with dates.
Do not soften this slide. It is the strongest evidence of good faith in the deck.
Closure testing turns disagreement into a verdict
T8 covers 38%, failing 4 of 5 draws and reaching only 32% at the best convergence — the same sector the Fisher spectrum flagged independently. But the strange directions also carry the lowest per-direction sampling (ESS ≈ 10–58), so physical versus numerical is not yet fully separated.
If asked whether 57.9% vs 68% is a failure: no — it is inside the pre-declared gate and within draw scatter. The robust signal is T8, which fails 4 of 5 draws.
What stands, what does not, and what I need
- The three prescriptions genuinely diverge, 12–144× on real data.
- The cause is a measured degeneracy — 12 of 52 directions, worst in strangeness.
- Closure adjudicates: the converged Bayesian is the one to use.
- An over-claim of mine was caught and withdrawn by my own test.
Two transferable results, easy to miss: nested sampling cannot converge on this posterior — 300→2400 live points, 15.1k→56.3k iterations, ESS = 1 after five days, while Fisher-preconditioned NUTS converges in an hour; and single-draw closure coverage cannot be trusted — 47%↔76% from noise alone.
- The Hessian was never given its production form — real fits regularise; mine did not.
- The Tier-2 sweep came back under-converged (ESS ≈ 30), so "more data" is unresolved.
- One functional form only; prior widths unscanned.
- The real-data verdict is transferred from closure — there is no external anchor on real data.
1 · Finish the sweep — a converged whitened Tier-2 run, ~3 h of compute, settles whether DY/W/Z cures the strange under-coverage.
2 · A fair Hessian rematch — score the regularised version production fits actually use.
3 · The ask: ν-DIS dimuon and W+charm FK tables, to constrain strangeness directly rather than inferring it.
What I would most value from you: whether the strange result is worth pursuing, and which of these to run first.
End on the ask, not the result. She is the one who can judge whether the strange residual is physically interesting.