The oldest continuous line in this field. It began in 1987 as MRS, and the same collaboration has been producing proton distributions ever since, through four changes of name as its membership changed. MSHT20 is the current version: a polynomial description of the proton fitted to 4,363 measurements, with uncertainties from the Hessian method and a distinctive twist — the size of the error bar is not chosen by hand, but worked out from how much the datasets disagree with each other.
If you read nothing else on this page.
MSHT20 writes the proton down as a handful of smooth mathematical curves — one for the gluons, one for the up quarks, and so on — each described by about fifty adjustable numbers. It then tunes those numbers until the predictions match 4,363 real measurements from accelerators around the world.
For the uncertainty, it asks: how far can I push these numbers before the fit visibly gets worse? That is the Hessian method. Their particular contribution is deciding how much worse counts as too much — and rather than picking a number, they derive it from the fact that different experiments pull in different directions. The more the data disagree, the wider the error bar becomes.
The flagship paper is Bailey, Cridge, Harland-Lang, Martin and Thorne (2021), and its headline finding is that for the first time the higher-order (NNLO) calculation is strongly favoured over the lower-order one — the fit improves by around 700 units of χ², against roughly 200 in their previous release.
Everyone assumes MSHT is four surnames. It is not — and that change was deliberate.
Every previous name in this lineage was the authors' initials. As people joined and left, the acronym changed. By 2020 the group decided this had run out of road and picked a name that describes the method instead of the people:
| set | from | what the letters mean |
|---|---|---|
| MRS | 1987 | Martin, Roberts, Stirling |
| MRST | 1998 | Martin, Roberts, Stirling, Thorne |
| MSTW08 | 2009 | Martin, Stirling, Thorne, Watt |
| MMHT14 | 2015 | Martin, Motylinski, Harland-Lang, Thorne |
| MSHT20 | 2021 | Mass Scheme Hessian Tolerance — not surnames |
Each word is a claim about what makes their analysis theirs. Mass Scheme: they were the first group to obtain PDFs using a general-mass variable-flavour-number scheme, which handles the fact that quarks have mass and that heavier ones only participate above a certain energy. Hessian: they have always used curvature at the best fit for uncertainties. Tolerance: the dynamic procedure described in §06.
The MSHT20 paper carries a dedication: to James Stirling and Dick Roberts, both of whom died in the two years before publication. They were founding members of the collaboration, which began with the presentation of the first ever NLO PDF sets in 1987.
Five authors across three British universities.
| author | affiliation on the MSHT20 paper |
|---|---|
| S. Bailey | Rudolf Peierls Centre for Theoretical Physics, Oxford |
| T. Cridge | Department of Physics and Astronomy, University College London |
| L. A. Harland-Lang | Rudolf Peierls Centre for Theoretical Physics, Oxford |
| A. D. Martin | Institute for Particle Physics Phenomenology (IPPP), Durham |
| R. S. Thorne | Department of Physics and Astronomy, University College London |
Those are the affiliations as printed in 2021. People have since moved — Harland-Lang is now at UCL, and Cridge has appeared at DESY and Antwerp on more recent papers. The group's home is hep.ucl.ac.uk/msht.
61 datasets, 4,363 measurements, four decades of accelerators.
Not 4,363 collisions — accelerators record billions. Each "point" is a bin: a published number summarising a huge number of collisions in one small window of energy and angle, together with its error bar. The fit compares its prediction against each of those numbers.
| category | sets | what it is |
|---|---|---|
| Fixed-target lepton scattering | 10 | Firing muons or electrons at hydrogen and deuterium targets — BCDMS, NMC, E665, SLAC. The oldest and still among the most constraining data. |
| Neutrino scattering | 6 | NuTeV, CHORUS, CCFR — including "dimuon" events, one of the few direct handles on the strange quark. |
| Fixed-target Drell–Yan | 2 | E866/NuSea — sensitive to the difference between anti-down and anti-up quarks. |
| HERA | 8 | The electron–proton collider at DESY. Combined H1+ZEUS measurements, including charm production. |
| Tevatron | 8 | CDF and DØ — jets, W asymmetry, Z rapidity. |
| LHC | 27 | ATLAS, CMS, LHCb at 7 and 8 TeV: W and Z production, high-mass Drell–Yan, inclusive jets, Z transverse momentum, W+charm, and top-quark pair production. |
The split is 1,328 LHC points and 3,035 non-LHC points. Fit quality at NNLO is χ²/N = 1.17 (5,121.9 over 4,363); at NLO it is 1.33. That gap is the paper's headline: the LHC subset in particular goes from 1.79 at NLO to 1.33 at NNLO, which is what makes the higher-order calculation "strongly favoured" for the first time.
Polynomials — chosen, tested, and deliberately not a neural network.
Each distribution is written at a starting energy Q₀ = 1 GeV in the form
x·f(x, Q₀²) = A (1−x)η xδ × [ 1 + Σ aᵢ Tᵢ(y) ] , y = 1 − 2√x
The first part encodes long-known physical behaviour: PDFs fall off like a power of x at small momentum fraction and vanish like a power of (1−x) at large. The bracket is a Chebyshev polynomial series that lets the curve take whatever shape the data want on top of that. Chebyshev polynomials are used because they are numerically well-behaved and don't develop wild oscillations.
| starting scale | Q₀² = 1 GeV² |
| polynomial order | 6 Chebyshevs by default (up from 4 in MMHT14) — enough for better than 1% precision across most of the x range |
| free parameters | 52, after imposing the sum rules |
| sum rules imposed | valence-quark counting, total momentum, and zero net strangeness |
| heavy quarks | generated perturbatively, not fitted. mc = 1.4 GeV, mb = 4.75 GeV, in their "optimal TR′" general-mass scheme |
| strong coupling | αS(MZ²) = 0.118 default; free-fit best values 0.1175 (NNLO), 0.1203 (NLO) |
This is the same parametrisation, and the same 52 free parameters, that our own three-method comparison uses. When the Explorer's Tier 2 says "MSHT20 parametrisation, 52 free parameters", it means this — we adopted their model of the proton and changed only how the uncertainty is computed. See the Colibri page for why.
The "T" in the name. This is the part that distinguishes them.
The Hessian method measures the curvature of the fit quality around its best point. Steep curvature in some direction means the data pin that combination of parameters down tightly; shallow curvature means it is poorly determined. Inverting the curvature matrix turns this into an uncertainty.
The textbook rule says the 1σ error bar is where χ² rises by 1. MSHT do not use that rule, and the reason is important:
The textbook Δχ² = 1 assumes every dataset is mutually consistent and every error bar is correct. In a global fit they are not: different experiments genuinely pull the answer in different directions. If you used Δχ² = 1 you would quote an uncertainty far smaller than the actual spread of what the data support.
So MSHT inflate it — by a tolerance factor T, where the 1σ band is at Δχ² = T². The innovation is that T is not chosen by hand. For each direction in parameter space they ask how far they can move before any individual dataset becomes unacceptably badly described, and take that as the limit. The tolerance is therefore read off from the tensions in the data.
| eigenvector pairs | 32 (up from 20 in MSTW08, 25 in MMHT14) |
| total members | 65 — 64 eigenvector directions plus the central fit |
| confidence level | 68% |
| tolerance in practice | T typically ≈ 2–5 per direction; mean T² ≈ 13 at NNLO (range of T: 0.52 – 6.65) |
You will see "MSHT uses T² = 10" in secondary sources. That number does not appear in the MSHT20 paper. The tolerance is per-eigenvector and tabulated direction by direction; the figures above are computed from those published tables. Quote it as a range, not a constant.
Why stop at 32 pairs? Because the Hessian approach relies on the fit behaving quadratically near its minimum. Beyond 32 directions that approximation visibly breaks down — the paper notes it would "quickly result in mismatches of much more than a factor of 2". That limitation is precisely the subject of our own work, described on the Colibri page.
The actual released MSHT20 grid, drawn in your browser.
Pick an energy. Watch the gluon grow at small x as the energy rises: that is DGLAP evolution, and it is calculated rather than fitted.
Same quantity, three independent analyses.
MSHT20 solid with the heavier band; the others dashed. Switch parton species — agreement is excellent for the gluon and much weaker for strangeness.
Note that all three bands here have been put on a common 68% footing. CT18 publishes at 90% confidence, so its band is rescaled by 1/1.645 before plotting; without that step CT18 would look about 64% more uncertain than it is. See the CT18 page.
Downloadable grids in LHAPDF format — the standard the whole field reads.
| grid | what it is |
|---|---|
| MSHT20nnlo_as118 | The default. NNLO, αS = 0.118, central + 64 eigenvectors. This is the set plotted on this site. |
| MSHT20nlo_as118 / _as120 | NLO. The αS = 0.120 variant exists because the NLO fit genuinely prefers a higher coupling. |
| MSHT20lo_as130 | Leading order, 30 eigenvector pairs. |
| MSHT20an3lo_as118 | Approximate N³LO — the first global PDF fit at this order. 52 eigenvector pairs: 32 for the PDFs plus 20 representing theory uncertainty. |
| MSHT20qed_nnlo | Includes a photon distribution inside the proton, built on the LUXqed formulation. Elastic/inelastic and neutron variants exist. |
| αS and quark-mass scans | Central-only sets scanning αS, mc and mb, plus 3- and 4-flavour variants. |
| MSHT20xNNPDF40_an3lo | A joint MSHT+NNPDF combined aN³LO set — the two groups collaborating despite the methodological argument in §10. |
The group documents two corrections: an αS interpolation-grid formatting problem fixed in data version 2 of the baseline sets, and an output error that duplicated 1–2 eigenvectors in MSHT20nnlo_as118 and three others, corrected in later data versions (stated impact ≈ 1% of the PDF error). Check your local data version if precision matters.
This is a live scientific argument, conducted in public and in print.
MSHT's own comparison, in their words, is that NNPDF "combine a Monte Carlo representation of the probability measure in the space of PDFs, with the use of neural networks", whereas MSHT "use parameterisations based on Chebyshev polynomials". Of CT18 they note it also uses a tolerance and a polynomial form, "albeit with a rather more restrictive parameterisation in general" — for example CT18 assumes the strangeness asymmetry s − s̄ is exactly zero, where MSHT fit it.
The sharpest disagreement is with NNPDF, over exactly the question this whole site is about: how big should the error bar be? In 2024 MSHT published a closure test that went further than argument — they refitted using NNPDF4.0's own data and theory settings, changing only the fitting methodology, so that the comparison was genuinely like-for-like. Their conclusion, quoted:
"The NNPDF4.0 uncertainties are found to be broadly in line with the MSHT results if a textbook T² = 1 tolerance is applied, but to be significantly smaller if a tolerance typical of the MSHT20 fit is applied. This points to an inherent inconsistency between these approaches."
They also found that their polynomial parametrisation reproduced a known input to well within its quoted uncertainties — evidence that "parameterisation inflexibility in the MSHT20 fit is not a significant issue in the data region", answering a long-standing criticism from the neural-network camp.
They cooperate as well as argue. MSHT20 is one of the three inputs to the PDF4LHC21 combination alongside CT18 and NNPDF3.1, and MSHT and NNPDF jointly produced combined approximate-N³LO sets in 2024. The disagreement is methodological and openly conducted, not factional.
They are candid about tensions inside their own fit too: between the precise ATLAS 7 TeV W and Z data and both the older BCDMS data and the DØ W asymmetry; and the long-running conflict between the strangeness preferred by ATLAS W/Z data and that preferred by neutrino dimuon data. Those tensions are exactly what the dynamic tolerance converts into a wider error bar.
The flagship, then the extensions that followed it.