partonmap
Study · 3.6

MSHT — fixed forms and the dynamic tolerance

The MRS→MRST→MSTW→MMHT→MSHT lineage, the Chebyshev fixed form, and the Hessian dynamic-tolerance method whose parametrization backs our fits.

Three collaborations have carried global PDF fitting for two decades, and they disagree — productively — on almost every methodological choice: how to parametrize the input, how to assign an uncertainty, how to treat heavy quarks. This section profiles the first of the three, the group whose exact input functional form is the backbone of the fits in this lab. Their initials have changed six times in thirty years, but the intellectual signature has not: a compact fixed functional form fitted by χ² minimisation, with uncertainties from the Hessian method and a home-grown fix — the dynamic tolerance — for the one place the textbook Hessian recipe breaks down on real global data.

The lineage — six acronyms, one thread

The story starts in Durham in the late 1980s. Alan Martin, Dick (R. G.) Roberts and James (W. J.) Stirling published the first NLO global analyses of parton distributions under the initials MRS. When Robert Thorne joined and the heavy-quark treatment matured, the sets became MRST (Martin–Roberts–Stirling–Thorne), running through the 1998–2006 vintages. The 2008 rewrite — new HERA combined data, a cleaner uncertainty prescription — was released as MSTW08 (Martin–Stirling–Thorne–Watt, with Graeme Watt). The next generation, MMHT14 (Harland-Lang–Martin–Motylinski–Thorne), added LHC data and, crucially, switched the input parametrisation to Chebyshev polynomials. The current flagship is MSHT20 (Bailey–Cridge–Harland-Lang–Martin–Thorne, Eur. Phys. J. C 81, 341 (2021)), and the frontier extension MSHT20aN3LO adds approximate N³LO splitting and coefficient functions with an honest theory-uncertainty estimate.

Why does the acronym keep changing?

Because it is a roster, not a brand. Each set is named for its living authors at publication time, so the letters track a UK-centred phenomenology school across generations — Durham's Institute for Particle Physics Phenomenology and University College London are the two poles. When you read "MSHT" you should hear "the MRST/MSTW tradition, current authors" — the methods are the continuity, and those methods are what this section is really about.

The parametrization philosophy — a fixed form, deliberately compact

The defining choice of this school is to write each input distribution at the starting scale $Q_0$ as an explicit algebraic function of $x$ with a modest number of free parameters, and to let DGLAP evolution (§3.1) carry it to all other scales. The generic template is the familiar

$$ x f(x, Q_0^2) \;=\; A\, x^{\delta} (1-x)^{\eta}\; \times\; \big[\text{a smooth shape factor}\big] $$

where the $x^{\delta}$ controls the small-$x$ (Regge-like) behaviour, the $(1-x)^{\eta}$ enforces the correct large-$x$ fall-off, and the bracket is the interesting part. We do not re-derive this form here — §3.1 built it from the physics of the endpoints. What matters for the group profile is how MSHT fills the bracket.

Since MMHT14, that shape factor is a sum of Chebyshev polynomials $T_i(y)$ in a transformed variable, roughly $y = 1 - 2\sqrt{x}$, mapping $x\in[0,1]$ onto the Chebyshev interval $y\in[-1,1]$:

$$ \big[\text{shape}\big] \;=\; 1 + \sum_{i=1}^{n} a_i\, T_i\!\big(y(x)\big), \qquad T_i(\cos\theta) = \cos(i\theta). $$

The motivation is numerical, and specific: on $[-1,1]$ the Chebyshev polynomials equi-oscillate — not only do their endpoints reach magnitude 1, but every interior maximum and minimum does too. A Legendre or plain-power basis has interior extrema that shrink, so adding terms buys you flexibility unevenly across the $x$ range. Chebyshev terms buy flexibility evenly, which is exactly what you want when the shape of $xf(x)$ must be pinned equally well from the valence peak to the sea. In practice a handful of Chebyshev coefficients per distribution — on top of the $A,\delta,\eta$ trio — suffices; the whole global fit carries of order a few dozen free parameters across all the input flavours combined.

y Tᵢ −101 +1−1 T₂ T₅ (equi-oscillating) Legendre-like: interior extrema shrink
Why Chebyshev. On $[-1,1]$ the Chebyshev polynomials touch $\pm 1$ at every extremum (orange, $T_5$), so successive terms add resolving power uniformly across the interval. A Legendre-style basis (grey dashed) has interior extrema that decay, giving less handle in the middle of the $x$ range. MSHT parametrizes the input shape in $y \approx 1 - 2\sqrt{x}$ so this uniform flexibility maps onto $xf(x)$.
The fixed-form bet: a compact, smooth, human-readable function with tens of parameters. Its strength is economy and control — the fit cannot wander into wiggly overfitted shapes, uncertainties are cheap to propagate, and every parameter has a rough physical meaning. Its risk is parametrization bias: the true PDF might not live inside the chosen function space, and no amount of data uncertainty captures a shape the form simply cannot represent. The Chebyshev switch and periodic form-flexibility studies are precisely the group's ongoing hedge against that risk. This is the exact axis on which they differ from NNPDF (§3.8), which replaces the fixed form with a neural network.

Hessian uncertainties and the dynamic tolerance

Given a fixed form, the fit is a χ² minimisation over the parameter vector $\mathbf{a}$. The uncertainty machinery is the Hessian method, and this school's signature refinement of it — the dynamic tolerance — is arguably their most cited methodological contribution.

Derivation — quadratic expansion and the eigenvector basis

Let $\mathbf{a}_0$ be the best-fit parameter vector, the minimum of the global $\chi^2(\mathbf{a})$. Expand around it. The gradient vanishes at the minimum, so to quadratic order

$$ \Delta\chi^2(\mathbf{a}) \;\equiv\; \chi^2(\mathbf{a}) - \chi^2(\mathbf{a}_0) \;=\; \tfrac12 \sum_{i,j} \frac{\partial^2 \chi^2}{\partial a_i\, \partial a_j}\bigg|_{\mathbf{a}_0} (a_i - a_{0,i})(a_j - a_{0,j}) \;\equiv\; \tfrac12\, \delta\mathbf{a}^{\!\top} H\, \delta\mathbf{a}, $$

where $H$ is the Hessian matrix of second derivatives. Surfaces of constant $\Delta\chi^2$ are ellipsoids; their axes are misaligned with the parameter axes and can be enormously elongated, because different parameter combinations are constrained by the data to wildly different precision.

Diagonalise: $H = V\, \Lambda\, V^{\!\top}$ with $\Lambda = \mathrm{diag}(\lambda_1,\dots,\lambda_N)$ and $V$ orthogonal. Define rescaled eigenvector coordinates $z_k = \sqrt{\lambda_k}\,\big(V^{\!\top}\delta\mathbf{a}\big)_k$. In these coordinates the ellipsoid becomes a sphere:

$$ \Delta\chi^2 \;=\; \sum_{k=1}^{N} z_k^2. $$

The $z_k$ are orthogonal, uncorrelated, and dimensionless. The error sets are built by stepping a distance $t_k$ along each eigendirection in turn — $z_k = \pm t_k$, all others zero — giving $2N$ PDF members $S_k^{\pm}$. Any observable's uncertainty is then the quadrature sum of its shifts across the pairs, $(\Delta\mathcal{O})^2 = \tfrac14\sum_k \big(\mathcal{O}(S_k^{+}) - \mathcal{O}(S_k^{-})\big)^2$. This is the Hessian recipe that LHAPDF ships and that experiments propagate. Everything hinges on one number per direction: the step size $t_k$. $\blacksquare$

How far is "$1\sigma$" along an eigenvector? The dogmatic answer from textbook statistics is $\Delta\chi^2 = 1$ (i.e. $t_k = 1$ for all $k$): move until the total χ² worsens by one unit. This is correct when the fit is to a single, mutually consistent dataset with well-understood Gaussian errors. It is badly wrong for a global PDF fit, and the reason is physical, not statistical bookkeeping.

Why is $\Delta\chi^2 = 1$ too aggressive for a global fit?

A global fit stitches together dozens of datasets — DIS, Drell–Yan, jets, top — taken by different experiments with independent systematics, and they are in mild tension: no single PDF pleases all of them at once, and each has a slightly different preferred region of parameter space. There are also stresses the χ² does not model: fixed-order perturbation theory is truncated, the input form has finite flexibility, and some quoted experimental errors are optimistic. Under $\Delta\chi^2 = 1$ the error band would be so tight that neighbouring datasets' central fits fall outside each other's uncertainties — a self-contradiction. The band is smaller than the spread of the data it is meant to describe. So the tolerance must be inflated; the only honest question is by how much, and determined how.

The MSTW08 answer — carried into MMHT14 and MSHT20 — is the dynamic tolerance: instead of one global inflated $\Delta\chi^2$ chosen by hand, each eigenvector direction gets its own tolerance $t_k$, read off from the data itself. The procedure is per-direction and per-dataset. Walk outward along eigenvector $k$. Track, for every individual dataset $n$, how its $\chi^2_n$ degrades. Each dataset defines a boundary — the point along that direction where the fit to that dataset becomes unacceptable at a chosen confidence level (68% or 90%, via the appropriate $\Delta\chi^2_n$ percentile for its number of points). The tolerance $t_k^{\pm}$ for stepping in the $+$ and $-$ senses of direction $k$ is set by the first dataset to object — the most constraining one in that direction. Directions tightly pinned by one experiment get a small tolerance; loosely constrained directions get a large one. The tolerances emerge from the fit rather than being imposed, and in MSTW08 they came out around $\Delta\chi^2 \sim 3\text{–}4$ on average (i.e. $t_k$ of a few), varying direction by direction — a far cry from $\Delta\chi^2 = 1$.

a₂a₁ Δχ² surface a₀ (min) z₁ (loose) large t₁ z₂ (tight) small t₂ dataset A objects → t₁⁺ distance along z₁ χ²ₙ t₁ set by first crossing
Dynamic tolerance, geometrically. Left: the $\Delta\chi^2$ ellipse in parameter space, with the Hessian eigenvectors $z_1$ (loose, long axis) and $z_2$ (tight, short axis). Right: walking outward along $z_1$, every dataset's own $\chi^2_n$ rises at its own rate; the tolerance $t_1^{+}$ is set where the first dataset crosses its acceptability boundary (red). Each direction, and each sign, gets its own tolerance — hence "dynamic". A flat $\Delta\chi^2 = 1$ would draw the same tiny circle in every direction and grossly under-cover.

Heavy flavours — the Thorne–Roberts GM-VFNS

Charm and bottom quarks are a genuine complication: near threshold, $Q^2 \sim m_h^2$, the heavy quark is best treated as a massive final-state particle (a fixed-flavour-number scheme, FFNS, with $n_f$ light flavours), whereas at $Q^2 \gg m_h^2$ the large logs $\ln(Q^2/m_h^2)$ must be resummed by treating the heavy quark as an active massless parton (a zero-mass VFNS). Neither limit is right across the whole range. A general-mass variable-flavour-number scheme (GM-VFNS) interpolates smoothly between them, keeping the full mass dependence near threshold and matching onto the massless resummed description at high scale.

This group uses its own construction, the Thorne–Roberts scheme, in its refined "TR′" form. TR′ fixes the number of active flavours to change by one as $Q^2$ crosses each heavy-quark threshold, adds the higher-order terms needed to make the structure functions and their slopes continuous through the transition, and thereby avoids the kinks a naïve scheme change would leave in the observables. The heavy-quark masses $m_c, m_b$ enter as fit inputs, and the scheme is one of the standard, well-tested GM-VFNS choices alongside ACOT and its variants used by other groups.

Why this group matters to our lab

Every fit in this project sits on MSHT's shoulders in a concrete way: the exact input functional form used here is the MSHT parametrization, ported into the Colibri framework. Colibri is model-agnostic — a user supplies any differentiable map from parameters to $xf(x, Q_0)$ — and the model we supply is the MSHT Chebyshev-based form. So when our Results pages show a Bayesian posterior or a Monte-Carlo ensemble of PDFs, the shape space being explored is MSHT's fixed form; what changes between our three methods is only how that parameter space is sampled and how the uncertainty is defined, not the functional backbone.

That makes MSHT's uncertainty method a live participant in our own comparison, not just history. The Hessian method derived above reappears in §5.1 as one of the three legs of our headline methodological study (PPDF-15 / PPDF-17): the same fixed form is fitted and its uncertainty read out three ways — the Hessian (with tolerance) profiled here, a Monte-Carlo replica ensemble, and full Bayesian nested sampling. Because the form is held fixed across all three, any difference in the resulting error bands is a clean measurement of the uncertainty-estimation method itself, with the parametrization-bias axis controlled. Understanding what the dynamic tolerance is doing — and where a flat $\Delta\chi^2 = 1$ would mislead — is therefore a prerequisite for reading our own three-way verdict.

Hold on to four things. (1) MSHT is the current name of the MRS→MRST→MSTW→MMHT lineage — a UK phenomenology school defined by its methods, not its letters. (2) Its parametrization is a compact fixed form, Chebyshev polynomials in $y\approx 1-2\sqrt{x}$, prized for even flexibility and control, at the cost of possible form-bias. (3) Its signature is the dynamic tolerance: build Hessian eigenvector error sets, but let each direction's step size be set by when the first individual dataset objects, because global-fit tensions make the textbook $\Delta\chi^2 = 1$ far too tight. (4) This is not background reading — MSHT's exact form is our fit backbone via Colibri, and its Hessian method is one of the three legs we weigh against each other in §5.1.