partonmap
Map ▸ collaboration 3 of 4 ▸ the machine-learning line

NNPDF4.0

Neural Network PDFs — Edinburgh · Milan · Cambridge · Amsterdam · and beyond

The methodological outlier, and deliberately so. Where MSHT20 and CT18 write the proton as a polynomial with a few dozen tunable numbers, NNPDF replaces the formula with a neural network of 763 parameters — on the argument that any chosen formula smuggles in an assumption about the answer. And where the other two inflate their error bars to account for datasets that disagree, NNPDF does not. That single choice is the sharpest open argument in the field, and it is why their bands are the narrowest on this site.

basedEurope + Singapore, UAE uncertaintyMonte-Carlo replicas tolerancenone — by design data4,618 points (NNLO) parameters763 (neural network) members101 (100 replicas)

In one paragraph

If you read nothing else on this page.

plain English

Every group has to guess a shape for the proton's distributions before fitting. MSHT and CT guess a polynomial. NNPDF's founding argument is that the guess itself biases the answer, and that the uncertainty you quote will not include the error you made by guessing. So they use a neural network instead — flexible enough, they argue, that it imposes no shape at all.

For uncertainties they do something different too. Instead of measuring curvature, they generate a few hundred fake copies of the entire dataset, each jiggled within its error bars, and fit every one separately. The spread of the resulting protons is the uncertainty. No inflation factor is applied — because, they argue, if you clean the data properly beforehand, none is needed.

That last clause is the fault line. MSHT and CT both inflate; NNPDF does not; and the resulting bands differ by more than a factor of two in places. Nobody has settled it — which is precisely why our own work exists.

The name, and the idea

Unlike the other two, the acronym is never officially expanded.

NNPDF is universally read as Neural Network Parton Distribution Functions. Notably, neither the collaboration's website nor its documentation actually spells this out anywhere — the reading traces back to the founding methodology paper of 2007, on the non-singlet case. The collaboration describes itself as determining "the structure of the proton using contemporary methods of artificial intelligence."

releasewhat it added
NNPDF1.0 (2008)"A determination of parton distributions with faithful uncertainty estimation" — the word faithful is the programme in one adjective
NNPDF2.x (2010–12)First unbiased global determination; heavy-quark masses; LHC data
NNPDF3.0 (2014)The first PDF set validated by a closure test — fitting data generated from a known answer, to check the method recovers it
NNPDF3.1 (2017)Charm fitted rather than computed, taking the independent distributions from seven to eight
NNPDF4.0 (2021)New machine-learning methodology derived by hyperparameter optimisation; 44 new datasets

Who they are

A genuine multi-national collaboration — 17 authors on the flagship, ten institutes today.

This is the structural opposite of Colibri: where that is a five-author paper from one group, NNPDF4.0 is a 139-page collaboration paper spanning three continents.

countryinstitutes
United KingdomUniversity of Edinburgh (Higgs Centre); University of Cambridge (DAMTP and Cavendish)
ItalyUniversità di Milano & INFN Milano (TIF Lab); INFN Torino
NetherlandsNikhef, Amsterdam; Vrije Universiteit Amsterdam
Finland · SpainUniversity of Jyväskylä; University of Seville
Singapore · UAENational University of Singapore; Technology Innovation Institute, Abu Dhabi

Leadership: Stefano Forte is Spokesperson, Juan Rojo Physics Coordinator, Stefano Carrazza R&D Coordinator. Richard D. Ball and Forte head the author list.

a connection worth noticing

Maria Ubiali is an author on NNPDF4.0, and Cambridge DAMTP is an NNPDF institute. The Colibri framework is therefore not an outside critique of the Monte-Carlo approach — it is built by someone inside that tradition, deliberately constructing the tools to test it. Colibri also draws its experimental data and fast theory predictions from NNPDF's public code.

What data they fit

The largest dataset of the three — 4,618 points at NNLO.

points fitted (NNLO)4,618 — up from 4,285 in NNPDF3.1
points fitted (NLO)4,426 — up from 4,295 in NNPDF3.1
new datasets vs 3.144, mostly from the LHC
dataset entries in the release runcard76
kinematic cutsQ² > 3.49 GeV², W² > 12.5 GeV²

Coverage spans fixed-target DIS (NMC, SLAC, BCDMS), neutrino DIS (CHORUS, NuTeV dimuon), the full HERA combined dataset including charm, fixed-target Drell–Yan (E605, E866 and E906/SeaQuest), Tevatron W and Z, and an unusually broad LHC selection: ATLAS, CMS and LHCb W/Z, Drell–Yan, Z transverse momentum, inclusive jets and dijets, direct photon production, single top, and top-pair production both total and differential.

what we could not verify

The per-process breakdown of data points — how many of the 4,618 come from DIS versus jets versus top — is tabulated in the paper but we could not retrieve those tables intact. We have deliberately not reproduced numbers we could not check.

One methodological difference worth flagging here: NNPDF treat nuclear effects (from deuteron and heavy-target data) as a theory covariance matrix rather than fitting a correction model as MSHT does, or cutting the affected region as CT does.

The neural network

One network, eight outputs, 763 parameters — and every setting chosen by machine.

architecture2 → 25 → 20 → 8, dense feed-forward
inputsx and log x
outputsthe 8 fitted combinations: Σ, g, V, V₃, V₈, T₃, T₈, T₁₅
activationstanh, tanh, linear
free parameters763 — against 52 in MSHT20 and 29 in CT18
starting scaleQ₀ = 1.65 GeV
optimiserNadam, up to 17,000 epochs

A notable change from NNPDF3.1: that release used one network per flavour; 4.0 uses a single network producing all eight outputs at once.

what "hyperparameter optimisation" means, and why it matters

The architecture above — how many layers, how many nodes, which activation, which optimiser, when to stop — was not chosen by hand. It is the output of an automated search over thousands of candidate configurations, scored on held-out data. The paper's own headline is that this is a "novel methodology derived through hyperparameter optimisation".

The intent is to remove the last place where a human could smuggle in a preference. The counter-argument is that it moves the choice rather than eliminating it — the scoring function is still a human choice.

Constraints. Momentum and valence sum rules are imposed analytically, as a network layer that computes the normalisation — not as a penalty. Positivity is imposed by Lagrange multiplier on 15 pseudo-observables, including strict positivity of the MS̄ distributions themselves. Integrability of the non-singlet T₃ and T₈ adds two more constraints. The systematic treatment of positivity and integrability is one of the things 4.0 introduced.

How they get the error bar

No curvature, no eigenvectors, no tolerance. An ensemble instead.

The procedure is conceptually simple and computationally brutal:

  • Take the real dataset. Generate a replica by shifting every measurement randomly within its quoted uncertainties, respecting correlations.
  • Fit a complete neural network to that replica — a full fit, from scratch.
  • Repeat 100 (or 1,000) times.
  • The spread of the resulting ensemble is the uncertainty. The mean is the central value.
using these sets correctly

Because this is not a Hessian set, the arithmetic is different. The uncertainty is the standard deviation over members 1–100not the Hessian sum over ± eigenvector pairs. Member 0 is the precomputed average, not a fit in its own right. The grid metadata says so: ErrorType: replicas. Applying the Hessian formula to a replica set is a real and common error.

Overfitting is the obvious hazard with 763 parameters, and it is controlled by cross-validation: 75% of the points train the network, 25% are held back to validate it, and training stops when the held-back quality stops improving (a patience of 1,700 epochs). This is what replaces the polynomial groups' restriction on parameter count.

Validation is by closure test — fit data generated from a known input and check the method recovers it within its quoted uncertainty — an approach NNPDF introduced in 3.0 and which this project uses too. They add future tests: checking backward and forward compatibility with data outside the fitted range.

a version trap we had to avoid

NNPDF's public documentation describes the current development code, which differs from the 2021 release in ways that are easy to misattribute. Two examples: generating the replicas from the t0 covariance matrix, and splitting training/validation in a basis that diagonalises the correlation matrix, are both later changes — neither is what NNPDF4.0 did. The description above is the 4.0 release behaviour, taken from its own published runcard.

Their proton, live

The released NNPDF4.0 NNLO grid, 100 replicas.

NNPDF4.0 — every parton species

Bands are the standard deviation over the 100-replica ensemble. Note how narrow they are at small x compared with the other two groups.

Against the other groups

This is where the argument becomes visible.

NNPDF4.0 versus MSHT20 and CT18

NNPDF4.0 solid with the heavier band; the others dashed. All three at a common 68% confidence level.

For the gluon and the quark singlet in the data region, the NNPDF4.0 band is roughly half the width of MSHT20's and about a third of CT18's. That is not a rendering artefact and not a confidence-level mismatch — it is the direct consequence of applying no tolerance where the other two apply one.

The independent PDF4LHC21 benchmarking exercise found the same thing under controlled conditions: refitting all three methods on common data and theory, the gluon–gluon luminosity uncertainty at the Higgs mass came out at 2.3% (MSHT), 2.1% (CT) and 1.2% (NNPDF). Its conclusion was that methodological uncertainties "can be as large or even larger than the PDF uncertainties associated with the fitted data".

a detail people get wrong

PDF4LHC21 combined CT18′, MSHT20 and NNPDF3.1′ — not NNPDF4.0, because 4.0 was released after the benchmarking finished. The same report warns that NNPDF4.0 "may fall outside the 68% CL uncertainty bands of the PDF4LHC21 sets".

What they release

Including — usefully — an official Hessian conversion.

gridmemberswhat it is
NNPDF40_nnlo_as_01180101The default. Plotted on this site.
NNPDF40_nlo_as_01180101NLO.
NNPDF40_nnlo_as_01180_10001001The 1,000-replica ensemble, for work needing finer statistics.
NNPDF40_nnlo_as_01180_hessian51An official Hessian conversion — 50 symmetric eigenvectors, derived from the 1,000-replica set.
..._mhou101With missing-higher-order theory uncertainties.
..._qed101Adds the photon distribution (and top) to the flavour list.
NNPDF40_an3lo_as_01180101Approximate N³LO.
αS and charm variantsαS = 0.117, 0.118, 0.119 at NLO and NNLO, plus perturbative-charm (_pch) variants.

The Hessian conversion is genuinely useful: many experimental analyses require Hessian-format inputs, and it lets NNPDF be used in those pipelines. Note that it is built from the 1,000-replica set, not the 100. The flagship grid spans x from 10⁻⁹ to 1 and Q from 1.65 GeV to 100 TeV.

The central argument

The most important section on this page — and the reason this project exists.

NNPDF state their position plainly: their approach is "based on a pre-processing of the data that yields a consistent input dataset, rather than including an explicit tolerance factor in the fit", and that this is what distinguishes them from CT and MSHT. The claim is that inconsistent data should be fixed before fitting, not compensated for afterwards by widening the answer.

Round 1 — the sampling critique

CTEQ-affiliated authors used NNPDF's own open-source code to search for alternative solutions, and reported finding fits with χ² equal to or better than the NNPDF4.0 central replica that nonetheless predict LHC cross sections outside the quoted bands. Their conclusion: the uncertainties on large-x charm and small-x gluon are understated.

NNPDF replied that the ensemble "is not a random sampling, but rather an importance sampling of the PDF space", that the critique misidentifies the quantity being minimised, and that the alternative solutions "correspond to overfitted solutions" — reproducible but improbable. Two status points a specialist will check: that reply carries no journal reference, and no counter-reply was published.

Round 2 — the MSHT closure test

MSHT refitted using NNPDF's own data and theory settings, isolating methodology, and found NNPDF4.0's uncertainties "broadly in line with the MSHT results if a textbook T² = 1 tolerance is applied, but significantly smaller if a tolerance typical of the MSHT20 fit is applied" — pointing to "an inherent inconsistency between these approaches". To their credit MSHT list three readings neutrally, including "the MSHT uncertainties may be too large". No published NNPDF reply to this paper was found.

Round 3 — can MS̄ distributions be negative?

NNPDF4.0 imposes strict positivity, justified by a 2020 paper from within the collaboration. That was rebutted by Collins, Rogers and Sato, who showed the underlying assumption fails and warned that imposing positivity "especially at a low initial scale… is likely to introduce excessive theoretical bias". NNPDF conceded ground in a 2024 reply, bounding the claim to Q² ≳ 5 GeV².

an unresolved tension worth knowing

NNPDF's own starting scale is Q₀ = 1.65 GeV, i.e. Q₀² ≈ 2.7 GeV² — which sits below the 5 GeV² bound their own reply establishes. NNPDF have since reported that a closure-test failure "is explained by the requirement of PDF and observable positivity in NNPDF4.0, which necessarily breaks the assumed Gaussianity".

The argument does not run one way

It would be easy to read this page as "NNPDF's errors are too small". The literature does not support that conclusion cleanly:

  • CTEQ-affiliated authors argue NNPDF uncertainties may be understated;
  • MSHT argue similarly, while listing "MSHT uncertainties may be too large" as equally live;
  • Hunt-Smith and collaborators argue neural-network cross-validation can inflate uncertainties;
  • and the Cambridge group has shown the Monte-Carlo replica method can err in either direction, with the sign determined by the residual-weighted Hessian.

PDF4LHC21 — a joint document of all three groups — deliberately declines to adjudicate: "The differences, in general, represent valid physics choices." There are also real convergence signals: CT25 is formalising parametrisation bias as an explicit uncertainty, and NNPDF and MSHT authors have jointly released code running neural-network and fixed-polynomial parametrisations on identical inputs.

why this is our project

Every round above is an argument between groups that differ in data, theory, code and method at once. That is exactly the confound Colibri was built to remove, and what our own Tier-2 comparison does: one fit, one dataset, one parametrisation, three uncertainty methods. Our answer — that they diverge by factors of 8 to 30 million on a realistic 52-parameter fit — does not settle who is right, but it does show the disagreement is not a rounding detail.

What comes next

NNPDF4.1 is announced but not released.

NNPDF describe 4.1 as forthcoming: full LHC Run II luminosity, QED and missing-higher-order effects and approximate N³LO folded into the baseline, exact NNLO grids removing the current K-factor approximation, new hyperoptimisation criteria, and GPU support. No NNPDF4.1 paper or grid exists yet — we checked the LHAPDF server directly and the set returns a 404. Treat any reference to it as in-progress.

Since 4.0 the collaboration has released missing-higher-order uncertainties, a photon distribution, approximate N³LO, nuclear PDFs (nNNPDF3.0) and polarised PDFs (NNPDFpol2.0). Their highest-profile result in this period is not a PDF set at all: 3σ evidence for intrinsic charm in the proton, published in Nature in 2022.

Papers

The flagship, the code, and the argument.

  • The path to proton structure at 1% accuracy NNPDF Collaboration (17 authors; Ball, Carrazza, Cruz-Martinez, Del Debbio, Forte, Giani, Iranipour, Kassabov, Latorre, Nocera, Pearson, Rojo, Stegeman, Schwan, Ubiali, Voisey, Wilson) — Eur. Phys. J. C 82 (2022) 428. 139 pages. Note the arXiv version is titled "The Path to Proton Structure at One-Percent Accuracy". arXiv:2109.02653
  • An open-source machine learning framework for global analyses of parton distributions The companion code paper, released the same day. It is what made the independent audits in §10 possible at all. arXiv:2109.02671
  • Evidence for intrinsic charm quarks in the proton NNPDF Collaboration — Nature 608 (2022) 483. The collaboration's highest-profile result. arXiv:2208.08372
  • Parton distributions need representative sampling Courtoy, Huston, Nadolsky, Xie, Yan, Yuan — Phys. Rev. D 107 (2023) 034008. The sampling critique. arXiv:2205.10444
  • Response to "Parton distributions need representative sampling" NNPDF Collaboration. Preprint; no journal reference. arXiv:2211.12961
  • Can MS̄ parton distributions be negative? — and the exchange that followed Candido, Forte, Hekhorn, JHEP 11 (2020) 129; rebutted by Collins, Rogers, Sato, Phys. Rev. D 105 (2022) 076010; reply by Candido, Forte, Giani, Hekhorn, Eur. Phys. J. C 84 (2024) 335. arXiv:2006.07377 · arXiv:2111.01170 · arXiv:2308.00025
  • The PDF4LHC21 combination of global PDF fits for the LHC Run III All three groups jointly — J. Phys. G 49 (2022) 080501. Uses NNPDF3.1′, not 4.0. arXiv:2203.05506
  • Collaboration home Members, institutes, code and documentation. nnpdf.mi.infn.it