partonmap
Study · 3.7

CT (CTEQ-TEA) — Hessian errors and tolerance

The CTEQ tradition, Bernstein forms, the two-tier Δχ²≈100 tolerance, and the Lagrange-multiplier cross-check.

The second of our two Hessian-method groups is the North-American one: the CTEQ collaboration and its parton-distribution arm, CTEQ-TEA. Where MSHT (§3.6) grew out of the British MRST/CTEQ dialogue, CTEQ is the group that invented much of the modern PDF-uncertainty machinery — the Hessian eigenvector sets, the Lagrange-multiplier cross-check, and the ACOT heavy-quark scheme all carry a CTEQ pedigree. This section profiles the group, walks its set lineage up to the current CT18 family, and then dwells on the two things that make CT distinctive as an uncertainty engine: its two-tier tolerance criterion and its Lagrange-multiplier scan. Both matter directly to us, because CT is one leg of this lab's three-method comparison (§5.1, PPDF-15/17).

Who they are

CTEQ stands for the Coordinated Theoretical-Experimental project on QCD — a broad US-based collaboration founded in the early 1990s by Wu-Ki Tung (together with Jorge Morfín, who with Tung is credited with coining the very phrase "global analysis"). CTEQ is bigger than PDFs: it runs summer schools and a handbook. The PDF-fitting sub-group carries the label CTEQ-TEA — "Tung Et Al." — a memorial to Tung, who died in 2009. The set names shortened from "CTEQ" to just "CT" starting with CT10, which is why the current standard is CT18, not "CTEQ18".

The collaboration is spread across many institutions. Anchor nodes are Michigan State University (C.-P. Yuan, Joey Huston, Dan Stump, Carl Schmidt, historically Huey-Wen Lin) and Southern Methodist University (Pavel Nadolsky, Fred Olness). Others contributing to recent CT18 work include Tie-Jiun Hou, Keping Xie, Sayipjamal Dulat, Marco Guzzi (Kennesaw State), Aurore Courtoy (UNAM, Mexico), and Tim Hobbs (Argonne). Nadolsky, Huston and Yuan are the names you will most often see leading a CT release.

Why does the same handful of names appear on both CT and MSHT history?

Because the field is small and inbred in the best sense. The Hessian-with-tolerance philosophy was worked out jointly across the CTEQ–MRST community around 2000–2002; the two groups then made different engineering choices from a shared starting point. Reading §3.6 and §3.7 back-to-back, you are watching one idea — "a naïve $\Delta\chi^2=1$ is far too optimistic for a global fit with mutually-tense datasets" — fork into two dialects.

The set lineage

Each generation is a snapshot: same fitting technology, more data, occasionally a methodological upgrade. The spine our field cites:

Set~YearWhat it added
CTEQ62002First modern CTEQ Hessian set with published error PDFs.
CTEQ6.62008General-mass VFNS (ACOT-based) heavy-quark treatment throughout.
CT102010 (NLO) / 2013 (NNLO)First combined HERA-I data; name shortened to "CT".
CT142015Early LHC jet/vector-boson data; refreshed parametrization.
CT182019–2021HERA-I+II combined; precision 7–8 TeV LHC data. Current standard.

CT18 is not a single set but a small family, because some new LHC data pull against the older DIS constraints and the group refused to bury that tension inside one averaged answer:

  • CT18 — the default, recommended general-purpose set.
  • CT18A — adds the high-precision ATLAS 7 TeV $W/Z$ data, which prefer a larger strangeness than the DIS experiments alone.
  • CT18X — uses an $x$-dependent factorization scale in low-$x$ DIS to mimic enhanced small-$x$ logarithms; raises the small-$x$ gluon.
  • CT18Z — all of the above plus a slightly higher charm mass (pole mass $1.4$ GeV vs $1.3$ GeV in CT18); built to maximize the departure from CT18 while keeping a comparable fit quality. It is the "how different could the answer legitimately be?" bracket.

Publishing A/X/Z alongside the default is itself a scientific statement: the spread between CT18 and CT18Z is an honest map of methodology-dependence that a single number would hide.

Parametrization — Bernstein polynomials

CT is squarely in the fixed-functional-form camp (contrast the neural-network approach of §3.5; the general form-vs-flexibility trade-off is laid out in §3.1, which we will not re-derive here). At a low starting scale $Q_0 \sim 1.3$ GeV, each flavour combination is written

$$ x f(x, Q_0) \;=\; x^{a_1}\,(1-x)^{a_2}\; P_n(x), $$

where the $x^{a_1}$ factor fixes the small-$x$ behaviour, $(1-x)^{a_2}$ fixes the large-$x$ fall-off, and the smooth shape in between is carried by $P_n(x)$. Since CT14, CTEQ builds $P_n$ from Bernstein polynomials (the basis behind Bézier curves):

$$ P_n(y) \;=\; \sum_{k=0}^{n} c_k\, B_{k,n}(y), \qquad B_{k,n}(y) = \binom{n}{k}\, y^{k}(1-y)^{n-k}, \qquad y = \sqrt{x}. $$

Bernstein bases are the polished choice for a reason worth stating: each basis function is non-negative and the whole family is a partition of unity ($\sum_k B_{k,n} = 1$), so the coefficients $c_k$ act like "control points" that nudge the curve locally without the wild oscillations a high-order monomial fit suffers. Each parton flavour carries roughly five to eight free parameters, giving CT18 on the order of $\sim\!30$ fit parameters overall — a deliberately parsimonious form. That parsimony is exactly what makes the next ingredient, the Hessian, tractable.

The Hessian and the two-tier tolerance

Like MSHT, CT quantifies uncertainty by expanding $\chi^2$ to second order about the best fit and diagonalizing the Hessian into orthogonal eigenvector directions. The construction is worth doing once explicitly, because the error sets we actually load into this lab are its direct output.

Derivation — from the Hessian to eigenvector error sets

Let $\vec a$ be the $N$ fit parameters and $\hat a$ the best-fit point. Near the minimum,

$$ \Delta\chi^2 \equiv \chi^2(\vec a) - \chi^2(\hat a) \;\approx\; \tfrac{1}{2}\sum_{i,j} H_{ij}\,(a_i - \hat a_i)(a_j - \hat a_j), \qquad H_{ij} = \left.\frac{\partial^2\chi^2}{\partial a_i\,\partial a_j}\right|_{\hat a}. $$

$H$ is symmetric and (at a genuine minimum) positive-definite, so it has an orthonormal eigenbasis $H\,\vec v_k = \lambda_k\,\vec v_k$. Rescale into that basis, $z_k = \sqrt{\lambda_k/2}\;\vec v_k\!\cdot\!(\vec a - \hat a)$, and the paraboloid becomes a plain sphere:

$$ \Delta\chi^2 \;=\; \sum_k z_k^2. $$

In this "$z$-space" the confidence region $\Delta\chi^2 \le T^2$ is the ball of radius $T$. CTEQ then produces a pair of error PDFs for each eigenvector direction by stepping to $z_k = \pm T$ and transforming back to parameter space — giving $2N$ error sets $S_k^{\pm}$ flanking the central set $S_0$. Any observable's uncertainty is the symmetric (or asymmetric) sum over these displacements,

$$ \delta X \;=\; \tfrac{1}{2}\sqrt{\sum_k \big(X(S_k^{+}) - X(S_k^{-})\big)^2}. $$

Everything now hinges on one number: the tolerance $T$. Set it wrong and every error bar downstream is wrong by the same factor. $\blacksquare$

The naïve textbook choice $\Delta\chi^2 = 1$ (i.e. $T=1$) is right only for a single dataset with Gaussian, perfectly-compatible errors. A global fit is neither: dozens of experiments, each with its own systematics, sit in mild mutual tension. CTEQ's answer is a two-tier criterion:

  • Tier 1 — a global tolerance. The error sets are placed at a global $\Delta\chi^2 \approx 100$. With $\gtrsim 2000$ data points, $\Delta\chi^2 \sim 100$ corresponds to roughly a $10\sigma$ excursion in raw $\chi^2$ terms — far beyond $\Delta\chi^2=1$ — reflecting that moving the PDFs by an amount that keeps all experiments jointly acceptable costs far more than one unit of $\chi^2$.
  • Tier 2 — a per-experiment penalty. On top of the global bound, CTEQ imposes that no individual experiment be pushed outside its own $\sim$90% confidence region. An additional penalty switches on if any single dataset's $\chi^2$ degrades past its 90%-CL threshold, so the eigenvector displacement stops as soon as it would make one experiment "unhappy," even if the global $\Delta\chi^2=100$ ball would have allowed further travel.

The two tiers pull in complementary directions: Tier 1 asks "is the fit as a whole still acceptable?"; Tier 2 asks "is any one experiment being sacrificed to buy that?" Later CT analyses refined this into a more dynamic criterion, allowing the effective tolerance to differ direction-by-direction rather than sitting at a flat $T\approx10$.

The two-tier tolerance in eigenvector (z) space z₂z₁ Tier 1 — global Δχ² ≈ 100 (radius T ≈ 10) expt A · 90% CL expt B · 90% CL best fit travel along z₁ … stops here (Tier 2 binds) Tier 1 edge
Two tiers, whichever binds first. In the sphere-ized eigenvector space, Tier 1 is the global Δχ²≈100 ball (teal, radius $T\approx10$). Tier 2 adds a 90%-CL wall for each individual experiment (red/orange). Walking out along an eigenvector, the error set stops at the first boundary reached — here experiment A's Tier-2 wall bites before the Tier-1 circle, so the tolerance is tighter than $T\approx10$ in that direction; in another direction the circle itself binds. This is how CT keeps any one experiment from being sacrificed to the global fit.
CT two-tier vs MSHT dynamic vs the naïve $\Delta\chi^2 = 1$ — how do they differ?

$\Delta\chi^2 = 1$: statistically pure, but assumes all data are mutually consistent and Gaussian. For a global fit it badly under-estimates the uncertainty — it trusts every experiment's quoted errors literally.

CT two-tier: a fixed global $\Delta\chi^2 \approx 100$ ($T\approx10$) plus a 90%-CL guard rail on each experiment. Conservative and transparent — one headline number, one veto rule.

MSHT dynamic (§3.6): no single global $T$; instead the tolerance is set separately for each eigenvector direction by the point at which the most-constraining experiments along that direction start to complain — so $T$ varies (typically $\sim$1–6 per direction) across the $z$-space.

They converge in spirit: all three encode "believe the data, but don't believe any one dataset too much." CT hard-codes the compromise into a global constant with an override; MSHT lets the data recompute the compromise on every axis. That the two groups land on comparable error bands by different roads is one of the quiet reassurances of the whole enterprise.

The Lagrange-multiplier cross-check

The Hessian recipe rests on one assumption: that $\chi^2$ is well-approximated by a paraboloid out to $\Delta\chi^2 \approx 100$. For a stiff, well-constrained observable that is fine. For something the data barely pin down — where $\chi^2$ is flat, skewed, or bounded — the quadratic approximation can lie. CTEQ's own answer, and arguably its most elegant contribution, is the Lagrange-multiplier (LM) method (Stump, Pumplin et al., 2001–2002): instead of trusting the local curvature, directly map how $\chi^2$ responds when you force a chosen observable to take a fixed value.

Derivation — the Lagrange-multiplier scan

Pick any observable $X(\vec a)$ — a Higgs cross-section, a high-$x$ gluon, a moment of a PDF. Define a modified objective

$$ \Phi(\vec a; \lambda) \;=\; \chi^2(\vec a) \;+\; \lambda\, X(\vec a), $$

and minimize $\Phi$ over the full parameter space at fixed $\lambda$. The stationarity condition

$$ \nabla_{\vec a}\,\chi^2 \;=\; -\lambda\,\nabla_{\vec a}\,X $$

is exactly the statement "find the smallest $\chi^2$ consistent with $X$ held at whatever value $\lambda$ drives it to." Sweeping $\lambda$ from negative to positive traces out a curve of best-fit points, each with a value of $X$ and a value of $\chi^2$. Plotting $\chi^2$ against $X$ gives the exact profile — no quadratic assumption anywhere. The uncertainty on $X$ is then read straight off this curve: it is the range of $X$ over which $\Delta\chi^2$ stays within the tolerance (the same $\Delta\chi^2 \approx 100$ / per-experiment criterion as before). $\blacksquare$

A Lagrange-multiplier scan: χ² profiled against an observable X = gluon momentum fraction → Δχ² tolerance Δχ² ≈ 100 Hessian (symmetric) true χ²(X) best fit X = 0.382 0.363 0.412 +0.030 / −0.019 → asymmetric
The LM profile is exact — and asymmetric. Re-minimizing $\chi^2 + \lambda X$ along a ladder of $\lambda$ traces the true $\chi^2(X)$ curve (teal); the uncertainty is where it crosses the $\Delta\chi^2\approx100$ tolerance (red), giving $X = 0.382^{+0.030}_{-0.019}$. A pure Hessian error assumes the symmetric parabola (grey dashed) and would report a misleadingly symmetric $\pm0.025$, missing the steeper low-side wall. This is why CTEQ validates Hessian bands against LM scans for the observables that matter.
Reading a Lagrange-multiplier scan

Setup: suppose we want the uncertainty on the gluon momentum fraction $X = \int_0^1 x\,g(x)\,dx$ at $Q = 100$ GeV. Run the LM scan: for a ladder of $\lambda$ values, re-minimize $\chi^2 + \lambda X$ and record the pair $(X, \chi^2)$.

Result (schematic): we get a curve like

$X$ (gluon frac.)$\Delta\chi^2$
0.35+140
0.37+40
0.3820 (best fit)
0.40+55
0.42+165

Read-off: the tolerance $\Delta\chi^2 \approx 100$ is crossed near $X \approx 0.363$ below and $X \approx 0.412$ above, giving $X = 0.382^{+0.030}_{-0.019}$. Two lessons jump out. First, the interval is asymmetric — the profile rises faster on the low side. A pure Hessian error, being symmetric by construction, would have reported $\pm 0.025$ and quietly mislabelled the shape. Second, if the curve had a long flat shoulder, the LM method would expose it whereas the Hessian would not. This is why CTEQ uses LM as the gold standard against which the (cheaper, reusable) Hessian error sets are validated.

Heavy flavour — the ACOT scheme

One more CTEQ invention deserves a mention because it underpins how every modern fit treats charm and bottom. The ACOT scheme (Aivazis–Collins–Olness–Tung, 1994) is a general-mass variable-flavour-number scheme (GM-VFNS). The problem it solves: near the heavy-quark threshold $Q \sim m_c$ you should keep the full charm mass in the kinematics (a fixed-flavour picture); far above it, $Q \gg m_c$, the mass logs $\ln(Q^2/m_c^2)$ must be resummed by treating charm as a massless active parton (a variable-flavour picture). ACOT stitches these two limits into one consistent calculation by combining massive Wilson coefficients with DGLAP-evolved PDFs and a subtraction term that removes the double-counting where the two pictures overlap. The result interpolates smoothly from threshold to the asymptotic regime. CTEQ ships a computationally lighter variant, S-ACOT (simplified ACOT), used from CTEQ6.6 onward. The upshot for us: a "5-flavour" PDF like CT18 is not naïvely massless — its charm and bottom were switched on through this careful threshold treatment.

Connecting to this lab

CT is the second of our two Hessian-method groups, standing beside MSHT (§3.6). Together they define the Hessian leg of this lab's three-method comparison — Hessian error sets vs Monte-Carlo replicas (§3.5) vs our own Bayesian nested-sampling posteriors — pursued in Topic 5.1 and logged as PPDF-15 and PPDF-17. When we quote a "CT18 uncertainty band" in a Results plot, that band is the eigenvector-set sum of the Derivation above, placed at the two-tier tolerance, and (for the observables that matter) cross-checked against a Lagrange-multiplier scan. Knowing how that band was constructed is what lets us compare it fairly against a posterior credible interval that was built on entirely different statistical foundations — the whole point of the three-method exercise.

Hold on to four things: (1) CTEQ-TEA is the US Hessian group — Nadolsky, Huston, Yuan and collaborators across MSU, SMU and beyond — with lineage CTEQ6 → CTEQ6.6 → CT10 → CT14 → CT18 (plus the A/X/Z variants), using a parsimonious Bernstein-polynomial × $x^{a}(1-x)^{b}$ form. (2) CT's uncertainty is a Hessian eigenvector construction placed at a two-tier tolerance: a global $\Delta\chi^2 \approx 100$ and a 90%-CL guard on each experiment — far more conservative than the naïve $\Delta\chi^2=1$, cousin to MSHT's dynamic tolerance. (3) The Lagrange-multiplier scan maps $\chi^2$ vs any observable exactly, catching the asymmetries and flat directions the quadratic Hessian misses. (4) CT + MSHT are the Hessian leg of our three-method comparison (§5.1, PPDF-15/17); the ACOT scheme is the CTEQ-born heavy-flavour treatment beneath it all.