Path-Integral Quantization

Contents
  1. The path integral in quantum mechanics
  2. Euclidean formulation
  3. Functional quantization of fields
  4. Gauge fields
  5. The lattice
  6. Topology of the gauge vacuum
  7. Anomalies

The canonical quantization of Canonical Quantization of Fields promotes coordinates to operators and privileges a time coordinate; the path integral does neither. It computes an amplitude as a sum over every history weighted by \(\exp(\ii S/\hbar)\), with \(S\) the classical action of Calculus of Variations, so the Lorentz invariance of the Lagrangian is manifest at every step. The idea is Dirac's [Dirac:1933] and its execution Feynman's [Feynman:1948] [Feynman:1949b]. This chapter develops it from non-relativistic quantum mechanics through the functional quantization of fields: generating functionals, the Euclidean continuation, the LSZ reduction [Lehmann:1955], Feynman rules and the one-particle-irreducible effective action. It therefore follows Generalized Classical Field Theory and supplies the machinery used throughout Quantum Chromodynamics and Electroweak Unification and the Higgs Boson and The Renormalization Group.

Three things make the path integral more than a restatement. Gauge theories require it: the Faddeev–Popov construction [Faddeev:1967] and the BRST symmetry that organizes it [Becchi:1976] [Tyutin:1975] have no simple canonical counterpart, and Gribov showed that the gauge-fixing is not globally well posed [Gribov:1978]. Non-perturbative physics requires it: Wilson's lattice formulation [Wilson:1974] turns the Euclidean integral into a finite-dimensional one that a computer can sample, which is how the hadron spectrum of Quantum Chromodynamics is computed from first principles [Duerr:2008]. And it explains anomalies: a classical symmetry can fail to survive quantization because the functional measure is not invariant [Fujikawa:1979] — a statement with a measured consequence, the decay rate of the neutral pion [Adler:1969] [Bell:1969] [Larin:2020]. The canonical monographs are [Feynman:1965] [Coleman:1985] [Weinberg:1996].

Conventions are those of Section 100.1.1 without exception. The metric is \(\eta_{\mu\nu}=\diag(+1,-1,-1,-1)\); coordinates are \(x^{\mu}=(ct,\vect{x})\), so \(\dd^{4}x\) has SI dimension \(\mathrm{m}^{4}\); four-momenta are \(p^{\mu}=(E/c,\vect{p})\) of dimension \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), and the mass shell is \(p^{2}=m^{2}c^{2}\) as in Equation (100.1), so every propagator denominator reads \(p^{2}-m^{2}c^{2}\) and never \(p^{2}-m^{2}\). The action is \(S=c^{-1}\int\dd^{4}x\,\Lag\) as in Equation (100.2), with \(\Lag\) an energy density; the Fourier convention is Equation (100.14). Every \(\hbar\) and every \(c\) is written out, and here that is not a matter of taste: the expansion parameter of this whole chapter is \(\hbar\). The loop expansion of Section 106.3.4 is an expansion in \(\hbar\) and in nothing else, and the semiclassical limit of Section 106.1.3 is the statement that \(\hbar\) is small compared with the classical action. A treatment that puts the reduced Planck constant equal to unity erases the very quantity whose smallness is being exploited. Readers coming from the standard references should use the translation rule of Remark 100.4, which is the only place in this part where that convention is discussed.

Five symbols carry more than one standard meaning in the material of this chapter, and the collisions are settled here rather than in passing. The electric charge of a particle or a field is written \(Q\) throughout, never \(q\): the letter \(q\) is the mechanical coordinate of Section 106.1 and the topological charge density of Definition 106.79, and both meet a charge on the same page. The BRST charge of Definition 106.59 is therefore written \(Q_{B}\). The ghost fields are \(c^{a}\) and \(\bar c^{a}\), always carrying an adjoint index or an explicit generator, so a bare \(c\) is the speed of light and nothing else. The lattice coupling \(\beta\) of Equation (106.81) is not the inverse temperature \(\beta\) of Equation (106.27); each name is fixed by its own literature, and they never enter the same equation here. And \(\Gamma\) is the effective action of Definition 106.41 everywhere except in Section 106.7.2, where it is a decay width, and at one step of Proposition 106.45, where it is the Euler function and is flagged as such on the spot.

The path integral in quantum mechanics

Dirac's Lagrangian formulation

Take one degree of freedom \(q\) with Hamiltonian \(H=p^{2}/2m+V(q)\), both operators, and write the finite-time propagator

\begin{equation}\tag{106.1} K(q'',t'';q',t') :=\bra{q''}U(t'',t')\ket{q'}\ec\qquad U(t'',t')=\exp\!\left[-\frac{\ii}{\hbar}H(t''-t')\right]\ec \end{equation}

for a time-independent \(H\). Since \(\braket{q''}{q'}=\delta(q''-q')\), the states \(\ket{q}\) carry dimension \(\mathrm{m}^{-1/2}\) and \(K\) carries \(/\mathrm{m}\). Everything about the dynamics is in \(K\): given \(\psi(q',t')\), the wavefunction at any later time is \(\psi(q'',t'')=\int\dd q'\,K(q'',t'';q',t')\psi(q',t')\).

The classical object that propagates a mechanical system from \((q',t')\) to \((q'',t'')\) is Hamilton's principal function — the value of the action along the classical path,

\begin{equation}\tag{106.2} S_{\mathrm{cl}}(q'',t'';q',t') =\int_{t'}^{t''}\dd t\;L\bigl(q_{\mathrm{cl}}, \dot q_{\mathrm{cl}},t\bigr)\ec \end{equation}

of dimension \(\mathrm{J}\,\mathrm{s}\). It is the type-one generating function of the canonical transformation that maps the state at \(t'\) onto the state at \(t''\) (Hamilton–Jacobi Theory and the Optical–Mechanical Analogy), obeying

\begin{equation}\tag{106.3} \pdv{S_{\mathrm{cl}}}{q''}=p''\ec\qquad \pdv{S_{\mathrm{cl}}}{q'}=-p'\ec\qquad \pdv{S_{\mathrm{cl}}}{t''} +H\!\left(q'',\pdv{S_{\mathrm{cl}}}{q''}\right)=0\ep \end{equation}

Dirac's observation [Dirac:1933] was that the quantum object Equation (106.1) and the classical object Equation (106.2) are related by

\begin{equation}\tag{106.4} \bra{q''}U(t'',t')\ket{q'}\quad\text{``corresponds to''}\quad \exp\!\left[\frac{\ii}{\hbar}S_{\mathrm{cl}}(q'',t'';q',t')\right]\ep \end{equation}

The scare quotes are Dirac's own. He did not claim an equality, and he was right not to: the two sides do not even have the same dimension, the left-hand side carrying \(/\mathrm{m}\) and the right-hand side being a pure phase. What he had was the observation that a canonical transformation is implemented in quantum mechanics by a unitary operator, that the generating function of the classical transformation therefore ought to control the phase of the quantum kernel, and that in the classical limit Equation (106.3) is exactly the eikonal relation satisfied by the phase of a rapidly oscillating wave.

What Feynman added is that Equation (106.4) becomes an equality in the infinitesimal case, up to a normalization that carries the missing dimension, and that finite times then follow by composition. The infinitesimal statement is a short computation.

Lemma 106.1 (The short-time kernel).

Let \(H=p^{2}/2m+V(q)\) with \(V\) bounded and continuous. Then, for \(\Delta t\to0\),

\begin{equation}\tag{106.5} \bra{q''}\ee^{-\ii H\Delta t/\hbar}\ket{q'} =\sqrt{\frac{m}{2\pi\ii\hbar\Delta t}}\; \exp\!\left\{\frac{\ii\Delta t}{\hbar}\left[ \frac{m}{2}\left(\frac{q''-q'}{\Delta t}\right)^{2} -V\!\left(\frac{q''+q'}{2}\right)\right]\right\} \left[1+O((\Delta t)^{2})\right]\ep \end{equation}

The prefactor carries \(/\mathrm{m}\), the exponent is dimensionless, and the bracket in the exponent is the Lagrangian \(L=\frac{1}{2}m\dot q^{2}-V\) evaluated on the straight segment joining the endpoints. Rests on Lemma 106.2 and Corollary 8.15.

Derivation. Derives Lemma 106.1. The two terms of \(H\) do not commute, but over a time \(\Delta t\) the error of separating them is second order: by the Baker–Campbell–Hausdorff expansion,

\[ \ee^{-\ii(T+V)\Delta t/\hbar} =\ee^{-\ii T\Delta t/\hbar}\,\ee^{-\ii V\Delta t/\hbar} \exp\!\left[-\frac{(\Delta t)^{2}}{2\hbar^{2}} \comm{T}{V}+O((\Delta t)^{3})\right]\ec \]

with \(T=p^{2}/2m\). Dropping the correction and inserting the momentum-space resolution of the identity \(\identity=\int\dd p\,\ketbra{p}{p}\) with \(\braket{q}{p}=(2\pi\hbar)^{-1/2}\ee^{\ii pq/\hbar}\),

\begin{align} \bra{q''}\ee^{-\ii T\Delta t/\hbar}\ee^{-\ii V\Delta t/\hbar} \ket{q'} &=\ee^{-\ii\Delta t V(q')/\hbar} \int\frac{\dd p}{2\pi\hbar}\, \exp\!\left[\frac{\ii}{\hbar}p(q''-q') -\frac{\ii\Delta t p^{2}}{2m\hbar}\right]\nn\\ &=\ee^{-\ii\Delta t V(q')/\hbar} \sqrt{\frac{m}{2\pi\ii\hbar\Delta t}}\, \exp\!\left[\frac{\ii m(q''-q')^{2}}{2\hbar\Delta t}\right]\ec \tag{106.6} \end{align}

the \(p\) integral being the Gaussian \(\int\dd p\,\ee^{-ap^{2}/2+bp}=\sqrt{2\pi/a}\,\ee^{b^{2}/2a}\) of Lemma 106.2 below, analytically continued to \(a=\ii\Delta t/(m\hbar)\) on the boundary of its domain of convergence — legitimate because the integrand is the boundary value of a function holomorphic in \(\Re a>0\) and the contour may be rotated by Corollary 8.15. The square root is the principal branch, \(\sqrt{\ii}=\ee^{\ii\pi/4}\).

Finally \(V(q')\) may be replaced by \(V\bigl((q''+q')/2\bigr)\): the two differ by \(\frac{1}{2}(q''-q')V'+O((q''-q')^{2})\), and the prefactor confines \(\abs{q''-q'}\) to order \(\sqrt{\hbar\Delta t/m}\), so the change in the exponent is of order \((\Delta t)^{3/2}\); it cancels against its mirror image on symmetrizing, leaving an error \(O((\Delta t)^{2})\). The midpoint form is not an aesthetic choice, and Section 106.1.3 shows where it becomes compulsory.

The prefactor's dimension is \([\,m/(\hbar\Delta t)\,]^{1/2} =[\mathrm{kg}/\mathrm{J}/\mathrm{s}^{2}]^{1/2} =/\mathrm{m}\), as \(K\) requires.

So Dirac's correspondence is an equality once the normalization \((m/2\pi\ii\hbar\Delta t)^{1/2}\) is supplied, and once the principal function is read as the action of the straight segment, which for infinitesimal times is the classical path to the accuracy retained. The whole of the next subsection is the observation that finite times are built out of infinitesimal ones, and that when they are, the straight segments become arbitrary paths.

Feynman's sum over histories

The Gaussian integral, once and for all

Every explicit evaluation in this chapter reduces to one integral, so it is worth recording in the generality that will be needed.

Lemma 106.2 (Gaussian integrals).

Let \(A\) be a real symmetric positive-definite \(n\times n\) matrix and \(b\in\R^{n}\). Then

\begin{equation}\tag{106.7} \int_{\R^{n}}\dd^{n}x\; \exp\!\left(-\tfrac{1}{2}x\transpose Ax+b\transpose x\right) =\frac{(2\pi)^{n/2}}{\sqrt{\det A}}\; \exp\!\left(\tfrac{1}{2}b\transpose A^{-1}b\right)\ep \end{equation}

The identity extends by analytic continuation to every complex symmetric \(A\) with positive-definite real part, and to the boundary case \(A=-\ii M\) with \(M\) real symmetric and non-singular, where

\begin{equation}\tag{106.8} \int_{\R^{n}}\dd^{n}x\; \exp\!\left(\tfrac{\ii}{2}x\transpose Mx\right) =\frac{(2\pi)^{n/2}}{\abs{\det M}^{1/2}}\; \ee^{\ii\pi\sigma(M)/4}\ec \end{equation}

\(\sigma(M)\) being the signature of \(M\), i.e. the number of positive minus the number of negative eigenvalues. Rests on Equation (17.42) and Lemma 8.9.

Derivation. Derives Lemma 106.2. Diagonalize: \(A=O\transpose DO\) with \(O\) orthogonal and \(D=\diag(a_{1},\dots,a_{n})\), \(a_{i}>0\), which is possible by the finite-dimensional spectral theorem for real symmetric matrices. That theorem is quoted here from textbook linear algebra and is not proved anywhere in this book: the subsection it belongs to is Section 12.4.2, and Remark 11.86 records the same import where it is made in the probability chapter. The substitution \(y=Ox\) has unit Jacobian and factorizes the integral into \(n\) copies of the one-dimensional case. Completing the square in one dimension, \(-\frac{1}{2}ay^{2}+\beta y =-\frac{1}{2}a(y-\beta/a)^{2}+\beta^{2}/(2a)\), and \(\int\dd u\,\ee^{-au^{2}/2}=\sqrt{2\pi/a}\) — the Gauss integral, which is Equation (17.42) of Fourier Analysis and Integral Transforms evaluated at zero frequency. Reassembling, \(\prod_{i}a_{i}=\det A\) and \(\sum_{i}\beta_{i}^{2}/a_{i}=b\transpose A^{-1}b\), which is Equation (106.7).

Both sides of Equation (106.7) are holomorphic in the entries of \(A\) on the domain \(\Re A>0\) — the left-hand side because the integral converges absolutely and uniformly there, the right-hand side because \(\det A\neq0\) — so they agree throughout it by the identity theorem (Complex Analysis). For the purely imaginary boundary case, diagonalize \(M\) and rotate each contour: for \(a>0\),

\[ \int_{-\infty}^{\infty}\dd u\;\ee^{\ii au^{2}/2} =\lim_{\delta\to0^{+}}\int_{-\infty}^{\infty}\dd u\; \ee^{-(\delta-\ii a)u^{2}/2} =\lim_{\delta\to0^{+}}\sqrt{\frac{2\pi}{\delta-\ii a}} =\sqrt{\frac{2\pi}{a}}\;\ee^{\ii\pi/4}\ec \]

the limit being uniform on compacta by Cauchy's theorem applied to the wedge \(0\leq\arg u\leq\pi/4\), whose arc contribution vanishes by the estimation lemma Lemma 8.9. A negative eigenvalue rotates the other way and contributes \(\ee^{-\ii\pi/4}\); collecting the phases gives the signature.

The phase \(\ee^{\ii\pi\sigma/4}\) in Equation (106.8) is not decoration. It is the origin of the Maslov index, which is what makes the semiclassical propagator of Section 106.1.3 agree with the turning-point counting of Remark 9.79.

Composition

Theorem 106.3 (Feynman's sum over histories).

Let \(H=p^{2}/2m+V(q)\) as above and let \(t''>t'\). Divide the interval into \(N\) steps of length \(\Delta t=(t''-t')/N\) and write \(q_{k}=q(t'+k\Delta t)\), \(q_{0}=q'\), \(q_{N}=q''\). Then

\begin{equation}\tag{106.9} K(q'',t'';q',t') =\lim_{N\to\infty} \left(\frac{m}{2\pi\ii\hbar\Delta t}\right)^{N/2} \int\prod_{k=1}^{N-1}\dd q_{k}\; \exp\!\left[\frac{\ii}{\hbar}\sum_{k=1}^{N} \Delta t L_{k}\right]\ec \end{equation}

with \(L_{k}=\frac{m}{2}\bigl((q_{k}-q_{k-1})/\Delta t\bigr)^{2} -V\bigl((q_{k}+q_{k-1})/2\bigr)\). Written compactly,

\begin{equation}\tag{106.10} K(q'',t'';q',t') =\int_{q(t')=q'}^{q(t'')=q''}\!\!\mathcal{D}q\; \exp\!\left(\frac{\ii}{\hbar}S[q]\right)\ec \end{equation}

where \(\mathcal{D}q\) denotes the limit in Equation (106.9) and nothing else. Rests on Equation (106.1) and Lemma 106.1.

Derivation. Derives Theorem 106.3. The evolution operator composes exactly, \(U(t'',t')=U(t'',t_{N-1})\cdots U(t_{1},t')\), because \(H\) is time-independent. Insert \(\identity=\int\dd q_{k}\ketbra{q_{k}}{q_{k}}\) between consecutive factors:

\[ K=\int\prod_{k=1}^{N-1}\dd q_{k}\; \prod_{k=1}^{N} \bra{q_{k}}\ee^{-\ii H\Delta t/\hbar}\ket{q_{k-1}}\ec \]

which is exact for every \(N\). Now apply Lemma 106.1 to each factor. Each contributes the normalization \((m/2\pi\ii\hbar\Delta t)^{1/2}\) — there are \(N\) of them, one per step, against \(N-1\) integrations — and the phase \(\ee^{\ii\Delta t L_{k}/\hbar}\), and the individual errors \(O((\Delta t)^{2})\) accumulate to \(N\cdot O((\Delta t)^{2}) =O(\Delta t)\), which vanishes with \(N\to\infty\). The exponents add, giving the Riemann sum \(\sum_{k}\Delta t L_{k}\) for the action \(S[q]=\int_{t'}^{t''}\dd t\,L\) of the piecewise-linear path through the sampled points.

Three features of Equation (106.9) deserve emphasis, because they are what makes the formalism different from Canonical Quantization of Fields rather than equivalent to it in disguise.

First, every path contributes, and every path contributes with the same magnitude. The integrand's modulus is one; only the phase \(S[q]/\hbar\) distinguishes paths. There is no sense in which the classical path is weighted more heavily. Its distinction is that its neighbours share its phase — \(\delta S=0\) means precisely that \(S\) is stationary — so that contributions near it add coherently while contributions elsewhere cancel. That is the content of Section 106.1.3.

Second, the paths summed over are not classical trajectories and are not even differentiable. The measure concentrates \(\abs{q_{k}-q_{k-1}}\) on \(\sqrt{\hbar\Delta t/m}\), so the typical increment scales as \((\Delta t)^{1/2}\) and the typical “velocity” as \((\Delta t)^{-1/2}\), which diverges. The paths dominating Equation (106.9) are continuous and nowhere differentiable, the same class that carries Wiener measure in Section 106.2.2.

Third, the object \(S[q]\) is the classical action — a Lorentz scalar in the relativistic case, a functional of the path and of nothing else. No operator ordering appears in Equation (106.10), no Hilbert space, no choice of time slicing beyond the one used to define the limit. That is why the formalism generalizes to gauge fields, where the operator formulation must first solve a constraint problem (Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism) before it can begin.

Equivalence with the Schrödinger equation

Proposition 106.4 (From the kernel to the Schrödinger equation).

Let \(\psi\) evolve by \(\psi(q,t+\Delta t)=\int\dd q'\,K(q,t+\Delta t;q',t)\psi(q',t)\) with \(K\) given by Equation (106.5). Then to first order in \(\Delta t\),

\begin{equation}\tag{106.11} \ii\hbar\,\pdv{\psi}{t} =-\frac{\hbar^{2}}{2m}\,\frac{\pp^{2}\psi}{\pp q^{2}} +V(q)\psi\ep \end{equation}

Rests on Equations (106.5) and (106.8).

Derivation. Derives Proposition 106.4. Write \(q'=q+\eta\). Then

\[ \psi(q,t+\Delta t) =\sqrt{\frac{m}{2\pi\ii\hbar\Delta t}} \int\dd\eta\; \exp\!\left[\frac{\ii m\eta^{2}}{2\hbar\Delta t} -\frac{\ii\Delta t}{\hbar}V\!\left(q+\tfrac{\eta}{2}\right) \right]\psi(q+\eta,t)\ep \]

The Gaussian factor suppresses all but \(\eta=O(\sqrt{\hbar \Delta t/m})\), so expanding to the order that survives means keeping \(\eta^{0},\eta^{1},\eta^{2}\) in \(\psi\) and \((\Delta t)^{0}, (\Delta t)^{1}\) elsewhere:

\[ \ee^{-\ii\Delta t V/\hbar}\approx1-\frac{\ii\Delta t}{\hbar}V(q)\ec \qquad \psi(q+\eta)\approx\psi+\eta\,\pdv{\psi}{q} +\frac{\eta^{2}}{2}\frac{\pp^{2}\psi}{\pp q^{2}}\ep \]

By Equation (106.8) with \(M=m/(\hbar\Delta t)\),

\[ \sqrt{\frac{m}{2\pi\ii\hbar\Delta t}} \int\dd\eta\,\ee^{\ii m\eta^{2}/2\hbar\Delta t}=1\ec\qquad \int\dd\eta\;\eta\,\ee^{\ii m\eta^{2}/2\hbar\Delta t}=0\ec \]

by parity, and differentiating the first with respect to \(m/(\hbar\Delta t)\) gives the second moment

\[ \sqrt{\frac{m}{2\pi\ii\hbar\Delta t}} \int\dd\eta\;\eta^{2}\,\ee^{\ii m\eta^{2}/2\hbar\Delta t} =\frac{\ii\hbar\Delta t}{m}\ep \]

Collecting,

\[ \psi(q,t+\Delta t) =\psi-\frac{\ii\Delta t}{\hbar}V\psi +\frac{\ii\hbar\Delta t}{2m}\frac{\pp^{2}\psi}{\pp q^{2}} +O((\Delta t)^{2})\ec \]

and \((\psi(q,t+\Delta t)-\psi(q,t))/\Delta t\to\pp\psi/\pp t\) gives Equation (106.11) after multiplication by \(\ii\hbar\). Note where the two derivatives came from: the potential from the first-order term of the exponential, the kinetic operator from the second moment of the Gaussian, which is why the dimensionless combination controlling the expansion is \(\hbar\Delta t/m\) and why the paths have the roughness they do.

Interference, and the evidence that the action is primitive

The two-slit experiment is the elementary case of Theorem 106.3: the paths from source to detector fall into two homotopy classes, one through each slit, and within each class the contributions add coherently around the stationary path. Writing the two class sums as \(A_{1}\) and \(A_{2}\), the probability density on the screen is \(\abs{A_{1}+A_{2}}^{2} =\abs{A_{1}}^{2}+\abs{A_{2}}^{2} +2\Re(A_{1}^{*}A_{2})\), and the phase of \(A_{1}^{*}A_{2}\) is \((S_{2}-S_{1})/\hbar\). For a free particle of momentum \(p\) the stationary paths are straight, \(S=p\ell\) along a segment of length \(\ell\), so slits separated by \(d\) and a screen at distance \(L\gg d\) give a path difference \(\ell_{2}-\ell_{1}\approx dx/L\) and fringes wherever \(p\,dx/(L\hbar)\) is a multiple of \(2\pi\): the spacing is \(2\pi\hbar L/(pd)=hL/(pd)\), the de Broglie result of Matter Waves. That the pattern is built out of individual particles, each of which arrives at one point and none of which “goes through both slits”, was shown directly by Tonomura and co-workers, who recorded the electron-by-electron accumulation of the interference pattern from a biprism [Tonomura:1989].

The sharpest evidence that the path integral's primitive quantity is the action, and not the force, is the Aharonov–Bohm effect.

Phenomenon 106.5 (The Aharonov–Bohm effect).

An electron beam split into two arms that enclose a magnetic flux \(\Phi\) confined entirely to a region the beam never enters shows an interference pattern whose fringes shift with \(\Phi\), periodically with period

\begin{equation}\tag{106.12} \Phi_{0}=\frac{h}{e} =4.135667696\times 10^{-15}\,\mathrm{Wb}\ec \end{equation}

even though the electric and magnetic fields vanish identically along both arms [Aharonov:1959]. Rests on Theorem 106.3 and Equation (100.8).

Derivation. Derives Phenomenon 106.5. The Lagrangian of a particle of charge \(Q\) in an electromagnetic field is \(L=\frac{1}{2}m\dot{\vect{x}}^{2} +Q\,\dot{\vect{x}}\cdot\vect{A}-Q\phi\), whose Euler–Lagrange equation is \(m\ddot{\vect{x}}=Q(\vect{E}+\dot{\vect{x}}\times\vect{B})\), the Lorentz force; that derivation belongs to Generalized Classical Field Theory, with the covariant coupling \(-A_{\mu}J^{\mu}\), and is not repeated here. Along any path \(\gamma\) the action therefore contains the term

\[ S_{\mathrm{mag}}[\gamma]=Q\int_{\gamma}\dd t\; \dot{\vect{x}}\cdot\vect{A} =Q\int_{\gamma}\vect{A}\cdot\dd\vect{x}\ec \]

which depends on the path but not on its parametrization. For two paths \(\gamma_{1},\gamma_{2}\) with the same endpoints, the difference of the two actions is the circulation around the closed loop \(\gamma_{2}-\gamma_{1}\),

\[ S_{\mathrm{mag}}[\gamma_{2}]-S_{\mathrm{mag}}[\gamma_{1}] =Q\oint\vect{A}\cdot\dd\vect{x} =Q\int_{\Sigma}\vect{B}\cdot\dd\vect{S}=Q\Phi\ec \]

by Stokes's theorem, \(\Sigma\) being any surface spanning the loop. The interference term in Theorem 106.3 therefore acquires the phase

\begin{equation}\tag{106.13} \Delta\varphi=\frac{Q\Phi}{\hbar}\ec \end{equation}

and for the electron, \(Q=-e\), the pattern repeats when \(e\Phi/\hbar=2\pi\), that is when \(\Phi\) advances by \(h/e\). The numerical value in Equation (106.12) is exact in the present SI, both \(h\) and \(e\) being defined constants [Mohr:2025].

The point is the one Aharonov and Bohm made. The surface \(\Sigma\) necessarily passes through the solenoid, where \(\vect{B}\neq0\); the paths do not. In a formulation whose primitive quantities are the fields, the effect is impossible; in a formulation whose primitive quantity is \(\int\vect{A}\cdot\dd\vect{x}\), it is the first thing one would predict. What is measurable is not \(\vect{A}\), which is gauge dependent, but the loop integral \(\oint\vect{A}\cdot\dd\vect{x}\), which is not — a gauge transformation Equation (100.8) changes it by \(\oint\pp_{\mu}\Lambda\,\dd x^{\mu}=0\) for a single-valued \(\Lambda\). The gauge-invariant content of the potential is exactly its holonomy, and Section 106.5.1 is built on that observation.

Remark 106.6 (Two flux quanta, and which is which).

The period Equation (106.12) is \(h/e\) because the interfering object carries charge \(e\). The superconducting flux quantum is \(h/2e=2.067833848\times 10^{-15}\,\mathrm{Wb}\), half as large, because the interfering object there is a Cooper pair of charge \(2e\). Neither number is a property of the electromagnetic field; both are properties of the charge of the thing whose phase is being compared.

Time slicing, the measure and the classical limit

Convergence of the product formula

Theorem 106.3 took the splitting \(\ee^{-\ii(T+V)\Delta t/\hbar}\approx \ee^{-\ii T\Delta t/\hbar}\ee^{-\ii V\Delta t/\hbar}\) on the strength of a formal expansion. That the product converges as \(N\to\infty\) is a theorem about operator semigroups, not about exponentials of numbers, and it is due to Trotter [Trotter:1959].

Theorem 106.7 (Trotter product formula).

Let \(T\) and \(V\) be self-adjoint operators on a Hilbert space such that \(T+V\) is essentially self-adjoint on the intersection of their domains. Then, in the strong operator topology,

\begin{equation}\tag{106.14} \ee^{-\ii(T+V)t/\hbar} =\lim_{N\to\infty}\left(\ee^{-\ii Tt/N\hbar}\, \ee^{-\ii Vt/N\hbar}\right)^{N}\ep \end{equation}

Rests on Equation (106.1).

Derivation, bounded case. Derives Theorem 106.7. Let \(T\) and \(V\) be bounded operators, write \(\Delta t=t/N\), \(X_{N}=\ee^{-\ii(T+V)\Delta t/\hbar}\) and \(Y_{N}=\ee^{-\ii T\Delta t/\hbar}\ee^{-\ii V\Delta t/\hbar}\). Both are contractions, \(\norm{X_{N}}\leq1\) and \(\norm{Y_{N}}\leq1\), and expanding both exponentials gives \(\norm{X_{N}-Y_{N}}\leq K(\Delta t)^{2}\) with \(K=\norm{\comm{T}{V}}/2\hbar^{2}\), since the terms of order \(\Delta t^{0}\) and \(\Delta t^{1}\) agree. The telescoping identity

\[ X_{N}^{N}-Y_{N}^{N} =\sum_{k=0}^{N-1}X_{N}^{\,N-1-k} \bigl(X_{N}-Y_{N}\bigr)Y_{N}^{\,k} \]

then bounds \(\norm{X_{N}^{N}-Y_{N}^{N}}\leq N\norm{X_{N}-Y_{N}} \leq Kt^{2}/N\to0\). Unbounded \(T\) and \(V\) — the physical case, since \(p^{2}\) is unbounded — require the strong-topology argument of Trotter [Trotter:1959], whose extension to operators bounded relative to \(T\) is Kato's [Kato:1966]; the mechanism is the same telescoping, applied to the resolvents.

The hypothesis is met for \(T=p^{2}/2m\) and any \(V\) that is Kato-small with respect to \(T\), which includes every potential of physical interest here, the Coulomb potential among them [Kato:1966]. This matters more than it looks: without Theorem 106.7 the sliced expression Equation (106.9) would be a mnemonic, and its limit would have no reason to be the propagator of the theory one started from.

What $\mathcal{D}q$ is not

Remark 106.8 (There is no Feynman measure).

The symbol \(\mathcal{D}q\) in Equation (106.10) is not a measure. There is no countably additive complex measure of finite total variation on the space of continuous paths that reproduces Equation (106.9), and the obstruction is not technical: a measure with the required translation and scaling behaviour would have to assign every cylinder set a weight of unit modulus, and the total variation of such an object over the \(N\)-fold product diverges like \(\abs{2\pi\ii\hbar\Delta t/m}^{-N/2}\). The precise statement and its consequences are Nelson's [Nelson:1964]. The Euclidean continuation of Section 106.2 repairs exactly this: the weight there is real and positive, the total variation is finite, and the measure exists — it is Wiener's [Wiener:1923].

This is not a licence to be careless. It means that every manipulation of \(\mathcal{D}q\) in this chapter is shorthand for a manipulation of the finite-dimensional integral in Equation (106.9), to be justified there. The two places where the difference bites are the ordering ambiguity below and the Jacobian of a change of variables, which is the subject of Section 106.7.3.

Ordering, and the midpoint rule

Lemma 106.1 placed \(V\) at the midpoint and remarked that for a potential the choice is immaterial at the order retained. It is not immaterial when the Hamiltonian depends on momentum beyond \(p^{2}\), and the case that matters physically is the magnetic one.

Proposition 106.9 (The midpoint rule is forced by gauge invariance).

For \(H=\bigl(\vect{p}-Q\vect{A}(\vect{x})\bigr)^{2}/2m\) the short-time kernel is

\begin{equation}\tag{106.15} \bra{\vect{x}''}\ee^{-\ii H\Delta t/\hbar}\ket{\vect{x}'} =\left(\frac{m}{2\pi\ii\hbar\Delta t}\right)^{3/2} \exp\!\left\{\frac{\ii}{\hbar}\left[ \frac{m(\vect{x}''-\vect{x}')^{2}}{2\Delta t} +Q\,(\vect{x}''-\vect{x}')\cdot \vect{A}\!\left(\frac{\vect{x}''+\vect{x}'}{2}\right) \right]\right\}\ec \end{equation}

to the accuracy of Equation (106.5). Evaluating \(\vect{A}\) at either endpoint instead gives an answer that differs at order \(\Delta t\) and does not reproduce the Schrödinger equation. Rests on Equations (106.5) and (106.6).

Derivation. Derives Proposition 106.9. Repeat the momentum-space computation Equation (106.6) with \(H\) quadratic in \(\vect{p}-Q\vect{A}\). Writing \(\vect{A}\) at a fixed point \(\vect{x}_{*}\) inside the segment, the \(\vect{p}\) integral is Gaussian with a shifted centre and gives Equation (106.15) with \(\vect{A}(\vect{x}_{*})\) in place of the midpoint value. To fix \(\vect{x}_{*}\), demand that the kernel transform correctly under a gauge transformation \(\vect{A}\to\vect{A}+\nabla\Lambda\), under which the wavefunction acquires the phase \(\ee^{\ii Q\Lambda(\vect{x})/\hbar}\) and hence the kernel must acquire \(\ee^{\ii Q[\Lambda(\vect{x}'')-\Lambda(\vect{x}')]/\hbar}\). The exponent in Equation (106.15) changes by \((\ii Q/\hbar)(\vect{x}''-\vect{x}')\cdot\nabla\Lambda(\vect{x}_{*})\), and

\[ \Lambda(\vect{x}'')-\Lambda(\vect{x}') =(\vect{x}''-\vect{x}')\cdot\nabla\Lambda(\vect{x}_{*}) +O\bigl(\abs{\vect{x}''-\vect{x}'}^{3}\bigr) \]

holds only for the midpoint \(\vect{x}_{*}=(\vect{x}''+ \vect{x}')/2\); at any other point the error is quadratic, i.e. of order \(\hbar\Delta t/m\) in the exponent, which survives the limit. Endpoint prescriptions therefore quantize a different Hamiltonian, differing from \(H\) by a term proportional to \(\nabla\cdot\vect{A}\).

The moral generalizes: the time-sliced integral is a definition, and different discretizations of the same classical Lagrangian can define different quantum theories. What selects the right one is a symmetry that must survive — gauge invariance here, and in Section 106.7.3 chiral symmetry, where the answer turns out to be that no discretization preserves it and the failure is observable.

The classical limit

Theorem 106.10 (Stationary phase).

Let \(f:\R^{n}\to\R\) be smooth with a single non-degenerate critical point \(x_{0}\), \(\nabla f(x_{0})=0\) and \(\det f''(x_{0})\neq0\), and let \(g\) be smooth and compactly supported near \(x_{0}\). Then as \(\hbar\to0^{+}\)

\begin{equation}\tag{106.16} \int\dd^{n}x\;g(x)\,\ee^{\ii f(x)/\hbar} =\frac{(2\pi\hbar)^{n/2}\,g(x_{0})} {\abs{\det f''(x_{0})}^{1/2}}\, \ee^{\ii f(x_{0})/\hbar}\, \ee^{\ii\pi\sigma/4}\bigl[1+O(\hbar)\bigr]\ec \end{equation}

with \(\sigma\) the signature of \(f''(x_{0})\). Rests on Equation (106.8).

Derivation. Derives Theorem 106.10. Away from \(x_{0}\) the integrand oscillates: writing \(\ee^{\ii f/\hbar}=(\hbar/\ii)\nabla f\cdot\nabla \ee^{\ii f/\hbar}/\abs{\nabla f}^{2}\) and integrating by parts \(k\) times produces a factor \(\hbar^{k}\), so those regions contribute beyond all orders. Near \(x_{0}\), substitute \(x=x_{0}+\sqrt{\hbar}\,u\) and expand \(f(x)=f(x_{0})+\frac{\hbar}{2}u\transpose f''(x_{0})u+O(\hbar^{3/2})\). The measure supplies \(\hbar^{n/2}\), the exponent becomes \(\ii f(x_{0})/\hbar+\frac{\ii}{2}u\transpose f''u\), and Equation (106.8) evaluates the remaining integral.

Apply this to Equation (106.9), whose integration variables are the \(N-1\) intermediate positions and whose phase is the discretized action. The critical point of \(S\) is the path satisfying \(\pp S/\pp q_{k}=0\) for every \(k\), which is the discretized Euler–Lagrange equation, i.e. Hamilton's principle (Calculus of Variations); in the limit it is the classical trajectory. Therefore

\begin{equation}\tag{106.17} K(q'',t'';q',t') \;\underset{\hbar\to0}{\longrightarrow}\; \sqrt{\frac{1}{2\pi\ii\hbar} \left(-\pdv{^{2}S_{\mathrm{cl}}}{q''\pp q'}\right)}\; \exp\!\left[\frac{\ii}{\hbar}S_{\mathrm{cl}}(q'',t'';q',t') -\frac{\ii\pi\nu}{2}\right]\ec \end{equation}

the Van Vleck form, with \(\nu\) the number of conjugate points passed — the Maslov index, the accumulated signature phase of Equation (106.8). Dirac's correspondence Equation (106.4) is recovered as the leading term, with the normalization now derived rather than guessed.

Proposition 106.11 (The semiclassical propagator is the WKB wavefunction).

For a particle of energy \(E\) in one dimension, the Fourier transform of Equation (106.17) with respect to \(t''-t'\) has stationary phase at the classical time of flight and gives, up to normalization,

\begin{equation}\tag{106.18} \psi(q)\propto\frac{1}{\sqrt{p(q)}}\, \exp\!\left[\pm\frac{\ii}{\hbar}\int^{q}p(u)\,\dd u\right]\ec \qquad p(q)=\sqrt{2m\bigl(E-V(q)\bigr)}\ec \end{equation}

which is the semiclassical wave Equation (9.69) of Corollary 9.62, where the local momentum is written \(P\) and the amplitude factor \(P^{-1/2}\) is derived from the transport equation rather than from a stationary phase. Rests on Equation (106.3), Equation (106.17) and Theorem 106.10.

Derivation. Derives Proposition 106.11. The energy eigenfunctions are obtained from \(K\) by \(\int\dd t\,\ee^{\ii Et/\hbar}K(q,t;q',0)\), a Fourier transform of the type of Fourier Analysis and Integral Transforms. Applying Theorem 106.10 in the single variable \(t\), the stationary point is where \(\pp S_{\mathrm{cl}}/\pp t=-E\) by Equation (106.3), i.e. at the classical time of flight at energy \(E\), and the phase there becomes the Legendre transform \(W(q,q'';E)=S_{\mathrm{cl}}+Et=\int p\,\dd q\), Hamilton's characteristic function. The prefactor collects \((\pp^{2}W/\pp q\pp q'')^{1/2}\), which for one degree of freedom is \((\pp p/\pp q'')^{1/2}\propto p(q)^{-1/2}\) after the reciprocity relation. This is Equation (106.18).

Proposition 106.11 is the exact junction between this chapter and Section 9.6: the WKB ansatz \(\psi=\exp(\ii S/\hbar)\) of Equation (9.62) and the path-integral weight \(\exp(\ii S/\hbar)\) of Equation (106.10) are the same object seen from two sides, and the expansion of the exponent in powers of \(\hbar\) in Equation (9.64) is the loop expansion of Section 106.3.4. In particular the \(\hbar^{1}\) term of the WKB series, which produces the \(p^{-1/2}\) amplitude, is the one-loop determinant, and the failure of the expansion at a turning point (Remark 9.66) is the failure of Theorem 106.10 at a degenerate critical point.

Euclidean formulation

Wick rotation

Remark 106.8 recorded that the oscillatory weight \(\ee^{\ii S/\hbar}\) defines no measure. The remedy is to continue the time variable into the complex plane until the weight becomes real and decaying. The substitution is Wick's [Wick:1954].

Definition 106.12 (Euclidean time).

Set \(t=-\ii\tau\) with \(\tau\) real, of SI dimension \(\mathrm{s}\). Equivalently \(x^{0}=ct\mapsto-\ii x^{4}\) with \(x^{4}=c\tau\), and the metric \(\eta_{\mu\nu}=\diag(+1,-1,-1,-1)\) becomes \(-\delta_{\mu\nu}\): squared four-vectors change sign, \(p^{2}\mapsto-p_{E}^{2}\) with \(p_{E}^{2}=(p^{4})^{2}+\vect{p}^{2}\geq0\).

Proposition 106.13 (The Euclidean weight).

Under Definition 106.12 the action of a particle in a potential becomes \(S=\ii S_{E}\) with

\begin{equation}\tag{106.19} S_{E}[q]=\int\dd\tau\left[ \frac{m}{2}\left(\dv{q}{\tau}\right)^{2}+V(q)\right]\ec \end{equation}

of dimension \(\mathrm{J}\,\mathrm{s}\), so the path-integral weight is

\begin{equation}\tag{106.20} \exp\!\left(\frac{\ii}{\hbar}S\right) \longmapsto\exp\!\left(-\frac{1}{\hbar}S_{E}\right)\ec \end{equation}

real, positive and bounded by one, since \(S_{E}\geq0\) for \(V\geq0\). The Euclidean action is the sum, not the difference, of kinetic and potential terms: the Euclidean motion takes place in the inverted potential \(-V\). Rests on Definition 106.12.

Derivation. Derives Proposition 106.13. With \(t=-\ii\tau\), \(\dd t=-\ii\,\dd\tau\) and \(\dd q/\dd t=\ii\,\dd q/\dd\tau\), so \(\frac{m}{2}(\dd q/\dd t)^{2}=-\frac{m}{2}(\dd q/\dd\tau)^{2}\) and

\[ S=\int\dd t\left[\frac{m}{2}\left(\dv{q}{t}\right)^{2}-V\right] =\int(-\ii\,\dd\tau)\left[-\frac{m}{2} \left(\dv{q}{\tau}\right)^{2}-V\right] =\ii\int\dd\tau\left[\frac{m}{2} \left(\dv{q}{\tau}\right)^{2}+V\right]\ec \]

whence \(\ii S/\hbar=\ii\cdot\ii S_{E}/\hbar=-S_{E}/\hbar\).

The Euclidean weight is a genuine probability density on path space once normalized — that is the subject of Section 106.2.2 — and it is what makes the numerical work of Section 106.5.2 possible at all. What must be established is that the continuation is legitimate: that the Minkowski amplitude may be recovered from the Euclidean one by rotating back.

The pole structure that licenses the rotation

Proposition 106.14 (The Feynman contour permits the rotation).

Let \(f(p^{0})\) be a product of Feynman propagators Equation (100.23) and scalar propagators \(\ii\hbar^{3}c/(p^{2}-m^{2}c^{2}+\ii\epsilon)\) and let the integrand fall faster than \(\abs{p^{0}}^{-1}\) at infinity. Then

\begin{equation}\tag{106.21} \int_{-\infty}^{\infty}\dd p^{0}\,f(p^{0}) =\ii\int_{-\infty}^{\infty}\dd p^{4}\,f(\ii p^{4})\ec \end{equation}

the contour being rotated counterclockwise from the real to the imaginary axis. Rests on Equation (100.23), Theorem 8.12 and Lemma 8.9.

Derivation. Derives Proposition 106.14. A propagator denominator \(p^{2}-m^{2}c^{2}+\ii\epsilon =(p^{0})^{2}-\vect{p}^{2}-m^{2}c^{2}+\ii\epsilon\) vanishes at

\[ p^{0}=\pm\left(\sqrt{\vect{p}^{2}+m^{2}c^{2}} -\ii\epsilon'\right)\ec\qquad\epsilon'>0\ec \]

to first order in \(\epsilon\): the positive-energy pole sits just below the positive real axis and the negative-energy pole just above the negative real axis, i.e. in the fourth and second quadrants respectively. The first and third quadrants are therefore free of singularities, and the closed contour consisting of the real axis, the imaginary axis traversed downwards, and the two connecting arcs in those quadrants encloses no pole. By Cauchy's theorem (Theorem 8.12) its integral vanishes; the arcs contribute nothing by the estimation lemma Lemma 8.9 given the assumed fall-off. Rotating \(p^{0}=\ee^{\ii\theta}\abs{p^{0}}\) from \(\theta=0\) to \(\theta=\pi/2\) and writing \(p^{0}=\ii p^{4}\) gives \(\dd p^{0}=\ii\,\dd p^{4}\), which is Equation (106.21). Under it \(p^{2}=(p^{0})^{2}-\vect{p}^{2}\mapsto-(p^{4})^{2}-\vect{p}^{2} =-p_{E}^{2}\), so the denominator becomes \(-(p_{E}^{2}+m^{2}c^{2})\), which never vanishes for real \(p_{E}\): the Euclidean propagator has no poles at all, and no \(\ii\epsilon\) is needed.

The \(\ii\epsilon\) prescription of Equation (100.23) is thus not a computational convenience but the statement of which boundary value of an analytic function the theory picks — and it is exactly the choice that puts the poles out of the way of the rotation. That is the same fact seen twice: the Feynman propagator is the one whose Euclidean continuation exists.

Remark 106.15 (What the rotation costs).

Three things are lost or endangered.

Real-time information. The Euclidean correlator determines the Minkowski one only through an analytic continuation, and continuing numerical data from a finite set of Euclidean times back to real time is an ill-posed inverse problem: two spectral functions differing by an oscillation of high frequency give Euclidean correlators differing by less than the statistical error. This is why the lattice computes masses and matrix elements well and transport coefficients badly.

Positivity of the weight, at finite chemical potential. Introducing a chemical potential \(\mu\) for a conserved charge adds \(\mu\) to the time component of the covariant derivative in the fermion operator; after the rotation this term is imaginary relative to the rest, the determinant \(\det(\ii\hbar\gamma^{\mu}D_{\mu}-mc)\) of Section 106.3.3 becomes complex, and \(\ee^{-S_{E}/\hbar}\) is no longer a probability. This is the sign problem. It is not a difficulty of technique: no reweighting scheme has been found whose cost is not exponential in the volume, and the phase diagram of QCD at large baryon density is consequently not known from first principles. The honest position is that Section 106.5.2's successes are successes at zero or small chemical potential.

Nothing else. At zero chemical potential the rotation is exact, and Section 106.2.3 states the conditions under which the Euclidean theory contains the whole of the Minkowski one.

The Feynman–Kac formula

Theorem 106.16 (Feynman–Kac).

Let \(H=p^{2}/2m+V(q)\) with \(V\) bounded below and continuous. Then for \(\tau>0\)

\begin{equation}\tag{106.22} \bra{q''}\ee^{-H\tau/\hbar}\ket{q'} =\int_{q(0)=q'}^{q(\tau)=q''}\!\!\mathcal{D}q\; \ee^{-S_{E}[q]/\hbar} =\rho_{\tau}(q''-q')\; \mathbb{E}_{q'\to q''}\!\left[ \exp\!\left(-\frac{1}{\hbar}\int_{0}^{\tau}V(q(s))\,\dd s\right) \right]\ec \end{equation}

where the expectation is over the Wiener measure of Brownian bridges from \(q'\) to \(q''\) with diffusion constant

\begin{equation}\tag{106.23} D=\frac{\hbar}{2m}\ec\qquad [D]=\mathrm{m}^{2}/\mathrm{s}\ec \end{equation}

and \(\rho_{\tau}(x)=(4\pi D\tau)^{-1/2}\ee^{-x^{2}/4D\tau}\) is the free heat kernel. Rests on Theorem 106.3 and Proposition 106.13.

Derivation. Derives Theorem 106.16. Repeat the slicing of Theorem 106.3 with \(\ee^{-H\tau/\hbar}\) in place of \(\ee^{-\ii Ht/\hbar}\); every step goes through with \(\ii\varepsilon\mapsto\Delta\tau\), giving

\[ \bra{q''}\ee^{-H\tau/\hbar}\ket{q'} =\lim_{N\to\infty} \left(\frac{m}{2\pi\hbar\Delta\tau}\right)^{N/2} \int\prod_{k=1}^{N-1}\dd q_{k}\; \exp\!\left\{-\frac{1}{\hbar}\sum_{k=1}^{N}\Delta\tau \left[\frac{m}{2} \left(\frac{q_{k}-q_{k-1}}{\Delta\tau}\right)^{2} +V\right]\right\}\ep \]

Now the prefactor and the kinetic term together are \(\prod_{k}\rho_{\Delta\tau}(q_{k}-q_{k-1})\) with \(\rho\) as stated, because \(m/(2\hbar\Delta\tau)=1/(4D\Delta\tau)\) with \(D\) of Equation (106.23). That product is precisely the finite-dimensional distribution of Brownian motion of diffusion constant \(D\) sampled at the times \(k\Delta\tau\), and by Kolmogorov's extension theorem it defines a genuine countably additive probability measure on continuous paths — Wiener's measure [Wiener:1923]. What remains in the exponent is \(-\frac{1}{\hbar}\sum\Delta\tau\,V\), the Riemann sum for \(\frac{1}{\hbar}\int V\), and the identification of the limit with the expectation is Kac's theorem [Kac:1949]. Conditioning on the endpoint \(q(\tau)=q''\) produces the bridge and the explicit factor \(\rho_{\tau}(q''-q')\).

That \(D=\hbar/2m\) is worth pausing over. The Euclidean quantum mechanics of a particle of mass \(m\) is a diffusion process whose diffusion constant is set by \(\hbar\); the classical limit \(\hbar\to0\) is the limit of vanishing diffusion, in which the paths collapse onto the single deterministic solution. For an electron, \(D=\hbar/2m_{e}=5.79\times 10^{-5}\,\mathrm{m}^{2}/\mathrm{s}\). The roughness noted after Theorem 106.3 is now a theorem rather than an estimate: Wiener paths are almost surely continuous, almost surely nowhere differentiable, and almost surely Hölder continuous of every exponent below \(1/2\).

Spectral consequences

Proposition 106.17 (Euclidean time projects onto the ground state).

Let \(H\) have discrete spectrum \(E_{0}<E_{1}\leq\cdots\) with eigenfunctions \(\psi_{n}\). Then

\begin{equation}\tag{106.24} \bra{q''}\ee^{-H\tau/\hbar}\ket{q'} =\sum_{n}\psi_{n}(q'')\psi_{n}^{*}(q')\, \ee^{-E_{n}\tau/\hbar} \;\underset{\tau\to\infty}{\longrightarrow}\; \psi_{0}(q'')\psi_{0}^{*}(q')\,\ee^{-E_{0}\tau/\hbar}\ec \end{equation}

so the large-\(\tau\) decay rate of a Euclidean correlator measures an energy, and the gap \(E_{1}-E_{0}\) governs the leading correction. Rests on Theorem 106.16.

Derivation. Derives Proposition 106.17. Insert the resolution of the identity in energy eigenstates into \(\ee^{-H\tau/\hbar}\); each term carries \(\ee^{-E_{n}\tau/\hbar}\) and the smallest exponent dominates. The subleading term is suppressed by \(\ee^{-(E_{1}-E_{0})\tau/\hbar}\).

Equation (106.24) is the entire basis of hadron spectroscopy on the lattice. A field theory correlator of an operator carrying the quantum numbers of a particle of mass \(m\) decays in Euclidean time as \(\ee^{-mc^{2}\tau/\hbar}\), and fitting that exponential to Monte Carlo data is how the numbers of Section 106.5.2 are obtained. The correlation length in the fourth direction is

\begin{equation}\tag{106.25} \xi=c\tau_{\mathrm{corr}}=\frac{\hbar}{mc}\ec \end{equation}

the Compton wavelength: a mass gap in the quantum theory is a finite correlation length in the statistical one, and a massless particle is a critical point.

Finite temperature

Theorem 106.18 (The partition function is a periodic path integral).

The canonical partition function of a system with Hamiltonian \(H\) at temperature \(T\) is

\begin{equation}\tag{106.26} Z(T)=\tr\,\ee^{-H/k_{B}T} =\oint_{q(0)=q(\beta\hbar)}\!\!\mathcal{D}q\; \ee^{-S_{E}[q]/\hbar}\ec\qquad \beta=\frac{1}{k_{B}T}\ec \end{equation}

the integral running over paths periodic in Euclidean time with period

\begin{equation}\tag{106.27} \tau_{\beta}=\beta\hbar=\frac{\hbar}{k_{B}T}\ec\qquad [\tau_{\beta}]=\mathrm{s}\ep \end{equation}

Rests on Equation (106.22).

Derivation. Derives Theorem 106.18. Take Equation (106.22) at \(\tau=\beta\hbar\), set \(q''=q'=q\) and integrate over \(q\): the left-hand side becomes \(\int\dd q\,\bra{q}\ee^{-\beta H}\ket{q}=\tr\ee^{-\beta H}\), and the right-hand side becomes the path integral restricted to closed paths. Fermionic degrees of freedom acquire antiperiodic boundary conditions instead, because the trace of a Grassmann bilinear picks up a sign; that computation is Section 106.3.3.

The thermal circumference Equation (106.27) is a length in Euclidean spacetime, \(c\tau_{\beta}=\hbar c/k_{B}T\). At room temperature, \(T=300\,\mathrm{K}\),

\begin{equation}\tag{106.28} \tau_{\beta}=\frac{\hbar}{k_{B}T} =\frac{1.054571817\times 10^{-34}\,\mathrm{J}\,\mathrm{s}} {1.380649\times 10^{-23}\,\mathrm{J}/\mathrm{K}\times300\,\mathrm{K}} =2.55\times 10^{-14}\,\mathrm{s}\ec \end{equation}

using the exactly defined Boltzmann constant [Mohr:2025]; at the temperature \(k_{B}T=150\,\mathrm{MeV}\) characteristic of the quark–hadron transition, \(\tau_{\beta}=4.39\times 10^{-24}\,\mathrm{s}\) and \(c\tau_{\beta}=1.32\,\mathrm{fm}\), comparable with the size of a hadron — which is why finite-temperature QCD is a strongly coupled problem and not a small correction.

Euclidean field theoryClassical statistical mechanics
Weight $\ee^{-S_{E}/\hbar}$Boltzmann factor $\ee^{-\Ham/k_{B}T}$
$\hbar$$k_{B}T$
Euclidean action $S_{E}$energy $\Ham$ of a configuration
Generating functional $Z[0]$partition function
Mass $m$ of the lightest stateinverse correlation length, $\xi=\hbar/mc$
Massless particlecritical point, $\xi\to\infty$
Continuum limit $a\to0$second-order critical point
Vacuum expectation valuethermal average
Temperature $T$ of the quantum systemcircumference $\hbar/k_{B}T$ of the fourth direction
The Euclidean dictionary. A quantum field theory in $3+1$ spacetime dimensions, continued to imaginary time, is a classical statistical system in four Euclidean dimensions; a quantum system at temperature $T$ is the same system on a slab of thickness $\hbar/k_{B}T$. Every entry is an identity between the two columns after Theorems 106.16 and 106.18, not an analogy.

Table 106.1 is the reason this chapter sits next to Statistical Mechanics in the architecture of the book, and why the renormalization group of The Renormalization Group is the same construction as the one applied to critical phenomena in Phase Transitions and Critical Phenomena. The identification \(\hbar\leftrightarrow k_{B}T\) fixes the direction of the loop expansion of Section 106.3.4: it is an expansion in powers of \(\hbar\), hence in powers of \(k_{B}T\), and so a low-temperature expansion of the statistical system — the saddle-point expansion about the configuration of least energy, with each loop one power of \(k_{B}T\) smaller than the last. The classical limit \(\hbar\to0\) is correspondingly its zero-temperature limit, in which fluctuations are frozen and the configuration of least energy is the only one. A genuine high-temperature expansion of the statistical system is an expansion in \(1/k_{B}T\) and is the strong-coupling expansion of Theorem 106.72, at the opposite end of the same axis.

Reconstruction of the real-time theory

The Euclidean theory is a computational device only if the Minkowski theory can be recovered from it. Osterwalder and Schrader gave necessary and sufficient conditions [Osterwalder:1973] [Osterwalder:1975].

Definition 106.19 (Schwinger functions).

The Euclidean correlation functions of a scalar field,

\begin{equation}\tag{106.29} S_{n}(x_{1},\dots,x_{n}) =\frac{1}{Z}\int\mathcal{D}\phi\; \phi(x_{1})\cdots\phi(x_{n})\,\ee^{-S_{E}[\phi]/\hbar}\ec \end{equation}

with all \(x_{i}\) in Euclidean \(\R^{4}\) and distinct.

Theorem 106.20 (Osterwalder–Schrader reconstruction).

Let a family \(\set{S_{n}}\) satisfy

  1. regularity: each \(S_{n}\) is a tempered distribution, with growth in \(n\) bounded so that the reconstructed fields are operator-valued distributions;

  2. Euclidean invariance: \(S_{n}\) is invariant under the rotations and translations of \(\SO(4)\ltimes\R^{4}\);

  3. reflection positivity: for the reflection \(\theta:(x^{4},\vect{x})\mapsto(-x^{4},\vect{x})\) and any finite family of test functions \(f_{k}\) supported in \(x^{4}>0\),

    \begin{equation}\tag{106.30} \sum_{j,k}\bar{c}_{j}c_{k}\, S\bigl(\theta \bar{f}_{j}\otimes f_{k}\bigr)\geq0\ec \end{equation}
  4. symmetry: \(S_{n}\) is invariant under permutations of its arguments;

  5. cluster decomposition: \(S_{n+m}(x_{1},\dots,x_{n}, y_{1}+a,\dots)\to S_{n}S_{m}\) as the Euclidean separation \(\abs{a}\to\infty\);

then there is a Wightman quantum field theory in \(3+1\) Minkowski dimensions — a Hilbert space with a positive-definite inner product, a unitary representation of the Poincaré group with spectrum in the forward cone, a unique vacuum, and local fields — whose vacuum correlators continue to the \(S_{n}\). Conversely, the Schwinger functions of any Wightman theory satisfy (i)–(v). The Wightman axioms so reconstructed are the subject of Axiomatic Quantum Field Theory. Rests on Definition 106.19.

Of the five, reflection positivity is the one that carries the physics, and it is worth seeing why in the simplest case, where the whole content is visible.

Proposition 106.21 (Reflection positivity is the positivity of the Hilbert space).

In Euclidean quantum mechanics, let \(F\) be any functional of the path depending only on \(q(s)\) for \(s>0\), and let \(\theta F\) be its reflection, depending only on \(s<0\). Then

\begin{equation}\tag{106.31} \avg{\overline{\theta F}\,F}=\norm{\Psi_{F}}^{2}\geq0\ec \end{equation}

where \(\Psi_{F}\) is a vector in the Hilbert space of the theory built from \(F\) and the ground state \(\ket{\Omega}\) by Theorem 106.16. Rests on Theorem 106.16 and Equation (106.30).

Derivation. Derives Proposition 106.21. Take \(F=\prod_{i}f_{i}(q(s_{i}))\) with \(0<s_{1}<\cdots<s_{n}\). By Theorem 106.16 the Euclidean expectation of a product of such factors is a time-ordered chain of operators separated by Euclidean evolutions,

\[ \avg{\overline{\theta F}\,F} =\bra{\Omega}\hat f_{n}^{\dagger}\ee^{-H(s_{n}-s_{n-1})/\hbar} \cdots\hat f_{1}^{\dagger}\,\hat f_{1} \cdots\ee^{-H(s_{n}-s_{n-1})/\hbar}\hat f_{n}\ket{\Omega}\ec \]

because reflecting \(s\mapsto-s\) reverses the order of the factors and conjugates them, while the intervening evolution operators \(\ee^{-H\Delta s/\hbar}\) are positive and self-adjoint. The whole expression is therefore \(\norm{\Psi}^{2}\) with \(\Psi=\hat f_{1}\ee^{-H(s_{2}-s_{1})/\hbar}\cdots\ket{\Omega}\), manifestly non-negative. Linear combinations give Equation (106.30).

Read backwards, this is the content of the axiom: the Euclidean functional integral supplies a bilinear form on functionals of the positive-time half-path, and reflection positivity says exactly that this form is a non-negative inner product. The physical Hilbert space is the completion of that space modulo null vectors, the operator \(\ee^{-H\tau/\hbar}\) is the translation in Euclidean time, and \(H\) is self-adjoint and bounded below because it generates a contraction semigroup. Without Equation (106.30) one obtains a vector space with an indefinite form and no probability interpretation — and no amount of Euclidean invariance repairs it.

Remark 106.22 (Reflection positivity on the lattice).

The lattice discretization of Section 106.5.1 must inherit Equation (106.30) or it is not the Euclidean form of any quantum theory, however well it reproduces the classical action. Wilson's plaquette action does inherit it, because it couples only nearest-neighbour time slices, so the transfer operator constructed above exists and is positive. Actions “improved” by adding couplings between next-to-nearest time slices generally do not: they can produce a transfer matrix with negative eigenvalues, whose signature is a correlator that oscillates in Euclidean time and which cannot be fitted by Equation (106.24) with any real spectrum. Improvement schemes are therefore constructed to act within a time slice, or their violation is arranged to vanish faster than the discretization errors they remove. This is a working constraint on the design of a lattice action, not a foundational nicety.

Derivation pending.

Osterwalder–Schrader reconstruction: the derivation above establishes the easy half in quantum mechanics — that reflection positivity is the positivity of the reconstructed inner product — and states the field-theoretic theorem. The construction of the Wightman functions from the Schwinger functions by analytic continuation in the time arguments, and the verification of the spectrum condition and locality, are not carried out here; they belong in Appendix A, alongside the axiomatic material of the axiomatic quantum field theory chapter of this part.

Functional quantization of fields

Generating functionals

A field is a mechanical system with one degree of freedom per point of space, so Theorem 106.3 carries over with \(q(t)\) replaced by \(\phi(x)\) and the time slicing replaced by a slicing in \(x^{0}\). The one thing that must be fixed first is dimensions, because the whole chapter turns on the exponent \(S/\hbar\) being a pure number.

Notation 106.23 (The scalar field and its source).

For a single real scalar field the Lagrangian density is

\begin{equation}\tag{106.32} \Lag=\frac{1}{2}\pp_{\mu}\phi\,\pp^{\mu}\phi -\frac{1}{2}\left(\frac{mc}{\hbar}\right)^{2}\phi^{2} -\frac{g}{4!}\phi^{4}\ec \end{equation}

of dimension \(\mathrm{J}/\mathrm{m}^{3}\) by Equation (100.2). Since \(\pp_{\mu}\) carries \(/\mathrm{m}\), the field carries

\begin{equation}\tag{106.33} [\phi]=\mathrm{J}^{1/2}\,\mathrm{m}^{-1/2} =\mathrm{kg}^{1/2}\,\mathrm{m}^{1/2}/\mathrm{s}\ec \end{equation}

the mass term carries \((mc/\hbar)^{2}\) of dimension \(/\mathrm{m}^{2}\) — the inverse squared Compton wavelength, so that the field equation is \(\bigl(\Box+(mc/\hbar)^{2}\bigr)\phi=0\) — and the quartic coupling carries

\begin{equation}\tag{106.34} [g]=/\mathrm{J}/\mathrm{m}\ec\qquad \hat g:=g\hbar c\;\text{ dimensionless}\ep \end{equation}

The source \(J\) is coupled by adding \(c^{-1}\int\dd^{4}x\,J\phi\) to the action, so \([J]=\mathrm{J}^{1/2}\,\mathrm{m}^{-5/2}\).

That \(\hat g=g\hbar c\) and not \(g\) is the dimensionless coupling is not bookkeeping. It says that the strength of a quartic self-interaction is measured against \(\hbar c=3.16153\times 10^{-26}\,\mathrm{J}\,\mathrm{m}\), the only quantity in the theory with the dimension of an energy times a length, exactly as the electromagnetic coupling is measured against it in Equation (100.3). A “dimensionless coupling” is always a ratio to \(\hbar c\), and saying so is what makes the classical limit visible: as \(\hbar\to0\) at fixed \(g\), every dimensionless coupling goes to zero and the theory becomes free and classical.

Definition 106.24 (The generating functional).
\begin{equation}\tag{106.35} Z[J]=\int\mathcal{D}\phi\; \exp\!\left\{\frac{\ii}{\hbar}\left(S[\phi] +\frac{1}{c}\int\dd^{4}x\;J(x)\phi(x)\right)\right\}\ec \end{equation}

with \(S[\phi]=c^{-1}\int\dd^{4}x\,\Lag\) and \(\mathcal{D}\phi\) the slicing limit of Equation (106.9) applied to each mode. It is dimensionless, and \(Z[0]\) is the vacuum persistence amplitude \(\braket{0,+\infty}{0,-\infty}\): \(\abs{Z[0]}^{2}\) is the probability that a system prepared in the vacuum is found in the vacuum, which is unity for a stable theory and less than unity when the external conditions can create particles.

Schwinger's insight, which is what makes Equation (106.35) the master object of the subject, is that all correlation functions are encoded in the response of the vacuum amplitude to an external source [Schwinger:1951b].

Theorem 106.25 (Correlators as functional derivatives).

The time-ordered vacuum correlation functions are

\begin{equation}\tag{106.36} \bra{0}T\,\phi(x_{1})\cdots\phi(x_{n})\ket{0} =\frac{1}{Z[0]}\prod_{k=1}^{n} \left(\frac{\hbar c}{\ii}\, \frac{\delta}{\delta J(x_{k})}\right) Z[J]\bigg|_{J=0}\ec \end{equation}

each functional derivative carrying the factor \(\hbar c/\ii\) so that the result has dimension \([\phi]^{n}\). Rests on Equations (106.9) and (106.35).

Derivation. Derives Theorem 106.25. Differentiating Equation (106.35),

\[ \frac{\delta Z}{\delta J(x)} =\int\mathcal{D}\phi\;\frac{\ii}{\hbar c}\,\phi(x)\, \ee^{\ii(S+c^{-1}\!\int J\phi)/\hbar}\ec \]

so multiplying by \(\hbar c/\ii\) inserts one factor of \(\phi(x)\) into the integrand. Repeating gives \(n\) insertions; dividing by \(Z[0]\) normalizes. That the insertions come out time-ordered is a property of the slicing: the fields in Equation (106.9) are \(c\)-numbers attached to time slices, and reassembling them into operators reinstates them in the order of their time arguments, whatever order they were written in. This is the technical advantage of the formalism — the object it computes naturally is the object that appears in the Dyson series Equation (100.25) — and also its limitation, since it does not directly give an unordered product.

Proposition 106.26 (Connected correlators and the linked-cluster theorem).

Define

\begin{equation}\tag{106.37} W[J]:=-\ii\hbar\ln Z[J]\ec \end{equation}

of dimension \(\mathrm{J}\,\mathrm{s}\). Then \(W\) generates the connected correlation functions: in particular

\begin{equation}\tag{106.38} \avg{\phi(x)}_{J}=c\,\frac{\delta W}{\delta J(x)}\ec\qquad \avg{T\phi(x)\phi(y)}_{\mathrm{conn}} =\frac{\hbar c^{2}}{\ii}\, \frac{\delta^{2}W}{\delta J(x)\,\delta J(y)}\ep \end{equation}

Moreover \(Z[J]=\exp(\ii W[J]/\hbar)\) states that the sum of all diagrams is the exponential of the sum of the connected ones. Rests on Equations (106.35) and (106.36).

Derivation. Derives Proposition 106.26. From \(Z=\ee^{\ii W/\hbar}\), \(\delta Z/\delta J=(\ii/\hbar)(\delta W/\delta J)Z\), and combining with Equation (106.36) for \(n=1\) gives \(\avg{\phi}=(\hbar c/\ii)(\ii/\hbar)\delta W/\delta J =c\,\delta W/\delta J\). Differentiating once more,

\[ \frac{\delta^{2}Z}{\delta J_{1}\delta J_{2}} =\left[\frac{\ii}{\hbar}\frac{\delta^{2}W}{\delta J_{1}\delta J_{2}} +\left(\frac{\ii}{\hbar}\right)^{2} \frac{\delta W}{\delta J_{1}}\frac{\delta W}{\delta J_{2}}\right]Z\ec \]

so that, multiplying by \((\hbar c/\ii)^{2}=-\hbar^{2}c^{2}\),

\[ \avg{T\phi_{1}\phi_{2}} =-\ii\hbar c^{2}\frac{\delta^{2}W}{\delta J_{1}\delta J_{2}} +\avg{\phi_{1}}\avg{\phi_{2}}\ec \]

and the first term is by definition the connected part.

For the general statement, argue combinatorially. Let the connected diagrams (with sources attached) be indexed by \(i\) with values \(C_{i}\). A general diagram is a disjoint union of connected pieces; if it contains \(n_{i}\) copies of the connected diagram \(i\), its symmetry factor includes \(1/n_{i}!\) from the permutations of identical components, because the Wick contractions that produce it differ only by relabelling those components. Summing over all multiplicities independently,

\[ Z=Z_{\mathrm{free}}\sum_{\set{n_{i}}}\prod_{i} \frac{C_{i}^{\,n_{i}}}{n_{i}!} =Z_{\mathrm{free}}\prod_{i}\ee^{C_{i}} =Z_{\mathrm{free}}\,\exp\!\left(\sum_{i}C_{i}\right)\ec \]

so \(\ln Z\) is the sum of connected diagrams alone. That is the linked-cluster theorem, and it is why the vacuum diagrams — those with no external source — cancel between numerator and denominator in Equation (106.36): they exponentiate into an overall factor \(Z[0]\).

The linked-cluster theorem has a consequence worth stating separately, because it is the reason the formalism is usable at all. The connected correlators are extensive: \(\ln Z\) is proportional to the spacetime volume, whereas \(Z\) itself is exponentially large in it. Any expansion organized around \(W\) rather than \(Z\) therefore has coefficients with a finite thermodynamic limit. In the Euclidean language of Table 106.1 this is the statement that the free energy is extensive while the partition function is not.

Free fields and the perturbative expansion

The Gaussian theory, solved exactly

Theorem 106.27 (The free generating functional and the Feynman propagator).

For \(g=0\) in Equation (106.32),

\begin{equation}\tag{106.39} Z_{0}[J]=Z_{0}[0]\, \exp\!\left\{-\frac{\ii}{2\hbar c} \int\dd^{4}x\,\dd^{4}y\;J(x)\,G(x-y)\,J(y)\right\}\ec \end{equation}

where \(G\) is the Green function of the Klein–Gordon operator,

\begin{equation}\tag{106.40} -\left(\Box_{x}+\left(\frac{mc}{\hbar}\right)^{2}\right)G(x-y) =\delta^{4}(x-y)\ec\qquad [G]=/\mathrm{m}^{2}\ep \end{equation}

The two-point function is

\begin{equation}\tag{106.41} \Delta_{F}(x-y):=\bra{0}T\phi(x)\phi(y)\ket{0}=\ii\hbar c\,G(x-y)\ec \qquad \widetilde{\Delta}_{F}(p) =\frac{\ii\hbar^{3}c}{p^{2}-m^{2}c^{2}+\ii\epsilon}\ec \end{equation}

of dimension \(\mathrm{J}/\mathrm{m}\) in position space and \(\mathrm{J}\,\mathrm{m}^{3}\) in momentum space. Rests on Equation (106.32), Equation (106.35) and Lemma 106.2.

Derivation. Derives Theorem 106.27. Integrate the kinetic term by parts, discarding the surface term:

\[ S_{0}[\phi]=\frac{1}{c}\int\dd^{4}x\;\frac{1}{2}\left[ \pp_{\mu}\phi\pp^{\mu}\phi-\left(\frac{mc}{\hbar}\right)^{2} \phi^{2}\right] =\frac{1}{2c}\int\dd^{4}x\;\phi\,\mathcal{K}\phi\ec\qquad \mathcal{K}=-\left(\Box+\left(\frac{mc}{\hbar}\right)^{2}\right)\ep \]

The exponent of Equation (106.35) is then \(\frac{\ii}{\hbar c}\bigl[\frac{1}{2}\int\phi\mathcal{K}\phi +\int J\phi\bigr]\), a quadratic form in \(\phi\), and Lemma 106.2 applies once the square is completed. Put \(\phi=\phi_{0}+\chi\) with \(\phi_{0}(x)=-\int\dd^{4}y\;G(x-y)J(y)\), so that \(\mathcal{K}\phi_{0}=-J\) by Equation (106.40). Then

\[ \frac{1}{2}\phi\mathcal{K}\phi+J\phi =\frac{1}{2}\chi\mathcal{K}\chi +\chi(\mathcal{K}\phi_{0}+J) +\frac{1}{2}\phi_{0}\mathcal{K}\phi_{0}+J\phi_{0} =\frac{1}{2}\chi\mathcal{K}\chi -\frac{1}{2}\int\!\!\int J\,G\,J\ec \]

the cross term vanishing by construction. The shift \(\phi\mapsto\chi\) is a translation of the integration variable and leaves \(\mathcal{D}\phi\) invariant — this is a property of the sliced definition, each \(\dd\phi_{k}\) being a Lebesgue measure on \(\R\) — so the \(\chi\) integral is the \(J\)-independent constant \(Z_{0}[0]\) and what remains is Equation (106.39).

By Equation (106.37), \(W_{0}[J]=-\frac{1}{2c}\int\!\!\int JGJ\), and Equation (106.38) gives \(\avg{T\phi_{1}\phi_{2}} =(\hbar c^{2}/\ii)(-G/c)=\ii\hbar cG\).

In momentum space, with the convention Equation (100.14), \(\pp_{\mu}\mapsto-\ii p_{\mu}/\hbar\) and \(\Box\mapsto-p^{2}/\hbar^{2}\), so \(\mathcal{K}\mapsto(p^{2}-m^{2}c^{2})/\hbar^{2}\) and \(\widetilde{G}(p)=\hbar^{2}/(p^{2}-m^{2}c^{2})\). The pole on the real axis is resolved by the Feynman prescription \(m^{2}c^{2}\to m^{2}c^{2}-\ii\epsilon\), which is the choice singled out by Proposition 106.14. Multiplying by \(\ii\hbar c\) gives Equation (106.41). The result has exactly the structure of the photon propagator Equation (100.17), the factor \(\hbar^{3}c\) being common to both.

Corollary 106.28 (Wick's theorem).

In the free theory, correlators of an odd number of fields vanish, and

\begin{equation}\tag{106.42} \bra{0}T\phi(x_{1})\cdots\phi(x_{2n})\ket{0} =\sum_{\text{pairings}}\;\prod_{\text{pairs }(i,j)} \Delta_{F}(x_{i}-x_{j})\ec \end{equation}

the sum running over all \((2n-1)!!\) ways of partitioning the arguments into pairs. Rests on Equations (106.36) and (106.39).

Derivation. Derives Corollary 106.28. Apply Equation (106.36) to Equation (106.39). Each derivative \(\hbar c\,\delta/(\ii\delta J)\) acting on the exponential brings down one factor \(-(\ii/\hbar c)\int G J\) times \(\hbar c/\ii\), i.e. \(-\int GJ\), which is set to zero at \(J=0\) unless a later derivative acts on it. A non-vanishing term therefore requires the derivatives to be used in pairs: one to bring down the bilinear, one to remove the surviving \(J\). The pair \((i,j)\) contributes \((\hbar c/\ii)^{2}\times(-\ii/\hbar c)G=\ii\hbar cG=\Delta_{F}\), and the number of ways of pairing \(2n\) objects is \((2n-1)!!\). An odd number of derivatives always leaves one \(J\) unpaired.

Corollary 106.28 is the same statement as the operator Wick theorem [Wick:1950] invoked in Proposition 100.14. In the operator formalism it is a statement about normal-ordering products of creation and annihilation operators, and that chapter quotes it from Wick's paper rather than proving it; here it is a property of the Gaussian integral, derived above, and no operators appear. That is the second road promised in the introduction.

Perturbation theory and the Feynman rules

Proposition 106.29 (Feynman rules from the functional integral).

Write \(S=S_{0}+S_{\mathrm{int}}\) with \(S_{\mathrm{int}}=-\frac{1}{c}\int\dd^{4}x\,\frac{g}{4!}\phi^{4}\). Then

\begin{equation}\tag{106.43} Z[J]=\exp\!\left\{\frac{\ii}{\hbar} S_{\mathrm{int}}\!\left[\frac{\hbar c}{\ii} \frac{\delta}{\delta J}\right]\right\}Z_{0}[J]\ec \end{equation}

and expanding the outer exponential generates, order by order in \(\hat g=g\hbar c\), the diagrammatic rules: a line for each propagator Equation (106.41), a four-point vertex carrying

\begin{equation}\tag{106.44} -\frac{\ii g}{\hbar c}\int\dd^{4}x \quad\leftrightarrow\quad -\frac{\ii g}{\hbar c}\,(2\pi\hbar)^{4}\, \delta^{4}\!\left(\textstyle\sum_{k}p_{k}\right) \quad\text{in momentum space,} \end{equation}

one integration \(\int\dd^{4}\ell/(2\pi\hbar)^{4}\) per loop, and the symmetry factor of the diagram. Rests on Equation (106.35), Theorem 106.25 and Corollary 106.28.

Derivation. Derives Proposition 106.29. Every occurrence of \(\phi(x)\) in \(S_{\mathrm{int}}\) may be replaced by \((\hbar c/\ii)\delta/\delta J(x)\) acting on the integrand of Equation (106.35), by the computation in Theorem 106.25; the resulting operator may then be pulled outside the \(\phi\) integral, leaving \(Z_{0}[J]\). That is Equation (106.43), which is exact. Expanding \(\exp(\ii S_{\mathrm{int}}/\hbar)\) in powers and applying Corollary 106.28 to each term produces sums of products of propagators with four lines meeting at each vertex position, each vertex carrying \((\ii/\hbar)\times(-1/c)\times(g/4!)\) and an integration \(\int\dd^{4}x\); the \(4!\) is cancelled by the number of ways of assigning the four lines to the four fields, except when a symmetry of the diagram leaves some of those assignments equivalent, which is the origin of the symmetry factor. Fourier transforming each propagator by Equation (100.14) turns the position integrals into the momentum-conserving delta functions of Equation (106.44), leaving one free integration per independent loop.

Theorem 106.30 (The loop expansion is the expansion in $\hbar$).

A connected diagram with \(V\) vertices, \(I\) internal lines and \(L\) independent loops carries the factor

\begin{equation}\tag{106.45} \hbar^{I-V}=\hbar^{L-1}\ec\qquad L=I-V+1\ec \end{equation}

counting the vertices and the internal lines alone. A tree graph has \(L=0\) and therefore carries \(\hbar^{-1}\); relative to it an \(L\)-loop graph is suppressed by \(\hbar^{L}\). The two statements are the same one and must not be added: \(\hbar^{L-1}\) is the absolute power, \(\hbar^{L}\) the power relative to the classical contribution. The perturbative expansion in the number of loops is therefore an expansion in powers of \(\hbar\), and in nothing else. Rests on Proposition 106.29 and Equation (106.40).

Derivation. Derives Theorem 106.30. By Proposition 106.29 each vertex carries an explicit \(\hbar^{-1}\) (from \(\ii S_{\mathrm{int}}/\hbar\)) and each internal line one factor of \(\hbar\) (the propagator \(\ii\hbar cG\) is linear in \(\hbar\), \(G\) being \(\hbar\)-independent by Equation (106.40) once \(mc/\hbar\) is regarded as the fixed inverse Compton wavelength). A diagram therefore carries \(\hbar^{I-V}\). The second relation of Equation (106.45) is Euler's relation for a connected graph: each internal line carries one momentum, the \(V\) vertices impose \(V\) delta functions of which one is the overall momentum conservation, leaving \(I-(V-1)\) undetermined loop momenta. Hence \(\hbar^{I-V}=\hbar^{L-1}\), and a tree, having \(L=0\), sets the reference power \(\hbar^{-1}\) against which the loops are counted.

Theorem 106.30 is the reason this chapter refuses to drop \(\hbar\) from its formulae. In the standard presentation the loop counting is a combinatorial fact about graphs; here it is a physical statement — a one-loop correction is the first quantum correction, suppressed relative to the classical answer by one power of the quantity whose smallness defines the classical limit, and the tree diagrams alone reproduce classical field theory exactly. This will be made precise in Section 106.3.4, where the tree approximation to \(\Gamma\) is shown to be the classical action.

The series does not converge

Proposition 106.31 (Dyson's argument).

The perturbation series of quantum electrodynamics in powers of \(e^{2}\) has zero radius of convergence [Dyson:1952]. Rests on Proposition 106.29.

Derivation. Derives Proposition 106.31. Suppose the series converged in some disc \(\abs{e^{2}}<r\). Then it would converge for real negative \(e^{2}\), that is for an imaginary charge, and by analyticity the sum would be the physical quantity of that theory. But with \(e^{2}<0\) like charges attract: the vacuum is unstable against the spontaneous creation of a large number of electron–positron pairs, which segregate into two clouds of like charge whose negative potential energy exceeds their rest energy once enough pairs are made. No lowest-energy state exists, so the theory has no ground state, no \(S\)-matrix, and no analytic continuation of the physical quantities to compare with. Hence the radius of convergence is zero.

Dyson's argument is physical, not a theorem: it establishes that the theory at \(e^{2}<0\) is pathological, and infers from this that the function of \(e^{2}\) cannot be analytic at the origin.

The concrete form of the divergence is visible in a single ordinary integral, which is the whole of the phenomenon with none of the field theory.

Example 106.32 (The zero-dimensional model).

Let \(Z(\hat g)=\int_{-\infty}^{\infty}\dd x\; \ee^{-x^{2}/2-\hat gx^{4}/4!}\), the \(0+0\)-dimensional version of Equation (106.32) after Wick rotation. Expanding the quartic term and using \(\avg{x^{4n}}=(4n-1)!!\) from Lemma 106.2,

\begin{equation}\tag{106.46} \frac{Z(\hat g)}{Z(0)} =\sum_{n=0}^{\infty}\frac{(-\hat g)^{n}}{n!\,(4!)^{n}}\,(4n-1)!! =1-\frac{\hat g}{8}+\frac{35\hat g^{2}}{384} -\frac{385\hat g^{3}}{3072}+\cdots\ep \end{equation}

The coefficients grow factorially. Writing \((4n-1)!!=(4n)!/\bigl(2^{2n}(2n)!\bigr)\) and applying Stirling's formula to each factorial gives

\begin{equation}\tag{106.47} \abs{a_{n}}=\frac{(4n-1)!!}{n!\,(4!)^{n}} \sim\frac{n!}{\sqrt{2}\,\pi n}\left(\frac{2}{3}\right)^{n} \qquad(n\to\infty)\ec \end{equation}

so the series diverges for every \(\hat g\neq0\) although the integral it came from is perfectly finite and analytic in \(\hat g\) on the cut plane. The best one can do is truncate at the smallest term, \(n\approx3/(2\hat g)\), with an error of order \(\ee^{-3/(2\hat g)}\): an asymptotic series.

The same estimate applied to quantum electrodynamics with \(\alpha=1/137.036\) puts the optimal truncation near order \(1/\alpha\approx137\) and the irreducible error near \(\ee^{-1/\alpha}\approx10^{-60}\) — some fifty orders of magnitude below the precision of the most accurate measurement in physics (Section 100.8.2). The non-convergence of the series is therefore a statement about the far tail of a computation nobody will ever do, not a defect of the predictions. What it does mean is that perturbation theory alone cannot define the theory: quantities of the size \(\ee^{-1/\alpha}\), or \(\ee^{-1/\hat g}\), are invisible to it, and those are exactly the instanton effects of Section 106.6.1 and the confining physics of Section 106.5. The Euclidean path integral, defined non-perturbatively on a lattice, is at present the only construction that reaches them.

Fermions and Grassmann integration

Why the classical fields must anticommute

The construction of Theorem 106.3 inserted a resolution of the identity built from eigenstates of the coordinate operator. For a field theory the corresponding step inserts eigenstates of the field operator, \(\hat\phi(\vect{x})\ket{\phi} =\phi(\vect{x})\ket{\phi}\) at fixed time, and the eigenvalues \(\phi(\vect{x})\) are the integration variables. For a Dirac field this is impossible with ordinary numbers.

Proposition 106.33 (Field eigenvalues inherit the statistics).

Let \(\hat\psi_{\alpha}(\vect{x})\) satisfy the equal-time anticommutation relations \(\acomm{\hat\psi_{\alpha}(\vect{x})}{\hat\psi_{\beta}(\vect{y})}=0\) required by the spin–statistics connection of Identical Particles, and let \(\ket{\psi}\) be a simultaneous eigenstate, \(\hat\psi_{\alpha}(\vect{x})\ket{\psi} =\psi_{\alpha}(\vect{x})\ket{\psi}\). Then the eigenvalues satisfy

\begin{equation}\tag{106.48} \psi_{\alpha}(\vect{x})\psi_{\beta}(\vect{y}) =-\psi_{\beta}(\vect{y})\psi_{\alpha}(\vect{x})\ec \end{equation}

and in particular \(\psi_{\alpha}(\vect{x})^{2}=0\): they are Grassmann numbers, generators of an exterior algebra, not elements of \(\R\) or \(\C\). Rests on Theorem 99.46.

Derivation. Derives Proposition 106.33. Apply the two operators in both orders to the eigenstate: \(\hat\psi_{\alpha}(\vect{x})\hat\psi_{\beta}(\vect{y})\ket{\psi} =\psi_{\beta}(\vect{y})\psi_{\alpha}(\vect{x})\ket{\psi}\) — reading the operators from right to left — while the anticommutator gives \(\hat\psi_{\alpha}\hat\psi_{\beta} =-\hat\psi_{\beta}\hat\psi_{\alpha}\), whose eigenvalue is \(-\psi_{\alpha}(\vect{x})\psi_{\beta}(\vect{y})\). Equating the two gives Equation (106.48). Setting \(\alpha=\beta\), \(\vect{x}=\vect{y}\) gives \(\psi^{2}=-\psi^{2}=0\).

Nothing was assumed here beyond the anticommutation relations themselves. The Grassmann variables of the fermionic path integral are not a formal device chosen for convenience: they are what the eigenvalues of an anticommuting operator are. What remains is to say what it means to integrate over them, and Berezin's answer [Berezin:1966] is not a postulate either — it is forced by two requirements.

Berezin's rules, derived

Definition 106.34 (Grassmann algebra).

Let \(\theta_{1},\dots,\theta_{n}\) generate an associative algebra over \(\C\) with \(\theta_{i}\theta_{j}=-\theta_{j}\theta_{i}\), hence \(\theta_{i}^{2}=0\). Every element is a finite sum \(f=\sum_{S}c_{S}\prod_{i\in S}\theta_{i}\) over subsets \(S\subseteq\set{1,\dots,n}\); there are \(2^{n}\) terms and the algebra is finite-dimensional. In one generator the general element is simply \(f(\theta)=a+b\theta\).

Theorem 106.35 (The Berezin integral).

Suppose a linear functional \(f\mapsto\int\dd\theta\,f(\theta)\) on the one-generator algebra is required to be

  1. linear, and

  2. translation invariant: \(\int\dd\theta\,f(\theta+\eta)=\int\dd\theta\,f(\theta)\) for every Grassmann \(\eta\).

Then it is determined up to normalization, and with the normalization \(\int\dd\theta\,\theta=1\),

\begin{equation}\tag{106.49} \int\dd\theta\;1=0\ec\qquad\int\dd\theta\;\theta=1\ec \qquad\text{hence}\qquad \int\dd\theta\,f(\theta)=\pdv{f}{\theta}\ep \end{equation}

Integration over a Grassmann variable is differentiation with respect to it. Rests on Definition 106.34.

Derivation. Derives Theorem 106.35. By linearity the functional is fixed by the two numbers \(I_{0}=\int\dd\theta\,1\) and \(I_{1}=\int\dd\theta\,\theta\). Apply translation invariance to \(f=a+b\theta\):

\[ \int\dd\theta\,\bigl(a+b\theta+b\eta\bigr) =aI_{0}+bI_{1}+b\eta I_{0} \quad\text{must equal}\quad aI_{0}+bI_{1}\ec \]

for every \(b\) and every \(\eta\), which forces \(I_{0}=0\). Normalizing \(I_{1}=1\) then gives Equation (106.49), and since \(\pp(a+b\theta)/\pp\theta=b=aI_{0}+bI_{1}\), integration and differentiation coincide.

Translation invariance is not a technical convenience: it is exactly the property used in Theorem 106.27 to complete the square, and without it there would be no fermionic Gaussian integral and hence no fermionic path integral. It is also the property whose failure under a chiral rotation produces the anomaly of Section 106.7.3 — there the transformation is linear but not a translation, and its Jacobian is the whole story.

Corollary 106.36 (Change of variables inverts).

For an invertible linear change of Grassmann variables \(\theta'_{i}=A_{ij}\theta_{j}\),

\begin{equation}\tag{106.50} \dd^{n}\theta'=(\det A)^{-1}\dd^{n}\theta\ec \end{equation}

the determinant appearing with the opposite power to the ordinary case. Rests on Definition 106.34 and Theorem 106.35.

Derivation. Derives Corollary 106.36. Only the top monomial survives an \(n\)-fold integration, so the rule is fixed by \(\int\dd^{n}\theta'\,\theta'_{1}\cdots\theta'_{n}=1\). Expanding, \(\theta'_{1}\cdots\theta'_{n} =A_{1j_{1}}\cdots A_{nj_{n}}\theta_{j_{1}}\cdots\theta_{j_{n}} =(\det A)\,\theta_{1}\cdots\theta_{n}\), since reordering the \(\theta\)'s into canonical order produces the sign of the permutation and the sum over permutations with signs is the Leibniz formula for the determinant. To keep the normalization, \(\dd^{n}\theta'\) must carry \((\det A)^{-1}\).

Theorem 106.37 (The Grassmann Gaussian is a determinant).

Let \(\theta_{i},\bar\theta_{i}\), \(i=1,\dots,n\), be independent Grassmann generators, with the convention \(\int\dd\bar\theta_{i}\,\dd\theta_{i}\;\theta_{i}\bar\theta_{i}=1\). Then for any \(n\times n\) matrix \(A\),

\begin{equation}\tag{106.51} \int\prod_{i=1}^{n}\dd\bar\theta_{i}\,\dd\theta_{i}\; \exp\!\left(-\bar\theta_{i}A_{ij}\theta_{j}\right)=\det A\ec \end{equation}

to be compared with the bosonic \(\int\dd^{n}x\,\ee^{-x\transpose Ax/2} =(2\pi)^{n/2}(\det A)^{-1/2}\) of Equation (106.7). With sources,

\begin{equation}\tag{106.52} \int\prod_{i}\dd\bar\theta_{i}\,\dd\theta_{i}\; \exp\!\left(-\bar\theta A\theta +\bar\eta\theta+\bar\theta\eta\right) =\det A\;\exp\!\left(\bar\eta A^{-1}\eta\right)\ep \end{equation}

Rests on Equation (106.7) and Theorem 106.35.

Derivation. Derives Theorem 106.37. Expand the exponential. Because each generator squares to zero, the only term that survives the \(2n\)-fold integration is the one containing each of \(\theta_{1},\dots,\theta_{n}\) and each of \(\bar\theta_{1},\dots,\bar\theta_{n}\) exactly once, namely

\[ \frac{(-1)^{n}}{n!}\left(\bar\theta_{i}A_{ij}\theta_{j}\right)^{n} =(-1)^{n}\sum_{\sigma\in S_{n}} A_{1\sigma(1)}\cdots A_{n\sigma(n)}\, \bar\theta_{1}\theta_{\sigma(1)}\cdots \bar\theta_{n}\theta_{\sigma(n)}\ec \]

the \(n!\) cancelling against the \(n!\) orderings of the identical factors. Reordering each \(\bar\theta_{k}\theta_{\sigma(k)}\) pair into the canonical order \(\theta_{k}\bar\theta_{k}\) produces \(\sgn(\sigma)\) together with \((-1)^{n}\), and the surviving sum is \(\sum_{\sigma}\sgn(\sigma)\prod_{k}A_{k\sigma(k)}=\det A\). The sourced version follows by the shift \(\theta\to\theta+A^{-1}\eta\), \(\bar\theta\to\bar\theta+\bar\eta A^{-1}\), which is legitimate by Theorem 106.35(ii) and has unit Jacobian, being a translation.

Equation (106.51) and Equation (106.7) differ in exactly one place — the sign of the exponent of the determinant — and every distinctive feature of fermions in the path integral follows from it.

The fermion determinant

Proposition 106.38 (Integrating out the fermions).

For a Dirac field coupled to a gauge field \(A_{\mu}\) with the Lagrangian density \(\Lag_{\psi}=\bar\psi\bigl(\ii\hbar c\,\gamma^{\mu}D_{\mu} -mc^{2}\bigr)\psi\) of Equation (100.5),

\begin{equation}\tag{106.53} \int\mathcal{D}\bar\psi\,\mathcal{D}\psi\; \exp\!\left\{\frac{\ii}{\hbar c}\int\dd^{4}x\; \bar\psi\bigl(\ii\hbar c\gamma^{\mu}D_{\mu}-mc^{2}\bigr)\psi\right\} \;\propto\;\det\!\left(\ii\hbar\gamma^{\mu}D_{\mu}-mc\right)\ec \end{equation}

the constant of proportionality being field-independent. Each entry of the operator carries the dimension of a momentum, so the determinant must be normalized against its free value to be a number; only ratios appear in physical quantities. Rests on Equations (100.5) and (106.51).

Derivation. Derives Proposition 106.38. The exponent is \(\ii S/\hbar\) with \(S=c^{-1}\int\dd^{4}x\,\Lag_{\psi}\), which is a Grassmann bilinear \(-\bar\psi\mathcal{A}\psi\) with \(\mathcal{A}=-(\ii/\hbar c)(\ii\hbar c\gamma^{\mu}D_{\mu}-mc^{2}) =-(\ii/\hbar)(\ii\hbar\gamma^{\mu}D_{\mu}-mc)\) after extracting the common factor \(c\). Apply Equation (106.51) mode by mode: the answer is \(\det\mathcal{A}\), and the field-independent factors \(-\ii/\hbar\) per mode contribute an infinite but \(A\)-independent constant, absorbed into the normalization.

Corollary 106.39 (The minus sign of a closed fermion loop).

Write \(\det\mathcal{A}=\exp\tr\ln\mathcal{A}\) and expand in powers of the gauge field. The resulting contribution to the effective action is \(-\ii\hbar\,\tr\ln\mathcal{A}\), whereas a boson of the same coupling gives \(+\frac{\ii\hbar}{2}\tr\ln\) by Equation (106.7). Hence every closed fermion loop carries a factor \((-1)\) relative to the corresponding boson loop, together with a Dirac trace — which is rule five of Proposition 100.14, here derived rather than imposed. Rests on Proposition 106.38, Equation (106.7) and Equation (106.37).

Derivation. Derives Corollary 106.39. By Proposition 106.38, integrating out the fermions contributes \(\det\mathcal{A}\) to \(Z\), hence \(\tr\ln\mathcal{A}\) to \(\ln Z\) and \(-\ii\hbar\tr\ln\mathcal{A}\) to \(W\) by Equation (106.37). Expanding \(\tr\ln(\mathcal{A}_{0}+\delta\mathcal{A}) =\tr\ln\mathcal{A}_{0} +\sum_{k\geq1}\frac{(-1)^{k+1}}{k} \tr\bigl[(\mathcal{A}_{0}^{-1}\delta\mathcal{A})^{k}\bigr]\) produces, at order \(k\), a single trace around a closed chain of \(k\) propagators and \(k\) vertices: a one-loop diagram with \(k\) external legs. For a real boson the same manipulation applied to \((\det)^{-1/2}\) gives \(+\frac{1}{2}\tr\ln\), so the fermion loop differs by the factor \(-2\times\frac{1}{2}=-1\) per loop and by the appearance of a Dirac trace, the boson having no spinor index. The factor \(1/2\) absent in the fermionic case is the symmetry factor of a loop of a real field, which a complex fermion loop does not have because its two orientations are distinct.

Remark 106.40 (Quenching, and why it was abandoned).

The determinant Equation (106.53) is a non-local functional of the gauge field: changing a link far away changes it. Early lattice computations therefore replaced it by a constant — the quenched approximation, equivalent to omitting all internal quark loops — because the cost of evaluating it repeatedly exceeded the machines of the time by orders of magnitude. The approximation is not controlled: it has no small parameter, its error was estimated empirically at ten to twenty percent for typical hadronic quantities, and in some channels it is qualitatively wrong, since a quenched theory has no two-particle thresholds from sea quarks and hence a different analytic structure. It is no longer used for results that are quoted with error bars; the computation of Section 106.5.2 includes the determinant of two degenerate light flavours and one strange flavour [Duerr:2008]. Squaring the determinant of the degenerate pair is what keeps the Euclidean weight positive, which is the practical form of Remark 106.15.

The effective action

Definition 106.41 (Effective action).

Let \(\phi_{c}(x):=c\,\delta W/\delta J(x)=\avg{\phi(x)}_{J}\) be the field expectation value in the presence of the source. The effective action is the Legendre transform

\begin{equation}\tag{106.54} \Gamma[\phi_{c}]:=W[J]-\frac{1}{c}\int\dd^{4}x\;J(x)\phi_{c}(x)\ec \end{equation}

with \(J\) eliminated in favour of \(\phi_{c}\). It has the dimension of an action, \(\mathrm{J}\,\mathrm{s}\).

The construction is Jona-Lasinio's [JonaLasinio:1964], who introduced it precisely to have a functional whose stationary points locate the vacuum: a symmetry of the action that is not a symmetry of the minimising \(\phi_{c}\) is a symmetry broken by the state and not by the dynamics, which is the situation Electroweak Unification and the Higgs Boson turns into masses.

Proposition 106.42 (The quantum equation of motion).
\begin{equation}\tag{106.55} \frac{\delta\Gamma}{\delta\phi_{c}(x)}=-\frac{1}{c}J(x)\ec \end{equation}

so that in the absence of a source the physical field configurations are the stationary points of \(\Gamma\), exactly as the classical ones are the stationary points of \(S\). Rests on Equation (106.54).

Derivation. Derives Proposition 106.42. Vary Equation (106.54) with respect to \(\phi_{c}\), allowing \(J\) to vary with it:

\[ \delta\Gamma=\int\dd^{4}x\left[ \frac{\delta W}{\delta J}\,\delta J -\frac{1}{c}\phi_{c}\,\delta J -\frac{1}{c}J\,\delta\phi_{c}\right] =-\frac{1}{c}\int\dd^{4}x\;J\,\delta\phi_{c}\ec \]

the first two terms cancelling by the definition of \(\phi_{c}\).

Theorem 106.43 (The effective action generates 1PI vertices).

Write \(W_{x_{1}\cdots x_{n}}\) and \(\Gamma_{x_{1}\cdots x_{n}}\) for the functional derivatives of \(W\) with respect to \(J\) and of \(\Gamma\) with respect to \(\phi_{c}\). Then

\begin{equation}\tag{106.56} \int\dd^{4}z\;W_{xz}\,\Gamma_{zy}=-\frac{1}{c^{2}}\, \delta^{4}(x-y)\ec \end{equation}

so \(\Gamma_{xy}\) is, up to the factor \(-c^{-2}\), the inverse of the full connected propagator; and

\begin{equation}\tag{106.57} W_{xyz}=c^{3}\int\dd^{4}a\,\dd^{4}b\,\dd^{4}d\; W_{xa}W_{yb}W_{zd}\,\Gamma_{abd}\ec \end{equation}

so the connected three-point function is the third derivative of \(\Gamma\) dressed with a full propagator on each leg. In general the \(n\)-th derivative of \(\Gamma\) is the amputated one-particle-irreducible \(n\)-point vertex. Rests on Definition 106.41 and Equation (106.55).

Derivation. Derives Theorem 106.43. By definition \(\phi_{c}=c\,\delta W/\delta J\) and, by Equation (106.55), \(J=-c\,\delta\Gamma/\delta\phi_{c}\). The two are inverse maps, so by the chain rule

\[ \delta^{4}(x-y)=\int\dd^{4}z\; \frac{\delta\phi_{c}(x)}{\delta J(z)}\, \frac{\delta J(z)}{\delta\phi_{c}(y)} =\int\dd^{4}z\;\bigl(cW_{xz}\bigr)\bigl(-c\Gamma_{zy}\bigr)\ec \]

which is Equation (106.56). Differentiate it with respect to \(J(w)\), using \(\delta/\delta J(w)=c\int\dd^{4}u\;W_{uw}\,\delta/\delta\phi_{c}(u)\):

\[ \int\dd^{4}z\;W_{xzw}\Gamma_{zy} =-c\int\dd^{4}z\,\dd^{4}u\;W_{xz}W_{uw}\Gamma_{zyu}\ep \]

Contract with \(-c^{2}W_{yb}\) and integrate over \(y\); the left-hand side collapses by Equation (106.56) to \(W_{xbw}\), and the right-hand side becomes \(c^{3}\int W_{xz}W_{by}W_{wu}\Gamma_{zyu}\), which is Equation (106.57) after renaming.

A connected three-point function with a full propagator removed from each external leg cannot be cut into two pieces by severing one internal line, since any such cut would isolate one external leg and the piece so isolated is precisely the propagator already amputated. Hence \(\Gamma_{abd}\) is one-particle irreducible. Differentiating Equation (106.57) once more produces

\[ W_{4}=c^{4}\Bigl[WWWW\,\Gamma_{4} +\bigl(\text{three channels}\bigr)WWWWW\,\Gamma_{3}\Gamma_{3}\Bigr]\ec \]

in which the second group of terms is exactly the set of one-particle-reducible tree graphs built from two cubic vertices joined by one full propagator. The pattern repeats: at each order the new derivative of \(\Gamma\) supplies the irreducible part and the terms generated by differentiating the propagators supply all the reducible trees. Induction on \(n\) gives the general statement.

Theorem 106.43 is what makes \(\Gamma\) the practical object. Since it generates the irreducible vertices and the exact propagator, knowing \(\Gamma\) is knowing the theory: every connected amplitude is assembled from it by tree graphs alone.

The loop expansion of $\Gamma$

Theorem 106.44 (One loop).
\begin{equation}\tag{106.58} \Gamma[\phi_{c}]=S[\phi_{c}] +\frac{\ii\hbar}{2}\,\tr\ln \left(\frac{\delta^{2}S}{\delta\phi\,\delta\phi} \bigg|_{\phi=\phi_{c}}\right)+O(\hbar^{2})\ec \end{equation}

so the classical action is the tree approximation to \(\Gamma\) and the first quantum correction is one power of \(\hbar\) smaller, in accord with Theorem 106.30. Rests on Equations (106.7), (106.54) and (106.55).

Derivation. Derives Theorem 106.44. Combining Equations (106.35), (106.37) and (106.54),

\[ \exp\!\left(\frac{\ii}{\hbar}\Gamma[\phi_{c}]\right) =\int\mathcal{D}\phi\;\exp\!\left\{\frac{\ii}{\hbar}\left( S[\phi]-\frac{1}{c}\int\dd^{4}x\; \frac{\delta\Gamma}{\delta\phi_{c}}\, \bigl(\phi-\phi_{c}\bigr)\right)\right\}\ec \]

using Equation (106.55) to eliminate \(J\). Substitute \(\phi=\phi_{c}+\sqrt{\hbar}\,\chi\) and expand:

\[ S[\phi_{c}+\sqrt{\hbar}\chi] =S[\phi_{c}]+\sqrt{\hbar}\int\frac{\delta S}{\delta\phi}\chi +\frac{\hbar}{2}\int\!\!\int\chi\, \frac{\delta^{2}S}{\delta\phi\delta\phi}\,\chi +O(\hbar^{3/2})\ep \]

At order \(\sqrt{\hbar}\) the term linear in \(\chi\) cancels against the source term, because \(\delta\Gamma/\delta\phi_{c} =\delta S/\delta\phi_{c}+O(\hbar)\); what remains is a Gaussian integral over \(\chi\), evaluated by Equation (106.7) in the form \((\det)^{-1/2}=\exp(-\frac{1}{2}\tr\ln)\). Multiplying by \(\hbar/\ii\) to extract \(\Gamma\) turns \(-\frac{1}{2}\tr\ln\) into \(+\frac{\ii\hbar}{2}\tr\ln\), which is Equation (106.58). The substitution \(\phi=\phi_{c}+\sqrt{\hbar}\chi\) is what makes the \(\hbar\) counting automatic: each additional vertex in the expansion brings \(\hbar^{1/2}\) per field beyond the second, and the Gaussian contractions pair them up.

The effective potential

For a constant \(\phi_{c}\) the effective action reduces to a function of one variable. Writing \(V_{4}=\int\dd^{4}x\),

\begin{equation}\tag{106.59} \Gamma[\phi_{c}]=-\frac{V_{4}}{c}\,U(\phi_{c})\ec \end{equation}

which defines the effective potential \(U\), of dimension \(\mathrm{J}/\mathrm{m}^{3}\); at tree level \(U\) is the classical potential \(U_{0}=\frac{1}{2}(mc/\hbar)^{2}\phi_{c}^{2} +\frac{g}{4!}\phi_{c}^{4}\).

Proposition 106.45 (One-loop effective potential).

Define the field-dependent mass \(M(\phi_{c})\) by \(\bigl(M c/\hbar\bigr)^{2}=(mc/\hbar)^{2}+\frac{g}{2}\phi_{c}^{2}\). Then, in the minimal-subtraction scheme of Definition 100.54 at renormalization energy \(\mathcal{E}\),

\begin{equation}\tag{106.60} U_{1}(\phi_{c}) =\frac{\bigl(Mc^{2}\bigr)^{4}}{64\pi^{2}(\hbar c)^{3}} \left[\ln\frac{\bigl(Mc^{2}\bigr)^{2}}{\mathcal{E}^{2}} -\frac{3}{2}\right]\ec \end{equation}

of dimension \(\mathrm{J}/\mathrm{m}^{3}\) as required, since \((Mc^{2})^{4}/(\hbar c)^{3}\) is an energy divided by a volume. Rests on Equation (106.58), Equation (106.59) and Lemma 100.28.

Derivation. Derives Proposition 106.45. By Equation (106.58) the one-loop term is \(\frac{\ii\hbar}{2}\tr\ln\mathcal{K}_{\phi}\) with \(\mathcal{K}_{\phi}=-(\Box+M^{2}c^{2}/\hbar^{2})\) read off from the second variation of Equation (106.32). For a translation-invariant kernel the trace is \(V_{4}\int\dd^{4}p\,(2\pi\hbar)^{-4}\) of the momentum-space symbol, so

\[ \Gamma_{1}=\frac{\ii\hbar V_{4}}{2}\int\frac{\dd^{4}p}{(2\pi\hbar)^{4}} \,\ln\frac{p^{2}-M^{2}c^{2}+\ii\epsilon}{\hbar^{2}}\ec \]

and Equation (106.59) gives \(U_{1}=-c\Gamma_{1}/V_{4}\). Rotate by Proposition 106.14, \(\dd^{4}p=\ii\,\dd^{4}p_{E}\) and \(p^{2}=-p_{E}^{2}\):

\[ U_{1}=\frac{\hbar c}{2}\int\frac{\dd^{4}p_{E}}{(2\pi\hbar)^{4}}\; \ln\frac{p_{E}^{2}+M^{2}c^{2}}{\hbar^{2}}\ec \]

up to a field-independent constant. The integral is quartically divergent and is evaluated in \(d=4-2\varepsilon\) dimensions by Lemma 100.28. Writing \(\Lambda^{2}=M^{2}c^{2}\) for the mass-squared parameter, the standard \(d\)-dimensional result is

\[ \int\frac{\dd^{d}p_{E}}{(2\pi)^{d}}\, \ln\bigl(p_{E}^{2}+\Lambda^{2}\bigr) =-\frac{1}{(4\pi)^{d/2}}\, \frac{\bigl(\Lambda^{2}\bigr)^{d/2}}{\Gamma(1)}\, \Gamma\!\left(-\frac{d}{2}\right)\ec \]

\(\Gamma\) here denoting the Euler gamma function and not the effective action. Expanding its pole at \(d=4\) by Lemma 100.29 and discarding that pole together with the \(-\gamma_{\mathrm{E}} +\ln4\pi\) which accompanies it — the definition of minimal subtraction, Equation (100.75) — leaves \(\bigl(\Lambda^{2}\bigr)^{2} \bigl[\ln(\Lambda^{2}/\mu^{2})-3/2\bigr]/(32\pi^{2})\). Restoring the \(\hbar^{-4}\) of the measure and the prefactor \(\hbar c/2\) gives Equation (106.60) with \(\mathcal{E}=\mu c\).

The physical content of Equation (106.60) is Coleman and Weinberg's [Coleman:1973]: a theory whose classical potential has its minimum at the origin can have its effective potential minimized away from it, so that a symmetry unbroken classically is broken by radiative corrections alone. The mechanism is real and is used in Electroweak Unification and the Higgs Boson as one of the two possible origins of a scalar expectation value; whether it is the origin of the electroweak one is settled by measurement, and the measured Higgs and top masses say that it is not. The one-loop potential also generates the vacuum-stability question of that chapter, since Equation (106.60) turns negative at large field when the fermion loops dominate.

Remark 106.46 (The exact effective potential is convex).

The one-loop potential can have two minima; the exact one cannot have the shape between them that Equation (106.60) suggests. In the Euclidean formulation \(W[J]\) is a convex functional of \(J\), because

\[ \frac{\delta^{2}W_{E}}{\delta J(x)\delta J(y)} \propto\avg{\phi(x)\phi(y)}-\avg{\phi(x)}\avg{\phi(y)} \]

is a covariance matrix, hence positive semidefinite — and it is positive by reflection positivity Equation (106.30). The Legendre transform of a convex function is convex, so \(\Gamma\), and with it \(U\), is convex exactly. The apparent non-convex region between two degenerate minima is therefore replaced in the exact theory by the straight segment joining them — the Maxwell equal-area construction, the same object that replaces the unphysical loop of the van der Waals isotherm in Phase Transitions and Critical Phenomena. Physically, a constrained expectation value \(\phi_{c}\) intermediate between the two minima is realized not by a homogeneous state of that field value but by a mixture — domains of one vacuum and of the other, in the proportion that gives the average. The loop expansion does not see this because it expands about a single homogeneous configuration. The non-convex one-loop potential is nonetheless the right object for computing the physics of a single phase, which is what it is used for in Electroweak Unification and the Higgs Boson; the convexity statement is about a question the loop expansion was not asked.

The LSZ reduction formula

The path integral computes time-ordered correlation functions of fields. Experiments measure the probabilities of scattering processes. Lehmann, Symanzik and Zimmermann supplied the bridge [Lehmann:1955], and its content is that the \(S\)-matrix lives in the poles of the correlators.

Proposition 106.47 (Källén–Lehmann representation).

For an interacting scalar theory the exact two-point function is

\begin{equation}\tag{106.61} \widetilde{\Delta}(p) =\frac{\ii\hbar^{3}cZ}{p^{2}-m^{2}c^{2}+\ii\epsilon} +\int_{(2mc)^{2}}^{\infty}\!\!\dd s\;\rho(s)\, \frac{\ii\hbar^{3}c}{p^{2}-s+\ii\epsilon}\ec \end{equation}

where \(m\) is the physical mass of the one-particle state, \(s\) is a squared momentum, \(\rho\geq0\) is the spectral density of the multiparticle continuum, and the field-strength renormalization \(Z=\abs{\bra{0}\phi(0)\ket{p}}^{2}/\abs{\bra{0}\phi_{\mathrm{free}}(0) \ket{p}}^{2}\) is a dimensionless number with \(0\leq Z\leq1\). Rests on Theorem 106.20 and Equation (106.30).

Derivation. Derives Proposition 106.47. Insert a complete set of eigenstates of the four-momentum operator between the two fields in \(\bra{0}T\phi(x)\phi(y)\ket{0}\). Poincaré invariance and the spectrum condition (Theorem 105.34) organize them into the vacuum, which contributes only \(\abs{\bra{0}\phi\ket{0}}^{2}\) and is removed by normal-ordering the field; the one-particle states of mass \(m\), whose contribution is fixed by Lorentz invariance to be that of a free field of mass \(m\) times a constant \(Z\); and states whose invariant mass is at least \(2m\), summarized by a density \(\rho(s)\geq0\) supported on \(s\geq(2mc)^{2}\). The positivity of \(\rho\) is the positivity of the Hilbert-space norm, i.e. Equation (106.30) once more, and the sum of all contributions must reproduce the canonical equal-time commutator, which gives \(Z+\int\rho=1\) and hence \(Z\leq1\).

Theorem 106.48 (LSZ reduction).

Let \(\widetilde{G}_{n}(p_{1},\dots,p_{n})\) be the momentum-space connected correlator of \(n\) fields, all momenta incoming. Then the \(S\)-matrix element for the process in which the momenta with \(p_{k}^{0}<0\) are incoming particles and the rest outgoing is

\begin{equation}\tag{106.62} \ii\mathcal{M}\,(2\pi\hbar)^{4} \delta^{4}\!\left(\textstyle\sum_{k}p_{k}\right) =\lim_{p_{k}^{2}\to m^{2}c^{2}}\; \prod_{k=1}^{n} \frac{p_{k}^{2}-m^{2}c^{2}}{\ii\hbar^{3}c\sqrt{Z}}\; \widetilde{G}_{n}(p_{1},\dots,p_{n})\ec \end{equation}

that is: put each external leg on its mass shell, divide out the full propagator, and take one factor \(\sqrt{Z}\) off each leg. Rests on Proposition 106.47.

Derivation. Derives Theorem 106.48. Consider the dependence of \(\widetilde{G}_{n}\) on one momentum \(p_{1}\). By Proposition 106.47 applied to the leg carrying \(p_{1}\), the correlator has a simple pole at \(p_{1}^{2}=m^{2}c^{2}\) whose residue is \(\ii\hbar^{3}cZ\) times the matrix element of the remaining fields between the vacuum and the one-particle state \(\ket{p_{1}}\), divided by \(\sqrt{Z}\) — one factor of \(\sqrt{Z}\) being consumed by the overlap \(\bra{0}\phi(0)\ket{p_{1}}\) and the other remaining in the residue. Multiplying by \((p_{1}^{2}-m^{2}c^{2})\) and taking the limit isolates that residue and discards the multiparticle continuum, which has only a branch cut starting at \((2mc)^{2}\) and is regular at the pole. Repeating for each leg replaces every field by an asymptotic one-particle state; what remains between them is by definition the \(S\)-matrix element, which carries the overall momentum-conserving delta function.

The physical statement behind the algebra is that an asymptotic particle is a pole in a correlation function, and nothing else: the interacting field \(\phi\) need not create a one-particle state cleanly, and generically it creates a superposition of one-particle and multiparticle states. The prescription discards the second because it is analytically distinguishable from the first.

Corollary 106.49 (Any interpolating field will do).

Let \(\mathcal{O}(x)\) be any local operator with the quantum numbers of the particle and with \(\bra{0}\mathcal{O}(0)\ket{p}\neq0\). Then Equation (106.62) holds with \(\phi\) replaced by \(\mathcal{O}\) and \(\sqrt{Z}\) replaced by \(\bra{0}\mathcal{O}(0)\ket{p}\) appropriately normalized, and the \(S\)-matrix element obtained is the same. Rests on Theorem 106.48 and Equation (106.61).

Derivation. Derives Corollary 106.49. The derivation above used only two properties of \(\phi\): that it is local, so the spectral decomposition Equation (106.61) applies to it; and that it has non-zero overlap with the one-particle state, so the pole is present. Any \(\mathcal{O}\) with those properties has the same pole at \(p^{2}=m^{2}c^{2}\), differing only in its residue, and the prescription divides that residue out. The \(S\)-matrix element is a property of the states, not of the operator used to reach them.

Corollary 106.49 is why lattice computations may use whatever operator is numerically convenient to create a proton — any gauge-invariant combination of three quark fields with the right quantum numbers — and still measure the same mass. It is also the reason the composite pion field and the quark bilinear \(\bar d\gamma^{5}u\) give the same \(\pi\to\mu\nu\) amplitude, which the weak-interaction phenomenology of Weak Interactions takes for granted. The scattering formalism into which Equation (106.62) feeds is Scattering Theory, and the analyticity assumptions under which the pole structure used here is guaranteed belong to Axiomatic Quantum Field Theory.

Gauge fields

Faddeev–Popov ghosts

Notation for the non-abelian case

Notation 106.50 (Absorbing the coupling into the potential).

Let \(T^{a}\) be Hermitian generators of a compact simple Lie group with \(\comm{T^{a}}{T^{b}}=\ii f^{abc}T^{c}\) and \(\tr(T^{a}T^{b})=\frac{1}{2}\delta^{ab}\) (Section 14.3). Write

\begin{equation}\tag{106.63} \mathcal{A}_{\mu}:=\frac{g}{\hbar}A^{a}_{\mu}T^{a}\ec\qquad [\mathcal{A}_{\mu}]=/\mathrm{m}\ec \end{equation}

so that the covariant derivative is \(D_{\mu}=\pp_{\mu}+\ii\mathcal{A}_{\mu}\) — which for one abelian generator is exactly Equation (100.9) with \(\mathcal{A}_{\mu}=QA_{\mu}/\hbar\) for a field of charge \(Q\) — and the field strength is

\begin{equation}\tag{106.64} \mathcal{F}_{\mu\nu}:=\frac{1}{\ii}\comm{D_{\mu}}{D_{\nu}} =\pp_{\mu}\mathcal{A}_{\nu}-\pp_{\nu}\mathcal{A}_{\mu} +\ii\comm{\mathcal{A}_{\mu}}{\mathcal{A}_{\nu}}\ec\qquad [\mathcal{F}_{\mu\nu}]=/\mathrm{m}^{2}\ep \end{equation}

The Yang–Mills Lagrangian density is then

\begin{equation}\tag{106.65} \Lag_{\mathrm{YM}} =-\frac{\hbar c}{2\hat g^{2}}\, \tr\!\left(\mathcal{F}_{\mu\nu}\mathcal{F}^{\mu\nu}\right) =-\frac{\hbar c}{4\hat g^{2}}\, \mathcal{F}^{a}_{\mu\nu}\mathcal{F}^{a\,\mu\nu}\ec \end{equation}

of dimension \(\mathrm{J}/\mathrm{m}^{3}\), since \(\hbar c=3.16153\times 10^{-26}\,\mathrm{J}\,\mathrm{m}\) and \(\mathcal{F}^{2}\) carries \(\mathrm{m}^{-4}\). The number \(\hat g^{2}\) is the dimensionless coupling,

\begin{equation}\tag{106.66} \hat g^{2}:=4\pi\alpha_{s}\ec \end{equation}

in exact parallel with Equation (100.3): for electromagnetism \(\hat g^{2}=\mu_{0}ce^{2}/\hbar=4\pi\alpha\) by Equation (100.4), \(\mathcal{F}_{\mu\nu} =eF_{\mu\nu}/\hbar\), and Equation (106.65) collapses to \(-F_{\mu\nu}F^{\mu\nu}/4\mu_{0}\), which is Equation (100.12). For the strong interaction \(\alpha_{s}\) is the running coupling measured in Equation (102.47). Infinitesimal gauge transformations are

\begin{equation}\tag{106.67} \delta\mathcal{A}_{\mu}=D_{\mu}\omega :=\pp_{\mu}\omega+\ii\comm{\mathcal{A}_{\mu}}{\omega}\ec \end{equation}

with \(\omega=\omega^{a}T^{a}\) dimensionless (\(\omega^{a}=g\Lambda^{a}/\hbar\) in the notation of Equation (100.8)).

Every formula of this section is therefore free of \(g\) and \(\hbar\) except through the single dimensionless number \(\hat g^{2}\), and the factor \(\hbar c/\hat g^{2}\) multiplying the action. That is not a retreat into natural units: \(\mathcal{A}\), \(\mathcal{F}\) and \(\omega\) have been given SI dimensions above, and the action \(S_{\mathrm{YM}}=c^{-1}\int\dd^{4}x\,\Lag_{\mathrm{YM}} =-\hbar\hat g^{-2}\int\dd^{4}x\, \mathcal{F}^{a}_{\mu\nu}\mathcal{F}^{a\mu\nu}/4\) carries \(\mathrm{J}\,\mathrm{s}\) because \(\hbar\) stands in front of a dimensionless integral. It is precisely this form that makes \(S/\hbar\) visibly equal to a pure number divided by \(\hat g^{2}\), which is the origin of the \(\ee^{-8\pi^{2}/\hat g^{2}}\) of Section 106.6.1.

The redundancy, and the divergence it causes

Section 100.1.3 showed that the quadratic form of the photon action is singular: \(\widetilde{M}_{\mu\nu}p^{\nu}=0\) in Equation (100.15). In the functional integral the same fact appears as an infinite factor.

Proposition 106.51 (The naive integral diverges by the volume of the gauge group).

The integrand of \(\int\mathcal{D}\mathcal{A}\, \ee^{\ii S_{\mathrm{YM}}/\hbar}\) is constant along each gauge orbit \(\set{\mathcal{A}^{\omega}:\omega\in\mathcal{G}}\), so the integral contains an overall factor equal to the volume of the group \(\mathcal{G}\) of gauge transformations — an infinite constant, independent of the physics but not factorizable in the presence of sources, since the source term is not gauge invariant. Rests on Equations (106.65) and (106.67).

Derivation. Derives Proposition 106.51. \(S_{\mathrm{YM}}\) is invariant under Equation (106.67) by construction, and the measure \(\mathcal{D}\mathcal{A}\) is invariant because Equation (106.67) is, at each point, a translation of \(\mathcal{A}_{\mu}\) by \(\pp_{\mu}\omega\) plus a rotation in the adjoint index, and both preserve the Lebesgue measure on the fibre. The integral therefore factorizes as \(\bigl(\int_{\mathcal{G}}\mathcal{D}\omega\bigr)\times \bigl(\int_{\mathcal{A}/\mathcal{G}}\ee^{\ii S/\hbar}\bigr)\), the first factor being the divergent volume.

Faddeev and Popov's device [Faddeev:1967] is to insert a resolution of unity that selects one representative from each orbit and supplies the Jacobian of that choice.

Theorem 106.52 (Faddeev–Popov).

Let \(G^{a}[\mathcal{A}](x)\) be a gauge-fixing function — for the covariant family, \(G^{a}=\pp^{\mu}\mathcal{A}^{a}_{\mu}\) — such that \(G^{a}[\mathcal{A}^{\omega}]=\sigma^{a}\) has exactly one solution \(\omega\) on each orbit. Then

\begin{equation}\tag{106.68} \int\mathcal{D}\mathcal{A}\;\ee^{\ii S_{\mathrm{YM}}/\hbar} =\left(\int_{\mathcal{G}}\mathcal{D}\omega\right) \int\mathcal{D}\mathcal{A}\; \delta\bigl[G[\mathcal{A}]-\sigma\bigr]\, \det\mathcal{M}[\mathcal{A}]\; \ee^{\ii S_{\mathrm{YM}}/\hbar}\ec \end{equation}

with the Faddeev–Popov operator

\begin{equation}\tag{106.69} \mathcal{M}^{ab}(x,y) =\frac{\delta G^{a}[\mathcal{A}^{\omega}](x)} {\delta\omega^{b}(y)}\bigg|_{\omega=0} =\pp^{\mu}\bigl(D_{\mu}\bigr)^{ab}\delta^{4}(x-y)\ec\qquad \bigl(D_{\mu}\bigr)^{ab} =\delta^{ab}\pp_{\mu}+f^{abc}\mathcal{A}^{c}_{\mu}\ec \end{equation}

for the covariant gauge condition; the sign of the second term is fixed by \(\comm{T^{a}}{T^{b}}=\ii f^{abc}T^{c}\) together with \(D_{\mu}=\pp_{\mu}+\ii\mathcal{A}_{\mu}\) and is derived below. The divergent factor now stands outside and cancels in every normalized expectation value. Rests on Proposition 106.51 and Equation (106.67).

Derivation. Derives Theorem 106.52. For each \(\mathcal{A}\) define \(\Delta[\mathcal{A}]\) by

\[ \Delta[\mathcal{A}]^{-1} :=\int_{\mathcal{G}}\mathcal{D}\omega\; \delta\bigl[G[\mathcal{A}^{\omega}]-\sigma\bigr]\ep \]

By hypothesis the delta functional is supported at a single \(\omega_{*}\), and the one-dimensional identity \(\int\dd x\,\delta(f(x))=\abs{f'(x_{*})}^{-1}\) generalizes to \(\Delta[\mathcal{A}]=\abs{\det(\delta G/\delta\omega)}_{\omega_{*}} =\abs{\det\mathcal{M}}\). Also \(\Delta\) is gauge invariant, since \(\mathcal{D}\omega\) is a left- and right-invariant Haar measure on \(\mathcal{G}\) and the defining integral is unchanged by relabelling the orbit representative. Insert \(1=\Delta[\mathcal{A}]\int\mathcal{D}\omega\, \delta[G[\mathcal{A}^{\omega}]-\sigma]\) into the left-hand side of Equation (106.68), exchange the order of integration, and change variables \(\mathcal{A}\to\mathcal{A}^{-\omega}\); the action, the measure and \(\Delta\) are all invariant, so the \(\omega\) integral factorizes out as the group volume and what remains is the right-hand side.

For the covariant condition, \(G^{a}[\mathcal{A}^{\omega}] =\pp^{\mu}\bigl(\mathcal{A}^{a}_{\mu}+(D_{\mu}\omega)^{a}\bigr)\) by Equation (106.67), whose derivative with respect to \(\omega^{b}\) is \(\pp^{\mu}(D_{\mu})^{ab}\), i.e.\ Equation (106.69). The component form of the adjoint covariant derivative follows from \(\ii\comm{\mathcal{A}_{\mu}}{\omega} =\ii\mathcal{A}^{c}_{\mu}\omega^{b}\ii f^{cba}T^{a} =-f^{cba}\mathcal{A}^{c}_{\mu}\omega^{b}T^{a}\), and \(f^{cba}=-f^{bca}=-f^{abc}\) by total antisymmetry and cyclicity, so that \((D_{\mu}\omega)^{a} =\pp_{\mu}\omega^{a}+f^{abc}\mathcal{A}^{c}_{\mu}\omega^{b}\) — the plus sign of Equation (106.69). Equivalently, the adjoint generators are \((T^{a}_{\mathrm{adj}})^{bc}=-\ii f^{abc}\), as they are for \(\SU(3)\) in Proposition 14.55, and \(\delta^{bc}\pp_{\mu}+\ii\mathcal{A}^{a}_{\mu} (T^{a}_{\mathrm{adj}})^{bc} =\delta^{bc}\pp_{\mu}+f^{abc}\mathcal{A}^{a}_{\mu}\) is the same operator with the summed index moved by one cyclic step.

Corollary 106.53 (Why electrodynamics has no ghosts).

For an abelian group \(f^{abc}=0\), so \(\mathcal{M}=\pp^{\mu}\pp_{\mu}=\Box\) is independent of \(\mathcal{A}\); \(\det\mathcal{M}\) is then a constant that cancels between numerator and denominator, and the Faddeev–Popov procedure reduces to the gauge-fixing term Equation (100.16) alone. This is the derivation of a fact that Section 100.1.3 could only assert. Rests on Theorem 106.52, Equation (106.69) and Equation (100.16).

Exponentiating the determinant: ghosts

The delta functional is inconvenient and the determinant is not a local object. Both are repaired at once, and it is Section 106.3.3 that repairs them.

Theorem 106.54 (The gauge-fixed action).

Averaging Equation (106.68) over the arbitrary function \(\sigma\) with the Gaussian weight \(\exp\bigl\{-\frac{\ii}{2\hat g^{2}\xi} \int\dd^{4}x\,\sigma^{a}\sigma^{a}\bigr\}\) — dimensionless, since \(\sigma\) carries \(/\mathrm{m}^{2}\) — and writing the determinant as a Grassmann integral by Equation (106.51) gives

\begin{equation}\tag{106.70} Z=\int\mathcal{D}\mathcal{A}\,\mathcal{D}\bar c\,\mathcal{D}c\; \exp\!\left\{\frac{\ii}{\hbar} \bigl(S_{\mathrm{YM}}+S_{\mathrm{gf}}\bigr)\right\}\, \exp\!\left\{-\int\dd^{4}x\; \bar c^{a}\bigl(\pp^{\mu}D_{\mu}\bigr)^{ab}c^{b}\right\}\ec \end{equation}

with

\begin{equation}\tag{106.71} S_{\mathrm{gf}}=-\frac{\hbar}{2\hat g^{2}\xi}\int\dd^{4}x\; \bigl(\pp^{\mu}\mathcal{A}^{a}_{\mu}\bigr)^{2}\ec \end{equation}

the \(R_{\xi}\) family, \(\xi\) dimensionless. The ghost fields \(c^{a},\bar c^{a}\) are anticommuting, carry dimension \(/\mathrm{m}\), are scalars under the Lorentz group, and transform in the adjoint representation of the gauge group. Rests on Equation (106.68), Equation (106.51) and Lemma 106.2.

Derivation. Derives Theorem 106.54. The left-hand side of Equation (106.68) does not depend on \(\sigma\), so it may be averaged over \(\sigma\) with any normalized weight; the Gaussian one converts \(\delta[G-\sigma]\) into \(\exp\bigl(\ii S_{\mathrm{gf}}/\hbar\bigr)\) with \(S_{\mathrm{gf}}\) as stated, by Lemma 106.2. The determinant is written as a Grassmann integral by reading Equation (106.51) from right to left with \(A=\mathcal{M}\) — and it is worth noting that this is the one place in physics where the exponent of a determinant being \(+1\) rather than \(-\frac{1}{2}\) decides the statistics of a field: the Jacobian sits in the numerator, so the field that represents it must be the one whose Gaussian integral gives \(\det\) and not \((\det)^{-1/2}\), i.e. an anticommuting one, by Theorem 106.37. The ghosts are spin-zero fields obeying Fermi statistics, which is why they cannot be physical: they violate the spin–statistics connection of Identical Particles.

Written as a single exponent \(\ii S_{\mathrm{tot}}/\hbar\), the ghost term is \(S_{\mathrm{gh}}=\ii\hbar\int\dd^{4}x\, \bar c^{a}(\pp^{\mu}D_{\mu})^{ab}c^{b}\), whose explicit factor of \(\ii\) records that the ghost action is not real. Nothing is wrong: an action is required to be real only when its exponential must be a phase for physical configurations, and the ghosts have no physical configurations. The dimension assignment follows from requiring the exponent in Equation (106.70) to be a pure number: \(\int\dd^{4}x\) supplies \(\mathrm{m}^{4}\), \(\pp^{\mu}D_{\mu}\) supplies \(/\mathrm{m}^{2}\), so \(\bar cc\) must supply \(/\mathrm{m}^{2}\).

Proposition 106.55 (Ghosts are required by unitarity).

In a non-abelian theory the ghost loop is not optional. Cutting a one-loop gluon self-energy diagram with the covariant polarization sum \(-\eta^{\mu\nu}\) counts four intermediate polarizations where the physical theory has two; the ghost loop, carrying \((-1)\) by Corollary 106.39 and running over a pair \((c,\bar c)\), subtracts exactly the two unphysical ones, and the optical theorem is restored. Rests on Theorem 106.54 and Corollary 106.39.

Derivation. Derives Proposition 106.55. A massless vector of momentum \(k\) has two physical transverse polarizations, whose sum is \(\sum_{\lambda=1,2}\varepsilon^{\mu}_{(\lambda)} \varepsilon^{*\nu}_{(\lambda)} =-\eta^{\mu\nu}+(k^{\mu}n^{\nu}+n^{\mu}k^{\nu})/(k\cdot n)\) for a lightlike auxiliary \(n\). In a covariant gauge the propagator supplies \(-\eta^{\mu\nu}\), which is the same object plus the longitudinal and timelike pair; those two do not cancel of their own accord when the external current is not conserved, and in a non-abelian theory it is not conserved diagram by diagram, because the gluon carries colour and its own current is only covariantly conserved. The discrepancy is a two-state contribution of positive norm. The ghost loop supplies two scalar states (\(c\) and \(\bar c\)) with the sign \((-1)\) of a closed anticommuting loop, i.e. \(-2\) states, which cancels it identically. The all-orders statement — that the cancellation holds to every order and for every process — is the Kugo–Ojima quartet mechanism of Section 106.4.2, and it is a theorem about the BRST cohomology rather than a diagrammatic accident.

The same construction can be reached from the Hamiltonian side. Gauge invariance makes the Yang–Mills Lagrangian singular in Dirac's sense: the momentum conjugate to \(\mathcal{A}_{0}\) vanishes identically, so the system carries first-class constraints and its quantization requires the Dirac-bracket apparatus of Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism. Carrying that programme through and then integrating out the non-dynamical components reproduces Equation (106.70), ghosts and all. The functional route is shorter and manifestly covariant, which is why it is the one used; the Hamiltonian route is what shows that the answer is not an artefact of the trick.

BRST symmetry

The gauge-fixed action Equation (106.70) is not gauge invariant — that was the point of fixing the gauge. Becchi, Rouet and Stora [Becchi:1974] [Becchi:1976], and independently Tyutin [Tyutin:1975], found that it retains an exact global symmetry whose parameter is anticommuting, and that this residual symmetry carries everything the lost gauge invariance was doing.

Definition 106.56 (The BRST operator).

Introduce the auxiliary Nakanishi–Lautrup field \(B=B^{a}T^{a}\), of dimension \(/\mathrm{m}^{2}\), and define a Grassmann-odd derivation \(s\) by

\begin{equation}\tag{106.72} s\mathcal{A}_{\mu}=D_{\mu}c\ec\qquad sc=-\ii c^{2}\ec\qquad s\bar c=B\ec\qquad sB=0\ec \end{equation}

where \(c=c^{a}T^{a}\), \(\bar c=\bar c^{a}T^{a}\), and \(c^{2}=\frac{1}{2}\acomm{c}{c}\) is non-zero because the components anticommute while the matrices do not. On a matter field \(\psi\) transforming in a representation \(R\), \(s\psi=-\ii c\psi\). The operator acts on products by \(s(XY)=(sX)Y+(-1)^{\abs{X}}X(sY)\).

Theorem 106.57 (Nilpotency).

\(s^{2}=0\) on every field of Equation (106.72). Rests on Equation (106.72).

Derivation. Derives Theorem 106.57. On the ghost: \(s^{2}c=-\ii\bigl[(sc)c-c(sc)\bigr] =-\ii\bigl[(-\ii c^{2})c-c(-\ii c^{2})\bigr] =-c^{3}+c^{3}=0\), the sign in the middle being that of \(s\) passing the odd \(c\). Associativity of matrix multiplication has done here exactly the work that the Jacobi identity does in component notation, which is the reason for writing the algebra in matrix form.

On the connection, expand \(s^{2}\mathcal{A}_{\mu}=s\bigl(\pp_{\mu}c +\ii\comm{\mathcal{A}_{\mu}}{c}\bigr)\). Since \(\mathcal{A}\) is even and \(c\) odd, \(s\comm{\mathcal{A}_{\mu}}{c} =\acomm{s\mathcal{A}_{\mu}}{c}+\comm{\mathcal{A}_{\mu}}{sc}\), so

\[ s^{2}\mathcal{A}_{\mu} =\pp_{\mu}(sc)+\ii\acomm{D_{\mu}c}{c} +\ii\comm{\mathcal{A}_{\mu}}{sc}\ep \]

Insert \(sc=-\ii c^{2}\) and use \(\pp_{\mu}(c^{2})=\acomm{\pp_{\mu}c}{c}\) and \(\acomm{\comm{\mathcal{A}_{\mu}}{c}}{c} =\comm{\mathcal{A}_{\mu}}{c^{2}}\):

\[ s^{2}\mathcal{A}_{\mu} =-\ii\acomm{\pp_{\mu}c}{c} +\ii\acomm{\pp_{\mu}c}{c} -\comm{\mathcal{A}_{\mu}}{c^{2}} +\comm{\mathcal{A}_{\mu}}{c^{2}}=0\ep \]

On \(\bar c\) and \(B\) nilpotency is immediate from \(sB=0\). Since \(s\) is a derivation and vanishes on the square of every generator, \(s^{2}=0\) on any polynomial in the fields.

Theorem 106.58 (The gauge-fixing sector is BRST-exact).

With the gauge fermion

\begin{equation}\tag{106.73} \Psi=-\frac{\hbar}{\hat g^{2}}\int\dd^{4}x\; \bar c^{a}\left(\pp^{\mu}\mathcal{A}^{a}_{\mu} -\frac{\xi}{2}B^{a}\right)\ec \end{equation}

the total action is

\begin{equation}\tag{106.74} S_{\mathrm{tot}}=S_{\mathrm{YM}}+s\Psi\ec \end{equation}

and \(sS_{\mathrm{tot}}=0\), since \(S_{\mathrm{YM}}\) is gauge invariant and \(s^{2}=0\). Rests on Equation (106.72), Equation (106.71) and Theorem 106.57.

Derivation. Derives Theorem 106.58. Applying \(s\) to Equation (106.73) and using Equation (106.72),

\[ s\Psi=-\frac{\hbar}{\hat g^{2}}\int\dd^{4}x\left[ B^{a}\left(\pp^{\mu}\mathcal{A}^{a}_{\mu} -\frac{\xi}{2}B^{a}\right) -\bar c^{a}\,\pp^{\mu}\bigl(D_{\mu}c\bigr)^{a}\right]\ec \]

the sign of the second term coming from \(s\) passing the odd \(\bar c\). The auxiliary field is algebraic and may be integrated out, its equation of motion being \(B^{a}=\pp^{\mu}\mathcal{A}^{a}_{\mu}/\xi\); the first bracket then becomes \((\pp^{\mu}\mathcal{A}^{a}_{\mu})^{2}/2\xi\) and, with the prefactor, is exactly \(S_{\mathrm{gf}}\) of Equation (106.71). The second term is \(+(\hbar/\hat g^{2})\int\dd^{4}x\, \bar c^{a}\pp^{\mu}(D_{\mu}c)^{a}\), which enters the weight \(\exp(\ii S_{\mathrm{tot}}/\hbar)\) as \(\exp\bigl[(\ii/\hat g^{2})\int\bar c\,\pp^{\mu}D_{\mu}c\bigr]\), whereas Equation (106.70) carries \(\exp\bigl[-\int\bar c\,\pp^{\mu}D_{\mu}c\bigr]\). The two differ by the constant factor \(-\ii/\hat g^{2}\) multiplying the ghost bilinear, and are brought into coincidence by rescaling the pair, \(c\mapsto\kappa c\) and \(\bar c\mapsto\kappa\bar c\) with \(\kappa^{2}=\ii\hat g^{2}\). It must be the pair and not one against the other: the ghosts enter only through the product \(\bar cc\), so a rescaling \(c\mapsto\kappa c\), \(\bar c\mapsto\bar c/\kappa\) leaves the ghost term exactly as it was and cannot fix a normalization. The rescaling used here multiplies the Grassmann measure by a field-independent constant, by Corollary 106.36, which cancels between numerator and denominator in every normalized expectation value — equivalently, \(\det(\lambda\mathcal{M})=\lambda^{n}\det\mathcal{M}\) and only ratios of determinants are ever used. Finally \(sS_{\mathrm{YM}}=0\) because \(s\) acts on \(\mathcal{A}\) as a gauge transformation with parameter \(c\), and \(s(s\Psi)=s^{2}\Psi=0\) by Theorem 106.57.

Definition 106.59 (BRST charge and physical states).

Let \(Q_{B}\) be the conserved charge generating \(s\) through the graded commutator, \(sX=\ii\hbar^{-1}\comm{Q_{B}}{X}_{\pm}\); nilpotency of \(s\) makes \(Q_{B}^{2}=0\). A state is physical when it is annihilated by \(Q_{B}\), and two physical states differing by \(Q_{B}\) applied to something are identified:

\begin{equation}\tag{106.75} \mathcal{H}_{\mathrm{phys}} :=\frac{\ker Q_{B}}{\im Q_{B}}\ep \end{equation}

The construction is the standard one for any nilpotent operator: \(Q_{B}\) is a differential, \(\ker Q_{B}/\im Q_{B}\) its cohomology, and the physical Hilbert space is a cohomology class rather than a subspace. The differential-geometric setting in which that pattern is usually met — connections and curvature on a principal bundle — belongs to Section 14.6, which at present fixes the notation used by Theorem 14.111 and leaves the underlying definitions owed. Two facts make the construction work, and both are checkable.

Proposition 106.60 (The quartet mechanism).

The unphysical modes of a gauge field organize into quartets: for each colour and momentum, the longitudinal and timelike polarizations together with the ghost and antighost form a set of four states on which \(Q_{B}\) acts by mapping two into the other two. Every such quartet contributes zero to any matrix element between states of \(\ker Q_{B}\), so unphysical states cancel in pairs [Kugo:1979]. Rests on Equation (106.72) and Definition 106.59.

Derivation. Derives Proposition 106.60. From Equation (106.72) the free-field limit of the transformations is \(s\mathcal{A}_{\mu}=\pp_{\mu}c\), \(sc=0\), \(s\bar c=B\), \(sB=0\): the ghost \(c\) is \(Q_{B}\)-closed and maps the longitudinal polarization; the antighost \(\bar c\) maps onto \(B\), which is the timelike combination on shell. So among the four states \(\set{\mathcal{A}_{L},\mathcal{A}_{0},c,\bar c}\), two are \(Q_{B}\)-exact and two are not \(Q_{B}\)-closed. A physical state, being \(Q_{B}\)-closed, has zero inner product with a \(Q_{B}\)-exact one (since \(\braket{\Phi}{Q_{B}\Lambda}=\braket{Q_{B}\Phi}{\Lambda}=0\) by the hermiticity of \(Q_{B}\)), and the states that are not \(Q_{B}\)-closed are excluded from \(\mathcal{H}_{\mathrm{phys}}\) by definition. Hence the quartet decouples entirely, leaving the two transverse polarizations, which is Proposition 106.55 at all orders. The inner product induced on \(\mathcal{H}_{\mathrm{phys}}\) is therefore positive definite and the \(S\)-matrix restricted to it is unitary [Kugo:1979].

Theorem 106.61 (Gauge-parameter independence).

Any expectation value of a \(Q_{B}\)-closed operator is independent of the gauge parameter \(\xi\) and of the choice of gauge-fixing function altogether. Rests on Equation (106.74) and Theorem 106.58.

Derivation. Derives Theorem 106.61. By Equation (106.74) the entire dependence on the gauge-fixing choice is in \(\Psi\). Let \(\mathcal{O}\) satisfy \(s\mathcal{O}=0\) and vary \(\Psi\to\Psi+\delta\Psi\). Then

\[ \delta\avg{\mathcal{O}} =\frac{\ii}{\hbar}\avg{\mathcal{O}\,s\,\delta\Psi} =\frac{\ii}{\hbar}\avg{s\bigl(\mathcal{O}\,\delta\Psi\bigr)}=0\ec \]

the second step because \(s\mathcal{O}=0\) and the last because the expectation value of a BRST variation vanishes: the measure and the action are \(s\)-invariant, so \(\avg{sX}=0\) for every \(X\), by the same argument that gives a Ward identity from any symmetry of the functional integral. In particular the \(\xi\)-dependence of Equation (106.71) cancels out of every physical quantity, which is the checkable statement quoted at Equation (100.18).

That is what BRST buys, and it is worth listing plainly because the formalism can look like machinery for its own sake. It supplies (i) a definition of the physical state space in a manifestly covariant gauge, Equation (106.75); (ii) a proof that the unphysical states decouple to all orders, Proposition 106.60; (iii) a one-line proof of gauge-parameter independence, Theorem 106.61, replacing an order-by-order diagrammatic check; and (iv) the functional identities of Section 106.4.3, which are what makes renormalizability provable. The generalization to gauge algebras that close only on shell is the antifield construction of Batalin and Vilkovisky [Batalin:1981], in which \(s\) is generated by an antibracket and nilpotency becomes a master equation; it is not needed for the Standard Model, whose gauge algebra closes off shell, and it is mentioned here only because Section 106.4.3 states its identity in that language.

Slavnov–Taylor identities

Theorem 106.62 (Slavnov–Taylor).

For every functional \(X\) of the fields,

\begin{equation}\tag{106.76} \avg{sX}=0\ec \end{equation}

and the resulting relations among Green functions are the Slavnov–Taylor identities [Slavnov:1972] [Taylor:1971]. In the abelian case they reduce to the Ward–Takahashi identity Equation (100.71). Rests on Equation (106.72), Equation (106.70) and Theorem 106.58.

Derivation. Derives Theorem 106.62. Change variables in Equation (106.70) by an infinitesimal BRST transformation with a constant anticommuting parameter \(\zeta\): \(\Phi\to\Phi+\zeta\,s\Phi\) for every field \(\Phi\). The action is invariant by Theorem 106.58. The measure is invariant too, and this requires proof rather than assertion: the Jacobian of the transformation is \(\det\bigl(1+\zeta\,\pp(s\Phi)/\pp\Phi\bigr) =1+\zeta\,\tr\bigl(\pp(s\Phi)/\pp\Phi\bigr)\) with a sign for each Grassmann direction, and the bosonic and fermionic traces cancel — the \(\mathcal{A}\)–\(c\) block against the \(c\)–\(c\) block — because \(s\) maps a boson to a ghost and a ghost to a ghost bilinear with matched coefficients. (This cancellation is exactly what fails for a chiral rotation of the fermion measure, and its failure is the anomaly of Section 106.7.3; a gauge anomaly is the statement that it fails here too, and Section 106.7.4 is the condition that it does not.) With action and measure invariant, the integral of a total BRST variation vanishes, which is Equation (106.76).

In the abelian case the ghosts decouple by Corollary 106.53 and Equation (106.72) collapses to \(s\mathcal{A}_{\mu}=\pp_{\mu}c\) with \(c\) a free field, so that a BRST variation is a gauge transformation whose parameter is an anticommuting constant times that free field. Taking \(X=\bar c(y)\,\psi(x_{1})\bar\psi(x_{2})\) then turns Equation (106.76) into the relation between the vertex function and the inverse propagator, \(q_{\mu}\Gamma^{\mu}(p',p)=S^{-1}(p')-S^{-1}(p)\), whose derivation is carried out in full at Theorem 100.48.

Corollary 106.63 (The Zinn-Justin equation).

Coupling sources \(K^{\mu}_{a},L_{a}\) to the composite operators \(s\mathcal{A}^{a}_{\mu}\) and \(sc^{a}\) and Legendre-transforming as in Definition 106.41, the whole family Equation (106.76) collapses into the single bilinear equation

\begin{equation}\tag{106.77} \int\dd^{4}x\left[ \frac{\delta\Gamma}{\delta K^{\mu}_{a}} \frac{\delta\Gamma}{\delta\mathcal{A}^{a}_{\mu}} +\frac{\delta\Gamma}{\delta L_{a}} \frac{\delta\Gamma}{\delta c^{a}}\right]=0\ec \end{equation}

which is the statement that the effective action is BRST invariant with the quantum-corrected transformation law. Rests on Definition 106.41, Equation (106.76) and Theorem 106.57.

Derivation. Derives Corollary 106.63. Add to the action the terms \(\int\dd^{4}x\,\bigl(K^{\mu}_{a}\,s\mathcal{A}^{a}_{\mu} +L_{a}\,sc^{a}\bigr)\). They are BRST invariant by Theorem 106.57, since \(s\) applied to them gives \(K\,s^{2}\mathcal{A}+L\,s^{2}c=0\), so the extended action is still \(s\)-invariant and Equation (106.76) still holds. Take \(X=\Phi\) for each field in turn: the resulting identities say that \(W\) is annihilated by the operator \(\int\dd^{4}x\,\bigl[(\delta W/\delta K^{\mu}_{a}) (\delta/\delta j^{a}_{\mu})+\cdots\bigr]\) acting on the sources. Legendre-transforming in the field variables only — \(K\) and \(L\) are carried along untouched, so that \(\delta\Gamma/\delta K=\delta W/\delta K\) — converts the source derivatives into field derivatives of \(\Gamma\) and produces Equation (106.77). The equation is bilinear in \(\Gamma\) because the composite operators \(s\mathcal{A}\) and \(sc\) are themselves products of two fields, so their sources cannot be eliminated linearly.

The identities do three jobs. First, they constrain the divergences: the counterterms needed to renormalize a gauge theory must themselves satisfy Equation (106.77), which forces the gauge coupling appearing in the three-gluon, four-gluon, ghost and matter vertices to renormalize by the same factor — the non-abelian generalization of \(Z_{1}=Z_{2}\) (Corollary 100.49). Without that, a single coupling constant would not stay single, and the theory would not be renormalizable with finitely many parameters. This is the role the identities play in 't Hooft's proof [tHooft:1971b] [tHooft:1971a], completed with the dimensional-regularization technique that respects them [tHooft:1972] — and the fact that dimensional regularization respects gauge invariance, where a momentum cutoff does not, is why Definition 100.30 is the scheme of choice throughout this part.

Second, they give the practical check of Theorem 106.61: a physical quantity computed in Feynman gauge \(\xi=1\) and in Landau gauge \(\xi=0\) must agree, while the individual propagators and vertices do not. Every serious calculation in this part is done twice for that reason.

Third, they are the object that a gauge anomaly destroys, which is the subject of Section 106.7.4.

The Gribov ambiguity

Theorem 106.52 assumed that \(G[\mathcal{A}^{\omega}]=\sigma\) has exactly one solution on each orbit. Gribov showed that for the Coulomb and Landau conditions it does not [Gribov:1978].

Proposition 106.64 (Gribov copies exist).

Let \(\mathcal{A}\) satisfy the Landau condition \(\pp^{\mu}\mathcal{A}_{\mu}=0\). A gauge-equivalent field \(\mathcal{A}^{\omega}\) satisfies the same condition, to first order in \(\omega\), if and only if

\begin{equation}\tag{106.78} \mathcal{M}^{ab}\omega^{b} =\pp^{\mu}\bigl(D_{\mu}\bigr)^{ab}\omega^{b}=0\ec \end{equation}

i.e. if and only if the Faddeev–Popov operator has a normalizable zero mode. Since \(\mathcal{M}=\Box+O(\mathcal{A})\) is positive for small \(\mathcal{A}\) and its lowest eigenvalue decreases continuously as \(\mathcal{A}\) grows, there is a surface in field space — the Gribov horizon — on which the lowest eigenvalue first vanishes, and beyond it copies exist. Rests on Equations (106.67) and (106.69).

Derivation. Derives Proposition 106.64. Under Equation (106.67), \(\pp^{\mu}\mathcal{A}^{\omega}_{\mu} =\pp^{\mu}\mathcal{A}_{\mu}+\pp^{\mu}(D_{\mu}\omega)\), so preserving the condition requires Equation (106.78). In the Euclidean theory the Landau condition \(\pp^{\mu}\mathcal{A}_{\mu}=0\) makes \(-\mathcal{M}\) Hermitian on square-integrable functions, with \(-\mathcal{M}^{ab} =-\delta^{ab}\pp^{2}-f^{abc}\mathcal{A}^{c\mu}\pp_{\mu}\) by Equation (106.69), so its spectrum is real; at \(\mathcal{A}=0\) it is the positive operator \(-\pp^{2}\) on non-constant modes. As the amplitude of \(\mathcal{A}\) is increased along any ray in field space the eigenvalues move continuously, and Gribov exhibited explicit configurations for which the lowest one crosses zero, at which point Equation (106.78) has a solution and the orbit meets the gauge-fixing surface again.

The failure is therefore not perturbative: it occurs only for fields of amplitude of order \(1/\hat g\), which is invisible to any finite order in \(\hat g^{2}\). Everything computed in Section 100.4 and Section 106.4.3 is unaffected. What is affected is any statement about the infrared behaviour of the non-abelian theory, which is where the large-amplitude configurations live.

Theorem 106.65 (Singer).

For a non-abelian structure group there is no continuous gauge-fixing prescription that works globally: no continuous section of the principal bundle of gauge fields over the space of orbits exists, for gauge fields on \(S^{4}\) or on \(S^{3}\) [Singer:1978]. Rests on Proposition 106.64.

Singer's statement is not about the particular choice \(\pp^{\mu}\mathcal{A}_{\mu}=0\); it says that every choice fails. The obstruction is topological: a gauge-fixing prescription is a section of the principal bundle whose base is the space of orbits, a principal bundle admits a global section if and only if it is trivial, and this one is not. Both statements are quoted here from [Singer:1978]; the bundle language they are written in belongs to Section 14.6, which so far carries only the notation that Theorem 14.111 needs and reserves the definition of a section itself. The Faddeev–Popov construction is therefore a local construction that has been used globally, and Theorem 106.52 is exactly true only in a neighbourhood of \(\mathcal{A}=0\).

Two responses exist, and the honest position is that neither is closed.

Restrict the integration region. The Gribov region \(\Omega\) is the set of transverse fields on which \(\mathcal{M}\) is positive; it is convex and bounded in every direction, and every orbit meets it. It still contains copies, and the fundamental modular region \(\Lambda\subset\Omega\), defined by minimizing \(\int\dd^{4}x\,\mathcal{A}^{a}_{\mu}\mathcal{A}^{a\mu}\) along the orbit, contains exactly one representative of each orbit except on its boundary. Restricting the functional integral to \(\Omega\) modifies the infrared behaviour of the gluon propagator, and the two candidate answers in the literature — a propagator vanishing at zero momentum and one approaching a finite constant — correspond to different treatments of the boundary. Which is right is not settled.

Do not fix a gauge at all. On the lattice (Section 106.5.1) the integration variables are group elements and the group volume per link is finite, so Proposition 106.51 does not arise: the integral over gauge orbits contributes a finite constant and is simply left in. No gauge fixing means no Faddeev–Popov determinant, no ghosts and no Gribov copies. This is the strongest practical argument for the lattice formulation, and it is why the results of Section 106.5.2 are free of the ambiguity entirely.

Derivation pending.

The Gribov horizon and Singer's obstruction: the derivation above establishes the criterion — a copy exists exactly when the Faddeev–Popov operator acquires a normalizable zero mode — and states that configurations reaching it exist. Gribov's explicit spherically symmetric family, the proof that the Gribov region is convex and bounded, and Singer's theorem that the bundle of gauge fields over the orbit space admits no continuous global section are stated here and not proved; all three belong in Appendix A.

The lattice

Wilson's lattice gauge theory

Everything so far has been perturbative, and Proposition 106.31 showed that perturbation theory cannot be the definition of the theory. Wilson supplied a definition [Wilson:1974]: replace Euclidean spacetime by a finite hypercubic lattice of spacing \(a\), so that the functional integral becomes an ordinary integral of finite dimension, and choose the variables so that gauge invariance is exact at finite \(a\) rather than only in the limit.

Proposition 106.67 (Exact gauge invariance at finite spacing).

Under \(U_{\mu}(n)\mapsto\Omega(n)U_{\mu}(n)\Omega^{\dagger} (n+\hat\mu)\) with \(\Omega(n)\in G\) arbitrary at each site, the trace of any closed loop of links is invariant. In particular \(\tr U_{p}\) is invariant. Rests on Definition 106.66 and Equation (106.80).

Derivation. Derives Proposition 106.67. Each link carries \(\Omega\) at its start and \(\Omega^{\dagger}\) at its end, so in a product of links forming a path the factors at every interior site cancel, \(\Omega^{\dagger}(m)\Omega(m)=1\). For a closed path the two remaining factors sit at the same site and are removed by the cyclicity of the trace.

This is the decisive structural point. The gauge field lives in the group, not in the algebra, and gauge invariance is therefore an exact symmetry of a finite-dimensional integral rather than a property that survives only in the continuum. There is no gauge fixing, hence no Faddeev–Popov determinant, no ghosts and — as Section 106.4.4 noted — no Gribov problem: the volume of the gauge group is \(\bigl(\text{vol}\,G\bigr)^{\text{sites}}\), a finite constant.

Theorem 106.68 (The Wilson action and its continuum limit).

Let

\begin{equation}\tag{106.81} S_{W}=\beta\sum_{p}\left(1-\frac{1}{N}\Re\tr U_{p}\right)\ec \qquad\beta=\frac{2N}{\hat g^{2}}\ec \end{equation}

the sum running over all plaquettes of the lattice, \(N\) the dimension of the fundamental representation, and \(\beta\) dimensionless. Then as \(a\to0\)

\begin{equation}\tag{106.82} S_{W}\longrightarrow\frac{1}{4\hat g^{2}} \int\dd^{4}x_{E}\; \mathcal{F}^{a}_{\mu\nu}\mathcal{F}^{a}_{\mu\nu} =\frac{S_{E}}{\hbar}\ec \end{equation}

with \(S_{E}\) the Euclidean Yang–Mills action of Equation (106.65). For \(\SU(3)\), \(\beta=6/\hat g^{2} =6/(4\pi\alpha_{s})\). This \(\beta\) is the lattice coupling and not the inverse temperature of Equation (106.27); under the dictionary of Table 106.1 it does play the role of an inverse temperature for the four-dimensional statistical system, which is why strong coupling is that system's high-temperature regime. Rests on Equations (106.64), (106.65) and (106.80).

Derivation. Derives Theorem 106.68. Multiplying the four exponentials of Equation (106.80) and applying the Baker–Campbell–Hausdorff formula, the leading term of the exponent is the discrete curl of \(\mathcal{A}\), so \(U_{p}=\exp\bigl[\ii a^{2}\mathcal{F}_{\mu\nu}+O(a^{3})\bigr]\) with \(\mathcal{F}_{\mu\nu}=\mathcal{F}^{a}_{\mu\nu}T^{a}\) evaluated at the centre of the plaquette — this is the finite version of Equation (106.64), and the commutator term of Equation (106.64) arises here from the non-commutativity of the four link matrices. Expanding,

\[ 1-\frac{1}{N}\Re\tr U_{p} =\frac{1}{2N}\tr\!\left(a^{2}\mathcal{F}_{\mu\nu}\right)^{2} +O(a^{6}) =\frac{a^{4}}{4N}\, \mathcal{F}^{a}_{\mu\nu}\mathcal{F}^{a}_{\mu\nu}+O(a^{6})\ec \]

with no sum over \(\mu,\nu\), using \(\tr(T^{a}T^{b})=\frac{1}{2}\delta^{ab}\). The sum over plaquettes is a sum over sites and over \(\mu<\nu\), which becomes \(a^{-4}\int\dd^{4}x_{E}\,\frac{1}{2}\sum_{\mu\nu}\); hence \(S_{W}\to(\beta/8N)\int\dd^{4}x_{E}\, \mathcal{F}^{a}_{\mu\nu}\mathcal{F}^{a}_{\mu\nu}\), and matching to Equation (106.82) gives \(\beta=2N/\hat g^{2}\).

Remark 106.69 (A regulator, and a critical continuum limit).

Momenta on the lattice are bounded, \(\abs{p_{\mu}}\leq\pi\hbar/a\), so the theory is finite: the spacing is an ultraviolet cutoff and no divergence occurs at any order. Recovering continuum physics requires the physical correlation length to be large in lattice units, \(\xi/a\to\infty\), and by Equation (106.25) \(\xi=\hbar/mc\), so \(a\to0\) at fixed physical mass means \(ma c/\hbar\to0\): in the language of Table 106.1, the continuum limit is a second-order critical point of the four-dimensional statistical system defined by Equation (106.81). Asymptotic freedom says where that point is — at \(\hat g\to0\), i.e. \(\beta\to\infty\) — and the rate of approach is fixed by the beta function of The Renormalization Group, which is what makes the extrapolation controlled rather than a guess.

Confinement as an area law

Definition 106.70 (Wilson loop).

For a closed contour \(C\) of links, \(W(C)=\frac{1}{N}\avg{\tr\prod_{\ell\in C}U_{\ell}}\), the continuum form being the trace of the path-ordered exponential \(\frac{1}{N}\avg{\tr P\exp\bigl(\ii\oint_{C} \mathcal{A}_{\mu}\dd x^{\mu}\bigr)}\).

Proposition 106.71 (The Wilson loop measures the static potential).

For a rectangular contour of spatial extent \(R\) and Euclidean time extent \(\tau\), with \(\tau\gg R/c\),

\begin{equation}\tag{106.83} W(R,\tau)\;\sim\;\exp\!\left[-\frac{V(R)\,\tau}{\hbar}\right]\ec \end{equation}

where \(V(R)\) is the energy of a static quark–antiquark pair at separation \(R\). An area law, \(W\sim\exp[-\sigma R c\tau/(\hbar c)]\), is therefore equivalent to a linearly rising potential \(V(R)=\sigma R\), i.e. to confinement; a perimeter law corresponds to a potential approaching a constant, i.e. to a screened, deconfined phase. Rests on Definition 106.70 and Proposition 106.17.

Derivation. Derives Proposition 106.71. The two timelike sides of the rectangle are the parallel transporters that carry a static colour source forward in Euclidean time; the two spacelike sides create and annihilate the pair. By Proposition 106.17 the correlator of a creation and an annihilation operator separated by Euclidean time \(\tau\) decays as \(\ee^{-E_{0}\tau/\hbar}\), with \(E_{0}\) the lowest energy in the sector created — here the energy of the static pair, \(V(R)\), measured from the vacuum. Substituting \(A=Rc\tau\) for the enclosed area converts an exponential in the area into an exponential in \(V\tau\) with \(V=\sigma R\). Dimensions: \(\sigma\) carries \(\mathrm{N}\), so \(\sigma A/(\hbar c)\) is dimensionless.

Theorem 106.72 (Confinement at strong coupling).

For \(\beta\ll1\) the Wilson loop of a planar rectangle obeys an area law. For \(\SU(N)\) with \(N\geq3\),

\begin{equation}\tag{106.84} W(C)=\left(\frac{\beta}{2N^{2}}\right)^{A/a^{2}} \bigl[1+O(\beta)\bigr]\ec\qquad \sigma=\frac{\hbar c}{a^{2}}\, \ln\!\frac{2N^{2}}{\beta}\ec \end{equation}

\(A\) being the minimal area spanned by \(C\). For \(\SU(2)\) every plaquette contributes twice as much,

\begin{equation}\tag{106.85} W(C)=\left(\frac{\beta}{4}\right)^{A/a^{2}} \bigl[1+O(\beta)\bigr]\ec\qquad \sigma=\frac{\hbar c}{a^{2}}\,\ln\!\frac{4}{\beta}\ec \end{equation}

and not the \(\beta/2N^{2}=\beta/8\) that Equation (106.84) would give: the fundamental representation of \(\SU(2)\) is pseudo-real, so \(U\) and \(U^{\dagger}\) are not independent and the two orientations of a plaquette supply the same term rather than two different ones. The distinction is not academic here — \(\SU(2)\) is the group of the measurement quoted in Section 106.5.2. Rests on Equation (106.81), Equation (106.83) and Theorem 5.153.

Derivation. Derives Theorem 106.72. Expand \(\ee^{-S_{W}}\) in powers of \(\beta\): for each plaquette, \(\exp\bigl[(\beta/N)\Re\tr U_{p}\bigr] =1+\frac{\beta}{2N}\bigl(\tr U_{p}+\tr U_{p}^{\dagger}\bigr)+O(\beta^{2})\). The integration is over independent Haar measures on each link, and the two identities that are needed are

\[ \int\dd U\;U_{ij}=0\ec\qquad \int\dd U\;U_{ij}\bigl(U^{\dagger}\bigr)_{kl} =\frac{1}{N}\delta_{il}\delta_{jk}\ec \]

the first because the integrand transforms in a non-trivial representation and the Haar measure is invariant, the second by Schur's first lemma (Theorem 5.153) applied to the invariant tensor. Every link of the contour must therefore be paired with a link of opposite orientation supplied by some plaquette; the cheapest way to do this is to tile the minimal surface spanned by \(C\) with \(A/a^{2}\) plaquettes, each bringing \(\beta/2N\) from the expansion and each internal link contributing \(1/N\) from the second identity. The single-plaquette case is exact and does the counting transparently:

\[ \avg{\tfrac{1}{N}\tr U_{p}} =\frac{\beta}{2N}\cdot\frac{1}{N} \int\dd U\;\tr U\,\tr U^{\dagger}+O(\beta^{2}) =\frac{\beta}{2N^{2}}+O(\beta^{2})\ec \]

since \(\int\dd U\,\tr U\tr U^{\dagger} =\sum_{ij}\delta_{ij}\delta_{ij}/N=1\). Iterating over the tiling gives Equation (106.84), and comparing with Equation (106.83) gives \(\sigma\).

That counting used \(\int\dd U\,U_{ij}U_{kl}=0\), i.e.\ \(\int\dd U\,(\tr U)^{2}=0\), which holds only when \(\vect{N}\otimes\vect{N}\) contains no singlet. For \(N\geq3\) it splits into the symmetric and antisymmetric two-index representations and neither is trivial, so the identity holds. For \(\SU(2)\), \(\vect{2}\otimes\vect{2}=\vect{3}\oplus\vect{1}\) does contain the singlet, \(\tr U\) is real, and orthogonality of characters gives \(\int\dd U\,(\tr U)^{2}=1\). The plaquette factor is then \((\beta/2N)(\tr U_{p}+\tr U_{p}^{\dagger})=\tfrac{\beta}{2}\tr U_{p}\) and

\[ \avg{\tfrac{1}{2}\tr U_{p}} =\frac{1}{2}\cdot\frac{\beta}{2}\int\dd U\;(\tr U)^{2}+O(\beta^{2}) =\frac{\beta}{4}+O(\beta^{2})\ec \]

which is Equation (106.85): twice the value the generic formula assigns, hence \(\ln(4/\beta)\) in place of \(\ln(8/\beta)\) in the string tension.

Remark 106.73 (What the strong-coupling result does and does not prove).

Theorem 106.72 is a genuine derivation of confinement — of an area law with a computable string tension — but for the lattice theory at large \(\hat g\), not for the continuum theory. The continuum limit lives at \(\beta\to\infty\) (Remark 106.69), which is the opposite end, and the expansion Equation (106.84) says nothing about whether the area law survives to it. Worse, the expansion converges for compact \(\U(1)\) too, and there confinement is known not to survive: the theory has a phase transition at intermediate \(\beta\) and its weak-coupling phase is ordinary massless electrodynamics. Whether \(\SU(3)\) has such a transition — and the evidence is that it does not — is a question that only a numerical computation covering the whole range of \(\beta\) can answer, which is Section 106.5.2. The confinement of quarks remains, in the strict sense, unproved; what exists is a definition of the theory in which the question is well posed, and overwhelming numerical evidence about the answer.

For orientation, the customary phenomenological value of the string tension, obtained by matching the linear potential to the slope of the observed Regge trajectories, is \(\sqrt{\sigma\hbar c}\approx0.44\,\mathrm{GeV}\), i.e.

\begin{equation}\tag{106.86} \sigma=\frac{(0.44\,\mathrm{GeV})^{2}}{\hbar c} \approx1.6\times 10^{5}\,\mathrm{N}\ec \end{equation}

about the weight of sixteen tonnes, independent of separation. That number is the reason no free quark has been seen: separating a pair by one femtometre costs \(\sigma\times1\,\mathrm{fm} =1.6\times 10^{-10}\,\mathrm{J}\approx1\,\mathrm{GeV}\), and by one metre would cost \(10^{15}\) times more — long before which it is cheaper to create a quark–antiquark pair out of the vacuum and break the string.

Fermions on a lattice

Proposition 106.74 (Naive discretization doubles the species sixteen times).

Replacing \(\gamma^{\mu}\pp_{\mu}\) by the symmetric difference \(\gamma^{\mu}\bigl[\psi(n+\hat\mu)-\psi(n-\hat\mu)\bigr]/2a\) gives a momentum-space operator whose massless part is

\begin{equation}\tag{106.87} \widetilde{D}(p)=\frac{\ii}{a}\sum_{\mu}\gamma^{\mu} \sin\!\left(\frac{ap_{\mu}}{\hbar}\right)\ec \end{equation}

which vanishes not only at \(p_{\mu}=0\) but at every corner of the Brillouin zone with \(ap_{\mu}/\hbar\in\set{0,\pi}\): \(2^{4}=16\) species where one was intended. Rests on Definition 106.66 and Equation (100.5).

Derivation. Derives Proposition 106.74. Fourier transform the difference operator: the shift \(\psi(n\pm\hat\mu)\) becomes \(\ee^{\pm\ii ap_{\mu}/\hbar}\psi\), so the symmetric difference becomes \(\ii\sin(ap_{\mu}/\hbar)/a\), giving Equation (106.87). The propagator is \(\widetilde{D}^{-1}\), and a pole occurs wherever \(\widetilde{D}\) vanishes; \(\sin(ap_{\mu}/\hbar)=0\) has two solutions in the Brillouin zone \(-\pi<ap_{\mu}/\hbar\leq\pi\) for each of the four directions, hence \(2^{4}\) poles, each of which is a physical particle in the continuum limit.

Theorem 106.75 (Nielsen–Ninomiya).

No lattice Dirac operator can simultaneously (i) be local, (ii) possess the continuum chiral symmetry \(\acomm{D}{\gamma^{5}}=0\), (iii) have the correct continuum limit, and (iv) describe a single species. One of the four must be given up [Nielsen:1981]. Rests on Proposition 106.74.

The choices actually made are: Wilson's, which adds \(-\frac{ra}{2}\pp^{2}\) to the operator, lifting the fifteen doublers to masses of order \(\hbar/(ac)\) at the price of breaking chiral symmetry explicitly at finite \(a\) — the symmetry is recovered in the continuum limit, but only after an additive renormalization of the quark mass; and the staggered construction, which distributes the spinor components over the corners of a hypercube, retaining a remnant of chiral symmetry and reducing the doubling from sixteen to four. The computation of Section 106.5.2 uses the Wilson operator with smeared links [Duerr:2008]. That a theorem of this kind exists at all is the sharpest illustration of the point made after Proposition 106.9: a discretization is a definition, and a symmetry of the classical Lagrangian need not survive it. Section 106.7.3 shows that for the axial symmetry the failure is not an artefact of any discretization but a property of the continuum theory itself, with a measured consequence.

Derivation pending.

The Nielsen–Ninomiya theorem: the fermion doubling of the naive operator is derived above, but the no-go theorem itself — that no local, chirally symmetric, doubler-free operator exists, proved by counting the zeros of the energy–momentum relation as a topological invariant of the Brillouin torus — is stated and not proved here. It belongs in Appendix A.

Monte Carlo evidence

Why sampling, and why it works

A lattice of \(64^{3}\times128\) sites carries \(33554432\) sites, \(134217728\) links and, for \(\SU(3)\) with eight parameters per link, about \(1.1\times10^{9}\) real integration variables. No deterministic quadrature can touch an integral of that dimension: a product rule with \(k\) points per axis costs \(k^{d}\) evaluations. Monte Carlo integration does not degrade with dimension at all.

Proposition 106.76 (Importance sampling).

Let configurations \(\set{U^{(i)}}_{i=1}^{M}\) be drawn with probability proportional to \(\ee^{-S_{E}[U]/\hbar}\). Then

\begin{equation}\tag{106.88} \avg{\mathcal{O}} =\frac{\int\mathcal{D}U\;\mathcal{O}[U]\,\ee^{-S_{E}/\hbar}} {\int\mathcal{D}U\;\ee^{-S_{E}/\hbar}} =\lim_{M\to\infty}\frac{1}{M}\sum_{i=1}^{M} \mathcal{O}\bigl[U^{(i)}\bigr]\ec \end{equation}

with a statistical error falling as \(M^{-1/2}\) independently of the dimension of the integral, by the central limit theorem. Rests on Proposition 106.13 and Theorem 106.68.

Derivation. Derives Proposition 106.76. The estimator is the sample mean of a random variable whose expectation is \(\avg{\mathcal{O}}\) by construction, so it is unbiased; its variance is \(\mathrm{Var}(\mathcal{O})/M\), and the central limit theorem gives the distribution of the error. Nothing in either statement refers to the dimension. What replaces the dimension in the cost is the autocorrelation of the algorithm that generates the sample, since successive configurations are not independent.

Two facts make Equation (106.88) legitimate, and both were established earlier in this chapter rather than assumed. Proposition 106.13 makes the weight real and positive, so that “probability proportional to \(\ee^{-S_{E}/\hbar}\)” means something; and Remark 106.40 explains why the fermion determinant is included as \(\det^{2}\) for the degenerate light pair, which keeps it positive. Where positivity fails — at non-zero baryon chemical potential — the method fails with it (Remark 106.15).

Confinement and asymptotic scaling

Creutz's computation [Creutz:1980b] was the first to show that the lattice theory and the continuum theory are the same theory. He extracted the string tension of \(\SU(2)\) gauge theory from ratios of Wilson loops,

\begin{equation}\tag{106.89} \chi(I,J)=-\ln\frac{W(I,J)\,W(I-1,J-1)} {W(I-1,J)\,W(I,J-1)}\ec \end{equation}

a combination in which the perimeter and corner contributions cancel and only the area term survives, so that \(\chi\to\sigma a^{2}/(\hbar c)\) for large loops. Measuring \(\chi\) over a range of \(\beta\), he found that it followed the asymptotic-freedom scaling law

\begin{equation}\tag{106.90} \sqrt{\chi}\;\propto\; \exp\!\left[-\frac{4\pi^{2}\beta}{Nb_{0}}\right]\ec\qquad b_{0}=\frac{11N}{3}-\frac{2n_{f}}{3}\ec \end{equation}

in which \(b_{0}\) is the one-loop beta-function coefficient in the convention \(\mu\,\dd\hat g/\dd\mu=-b_{0}\hat g^{3}/16\pi^{2}\) — \(b_{0}=22/3\) for the pure \(\SU(2)\) theory Creutz simulated, so that the exponent is \(-3\pi^{2}\beta/11\). The exponent follows from \(a\Lambda=\exp[-8\pi^{2}/(b_{0}\hat g^{2})]\) with \(\hat g^{2}=2N/\beta\), and the scaling law itself is the prediction of the renormalization group of The Renormalization Group — that is, that the dimensionless lattice quantity ran with \(\beta\) exactly as a physical quantity held fixed in units of a continuum scale must. Confinement at strong coupling was already known from Theorem 106.72; what Creutz showed is that it persists into the region where the continuum limit is taken.

The hadron spectrum from three inputs

Phenomenon 106.77 (The light hadron spectrum is a prediction of QCD).

With the three free parameters of the light-quark sector fixed by three measured masses — those of the pion, the kaon and one baryon, which set the light-quark mass, the strange-quark mass and the overall scale — the Euclidean lattice path integral reproduces the masses of the remaining light hadrons, including the nucleon, to a few percent [Duerr:2008]. Nothing else is put in: no potential, no constituent masses, no fitted form factors. Rests on Equation (106.24), Corollary 106.49 and Proposition 106.17.

Derivation. Derives Phenomenon 106.77. The procedure is Equation (106.24) applied to a lattice correlator. Choose an interpolating operator \(\mathcal{O}\) with the quantum numbers of the hadron — for the nucleon, a colour-singlet product of three quark fields; Corollary 106.49 guarantees that the choice does not affect the answer. Measure

\[ C(\tau)=\avg{\mathcal{O}(\tau)\,\mathcal{O}^{\dagger}(0)} \;\underset{\tau\to\infty}{\longrightarrow}\; \abs{\bra{0}\mathcal{O}\ket{h}}^{2}\, \ee^{-m_{h}c^{2}\tau/\hbar}\ec \]

by Proposition 106.17, and read the mass off the exponential decay rate. Repeat at several lattice spacings, volumes and quark masses, and extrapolate.

The counting of inputs is what makes this a prediction rather than a fit. The Lagrangian of the light sector contains exactly three parameters that are not fixed by symmetry: \(\hat g^{2}\), the degenerate up–down mass, and the strange mass. Measuring \(m_{\pi}\), \(m_{K}\) and one baryon mass fixes all three (the coupling enters only through the overall scale, by dimensional transmutation — The Renormalization Group). Every other hadron mass is then output. The masses to be reproduced include, from the shipped Particle Data Group compilation [Navas:2024], \(m_{p}c^{2}=938.272\,\mathrm{MeV}\), \(m_{n}c^{2}=939.565\,\mathrm{MeV}\), \(m_{\Lambda}c^{2}=1115.683\,\mathrm{MeV}\) and \(m_{\Delta}c^{2}=1232.0(2.0)\,\mathrm{MeV}\), together with the \(\Sigma\), \(\Xi\), \(\Sigma^{*}\), \(\Xi^{*}\) and \(\Omega\); the computed values agree with all of them within the combined statistical and systematic uncertainties, which are at the level of a few percent [Duerr:2008].

This is the strongest single piece of evidence that the quantization scheme of this chapter describes Nature and not merely a formal expansion. The perturbative successes of Quantum Electrodynamics and Renormalization test the theory where the coupling is small and the series is asymptotic; here the coupling is of order unity, no series is involved, and the answer still comes out right. It is also the answer to Remark 106.73: no phase transition intervenes between the strong-coupling region and the continuum limit, because if one did the spectrum would not come out.

Remark 106.78 (The systematics, and how they are quantified).

A lattice result is not a measurement and its error budget is not statistical alone. Three systematic effects dominate, and each is controlled by computing at several values of the corresponding parameter and extrapolating rather than by estimating.

Finite spacing. The discretization error vanishes as a power of \(a\) — \(O(a)\) for the unimproved Wilson operator, \(O(a^{2})\) after improvement. Computations are repeated at three or more spacings, in the range \(a\approx0.065\text{–}0.125\,\mathrm{fm}\) for [Duerr:2008], and extrapolated to \(a=0\).

Finite volume. A box of side \(L\) with periodic boundaries distorts a hadron by the exchange of the lightest particle around the box, an effect falling as \(\ee^{-m_{\pi}cL/\hbar}\). The working rule is \(m_{\pi}cL/\hbar\gtrsim4\), i.e. \(L\gtrsim4\hbar/(m_{\pi}c)\), four reduced Compton wavelengths of the pion. At the physical charged-pion mass \(m_{\pi}c^{2}=139.57039\,\mathrm{MeV}\) [Navas:2024] that wavelength is \(\hbar c/(m_{\pi}c^{2})=1.414\,\mathrm{fm}\), so the rule demands \(L\gtrsim5.7\,\mathrm{fm}\) — a demanding box, which is why the requirement is usually met at a pion mass above the physical one, where the same rule costs less: at \(m_{\pi}c^{2}=190\,\mathrm{MeV}\) it asks only for \(L\gtrsim4.2\,\mathrm{fm}\). The dependence is checked by running at two volumes.

Quark mass. Simulating at the physical light-quark mass is expensive because the algorithms slow down as the pion becomes light, so computations are done at larger masses and extrapolated using the functional form supplied by chiral perturbation theory. This was historically the largest uncertainty, and it is the one that has shrunk most as machines have grown.

Note that these are all controlled extrapolations in parameters the computation itself can vary, which is what distinguishes them from the uncontrolled quenched approximation of Remark 106.40. It is that difference, and not raw computer power, that turned lattice QCD from an illustration into a source of numbers with error bars.

Topology of the gauge vacuum

Instantons

The Euclidean action of Equation (106.82) has non-trivial finite-action stationary points, found by Belavin, Polyakov, Schwartz and Tyupkin [Belavin:1975]. They contribute to the path integral factors of the form \(\ee^{-S_{E}/\hbar}\) which are invisible to every order of perturbation theory, and they are the mechanism by which the vacuum acquires structure.

Definition 106.79 (Topological charge).

For a Euclidean gauge field of finite action define the dual \(\widetilde{\mathcal{F}}^{a}_{\mu\nu} =\frac{1}{2}\epsilon_{\mu\nu\rho\sigma} \mathcal{F}^{a}_{\rho\sigma}\) and

\begin{equation}\tag{106.91} \nu:=\frac{1}{32\pi^{2}}\int\dd^{4}x_{E}\; \mathcal{F}^{a}_{\mu\nu}\widetilde{\mathcal{F}}^{a}_{\mu\nu} =\frac{1}{32\pi^{2}}\int\dd^{4}x_{E}\; \epsilon_{\mu\nu\rho\sigma}\, \tr\!\left(\mathcal{F}_{\mu\nu}\mathcal{F}_{\rho\sigma}\right)\ec \end{equation}

dimensionless, since \(\mathcal{F}\) carries \(/\mathrm{m}^{2}\) by Equation (106.64). This is the quantity denoted \(n\) in Equation (105.109), and the corresponding density \(q(x)=(c/32\pi^{2})\epsilon_{\mu\nu\rho\sigma} \tr(\mathcal{F}_{\mu\nu}\mathcal{F}_{\rho\sigma})\), of dimension \(/\mathrm{m}^{3}/\mathrm{s}\), is the \(q\) of Definition 105.43.

Theorem 106.80 (Bogomolny bound and self-duality).

For any Euclidean configuration,

\begin{equation}\tag{106.92} S_{E}\geq\frac{8\pi^{2}\hbar}{\hat g^{2}}\,\abs{\nu}\ec \end{equation}

with equality if and only if the field is self-dual (\(\mathcal{F}=\widetilde{\mathcal{F}}\), for \(\nu>0\)) or anti-self-dual (\(\mathcal{F}=-\widetilde{\mathcal{F}}\), for \(\nu<0\)). A configuration saturating the bound with \(\nu=1\) is an instanton, and its contribution to the path integral is weighted by

\begin{equation}\tag{106.93} \ee^{-S_{E}/\hbar} =\exp\!\left(-\frac{8\pi^{2}}{\hat g^{2}}\right) =\exp\!\left(-\frac{2\pi}{\alpha_{s}}\right)\ec \end{equation}

which at \(\alpha_{s}=0.1180\) (Equation (102.47)) is \(8\times10^{-24}\) — a number with no expansion in powers of \(\alpha_{s}\) at all. Rests on Equations (106.82) and (106.91).

Derivation. Derives Theorem 106.80. The Euclidean action is \(S_{E}=(\hbar/4\hat g^{2})\int\dd^{4}x_{E}\, \mathcal{F}^{a}_{\mu\nu}\mathcal{F}^{a}_{\mu\nu}\) by Equation (106.82), and the integrand is a sum of squares, hence non-negative. Consider

\[ 0\leq\int\dd^{4}x_{E}\; \left(\mathcal{F}^{a}_{\mu\nu} \mp\widetilde{\mathcal{F}}^{a}_{\mu\nu}\right)^{2} =2\int\dd^{4}x_{E}\; \mathcal{F}^{a}_{\mu\nu}\mathcal{F}^{a}_{\mu\nu} \mp2\int\dd^{4}x_{E}\; \mathcal{F}^{a}_{\mu\nu}\widetilde{\mathcal{F}}^{a}_{\mu\nu}\ec \]

using \(\widetilde{\mathcal{F}}\widetilde{\mathcal{F}} =\mathcal{F}\mathcal{F}\), which follows from \(\epsilon_{\mu\nu\rho\sigma}\epsilon_{\mu\nu\alpha\beta} =2(\delta_{\rho\alpha}\delta_{\sigma\beta} -\delta_{\rho\beta}\delta_{\sigma\alpha})\) in four Euclidean dimensions. Hence \(\int\mathcal{F}\mathcal{F}\geq \abs{\int\mathcal{F}\widetilde{\mathcal{F}}}=32\pi^{2}\abs{\nu}\) by Equation (106.91), with equality iff \(\mathcal{F}=\pm\widetilde{\mathcal{F}}\). Multiplying by \(\hbar/4\hat g^{2}\) gives Equation (106.92). A self-dual field automatically solves the Euclidean field equations, since \(D_{\mu}\mathcal{F}_{\mu\nu} =D_{\mu}\widetilde{\mathcal{F}}_{\mu\nu}=0\) by the Bianchi identity; self-duality, a first-order condition, therefore replaces the second-order equations of motion.

Numerically, \(2\pi/0.1180=53.24\) and \(\ee^{-53.24}=8\times10^{-24}\).

Proposition 106.81 (The topological charge is an integer).

Finiteness of the action forces \(\mathcal{F}\to0\) at infinity, hence \(\mathcal{A}_{\mu}\to-\ii\,g^{-1}\pp_{\mu}g\) for some map \(g:S^{3}\to G\) on the sphere at infinity. Then

\begin{equation}\tag{106.94} \nu=\frac{1}{24\pi^{2}}\int_{S^{3}}\dd^{3}S_{\mu}\; \epsilon_{\mu\nu\rho\sigma}\, \tr\!\left(g^{-1}\pp_{\nu}g\; g^{-1}\pp_{\rho}g\;g^{-1}\pp_{\sigma}g\right)\ec \end{equation}

which is the degree of \(g\) as a map \(S^{3}\to S^{3}\) — for \(G=\SU(2)\), whose manifold is \(S^{3}\) — and is therefore an integer. Rests on Equation (106.91), Equation (14.155) and Theorem 14.111.

Derivation. Derives Proposition 106.81. The integrand of Equation (106.91) is a total divergence, \(\epsilon\tr(\mathcal{F}\mathcal{F})=\pp_{\mu}K_{\mu}\) with the Chern–Simons current \(K_{\mu}=2\epsilon_{\mu\nu\rho\sigma} \tr\bigl(\mathcal{A}_{\nu}\pp_{\rho}\mathcal{A}_{\sigma} +\frac{2\ii}{3}\mathcal{A}_{\nu}\mathcal{A}_{\rho} \mathcal{A}_{\sigma}\bigr)\); this is the transgression form of the second Chern class, and it is Equation (14.155) of Theorem 14.111 specialized to \(r=2\). The volume integral is therefore a surface integral over the sphere at infinity, where \(\mathcal{A}\) is pure gauge; substituting \(\mathcal{A}_{\mu}=-\ii g^{-1}\pp_{\mu}g\) makes the first term of \(K_{\mu}\) combine with the second into Equation (106.94).

That the resulting number is an integer is the degree theorem. Let \(y\) be a regular value of the smooth map \(g:S^{3}\to S^{3}\); its preimage is a finite set of points, and Equation (106.94) evaluates to the sum over those points of the sign of the Jacobian determinant of \(g\), i.e. to the number of times \(g\) covers the target counted with orientation. That count is an integer, and it is unchanged under continuous deformation of \(g\), since the preimages can only be created or destroyed in pairs of opposite sign. The map \(g(x)=\bigl(x_{4}+\ii\,\vect{x}\cdot\vect{\tau}\bigr)/\abs{x}\), with \(\vect{\tau}\) the Pauli matrices, covers \(\SU(2)\cong S^{3}\) exactly once and has \(\nu=1\): it is the identity map, whose Jacobian is everywhere positive.

Remark 106.82 (The dilute-gas approximation and its failure).

Instantons come in a family: the classical equations are scale invariant, so a solution of size \(\rho\) costs the same action for every \(\rho\). Integrating over the collective coordinates — position, size and global colour orientation — requires the determinant of the fluctuation operator about the instanton, computed by 't Hooft [tHooft:1976b]. The result is a density of instantons per unit four-volume of the form

\[ \dd n\;\propto\;\frac{\dd\rho}{\rho^{5}}\; \left(\frac{1}{\hat g^{2}(\rho)}\right)^{2N} \exp\!\left(-\frac{8\pi^{2}}{\hat g^{2}(\rho)}\right)\ec \]

in which the running of \(\hat g\) with the scale \(1/\rho\) turns the exponential into a power, \(\rho^{b_{0}}\) with \(b_{0}=11N/3-2n_{f}/3\) the one-loop beta-function coefficient in the same convention as Equation (106.90), so that \(8\pi^{2}/\hat g^{2}(\rho)=b_{0}\ln\bigl(1/\rho\Lambda\bigr)\). For \(\SU(3)\) with three light flavours \(b_{0}=9\), so the density behaves as \(\rho^{4}\dd\rho\): the integral is dominated by large instantons, where the coupling is strong and the semiclassical approximation that produced the formula has failed. The honest conclusion is that the dilute-gas approximation is not a controlled calculation in QCD. What survives is the qualitative structure — the vacuum has topological sectors, they are populated, and the population is a non-perturbative effect — together with the one place where the counting is protected by an index theorem, which is Section 106.7.3. The quantitative statements about the topological susceptibility of the QCD vacuum come from the lattice, not from instanton calculus.

The physical use of all this in Quantum Chromodynamics is the mass of the \(\eta'\). The nine light pseudoscalars would be the Goldstone bosons of a spontaneously broken \(\U(3)_{L}\times\U(3)_{R}\), and eight of them — the pions, kaons and \(\eta\) — are light for that reason. The ninth is not: the singlet axial symmetry is broken by the anomaly of Section 106.7.1, and the topological configurations of this subsection are what make the anomaly have consequences rather than being a total derivative one can discard. That is the resolution of what was called the \(\U(1)_{A}\) problem, and it is the reason the \(\eta'\) is heavier than the nucleon-scale expectation for a Goldstone boson.

Derivation pending.

The instanton fluctuation determinant: the Bogomolny bound, the integrality of the winding number and the exponential weight are derived above, but the explicit self-dual solution of Belavin, Polyakov, Schwartz and Tyupkin and the one-loop determinant about it, which supply the measure over collective coordinates quoted in the dilute-gas remark, are not computed here. They belong in Appendix A.

Theta vacua

Proposition 106.83 (The vacuum is a superposition of winding sectors).

Let \(\ket{n}\) denote the state built on the classical vacuum of Chern–Simons number \(n\in\Z\), and let \(T\) be the operator implementing a large gauge transformation, \(T\ket{n}=\ket{n+1}\). Since \(T\) commutes with the Hamiltonian and with all local observables, the energy eigenstates are simultaneous eigenstates of \(T\):

\begin{equation}\tag{106.95} \ket{\theta}=\sum_{n\in\Z}\ee^{\ii n\theta}\ket{n}\ec\qquad T\ket{\theta}=\ee^{-\ii\theta}\ket{\theta}\ec \end{equation}

with \(\theta\in[0,2\pi)\) a superselection parameter: no local operator connects different \(\theta\) [Callan:1976] [Jackiw:1976]. Rests on Proposition 106.81.

Derivation. Derives Proposition 106.83. Gauge transformations that tend to the identity at infinity are generated by the Gauss constraint and act trivially on physical states. Those that tend to a non-trivial element — the large transformations, classified by Proposition 106.81 — are symmetries of the theory but are not required to act trivially, exactly as a translation by a lattice vector in a crystal is a symmetry that acts non-trivially on Bloch states. \(T\) is unitary and commutes with \(H\), so \(H\) and \(T\) may be diagonalized together; the eigenvalues of a unitary operator with \(T\ket{n}=\ket{n+1}\) are phases \(\ee^{-\ii\theta}\), and the corresponding eigenvector is Equation (106.95). Because a local operator changes \(n\) by a finite amount, its matrix elements between \(\ket{\theta}\) and \(\ket{\theta'}\) are proportional to \(\delta(\theta-\theta')\): the \(\theta\) sectors are superselected and \(\theta\) is a parameter of the theory, not a state to be chosen.

Corollary 106.84 (The theta term).

Working in the \(\theta\) vacuum is equivalent to adding to the Lagrangian density

\begin{equation}\tag{106.96} \Lag_{\theta}=\theta\hbar\,q(x) =\frac{\theta\hbar c}{32\pi^{2}}\, \epsilon^{\mu\nu\rho\sigma}\, \tr\!\left(\mathcal{F}_{\mu\nu}\mathcal{F}_{\rho\sigma}\right)\ec \end{equation}

which is Equation (105.110), and which contributes the phase \(\ee^{\ii\theta\nu}\) to the weight of a configuration of winding number \(\nu\). Rests on Equation (106.95) and Definition 106.79.

Derivation. Derives Corollary 106.84. Summing over sectors with the weights of Equation (106.95) inserts \(\ee^{\ii n\theta}\) into the path integral for each topological class \(n\). By Definition 106.79 that class is labelled by \(\nu=\int\dd t\,\dd^{3}x\;q\), so the insertion is \(\exp\bigl(\ii\theta\int\dd t\,\dd^{3}x\,q\bigr) =\exp\bigl(\tfrac{\ii}{\hbar}\cdot\tfrac{1}{c}\int\dd^{4}x\; \theta\hbar q\bigr)\), which is the addition of Equation (106.96) to \(\Lag\). That \(\Lag_{\theta}\) has the dimension of an energy density follows from \([q]=/\mathrm{m}^{3}/\mathrm{s}\) and \([\hbar] =\mathrm{J}\,\mathrm{s}\).

Three consequences follow, and each is taken up elsewhere in this part rather than repeated here.

Baryon number is not conserved. The baryon and lepton currents of the Standard Model are anomalous in the sense of Section 106.7.1, with the topological density of Equation (106.91) of the \(\SU(2)_{L}\) field on the right-hand side. Integrating gives \(\Delta B=\Delta L=n_{f}\nu\) with \(n_{f}=3\) generations, which is Equation (105.103); \(B-L\) is conserved. At zero temperature the transition rate carries the factor Equation (106.93) with \(\alpha_{W}\) in place of \(\alpha_{s}\) and is utterly negligible, but above the electroweak scale the barrier — the sphaleron of Equation (105.104) — can be crossed thermally and the rate is unsuppressed [tHooft:1976a] [Kuzmin:1985]. This is how the Standard Model satisfies the first of Sakharov's conditions, Theorem 105.41.

The theta term violates \(P\) and \(CP\). That Equation (106.96) is \(P\)-odd and \(T\)-odd is Proposition 105.44; that only the combination \(\bar\theta=\theta+\arg\det M_{q}\) is physical, because a chiral rotation of the quark fields shifts \(\theta\) through the non-invariance of the measure, is Equation (105.111) — and the mechanism of that shift is precisely Section 106.7.3. The experimental bound from the neutron electric dipole moment, \(\abs{\bar\theta}\lesssim10^{-10}\), is Equation (105.113) [Baker:2006] [Abel:2020]. The strong \(CP\) problem is the observation that a parameter which could have been anything is measured to be this small, and it is recorded as an open question in Section 105.7. The best-known proposed resolution promotes \(\bar\theta\) to a dynamical field [Peccei:1977] [Wilczek:1978]; the particle it predicts has not been observed, and this book records it as a proposal, not as physics.

The \(\theta\) dependence is invisible in perturbation theory. \(q\) is a total derivative, so Equation (106.96) contributes nothing to any Feynman diagram; its effects exist only because the boundary term is non-zero on configurations of non-trivial topology. A parameter that is invisible perturbatively and physical non-perturbatively is the worst possible combination for anyone hoping to argue it away, and that is the whole difficulty.

Anomalies

The Adler–Bell–Jackiw anomaly

A classical conservation law

Proposition 106.85 (Classical conservation of the axial current).

For a Dirac field of mass \(m\) coupled to an external gauge field, the axial current

\begin{equation}\tag{106.97} j_{5}^{\mu}=c\,\bar\psi\gamma^{\mu}\gamma^{5}\psi\ec\qquad [j_{5}^{\mu}]=/\mathrm{m}^{2}/\mathrm{s}\ec \end{equation}

satisfies, by the classical field equations,

\begin{equation}\tag{106.98} \pp_{\mu}j_{5}^{\mu} =\frac{2\ii mc^{2}}{\hbar}\,\bar\psi\gamma^{5}\psi\ec \end{equation}

which vanishes identically for \(m=0\): the axial charge of a massless Dirac field is conserved. Rests on Equation (100.5).

Derivation. Derives Proposition 106.85. The Dirac equation in an external field is \((\ii\hbar\gamma^{\mu}D_{\mu}-mc)\psi=0\), and its conjugate reads \(\ii\hbar\,(D_{\mu}\bar\psi)\gamma^{\mu}+mc\,\bar\psi=0\), the derivative acting on \(\bar\psi\). Then

\[ \pp_{\mu}\bigl(\bar\psi\gamma^{\mu}\gamma^{5}\psi\bigr) =\bigl(D_{\mu}\bar\psi\bigr)\gamma^{\mu}\gamma^{5}\psi +\bar\psi\gamma^{\mu}\gamma^{5}\bigl(D_{\mu}\psi\bigr) =\frac{\ii mc}{\hbar}\bar\psi\gamma^{5}\psi +\frac{\ii mc}{\hbar}\bar\psi\gamma^{5}\psi\ec \]

using \(\acomm{\gamma^{\mu}}{\gamma^{5}}=0\) to move \(\gamma^{5}\) past \(\gamma^{\mu}\) in the second term, which flips the sign of the mass contribution and makes the two add rather than cancel. Multiplying by \(c\) gives Equation (106.98). The gauge field drops out because the covariant derivatives combine into an ordinary one in a neutral bilinear.

By Noether's theorem [Noether:1918] this conservation law is the statement that the massless action is invariant under the global chiral rotation

\begin{equation}\tag{106.99} \psi\longmapsto\ee^{\ii\alpha\gamma^{5}}\psi\ec\qquad \bar\psi\longmapsto\bar\psi\,\ee^{\ii\alpha\gamma^{5}}\ec \end{equation}

the second transformation carrying \(\ee^{+\ii\alpha\gamma^{5}}\) rather than its inverse because \(\gamma^{5}\) anticommutes with the \(\gamma^{0}\) hidden in \(\bar\psi\). Everything above is classical. What Adler [Adler:1969] and Bell and Jackiw [Bell:1969] found is that Equation (106.98) is false in the quantum theory even at \(m=0\).

Where the failure comes from

The one-loop amplitude with one axial and two vector vertices — the triangle — is linearly divergent, and a linearly divergent integral is not invariant under a shift of its integration variable.

Lemma 106.86 (A shift of a divergent integral leaves a surface term).

For a smooth \(f:\R^{4}\to\C\) and constant \(a\),

\begin{equation}\tag{106.100} \int\dd^{4}\ell\;\bigl[f(\ell+a)-f(\ell)\bigr] =\lim_{R\to\infty}\oint_{\abs{\ell}=R} a_{\mu}\,f(\ell)\;\dd S^{\mu}\ec \end{equation}

which vanishes when \(f\) falls faster than \(\abs{\ell}^{-3}\) and is a finite non-zero constant, linear in \(a\), when \(f\) falls exactly as \(\abs{\ell}^{-3}\) — that is, when the integral is linearly divergent. Rests on Theorem 7.101.

Derivation. Derives Lemma 106.86. Expand \(f(\ell+a)-f(\ell)=a_{\mu}\pp^{\mu}f+O(a^{2})\) — the higher terms give integrals of higher divergences of \(f\) and are handled the same way — and apply the divergence theorem to the ball \(\abs{\ell}\leq R\). The surface element grows as \(R^{3}\), so the limit is zero, finite, or infinite according as \(f\) falls faster than, exactly as, or slower than \(R^{-3}\).

Theorem 106.87 (The axial anomaly).

In the quantum theory of a massless Dirac field of charge \(Q\) in an external electromagnetic field, the axial current is not conserved:

\begin{equation}\tag{106.101} \pp_{\mu}j_{5}^{\mu} =\frac{Q^{2}c}{16\pi^{2}\hbar^{2}}\, \epsilon^{\mu\nu\rho\sigma}F_{\mu\nu}F_{\rho\sigma} =\frac{\alpha}{4\pi\mu_{0}\hbar}\, \epsilon^{\mu\nu\rho\sigma}F_{\mu\nu}F_{\rho\sigma}\ec \end{equation}

the second form for \(Q=e\), with \(\epsilon^{0123}=+1\). Both sides carry \(/\mathrm{m}^{3}/\mathrm{s}\). The vector current \(j^{\mu}=Qc\bar\psi\gamma^{\mu}\psi\) of Equation (100.6) remains exactly conserved. Rests on Proposition 106.85 and Equation (106.99).

The coefficient is derived in Section 106.7.3, where the computation is short and its topological character is manifest. What the triangle diagram contributes is the reason: the anomaly is not a choice but an obstruction.

Proposition 106.88 (The obstruction).

The triangle amplitude \(T^{\mu\nu\rho}\) with one axial and two vector vertices depends on how the loop momentum is routed, by Lemma 106.86. Fixing the routing so that the two vector currents are conserved, \(k_{1\nu}T^{\mu\nu\rho} =k_{2\rho}T^{\mu\nu\rho}=0\), forces \((k_{1}+k_{2})_{\mu}T^{\mu\nu\rho}\neq0\), and conversely. No regulator preserves both. Rests on Lemma 106.86 and Theorem 100.17.

Derivation. Derives Proposition 106.88. The integrand of the triangle falls as \(\abs{\ell}^{-3}\) at large loop momentum, since it is a product of three fermion propagators (one power of \(\ell\) each in the denominator, one in the numerator from each \(\gamma\cdot\ell\)) — the superficial degree of divergence is \(D=+1\) by the counting of Theorem 100.17 applied to three external boson lines and a closed fermion loop. By Lemma 106.86 a shift of the routing therefore changes the amplitude by a finite, non-zero, momentum-dependent constant. The three Ward identities — one per vertex — are each linear conditions on that constant, and the constant has fewer free parameters than there are conditions: their sum is fixed independently of the routing and is non-zero. The routing can distribute the failure among the three vertices but cannot remove it. Since the vector current is coupled to the photon and its non-conservation would destroy gauge invariance and unitarity (Section 106.7.4), the failure is assigned entirely to the axial vertex.

The same statement in the language of regulators: Pauli–Villars requires a heavy fermion of mass \(M\), and a Dirac mass term \(M\bar\psi\psi\) is not chirally invariant, so the regulator breaks the axial symmetry by construction and the breaking survives \(M\to\infty\) by Equation (106.98). Dimensional regularization requires \(\gamma^{5}\) in \(d\neq4\) dimensions, where a totally antisymmetric product of four gamma matrices is not available. A lattice regulator is defeated by Theorem 106.75. Every route fails in a different way, which is itself the evidence that the obstruction is physical rather than technical.

Theorem 106.89 (Adler–Bardeen).

The one-loop result Equation (106.101) is exact: the coefficient of the anomaly receives no corrections at any higher order in the coupling [Adler:1969b]. Rests on Equation (106.101).

The reason is visible in Section 106.7.3 and not in the diagrams: the integrated anomaly is an integer — the index of the Dirac operator — and an integer cannot depend continuously on a coupling constant. This is the second time in this chapter that a topological argument makes a statement that a perturbative one can only verify order by order, the first being Proposition 106.81.

Proposition 106.90 (Wess–Zumino consistency).

Let \(G_{a}(x)\) generate gauge transformations, satisfying the algebra \(\comm{G_{a}(x)}{G_{b}(y)} =\ii f_{abc}\,\delta^{4}(x-y)\,G_{c}(x)\), and let \(\mathsf{A}_{a}(x):=G_{a}(x)\,W\) be the anomalous variation of the effective action. Then

\begin{equation}\tag{106.102} G_{a}(x)\,\mathsf{A}_{b}(y)-G_{b}(y)\,\mathsf{A}_{a}(x) =\ii f_{abc}\,\delta^{4}(x-y)\,\mathsf{A}_{c}(x)\ec \end{equation}

a strong constraint on the possible form of an anomaly [Wess:1971]. Rests on Equations (106.37) and (106.67).

Derivation. Derives Proposition 106.90. Apply the commutator of two gauge generators to the single functional \(W\) and use the algebra: \(\comm{G_{a}(x)}{G_{b}(y)}W =\ii f_{abc}\delta^{4}(x-y)G_{c}(x)W\). Writing out the left-hand side as \(G_{a}(x)\bigl(G_{b}(y)W\bigr)-G_{b}(y)\bigl(G_{a}(x)W\bigr)\) and substituting the definition of \(\mathsf{A}\) gives Equation (106.102). The content is that the anomaly is a cocycle of the gauge group: it is not an arbitrary functional but one constrained by the group structure, which is why the possible anomalies in four dimensions form a one-parameter family and the coefficient is the only thing left to compute.

Neutral pion decay as the measurement

The anomaly is not a formal curiosity. It fixes, with no free parameter beyond one measured decay constant, the rate of a decay that has been measured to about one and a half percent — and the answer depends on the number of colours.

Phenomenon 106.91 (The neutral pion decays to two photons at the anomaly rate).

The two-photon width of the neutral pion is fixed by the axial anomaly to be

\begin{equation}\tag{106.103} \Gamma(\pi^{0}\to\gamma\gamma) =\left(\frac{N_{c}}{3}\right)^{2} \frac{\alpha^{2}\,(m_{\pi}c^{2})^{3}} {64\pi^{3}F_{\pi}^{2}}\ec \qquad \frac{1}{\tau_{\gamma\gamma}} =\frac{\Gamma(\pi^{0}\to\gamma\gamma)}{\hbar}\ec \end{equation}

with \(F_{\pi}\) the pion decay constant. The first quantity is an energy — the width as the Particle Data Group quotes it — and \(\hbar\) enters only in the second, which converts it to a rate; the right-hand side of the first carries \((\text{energy})^{3}/ (\text{energy})^{2}\) and no \(\hbar\) at all. Using \(m_{\pi^{0}}c^{2}=134.9768(5)\,\mathrm{MeV}\) [Navas:2024], \(F_{\pi}=f_{\pi}/\sqrt2=92.07\,\mathrm{MeV}\) from \(f_{\pi}=130.2(1.2)\,\mathrm{MeV}\) [Navas:2024], and \(\alpha=7.2973525643\times 10^{-3}\) [Mohr:2025], the prediction for \(N_{c}=3\) is

\begin{equation}\tag{106.104} \Gamma=7.79(14)\,\mathrm{eV}\ec\qquad \frac{1}{\tau_{\gamma\gamma}}=1.18\times 10^{16}\,/\mathrm{s}\ec \end{equation}

the quoted uncertainty being that of \(f_{\pi}\) alone, which enters squared and so contributes twice its own \(0.9\,\mathrm{\%}\). Against this stands the PrimEx-II measurement

\begin{equation}\tag{106.105} \Gamma_{\mathrm{exp}}=7.802(117)\,\mathrm{eV} \end{equation}

[Larin:2020]: the two agree to \(0.2\,\mathrm{\%}\), which is \(0.09\) of the combined uncertainty — and the prediction is now the less precise of the two, its error being dominated by the measured decay constant rather than by anything in the anomaly. Were there one colour instead of three, the prediction would be nine times smaller. Rests on Equation (106.101) and Theorem 106.89.

Derivation. Derives Phenomenon 106.91. The anomaly Equation (106.101), applied to the third component of the axial isospin current of the light quarks and combined with the partial conservation of that current, produces an effective interaction between the neutral pion field and two photons,

\begin{equation}\tag{106.106} \Lag_{\pi\gamma\gamma} =-\frac{g_{\pi\gamma\gamma}}{4}\, \pi^{0}\,F_{\mu\nu}\widetilde{F}^{\mu\nu}\ec \qquad g_{\pi\gamma\gamma}=\frac{N_{c}}{3}\cdot\frac{\alpha}{\pi F_{\pi}}\ec \end{equation}

with \(\widetilde{F}^{\mu\nu} =\frac{1}{2}\epsilon^{\mu\nu\rho\sigma}F_{\rho\sigma}\) and \(g_{\pi\gamma\gamma}\) of dimension (energy)\(^{-1}\) once the fields are normalized as in Notation 106.23. The colour factor is \(N_{c}\bigl(e_{u}^{2}-e_{d}^{2}\bigr) =N_{c}\bigl(\tfrac{4}{9}-\tfrac{1}{9}\bigr)=\tfrac{N_{c}}{3}\): each quark flavour runs around the triangle once for each colour, weighted by the square of its electric charge, and the two light flavours enter with opposite sign because the current is the third isospin component. That is where the sensitivity to \(N_{c}\) comes from, and it is the only place in this book where a decay rate counts colours directly.

Given Equation (106.106), the width follows from kinematics. The amplitude for \(\pi^{0}(p)\to\gamma(k_{1},\varepsilon_{1})\gamma(k_{2}, \varepsilon_{2})\) is \(\mathcal{M}=g_{\pi\gamma\gamma}\,\epsilon^{\mu\nu\rho\sigma} k_{1\mu}\varepsilon^{*}_{1\nu}k_{2\rho}\varepsilon^{*}_{2\sigma}\). Sum over photon polarizations with \(\sum\varepsilon^{*}_{\mu}\varepsilon_{\nu}\to-\eta_{\mu\nu}\) — legitimate because \(\mathcal{M}\) vanishes when either \(\varepsilon\) is replaced by its momentum. The two metric factors enter with \((-1)^{2}=+1\), and the two Levi-Civita symbols contract by

\[ \epsilon^{\mu\nu\rho\sigma}\, \epsilon^{\alpha}{}_{\nu}{}^{\gamma}{}_{\sigma} =-2\bigl(\eta^{\mu\alpha}\eta^{\rho\gamma} -\eta^{\mu\gamma}\eta^{\rho\alpha}\bigr) \]

in the signature \(\diag(+1,-1,-1,-1)\) with \(\epsilon^{0123}=+1\), so that

\[ \sum_{\mathrm{pol}}\abs{\mathcal{M}}^{2} =-2g_{\pi\gamma\gamma}^{2} \bigl[k_{1}^{2}k_{2}^{2}-(k_{1}\cdot k_{2})^{2}\bigr] =2g_{\pi\gamma\gamma}^{2}\,(k_{1}\cdot k_{2})^{2}\ec \]

the first term vanishing because both photons are massless. The coefficient is \(2\), not \(8\): the same number comes out of summing \(\abs{\mathcal{M}}^{2}\) over the two physical transverse polarizations of each photon directly, without the covariant substitution. On the mass shell \(2k_{1}\cdot k_{2}=p^{2}=m_{\pi}^{2}c^{2}\). Measuring every momentum of this kinematic step as an energy, \(k\mapsto ck\), that reads \(k_{1}\cdot k_{2}=(m_{\pi}c^{2})^{2}/2\) and makes \(g_{\pi\gamma\gamma}\) an inverse energy throughout, whence \(\sum\abs{\mathcal{M}}^{2} =\tfrac{1}{2}g_{\pi\gamma\gamma}^{2}(m_{\pi}c^{2})^{4}\). The two-body massless phase-space volume is \(1/8\pi\), the flux factor for a decay at rest is \(1/(2m_{\pi}c^{2})\), and the two photons are identical so the result is halved:

\[ \frac{1}{\tau_{\gamma\gamma}} =\frac{1}{2}\cdot\frac{1}{2m_{\pi}c^{2}}\cdot \frac{1}{8\pi}\cdot \frac{g_{\pi\gamma\gamma}^{2}(m_{\pi}c^{2})^{4}}{2\hbar} =\frac{g_{\pi\gamma\gamma}^{2}(m_{\pi}c^{2})^{3}}{64\pi\hbar}\ec \]

a rate, since \(g_{\pi\gamma\gamma}^{2}(m_{\pi}c^{2})^{3}\) is an energy and \(\hbar\) divides it. The width is \(\hbar\) times that rate, \(\Gamma=g_{\pi\gamma\gamma}^{2}(m_{\pi}c^{2})^{3}/64\pi\), and substituting \(g_{\pi\gamma\gamma}\) from Equation (106.106) gives Equation (106.103). That is the only place \(\hbar\) enters, and it is why the displayed width carries none.

Arithmetic, in \(\mathrm{MeV}\): \(\alpha^{2}=5.32514\times 10^{-5}\); \((m_{\pi}c^{2})^{3}=2.4591\times 10^{6}\,\mathrm{MeV}^{3}\), the cube of \(134.9768\,\mathrm{MeV}\); \(64\pi^{3}=1984.40\); and \(F_{\pi}^{2}=8476.0\,\mathrm{MeV}^{2}\), the square of \(130.2\,\mathrm{MeV}/\sqrt2\). Then

\[ \Gamma=\frac{5.32514\times 10^{-5}\times2.4591\times 10^{6}} {1984.40\times 8476.0}\,\mathrm{MeV} =7.786\times 10^{-6}\,\mathrm{MeV} =7.79\,\mathrm{eV}\ec \]

which is Equation (106.104), and \(\Gamma/\hbar=7.79\,\mathrm{eV}/6.582119569\times 10^{-16}\,\mathrm{eV}\,\mathrm{s} =1.18\times 10^{16}\,/\mathrm{s}\). The prediction and the measurement differ by \(0.2\,\mathrm{\%}\) — well inside either uncertainty separately, and \(0.09\) of the two combined in quadrature.

Remark 106.92 (Two lifetimes that are both right).

The shipped Particle Data Group compilation gives the total width \(\Gamma_{\mathrm{tot}}=7.81(12)\,\mathrm{eV}\) [Navas:2024], whence a mean life

\[ \tau=\frac{\hbar}{\Gamma_{\mathrm{tot}}} =\frac{1.054571817\times 10^{-34}\,\mathrm{J}\,\mathrm{s}} {7.81\,\mathrm{eV}\times1.602176634\times 10^{-19}\,\mathrm{J}/\mathrm{eV}} =8.43\times 10^{-17}\,\mathrm{s}\ep \]

The PrimEx-II paper quotes \(\tau=8.34\times 10^{-17}\,\mathrm{s}\) [Larin:2020], which is a different number and not a discrepancy: what PrimEx-II measures is the two-photon partial width Equation (106.105), and converting it to a lifetime requires dividing by the branching fraction \(B(\pi^{0}\to\gamma\gamma)=0.98823\), giving \(\Gamma_{\mathrm{tot}}=7.895\,\mathrm{eV}\) and \(\tau=\hbar/\Gamma_{\mathrm{tot}}=8.34\times 10^{-17}\,\mathrm{s}\). The two lifetimes differ by about one percent because the PrimEx-II partial width is about one percent above the world average it is now the dominant input to. Quoting either without saying which is which is how a one-percent inconsistency gets into a book.

This is the cleanest existing evidence that the number of colours is three, and it is independent of the other route — the ratio \(R\) of hadronic to muonic annihilation cross-sections at an electron–positron collider, Equation (102.14), which counts colours through a sum over quark charges rather than through an anomaly. The two are worth having both, because they fail in different ways: \(R\) requires the perturbative treatment of the final state and carries QCD corrections of order \(\alpha_{s}/\pi\), while Equation (106.103) is protected from corrections by Theorem 106.89 but requires the chiral treatment of the pion.

Derivation pending.

The chiral matching for the pion two-photon coupling: the phase space, the polarization sum and the colour counting of the neutral-pion width are derived above, but the step from the anomalous divergence of the axial isospin current to the effective coupling \(g_{\pi\gamma\gamma}\) — the partial conservation of the axial current, the pion pole dominance and the normalization by the decay constant — is quoted rather than derived. It belongs in Appendix A, alongside the chiral material of the QCD chapter of this part.

Fujikawa's derivation from the measure

Every derivation so far has located the anomaly in a divergent loop integral. Fujikawa showed that it is in the measure [Fujikawa:1979]: the classical action of a massless fermion is chirally invariant, and so is the classical field equation, but \(\mathcal{D}\bar\psi\,\mathcal{D}\psi\) is not.

Theorem 106.93 (Fujikawa).

Under the local chiral rotation Equation (106.99) with \(\alpha=\alpha(x)\), the Euclidean fermion measure transforms with the Jacobian

\begin{equation}\tag{106.107} \mathcal{D}\bar\psi'\,\mathcal{D}\psi' =\mathcal{D}\bar\psi\,\mathcal{D}\psi\; \exp\!\left[-\frac{\ii}{16\pi^{2}}\int\dd^{4}x\;\alpha(x)\, \epsilon_{\mu\nu\rho\sigma} \tr\!\left(\mathcal{F}_{\mu\nu}\mathcal{F}_{\rho\sigma}\right) \right]\ec \end{equation}

which for constant \(\alpha\) is \(\ee^{-2\ii\alpha\nu}\) with \(\nu\) the topological charge of Equation (106.91). The anomaly Equation (106.101) follows. Rests on Equation (106.99), Corollary 106.36 and Equation (106.91).

Derivation. Derives Theorem 106.93. Expand the fields in eigenfunctions of the Hermitian Euclidean Dirac operator \(\mathcal{D}=\gamma^{\mu}D_{\mu}\), \(\mathcal{D}\varphi_{n}=\lambda_{n}\varphi_{n}\), with the \(\varphi_{n}\) orthonormal: \(\psi(x)=\sum_{n}a_{n}\varphi_{n}(x)\), \(\bar\psi(x)=\sum_{n}\bar b_{n}\varphi_{n}^{\dagger}(x)\), the coefficients \(a_{n},\bar b_{n}\) being Grassmann. The measure is \(\prod_{n}\dd a_{n}\,\dd\bar b_{n}\).

Under Equation (106.99) the coefficients transform linearly, \(a'_{m}=\sum_{n}C_{mn}a_{n}\) with

\[ C_{mn}=\int\dd^{4}x\;\varphi_{m}^{\dagger}(x)\, \ee^{\ii\alpha(x)\gamma^{5}}\varphi_{n}(x) =\delta_{mn}+\ii\int\dd^{4}x\;\alpha(x)\, \varphi_{m}^{\dagger}\gamma^{5}\varphi_{n}+O(\alpha^{2})\ec \]

and \(\bar b\) transforms with the same matrix, because Equation (106.99) carries the same sign for \(\bar\psi\). By Corollary 106.36 a Grassmann measure transforms with the inverse determinant, so

\[ \prod_{n}\dd a'_{n}\,\dd\bar b'_{n} =\bigl(\det C\bigr)^{-2}\prod_{n}\dd a_{n}\,\dd\bar b_{n}\ec \qquad \bigl(\det C\bigr)^{-2} =\exp\!\left[-2\ii\int\dd^{4}x\;\alpha(x)\, \mathcal{J}(x)\right]\ec \]

using \(\det=\exp\tr\ln\) to first order in \(\alpha\), with \(\mathcal{J}(x)=\sum_{n}\varphi_{n}^{\dagger}(x) \gamma^{5}\varphi_{n}(x)\).

The sum \(\mathcal{J}\) is formally \(\tr\gamma^{5}\) times a delta function — zero times infinity — and must be regularized. The regulator that respects gauge invariance is the operator's own spectrum:

\[ \mathcal{J}(x)=\lim_{M\to\infty}\sum_{n} \varphi_{n}^{\dagger}(x)\gamma^{5} \ee^{-\lambda_{n}^{2}/M^{2}}\varphi_{n}(x) =\lim_{M\to\infty}\tr\bra{x}\gamma^{5} \ee^{-\mathcal{D}^{2}/M^{2}}\ket{x}\ec \]

where the trace is over Dirac and gauge indices. Insert plane waves, \(\bra{x}\cdots\ket{x}=\int\dd^{4}k\,(2\pi)^{-4}\, \ee^{-\ii k\cdot x}(\cdots)\ee^{\ii k\cdot x}\), and use

\[ \mathcal{D}^{2}=(\gamma^{\mu}D_{\mu})^{2} =D^{2}+\frac{1}{2}\gamma^{\mu}\gamma^{\nu} \comm{D_{\mu}}{D_{\nu}} =D^{2}+\frac{\ii}{2}\gamma^{\mu}\gamma^{\nu} \mathcal{F}_{\mu\nu}\ec \]

by Equation (106.64). Expanding the exponential, the Dirac trace with \(\gamma^{5}\) vanishes unless at least four gamma matrices are present, so the first surviving term is the second order in \(\mathcal{F}\):

\[ \mathcal{J}=\lim_{M\to\infty}\int\frac{\dd^{4}k}{(2\pi)^{4}}\; \ee^{-k^{2}/M^{2}}\,\frac{1}{2!}\, \frac{1}{M^{4}}\left(\frac{\ii}{2}\right)^{2} \tr\!\left(\gamma^{5}\gamma^{\mu}\gamma^{\nu} \gamma^{\rho}\gamma^{\sigma}\right) \tr\!\left(\mathcal{F}_{\mu\nu}\mathcal{F}_{\rho\sigma}\right)\ep \]

Terms with fewer than four gammas vanish on the Dirac trace; terms with more carry additional inverse powers of \(M\) and vanish in the limit. Now \(\int\dd^{4}k\,(2\pi)^{-4}\ee^{-k^{2}/M^{2}}=M^{4}/(16\pi^{2})\) by Lemma 106.2, which cancels the \(M^{-4}\) exactly — so the limit is finite and independent of the regulator scale — and \(\tr(\gamma^{5}\gamma^{\mu}\gamma^{\nu}\gamma^{\rho}\gamma^{\sigma}) =-4\epsilon^{\mu\nu\rho\sigma}\). Collecting the numbers, \(\frac{1}{2}\times(-\frac{1}{4})\times(-4)\times\frac{1}{16\pi^{2}} =\frac{1}{32\pi^{2}}\), so

\begin{equation}\tag{106.108} \mathcal{J}(x)=\frac{1}{32\pi^{2}}\, \epsilon_{\mu\nu\rho\sigma} \tr\!\left(\mathcal{F}_{\mu\nu}\mathcal{F}_{\rho\sigma}\right)\ec \end{equation}

which is Equation (106.107). Comparison with Equation (106.91) gives \(\int\dd^{4}x\,\mathcal{J}=\nu\), so for constant \(\alpha\) the Jacobian is \(\ee^{-2\ii\alpha\nu}\).

Finally, the anomalous divergence. Invariance of \(Z\) under a change of integration variables requires that the variation of the action cancel the Jacobian. The action's variation under a local rotation produces \(\int\alpha\,\pp_{\mu}j_{5}^{\mu}\) after integrating by parts, so \(\pp_{\mu}j_{5}^{\mu}=2c\,\mathcal{J}\) for a massless field, i.e.\ \(\pp_{\mu}j_{5}^{\mu}=2q(x)\) with the topological charge density \(q\) of Definition 106.79 — not to be confused with the electric charge, which is why the latter is written \(Q\) throughout this chapter. For the abelian case, \(\mathcal{F}_{\mu\nu}=QF_{\mu\nu}/\hbar\) with no trace factor, and \(2c/(32\pi^{2})\times Q^{2}/\hbar^{2} =Q^{2}c/(16\pi^{2}\hbar^{2})\), which is Equation (106.101).

Three things about this derivation are worth separating out.

It uses the Grassmann Jacobian, and only the Grassmann Jacobian. Corollary 106.36 — that a fermionic measure transforms with \((\det)^{-1}\) where a bosonic one transforms with \(\det\) — was derived in Section 106.3.3 from Berezin's rules, which were themselves derived from translation invariance. That chain is the whole content of the anomaly: a symmetry of the action that fails to be a symmetry of the measure.

It is manifestly topological. By Equation (106.108) and Proposition 106.81, \(\int\mathcal{J}\) is an integer. The Atiyah–Singer index theorem identifies it further: the eigenvalue sum \(\sum_{n}\varphi_{n}^{\dagger}\gamma^{5}\varphi_{n}\) receives contributions only from zero modes, because \(\mathcal{D}\) maps the \(\gamma^{5}=+1\) and \(\gamma^{5}=-1\) subspaces into each other and therefore pairs every non-zero eigenvalue with a partner of opposite chirality, whose contributions cancel. Hence

\begin{equation}\tag{106.109} \nu=n_{+}-n_{-}\ec \end{equation}

the difference between the numbers of right- and left-handed zero modes of the Dirac operator. A decay rate — the measured \(7.802(117)\,\mathrm{eV}\) of Equation (106.105) — is thereby tied to an index. It is also the reason for Theorem 106.89: a quantity that equals an integer cannot vary continuously with a coupling constant.

It explains the third failure. The chiral rotation that would remove the phase of the quark mass matrix does not remove it: it shifts \(\theta\) by the Jacobian, which is why only \(\bar\theta=\theta+\arg\det M_{q}\) is observable (Equation (105.111)). The mechanism promised there is Equation (106.107), exactly.

Anomaly cancellation

An anomaly in a global symmetry is a prediction, and Section 106.7.2 measured one. An anomaly in a gauge symmetry is fatal.

Theorem 106.94 (A gauge anomaly destroys the theory).

If the currents coupled to gauge fields are anomalous, the Slavnov–Taylor identities Equation (106.76) fail. Then the unphysical polarizations no longer decouple (Proposition 106.60), the \(S\)-matrix is not unitary on \(\mathcal{H}_{\mathrm{phys}}\), and the divergences no longer organize into a renormalization of finitely many parameters [Gross:1972]. Rests on Theorem 106.62, Proposition 106.60 and Equation (106.107).

Derivation. Derives Theorem 106.94. Theorem 106.62 derived Equation (106.76) from the invariance of the action and of the measure under BRST. An anomaly is precisely the statement that the measure is not invariant: by Equation (106.107) the Jacobian contributes an extra term to every identity, so \(\avg{sX}=\avg{X\,\mathsf{A}}\neq0\). The chain of consequences then runs backwards through Section 106.4.2: the decoupling of the quartet in Proposition 106.60 used \(\braket{\Phi}{Q_{B}\Lambda}=0\), which requires \(Q_{B}\) to be conserved, and gauge-parameter independence Theorem 106.61 used \(\avg{sX}=0\) directly. Losing them means that a gauge-dependent quantity appears in a physical answer and that states of negative norm contribute to the optical theorem, which is a loss of probability, not a technical inconvenience.

Theorem 106.95 (The cancellation condition).

For chiral fermions in a representation with generators \(T^{a}\), the gauge anomaly is proportional to

\begin{equation}\tag{106.110} \mathsf{A}^{abc} =\tr\!\left(\acomm{T^{a}}{T^{b}}T^{c}\right)_{L} -\tr\!\left(\acomm{T^{a}}{T^{b}}T^{c}\right)_{R}\ec \end{equation}

the traces running over the left- and right-handed multiplets separately, and the theory is consistent if and only if \(\mathsf{A}^{abc}=0\) for every triple [Bouchiat:1972]. Rests on Propositions 106.88 and 106.90.

Derivation. Derives Theorem 106.95. The triangle of Proposition 106.88 now carries a group generator at each of its three vertices, and the two orientations of the fermion loop contribute the two orderings, so the group factor is the symmetrized trace \(\tr(T^{a}T^{b}T^{c}+T^{a}T^{c}T^{b}) =\tr(\acomm{T^{b}}{T^{c}}T^{a})\). A right-handed fermion contributes with the opposite sign, because \(\gamma^{5}\) has the opposite eigenvalue and the anomalous piece of the triangle is the one proportional to \(\gamma^{5}\). Summing over all fermions in the theory gives Equation (106.110), and by Proposition 106.90 nothing else can appear.

Proposition 106.96 (The Standard Model cancels, and only with three colours).

Write each generation as left-handed Weyl fields with hypercharges \(Y\) normalized by \(Q=T^{3}+Y\): \(Q_{L}=(3,2)_{1/6}\), \(u^{c}_{L}=(\bar3,1)_{-2/3}\), \(d^{c}_{L}=(\bar3,1)_{1/3}\), \(L_{L}=(1,2)_{-1/2}\), \(e^{c}_{L}=(1,1)_{1}\). Then every anomaly coefficient of \(\SU(3)\times\SU(2)\times\U(1)\) vanishes, and two of them vanish only because the number of colours equals three. Rests on Theorem 106.95 and Equation (106.110).

Derivation. Derives Proposition 106.96. There are four independent conditions, since \(\SU(3)^{3}\) vanishes identically (the quarks appear in a vector-like pair of triplet and antitriplet with respect to colour, once left and right are combined) and \(\SU(2)^{3}\) vanishes because the symmetrized trace of three \(\SU(2)\) generators is zero for any representation, \(\SU(2)\) having only real representations.

(i) \(\SU(3)^{2}\,\U(1)\). Sum \(Y\) over the colour-triplet Weyl fields, weighted by the number of \(\SU(2)\) components:

\[ 2\times\tfrac{1}{6}+1\times\left(-\tfrac{2}{3}\right) +1\times\tfrac{1}{3} =\tfrac{1}{3}-\tfrac{2}{3}+\tfrac{1}{3}=0\ep \]

(ii) \(\SU(2)^{2}\,\U(1)\). Sum \(Y\) over the \(\SU(2)\) doublets, weighted by their colour multiplicity:

\[ N_{c}\times\tfrac{1}{6}+1\times\left(-\tfrac{1}{2}\right) =\frac{N_{c}}{6}-\frac{1}{2}\ec \]

which vanishes if and only if \(N_{c}=3\).

(iii) \(\U(1)^{3}\). Sum \(Y^{3}\) over all Weyl fields with their multiplicities \(6,3,3,2,1\):

\[ 6\left(\tfrac{1}{6}\right)^{3} +3\left(-\tfrac{2}{3}\right)^{3} +3\left(\tfrac{1}{3}\right)^{3} +2\left(-\tfrac{1}{2}\right)^{3} +1\cdot1^{3} =\tfrac{1}{36}-\tfrac{32}{36}+\tfrac{4}{36} -\tfrac{9}{36}+\tfrac{36}{36}=0\ep \]

Carrying \(N_{c}\) instead of \(3\) through the quark multiplicities gives \(N_{c}\bigl[\tfrac{1}{108}-\tfrac{8}{27}+\tfrac{1}{27}\bigr] +\tfrac{3}{4}=\tfrac{3-N_{c}}{4}\), so this condition too holds only for \(N_{c}=3\).

(iv) Mixed gauge–gravitational. The same triangle with two insertions of the energy–momentum tensor requires \(\sum Y=0\):

\[ 6\left(\tfrac{1}{6}\right)+3\left(-\tfrac{2}{3}\right) +3\left(\tfrac{1}{3}\right)+2\left(-\tfrac{1}{2}\right)+1 =1-2+1-1+1=0\ep \]

This one is a statement about the Standard Model in a fixed classical background metric — a one-loop consistency condition on the gauge currents, not a theory of quantum gravity.

In every case the quarks alone and the leptons alone fail; only the complete generation cancels, and only with three colours. This is the theoretical counterpart of the measurement in Section 106.7.2, and the reason Electroweak Unification and the Higgs Boson requires complete generations: a fourth lepton doublet without its quark partners would make the theory inconsistent, so the discovery of a new charged lepton is a prediction that its quarks exist.

Remark 106.97 (The surviving non-perturbative constraint).

Theorem 106.95 is a perturbative condition, derived from a triangle diagram. Witten found a further one that no diagram sees [Witten:1982]: because \(\pi_{4}(\SU(2))=\Z_{2}\), there are gauge transformations on \(S^{4}\) not continuously connected to the identity, and the fermion determinant changes sign under them unless the number of \(\SU(2)\) doublets is even. The functional integral then vanishes identically — the theory has no non-vanishing amplitudes at all, which is a more complete failure than a broken Ward identity.

The Standard Model passes. Each generation contributes \(N_{c}=3\) quark doublets and one lepton doublet, four in all, and three generations give twelve: even. Had the leptons been counted without the quarks, or the colours been two, the count would have been odd and the theory empty. That two independent consistency conditions — one perturbative and one topological — and one measurement, Equation (106.105), all point at the same integer is the kind of redundancy that makes the Standard Model believable.