Weak Interactions

Contents
  1. The beta-decay crisis
  2. Fermi's theory
  3. Parity violation
  4. The $V-A$ interaction
  5. Strangeness-changing currents
  6. Neutral currents and the intermediate bosons
  7. Why Fermi theory had to be replaced

The weak interaction was forced on physics by an anomaly nobody wanted: the electrons emitted in nuclear beta decay come out with a continuous range of energies, although the initial and final nuclei have sharp masses. Everything in this chapter follows from the two responses to that single observation — a new, undetected neutral particle, and a new, extremely short-ranged interaction that does not conserve the quantum numbers the other forces conserve. The chapter runs from Chadwick's spectrum of 1914 to the discovery of the \(W\) and \(Z\) bosons in 1983, and it is deliberately told in that order, because the weak interaction is the one force whose theoretical structure was extracted almost entirely from experiment: parity violation, the \(V-A\) current, universality and the charm quark were each read off measurements before any principle demanded them.

It sits after Quantum Electrodynamics and Renormalization because it is where the QED template first fails. Fermi's four-fermion theory is non-renormalizable and violates unitarity at energies of a few hundred \(\mathrm{GeV}\); the resolution — a massive intermediate vector boson, and hence a spontaneously broken gauge symmetry — is the subject of Electroweak Unification and the Higgs Boson, which this chapter exists to motivate. Its immediate neighbours are the parity experiment of Experiment: Parity Violation, the neutrino-helicity measurement of Experiment: Neutrino Helicity, the quark-mixing and neutrino-mass material of Flavour Physics and Neutrinos, and the nuclear phenomenology of Nuclear Forces and Nuclear Structure. Standard treatments are [Weinberg:1996] [Halzen:1984]; current values throughout are from [Navas:2024].

Conventions are those of Section 100.1.1 without exception: the metric is \(\eta_{\mu\nu}=\diag(+1,-1,-1,-1)\), four-momenta are \(p^{\mu}=(E/c,\vect{p})\) of SI dimension \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), the mass shell is \(p^{2}=m^{2}c^{2}\) as in Equation (100.1), so that a propagator denominator reads \(p^{2}-m^{2}c^{2}\) and never \(p^{2}-m^{2}\), and the Dirac field carries \(\mathrm{m}^{-3/2}\) by Definition 100.2. Every \(\hbar\) and every \(c\) is written out. The one place where this chapter's arithmetic differs in appearance from its sources is the Fermi constant, which the literature quotes divided by \((\hbar c)^{3}\); both forms are given in Definition 101.8 and the difference between them is stated there rather than left to be guessed. Readers coming from the standard references, which drop both constants from every formula, should use the translation rule of Remark 100.4, which is the only place in this part where that convention is discussed.

The beta-decay crisis

The continuous spectrum

Alpha and gamma emission produce line spectra: a definite parent nucleus and a definite daughter nucleus differ by a definite energy, and the emitted quantum carries it. Beta emission was expected to do the same, and for two decades the experimental record was read as though it did — the sharp internal-conversion lines that accompany the beta continuum were mistaken for the primary spectrum, and the continuous background was ascribed to electrons that had lost energy on their way out of the source.

Chadwick settled the question in 1914 with a magnetic spectrometer read by an ionization chamber rather than by a photographic plate [Chadwick:1914]. A plate saturates, and a sharp line on a saturated plate looks like all there is; a counter measures intensity, and Chadwick's counter showed a broad continuous distribution on which the conversion lines sat as small peaks. The primary beta spectrum of radium B and C was continuous.

The escape route left open was that the electron leaves with the full energy and loses part of it in the source or in the surrounding material, so that the missing energy reappears as secondary radiation. Ellis and Wooster closed it by calorimetry [Ellis:1927]. They enclosed a radium-E source in a thick-walled calorimeter that would absorb everything except a neutrino, and measured the heat per disintegration: about \(0.35\,\mathrm{MeV}\), against an endpoint near \(1\,\mathrm{MeV}\). Had the missing energy been carried by any form of radiation, the calorimeter would have found it. Meitner and Orthmann repeated the measurement with a different calorimeter and a different absolute calibration and obtained the same answer [Meitner:1930], which is what turned a surprising result into an accepted one.

The seriousness of the anomaly is measured by the responses it drew. Bohr was for several years willing to entertain the possibility that energy conservation holds only statistically in nuclear processes — that is, to give up the conservation law rather than postulate an unseen particle. Pauli took the opposite view, and this chapter is the record of his being right; but the point worth keeping is that at the time neither option was obviously the cheaper one.

Phenomenon 101.1 (The beta spectrum is continuous).

The electrons emitted in nuclear beta decay do not carry a single energy. They are distributed continuously from zero up to a sharp endpoint fixed by the mass difference of the initial and final nuclei, although both of those masses are sharp [Chadwick:1914]. The missing energy is not recovered as secondary radiation: a calorimeter enclosing a radium-E source records a mean energy per disintegration of about \(0.35\,\mathrm{MeV}\) against an endpoint near \(1\,\mathrm{MeV}\) [Ellis:1927].

Derivation. Derives Phenomenon 101.1. Suppose the decay had only two products, a daughter nucleus of mass \(M'\) and an electron of mass \(m\), and let the parent of mass \(M\) be at rest. Momentum conservation makes the two momenta equal and opposite, \(\vect{p}'=-\vect{p}\), and energy conservation then reads

\begin{equation}\tag{101.1} Mc^{2}=\sqrt{p^{2}c^{2}+M'^{2}c^{4}} +\sqrt{p^{2}c^{2}+m^{2}c^{4}}\ep \end{equation}

The right-hand side is a strictly increasing function of \(p\), so it attains the value \(Mc^{2}\) for exactly one \(p\): a two-body decay of a state of definite mass yields a line, and the electron energy \(E=\sqrt{p^{2}c^{2}+m^{2}c^{4}}\) is fixed. The observed continuum is therefore, by itself, a proof that at least three bodies share the released energy. Since the daughter and the electron are both detected and account between them for less than the release, the third is not detected: it must be neutral, must interact far more weakly than the electron, and must be light enough to leave the endpoint where the two-body kinematics puts it. That is precisely the particle Pauli proposed in 1930 [Pauli:1930] and Cowan and Reines detected twenty-six years later [Cowan:1956]. The argument also shows what would have followed had the spectrum been a line with a lower endpoint — energy conservation itself would have been at stake, which is why Bohr was prepared to give it up.

Remark 101.2 (What the calorimeter measures, quantitatively).

The calorimetric result is stronger than “some energy is missing”. Write \(\avg{T_{e}}\) for the mean electron kinetic energy and \(T_{0}\) for its endpoint. A three-body decay with a massless third particle distributes the release between electron and neutrino according to the phase-space factor derived in Proposition 101.11; for an allowed transition with \(T_{0}\approx1\,\mathrm{MeV}\) that factor gives \(\avg{T_{e}}\approx0.40\,T_{0}\), and the measured ratio for radium E is about \(0.33\) [Ellis:1927]. A calorimeter that recorded \(T_{0}\) per disintegration would have proved the missing energy to be radiation; one that records a third of it proves the energy is carried away by something the calorimeter cannot stop. The residual discrepancy between \(0.40\) and \(0.33\) is itself instructive and was not understood until much later: the decay of radium E, that is \(^{210}\mathrm{Bi}\), is first-forbidden and its spectrum is correspondingly softer than the allowed shape. Nothing in the conclusion depends on that refinement — the alternative on offer in 1927 was that the calorimeter should have found the whole release, and it found a third.

Pauli's neutrino

On 4 December 1930 Pauli wrote to a conference at Tübingen he could not attend, addressing the “radioactive ladies and gentlemen” and proposing what he called a “desperate remedy” [Pauli:1930]: a neutral particle of spin \(\tfrac{1}{2}\), obeying the exclusion principle, of mass “of the same order of magnitude as the electron mass” and in any case not larger than about one per cent of the proton mass, emitted together with the electron in every beta decay. He called it the neutron; the name neutrino is Fermi's, coined after Chadwick's discovery of the real neutron in 1932 had taken the first name.

The proposal is remarkable for how much it buys at once. Three independent difficulties, which had no reason to be connected, are removed by one hypothesis.

  1. Energy. The release is shared between two light particles, so the electron energy is continuous with an endpoint at the full release, which is Phenomenon 101.1.

  2. Angular momentum. Beta decay changes the mass number \(A\) by zero, so it changes the number of nucleons not at all and the number of fermions in the nucleus not at all; emitting a single spin-\(\tfrac{1}{2}\) electron would change the total angular momentum of the system by a half-integer, which no rotation can do. A second spin-\(\tfrac{1}{2}\) particle restores integrality.

  3. Statistics. The nitrogen-14 nucleus was known from the band spectrum of \(\mathrm{N}_{2}\) to be a boson of spin \(1\). On the proton–electron model of the nucleus, \(^{14}\mathrm{N}\) contains \(14\) protons and \(7\) electrons, that is \(21\) fermions, and must be a fermion. The contradiction is what Pauli's letter calls the “wrong statistics”.

Proposition 101.3 (The endpoint measures the neutrino mass).

Let a nucleus of mass \(M\) at rest decay to a daughter of mass \(M'\), an electron of mass \(m_{e}\) and a neutral particle of mass \(m_{\nu}\). Write \(Q=(M-M'-m_{e})c^{2}\) for the energy released as kinetic energy. Then the maximum kinetic energy of the electron is

\begin{equation}\tag{101.2} T_{e}^{\max} =\frac{\left(Q-m_{\nu}c^{2}\right) \left(2M'c^{2}+Q+m_{\nu}c^{2}\right)} {2\left(M'c^{2}+m_{e}c^{2}+Q\right)} =\left(Q-m_{\nu}c^{2}\right) \left[1-\frac{2m_{e}c^{2}+Q-m_{\nu}c^{2}}{2M'c^{2}} +\cdots\right]\ec \end{equation}

and in the limit \(M'c^{2}\gg Q\) in which the daughter recoil is negligible, \(T_{e}^{\max}=Q-m_{\nu}c^{2}\). Rests on Equation (101.1).

Proof.

Derives Proposition 101.3. The electron takes its greatest energy when the daughter and the neutrino move together, so that the pair may be treated as a single body of mass \(M'+m_{\nu}\) recoiling against the electron. That is a two-body decay, and Equation (101.1) with \(M'\to M'+m_{\nu}\) fixes the electron energy uniquely:

\[ E_{e}^{\max}=\frac{M^{2}c^{4}+m_{e}^{2}c^{4} -\left(M'+m_{\nu}\right)^{2}c^{4}}{2Mc^{2}}\ep \]

Subtract \(m_{e}c^{2}\) and use \(Mc^{2}-m_{e}c^{2}=M'c^{2}+Q\) to factor the numerator as \(\left(Q-m_{\nu}c^{2}\right)\left(2M'c^{2}+Q+m_{\nu}c^{2}\right)\); writing the denominator as \(2(M'c^{2}+m_{e}c^{2}+Q)\) gives the exact form in Equation (101.2), and expanding in \(1/M'c^{2}\) gives the second. The recoil correction is of relative order \(\left(2m_{e}c^{2}+Q\right)/2M'c^{2}\) — dominated by the electron mass, not by \(Q\), which is the trap in this expansion — and for tritium (\(Q=18.6\,\mathrm{keV}\), \(M'c^{2}=2808\,\mathrm{MeV}\)) it is \(1.85\times 10^{-4}\), that is \(3.4\,\mathrm{eV}\) of recoil energy taken by the daughter. That is four times the whole mass sensitivity of Equation (101.23), so it is computed and subtracted rather than neglected. The leading statement is the one that matters here: the endpoint sits a distance \(m_{\nu}c^{2}\) below the release, so a measured endpoint bounds the neutrino mass without any assumption about the interaction that produced it.

Pauli's own bound came from exactly this argument applied to the spectra then available, and gave \(m_{\nu}\lesssim m_{e}\); the modern version of the same measurement, on tritium, gives \(m_{\nu}<0.8\,\mathrm{eV}/c^{2}\) [Aker:2022], an improvement of six orders of magnitude in the same observable.

Pauli's caution was well judged in one respect. He estimated that his particle, being neutral and weakly coupled, would have a penetrating power far greater than a photon's, and doubted it could ever be observed. That doubt is quantified in Section 101.1.3 and it was very nearly right.

Detection

A particle postulated to explain a missing quantity is a bookkeeping device until it is detected in its own right. The detection reaction is beta decay run backwards,

\begin{equation}\tag{101.3} \bar{\nu}_{e}+p\longrightarrow n+e^{+}\ec \end{equation}

which has two attractive features: free protons are available in bulk as the hydrogen of water, and the final state carries its own signature in time.

Proposition 101.4 (Threshold of the inverse beta reaction).

Reaction Equation (101.3) on a proton at rest requires

\begin{equation}\tag{101.4} E_{\bar{\nu}}\geq \frac{\left(m_{n}+m_{e}\right)^{2}c^{4}-m_{p}^{2}c^{4}} {2m_{p}c^{2}}=1.806\,\mathrm{MeV}\ep \end{equation}

Rests on Equation (101.3).

Proof.

Derives Proposition 101.4. The invariant \(s=(p_{\bar{\nu}}+p_{p})^{2}\) is conserved, and the final state exists only if \(\sqrt{s}\,c\geq(m_{n}+m_{e})c^{2}\), the least energy the pair can have. With the proton at rest and the antineutrino massless, \(s c^{2}=m_{p}^{2}c^{4}+2m_{p}c^{2}E_{\bar{\nu}}\), and the inequality rearranges to Equation (101.4). Inserting \(m_{n}c^{2}=939.565\,\mathrm{MeV}\), \(m_{p}c^{2}=938.272\,\mathrm{MeV}\) and \(m_{e}c^{2}=0.511\,\mathrm{MeV}\) [Navas:2024] gives \(1.806\,\mathrm{MeV}\). A fission reactor emits antineutrinos with a spectrum extending to about \(8\,\mathrm{MeV}\), so a useful fraction of them are above threshold; this is why the experiment was done at a reactor and not with a natural source.

The signature is a delayed coincidence. The positron annihilates within nanoseconds against an electron of the target, giving two \(511\,\mathrm{keV}\) photons back to back; the neutron wanders, thermalizes, and is captured microseconds later on a nucleus with a large capture cross-section — cadmium dissolved in the water — which de-excites by emitting several megaelectronvolts of gamma rays. Two flashes separated by a characteristic time, in the right energy ratio, in a detector shielded against everything else, is a signature no ordinary background reproduces. Cowan, Reines and their collaborators built it at the Savannah River reactor and reported the observation in 1956 [Cowan:1956] [Reines:1956], twenty-six years after Pauli's letter.

Phenomenon 101.5 (The neutrino interacts, and almost never).

Free antineutrinos from a reactor induce reaction Equation (101.3) at a rate corresponding to a cross-section of order \(10^{-47}\,\mathrm{m}^{2}\) for a few-\(\mathrm{MeV}\) antineutrino [Cowan:1956] [Reines:1956]. A block of lead containing that many protons per unit volume would have to be of order a light-year thick to stop half the beam. Rests on Equation (101.3) and Proposition 101.4.

Derivation. Derives Phenomenon 101.5. The cross-section may be predicted from the coupling that Section 101.2.1 extracts from beta decay itself, which is what makes the number a test rather than a fit. At energies far below the mass of the mediating boson the interaction is a contact one, the amplitude is a product of two currents at a point, and by the derivation of Proposition 101.54 — which is performed there in full and is common to every charged-current process of this kind — the cross-section for the nucleon reaction is

\begin{equation}\tag{101.5} \sigma\left(\bar{\nu}_{e}p\to ne^{+}\right) =\frac{G_{F}^{2}\abs{V_{ud}}^{2}\left(1+3\lambda^{2}\right)} {\pi\left(\hbar c\right)^{4}}\,p_{e}c\,E_{e}\ec \end{equation}

with \(E_{e}=E_{\bar{\nu}}-(m_{n}-m_{p})c^{2}\) the positron energy, \(p_{e}c=\sqrt{E_{e}^{2}-m_{e}^{2}c^{4}}\) and \(\lambda=g_{A}/g_{V}=1.2754\) the ratio of the axial and vector nucleon couplings measured in neutron decay [Navas:2024]; the factor \(1+3\lambda^{2}\) is the sum of one Fermi and three Gamow–Teller spin channels, counted in Section 101.2.2. The factor \(\abs{V_{ud}}^{2}=\cos^{2}\theta_{C}=0.949\) is there because this is a hadronic charged current and not a leptonic one: the reaction turns a \(d\) quark into a \(u\), and by Equation (101.63) that vertex carries \(\cos\theta_{C}\). Omitting it overstates the rate by five per cent. At \(E_{\bar{\nu}}=3\,\mathrm{MeV}\) this gives \(E_{e}=1.707\,\mathrm{MeV}\), \(p_{e}c=1.628\,\mathrm{MeV}\) and

\[ \sigma=2.6\times 10^{-47}\,\mathrm{m}^{2}\ec \]

which is the observed order of magnitude. For the penetrating power, lead has \(11.34\times 10^{3}\,\mathrm{kg}/\mathrm{m}^{3}\) and molar mass \(0.2072\,\mathrm{kg}/\mathrm{mol}\), hence \(3.30\times10^{28}\) nuclei and \(2.70\times10^{30}\) protons per cubic metre; the mean free path is

\[ \ell=\frac{1}{n\sigma} =\frac{1}{\left(2.70\times10^{30}\,/\mathrm{m}^{3}\right) \left(2.6\times 10^{-47}\,\mathrm{m}^{2}\right)} =1.4\times 10^{16}\,\mathrm{m}\ec \]

which is \(1.5\,\mathrm{yr}\) of travel at the speed of light. The idealization is explicit and is worth stating: a proton bound in a lead nucleus is not a free proton and the reaction on it is not Equation (101.3), so the figure counts protons, not targets. As an order of magnitude for the transparency of matter to neutrinos it is nevertheless the right one, and it is the reason Pauli thought his particle undetectable.

Two later measurements complete the picture of what a neutrino is. The first is that there is more than one kind. Firing a beam of neutrinos produced in pion decay, \(\pi^{+}\to\mu^{+}\nu\), at a spark-chamber detector, Danby and collaborators at the Brookhaven AGS found the reaction products to be muons and never electrons [Danby:1962]: the neutrino accompanying a muon makes a muon, not an electron. Since nothing in the kinematics forbids the electron channel — it is in fact the more open of the two — the null result establishes a conserved quantum number distinguishing the two families, which is lepton flavour. The third family's neutrino was observed directly at Fermilab forty years later by the DONUT collaboration, which identified the short kinked track of a tau produced in a nuclear emulsion [Kodama:2001]. The second completion is the counting: Section 101.6.4 shows that the number of light neutrino species that couple to the \(Z\) is three, and that this is measured, not assumed.

Fermi's theory

The four-fermion interaction

Fermi's 1934 paper [Fermi:1934] is the origin of every equation in this chapter. Its construction is an analogy, and the analogy is with the electromagnetic interaction of Section 100.1. In QED the interaction Lagrangian density is a current contracted with a field, \(\Lag_{\mathrm{int}}=-j^{\mu}A_{\mu}\), and an elastic scattering amplitude is a product of two currents joined by the photon propagator: the interaction is current–current, with the two currents at different points and a field carrying the influence between them. Fermi kept the current–current form, kept the vector character of the currents, and threw away the field: his two currents sit at the same spacetime point.

The reasons were empirical. Beta decay had a range too short to measure, and no quantum of the required kind was known; the electron and the neutrino are created in the decay, exactly as a photon is created in an atomic transition, which is what suggested a field theory rather than a rearrangement of pre-existing constituents; and the vector form was the one that reproduced the observed spectrum shape. Nature rejected the paper as “too remote from physical reality”. It appeared in Zeitschrift für Physik and contains, in the space of a few pages, the coupling constant, the spectrum shape, and the lifetime systematics.

Definition 101.6 (Fermi's interaction).

Let \(\psi_{p},\psi_{n},\psi_{e},\psi_{\nu}\) be the Dirac fields of the proton, neutron, electron and neutrino, each of SI dimension \(\mathrm{m}^{-3/2}\) by Definition 100.2. Fermi's interaction is the Lorentz-invariant contact term

\begin{equation}\tag{101.6} \Lag_{\mathrm{F}}=-G\left(\bar{\psi}_{p}\gamma^{\mu}\psi_{n}\right) \left(\bar{\psi}_{e}\gamma_{\mu}\psi_{\nu}\right) +\text{h.c.}\ec \end{equation}

in which each factor in parentheses is a current, the two are evaluated at the same point \(x\), and \(G\) is a constant.

Proposition 101.7 (The Fermi coupling has the dimension of an energy times a volume).

For Equation (101.6) to be an energy density, \(\mathrm{J}/\mathrm{m}^{3}\), the constant \(G\) must carry

\begin{equation}\tag{101.7} [G]=\mathrm{J}\,\mathrm{m}^{3} =\mathrm{J}^{-2}\times\left[\left(\hbar c\right)^{3}\right]\ec \end{equation}

that is, \(G/(\hbar c)^{3}\) is an inverse squared energy. Rests on Equation (101.6) and Definition 100.2.

Proof.

Derives Proposition 101.7. By Definition 100.2 the bilinear \(\bar{\psi}\gamma^{\mu}\psi\) has dimension \(/\mathrm{m}^{3}\), so the product of two currents carries \(\mathrm{m}^{-6}\). An energy density is \(\mathrm{J}/\mathrm{m}^{3}\). Hence \([G]=\mathrm{J}/\mathrm{m}^{3}\times\mathrm{m}^{6} =\mathrm{J}\,\mathrm{m}^{3}\). Since \(\hbar c\) has dimension \(\mathrm{J}\,\mathrm{m}\), dividing by \((\hbar c)^{3}\) leaves \(\mathrm{J}^{-2}\).

This one line is the most consequential dimensional statement in the chapter. Every coupling in Quantum Electrodynamics and Renormalization is dimensionless once expressed as \(\alpha\); Fermi's is not, and cannot be made so. Section 101.7.2 shows that this alone condemns the theory to non-renormalizability, and Section 101.7.1 shows that it also fixes, in advance, the energy at which the theory must fail.

Definition 101.8 (The Fermi constant).

The Fermi constant is defined by the muon lifetime through Equation (101.53), and its measured value is

\begin{equation}\tag{101.8} \frac{G_{F}}{\left(\hbar c\right)^{3}} =1.1663787(6)\times 10^{-5}\,/\mathrm{GeV}^{2}\ec\qquad G_{F}=1.43585\times 10^{-62}\,\mathrm{J}\,\mathrm{m}^{3}\ep \end{equation}

The first form is the one universally quoted, and it is not \(G_{F}\) but \(G_{F}\) divided by the cube of \(\hbar c\); the second is \(G_{F}\) itself in SI base units. Confusing them costs a factor \((\hbar c)^{3}=3.1600\times 10^{-77}\,\mathrm{J}^{3}\,\mathrm{m}^{3}\), so the distinction is stated here once and relied on thereafter.

Derivation of the conversion. Derives Definition 101.8. \(\hbar c=3.161527\times 10^{-26}\,\mathrm{J}\,\mathrm{m}\) from the defined values of \(\hbar\) and \(c\), so \((\hbar c)^{3}=3.16003\times 10^{-77}\,\mathrm{J}^{3}\,\mathrm{m}^{3}\). One gigaelectronvolt is \(1.602176634\times 10^{-10}\,\mathrm{J}\) exactly, since the elementary charge is fixed by definition since 2019, so \(/\mathrm{GeV}^{2}=3.89566\times 10^{19}\,/\mathrm{J}^{2}\). Then

\[ G_{F}=1.1663787\times 10^{-5}\times3.89566\times 10^{19}\,/\mathrm{J}^{2} \times3.16003\times 10^{-77}\,\mathrm{J}^{3}\,\mathrm{m}^{3} =1.43585\times 10^{-62}\,\mathrm{J}\,\mathrm{m}^{3}\ep \]

Two derived numbers used repeatedly below follow at once: the energy at which the dimensionless combination \(G_{F}E^{2}/(\hbar c)^{3}\) reaches unity is

\begin{equation}\tag{101.9} E_{F}=\left(\frac{\left(\hbar c\right)^{3}}{G_{F}}\right)^{1/2} =292.8\,\mathrm{GeV}\ec \end{equation}

and the combination \((\sqrt{2}G_{F}/(\hbar c)^{3})^{-1/2} =246.2\,\mathrm{GeV}\) is the electroweak scale \(v\) that Electroweak Unification and the Higgs Boson identifies with the vacuum expectation value of the scalar field.

Remark 101.9 (Why the customary units are not abandoned here).

Editorial rule 3 requires SI, and Equation (101.8) gives it. The reciprocal-squared-energy form is nevertheless retained alongside, for a reason that is physics and not deference: the quantity that controls the size of every weak amplitude is the dimensionless ratio \(G_{F}E^{2}/(\hbar c)^{3}\), and quoting \(G_{F}/(\hbar c)^{3}\) in \(/\mathrm{GeV}^{2}\) makes that ratio readable at a glance — at \(E=1\,\mathrm{GeV}\) it is \(1.17\times 10^{-5}\), at \(E=100\,\mathrm{GeV}\) it is \(0.117\), and at Equation (101.9) it is one. The bare SI number \(1.43585\times 10^{-62}\,\mathrm{J}\,\mathrm{m}^{3}\) conveys none of that. Both are correct; each is quoted where it does work.

The rest of this subsection extracts from Equation (101.6) the two results Fermi's paper is remembered for: the shape of the electron spectrum and the systematics of the lifetime. Both follow from one tool, which is used again in Section 101.4.3 in its relativistic form and is therefore stated here once with its derivation.

Proposition 101.10 (The golden rule).

Let a system be prepared at \(t=0\) in a discrete state \(\ket{i}\) of energy \(E_{i}\) and let a time-independent perturbation \(V\) connect it to a continuum of final states \(\ket{f}\) with density \(\rho(E_{f})\) states per unit energy. Then, for times long compared with \(\hbar/\abs{E_{f}-E_{i}}\) over the width of the continuum and short compared with the lifetime, the transition rate is constant and equal to

\begin{equation}\tag{101.10} \Gamma=\frac{2\pi}{\hbar}\sum_{f} \abs{\bra{f}V\ket{i}}^{2}\, \delta\left(E_{f}-E_{i}\right)\ec \end{equation}

the sum being an integral \(\int\dd E_{f}\,\rho(E_{f})\) over the continuum. Rests on Postulate 77.4.

Proof.

Derives Proposition 101.10. To first order in \(V\) the amplitude for finding the system in \(\ket{f}\) at time \(t\) is

\[ c_{f}(t)=-\frac{\ii}{\hbar}\int_{0}^{t}\dd t'\, \bra{f}V\ket{i}\,\ee^{\ii\left(E_{f}-E_{i}\right)t'/\hbar} =\bra{f}V\ket{i}\, \frac{1-\ee^{\ii\omega_{fi}t}}{\hbar\omega_{fi}}\ec \]

with \(\hbar\omega_{fi}=E_{f}-E_{i}\). Hence

\[ \abs{c_{f}(t)}^{2}=\abs{\bra{f}V\ket{i}}^{2}\, \frac{4\sin^{2}\left(\omega_{fi}t/2\right)} {\hbar^{2}\omega_{fi}^{2}}\ep \]

The kernel \(4\sin^{2}(\omega t/2)/(\hbar\omega)^{2}\) has total integral \(\int\dd\omega\,4\sin^{2}(\omega t/2)/(\hbar\omega)^{2} =2\pi t/\hbar^{2}\) and a width in \(\omega\) of order \(1/t\), so for \(t\) large compared with the inverse spread of the continuum it acts as \((2\pi t/\hbar^{2})\delta(\omega)=(2\pi t/\hbar)\delta(E_{f}-E_{i})\). Summing over \(f\) and dividing by \(t\) gives Equation (101.10); the probability grows linearly in \(t\), which is what makes the rate well defined and the decay law exponential once the depletion of \(\ket{i}\) is fed back in. The result is Dirac's [Dirac:1927]; the name is Fermi's [Fermi:1950], and it is the tool with which he wrote [Fermi:1934].

Proposition 101.11 (The allowed beta spectrum).

Let a nucleus decay by an allowed transition — one in which the electron and neutrino are emitted in relative \(s\) waves, so that the nuclear matrix element does not depend on their momenta. Then the electron energy spectrum is

\begin{equation}\tag{101.11} \frac{\dd\Gamma}{\dd E_{e}}\propto F(Z,E_{e})\;p_{e}E_{e} \left(E_{0}-E_{e}\right) \sqrt{\left(E_{0}-E_{e}\right)^{2}-m_{\nu}^{2}c^{4}}\ec \end{equation}

where \(E_{0}\) is the total energy available to the lepton pair, \(p_{e}=\sqrt{E_{e}^{2}/c^{2}-m_{e}^{2}c^{2}}\), and \(F(Z,E_{e})\) is the Coulomb correction for the motion of the emitted electron in the field of the daughter nucleus. For \(m_{\nu}=0\) this reduces to

\begin{equation}\tag{101.12} \frac{\dd\Gamma}{\dd E_{e}}\propto F(Z,E_{e})\,p_{e}E_{e} \left(E_{0}-E_{e}\right)^{2}\ep \end{equation}

Rests on Proposition 101.10.

Proof.

Derives Proposition 101.11. Apply the golden rule of Proposition 101.10 with two light particles in the final state and the recoiling daughter treated as infinitely heavy, so that it absorbs momentum but no energy. With the matrix element independent of the lepton momenta, the entire energy dependence is the density of final states,

\[ \dd\rho\propto\dd^{3}p_{e}\,\dd^{3}p_{\nu}\, \delta\left(E_{0}-E_{e}-E_{\nu}\right) =16\pi^{2}p_{e}^{2}\,\dd p_{e}\;p_{\nu}^{2}\,\dd p_{\nu}\, \delta\left(E_{0}-E_{e}-E_{\nu}\right)\ec \]

the angular integrations giving \(4\pi\) each because an \(s\)-wave matrix element is isotropic. Change variables with \(p^{2}\,\dd p/\dd E=pE/c^{2}\), valid for either particle, and use the delta function to set \(E_{\nu}=E_{0}-E_{e}\):

\[ \frac{\dd\rho}{\dd E_{e}}\propto \frac{p_{e}E_{e}}{c^{2}}\cdot\frac{p_{\nu}E_{\nu}}{c^{2}} =\frac{p_{e}E_{e}}{c^{4}}\left(E_{0}-E_{e}\right) \sqrt{\left(E_{0}-E_{e}\right)^{2}-m_{\nu}^{2}c^{4}}\ec \]

where the last step used \(p_{\nu}c=\sqrt{E_{\nu}^{2}-m_{\nu}^{2}c^{4}}\). That is Equation (101.11). The Coulomb factor is a separate effect: the emitted electron is not a plane wave but a continuum Coulomb wave in the field of the daughter's charge \(Z\), and the rate is proportional to the squared modulus of that wave at the nucleus. Its non-relativistic form is

\begin{equation}\tag{101.13} F(Z,E_{e})=\frac{2\pi\eta}{1-\ee^{-2\pi\eta}}\ec\qquad \eta=\pm\frac{Z\alpha c}{v_{e}}\ec \end{equation}

with the upper sign for \(\beta^{-}\) emission and the lower for \(\beta^{+}\), \(Z\) being the charge number of the daughter. The sign is the entire physical content of the factor, and the two limits fix it. For \(\beta^{-}\) the electron is attracted, so \(\eta>0\), and \(F\to2\pi\eta>1\) as \(v_{e}\) falls: the Coulomb field draws the outgoing wave onto the nucleus and enhances the low-energy end of the spectrum. For \(\beta^{+}\) the positron is repelled, so \(\eta<0\), and \(F\to2\pi\abs{\eta}\ee^{-2\pi\abs{\eta}}\ll1\), which is the Gamow barrier factor and suppresses the same end. Here \(\alpha\) is the fine-structure constant of Equation (100.3) and \(v_{e}\) the electron speed; \(\eta\) is dimensionless, as it must be.

Proposition 101.12 (Sargent's rule).

For an allowed transition with \(E_{0}\gg m_{e}c^{2}\) and the Coulomb factor neglected, the total rate grows as the fifth power of the energy release,

\begin{equation}\tag{101.14} \Gamma\propto E_{0}^{5}\ep \end{equation}

Rests on Equation (101.12).

Proof.

Derives Proposition 101.12. Integrate Equation (101.12) with \(F=1\) and \(p_{e}c\simeq E_{e}\):

\[ \Gamma\propto\int_{0}^{E_{0}}E_{e}^{2} \left(E_{0}-E_{e}\right)^{2}\dd E_{e} =E_{0}^{5}\left(\frac{1}{3}-\frac{1}{2}+\frac{1}{5}\right) =\frac{E_{0}^{5}}{30}\ep \]

The fifth power is pure phase space — two particles sharing a release, each contributing a \(p^{2}\dd p\) — and it is what Sargent had found empirically in 1933 by plotting the decay constants of the naturally occurring beta emitters against their endpoints on logarithmic axes and obtaining two straight lines of slope five [Sargent:1933]. The two lines, rather than one, are the first appearance in this chapter of a distinction that Section 101.2.2 identifies: allowed and first-forbidden transitions have different matrix elements and therefore different intercepts, but the same slope, because the slope is kinematics.

Remark 101.13 (Fermi's theory predicted a shape, and was believed because of it).

It is worth being explicit about the logic, because it recurs. Equation (101.12) contains no free function: once \(E_{0}\) is read off the endpoint and \(Z\) off the daughter, the entire spectrum is fixed, with the overall normalization the only adjustable number. A theory that merely permitted a continuous spectrum would have explained nothing — the continuum was the datum. Fermi's theory predicted which continuum, and the prediction is a one-parameter fit to a curve of several hundred points. That is why the vector form survived until parity violation forced its modification twenty-three years later, and why the modification, when it came, had to preserve the spectrum shape — as Section 101.4.1 shows the \(V-A\) form does.

Allowed transitions: Fermi and Gamow–Teller

Fermi's choice of a vector current in Equation (101.6) was one of five possibilities, and Nature uses two of them. The enumeration is fixed by Lorentz invariance and is worth doing, because the same list organizes every experiment in Section 101.4.1 that fixed the interaction empirically.

Proposition 101.14 (There are exactly five bilinear covariants).

The sixteen matrices

\begin{equation}\tag{101.15} \Gamma_{i}\in\set{\identity,\ \gamma^{\mu},\ \sigma^{\mu\nu},\ \gamma^{\mu}\gamma^{5},\ \gamma^{5}}\ec\qquad \sigma^{\mu\nu}=\frac{\ii}{2} \comm{\gamma^{\mu}}{\gamma^{\nu}}\ec \end{equation}

form a basis of the complex \(4\times4\) matrices, and the bilinears \(\bar{\psi}\Gamma_{i}\chi\) built from them transform under the Lorentz group as a scalar (S), a vector (V), an antisymmetric tensor (T), an axial vector (A) and a pseudoscalar (P) respectively. Counting components, \(1+4+6+4+1=16\). Rests on Lemma 100.23.

Proof.

Derives Proposition 101.14. Linear independence: with \(\acomm{\gamma^{\mu}}{\gamma^{\nu}}=2\eta^{\mu\nu}\identity\), every product of gamma matrices reduces to one of Equation (101.15) up to a sign, and the trace relations of Lemma 100.23 give \(\tr(\Gamma_{i}\Gamma_{j}^{-1}) =4\delta_{ij}\) for the sixteen listed matrices, from which \(\sum_{i}c_{i}\Gamma_{i}=0\) forces every \(c_{i}=0\). Sixteen independent matrices in a sixteen-dimensional space are a basis. The transformation laws follow from the defining property of the spinor representation, \(S^{-1}(\Lambda)\gamma^{\mu}S(\Lambda) =\Lambda^{\mu}{}_{\nu}\gamma^{\nu}\), established in The Dirac Equation: \(\bar{\psi}\chi\) is invariant, \(\bar{\psi}\gamma^{\mu}\chi\) carries one vector index, and so on. That the last two are pseudo-quantities is the statement that \(\gamma^{5}\) anticommutes with every \(\gamma^{\mu}\) and therefore picks up a sign under the improper transformation \(P\), which is proved in Proposition 101.23.

The most general Lorentz-invariant contact interaction between two pairs of spin-\(\tfrac{1}{2}\) fields is then a sum of five terms, one for each covariant, with five complex coefficients — and, once parity is not assumed, ten, since each term may be accompanied by its parity-reflected partner. Determining those ten numbers from experiment is the programme that occupies Sections 101.3 and 101.4.

For nuclear beta decay only two of the five survive, and the reason is angular momentum. In an allowed transition the leptons carry away no orbital angular momentum, so the total change in nuclear spin equals the total spin of the emitted pair, which is \(0\) or \(1\).

Proposition 101.15 (Allowed selection rules).

For an allowed transition,

  1. the S and V couplings emit the lepton pair in a spin singlet, giving \(\Delta J=0\) with no change of nuclear parity — the Fermi selection rule;

  2. the T and A couplings emit it in a spin triplet, giving \(\Delta J=0,\pm1\) with no change of parity and with \(0\to0\) forbidden — the Gamow–Teller rule [Gamow:1936];

  3. the P coupling contributes nothing at leading order, its nuclear matrix element being of order \(v_{\mathrm{nucleon}}/c\).

The parity of the nucleus cannot change in an allowed transition, because the leptons carry away no orbital angular momentum and the nuclear operators that survive, \(\identity\) and \(\vect{\sigma}\), are both even under reflection. Rests on Proposition 101.14.

Proof.

Derives Proposition 101.15. Take the non-relativistic limit of the nucleon current, in which the large components of the nucleon spinors dominate. For the scalar and the time component of the vector, \(\bar{\psi}_{p}\psi_{n}\) and \(\bar{\psi}_{p}\gamma^{0}\psi_{n}\) both reduce to \(\phi_{p}^{\dagger}\phi_{n}\) acting on the two-component nucleon spinors: the operator is the identity in spin space, so the nucleon spin is untouched and by conservation the pair carries spin zero, hence \(\Delta J=0\). For the axial spatial components and the space components of the tensor, \(\bar{\psi}_{p}\gamma^{i}\gamma^{5}\psi_{n}\) and \(\bar{\psi}_{p}\sigma^{ij}\psi_{n}\) reduce to \(\phi_{p}^{\dagger}\sigma^{i}\phi_{n}\): the operator is a Pauli matrix, which flips the nucleon spin, so the pair carries spin one and \(\Delta J\in\set{0,\pm1}\). A spin-one pair cannot be emitted in a \(0\to0\) transition, since there is nothing to absorb its unit of angular momentum. The remaining components — the space part of the vector, the time part of the axial vector and the pseudoscalar — carry an explicit factor of the nucleon velocity in the non-relativistic reduction and are suppressed by \(v_{\mathrm{nucleon}}/c\sim10^{-1}\); they are the leading terms of the forbidden transitions.

Parity: in the allowed approximation the lepton wave functions are constants evaluated at the nucleus, so the whole dependence on the nuclear coordinates sits in the two operators just obtained, \(\identity\) and \(\sigma^{i}\), and both are even under spatial reflection — the identity trivially, the Pauli matrix because a spin is an axial vector. A parity-even operator has vanishing matrix element between nuclear states of opposite parity, so the nuclear parity cannot change. The created pair does carry intrinsic parity \(-1\), a fermion and an antifermion having opposite intrinsic parities, but that factor is the same in every beta decay and multiplies every amplitude alike; what distinguishes one transition from another is the orbital factor \((-1)^{\ell}\), which is \(+1\) here and is exactly what changes sign in a first-forbidden transition.

The decisive datum is that both rules are needed. The decay

\[ {}^{6}\mathrm{He}\left(0^{+}\right)\longrightarrow {}^{6}\mathrm{Li}\left(1^{+}\right)+e^{-}+\bar{\nu}_{e} \]

has \(\Delta J=1\) and is fast, which the Fermi rule forbids outright; the decay \({}^{14}\mathrm{O}(0^{+})\to{}^{14}\mathrm{N}^{*}(0^{+})\) is fast and has \(\Delta J=0\) between two \(0^{+}\) states, which the Gamow–Teller rule forbids outright. Gamow and Teller proposed the second coupling in 1936 for exactly this reason [Gamow:1936]: the observed decays do not fit one interaction, and the pair of them is the minimum the data admit.

Definition 101.16 (Comparative half-life).

The measured half-life \(t_{1/2}\) of a beta transition depends on the energy release through the fifth power of Equation (101.14) and on the nuclear structure through the matrix elements. Dividing out the first isolates the second. Define the dimensionless phase-space integral

\begin{equation}\tag{101.16} f\left(Z,E_{0}\right)=\frac{1}{\left(m_{e}c^{2}\right)^{5}} \int_{m_{e}c^{2}}^{E_{0}}F\left(Z,E\right)\,p c\,E\, \left(E_{0}-E\right)^{2}\dd E\ec \end{equation}

and call \(ft_{1/2}\), of dimension \(\mathrm{s}\), the comparative half-life. A small \(ft\) means a fast transition for its energy release, that is a large nuclear matrix element.

Proposition 101.17 (The comparative half-life of an allowed transition).

For an allowed transition,

\begin{equation}\tag{101.17} ft_{1/2}=\frac{K}{G_{F}^{2}\abs{V_{ud}}^{2} \left(\abs{M_{\mathrm{F}}}^{2} +g_{A}^{2}\abs{M_{\mathrm{GT}}}^{2}\right)}\ec\qquad K=\frac{2\pi^{3}\hbar\ln 2\,\left(\hbar c\right)^{6}} {\left(m_{e}c^{2}\right)^{5}}\ec \end{equation}

where \(M_{\mathrm{F}}\) and \(M_{\mathrm{GT}}\) are the Fermi and Gamow–Teller nuclear matrix elements, \(g_{A}\) is the axial coupling of the nucleon and \(V_{ud}\) the quark-mixing element of Section 101.5. Numerically

\begin{equation}\tag{101.18} \frac{K}{\left(\hbar c\right)^{6}} =8.120276\times 10^{-7}\,\mathrm{GeV}^{-4}\,\mathrm{s}\ec\qquad K=1.23058\times 10^{-120}\,\mathrm{J}^{2}\,\mathrm{m}^{6}\,\mathrm{s}\ep \end{equation}

Rests on Proposition 101.11, Equation (101.10) and Equation (101.16).

Proof.

Derives Proposition 101.17. Collect the factors of Proposition 101.11. The rate is the golden-rule expression Equation (101.10) with the squared coupling \(G_{F}^{2}\abs{V_{ud}}^{2}\) in front, the nuclear matrix elements squared and summed over the two spin channels with their multiplicities — one for the singlet, three for the triplet, which is the origin of the factor \(3\) in Equation (101.5) — and the lepton phase space integrated, which is \((m_{e}c^{2})^{5}f(Z,E_{0})\) by Equation (101.16). Half-life and rate are related by \(t_{1/2}=\hbar\ln 2/(\hbar\Gamma)=\ln2/\Gamma\). Assembling the numerical factors of the two-particle phase space, which are \((2\pi)^{-3}\) per particle from the density of states and \(2\pi/\hbar\) from Equation (101.10), leaves \(K=2\pi^{3}\hbar\ln2\,(\hbar c)^{6}/(m_{e}c^{2})^{5}\). Its value follows from \(\hbar=6.582120\times 10^{-25}\,\mathrm{GeV}\,\mathrm{s}\) and \(m_{e}c^{2}=0.51099895\,\mathrm{MeV}\):

\[ \frac{K}{\left(\hbar c\right)^{6}} =\frac{2\pi^{3}\left(\ln 2\right) \left(6.582120\times 10^{-25}\,\mathrm{GeV}\,\mathrm{s}\right)} {\left(5.1099895\times 10^{-4}\,\mathrm{GeV}\right)^{5}} =8.120276\times 10^{-7}\,\mathrm{GeV}^{-4}\,\mathrm{s}\ec \]

and multiplying by \((\hbar c)^{6}=9.98576\times 10^{-154}\,\mathrm{J}^{6}\,\mathrm{m}^{6}\) after converting \(\mathrm{GeV}^{-4}\) to \(\mathrm{J}^{-4}\) gives the SI form in Equation (101.18).

Phenomenon 101.18 (Superallowed decays measure one number).

The half-lives, energy releases and branching ratios of the \(0^{+}\to0^{+}\) beta transitions between members of an isospin triplet — \(^{10}\mathrm{C}\), \(^{14}\mathrm{O}\), \(^{26m}\mathrm{Al}\), \(^{34}\mathrm{Cl}\) and about a dozen others, spanning a factor of \(10^{4}\) in \(f\) — yield, after small and calculable corrections for isospin breaking and radiative effects, a single common value

\begin{equation}\tag{101.19} \mathcal{F}t=3072\,\mathrm{s} \end{equation}

[Navas:2024]. Nuclei of different mass, different structure and different lifetime give the same number to a part in \(10^{4}\). Rests on Proposition 101.15 and Equation (101.17).

Derivation. Derives Phenomenon 101.18. A \(0^{+}\to0^{+}\) transition is pure Fermi by Proposition 101.15: the Gamow–Teller operator cannot connect two spinless states. The Fermi operator is the isospin raising or lowering operator, whose matrix element between members of an isospin multiplet is fixed by the algebra and not by the nuclear wave function at all: for a triplet, \(\abs{M_{\mathrm{F}}}^{2}=2\). Hence Equation (101.17) loses every trace of nuclear structure and reduces to

\begin{equation}\tag{101.20} \mathcal{F}t=\frac{K}{2G_{F}^{2}\abs{V_{ud}}^{2} \left(1+\Delta_{R}^{V}\right)}\ec \end{equation}

where \(\Delta_{R}^{V}\) is the transition-independent radiative correction and \(\mathcal{F}t\) denotes \(ft\) after the small nucleus-dependent corrections have been removed. With \(G_{F}/(\hbar c)^{3}=1.1663787\times 10^{-5}\,/\mathrm{GeV}^{2}\), \(\abs{V_{ud}}=0.97435\) and \(\Delta_{R}^{V}=0.0240\) [Navas:2024],

\[ \mathcal{F}t=\frac{8.120276\times 10^{-7}\,\mathrm{GeV}^{-4}\,\mathrm{s}} {2\left(1.1663787\times 10^{-5}\right)^{2}\left(0.97435\right)^{2} \left(1.0240\right)\,\mathrm{GeV}^{-4}} =3070\,\mathrm{s}\ec \]

against the measured \(3072\,\mathrm{s}\). The agreement to a part in \(10^{3}\) is what turns Equation (101.20) around: read backwards it is the most precise determination of \(\abs{V_{ud}}\) available, and hence the first-row test of the unitarity of the mixing matrix in Section 101.5. The constancy of \(\mathcal{F}t\) across the multiplet is separately the evidence for the conserved vector current of Section 101.4.2: if the strong interaction renormalized the vector coupling, it would do so differently in a carbon nucleus and an aluminium one, and the fourteen numbers would not agree.

Away from the superallowed set the same formula is a measurement of nuclear structure rather than of the coupling. Allowed transitions have \(\log_{10}(ft/\mathrm{s})\) between about \(3\) and \(6\); each unit of forbiddenness — one unit of orbital angular momentum carried by the leptons, costing a factor \((p_{e}R/\hbar)^{2}\) in rate, with \(R\) the nuclear radius — adds between three and four to the logarithm, which is why \(^{40}\mathrm{K}\) lives for \(1.25\times 10^{9}\,\mathrm{yr}\) and \(^{6}\mathrm{He}\) for \(0.8\,\mathrm{s}\). The systematics belong to Nuclear Forces and Nuclear Structure; what belongs here is that the ladder exists and is explained, and that the neutron itself sits on the allowed rung with \(ft\approx1.05\times 10^{3}\,\mathrm{s}\), a Fermi and a Gamow–Teller amplitude combining as \(\abs{M_{\mathrm{F}}}^{2}+g_{A}^{2}\abs{M_{\mathrm{GT}}}^{2} =1+3\lambda^{2}\) with \(\lambda=1.2754\).

The Kurie plot and the neutrino mass

Equation (101.11) is a curve, and curves are hard to compare by eye. Kurie, Richardson and Paxton observed that dividing out the known factors linearizes it [Kurie:1936], turning the comparison of theory with data into the inspection of a straight line — and, more importantly, turning the neutrino mass into a departure from straightness near one end.

Definition 101.19 (Kurie function).

For a measured spectrum \(\dd\Gamma/\dd E_{e}\), the Kurie function is

\begin{equation}\tag{101.21} K\left(E_{e}\right)= \left[\frac{\dd\Gamma/\dd E_{e}} {F\left(Z,E_{e}\right)p_{e}E_{e}}\right]^{1/2}\ep \end{equation}
Proposition 101.20 (The Kurie plot is a straight line if and only if the neutrino is massless).

With Equation (101.11),

\begin{equation}\tag{101.22} K\left(E_{e}\right)\propto \left[\left(E_{0}-E_{e}\right) \sqrt{\left(E_{0}-E_{e}\right)^{2} -m_{\nu}^{2}c^{4}}\right]^{1/2}\ec \end{equation}

which is \(E_{0}-E_{e}\) exactly — a straight line meeting the axis at \(E_{e}=E_{0}\) — if and only if \(m_{\nu}=0\). For \(m_{\nu}\neq0\) the line bends downwards and meets the axis vertically at \(E_{e}=E_{0}-m_{\nu}c^{2}\). Rests on Equations (101.11) and (101.21).

Proof.

Derives Proposition 101.20. Substitute Equation (101.11) into Equation (101.21); the factors \(F\), \(p_{e}\) and \(E_{e}\) cancel identically, leaving Equation (101.22). Writing \(\epsilon=E_{0}-E_{e}\) for the energy left to the neutrino, the bracket is \(\epsilon\sqrt{\epsilon^{2}-m_{\nu}^{2}c^{4}}\), which vanishes at \(\epsilon=m_{\nu}c^{2}\) rather than at \(\epsilon=0\), and whose derivative there is infinite. Hence the endpoint moves down by \(m_{\nu}c^{2}\) — which is Proposition 101.3 recovered from the spectrum instead of from the kinematics — and the approach to it becomes vertical. Both features are confined to the last \(m_{\nu}c^{2}\) of the spectrum, and the fraction of decays falling there is, from Equation (101.12),

\[ \frac{\int_{E_{0}-\Delta}^{E_{0}}\left(E_{0}-E\right)^{2}\dd E} {\int^{E_{0}}\left(E_{0}-E\right)^{2}\dd E} \approx\left(\frac{\Delta}{E_{0}}\right)^{3}\ec \]

so a neutrino mass of \(1\,\mathrm{eV}/c^{2}\) affects only \(2\times 10^{-13}\) of the decays of a nucleus with \(E_{0}=18.6\,\mathrm{keV}\). That cube is the whole difficulty of the measurement and the whole reason for choosing the smallest available \(E_{0}\).

Tritium, \({}^{3}\mathrm{H}\to{}^{3}\mathrm{He}+e^{-}+\bar{\nu}_{e}\), is the laboratory of choice on four counts: its release \(E_{0}=18.6\,\mathrm{keV}\) is among the smallest known, which maximizes the fraction \((\Delta/E_{0})^{3}\); the transition is superallowed, so the rate is high and the matrix element is structureless; the daughter is a two-electron atom whose final-state excitation spectrum can be computed exactly, which matters because molecular final states smear the endpoint by several electronvolts; and the half-life of \(12.3\,\mathrm{yr}\) is short enough for a usable source strength and long enough for the source to be stable.

Phenomenon 101.21 (The neutrino mass is smaller than any other fermion mass by six orders of magnitude).

The endpoint of the tritium beta spectrum bounds the electron antineutrino mass at

\begin{equation}\tag{101.23} m_{\nu}<0.8\,\mathrm{eV}/c^{2}\qquad\left(90\,\mathrm{\%} \text{ confidence}\right) \end{equation}

[Aker:2022]. The next lightest fermion, the electron, has \(m_{e}c^{2}=0.511\,\mathrm{MeV}\): the ratio is below \(1.6\times 10^{-6}\). Rests on Proposition 101.20 and Equation (101.22).

Derivation. Derives Phenomenon 101.21. The bound is Proposition 101.20 applied to a measured spectrum. KATRIN scans the integral spectrum with a magnetic-adiabatic collimation spectrometer of \(1\,\mathrm{eV}\) resolution, fits the counting rate above a variable retarding potential to the predicted shape convolved with the response function and the calculable molecular final-state distribution, and treats \(m_{\nu}^{2}\) as the fit parameter — \(m_{\nu}^{2}\), and not \(m_{\nu}\), because Equation (101.22) depends on the square and because a fit constrained to positive values is biased. The result is consistent with zero, and the interval is converted to the bound Equation (101.23) [Aker:2022]. What the experiment measures is strictly an incoherent average, \(m_{\beta}^{2}=\sum_{i}\abs{U_{ei}}^{2}m_{i}^{2}\), over the mass eigenstates that make up the electron antineutrino, because the individual masses are far too close together to be resolved as separate kinks; the mixing matrix \(U\) is that of Flavour Physics and Neutrinos.

The bound is an upper limit, and an upper limit is compatible with zero. That the masses are not zero is known from a different kind of experiment entirely — neutrino oscillation, which measures differences of squared masses through an interference phase and is insensitive to the absolute scale. Oscillation gives \(\abs{\Delta m^{2}_{32}}\approx2.5\times 10^{-3}\,\mathrm{eV}^{2}/c^{4}\) and hence at least one mass above \(0.05\,\mathrm{eV}/c^{2}\); the apparatus and the analysis are Experiment: Neutrino Oscillations and the theory is Flavour Physics and Neutrinos. The two measurements are complementary and neither replaces the other: the endpoint bounds the scale from above, oscillation bounds it from below, and the gap between \(0.05\,\mathrm{eV}/c^{2}\) and \(0.8\,\mathrm{eV}/c^{2}\) is where the answer lies.

Remark 101.22 (What a nonzero neutrino mass costs this chapter).

Nothing in Section 101.1 or Section 101.2 assumed the neutrino massless. Two later results do, and they are flagged where they occur: the exact identification of chirality with helicity in Proposition 101.32, which acquires a correction of relative size \(m_{\nu}^{2}c^{4}/4E^{2}\), and the two-component description of Section 101.3.3, which is exact only for a massless field. Both corrections are numerically negligible at any energy at which a neutrino has ever been detected — at \(E=1\,\mathrm{MeV}\) and \(m_{\nu}c^{2}=0.8\,\mathrm{eV}\) the correction is \(1.6\times 10^{-13}\) — and both are conceptually decisive, because a Standard Model with massless neutrinos has no room for the oscillations that are observed.

Parity violation

The tau–theta puzzle and the Lee–Yang question

Parity is the operation \(P:\vect{x}\mapsto-\vect{x}\). Its significance is that it is a symmetry of the Coulomb interaction, of the nuclear force as then understood, and of every optical and spectroscopic selection rule that had ever been checked; atomic states carry a definite parity, and transitions between states of the wrong relative parity are absent to the precision of the measurement. The inference universally drawn was that parity is a symmetry of Nature. The inference is a generalization from two forces to all of them, and nobody had tested it in the third.

Proposition 101.23 (Behaviour of the bilinears under parity).

Under the parity operation, implemented on a Dirac field by \(\psi(t,\vect{x})\mapsto\gamma^{0}\psi(t,-\vect{x})\), the covariants of Proposition 101.14 transform as

\begin{align} \bar{\psi}\chi&\longmapsto+\bar{\psi}\chi\ec & \bar{\psi}\gamma^{5}\chi&\longmapsto-\bar{\psi}\gamma^{5}\chi\ec \tag{101.24}\\ \bar{\psi}\gamma^{\mu}\chi&\longmapsto +\bar{\psi}\gamma_{\mu}\chi\ec & \bar{\psi}\gamma^{\mu}\gamma^{5}\chi&\longmapsto -\bar{\psi}\gamma_{\mu}\gamma^{5}\chi\ec \tag{101.25} \end{align}

with the arguments evaluated at the reflected point. Hence a current–current interaction is invariant under \(P\) if and only if each of its two currents has a definite and equal parity type: a product \(V\!\times\!V\) or \(A\!\times\!A\) is even, and a cross term \(V\!\times\!A\) is odd. Rests on Proposition 101.14.

Proof.

Derives Proposition 101.23. Use \(\gamma^{0}\gamma^{\mu}\gamma^{0}=\gamma_{\mu}\) — no sum, and the index is lowered, which is the statement that the time component is unchanged and the space components reverse — together with \((\gamma^{0})^{2}=\identity\) and \(\gamma^{0}\gamma^{5}\gamma^{0}=-\gamma^{5}\), the last because \(\gamma^{5}\) anticommutes with \(\gamma^{0}\). Then \(\bar{\psi}\chi=\psi^{\dagger}\gamma^{0}\chi\mapsto \psi^{\dagger}\gamma^{0}\gamma^{0}\gamma^{0}\chi=\bar{\psi}\chi\), and \(\bar{\psi}\gamma^{5}\chi\mapsto \psi^{\dagger}\gamma^{0}\gamma^{0}\gamma^{5}\gamma^{0}\chi =-\bar{\psi}\gamma^{5}\chi\), and similarly for the two vectors. For the last statement, contract two currents: the product of two quantities that both pick up \(+1\), or both \(-1\), is invariant, while a product of one of each changes sign. A theory built only from \(V\times V\) and \(A\times A\) conserves parity; one containing \(V\times A\) does not.

Proposition 101.23 says that a parity-violating interaction is not exotic — it is simply one whose two currents are not both of the same type — and that testing parity means looking for the interference between the even and the odd pieces. Such an interference shows up as a pseudoscalar observable: a quantity built from measured vectors that changes sign under reflection, such as \(\avg{\vect{J}}\cdot\hat{\vect{p}}\), the projection of a momentum on a spin. Before 1956 no experiment on a weak process had ever measured one.

What forced the question was a puzzle in the cosmic-ray and accelerator data. Two particles were seen, called at the time \(\theta^{+}\) and \(\tau^{+}\) — the second has nothing to do with the tau lepton, which was discovered twenty years later — decaying as

\[ \theta^{+}\longrightarrow\pi^{+}\pi^{0}\ec\qquad \tau^{+}\longrightarrow\pi^{+}\pi^{+}\pi^{-}\ep \]

As the measurements improved, the two came to have the same mass and the same lifetime, to a precision at which a coincidence became absurd. But their decay products have opposite parity.

Proposition 101.24 (The parities of the two final states are opposite).

A state of \(n\) pions in relative \(s\) waves has parity \((-1)^{n}\). Hence the \(\pi\pi\) final state has parity \(+1\) and the \(\pi\pi\pi\) final state parity \(-1\), and no single particle of definite parity can decay to both if parity is conserved. Rests on Phenomenon 105.5 and Proposition 105.15.

Proof.

Derives Proposition 101.24. The pion is a pseudoscalar, \(J^{P}=0^{-}\), established from the capture of a negative pion at rest on deuterium: the capture proceeds from an atomic \(s\) state to \(nn\), the two neutrons must be in a state antisymmetric overall, and the only such state with \(J=1\) is \(^{3}P_{1}\), of orbital parity \(-1\); equating initial and final parities gives \(P_{\pi}=-1\). The parity of a system of \(n\) such particles with no relative orbital angular momentum is the product of the intrinsic parities, \((-1)^{n}\). For \(n=2\) this is \(+1\); for \(n=3\), \(-1\). If the decaying particle has a definite parity \(P_{X}\) and parity is conserved, then \(P_{X}=+1\) from the first mode and \(P_{X}=-1\) from the second, a contradiction. The escape that the three-pion state carries orbital angular momentum, and so is not in relative \(s\) waves, is closed by the observed distribution of the three pions over their kinematically allowed region: the density is uniform up to the edges, which is the signature of \(\ell=0\), since any nonzero orbital angular momentum would suppress the configurations in which two pions are collinear.

The way out was proposed by Lee and Yang in 1956 [Lee:1956], and the paper is remarkable less for what it proposed than for what it established first. They surveyed the entire experimental literature of weak decays and found that no measurement had ever been sensitive to a pseudoscalar observable — that the parity invariance of the weak interaction was not a tested fact but an untested assumption inherited from the strong and electromagnetic sectors. Having established that, they listed experiments that could decide, of which the most direct was to align nuclear spins at low temperature and look for a front–back asymmetry in the emitted electrons.

Remark 101.25 (The methodological point).

This is worth stating flatly because it recurs throughout physics: a symmetry can hold to extraordinary precision in one interaction and fail completely in another, and the only way to know is to measure it where it matters. The strong and electromagnetic interactions conserve parity, in the second case to a relative precision better than \(10^{-8}\); the weak interaction violates it maximally. Nothing about the first two facts implies anything about the third, and the inference that was drawn for twenty-five years was a habit, not an argument. The same lesson is recorded for \(CP\) in Discrete Symmetries and CPT and for \(CPT\), which unlike the others is a theorem and not an assumption.

Observation

Wu and her collaborators performed the experiment within months [Wu:1957]. A thin layer of \(^{60}\mathrm{Co}\) was polarized by adiabatic demagnetization of a cerium magnesium nitrate crystal at a few millikelvin, the degree of polarization being monitored by the anisotropy of the gamma rays from the daughter, and the beta electrons counted in an anthracene scintillator whose light was piped out of the cryostat. Reversing the polarizing field reverses \(\avg{\vect{J}}\) and so must, if parity holds, leave the counting rate unchanged. It did not: the rate changed by tens of per cent, and the electrons came out preferentially into the hemisphere opposite to the nuclear spin. Experiment: Parity Violation is the experiment chapter this measurement belongs to. It states the result and proves the parity argument in full; its apparatus, procedure and systematic-check sections — among them the warming curve that shows the asymmetry dying with the polarization — carry their headings and their sources but are not yet written, so the sketch above is for the moment the fuller of the two accounts.

Garwin, Lederman and Weinrich obtained the same conclusion within days by a completely different route [Garwin:1957]. Stopping positive muons from the decay \(\pi^{+}\to\mu^{+}\nu_{\mu}\) in carbon and observing the decay positrons, they found the positrons emitted asymmetrically about the muon spin direction — and, since the muons had been polarized by the pion decay itself, the same data show parity violation twice over, once in \(\pi\to\mu\) and once in \(\mu\to e\). They also measured the muon magnetic moment in the same apparatus by precessing the spin, which is the ancestor of the experiment in Experiment: The Electron Anomalous Magnetic Moment.

Phenomenon 101.26 (Parity is violated in weak decays).

In the beta decay of polarized \(^{60}\mathrm{Co}\) nuclei the electrons are emitted preferentially into the hemisphere opposite to the nuclear spin. The angular distribution has the form

\begin{equation}\tag{101.26} W(\theta)\propto 1+A\,\frac{v}{c}\, \frac{\langle\vect{J}\rangle\cdot\hat{\vect{p}}} {\abs{\langle\vect{J}\rangle}}\ec \end{equation}

with \(A\) negative and of order unity [Wu:1957], and the same asymmetry appears twice over in the chain \(\pi\to\mu\to e\) [Garwin:1957]. The apparatus is Experiment: Parity Violation. No comparable asymmetry occurs in any strong or electromagnetic process. Rests on Proposition 101.23.

Derivation. Derives Phenomenon 101.26. Under the parity operation a polar vector such as the momentum \(\vect{p}\) reverses while an axial vector such as the angular momentum \(\vect{J}\) does not, so the scalar product \(\langle\vect{J}\rangle\cdot\hat{\vect{p}}\) appearing in Equation (101.26) is a pseudoscalar and changes sign. The mirror-image experiment therefore predicts \(W(\pi-\theta)\) wherever the original predicts \(W(\theta)\), and the two disagree unless \(A=0\). A measured \(A\neq0\) is thus in itself a proof that the interaction responsible is not invariant under \(P\): no dynamical input whatever is required, which is why one asymmetry measurement settled a question that decades of spectroscopy had not thought to ask [Lee:1956] [Wu:1957]. What the further observation \(\abs{A}\approx1\) adds is that the violation is as large as the kinematics permits — the interaction does not merely prefer one handedness, it uses one and ignores the other — and that statement is delivered by Proposition 101.27 below, once the \(V-A\) structure of Section 101.4 is in place.

Proposition 101.27 (The sign and size of the cobalt asymmetry).

For the pure Gamow–Teller transition \(^{60}\mathrm{Co}\,(5^{+})\to{}^{60}\mathrm{Ni}\,(4^{+})\), an interaction that emits only left-chiral electrons and right-chiral antineutrinos gives \(A=-1\) in Equation (101.26): the electrons are emitted opposite to the nuclear spin, with an asymmetry equal in magnitude to the electron's speed in units of \(c\). Rests on Equation (101.26), Proposition 101.32 and Proposition 101.34.

Proof.

Derives Proposition 101.27. The transition lowers the nuclear angular momentum by one unit along the polarization axis \(\hat{\vect{z}}\), so the lepton pair must carry away \(J_{z}=+1\): both leptons have their spin projections \(+\tfrac {1}{2}\) along \(\hat{\vect{z}}\). Now apply the chirality rule. The antineutrino, being massless to the accuracy of Phenomenon 101.21, has positive helicity exactly by Proposition 101.32, so its spin is along its momentum: with \(S_{z}=+\tfrac{1}{2}\) it must travel along \(+\hat{\vect{z}}\). The electron is emitted in the left-chiral combination, whose helicity expectation value is \(-v/c\) by Proposition 101.34; with its spin also along \(+\hat{\vect{z}}\), a negative helicity means momentum along \(-\hat{\vect{z}}\). The electron therefore prefers the hemisphere opposite the nuclear spin, and the preference is complete only in the limit \(v\to c\), being diluted by the factor \(v/c\) at lower speeds. Writing the distribution as Equation (101.26) identifies \(A=-1\) for this transition.

Comparing that with Wu's number requires one factor Equation (101.26) does not contain, and it is worth being careful, because the arithmetic does not otherwise close. The distribution is written for a fully polarized ensemble; a real source has a degree of nuclear polarization \(f=\abs{\avg{\vect{J}}}/J <1\), and only that fraction of the nuclei contributes to the front–back difference, so the measured counting-rate asymmetry between the two senses of the polarizing field is

\begin{equation}\tag{101.27} \varepsilon=\frac{W(0)-W(\pi)}{W(0)+W(\pi)} =A\,\frac{v}{c}\,f\ep \end{equation}

Wu's electrons had \(v/c\approx0.6\) and the source reached \(f\approx0.6\), giving \(\abs{\varepsilon}\approx0.36\), which is the observed asymmetry of about \(0.4\). The product \(\abs{A}\,v/c\) alone would predict \(0.6\), half again too large: the polarization factor is not optional, and it is why the experiment had to measure \(f\) independently, from the anisotropy of the daughter's gamma cascade. The same factor supplies the control — the asymmetry disappears as the crystal warms and \(f\) decays to zero, on the same time constant as the gamma anisotropy, which is what makes the result an asymmetry and not an instrumental drift [Wu:1957].

Two further consequences of the same data deserve separate statement.

The first is that charge conjugation is violated as well. If \(C\) were a symmetry, the mirror process for antimatter would be obtained by conjugating charges; the \(\pi^{+}\to\mu^{+}\) and \(\pi^{-}\to\mu^{-}\) chains would then show asymmetries of the same sign. They show opposite ones [Garwin:1957]. What survives is the combination \(CP\): the mirror image of a left-handed neutrino is a right-handed neutrino, which does not exist, but the \(CP\) image is a right-handed antineutrino, which does. That \(CP\) is itself violated, by a small amount and in a different sector, is the subject of Discrete Symmetries and CPT and Experiment: CP Violation.

The second is quantitative and constrains the theory sharply. A parity-violating interaction has, in general, both a \(P\)-even and a \(P\)-odd part, and the size of the asymmetry measures their relative magnitude. An asymmetry of order unity — not of order \(10^{-2}\), and not of order \(10^{-6}\) — says that the two parts are of equal size. An interaction whose \(P\)-even and \(P\)-odd parts are equal is one that couples to a single chirality and not at all to the other. That is the empirical origin of the \(V-A\) structure, and it was read off the data in exactly this way.

Neutrino helicity

The most direct statement the parity data suggest is that the neutrino comes in one handedness only. Landau, Lee and Yang, and Salam each proposed independently in 1957 that the neutrino is described by a two-component field — the Weyl equation of The Weyl Equation and Neutrinos — rather than by a four-component Dirac field with one half of it inert [Landau:1957] [Lee:1957] [Salam:1957]. A two-component field has no mass term available to it and no partner of the opposite handedness, so the proposal makes handedness a defining property rather than a dynamical accident, and it predicts a definite sign which experiment can check.

Checking it means measuring the helicity of a particle that is never detected. Goldhaber, Grodzins and Sunyar found a way [Goldhaber:1958]: in the orbital electron capture \(^{152m}\mathrm{Eu}+e^{-}\to{}^{152}\mathrm{Sm}^{*}+\nu_{e}\) followed by \(^{152}\mathrm{Sm}^{*}\to{}^{152}\mathrm{Sm}+\gamma\), the photons emitted opposite to the neutrino carry the neutrino's helicity, and their circular polarization can be measured by transmission through magnetized iron. Selecting those photons is done by resonant scattering, which only the ones emitted against the neutrino have enough energy to undergo. Experiment: Neutrino Helicity is the experiment chapter for this measurement; it names the decay scheme, the magnetized-iron analyser and the resonant scatterer, and its apparatus and systematic-check sections are still owed. The angular-momentum argument is short enough to give here, and is.

Phenomenon 101.28 (Neutrinos are left-handed).

The neutrino emitted in orbital electron capture on \(^{152}\mathrm{Eu}\) has negative helicity — its spin points opposite to its momentum — with an uncertainty that leaves no room for an appreciable admixture of the other state [Goldhaber:1958]; the antineutrino emitted in beta decay has positive helicity. No right-handed neutrino, and no left-handed antineutrino, has ever been observed to take part in a weak interaction. Rests on Definition 101.29.

Derivation. Derives Phenomenon 101.28. What has to be established is that the photon the apparatus selects carries the neutrino's helicity, since the photon's circular polarization is the only thing measured. Both steps are angular-momentum bookkeeping along a single axis.

Take \(\hat{\vect{z}}\) along the photon's momentum. The capture \(^{152m}\mathrm{Eu}\,(0^{-})+e^{-}\to{}^{152}\mathrm{Sm}^{*}(1^{-}) +\nu_{e}\) proceeds from a \(K\)-shell \(s\) orbital, so nothing brings orbital angular momentum in, and the two products leave back to back, the neutrino along \(-\hat{\vect{z}}\) and the recoiling \(^{152}\mathrm{Sm}^{*}\) along \(+\hat{\vect{z}}\). A two-body state collinear with the quantization axis has \(L_{z}=0\), so along that axis

\begin{equation}\tag{101.28} S_{z}^{e}=S_{z}^{\nu}+J_{z}^{\mathrm{Sm}^{*}}\ep \end{equation}

Take \(S_{z}^{e}=-\tfrac{1}{2}\). A negative-helicity neutrino moving along \(-\hat{\vect{z}}\) has its spin along \(+\hat{\vect{z}}\), that is \(S_{z}^{\nu}=+\tfrac{1}{2}\), and Equation (101.28) then forces \(J_{z}^{\mathrm{Sm}^{*}}=-1\). The excited state decays to a spinless ground state, \(^{152}\mathrm{Sm}^{*}(1^{-})\to {}^{152}\mathrm{Sm}\,(0^{+})+\gamma\), so the photon carries away the whole of \(J_{z}=-1\); and a photon travelling along \(+\hat{\vect{z}}\) with \(J_{z}=-1\) has helicity \(-1\). Reversing the neutrino's helicity reverses every sign in the chain, so the photon's helicity is not merely correlated with the neutrino's but equal to it.

That fixes the photons emitted opposite to the neutrino, and the resonance condition selects exactly those and no others. A photon emitted by a nucleus at rest falls short of the absorption resonance by the recoil energies of emitter and absorber; only a nucleus recoiling towards the emission direction — that is, one whose neutrino went the other way — is moving fast enough for the Doppler shift to make the deficit up, so only those photons scatter resonantly on a \(^{152}\mathrm{Sm}\) target. Their circular polarization, measured by transmission through magnetized iron, is therefore the neutrino's helicity, and it came out negative [Goldhaber:1958]. The antineutrino statement follows from \(CP\): the \(CP\) image of a negative-helicity neutrino is a positive-helicity antineutrino, and \(CP\) survives these decays even though \(C\) and \(P\) separately do not.

To say what such a measurement establishes, and what it does not, the two notions it is easy to conflate must be separated. They are different in kind: one is a property of a state, the other a property of a field; one is conserved but frame-dependent, the other frame-independent but not conserved; and they coincide only in a limit.

Definition 101.29 (Helicity).

The helicity operator for a spin-\(\tfrac{1}{2}\) particle of momentum \(\vect{p}\) is

\begin{equation}\tag{101.29} h=\frac{\vect{\Sigma}\cdot\vect{p}}{\abs{\vect{p}}}\ec\qquad \vect{\Sigma}=\begin{pmatrix} \vect{\sigma}&0\\0&\vect{\sigma}\end{pmatrix}\ec \end{equation}

with eigenvalues \(\pm1\). It is the projection of the spin on the direction of motion, and it is a genuine observable: both factors are measurable.

Definition 101.30 (Chirality).

The chirality operator is \(\gamma^{5} =\ii\gamma^{0}\gamma^{1}\gamma^{2}\gamma^{3}\), with \((\gamma^{5})^{2}=\identity\) and \(\acomm{\gamma^{5}}{\gamma^{\mu}}=0\). Its eigenvalues are \(\pm1\) and the associated projectors are

\begin{equation}\tag{101.30} P_{\mathrm{L}}=\frac{1}{2}\left(\identity-\gamma^{5}\right)\ec\qquad P_{\mathrm{R}}=\frac{1}{2}\left(\identity+\gamma^{5}\right)\ec \end{equation}

so that \(\psi=\psi_{\mathrm{L}}+\psi_{\mathrm{R}}\) with \(\psi_{\mathrm{L}}=P_{\mathrm{L}}\psi\). Chirality is a property of the field, defined without reference to any momentum.

Proposition 101.31 (Chirality against helicity).

Chirality is Lorentz invariant but not conserved; helicity is conserved in free motion but not Lorentz invariant. Precisely:

  1. \(\gamma^{5}\) commutes with the generators \(\sigma^{\mu\nu}\) of the spinor representation, so \(P_{\mathrm{L}}\) and \(P_{\mathrm{R}}\) commute with every Lorentz transformation: chirality is frame-independent.

  2. A Dirac mass term couples the two chiralities, \(\bar{\psi}\psi=\bar{\psi}_{\mathrm{L}}\psi_{\mathrm{R}} +\bar{\psi}_{\mathrm{R}}\psi_{\mathrm{L}}\), so chirality is conserved by the free equation of motion only if \(m=0\).

  3. \(h\) commutes with the free Dirac Hamiltonian, so helicity is conserved in free motion.

  4. For \(m\neq0\) a particle moves slower than light, so an observer may overtake it; in the overtaking frame \(\vect{p}\) reverses and \(\vect{\Sigma}\) does not, and the helicity changes sign. Helicity is frame-dependent.

Rests on Definitions 101.29 and 101.30.

Proof.

Derives Proposition 101.31. (i) \(\sigma^{\mu\nu}\) is a product of two distinct gamma matrices, and \(\gamma^{5}\) anticommutes with each, hence commutes with the pair; therefore \(\comm{\gamma^{5}}{\sigma^{\mu\nu}}=0\) and \(\gamma^{5}\) commutes with \(\exp(-\tfrac{\ii}{4}\omega_{\mu\nu}\sigma^{\mu\nu})\). (ii) \(\bar{\psi}\psi=\bar{\psi}(P_{\mathrm{L}}+P_{\mathrm{R}})\psi\), and \(\bar{\psi}P_{\mathrm{L}}=\overline{P_{\mathrm{R}}\psi}\) because \(\gamma^{0}\gamma^{5}=-\gamma^{5}\gamma^{0}\); the diagonal terms therefore vanish and only the cross terms survive. (iii) The free Hamiltonian is \(H=c\vect{\alpha}\cdot\vect{p}+\beta mc^{2}\) with \(\vect{\alpha}=\gamma^{0}\vect{\gamma}\); a direct computation using \(\comm{\Sigma^{i}}{\alpha^{j}}=2\ii\varepsilon^{ijk}\alpha^{k}\) and \(\comm{\vect{\Sigma}}{\beta}=0\) gives \(\comm{\vect{\Sigma}\cdot\vect{p}}{H}=0\). (iv) A boost of speed \(u>v\) along \(\hat{\vect{p}}\) sends \(\vect{p}\mapsto-\abs{\vect{p}'} \hat{\vect{p}}\) while the spin projection on \(\hat{\vect{p}}\) is unchanged, since a boost along an axis does not rotate spin about that axis; hence \(h\mapsto-h\). For \(m=0\) no such boost exists and the obstruction disappears.

The two therefore cannot be the same thing. The following proposition says exactly how close they are, and it is the quantitative content of every statement in this chapter of the form “only left-handed particles participate”.

Proposition 101.32 (Chirality and helicity in a state of definite momentum).

Let \(u_{\lambda}(p)\) be a positive-energy Dirac spinor of momentum \(\vect{p}\), energy \(E\) and helicity \(\lambda=\pm1\). Then

\begin{equation}\tag{101.31} \frac{\abs{P_{\mathrm{L}}u_{\lambda}}^{2}}{\abs{u_{\lambda}}^{2}} =\frac{1}{2}\left(1-\lambda\beta\right)\ec\qquad \beta=\frac{\abs{\vect{p}}c}{E}=\frac{v}{c}\ep \end{equation}

In particular, for \(\beta\to1\) the left-chiral projection of a positive-helicity state vanishes as

\begin{equation}\tag{101.32} \frac{1}{2}\left(1-\beta\right) \simeq\frac{m^{2}c^{4}}{4E^{2}}\ec \end{equation}

so that chirality and helicity agree exactly in the massless limit and to relative accuracy \(m^{2}c^{4}/4E^{2}\) otherwise. Rests on Definition 101.29 and Equation (101.30).

Proof.

Derives Proposition 101.32. Work in the Dirac representation, in which \(\gamma^{5}=\begin{pmatrix}0&\identity\\ \identity&0\end{pmatrix}\) in \(2\times2\) blocks. A positive-energy solution of definite helicity is

\begin{equation}\tag{101.33} u_{\lambda}(p)=N\begin{pmatrix} \sqrt{E+mc^{2}}\;\chi_{\lambda}\\ \lambda\sqrt{E-mc^{2}}\;\chi_{\lambda}\end{pmatrix}\ec\qquad \left(\vect{\sigma}\cdot\hat{\vect{p}}\right)\chi_{\lambda} =\lambda\chi_{\lambda}\ec \end{equation}

as follows from the Dirac equation with the lower two components given by \(c\,\vect{\sigma}\cdot\vect{p}/(E+mc^{2})\) times the upper ones and \(c\abs{\vect{p}}=\sqrt{(E+mc^{2})(E-mc^{2})}\). Its norm is \(\abs{u_{\lambda}}^{2}=N^{2}\left[(E+mc^{2})+(E-mc^{2})\right] =2N^{2}E\). Applying \(P_{\mathrm{L}}=\tfrac{1}{2}\begin{pmatrix} \identity&-\identity\\-\identity&\identity\end{pmatrix}\),

\[ P_{\mathrm{L}}u_{\lambda}=\frac{N}{2} \left(\sqrt{E+mc^{2}}-\lambda\sqrt{E-mc^{2}}\right) \begin{pmatrix}\chi_{\lambda}\\-\chi_{\lambda}\end{pmatrix}\ec \]

whose squared norm is

\[ \abs{P_{\mathrm{L}}u_{\lambda}}^{2} =\frac{N^{2}}{2}\left(\sqrt{E+mc^{2}} -\lambda\sqrt{E-mc^{2}}\right)^{2} =N^{2}\left(E-\lambda\sqrt{E^{2}-m^{2}c^{4}}\right) =N^{2}E\left(1-\lambda\beta\right)\ec \]

using \(\lambda^{2}=1\) and \(\sqrt{E^{2}-m^{2}c^{4}}=\abs{\vect{p}}c=\beta E\). Dividing gives Equation (101.31). For the limit, expand \(\beta=\sqrt{1-m^{2}c^{4}/E^{2}}\simeq1-m^{2}c^{4}/2E^{2}\).

Corollary 101.33 (What Goldhaber's experiment measures).

For a neutrino of mass below \(0.8\,\mathrm{eV}/c^{2}\) (Phenomenon 101.21) emitted with an energy near \(0.9\,\mathrm{MeV}\), as in the europium capture, the wrong-helicity admixture in a purely left-chiral coupling is, by Equation (101.32), at most

\[ \frac{m_{\nu}^{2}c^{4}}{4E^{2}} <\frac{\left(0.8\,\mathrm{eV}\right)^{2}} {4\left(0.9\,\mathrm{MeV}\right)^{2}} =2\times 10^{-13}\ep \]

A measurement of the helicity is therefore, to that accuracy, a measurement of the chirality of the coupling — which is the quantity the theory contains and the experiment cannot see directly. Rests on Phenomenon 101.21 and Equation (101.32).

Proposition 101.34 (The longitudinal polarization of a beta electron).

An interaction that couples only to \(\psi_{\mathrm{L}}\) emits electrons with mean helicity

\begin{equation}\tag{101.34} \avg{h}=-\frac{v}{c}\ec \end{equation}

and positrons with mean helicity \(+v/c\). Rests on Equation (101.31).

Proof.

Derives Proposition 101.34. By Equation (101.31) the left-chiral projection contains helicity \(-1\) with weight \(\tfrac{1}{2}(1+\beta)\) and helicity \(+1\) with weight \(\tfrac{1}{2}(1-\beta)\), and these are the relative probabilities with which the interaction produces the two. Hence \(\avg{h}=\tfrac{1}{2}(1-\beta)-\tfrac{1}{2}(1+\beta)=-\beta\). For the antiparticle the roles of the projectors are exchanged, because the negative-energy spinor \(v_{\lambda}\) carries the opposite relation between chirality and helicity, and the sign reverses. Equation (101.34) is a prediction with no free parameter: it says that a beta electron of \(v/c=0.5\) is \(50\) per cent longitudinally polarized and one of \(v/c=0.9\) is \(90\) per cent polarized, and it is measured directly, by passing the emitted electrons through a polarization-sensitive scatterer. Frauenfelder and collaborators did it for \(^{60}\mathrm{Co}\) within months of Wu's result and found the electrons longitudinally polarized, with the predicted sign and a magnitude consistent with \(-v/c\) [Frauenfelder:1957]. It is also what supplies the factor \(v/c\) in Equation (101.26), which is therefore not an empirical fudge but the electron mass making itself felt.

Finally, the same arithmetic explains a rate that would otherwise be inexplicable. The charged pion decays to a muon and not to an electron, although the electronic channel has far more phase space, because the pion is spinless and forces the charged lepton into the helicity that the left-chiral coupling suppresses by \(\tfrac{1}{2}(1-\beta)\); the suppression is governed by the lepton mass and therefore hurts the electron far more than the muon. The derivation and the measured ratio \(1.23\times 10^{-4}\) are Phenomenon 96.3. It is the sharpest confirmation of Proposition 101.32 available, because it turns a factor of \((m_{e}/m_{\mu})^{2}\) into a directly counted branching ratio.

The $V-A$ interaction

Chirality and the left-handed current

Section 101.3 leaves the interaction determined up to ten coefficients, and the observations of maximal asymmetry and of a purely left-handed neutrino determine them. Feynman and Gell-Mann [Feynman:1958] and, independently and in the same year, Sudarshan and Marshak [Sudarshan:1958] proposed that the answer is

\begin{equation}\tag{101.35} J^{\mu}=\bar{\psi}\,\gamma^{\mu} \left(\identity-\gamma^{5}\right)\chi\ec \end{equation}

one combination out of the five, with the vector and axial pieces of equal magnitude and opposite sign. Every charged-current process in the Standard Model is built from currents of this form.

Proposition 101.35 (The $V-A$ current is the left-chiral current).
\begin{equation}\tag{101.36} \bar{\psi}\gamma^{\mu}\left(\identity-\gamma^{5}\right)\chi =2\,\bar{\psi}_{\mathrm{L}}\gamma^{\mu}\chi_{\mathrm{L}}\ec \end{equation}

so an interaction built from Equation (101.35) involves the right-chiral components of the fields not at all. Rests on Equations (101.30) and (101.35).

Proof.

Derives Proposition 101.35. From Equation (101.30), \(P_{\mathrm{L}}^{2}=P_{\mathrm{L}}\) and \(\gamma^{\mu}P_{\mathrm{L}}=P_{\mathrm{R}}\gamma^{\mu}\), the latter because \(\gamma^{5}\) anticommutes with \(\gamma^{\mu}\). Also \(\overline{P_{\mathrm{L}}\psi}=\psi^{\dagger}P_{\mathrm{L}}\gamma^{0} =\bar{\psi}P_{\mathrm{R}}\), using \(P_{\mathrm{L}}^{\dagger}=P_{\mathrm{L}}\) and \(\gamma^{0}\gamma^{5}=-\gamma^{5}\gamma^{0}\). Hence

\[ \bar{\psi}_{\mathrm{L}}\gamma^{\mu}\chi_{\mathrm{L}} =\bar{\psi}P_{\mathrm{R}}\gamma^{\mu}P_{\mathrm{L}}\chi =\bar{\psi}\gamma^{\mu}P_{\mathrm{L}}P_{\mathrm{L}}\chi =\bar{\psi}\gamma^{\mu}P_{\mathrm{L}}\chi =\frac{1}{2}\bar{\psi}\gamma^{\mu} \left(\identity-\gamma^{5}\right)\chi\ep \]

The same computation shows that the scalar bilinear does the opposite: \(\bar{\psi}_{\mathrm{L}}\chi_{\mathrm{L}} =\bar{\psi}P_{\mathrm{R}}P_{\mathrm{L}}\chi=0\), so a chirally projected theory admits vector and axial currents and forbids scalar, pseudoscalar and tensor couplings between fields of the same chirality. Equation (101.35) is therefore not one arbitrary choice among five: it is the only current a theory of purely left-chiral fields can have.

Proposition 101.36 (Maximal parity violation is exactly $V-A$).

Write a general vector-and-axial current as \(\bar{\psi}\gamma^{\mu}(g_{V}-g_{A}\gamma^{5})\chi\). Then the \(P\)-violating interference between the two, measured by any pseudoscalar observable, is proportional to \(g_{V}g_{A}\) and the \(P\)-conserving part to \(g_{V}^{2}+g_{A}^{2}\). The ratio of the two is extremal, and equal to unity, if and only if \(\abs{g_{V}}=\abs{g_{A}}\). Rests on Proposition 101.23 and Equation (101.35).

Proof.

Derives Proposition 101.36. By Proposition 101.23 the \(VV\) and \(AA\) terms of the squared amplitude are \(P\)-even and the \(VA\) cross terms \(P\)-odd, and the coefficients are \(g_{V}^{2}\), \(g_{A}^{2}\) and \(2g_{V}g_{A}\) respectively. The asymmetry is therefore \(2g_{V}g_{A}/(g_{V}^{2}+g_{A}^{2})\), which by the arithmetic–geometric mean inequality is at most \(1\) in modulus with equality if and only if \(g_{V}=\pm g_{A}\). The measured asymmetries of Section 101.3.2 are of order unity, which forces the equality, and the observed sign — electrons opposite to the nuclear spin, neutrinos left-handed — selects \(g_{A}=+g_{V}\) in the convention of Equation (101.35), that is \(V-A\) and not \(V+A\).

Remark 101.37 (Where the coefficients actually came from).

The account above is the logical order, not the historical one. The coefficients were fixed by a decade of experiments on the angular correlation between the electron and the recoiling nucleus, the observable most sensitive to the Lorentz structure and the one that needs no polarization at all. For an allowed transition it has the form

\begin{equation}\tag{101.37} \frac{\dd\Gamma}{\dd\Omega_{e\nu}}\propto 1+a\,\frac{v_{e}}{c}\cos\theta_{e\nu}\ec \end{equation}

with \(\theta_{e\nu}\) the angle between the electron and antineutrino momenta, inferred in practice from the recoil of the daughter since the antineutrino is not seen. The coefficient \(a\) takes a different value for each of the four couplings, so one number decides among them:

couplingtransition type$a$
scalar (S)Fermi$-1$
vector (V)Fermi$+1$
tensor (T)Gamow–Teller$+1/3$
axial (A)Gamow–Teller$-1/3$

The values follow from Proposition 101.32 and angular momentum alone; the derivation is Proposition 101.38.

The history is cautionary, and is recorded because the answer was wrong for five years. Feynman and Gell-Mann list the measurements their proposal contradicted — among them a recoil experiment on \(^{6}\mathrm{He}\), a pure Gamow–Teller decay, whose reported correlation favoured tensor over axial — and say in print that they believe those experiments to be wrong [Feynman:1958]. They were, and the \(^{6}\mathrm{He}\) value was later revised to the axial \(-1/3\); but the papers reporting and revising it are not in this bibliography, so what is said here rests on Feynman and Gell-Mann's account of them and not on the sources. The lesson is not that theory ran ahead of experiment: it is that one measurement of a small correlation, with an uncontrolled systematic, decided a structural question for five years, and that the theorists who disbelieved it did so on the strength of other data — the maximal asymmetries.

Proposition 101.38 (The electron–neutrino correlation coefficients).

For a \(V-A\) interaction the coefficient \(a\) in Equation (101.37) is \(+1\) for a Fermi transition and \(-1/3\) for a Gamow–Teller transition; for the parity-conjugate choices \(S\) and \(T\) the signs reverse. Rests on Equation (101.37), Proposition 101.32 and Proposition 101.35.

Proof.

Derives Proposition 101.38. Work in the limit \(v_{e}\to c\), where by Proposition 101.32 the electron emitted by a left-chiral coupling has helicity \(-1\) exactly and the antineutrino has helicity \(+1\) exactly; the factor \(v_{e}/c\) in Equation (101.37) is what restores the general case, being the electron's helicity by Equation (101.34). Then the spin of the antineutrino points along its momentum and the spin of the electron points against its own.

A Fermi transition emits the pair in the spin singlet, whose two spins are antiparallel. Let the antineutrino spin be \(+\hat{\vect{z}}\) and the electron spin \(-\hat{\vect{z}}\). The antineutrino momentum is then along \(+\hat{\vect{z}}\) and the electron momentum, being opposite its spin, is also along \(+\hat{\vect{z}}\): the two particles go the same way, \(\cos\theta_{e\nu}=+1\), and the distribution is \(1+\cos\theta_{e\nu}\), that is \(a=+1\).

A Gamow–Teller transition emits the pair in the spin triplet, and the three substates must be averaged because an unpolarized nucleus populates them equally. For \(m=+1\) both spins are \(+\hat{\vect{z}}\); the antineutrino goes along \(+\hat{\vect{z}}\) and the electron along \(-\hat{\vect{z}}\), giving \(\cos\theta_{e\nu}=-1\); the same for \(m=-1\). The state \(m=0\) is \((\ket{\uparrow\downarrow}+\ket{\downarrow\uparrow})/\sqrt{2}\), which is the singlet's partner with antiparallel spins and gives \(\cos\theta_{e\nu}=+1\). Averaging, \(a=(-1-1+1)/3=-1/3\).

Replacing \(V\) by \(S\), or \(A\) by \(T\), replaces the left-chiral projector on the electron by the right-chiral one — this is Proposition 101.35 read backwards, a scalar bilinear connecting opposite chiralities — so the electron's momentum reverses relative to its spin and every sign above flips. The four entries of the table follow.

Remark 101.39 (Why the spectrum shape did not change).

Proposition 101.11 was derived from Fermi's vector coupling and survives its replacement by \(V-A\) unaltered, which is why the 1934 spectrum shape remained good evidence after 1957. The reason is that the shape came entirely from phase space, the matrix element being momentum-independent for an allowed transition; and \(V-A\), like \(V\), has a momentum-independent allowed matrix element. What changes is the correlations — the asymmetry with respect to a polarization, the electron's own polarization, the electron–neutrino angle — none of which anybody had measured before 1956. A theory can be right about everything that has been measured and wrong about its structure, and the way to tell is to measure something else.

Conserved vector current and universality

Two facts about Equation (101.35) are surprising once the strong interaction is remembered. The nucleon is not an elementary field but a bound state of strongly interacting constituents, so the current a \(W\) sees is not the quark current but a dressed object; one would expect the dressing to change the coupling by a factor of order one, and to change it differently for the vector and the axial parts, and differently again for a nucleon and a nucleus. In fact the vector coupling is not renormalized at all, and the axial one is renormalized by only \(27\) per cent.

Phenomenon 101.40 (The vector coupling is not renormalized).

The vector coupling constant extracted from superallowed nuclear beta decay, where the decaying object is a bound state of strongly interacting quarks, agrees with the coupling extracted from muon decay, where nothing strong is involved, once the quark-mixing factor of Section 101.5 is applied. The axial coupling, by contrast, is renormalized to \(\lambda=g_{A}/g_{V}=1.2754(13)\) [Navas:2024]. Rests on Proposition 100.6 and Phenomenon 101.18.

Derivation. Derives Phenomenon 101.40. The hypothesis that explains the first half is that of Feynman and Gell-Mann [Feynman:1958]: the weak vector current \(V^{\mu}=\bar{\psi}_{p}\gamma^{\mu}\psi_{n}\) and the isovector part of the electromagnetic current are three components of one isospin triplet,

\begin{equation}\tag{101.38} V^{\mu}_{a}=\bar{\psi}_{N}\gamma^{\mu}\frac{\tau_{a}}{2}\psi_{N}\ec \qquad a=1,2,3\ec \end{equation}

with the weak charged currents the combinations \(V^{\mu}_{1}\pm\ii V^{\mu}_{2}\) and the isovector electromagnetic current \(V^{\mu}_{3}\). Isospin is a symmetry of the strong interaction, so all three currents are conserved, \(\pp_{\mu}V^{\mu}_{a}=0\), and their charges \(T_{a}=\frac{1}{c}\int\dd^{3}x\,V^{0}_{a}\) generate the algebra \(\comm{T_{a}}{T_{b}}=\ii\varepsilon_{abc}T_{c}\).

The argument is now exactly that of Proposition 100.6, transposed. A charge that generates a symmetry algebra takes eigenvalues fixed by the representation theory of that algebra — half-integers, in units set by the commutation relations — and a continuous dynamical effect cannot shift a discrete eigenvalue. The strong interaction commutes with \(T_{a}\), so it cannot alter the matrix element of \(V^{\mu}_{a}\) between states of definite isospin at zero momentum transfer, any more than the strong interaction alters the electric charge of the proton, which by Proposition 100.6 and the Ward–Takahashi identity of Section 100.5.2 is exactly the positron's. The vector coupling of a nucleon is therefore the same as that of a bare quark, and the constancy of \(\mathcal{F}t\) across the superallowed multiplet in Phenomenon 101.18 is the measurement that confirms it: a renormalization that depended on nuclear structure would show up as a spread in those fourteen numbers, and the spread is a part in \(10^{4}\).

The axial current is a different matter, because it is not conserved: nothing in the strong interaction is invariant under a chiral rotation once the quarks have mass, and, more importantly, the axial symmetry is spontaneously broken. There is accordingly no protection theorem and \(g_{A}\) is renormalized. The measured \(27\) per cent is small only because the quark masses are small, and its theoretical status belongs to the chiral symmetry of Quantum Chromodynamics. That the vector coupling is protected and the axial one is not is a structural prediction, and the two measured numbers — \(1\) and \(1.2754\) — are the test of it.

The partial conservation that replaces the exact conservation has a sharp consequence, which is one of the very few quantitative relations in this chapter that connects a weak amplitude to a purely strong one.

Definition 101.41 (The pion decay constant and PCAC).

Define \(f_{\pi}\) by the matrix element of the axial current between the vacuum and a one-pion state,

\begin{equation}\tag{101.39} \bra{0}A^{\mu}_{a}(0)\ket{\pi_{b}(q)} =\ii f_{\pi}\,\delta_{ab}\,q^{\mu}\ec \end{equation}

with \(f_{\pi}c=92.3\,\mathrm{MeV}\) in the normalization used here [Navas:2024], so that \(f_{\pi}\) has the dimension of a momentum. The hypothesis of a partially conserved axial current is that \(\pp_{\mu}A^{\mu}_{a}\) is not zero but is proportional to the pion field,

\begin{equation}\tag{101.40} \pp_{\mu}A^{\mu}_{a} =f_{\pi}\left(\frac{m_{\pi}c}{\hbar}\right)^{2}\hbar\,\phi_{a}\ec \end{equation}

which is consistent with Equation (101.39) and vanishes in the limit \(m_{\pi}\to0\): the pion is the almost-Goldstone boson of the broken chiral symmetry.

Proposition 101.42 (Goldberger–Treiman).

Let \(g_{\pi N}\) be the pion–nucleon coupling constant measured in strong-interaction scattering and \(g_{A}\) the axial coupling of the nucleon measured in beta decay. Then

\begin{equation}\tag{101.41} f_{\pi}\,g_{\pi N}=g_{A}\,m_{N}c\ec \end{equation}

up to corrections of order \(m_{\pi}^{2}/m_{N}^{2}\). Rests on Equations (101.39) and (101.40).

Proof.

Derives Proposition 101.42. Take the matrix element of \(\pp_{\mu}A^{\mu}_{a}\) between two nucleon states of momenta \(p\) and \(p'\), and evaluate it in two ways.

On the left, use the general decomposition of the axial matrix element consistent with Lorentz invariance and parity,

\begin{equation}\tag{101.42} \bra{N(p')}A^{\mu}_{a}\ket{N(p)} =\bar{u}(p')\left[g_{A}\!\left(q^{2}\right)\gamma^{\mu}\gamma^{5} +\frac{h_{A}\!\left(q^{2}\right)}{m_{N}c}\,q^{\mu}\gamma^{5}\right] \frac{\tau_{a}}{2}u(p)\ec\qquad q=p'-p\ec \end{equation}

with \(g_{A}\) and \(h_{A}\) dimensionless. Contracting with \(q_{\mu}\), and using \(\bar{u}(p')\gamma^{\mu}q_{\mu}\gamma^{5}u(p) =\bar{u}(p')\left(\gamma\cdot p'-\gamma\cdot p\right)\gamma^{5}u(p) =2m_{N}c\,\bar{u}(p')\gamma^{5}u(p)\) — the Dirac equation applied to each spinor, the sign of the second term flipped by \(\gamma^{5}\gamma\cdot p=-\gamma\cdot p\,\gamma^{5}\) so that the two add — gives

\begin{equation}\tag{101.43} q_{\mu}\bra{N(p')}A^{\mu}_{a}\ket{N(p)} =\left[2m_{N}c\,g_{A}\!\left(q^{2}\right) +\frac{q^{2}}{m_{N}c}h_{A}\!\left(q^{2}\right)\right] \bar{u}(p')\gamma^{5}\frac{\tau_{a}}{2}u(p)\ep \end{equation}

On the right, use Equation (101.40): the divergence is \(f_{\pi}(m_{\pi}c/\hbar)^{2}\hbar\phi_{a}\), and the nucleon matrix element of the pion field is the pion propagator times the pion–nucleon vertex, which with the coupling normalized as \(\Lag_{\pi N}=-\ii g_{\pi N} \bar{\psi}_{N}\gamma^{5}\tau_{a}\psi_{N}\phi_{a}\) reads

\[ \bra{N(p')}\hbar\phi_{a}\ket{N(p)} =\frac{2\hbar\,g_{\pi N}} {q^{2}-m_{\pi}^{2}c^{2}}\; \bar{u}(p')\gamma^{5}\frac{\tau_{a}}{2}u(p)\cdot\hbar\ep \]

Equating the two expressions and cancelling the common spinor factor,

\begin{equation}\tag{101.44} 2m_{N}c\,g_{A}\!\left(q^{2}\right) +\frac{q^{2}}{m_{N}c}h_{A}\!\left(q^{2}\right) =\frac{2f_{\pi}g_{\pi N}\,m_{\pi}^{2}c^{2}} {q^{2}-m_{\pi}^{2}c^{2}}\ep \end{equation}

Now set \(q^{2}=0\), where the second term on the left vanishes provided \(h_{A}\) has no pole at the origin — which it has not, the pion pole sitting at \(q^{2}=m_{\pi}^{2}c^{2}\) — and the right-hand side becomes \(-2f_{\pi}g_{\pi N}\). Hence \(m_{N}c\,g_{A}(0)=-f_{\pi}g_{\pi N}\), which is Equation (101.41) up to the sign convention for \(g_{\pi N}\). The mechanism is worth naming: the right-hand side of Equation (101.44) is proportional to \(m_{\pi}^{2}\) and would vanish in the chiral limit, but the pion propagator supplies a compensating \(1/m_{\pi}^{2}\) at \(q^{2}=0\), so the product survives. That is why an exactly conserved axial current would still give a nonzero relation, and why the corrections are of relative order \(m_{\pi}^{2}/m_{N}^{2}\) rather than of order one. Numerically, with \(f_{\pi}c=92.3\,\mathrm{MeV}\), \(g_{\pi N}=13.05\), \(m_{N}c^{2}=938.9\,\mathrm{MeV}\) and \(g_{A}=1.2754\) [Navas:2024],

\[ f_{\pi}c\,g_{\pi N}=1204\,\mathrm{MeV}\ec\qquad g_{A}m_{N}c^{2}=1197\,\mathrm{MeV}\ec \]

agreeing to \(0.6\) per cent, which is of the expected order \(m_{\pi}^{2}/m_{N}^{2}\approx0.02\). The relation is [Goldberger:1958]. What makes it worth the page is that it connects three quantities measured in three different ways — a weak decay rate, a strong scattering length and a leptonic pion decay — and is right to one per cent.

Phenomenon 101.43 (One coupling governs every charged-current process).

The strength of the charged-current interaction extracted from the purely leptonic decay \(\mu^{-}\to e^{-}\bar{\nu}_{e}\nu_{\mu}\), from superallowed nuclear beta decay and from \(\tau\) decay is one and the same number. The muon lifetime is measured to a part per million, \(\tau_{\mu}=2.1969811\times 10^{-6}\,\mathrm{s}\) [Tishchenko:2013], and defines

\begin{equation}\tag{101.45} \frac{G_{F}}{(\hbar c)^{3}} =1.1663787\times 10^{-5}\,/\mathrm{GeV}^{2} \end{equation}

[Navas:2024]; the same constant reproduces the nuclear rates once the rotation of Section 101.5 is applied to the strangeness-changing part. Universality was read off the data, not imposed: nothing in Fermi's construction required the coupling of a muon and the coupling of a nucleon to agree. Rests on Equation (101.20), Phenomenon 101.40 and Equation (101.53).

Derivation. Derives Phenomenon 101.43. The three determinations are as follows. Muon decay gives \(G_{F}\) directly by Equation (101.53), with no hadron anywhere in the problem and hence no strong correction beyond the electromagnetic ones. Superallowed nuclear decay gives \(G_{F}\abs{V_{ud}}\) by Equation (101.20), and the vector coupling entering it is unrenormalized by Phenomenon 101.40; dividing out \(\abs{V_{ud}}=0.97435\) returns the muon-decay value. Tau decay to \(\nu_{\tau}e\bar{\nu}_{e}\) is muon decay with a heavier lepton, so Equation (101.53) applies with \(m_{\mu}\to m_{\tau}\); the predicted partial width, using \(m_{\tau}c^{2}=1776.93\,\mathrm{MeV}\) and the same \(G_{F}\), reproduces the measured branching fraction of \(17.8\,\mathrm{\%}\) together with the measured lifetime \(290.3\times 10^{-15}\,\mathrm{s}\) [Navas:2024]. Three processes whose energy scales differ by a factor \(10^{4}\) and whose participants differ in every other respect share one constant. Puppi's observation that the couplings of \((\mu,\nu_{\mu})\), \((e,\nu_{e})\) and \((n,p)\) to one another form a triangle of equal sides, and Konopinski and Mahmoud's version of the same statement [Konopinski:1953], are the origin of the phrase universal Fermi interaction; the residual deficit in the strangeness-changing sector, which is what stops the triangle from closing exactly, is Section 101.5.

Muon decay and the Fermi constant

The muon is the cleanest weak decay in existence. Nothing that feels the strong interaction appears in either the initial or the final state; the final state has three particles, two of which are essentially massless; and the lifetime can be measured by simply counting. It is therefore the process that defines \(G_{F}\), and this subsection carries out the calculation in full.

Definition 101.44 (The effective Lagrangian for muon decay).

At energies far below the mass of the \(W\), the decay \(\mu^{-}\to e^{-}\bar{\nu}_{e}\nu_{\mu}\) is governed by the contact interaction

\begin{equation}\tag{101.46} \Lag_{\mu}=-\frac{G_{F}}{\sqrt{2}} \left[\bar{\psi}_{\nu_{\mu}}\gamma^{\alpha} \left(\identity-\gamma^{5}\right)\psi_{\mu}\right] \left[\bar{\psi}_{e}\gamma_{\alpha} \left(\identity-\gamma^{5}\right)\psi_{\nu_{e}}\right] +\text{h.c.}\ec \end{equation}

whose coefficient \(G_{F}/\sqrt{2}\) carries \(\mathrm{J}\,\mathrm{m}^{3}\) by Proposition 101.7; the factor \(\sqrt{2}\) is a convention, fixed once and for all here, and Section 101.6.1 shows what it becomes when the \(W\) is restored. The corresponding invariant amplitude is

\begin{equation}\tag{101.47} \mathcal{M}=\frac{G_{F}}{\sqrt{2}\,\hbar c}\,\mathcal{A}\ec\qquad \mathcal{A}=\left[\bar{u}_{\nu_{\mu}}\gamma^{\alpha} \left(\identity-\gamma^{5}\right)u_{\mu}\right] \left[\bar{u}_{e}\gamma_{\alpha} \left(\identity-\gamma^{5}\right)v_{\bar{\nu}_{e}}\right]\ec \end{equation}

of SI dimension \(\hbar^{2}\), in agreement with the general statement of Proposition 100.14: the four external spinors contribute a squared momentum, and \(G_{F}/\hbar c\) carries \(\mathrm{m}^{2}\).

Proposition 101.45 (Decay-rate master formula).

For the decay of a particle of energy \(E_{1}\) into \(n\) particles,

\begin{equation}\tag{101.48} \dd\Gamma=\frac{c^{2}}{2E_{1}}\, \overline{\abs{\mathcal{M}}^{2}}\;\dd\Phi_{n}\ec \end{equation}

with the invariant phase space \(\dd\Phi_{n}\) of Equation (100.39). In the rest frame the prefactor is \(1/2m_{1}\); in any other frame the rate is smaller by the Lorentz factor \(m_{1}c^{2}/E_{1}\), which is time dilation. Rests on Equation (100.39), Equation (101.10) and Proposition 100.14.

Proof.

Derives Proposition 101.45. Quantize in a box of volume \(V\) with periodic boundary conditions and normalize each one-particle state to one particle in the box,

\begin{equation}\tag{101.49} \psi(x)=\frac{1}{\sqrt{V}}\, \frac{u(p,s)}{\sqrt{2E/c}}\,\ee^{-\ii p\cdot x/\hbar}\ec \end{equation}

which is correctly normalized because \(u^{\dagger}u=2E/c\) by Proposition 100.14, so that \(\int_{V}\dd^{3}x\,\psi^{\dagger}\psi=1\); note that the prefactor is dimensionally forced, \(\mathrm{m}^{-3/2}\) on the left and \(u/\sqrt{2E/c}\) dimensionless. With \(\Ham_{\mathrm{int}}=-\Lag_{\mu}\), which holds because the interaction contains no derivatives, the matrix element between the one-muon state and a three-particle final state is

\[ \bra{f}\Ham_{\mathrm{int}}\ket{i} =\frac{G_{F}}{\sqrt{2}}\;\frac{1}{V^{2}}\; \frac{\mathcal{A}}{\prod_{k=1}^{4}\sqrt{2E_{k}/c}}\; \times V\delta_{\vect{p}_{1},\,\vect{p}_{2}+\vect{p}_{3}+\vect{p}_{4}}\ec \]

the four factors \(V^{-1/2}\) coming from the four fields and the factor \(V\) from \(\int_{V}\dd^{3}x\), which also enforces momentum conservation as a Kronecker delta. Squaring, using \(\delta_{\mathrm{K}}^{2}=\delta_{\mathrm{K}}\), and inserting into the golden rule Equation (101.10) with the continuum sum \(\sum_{f}\to\prod_{k=2}^{4}V\dd^{3}p_{k}/(2\pi\hbar)^{3}\) and the continuum form of the Kronecker delta, \(\delta_{\mathrm{K}}=(2\pi\hbar)^{3}V^{-1} \delta^{3}(\vect{p}_{1}-\sum\vect{p}_{f})\), gives

\[ \Gamma=\frac{2\pi}{\hbar}\,\frac{G_{F}^{2}}{2}\,\frac{c}{2E_{1}} \int\abs{\mathcal{A}}^{2} \left[\prod_{f}\frac{c\,\dd^{3}p_{f}}{(2\pi\hbar)^{3}\,2E_{f}}\right] \left(2\pi\hbar\right)^{3} \delta^{3}\!\left(\vect{p}_{1}-\sum_{f}\vect{p}_{f}\right) \delta\!\left(E_{1}-\sum_{f}E_{f}\right)\ec \]

every power of \(V\) having cancelled: \(V^{3}\) from the three measures against \(V^{-2}\) from the squared matrix element and \(V^{-1}\) from the delta. Now use \(\left(2\pi\hbar\right)^{4}\delta^{4}(p) =\left(2\pi\hbar\right)^{3}\delta^{3}(\vect{p})\cdot 2\pi\hbar c\,\delta(E)\), so that \(2\pi\delta(E)\left(2\pi\hbar\right)^{3}\delta^{3} =\left(2\pi\hbar\right)^{4}\delta^{4}/\hbar c\), to recognize \(\dd\Phi_{3}\) of Equation (100.39) entire:

\[ \Gamma=\frac{G_{F}^{2}}{4\hbar^{2}E_{1}} \int\abs{\mathcal{A}}^{2}\dd\Phi_{3}\ep \]

Finally substitute \(\abs{\mathcal{A}}^{2}=2\hbar^{2}c^{2}\abs{\mathcal{M}}^{2}/G_{F}^{2}\) from Equation (101.47), which cancels \(G_{F}\) and \(\hbar\) together and leaves Equation (101.48). Nothing in the last two steps used the specific interaction: the coupling and the spinor structure enter only through \(\mathcal{M}\), while the factors of \(V\) and of \(2E/c\) belong to the external legs. Hence the formula holds for any \(1\to n\) decay, which is how it is used again in Section 101.6.4.

Lemma 101.46 (The spin-summed square).

With the neutrinos massless and the muon spin averaged,

\begin{equation}\tag{101.50} \overline{\abs{\mathcal{A}}^{2}} =128\left(p_{\mu}\cdot p_{\bar{\nu}_{e}}\right) \left(p_{e}\cdot p_{\nu_{\mu}}\right)\ep \end{equation}

The muon pairs with the electron antineutrino and the electron with the muon neutrino; the pairing is a consequence of the current structure and not a choice. Rests on Equation (100.33) and Lemma 100.23.

Proof.

Derives Lemma 101.46. Write \(\mathcal{A}=L^{\alpha}M_{\alpha}\) with \(L^{\alpha}=\bar{u}_{\nu_{\mu}}\gamma^{\alpha}(1-\gamma^{5})u_{\mu}\) and \(M_{\alpha}=\bar{u}_{e}\gamma_{\alpha}(1-\gamma^{5}) v_{\bar{\nu}_{e}}\). Summing over final spins and averaging over the two muon spins, \(\overline{\abs{\mathcal{A}}^{2}} =\tfrac{1}{2}L^{\alpha\beta}M_{\alpha\beta}\) with

\[ L^{\alpha\beta}=\tr\left[\gamma\cdot p_{\nu_{\mu}}\, \gamma^{\alpha}\left(1-\gamma^{5}\right) \left(\gamma\cdot p_{\mu}+m_{\mu}c\right) \gamma^{\beta}\left(1-\gamma^{5}\right)\right]\ec \]

the spin sums being those of Equation (100.33) and the conjugation using \(\gamma^{0}\left[\gamma^{\beta}(1-\gamma^{5})\right]^{\dagger} \gamma^{0}=\gamma^{\beta}(1-\gamma^{5})\). The mass term drops, because it leaves an odd number of gamma matrices in the trace. Using \(\left(1-\gamma^{5}\right)^{2}=2\left(1-\gamma^{5}\right)\) and the trace theorems of Lemma 100.23 together with \(\tr\left(\gamma^{\mu}\gamma^{\nu}\gamma^{\rho}\gamma^{\sigma} \gamma^{5}\right)=-4\ii\varepsilon^{\mu\nu\rho\sigma}\),

\[ L^{\alpha\beta}=8\left[ p_{\nu_{\mu}}^{\alpha}p_{\mu}^{\beta} +p_{\nu_{\mu}}^{\beta}p_{\mu}^{\alpha} -\eta^{\alpha\beta}\left(p_{\nu_{\mu}}\cdot p_{\mu}\right) +\ii\varepsilon^{\rho\alpha\sigma\beta} p_{\nu_{\mu}\rho}p_{\mu\sigma}\right]\ec \]

and \(M_{\alpha\beta}\) is the same expression with \(p_{\nu_{\mu}}\to p_{e}\) and \(p_{\mu}\to p_{\bar{\nu}_{e}}\) and the indices lowered. Abbreviate \(P_{1}=p_{\nu_{\mu}}\), \(P_{2}=p_{\mu}\), \(P_{3}=p_{e}\), \(P_{4}=p_{\bar{\nu}_{e}}\) — numbered rather than lettered, since \(c\) is the speed of light throughout this book — and contract.

The symmetric parts contract to

\[ \begin{aligned} 64\bigl[&2\left(P_{1}\cdot P_{3}\right)\left(P_{2}\cdot P_{4}\right) +2\left(P_{1}\cdot P_{4}\right)\left(P_{2}\cdot P_{3}\right)\\ &-4\left(P_{1}\cdot P_{2}\right)\left(P_{3}\cdot P_{4}\right) +4\left(P_{1}\cdot P_{2}\right)\left(P_{3}\cdot P_{4}\right)\bigr]\\ &=128\left[\left(P_{1}\cdot P_{3}\right)\left(P_{2}\cdot P_{4}\right) +\left(P_{1}\cdot P_{4}\right)\left(P_{2}\cdot P_{3}\right)\right]\ec \end{aligned} \]

the two trace terms cancelling exactly against the term in which the two metrics meet. The product of the two Levi-Civita terms, using \(\varepsilon^{\rho\alpha\sigma\beta} \varepsilon_{\kappa\alpha\lambda\beta} =-2\left(\delta^{\rho}_{\kappa}\delta^{\sigma}_{\lambda} -\delta^{\rho}_{\lambda}\delta^{\sigma}_{\kappa}\right)\) and \(\ii^{2}=-1\), contributes

\[ 128\left[\left(P_{1}\cdot P_{3}\right)\left(P_{2}\cdot P_{4}\right) -\left(P_{1}\cdot P_{4}\right)\left(P_{2}\cdot P_{3}\right)\right]\ec \]

while the cross terms between a symmetric and an antisymmetric tensor vanish identically. Adding, the \(\left(P_{1}\cdot P_{4}\right) \left(P_{2}\cdot P_{3}\right)\) pieces cancel and the others double:

\begin{equation}\tag{101.51} L^{\alpha\beta}M_{\alpha\beta} =256\left(p_{\nu_{\mu}}\cdot p_{e}\right) \left(p_{\mu}\cdot p_{\bar{\nu}_{e}}\right)\ec \end{equation}

which halved is Equation (101.50). This is the single most useful identity in weak-interaction physics: any two \(V-A\) currents, contracted and summed over spins, collapse to a product of two dot products with the coefficient \(256\), and the cancellation of the other pairing is what fixes which particle goes with which.

Lemma 101.47 (The two-neutrino phase space).

For two massless particles sharing a total four-momentum \(q\),

\begin{equation}\tag{101.52} \int\dd\Phi_{2}\;k^{\alpha}k'^{\beta} =\frac{1}{96\pi\hbar^{2}} \left(q^{2}\eta^{\alpha\beta}+2q^{\alpha}q^{\beta}\right)\ec \end{equation}

where \(\dd\Phi_{2}\) is the two-body element of Equation (100.39) with \(k+k'=q\). Rests on Equation (100.39) and Proposition 100.24.

Proof.

Derives Lemma 101.47. The integral is a Lorentz tensor built from \(q\) alone, so \(I^{\alpha\beta}=Aq^{2}\eta^{\alpha\beta}+Bq^{\alpha}q^{\beta}\). Contract with \(\eta_{\alpha\beta}\): since \(k^{2}=k'^{2}=0\) gives \(2k\cdot k'=q^{2}\), the left side is \(\tfrac{1}{2}q^{2}\Phi_{2}\) with \(\Phi_{2}=\int\dd\Phi_{2}\), so \(4A+B=\tfrac{1}{2}\Phi_{2}\). Contract with \(q_{\alpha}q_{\beta}\): \(q\cdot k=k'\cdot k=\tfrac{1}{2}q^{2}\) and likewise \(q\cdot k'\), so the left side is \(\tfrac{1}{4}q^{4}\Phi_{2}\) and \(A+B=\tfrac{1}{4}\Phi_{2}\). Hence \(A=\Phi_{2}/12\) and \(B=\Phi_{2}/6\). The total two-body phase space for massless products follows from Proposition 100.24, where \(\dd\Phi_{2}=\hbar^{-2}\abs{\vect{p}_{f}} \dd\Omega/16\pi^{2}\sqrt{s}\) with \(\abs{\vect{p}_{f}} =\tfrac{1}{2}\sqrt{s}\) here, giving \(\Phi_{2}=4\pi/32\pi^{2}\hbar^{2}=1/8\pi\hbar^{2}\); note that it carries \(\hbar^{-2}\), as every element of \(\dd\Phi\) must. Inserting gives Equation (101.52).

Theorem 101.48 (The muon decay rate).

Neglecting the electron mass and radiative corrections, the \(V-A\) interaction Equation (101.46) gives

\begin{equation}\tag{101.53} \Gamma_{\mu}=\frac{1}{\tau_{\mu}} =\frac{1}{192\pi^{3}\hbar} \left(\frac{G_{F}}{\left(\hbar c\right)^{3}}\right)^{2} \left(m_{\mu}c^{2}\right)^{5} =\frac{G_{F}^{2}\left(m_{\mu}c^{2}\right)^{5}} {192\pi^{3}\hbar\left(\hbar c\right)^{6}}\ec \end{equation}

and the normalized electron energy spectrum, with \(x=2E_{e}/m_{\mu}c^{2}\in[0,1]\),

\begin{equation}\tag{101.54} \frac{1}{\Gamma_{\mu}}\frac{\dd\Gamma_{\mu}}{\dd x} =2x^{2}\left(3-2x\right)\ep \end{equation}

Rests on Equations (101.46), (101.48) and (101.50).

Proof.

Derives Theorem 101.48. Combine Equation (101.48), Equation (101.47) and Equation (101.50):

\[ \overline{\abs{\mathcal{M}}^{2}} =\frac{G_{F}^{2}}{2\hbar^{2}c^{2}}\cdot 128\left(p_{\mu}\cdot p_{\bar{\nu}_{e}}\right) \left(p_{e}\cdot p_{\nu_{\mu}}\right) =\frac{64G_{F}^{2}}{\hbar^{2}c^{2}} \left(p_{\mu}\cdot p_{\bar{\nu}_{e}}\right) \left(p_{e}\cdot p_{\nu_{\mu}}\right)\ep \]

The two neutrinos are not observed, so integrate over them first. Writing \(q=p_{\mu}-p_{e}\) for their total four-momentum and using Equation (101.52) with \(k=p_{\nu_{\mu}}\) and \(k'=p_{\bar{\nu}_{e}}\),

\[ \int\dd\Phi_{2}\left(p_{\mu}\cdot p_{\bar{\nu}_{e}}\right) \left(p_{e}\cdot p_{\nu_{\mu}}\right) =p_{e\alpha}p_{\mu\beta}\int\dd\Phi_{2}\, k^{\alpha}k'^{\beta} =\frac{q^{2}\left(p_{e}\cdot p_{\mu}\right) +2\left(q\cdot p_{e}\right)\left(q\cdot p_{\mu}\right)} {96\pi\hbar^{2}}\ep \]

Put the muon at rest, \(p_{\mu}=(m_{\mu}c,\vect{0})\), and neglect the electron mass, \(p_{e}^{2}=0\). Then, writing \(y=p_{\mu}\cdot p_{e}=m_{\mu}E_{e}\),

\[ q^{2}=m_{\mu}^{2}c^{2}-2y\ec\qquad q\cdot p_{e}=y\ec\qquad q\cdot p_{\mu}=m_{\mu}^{2}c^{2}-y\ec \]

so the numerator is \(\left(m_{\mu}^{2}c^{2}-2y\right)y+2y\left(m_{\mu}^{2}c^{2}-y\right) =m_{\mu}^{2}E_{e}\left(3m_{\mu}c^{2}-4E_{e}\right)\). The remaining electron integral is the one-particle element of Equation (100.39) with \(\dd^{3}p_{e}=4\pi E_{e}^{2} \dd E_{e}/c^{3}\) for a massless electron:

\[ \int\frac{c\,\dd^{3}p_{e}}{(2\pi\hbar)^{3}2E_{e}}\, m_{\mu}^{2}E_{e}\left(3m_{\mu}c^{2}-4E_{e}\right) =\frac{4\pi m_{\mu}^{2}}{2c^{2}\left(2\pi\hbar\right)^{3}} \int_{0}^{m_{\mu}c^{2}/2}E_{e}^{2} \left(3m_{\mu}c^{2}-4E_{e}\right)\dd E_{e}\ec \]

the upper limit being the greatest energy an electron can take when the two neutrinos go together against it. The integral is \(m_{\mu}c^{2}W^{3}-W^{4}\) with \(W=m_{\mu}c^{2}/2\), that is \(\left(m_{\mu}c^{2}\right)^{4}/16\), and \(\left(2\pi\hbar\right)^{3}=8\pi^{3}\hbar^{3}\), so the electron integral equals \(m_{\mu}^{2}\left(m_{\mu}c^{2}\right)^{4}/64\pi^{2}\hbar^{3}c^{2}\). Assembling with the prefactor \(c^{2}/2m_{\mu}c^{2}\) of Equation (101.48),

\[ \Gamma_{\mu}=\frac{1}{2m_{\mu}}\cdot \frac{64G_{F}^{2}}{\hbar^{2}c^{2}}\cdot\frac{1}{96\pi\hbar^{2}} \cdot\frac{m_{\mu}^{2}\left(m_{\mu}c^{2}\right)^{4}} {64\pi^{2}\hbar^{3}c^{2}} =\frac{G_{F}^{2}m_{\mu}\left(m_{\mu}c^{2}\right)^{4}} {192\pi^{3}\hbar^{7}c^{4}}\ec \]

since \(64/(2\cdot96\cdot64)=1/192\). Writing \(m_{\mu}=\left(m_{\mu}c^{2}\right)/c^{2}\) and \(\hbar^{7}c^{6}=\hbar\left(\hbar c\right)^{6}\) gives Equation (101.53). The spectrum Equation (101.54) is the integrand before the last integration: with \(E_{e}=xm_{\mu}c^{2}/2\) the integrand is proportional to \(E_{e}^{2}\left(3m_{\mu}c^{2}-4E_{e}\right)\propto x^{2}\left(3-2x\right)\), and \(\int_{0}^{1}2x^{2}(3-2x)\dd x=2\left(1-\tfrac{1}{2}\right)=1\) fixes the normalization.

Remark 101.49 (What the fifth power is).

The \(\left(m_{\mu}c^{2}\right)^{5}\) in Equation (101.53) is forced and could have been written down without the calculation. \(\Gamma\) has dimension \(/\mathrm{s}\); the only quantities available are \(G_{F}\), \(m_{\mu}c^{2}\), \(\hbar\) and \(c\); and \(G_{F}^{2}/(\hbar c)^{6}\) carries \(\mathrm{J}^{-4}\), so exactly four powers of energy are needed to make it dimensionless and a fifth to convert \(\hbar^{-1}\) into a rate. What the calculation supplies is the pure number \(1/192\pi^{3}\), and that number is where the \(V-A\) structure and the three-body phase space actually live.

The same argument applies to the tau, and it is worth being exact about how much of the answer it supplies. The tau is \(16.82\) times heavier than the muon [Navas:2024], so the fifth power predicts its partial width to \(e^{-}\bar{\nu}_{e}\nu_{\tau}\) to be \(16.82^{5}=1.35\times 10^{6}\) times the muon's total width. That alone would give the tau a lifetime of \(1.63\times 10^{-12}\,\mathrm{s}\). The measured lifetime is \(2.903\times 10^{-13}\,\mathrm{s}\) [Navas:2024], shorter by a further factor \(5.6\), and that factor is not dimensional analysis: it is \(1/\mathcal{B}(\tau^{-}\to e^{-}\bar{\nu}_{e}\nu_{\tau})=1/0.178\), the hadronic and muonic channels open to a tau and closed to a muon. Putting the two together, \(2.1969811\times 10^{-6}\,\mathrm{s}\times0.178/1.35\times 10^{6} =2.91\times 10^{-13}\,\mathrm{s}\), which is the measurement. The ratio of total lifetimes is \(7.6\times 10^{6}\), and quoting that as what the fifth power predicts would credit dimensional analysis with a branching fraction it knows nothing about. It is the same fifth power as Sargent's rule Equation (101.14): two light particles sharing a release.

Phenomenon 101.50 (The muon lifetime measures the Fermi constant).

The measured positive-muon lifetime is

\begin{equation}\tag{101.55} \tau_{\mu}=2.1969811(22)\times 10^{-6}\,\mathrm{s} \end{equation}

[Tishchenko:2013], a determination to one part per million from \(2\times 10^{12}\) recorded decays. Inserted into Equation (101.53), corrected for the electron mass and for the one-loop electromagnetic radiative correction, it yields Equation (101.45); conversely, that value of \(G_{F}\) predicts the lifetime. Rests on Equations (101.45) and (101.53).

Derivation. Derives Phenomenon 101.50. Evaluate Equation (101.53) with \(G_{F}/(\hbar c)^{3}=1.1663787\times 10^{-5}\,/\mathrm{GeV}^{2}\), \(m_{\mu}c^{2}=105.6583755\,\mathrm{MeV}\) and \(\hbar=6.582120\times 10^{-25}\,\mathrm{GeV}\,\mathrm{s}\):

\[ \Gamma_{\mu}^{\mathrm{tree}} =\frac{\left(1.1663787\times 10^{-5}\right)^{2} \left(0.1056583755\right)^{5}} {192\pi^{3}\left(6.582120\times 10^{-25}\,\mathrm{GeV}\,\mathrm{s}\right)} \,\mathrm{GeV}^{-4}\,\mathrm{GeV}^{5} =4.5717\times 10^{5}\,/\mathrm{s}\ec \]

that is \(\tau=2.1873\times 10^{-6}\,\mathrm{s}\), which is \(0.44\) per cent below the measurement. Two known corrections account for the difference and for nothing else.

The first is kinematic. Retaining the electron mass multiplies Equation (101.53) by

\begin{equation}\tag{101.56} f(z)=1-8z+8z^{3}-z^{4}-12z^{2}\ln z\ec\qquad z=\frac{m_{e}^{2}}{m_{\mu}^{2}}=2.339\times 10^{-5}\ec \end{equation}

whose value here is \(f=0.99981\), a reduction of \(1.87\times 10^{-4}\): the electron mass eats into the available phase space, and the effect is second order in \(m_{e}/m_{\mu}\) because the spectrum vanishes quadratically at low \(x\) by Equation (101.54).

The second is electromagnetic. The one-loop QED correction to the contact interaction, computed once and for all by the methods of Section 100.4, multiplies the rate by

\begin{equation}\tag{101.57} 1+\Delta q\ec\qquad \Delta q=\frac{\alpha\left(m_{\mu}c^{2}\right)}{2\pi} \left(\frac{25}{4}-\pi^{2}\right)=-4.24\times 10^{-3}\ec \end{equation}

using \(\alpha(m_{\mu}c^{2})^{-1}=135.9\) for the fine-structure constant of Equation (100.3) run to the muon mass by Proposition 100.63. The combination \(f(z)\left(1+\Delta q\right)=0.99557\) gives

\[ \Gamma_{\mu}=4.5515\times 10^{5}\,/\mathrm{s}\ec\qquad \tau_{\mu}=2.1971\times 10^{-6}\,\mathrm{s}\ec \]

against the measured \(2.1969811\times 10^{-6}\,\mathrm{s}\): agreement to \(4\times10^{-5}\), which is the accuracy of the two-loop terms not included here and of the fourth digit of \(\alpha(m_{\mu}c^{2})\) used above. It is worth being explicit about the logic of the comparison. \(G_{F}\) is defined by inverting this calculation, so the agreement of the central values is a tautology; what is not a tautology, and is the actual content, is that the correction factors \(f(z)\) and \(1+\Delta q\) — computed from the electron mass and from QED, with no adjustable parameter — move the tree-level answer by exactly the \(0.44\) per cent by which it was wrong.

Derivation pending.

Muon decay, the parts not derived here: the phase-space function \(f(z)\) of the electron-mass correction, obtained by repeating the phase-space integral of this section with \(p_{e}^{2}=m_{e}^{2}c^{2}\) retained; the one-loop electromagnetic correction \({[}25/4-\pi^{2}{]}\), which is a QED calculation of the kind carried out in the QED chapter but applied to a contact operator; and the spin-dependent decay distribution \(x^{2}{[}(3-2x)+P_{\mu}\cos\theta\,(2x-1){]}\) obtained by keeping the muon spin projector in place of the average taken in the trace lemma. All three are Appendix A material and all three are quoted with their numerical values in the derivation above.

Definition 101.51 (Michel parameters).

The most general local, lepton-number-conserving, four-fermion interaction gives an electron spectrum of the form

\begin{equation}\tag{101.58} \frac{\dd\Gamma}{\dd x}\propto x^{2}\left[3\left(1-x\right) +\frac{2}{3}\rho\left(4x-3\right)\right]\ec \end{equation}

together with three further parameters \(\eta\), \(\xi\) and \(\delta\) governing the electron-mass term and the dependence on the muon polarization [Michel:1950]. The \(V-A\) interaction predicts \(\rho=\tfrac{3}{4}\), \(\eta=0\), \(\xi=1\), \(\delta=\tfrac{3}{4}\).

Phenomenon 101.52 (The muon decay spectrum has the $V-A$ shape).

The measured Michel parameters are \(\rho=0.74979(26)\) and \(\delta=0.75047(34)\) [Navas:2024], agreeing with the \(V-A\) prediction \(\tfrac{3}{4}\) to a few parts in \(10^{4}\); measurements with polarized muons bound the admixture of a right-handed charged current at the same level [Bayes:2011]. Rests on Equations (101.54) and (101.58).

Derivation. Derives Phenomenon 101.52. Setting \(\rho=\tfrac{3}{4}\) in Equation (101.58) gives \(x^{2}\left[3-3x+\tfrac{1}{2}(4x-3)\right] =\tfrac{1}{2}x^{2}\left(3-2x\right)\), which is Equation (101.54) up to normalization; so the derived spectrum is the \(\rho=\tfrac{3}{4}\) member of the family and the measurement is a direct test of the Lorentz structure, independent of the rate and therefore independent of \(G_{F}\). The value \(\tfrac{3}{4}\) is not generic: a scalar or tensor contact interaction gives \(\rho=0\), and a \(V+A\) admixture moves \(\xi\) away from \(1\). What the agreement excludes is stated most sharply in the language of left–right symmetric models, where a right-handed \(W\) of mass \(M_{2}\) mixing with the ordinary one shifts \(P_{\mu}\xi\delta/\rho\) below unity; the measured value of that combination is consistent with \(1\) and bounds the mixing [Bayes:2011].

The last property of the muon that belongs here is a null result, and it is one of the most stringent in physics.

Phenomenon 101.53 (Charged-current lepton flavour is conserved).

The decay \(\mu^{+}\to e^{+}\gamma\), which conserves electric charge, total lepton number, energy and angular momentum, and which would occur if the muon and the electron differed only in mass, has never been observed. The branching ratio is bounded by

\begin{equation}\tag{101.59} \mathcal{B}\left(\mu^{+}\to e^{+}\gamma\right)<1.5\times 10^{-13} \qquad\left(90\,\mathrm{\%}\text{ confidence}\right) \end{equation}

from the first dataset of MEG II combined with the full MEG sample; that dataset alone gives \(<7.5\times 10^{-13}\) [Afanaciev:2024]. Rests on Equation (101.46).

Derivation. Derives Phenomenon 101.53. What the bound establishes is that the quantum numbers \(L_{e}\) and \(L_{\mu}\) are separately conserved in charged-current processes, and not merely their sum. Were they not, the same \(W\) exchange that produces \(\mu\to e\nu\bar{\nu}\) would, with the two neutrinos joined into a loop and a photon attached, produce \(\mu\to e\gamma\) at a rate suppressed only by \(\alpha\) and by loop factors — of order \(10^{-4}\) of the total, not \(10^{-13}\). The observed suppression is therefore a statement about the structure of the current and not about the size of a coupling. In the Standard Model with massive neutrinos the amplitude is not exactly zero but is suppressed by \(\left(\Delta m_{\nu}^{2}c^{4}/M_{W}^{2}c^{4}\right)^{2}\sim 10^{-50}\), an unobservably small number, so any signal at all at the level of Equation (101.59) would be new physics; this is why the search continues. The experiment stops positive muons in a thin target and looks for a positron and a photon, back to back, each of energy \(m_{\mu}c^{2}/2=52.8\,\mathrm{MeV}\), in time coincidence [Afanaciev:2024].

The crossed version of muon decay is a scattering process, and its cross-section is needed twice below — once in Phenomenon 101.5 and once in Section 101.7.1.

Proposition 101.54 (Inverse muon decay).

For \(\nu_{\mu}e^{-}\to\mu^{-}\nu_{e}\) at centre-of-mass energy \(E_{\mathrm{cm}}\) large compared with the muon mass, the \(V-A\) contact interaction gives an isotropic cross-section

\begin{equation}\tag{101.60} \sigma=\frac{G_{F}^{2}E_{\mathrm{cm}}^{2}} {\pi\left(\hbar c\right)^{4}}\ep \end{equation}

Rests on Equations (100.40), (101.47) and (101.51).

Proof.

Derives Proposition 101.54. The amplitude is Equation (101.47) with the external legs crossed, and Equation (101.51) applies unchanged: summing over final spins and averaging over the two electron spins — the neutrino has only one helicity, so nothing is averaged there —

\[ \overline{\abs{\mathcal{A}}^{2}} =\frac{1}{2}\cdot256\left(p_{\nu_{\mu}}\cdot p_{e}\right) \left(p_{\mu}\cdot p_{\nu_{e}}\right) =128\cdot\frac{s}{2}\cdot\frac{s}{2}=32s^{2}\ec \]

using \(2p_{\nu_{\mu}}\cdot p_{e}=s\) for massless incoming particles and the same for the outgoing pair, with \(s=E_{\mathrm{cm}}^{2}/c^{2}\) a squared momentum in the conventions of Notation 100.1. Then \(\overline{\abs{\mathcal{M}}^{2}} =16G_{F}^{2}s^{2}/\hbar^{2}c^{2}\), and Equation (100.40) with \(\abs{\vect{p}_{f}}=\abs{\vect{p}_{i}}\) gives

\[ \frac{\dd\sigma}{\dd\Omega} =\frac{1}{64\pi^{2}\hbar^{2}s}\cdot \frac{16G_{F}^{2}s^{2}}{\hbar^{2}c^{2}} =\frac{G_{F}^{2}s}{4\pi^{2}\hbar^{4}c^{2}}\ec \]

independent of angle. Multiplying by \(4\pi\) and writing \(s=E_{\mathrm{cm}}^{2}/c^{2}\) gives Equation (101.60). The isotropy is not an accident: the contact interaction has no length scale, so it can produce only \(s\)-wave scattering, and that fact is what Section 101.7.1 turns against the theory. Applied to a nucleon rather than to an electron, with the vector and axial nucleon couplings inserted, the same calculation gives Equation (101.5): the factor \(1+3\lambda^{2}\) is the one Fermi channel plus the three Gamow–Teller channels of Proposition 101.15, each weighted by its coupling, and the factor \(\abs{V_{ud}}^{2}\) is the quark mixing of Section 101.5, which has no counterpart in the purely leptonic process computed here.

Strangeness-changing currents

The strangeness deficit and the Cabibbo rotation

Phenomenon 101.43 is not quite true, and the size of the discrepancy is the subject of this section. When the same \(G_{F}\) that governs muon decay and superallowed nuclear decay is used to predict the rate of a decay that changes strangeness — \(K^{+}\to\mu^{+}\nu_{\mu}\) against \(\pi^{+}\to\mu^{+}\nu_{\mu}\), or \(\Lambda\to p e^{-}\bar{\nu}_{e}\) against neutron decay — the prediction is too large by a factor of about twenty. Universality holds to a per cent for the strangeness-conserving processes and fails by a factor of twenty for the others.

Phenomenon 101.55 (Strangeness-changing decays are suppressed by one angle).

Semileptonic decays with \(\Delta S=1\) proceed at roughly one twentieth of the rate predicted by the coupling measured in \(\Delta S=0\) decays; and the deficit of the \(\Delta S=0\) rate below the purely leptonic strength is exactly what is needed to make up the difference. One angle,

\begin{equation}\tag{101.61} \theta_{C}=12.96^\circ\ec\qquad \sin\theta_{C}=0.2243\ec\qquad \cos\theta_{C}=0.9745\ec \end{equation}

accounts for both [Cabibbo:1963] [Navas:2024]. Rests on Phenomenon 101.43 and Equation (101.20).

Derivation. Derives Phenomenon 101.55. Cabibbo's hypothesis is that the state entering the hadronic charged current is not the down quark but a rotation of the down-type mass eigenstates,

\begin{equation}\tag{101.62} d'=d\cos\theta_{C}+s\sin\theta_{C}\ec \end{equation}

so that the hadronic current is

\begin{equation}\tag{101.63} J^{\mu}_{\mathrm{had}} =\bar{\psi}_{u}\gamma^{\mu}\left(\identity-\gamma^{5}\right) \psi_{d'} =\cos\theta_{C}\,\bar{\psi}_{u}\gamma^{\mu} \left(\identity-\gamma^{5}\right)\psi_{d} +\sin\theta_{C}\,\bar{\psi}_{u}\gamma^{\mu} \left(\identity-\gamma^{5}\right)\psi_{s}\ec \end{equation}

and that the same \(G_{F}\) multiplies this current and the leptonic one. The two consequences are then not independent. A \(\Delta S=0\) amplitude carries \(\cos\theta_{C}\) and a \(\Delta S=1\) amplitude carries \(\sin\theta_{C}\), so their rates stand in the ratio

\[ \frac{\Gamma\left(\Delta S=1\right)}{\Gamma\left(\Delta S=0\right)} =\tan^{2}\theta_{C}=0.0530=\frac{1}{18.9}\ec \]

which is the observed factor of about twenty; and the strangeness-conserving hadronic rate is suppressed relative to the purely leptonic one by \(\cos^{2}\theta_{C}=0.9496\), a deficit of about five per cent, which is precisely what Equation (101.20) sees as \(\abs{V_{ud}}^{2}\). The angle is therefore measured twice, in unrelated experiments: from the kaon rates as \(\sin\theta_{C}=\abs{V_{us}}=0.2243\), and from the superallowed nuclear rates as \(\cos\theta_{C}=\abs{V_{ud}}=0.97435\). The two determinations satisfy \(\sin^{2}+\cos^{2}=1\) to about a part in \(10^{3}\), and that residual is presently the sharpest unexplained tension in flavour physics; it is recorded, with the current numbers, in Phenomenon 103.7.

The conceptual step in Equation (101.62) is larger than its algebra suggests, and it is the one that the whole of Flavour Physics and Neutrinos elaborates: the states that couple to the \(W\) are not the states of definite mass. There are two bases for the down-type quarks, one diagonalizing the mass matrix and one diagonalizing the charged-current coupling, and they are related by a rotation. Nothing requires them to coincide, and they do not. Universality is restored not by making the couplings equal but by recognizing that a single coupling is shared out between channels, with the shares summing to one — which is a statement of orthogonality, and therefore the first appearance of the unitarity that Section 101.5.4 turns into a testable constraint.

Selection rules

Equation (101.63) carries two selection rules that were established empirically before it was written and that are among the sharpest tests of the current structure.

Proposition 101.56 (The $\Delta S=\Delta Q$ rule).

A semileptonic decay of hadrons mediated by Equation (101.63) satisfies \(\Delta S=\Delta Q\), where both refer to the hadronic system. Rests on Equation (101.63).

Proof.

Derives Proposition 101.56. The strangeness-changing part of Equation (101.63) is \(\bar{\psi}_{u}\Gamma\psi_{s}\), which annihilates an \(s\) quark and creates a \(u\). The \(s\) quark carries \(S=-1\) and charge \(-\tfrac{1}{3}e\); the \(u\) carries \(S=0\) and \(+\tfrac{2}{3}e\). Hence \(\Delta S=0-(-1)=+1\) and \(\Delta Q=+\tfrac{2}{3}-(-\tfrac{1}{3})=+1\) in units of \(e\), and the two are equal. The Hermitian conjugate term gives \(\Delta S=\Delta Q=-1\). The rule is a prediction and not a convention: it forbids \(\Sigma^{+}\to n e^{+}\nu_{e}\), which has \(\Delta S=+1\) and \(\Delta Q=-1\), while allowing \(\Sigma^{-}\to n e^{-}\bar{\nu}_{e}\). The forbidden mode is bounded at the \(10^{-5}\) level and the allowed one occurs at \(10^{-3}\) [Navas:2024]: two decays of the same multiplet, differing by two orders of magnitude, with the sign of a charge as the only distinction.

Proposition 101.57 (The $\Delta I=\tfrac{1}{2}$ rule, and its limits).

The \(\Delta S=1\) current \(\bar{\psi}_{u}\Gamma\psi_{s}\) carries isospin \(\tfrac{1}{2}\), since \(u\) is a member of an isospin doublet and \(s\) an isospin singlet. Semileptonic \(\Delta S=1\) decays therefore obey \(\Delta I=\tfrac{1}{2}\) exactly. For nonleptonic decays the current–current product contains both \(\Delta I=\tfrac{1}{2}\) and \(\Delta I=\tfrac{3}{2}\), so the rule is not a consequence of the structure; the observed dominance of the first by a factor of about \(20\) in amplitude is a strong-interaction effect and is not explained here. Rests on Equation (101.63).

Proof.

Derives Proposition 101.57. The isospin assignment is immediate. For the nonleptonic case, the effective operator is a product of two hadronic currents, each of isospin \(\tfrac{1}{2}\) in the \(\Delta S=1\) sector and each of isospin \(1\) or \(0\) in the \(\Delta S=0\) sector; combining \(\tfrac{1}{2}\otimes1\) gives \(\tfrac{1}{2}\oplus\tfrac{3}{2}\), so both appear with comparable coefficients at the level of the operator. The measurement is the ratio of the two kaon decay amplitudes \(K^{0}\to\pi^{+}\pi^{-}\) and \(K^{+}\to\pi^{+}\pi^{0}\), the second of which is pure \(\Delta I=\tfrac{3}{2}\) because two pions in the \(I=0\) state cannot carry unit charge; the measured ratio of amplitudes is about \(1/22\) [Navas:2024]. Honesty requires the statement that this factor is a QCD effect — short-distance renormalization of the operators plus long-distance hadronic dynamics — and is not derived in this book; it is the reason the rule is quoted here as an observation with its magnitude, rather than proved. The related but different question of why the operator basis is what it is belongs to The Renormalization Group.

Derivation pending.

The \({[}\Delta I=1/2{]}\) enhancement in nonleptonic strange decays: why the isospin-one-half amplitude dominates the isospin-three-halves one by a factor near twenty, when the current–current operator produces both with comparable coefficients. The explanation combines short-distance operator mixing under the renormalization group with long-distance hadronic matrix elements, so the derivation belongs with the strong-interaction and renormalization-group material rather than here; the measured ratio is quoted with its source above.

The GIM mechanism

Equation (101.62) solves one problem and creates another, and the second is severe. If the state \(d'\) couples to the \(W\), it also appears in whatever neutral current the theory possesses — and it appears bilinearly.

Proposition 101.58 (A single rotated doublet produces a flavour-changing neutral current).

With the Cabibbo rotation Equation (101.62) and one down-type state \(d'\),

\begin{equation}\tag{101.64} \bar{\psi}_{d'}\Gamma\psi_{d'} =\cos^{2}\theta_{C}\,\bar{\psi}_{d}\Gamma\psi_{d} +\sin^{2}\theta_{C}\,\bar{\psi}_{s}\Gamma\psi_{s} +\sin\theta_{C}\cos\theta_{C} \left(\bar{\psi}_{d}\Gamma\psi_{s} +\bar{\psi}_{s}\Gamma\psi_{d}\right)\ec \end{equation}

whose last term changes strangeness without changing charge, with a coefficient \(\sin\theta_{C}\cos\theta_{C}=0.219\) — of the same order as everything else in the theory. Rests on Equation (101.62).

Proof.

Derives Proposition 101.58. Substitute Equation (101.62) and expand. The cross terms are the last bracket.

Such a term would drive \(K^{0}\to\mu^{+}\mu^{-}\), \(K^{0}\)–\(\bar{K}^{0}\) mixing and \(K^{+}\to\pi^{+}\nu\bar{\nu}\) at rates comparable with ordinary weak processes. The measured branching ratio is \(\mathcal{B}(K_{L}\to\mu^{+}\mu^{-})=6.84\times 10^{-9}\) [Navas:2024], against \(\mathcal{B}(K^{+}\to\mu^{+}\nu_{\mu}) =0.636\): the neutral mode is suppressed by some eight orders of magnitude relative to the charged one. Equation (101.64) is therefore not merely unobserved but excluded by a very large factor, and something must cancel it.

Theorem 101.59 (Glashow–Iliopoulos–Maiani).

Let the down-type states be rotated by an orthogonal matrix and supplied to two up-type partners, so that the two charged-current doublets are

\begin{equation}\tag{101.65} \begin{pmatrix}u\\ d'\end{pmatrix}\ec\qquad \begin{pmatrix}c\\ s'\end{pmatrix}\ec\qquad \begin{aligned} d'&=\ \ \,d\cos\theta_{C}+s\sin\theta_{C}\ec\\ s'&=-d\sin\theta_{C}+s\cos\theta_{C}\ec \end{aligned} \end{equation}

with \(c\) a fourth quark. Then the neutral current is diagonal:

\begin{equation}\tag{101.66} \bar{\psi}_{d'}\Gamma\psi_{d'}+\bar{\psi}_{s'}\Gamma\psi_{s'} =\bar{\psi}_{d}\Gamma\psi_{d}+\bar{\psi}_{s}\Gamma\psi_{s}\ec \end{equation}

and no flavour-changing neutral current survives at tree level [Glashow:1970]. Rests on Proposition 101.58 and Equation (101.62).

Proof.

Derives Theorem 101.59. The two rotated states are the components of \(\psi_{d'}=R\psi_{d}\) with \(R\) the orthogonal matrix of Equation (101.65). Then \(\sum_{i}\bar{\psi}_{i'}\Gamma\psi_{i'} =\bar{\psi}_{j}\left(R\transpose R\right)_{jk}\Gamma\psi_{k} =\bar{\psi}_{j}\delta_{jk}\Gamma\psi_{k}\), which is Equation (101.66). Explicitly, the cross terms are \(+\sin\theta_{C}\cos\theta_{C}\) from \(d'\) and \(-\sin\theta_{C}\cos\theta_{C}\) from \(s'\), and they cancel exactly. Nothing about the value of \(\theta_{C}\) was used: the cancellation is the orthogonality of the rotation, not a numerical accident, which is why it survives to all orders in \(\theta_{C}\) and generalizes without change to the unitary matrix of Section 101.5.4.

Remark 101.60 (A particle predicted by a null result).

Theorem 101.59 is worth pausing on as a piece of scientific reasoning. The input is the absence of a process; the output is the existence of a particle. Glashow, Iliopoulos and Maiani proposed the fourth quark in 1970 [Glashow:1970] on no other evidence, and it was found four years later as the \(J/\psi\) resonance.

The cancellation is exact only at zero momentum transfer and only if the two up-type quarks were degenerate. At one loop, in the box diagram that generates \(K^{0}\)–\(\bar{K}^{0}\) mixing, the two contributions differ by the difference of the squared quark masses, and the surviving amplitude is suppressed by \(\left(m_{c}^{2}-m_{u}^{2}\right)c^{4}/\left(M_{W}c^{2}\right)^{2}\) rather than cancelling outright. The measured \(K_{L}\)–\(K_{S}\) mass difference, \(\Delta m_{K}c^{2}=3.484\times 10^{-12}\,\mathrm{MeV}\) [Navas:2024] — a fractional mass splitting of \(7\times10^{-15}\), one of the smallest measured quantities in physics — therefore bounds the charm mass, and the bound obtained before the discovery was of order a few \(\mathrm{GeV}/c^{2}\), which is where the charm quark was found. The estimate is due to Gaillard and Lee, whose paper is not in this book's bibliography and is therefore named without a citation rather than cited to a key that does not exist; the mechanism itself is [Glashow:1970], and the loop calculation belongs to Flavour Physics and Neutrinos.

Three generations: the CKM matrix

Kobayashi and Maskawa observed in 1973, before the fourth quark had been seen and while the third generation was entirely hypothetical, that the rotation of Equation (101.65) generalizes to \(N\) generations as an \(N\times N\) unitary matrix, and that the number of physically meaningful parameters it contains depends on \(N\) in a way that has an observable consequence [Kobayashi:1973].

Definition 101.61 (The quark mixing matrix).

With three generations the charged-current interaction of the quarks is

\begin{equation}\tag{101.67} J^{\mu}_{\mathrm{had}} =\sum_{i,k}\bar{\psi}_{u_{i}}\gamma^{\mu} \left(\identity-\gamma^{5}\right)V_{ik}\,\psi_{d_{k}}\ec \end{equation}

where \(i\) runs over \((u,c,t)\), \(k\) over \((d,s,b)\), and \(V\) is the Cabibbo–Kobayashi–Maskawa matrix [Cabibbo:1963] [Kobayashi:1973]. Its entries are dimensionless pure numbers.

Theorem 101.62 (The mixing matrix is unitary).

\(V\) is unitary, \(V^{\dagger}V=VV^{\dagger}=\identity\). Rests on Propositions 5.148 and 5.150.

Proof.

Derives Theorem 101.62. \(V\) is the matrix relating two orthonormal bases of one three-dimensional complex internal space — the basis of definite mass and the basis that couples to the \(W\) — and a change between orthonormal bases is unitary. That the space carries an inner product for which both bases are orthonormal is not automatic and is what has to be supplied: it holds because the flavour symmetry acting on the space can be realized by inner-product-preserving operators, which is Proposition 5.150, and because the space then decomposes cleanly, which is Proposition 5.148. Those two propositions are exactly where the licence to write the word “unitary” here comes from, and the argument is set out in Remark 5.152, which states it for this case and records the hypotheses under which each of the two fails. What unitarity buys is set out there as well: two families of relations, the row relations \(\sum_{k}V_{ik}V_{jk}^{*}=\delta_{ij}\) and the column relations \(\sum_{i}V_{ik}^{*}V_{il}=\delta_{kl}\), whose off-diagonal cases are sums of three complex numbers equal to zero and therefore closed triangles in the complex plane. The one that is measured is the subject of Experiment: CP Violation.

Proposition 101.63 (Parameter counting, and why three generations permit $CP$ violation).

An \(N\times N\) unitary mixing matrix contains

\begin{equation}\tag{101.68} \frac{N\left(N-1\right)}{2}\ \text{angles}\qquad\text{and}\qquad \frac{\left(N-1\right)\left(N-2\right)}{2}\ \text{phases} \end{equation}

that cannot be removed by redefining the phases of the quark fields. For \(N=2\) there is one angle and no phase, and the interaction is automatically invariant under \(CP\); for \(N=3\) there are three angles and one phase, and \(CP\) invariance is violated unless that phase happens to vanish. Rests on Theorem 101.62, Equation (101.67) and Theorem 101.59.

Proof.

Derives Proposition 101.63. A unitary \(N\times N\) matrix has \(N^{2}\) real parameters, of which \(N(N-1)/2\) are the angles of the real orthogonal subgroup and the remaining \(N^{2}-N(N-1)/2=N(N+1)/2\) are phases. The quark fields may be rephased freely, \(\psi_{u_{i}}\mapsto\ee^{\ii\alpha_{i}} \psi_{u_{i}}\) and \(\psi_{d_{k}}\mapsto\ee^{\ii\beta_{k}} \psi_{d_{k}}\), which changes \(V_{ik}\mapsto \ee^{\ii(\beta_{k}-\alpha_{i})}V_{ik}\) and leaves every other term of the Lagrangian invariant, the mass terms because they are diagonal and the neutral currents because they are diagonal by Theorem 101.59. There are \(2N\) such phases, but the overall common phase \(\alpha_{i}=\beta_{k}\) acts trivially on \(V\), so \(2N-1\) of the phases in \(V\) can be removed. The number surviving is \(N(N+1)/2-(2N-1)=(N-1)(N-2)/2\), which with the angles gives Equation (101.68).

For the \(CP\) statement, note that the \(CP\) conjugate of Equation (101.67) is the same expression with \(V\) replaced by \(V^{*}\). If every entry of \(V\) can be made real by rephasing, the interaction is \(CP\) invariant; if a phase survives, it cannot, and \(CP\) is violated. With \(N=2\), \((N-1)(N-2)/2=0\): the Cabibbo matrix is real and Equation (101.62) conserves \(CP\) however \(\theta_{C}\) is chosen. With \(N=3\) one phase survives. Kobayashi and Maskawa's inference [Kobayashi:1973] was made in the reverse direction: \(CP\) violation had been observed in the neutral kaon system nine years earlier [Christenson:1964], no two-generation model could accommodate it without additional structure, and a third generation could. Two of its three quarks were subsequently found. It is one of the few occasions in this chapter on which theory ran ahead of experiment, and it is worth noting that it did so by counting parameters.

Remark 101.64 (What this chapter does and does not claim about flavour).

The measured magnitudes of the mixing elements, the closure of the unitarity triangle, the size of the \(CP\)-violating phase and the whole neutrino sector belong to Flavour Physics and Neutrinos, and the experiments that establish them to Experiment: CP Violation and Experiment: Neutrino Oscillations. What has been established here is narrower and prior to all of it: that the charged-current interaction has one universal strength; that the hadronic states it couples are rotations of the mass eigenstates; that the rotation is unitary, with the licence for the word supplied by Remark 5.152; that unitarity is what removes flavour-changing neutral currents; and that the number of surviving parameters, which is pure counting, decides whether \(CP\) can be violated at all. Nothing in the list is a fit. Every one of the five statements is a structural consequence of Equation (101.35) plus the observation that the two bases differ.

Neutral currents and the intermediate bosons

The intermediate vector boson hypothesis

Fermi's contact term was always understood as an approximation. The electromagnetic interaction it was modelled on is not a contact term: two currents are joined by a photon, and the interaction has infinite range because the photon is massless. Yukawa's insight [Yukawa:1935], made for the nuclear force, was that the range of an exchange force is set by the mass of what is exchanged.

Proposition 101.65 (Range and mass).

A field \(\phi\) obeying the Klein–Gordon equation with mass \(M\) and a static point source produces the potential

\begin{equation}\tag{101.69} V(r)\propto\frac{\ee^{-r/R}}{r}\ec\qquad R=\frac{\hbar}{Mc}\ec \end{equation}

so an interaction of observed range \(R\) requires a quantum of mass \(M=\hbar/Rc\); equivalently, a momentum transfer \(q\) small compared with \(Mc\) cannot resolve the exchange and the interaction appears pointlike. Rests on Equation (101.6).

Proof.

Derives Proposition 101.65. The static field equation is \(\left(\nabla^{2}-\left(Mc/\hbar\right)^{2}\right)\phi =-g\,\delta^{3}(\vect{x})\), whose spherically symmetric solution is obtained by writing \(\phi=u(r)/r\), giving \(u''=(Mc/\hbar)^{2}u\) and hence \(u\propto\ee^{-Mcr/\hbar}\) for the solution decaying at infinity. In momentum space the same statement is that the propagator is \(\left[q^{2}+(Mc)^{2}\right]^{-1}\) in the Euclidean region, which for \(q\ll Mc\) is the constant \((Mc)^{-2}\): the propagator loses all memory of the momentum transfer, and the amplitude reduces to a product of two currents at one point, which is Equation (101.6). That is why Fermi's theory works, and it also says exactly how well: the neglected terms are of relative order \(q^{2}/(Mc)^{2}\).

The weak interaction has no measurable range, which by Proposition 101.65 means a very heavy quantum; but that argument alone gives only a lower bound. The quantitative statement comes from \(G_{F}\) itself, and requires knowing the coupling of the current to the boson. Glashow supplied the structure in 1961 [Glashow:1961]: the charged currents \(J^{\mu}_{1}\pm\ii J^{\mu}_{2}\) of Equation (101.35) do not close on themselves under commutation but generate a third, neutral current, so a gauge theory of the weak interaction must have a group at least as large as \(\SU(2)\), and once the electromagnetic \(\U(1)\) is included the smallest possibility is \(\SU(2)_{\mathrm{L}}\times\U(1)_{Y}\). Weinberg [Weinberg:1967] and Salam [Salam:1968] completed it with the mechanism that gives the bosons mass, and 't Hooft proved that the result is renormalizable [tHooft:1971a] [tHooft:1972] — which is what made the whole construction believable, since without that proof it would have been no better off than the theory it replaced. The derivations belong to Electroweak Unification and the Higgs Boson; what is needed here is the structure of the couplings and their SI dimensions.

Definition 101.66 (The electroweak couplings in SI).

Let \(W^{a}_{\mu}\) (\(a=1,2,3\)) and \(B_{\mu}\) be the gauge fields of \(\SU(2)_{\mathrm{L}}\) and \(\U(1)_{Y}\), each of the dimension of the electromagnetic four-potential, \(\mathrm{V}\,\mathrm{s}/\mathrm{m}\), and let the covariant derivative be

\begin{equation}\tag{101.70} D_{\mu}=\pp_{\mu}+\frac{\ii}{\hbar} \left(g\,W^{a}_{\mu}T^{a}+g'\,B_{\mu}\frac{Y}{2}\right)\ec \end{equation}

in exact parallel with Equation (100.9). Then \(g\) and \(g'\) carry the dimension of electric charge, \(\mathrm{C}\), and the dimensionless couplings are

\begin{equation}\tag{101.71} \alpha_{W}=\frac{g^{2}}{4\pi\varepsilon_{0}\hbar c}\ec\qquad \alpha'=\frac{g'^{2}}{4\pi\varepsilon_{0}\hbar c}\ec \end{equation}

built exactly as \(\alpha\) is in Equation (100.3). The generators are \(T^{a}=\tfrac{1}{2}\sigma^{a}\) acting on left-chiral doublets and zero on right-chiral singlets, and \(Y\) is the weak hypercharge.

Proposition 101.67 (The mixing angle and the fermion couplings).

Let the neutral fields be rotated by an angle \(\theta_{W}\),

\begin{equation}\tag{101.72} W^{3}_{\mu}=\cos\theta_{W}\,Z_{\mu}+\sin\theta_{W}\,A_{\mu}\ec\qquad B_{\mu}=-\sin\theta_{W}\,Z_{\mu}+\cos\theta_{W}\,A_{\mu}\ep \end{equation}

Requiring \(A_{\mu}\) to couple to the electric charge \(Q\) and to nothing else fixes

\begin{equation}\tag{101.73} e=g\sin\theta_{W}=g'\cos\theta_{W}\ec\qquad Q=T^{3}+\frac{Y}{2}\ec \end{equation}

and then \(Z_{\mu}\) couples to

\begin{equation}\tag{101.74} \frac{e}{\sin\theta_{W}\cos\theta_{W}} \left(T^{3}-\sin^{2}\theta_{W}\,Q\right) =\frac{g}{2\cos\theta_{W}}\,\gamma^{\mu} \left(g_{V}-g_{A}\gamma^{5}\right)\ec \end{equation}

with, for each fermion,

\begin{equation}\tag{101.75} g_{V}=T^{3}-2Q\sin^{2}\theta_{W}\ec\qquad g_{A}=T^{3}\ec \end{equation}

\(T^{3}\) being the third component of weak isospin of the fermion's left-chiral part. Also \(\alpha_{W}=\alpha/\sin^{2}\theta_{W}\). Rests on Equation (101.70).

Proof.

Derives Proposition 101.67. Substituting Equation (101.72) into the neutral part of Equation (101.70), the coefficient of \(A_{\mu}\) is \(g\sin\theta_{W}T^{3}+g'\cos\theta_{W}Y/2\). For this to be proportional to a single generator \(Q\) — which is what it means for the photon to couple to electric charge — the two coefficients must agree, \(g\sin\theta_{W}=g'\cos\theta_{W}=:e\), and then the coefficient is \(e\left(T^{3}+Y/2\right)\), giving Equation (101.73). The coefficient of \(Z_{\mu}\) is \(g\cos\theta_{W}T^{3}-g'\sin\theta_{W}Y/2\), which on eliminating \(g\) and \(g'\) in favour of \(e\) and substituting \(Y/2=Q-T^{3}\) becomes

\[ \frac{e}{\sin\theta_{W}\cos\theta_{W}} \left[\cos^{2}\theta_{W}T^{3} -\sin^{2}\theta_{W}\left(Q-T^{3}\right)\right] =\frac{e}{\sin\theta_{W}\cos\theta_{W}} \left(T^{3}-\sin^{2}\theta_{W}Q\right)\ep \]

Since \(T^{3}\) acts only on the left-chiral part, write \(T^{3}\to T^{3}P_{\mathrm{L}} =\tfrac{1}{2}T^{3}\left(\identity-\gamma^{5}\right)\); then \(T^{3}P_{\mathrm{L}}-\sin^{2}\theta_{W}Q =\tfrac{1}{2}\left[\left(T^{3}-2Q\sin^{2}\theta_{W}\right) -T^{3}\gamma^{5}\right]\), which with \(e/\sin\theta_{W}=g\) is Equation (101.74) and Equation (101.75). Finally \(e=g\sin\theta_{W}\) squared and divided by \(4\pi\varepsilon_{0}\hbar c\) gives \(\alpha=\alpha_{W}\sin^{2}\theta_{W}\).

Proposition 101.68 (The low-energy limit reproduces Fermi).

Let the charged-current interaction be

\begin{equation}\tag{101.76} \Lag_{\mathrm{CC}}=-\frac{gc}{2\sqrt{2}} \left[\bar{\psi}_{\nu}\gamma^{\mu} \left(\identity-\gamma^{5}\right)\psi_{\ell}\,W^{+}_{\mu} +\text{h.c.}\right]\ec \end{equation}

and let the \(W\) propagator be that of a massive vector field,

\begin{equation}\tag{101.77} \widetilde{D}^{W}_{\mu\nu}(q) =\frac{-\ii\mu_{0}\hbar^{3}c} {q^{2}-M_{W}^{2}c^{2}+\ii\epsilon} \left(\eta_{\mu\nu} -\frac{q_{\mu}q_{\nu}}{M_{W}^{2}c^{2}}\right)\ec \end{equation}

which is Equation (100.17) with the mass restored. Then for \(\abs{q^{2}}\ll M_{W}^{2}c^{2}\) the exchange of a \(W\) between two such currents is indistinguishable from Equation (101.46) with

\begin{equation}\tag{101.78} \frac{G_{F}}{\sqrt{2}}=\frac{g^{2}\mu_{0}\hbar^{2}}{8M_{W}^{2}}\ec \qquad\text{equivalently}\qquad \frac{G_{F}}{\sqrt{2}\left(\hbar c\right)^{3}} =\frac{\pi\alpha_{W}}{2\left(M_{W}c^{2}\right)^{2}} =\frac{\pi\alpha} {2\sin^{2}\theta_{W}\left(M_{W}c^{2}\right)^{2}}\ep \end{equation}

Rests on Equation (101.46), Equation (100.17) and Proposition 100.14.

Proof.

Derives Proposition 101.68. The vertex factor read off Equation (101.76) by the rule of Proposition 100.14 — expand \(\exp(\ii S/\hbar)\) with \(S=c^{-1}\int\dd^{4}x\,\Lag\) — is \(-\ii g\gamma^{\mu}(\identity-\gamma^{5})/2\sqrt{2}\hbar\), of the same SI dimension \(\mathrm{C}/\mathrm{J}/\mathrm{s}\) as the QED vertex Equation (100.24), since \(g\) is a charge by Definition 101.66. Two vertices joined by Equation (101.77) contribute, in the limit \(q^{2} \to 0\),

\[ \left(\frac{-\ii g}{2\sqrt{2}\hbar}\right)^{2} \frac{-\ii\mu_{0}\hbar^{3}c\,\eta_{\mu\nu}} {q^{2}-M_{W}^{2}c^{2}} \;\longrightarrow\; -\frac{\ii g^{2}\mu_{0}\hbar}{8M_{W}^{2}c}\,\eta_{\mu\nu}\ec \]

the \(q_{\mu}q_{\nu}\) term being dropped because it is contracted with currents of light external fermions, on which it gives a factor \(m_{\ell}c/M_{W}c\) squared by the argument of Equation (100.18). The contact interaction Equation (101.46), treated the same way, contributes \(-\ii G_{F}/\sqrt{2}\hbar c\) times the same spinor structure. Setting the two equal,

\[ \frac{G_{F}}{\sqrt{2}\,\hbar c} =\frac{g^{2}\mu_{0}\hbar}{8M_{W}^{2}c}\ec \]

which is the first form of Equation (101.78). For the second, substitute \(g^{2}=4\pi\varepsilon_{0}\hbar c\,\alpha_{W}\) from Equation (101.71) and \(\mu_{0}=1/\varepsilon_{0}c^{2}\):

\[ \frac{G_{F}}{\sqrt{2}} =\frac{4\pi\varepsilon_{0}\hbar c\,\alpha_{W}\hbar^{2}} {8M_{W}^{2}\varepsilon_{0}c^{2}} =\frac{\pi\alpha_{W}\hbar^{3}}{2M_{W}^{2}c}\ec \]

and dividing by \((\hbar c)^{3}\) gives \(\pi\alpha_{W}/2\left(M_{W}c^{2}\right)^{2}\). The dimensional check is worth doing once: \(g^{2}\mu_{0}\hbar^{2}/M_{W}^{2}\) carries \(\mathrm{C}^{2}\times \mathrm{kg}\,\mathrm{m}/\mathrm{C}^{2}\times \mathrm{J}^{2}\,\mathrm{s}^{2}/\mathrm{kg}^{2} =\mathrm{J}\,\mathrm{m}^{3}\), which is Equation (101.7).

Remark 101.69 (The problem has moved, not gone).

Equation (101.78) contains two unknowns where Fermi's theory had one, and it makes the dimensionful \(G_{F}\) into the ratio of a dimensionless coupling to a mass. That is progress only if the mass is itself explained, and a massive vector field is not innocent: the term \(\tfrac{1}{2}M^{2}W_{\mu}W^{\mu}\) is not invariant under the gauge transformation Equation (100.8), and the \(q_{\mu}q_{\nu}/M^{2}c^{2}\) in Equation (101.77) does not fall off at large \(q\), which destroys the power counting of Theorem 100.17 exactly as \(G_{F}\) does. Inserting the boson by hand therefore buys a finite range and loses renormalizability. What buys both is spontaneous breaking of a gauge symmetry, and that is the subject of Electroweak Unification and the Higgs Boson; 't Hooft's theorems [tHooft:1971a] [tHooft:1972] are the statement that it works.

Gargamelle and the weak neutral current

The most immediate prediction of Proposition 101.67 is that a fourth current exists, coupling to the \(Z\), which changes neither charge nor flavour. Nothing in Fermi's theory contains such a thing — the observed weak processes all involve a charge change — so its existence or absence was a decisive test.

Proposition 101.70 (The effective neutral-current interaction).

At \(\abs{q^{2}}\ll M_{Z}^{2}c^{2}\), and given the tree-level relation \(M_{W}=M_{Z}\cos\theta_{W}\) of Phenomenon 104.35, the neutral-current interaction of a neutrino with a fermion \(f\) is

\begin{equation}\tag{101.79} \Lag_{\mathrm{NC}}=-\frac{G_{F}}{\sqrt{2}} \left[\bar{\psi}_{\nu}\gamma^{\alpha} \left(\identity-\gamma^{5}\right)\psi_{\nu}\right] \left[\bar{\psi}_{f}\gamma_{\alpha} \left(g_{V}^{f}-g_{A}^{f}\gamma^{5}\right)\psi_{f}\right]\ec \end{equation}

with the same \(G_{F}\) as the charged current and the couplings of Equation (101.75). The neutral-current strength is therefore not a new parameter; the only new number is \(\sin^{2}\theta_{W}\). Rests on Proposition 101.68, Equation (101.74) and Equation (101.75).

Proof.

Derives Proposition 101.70. Repeat the matching of Proposition 101.68 with the \(Z\) vertex \(-\ii g\gamma^{\mu}(g_{V}-g_{A}\gamma^{5})/2\hbar \cos\theta_{W}\) of Equation (101.74) and the \(Z\) propagator, which is Equation (101.77) with \(M_{W}\to M_{Z}\). The effective coefficient is \(\left(g/2\cos\theta_{W}\right)^{2}\mu_{0}\hbar^{2}/M_{Z}^{2}\) times the two current structures. For the neutrino, \(T^{3}=+\tfrac{1}{2}\) and \(Q=0\) give \(g_{V}^{\nu}=g_{A}^{\nu}=\tfrac{1}{2}\) by Equation (101.75), so its factor \(\tfrac{1}{2}(\identity-\gamma^{5})\) contributes an extra \(\tfrac{1}{2}\). Hence the coefficient is

\[ \frac{1}{2}\cdot\frac{g^{2}\mu_{0}\hbar^{2}} {4\cos^{2}\theta_{W}M_{Z}^{2}} =\frac{g^{2}\mu_{0}\hbar^{2}}{8M_{W}^{2}} =\frac{G_{F}}{\sqrt{2}}\ec \]

using \(M_{Z}^{2}\cos^{2}\theta_{W}=M_{W}^{2}\) and then Equation (101.78). Two independent factors have conspired: the \(Z\) is heavier than the \(W\), which weakens the interaction, and it couples more strongly by \(1/\cos\theta_{W}\), which strengthens it by exactly the same amount. That the two cancel is the content of the relation \(\rho=1\) of Equation (104.45), and it is a prediction of the doublet structure of the symmetry breaking rather than a fitted coincidence.

Proposition 101.71 (Neutrino–electron elastic scattering).

For \(E_{\nu}\gg m_{e}c^{2}\), with the squared centre-of-mass energy \(E_{\mathrm{cm}}^{2}=2m_{e}c^{2}E_{\nu}\) for a neutrino of laboratory energy \(E_{\nu}\) on an electron at rest,

\begin{equation}\tag{101.80} \sigma\left(\nu_{\mu}e^{-}\to\nu_{\mu}e^{-}\right) =\frac{G_{F}^{2}E_{\mathrm{cm}}^{2}}{\pi\left(\hbar c\right)^{4}} \left(g_{L}^{2}+\frac{g_{R}^{2}}{3}\right)\ec\qquad \sigma\left(\bar{\nu}_{\mu}e^{-}\to\bar{\nu}_{\mu}e^{-}\right) =\frac{G_{F}^{2}E_{\mathrm{cm}}^{2}}{\pi\left(\hbar c\right)^{4}} \left(g_{R}^{2}+\frac{g_{L}^{2}}{3}\right)\ec \end{equation}

with \(g_{L}=\tfrac{1}{2}(g_{V}+g_{A})=-\tfrac{1}{2} +\sin^{2}\theta_{W}\) and \(g_{R}=\tfrac{1}{2}(g_{V}-g_{A})=\sin^{2}\theta_{W}\) for the electron. The ratio to the charged-current process Equation (101.60) is \(g_{L}^{2}+g_{R}^{2}/3\), which for \(\sin^{2}\theta_{W}=0.231\) is \(0.090\). Rests on Equations (101.60) and (101.79).

Proof.

Derives Proposition 101.71. Split the electron current into its chiral parts, \(g_{V}-g_{A}\gamma^{5} =g_{L}\left(\identity-\gamma^{5}\right) +g_{R}\left(\identity+\gamma^{5}\right)\), which is the definition of \(g_{L}\) and \(g_{R}\). The \(g_{L}\) piece is Equation (101.46) with a different coefficient, so Equation (101.60) applies to it directly and gives an isotropic cross-section proportional to \(g_{L}^{2}\).

The \(g_{R}\) piece requires one new ingredient, and it is angular momentum rather than algebra. In the centre-of-mass frame let the neutrino move along \(+\hat{\vect{z}}\). It is left-handed, so \(S_{z}^{\nu}=-\tfrac{1}{2}\). A left-chiral electron moving along \(-\hat{\vect{z}}\) has its spin along \(+\hat{\vect{z}}\), \(S_{z}^{e}=+\tfrac{1}{2}\), so the pair has \(J_{z}=0\) and can be in a \(J=0\) state: the amplitude may be isotropic, and is. A right-chiral electron moving along \(-\hat{\vect{z}}\) has \(S_{z}^{e}=-\tfrac{1}{2}\), so \(J_{z}=-1\) and the state must have \(J\geq1\); the amplitude is then proportional to the rotation matrix element \(d^{1}_{-1,-1}(\theta^{*})=\tfrac{1}{2} \left(1+\cos\theta^{*}\right)\), which vanishes in the backward direction because a backward scatter would reverse \(J_{z}\) without anything to absorb the two units. Hence

\begin{equation}\tag{101.81} \frac{\dd\sigma}{\dd\cos\theta^{*}} \propto g_{L}^{2}+g_{R}^{2} \left(\frac{1+\cos\theta^{*}}{2}\right)^{2}\ec \end{equation}

and integrating over \(\cos\theta^{*}\in[-1,1]\) gives \(2g_{L}^{2}\) and \(2g_{R}^{2}/3\) respectively, which is the factor \(g_{L}^{2}+g_{R}^{2}/3\) in Equation (101.80) once the normalization is fixed against Equation (101.60). For the antineutrino every helicity reverses, so \(g_{L}\) and \(g_{R}\) exchange roles. Numerically, at \(\sin^{2}\theta_{W}=0.231\), \(g_{L}=-0.269\) and \(g_{R}=0.231\), giving \(0.0901\) for the neutrino and \(0.0775\) for the antineutrino, and

\[ \sigma\left(\nu_{\mu}e^{-}\to\nu_{\mu}e^{-}\right) =1.55\times 10^{-46}\,\mathrm{m}^{2} \times\frac{E_{\nu}}{1\,\mathrm{GeV}}\ec\qquad \sigma\left(\bar{\nu}_{\mu}e^{-}\to\bar{\nu}_{\mu}e^{-}\right) =1.34\times 10^{-46}\,\mathrm{m}^{2} \times\frac{E_{\bar{\nu}}}{1\,\mathrm{GeV}}\ec \]

which is the size of the signal Gargamelle was looking for. Note that the \(g_{V}\) and \(g_{A}\) enter the two cross-sections in different combinations, so measuring both determines the sign of \(g_{A}\) as well as the value of \(\sin^{2}\theta_{W}\) — a single measurement would not.

Phenomenon 101.72 (Neutrinos scatter without becoming charged leptons).

A muon neutrino can transfer energy and momentum to matter and leave the apparatus still a neutrino. Both the leptonic channel \(\bar{\nu}_{\mu}e^{-}\to\bar{\nu}_{\mu}e^{-}\) [Hasert:1973a] and the hadronic channel \(\nu_{\mu}N\to\nu_{\mu}X\) [Hasert:1973b] were seen in the Gargamelle bubble chamber as events carrying hadrons or a recoil electron but no outgoing muon, at a rate a substantial fraction of the charged-current rate. Such a current changes neither electric charge nor flavour, and Fermi's theory contains nothing of the kind. Rests on Proposition 101.70, Proposition 101.71 and Phenomenon 101.5.

Derivation. Derives Phenomenon 101.72. That the process exists at all is Proposition 101.70: a gauge theory containing the observed charged currents cannot avoid a neutral one, because the commutator of the two charged generators is a third generator. That it occurs at the observed rate is Proposition 101.71 applied to the leptonic channel Gargamelle actually saw, which was the antineutrino one: its predicted cross-section per unit beam energy is \(1.34\times 10^{-46}\,\mathrm{m}^{2}/\mathrm{GeV}\) at \(\sin^{2}\theta_{W}=0.231\), the neutrino channel being the larger of the two at \(1.55\times 10^{-46}\,\mathrm{m}^{2}/\mathrm{GeV}\). Gargamelle recorded a single such event in the whole exposure, which was the entire leptonic sample and which no background could produce, since an electron recoiling almost exactly along the beam direction with nothing else in the chamber has no other source [Hasert:1973a]. The hadronic channel supplied the statistics: some hundreds of muonless events, with a ratio to the charged-current rate of roughly one to five [Hasert:1973b].

The obstacle was background, and it is worth recording because the signal is defined by an absence. A neutron entering the chamber from the surrounding material produces hadrons and no muon, which is exactly the signature. The argument that settled it was the spatial distribution: neutrons are attenuated as they penetrate the liquid, so a neutron-induced sample is concentrated near the walls and falls off with depth, whereas a neutrino-induced sample is uniform because the chamber is transparent to neutrinos by Phenomenon 101.5. The observed distribution was flat. The same reasoning applied to the ratio of muonless to muonful events as a function of position gave a rate that could not be produced by neutrons at any plausible flux.

Derivation pending.

The hadronic neutral-current ratio: the prediction for the ratio of neutral-current to charged-current rates on an isoscalar nuclear target, which is what the Gargamelle hadronic sample actually measured. It requires the quark-parton content of the nucleon and the quark neutral-current couplings derived in this section, and therefore belongs with the deep-inelastic material rather than here; the leptonic ratio, which needs nothing beyond this chapter, is derived above.

Discovery of the $W$ and $Z$

The bosons themselves were produced ten years later. The difficulty was energy: Equation (101.78) predicted masses near \(80\,\mathrm{GeV}\)\(/c^{2}\) and \(90\,\mathrm{GeV}\)\(/c^{2}\), and no machine reached the required centre-of-mass energy. Rubbia's proposal was to convert the CERN Super Proton Synchrotron into a proton — antiproton collider, which requires no second ring but does require enough antiprotons; van der Meer's stochastic cooling supplied them [vanderMeer:1985], by sensing the deviation of a sample of the beam at one point in the ring and applying a correcting kick at another, thereby reducing the phase-space volume of a stored beam — an operation that looks at first sight like a violation of Liouville's theorem and is not, because the feedback system carries away the entropy.

Phenomenon 101.73 (The weak interaction is carried by heavy bosons).

Two massive vector particles are produced in \(p\bar{p}\) collisions and observed as resonances: the \(W\), as an isolated high-transverse-momentum electron recoiling against missing transverse energy [Arnison:1983a] [Banner:1983], and the \(Z\), as a narrow peak in the invariant mass of charged lepton pairs [Arnison:1983b] [Bagnaia:1983]. Their masses are

\begin{equation}\tag{101.82} M_{W}c^{2}=80.37\,\mathrm{GeV}\ec\qquad M_{Z}c^{2}=91.1880\,\mathrm{GeV} \end{equation}

[Navas:2024]. Because these are enormous compared with the energy released in any nuclear decay, the exchange is indistinguishable from a contact interaction at every energy at which beta decay was ever studied. Rests on Equation (101.78), Proposition 101.71 and Proposition 100.63.

Derivation. Derives Phenomenon 101.73. The masses were predicted before they were measured, from two quantities determined in entirely different experiments. Inverting Equation (101.78),

\begin{equation}\tag{101.83} \left(M_{W}c^{2}\right)^{2} =\frac{\pi\alpha}{\sqrt{2}\sin^{2}\theta_{W}} \cdot\frac{\left(\hbar c\right)^{3}}{G_{F}}\ec \end{equation}

and inserting \(G_{F}/(\hbar c)^{3} =1.1663787\times 10^{-5}\,/\mathrm{GeV}^{2}\) from the muon lifetime, \(\alpha^{-1}=127.95\) at this scale from Proposition 100.63, and \(\sin^{2}\theta_{W}=0.2312\) from the neutral-current measurements of Proposition 101.71 and Phenomenon 104.55,

\[ \left(M_{W}c^{2}\right)^{2} =\frac{\pi\left(1/127.95\right)} {\sqrt{2}\left(0.2312\right)\left(1.1663787\times 10^{-5}\right)} \,\mathrm{GeV}^{2} =6438\,\mathrm{GeV}^{2}\ec \]

that is \(M_{W}c^{2}=80.2\,\mathrm{GeV}\), and \(M_{Z}=M_{W}/\cos\theta_{W}\) gives \(M_{Z}c^{2}=91.5\,\mathrm{GeV}\). The measured values are Equation (101.82): the tree-level prediction is right to a few parts in a thousand, the residual being the electroweak radiative correction that Section 101.6.4 turns into a measurement in its own right.

Read the other way, the same relation shows why Fermi's contact theory worked so well for so long. Writing \(Q^{2}=-q^{2}c^{2}\) for the squared four-momentum transfer measured as a squared energy, the exchange contributes the propagator factor \(\left[Q^{2}+(M_{W}c^{2})^{2}\right]^{-1}\), and in a nuclear decay \(Q\) is a few \(\mathrm{MeV}\), so the neglected ratio \(Q^{2}/(M_{W}c^{2})^{2}\) is of order \(10^{-9}\). The corresponding range is \(\hbar/M_{W}c=2.5\times 10^{-18}\,\mathrm{m}\), a thousandth of the diameter of a proton, which is why no experiment before 1983 had resolved the exchange. Where it must fail is at \(Q\) approaching \(M_{W}c^{2}\), which is the unitarity problem of Section 101.7.1.

Remark 101.74 (Why the $W$ has no mass peak and the $Z$ does).

The difference is kinematic. The \(Z\) decays to \(e^{+}e^{-}\) or \(\mu^{+}\mu^{-}\), both measured, so the invariant mass of the pair is formed directly and the \(Z\) is a peak in it [Arnison:1983b] [Bagnaia:1983]. The \(W\) decays to a charged lepton and a neutrino, which is not measured, and at a hadron collider the longitudinal balance is unusable because the colliding partons carry unknown longitudinal momenta and most of the debris escapes down the beam pipe. The transverse balance is usable, the initial transverse momentum being zero: the vector sum of the observed transverse momenta, negated, is the neutrino's. The variable that then carries the mass is the transverse mass

\begin{equation}\tag{101.84} M_{\mathrm{T}}^{2}c^{4}=2p_{\mathrm{T}}^{\ell}c\, p_{\mathrm{T}}^{\nu}c \left(1-\cos\Delta\varphi\right)\ec \end{equation}

with \(\Delta\varphi\) the azimuthal angle between the two. It satisfies \(M_{\mathrm{T}}\leq M_{W}\), with equality for a transverse decay, so its distribution has a sharp edge at \(M_{W}\) — a Jacobian peak, from the vanishing of \(\dd M_{\mathrm{T}}/\dd\theta^{*}\) there — and equivalently the lepton's own transverse-momentum spectrum has an edge at \(M_{W}c/2\). UA1 and UA2 report exactly this: isolated electrons of \(p_{\mathrm{T}}\approx40\,\mathrm{GeV}/c\) against an equal and opposite missing transverse energy [Arnison:1983a] [Banner:1983], which is still, forty years later, the operational signature of a neutrino and of anything else that leaves a detector without interacting.

Precision electroweak measurements

Producing the \(Z\) in \(e^{+}e^{-}\) annihilation, where the initial state is known exactly, converts it from a discovery into a metrological instrument. LEP ran at and around the resonance from 1989 to 1995 and recorded some seventeen million \(Z\) decays; the combined analysis of the four experiments [Schael:2006] fixes \(M_{Z}\) to two parts in \(10^{5}\) and every partial width to a fraction of a per cent. One of those measurements counts particles that were never detected.

Proposition 101.75 (Ratios of $Z$ partial widths).

For final-state fermions light compared with the \(Z\),

\begin{equation}\tag{101.85} \frac{\Gamma\left(Z\to f\bar{f}\right)} {\Gamma\left(Z\to f'\bar{f}'\right)} =\frac{N_{c}^{f}\left[\left(g_{V}^{f}\right)^{2} +\left(g_{A}^{f}\right)^{2}\right]} {N_{c}^{f'}\left[\left(g_{V}^{f'}\right)^{2} +\left(g_{A}^{f'}\right)^{2}\right]}\ec \end{equation}

with \(N_{c}=1\) for leptons and \(3\) for quarks. In absolute terms

\begin{equation}\tag{101.86} \Gamma\left(Z\to f\bar{f}\right) =\frac{N_{c}^{f}}{6\sqrt{2}\pi\hbar} \frac{G_{F}}{\left(\hbar c\right)^{3}} \left(M_{Z}c^{2}\right)^{3} \left[\left(g_{V}^{f}\right)^{2} +\left(g_{A}^{f}\right)^{2}\right]\ep \end{equation}

Rests on Equation (101.48), Lemma 101.47 and Equation (101.74).

Proof.

Derives Proposition 101.75. By Equation (101.48) the width is \(\left(c^{2}/2M_{Z}c^{2}\right)\int \overline{\abs{\mathcal{M}}^{2}}\dd\Phi_{2}\), and for massless final states \(\int\dd\Phi_{2}=1/8\pi\hbar^{2}\) by Lemma 101.47. The squared amplitude, summed over final spins and averaged over the three \(Z\) polarizations, involves the trace

\[ T^{\mu\nu}=\tr\left[\gamma\cdot p'\,\gamma^{\mu} \left(g_{V}-g_{A}\gamma^{5}\right)\gamma\cdot p\,\gamma^{\nu} \left(g_{V}-g_{A}\gamma^{5}\right)\right]\ec \]

contracted with the polarization sum \(\sum_{\lambda}\varepsilon^{*}_{\mu}\varepsilon_{\nu} =-\eta_{\mu\nu}+q_{\mu}q_{\nu}/M_{Z}^{2}c^{2}\). Moving the first factor \(\left(g_{V}-g_{A}\gamma^{5}\right)\) rightwards through \(\gamma\cdot p\) and then through \(\gamma^{\nu}\) — each passage flipping the sign of the \(\gamma^{5}\) term — collects the two into \(\left(g_{V}-g_{A}\gamma^{5}\right)^{2} =g_{V}^{2}+g_{A}^{2}-2g_{V}g_{A}\gamma^{5}\). The \(\gamma^{5}\) piece produces a Levi-Civita tensor, antisymmetric in \(\mu\nu\), which vanishes against the symmetric polarization sum. Hence the entire dependence on the fermion species is the factor \(g_{V}^{2}+g_{A}^{2}\), together with the colour multiplicity \(N_{c}\), and Equation (101.85) follows without evaluating anything else. Carrying the remaining common factor through, with the coupling \(g/2\cos\theta_{W}\) of Equation (101.74) re-expressed through Equation (101.78) and \(M_{W}=M_{Z}\cos\theta_{W}\), gives Equation (101.86). As a check, for a neutrino \(g_{V}=g_{A}=\tfrac{1}{2}\) and Equation (101.86) gives, as an energy width,

\[ \hbar\Gamma_{\nu\bar{\nu}} =\frac{\left(1.1663787\times 10^{-5}\right) \left(91.1880\right)^{3}\left(0.5\right)} {6\sqrt{2}\pi}\,\mathrm{GeV} =166\,\mathrm{MeV}\ec \]

against the measured \(167.2\,\mathrm{MeV}\) [Schael:2006]: a tree-level formula good to one per cent, the difference being the radiative corrections. Widths are quoted throughout as energies, \(\hbar\Gamma\), which is the convention of the source; the rate itself is \(\Gamma\) and carries \(/\mathrm{s}\).

Phenomenon 101.76 (There are three light neutrino species).

The \(Z\) resonance is wider than the sum of its visible decay channels, and the invisible remainder counts the neutrino species light enough for the \(Z\) to decay into. The first scan already gave three to within a few per cent [Decamp:1989]; the final four-experiment combination gives

\begin{equation}\tag{101.87} N_{\nu}=2.984\pm0.008 \end{equation}

[Schael:2006]. A fourth generation with a neutrino lighter than \(M_{Z}c^{2}/2\) is excluded at more than fifteen standard deviations. Rests on Equations (101.75) and (101.85).

Derivation. Derives Phenomenon 101.76. Measure the total width \(\Gamma_{Z}\) from the shape of the resonance in \(e^{+}e^{-}\to\text{anything visible}\), the hadronic width from the hadronic cross-section, and the leptonic width from the lepton-pair cross-sections. The invisible width is the difference, \(\Gamma_{\mathrm{inv}}=\Gamma_{Z}-\Gamma_{\mathrm{had}} -3\Gamma_{\ell\ell}\). Then

\begin{equation}\tag{101.88} N_{\nu}=\frac{\Gamma_{\mathrm{inv}}}{\Gamma_{\ell\ell}} \left(\frac{\Gamma_{\ell\ell}}{\Gamma_{\nu\bar{\nu}}}\right)_ {\mathrm{SM}}\ec \end{equation}

which is arranged so that the absolute normalization of Equation (101.86) cancels and only the ratio Equation (101.85) is needed. That ratio is pure algebra: for the neutrino \(g_{V}=g_{A}=\tfrac{1}{2}\) gives \(g_{V}^{2}+g_{A}^{2}=\tfrac{1}{2}\), while for a charged lepton \(T^{3}=-\tfrac{1}{2}\) and \(Q=-1\) give, by Equation (101.75), \(g_{A}=-\tfrac{1}{2}\) and \(g_{V}=-\tfrac{1}{2}+2\sin^{2}\theta_{W}=-0.0376\) at \(\sin^{2}\theta_{W}=0.23122\), so \(g_{V}^{2}+g_{A}^{2}=0.2514\) and

\[ \left(\frac{\Gamma_{\nu\bar{\nu}}}{\Gamma_{\ell\ell}}\right)_ {\mathrm{SM}}=\frac{0.5}{0.2514}=1.989\ep \]

The measured ratio is \(\Gamma_{\mathrm{inv}}/\Gamma_{\ell\ell}=5.943\pm0.016\) [Schael:2006], whence \(N_{\nu}=5.943/1.989=2.988\), and the full treatment with radiative corrections gives Equation (101.87). The near-cancellation of \(g_{V}\) for the charged lepton — \(-\tfrac{1}{2}+2(0.231)\) is small because \(\sin^{2}\theta_{W}\) happens to lie near \(\tfrac{1}{4}\) — is why \(\Gamma_{\ell\ell}\) is almost exactly half \(\Gamma_{\nu\bar{\nu}}\), and it is a numerical accident with no significance beyond making the arithmetic memorable.

Two caveats belong with the number. It counts only neutrinos with \(m_{\nu}<M_{Z}/2\) that couple to the \(Z\) with standard strength: a heavier or a sterile neutrino is invisible to this measurement, and Flavour Physics and Neutrinos records what other evidence says about those. And it is a counting measurement whose result is not an integer, \(2.984\pm0.008\); the deficit below three is about two standard deviations and is understood as a residual effect in the Bhabha-scattering normalization, which is exactly the kind of statement Probability and Statistics is needed to assess.

The rest of the LEP programme, and the SLC programme beside it, turned the same resonance into a test of the theory's quantum corrections. The weak mixing angle is measured in half a dozen mutually independent ways — from the forward–backward asymmetries of the \(Z\) decays, from the left–right asymmetry with a polarized electron beam, from neutrino–nucleon scattering, from parity violation in atomic caesium [Bouchiat:1974], and from polarized electron–deuteron scattering [Prescott:1978], whose apparatus is Experiment: Parity Violation — at momentum transfers spanning six orders of magnitude, and one running, scheme-defined angle describes them all [Navas:2024]. That consistency, and not any single number, is the measurement. One entry does not fit — a measurement of the \(W\) mass sitting several standard deviations from the rest of the electroweak fit — and it is recorded here, because it is unresolved and this is the chapter that owns the quantity. What We Observe but Do Not Understand is not the place to look for it: that chapter collects the observations with no theoretical explanation at all — the flavour hierarchy, the strong \(CP\) problem, dark matter, dark energy, the baryon asymmetry — and does not treat the electroweak fit.

Why Fermi theory had to be replaced

Unitarity violation

Proposition 101.7 said that \(G_{F}\) carries \(\mathrm{J}\,\mathrm{m}^{3}\), so that the dimensionless number governing a weak amplitude at energy \(E\) is \(G_{F}E^{2}/(\hbar c)^{3}\). A dimensionless expansion parameter that grows with energy is a theory with an expiry date, and the date can be computed.

Phenomenon 101.77 (The contact cross-section grows without bound).

The measured cross-section for neutrino scattering rises linearly with the beam energy from \(1\,\mathrm{MeV}\) to several hundred \(\mathrm{GeV}\), as Equation (101.60) requires. The linear rise is a statement about a range and not about all energies, and the range is fixed by this chapter's own formulae: Equation (101.60) is the \(M_{W}\to\infty\) limit, and restoring the \(W\) propagator of Equation (101.77) damps the growth once the momentum transfer approaches \(M_{W}c\). For a nucleon target \(E_{\mathrm{cm}}^{2}=2m_{N}c^{2}E_{\nu}\), so the linear regime is \(E_{\nu}\ll\left(M_{W}c^{2}\right)^{2}/2m_{N}c^{2} =3.4\,\mathrm{TeV}\); already at the top of the accelerator range, \(E_{\nu}=350\,\mathrm{GeV}\), the ratio \(E_{\mathrm{cm}}^{2}/(M_{W}c^{2})^{2}\) has reached \(0.10\) and the neglect is no longer safe. The departure is itself observed, in the momentum-transfer dependence of deep-inelastic neutrino scattering (Experiment: Deep Inelastic Scattering). Rests on Equation (101.60), Equation (101.77) and Proposition 101.54.

Derivation. Derives Phenomenon 101.77. This is Proposition 101.54: the cross-section is \(G_{F}^{2}E_{\mathrm{cm}}^{2}/\pi(\hbar c)^{4}\), and \(E_{\mathrm{cm}}^{2}=2m_{\mathrm{target}}c^{2}E_{\mathrm{beam}}\) for a fixed target, so \(\sigma\) is proportional to the beam energy. The growth is not a defect of the calculation but a consequence of the dimension of the coupling: with only \(G_{F}\), \(\hbar\), \(c\) and \(E_{\mathrm{cm}}\) available, and \(\sigma\) an area, the answer must be \(G_{F}^{2}E_{\mathrm{cm}}^{2}/(\hbar c)^{4}\) times a pure number. Restoring the propagator replaces this by \(\sigma\propto G_{F}^{2}E_{\mathrm{cm}}^{2}/ \left[1+E_{\mathrm{cm}}^{2}/(M_{W}c^{2})^{2}\right]\), which stops rising once \(E_{\mathrm{cm}}\) reaches \(M_{W}c^{2}\) and tends to the constant \(G_{F}^{2}(M_{W}c^{2})^{2}\) instead; the turnover is the direct observation of the propagator, and hence of the \(W\) in a scattering experiment rather than as a resonance.

Theorem 101.78 (The contact theory violates unitarity).

Probability conservation bounds a cross-section proceeding through a single partial wave \(J\) by

\begin{equation}\tag{101.89} \sigma_{J}\leq\frac{4\pi\left(2J+1\right)}{k^{2}}\ec\qquad k=\frac{\abs{\vect{p}}}{\hbar}= \frac{E_{\mathrm{cm}}}{2\hbar c}\ \ \text{(massless)}\ec \end{equation}

that is \(\sigma_{J=0}\leq16\pi(\hbar c)^{2}/E_{\mathrm{cm}}^{2}\). Equation (101.60) is pure \(J=0\) and grows as \(E_{\mathrm{cm}}^{2}\), so it breaches the bound at

\begin{equation}\tag{101.90} E_{\mathrm{cm}}=\left(16\pi^{2}\frac {\left(\hbar c\right)^{6}}{G_{F}^{2}}\right)^{1/4} =1.04\,\mathrm{TeV}\ec \end{equation}

and the theory is inconsistent above it. Rests on Equation (101.60) and Proposition 101.54.

Proof.

Derives Theorem 101.78. The bound is the statement that the \(S\)-matrix is unitary. Writing the elastic amplitude in partial waves, \(f(\theta)=k^{-1}\sum_{J}(2J+1)a_{J}P_{J}(\cos\theta)\), unitarity of \(S_{J}=1+2\ii a_{J}\) forces \(\abs{a_{J}}\leq1\), whence \(\sigma_{J}=4\pi(2J+1)\abs{a_{J}}^{2}/k^{2} \leq4\pi(2J+1)/k^{2}\); the partial-wave apparatus is that of Scattering Theory. That the contact interaction produces only \(J=0\) is Proposition 101.54: the cross-section came out isotropic, and an isotropic amplitude is by definition pure \(s\)-wave. Setting \(G_{F}^{2}E_{\mathrm{cm}}^{2}/\pi(\hbar c)^{4} =16\pi(\hbar c)^{2}/E_{\mathrm{cm}}^{2}\) and solving gives Equation (101.90); with \(G_{F}/(\hbar c)^{3}=1.1663787\times 10^{-5}\,/\mathrm{GeV}^{2}\), \(E_{\mathrm{cm}}^{4}=16\pi^{2}/(1.1663787\times 10^{-5})^{2} \mathrm{GeV}^{4}\) and \(E_{\mathrm{cm}}=1038\,\mathrm{GeV}\).

Remark 101.79 (Which number to quote, and why they differ).

The scale at which Fermi theory fails is quoted in the literature as anything between \(300\,\mathrm{GeV}\) and \(1\,\mathrm{TeV}\), and the spread is not carelessness: the two ends of the range are different statements.

  • Equation (101.90), about \(1\,\mathrm{TeV}\), is the energy at which this particular channel saturates the \(J=0\) bound exactly. It depends on the channel through the spin and colour counting and on the convention for the bound, and moving either shifts it by tens of per cent.

  • Equation (101.9), \(E_{F}=\left((\hbar c)^{3}/G_{F}\right)^{1/2}=292.8\,\mathrm{GeV}\), is the energy at which the dimensionless expansion parameter \(G_{F}E^{2}/(\hbar c)^{3}\) reaches unity. That is where the perturbative expansion stops converging, which happens before the leading term saturates the bound; it is channel-independent, and it is the number the phrase “a few hundred \(\mathrm{GeV}\)” refers to.

Both say the same physics: the theory predicts its own replacement, and predicts it at a definite scale. What is remarkable is that the prediction was right — the \(W\) was found at \(80.4\,\mathrm{GeV}\)\(/c^{2}\), within a factor of a few of \(E_{F}\), and the Higgs boson at \(125\,\mathrm{GeV}\)\(/c^{2}\) (Experiment: The Higgs Boson Discovery) — and the argument is now the standard template for looking beyond the Standard Model, as Section 101.7.2 records.

Remark 101.80 (The cure, and its cost).

Replacing the contact vertex by \(W\) exchange cures the growth: by Equation (101.78) the amplitude acquires the factor \(\left(M_{W}c^{2}\right)^{2}/ \left[Q^{2}+\left(M_{W}c^{2}\right)^{2}\right]\), so the cross-section stops rising at \(Q\sim M_{W}c^{2}\) and falls thereafter, and Equation (101.89) is never approached. It does not, however, cure everything: the process \(\nu\bar{\nu}\to W^{+}W^{-}\), with longitudinally polarized \(W\)s whose polarization vectors grow as \(q^{\mu}/M_{W}c\), has an amplitude that again rises with energy, and the cancellation of that growth requires the scalar of Electroweak Unification and the Higgs Boson and fixes the shape of its couplings. The unitarity argument therefore does not stop at the \(W\); it is what made a Higgs-like state mandatory below about \(1\,\mathrm{TeV}\), which is why the LHC was built to that energy and what it found is Experiment: The Higgs Boson Discovery.

Non-renormalizability and the effective-theory reading

The second failure is not about any particular energy but about the perturbation series itself, and it was diagnosed by Dyson's power counting [Dyson:1949b] — the same analysis that established QED to be renormalizable, applied to a different vertex, giving the opposite answer.

Theorem 101.81 (Superficial degree of divergence of the contact theory).

For a connected diagram built from \(V\) four-fermion vertices Equation (101.6) with \(E\) external fermion lines, the superficial degree of divergence is

\begin{equation}\tag{101.91} D=4-\frac{3}{2}E+2V\ep \end{equation}

It grows with the order of perturbation theory, so infinitely many amplitudes diverge and infinitely many counterterms are needed. Rests on Equations (100.23), (100.26) and (101.6).

Proof.

Derives Theorem 101.81. Let \(L\) be the number of loops and \(P\) the number of internal fermion lines. Each loop contributes \(\dd^{4}\ell\) and each internal fermion propagator falls as \(\ell^{-1}\) by Equation (100.23), so \(D=4L-P\). Every vertex has four fermion ends, every internal line uses two and every external line one, giving \(4V=2P+E\); and the number of independent loops is \(L=P-V+1\). Eliminating,

\[ D=4\left(P-V+1\right)-P=3P-4V+4 =3\left(2V-\frac{E}{2}\right)-4V+4 =2V+4-\frac{3}{2}E\ec \]

which is Equation (101.91).

The contrast with Equation (100.26) is the whole point. There, \(D=4-\tfrac{3}{2}E_{e}-E_{\gamma}\) and \(V\) cancelled: the degree of divergence is a property of the amplitude, not of the order, so only finitely many amplitudes diverge and finitely many parameters absorb them, which is Definition 100.18. Here \(V\) does not cancel, and it enters with a positive sign. The four-point function has \(E=4\) and \(D=2V-2\): convergent at \(V=1\), quadratically divergent at \(V=2\), quartically at \(V=3\), and so on without end. Each new order requires a new counterterm of a new operator structure, and each new counterterm brings a new undetermined constant, so the theory has no predictive power beyond the order at which it is first fitted.

Corollary 101.82 (The diagnosis is dimensional).

The extra \(2V\) in Equation (101.91) is exactly the \(-2\) mass dimension of the coupling, counted once per vertex. A coupling of negative mass dimension is non-renormalizable and a coupling of non-negative mass dimension is not, and no calculation beyond dimensional analysis is required to tell which. Rests on Equation (101.91) and Proposition 101.7.

Proof.

Derives Corollary 101.82. An amplitude of fixed external content has a fixed overall dimension. If a vertex carries a coupling of mass dimension \(-\delta\), then \(V\) vertices carry \(-V\delta\), and the loop integrals must supply \(+V\delta\) more powers of momentum to compensate, which is exactly \(D\to D+V\delta\). With \(\delta=2\) for \(G_{F}\) this is the \(2V\) of Equation (101.91). Conversely the QED coupling \(e/\hbar\) has \(\delta=0\) once the \(\hbar\)'s are counted — which is the content of \(\alpha\) being dimensionless, Equation (100.3) — and no such term appears.

Remark 101.83 (The modern reading: Fermi's theory is not wrong, it is effective).

The verdict of Theorem 101.81 was, for thirty years, that Fermi's theory was a stopgap. The modern reading is different, and it is the frame in which every current search for physics beyond the Standard Model is conducted.

Equation (101.6) is the leading term of a systematic expansion. Start from the theory with the \(W\) and integrate out the heavy field — expand the propagator Equation (101.77) in powers of \(q^{2}/M_{W}^{2}c^{2}\):

\[ \frac{1}{q^{2}-M_{W}^{2}c^{2}} =-\frac{1}{M_{W}^{2}c^{2}} \left(1+\frac{q^{2}}{M_{W}^{2}c^{2}} +\frac{q^{4}}{M_{W}^{4}c^{4}}+\cdots\right)\ep \]

The leading term is the contact interaction with the coefficient Equation (101.78); the next is a dimension-eight operator with two extra derivatives and two further powers of \(M_{W}\) in the denominator; and so on. The expansion is in \(E/M_{W}c^{2}\), it is systematic, and its accuracy at any energy can be estimated in advance — which is why Fermi's theory reproduces nuclear beta decay with the margin Phenomenon 101.73 computed.

Non-renormalizability is then not a disease but a measurement. An operator whose coefficient has mass dimension \(-2\) announces a scale, and that scale is where the description stops: Fermi's coupling announced Equation (101.9), \(292.8\,\mathrm{GeV}\), and the \(W\) was found at \(80.4\,\mathrm{GeV}\)\(/c^{2}\). How such operators mix under a change of scale, and why the leading ones dominate, is The Renormalization Group.

The same logic now runs in the other direction: the Standard Model is treated as the leading term, and searches are reported as bounds on the coefficients of the higher-dimension operators that could be added to it, equivalently as bounds on a scale \(\Lambda\). That is how a null result at a collider becomes a quantitative statement, and it is why the failure recorded in this section is now regarded as the most useful thing Fermi's theory ever did.

Remark 101.84 (What this chapter has and has not established).

The chapter began with a continuous spectrum and ends with two massive bosons. Everything between was extracted from measurement: the neutrino from kinematics (Phenomenon 101.1), the contact form and the size of the coupling from spectra and lifetimes (Proposition 101.11 and Phenomenon 101.50), the two nuclear couplings from the observed selection rules (Proposition 101.15), the violation of parity from one asymmetry (Phenomenon 101.26), the \(V-A\) structure from the size of that asymmetry together with the neutrino helicity (Proposition 101.36 and Phenomenon 101.28), the quark rotation from a deficit in rate (Phenomenon 101.55), a fourth quark from a null result (Theorem 101.59), a third generation from a parameter count (Proposition 101.63), and the neutral current from events with a missing muon (Phenomenon 101.72).

What has not been established is why any of it is so. The gauge group was quoted, not derived; the boson masses were matched to \(G_{F}\), not explained; the mixing angles and the fermion masses are inputs. Those are the business of Electroweak Unification and the Higgs Boson and Flavour Physics and Neutrinos — and, in the case of the mixing parameters, of nobody at present: they are measured numbers with no known origin, recorded as such in What We Observe but Do Not Understand and in the free-parameter count of the appendices.