The Lindeberg–Feller Central Limit Theorem
This appendix proves Theorem 15.48 of Probability and Statistics: a triangular array of independent, individually negligible summands whose tails satisfy Lindeberg's condition has an asymptotically Gaussian normalised sum. It is the form of the central limit theorem the error theory of Measurement, SI Units, and the Theory of Errors actually needs, because the disturbances that perturb a measurement — a thermal drift, a vibration, a quantisation step, a reading error — are independent but certainly not identically distributed, so the Lindeberg–Lévy theorem Theorem 15.46 does not apply to them.
The argument is the characteristic-function argument of Section 15.3.2 carried out uniformly in the summand index. Three of its four ingredients are already in the chapter: the elementary properties of characteristic functions (Proposition 15.43), the exponential remainder bound Equation (15.31), and the telescoping product estimate Equation (15.36). The fourth is Lévy's continuity theorem, proved in Lévy's Continuity Theorem. Nothing else is assumed, and in particular the Lindeberg condition is shown here to imply the uniform asymptotic negligibility that the expansion needs — that implication is not an extra hypothesis.
Statement and notation
For each \(n\) let \(X_{n,1},\dots,X_{n,k_{n}}\) be independent random variables with \(\avg{X_{n,k}}=0\) and finite variances \(\sigma_{n,k}^{2}=\operatorname{var}X_{n,k}\), and put \(s_{n}^{2}=\sum_{k=1}^{k_{n}}\sigma_{n,k}^{2}\), assumed strictly positive. If for every \(\varepsilon>0\)
then \(s_{n}^{-1}\sum_{k=1}^{k_{n}}X_{n,k}\rightarrow\mathcal{N}(0,1)\) in distribution. Rests on Definition 15.42, Proposition 15.43 and Definition 15.28.
It is convenient to normalise once and for all. Put
and write \(\varphi_{n,k}\) for the characteristic function of \(Y_{n,k}\). The event \(\set{\abs{X_{n,k}}>\varepsilon s_{n}}\) is the event \(\set{\abs{Y_{n,k}}>\varepsilon}\), so Equation (A19.1) reads
Every quantity in Equation (A19.2) is a pure number: the \(X_{n,k}\) carry whatever SI unit the measurement error carries, and \(s_{n}\) carries the same one, so the ratio is dimensionless and so is the argument \(t\) of the characteristic functions below.
Lindeberg's condition implies uniform asymptotic negligibility
Under Equation (A19.3),
Rests on Equation (A19.1).
Derives Lemma A19.2. Fix \(\varepsilon>0\) and \(k\). Splitting the expectation at \(\abs{Y_{n,k}}=\varepsilon\) and bounding \(Y_{n,k}^{2}\) by \(\varepsilon^{2}\) on the inner piece,
The last term is one summand of the non-negative sum \(L_{n}(\varepsilon)\), hence at most \(L_{n}(\varepsilon)\), and the bound is the same for every \(k\):
Letting \(n\rightarrow\infty\) gives \(\limsup_{n}\max_{k}\tau_{n,k}^{2}\le\varepsilon^{2}\), and \(\varepsilon>0\) was arbitrary.
∎This is Equation (15.39) of Probability and Statistics, proved. It is what licenses treating each factor of the product below as close to \(1\), and it is used twice: to know \(1-\tfrac12 t^{2}\tau_{n,k}^{2}\) lies in \([0,1]\), and to control the sum of squares in Lemma A19.5.
The uniform second-order expansion
Lemma 15.44 expands one characteristic function to second order, with a remainder \(o(t^{2})\) that depends on the distribution. Here there is a different distribution for every \((n,k)\) and the remainders must be summed, so what is needed is a bound on the total remainder, uniform in the array. The Lindeberg condition is exactly what supplies it.
For every \(t\in\R\),
Derives Lemma A19.3. Apply Equation (15.31) with \(n=2\) and \(x=tY_{n,k}\):
Take expectations and use \(\abs{\avg{Z}}\le\avg{\abs{Z}}\); since \(\avg{Y_{n,k}}=0\) and \(\avg{Y_{n,k}^{2}}=\tau_{n,k}^{2}\) the left side becomes the \(k\)th term of \(\Delta_{n}(t)\), so
Fix \(\varepsilon>0\) and split the expectation at \(\abs{Y_{n,k}}=\varepsilon\), using a different branch of the minimum on each piece. Where \(\abs{Y_{n,k}}\le\varepsilon\) the first branch gives \(\tfrac16\abs{t}^{3}\abs{Y_{n,k}}^{3} \le\tfrac16\abs{t}^{3}\varepsilon\,Y_{n,k}^{2}\); where \(\abs{Y_{n,k}}>\varepsilon\) the second gives \(t^{2}Y_{n,k}^{2}\). Hence
From Equation (A19.7) (split the expectation at \(\abs{Y_{n,k}}=\varepsilon\) and use one branch of the minimum on each piece). Summing over \(k\) and using \(\sum_{k}\tau_{n,k}^{2}=1\) from Equation (A19.2) and the definition Equation (A19.3),
For fixed \(t\) and \(\varepsilon\) the second term tends to \(0\), so \(\limsup_{n}\Delta_{n}(t)\le\varepsilon\abs{t}^{3}/6\); and \(\varepsilon>0\) was arbitrary.
∎The contrast with Lemma 15.44 is worth stating plainly. There the remainder was \(o(t^{2})\) as \(t\rightarrow0\) for a single distribution, and the identically distributed case could rescale one such estimate \(n\) times. Here \(t\) is fixed and \(n\) grows, and the estimate Equation (A19.9) is a statement about the array as a whole: the cubic branch of the minimum is used on the bulk of each summand and the quadratic branch on its tail, and Equation (A19.1) is precisely the hypothesis that the tails contribute nothing in total.
Two product estimates
For every \(t\in\R\) there is \(N(t)\) such that for \(n\ge N(t)\)
Rests on Equation (15.36), Lemma A19.3 and Lemma A19.2.
Derives Lemma A19.4. Set \(a_{k}=\varphi_{n,k}(t)\) and \(b_{k}=1-\tfrac12 t^{2}\tau_{n,k}^{2}\). By Proposition 15.43(i), \(\abs{a_{k}}\le1\). For the \(b_{k}\): by Lemma A19.2 there is \(N(t)\) with \(\tfrac12 t^{2}\max_{k}\tau_{n,k}^{2}\le1\) for \(n\ge N(t)\), and then every \(b_{k}\) lies in \([0,1]\). The telescoping inequality Equation (15.36) applies to such factors and gives \(\abs{\prod a_{k}-\prod b_{k}}\le\sum_{k}\abs{a_{k}-b_{k}}\), which is \(\Delta_{n}(t)\) by Equation (A19.6).
∎For every \(t\in\R\),
Rests on Equation (15.36), Lemma A19.2 and Theorem 11.38.
Derives Lemma A19.5. Write \(z_{k}=\tfrac12 t^{2}\tau_{n,k}^{2}\ge0\), so that \(\sum_{k}z_{k}=\tfrac12t^{2}\) by Equation (A19.2) and \(\max_{k}z_{k}\rightarrow0\) by Lemma A19.2. Take \(n\) large enough that \(\max_{k}z_{k}\le1\); then both \(1-z_{k}\) and \(\ee^{-z_{k}}\) lie in \([0,1]\) and Equation (15.36) applies:
Taylor's theorem with Lagrange remainder (Theorem 11.38) applied to \(z\mapsto\ee^{-z}\) gives \(\ee^{-z}=1-z+\tfrac12z^{2}\ee^{-\xi}\) with \(\xi\) between \(0\) and \(z\), so for \(z\ge0\) one has \(\abs{\ee^{-z}-1+z}\le\tfrac12z^{2}\). Hence
the middle step bounding each term by \(z_{k}^{2}/2\), factoring out the largest \(z_{k}\) and using \(\sum_{k}z_{k}=t^{2}/2\); while \(\prod_{k}\ee^{-z_{k}}=\ee^{-\sum_{k}z_{k}}=\ee^{-t^{2}/2}\) exactly.
∎Proof of the theorem
Proof of Theorem A19.1. Derives Theorem A19.1. Write \(T_{n}=s_{n}^{-1}\sum_{k}X_{n,k}=\sum_{k}Y_{n,k}\). The \(Y_{n,k}\), \(k=1,\dots,k_{n}\), are independent, so Proposition 15.43(iii), extended from two factors to \(k_{n}\) by induction, gives
Fix \(t\in\R\). By the triangle inequality,
whose first term tends to \(0\) by Lemma A19.4 and whose second tends to \(0\) by Lemma A19.5. Hence
From Equations (A19.10), (A19.11) and (A19.13) (replace each factor by its quadratic approximation, then the product of those by the exponential of the sum). The limit is continuous at \(t=0\), and it is the characteristic function of the standard Gaussian: as shown in the derivation of Theorem 15.46, the function \(g(t)=\int\ee^{\ii tu}\ee^{-u^{2}/2}\dd u/\sqrt{2\pi}\) satisfies \(g'(t)=-t\,g(t)\) with \(g(0)=1\), whence \(g(t)=\ee^{-t^{2}/2}\). Theorem A18.1 therefore gives \(T_{n}\rightarrow\mathcal{N}(0,1)\) in distribution.
∎Theorem 15.46 follows from Theorem A19.1. Rests on Theorem A19.1.
Derives Corollary A19.6. Let \(X_{1},X_{2},\dots\) be independent and identically distributed with mean \(\mu\) and variance \(\sigma^{2}\in(0,\infty)\), and set \(k_{n}=n\), \(X_{n,k}=X_{k}-\mu\). Then \(\sigma_{n,k}^{2}=\sigma^{2}\) and \(s_{n}^{2}=n\sigma^{2}\), so \(Z_{n}=s_{n}^{-1}\sum_{k}X_{n,k}\) is the statistic of Equation (15.33). The Lindeberg sum is
which tends to \(0\) as \(n\rightarrow\infty\) because \(\avg{(X-\mu)^{2}}=\sigma^{2}\) is finite, a finite integral being one whose tail contribution vanishes — the same elementary fact used at Equation (15.32). The threshold \(\varepsilon\sigma\sqrt{n}\) grows without bound, so the truncation level recedes and the tail expectation goes to zero for each fixed \(\varepsilon\).
∎If there is \(\delta>0\) with
then Equation (A19.1) holds, and hence so does the conclusion of Theorem A19.1. Rests on Theorem A19.1 and Equation (A19.1).
Derives Corollary A19.7. On the event \(\abs{Y_{n,k}}>\varepsilon\) one has \(1<\abs{Y_{n,k}}^{\delta}/\varepsilon^{\delta}\), so \(Y_{n,k}^{2}\indic{\abs{Y_{n,k}}>\varepsilon} \le\varepsilon^{-\delta}\abs{Y_{n,k}}^{2+\delta}\) pointwise. Taking expectations and summing over \(k\),
which tends to \(0\) for each fixed \(\varepsilon>0\) by Equation (A19.15).
∎The condition is sufficient and not necessary, and Probability and Statistics exhibits both failures of the converse: summands equal to \(\pm\sqrt{n}\) with probability \(1/(2n)\) have vanishing variance shares yet Lindeberg sum \(1\) for every \(n\), and adjoining one \(\mathcal{N}(0,n)\) summand to \(n\) standard Gaussians breaks Equation (A19.1) while leaving the normalised sum exactly \(\mathcal{N}(0,1)\). What Lemma A19.2 establishes is the one implication the error theory uses — no single disturbance may carry a fixed fraction of the total variance — and it is an implication, not a converse: a dominant source violates Equation (A19.1), but the absence of a dominant source does not by itself deliver a Gaussian. Feller's companion theorem, which closes the loop inside the negligible class, is not proved here [Feller:1971].
The Lindeberg–Feller Central Limit Theorem discharges the proof obligation of Theorem 15.48, stated in Section 15.3.2 of Probability and Statistics and used there to justify treating the aggregate of many small, non-identically distributed measurement disturbances as Gaussian — the licence on which the standard uncertainty of Measurement, SI Units, and the Theory of Errors rests. Two by-products are worth noting: Equation (15.39) of the chapter is Lemma A19.2 here, proved rather than asserted, and Lyapunov's condition, which the chapter quotes, is Corollary A19.7. The only input this appendix does not derive is Lévy's continuity theorem, which is Lévy's Continuity Theorem.