Wilks' Theorem in the Multi-Parameter Case
This appendix completes the derivation of Theorem 15.88 of Probability and Statistics in \(k\) parameters with \(r\) constraints. The chapter's argument is correct at the level of the quadratic approximation and two of its steps are asserted rather than proved:
-
[(i)] the remainder in the \(k\)-dimensional Taylor expansion of the log-likelihood must be shown uniform on a neighbourhood of the true parameter that shrinks at rate \(n^{-1/2}\), since the expansion is evaluated not at a fixed point but at two estimators that move with the data;
-
[(ii)] both estimators — the unconstrained one and the one constrained to \(\Theta_{0}\) — must be shown to lie in that neighbourhood with probability tending to one, since that is what licenses replacing the constraint surface \(\Theta_{0}\) by its tangent space at \(\theta_{0}\), the step written “\(\Theta_{0}\simeq\theta_{0}+E\)” at Equation (15.85).
Both are proved below, from exactly the regularity conditions (R1)–(R5) of Definition 15.59 and nothing stronger, and the chi-squared limit with \(r\) degrees of freedom is then assembled from them. The one-parameter case with a simple null needs none of this — there \(\Theta_{0}\) is a point, there is no surface to replace — which is why Probability and Statistics carries that derivation in full and defers only this one.
Setting and notation
Let \(X_{1},\dots,X_{n}\) be independent and identically distributed with density \(f(\,\cdot\,;\theta_{0})\), let \(\Theta\subseteq\R^{k}\) and let (R1)–(R5) hold at \(\theta_{0}\), read in \(k\) parameters: (R4) asks that \(\theta\mapsto f(x;\theta)\) be three times continuously differentiable near \(\theta_{0}\) with the first two derivatives available under the integral, and (R5) that the information matrix \(I=I_{1}(\theta_{0})\) of Definition 15.60 be finite and positive definite and that a single integrable \(M(x)\) dominate every third partial derivative \(\pp^{3}\ln f(x;\theta)/\pp\theta^{a}\pp\theta^{b}\pp\theta^{c}\) throughout a ball \(\bar B(\theta_{0},a_{0})\subseteq\Theta\). Write \(\lambda>0\) for the smallest eigenvalue of \(I\) and \(\norm{A}=\sup_{\norm{\vect v}=1}\norm{A\vect v}\) for the operator norm of a matrix.
Put \(\ell_{n}(\theta)=\sum_{i=1}^{n}\ln f(X_{i};\theta)\) and
All three are averages of independent identically distributed terms: \(\vect{Z}_{n}\) of the scores, whose common mean is zero and whose covariance is \(I\) (Lemma 15.61); \(J_{n}\) of the observation Hessians, whose common mean is \(-I\) by Equation (15.48); and \(\bar{M}_{n}\) of \(M\), whose common mean \(\bar{M}=\avg{M}\) is finite by (R5). The statistic is a pure number, the SI dimensions of \(\theta\) cancelling in the likelihood ratio of Equation (15.80), and so is every quantity above: \(\nabla\ell_{n}\) carries the reciprocal of the unit of \(\theta\) and \(I\) its square, so the combinations formed below — \(\vect{u}\transpose I\vect{u}\) with \(\vect{u}=\sqrt{n}(\theta-\theta_{0})\), and \(\vect{Z}_{n}\transpose I^{-1}\vect{Z}_{n}\) — are dimensionless.
The two estimators are the local ones, defined exactly as Probability and Statistics defines its root of the score equation. Fix \(a_{1}\in(0,a_{0}]\), to be chosen in Lemma A21.9, and let
Both maxima are attained: \(\ell_{n}\) is continuous and both sets are closed and bounded, hence compact, so the extreme value theorem (Theorem 11.24) applies. Both sets contain \(\theta_{0}\), so
the second because \(\theta_{0}\in\Theta_{0}\) by hypothesis. The statistic under study is \(q=-2\left[\ell_{n}(\hat{\theta}_{0,n}) -\ell_{n}(\hat{\theta}_{n})\right]\ge0\).
The null set is a \(C^{2}\) submanifold near \(\theta_{0}\), which is used in exactly one form and is therefore stated in that form.
\(\Theta_{0}\) is a \(C^{2}\) submanifold of dimension \(k-r\) through \(\theta_{0}\) with tangent space \(E\) when there are a linear subspace \(E\subseteq\R^{k}\) with \(\dim E=k-r\), a neighbourhood \(W\) of \(\theta_{0}\), an open \(V\ni\vect{0}\) in \(E\), and a twice continuously differentiable map \(h:V\rightarrow E^{\perp}\) with \(h(\vect{0})= \vect{0}\) and \(Dh(\vect{0})=0\), such that
Rests on Definition 15.87.
Theorem 15.88 states the hypothesis as “\(H_{0}\) imposes \(r\) smooth, functionally independent constraints”, i.e. \(\Theta_{0}\) is the level set \(\set{g(\theta)=\vect{0}}\) of a \(C^{2}\) map \(g:\Theta\rightarrow\R^{r}\) with \(Dg(\theta_{0})\) of rank \(r\). The equivalence of that description with Equation (A21.4) is the implicit function theorem in several variables, which Differentiable Manifolds, Tensors, and Curvature records as owed rather than proved. This appendix therefore takes Definition A21.1 as the hypothesis: it is the form the proof uses, it is one of the standard equivalent definitions of a submanifold, and adopting it keeps the argument free of a theorem the treatise has not yet built. A reader who has the implicit function theorem may read the two hypotheses as the same one.
What is quoted
Let \(W_{1},W_{2},\dots\) be independent and identically distributed with \(\avg{\abs{W}}<\infty\). Then \(n^{-1}\sum_{i=1}^{n}W_{i}\rightarrow\avg{W}\) in probability [Billingsley:1995]. Rests on Corollary 15.51.
Let \(\vect{W}_{n},\vect{W}\) be random vectors in \(\R^{k}\). Then \(\vect{W}_{n}\rightarrow\vect{W}\) in distribution if and only if \(\vect{a}\transpose\vect{W}_{n}\rightarrow \vect{a}\transpose\vect{W}\) in distribution for every \(\vect{a}\in\R^{k}\). Moreover, if \(\vect{W}_{n}\rightarrow\vect{W}\) in distribution and \(g:\R^{k}\rightarrow\R\) is continuous, then \(g(\vect{W}_{n})\rightarrow g(\vect{W})\) in distribution [Billingsley:1995]. Rests on Definition 15.21 and Lemma 15.52.
Three imports, all of them already relied on by Probability and Statistics itself, and no more. Theorem A21.3 is the weak law in the form that assumes only integrability; Corollary 15.51 is the Chebyshev form and needs a finite variance, which (R5) does not supply for the Hessian summands or for \(M\) — the chapter names this gap explicitly in the derivation of Theorem 15.67 and quotes the same source. Theorem A21.4 is what promotes the scalar central limit theorem Theorem 15.46 to the vector statement \(\vect{Z}_{n}\rightarrow\mathcal{N}(0,I)\) and lets a continuous function be applied under a limit in distribution; the chapter uses both silently in the \(k\)-parameter clause of Theorem 15.67 and proves the one-dimensional special case of the second at Equation (15.44). The third import is the implicit function theorem, and it is avoided rather than used, by taking Definition A21.1 as the hypothesis (Remark A21.2). Everything else below — the uniform remainder, the localisation of both estimators, the tangent-space replacement and the degrees of freedom — is derived here.
\(\norm{J_{n}+I}\rightarrow0\) and \(\bar{M}_{n}\rightarrow\bar{M}\) in probability, and \(\vect{Z}_{n}\rightarrow\mathcal{N}(0,I)\) in distribution. In particular each component of \(\vect{Z}_{n}\), and hence \(\norm{\vect{Z}_{n}}\), is bounded in probability. Rests on Lemma 15.61, Theorem A21.3 and Theorem A21.4.
Derives Lemma A21.6. Each entry of \(J_{n}\) is the average of \(n\) independent copies of \(\pp^{2}\ln f(X;\theta_{0})/\pp\theta^{a}\pp\theta^{b}\), integrable by (R4) with mean \(-I_{ab}\) by Equation (15.48), so Theorem A21.3 gives \(J_{n}\rightarrow-I\) entrywise in probability; a matrix norm is a continuous function of finitely many entries, so \(\norm{J_{n}+I}\rightarrow0\) in probability. The same theorem applied to \(M\), integrable by (R5), gives the second claim.
For the third, fix \(\vect{a}\in\R^{k}\), \(\vect{a}\neq\vect{0}\). Then \(\vect{a}\transpose\vect{Z}_{n} =n^{-1/2}\sum_{i}\vect{a}\transpose\vect{U}(\theta_{0};X_{i})\) with \(\vect{U}=\nabla\ln f\) the score, a normalised sum of independent identically distributed variables with mean zero (Lemma 15.61) and variance \(\vect{a}\transpose I\vect{a}\), finite and strictly positive by (R5). Theorem 15.46 gives \(\vect{a}\transpose\vect{Z}_{n}\rightarrow \mathcal{N}(0,\vect{a}\transpose I\vect{a})\), and Theorem A21.4 converts this into \(\vect{Z}_{n}\rightarrow\mathcal{N}(0,I)\). Boundedness in probability of each component is Equation (15.43), and \(\norm{\vect{Z}_{n}}\le\sum_{a}\abs{Z_{n}^{a}}\) is a finite sum of variables each bounded in probability.
∎The uniform quadratic expansion
This is item (i). The content is not that a remainder exists — Taylor gives that at every fixed \(\theta\) — but that a single random variable, not depending on \(\theta\), bounds it throughout the ball.
For every \(\theta\in\bar B(\theta_{0},a_{0})\), writing \(\delta=\theta-\theta_{0}\),
The bound is uniform in \(\theta\): apart from the explicit factor \(\norm{\delta}^{3}\) it involves only \(\bar{M}_{n}\), which does not depend on \(\theta\). Rests on Definition 15.59 and Theorem 11.38.
Derives Lemma A21.7. Fix \(\theta\) and set \(g(u)=\ell_{n}(\theta_{0}+u\delta)\) for \(u\in[0,1]\); the segment lies in \(\bar B(\theta_{0},a_{0})\) by convexity of the ball, so \(g\) is three times continuously differentiable by (R4). Taylor's theorem with Lagrange remainder (Theorem 11.38) gives \(g(1)=g(0)+g'(0)+\tfrac12g''(0)+\tfrac16g'''(\xi)\) for some \(\xi\in(0,1)\). By the chain rule (Proposition 11.104),
which are the first two displayed terms, and
Each third partial derivative of \(\ell_{n}\) is a sum of \(n\) terms, one per observation, and (R5) bounds each of them by \(M(X_{i})\) throughout \(\bar B(\theta_{0},a_{0})\) — in particular at the unknown intermediate point \(\theta_{0}+\xi\delta\), which is what makes the estimate independent of \(\xi\) and hence of \(\theta\). Since \(\abs{\delta^{a}}\le\norm{\delta}\) for each component and the sum Equation (A21.6) has \(k^{3}\) terms,
and \(R_{n}(\theta)=\tfrac16g'''(\xi)\).
∎Define, for \(\vect{u}\in\R^{k}\),
Then for every \(K>0\) and every \(n\) with \(K/\sqrt{n}\le a_{0}\),
where \(\epsilon_{n}(K)=K^{2}\norm{J_{n}+I} +\tfrac13k^{3}K^{3}\bar{M}_{n}/\sqrt{n}\) tends to \(0\) in probability. Rests on Lemmas A21.6 and A21.7.
Derives Corollary A21.8. Put \(\delta=\vect{u}/\sqrt{n}\) in Equation (A21.5) and multiply by \(-2\):
because \(n\delta\transpose J_{n}\delta =\vect{u}\transpose J_{n}\vect{u}\) and \(\sqrt{n}\vect{Z}_{n}\transpose\delta =\vect{Z}_{n}\transpose\vect{u}\). Subtracting Equation (A21.7) leaves \(-\vect{u}\transpose\left(J_{n}+I\right)\vect{u}-2R_{n}\), whose first part is at most \(\norm{J_{n}+I}\norm{\vect{u}}^{2}\) in modulus and whose second is at most \(\tfrac13k^{3}\bar{M}_{n}\norm{\vect{u}}^{3}/\sqrt{n}\) by Equation (A21.5) with \(\norm{\delta}=\norm{\vect{u}}/ \sqrt{n}\). Both are increasing in \(\norm{\vect{u}}\), so the supremum over the ball is attained at the bound. That \(\epsilon_{n}(K)\rightarrow0\) in probability is Lemma A21.6: the first term is \(K^{2}\) times something tending to \(0\), and the second is \(\bar{M}_{n}\), bounded in probability, times \(\tfrac13k^{3}K^{3}n^{-1/2}\rightarrow0\).
∎Localisation of both estimators
This is item (ii). Note that it is proved for any closed subset of the ball containing \(\theta_{0}\), so that the unconstrained and the constrained estimators are covered by one argument; nothing about the geometry of \(\Theta_{0}\) is used here.
Choose \(a_{1}\in(0,a_{0}]\) with \(\tfrac16k^{3}\left(\bar{M}+1\right)a_{1}\le\lambda/8\), and let \(A\subseteq\bar B(\theta_{0},a_{1})\) be any closed set with \(\theta_{0}\in A\). Let \(\hat{\theta}\) maximise \(\ell_{n}\) over \(A\). Then
and consequently \(\sqrt{n}\norm{\hat{\theta}-\theta_{0}}\) is bounded in probability and \(\hat{\theta}\rightarrow\theta_{0}\) in probability. In particular both estimators Equation (A21.2) satisfy this. Rests on Lemmas A21.6 and A21.7.
Derives Lemma A21.9. Write \(\hat{\delta}=\hat{\theta}-\theta_{0}\) and \(d=\norm{\hat{\delta}}\le a_{1}\). Since \(\theta_{0}\in A\) we have \(\ell_{n}(\hat{\theta})\ge\ell_{n}(\theta_{0})\), so Equation (A21.5) gives
Bound the three terms. The first is at most \(\sqrt{n}\norm{\vect{Z}_{n}}d\) by the Cauchy–Schwarz inequality for vectors. For the second, write \(\hat{\delta}\transpose J_{n}\hat{\delta} =-\hat{\delta}\transpose I\hat{\delta} +\hat{\delta}\transpose\left(J_{n}+I\right)\hat{\delta}\); the first piece is at most \(-\lambda d^{2}\), since \(\lambda\) is the smallest eigenvalue of the positive-definite \(I\), and the second is at most \(\norm{J_{n}+I}d^{2}\). The third is at most \(\tfrac16nk^{3}\bar{M}_{n}d^{3}\). Hence
From Equations (A21.5) and (A21.10) (bound the linear term by Cauchy–Schwarz, split the Hessian term into \(-I\) plus its deviation, and insert the uniform remainder bound).
Let \(E_{n}\) be the event on which both \(\norm{J_{n}+I}\le\lambda/4\) and \(\bar{M}_{n}\le\bar{M}+1\); by Lemma A21.6, \(\Pr(E_{n})\rightarrow1\). On \(E_{n}\), and using \(d\le a_{1}\) with the choice of \(a_{1}\),
Substituting in Equation (A21.11) and cancelling \(n\lambda d^{2}/4\) from both sides leaves \(\tfrac14n\lambda d^{2}\le\sqrt{n}\norm{\vect{Z}_{n}}d\); dividing by \(\tfrac14n\lambda d\) when \(d>0\) — and noting the conclusion is trivial when \(d=0\) — gives \(\sqrt{n}\,d\le4\norm{\vect{Z}_{n}}/\lambda\), which is Equation (A21.9). Boundedness in probability follows because \(\norm{\vect{Z}_{n}}\) is bounded in probability (Lemma A21.6), and consistency because \(\sqrt{n}d=O_{P}(1)\) forces \(d\rightarrow0\) in probability.
∎Replacing the surface by its tangent space
Under Definition A21.1 there are \(c_{0}<\infty\) and \(a_{2}\in(0,a_{1}]\) such that
-
if \(\theta\in\Theta_{0}\) and \(\norm{\theta-\theta_{0}}\le a_{2}\), then the orthogonal projection \(\vect{w}\) of \(\theta-\theta_{0}\) onto \(E\) satisfies \(\norm{\left(\theta-\theta_{0}\right)-\vect{w}} \le c_{0}\norm{\theta-\theta_{0}}^{2}\);
-
if \(\vect{e}\in E\) and \(\norm{\vect{e}}\le a_{2}\), there is \(\theta\in\Theta_{0}\) with \(\norm{\left(\theta-\theta_{0}\right)-\vect{e}} \le c_{0}\norm{\vect{e}}^{2}\).
Rests on Definition A21.1 and Theorem 11.38.
Derives Lemma A21.10. Choose \(a_{2}\in(0,a_{1}]\) small enough that the closed ball \(\bar B(\theta_{0},2a_{2})\) lies in \(W\) and that \(\set{\vect{e}\in E\mid\norm{\vect{e}}\le2a_{2}}\subseteq V\); let \(c_{0}=\tfrac12\sup\norm{D^{2}h}\) over that compact set, finite because \(h\) is twice continuously differentiable and the supremum of a continuous function on a compact set is attained (Theorem 11.24). Taylor's theorem applied to \(u\mapsto h(u\vect{e})\) on \([0,1]\), with \(h(\vect{0})=\vect{0}\) and \(Dh(\vect{0})=0\) killing the first two terms, gives
(ii) Take \(\theta=\theta_{0}+\vect{e}+h(\vect{e})\), which lies in \(\Theta_{0}\) by Equation (A21.4); then \((\theta-\theta_{0})-\vect{e}=h(\vect{e})\) and Equation (A21.12) applies.
(i) Such a \(\theta\) lies in \(W\), so \(\theta-\theta_{0}=\vect{e}+h(\vect{e})\) for some \(\vect{e}\in V\); the two summands are orthogonal, so \(\vect{w}=\vect{e}\) and \(\norm{\vect{e}}\le\norm{\theta-\theta_{0}}\le a_{2}\). Then \((\theta-\theta_{0})-\vect{w}=h(\vect{e})\), and Equation (A21.12) together with \(\norm{\vect{e}}\le\norm{\theta-\theta_{0}}\) gives the bound.
∎Write \(\hat{\vect{u}}=\sqrt{n}(\hat{\theta}_{n}-\theta_{0})\) and \(\hat{\vect{u}}_{0}=\sqrt{n}(\hat{\theta}_{0,n}-\theta_{0})\). Then
both in probability. Rests on Corollary A21.8, Lemma A21.9 and Lemma A21.10.
Derives Lemma A21.11. Since \(I\) is positive definite, \(\Lambda_{n}\) is a strictly convex quadratic; completing the square with \(\vect{m}=I^{-1}\vect{Z}_{n}\),
so the unconstrained minimum is attained at \(\vect{m}\) and the minimum over the subspace \(E\) at the \(I\)-orthogonal projection \(\vect{p}\) of \(\vect{m}\) onto \(E\). Both \(\norm{\vect{m}}\) and \(\norm{\vect{p}}\) are bounded in probability, being bounded linear images of \(\vect{Z}_{n}\). Also \(\Lambda_{n}\) is Lipschitz on each ball: for \(\norm{\vect{u}},\norm{\vect{u}'}\le K\),
which is bounded in probability for fixed \(K\).
Fix \(\varepsilon>0\). By Lemma A21.9 and the boundedness just noted, there is \(K\) such that with probability at least \(1-\varepsilon\) for all large \(n\) the four vectors \(\hat{\vect{u}},\hat{\vect{u}}_{0},\vect{m},\vect{p}\) all lie in \(\bar B(\vect{0},K)\); call that event \(A_{n}\) and intersect it with the event that \(\epsilon_{n}(K+1)\le\varepsilon\), which also has probability approaching \(1\) by Corollary A21.8. Work on the intersection, and abbreviate \(D_{n}(\theta)=-2\left[\ell_{n}(\theta)-\ell_{n}(\theta_{0})\right]\), so that Equation (A21.8) reads \(\abs{D_{n}(\theta_{0}+\vect{u}/\sqrt{n})-\Lambda_{n}(\vect{u})} \le\varepsilon\) for every \(\norm{\vect{u}}\le K+1\). The extra unit of radius is there because one auxiliary point below, the image in \(\Theta_{0}\) of the tangential minimiser, misses \(\bar B(\vect{0},K)\) by \(O(n^{-1/2})\); all other points used lie in \(\bar B(\vect{0},K)\). Write \(L_{n}=L_{n}(K+1)\) for the Lipschitz constant Equation (A21.15) at that radius.
First limit. That \(\Lambda_{n}(\hat{\vect{u}})\ge \min_{\R^{k}}\Lambda_{n}\) is immediate. For the reverse, put \(\theta_{\mathrm{m}}=\theta_{0}+\vect{m}/\sqrt{n}\), which lies in \(\bar B(\theta_{0},a_{1})\) once \(K/\sqrt{n}\le a_{1}\), so \(\ell_{n}(\hat{\theta}_{n})\ge\ell_{n}(\theta_{\mathrm{m}})\) and hence \(D_{n}(\hat{\theta}_{n})\le D_{n}(\theta_{\mathrm{m}})\). Therefore
so the first difference in Equation (A21.13) is at most \(2\varepsilon\) on an event of probability approaching \(1-\varepsilon\).
Second limit. For the lower bound let \(\vect{w}\) be the orthogonal projection of \(\hat{\vect{u}}_{0}\) onto \(E\). By Lemma A21.10(i) applied to \(\theta=\hat{\theta}_{0,n}\in\Theta_{0}\), whose distance to \(\theta_{0}\) is at most \(K/\sqrt{n}\le a_{2}\) for large \(n\),
and \(\norm{\vect{w}}\le\norm{\hat{\vect{u}}_{0}}\le K\), so Equation (A21.15) gives \(\Lambda_{n}(\hat{\vect{u}}_{0})\ge\Lambda_{n}(\vect{w}) -L_{n}c_{0}K^{2}/\sqrt{n}\ge\min_{E}\Lambda_{n} -L_{n}c_{0}K^{2}/\sqrt{n}\).
For the upper bound apply Lemma A21.10(ii) to \(\vect{p}/\sqrt{n}\), of norm at most \(K/\sqrt{n}\le a_{2}\): there is \(\theta_{\mathrm{p}}\in\Theta_{0}\) with \(\norm{\left(\theta_{\mathrm{p}}-\theta_{0}\right) -\vect{p}/\sqrt{n}}\le c_{0}K^{2}/n\), hence \(\norm{\sqrt{n}\left(\theta_{\mathrm{p}}-\theta_{0}\right)-\vect{p}} \le c_{0}K^{2}/\sqrt{n}\) and \(\theta_{\mathrm{p}}\in \Theta_{0}\cap\bar B(\theta_{0},a_{1})\) for large \(n\). Since \(\hat{\theta}_{0,n}\) maximises \(\ell_{n}\) over that set, \(D_{n}(\hat{\theta}_{0,n})\le D_{n}(\theta_{\mathrm{p}})\), and running the same three-step chain as before, then moving from \(\sqrt{n}(\theta_{\mathrm{p}}-\theta_{0})\) to \(\vect{p}\) by Equation (A21.15),
Both correction terms tend to \(0\) in probability, \(L_{n}\) being bounded in probability, and \(\varepsilon>0\) was arbitrary.
∎The chi-squared limit
Let (R1)–(R5) hold at \(\theta_{0}\) in \(k\) parameters, let the sample be independent and identically distributed of size \(n\), and let \(\Theta_{0}\ni\theta_{0}\) satisfy Definition A21.1 with \(\dim E=k-r\). Then the local likelihood-ratio statistic \(q=-2\left[\ell_{n}(\hat{\theta}_{0,n}) -\ell_{n}(\hat{\theta}_{n})\right]\) built from the estimators Equation (A21.2) converges in distribution to \(\chi^{2}_{r}\). Rests on Lemma A21.11, Lemma 15.83 and Theorem A21.4.
Derives Theorem A21.12. Write \(D_{n}\) as in Lemma A21.11, so that \(q=D_{n}(\hat{\theta}_{0,n})-D_{n}(\hat{\theta}_{n})\) — the common reference value \(\ell_{n}(\theta_{0})\) cancels. By Corollary A21.8 each \(D_{n}\) equals the corresponding \(\Lambda_{n}\) up to \(\epsilon_{n}(K)\), and by Lemma A21.11 each \(\Lambda_{n}\) equals the corresponding minimum up to a quantity tending to zero in probability. Hence
From Equations (A21.13) and (A21.14) (replace each statistic by the minimum of the quadratic form and cancel the common constant).
The second equality is Equation (A21.14), whose additive constant \(-\vect{Z}_{n}\transpose I^{-1}\vect{Z}_{n}\) is the same in both minima and cancels.
Now substitute \(\vect{V}_{n}=I^{1/2}\vect{m}=I^{-1/2}\vect{Z}_{n}\) and \(S=I^{1/2}E\), where \(I^{1/2}\) is any invertible matrix with \(I^{1/2}\left(I^{1/2}\right)\transpose=I\), symmetric here; the existence of one is the weak form of the spectral theorem recorded in Remark 15.86. Since \(I^{1/2}\) is invertible, \(\dim S=\dim E=k-r\), and
with \(P\) the orthogonal projector onto \(S^{\perp}\), of dimension \(r\): the nearest point of a subspace to a given vector is its orthogonal projection, and what is left over is the projection onto the orthogonal complement.
By Lemma A21.6, \(\vect{Z}_{n}\rightarrow \mathcal{N}(0,I)\) in distribution; applying Theorem A21.4 to the linear map \(\vect{z}\mapsto I^{-1/2}\vect{z}\) — for each \(\vect{a}\), \(\vect{a}\transpose I^{-1/2}\vect{Z}_{n}\rightarrow \mathcal{N}(0,\vect{a}\transpose I^{-1/2}II^{-1/2}\vect{a}) =\mathcal{N}(0,\norm{\vect{a}}^{2})\) — gives \(\vect{V}_{n}\rightarrow\vect{V}\sim\mathcal{N}(0,\identity)\) in distribution. The map \(\vect{v}\mapsto\norm{P\vect{v}}^{2}\) is continuous, so the continuous-mapping clause of Theorem A21.4 gives \(\norm{P\vect{V}_{n}}^{2}\rightarrow\norm{P\vect{V}}^{2}\) in distribution, and \(\norm{P\vect{V}}^{2}\sim\chi^{2}_{r}\) by Lemma 15.83. Finally, adding an \(o_{P}(1)\) term does not change a limit in distribution, by the sum clause of Lemma 15.52. Hence \(q\rightarrow\chi^{2}_{r}\).
∎Equation (15.80) defines \(q\) through the global suprema of the likelihood over \(\Theta\) and over \(\Theta_{0}\), while Theorem A21.12 is proved for the local maximisers Equation (A21.2). The gap is the same one Remark 15.69 records for Theorem 15.67: identifying a consistent local maximiser with the global one is Wald's theorem [Wald:1949], under conditions this treatise does not assume, and a likelihood with several local maxima can have a consistent root far from its highest peak at any finite \(n\). Stating the theorem for the local statistic is therefore not a weakening introduced here — it is the same statement the chapter's one-parameter derivation proves, made explicit. The number \(r\) is unaffected: it is \(\operatorname{codim}E\), a property of the constraint alone.
It is worth marking the two places, because they are the whole content of this appendix. Uniformity — item (i) — is Equation (A21.5): the remainder is bounded by \(\tfrac16nk^{3}\bar{M}_{n}\norm{\delta}^{3}\) with a coefficient that does not depend on \(\theta\), which is what (R5)'s single dominating function \(M\) buys and what a pointwise Taylor expansion would not. It is used at two moving points, \(\hat{\theta}_{n}\) and \(\hat{\theta}_{0,n}\), neither of which is known in advance; a \(\theta\)-dependent remainder would say nothing about either. Localisation — item (ii) — is Equation (A21.9): both estimators sit within \(O_{P}(n^{-1/2})\) of \(\theta_{0}\), so the surface is only ever probed at distance \(O(n^{-1/2})\), where Lemma A21.10 bounds its departure from the tangent space by \(O(n^{-1})\) — one order smaller, which after the \(\sqrt{n}\) rescaling of Equation (A21.16) is the \(O(n^{-1/2})\) that vanishes. That is the honest content of writing \(\Theta_{0}\simeq\theta_{0}+E\), and it is why second-order smoothness of the constraint, and not merely differentiability, is assumed.
Wilks' Theorem in the Multi-Parameter Case discharges the proof obligation left at Section 15.5.4 of Probability and Statistics. With it, Corollary 15.89 — one parameter of interest, any number of nuisance parameters, one degree of freedom — is a corollary of a proved theorem rather than of an argued one, and so is the practice it licenses: a collider search reads its significance off a one-degree-of-freedom distribution however many nuisance parameters describe the background and the detector [Cowan:2011]. The two failure modes remain exactly as Remark 15.90 and Remark 15.91 describe them, and both are visible in the proof above: a parameter on the boundary violates (R3), so Equation (A21.2) is no longer an interior maximisation and Lemma A21.9 loses the two-sided estimate; a parameter unidentified under the null makes \(I\) singular, so \(\lambda=0\) and every step from Equation (A21.11) onward fails at once.