The Hilbert–Schmidt Theorem for Compact Self-Adjoint Operators

Contents
  1. Statement
  2. Compactness in sequential form
  3. The supremum is attained: the first eigenvector
  4. The recursion
  5. Passage to the limit

This appendix proves Theorem 16.44 of Hilbert Spaces: a compact self-adjoint operator on a Hilbert space has an orthonormal system of eigenvectors, with real eigenvalues accumulating only at \(0\), in terms of which the operator is a norm- convergent sum of rank-one pieces. It is the one infinite-dimensional theorem that reproduces the finite-dimensional spectral theorem Theorem 9.79 without weakening any of its conclusions, and it is the theorem Hilbert proved for symmetric integral kernels [Hilbert:1912] and Schmidt recast in the language used here [Schmidt:1907].

Everything below rests on material already available: the numerical-radius formula \(\norm{A}=\sup_{\norm{x}=1}\abs{\braket{x}{Ax}}\) (Proposition 16.43), Bessel's inequality (Proposition 16.27), the projection theorem (Theorem 16.18), and the equivalence of compactness and sequential compactness in a metric space (Theorem 10.31). No weak topology is used; the closing Remark A23.8 says precisely what the standard weak-compactness argument would have supplied and why the elementary route replaces it.

Statement

Theorem A23.1 (Hilbert–Schmidt).

Let \(\mathcal{H}\neq\set{0}\) be a Hilbert space and let \(A\in\mathcal{B}(\mathcal{H})\) be compact and self-adjoint (Definition 16.41). Then there is a finite or countable orthonormal system \(\set{e_{n}}\) in \(\mathcal{H}\) and real numbers \(\lambda_{n}\neq0\) with \(\abs{\lambda_{1}}\geq\abs{\lambda_{2}}\geq\cdots\) such that

\begin{equation}\tag{A23.1} Ax=\sum_{n}\lambda_{n}\braket{e_{n}}{x}\,e_{n}\ec\qquad \forall\,x\in\mathcal{H}\ec \end{equation}

the series converging in the norm of \(\mathcal{H}\). If the system is infinite then \(\lambda_{n}\longrightarrow0\). Every eigenspace belonging to a non-zero eigenvalue is finite dimensional, \(Ae_{n}=\lambda_{n}e_{n}\), and \(\norm{A}=\max_{n}\abs{\lambda_{n}}\) when \(A\neq0\). Finally

\begin{equation}\tag{A23.2} \mathcal{H}=\ker A\oplus\overline{\text{span}\set{e_{n}}}\ec \end{equation}

so that adjoining any orthonormal basis of \(\ker A\) to \(\set{e_{n}}\) produces an orthonormal basis of \(\mathcal{H}\) consisting of eigenvectors of \(A\). Rests on Definition 16.41, Proposition 16.43 and Theorem 16.30.

The proof occupies the rest of the section: the sequential form of compactness (Compactness in sequential form), the attainment of the supremum and the first eigenvector (The supremum is attained: the first eigenvector), the recursion (The recursion), and the passage to the limit (Passage to the limit).

Compactness in sequential form

Lemma A23.2 (Sequential characterisation).

\(A\in\mathcal{B}(\mathcal{H})\) is compact if and only if every bounded sequence \((x_{n})\) in \(\mathcal{H}\) admits a subsequence \((x_{n_{k}})\) for which \((Ax_{n_{k}})\) converges in \(\mathcal{H}\). Rests on Definition 16.41 and Theorem 10.31.

Proof.

Derives Lemma A23.2. Suppose \(A\) is compact and let \(\norm{x_{n}}\leq c\) for all \(n\); we may assume \(c>0\). The vectors \(Ax_{n}/c\) lie in the image of the closed unit ball, hence in the set \(K=\overline{A\set{x\mid\norm{x}\leq1}}\), which is compact by hypothesis. A compact subset of the metric space \(\mathcal{H}\) is sequentially compact (Theorem 10.31), so some subsequence \(Ax_{n_{k}}/c\) converges in \(K\), and therefore \(Ax_{n_{k}}\) converges.

Conversely, assume the sequential property and let \((y_{m})\) be a sequence in \(K\). Each \(y_{m}\) is a limit of points \(Ax\) with \(\norm{x}\leq1\), so we may choose \(x_{m}\) with \(\norm{x_{m}}\leq1\) and \(\norm{y_{m}-Ax_{m}}<1/m\). By hypothesis some \(Ax_{m_{k}}\) converges, and then \(y_{m_{k}}\) converges to the same limit, which lies in the closed set \(K\). Hence \(K\) is sequentially compact, hence compact by Theorem 10.31.

Lemma A23.3 (Restriction to an invariant closed subspace).

Let \(A\) be compact and self-adjoint and let \(M\subseteq\mathcal{H}\) be a closed subspace with \(AM\subseteq M\). Then \(M\) is a Hilbert space, and the restriction \(A|_{M}\) is a compact self-adjoint operator on \(M\) with \(\norm{A|_{M}}\leq\norm{A}\). Moreover \(AM^{\perp}\subseteq M^{\perp}\). Rests on Lemma A23.2, Theorem 16.18 and Definition 16.41.

Proof.

Derives Lemma A23.3. A closed subspace of a complete metric space is complete, so \(M\) is a Hilbert space with the inherited inner product. Symmetry of \(A|_{M}\) is inherited, and by Proposition 16.42 a symmetric operator defined on all of the Hilbert space \(M\) is self-adjoint there. Compactness is Lemma A23.2: a bounded sequence in \(M\) is a bounded sequence in \(\mathcal{H}\), so \((Ax_{n})\) has a convergent subsequence, whose limit lies in the closed set \(M\). The norm bound is immediate, the supremum being taken over a smaller set. For the last claim, let \(y\in M^{\perp}\) and \(z\in M\); then \(Az\in M\), so \(\braket{z}{Ay}=\braket{Az}{y}=0\), and \(Ay\in M^{\perp}\).

The supremum is attained: the first eigenvector

This is the step Hilbert Spaces left open. The numerical-radius formula Equation (16.26) exhibits \(\norm{A}\) as a supremum; for a compact operator the supremum is a maximum, and the maximising vector is an eigenvector.

Lemma A23.4 (Attainment).

Let \(A\in\mathcal{B}(\mathcal{H})\) be compact and self-adjoint with \(A\neq0\). Then there are \(e\in\mathcal{H}\) with \(\norm{e}=1\) and \(\lambda\in\R\) with \(\abs{\lambda}=\norm{A}\) such that \(Ae=\lambda e\). Rests on Proposition 16.43, Lemma A23.2 and Proposition 16.42.

Proof.

Derives Lemma A23.4. By Proposition 16.43 there are unit vectors \(x_{n}\) with

\begin{equation}\tag{A23.3} \abs{\braket{x_{n}}{Ax_{n}}}\longrightarrow\norm{A}\ep \end{equation}

Each number \(\braket{x_{n}}{Ax_{n}}\) is real, because \(A\) is self-adjoint (Proposition 16.42). A real sequence whose absolute values converge to \(\norm{A}\) has a subsequence converging to \(+\norm{A}\) or to \(-\norm{A}\): infinitely many of its terms are non-negative or infinitely many are negative, and along the corresponding indices the absolute value carries the sign. Relabel, and let

\begin{equation}\tag{A23.4} \braket{x_{n}}{Ax_{n}}\longrightarrow\lambda\ec\qquad \lambda\in\R\ec\qquad\abs{\lambda}=\norm{A}\neq0\ep \end{equation}

Expand, using \(\lambda\) real, \(\norm{x_{n}}=1\), and \(\braket{Ax_{n}}{x_{n}}=\braket{x_{n}}{Ax_{n}}\):

\begin{equation}\tag{A23.5} \norm{Ax_{n}-\lambda x_{n}}^{2} =\norm{Ax_{n}}^{2}-2\lambda\braket{x_{n}}{Ax_{n}} +\lambda^{2}\ep \end{equation}

Now \(\norm{Ax_{n}}\leq\norm{A}=\abs{\lambda}\), so the right-hand side of Equation (A23.5) is at most \(2\lambda^{2}-2\lambda\braket{x_{n}}{Ax_{n}}\), which tends to \(0\) by Equation (A23.4). Hence

\begin{equation}\tag{A23.6} Ax_{n}-\lambda x_{n}\longrightarrow0\ep \end{equation}

This much uses only self-adjointness; it says that \(\lambda\) is an approximate eigenvalue. Compactness converts the approximation into an eigenvector.

The sequence \((x_{n})\) is bounded, so by Lemma A23.2 there is a subsequence with \(Ax_{n_{k}}\longrightarrow y\) for some \(y\in\mathcal{H}\). Since \(\lambda\neq0\) we may write

\begin{equation*} x_{n_{k}}=\frac{1}{\lambda}\Bigl(Ax_{n_{k}} -\bigl(Ax_{n_{k}}-\lambda x_{n_{k}}\bigr)\Bigr) \longrightarrow\frac{y}{\lambda}=:e\ec \end{equation*}

by Equation (A23.6). The norm is continuous (Corollary 16.5), so \(\norm{e}=\lim\norm{x_{n_{k}}} =1\); and \(A\) is continuous (Proposition 16.36), so \(Ax_{n_{k}}\longrightarrow Ae\). But \(Ax_{n_{k}}\longrightarrow y=\lambda e\). Limits in a metric space are unique, so \(Ae=\lambda e\).

The recursion

Lemma A23.5 (Construction of the system).

Let \(A\) be compact and self-adjoint. There are a finite or countable orthonormal system \(\set{e_{n}}\) and real numbers \(\lambda_{n}\neq0\) with \(Ae_{n}=\lambda_{n}e_{n}\) and \(\abs{\lambda_{1}}\geq\abs{\lambda_{2}}\geq\cdots\), such that, writing \(\mathcal{H}_{1}=\mathcal{H}\) and

\begin{equation}\tag{A23.7} \mathcal{H}_{N+1}=\set{e_{1},\dots,e_{N}}^{\perp}\ec\qquad A_{N+1}=A|_{\mathcal{H}_{N+1}}\ec \end{equation}

one has \(\norm{A_{N+1}}=\abs{\lambda_{N+1}}\) at every stage at which the construction continues, and the construction stops at stage \(N\) exactly when \(A_{N}=0\), i.e. when \(\mathcal{H}_{N}\subseteq\ker A\). Rests on Lemma A23.4, Lemma A23.3 and Theorem 16.18.

Proof.

Derives Lemma A23.5. Set \(\mathcal{H}_{1}=\mathcal{H}\) and \(A_{1}=A\). Suppose \(\mathcal{H}_{N}\) and \(A_{N}\) have been defined, with \(\mathcal{H}_{N}\) a closed \(A\)-invariant subspace and \(A_{N}=A|_{\mathcal{H}_{N}}\) compact and self-adjoint on it (Lemma A23.3). If \(A_{N}=0\) the construction stops, and then every \(x\in\mathcal{H}_{N}\) has \(Ax=0\). If \(A_{N}\neq0\), apply Lemma A23.4 inside the Hilbert space \(\mathcal{H}_{N}\): it produces a unit vector \(e_{N}\in\mathcal{H}_{N}\) and a real \(\lambda_{N}\) with \(\abs{\lambda_{N}}=\norm{A_{N}}\neq0\) and \(Ae_{N}=A_{N}e_{N}=\lambda_{N}e_{N}\).

The vector \(e_{N}\) lies in \(\mathcal{H}_{N}\), hence is orthogonal to \(e_{1},\dots,e_{N-1}\), so the system stays orthonormal. Put \(\mathcal{H}_{N+1}=\mathcal{H}_{N}\cap\set{e_{N}}^{\perp} =\set{e_{1},\dots,e_{N}}^{\perp}\), which is closed (Proposition 16.17). It is \(A\)-invariant: for \(x\in\mathcal{H}_{N+1}\) and \(j\leq N\),

\begin{equation}\tag{A23.8} \braket{e_{j}}{Ax}=\braket{Ae_{j}}{x} =\lambda_{j}^{\ast}\braket{e_{j}}{x} =\lambda_{j}\braket{e_{j}}{x}=0\ec \end{equation}

using self-adjointness and \(\lambda_{j}\in\R\); and \(Ax\in\mathcal{H}_{N}\) because \(\mathcal{H}_{N}\) is invariant. Finally \(\norm{A_{N+1}}\leq\norm{A_{N}}\) because \(A_{N+1}\) is a restriction of \(A_{N}\), whence \(\abs{\lambda_{N+1}}\leq\abs{\lambda_{N}}\); the construction therefore delivers a non-increasing sequence of moduli, and it is exactly the statement \(\norm{A_{N}}=\abs{\lambda_{N}}\) that Lemma A23.4 supplies at each stage.

Lemma A23.6 (The eigenvalues tend to zero, with finite multiplicity).

If the construction of Lemma A23.5 does not stop, then \(\lambda_{n}\longrightarrow0\). Moreover, for every \(\lambda\neq0\) the eigenspace \(\ker(A-\lambda\identity)\) is finite dimensional. Rests on Lemmas A23.2 and A23.5.

Proof.

Derives Lemma A23.6. Both statements follow from one computation. Let \((f_{n})\) be an orthonormal sequence with \(Af_{n}=\nu_{n}f_{n}\). For \(n\neq m\), orthogonality gives

\begin{equation}\tag{A23.9} \norm{Af_{n}-Af_{m}}^{2} =\norm{\nu_{n}f_{n}-\nu_{m}f_{m}}^{2} =\abs{\nu_{n}}^{2}+\abs{\nu_{m}}^{2}\ep \end{equation}

Eigenvalues. Apply this to \(f_{n}=e_{n}\), \(\nu_{n}=\lambda_{n}\). If \((\lambda_{n})\) does not tend to \(0\) there are \(\varepsilon>0\) and infinitely many indices with \(\abs{\lambda_{n}}\geq\varepsilon\); along them Equation (A23.9) gives \(\norm{Ae_{n}-Ae_{m}}^{2}\geq2\varepsilon^{2}\), so \((Ae_{n})\) has no Cauchy subsequence, hence no convergent subsequence. Since \((e_{n})\) is bounded, this contradicts Lemma A23.2. As the moduli are non-increasing, \(\lambda_{n}\longrightarrow0\).

Multiplicity. If \(\ker(A-\lambda\identity)\) were infinite dimensional for some \(\lambda\neq0\), Gram–Schmidt (Proposition 16.23) would produce an infinite orthonormal sequence \((f_{n})\) inside it, all with \(\nu_{n}=\lambda\), and Equation (A23.9) would give \(\norm{Af_{n}-Af_{m}}^{2}=2\abs{\lambda}^{2}>0\): the same contradiction.

Passage to the limit

Proof of Theorem A23.1. Derives Theorem A23.1. Let \(\set{e_{n}}\), \(\set{\lambda_{n}}\) be the system of Lemma A23.5. Fix \(x\in\mathcal{H}\) and put

\begin{equation}\tag{A23.10} x_{N}=x-\sum_{n=1}^{N}\braket{e_{n}}{x}\,e_{n}\ep \end{equation}

Then \(\braket{e_{j}}{x_{N}}=\braket{e_{j}}{x}-\braket{e_{j}}{x}=0\) for \(j\leq N\), so \(x_{N}\in\mathcal{H}_{N+1}\), and Bessel's inequality (Proposition 16.27) gives

\begin{equation}\tag{A23.11} \norm{x_{N}}^{2}=\norm{x}^{2} -\sum_{n=1}^{N}\abs{\braket{e_{n}}{x}}^{2}\leq\norm{x}^{2}\ep \end{equation}

Applying \(A\) to Equation (A23.10) and using \(Ae_{n}=\lambda_{n}e_{n}\),

\begin{equation}\tag{A23.12} Ax-\sum_{n=1}^{N}\lambda_{n}\braket{e_{n}}{x}\,e_{n} =Ax_{N}=A_{N+1}x_{N}\ec \end{equation}

the last equality because \(x_{N}\in\mathcal{H}_{N+1}\). Therefore, by Lemma A23.5 and Equation (A23.11),

\begin{equation}\tag{A23.13} \norm{Ax-\sum_{n=1}^{N}\lambda_{n}\braket{e_{n}}{x}\,e_{n}} \leq\norm{A_{N+1}}\,\norm{x_{N}} \leq\abs{\lambda_{N+1}}\,\norm{x}\ep \end{equation}

If the construction stops at stage \(N\) then \(A_{N}=0\) and Equation (A23.12) already gives Equation (A23.1) as a finite sum. If it does not stop, \(\abs{\lambda_{N+1}}\longrightarrow0\) by Lemma A23.6, so the right-hand side of Equation (A23.13) tends to \(0\) and Equation (A23.1) holds with convergence in norm. The eigenvalue relation \(Ae_{n}=\lambda_{n}e_{n}\), the ordering of the moduli and the finiteness of the non-zero eigenspaces are Lemmas A23.5 and A23.6. For \(A\neq0\), \(\norm{A}=\norm{A_{1}}=\abs{\lambda_{1}} =\max_{n}\abs{\lambda_{n}}\), the maximum being attained because the moduli decrease.

It remains to prove Equation (A23.2). Write \(M=\overline{\text{span}\set{e_{n}}}\), a closed subspace, so that \(\mathcal{H}=M\oplus M^{\perp}\) by Theorem 16.18. If \(x\in M^{\perp}\) then \(\braket{e_{n}}{x}=0\) for every \(n\), and Equation (A23.1) gives \(Ax=0\): thus \(M^{\perp}\subseteq\ker A\). Conversely if \(Ax=0\) then for every \(n\),

\begin{equation*} 0=\braket{e_{n}}{Ax}=\braket{Ae_{n}}{x} =\lambda_{n}\braket{e_{n}}{x}\ec \end{equation*}

and \(\lambda_{n}\neq0\) forces \(\braket{e_{n}}{x}=0\), so \(x\in M^{\perp}\); hence \(\ker A=M^{\perp}\), which is Equation (A23.2). Every vector of \(\ker A\) is an eigenvector with eigenvalue \(0\), so an orthonormal basis of \(\ker A\) adjoined to \(\set{e_{n}}\) is an orthonormal system of eigenvectors whose closed span is \(M\oplus\ker A=\mathcal{H}\): an orthonormal basis of \(\mathcal{H}\) in the sense of Definition 16.29.

Remark A23.7 (Where the orthonormal basis of the kernel comes from).

The last sentence of Theorem A23.1 needs an orthonormal basis of the closed subspace \(\ker A\), and that is the only point at which the theorem is not constructive. If \(\mathcal{H}\) is separable — the case of every physical application, and the standing assumption of Theorem 16.33 — so is \(\ker A\), and Gram–Schmidt applied to a countable dense set produces the basis (Corollary 16.24). For a non-separable \(\mathcal{H}\) the existence of an orthonormal basis of \(\ker A\) is a maximality statement proved with Zorn's lemma (Logic, Sets, and Maps); the system \(\set{e_{n}}\) itself, on which the whole content of the theorem rests, is always countable and always constructed, never chosen.

Remark A23.8 (What the weak topology would have bought, and why it is not needed).

The route usually taken to Lemma A23.4 [Reed:1972] passes through two statements about the weak topology, neither of which this treatise develops: that the closed unit ball of a Hilbert space is weakly sequentially compact, so that a maximising sequence has a weakly convergent subsequence \(x_{n}\rightharpoonup e\); and that a compact operator carries weakly convergent sequences to norm convergent ones, so that \(Ax_{n}\longrightarrow Ae\) and the supremum is attained at \(e\). Both are true and both are proved in [Reed:1972], chapter VI.

The proof given above needs neither. The reason is Equation (A23.6): self-adjointness alone forces the maximising sequence to satisfy \(Ax_{n}-\lambda x_{n}\longrightarrow0\) in norm, and once that is known the sequential form of compactness, Lemma A23.2, is enough to extract a genuine eigenvector — the vector \(e\) is recovered from \(Ax_{n_{k}}\) by dividing by \(\lambda\neq0\), not by any weak limit. The argument therefore quotes nothing beyond Proposition 16.43 and Theorem 10.31, and Hilbert Spaces may drop the reservation with which Theorem 16.44 was first stated.

Remark A23.9.

The Hilbert–Schmidt Theorem for Compact Self-Adjoint Operators discharges the proof obligation of Theorem 16.44 (Section 16.3.1). The theorem is what makes the integral operators Equation (16.28) of Partial Differential Equations — the Green functions of a boundary-value problem — carry a discrete spectrum of modes with a complete set of eigenfunctions, and it is the exact infinite-dimensional counterpart of Equation (9.105). It is also the last point at which the counterpart is exact: Example 16.56 exhibits a bounded self-adjoint operator, not compact, with no eigenvector at all, and the repair for that case is the projection-valued measure of The Spectral Theorem for a Bounded Self-Adjoint Operator.