Completeness of the Sturm–Liouville Eigenfunctions in the Energy Norm

Contents
  1. Statement
  2. The energy space
  3. The Green function and the inverse operator
  4. Compactness and the spectral decomposition
  5. From the weighted norm to the energy norm

This appendix proves the statement quoted in Remark 20.82 of Calculus of Variations: the eigenfunctions of a regular Sturm–Liouville problem are complete not merely in the weighted \(L^{2}\) space, but in the energy norm built from the quadratic form of the problem, and with that completeness the inequality Equation (20.74) sharpens into the identity Equation (20.75) for the energy form. The identity is what the Rayleigh–Ritz error estimate Corollary 20.83 needs; the inequality, which is what the variational characterisation Theorem 20.81 needs, is proved in the chapter from Bessel's inequality alone and is independent of everything below — a point Remark A42.13 returns to, since a circular appendix would silently destroy the chapter's argument.

The route is the one the subject was born from [Hilbert:1912]: invert the differential operator. The inverse of a regular Sturm–Liouville operator with Dirichlet data is an integral operator whose kernel is the Green function of the problem; that kernel is continuous, so the operator is compact and self-adjoint, and the Hilbert–Schmidt theorem Theorem 16.44, proved in The Hilbert–Schmidt Theorem for Compact Self-Adjoint Operators, hands back an orthonormal basis of eigenfunctions. Completeness in the weighted \(L^{2}\) space — which Theorem 20.81 assumes — then comes out as a theorem, and completeness in the energy norm follows from it by one identity relating the two inner products.

Throughout, the data are those of Definition 20.80: \(p\), \(q\), \(r\) are continuous and real on \([a,b]\) with \(p>0\) and \(r>0\), and the boundary conditions are \(y(a)=y(b)=0\). As in the proof of Theorem 20.81, fix once and for all a constant \(K\) with

\begin{equation}\tag{A42.1} q+Kr\ge0\qquad\text{on }[a,b]\ec \end{equation}

which is possible because \(q\) is continuous and \(r>0\) on a compact interval, and write

\begin{equation}\tag{A42.2} \avg{u,v}_{r}=\int_{a}^{b}r\,u\,v\,\dd x\ec\qquad B_{K}[u,v]=\int_{a}^{b}\left[p\,u'v' +\left(q+Kr\right)u\,v\right]\dd x\ep \end{equation}

The operator of the problem, shifted by \(K\), is

\begin{equation}\tag{A42.3} L_{K}[u]=-\dv{}{x}\left(p\,\dv{u}{x}\right) +\left(q+Kr\right)u\ec \end{equation}

so that \(y\) solves Equation (20.70) with eigenvalue \(\lambda\) if and only if \(L_{K}[y]=\left(\lambda+K\right)r\,y\).

Statement

Theorem A42.1 (Completeness in the weighted and in the energy norm).

Let \(p,q,r\in C^{0}[a,b]\) with \(p>0\) and \(r>0\). Then:

  1. the Dirichlet problem Equation (20.70) has an infinite sequence of real eigenvalues \(\lambda_{1}\le\lambda_{2}\le\cdots\) with \(\lambda_{n}\longrightarrow+\infty\), each of finite multiplicity, and eigenfunctions \(y_{n}\) that are orthonormal in \(\avg{\cdot,\cdot}_{r}\) and form an orthonormal basis of the weighted space \(L^{2}_{r}(a,b)\);

  2. the rescaled eigenfunctions \(y_{n}/\sqrt{\lambda_{n}+K}\) form an orthonormal basis of the energy space \(H_{E}\) of Definition A42.2 with the inner product \(B_{K}\);

  3. consequently, for every \(y\in H_{E}\) — in particular for every \(y\in C^{1}_{0}[a,b]\) — the partial sums \(S_{N}=\sum_{n\le N}c_{n}y_{n}\), \(c_{n}=\avg{y,y_{n}}_{r}\), satisfy \(B_{K}\left[y-S_{N},y-S_{N}\right]\longrightarrow0\), and

    \begin{equation}\tag{A42.4} \int_{a}^{b}\left(p\,y'^{2}+q\,y^{2}\right)\dd x =\sum_{n\ge1}\lambda_{n}\,c_{n}^{2}\ec \end{equation}

    the series on the right converging.

Rests on Definition 20.80, Theorem 16.44 and Equation (20.70).

Equation (A42.4) is Equation (20.75). The proof occupies the rest of the section: the energy space (The energy space), the Green function and the inverse operator (The Green function and the inverse operator), its compactness and the spectral decomposition (Compactness and the spectral decomposition), and the transfer to the energy norm (From the weighted norm to the energy norm).

The energy space

Definition A42.2 (The weighted space and the energy space).

\(L^{2}_{r}(a,b)\) is the space of square-integrable functions with the inner product \(\avg{\cdot,\cdot}_{r}\) of Equation (A42.2), and

\begin{equation}\tag{A42.5} H_{E}=\set{u\in W^{1,2}(a,b)\ \mid\ u(a)=u(b)=0} \end{equation}

is the energy space, where \(W^{1,2}(a,b)\) is the space of Definition A41.3 with \(r=2\): the continuous functions \(u\) for which there is \(u'\in L^{2}(a,b)\) with \(u(x)=u(a)+\int_{a}^{x}u'\). On \(H_{E}\) the energy form \(B_{K}\) of Equation (A42.2) is used as an inner product, with norm \(\norm{u}_{E}=B_{K}[u,u]^{1/2}\). Rests on Definitions 20.80 and A41.3.

Lemma A42.3 (Poincaré inequality; $B_{K}$ is an inner product).

Write \(p_{-}=\min_{[a,b]}p>0\), \(p_{+}=\max_{[a,b]}p\), \(r_{-}=\min_{[a,b]}r>0\), \(r_{+}=\max_{[a,b]}r\) and \(\kappa_{+}=\max_{[a,b]}\left(q+Kr\right)\). For every \(u\in H_{E}\),

\begin{equation}\tag{A42.6} \max_{[a,b]}\abs{u} \le\left(b-a\right)^{1/2}\norm{u'}_{L^{2}}\ec\qquad \avg{u,u}_{r}\le r_{+}\left(b-a\right)^{2}\norm{u'}_{L^{2}}^{2}\ec \end{equation}

and consequently

\begin{equation}\tag{A42.7} p_{-}\norm{u'}_{L^{2}}^{2} \le B_{K}[u,u] \le\left[p_{+}+\kappa_{+}\left(b-a\right)^{2}\right] \norm{u'}_{L^{2}}^{2}\ep \end{equation}

In particular \(B_{K}\) is an inner product on \(H_{E}\): it is bilinear and symmetric by inspection, and \(B_{K}[u,u]=0\) forces \(u=0\). Rests on Definition A42.2 and Lemma A41.5.

Proof.

Derives Lemma A42.3. For \(u\in H_{E}\) and \(x\in[a,b]\), \(u(a)=0\) gives \(u(x)=\int_{a}^{x}u'\), and the Cauchy–Schwarz case \(r=2\) of Hölder's inequality Equation (A41.4) bounds this by \(\left(x-a\right)^{1/2}\norm{u'}_{L^{2}} \le\left(b-a\right)^{1/2}\norm{u'}_{L^{2}}\), which is the first inequality of Equation (A42.6). Integrating its square against \(r\) over \([a,b]\) gives the second.

For Equation (A42.7), the lower bound drops the term \(\left(q+Kr\right)u^{2}\ge0\), which is nonnegative by Equation (A42.1), and bounds \(p\ge p_{-}\); the upper bound uses \(p\le p_{+}\), \(q+Kr\le\kappa_{+}\) and the second inequality of Equation (A42.6) with \(r\) replaced by the constant \(1\), i.e.\ \(\int u^{2}\le\left(b-a\right)^{2}\norm{u'}^{2}_{L^{2}}\). If \(B_{K}[u,u]=0\) then \(\norm{u'}_{L^{2}}=0\) by the lower bound, so \(u'=0\) almost everywhere and \(u(x)=u(a)+\int_{a}^{x}u'=0\).

Lemma A42.4 ($H_{E}$ is a Hilbert space).

\(\left(H_{E},B_{K}\right)\) is complete, hence a Hilbert space in the sense of Definition 16.2, and the inclusion \(H_{E}\subset L^{2}_{r}(a,b)\) is continuous. Rests on Lemma A42.3 and Theorem 16.12.

Proof.

Derives Lemma A42.4. Continuity of the inclusion is the second inequality of Equation (A42.6) combined with the lower bound in Equation (A42.7): \(\avg{u,u}_{r}\le r_{+}\left(b-a\right)^{2}p_{-}^{-1}B_{K}[u,u]\).

Let \(\left(u_{k}\right)\) be Cauchy in \(\norm{\cdot}_{E}\). By the lower bound in Equation (A42.7) the sequence \(\left(u_{k}'\right)\) is Cauchy in \(L^{2}(a,b)\), which is complete (Theorem 16.12), so \(u_{k}'\longrightarrow v\) in \(L^{2}\). By the first inequality of Equation (A42.6) applied to \(u_{k}-u_{l}\), the sequence \(\left(u_{k}\right)\) is uniformly Cauchy on \([a,b]\), hence converges uniformly to a continuous \(u\) with \(u(a)=u(b)=0\). Passing to the limit in \(u_{k}(x)=\int_{a}^{x}u_{k}'\) — legitimate because \(\abs{\int_{a}^{x}\left(u_{k}'-v\right)} \le\left(b-a\right)^{1/2}\norm{u_{k}'-v}_{L^{2}} \longrightarrow0\) by Equation (A41.4) — gives \(u(x)=\int_{a}^{x}v\), so \(u\in H_{E}\) with \(u'=v\). Finally the upper bound in Equation (A42.7) applied to \(u_{k}-u\) gives \(\norm{u_{k}-u}_{E}\longrightarrow0\).

Remark A42.5 (This is the space the chapter's trial functions live in).

\(C^{1}_{0}[a,b]\subset H_{E}\), since a continuously differentiable function vanishing at both ends satisfies Equation (A41.2) with \(u'\) its classical derivative. So every statement below applies to the trial functions of Theorem 20.81 and Proposition 20.84 without further comment. The space \(H_{E}\) is the \(H^{1}_{0}(a,b)\) of Definition 14.85: the identification of the two descriptions is Remark A41.4 together with the fact that a \(W^{1,2}\) function vanishing at both endpoints is a uniform limit of smooth functions of compact support, which is not needed here and is not proved.

The Green function and the inverse operator

Lemma A42.6 (The homogeneous Dirichlet problem is trivial).

Let \(\varphi\) solve \(L_{K}[\varphi]=0\) on \([a,b]\) with \(\varphi(a)=\varphi(b)=0\). Then \(\varphi\equiv0\). Rests on Lemma A42.3 and Equation (A42.3).

Proof.

Derives Lemma A42.6. Multiply \(L_{K}[\varphi]=0\) by \(\varphi\) and integrate over \([a,b]\); one integration by parts (Equation (11.28)), whose boundary term \(\left[p\,\varphi\,\varphi'\right]_{a}^{b}\) vanishes because \(\varphi\) does at both ends, gives \(B_{K}[\varphi,\varphi]=0\). Since \(\varphi\in C^{1}[a,b]\) vanishes at \(a\), it lies in \(H_{E}\), and Lemma A42.3 forces \(\varphi=0\).

Definition A42.7 (Green function of the shifted problem).

Let \(\phi\) and \(\psi\) be the solutions of \(L_{K}[u]=0\) determined by

\begin{equation}\tag{A42.8} \phi(a)=0\ec\quad\left(p\,\phi'\right)(a)=1\ec\qquad \psi(b)=0\ec\quad\left(p\,\psi'\right)(b)=1\ec \end{equation}

which exist, are unique and are defined on the whole of \([a,b]\) by Lemma A40.4 applied to the system form of \(L_{K}[u]=0\) — the equation \(\left(p\,u'\right)'=\left(q+Kr\right)u\) is Equation (A40.1) with \(P=p\) and \(Q=q+Kr\). Put

\begin{equation}\tag{A42.9} C=p\left(\phi\,\psi'-\psi\,\phi'\right)\ec \end{equation}

a nonzero constant, and define the Green function

\begin{equation}\tag{A42.10} G\left(x,s\right)=-\frac{1}{C}\times \begin{cases} \phi(x)\,\psi(s) & a\le x\le s\le b\ec\\ \phi(s)\,\psi(x) & a\le s\le x\le b\ep \end{cases} \end{equation}

Rests on Lemmas A40.7 and A42.6.

Lemma A42.8 (The kernel is well defined, symmetric and Lipschitz).

The constant \(C\) of Equation (A42.9) is independent of \(x\) and nonzero; \(G\) is well defined by Equation (A42.10) (the two branches agree on \(x=s\)), symmetric, \(G\left(x,s\right)=G\left(s,x\right)\), continuous on \([a,b]\times[a,b]\), and there is a constant \(L_{G}\) with

\begin{equation}\tag{A42.11} \abs{G\left(x,s\right)-G\left(x',s\right)} \le L_{G}\,\abs{x-x'} \qquad\text{for all }x,x',s\in[a,b]\ep \end{equation}

Rests on Definition A42.7 and Theorem 11.35.

Proof.

Derives Lemma A42.8. That \(C\) is constant is Lemma A40.7 applied with \(P=p\) and \(Q=q+Kr\). If \(C=0\) the same lemma makes \(\phi\) and \(\psi\) linearly dependent, so \(\phi\) would vanish at \(b\) as \(\psi\) does, and \(\phi\) would be a nontrivial solution of the homogeneous Dirichlet problem — nontrivial because \(\left(p\phi'\right)(a)=1\) — contradicting Lemma A42.6. Hence \(C\neq0\).

The two branches of Equation (A42.10) coincide when \(x=s\), so \(G\) is well defined, and exchanging \(x\) and \(s\) exchanges the two branches, which is the symmetry. Both \(\phi\) and \(\psi\) are of class \(C^{1}\) on the compact interval, hence Lipschitz there by the mean value theorem (Theorem 11.35) with constants \(\max\abs{\phi'}\) and \(\max\abs{\psi'}\); each branch of Equation (A42.10) is therefore Lipschitz in \(x\) uniformly in \(s\), with \(L_{G}=\abs{C}^{-1}\max\set{\max\abs{\phi'}\max\abs{\psi}, \max\abs{\phi}\max\abs{\psi'}}\), and since the branches agree where they meet, the bound Equation (A42.11) holds for \(x,x'\) on opposite sides of \(s\) as well, by splitting the increment at \(s\). Continuity follows, jointly in \(\left(x,s\right)\), from Equation (A42.11) and the symmetry.

Theorem A42.9 (The Green operator inverts $L_{K}$).

Define

\begin{equation}\tag{A42.12} \left(Tf\right)(x)=\int_{a}^{b}G\left(x,s\right)f(s)\,r(s)\,\dd s\ec \qquad f\in L^{2}_{r}(a,b)\ep \end{equation}

Then:

  1. for continuous \(f\), \(u=Tf\) is of class \(C^{1}\) with \(p\,u'\) of class \(C^{1}\), and it is the unique solution of

    \begin{equation}\tag{A42.13} L_{K}[u]=r\,f\ec\qquad u(a)=u(b)=0\ep \end{equation}
  2. \(T\) maps \(L^{2}_{r}(a,b)\) into \(H_{E}\), with \(\norm{Tf}_{E}\le C_{P}\norm{f}_{r}\) where \(C_{P}=\left(r_{+}\left(b-a\right)^{2}/p_{-}\right)^{1/2}\), and

    \begin{equation}\tag{A42.14} B_{K}\left[Tf,\chi\right]=\avg{f,\chi}_{r} \qquad\text{for every }\chi\in H_{E}\ep \end{equation}
  3. \(T\) is self-adjoint and positive on \(L^{2}_{r}(a,b)\), and injective.

Rests on Definition A42.7, Lemma A42.3 and Lemma A42.6.

Proof.

Derives Theorem A42.9. (1) Split Equation (A42.12) at \(x\) using Equation (A42.10):

\begin{equation}\tag{A42.15} -C\,u(x)=\psi(x)\int_{a}^{x}\phi\,f\,r\,\dd s +\phi(x)\int_{x}^{b}\psi\,f\,r\,\dd s\ec \end{equation}

the first branch of Equation (A42.10) governing \(s\ge x\) and the second \(s\le x\). Both integrals have continuous integrands, so by the fundamental theorem of calculus (Theorem 11.42) the right side is differentiable and

\begin{equation*} -C\,u'(x)=\psi'(x)\int_{a}^{x}\phi\,f\,r\,\dd s +\phi'(x)\int_{x}^{b}\psi\,f\,r\,\dd s +\left[\psi(x)\phi(x)-\phi(x)\psi(x)\right]f(x)r(x)\ec \end{equation*}

the last bracket vanishing identically: the two boundary terms produced by differentiating the limits of integration cancel. Hence

\begin{equation*} -C\,p(x)u'(x) =\left(p\psi'\right)(x)\int_{a}^{x}\phi\,f\,r\,\dd s +\left(p\phi'\right)(x)\int_{x}^{b}\psi\,f\,r\,\dd s\ec \end{equation*}

which is again differentiable because \(p\phi'\) and \(p\psi'\) are of class \(C^{1}\), with \(\left(p\phi'\right)'=\left(q+Kr\right)\phi\) and likewise for \(\psi\). Differentiating,

\begin{align*} -C\left(p\,u'\right)'(x) &=\left(q+Kr\right)\psi(x)\int_{a}^{x}\phi\,f\,r\,\dd s +\left(q+Kr\right)\phi(x)\int_{x}^{b}\psi\,f\,r\,\dd s\\ &\qquad+\left[\left(p\psi'\right)(x)\phi(x) -\left(p\phi'\right)(x)\psi(x)\right]f(x)\,r(x)\\ &=-C\left(q+Kr\right)u(x)+C\,f(x)\,r(x)\ec \end{align*}

using Equation (A42.15) for the first two terms and Equation (A42.9) for the bracket, which is \(\phi\left(p\psi'\right)-\psi\left(p\phi'\right) =p\left(\phi\psi'-\psi\phi'\right)=C\). Dividing by \(-C\) gives \(\left(p\,u'\right)'=\left(q+Kr\right)u-r\,f\), that is Equation (A42.13). The boundary conditions hold because \(x=a\) kills the first integral in Equation (A42.15) and leaves \(\phi(a)=0\) in the second, and symmetrically at \(x=b\) with \(\psi(b)=0\). Uniqueness is Lemma A42.6 applied to the difference of two solutions.

(2) Let first \(f\) be continuous and \(u=Tf\). Then \(u\in C^{1}[a,b]\) with \(u(a)=u(b)=0\), so \(u\in H_{E}\), and for \(\chi\in C^{1}_{0}[a,b]\) one integration by parts (Equation (11.28)) applied to Equation (A42.13) gives \(B_{K}[u,\chi]=\avg{f,\chi}_{r}\), the boundary term \(\left[p\,u'\chi\right]_{a}^{b}\) vanishing because \(\chi\) does. The same identity for arbitrary \(\chi\in H_{E}\) follows because both sides are continuous in \(\chi\) for \(\norm{\cdot}_{E}\) — the left side by the Cauchy–Schwarz inequality for the inner product \(B_{K}\), the right by Lemma A42.4 — and, by the argument of Lemma A42.4 run backwards, each \(\chi\in H_{E}\) is an \(\norm{\cdot}_{E}\)-limit of members of \(C^{1}_{0}[a,b]\): take \(\chi_{k}(x)=\int_{a}^{x}v_{k}\) with \(v_{k}\) continuous, \(\int_{a}^{b}v_{k}=0\) and \(v_{k}\longrightarrow\chi'\) in \(L^{2}\), which is possible because the continuous functions are dense in \(L^{2}\) (Remark A42.14) and the mean may be subtracted off with a loss tending to zero.

Taking \(\chi=u\) in Equation (A42.14) and using Equation (A42.6),

\begin{equation}\tag{A42.16} \norm{u}_{E}^{2}=\avg{f,u}_{r} \le\norm{f}_{r}\norm{u}_{r} \le C_{P}\norm{f}_{r}\norm{u}_{E}\ec \end{equation}

so \(\norm{Tf}_{E}\le C_{P}\norm{f}_{r}\) for continuous \(f\). For general \(f\in L^{2}_{r}\) take continuous \(f_{k}\longrightarrow f\) in \(L^{2}_{r}\); by Equation (A42.16) the sequence \(\left(Tf_{k}\right)\) is Cauchy in \(H_{E}\), hence convergent there (Lemma A42.4), while by Equation (A42.12) and the Cauchy–Schwarz inequality \(\abs{Tf_{k}(x)-Tf(x)}\le\max\abs{G} \left(r_{+}\left(b-a\right)\right)^{1/2}\norm{f_{k}-f}_{r}\) tends to zero uniformly. The two limits agree, so \(Tf\in H_{E}\), the bound persists, and Equation (A42.14) passes to the limit.

(3) For \(f,g\in L^{2}_{r}\) put \(u=Tf\), \(w=Tg\). Then Equation (A42.14) twice, with the symmetry of \(B_{K}\) in between, gives

\begin{equation*} \avg{Tf,g}_{r}=\avg{g,u}_{r}=B_{K}[w,u]=B_{K}[u,w] =\avg{f,w}_{r}=\avg{f,Tg}_{r}\ec \end{equation*}

so \(T\) is self-adjoint; and \(\avg{Tf,f}_{r}=B_{K}[u,u]\ge0\), so \(T\) is positive. If \(Tf=0\) then \(B_{K}[0,\chi]=\avg{f,\chi}_{r}=0\) for every \(\chi\in H_{E}\), and since \(H_{E}\) is dense in \(L^{2}_{r}\) — it contains the functions \(\chi_{k}\) built above from an arbitrary continuous \(v\), and the continuous functions are dense — this forces \(f=0\). Hence \(T\) is injective.

Compactness and the spectral decomposition

Theorem A42.10 ($T$ is compact).

\(T\) maps every bounded sequence of \(L^{2}_{r}(a,b)\) to a sequence with a convergent subsequence in \(L^{2}_{r}(a,b)\); that is, \(T\) is compact in the sense of Definition 16.41. Rests on Lemmas A41.7 and A42.8.

Proof.

Derives Theorem A42.10. Let \(\norm{f_{n}}_{r}\le M\). By the Cauchy–Schwarz inequality in \(L^{2}_{r}\) applied to Equation (A42.12),

\begin{equation*} \abs{\left(Tf_{n}\right)(x)} \le\left(\int_{a}^{b}G\left(x,s\right)^{2}r(s)\,\dd s\right)^{1/2} \norm{f_{n}}_{r} \le M\max\abs{G}\left(r_{+}\left(b-a\right)\right)^{1/2}\ec \end{equation*}

so the sequence \(\left(Tf_{n}\right)\) is uniformly bounded, and by the Lipschitz estimate Equation (A42.11),

\begin{equation*} \abs{\left(Tf_{n}\right)(x)-\left(Tf_{n}\right)\left(x'\right)} \le\int_{a}^{b}\abs{G\left(x,s\right)-G\left(x',s\right)} \abs{f_{n}(s)}r(s)\,\dd s \le L_{G}\left(r_{+}\left(b-a\right)\right)^{1/2}M\,\abs{x-x'}\ec \end{equation*}

so it is equicontinuous, with a modulus independent of \(n\). By the Arzelà–Ascoli theorem Lemma A41.7 a subsequence converges uniformly on \([a,b]\), and uniform convergence implies convergence in \(L^{2}_{r}\), since \(\norm{w}_{r}^{2}\le r_{+}\left(b-a\right)\max\abs{w}^{2}\).

Theorem A42.11 (Spectral decomposition and completeness in $L^{2}_{r}$).

There is an orthonormal basis \(\set{y_{n}}_{n\ge1}\) of \(L^{2}_{r}(a,b)\) and numbers \(\mu_{n}>0\) with \(\mu_{n}\longrightarrow0\) such that \(Ty_{n}=\mu_{n}y_{n}\). Setting \(\lambda_{n}=\mu_{n}^{-1}-K\), each \(y_{n}\) is a classical solution of the Sturm–Liouville problem Equation (20.70) with eigenvalue \(\lambda_{n}\), and \(\lambda_{n}\longrightarrow+\infty\). Every eigenvalue has finite multiplicity, and the \(\lambda_{n}\) may be enumerated in nondecreasing order. Rests on Theorems 16.44, A42.9 and A42.10.

Proof.

Derives Theorem A42.11. \(L^{2}_{r}(a,b)\) is a Hilbert space: its inner product differs from that of \(L^{2}(a,b)\) by the factor \(r\), which is bounded between the positive constants \(r_{-}\) and \(r_{+}\), so the two norms are equivalent and completeness transfers from Theorem 16.12. On it, \(T\) is bounded (by Equation (A42.16) and Equation (A42.6)), self-adjoint and compact, by Theorems A42.9 and A42.10. The Hilbert–Schmidt theorem Theorem 16.44, proved in The Hilbert–Schmidt Theorem for Compact Self-Adjoint Operators, therefore supplies real numbers \(\mu_{n}\longrightarrow0\) and an orthonormal system \(\set{y_{n}}\) with

\begin{equation}\tag{A42.17} Tf=\sum_{n}\mu_{n}\avg{y_{n},f}_{r}\,y_{n} \qquad\text{for every }f\in L^{2}_{r}(a,b)\ep \end{equation}

The system is a basis. If \(f\) is orthogonal to every \(y_{n}\) then Equation (A42.17) gives \(Tf=0\), and \(T\) is injective by Theorem A42.9(3), so \(f=0\). By Definition 16.29 the system is maximal, i.e. an orthonormal basis, and Theorem 16.30 applies to it.

The eigenvalues are positive. Taking \(f=y_{n}\) in Equation (A42.17) gives \(Ty_{n}=\mu_{n}y_{n}\), and \(\mu_{n}=\avg{Ty_{n},y_{n}}_{r}\ge0\) by positivity; \(\mu_{n}=0\) is excluded by injectivity. Hence \(\mu_{n}>0\) and \(\lambda_{n}=\mu_{n}^{-1}-K\) is well defined, with \(\lambda_{n}\longrightarrow+\infty\) because \(\mu_{n}\longrightarrow0^{+}\).

The eigenfunctions solve the differential equation. From \(y_{n}=\mu_{n}^{-1}Ty_{n}\) and Theorem A42.9(2), \(y_{n}\in H_{E}\); in particular \(y_{n}\) is continuous, so Theorem A42.9(1) applies to \(f=y_{n}\) and makes \(u=Ty_{n}=\mu_{n}y_{n}\) a classical solution of \(L_{K}[u]=r\,y_{n}\) with \(u(a)=u(b)=0\). Dividing by \(\mu_{n}\),

\begin{equation*} L_{K}\left[y_{n}\right]=\mu_{n}^{-1}r\,y_{n} =\left(\lambda_{n}+K\right)r\,y_{n}\ec \end{equation*}

which, on subtracting \(K\,r\,y_{n}\) from both sides, is Equation (20.70) with eigenvalue \(\lambda_{n}\) and Dirichlet data.

Finite multiplicity and ordering. An eigenvalue \(\lambda\) of Equation (20.70) corresponds to the eigenvalue \(\mu=\left(\lambda+K\right)^{-1}\) of \(T\), whose eigenspace is finite dimensional by Theorem 16.44. Since \(\lambda_{n}\longrightarrow+\infty\), only finitely many \(\lambda_{n}\) lie below any given bound, so the sequence may be reordered nondecreasingly, and the \(y_{n}\) with it.

From the weighted norm to the energy norm

Lemma A42.12 (The pairing identity).

For every \(\chi\in H_{E}\) and every \(n\),

\begin{equation}\tag{A42.18} B_{K}\left[y_{n},\chi\right] =\left(\lambda_{n}+K\right)\avg{y_{n},\chi}_{r}\ep \end{equation}

In particular \(B_{K}\left[y_{m},y_{n}\right] =\left(\lambda_{n}+K\right)\delta_{mn}\), and \(\lambda_{n}+K>0\) for every \(n\). Rests on Theorems A42.9 and A42.11.

Proof.

Derives Lemma A42.12. Apply Equation (A42.14) with \(f=y_{n}\): \(B_{K}\left[Ty_{n},\chi\right]=\avg{y_{n},\chi}_{r}\). Since \(Ty_{n}=\mu_{n}y_{n}\) and \(B_{K}\) is bilinear, the left side is \(\mu_{n}B_{K}\left[y_{n},\chi\right]\); dividing by \(\mu_{n}>0\) and writing \(\mu_{n}^{-1}=\lambda_{n}+K\) gives Equation (A42.18). Taking \(\chi=y_{m}\) and using orthonormality in \(\avg{\cdot,\cdot}_{r}\) gives the second statement, and \(\lambda_{n}+K=\mu_{n}^{-1}>0\).

Proof of Theorem A42.1. Derives Theorem A42.1. Part (1) is Theorem A42.11. For (2), put

\begin{equation}\tag{A42.19} \tilde y_{n}=\frac{y_{n}}{\sqrt{\lambda_{n}+K}}\ep \end{equation}

By Lemma A42.12 the system \(\set{\tilde y_{n}}\) is orthonormal for the inner product \(B_{K}\):

\begin{equation*} B_{K}\left[\tilde y_{m},\tilde y_{n}\right] =\frac{\left(\lambda_{n}+K\right)\delta_{mn}} {\sqrt{\left(\lambda_{m}+K\right)\left(\lambda_{n}+K\right)}} =\delta_{mn}\ep \end{equation*}

It is maximal in \(H_{E}\): if \(\chi\in H_{E}\) satisfies \(B_{K}\left[\tilde y_{n},\chi\right]=0\) for every \(n\), then Equation (A42.18) gives \(\avg{y_{n},\chi}_{r}=0\) for every \(n\), and \(\set{y_{n}}\) is an orthonormal basis of \(L^{2}_{r}\) by part (1), so \(\chi=0\) as an element of \(L^{2}_{r}\); being continuous, \(\chi\) vanishes identically. Since \(\left(H_{E},B_{K}\right)\) is a Hilbert space (Lemma A42.4), maximality is completeness: Theorem 16.30 applies and \(\set{\tilde y_{n}}\) is an orthonormal basis of \(H_{E}\).

For (3), let \(y\in H_{E}\) and \(c_{n}=\avg{y,y_{n}}_{r}\). The energy-Fourier coefficients of \(y\) are, by Equation (A42.18),

\begin{equation}\tag{A42.20} B_{K}\left[\tilde y_{n},y\right] =\frac{\left(\lambda_{n}+K\right)\avg{y_{n},y}_{r}} {\sqrt{\lambda_{n}+K}} =\sqrt{\lambda_{n}+K}\;c_{n}\ec \end{equation}

so the \(N\)-th partial sum of the expansion of \(y\) in the basis \(\set{\tilde y_{n}}\) is

\begin{equation*} \sum_{n\le N}B_{K}\left[\tilde y_{n},y\right]\tilde y_{n} =\sum_{n\le N}\sqrt{\lambda_{n}+K}\;c_{n}\, \frac{y_{n}}{\sqrt{\lambda_{n}+K}} =\sum_{n\le N}c_{n}y_{n}=S_{N}\ec \end{equation*}

exactly the partial sum formed in the chapter. By Equation (16.18) the expansion converges in the norm of \(H_{E}\), which says \(B_{K}\left[y-S_{N},y-S_{N}\right]\longrightarrow0\), and by Parseval's identity Equation (16.19) together with Equation (A42.20),

\begin{equation}\tag{A42.21} B_{K}[y,y]=\sum_{n\ge1}\left(\lambda_{n}+K\right)c_{n}^{2}\ec \end{equation}

the series converging because it is a Parseval sum. Finally, \(\set{y_{n}}\) is an orthonormal basis of \(L^{2}_{r}\), so Parseval's identity there gives \(\avg{y,y}_{r}=\sum_{n}c_{n}^{2}\); subtracting \(K\) times that from Equation (A42.21) and using \(B_{K}[y,y]=\int\left(p\,y'^{2}+q\,y^{2}\right)\dd x +K\avg{y,y}_{r}\) leaves Equation (A42.4). Both series converge separately — the first as a Parseval sum, the second because \(\sum_{n}c_{n}^{2}=\avg{y,y}_{r}<\infty\) — so the subtraction is legitimate.

Remark A42.13 (The chapter's inequality is untouched).

Remark 20.82 insists that the inequality Equation (20.74), which is what the variational characterisation Theorem 20.81 rests on, is independent of the identity proved here, and that remains true: nothing above is used in the chapter's proof of Theorem 20.81, which argues from Equation (20.72) and the pairing Equation (20.73) to \(B_{K}[y,y]\ge\sum_{n\le N}\left(\lambda_{n}+K\right)c_{n}^{2}\) for every finite \(N\) — Bessel's inequality in the energy inner product, which needs no completeness at all. Conversely nothing in this appendix uses Theorem 20.81, so there is no circle. What this appendix does remove is a hypothesis: Theorem 20.81 assumes that the eigenfunctions are complete in \(L^{2}_{r}\), and Theorem A42.11 proves it. The difference between the two statements is exactly the difference between Bessel's inequality Equation (16.15) and Parseval's identity Equation (16.19), which is the four-way criterion Theorem 16.30 in the one place where the distinction has a consequence a reader can feel: by Corollary 20.83 the Rayleigh–Ritz error is second order in the deviation measured in the energy norm, and the estimate is vacuous without Equation (A42.4).

Remark A42.14 (What is quoted here).

Two inputs are used and not proved in this section, and one further result is proved elsewhere in this appendix.

  1. The completeness of \(L^{2}\) (Theorem 16.12) and, with it, the density of the continuous functions in \(L^{2}(a,b)\). The first is quoted in Hilbert Spaces itself (Remark 16.1) as resting on the convergence theorems of Lebesgue integration, which this treatise does not develop; the second belongs to the same package and is used twice above, both times to pass from continuous \(f\) to \(f\in L^{2}_{r}\) in Theorem A42.9. Reed and Simon [Reed:1972] is the reference of record for both.

  2. Nothing else. The Hilbert–Schmidt theorem Theorem 16.44 is not an import: it is proved, from the numerical-radius formula and sequential compactness, in The Hilbert–Schmidt Theorem for Compact Self-Adjoint Operators of this appendix. The Arzelà–Ascoli theorem used for the compactness of \(T\) is proved in Lemma A41.7, and the existence and uniqueness theory for the second-order equation defining \(\phi\) and \(\psi\) is Lemma A40.4, both in this appendix.

Remark A42.15.

Theorem A42.1 discharges the derivation owed in Remark 20.82 of Section 20.5.3: the eigenfunctions of a regular Sturm–Liouville problem are complete in the energy norm, and Equation (20.75) holds for every admissible trial function. The identity is used in the chapter to make the error of a Ritz estimate second order (Corollary 20.83), and through that corollary it underwrites the convergence of the Ritz scheme Equation (20.76) and of its descendants — the finite element method, the linear variational method for molecular orbitals, and the variational principle for the ground-state energy of Approximation Methods. The eigenvalue problem itself is that of Ordinary Differential Equations and Sturm–Liouville Theory, whose orthogonality theorem Theorem 13.57 is here recovered as a by-product: eigenfunctions belonging to different eigenvalues are orthogonal in \(\avg{\cdot,\cdot}_{r}\) because they are eigenvectors of a self-adjoint operator for different eigenvalues.