The Strengthened Jacobi Condition and Conjugate Points
This appendix proves the equivalence asserted in Remark 20.62 of Calculus of Variations: for a regular problem, the hypothesis of Theorem 20.61 — the existence of a solution of the Jacobi accessory equation Equation (20.53) with no zero on the closed interval \([a,b]\) — holds if and only if no point of \(\left(a,b\right]\) is conjugate to \(a\) in the sense of Definition 20.58. That equivalence is what joins Jacobi's necessary condition (Theorem 20.60) to his sufficient one, and without it the two theorems test different hypotheses.
Three ingredients are needed and all three are proved here: global existence and uniqueness for the accessory equation, which is Corollary 13.9 of Ordinary Differential Equations and Sturm–Liouville Theory once the equation is written as a first-order system; the continuous dependence of a solution on the point at which its initial data are posed, which is proved from the integral equation Equation (13.7) underlying the Picard–Lindelöf theorem Theorem 13.8 together with a Grönwall estimate; and the Sturm separation theorem, which is proved from the Lagrange identity Lemma 13.56. The notation is that of Section 20.4.3 throughout: \(P\) and \(Q\) are the coefficients Equation (20.46) evaluated along the extremal, the problem is regular, meaning \(P>0\) on \([a,b]\), and the accessory equation is
Only the continuity of \(P\) and \(Q\) and the positivity of \(P\) are used; in particular \(P\) is never differentiated, which is why the results below hold for every \(C^{3}\) integrand and \(C^{2}\) extremal without a further regularity hypothesis.
Statement
Let \(P,Q\in C^{0}[a,b]\) with \(P>0\) on \([a,b]\), and let \(u_{a}\) denote the solution of Equation (A40.1) determined by
The following three statements are equivalent.
-
There exists a solution \(w\) of Equation (A40.1) with \(w(x)\neq0\) for every \(x\in[a,b]\) (the strengthened Jacobi condition).
-
No point of \(\left(a,b\right]\) is conjugate to \(a\) (Definition 20.58).
-
\(u_{a}>0\) on \(\left(a,b\right]\).
Rests on Definition 20.58 and Equation (20.53).
Definition 20.58 normalises the Jacobi solution by \(u'(a)=1\), whereas Equation (A40.2) normalises the quantity \(P\,\dd u/\dd x\), which is the natural momentum variable of Equation (A40.1). The two differ by the positive factor \(P(a)\), and a nonzero constant multiple of a solution has exactly the same zeros, so the set of points conjugate to \(a\) is the same under either normalisation. The momentum normalisation is used below because \(P\) is not assumed differentiable and \(P\,\dd u/\dd x\), unlike \(\dd u/\dd x\) alone, is a component of the first-order system Equation (A40.3).
The proof occupies the rest of the section: the system form and the elementary structure of the zeros (The accessory equation as a first-order system), continuous dependence on the initial point (Continuous dependence on the initial point), the Sturm separation theorem (The Sturm separation theorem), and the equivalence itself (Proof of the equivalence), after which The displaced initial point recovers the exact phrasing used in the chapter, in which the initial point is displaced to the left of \(a\).
The accessory equation as a first-order system
For a solution \(u\) of Equation (A40.1) put \(v=P\,\dd u/\dd x\) and \(\vect{Y}=\left(u,v\right)\). Then \(\vect{Y}\) solves the linear system
whose coefficient matrix is continuous on \([a,b]\) because \(P\) and \(Q\) are continuous and \(P\) does not vanish. Conversely, if \(\vect{Y}=(u,v)\) solves Equation (A40.3) then \(u\) is of class \(C^{1}\) with \(P\,\dd u/\dd x=v\) of class \(C^{1}\), and \(\dd\left(P\,\dd u/\dd x\right)/\dd x=Qu\), which is Equation (A40.1). Rests on Equations (13.3) and (20.53).
For every \(s\in[a,b]\) and every \(\vect{Y}_{0}\in\R^{2}\) the system Equation (A40.3) has exactly one solution on the whole of \([a,b]\) with \(\vect{Y}(s)=\vect{Y}_{0}\). Consequently, for a solution \(u\) of Equation (A40.1) that does not vanish identically:
-
at every zero \(x_{0}\) of \(u\) one has \(\left(P\,\dd u/\dd x\right)(x_{0})\neq0\), so the zero is simple;
-
the zeros of \(u\) in \([a,b]\) are isolated, and there are finitely many of them.
Rests on Definition A40.3 and Corollary 13.9.
Derives Lemma A40.4. The system Equation (A40.3) is linear with a continuous coefficient matrix on the compact interval \([a,b]\), so Corollary 13.9 applies verbatim with \(\vect{g}=\vect{0}\): there is exactly one solution on all of \([a,b]\) for each choice of \(s\) and \(\vect{Y}_{0}\).
(1) If \(u(x_{0})=0\) and \(\left(P\,\dd u/\dd x\right)(x_{0})=0\) then \(\vect{Y}(x_{0})=\vect{0}\); but \(\vect{Y}\equiv\vect{0}\) is a solution with those data, so uniqueness forces \(\vect{Y}\equiv\vect{0}\) and hence \(u\equiv0\), which was excluded.
(2) Let \(x_{0}\) be a zero. By (1) and \(P>0\) we have \(\left(\dd u/\dd x\right)(x_{0})\neq0\), so by the definition of the derivative there is a punctured neighbourhood of \(x_{0}\) in which \(u(x)/(x-x_{0})\) keeps the sign of \(\left(\dd u/\dd x\right)(x_{0})\) and in particular \(u(x)\neq0\): the zero is isolated. If there were infinitely many zeros in \([a,b]\), Bolzano–Weierstrass (Theorem 11.7) would give a sequence of distinct zeros converging to some \(x_{\ast}\in[a,b]\); continuity of \(u\) makes \(x_{\ast}\) a zero as well, and it is not isolated — a contradiction.
∎Continuous dependence on the initial point
The estimate below is the one place where a quantitative statement is needed rather than a qualitative one, and it is obtained from the integral form Equation (13.7) of the initial value problem, which is where the proof of Theorem 13.8 begins. Throughout, \(\norm{\cdot}\) is any fixed norm on \(\R^{2}\) and \(\norm{A(x)}\) the induced operator norm, so that \(\norm{A(x)\vect{Z}}\le\norm{A(x)}\,\norm{\vect{Z}}\); put
which is finite because \(A\) is continuous on a compact interval (Theorem 11.24).
Let \(\varphi\in C^{0}[a,b]\) be nonnegative, let \(x_{0}\in[a,b]\), and let \(C\ge0\) and \(L\ge0\) satisfy
Then \(\varphi(x)\le C\,\ee^{L\abs{x-x_{0}}}\) on \([a,b]\). Rests on Theorems 11.35 and 11.42.
Derives Lemma A40.5. Take first \(x\ge x_{0}\) and set \(\Phi(x)=C+L\int_{x_{0}}^{x}\varphi(t)\,\dd t\), which is of class \(C^{1}\) with \(\Phi'=L\varphi\) by the fundamental theorem of calculus (Theorem 11.42). Hypothesis Equation (A40.5) reads \(\varphi\le\Phi\) on \([x_{0},b]\), so \(\Phi'=L\varphi\le L\Phi\) there, and therefore
By the mean value theorem (Theorem 11.35), a function whose derivative is nowhere positive on an interval satisfies \(g(x)-g\left(x_{0}\right)=g'(\xi)\left(x-x_{0}\right)\le0\) for some interior \(\xi\) whenever \(x>x_{0}\); applied to \(g(x)=\Phi(x)\,\ee^{-L\left(x-x_{0}\right)}\) this gives \(\Phi(x)\,\ee^{-L\left(x-x_{0}\right)}\le\Phi(x_{0})=C\), that is \(\varphi(x)\le\Phi(x)\le C\,\ee^{L\left(x-x_{0}\right)}\).
For \(x\le x_{0}\) apply what has just been proved to \(\tilde\varphi(t)=\varphi\left(x_{0}-t\right)\) on \(\left[0,x_{0}-a\right]\), which satisfies the same hypothesis with the same constants because the substitution \(t\longmapsto x_{0}-t\) turns \(\left|\int_{x_{0}}^{x}\varphi\right|\) into \(\left|\int_{0}^{x_{0}-x}\tilde\varphi\right|\).
∎For \(s\in[a,b]\) and \(\vect{Y}_{0}\in\R^{2}\) write \(\vect{Y}\left(\cdot\,;s,\vect{Y}_{0}\right)\) for the solution of Equation (A40.3) with \(\vect{Y}(s)=\vect{Y}_{0}\). Then, with \(L\) as in Equation (A40.4), for all \(s,\tilde s\in[a,b]\) and all \(\vect{Y}_{0},\tilde{\vect{Y}}_{0}\),
In particular the map \(\left(x,s,\vect{Y}_{0}\right)\longmapsto \vect{Y}\left(x;s,\vect{Y}_{0}\right)\) is continuous on \([a,b]\times[a,b]\times\R^{2}\). Rests on Lemma A40.5, Lemma A40.4 and Equation (13.7).
Derives Theorem A40.6. Write \(\vect{Y}=\vect{Y}\left(\cdot\,;s,\vect{Y}_{0}\right)\) and \(\vect{Z}=\vect{Y}\left(\cdot\,;\tilde s,\tilde{\vect{Y}}_{0}\right)\). Both are continuous and, by the equivalence of the initial value problem with its integral form Equation (13.7) — which is the first step of the proof of Theorem 13.8 and uses only the fundamental theorem of calculus —
A bound for \(\vect{Z}\) alone. Taking norms in the second identity of Equation (A40.7),
so Lemma A40.5 with \(\varphi=\norm{\vect{Z}}\), \(C=\norm{\tilde{\vect{Y}}_{0}}\) and \(x_{0}=\tilde s\) gives
The difference. Subtracting the two identities of Equation (A40.7) and splitting the second integral at \(s\),
The last term is bounded in norm by \(L\,\abs{s-\tilde s}\,\max_{[a,b]}\norm{\vect{Z}}\), which Equation (A40.8) bounds by \(L\,\ee^{L\left(b-a\right)}\norm{\tilde{\vect{Y}}_{0}}\, \abs{s-\tilde s}\). Hence, writing \(\varphi=\norm{\vect{Y}-\vect{Z}}\) and \(C=\norm{\vect{Y}_{0}-\tilde{\vect{Y}}_{0}} +L\,\ee^{L\left(b-a\right)}\norm{\tilde{\vect{Y}}_{0}}\, \abs{s-\tilde s}\),
and Lemma A40.5 gives Equation (A40.6).
For the final assertion, fix \(\left(x_{1},s_{1},\vect{Y}_{1}\right)\). The estimate Equation (A40.6) shows that \(\vect{Y}\left(\cdot\,;s,\vect{Y}_{0}\right)\) converges uniformly in \(x\) to \(\vect{Y}\left(\cdot\,;s_{1},\vect{Y}_{1}\right)\) as \(\left(s,\vect{Y}_{0}\right)\longrightarrow \left(s_{1},\vect{Y}_{1}\right)\), and each solution is continuous in \(x\); a uniform limit of continuous functions is continuous, and the triangle inequality
makes both contributions small.
∎The Sturm separation theorem
Let \(u_{1},u_{2}\) solve Equation (A40.1) on \([a,b]\) and put
Then \(W\) is constant on \([a,b]\), and \(W=0\) if and only if \(u_{1}\) and \(u_{2}\) are linearly dependent. Rests on Lemmas 13.56 and A40.4.
Derives Lemma A40.7. Equation (A40.1) is the Sturm–Liouville equation \(L[u]=0\) of Lemma 13.56 with the data \(p=P\) and \(q=-Q\). The Lagrange identity Equation (13.56) therefore gives
so \(W\) is constant. In the notation of Definition A40.3, \(W\) is the determinant of the matrix with columns \(\vect{Y}_{1}=\left(u_{1},v_{1}\right)\) and \(\vect{Y}_{2}=\left(u_{2},v_{2}\right)\), since \(u_{1}v_{2}-u_{2}v_{1}=P\left(u_{1}u_{2}'-u_{2}u_{1}'\right)\). If \(W=0\), the two vectors are parallel at every point; picking a point \(x_{0}\) at which \(\vect{Y}_{1}\left(x_{0}\right)\neq\vect{0}\) — one exists unless \(u_{1}\equiv0\), in which case the dependence is trivial — gives \(\vect{Y}_{2}\left(x_{0}\right) =\kappa\,\vect{Y}_{1}\left(x_{0}\right)\) for a scalar \(\kappa\), and uniqueness in Lemma A40.4 then forces \(\vect{Y}_{2}=\kappa\vect{Y}_{1}\) identically, hence \(u_{2}=\kappa u_{1}\). Conversely \(u_{2}=\kappa u_{1}\) makes Equation (A40.9) vanish outright.
∎Let \(u_{1},u_{2}\) be linearly independent solutions of Equation (A40.1) on \([a,b]\). Then between any two consecutive zeros of \(u_{1}\) there lies exactly one zero of \(u_{2}\), and \(u_{1},u_{2}\) have no zero in common. In particular the zeros of the two solutions strictly interlace. Rests on Lemma A40.7 and Theorem 11.23.
Derives Theorem A40.8. By Lemma A40.7 the constant \(W=P\left(u_{1}u_{2}'-u_{2}u_{1}'\right)\) is nonzero. Neither solution vanishes identically, so Lemma A40.4 applies to both.
No common zero. If \(u_{1}\left(x_{0}\right) =u_{2}\left(x_{0}\right)=0\) then both terms of Equation (A40.9) vanish at \(x_{0}\) and \(W=0\), a contradiction.
At least one zero in between. Let \(x_{1}<x_{2}\) be consecutive zeros of \(u_{1}\) in \([a,b]\), so that \(u_{1}\neq0\) on \(\left(x_{1},x_{2}\right)\). Evaluating Equation (A40.9) at the two ends, where \(u_{1}\) vanishes,
Both derivatives are nonzero by Lemma A40.4(1), and they have opposite signs: \(u_{1}\) keeps one sign, say positive, throughout \(\left(x_{1},x_{2}\right)\), so \(\left(\dd u_{1}/\dd x\right)\left(x_{1}\right) =\lim_{x\to x_{1}^{+}}u_{1}(x)/\left(x-x_{1}\right)\ge0\) and hence is \(>0\), while \(\left(\dd u_{1}/\dd x\right)\left(x_{2}\right) =\lim_{x\to x_{2}^{-}}u_{1}(x)/\left(x-x_{2}\right)\le0\) and hence is \(<0\); if \(u_{1}<0\) in between, both signs reverse. Since \(W\neq0\) is the same number at both ends of Equation (A40.10) and \(P>0\), the values \(u_{2}\left(x_{1}\right)\) and \(u_{2}\left(x_{2}\right)\) are both nonzero and of opposite sign. The intermediate value theorem (Theorem 11.23) gives a zero of \(u_{2}\) in \(\left(x_{1},x_{2}\right)\).
At most one. Suppose \(u_{2}\) had two zeros \(y_{1}<y_{2}\) in \(\left(x_{1},x_{2}\right)\); choosing them consecutive among the zeros of \(u_{2}\) — possible because by Lemma A40.4(2) there are finitely many — and applying the previous paragraph with the roles of \(u_{1}\) and \(u_{2}\) exchanged produces a zero of \(u_{1}\) in \(\left(y_{1},y_{2}\right)\subset\left(x_{1},x_{2}\right)\), contradicting the assumption that \(x_{1}\) and \(x_{2}\) are consecutive zeros of \(u_{1}\).
∎Proof of the equivalence
Proof of Theorem A40.1. Derives Theorem A40.1. The three implications are proved in the cycle \((1)\Rightarrow(2)\Rightarrow(3)\Rightarrow(1)\).
\((1)\Rightarrow(2)\). Let \(w\) be a solution with no zero on \([a,b]\) and suppose, for contradiction, that some \(c\in\left(a,b\right]\) is conjugate to \(a\), i.e.\ \(u_{a}(c)=0\). The solution \(u_{a}\) does not vanish identically, since \(\left(P\,\dd u_{a}/\dd x\right)(a)=1\neq0\). Moreover \(u_{a}\) and \(w\) are linearly independent: a multiple of \(w\) cannot vanish at \(a\) as \(u_{a}\) does, unless the multiple is zero. By Lemma A40.4(2) the zeros of \(u_{a}\) in \([a,b]\) are finite in number, so among those lying in \(\left(a,c\right]\) there is a smallest one, \(c_{1}\); then \(a\) and \(c_{1}\) are consecutive zeros of \(u_{a}\). The Sturm separation theorem Theorem A40.8 places a zero of \(w\) in \(\left(a,c_{1}\right)\subset[a,b]\), contradicting the hypothesis on \(w\).
\((2)\Rightarrow(3)\). Since \(\left(P\,\dd u_{a}/\dd x\right)(a)=1\) and \(P(a)>0\), we have \(\left(\dd u_{a}/\dd x\right)(a)=1/P(a)>0\), and with \(u_{a}(a)=0\) this makes \(u_{a}(x)>0\) for \(x>a\) close enough to \(a\). By (2), \(u_{a}\) has no zero in \(\left(a,b\right]\); were \(u_{a}\) negative at some point of \(\left(a,b\right]\), the intermediate value theorem (Theorem 11.23) would produce a zero between that point and one where \(u_{a}>0\). Hence \(u_{a}>0\) throughout \(\left(a,b\right]\).
\((3)\Rightarrow(1)\). Let \(z\) be the solution of Equation (A40.1) with
which exists by Lemma A40.4, and consider the one-parameter family of solutions
each of which solves Equation (A40.1) by linearity. We show that \(w_{\delta}>0\) on \([a,b]\) for all small \(\delta>0\).
Write \(v_{a}=P\,\dd u_{a}/\dd x\), a continuous function with \(v_{a}(a)=1\), and recall \(z(a)=1\) with \(z\) continuous. Choose \(\eta>0\) so small that
On that subinterval \(\dd u_{a}/\dd x=v_{a}/P>0\), so \(u_{a}(x)=\int_{a}^{x}\left(v_{a}/P\right)\dd t\ge0\) by Theorem 11.43, and therefore
If \(a+\eta\ge b\) this already proves the claim for every \(\delta>0\). Otherwise, on the compact interval \(\left[a+\eta,b\right]\) the continuous function \(u_{a}\) is strictly positive by (3), so by the extreme value theorem (Theorem 11.24) it has a minimum \(m>0\) there; let \(M=\max_{[a,b]}\abs{z}\), also finite by Theorem 11.24. Choosing any \(\delta\) with \(0<\delta<m/\left(M+1\right)\) gives, for \(x\in\left[a+\eta,b\right]\),
So \(w=w_{\delta}\) has no zero on \([a,b]\), which is (1).
∎With Theorem A40.1 the two halves of the second-order theory close on each other. For a regular problem, Theorem 20.60 says that a weak minimum admits no conjugate point in the open interval \((a,b)\), and Theorem 20.61 says that a solution without zeros on the closed interval makes \(\delta^{2}J\) positive definite; by statement (2) of the theorem just proved, the hypothesis of the second is exactly the absence of conjugate points in \(\left(a,b\right]\). The one remaining gap between them is the single point \(x=b\): an extremal whose first conjugate point is the right endpoint itself satisfies the necessary condition and fails the sufficient one, and the second variation is then positive semidefinite with the Jacobi solution \(u_{a}\) in its null space — which is precisely the computation carried out in the proof of Theorem 20.60, where \(\delta^{2}J\left[y;\eta\right]=0\) for the broken variation built from \(u_{a}\). The threshold case is not a defect of the theory: it is the kinetic focus, and Remark 20.79 exhibits it as the exact instant \(T=\pi\sqrt{m/k}\) at which the harmonic oscillator's action stops being a minimum.
The displaced initial point
Remark 20.62 states the equivalence in the form in which it is usually met: the Jacobi solution vanishing slightly to the left of \(a\) has no zero in \([a,b]\) exactly when the solution vanishing at \(a\) has none in \(\left(a,b\right]\). That form presupposes that the coefficients are defined on a slightly larger interval, which happens whenever the extremal itself extends; under that hypothesis it is a corollary of Theorems A40.1 and A40.6, and it is the form in which continuous dependence on the initial point does real work.
Suppose \(P\) and \(Q\) extend continuously to \(\left[a-\varepsilon_{0},b\right]\) for some \(\varepsilon_{0}>0\), with \(P>0\) there, and for \(s\in\left[a-\varepsilon_{0},b\right)\) let \(u_{s}\) be the solution of Equation (A40.1) with \(u_{s}(s)=0\) and \(\left(P\,\dd u_{s}/\dd x\right)(s)=1\). Then the strengthened Jacobi condition holds on \([a,b]\) if and only if there is an \(\varepsilon\in\left(0,\varepsilon_{0}\right)\) such that \(u_{a-\varepsilon}\) has no zero in \([a,b]\) — and in that case no zero in \(\left[a,b\right]\) for every smaller \(\varepsilon\) as well. Rests on Theorems A40.1 and A40.6.
Derives Corollary A40.10. Sufficiency is immediate: \(u_{a-\varepsilon}\) restricted to \([a,b]\) is a solution of Equation (A40.1) without zeros there, which is statement (1) of Theorem A40.1.
Necessity. Assume the strengthened Jacobi condition, hence statement (3): \(u_{a}>0\) on \(\left(a,b\right]\). Work on the extended interval \(I=\left[a-\varepsilon_{0},b\right]\), on which Theorem A40.6 holds with the constant \(L\) of Equation (A40.4) computed over \(I\). Write \(v_{s}=P\,\dd u_{s}/\dd x\), so that the initial vector is \(\vect{Y}_{0}=\left(0,1\right)\) for every \(s\) and only the initial point varies; Equation (A40.6) then reads
where the norm on \(\R^{2}\) has been taken to be the sum of the moduli of the components and \(\norm{\vect{Y}_{0}}=1\).
Since \(v_{a}\) is continuous on \(I\) with \(v_{a}(a)=1\), fix \(\eta>0\) with \(\eta<\varepsilon_{0}\) and
If \(a+\eta<b\), the function \(u_{a}\) is continuous and strictly positive on the compact interval \(\left[a+\eta,b\right]\) and so has a minimum \(m>0\) there (Theorem 11.24); if \(a+\eta\ge b\) put \(m=+\infty\) and ignore the second condition below. Now choose \(\varepsilon>0\) subject to
By Equation (A40.13) the solution \(u=u_{a-\varepsilon}\) and its momentum \(v=v_{a-\varepsilon}\) then satisfy
On \(\left[a-\varepsilon,a+\eta\right]\) the first inequality gives \(\dd u/\dd x=v/P>0\), so \(u\) is strictly increasing there and, since \(u\left(a-\varepsilon\right)=0\), it is strictly positive on \(\left(a-\varepsilon,a+\eta\right]\) — in particular on \(\left[a,a+\eta\right]\), because \(a>a-\varepsilon\). Together with the second inequality, \(u>0\) on all of \([a,b]\), so \(u_{a-\varepsilon}\) has no zero there. The argument used only the three smallness conditions on \(\varepsilon\), each of which persists when \(\varepsilon\) is decreased, which proves the last clause.
∎Corollary A40.10 is the analytic content of the picture in Remark 20.59. A solution of Equation (A40.1) is the derivative \(\pp y/\pp\alpha\) of a one-parameter family of extremals through a common initial point, so \(u_{s}\) describes, to first order, the spreading of the pencil of extremals issuing from the point \(s\). The corollary says that the pencil issuing from \(a\) stays spread out over \([a,b]\) if and only if a pencil issuing from a slightly earlier point does, and Theorem A40.6 is what makes “slightly earlier” a legitimate perturbation: the focusing distance depends continuously on the point of emission. On the unit sphere, where the extremals are great circles and every pencil refocuses at the antipode, the corollary reproduces the familiar statement that an arc is a minimising geodesic exactly while it is shorter than half a great circle.
Theorem A40.1 discharges the equivalence asserted in Remark 20.62 of Section 20.4.3, and Corollary A40.10 supplies the displaced-point phrasing used there. The two conditions the equivalence links are Jacobi's necessary condition (Theorem 20.60) and the sufficiency of the second variation (Theorem 20.61); with the equivalence in hand, a regular extremal with no conjugate point in \(\left(a,b\right]\) has a positive definite second variation, which is the form in which the result is used in the field-theoretic sufficiency argument of Section 20.4.5 and in the discussion of the kinetic focus of Hamilton's principle in Section 20.7.1.