Calculus of Variations
Functionals, the Euler–Lagrange equations, constraints and Lagrange multipliers, and second-variation conditions — the mathematical engine behind every action principle in this treatise. Where the differential calculus of Real Analysis finds the points at which a function of finitely many variables is stationary, the calculus of variations finds the curves and fields at which a functional — a number assigned to a whole function, typically an integral \(J[y]=\int f(x,y,y')\,\dd x\) — is stationary. The subject was created for isolated geometric problems, systematised by Euler [Euler:1744] and Lagrange [Lagrange:1762], and became the common language of physics when mechanics, optics, field theory and general relativity were each recast as the statement that an action is stationary. The canonical modern exposition is Gelfand–Fomin [Gelfand:1963].
This chapter is consumed three times over: the Euler–Lagrange machinery feeds Lagrangian mechanics in Lagrangian Mechanics, the several-variable theory feeds classical field theory in Generalized Classical Field Theory and the gravitational action of The Einstein Field Equations, and Noether's theorem [Noether:1918] is the origin of every conservation law derived in this book.
The logical order of the chapter inverts its historical order. The three problems of Section 16.1 were posed and solved before any general method existed; they are stated first because they fix what the method must deliver, and each is solved here with the machinery of Section 16.2, to which the solutions forward-refer. A reader who wants the theory before the examples may begin at Section 16.2.1 and return.
The classical problems
The brachistochrone
In the June 1696 issue of the Acta Eruditorum Johann Bernoulli challenged the mathematicians of Europe to find, among all curves joining two given points in a vertical plane, the one down which a bead sliding without friction under gravity travels in least time [Bernoulli:1696]. Solutions by Newton, Leibniz, l'Hôpital and both Bernoullis appeared in the issue of May 1697. The problem is the origin of the subject: it asks for a minimum not over a finite list of numbers but over a whole class of curves.
Set the starting point at the origin, let \(x\) be measured horizontally and \(y\) vertically downwards, and let the bead of mass \(m\) start from rest. Energy conservation gives the speed at depth \(y\) as \(v=\sqrt{2gy}\), with \(g=9.80665\,\mathrm{m}/\mathrm{s}^{2}\) the standard acceleration of gravity; the mass cancels, which is why the answer is a purely geometric curve. Since the arc element is \(\dd s=\sqrt{1+y'^{2}}\,\dd x\), the transit time to the endpoint \((x_{1},y_{1})\) is the functional
measured in seconds, over curves with \(y(0)=0\), \(y(x_{1})=y_{1}>0\) and \(y>0\) on \((0,x_{1}]\).
Every extremal of Equation (16.1) is an arc of the cycloid
traced by a point on the rim of a circle of radius \(a\) rolling on the underside of the horizontal line \(y=0\). The radius \(a\), of dimension length, and the terminal parameter \(\theta_{1}\) are fixed by the endpoint condition. Rests on Equations (16.1) and (16.25).
Derives Proposition 16.1. The integrand \(f(y,y')=\bigl[(1+y'^{2})/(2gy)\bigr]^{1/2}\) carries no explicit \(x\), so the Beltrami identity Equation (16.25) applies: \(f-y'\,\pp f/\pp y'\) is constant along any extremal. With
the combination collapses to
that is, \(y\left(1+y'^{2}\right)=2a\) with \(2a=1/(2gC^{2})\), a length. Hence \(y'=\sqrt{(2a-y)/y}\), and \(y\) never exceeds \(2a\). Substituting \(y=a(1-\cos\theta)=2a\sin^{2}(\theta/2)\) gives \(2a-y=2a\cos^{2}(\theta/2)\) and therefore \(y'=\cos(\theta/2)/\sin(\theta/2)\), so that
Integrating from \(\theta=0\), where \(x=y=0\), gives Equation (16.2).
∎The cycloid leaves the origin vertically (\(y'\to\infty\) as \(\theta\to0\)), so the bead is accelerated hardest at the start; and for \(y_{1}\) small enough relative to \(x_{1}\) the optimal arc dips below the endpoint and climbs back, which no argument from symmetry would suggest. That the extremal is in fact a minimum, and not merely a stationary curve, follows from the sufficiency theory of Section 16.4.4: the integrand is convex in \(y'\), so the Weierstrass excess function is nonnegative.
Johann Bernoulli's own solution did not use any variational calculus. He replaced the falling bead by a light ray crossing a medium of continuously varying refractive index \(n\propto1/v\propto y^{-1/2}\), and applied the law of refraction in the form \(\sin\vartheta/v=\text{const}\) that Section 16.7.3 derives from Fermat's principle. Writing \(\sin\vartheta=1/\sqrt{1+y'^{2}}\) reproduces \(y(1+y'^{2})=\text{const}\) immediately. The step is not a trick: the identity of the two problems is the optical–mechanical analogy, made systematic by Hamilton and Jacobi in Hamilton–Jacobi Theory and the Optical–Mechanical Analogy, and in this chapter it reappears as the common form of Lemma 16.100, which serves both Maupertuis' principle and Fermat's.
The catenary and the isoperimetric problem
A uniform flexible chain of linear mass density \(\rho_{\ell}\) (in \(\mathrm{kg}/\mathrm{m}\)) and fixed total length \(\ell\) hangs from two points. Galileo asserted that the resulting curve is a parabola [Galilei:1638]; it is not, as Huygens showed while still a teenager, and the correct curve was found in 1691 by Johann Bernoulli, Leibniz and Huygens independently [Bernoulli:1691]. In the variational reading the chain minimises its gravitational potential energy
in joules, with \(y\) measured upwards, subject to the fixed-length constraint
Let \(\ell\) exceed the distance between the suspension points. Every extremal of Equation (16.3) subject to Equation (16.4) is a catenary
with \(a>0\) a length, \(x_{c}\) a horizontal offset, and \(\lambda\) the Lagrange multiplier of the length constraint. It is not a parabola. Rests on Equation (16.3), Theorem 16.43 and Equation (16.25).
Derives Proposition 16.4. By Theorem 16.43 the constrained extremals are the free extremals of \(U+\rho_{\ell}g\lambda L\), whose integrand, after division by the constant \(\rho_{\ell}g\), is \(F(y,y')=(y+\lambda)\sqrt{1+y'^{2}}\). It carries no explicit \(x\), so the Beltrami identity Equation (16.25) gives
a constant with the dimension of length. Writing \(u=y+\lambda\) this reads \(u=a\sqrt{1+u'^{2}}\), that is \(u'=\pm\sqrt{u^{2}/a^{2}-1}\), whose separated form \(\dd u/\sqrt{u^{2}-a^{2}}=\pm\dd x/a\) integrates to \(u=a\cosh\bigl((x-x_{c})/a\bigr)\). Restoring \(y=u-\lambda\) gives Equation (16.5). Expanding, \(a\cosh(z/a)=a+z^{2}/(2a)+z^{4}/(24a^{3})+\dots\), so the catenary agrees with a parabola only to second order about its vertex — which is why Galileo's conjecture survived casual inspection.
∎The catenary is the extremal of \(\int y\sqrt{1+y'^{2}}\,\dd x\) with the length fixed, and also, with no constraint at all, of the surface-area functional of Proposition 16.9; the two integrands differ only by the additive constant \(\lambda\) that the constraint supplies. This is not a coincidence but the multiplier rule at work: fixing the length of the generating curve and dropping the constant term are the same modification of the integrand.
Queen Dido's problem — to enclose the greatest area with a boundary of given length — is the oldest extremal problem on record and the prototype of the class that carries its name.
Among closed, simple, piecewise-\(C^{1}\) plane curves of fixed length \(\ell\), every extremal of the enclosed area is a circle of radius \(\ell/2\pi\). Rests on Theorems 16.35 and 16.43.
Derives Proposition 16.6. Parametrise the curve by arclength \(s\in[0,\ell]\), so that \(\dot x^{2}+\dot y^{2}=1\), and use the shoelace form of the area, \(A=\tfrac{1}{2}\oint\left(x\,\dd y-y\,\dd x\right)\). By Theorem 16.43 the constrained extremals are the free extremals of the integrand
The system Equation (16.30) gives, after imposing the arclength gauge \(\sqrt{\dot x^{2}+\dot y^{2}}=1\),
that is \(\lambda\ddot x=\dot y\) and \(\lambda\ddot y=-\dot x\). These are consistent with the gauge, since \(\dv{}{s}(\dot x^{2}+\dot y^{2}) =2\dot x\dot y/\lambda-2\dot y\dot x/\lambda=0\), and they give the signed curvature
constant. A plane curve of constant nonzero curvature is a circular arc, and a closed one of length \(\ell\) is the full circle of radius \(\abs{\lambda}=\ell/2\pi\).
∎Proposition 16.6 identifies the only candidate; it does not by itself show that a maximiser exists. Steiner's nineteenth-century symmetrisation arguments have the same gap, and Weierstrass's criticism of them is what forced the existence question into the open (Section 16.5.1). Existence for this problem follows from the direct method of Section 16.5.2; Euler's founding monograph [Euler:1744] treats the whole class systematically, and the general multiplier rule for integral constraints is Theorem 16.43.
Minimal surfaces and soap films
A surface spanning a given closed contour and minimising area is a minimal surface. Plateau realised such surfaces physically as soap films, whose surface tension makes the area the energy, and catalogued both the surfaces and the junction rules that films obey [Plateau:1873]: three films always meet along an edge at equal angles of \(120\) degrees, and four such edges meet at a vertex at the tetrahedral angle \(\arccos(-1/3)\approx109.47\) degrees. The physics of the surface tension that enforces this belongs to Fluid Dynamics; the geometry is here.
For a surface given as a graph \(z=u(x,y)\) over a region \(\Omega\subset\R^{2}\), the area functional is
The Euler–Lagrange equation of Equation (16.6) is
equivalently the vanishing of the mean curvature of the graph. Rests on Equations (16.6) and (16.32).
Derives Proposition 16.8. Apply Equation (16.32) with two independent variables and the integrand \(f=W\equiv\sqrt{1+u_{x}^{2}+u_{y}^{2}}\), which does not depend on \(u\) itself. Since \(\pp f/\pp u_{x}=u_{x}/W\) and \(\pp f/\pp u_{y}=u_{y}/W\), the equation reads
Carrying out the differentiations, with \(W_{x}=(u_{x}u_{xx}+u_{y}u_{xy})/W\) and \(W_{y}=(u_{x}u_{xy}+u_{y}u_{yy})/W\), and multiplying through by \(W^{3}\),
which is Equation (16.7). The left-hand side of Equation (16.8) is \(2H\), twice the mean curvature of the graph in the sense of Differentiable Manifolds, Tensors, and Curvature, so the equation says \(H=0\).
∎Among surfaces of revolution about the \(x\) axis generated by a positive profile \(y(x)\), the extremals of the area \(A=2\pi\int y\sqrt{1+y'^{2}}\,\dd x\) are the catenoids
Derives Proposition 16.9. The integrand \(f=y\sqrt{1+y'^{2}}\) carries no explicit \(x\), so Equation (16.25) gives \(f-y'f_{y'}=y/\sqrt{1+y'^{2}}=a\), a constant length. This is the equation solved in Proposition 16.4 with \(\lambda=0\), whose solution is Equation (16.9).
∎The catenoid and the helicoid \(\left(v\cos u,\,v\sin u,\,cu\right)\) are the two minimal surfaces known to Euler and Meusnier, and they are locally isometric. Neither existence nor uniqueness is automatic: for two coaxial rings whose separation exceeds about \(1.3255\) times their radius no catenoid joins them at all, and the soap film jumps to the two flat discs — a physical realisation of the failure of a minimising sequence to converge inside the admissible class. The general existence statement for an arbitrary contour is the Plateau problem of Section 16.5.4.
Functionals and the first variation
Function spaces and functionals
Fix \(a<b\) in \(\R\), boundary values \(y_{a},y_{b}\in\R\), and let
be the admissible class with fixed endpoints. A functional on \(\mathcal{A}\) is a map \(J:\mathcal{A}\longrightarrow\R\). The functionals of this chapter are integrals
with the integrand (or Lagrangian) \(f\) of class \(C^{2}\) on \([a,b]\times\R\times\R\). Rests on Definition 7.66 and Theorem 7.40.
On \(C^{1}[a,b]\) define
A strong \(\varepsilon\)-neighbourhood of \(y\) is \(\set{\tilde y\in\mathcal{A}\mid\norm{\tilde y-y}_{0}<\varepsilon}\); a weak one is \(\set{\tilde y\in\mathcal{A}\mid\norm{\tilde y-y}_{1}<\varepsilon}\). Both are metric notions in the sense of Definition 6.24, and since \(\norm{\cdot}_{0}\le\norm{\cdot}_{1}\) the weak \(\varepsilon\)-neighbourhood is contained in the strong one of the same radius, so every strong neighbourhood contains a weak one: the weak topology on the admissible class is the finer of the two. The containment is strict in the worst possible way — no strong neighbourhood of \(0\) is contained in any weak one, since \(y_{n}(x)=n^{-1}\sin\left(n^{2}x\right)\) has \(\norm{y_{n}}_{0}\to0\) while \(\norm{y_{n}}_{1}\to\infty\). Rests on Definitions 6.24 and 16.11.
\(y\in\mathcal{A}\) is a weak local minimum of \(J\) if \(J[y]\le J[\tilde y]\) for every \(\tilde y\) in some weak neighbourhood of \(y\), and a strong local minimum if the same holds in some strong neighbourhood. Every strong minimum is a weak minimum; the converse fails. Rests on Definition 16.12.
Take \(a=0\), \(b=1\), \(y_{a}=y_{b}=0\) and \(f(x,y,y')=y'^{2}+y'^{3}\), and let \(y\equiv0\), so that \(J[y]=0\). If \(\norm{\tilde y}_{1}<\tfrac{1}{2}\) then \(\abs{\tilde y'}<\tfrac12\) pointwise and the integrand is \(\tilde y'^{2}\left(1+\tilde y'\right)\ge\tfrac12\tilde y'^{2}\ge0\), so \(J[\tilde y]\ge0\): the zero function is a weak minimum. It is not a strong one. Fix \(A>1\) and \(0<\delta<1\), and let \(\tilde y\) be the piecewise-linear function of slope \(-A\) on a set of total length \(\delta\) and slope \(A\delta/(1-\delta)\) on the rest, so that \(\tilde y(0)=\tilde y(1)=0\) and \(\norm{\tilde y}_{0}\le A\delta\). Then
whose first term is negative and of order \(\delta A^{3}\) while the second is of order \(\delta^{2}A^{2}\); taking \(A=\delta^{-1/2}\) makes \(J[\tilde y]<0\) while \(\norm{\tilde y}_{0}\le\delta^{1/2}\to0\). So arbitrarily small strong perturbations lower the functional. The distinction is exactly what separates the Legendre and Jacobi conditions of Sections 16.4.1 and 16.4.2 from the Weierstrass condition of Section 16.4.3 [Gelfand:1963]. Rests on Definition 16.13.
A variation at \(y\in\mathcal{A}\) is a function \(\eta\in C^{1}[a,b]\) with \(\eta(a)=\eta(b)=0\); the set of such \(\eta\) is written \(C^{1}_{0}[a,b]\), and \(y+\varepsilon\eta\in\mathcal{A}\) for every \(\varepsilon\in\R\). The first variation (the Gâteaux variation) of \(J\) at \(y\) in the direction \(\eta\) is
when the limit exists. It is the functional derivative of Equation (7.138) paired with \(\eta\). Rests on Definition 16.11 and Equation (7.138).
If \(y\in\mathcal{A}\) is a weak local extremum of \(J\) and \(\delta J[y;\eta]\) exists, then \(\delta J[y;\eta]=0\) for every \(\eta\in C^{1}_{0}[a,b]\). Rests on Definition 16.15, Definition 16.13 and Lemma 7.33.
Derives Proposition 16.16. Fix \(\eta\ne0\) and set \(\varphi(\varepsilon)=J[y+\varepsilon\eta]\). Because \(\norm{\left(y+\varepsilon\eta\right)-y}_{1} =\abs{\varepsilon}\,\norm{\eta}_{1}\), every \(\varepsilon\) with \(\abs{\varepsilon}<\varepsilon_{0}/\norm{\eta}_{1}\) places \(y+\varepsilon\eta\) inside the weak \(\varepsilon_{0}\)-neighbourhood in which \(y\) extremises \(J\). Hence \(\varphi\) has a local extremum at \(\varepsilon=0\) and is differentiable there, so \(\varphi'(0)=\delta J[y;\eta]=0\) by Fermat's lemma (Lemma 7.33).
∎For \(J\) as in Equation (16.11) with \(f\) of class \(C^{2}\), and for every \(\eta\in C^{1}[a,b]\),
the partial derivatives being evaluated at \(\left(x,y(x),y'(x)\right)\). Rests on Equations (16.11) and (16.14).
Derives Proposition 16.17. The map \((\varepsilon,x)\longmapsto f\left(x,y+\varepsilon\eta,y'+\varepsilon\eta'\right)\) is \(C^{1}\) on the compact set \([-1,1]\times[a,b]\), so its \(\varepsilon\)-derivative is continuous there and differentiation under the integral sign is legitimate. Carrying it out and using the chain rule (Proposition 7.31),
which integrates to Equation (16.15).
∎The fundamental lemma
Equation (16.15) turns the stationarity condition into the statement that a certain integral vanishes for every variation. Converting that into a pointwise differential equation is the work of the following three lemmas. The first is the fundamental lemma of the calculus of variations; the second is du Bois-Reymond's sharpening [duBoisReymond:1879], which is what makes it unnecessary to assume in advance that the extremal is twice differentiable.
Let \(g\in C^{0}[a,b]\) satisfy
Then \(g\equiv0\) on \([a,b]\). Rests on Definition 16.15 and Theorem 7.40.
Derives Lemma 16.18. Suppose \(g(x_{0})\ne0\) for some \(x_{0}\); replacing \(g\) by \(-g\) if necessary, assume \(g(x_{0})>0\), and by continuity \(x_{0}\) may be taken interior. By continuity again (Definition 7.20) there is \(\delta>0\) with \([x_{0}-\delta,x_{0}+\delta]\subset(a,b)\) and \(g>\tfrac12g(x_{0})\) on that interval. Put
This \(\eta\) is of class \(C^{1}\) — both it and \(\eta'=4(x-x_{0})\left[(x-x_{0})^{2}-\delta^{2}\right]\) vanish at \(x=x_{0}\pm\delta\) — and lies in \(C^{1}_{0}[a,b]\). It is strictly positive on the open interval, so
contradicting Equation (16.16).
∎Let \(h\in C^{0}[a,b]\) satisfy
Then \(h\) is constant on \([a,b]\). Rests on Equation (16.18) and Theorem 7.42.
Derives Lemma 16.19. Let \(c=\left(b-a\right)^{-1}\int_{a}^{b}h\,\dd x\) and define
Then \(\eta(a)=0\), and \(\eta(b)=\int_{a}^{b}h\,\dd x-c(b-a)=0\) by the choice of \(c\), so \(\eta\in C^{1}_{0}[a,b]\); it is of class \(C^{1}\) with \(\eta'=h-c\) by the fundamental theorem of calculus (Theorem 7.42). Substituting into Equation (16.18),
the last integral vanishing by the definition of \(c\). A continuous nonnegative integrand with zero integral vanishes identically, so \(h\equiv c\).
∎Let \(M,N\in C^{0}[a,b]\) satisfy
Then \(N\) is of class \(C^{1}\) and \(N'=M\) on \([a,b]\). Rests on Lemma 16.19 and Corollary 7.44.
Derives Lemma 16.20. Set \(A(x)=\int_{a}^{x}M(t)\,\dd t\), which is \(C^{1}\) with \(A'=M\) (Theorem 7.42). Integrating by parts (Equation (7.28)) and using \(\eta(a)=\eta(b)=0\),
So Equation (16.19) says \(\int_{a}^{b}\left(N-A\right)\eta'\,\dd x=0\) for every admissible \(\eta\), whence \(N-A\) is a constant by Lemma 16.19. Thus \(N=A+\text{const}\) is of class \(C^{1}\) with \(N'=A'=M\).
∎Nothing above requires the competing curves to be smooth: the same three lemmas hold verbatim when \(C^{1}_{0}\) is enlarged to the piecewise-\(C^{1}\) functions vanishing at the endpoints, since the bump Equation (16.17) and the primitive built in Lemma 16.19 already lie in the smaller class. An extremal in the wider class may have corners, and the two conditions that must then hold across a corner are the Weierstrass–Erdmann conditions proved as Corollary 16.33 [Gelfand:1963].
The Euler–Lagrange equation
Let \(f\) be of class \(C^{2}\) and let \(y\in\mathcal{A}\) be a weak local extremum of Equation (16.11). Then \(x\longmapsto\pp f/\pp y'\left(x,y,y'\right)\) is of class \(C^{1}\) and
A solution of Equation (16.20) is called an extremal of \(J\). Rests on Proposition 16.16, Equation (16.15) and Lemma 16.20.
Derives Theorem 16.22. By Proposition 16.16 and Proposition 16.17,
Both coefficients are continuous in \(x\), because \(f\) is \(C^{2}\) and \(y\in C^{1}\). Lemma 16.20 with \(M=\pp f/\pp y\) and \(N=\pp f/\pp y'\) then says that \(N\) is \(C^{1}\) with \(N'=M\), which is Equation (16.20).
∎Two forms of the same statement are worth separating, because they carry different smoothness assumptions.
Under the hypotheses of Theorem 16.22 there is a constant \(C\) with
Rests on Equation (16.20) and Theorem 7.43.
Derives Corollary 16.23. Integrate Equation (16.20) from \(a\) to \(x\), using Theorem 7.43 on the \(C^{1}\) function \(\pp f/\pp y'\).
∎Expanding the total derivative in Equation (16.20) gives the second-order form
which presupposes \(y''\). It is legitimate wherever \(\pp^{2}f/\pp y'^{2}\ne0\): at such a point Equation (16.21) may be solved for \(y'\) as a \(C^{1}\) function of \(x\) by the implicit function theorem (Real Analysis), so \(y\in C^{2}\) there. Where \(\pp^{2}f/\pp y'^{2}\) vanishes the extremal may genuinely have a corner, and only the first-order form Equation (16.21) survives. This is why the du Bois-Reymond argument matters: the classical derivation, which integrates by parts a second time, assumes the conclusion.
Theorem 16.22 is a necessary condition only. An extremal may be a minimum, a maximum, or neither: the extremals of \(\int_{0}^{b}\left(y'^{2}-y^{2}\right)\dd x\) with \(y(0)=y(b)=0\) include \(y\equiv0\), which minimises for \(b<\pi\) and does not for \(b>\pi\), although the equation \(y''+y=0\) is the same in both cases. The extra information is second order and is the subject of Section 16.4; \(b=\pi\) is the first conjugate point in the sense of Section 16.4.2. In physics this is the reason Hamilton's principle is a principle of stationary, not least, action (Section 16.7.1).
Let \(y=\Psi(x,u)\) with \(\Psi\) of class \(C^{2}\) and \(\pp\Psi/\pp u\ne0\), and let
Then the Euler–Lagrange expressions are related by
so \(u\) is an extremal of \(\int F\,\dd x\) if and only if \(y=\Psi(x,u)\) is an extremal of \(\int f\,\dd x\). Rests on Equation (16.20) and Proposition 7.31.
Derives Proposition 16.26. Write \(\Psi_{x}\), \(\Psi_{u}\), \(\Psi_{xu}\), \(\Psi_{uu}\) for the partials of \(\Psi\) and note that \(y'=\Psi_{x}+\Psi_{u}u'\), whence \(\pp y'/\pp u'=\Psi_{u}\) and \(\pp y'/\pp u=\Psi_{xu}+\Psi_{uu}u'=\dd\Psi_{u}/\dd x\), the total derivative along the curve. The chain rule then gives
and therefore
Subtracting, the two \(\pp f/\pp y'\cdot\dd\Psi_{u}/\dd x\) terms cancel and Equation (16.23) remains. Since \(\Psi_{u}\ne0\), one side vanishes exactly when the other does.
∎Equation (16.23) says that the Euler–Lagrange expression transforms as a covector under changes of the dependent variables: the equation is coordinate independent even though its individual terms are not. This is what licenses the use of arbitrary generalized coordinates in Lagrangian Mechanics — polar, spherical, body-fixed — without rederiving the equations of motion each time, and what makes the same formalism usable on the curved configuration spaces of Differentiable Manifolds, Tensors, and Curvature and on the spacetimes of Geometric Formulation of Gravity. Newton's equations enjoy no such property: the components of \(m\ddot{\vect{x}}\) are not the components of anything in curvilinear coordinates.
First integrals and natural boundary conditions
The Euler–Lagrange equation is of second order, and integrating it once is exactly what makes the classical problems soluble. Two circumstances permit it, and both are instances of Noether's theorem (Section 16.6.1): a coordinate absent from the integrand, and the independent variable absent from the integrand.
If \(f\) does not depend on \(y\), that is \(\pp f/\pp y\equiv0\), then along every extremal
Rests on Equation (16.20).
Derives Proposition 16.28. Immediate from Equation (16.20): the first term vanishes, so the total derivative of \(\pp f/\pp y'\) is zero, and a \(C^{1}\) function with vanishing derivative on an interval is constant (Corollary 7.36).
∎If \(f\) does not depend explicitly on \(x\), that is \(\pp f/\pp x\equiv0\), then along every extremal of class \(C^{2}\)
Rests on Equation (16.20) and Proposition 7.31.
Derives Theorem 16.29. Differentiate the left-hand side along the curve, using the chain rule and \(\pp f/\pp x=0\):
the bracket vanishing by Equation (16.20).
∎The quantity \(h=y'\,\pp f/\pp y'-f\) conserved by Equation (16.25) is the Legendre transform of \(f\) with respect to \(y'\). When \(x\) is the time and \(f\) the Lagrangian of a mechanical system, \(h\) is the energy in joules and Equation (16.25) is its conservation (Lagrangian Mechanics); when \(f\) is convex in \(y'\) the same transform produces the Hamiltonian of Hamiltonian Mechanics, and the convexity hypothesis it needs is the strengthened Legendre condition of Section 16.4.1.
So far the endpoints have been held fixed. Releasing one of them costs nothing in the derivation and produces a boundary condition instead of a boundary datum.
Let \(J\) be extremised over \(\mathcal{A}_{b}=\set{y\in C^{1}[a,b]\mid y(a)=y_{a}}\), the value \(y(b)\) being free. Then a weak extremum satisfies Equation (16.20) on \([a,b]\) together with
Rests on Equation (16.15), Theorem 16.22 and Lemma 16.18.
Derives Theorem 16.31. Variations \(\eta\) now need only satisfy \(\eta(a)=0\). Those that also vanish at \(b\) form the class of Theorem 16.22, so Equation (16.20) holds and \(\pp f/\pp y'\) is \(C^{1}\). For a general admissible \(\eta\), integrating Equation (16.15) by parts now keeps the boundary term:
the integral vanishing by Equation (16.20) and the term at \(a\) by \(\eta(a)=0\). Choosing \(\eta\) with \(\eta(b)=1\) gives Equation (16.26).
∎Let the right-hand endpoint be required only to lie on a \(C^{1}\) curve \(y=\psi(x)\), the abscissa \(b\) being free as well. Then at the point \(\left(b,y(b)\right)\) where the extremal meets the curve,
Rests on Theorem 16.31 and Equation (16.20).
Derives Theorem 16.32. Embed the extremal in a one-parameter family of admissible curves \(y_{\varepsilon}\) with terminal abscissa \(b_{\varepsilon}=b+\varepsilon \beta+O(\varepsilon^{2})\) and terminal ordinate \(y_{\varepsilon}(b_{\varepsilon})=\psi(b_{\varepsilon})\), and write \(\eta\) for the vertical variation \(\left.\pp y_{\varepsilon}/\pp\varepsilon\right|_{0}\) at fixed \(x\), so \(\eta(a)=0\). Differentiating \(J_{\varepsilon}=\int_{a}^{b_{\varepsilon}}f\,\dd x\) at \(\varepsilon=0\), and using \(\dd\left(\int_{a}^{b_{\varepsilon}}\right)/\dd\varepsilon =\beta f(b)\) for the moving limit,
the integral collapsing as in Equation (16.27). The endpoint constraint relates \(\beta\) and \(\eta(b)\): differentiating \(y_{\varepsilon}(b_{\varepsilon})=\psi(b_{\varepsilon})\) gives \(\eta(b)+\beta\,y'(b)=\beta\,\psi'(b)\), that is \(\eta(b)=\beta\left(\psi'-y'\right)_{b}\). Substituting and dividing by the arbitrary \(\beta\ne0\) yields Equation (16.28).
∎Let \(y\) be a piecewise-\(C^{1}\) extremal with a corner at an interior point \(x_{0}\), and write \(y'(x_{0}^{-})\) and \(y'(x_{0}^{+})\) for the one-sided slopes. Then both
hold: the momentum \(\pp f/\pp y'\) and the energy \(y'\,\pp f/\pp y'-f\) are continuous across a corner even though the slope is not. Rests on Theorem 16.32 and Corollary 16.23.
Derives Corollary 16.33. The first condition follows from Equation (16.21), which holds on each smooth arc with the same constant \(C\) — the derivation of Lemma 16.20 never used differentiability of \(y\) — so \(\pp f/\pp y'\) equals a continuous function of \(x\) throughout \([a,b]\).
For the second, regard the corner \(\left(x_{0},y_{0}\right)\) as a common free endpoint of the two arcs, and vary it. Repeating the computation of Theorem 16.32 for the left arc, whose terminal point moves by \(\left(\beta,\gamma\right)\) with \(\gamma=\left.\pp y_{\varepsilon}(x_{0})/\pp\varepsilon\right|_{0} +\beta y'(x_{0}^{-})\), contributes \(\beta f|_{x_{0}^{-}} +\left(\pp f/\pp y'\right)_{x_{0}^{-}} \left(\gamma-\beta y'(x_{0}^{-})\right)\) to \(\delta J\); the right arc, for which the same point is an initial point, contributes the same expression at \(x_{0}^{+}\) with the opposite sign. Stationarity for all \(\left(\beta,\gamma\right)\) forces both brackets to match: the coefficient of \(\gamma\) gives the first condition again, and the coefficient of \(\beta\) gives \(\left[f-y'\,\pp f/\pp y'\right]\) continuous, the second.
∎The boundary term \(\left[\left(\pp f/\pp y'\right)\eta\right]_{a}^{b}\) discarded in Theorem 16.22 is not always discardable. In field theory it becomes a surface integral over \(\pp\Omega\) whose vanishing is a boundary condition on the fields (Generalized Classical Field Theory); in general relativity it does not vanish for the Einstein–Hilbert action, and a compensating surface term must be added by hand (The Einstein Field Equations). The elementary case here — a free end of a string forces \(\pp f/\pp y'=0\), which for a taut string is the statement that no transverse force is applied there — is the same phenomenon in one dimension.
Several functions, several variables, higher derivatives
Let \(J[y^{1},\dots,y^{n}] =\int_{a}^{b}f\left(x,y^{1},\dots,y^{n}, y'^{1},\dots,y'^{n}\right)\dd x\) with \(f\) of class \(C^{2}\), over \(n\)-tuples of \(C^{1}\) functions with all \(2n\) endpoint values fixed. At a weak extremum,
Rests on Theorem 16.22 and Lemma 16.20.
Derives Theorem 16.35. Fix \(i\) and vary only \(y^{i}\), holding the other \(n-1\) functions fixed: \(J\) restricted to that one-parameter family is a functional of the type of Equation (16.11), and Theorem 16.22 gives the \(i\)-th equation. Repeating for each \(i\) gives the system, and no further information is available from simultaneous variations, since the first variation is linear in \(\left(\eta^{1},\dots,\eta^{n}\right)\).
∎Let \(\Omega\subset\R^{D}\) be a bounded domain with piecewise-\(C^{1}\) boundary, and let
be extremised over fields \(\varphi^{A}\in C^{2}(\overline{\Omega})\) with prescribed values on \(\pp\Omega\). Then at a weak extremum
with the repeated index \(\mu\) summed. Rests on Equation (16.31), Theorem 7.101 and Lemma 16.18.
Derives Theorem 16.36. Let \(\eta^{A}\in C^{1}(\overline{\Omega})\) vanish on \(\pp\Omega\) and put \(\varphi_{\varepsilon}=\varphi+\varepsilon\eta\). Differentiating under the integral sign as in Proposition 16.17,
Write \(\pi_{A}^{\mu}=\pp\Lag/\pp(\pp_{\mu}\varphi^{A})\), which is \(C^{1}\) because \(\Lag\) is \(C^{2}\) and \(\varphi\in C^{2}\). Then \(\pi_{A}^{\mu}\,\pp_{\mu}\eta^{A} =\pp_{\mu}\left(\pi_{A}^{\mu}\eta^{A}\right) -\eta^{A}\,\pp_{\mu}\pi_{A}^{\mu}\), and the divergence theorem (Theorem 7.101) turns the first piece into a surface integral over \(\pp\Omega\), where \(\eta^{A}=0\). Hence
Stationarity makes this vanish for every such \(\eta\). The bracket is continuous, and the multidimensional form of Lemma 16.18 — proved by the same argument, with the bump Equation (16.17) replaced by its product over the \(D\) coordinates on a small cube inside \(\Omega\) — forces it to vanish identically, one component \(A\) at a time.
∎Take \(D=3\), one field \(u\), and \(\Lag=\tfrac12\abs{\nabla u}^{2}\), so that \(S[u]=\tfrac12\int_{\Omega}\abs{\nabla u}^{2}\dd^{3}x\) is the Dirichlet integral. Then \(\pp\Lag/\pp u=0\) and \(\pp\Lag/\pp(\pp_{\mu}u)=\pp_{\mu}u\), so Equation (16.32) reads
Laplace's equation of Partial Differential Equations. Multiplied by the permittivity of free space, \(\varepsilon_{0} =8.8541878128\times 10^{-12}\,\mathrm{F}/\mathrm{m}\), the same functional is the electrostatic field energy \(U=\tfrac{\varepsilon_{0}}{2}\int\abs{\nabla\Phi}^{2}\dd^{3}x\) in joules, and the statement becomes the one of Electrostatics: the electrostatic potential is the field of least energy compatible with the boundary data. This is the Dirichlet principle, whose existence gap is the subject of Section 16.5.1. Multiple-integral problems of this kind are the systematic subject of Courant–Hilbert [Courant:1962]. Rests on Equation (16.32).
Let \(J[y]=\int_{a}^{b} f\left(x,y,y',\dots,y^{(m)}\right)\dd x\) with \(f\) of class \(C^{m+1}\), be extremised over \(y\in C^{2m}[a,b]\) with \(y,y',\dots,y^{(m-1)}\) prescribed at both endpoints. Then
Rests on Equation (16.15), Lemma 16.18 and Corollary 7.44.
Derives Theorem 16.38. Admissible variations \(\eta\) satisfy \(\eta^{(j)}(a)=\eta^{(j)}(b)=0\) for \(j=0,\dots,m-1\). As in Proposition 16.17,
Integrate the \(k\)-th term by parts \(k\) times (Equation (7.28)). Each integration produces a boundary term carrying a factor \(\eta^{(j)}\) with \(j\le k-1\le m-1\), all of which vanish, and contributes one factor of \(-1\); after \(k\) steps the term is \(\left(-1\right)^{k}\int\eta\,\dd^{k} \left(\pp f/\pp y^{(k)}\right)/\dd x^{k}\). Summing over \(k\),
and Lemma 16.18 gives Equation (16.35).
∎Let \(m=2\), write \(q\) for the dependent variable and \(t\) for the independent one, and suppose the Lagrangian \(\Lag\left(q,\dot q,\ddot q\right)\) is nondegenerate, meaning \(\pp^{2}\Lag/\pp\ddot q^{2}\ne0\), so that \(\ddot q\) may be solved for in terms of \(\left(q,\dot q,\pp\Lag/\pp\ddot q\right)\). Then the conserved energy associated with time translation is linear in one of its momenta and is therefore unbounded below. Rests on Equation (16.35) and Theorem 16.29.
Derives Proposition 16.39. Introduce \(Q_{1}=q\), \(Q_{2}=\dot q\) and the Ostrogradsky momenta
Nondegeneracy lets \(P_{2}=\pp\Lag/\pp\ddot q\) be inverted for \(\ddot q=\mathcal{F}\left(Q_{1},Q_{2},P_{2}\right)\). The Beltrami constant of Theorem 16.29, extended to the higher-derivative case by the same computation, is
Only the first term contains \(P_{1}\), and it contains it linearly, so for fixed \(\left(Q_{1},Q_{2},P_{2}\right)\) with \(Q_{2}\ne0\) the value of \(\Ham\) runs over all of \(\R\) as \(P_{1}\) does.
∎Proposition 16.39 is Ostrogradsky's theorem [Ostrogradsky:1850]. Its physical content is that a nondegenerate theory whose Lagrangian contains second time derivatives has no stable ground state: the system can lower \(\Ham\) without bound by exciting the \(P_{1}\) direction, and once it is coupled to anything else the vacuum decays. This — not aesthetics — is why the Lagrangians of Lagrangian Mechanics, Generalized Classical Field Theory and The Einstein Field Equations are first order in time derivatives. The Einstein–Hilbert action is the instructive exception: it does contain second derivatives of the metric, but degenerately — they enter linearly and can be removed by a total divergence, which is exactly the boundary term of Remark 16.34 — so Ostrogradsky's hypothesis fails and general relativity has second-order field equations.
Constrained variation
Lagrange multipliers
Constraints come in two kinds, and they need different multipliers. A pointwise constraint must hold at every \(x\), and buys one multiplier function; an integral constraint holds once for the whole curve, and buys one multiplier constant. The rule for both descends from the finite-dimensional multiplier theorem of Real Analysis, of which the following is the functional analogue [Lagrange:1788].
Let \(J[y]=\int_{a}^{b}f\left(x,y^{i},y'^{i}\right)\dd x\) with \(i=1,\dots,n\), and let the competing curves be required to satisfy the \(m<n\) pointwise constraints
with \(g^{\alpha}\) of class \(C^{2}\) and the \(m\times n\) matrix \(\pp g^{\alpha}/\pp y^{i}\) of full rank \(m\) along the curve. If \(y\) is a weak extremum of \(J\) within the constrained class, there are functions \(\lambda_{\alpha}\in C^{0}[a,b]\) such that, with \(F=f+\lambda_{\alpha}g^{\alpha}\),
Rests on Equation (16.37), Theorem 16.35 and Lemma 16.18.
Derives Theorem 16.41. Admissible variations \(\eta^{i}\in C^{1}_{0}[a,b]\) are no longer arbitrary: differentiating Equation (16.37) along the family \(y+\varepsilon\eta\) at \(\varepsilon=0\) gives the \(m\) pointwise linear conditions
Stationarity gives, after the integration by parts of Theorem 16.22,
for every \(\eta\) obeying Equation (16.39). By the rank hypothesis the indices may be relabelled so that the leading \(m\times m\) block \(\pp g^{\alpha}/\pp y^{\beta}\), \(\beta=1,\dots,m\), is invertible at every \(x\); call \(y^{1},\dots,y^{m}\) dependent and \(y^{m+1},\dots,y^{n}\) free. Define \(\lambda_{\alpha}(x)\) as the unique solution of the linear system
which is continuous in \(x\) because the block and the \(E_{\beta}\) are. With this choice the first \(m\) of Equation (16.38) hold by construction, since \(g^{\alpha}\) does not contain \(y'^{i}\) and hence \(\pp F/\pp y'^{i}=\pp f/\pp y'^{i}\). Substituting back into Equation (16.40) and using Equation (16.39) to eliminate the dependent \(\eta^{\beta}\),
now for arbitrary \(\eta^{k}\in C^{1}_{0}[a,b]\), since the dependent components can always be recovered from Equation (16.39). Lemma 16.18, applied one free index at a time, gives the remaining \(n-m\) equations.
∎Comparing Equation (16.38) with the unconstrained equations, the effect of the constraint is the extra term \(\lambda_{\alpha}\,\pp g^{\alpha}/\pp y^{i}\) on the right-hand side — a force normal to the constraint surface, of magnitude fixed by \(\lambda_{\alpha}\). In mechanics this is exactly the reaction of the constraint, and \(\lambda_{\alpha}\) carries whatever SI unit makes \(\lambda_{\alpha}g^{\alpha}\) an energy: for a constraint written as a length, \(\lambda\) is a force in newtons. That reading is developed, with the Lagrange equations of the first kind, in Lagrangian Mechanics. It is also the practical reason to keep the multipliers rather than eliminate the constraint: the numbers one wants — the tension in the rod, the normal force on the track — are the multipliers themselves.
Isoperimetric constraints
Let \(J[y]=\int_{a}^{b}f\left(x,y,y'\right)\dd x\) and \(K[y]=\int_{a}^{b}g\left(x,y,y'\right)\dd x\) with \(f,g\) of class \(C^{2}\), and let \(y\in\mathcal{A}\) be a weak extremum of \(J\) within the class \(\set{\tilde y\in\mathcal{A}\mid K[\tilde y]=\ell}\). If \(y\) is not an extremal of \(K\) alone, there is a constant \(\lambda\in\R\) such that \(y\) is an unconstrained extremal of \(\int_{a}^{b}\left(f+\lambda g\right)\dd x\), that is
Rests on Definition 16.15, Proposition 16.17 and Lemma 16.20.
Derives Theorem 16.43. Because \(y\) is not an extremal of \(K\), Lemma 16.20 supplies some \(\eta_{2}\in C^{1}_{0}[a,b]\) with \(\delta K[y;\eta_{2}]\ne0\). Fix an arbitrary \(\eta_{1}\in C^{1}_{0}[a,b]\) and consider the two-parameter family \(y_{\varepsilon}=y+\varepsilon_{1}\eta_{1}+\varepsilon_{2}\eta_{2}\), setting
Both are \(C^{1}\) near the origin, with \(\pp\Phi/\pp\varepsilon_{i}\big|_{0}=\delta J[y;\eta_{i}]\) and similarly for \(\Psi\), by the argument of Proposition 16.17. Since \(\Psi(0,0)=\ell\) and \(\pp\Psi/\pp\varepsilon_{2}\big|_{0}=\delta K[y;\eta_{2}]\ne0\), the implicit function theorem (Real Analysis) provides a \(C^{1}\) function \(\varepsilon_{2}\left(\varepsilon_{1}\right)\) with \(\varepsilon_{2}(0)=0\) and \(\Psi\left(\varepsilon_{1},\varepsilon_{2}(\varepsilon_{1})\right) =\ell\) for small \(\varepsilon_{1}\). Every curve of this one-parameter family is admissible and lies in a weak neighbourhood of \(y\), so \(\varepsilon_{1}\longmapsto \Phi\left(\varepsilon_{1},\varepsilon_{2}(\varepsilon_{1})\right)\) has an extremum at \(0\):
the second identity coming from differentiating the constraint. Setting \(\lambda=-\delta J[y;\eta_{2}]/\delta K[y;\eta_{2}]\), a constant independent of \(\eta_{1}\), the display becomes \(\delta J[y;\eta_{1}]+\lambda\,\delta K[y;\eta_{1}]=0\), that is \(\delta\left(J+\lambda K\right)[y;\eta_{1}]=0\) for every \(\eta_{1}\in C^{1}_{0}[a,b]\). Lemma 16.20 then yields Equation (16.41).
∎If \(y\) is an extremal of \(K\) alone, the conclusion can fail, and the correct statement is the abnormal one: there are constants \(\left(\lambda_{0},\lambda\right)\ne(0,0)\) with \(y\) an extremal of \(\lambda_{0}f+\lambda g\), and \(\lambda_{0}\) may be zero. The finite-dimensional shadow of this is the failure of the multiplier rule at a critical point of the constraint function, and it is the same phenomenon: the constraint set has no tangent direction along which \(J\) can be tested.
Proposition 16.6 and Proposition 16.4 are both applications of Theorem 16.43, with \(\lambda\) playing two different physical roles: for Dido's problem \(\abs{\lambda}\) is the radius of the answer, and for the hanging chain \(\rho_{\ell}g\lambda\) is the horizontal tension in the chain, in newtons. Both illustrate Remark 16.42: the multiplier is not bookkeeping, it is the physically interesting number. Rests on Theorem 16.43.
Take \(f=p\,y'^{2}+q\,y^{2}\) and \(g=r\,y^{2}\) with the normalisation \(K[y]=\int_{a}^{b}r\,y^{2}\,\dd x=1\). Then Equation (16.41) reads
which is the Sturm–Liouville eigenvalue problem Equation (9.55) of Ordinary Differential Equations and Sturm–Liouville Theory with eigenvalue \(-\lambda\). The multiplier of the normalisation constraint is the eigenvalue. This is the identification the Rayleigh–Ritz method of Section 16.5.3 exploits, and it is why a variational estimate of a ground-state energy is available at all.
Holonomic and nonholonomic constraints
A constraint on a curve \(x\longmapsto\left(y^{1},\dots,y^{n}\right)\) is holonomic if it can be written as a finite set of relations \(g^{\alpha}\left(x,y^{1},\dots,y^{n}\right)=0\) among the coordinates and the independent variable alone, as in Equation (16.37). A constraint given in Pfaffian form,
is holonomic if the system Equation (16.43) is integrable — if there are functions \(g^{\alpha}\) and an invertible matrix \(\mu^{\alpha}{}_{\beta}\) with \(\mu^{\alpha}{}_{\beta}\left(A^{\beta}{}_{i}\dd y^{i} +A^{\beta}{}_{0}\dd x\right)=\dd g^{\alpha}\) — and nonholonomic otherwise. Integrability is decided by the Frobenius condition of Differentiable Manifolds, Tensors, and Curvature: the ideal generated by the one-forms must be closed under exterior differentiation. Rests on Equation (16.37) and Definition 21.2.
A thin disc of radius \(R\) rolls without slipping on a horizontal plane, staying vertical. Its configuration is \(\left(x,y,\varphi,\vartheta\right)\): contact point, heading angle and rotation angle. Rolling without slipping means the contact point has zero velocity, which in SI reads
with \(R\) in metres and the angles in radians. Neither one-form is integrable, alone or together: the disc can be rolled around a closed loop in the plane and returned to its starting point with \(\vartheta\) changed by an arbitrary amount, so no relation \(g\left(x,y,\varphi,\vartheta\right)=0\) can hold. The constraint removes two velocity directions but no configuration, and the accessible set is all four dimensions. This is the standard nonholonomic example [Goldstein:2002]; the dynamics belongs to Lagrangian Mechanics and Rigid Bodies and Rotating Frames. Rests on Definition 16.47.
Theorem 16.41 was proved for holonomic constraints, and the restriction is essential. Two inequivalent prescriptions exist for a Pfaffian constraint that is not integrable, and it is a matter of fact, not of convention, which describes a rolling body.
The vakonomic prescription applies the theorem literally: it restricts the class of comparison curves to those satisfying Equation (16.43) and makes the action stationary within that class. The d'Alembert–Chetaev prescription, which is the one mechanics uses, does not restrict the comparison curves at all; it restricts the virtual displacements to the annihilator of the constraint forms, \(A^{\alpha}{}_{i}\,\delta y^{i}=0\), and requires the constraint forces to do no virtual work. For a holonomic constraint the two prescriptions coincide, because a curve tangent to an integrable distribution stays on the leaf. For a nonholonomic one they do not: the vakonomic equations carry extra terms in the derivatives of the multipliers along the curve, and the two sets of trajectories differ already for the rolling disc of Example 16.48.
The observed motion of rolling bodies is the d'Alembert–Chetaev one; that is the content of the principle of virtual work as it is set up in Lagrangian Mechanics, and it is what the experiments on rolling and spinning rigid bodies in Rigid Bodies and Rotating Frames test. The vakonomic equations are not wrong — they are the correct equations of a different problem, namely the optimal-control problem in which the constraint is imposed on the admissible curves by fiat — but they are not the equations of a rolling disc. Degenerate constraint systems, where the constraints follow from the Lagrangian itself rather than being imposed, are a third matter, treated in Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism.
The second variation
The Legendre condition
Stationarity is a first-order condition and says nothing about whether an extremal minimises, maximises or does neither (Remark 16.25). The second-order information is carried by the second variation, which is a quadratic functional of the variation.
Let \(f\) be of class \(C^{3}\) and \(y\) an extremal of class \(C^{2}\). For \(\eta\in C^{1}_{0}[a,b]\),
where, with all partial derivatives evaluated along the extremal,
Rests on Equation (16.14), Theorem 7.38 and Corollary 7.44.
Derives Proposition 16.50. Differentiating twice under the integral sign, as in Proposition 16.17,
The cross term is a total derivative up to one integration by parts: \(2\eta\eta'=\left(\eta^{2}\right)'\), so by Equation (7.28) and \(\eta(a)=\eta(b)=0\),
which absorbed into the first term gives Equation (16.45) with the coefficients Equation (16.46).
∎If the extremal \(y\) is a weak local minimum of \(J\), then
Rests on Equation (16.45) and Proposition 16.16.
Derives Theorem 16.51. At a weak minimum \(\varepsilon\longmapsto J[y+\varepsilon\eta]\) has a minimum at \(\varepsilon=0\), so \(\delta^{2}J[y;\eta]\ge0\) for every \(\eta\in C^{1}_{0}[a,b]\). Suppose instead that \(P(x_{0})<0\) at some \(x_{0}\), which by continuity may be taken interior, and choose \(\delta>0\) and \(k>0\) with \([x_{0}-\delta,x_{0}+\delta]\subset(a,b)\) and \(P\le-k\) there. Take the localized variation
which lies in \(C^{1}_{0}[a,b]\) because \(\eta\) and \(\eta'\) both vanish at \(x_{0}\pm\delta\). Substituting \(t=(x-x_{0}+\delta)/2\delta\),
Hence, writing \(M=\max_{[a,b]}\abs{Q}\),
which is negative for \(\delta\) small enough — a contradiction. So \(P\ge0\) everywhere.
∎The strengthened Legendre condition is \(P>0\) on \([a,b]\); the problem is then called regular. Legendre's memoir on distinguishing maxima from minima [Legendre:1788] introduced the condition and attempted to prove that it is also sufficient, by completing the square in Equation (16.45) with the help of an auxiliary function. The step is correct wherever the auxiliary function exists, and Legendre did not notice that it may fail to exist on the whole interval. The scaling in Equation (16.48) shows why the \(P\) term dominates and hence why \(P\) alone controls the local behaviour: a variation squeezed into a short interval has a large \(\int\eta'^{2}\) and a small \(\int\eta^{2}\). The gap Legendre left is exactly the possibility that a variation spread over the whole interval defeats a positive \(P\), and closing it is Jacobi's work (Section 16.4.2). In physics \(P>0\) is the invertibility of the Legendre transform, that is, of the map from velocities to momenta (Hamiltonian Mechanics).
The Jacobi equation and conjugate points
The functional Equation (16.45) is itself a variational problem, and its own Euler–Lagrange equation is the key to the whole second-order theory [Jacobi:1837].
The Jacobi accessory equation of the extremal \(y\) is the Euler–Lagrange equation of the quadratic functional Equation (16.45),
A point \(c\in(a,b]\) is conjugate to \(a\) along \(y\) if the solution \(u\) of Equation (16.49) with \(u(a)=0\), \(u'(a)=1\) satisfies \(u(c)=0\). Rests on Equations (16.20) and (16.45).
Applying Equation (16.20) to the integrand \(P u'^{2}+Q u^{2}\) gives \(2\dd(Pu')/\dd x-2Qu=0\), which is Equation (16.49). Geometrically the equation is the linearisation of Equation (16.20) about the extremal: if \(y(x,\alpha)\) is a one-parameter family of extremals through the same initial point, then \(u=\pp y/\pp\alpha\) solves Equation (16.49). A conjugate point is therefore a place where neighbouring extremals leaving \(a\) come back together to first order — a focus. For geodesics on a sphere the conjugate point of any point is its antipode, which is why a great-circle arc longer than half the circumference is not the shortest path; in optics the same focusing is a caustic (Hamilton–Jacobi Theory and the Optical–Mechanical Analogy), and in mechanics a kinetic focus.
Let the problem be regular, \(P>0\) on \([a,b]\), and let the extremal \(y\) be a weak local minimum. Then no point of the open interval \((a,b)\) is conjugate to \(a\). Rests on Definition 16.53, Theorem 16.51 and Corollary 16.33.
Derives Theorem 16.55. Suppose \(c\in(a,b)\) is conjugate to \(a\), with \(u\) the solution of Equation (16.49) having \(u(a)=u(c)=0\) and \(u'(a)=1\); then \(u\) does not vanish identically, and \(u'(c)\ne0\), since a second-order linear solution vanishing together with its derivative at a point vanishes identically. Define the broken variation
which is continuous and piecewise \(C^{1}\) with \(\eta(a)=\eta(b)=0\). Multiplying Equation (16.49) by \(u\) and integrating by parts over \([a,c]\),
because \(u\) vanishes at both ends. Hence \(\delta^{2}J[y;\eta]=0\).
Because \(y\) is a weak minimum, \(\delta^{2}J[y;\cdot]\ge0\) on \(C^{1}_{0}[a,b]\), and hence on the larger class of piecewise-\(C^{1}\) variations vanishing at the endpoints: a corner may be rounded off inside an interval of length \(\varepsilon\), which changes \(\int\left(P\eta'^{2}+Q\eta^{2}\right)\dd x\) by \(O(\varepsilon)\), so the quadratic form is nonnegative on the closure by continuity. So \(\eta\) minimises \(\delta^{2}J\) over the wider class. A minimiser must satisfy the Euler–Lagrange equation Equation (16.49) on each smooth arc, and by the first Weierstrass–Erdmann condition of Corollary 16.33 — here the continuity of \(\pp/\pp\eta'\left(P\eta'^{2}+Q\eta^{2}\right) =2P\eta'\) — the product \(P\eta'\) must be continuous at \(c\). But \(P(c)\eta'(c^{-})=P(c)u'(c)\ne0\) while \(P(c)\eta'(c^{+})=0\), a contradiction.
∎Let \(P>0\) on \([a,b]\) and suppose the Jacobi equation Equation (16.49) has a solution \(w\) with no zero on \([a,b]\). Then \(\delta^{2}J[y;\eta]>0\) for every \(\eta\in C^{1}_{0}[a,b]\) that does not vanish identically. Rests on Equations (16.45) and (16.49).
Derives Theorem 16.56. Set \(W=-P\,w'/w\), which is \(C^{1}\) on \([a,b]\) because \(w\) has no zero there. Differentiating and using Equation (16.49) for \(w\),
For any \(\eta\in C^{1}_{0}[a,b]\) the quantity \(W\eta^{2}\) vanishes at both endpoints, so \(\int_{a}^{b}\left(W\eta^{2}\right)'\dd x=0\) and this may be added to Equation (16.45) free of charge:
the completion of the square being exact precisely because Equation (16.50) gives \(Q+W'=W^{2}/P\). The integrand is nonnegative, so \(\delta^{2}J\ge0\), and it vanishes only if \(\eta'=-W\eta/P\) throughout. That linear equation with \(\eta(a)=0\) has only the zero solution, so \(\delta^{2}J[y;\eta]>0\) unless \(\eta\equiv0\).
∎The hypothesis of Theorem 16.56 — a Jacobi solution without zeros on the closed interval — is the strengthened Jacobi condition, and it is equivalent, for a regular problem, to the absence of any point conjugate to \(a\) in \(\left(a,b\right]\): a solution vanishing at \(a-\varepsilon\) for small \(\varepsilon>0\) has, by continuous dependence of solutions on their initial data, no zero in \([a,b]\) exactly when the solution vanishing at \(a\) has none in \((a,b]\). The equivalence is proved in The Strengthened Jacobi Condition and Conjugate Points, where the continuous dependence on the initial point is derived from the integral form of the initial value problem by a Grönwall estimate and the Sturm separation theorem from the Lagrange identity Equation (9.56); with it, the necessary condition Theorem 16.55 and the sufficient condition Theorem 16.56 test the same hypothesis, and differ only at the single point \(x=b\).
Full derivation in Appendix A.
Derives Remark 16.57.
Conjugate points are not merely an obstruction to be avoided; their number is an invariant. Morse's index theorem states that the index of the second variation — the maximal dimension of a subspace on which \(\delta^{2}J\) is negative definite — equals the number of points conjugate to \(a\) in \((a,b)\), counted with multiplicity [Morse:1934]. The theorem is quoted here and not used; it returns for geodesics in Geometric Formulation of Gravity, where it converts a local calculation into a statement about the topology of the space of paths.
Strong variations and the Weierstrass condition
Everything in Sections 16.4.1 and 16.4.2 tests the extremal against curves whose slopes are close to its own, and Example 16.14 shows that this is not enough. A strong variation changes the slope by a finite amount over a short interval, and the quantity that measures its effect is Weierstrass's excess function [Weierstrass:1927].
For an integrand \(f\) of class \(C^{2}\),
It is the amount by which \(f\), as a function of the slope alone at fixed \((x,y)\), lies above its tangent line at \(p\): the excess of the value at \(q\) over the linear extrapolation from \(p\). Rests on Definition 16.11.
Let \(y\) be a strong local minimum of \(J\). Then
Rests on Equation (16.51) and Definition 16.13.
Derives Theorem 16.60. Fix an interior \(x_{0}\), write \(y_{0}=y(x_{0})\), fix \(q\in\R\) and fix \(\Delta>0\) small enough that \([x_{0},x_{0}+\Delta]\subset(a,b)\). For \(0\le h<\Delta\) define the comparison curve \(Y_{h}\): equal to \(y\) outside \(\left(x_{0},x_{0}+\Delta\right)\); the segment of slope \(q\) from \(\left(x_{0},y_{0}\right)\) on \([x_{0},x_{0}+h]\); and on \([x_{0}+h,x_{0}+\Delta]\) the straight segment rejoining \(\left(x_{0}+\Delta,y(x_{0}+\Delta)\right)\), whose slope is
Each \(Y_{h}\) is admissible and piecewise \(C^{1}\), and \(\norm{Y_{h}-y}_{0}\longrightarrow0\) as \(h\to0\) at fixed \(\Delta\), so for small \(h\) it lies in the strong neighbourhood on which \(y\) minimises. Hence \(G(h)\equiv J[Y_{h}]-J[y]\) satisfies \(G(0)=0\) and \(G(h)\ge0\), so \(G'(0^{+})\ge0\).
Only the interval \(\left(x_{0},x_{0}+\Delta\right)\) contributes. Writing \(\ell_{h}(x)=y_{0}+qh+s_{h}\left(x-x_{0}-h\right)\) for the rejoining segment,
Differentiating at \(h=0\), where \(\ell_{0}(x)=y_{0}+s_{\Delta}(x-x_{0})\), the two moving limits give \(f\left(x_{0},y_{0},q\right)-f\left(x_{0},y_{0},s_{\Delta}\right)\), while
so that
the partial derivatives being evaluated at \(\left(x,\ell_{0}(x),s_{\Delta}\right)\). Now let \(\Delta\to0\). The first part of the integral is \(O(\Delta)\); the second tends to \(\pp f/\pp y'\left(x_{0},y_{0},y'(x_{0})\right)\), since \(\Delta^{-1}\int_{x_{0}}^{x_{0}+\Delta}\) of a continuous function tends to its value at \(x_{0}\) and \(s_{\Delta}\to y'(x_{0})\) by the mean value theorem (Theorem 7.35). Writing \(p=y'(x_{0})\), the limit of Equation (16.53) is \(f(x_{0},y_{0},q)-f(x_{0},y_{0},p)-(q-p)\,\pp f/\pp y'(x_{0},y_{0},p) =E\left(x_{0},y_{0},p,q\right)\). Since \(G'(0^{+})\ge0\) for every \(\Delta\), the limit is nonnegative, which is Equation (16.52) at \(x_{0}\); continuity extends it to the endpoints.
∎By Taylor's theorem with Lagrange remainder (Theorem 7.38) applied to \(q\longmapsto f\left(x,y,q\right)\),
for some \(\bar q\) between \(p\) and \(q\). Hence \(E\ge0\) for all \(q\) if \(f\) is convex in the slope, and taking \(q\to p\) in Equation (16.54) recovers Legendre's condition \(\pp^{2}f/\pp y'^{2}\ge0\): Weierstrass' condition contains Legendre's, and the two agree exactly when the slope is varied infinitesimally. This is also the hypothesis under which the Legendre transform to the Hamiltonian formalism of Hamiltonian Mechanics is invertible, and under which the convexity requirement of Tonelli's existence theorem (Section 16.5.2) holds. Three separate questions — strong minimality, the passage to the Hamiltonian, and existence — turn on the same convexity of \(f\) in \(y'\).
Fields of extremals and sufficiency
All the conditions so far are necessary. Turning them into a sufficient condition needs one further construction: the extremal must be embedded in a family covering a whole neighbourhood of it.
A field of extremals on a simply connected region \(R\subset\R^{2}\) is a one-parameter family of extremals of \(J\) such that exactly one member passes through each point of \(R\), depending \(C^{1}\) on the point. Its slope function \(p:R\longrightarrow\R\) assigns to each \(\left(x,y\right)\in R\) the slope at that point of the unique member through it. An extremal is embedded in the field if its graph lies in \(R\). Rests on Theorem 16.22 and Definition 6.17.
Let \(p\) be the slope function of a field of extremals on \(R\). Then the line integral
depends only on the endpoints of the curve \(\Gamma\subset R\), not on the curve. Along any member of the field itself, \(I=\int f\left(x,y,y'\right)\dd x=J\). Rests on Definition 16.62, Proposition 7.104 and Equation (16.20).
Derives Theorem 16.63. Write the integrand as \(\mathcal{P}\,\dd x+\mathcal{Q}\,\dd y\) with
both \(C^{1}\) on \(R\). Path independence on a simply connected region is equivalent to \(\pp\mathcal{P}/\pp y=\pp\mathcal{Q}/\pp x\) (Proposition 7.104). Abbreviating \(f_{y}\), \(f_{y'}\), \(f_{xy'}\), \(f_{yy'}\), \(f_{y'y'}\) for the partial derivatives evaluated at \(\left(x,y,p(x,y)\right)\), the chain rule gives
Each member of the field satisfies Equation (16.20), and along it \(\dd/\dd x=\pp_{x}+p\,\pp_{y}+\left(\pp p/\pp x +p\,\pp p/\pp y\right)\pp_{y'}\), so the Euler–Lagrange equation reads
Subtracting the two displays,
by Equation (16.56). Finally, on a member of the field \(\dd y=p\,\dd x\), so the two \(f_{y'}\) terms in Equation (16.55) cancel and the integrand reduces to \(f\left(x,y,p\right)\dd x=f\left(x,y,y'\right)\dd x\).
∎Let the extremal \(y^{*}\) be embedded in a field of extremals on \(R\) with slope function \(p\), joining \(\left(a,y_{a}\right)\) to \(\left(b,y_{b}\right)\), and suppose
Then \(J[Y]\ge J[y^{*}]\) for every admissible piecewise-\(C^{1}\) curve \(Y\) whose graph lies in \(R\): \(y^{*}\) is a strong minimum relative to \(R\). Rests on Theorem 16.63, Equation (16.51) and Definition 16.62.
Derives Theorem 16.64. Let \(Y\) be admissible with graph in \(R\). By Theorem 16.63, \(I\) takes the same value on the graph of \(Y\) as on the graph of \(y^{*}\), and on the latter it equals \(J[y^{*}]\). Evaluating \(I\) along the graph of \(Y\), where \(\dd y=Y'\,\dd x\),
with \(p=p\left(x,Y(x)\right)\). Subtracting from \(J[Y]=\int_{a}^{b}f\left(x,Y,Y'\right)\dd x\) and comparing with Equation (16.51),
by Equation (16.57).
∎Let \(y^{*}\) be an extremal satisfying the strengthened Legendre condition \(P>0\) on \([a,b]\) and the strengthened Jacobi condition of Remark 16.57. Then \(y^{*}\) is a strict weak local minimum of \(J\). Rests on Theorems 16.51 and 16.56.
Derives Corollary 16.65. Theorem 16.56 gives \(\delta^{2}J[y^{*};\eta]>0\) for every admissible \(\eta\ne0\), and that by itself is not enough: the infimum of \(\delta^{2}J[y^{*};\cdot]\) over the sphere \(\norm{\eta}_{1}=1\) is zero. A smoothed tent of width \(\delta\) and unit maximum slope near an interior point has \(\delta^{2}J=O(\delta)\) while its \(C^{1}\) norm tends to \(1\), so no positive lower bound survives the passage to the limit and the positivity must be upgraded to coercivity first.
Step 1: a coercive lower bound. Let \(w\) be the zero-free Jacobi solution supplied by the strengthened Jacobi condition, and let \(W=-P\,w'/w\) and \(\mu(x)=\exp\int_{a}^{x}\left(W/P\right)\dd s\), so that \(\mu>0\) is \(C^{1}\) on \([a,b]\). The completion of the square in the proof of Theorem 16.56 is the exact identity
valid for every \(\eta\in C^{1}_{0}[a,b]\). Since \(\eta(a)=0\) we have \(\mu(x)\eta(x)=\int_{a}^{x}\mu\,v\,\dd s\), so the Cauchy–Schwarz inequality bounds \(\max\abs{\eta}\) by a constant depending only on \(\mu\) and \(b-a\) times \(\left(\int_{a}^{b}v^{2}\right)^{1/2}\); hence \(\int\eta^{2}\le\kappa_{2}\int v^{2}\), and \(\eta'=v-W\eta/P\) then gives \(\int\eta'^{2}\le\kappa_{3}\int v^{2}\) with \(\kappa_{3}\) likewise independent of \(\eta\). With \(P\ge P_{\min}>0\) this is the coercivity estimate
Step 2: the remainder. For \(\eta\in C^{1}_{0}[a,b]\), Taylor's theorem (Theorem 7.38) applied to \(f\) in its last two arguments, the vanishing of the first variation, and the integration by parts of Proposition 16.50 give
where \(\omega(t)\to0\) as \(t\to0\): the second derivatives of \(f\) are uniformly continuous on a compact neighbourhood of the graph of \(y^{*}\) because \(f\) is \(C^{3}\), and \(2\abs{\eta\eta'}\le\eta^{2}+\eta'^{2}\). This is the step the norm on the sphere could not supply — the remainder is now measured against \(\int(\eta^{2}+\eta'^{2})\), the same quantity Equation (16.59) controls from below.
Step 3. Because \(\eta(a)=0\), Cauchy–Schwarz also gives \(\int\eta^{2}\le\left(b-a\right)^{2}\int\eta'^{2}\). Writing \(\kappa=1+\left(b-a\right)^{2}\) and combining the two steps,
Choose \(\varepsilon>0\) with \(\kappa\,\omega(\varepsilon)<c/2\). For every \(\eta\ne0\) with \(\norm{\eta}_{1}<\varepsilon\) the bracket is positive and \(\int\eta'^{2}>0\), since an \(\eta\in C^{1}_{0}[a,b]\) with \(\eta'\equiv0\) vanishes identically. So \(y^{*}\) is a strict weak local minimum.
∎The slope function of a field is a gradient: on the region \(R\) the invariant integral Equation (16.55) defines a function \(S\left(x,y\right)\), unique up to a constant, with \(\pp S/\pp y=\pp f/\pp y'\left(x,y,p\right)\) and \(\pp S/\pp x=f-p\,\pp f/\pp y'\). Those two statements together are the Hamilton–Jacobi equation, and \(S\) is Hamilton's characteristic function (Hamilton–Jacobi Theory and the Optical–Mechanical Analogy). The Weierstrass sufficiency theory and Hamilton–Jacobi theory are therefore the same construction read in two directions: the physicist solves a first-order partial differential equation to generate the trajectories, the analyst builds the family of trajectories to certify a minimum. Hilbert announced the invariant integral in the lecture that also posed his twenty-third problem, a call for exactly this kind of systematic development of the calculus of variations [Hilbert:1900].
Direct methods and existence
Everything so far has assumed that a minimiser exists and asked what equations it must satisfy. The direct methods reverse the order: they establish that a minimiser exists before writing any differential equation, and then read the Euler–Lagrange equation off it. The nineteenth century learned the hard way that the first question cannot be skipped.
The Dirichlet principle and its failure
Let \(\Omega\subset\R^{3}\) be a bounded domain and \(h\) a continuous function on \(\pp\Omega\). The Dirichlet principle is the assertion that the Dirichlet integral
attains its infimum over \(\set{u\in C^{1}(\overline{\Omega})\mid u|_{\pp\Omega}=h}\), and that the minimiser is the harmonic function with boundary values \(h\). Rests on Equations (16.32) and (16.34).
The second half of the assertion is a theorem: by Example 16.37 the Euler–Lagrange equation of Equation (16.60) is Laplace's equation, so a minimiser, if there is one, is harmonic. Riemann used the principle in this form as the existence engine of his function theory — it is how he produced the conformal maps and the functions on Riemann surfaces that his geometry required — and named it after Dirichlet, from whose lectures he had it. The first half of the assertion is the problem: an infimum need not be attained.
Let
and
Then \(\inf_{\mathcal{C}}J=0\), and no \(y\in\mathcal{C}\) attains it. Rests on Equation (16.61) and Definition 16.11.
Derives Proposition 16.68. The integrand is nonnegative, so \(J\ge0\). For \(\varepsilon>0\) put
which lies in \(\mathcal{C}\), being smooth and odd with \(y_{\varepsilon}(\pm1)=\pm1\). Its derivative is \(y_{\varepsilon}'=\varepsilon \bigl[\arctan(1/\varepsilon)\bigr]^{-1} \left(\varepsilon^{2}+x^{2}\right)^{-1}\), so, using \(x^{2}\le\varepsilon^{2}+x^{2}\),
which tends to \(0\) as \(\varepsilon\to0^{+}\) because \(\arctan(1/\varepsilon)\to\pi/2\). Hence the infimum is \(0\).
Suppose now \(J[y]=0\) for some \(y\in\mathcal{C}\). The integrand \(x^{2}y'^{2}\) is continuous and nonnegative with vanishing integral, so it vanishes identically; hence \(y'(x)=0\) for every \(x\ne0\), and \(y\) is constant on \([-1,0)\) and on \((0,1]\). Continuity at \(0\) forces the two constants to agree, so \(y\) is constant on \([-1,1]\), contradicting \(y(-1)=-1\ne1=y(1)\).
∎Weierstrass presented this argument to the Berlin Academy in 1870 [Weierstrass:1895], and its effect was immediate: every existence proof resting on “the infimum is attained” was suspect, Riemann's function theory included. The gap is real and not an artefact of the example: the infimum of a functional bounded below is always approached by a minimising sequence, but nothing yet said that the sequence converges, nor that the functional is continuous along it.
Hilbert restored the Dirichlet principle thirty years later [Hilbert:1904], not by patching Riemann's argument but by supplying the two missing ingredients under explicit hypotheses on the domain and the boundary data: a minimising sequence with enough compactness to extract a convergent subsequence, and enough semicontinuity of \(D\) to conclude that the limit is a minimiser. This is the direct method, and its abstract form is Theorem 16.70. Hilbert's twentieth and twenty-third problems both grew out of it [Hilbert:1900], and the modern treatment is Courant–Hilbert [Courant:1962]. The stake for physics is not decorative: the existence of solutions to the boundary-value problems of Partial Differential Equations and Electrostatics is exactly the statement Riemann assumed and Weierstrass questioned.
Tonelli's existence theory
The direct method has three ingredients, and each answers one of the two objections above.
Let \(X\) be a reflexive Banach space and \(J:X\longrightarrow\R\cup\set{+\infty}\) a functional that is not identically \(+\infty\) and satisfies
-
coercivity: \(J[u]\longrightarrow+\infty\) whenever \(\norm{u}\longrightarrow\infty\);
-
sequential weak lower semicontinuity: whenever \(u_{n}\rightharpoonup u\) weakly in \(X\), \(J[u]\le\liminf_{n}J[u_{n}]\).
Then \(J\) attains its infimum on \(X\). Rests on Definitions 6.9 and 12.2.
Derives Theorem 16.70. Let \(m=\inf_{X}J<+\infty\) and pick a minimising sequence \((u_{n})\), so \(J[u_{n}]\longrightarrow m\). The sequence \(\left(J[u_{n}]\right)\) is bounded above, so by coercivity \(\left(\norm{u_{n}}\right)\) is bounded: otherwise a subsequence would have \(\norm{u_{n_{k}}}\to\infty\) and hence \(J[u_{n_{k}}]\to+\infty\). In a reflexive Banach space a bounded sequence has a weakly convergent subsequence — this is the compactness input, the Eberlein–Smulian theorem, whose Hilbert-space case is the one used in this book (Hilbert Spaces) — so \(u_{n_{k}}\rightharpoonup u\) for some \(u\in X\). Weak lower semicontinuity then gives
and \(J[u]\ge m\) by the definition of the infimum, so \(J[u]=m>-\infty\).
∎Coercivity and semicontinuity are demands on the same topology in opposite directions. Making the topology coarser makes compactness easier and semicontinuity harder; making it finer does the reverse. The weak topology of a reflexive space is the compromise that works, and it is why Sobolev spaces — not spaces of continuously differentiable functions — are the right home for the direct method: \(C^{1}[a,b]\) with the norm Equation (16.13) is not reflexive and its bounded sets are not weakly compact. That is the technical form of the objection Weierstrass raised.
The remaining question is which integrands make Equation (16.11) weakly lower semicontinuous. Tonelli identified the answer: convexity in the slope [Tonelli:1915].
Let \(1<r<\infty\), let \(f(x,q)\) be continuous, bounded below, of class \(C^{1}\) in \(q\) and convex in \(q\) for each \(x\), and suppose \(\abs{\pp f/\pp q\left(x,q\right)}\le C\left(1+\abs{q}^{r-1}\right)\). If \(u_{n}'\rightharpoonup u'\) weakly in \(L^{r}(a,b)\), then
Rests on Theorem 16.70 and Equation (16.11).
Derives Lemma 16.72. Convexity of a \(C^{1}\) function means its graph lies above every tangent:
Apply this pointwise with \(q_{1}=u'(x)\) and \(q_{2}=u_{n}'(x)\) and integrate:
The growth hypothesis puts \(\pp f/\pp q\left(\cdot,u'\right)\) in \(L^{r'}(a,b)\) with \(r'=r/(r-1)\), which is the dual of \(L^{r}\), so the last integral tends to \(0\) by the definition of weak convergence. Taking \(\liminf\) of both sides gives the claim.
∎Let \(1<r<\infty\) and let \(f\left(x,y,q\right)\) be continuous, of class \(C^{1}\) in \(q\), convex in \(q\) for each \(\left(x,y\right)\), and coercive in the sense \(f\left(x,y,q\right)\ge\alpha\abs{q}^{r}-\beta\) with \(\alpha>0\); assume further the growth bound \(\abs{\pp f/\pp q\left(x,y,q\right)} \le C_{R}\left(1+\abs{q}^{r-1}\right)\) for \(\abs{y}\le R\), which is the hypothesis of Lemma 16.72 made locally uniform in \(y\). Then the functional Equation (16.11) attains its minimum over the class of functions in the Sobolev space \(W^{1,r}(a,b)\) with the prescribed endpoint values. Rests on Theorem 16.70 and Lemma 16.72.
The derivation is carried out in Tonelli's Existence Theorem for an Integrand Depending on the Function. The coercivity estimate and the abstract argument are those of the direct-method theorem and the convexity lemma proved above; the passage from an integrand \(f(x,q)\) to one that also depends on \(y\) rests on the compact embedding of \(W^{1,r}(a,b)\) into the continuous functions, which the appendix proves from the fundamental theorem of calculus, Hölder's inequality and the Arzelà–Ascoli theorem. The growth bound is what lets that embedding replace the Scorza-Dragoni argument of the classical treatments; the appendix states precisely what dropping it would cost, and Tonelli's own hypotheses [Tonelli:1915] are the ones without it.
Full derivation in Appendix A.
Derives Theorem 16.73.
For a mechanical system with kinetic energy \(T=\tfrac12 m_{ij}(q)\dot q^{i}\dot q^{j}\) the Lagrangian \(\Lag=T-V\) is quadratic in the velocities with a positive definite coefficient matrix, hence convex in them, which is half of what Theorem 16.73 asks. The other half is coercivity, and it is a condition on the potential in the direction opposite to the one a physicist expects: \(\Lag\ge\alpha\abs{\dot q}^{2}-\beta\) uniformly in \(q\) demands that \(V\) be bounded above, not below. For a confining potential the action is in general not bounded below at all. The harmonic oscillator, \(m,k>0\) and \(q(0)=q(T)=0\), is the smallest example: on \(q(t)=A\sin\left(\pi t/T\right)\),
which for \(T>\pi\sqrt{m/k}=\pi/\omega\), with \(\omega=\sqrt{k/m}\), is negative and tends to \(-\infty\) as \(A\to\infty\): the infimum is \(-\infty\) and no minimiser exists in any class. That threshold is exactly the kinetic focus of Remark 16.98, and the same phenomenon as in the model functional of Remark 16.25. So Theorem 16.73 certifies existence for a mechanical Lagrangian whose potential is bounded above — a free or a repelled particle — and not for a confining one; past the first conjugate point the action of Section 16.7.1 is genuinely unbounded below, which is why that section states Hamilton's principle as one of stationary and not of least action. Convexity is not a technicality that can be dropped either. For
the integrand is not convex in \(y'\), and \(\inf J=0\): a sawtooth of slope \(\pm1\) and amplitude \(1/2n\) makes both terms as small as one likes. No function attains \(0\), since that would need \(y\equiv0\) and \(\abs{y'}\equiv1\) at once. Minimising sequences oscillate faster and faster — a fine structure with a physical counterpart in the twinned microstructure of martensitic phase transitions.
The Rayleigh–Ritz method
The direct method has a computational face. Instead of minimising over the whole admissible class, minimise over a finite-dimensional family of trial functions: the result is an upper bound on the true minimum, and a linear algebra problem in place of a differential equation.
For the Sturm–Liouville data \(p>0\), \(q\) and \(r>0\) continuous on \([a,b]\), the Rayleigh quotient of a nonzero \(y\in C^{1}_{0}[a,b]\) is
Rests on Equation (16.42) and Definition 9.54.
Let \(\lambda_{1}\le\lambda_{2}\le\cdots\) be the eigenvalues of the Sturm–Liouville problem
with eigenfunctions \(y_{n}\) orthonormal in the weighted inner product \(\avg{u,v}_{r}=\int_{a}^{b}r\,u\,v\,\dd x\) and complete in the corresponding Hilbert space (Ordinary Differential Equations and Sturm–Liouville Theory, Hilbert Spaces). Then
the minimum being attained exactly at the multiples of \(y_{1}\). Rests on Equations (16.65) and (16.66).
Derives Theorem 16.76. Write \(B[u,v]=\int_{a}^{b}\left(p\,u'v'+q\,uv\right)\dd x\) for the energy form, so that \(\mathcal{R}[y]=B[y,y]/\avg{y,y}_{r}\), and fix \(K>0\) with \(q+Kr\ge0\) on \([a,b]\), which is possible because \(q\) is continuous and \(r>0\) on a compact interval. Put \(B_{K}[u,v]=B[u,v]+K\avg{u,v}_{r}\). Since \(p>0\) and \(q+Kr\ge0\),
Expand \(y=\sum_{n\ge1}c_{n}y_{n}\) with \(c_{n}=\avg{y,y_{n}}_{r}\); by completeness the denominator of Equation (16.65) is \(\sum_{n}c_{n}^{2}\).
The one integration by parts needed is against the eigenfunctions, not against \(y\): for \(y\in C^{1}_{0}[a,b]\) and \(y_{n}\in C^{2}[a,b]\), Equation (7.28) with \(y(a)=y(b)=0\) moves the derivative onto \(y_{n}\),
by Equation (16.66). Only \(y_{n}\) is differentiated twice, so nothing beyond \(y\in C^{1}_{0}[a,b]\) is used — which matters, because the trial functions of Proposition 16.79 are only \(C^{1}\). Hence \(B_{K}[y,y_{n}]=\left(\lambda_{n}+K\right)c_{n}\), and for the partial sums \(S_{N}=\sum_{n\le N}c_{n}y_{n}\) bilinearity and orthonormality give \(B_{K}[y,S_{N}]=B_{K}[S_{N},S_{N}] =\sum_{n\le N}\left(\lambda_{n}+K\right)c_{n}^{2}\). Therefore
by Equation (16.68). Taking \(u=y_{1}\) there, together with Equation (16.69) for \(y=y_{1}\), shows \(B_{K}[y_{1},y_{1}]=\lambda_{1}+K\ge0\), so every term of the sum is nonnegative and \(N\to\infty\) is legitimate: \(B_{K}[y,y]\ge\sum_{n}\left(\lambda_{n}+K\right)c_{n}^{2} \ge\left(\lambda_{1}+K\right)\sum_{n}c_{n}^{2}\). Subtracting \(K\avg{y,y}_{r}=K\sum_{n}c_{n}^{2}\) from both sides leaves \(B[y,y]\ge\lambda_{1}\avg{y,y}_{r}\), that is
If \(\mathcal{R}[y]=\lambda_{1}\) then both inequalities in Equation (16.70) are equalities, and the second forces \(c_{n}=0\) whenever \(\lambda_{n}>\lambda_{1}\); the lowest eigenvalue of a regular Sturm–Liouville problem is simple, so \(y\) agrees with a multiple of \(y_{1}\) in the weighted inner product and, both being continuous, pointwise. Taking \(y=y_{1}\) in Equation (16.69) gives \(\mathcal{R}[y_{1}]=\lambda_{1}\), so the value is attained.
∎Only the inequality Equation (16.70) was proved, and deliberately so: it is all the variational characterisation needs, and it costs nothing beyond \(y\in C^{1}_{0}[a,b]\). The sharper statement
is equivalent to \(B_{K}[y-S_{N},y-S_{N}]\to0\), that is to completeness of the eigenfunctions in the energy norm \(\norm{u}_{E}=B_{K}[u,u]^{1/2}\) — a strictly stronger fact than the completeness in the weighted \(L^{2}\) space assumed in Theorem 16.76, and one proved in Completeness of the Sturm–Liouville Eigenfunctions in the Energy Norm by realising the inverse of the Sturm–Liouville operator as a compact integral operator with the Green function as its kernel and applying the Hilbert–Schmidt theorem Theorem 12.44 to it. That appendix also proves the weighted \(L^{2}\) completeness assumed here, so the hypothesis of Theorem 16.76 is discharged as well; the inequality Equation (16.70) proved above, which is what the variational characterisation uses, rests on Bessel's inequality alone and is independent of both, so there is no circle.
Full derivation in Appendix A.
Derives Equation (16.71). The distinction is not pedantry: it is exactly the identity that the error estimate in Corollary 16.78 needs, and the inequality alone bounds that error in the wrong direction.
For every admissible trial function \(\chi\ne0\), \(\lambda_{1}\le\mathcal{R}[\chi]\). A trial function accurate to first order in the deviation from \(y_{1}\) gives an eigenvalue accurate to second order. Rests on Equations (16.67) and (16.71).
Derives Corollary 16.78. The inequality is Equation (16.67). For the error statement, write \(\chi=y_{1}+\epsilon\) with \(\avg{\epsilon,y_{1}}_{r}=0\); then \(c_{1}=1\) and \(\sum_{n\ge2}c_{n}^{2}=\norm{\epsilon}_{r}^{2}\), so Equation (16.71) gives
whose numerator is a quadratic form in \(\epsilon\): the error is second order in the deviation. It must be measured in the energy norm of Remark 16.77 and not in \(\norm{\cdot}_{r}\), since \(\lambda_{n}-\lambda_{1}\) is unbounded and a deviation small in the weighted \(L^{2}\) norm may have a large derivative.
∎Let \(\chi_{1},\dots,\chi_{N}\) be linearly independent admissible trial functions and minimise \(\mathcal{R}\) over \(y=\sum_{k=1}^{N}c_{k}\chi_{k}\). Then the stationary values of \(\mathcal{R}\) are the roots \(\lambda\) of the generalized eigenvalue problem
with the symmetric matrices
and the lowest root satisfies \(\lambda_{1}\le\lambda^{(N)}\), decreasing as the trial family is enlarged. Rests on Equation (16.65) and Corollary 16.78.
Derives Proposition 16.79. Substituting the expansion into Equation (16.65) gives \(\mathcal{R}=\left(c_{j}K_{jk}c_{k}\right) /\left(c_{j}M_{jk}c_{k}\right)\), a ratio of quadratic forms in the \(N\) numbers \(c_{k}\). Stationarity with respect to \(c_{j}\) requires
which, after multiplying by \(\tfrac12 c_{l}M_{lm}c_{m}\) and recognising \(\mathcal{R}=\lambda\), is Equation (16.72). The bound \(\lambda_{1}\le\lambda^{(N)}\) is Corollary 16.78 applied to the minimising member of the trial family, and enlarging the family can only lower the minimum because the smaller family is a subset of the larger.
∎Rayleigh introduced the quotient to estimate the frequencies of vibrating systems [Rayleigh:1877]: for a string of length \(L\) under tension \(F\) (in newtons) and of linear density \(\rho_{\ell}\) (in \(\mathrm{kg}/\mathrm{m}\)) the data are \(p=F\), \(q=0\), \(r=\rho_{\ell}\), and \(\mathcal{R}\) has the SI dimension \(/\mathrm{s}^{2}\), so \(\lambda_{1}=\omega_{1}^{2}\) is the square of the fundamental angular frequency. Ritz turned the estimate into a systematic and convergent scheme by enlarging the trial family [Ritz:1909]; his 1908 spectroscopic paper [Ritz:1908] is a different work and should not be cited for this method. Equation (16.72) is the ancestor of the finite element method, of the linear variational method for molecular orbitals, and of the variational principle of quantum mechanics: replacing \(\mathcal{R}\) by \(\avg{\psi|\Ham|\psi}/\avg{\psi|\psi}\) gives an upper bound on the ground-state energy, which is the content of Approximation Methods. The normal-mode analyses of Oscillations and Mechanical Waves are the same computation with \(K\) the stiffness matrix and \(M\) the mass matrix. The minimax characterisation of the higher eigenvalues, which extends Equation (16.67) to \(\lambda_{n}\), belongs with the spectral theory of Hilbert Spaces.
The Plateau problem
The classical existence question of the subject is the one the soap films of Section 16.1.3 pose.
Let \(\Gamma\subset\R^{3}\) be a rectifiable Jordan curve. Then there exists a continuous map \(X\) from the closed unit disc \(\overline{B}\) into \(\R^{3}\), harmonic and conformal on the open disc, restricting to a monotone parametrisation of \(\Gamma\) on the boundary circle, and minimising the area among all such disc-type maps. Its image is a minimal surface spanning \(\Gamma\). Rests on Equation (16.7) and Theorem 16.70.
Douglas minimises a boundary-value functional over monotone boundary parametrisations, defeating the non-compactness of the conformal group of the disc by a three-point normalisation, and Radó argues instead through polygonal approximation and the Dirichlet integral. Either proof is monograph length and neither is reproduced here. The Plateau Problem: the Architecture of Douglas' Proof gives the architecture of Douglas' argument with every step this treatise can carry out carried out — the domination of the area by the Dirichlet integral with its equality case, the conformal invariance of that integral, the identity between Douglas' boundary functional and the Dirichlet integral of the harmonic extension, the lower semicontinuity, the three-point normalisation, the Courant–Lebesgue estimate and the compactness it buys, and the computation showing that a harmonic conformal map has vanishing mean curvature — and states as displayed theorems the results it does not prove. The theorem itself therefore remains quoted: the appendix reduces it to two named imports, the solvability of the Dirichlet problem for the disc and Douglas' inner-variation proof of conformality, and says so explicitly.
Full derivation in Appendix A.
Derives Theorem 16.81.
Theorem 16.81 was proved independently by Douglas [Douglas:1931] and Radó [Rado:1930], and Douglas' work was recognised with one of the first two Fields Medals. It is worth being exact about the claim, because the statement is often quoted more strongly than it was proved.
-
The surface produced is of disc type. A contour may bound minimal surfaces of higher genus or with several boundary components, and those require separate arguments.
-
The map is a minimiser of area among disc-type maps. It need not be an embedding: the image may have self-intersections, and it may carry isolated branch points where the differential drops rank, although interior branch points were later excluded for area-minimising discs.
-
Conformality is not an extra hypothesis but part of the conclusion, and it is what converts the geometric area functional into the Dirichlet integral to which the direct method applies. This is the technical heart of both proofs.
The result is the landmark application of the direct method in a genuinely two-dimensional problem, and the differential-geometric apparatus it needs — first and second fundamental forms, mean curvature — is that of Differentiable Manifolds, Tensors, and Curvature.
Symmetry and conservation: Noether's theorems
Emmy Noether's 1918 paper [Noether:1918] contains two theorems, and they answer opposite questions. The first says that a symmetry group of finitely many parameters produces one conserved quantity per parameter. The second says that a symmetry group depending on arbitrary functions produces no new conserved quantity at all, but instead identities among the field equations themselves. Both are consequences of the first variation, and both are proved here.
Noether's first theorem
A one-parameter family of transformations of the variables of Equation (16.11),
with \(\xi,\phi^{i}\) of class \(C^{1}\), is a variational symmetry of \(J\) if for every curve and every subinterval
for some function \(\Lambda\left(x,y\right)\). When \(\Lambda\) may be taken to vanish the symmetry is strict; the general case is the divergence symmetry of Bessel-Hagen [BesselHagen:1921], without which the Galilean boost \(\vect{x}\longmapsto\vect{x}+\vect{v}t\) escapes the theorem: it changes a free Lagrangian by the total time derivative of \(m\left(\vect{v}\cdot\vect{x}-\tfrac12\abs{\vect{v}}^{2}t\right)\), and the conserved quantity it then yields is the centre-of-mass theorem. Rests on Equation (16.11) and Definition 16.15.
Let Equation (16.74) be a variational symmetry of \(J\) with divergence term \(\Lambda\). Then along every solution of the Euler–Lagrange system Equation (16.30) the quantity
is constant: \(\dd I/\dd x=0\). Rests on Equation (16.75), Equation (16.30) and Proposition 16.17.
Derives Theorem 16.84. Write \(\bar y^{i}(\bar x)\) for the transformed curve as a function of the transformed independent variable, and compare it with the original curve at the same value of \(x\). From Equation (16.74), \(\bar y^{i}(\bar x)=y^{i}(x) +\varepsilon\phi^{i}\) and \(\bar y^{i}(x+\varepsilon\xi) =\bar y^{i}(x)+\varepsilon\xi\,y'^{i}(x)+O(\varepsilon^{2})\), so the vertical (evolutionary) variation is
Renaming the dummy variable \(\bar x\) back to \(x\), the left-hand side of Equation (16.75) is \(\int_{x_{1}+\varepsilon\xi_{1}}^{x_{2}+\varepsilon\xi_{2}} f\left(x,\bar y(x),\bar y'(x)\right)\dd x\), and moving the limits back costs one boundary term to first order:
By Proposition 16.17 applied to the vertical variation \(\bar\phi\), and integrating the \(\bar\phi'\) term by parts,
with \(E_{i}\) the Euler–Lagrange expression of Equation (16.40). Substituting both displays into Equation (16.75), dividing by \(\varepsilon\) and letting \(\varepsilon\to0\),
This holds for every subinterval \([x_{1},x_{2}]\subseteq[a,b]\) and every curve, so the integrand, which is continuous, vanishes identically. On a solution of Equation (16.30) the first term is zero, leaving \(\dd I/\dd x=0\) with \(I\) as in Equation (16.76), since \(\bar\phi^{i}=\phi^{i}-y'^{i}\xi\).
∎Equation (16.78) holds without any equation of motion: it says that for a variational symmetry the combination \(E_{i}\bar\phi^{i}\) is a total derivative, identically in the curve. That off-shell statement is what Section 16.6.2 exploits, and it is also the honest form of the converse: a conserved quantity of the form Equation (16.76) corresponds to a variational symmetry, but a conservation law of the equations of motion that does not arise from a symmetry of the action is not covered — Noether's theorem is a statement about Lagrangians, not about differential equations.
Equation (16.74) moves the point \(\left(x,y\right)\) only: \(\xi\) and \(\phi^{i}\) are functions of \(\left(x,y\right)\) and never of \(y'\), and no divergence term repairs that. The Laplace–Runge–Lenz vector of the Kepler problem is the standard witness. For \(f=\abs{\vect{p}}^{2}/2m-V\) with \(\vect{p}=m\dot{\vect{x}}\), the part of Equation (16.76) that is quadratic in \(\vect{p}\) is \(-\xi\left(t,\vect{x}\right)\abs{\vect{p}}^{2}/2m\), which is isotropic in \(\vect{p}\). The corresponding part of \(A_{j}=\abs{\vect{p}}^{2}x_{j} -\left(\vect{x}\cdot\vect{p}\right)p_{j} -mk\,x_{j}/\abs{\vect{x}}\) is not: it vanishes when \(\vect{p}\) is parallel to \(\vect{x}\) and equals \(\abs{\vect{p}}^{2}x_{j}\) when the two are perpendicular. No choice of \(\xi\) reproduces an anisotropy, so \(\vect{A}\) is the Noether charge of no transformation of the form Equation (16.74), divergence term or not. The generator that does produce it, \(\delta x_{i}=\varepsilon_{j}\left(2\dot x_{i}x_{j} -x_{i}\dot x_{j}-\delta_{ij}\,\vect{x}\cdot\dot{\vect{x}}\right)\), depends on the velocities: it is a generalized symmetry, outside the hypotheses proved here and belonging to the later extension of Noether's argument to transformations acting on the derivatives as well. The vector itself, and the hidden \(\SO(4)\) symmetry it generates, are put to work in The Hydrogen Atom.
If \(f\) has no explicit dependence on \(x\), then \(\xi=1\), \(\phi^{i}=0\), \(\Lambda=0\) is a strict variational symmetry and Equation (16.76) gives the conserved quantity
which for one dependent variable is the Beltrami identity Equation (16.25). Rests on Equations (16.25) and (16.76).
Derives Corollary 16.87. With \(\bar x=x+\varepsilon\) and \(\bar y=y\) the integrand is unchanged and the integration measure is unchanged, so Equation (16.75) holds with \(\Lambda=0\) precisely because \(f\) carries no explicit \(x\). Then \(I=\left(\pp f/\pp y'^{i}\right)\left(-y'^{i}\right)+f=-h\).
∎Let \(y^{1},y^{2},y^{3}\) be Cartesian coordinates of a particle in ordinary space and \(x=t\) the time.
-
If \(f\) does not depend on \(y^{k}\) for some fixed \(k\), then \(\xi=0\), \(\phi^{i}=\delta^{i}_{k}\) is a strict symmetry and \(\pp f/\pp y'^{k}\) is conserved: the momentum conjugate to a translation.
-
If \(f\) is invariant under rotations about the unit vector \(\vect{n}\), so that \(\phi^{i}=\left(\vect{n}\times\vect{y}\right)^{i}\) and \(\xi=0\) define a strict symmetry, then \(\vect{n}\cdot\left(\vect{y}\times\vect{p}\right)\) is conserved, with \(p_{i}=\pp f/\pp y'^{i}\).
Rests on Equation (16.76) and Proposition 16.28.
Derives Corollary 16.88. In both cases \(\xi=0\) and \(\Lambda=0\), so Equation (16.76) reduces to \(I=\left(\pp f/\pp y'^{i}\right)\phi^{i}=p_{i}\phi^{i}\). Case (1) gives \(I=p_{k}\), which is also Proposition 16.28. Case (2) gives \(I=\vect{p}\cdot\left(\vect{n}\times\vect{y}\right) =\vect{n}\cdot\left(\vect{y}\times\vect{p}\right)\) by the cyclic invariance of the scalar triple product.
∎With \(x=t\) in seconds and \(f=\Lag\) in joules, \(\pp\Lag/\pp\dot y^{i}\) is a momentum in \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\) when \(y^{i}\) is a length, and \(I\) in Equation (16.76) carries the dimension of action, \(\mathrm{J}\,\mathrm{s}\), divided by whatever dimension the group parameter \(\varepsilon\) has. For a dimensionless \(\varepsilon\) — an angle, say — \(I\) is an action: angular momentum in \(\mathrm{J}\,\mathrm{s}\) is the familiar case. For a time translation \(\varepsilon\) is a time and \(I\) is an energy in joules; for a spatial translation \(\varepsilon\) is a length and \(I\) a momentum. The rule is that the product \(I\,\varepsilon\) is always an action, which is the statement that the conserved charge generates the symmetry — the observation that becomes the Poisson-bracket structure of Hamiltonian Mechanics and the commutator structure of Linear Algebra and Representation Theory.
Let \(S[\varphi]\) be the field action Equation (16.31) on \(\Omega\subset\R^{D}\) and let
leave \(S\) invariant up to the divergence of a vector \(\Lambda^{\mu}\left(x,\varphi\right)\), in the sense that the analogue of Equation (16.75) holds on every subregion. Then on every solution of Equation (16.32) the Noether current
is conserved, \(\pp_{\mu}j^{\mu}=0\). Rests on Equation (16.32), Equation (16.33) and Theorem 16.84.
Derives Theorem 16.90. The argument of Theorem 16.84 carries over word for word with the divergence theorem (Theorem 7.101) in place of the one-dimensional integration by parts. The vertical variation is \(\bar\Phi^{A}=\Phi^{A}-\xi^{\nu}\pp_{\nu}\varphi^{A}\), obtained exactly as Equation (16.77). Moving the domain \(\bar\Omega\) back to \(\Omega\) contributes, to first order, the flux of \(\Lag\,\xi^{\mu}\) through \(\pp\Omega\), that is \(\varepsilon\int_{\Omega}\pp_{\mu}\left(\Lag\xi^{\mu}\right)\dd^{D}x\), because the Jacobian of Equation (16.80) is \(1+\varepsilon\,\pp_{\mu}\xi^{\mu}+O(\varepsilon^{2})\). The vertical part contributes, by Equation (16.33) before the boundary term is discarded,
Invariance up to \(\pp_{\mu}\Lambda^{\mu}\) therefore gives, for every subregion and hence pointwise,
On shell the bracket in the first term vanishes by Equation (16.32), leaving \(\pp_{\mu}j^{\mu}=0\) with \(j^{\mu}\) as in Equation (16.81).
∎Let \(D=4\) with coordinates \(x^{\mu}=\left(ct,x,y,z\right)\) on the spacetime of Minkowski Space and Its Symmetries, and let \(\Lag\) carry no explicit dependence on \(x^{\mu}\). Then \(\xi^{\mu}=a^{\mu}\) constant, \(\Phi^{A}=0\), \(\Lambda^{\mu}=0\) is a strict symmetry, and the four conserved currents assemble into the canonical energy–momentum tensor
Derives Corollary 16.91. Substituting \(\xi^{\mu}=a^{\mu}\) and \(\Phi^{A}=0\) into Equation (16.81) gives \(j^{\mu}=-a^{\nu}\left[\pi^{\mu}_{A}\,\pp_{\nu}\varphi^{A} -\delta^{\mu}_{\nu}\Lag\right]=-a^{\nu}T^{\mu}{}_{\nu}\), and \(\pp_{\mu}j^{\mu}=0\) for every constant \(a^{\nu}\) forces \(\pp_{\mu}T^{\mu}{}_{\nu}=0\) componentwise.
∎\(T^{0}{}_{0}\) is the energy density in \(\mathrm{J}/\mathrm{m}^{3}\) and \(T^{0}{}_{k}\) the momentum density; the conservation law Equation (16.83) integrates over a spatial region to the statement that energy and momentum leave it only through its boundary. The canonical tensor is not in general symmetric, and symmetrising it — adding an identically conserved improvement term — is what produces the tensor that couples to gravity in The Einstein Field Equations. An internal symmetry, \(\xi^{\mu}=0\) with \(\Phi^{A}\) a linear transformation of the fields, gives instead a current whose conserved charge is electric charge, or baryon number, or any other quantum number of that kind; the electric case is worked in Generalized Classical Field Theory, and the conservation laws of the mechanics chapters (Lagrangian Mechanics) are the one-dimensional specialisation Corollaries 16.87 and 16.88.
Noether's second theorem
Suppose the action Equation (16.31) is invariant under a group of transformations depending on \(m\) arbitrary functions \(\epsilon^{r}(x)\) of compact support in \(\Omega\), acting on the fields as
with \(R^{A}{}_{r}\) and \(R^{A\mu}{}_{r}\) of class \(C^{1}\). Then the Euler–Lagrange expressions \(E_{A}=\pp\Lag/\pp\varphi^{A}-\pp_{\mu}\pi^{\mu}_{A}\) obey the \(m\) identities
identically in the fields — not merely on solutions. Rests on Equation (16.33), Equation (16.32) and Theorem 16.90.
Derives Theorem 16.93. Because the \(\epsilon^{r}\) have compact support, all boundary terms in Equation (16.33) vanish and invariance of the action reads
for every choice of the \(\epsilon^{r}\), and for arbitrary field configurations, since the invariance is a property of the action and not of its extremals. Integrating the second term by parts, again with no boundary contribution,
The bracket is continuous and the \(\epsilon^{r}\) are arbitrary, so Lemma 16.18, in its \(D\)-dimensional form used in Theorem 16.36, gives Equation (16.85), one identity for each \(r\).
∎Under the hypotheses of Theorem 16.93 the \(n\) Euler–Lagrange equations \(E_{A}=0\) satisfy \(m\) differential identities, so at most \(n-m\) of them are independent. The Cauchy problem is correspondingly underdetermined: a solution stays a solution after any gauge transformation Equation (16.84) with \(\epsilon^{r}\) vanishing in the past. Rests on Equation (16.85).
Derives Corollary 16.94. Equation (16.85) is a linear relation among the \(E_{A}\) and their first derivatives holding identically, so the map \(\varphi\longmapsto\left(E_{1},\dots,E_{n}\right)\) has image in a set of codimension at least \(m\); equivalently, once \(n-m\) of the equations and suitable initial data are imposed, the remaining \(m\) follow by integration. That a gauge transform of a solution is a solution is immediate from invariance of the action.
∎Electrodynamics. The action of the electromagnetic potential is invariant under \(\delta A_{\mu}=\pp_{\mu}\epsilon\), one arbitrary function, so in Equation (16.84) \(R=0\) and \(R^{\nu}{}_{\mu}=\delta^{\nu}_{\mu}\). Equation (16.85) becomes \(\pp_{\mu}E^{\mu}=0\) identically, where \(E^{\mu}\) is the Euler–Lagrange expression of \(A_{\mu}\). Since \(E^{\mu}\) contains the source term, the field equations are consistent only if the source is itself conserved: charge conservation is not an extra postulate but the integrability condition of the Maxwell equations, which is the reading developed in Generalized Classical Field Theory [BesselHagen:1921].
General relativity. The Einstein–Hilbert action is invariant under diffeomorphisms, which in infinitesimal form are \(\delta g_{\mu\nu} =\nabla_{\mu}\epsilon_{\nu}+\nabla_{\nu}\epsilon_{\mu}\): four arbitrary functions, with \(R^{A}{}_{r}\) built from the connection and \(R^{A\mu}{}_{r}\) from the metric. The Euler–Lagrange expression is proportional to the Einstein tensor, and Equation (16.85) is the contracted Bianchi identity \(\nabla_{\mu}G^{\mu\nu}=0\), which follows from the second Bianchi identity Equation (13.314) of Proposition 13.157 by two contractions; that proposition states the uncontracted identities and does not perform the contraction. So the identity that makes the Einstein equations consistent with local energy–momentum conservation is not a happy accident of Riemannian geometry but a consequence of general covariance, exactly as The Einstein Field Equations presents it.
In both cases the degeneracy of the Lagrangian — the Hessian \(\pp^{2}\Lag/\pp\dot\varphi^{A}\pp\dot\varphi^{B}\) is singular, because some field velocities never appear — is the origin of the constraint analysis of Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism.
The variational principles of physics
Hamilton's principle
For a mechanical system with generalized coordinates \(q^{1},\dots,q^{n}\) and Lagrangian \(\Lag\left(q,\dot q,t\right)\) in joules, the action between two instants is
of SI dimension \(\mathrm{J}\,\mathrm{s}\). Hamilton's principle is the statement that the motion between prescribed configurations \(q(t_{1})\) and \(q(t_{2})\) makes Equation (16.86) stationary [Hamilton:1834]. Rests on Equation (16.11) and Definition 21.31.
Let a single particle of mass \(m\) move in ordinary space under a force derived from a potential \(V\left(\vect{x},t\right)\) in joules, and take \(\Lag=\tfrac12 m\abs{\dot{\vect{x}}}^{2}-V\). Then \(\vect{x}(t)\) makes Equation (16.86) stationary among curves with the prescribed endpoints if and only if it satisfies
Derives Theorem 16.97. The three Euler–Lagrange equations Equation (16.30) with \(x=t\) and \(y^{i}=x^{i}\) read
which is Equation (16.87) componentwise. Conversely a solution of Equation (16.87) satisfies the Euler–Lagrange system, hence makes the first variation vanish by Proposition 16.17 and the integration by parts of Theorem 16.22.
∎The name “principle of least action” is a historical accident. \(\Lag=T-V\) is convex in the velocities, so Equation (16.47) holds strictly and the second variation is positive definite over short enough intervals; but by Theorem 16.55 the motion ceases to minimise past the first conjugate point — the kinetic focus — and for a harmonic oscillator of angular frequency \(\omega\) that happens after a time \(\pi/\omega\), half a period. The action is then a saddle point, never a maximum. This is why Definition 16.96 says stationary. The standard mechanics treatments [Goldstein:2002] [Landau:1976] develop the principle in the form used throughout Lagrangian Mechanics, and the path-integral formulation of Path-Integral Quantization is what eventually explains why the stationary path is the one that matters: neighbouring paths interfere destructively except near a stationary point of \(S\).
Maupertuis' principle
Hamilton's principle fixes the two times and varies the path. The older principle fixes the energy instead, lets the transit time float, and varies only the geometric path.
For a system with momenta \(p_{i}=\pp\Lag/\pp\dot q^{i}\), the abbreviated action of a path traversed between two configurations is
also of dimension \(\mathrm{J}\,\mathrm{s}\). Rests on Equation (16.86) and Definition 16.96.
Let \(n\) be a positive \(C^{1}\) function on a region of ordinary space and consider the parametrisation-invariant functional
Its extremals, parametrised by arclength \(s\) so that \(\vect{T}=\dd\vect{x}/\dd s\) is a unit vector, satisfy
Derives Lemma 16.100. The integrand \(f=n\left(\vect{x}\right)\abs{\vect{x}'}\) is homogeneous of degree one in \(\vect{x}'\), so the functional does not depend on the parametrisation and the three Euler–Lagrange equations Equation (16.30) are not independent; the arclength gauge \(\abs{\vect{x}'}=1\) fixes the redundancy. With
setting \(\abs{\vect{x}'}=1\) after the differentiation turns Equation (16.30) into \(\dd\left(n\,T^{i}\right)/\dd s=\pp n/\pp x^{i}\), which is Equation (16.90).
∎Let a single particle of mass \(m\) move in a time-independent potential \(V\left(\vect{x}\right)\) with total energy \(E>V\left(\vect{x}\right)\) throughout the region considered. Then the unparametrised orbits of energy \(E\) are exactly the extremals of the geometric functional
that is, the geodesics of the Jacobi metric \(\dd\sigma^{2}=2m\left(E-V\right)\dd s^{2}\); and the time parametrisation is recovered from \(\dd t=\sqrt{m/\left[2\left(E-V\right)\right]}\,\dd s\). Rests on Equation (16.88), Lemma 16.100 and Equation (16.87).
Derives Theorem 16.101. Along a motion of energy \(E\) the speed is \(v=\abs{\dot{\vect{x}}}=\sqrt{2\left(E-V\right)/m}\), so the momentum has magnitude \(\abs{\vect{p}}=mv=\sqrt{2m\left(E-V\right)}\) and points along \(\vect{T}\). Hence \(p_{i}\,\dd q^{i}=\abs{\vect{p}}\,\dd s\) and Equation (16.88) coincides with Equation (16.91); this is Equation (16.89) with \(n=\sqrt{2m\left(E-V\right)}\), in \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\).
By Lemma 16.100 the extremals satisfy \(\dd\left(n\vect{T}\right)/\dd s=\nabla n\). Now \(n=\abs{\vect{p}}=mv\), so \(n\vect{T}=\vect{p}\), and
Since \(\dd/\dd s=v^{-1}\dd/\dd t\) along the motion, the ray equation reads \(v^{-1}\dd\vect{p}/\dd t=-v^{-1}\nabla V\), that is Equation (16.87). Every step is reversible, so the extremals of Equation (16.91) traversed with the stated time law are precisely the Newtonian orbits of energy \(E\).
∎The two principles are Legendre transforms of one another. For a motion of constant energy \(E\),
so making \(S\) stationary at fixed transit time and making \(W\) stationary at fixed energy are the same condition read with different quantities held fixed. The hypotheses must be stated exactly or the equivalence fails: Hamilton's principle varies the path with \(t_{1},t_{2}\) fixed and the energy free, Maupertuis' principle varies the path with \(E\) fixed and the transit time free, and \(E\) must exceed \(V\) so that \(n>0\) and the region is accessible. At a turning point, where \(E=V\), the Jacobi metric degenerates and the ray equation fails.
On the attribution: Maupertuis stated a principle of least action in 1744 for the refraction of light and in 1746 for mechanics, but as a metaphysical assertion about the economy of nature, with a quantity whose definition varied from case to case [Maupertuis:1746]. The mathematically correct statement, with the hypothesis of fixed energy and the integral \(\int mv\,\dd s\), is Euler's, in the second addendum to his founding monograph of the same period [Euler:1744]. Attributing the sharp principle to Maupertuis alone is the customary usage and is wrong; the name is kept here because it is the one every reader will look up, and the correction is recorded rather than the reference silently changed. The geometric reading, that orbits are geodesics of a conformally rescaled metric, is Jacobi's, and it is taken up again with Hamilton–Jacobi theory in Hamilton–Jacobi Theory and the Optical–Mechanical Analogy.
Fermat's principle
In a medium of refractive index \(n\left(\vect{x}\right)\) — dimensionless, equal to \(c/v_{\text{phase}}\) with \(c=299792458\,\mathrm{m}/\mathrm{s}\) the speed of light in vacuum — the optical path length of a curve is \(\mathcal{W}[\gamma]=\int_{\gamma}n\,\dd s\), in metres, and the transit time is \(\mathcal{W}/c\) in seconds. Fermat's principle is the statement that a light ray follows a path making \(\mathcal{W}\) stationary [Fermat:1894]. Rests on Equation (16.89) and Lemma 16.100.
Let two homogeneous media of indices \(n_{1}\) and \(n_{2}\) meet at a plane interface. Then a ray crossing the interface obeys
with \(\vartheta_{1},\vartheta_{2}\) the angles between the ray and the normal to the interface on the two sides, the incident ray, the refracted ray and the normal being coplanar. Rests on Definition 16.103 and Lemma 7.33.
Derives Theorem 16.104. Within each homogeneous medium \(\nabla n=0\), so Equation (16.90) gives \(\dd\vect{T}/\dd s=\vect{0}\): the ray is straight. The path is therefore determined by the single point at which it crosses the interface, and coplanarity follows because moving that point off the plane containing the endpoints and the normal lengthens both straight segments.
Place the interface at \(z=0\), the source at perpendicular distance \(a\) on one side and the target at perpendicular distance \(b\) on the other, their feet separated by \(d\), and let the crossing point be at distance \(u\) from the first foot. The optical length is
a smooth function of the single real variable \(u\), whose stationary points satisfy, by Lemma 7.33,
The two fractions are exactly \(\sin\vartheta_{1}\) and \(\sin\vartheta_{2}\), so the condition is Equation (16.93). Since \(\mathcal{W}''>0\) everywhere, the stationary point is the unique minimum.
∎Let \(n=n(z)\) depend on height alone. Then along a ray confined to a vertical plane,
with \(\vartheta\) the angle to the vertical. Rests on Theorem 16.29 and Equation (16.89).
Derives Proposition 16.105. Parametrise the ray in the vertical plane as \(x\longmapsto z(x)\), so that \(\mathcal{W}=\int f\,\dd x\) with \(f\left(z,z'\right)=n(z)\sqrt{1+z'^{2}}\). The integrand carries no explicit dependence on \(x\), so the Beltrami identity Equation (16.25) applies:
The unit tangent of the ray is \(\left(1,z'\right)/\sqrt{1+z'^{2}}\), so the angle \(\vartheta\) it makes with the vertical satisfies \(\sin\vartheta=\left(1+z'^{2}\right)^{-1/2}\), and the conserved quantity is \(n(z)\sin\vartheta\), which is Equation (16.94).
∎Equation (16.94) is the whole content of atmospheric refraction: warm air near a hot road has a smaller index, so \(\sin\vartheta\) must grow and a ray descending towards the road bends back upwards — the inferior mirage. The same identity governs gradient-index fibres and the apparent flattening of the setting Sun. Fermat's principle is not, however, an independent postulate. It is the short-wavelength limit of wave optics: substituting \(\vect{E}=\vect{E}_{0}\exp\left[\ii k_{0}L(\vect{x})\right]\) into the wave equation and dropping terms of order \(k_{0}^{-1}\) gives the eikonal equation \(\abs{\nabla L}^{2}=n^{2}\), whose characteristics are the rays of Equation (16.90); the derivation is that of Born and Wolf [Born:1999] and belongs to Electromagnetic Waves and Optics. Comparing Lemma 16.100 as used in Theorem 16.101 with its use here shows the optical–mechanical analogy in its sharpest form: a mechanical orbit of energy \(E\) is a light ray in a medium of index \(n=\sqrt{2m\left(E-V\right)}\), and the eikonal \(L\) is Hamilton's characteristic function. That correspondence, pursued in Hamilton–Jacobi Theory and the Optical–Mechanical Analogy, is the historical route to the Schrödinger equation.
Field actions
For fields \(\varphi^{A}\) on the spacetime of Minkowski Space and Its Symmetries the action is
with the Lagrangian density \(\Lag\) in \(\mathrm{J}/\mathrm{m}^{3}\), so that \(S\) carries the dimension \(\mathrm{J}\,\mathrm{s}\) of Equation (16.86). In covariant notation \(\dd^{4}x=c\,\dd t\,\dd^{3}x\), so \(S=c^{-1}\int\Lag\,\dd^{4}x\). Its Euler–Lagrange equations are Equation (16.32). Rests on Equations (16.31) and (16.86).
With the electric field \(\vect{E}\) in \(\mathrm{V}/\mathrm{m}\), the magnetic field \(\vect{B}\) in teslas, the charge density \(\rho\) in \(\mathrm{C}/\mathrm{m}^{3}\) and the current density \(\vect{J}\) in \(\mathrm{A}/\mathrm{m}^{2}\), the Lagrangian density
with \(\varepsilon_{0}=8.8541878128\times 10^{-12}\,\mathrm{F}/\mathrm{m}\) and \(\mu_{0}=1.25663706212\times 10^{-6}\,\mathrm{N}/\mathrm{A}^{2}\), and with \(\vect{E}=-\nabla\Phi-\pp\vect{A}/\pp t\) and \(\vect{B}=\nabla\times\vect{A}\), has the dimension \(\mathrm{J}/\mathrm{m}^{3}\). Its Euler–Lagrange equations in the potentials \(\left(\Phi,\vect{A}\right)\) are the two inhomogeneous Maxwell equations, the other two holding identically because the fields are built from potentials. The derivation, the covariant form
and the gauge analysis are the business of Generalized Classical Field Theory. The sign of the source term is fixed by the signature: in the \(\left(-,+,+,+\right)\) convention of Notation, with \(A^{\mu}=\left(\Phi/c,\vect{A}\right)\) and \(J^{\mu}=\left(c\rho,\vect{J}\right)\), one has \(A_{\mu}=\left(-\Phi/c,\vect{A}\right)\) and \(J^{\mu}A_{\mu}=\vect{J}\cdot\vect{A}-\rho\,\Phi\), which is exactly the last two terms of Equation (16.96); the opposite sign belongs to the mostly-minus convention, which this book does not use (see Equation (16.99) for the same choice). The \(F_{\mu\nu}F^{\mu\nu}\) term is the same in both. What belongs here is only that the Maxwell equations are the Euler–Lagrange equations of a multiple integral, so that everything proved in this chapter — the field equations, the Noether currents, the second theorem's identities — applies to them. Rests on Equations (16.32) and (16.95).
The gravitational action is
with \(x^{0}=ct\), \(R\) the Ricci scalar of Differentiable Manifolds, Tensors, and Curvature in \(/\mathrm{m}^{2}\), \(g\) the determinant of the metric, and \(G=6.67430\times 10^{-11}\,\mathrm{m}^{3}/\mathrm{kg}/\mathrm{s}^{2}\). The prefactor \(c^{3}/G\) has the SI dimension \(\mathrm{kg}/\mathrm{s}\), so \(S_{\text{EH}}\) carries \(\mathrm{J}\,\mathrm{s}\) as an action must; the frequently seen \(c^{4}\) goes with a different convention for \(\dd^{4}x\). Varying Equation (16.98) with respect to the metric gives the Einstein field equations, and varying the matter action gives the energy–momentum tensor that sources them [Hilbert:1915]; both computations, and the boundary term of Remark 16.34 that must be added because the integrand contains second derivatives of the metric, are carried out in The Einstein Field Equations. The diffeomorphism invariance of Equation (16.98) is the instance of Theorem 16.93 discussed in Remark 16.95. Rests on Equation (16.32) and Theorem 16.93.
A free test particle of mass \(m\) in a spacetime with metric \(g_{\mu\nu}\) extremises
in \(\mathrm{J}\,\mathrm{s}\), whose Euler–Lagrange equations are the geodesic equations of Geometric Formulation of Gravity. The integrand is homogeneous of degree one in \(\dd x^{\mu}\), so, exactly as in Lemma 16.100, the functional is parametrisation-invariant and the equations are degenerate until a gauge — here proper time, the analogue of arclength — is fixed. This is the same structure as the abbreviated action of Theorem 16.101: the free relativistic particle is the Maupertuis problem of a metric. Rests on Equations (16.89) and (16.91).
Three things, and they are the reason this chapter exists.
First, covariance for free: by Proposition 16.26 the Euler–Lagrange equations of a scalar action take the same form in every coordinate system, so a theory specified by an action is automatically consistent with whatever symmetry group leaves the action invariant.
Second, conservation laws on demand: by Theorem 16.90 every continuous symmetry of the action delivers a conserved current, computed from the Lagrangian by differentiation alone. Energy, momentum, angular momentum and electric charge are all obtained this way, and no separate postulate is needed for any of them — which is why Section 16.6.1 is the result the rest of this book leans on hardest.
Third, a criterion for admissible theories: the field equations of a Lagrangian theory are not arbitrary. They are second order when the Lagrangian is first order in derivatives, by Equation (16.32); they are unstable if it is genuinely higher order, by Proposition 16.39; and they are redundant, with constraints, precisely when a gauge group acts, by Corollary 16.94. Those three statements between them shape every fundamental theory in this treatise.