Hamilton–Jacobi Theory and the Optical–Mechanical Analogy
The canonical transformations of Section 22.2 were introduced as the transformations that preserve the form of Hamilton's equations. Nothing so far has been asked of them beyond that. Hamilton–Jacobi theory asks the extreme question: is there a canonical transformation that makes the equations of motion trivial—one whose new coordinates and momenta are all constants? The answer is yes, and the generating function that performs it satisfies a single first-order partial differential equation, the Hamilton–Jacobi equation. Solving that one equation solves the mechanical problem completely.
The reward is not merely computational. The function that appears in the equation is the action itself, and its level surfaces propagate through configuration space exactly as the wavefronts of a wave propagate through an optical medium, with the trajectories playing the part of the rays. This is Hamilton's optical–mechanical analogy: he built the characteristic-function method for optics first [Hamilton:1828] and then transplanted it whole into dynamics [Hamilton:1834], and Jacobi turned the transplant into the method of integration that carries both their names [Jacobi:1837]. It is the reason this chapter sits where it does: it is the last statement classical mechanics can make about itself, and the first sentence of quantum mechanics. The route from Equation (23.2) below to the Schrödinger equation of The Postulates of Quantum Mechanics is the shortest one in physics, and it is travelled in Section 23.5. Standard treatments are [Goldstein:2002] [Landau:1976].
The index conventions of Notation 22.1 remain in force: \(a,b=1, \ldots,f\) label the degrees of freedom, and repeated indices are summed. The \(f\) constants of integration that label a complete integral are written \(\alpha_{a}\), and the \(f\) constants conjugate to them \(\beta^{a}\). In the optical sections \(\vect{x}\) denotes a point of physical space and \(s\) the arclength along a ray.
The Hamilton–Jacobi equation
The transformation to constant coordinates
Let \(F_{2}=S(q,P,t)\) be a generating function of the second kind, in the sense of Equation (22.21), and demand of it that the transformed Hamiltonian vanish identically,
Hamilton's equations for the new variables then read \(\dot{Q}^{a}=0\) and \(\dot{P}_{a}=0\): the new coordinates and momenta are \(2f\) constants of the motion. Writing \(P_{a}=\alpha_{a}\) and \(Q^{a}=\beta^{a}\) for those constants, and substituting \(p_{a}=\pp S/\pp q^{a}\) from Equation (22.21) into \(K=\Ham+\pp S/\pp t\), the demand Equation (23.1) becomes a single equation for \(S\).
Hamilton's principal function is a function \(S(q^{a},\alpha_{a},t)\) of the coordinates, of \(f\) constants \(\alpha_{a}\), and of the time, satisfying the Hamilton–Jacobi equation
It carries the SI dimension of an action, \(\mathrm{J}\,\mathrm{s}\), since it generates the momenta by differentiation with respect to the coordinates. Rests on Equation (22.21), Equation (23.1) and Proposition 22.18.
Equation (23.2) is a first-order partial differential equation in the \(f+1\) variables \((q^{a},t)\), nonlinear in the derivatives whenever the Hamiltonian is not linear in the momenta. Its characteristic curves are precisely the trajectories of the mechanical system, which is the reading Remark 10.21 of Partial Differential Equations gives it.
Proposition 10.15 treats the quasilinear first-order equation \(a\,\pp_{x}u+b\,\pp_{y}u=c\), in which the derivatives enter linearly. Equation (23.2) is not of that form: the Hamiltonian is quadratic in \(\pp S/\pp q^{a}\) for every mechanical system of interest. The characteristic theory that does cover it—Cauchy's method of characteristic strips for a fully nonlinear equation \(F(q,S,\pp S/\pp q)=0\), in which the momenta are carried along as independent unknowns and the characteristic system is Hamilton's own—is not in Partial Differential Equations, and neither is the theorem that the envelope of a family of solutions is again a solution, which is how a complete integral generates the general one. Both are ordinary theory of first-order partial differential equations and belong there. This chapter does not use them: Theorem 23.6 below is proved directly from the non-degeneracy of the complete integral, which needs only the implicit function theorem Theorem A.286, and the envelope construction is never invoked. The gap is recorded rather than quietly filled. Rests on Equation (23.2), Proposition 10.15 and Theorem A.286.
The trade is \(2f\) coupled ordinary differential equations for one partial differential equation in \(f+1\) variables. Considered as a method of solution this is usually a bad bargain, and Hamilton–Jacobi theory is not the way most mechanical problems are solved. Its value is structural: it identifies the action as the object that propagates, it is the natural setting for separation of variables and for the action–angle variables of Section 23.3, and it is the classical limit of wave mechanics. Rests on Equation (23.2) and Theorem 22.6.
Jacobi's theorem
A complete integral of Equation (23.2) is a solution \(S(q^{a},\alpha_{a},t)\) depending on \(f\) constants \(\alpha_{a}\), none of which is purely additive, such that
Rests on Equation (23.2).
The constants \(\alpha_{a}\) are \(f\), not \(f+1\): Equation (23.2) contains \(S\) only through its derivatives, so an additive constant is always available and carries no information, which is what the exclusion of a purely additive constant says.
Let \(S(q^{a},\alpha_{a},t)\) be a complete integral of the Hamilton–Jacobi equation Equation (23.2). Then the \(2f\) relations
with \(\alpha_{a}\) and \(\beta^{a}\) arbitrary constants, define implicitly a \(2f\)-parameter family of curves \(\left(q^{a}(t),p_{a}(t)\right)\), and every such curve is a solution of Hamilton's equations Equations (22.3) and (22.4). Conversely every solution of Hamilton's equations is obtained in this way for a suitable choice of the constants. Rests on Definition 23.5, Theorem 22.6 and Theorem A.286.
Derives Theorem 23.6. The curves exist. Fix \(\alpha\) and \(\beta\) and apply the implicit function theorem Theorem A.286 to the map \(F_{b}(t,q):=\pp S/\pp\alpha_{b}(q,\alpha,t)-\beta^{b}\), whose derivative with respect to \(q\) is the matrix \(\pp^{2}S/\pp\alpha_{b}\pp q^{a}\), nonsingular by Equation (23.3). Near any solution point Equation (23.4) therefore determines \(q^{a}=q^{a}(\alpha,\beta, t)\) uniquely and differentiably, and Equation (23.5) then defines \(p_{a}(t)\) along that curve.
The first half of Hamilton's equations. Differentiate Equation (23.4) totally with respect to \(t\), holding \(\alpha\) and \(\beta\) fixed:
Now differentiate the Hamilton–Jacobi equation Equation (23.2) itself with respect to \(\alpha_{b}\), at fixed \(q\) and \(t\). The Hamiltonian depends on \(\alpha_{b}\) only through its momentum slots, so the chain rule Proposition 7.72 gives
Mixed partial derivatives commute (Proposition 7.73), so subtracting Equation (23.7) from Equation (23.6) leaves
and the matrix is nonsingular by Equation (23.3). Hence \(\dot{q}^{a}=\pp\Ham/\pp p_{a}\), which is Equation (22.3).
The second half. Differentiate Equation (23.5) totally with respect to \(t\):
Differentiate Equation (23.2) with respect to \(q^{a}\), now at fixed \(\alpha\) and \(t\), remembering that \(\Ham\) depends on \(q^{a}\) both directly and through the momenta:
Substituting \(\dot{q}^{b}=\pp\Ham/\pp p_{b}\), just proved, into Equation (23.9) and adding Equation (23.10) term by term, every second derivative of \(S\) cancels and
which is Equation (22.4).
The converse. Let \(\left(q^{a}(t),p_{a}(t)\right)\) solve Hamilton's equations with data \((q_{0},p_{0})\) at \(t_{0}\). By Equation (23.3) and Theorem A.286 again, the \(f\) equations \(p_{0a}=\pp S/\pp q^{a}(q_{0},\alpha,t_{0})\) determine \(\alpha\) near any point at which they hold; set \(\beta^{b}:=\pp S/\pp\alpha_{b}(q_{0},\alpha,t_{0})\). The curve the theorem produces from that \((\alpha,\beta)\) and the given solution obey the same first-order system with the same data at \(t_{0}\), so they coincide by the uniqueness half of Theorem 9.8.
∎The theorem is what makes the trade of Remark 23.4 worth considering: one solution of Equation (23.2) carrying \(f\) constants, and no other information, yields the complete solution of the mechanical problem by differentiation and elimination alone—no further integration is performed.
Along a solution of the equations of motion,
Hamilton's principal function is therefore the action of Calculus of Variations, regarded no longer as a functional of paths but as a function of the endpoint. Rests on Theorem 23.6, Equation (22.1) and Definition 16.96.
Derives Proposition 23.7. Along a solution the constants \(\alpha_{a}\) are fixed, so the total time derivative of \(S\left(q(t),\alpha,t\right)\) is
From Equations (23.2) and (23.5) (the first term from the momenta of the complete integral, the second from the equation itself). The second equality uses Equation (23.5) for the first term and Equation (23.2) for the second. The right-hand side is the inverse Legendre transformation Equation (22.1) of Hamiltonian Mechanics, that is the Lagrangian. Integrating between two times gives the second equality of Equation (23.12), the constant being the value of \(S\) at the lower limit.
∎Proposition 23.7 is the pivot of the whole chapter and it is easy to read past. In Calculus of Variations the action Equation (16.86) is a number assigned to an arbitrary path; here it is a field on configuration space, obtained by evaluating that number on the one extremal path that actually reaches the point \(q\) at time \(t\). The two readings are related exactly as Hilbert's invariant integral Theorem 16.63 is related to the functional it is built from: given a field of extremals in the sense of Definition 16.62, the value of the action along the member through each point defines a function whose gradient is the momentum, and Equation (23.5) is that statement. Rests on Proposition 23.7, Theorem 16.63 and Definition 16.62.
Time-independent systems and the characteristic function
If the Hamiltonian does not depend explicitly on the time, the Hamilton–Jacobi equation admits solutions of the form
where \(E=\alpha_{1}\) is the conserved value of the Hamiltonian and the function \(W\), called Hamilton's characteristic function, satisfies the time-independent Hamilton–Jacobi equation
Rests on Equation (23.2), Theorem 22.10 and Lemma 10.28.
Derives Proposition 23.9. Try the additive ansatz \(S(q,\alpha,t)=W(q,\alpha)-h(t)\), which is the general form in which the time enters additively. Then \(\pp S/\pp q^{a}= \pp W/\pp q^{a}\) and \(\pp S/\pp t=-h'(t)\), so Equation (23.2) reads
The left-hand side depends on \(q\) alone—this is where the absence of an explicit time dependence in \(\Ham\) is used—and the right-hand side on \(t\) alone, so by Lemma 10.28 both equal a constant. Call it \(E\); then \(h(t)=Et\) up to the additive constant excluded by Definition 23.5, and Equation (23.16) is Equation (23.15). That the constant is the conserved value of the Hamiltonian is immediate from Equation (23.15) and Equation (23.5), and the conservation itself is Theorem 22.10. Since \(E\) is one of the constants of the complete integral it may be named \(\alpha_{1}\).
∎The characteristic function is the abbreviated, or Maupertuis, action
and the variational principle it satisfies at fixed energy,
is Maupertuis' principle. For a Hamiltonian of the form \(\Ham=\abs{\vect{p}}^{2}/2m+V(\vect{x})\) it reads
\(\dd s\) being the element of arclength in configuration space. Rests on Proposition 23.9, Definition 16.99 and Theorem 16.101.
Derives Proposition 23.10. By Equation (23.14) and Proposition 23.7,
since \(\Ham=E\) along the motion; the last integral is \(\int p_{a}\,\dd q^{a}\), which is Equation (23.17) and is exactly the abbreviated action Equation (16.88) of Definition 16.99. The variational principle Equation (23.18) for that functional, with the energy fixed and the transit time free, is Theorem 16.101, which also supplies the metric form: for \(\Ham=\abs{\vect{p}}^{2}/2m+V\) one has \(\abs{\vect{p}}=\sqrt{2m(E-V)}\) and \(p_{a}\dd q^{a}=\abs{\vect{p}}\dd s\), so Equation (23.18) becomes Equation (16.91), which is Equation (23.19).
∎Two cautions carried over from Remark 16.102 and not to be lost here. The sharp statement of Equation (23.18), with the hypothesis of fixed energy, is Euler's [Euler:1744], not Maupertuis' [Maupertuis:1746]; the name is customary and is kept, the correction recorded. And Equation (23.19) requires \(E>V\) throughout the region: at a turning point the integrand vanishes, the Jacobi metric degenerates, and the principle says nothing. That degeneracy is not a technicality—it is the same failure that Section 23.5.2 identifies as the place where the whole geometrical description breaks down. Rests on Proposition 23.10 and Theorem 16.101.
\(S\) generates a transformation to variables that are constant; \(W\) generates a transformation to variables in which the new Hamiltonian is \(E\), so that the new momenta are constant and the new coordinates advance linearly in time. The second is the useful one for bounded motion, and it is what Section 23.3 builds on. Rests on Definition 23.2 and Proposition 23.9.
Complete integrals and separation of variables
The Hamilton–Jacobi equation is completely separable in the coordinates \(q^{a}\) if it admits a complete integral of the form
each summand depending on a single coordinate. Rests on Definition 23.5 and Equation (23.15).
Separability turns one partial differential equation into \(f\) ordinary ones, each solvable by a quadrature: from Equation (23.21) the momentum \(p_{a}=\dd W_{a}/\dd q^{a}\) is a function of its own coordinate and of the constants, so each \(W_{a}\) is an explicit integral. This is why every mechanical problem that can be “solved in closed form” is, in practice, a separable one.
If \(q^{1}\) is cyclic—that is, absent from the Hamiltonian, in the sense of Definition 21.35—then the characteristic function separates in it as
and \(\alpha_{1}=p_{1}\) is the conserved momentum conjugate to \(q^{1}\). Rests on Definition 23.13, Definition 21.35 and Theorem 21.36.
Derives Proposition 23.14. Substitute Equation (23.22) into Equation (23.15). Since \(\pp W/\pp q^{1}=\alpha_{1}\) and \(q^{1}\) occurs nowhere in \(\Ham\), the resulting equation
contains neither \(q^{1}\) nor any derivative of \(W'\) with respect to it: it is an equation for \(W'\) in \(f-1\) variables, and the ansatz is consistent. Conversely Equation (23.5) gives \(p_{1}=\pp W/\pp q^{1}=\alpha_{1}\), a constant, in agreement with Theorem 21.36.
∎Conditions for separability
Separability is a property of the pair (system, coordinate system), never of the system alone. The Kepler problem separates in spherical coordinates, in parabolic coordinates and in a rotated Cartesian system not at all; the harmonic oscillator separates in Cartesian coordinates and not in polar ones unless it is isotropic. The question this subsection answers is which coordinate systems work, for the class of Hamiltonians in which the kinetic energy is diagonal.
A Hamiltonian is orthogonal in the coordinates \(q^{a}\) if
with no cross terms \(p_{a}p_{b}\), \(a\neq b\). For a single particle of mass \(m\) moving in ordinary space with orthogonal curvilinear coordinates of scale factors \(h_{a}\), so that \(\dd s^{2}=\sum_{a}h_{a}^{2} \left(\dd q^{a}\right)^{2}\), one has \(g^{aa}=1/\left(m h_{a}^{2}\right)\). The SI dimension of \(g^{aa}\) depends on what dimension the coordinate \(q^{a}\) carries—\(/\mathrm{kg}\) when \(q^{a}\) is a length, \(/\mathrm{kg}/\mathrm{m}^{2}\) when it is an angle—and in every case \(g^{aa}p_{a}^{2}\) is an energy. Rests on Definitions 13.117 and 22.2.
A Stäckel matrix for the coordinates \(q^{a}\) is an \(f\times f\) matrix \(\Phi=\left(\phi_{ab}\right)\) whose \(a\)-th row depends on the single coordinate \(q^{a}\), with
The orthogonal Hamiltonian Equation (23.23) is a Stäckel system in those coordinates if there exist such a \(\Phi\) and \(f\) functions \(u_{a}(q^{a})\), each of one coordinate, with
Rests on Equation (23.23).
Equation (23.25) is a strong pair of requirements and it is worth reading twice. The first says the diagonal of the inverse metric is the first row of the inverse of a matrix built out of one-dimensional functions; it constrains the coordinate system alone. The second says the potential is a combination of one-dimensional potentials with exactly those coefficients; it constrains the system. Dimensionally, \(\phi_{a1}\) carries the reciprocal dimension of \(g^{aa}\) and \(u_{a}\) the dimension of \(p_{a}^{2}\), so that every term of Equation (23.27) below is a squared momentum.
If Equation (23.25) holds, then Equation (23.15) separates in the coordinates \(q^{a}\) and
is a complete integral of it, with \(\alpha_{1}=E\). Rests on Definition 23.16, Definition 23.13 and Equation (23.15).
Derives Theorem 23.17. The summand of Equation (23.26) depends on the single coordinate \(q^{a}\), so \(W\) has the separated form Equation (23.21) and
Insert this and Equation (23.25) into Equation (23.23). The potential terms cancel against the \(-2u_{a}\) and
From Equations (23.25) and (23.27) (the first row of the inverse matrix, times the matrix itself). The middle step is nothing but \(\Phi^{-1}\Phi=\identity\) read along its first row. So Equation (23.15) holds with \(E= \alpha_{1}\), identically in \(q\) and in the remaining constants.
It remains to check the non-degeneracy Equation (23.3). Differentiating Equation (23.27) with respect to \(\alpha_{b}\) gives \(\pp^{2}W/\pp q^{a}\pp\alpha_{b}=\phi_{ab}/W_{a}'\), that is the matrix \(\Phi\) with its \(a\)-th row divided by \(W_{a}'\). Hence
which is nonzero wherever no \(W_{a}'\) vanishes, by Equation (23.24). The zeros of \(W_{a}'\) are exactly the turning points of the \(a\)-th degree of freedom, and they are precisely where the construction is expected to fail; that failure is the subject of Section 23.5.2.
∎The implication of Theorem 23.17 runs the other way as well: if the orthogonal Hamiltonian Equation (23.23) admits a complete integral of the separated form Equation (23.21) in the coordinates \(q^{a}\), then a Stäckel matrix and functions \(u_{a}\) satisfying Equation (23.25) exist, so that separability of an orthogonal Hamiltonian is equivalent to the Stäckel conditions. That converse is Stäckel's own theorem in the form Levi-Civita gave it, and it is stated here on the authority of [Goldstein:2002]. This treatise does not prove it, and nothing in this chapter or later in the book uses it: every application runs through Theorem 23.17, in the direction that is proved.
It is worth saying where the difficulty sits, because the first two steps are easy and mislead. Write \(x_{a}:=\tfrac{1}{2}\left(W_{a}'\right)^{2}\), a function of the single coordinate \(q^{a}\) and of the constants. Differentiating the separated equation \(\sum_{a}g^{aa}x_{a}+V=\alpha_{1}\) with respect to \(\alpha_{b}\) gives \(\sum_{a}g^{aa}\phi_{ab}=\delta_{1b}\) with \(\phi_{ab}:=\pp x_{a}/\pp\alpha_{b}\), and Equation (23.29) identifies \(\det\Phi\neq0\) with the completeness Equation (23.3); so \(g^{aa}=\left(\Phi^{-1}\right)_{1a}\) already. What is not immediate is that \(\Phi\) may be taken independent of the constants—that is, that \(x_{a}\) is affine in \(\alpha\), which is exactly what makes \(\phi_{ab}\) a function of \(q^{a}\) alone and \(u_{a}:=\sum_{b}\alpha_{b}\phi_{ab}-x_{a}\) a function of \(q^{a}\) alone. A second differentiation returns only \(\sum_{a}g^{aa}\,\pp\phi_{ab}/\pp\alpha_{c}=0\), which says no more than that \(g^{aa}\) does not depend on the constants, and admits solutions with \(\pp\Phi/\pp\alpha\neq0\). Removing those requires a change of the separation constants \(\alpha\mapsto\tilde{\alpha}(\alpha)\), under which \(\Phi\mapsto\Phi\,\pp\alpha/\pp\tilde{\alpha}\), and it is the proof that such a change always exists—not the index bookkeeping—that this treatise leaves to its source. Rests on Theorem 23.17 and Definition 23.16.
Theorem 10.34 of Partial Differential Equations states that the Helmholtz equation separates in exactly eleven orthogonal coordinate systems of ordinary space, and that the criterion is the Stäckel condition on the scale factors; the derivation is The Stäckel Conditions and the Eleven Separable Systems, and Remark 10.35 there points forward to this chapter. The two statements are not the same statement, and the difference is instructive. The Helmholtz equation is linear and of second order and its separation constants are eigenvalues; Equation (23.15) is nonlinear and of first order and its separation constants are values of conserved quantities. What they share is Equation (23.25)'s first half, which involves only the coordinate system: the eleven systems that separate the Helmholtz equation are exactly the eleven that separate the free-particle Hamilton–Jacobi equation, because for \(V=0\) the second half of Equation (23.25) is vacuous. Adding a potential can only restrict the list further, never extend it. The relation between the two is made precise by the semiclassical expansion of Section 23.5.1, in which the first-order equation is the leading order of the second-order one. Rests on Theorems 10.34 and 23.17.
Let Equation (23.15) separate in the coordinates \(q^{a}\), with complete integral Equation (23.21). Then the \(f\) functions \(\alpha_{a}(q,p)\) obtained by solving Equation (23.5) for the constants are functionally independent constants of the motion in involution,
so a separable system is completely integrable in the sense of Theorem 23.31. For a Stäckel system they are
each quadratic in the momenta, and \(\alpha_{1}=\Ham\). Rests on Definition 23.13, Equation (22.45) and Equation (22.44).
Derives Proposition 23.20. \(W(q,\alpha)\) is a generating function of the second kind (Proposition 22.18) with old momenta \(p_{a}=\pp W/\pp q^{a}\) and new momenta \(P_{a}=\alpha_{a}\); the non-degeneracy Equation (23.3) is exactly what makes the transformation invertible, so the \(\alpha_{a}\) are genuine functions on phase space and are functionally independent. They are constants of the motion by Theorem 23.6. The Poisson bracket is invariant under a canonical transformation, Equation (22.45), so it may be computed in the new variables, where the fundamental relations Equation (22.44) give \(\pb{P_{a}}{P_{b}}_{Q,P}=0\); that is Equation (23.30).
For the explicit form, Equation (23.27) with \(p_{a}=W_{a}'\) reads \(\tfrac{1}{2}p_{a}^{2}+u_{a}(q^{a})=\sum_{b}\alpha_{b}\phi_{ab} (q^{a})\); multiplying by \(\left(\Phi^{-1}\right)_{ca}\) and summing over \(a\) inverts it into Equation (23.31). Setting \(c=1\) and using Equation (23.25) recovers Equation (23.23), so \(\alpha_{1}=\Ham\).
∎Equation (23.31) exhibits each separation constant as a quadratic form in the momenta,
That such a quantity is conserved along every trajectory of Equation (23.23) with \(V=0\) is, in the language of Differentiable Manifolds, Tensors, and Curvature, the statement that \(K_{b}^{\ cd}\) is a Killing tensor of rank two: a symmetric tensor field obeying \(\nabla^{(e}K_{b}^{\ cd)}=0\), the rank-two analogue of Killing's equation Equation (13.290). Separability is then the geometric condition that the configuration-space metric admit \(f\) commuting Killing tensors with simultaneously diagonalisable, pointwise-distinct eigenvectors—the coordinate system being the one whose surfaces those eigenvectors are normal to.
Part II does not carry this. Section 13.9.2 defines Killing vectors Definition 13.139 and derives Killing's equation Proposition 13.140, and stops there: there is no rank-two Killing tensor, no statement that \(\xi_{\mu}p^{\mu}\) or \(K_{\mu\nu}p^{\mu}p^{\nu}\) is constant along a geodesic, and no integrability condition for either. The definitions above are therefore stated here for use and are not treated as established; every result this chapter actually proves—Theorem 23.17, Proposition 23.20—is proved from Equation (23.25) directly, in coordinates, and needs none of it. The tensorial reading is what makes Carter's constant of Section 23.4.3 intelligible, and that is where the absence is felt. Rests on Propositions 13.140 and 23.20.
Worked separations
Six complete integrals, in the order of increasing structure. Each ends by recovering, from Theorem 23.6 alone, a result this treatise has already obtained by other means—which is the only honest way to test a method of this generality.
For \(\Ham=\abs{\vect{p}}^{2}/2m\) in Cartesian coordinates all three coordinates are cyclic, so Proposition 23.14 applies three times and
with \(\vect{\alpha}\) in \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\) and \(E= \abs{\vect{\alpha}}^{2}/2m\). Equation (23.4) gives \(\vect{\beta}=\pp S/\pp\vect{\alpha}=\vect{x}-\vect{\alpha}t/m\), that is
uniform rectilinear motion with \(\vect{p}=\vect{\alpha}\) constant by Equation (23.5). The surfaces \(S=\text{const}\) are the planes perpendicular to \(\vect{\alpha}\), advancing at speed \(E/\abs{\vect {\alpha}}=\abs{\vect{\alpha}}/2m\)—which is half the particle speed. That factor of two is not an error and is taken up in Remark 23.42. Rests on Proposition 23.14 and Theorem 23.6.
Take motion in a vertical plane with \(\Ham=\left(p_{x}^{2}+p_{z}^{2}\right)/2m+mgz\) and \(g=9.80665\,\mathrm{m}/\mathrm{s}^{2}\), the standard value fixed in Experiment: Free Fall and Projectile Motion. Here \(x\) is cyclic, so by Propositions 23.9 and 23.14
Write \(\alpha_{z}:=E-\alpha_{x}^{2}/2m\) for the energy of the vertical motion. Equation (23.4) with respect to \(\alpha_{x}\) and to \(\alpha_{z}\) gives
which is the parabolic trajectory derived from Newton's second law in Experiment: Free Fall and Projectile Motion, with the initial speeds \(u_{x}=\alpha_{x}/m\) and \(u_{z}=\sqrt{2\alpha_{z}/m}\) identified as the components of the launch velocity. The wavefronts \(W=\text{const}\) are the curves orthogonal to the family of parabolas of fixed energy—the envelope of that family being the safety parabola, a caustic in the sense of Section 23.5.2. Rests on Proposition 23.14 and Theorem 23.6.
For \(\Ham=p^{2}/2m+\tfrac{1}{2}m\omega_{0}^{2}x^{2}\) (Definition 28.3), Equation (23.15) gives
whose \(\beta=\pp W/\pp E-t\) reduces by the substitution \(x=\sqrt{2E/m\omega_{0}^{2}}\,\sin\varphi\) to \(\arcsin\left(x\sqrt{m\omega_{0}^{2}/2E}\right)=\omega_{0}(t+\beta)\), the solution Equation (28.4). The action variable Equation (23.49) is the phase-plane area of Proposition 28.6 divided by \(2\pi\),
in \(\mathrm{J}\,\mathrm{s}\), so the Hamiltonian in action–angle variables is simply \(\Ham(I)=\omega_{0}I\) and Equation (23.52) returns \(\dot{\theta}=\omega_{0}\): the frequency, read off without solving anything. This is the doorway of Remark 28.7. Rests on Proposition 28.6, Definition 23.28 and Theorem 23.30.
With \(\Ham=\left(p_{r}^{2}+p_{\varphi}^{2}/r^{2}\right)/2m+V(r)\) the coordinate \(\varphi\) is cyclic, so
with \(\alpha_{\varphi}=p_{\varphi}=L\) the angular momentum, in \(\mathrm{J}\,\mathrm{s}\). The two relations Equation (23.4) are
Equation (23.40) is the radial quadrature of Section 27.3, with \(\sqrt{2m(E-V)-L^{2}/r^{2}}/m\) the radial velocity Equation (27.34) and the bracket the effective potential Equation (27.32); Equation (23.41) is the orbit equation Equation (27.38). Nothing new is obtained, and that is the point: the two quadratures that Central Forces and Statics reaches by eliminating the time between two conservation laws fall out here as the two components of a single differentiation. Rests on Proposition 23.14, Theorem 23.6 and Equation (27.38).
In spherical coordinates with \(V=-k/r\) and \(k=e^{2}/4\pi\varepsilon_{0}\) in \(\mathrm{J}\,\mathrm{m}\) the Hamiltonian is \(\Ham=\left[p_{r}^{2}+p_{\theta}^{2}/r^{2}+p_{\varphi}^{2}/(r^{2} \sin^{2}\theta)\right]/2m-k/r\), and \(W=\alpha_{\varphi}\varphi+W_{\theta} (\theta)+W_{r}(r)\) separates it into
\(\alpha_{\theta}=L\) being the total angular momentum and \(\alpha_{\varphi}=L_{z}\) its component along the polar axis. For bounded motion, \(E<0\), the three circuits give action variables
the last two by the elementary quadratures \(\int_{\theta_{1}}^{\pi-\theta_{1}}\sqrt{L^{2}-L_{z}^{2}/\sin^{2}\theta} \,\dd\theta=\pi\left(L-\abs{L_{z}}\right)\) and \(\int_{r_{1}}^{r_{2}}r^{-1}\sqrt{-A^{2}r^{2}+Br-C}\,\dd r =\pi\left(B/2A-\sqrt{C}\right)\), both evaluated by the residue theorem Theorem 8.24 with branch points at the turning points. Adding the three,
so the energy depends on the three actions only through their sum: the Kepler problem is completely degenerate, which is why its orbits close. Imposing \(I_{a}=n_{a}\hbar\) (Proposition 23.38, in the uncorrected form the old quantum theory used) with \(n=n_{r}+n_{\theta} +n_{\varphi}\) gives
which is Equation (70.4): with \(R_{\infty}=1.0973731568\times 10^{7}\,/\mathrm{m}\) the ground state lies \(13.605693\,\mathrm{eV}\) below the ionization limit [Mohr:2025]. This is the single computation that made Hamilton–Jacobi theory a working tool of atomic physics for a decade, and the reason Atomic Models and Spectra cites this chapter. Rests on Definition 23.13, Equation (23.49) and Equation (70.4).
Let a particle of charge \(q\) and mass \(m\) move in a uniform field \(\vect{B}=B\hat{\vect{z}}\), described in the gauge \(\vect{A}=\left(0,Bx,0\right)\), so that \(\Ham=\left[p_{x}^{2}+\left(p_{y}-qBx\right)^{2}+p_{z}^{2}\right]/2m\). Both \(y\) and \(z\) are cyclic and
The transverse motion is bounded between the two roots of the radicand and its action variable is, substituting \(u=\alpha_{y}-qBx\) so that the circuit becomes a circle of radius \(\sqrt{2mE_{\perp}}\) in the \(u\)-plane,
with \(\omega_{\text{c}}=\abs{q}B/m\) the cyclotron angular frequency in \(\mathrm{rad}/\mathrm{s}\). For an electron in \(B=1.00\,\mathrm{T}\) this is \(\omega_{\text{c}}=1.7588\times 10^{11}\,\mathrm{rad}/\mathrm{s}\), from \(e/m_{\text{e}}\) built out of Physical Constants and SI Units. The magnetic moment of the gyrating charge is
in \(\mathrm{J}/\mathrm{T}\), and since \(\abs{q}/m\) is fixed, \(\mu\) is an adiabatic invariant exactly when \(I_{x}\) is—which Theorem 23.34 establishes. That identification is what Remark 23.37 rests on. Rests on Proposition 23.14, Equation (23.49) and Theorem 23.34.
Action–angle variables and integrable systems
For motion that is bounded, the useful canonical variables are not the constants \(\alpha_{a}\) themselves but a particular set of functions of them, chosen so that the conjugate coordinates are angles.
Let the motion be bounded and let the Hamilton–Jacobi equation separate in the coordinates \(q^{a}\), so that each pair \((q^{a},p_{a})\) traverses a closed curve \(C_{a}\) in its own plane of phase space. The action variables are
and the conjugate angle variables are
The factor \(1/2\pi\) in Equation (23.49) makes the angles \(2\pi\)-periodic and their rates ordinary angular frequencies. The older literature, and the quantization condition Equation (23.61) below in its original form, use \(J_{a}=\oint p_{a}\,\dd q^{a}=2\pi I_{a}\) with angles of period \(1\). Both appear in the sources of this treatise; the convention here is the first. Rests on Definition 23.28.
In action–angle variables the Hamiltonian depends on the actions alone, \(\Ham=\Ham(I)\), and the equations of motion integrate immediately:
The \(\omega^{a}\) are the angular frequencies of the motion, obtained without solving for the trajectory: one circuit of \(C_{b}\) advances \(\theta^{b}\) by \(2\pi\) and every other angle not at all, so the period of the \(b\)-th degree of freedom is \(T_{b}=2\pi/\omega^{b}\). Rests on Definition 23.28, Proposition 22.18 and Theorem 22.6.
Derives Theorem 23.30. By Equation (23.21) each momentum \(p_{a}=\dd W_{a}/\dd q^{a}\) is a function of \(q^{a}\) and of the constants \(\alpha\) alone, and the circuit \(C_{a}\) is fixed once \(\alpha\) is. Hence Equation (23.49) defines \(f\) functions \(I_{a}(\alpha)\), which are constants of the motion, and which may be inverted for \(\alpha\) wherever \(\det\left(\pp I_{a}/\pp\alpha_{b}\right)\neq0\). Substituting \(\alpha=\alpha(I)\) into \(W\) gives \(W(q,I)\), a generating function of the second kind (Proposition 22.18) with old momenta \(p_{a}=\pp W/\pp q^{a}\) and new coordinates \(\theta^{a}=\pp W/\pp I_{a}\), which is Equation (23.50).
\(W\) carries no explicit time dependence, so by Equation (22.21) the new Hamiltonian is the old one re-expressed; by Equation (23.15) that value is \(E=\alpha_{1}(\alpha)\), a function of the constants alone, hence of the actions alone. So \(\Ham=\Ham(I)\). Theorem 22.6 in the new variables then gives \(\dot{I}_{a}=-\pp\Ham/\pp\theta^{a}=0\), since \(\Ham\) does not contain the angles, and \(\dot{\theta}^{a}=\pp\Ham/\pp I_{a}\), a function of constants and therefore itself constant; integrating gives Equation (23.52).
For the periods, carry \(q\) once around \(C_{b}\) at fixed \(I\). Since \(\theta^{a}=\pp W/\pp I_{a}\) and the order of the two differentiations may be exchanged,
From Equations (23.49) and (23.50) (exchanging the two differentiations under the circuit integral). Comparing with Equation (23.52), one circuit takes a time \(T_{b}=2\pi/\omega^{b}\).
∎Let a system with \(f\) degrees of freedom possess \(f\) functionally independent constants of motion \(F_{1}=\Ham,F_{2},\ldots,F_{f}\) in involution, that is with
Then each compact connected component of a regular level set \(\set{F_{a}=c_{a}}\) is diffeomorphic to the \(f\)-torus \(T^{f}\), the motion on it is quasi-periodic with \(f\) frequencies, and there exist action–angle variables in a neighbourhood of it. Such a system is called completely integrable. Rests on Equation (22.41), Theorem 23.30 and Theorem 24.21.
Full derivation in Appendix A.
Derives Theorem 23.31.
The appendix argument is the classical one. The \(f\) commuting integrals generate \(f\) Hamiltonian vector fields that commute, are everywhere independent on a regular level set and are tangent to it; on a compact level set they are complete, so their joint flow is an action of \(\R^{f}\), which is transitive on each connected component because its orbits are open. The stabiliser is then a discrete subgroup of \(\R^{f}\) with compact quotient, hence a lattice of rank \(f\), and the component is \(\R^{f}\) modulo that lattice—the torus. The actions are the periods of \(p_{a}\,\dd q^{a}\) over a basis of its cycles, and the angles are conjugate to them.
One ingredient is imported rather than proved: that a discrete subgroup of \(\R^{f}\) with compact quotient is a lattice of full rank—equivalently, that a compact connected manifold carrying \(f\) commuting, independent, complete vector fields is diffeomorphic to \(T^{f}\). That is a theorem of the kind Part II owns and Part II does not carry it, so the debt is recorded here rather than discharged inside a physics appendix. Nothing this chapter proves depends on Theorem 23.31: it is used only through the statement that separable systems are integrable, and Proposition 23.20 establishes that on its own. Rests on Theorem 23.31 and Proposition 23.20.
The systems for which Theorem 23.31 applies form a set of measure zero among Hamiltonian systems. What happens to the invariant tori when an integrable system is perturbed—most of them survive, deformed, and the rest break up into the thin chaotic layers that carry the system's sensitivity to initial conditions—is the subject of KAM theory, treated in Section 32.4.2 and stated as Theorem 32.51. The observational foothold is the solar system: the Kirkwood gaps of the asteroid belt (Table 32.2) sit at the mean-motion resonances where the tori fail, and the \(3\!:\!1\) gap was shown to be chaotic rather than merely resonant by direct integration [Wisdom:1983]. Rests on Theorems 23.31 and 32.51.
Adiabatic invariance
Let a bounded one-dimensional system have Hamiltonian \(\Ham\left(q,p;\lambda\right)\) depending on a parameter \(\lambda(t)\) that varies slowly, in the sense that its fractional change over one period \(T\) of the motion is small,
Then the mean rate of change of the action variable \(I\) over one period vanishes,
so that \(I\) changes by \(\mathcal{O}(\varepsilon)\), not \(\mathcal{O}(1)\), over the whole time \(\sim T/\varepsilon\) in which \(\lambda\) changes appreciably. The energy, by contrast, is not conserved: it follows \(\lambda\) through \(E=E(I,\lambda)\) at fixed \(I\). Rests on Definition 23.28, Equation (22.10) and Theorem 23.30.
Derives Theorem 23.34. Freeze \(\lambda\) and let \(I(E,\lambda)=\left(2\pi\right)^{-1}\oint p\,\dd q\) be taken over the closed orbit \(\Ham(q,p;\lambda)=E\), with \(p=p(q,E, \lambda)\) the solution of that equation. Two partial derivatives are needed. Differentiating \(\Ham\left(q,p;\lambda\right)=E\) at fixed \(q\) and \(\lambda\) gives \(\left(\pp\Ham/\pp p\right)\pp p/\pp E=1\), and \(\pp\Ham/\pp p=\dot{q}\) by Equation (22.3), so
Differentiating the same relation at fixed \(q\) and \(E\) gives \(\pp\Ham/\pp\lambda+\dot{q}\,\pp p/\pp\lambda=0\), whence
the angle brackets denoting the average over one period of the frozen motion.
Along the true motion the energy is not conserved: by Equation (22.10), \(\dot{E}=\pp\Ham/\pp t=\dot{\lambda}\,\pp\Ham/\pp\lambda\). Hence
Both terms are of order \(\varepsilon\) separately. Averaging Equation (23.59) over one period, during which \(\dot{\lambda}\) may be held fixed at the cost of an error of relative order \(\varepsilon\), replaces \(\pp\Ham/\pp\lambda\) in the first term by its mean, and \(T=2\pi/\omega\) makes the two terms cancel identically:
This is Equation (23.56). Summing over the \(\sim1/\varepsilon\) periods of the slow evolution, the surviving change in \(I\) is the accumulated remainder, of order \(\varepsilon^{2}\times\varepsilon^{-1}=\varepsilon\), whereas an uncompensated first-order drift would have accumulated to order one. The energy statement is immediate: inverting \(I=I(E,\lambda)\) at fixed \(I\) makes \(E\) a function of \(\lambda\), and \(\pp E/\pp I=\omega\) by Equation (23.57), so \(E\) tracks \(\lambda\) through the frequency.
∎For the oscillator of Example 23.24 with a slowly varying frequency \(\omega_{0}(t)\), Equation (23.38) gives \(I=E/\omega_{0}\), so Theorem 23.34 says \(E/\omega_{0}\) is conserved and therefore \(E\propto\omega_{0}\): shortening a pendulum quasi-statically raises its energy in proportion to its frequency, while its amplitude falls as \(\omega_{0}^{-1/2}\). The energy is supplied by whatever does the shortening, and no amount of care in doing it slowly avoids that. Rests on Theorem 23.34 and Equation (23.38).
Theorem 23.34 is a first-order statement and is far weaker than what is true. If \(\lambda(t)\) is analytic and varies over a finite interval, returning to a constant at both ends, the total change in \(I\) is not \(\mathcal{O}(\varepsilon)\) but exponentially small, of order \(\ee^{-c/\varepsilon}\)—which is why adiabatic invariants are useful in practice rather than merely asymptotically, and why leaving the statement out would misrepresent how good the invariance is.
That sharper statement is not established in this treatise, and it is recorded as owed rather than asserted, for two reasons. No source in this bibliography states it, so it cannot honestly be cited. And the method it needs—a sequence of averaging transformations, each removing one more order in \(\varepsilon\) at the cost of shrinking a strip of analyticity, with Cauchy estimates controlling the remainder and the number of steps optimised against the shrinkage—rests on analytic machinery that Part II does not develop; the size of the exponent is then fixed by the distance from the real axis to the nearest complex singularity of \(\lambda\). Nothing in the book uses it: every application in this chapter and in Plasmas and Magnetohydrodynamics needs only the first-order result Equation (23.56) proved above. Rests on Theorem 23.34.
Adiabatic invariance is the reason the robust quantity of a slowly modulated oscillator is its action and not its energy. The magnetic moment Equation (23.48) of a charged particle gyrating in a slowly varying magnetic field is an adiabatic invariant of exactly this kind: \(\mu=E_{\perp}/B\) is fixed, so as a particle spirals into a region of stronger field its transverse energy must rise, and—the total energy being conserved in a static field—its parallel energy must fall until it reverses. That is the magnetic mirror, and it is what traps particles in the belts Van Allen identified from the first satellite instruments [VanAllen:1959]. The same invariant governs betatron acceleration and the confinement schemes of laboratory plasmas (Plasmas and Magnetohydrodynamics). Rests on Theorem 23.34 and Equation (23.48).
The old quantum theory
The rule that selected the observed atomic states in the old quantum theory is the quantization of the action variables,
one condition for each separable degree of freedom [Sommerfeld:1916]. Its correct semiclassical form carries an extra term,
where the integer \(\mu_{a}\)—the Maslov index of the \(a\)-th circuit—counts the caustics the circuit crosses, and equals \(2\) for a smooth one-dimensional well. Equation (23.61) is the case \(\mu_{a}=0\). Rests on Definition 23.28, Theorem 9.78 and Remark 9.79.
Derivation. Derives Proposition 23.38. Neither form can be derived within classical mechanics: both are selection rules imposed from outside, and what classical mechanics supplies is the object they quantize, the action variable Equation (23.49), defined by the geometry of phase space and by nothing else.
The justification is semiclassical and is already in this treatise. For one degree of freedom, Theorem 9.78 shows that a solution of the stationary Schrödinger equation Equation (9.59) decaying on both sides of a well exists only for those energies with \(\oint P\,\dd x=\left(n+\tfrac{1}{2}\right)h\), which is Equation (23.62) with \(\mu=2\), and Remark 9.79 identifies the two quarters as one \(\pi/4\) from each of the two turning points, contributed by the connection formulae Theorem 9.72. Keller extended the same accounting to several degrees of freedom by applying it to each independent circuit of an invariant torus and counting caustics rather than turning points [Keller:1958]; that count is \(\mu_{a}\), and its geometric content is Proposition 23.56 below. Substituting \(\oint p_{a}\dd q^{a}=2\pi I_{a}\) and \(h=2\pi\hbar\) turns the condition into Equation (23.62).
∎The empirical successes and the failures of the uncorrected rule—it gives the hydrogen spectrum Equation (23.45) and the fine structure, and fails for helium—are the subject of Atomic Models and Spectra. Two structural points survive the failure. First, the quantity being quantized is the action variable of Definition 23.28. Second, the missing \(\mu_{a}/4\) is not a patch but a statement about the classical trajectory: it counts where the projection of the invariant torus onto configuration space folds over. That is why Section 23.5.2 is part of this chapter and not an appendix to it.
The optical–mechanical analogy
The eikonal limit of wave optics
Let a monochromatic scalar wave of vacuum wavenumber \(k_{0}=\omega/c\), in \(/\mathrm{m}\), propagate in a medium of refractive index \(n(\vect{x})\), so that its amplitude satisfies the Helmholtz equation Equation (10.25) with local wavenumber \(k=k_{0}n(\vect{x})\),
Write it in the form
The real function \(\mathcal{S}\), in metres, is the eikonal; its level surfaces \(\mathcal{S}=\text{const}\) are the geometrical wavefronts. The name is Bruns' [Bruns:1895]; the function is Hamilton's characteristic function of [Hamilton:1828]. Rests on Equation (10.25) and Definition 16.103.
In the short-wavelength limit \(k_{0}\to\infty\), with \(A\) and \(n\) varying appreciably only over distances large compared with \(1/k_{0}\), the Helmholtz equation reduces at leading order to the eikonal equation
and at the next order to the transport equation
which states the conservation of energy flux along the tubes of rays. The neglected term is \(\nabla^{2}A/\left(k_{0}^{2}A\right)\), so the approximation is good where \(\abs{\nabla^{2}A}\ll k_{0}^{2}\abs{A}\). Rests on Definition 23.39 and Equation (23.63).
Derives Theorem 23.40. Differentiate Equation (23.64) twice:
Substituting into Equation (23.63) and cancelling the exponential leaves, with \(A\), \(\mathcal{S}\) and \(n\) all real,
and each bracket vanishes separately. The real one, divided by \(-k_{0}^{2}A\), reads \(\left(\vect{\nabla}\mathcal{S}\right)^{2}-n^{2} =\nabla^{2}A/\left(k_{0}^{2}A\right)\), whose right-hand side is the stated neglected term; dropping it gives Equation (23.65). The imaginary one, multiplied by \(A\), is \(2A\,\vect{\nabla}A\cdot\vect{\nabla}\mathcal{S} +A^{2}\nabla^{2}\mathcal{S} =\vect{\nabla}A^{2}\cdot\vect{\nabla}\mathcal{S} +A^{2}\nabla^{2}\mathcal{S}\), which is Equation (23.66).
∎The curves orthogonal to the wavefronts—the rays—satisfy
\(s\) being arclength, and they are the stationary curves of Fermat's principle
Rests on Theorem 23.40, Lemma 16.100 and Definition 16.103.
Derives Proposition 23.41. By Equation (23.65) the gradient of the eikonal has magnitude \(n\), so the unit normal to the wavefronts is \(\vect{T}=\vect{\nabla}\mathcal{S}/n\), and along a ray, parametrised by arclength, \(\vect{T}=\dd\vect{x}/\dd s\). Hence \(n\,\dd\vect{x}/\dd s=\vect{\nabla}\mathcal{S}\) and
From Equation (23.65) (a gradient field is curl-free, and its squared magnitude is fixed by the eikonal equation). The middle step holds because a gradient field has vanishing curl, so that \(\left(\vect{a}\cdot\vect{\nabla}\right)\vect{a} =\tfrac{1}{2}\vect{\nabla}\abs{\vect{a}}^{2}\) for \(\vect{a}=\vect{\nabla}\mathcal{S}\), and the last uses Equation (23.65). Equation (23.69) is Equation (16.90), the Euler–Lagrange equation of the optical path length established in Lemma 16.100; Fermat's principle Equation (23.70) is therefore equivalent to it, which is Definition 16.103 [Fermat:1894].
∎The dictionary
Compare Equation (23.65) with the time-independent Hamilton–Jacobi equation Equation (23.15) for a Hamiltonian \(\Ham=\abs{\vect{p}}^{2}/2m+V\), which reads
The two equations are the same equation. Every statement about wavefronts and rays therefore has a mechanical counterpart, obtained by the substitution \(n\mapsto\sqrt{2m(E-V)}\).
| Geometrical optics | Classical mechanics |
|---|---|
| eikonal $\mathcal{S}(\vect{x})$ | characteristic function $W(\vect{x})$ |
| refractive index $n(\vect{x})$ | $\sqrt{2m\left(E-V(\vect{x})\right)}$ |
| eikonal equation $(\vect{\nabla}\mathcal{S})^{2}=n^{2}$ | Hamilton–Jacobi equation $(\vect{\nabla}W)^{2}=2m(E-V)$ |
| wavefronts $\mathcal{S}=\text{const}$ | surfaces of constant action $W=\text{const}$ |
| rays, orthogonal to the wavefronts | trajectories, orthogonal to the surfaces $W=\text{const}$ |
| Fermat's principle $\delta\int n\,\dd s=0$ | Maupertuis' principle $\delta\int\sqrt{2m(E-V)}\,\dd s=0$ |
| Snell's law at an interface | refraction of a trajectory at a potential step |
| wavelength $\lambda=\lambda_{0}/n$ | de Broglie wavelength $\lambda=h/\abs{\vect{p}}$ |
| geometrical-optics limit $\lambda\to0$ | classical limit $\hbar\to0$ |
The Snell row of Table 23.1 is worth checking, because it is the row a reader is likeliest to take on trust. At a plane interface across which \(V\) jumps, the component of \(\vect{p}\) tangent to the interface is conserved—no force acts along it—so \(\abs{\vect{p}_{1}}\sin\vartheta_{1}=\abs{\vect{p}_{2}}\sin\vartheta_{2}\) with \(\vartheta\) measured from the normal, which under \(n=\abs{\vect{p}}\) is Equation (16.93) exactly (Theorem 16.104 and Phenomenon 61.3). A particle entering a region of lower potential speeds up and bends towards the normal, as light does entering a denser medium. This is the correspondence Newton used to argue that light is corpuscular, and the one that made the measured slowing of light in water Phenomenon 61.2 decisive against him.
The wavefronts \(S=\text{const}\) do not move at the speed of the particle. From Equation (23.14), a surface of constant \(S\) moves so that \(\vect{\nabla}W\cdot\dd\vect{x}=E\,\dd t\), giving a normal speed \(u=E/\abs{\vect{\nabla}W}=E/\abs{\vect{p}}\), whereas the particle moves at \(v=\abs{\vect{p}}/m\); for a free particle \(uv=E/m=v^{2}/2\), so \(u=v/2\), as Example 23.22 found. The wavefront speed is a phase speed and the particle speed a group speed, and their product is fixed. In the wave mechanics of Matter Waves the same two speeds reappear with the same relation, which is one of the checks de Broglie's hypothesis had to pass [deBroglie:1925]. Rests on Equation (23.14) and Example 23.22.
The dictionary of Table 23.1 is exact, and it is incomplete in one direction only. Geometrical optics is known to be an approximation: it is the short-wavelength limit of a wave theory, and where the limit fails—diffraction at an edge, interference, the caustics where neighbouring rays cross—the wave theory is required, and it is right. The last two rows say that classical mechanics stands to some wave theory in exactly the relation that geometrical optics stands to wave optics. Hamilton wrote the analogy in the 1830s and left it there. The wave theory whose short-wavelength limit is classical mechanics was found a century later, and the parameter that plays the part of the wavelength is \(h/\abs{\vect{p}}\). That is the content of Matter Waves, and the experiments that settled it—electron diffraction, neutron and atom interferometry—are its evidence. Rests on Equations (23.65) and (23.72).
Where the analogy is not an analogy
The dictionary of Table 23.1 pairs two theories that remain distinct. There is a regime in which the pairing collapses and the two columns become one column: relativistic mechanics, where the eikonal of a massless field and the characteristic function of a massive particle satisfy the same equation with the same metric, differing only in the value of one constant.
For a free particle of rest mass \(m\), with the signature and index conventions of Minkowski Space and Its Symmetries and \(\pp S/\pp x^{\mu}=-p_{\mu}\), the Hamilton–Jacobi equation is
both sides in \(\mathrm{kg}^{2}\,\mathrm{m}^{2}/\mathrm{s}^{2}\). In the presence of an electromagnetic four-potential \(A_{\mu}\) and for charge \(q\) it becomes
and in a spacetime of metric \(g_{\mu\nu}\) it is Equation (23.73) with \(\eta^{\mu\nu}\) replaced by \(g^{\mu\nu}\). Writing \(S=-mc^{2}t+S'\) and expanding in \(1/c^{2}\) recovers Equation (23.2) for \(S'\) with \(\Ham=\abs{\vect{p}}^{2}/2m\). Rests on Equation (40.5), Equation (23.2) and Definition 13.117.
Derives Proposition 23.44. The nonrelativistic relations \(p_{a}=\pp S/\pp q^{a}\) and \(\pp S/\pp t=-\Ham=-E\) are, in four-vector notation with \(x^{\mu}=(ct,\vect{x})\) and \(p^{\mu}=(E/c,\vect{p})\), the single statement \(\pp S/\pp x^{\mu}=-p_{\mu}\): the time component reads \(c^{-1}\pp_{t}S=-E/c=-p_{0}\) and the space components \(\pp_{i}S=p^{i}=-p_{i}\). Contracting with \(\eta^{\mu\nu}\) gives \(\eta^{\mu\nu}\pp_{\mu}S\,\pp_{\nu}S=\eta^{\mu\nu}p_{\mu}p_{\nu} =p^{2}=m^{2}c^{2}\) by Theorem 40.5, which is Equation (23.73); writing out the components with \(\eta=\diag(1,-1,-1,-1)\) gives the second form. Equation (23.74) follows by the minimal substitution \(p_{\mu}\mapsto p_{\mu}-qA_{\mu}\) in the same mass-shell condition, and the curved-space form by replacing the flat metric with \(g_{\mu\nu}\), which is what makes \(p^{2}\) the invariant square.
For the limit, put \(S=-mc^{2}t+S'\) in the second form of Equation (23.73). Then \(\pp_{t}S=-mc^{2}+\pp_{t}S'\) and
so Equation (23.73) becomes \(-2m\,\pp_{t}S'-\left(\vect{\nabla}S'\right)^{2} +c^{-2}\left(\pp_{t}S'\right)^{2}=0\). Dividing by \(-2m\) and dropping the last term, of relative order \(\left(\pp_{t}S'\right)/mc^{2}\), leaves \(\pp_{t}S'+\left(\vect{\nabla}S'\right)^{2}/2m=0\).
∎Set \(m=0\) in Equation (23.73): the result, \(g^{\mu\nu}\pp_{\mu}S\,\pp_{\nu}S=0\), is the eikonal equation for a massless field in that spacetime, and its characteristics are the null geodesics. Set \(m\neq0\) and the characteristics are the timelike geodesics of Geometric Formulation of Gravity. One equation, one geometry, one family of characteristic curves, and the only difference between the ray of light and the trajectory of the particle is the value on the right-hand side. In Table 23.1 the refractive index had to be manufactured out of \(E\) and \(V\) to make the two columns match; here nothing is manufactured, because there is only one column. The optical–mechanical analogy is, in the relativistic setting, an identity that was hiding inside a nonrelativistic approximation. Rests on Proposition 23.44 and Theorem 23.40.
The most consequential separation of Equation (23.73) outside flat space is Carter's. In the Kerr geometry—the rotating black hole of Schwarzschild Geometry and Black Holes—the equation admits a complete integral of the separated form \(S=-Et+L_{z}\varphi+S_{r}(r)+S_{\theta}(\theta)\) in Boyer–Lindquist coordinates, and the separation constant that couples the radial and polar equations is a fourth constant of the motion,
with \(a\) the spin parameter of the hole, a length in \(\mathrm{m}\) [Carter:1968]. Three constants come from symmetry—\(m\) from the mass shell, \(E\) and \(L_{z}\) from the two Killing vectors of the stationary axisymmetric metric—and the fourth comes from nowhere that Noether's theorem (Theorem 16.84) can see. In the language of Remark 23.21 it is a rank-two Killing tensor, and the separability of Equation (23.73) in that geometry is exactly the Stäckel condition of Definition 23.16 transplanted to a Lorentzian metric. That it exists at all is what makes geodesics around a rotating black hole integrable, and therefore computable; the orbits so computed are what the observed motions of Experiment: Black-Hole Observations are matched against. The separation itself is carried out in Part V and is not reproduced here, and Remark 23.21 has already recorded that the Killing-tensor machinery it deserves is missing from Differentiable Manifolds, Tensors, and Curvature. Rests on Proposition 23.44 and Definition 23.16.
From Hamilton–Jacobi to the Schrödinger equation
Write a solution of the Schrödinger equation of The Postulates of Quantum Mechanics in the polar form
with \(A\) and \(S\) real. Then the real and imaginary parts of the Schrödinger equation are, exactly and without approximation,
Equation (23.77) is the Hamilton–Jacobi equation Equation (23.2) with the single extra term \(Q\), of order \(\hbar^{2}\); Equation (23.78) is the continuity equation for the probability density \(A^{2}\), and is the exact counterpart of the transport equation Equation (23.66). Rests on Equations (10.26), (23.2) and (23.66).
Derives Proposition 23.47. With \(\psi=A\ee^{\ii S/\hbar}\),
Substituting into \(\ii\hbar\,\pp_{t}\psi=-\left(\hbar^{2}/2m\right)\nabla^{2}\psi+V\psi\) (Equation (10.26)) and cancelling the exponential, the real part is
which on division by \(A\) is Equation (23.77), and the imaginary part is
Multiplying the latter by \(2A/\hbar\) turns the left side into \(\pp\left(A^{2}\right)/\pp t\) and the right into \(-m^{-1}\left[\vect{\nabla}\left(A^{2}\right)\cdot\vect{\nabla}S +A^{2}\nabla^{2}S\right]\), which is \(-\vect{\nabla}\cdot\left(A^{2}\vect{\nabla}S/m\right)\): that is Equation (23.78). No approximation was made at any step; both equations are exact rewritings.
∎Equations (23.77) and (23.78) are an exact algebraic rewriting of the Schrödinger equation, not a different theory and not an interpretation of it. Read in one direction they show that classical mechanics is recovered when the term \(Q\) is negligible against the others, which is the precise sense in which the Hamilton–Jacobi equation is the \(\hbar\to0\) limit of wave mechanics. Read in the other direction, as a statement that a particle possesses a definite trajectory guided by \(Q\), they become an interpretive commitment that no experiment discussed in this treatise distinguishes from any other; the treatise's position on such commitments is stated in Interpretations (Evidence-Anchored). Rests on Equations (23.77) and (23.78).
The semiclassical expansion
Proposition 23.47 is exact but not systematic: it isolates one term of order \(\hbar^{2}\) and says nothing about what comes after it. The systematic version is the WKB expansion, and this treatise develops it in full—the expansion itself, its validity criterion, its failure at turning points, the connection formulae that carry a solution across one, the Gamow tunnelling factor and the quantisation condition with its Maslov correction—in one dimension, as mathematics, in Section 9.6 of Ordinary Differential Equations and Sturm–Liouville Theory. That material is not repeated here. What is established here is the one thing that section cannot state, because it is mechanics and not analysis: that the leading order of the expansion is the Hamilton–Jacobi equation of this chapter, and the next order the transport equation.
Write a solution of the stationary Schrödinger equation \(-\left(\hbar^{2}/2m\right)\nabla^{2}\psi+V\psi=E\psi\) in three space dimensions as a single exponential,
with \(\sigma\) complex and of SI dimension \(\mathrm{J}\,\mathrm{s}\). Then \(\sigma\) satisfies exactly
and the expansion \(\sigma=\sigma_{0}+\hbar\sigma_{1}+\hbar^{2}\sigma_{2}+\cdots\) of Equation (9.64) gives, at its first two orders,
Equation (23.82) is the time-independent Hamilton–Jacobi equation Equation (23.72), so \(\sigma_{0}=W\); and with \(A:=\ee^{\ii\sigma_{1}}\), Equation (23.83) is the transport equation \(\vect{\nabla}\cdot\left(A^{2}\vect{\nabla}W\right)=0\), that is Equation (23.66). Rests on Equation (23.72), Equation (23.66) and Proposition 9.61.
Derives Proposition 23.49. From Equation (23.80), \(\vect{\nabla}\psi=\left(\ii/\hbar\right)\vect{\nabla}\sigma\,\psi\) and \(\nabla^{2}\psi=\left[\left(\ii/\hbar\right)\nabla^{2}\sigma -\hbar^{-2}\left(\vect{\nabla}\sigma\right)^{2}\right]\psi\). Substituting into the stationary equation, cancelling the nowhere-vanishing \(\psi\) and multiplying by \(-2m\) gives Equation (23.81); this is the three-dimensional form of the exact Riccati equation Equation (9.63), and \(\hbar\) again appears in exactly one place, multiplying the highest derivative. Inserting the series, the squared term contributes \(\left(\vect{\nabla}\sigma_{0}\right)^{2} +2\hbar\,\vect{\nabla}\sigma_{0}\cdot\vect{\nabla}\sigma_{1} +\mathcal{O}(\hbar^{2})\) and the second term \(-\ii\hbar\,\nabla^{2}\sigma_{0}+\mathcal{O}(\hbar^{2})\), while the right side is of order \(\hbar^{0}\); equating coefficients of \(\hbar^{0}\) and \(\hbar^{1}\) gives Equations (23.82) and (23.83). Comparing Equation (23.82) with Equation (23.72) identifies \(\sigma_{0}\) with the characteristic function.
For the last statement, \(\psi=\ee^{\ii\sigma_{1}}\ee^{\ii\sigma_{0}/\hbar} =A\,\ee^{\ii W/\hbar}\) with \(A=\ee^{\ii\sigma_{1}}\), so \(\vect{\nabla}\left(A^{2}\right)=2\ii A^{2}\vect{\nabla}\sigma_{1}\) and
which vanishes precisely when \(2\vect{\nabla}\sigma_{0}\cdot\vect{\nabla}\sigma_{1} =\ii\nabla^{2}\sigma_{0}\), that is Equation (23.83).
∎Proposition 23.49 closes Table 23.1. The eikonal equation of wave optics and the Hamilton–Jacobi equation of mechanics were shown in Section 23.4.2 to be the same equation under \(n\mapsto\sqrt{2m(E-V)}\); they are now seen to occupy the same position in their respective theories—each is the leading term of a short-wavelength expansion of a linear wave equation, and each is accompanied at the next order by the same transport equation. The role of \(1/k_{0}\) in Theorem 23.40 is played by \(\hbar\) here, and the reduced de Broglie wavelength Equation (9.72) of Definition 9.64,
in metres, is the local wavelength whose smallness is the whole hypothesis. The criterion is \(\abs{\vect{\nabla}\bar{\lambda}}\ll1\), which is Equation (9.73) of Proposition 9.65 read in three dimensions: the wavelength must change by much less than itself over one wavelength. For an electron of kinetic energy \(100\,\mathrm{eV}\), \(\bar{\lambda}=1.95\times 10^{-11}\,\mathrm{m}\), so the criterion demands a potential smooth on the scale of a few tens of picometres—which is why the semiclassical description fails inside an atom and succeeds for an electron crossing a laboratory field. Rests on Propositions 9.65 and 23.49.
The criterion fails wherever \(\abs{\vect{P}}\to0\), that is on the classical turning surface \(V=E\), whatever the smoothness of \(V\): the amplitude \(\abs{\vect{P}}^{-1/2}\) of Equation (9.69) diverges there while the true solution stays finite (Remark 9.66). Within a distance of order
of the surface, which is Equation (9.74), the expansion is worthless and the exact local solution is an Airy function (Definition 9.68 and Theorem 9.70); the connection formulae Theorem 9.72 carry a solution through. That analysis is one-dimensional and transverse to the surface, and it applies here without change, because near a smooth turning surface the motion normal to it decouples from the motion along it at leading order. The two things it does not settle are what happens at a caustic which is not a turning surface, and how the phases accumulated at successive crossings add up around a closed circuit. Those are the subject of Section 23.5.2, and they are mechanics, not one-dimensional analysis. The series is in any case asymptotic and not convergent (Remark 9.67), so no amount of further orders repairs a place where the leading order is wrong. Rests on Proposition 9.65 and Theorem 9.72.
Caustics
A complete integral does not describe one trajectory but a family of them, and the geometrical description of Section 23.4.2 presumes that exactly one member of the family passes through each point. Where that fails—where neighbouring members cross—the description fails with it, and the failure is not a defect of the calculation but the signature of the wave theory underneath.
Let \(\gamma(\vect{a};\cdot)\) be a family of trajectories of one energy, smoothly parametrised by \(\vect{a}\in\R^{f-1}\) and issuing orthogonally from a common wavefront (or from a common point, for a point source), and let \(q^{b}(\vect{a},s)\) be the configuration reached along the member \(\vect{a}\) after arclength \(s\) of the Jacobi metric Equation (16.91), in which the trajectories of energy \(E\) are geodesics by Theorem 16.101. The Jacobian of the family is
and the caustic of the family is the set where \(J=0\): the locus at which the map from the family to configuration space ceases to be a local diffeomorphism. Equivalently, the caustic is the envelope of the family. Rests on Definitions 16.62 and 23.5.
A point \(q\left(\vect{a},s_{c}\right)\) lies on the caustic of the family issuing from a common source if and only if \(s_{c}\) is conjugate, in the sense of Definition 16.53, to the source along that trajectory. Consequently, by Theorem 16.55, a trajectory ceases to minimise the action beyond its first caustic, although it remains an extremal and a genuine solution of the equations of motion. Rests on Definition 23.52, Definition 16.53 and Theorem 16.55.
Derives Proposition 23.53. The variation fields of the family, \(u^{b}_{(i)}:=\pp q^{b}/\pp a^{i}\), are by Remark 16.54 solutions of the Jacobi accessory equation Equation (16.49)—the linearisation of the Euler–Lagrange equations about the trajectory—because the family consists of extremals depending differentiably on the parameter. All members issue from the same source, so \(u^{b}_{(i)}=0\) there.
Because \(s\) is the arclength of the Jacobi metric and the trajectories are its geodesics (Theorem 16.101), each \(u_{(i)}\) stays orthogonal to the tangent \(\vect{T}=\pp q/\pp s\): differentiating \(\vect{T}\cdot\vect{T}=1\) with respect to \(a^{i}\) gives \(\vect{T}\cdot\pp_{s}u_{(i)}=0\), and \(\pp_{s}\vect{T}=\vect{0}\) along a geodesic, so \(\vect{T}\cdot u_{(i)}\) is constant along the trajectory and vanishes at the source. The Jacobian Equation (23.86) therefore vanishes at \(s_{c}\) exactly when the \(f-1\) transverse vectors \(u_{(i)}\) are linearly dependent there, that is when some non-trivial constant combination \(u=c^{i}u_{(i)}\) vanishes at \(s_{c}\). Such a \(u\) is a non-trivial solution of Equation (16.49) vanishing at the source and at \(s_{c}\), which is the definition of a conjugate point. The converse reads the same equivalence backwards.
∎Along a tube of trajectories of one energy, the transport equation Equation (23.66) gives
\(\Sigma\) being the cross-sectional area of the tube and \(\abs{\vect{P}} =\sqrt{2m(E-V)}\) the local momentum. Since \(\Sigma\propto\abs{J}\), the geometrical amplitude behaves as \(A\propto\abs{J}^{-1/2}\) and diverges on the caustic. Rests on Equation (23.66), Definition 23.52 and Theorem 7.101.
Derives Proposition 23.54. Take the region bounded by two cross-sections \(\Sigma_{1},\Sigma_{2}\) of the tube and by the tube's side wall, which is built of trajectories. Integrate Equation (23.66) over it and apply the divergence theorem Theorem 7.101. The side wall contributes nothing, because \(\vect{\nabla}W\) is tangent to the trajectories and hence to the wall, so \(A^{2}\vect{\nabla}W\) has no normal component there. What remains is that the flux \(\int A^{2}\vect{\nabla}W\cdot\dd\vect{\Sigma}\) takes the same value on the two end faces. On a face orthogonal to the trajectories \(\vect{\nabla}W\cdot\dd\vect{\Sigma}=\abs{\vect{\nabla}W}\,\dd\Sigma =\abs{\vect{P}}\,\dd\Sigma\) by Equation (23.72), which for a thin tube is Equation (23.87). For the cross-section, the volume swept by the parameter increments \(\dd^{f-1}a\) over an arclength \(\dd s\) is \(\abs{J}\,\dd^{f-1}a\,\dd s\) by Equation (23.86), and it is also \(\Sigma\,\dd s\), the tangent being a unit vector orthogonal to the cross-section as shown in Proposition 23.53; hence \(\Sigma=\abs{J}\,\dd^{f-1}a\), proportional to \(\abs{J}\) at fixed increments. Solving Equation (23.87) for \(A\) gives \(A\propto\left(\abs{\vect{P}}\abs{J}\right)^{-1/2}\), and on a caustic that is not also a turning surface \(\abs{\vect{P}}\) is finite and non-zero while \(J\to0\).
One bookkeeping point: the flux tube is measured in ordinary space, while Definition 23.52 parametrises the family by the Jacobi arclength. The two differ by the reparametrisation \(\dd\sigma=\abs{\vect{P}}\,\dd s\), whose factor is positive and finite wherever \(E>V\), so the two Jacobians differ by a non-vanishing factor and have the same zero set. The caustic is the same locus in either parametrisation, which is what allows Proposition 23.53 and this proposition to speak of one \(J\).
∎Let the caustic be a fold: transversally to it, the map from the family parameter to position is quadratic,
\(\zeta\) being the signed distance from the caustic. Then two members of the family reach each point with \(c\zeta>0\) and none reaches a point with \(c\zeta<0\), so that the caustic separates a doubly illuminated side from a dark side, and the geometrical amplitude of each of the two obeys
Rests on Proposition 23.54 and Equation (23.86).
Derives Proposition 23.55. Solving Equation (23.88) for \(a\) gives \(a-a_{c}=\pm\sqrt{\zeta/c}\), real for \(c\zeta>0\) and not otherwise: hence two members on one side and none on the other. The transverse Jacobian is \(\dd\zeta/\dd a=2c\left(a-a_{c}\right)=\pm2\sqrt{c\zeta}\), so \(J\propto\abs{\zeta}^{1/2}\), and Proposition 23.54 gives \(A\propto\abs{J}^{-1/2}\propto\abs{\zeta}^{-1/4}\).
∎A semiclassical wave crossing a fold caustic acquires, relative to the geometrical continuation, an extra phase \(-\pi/2\). Around a closed circuit of an invariant torus the accumulated extra phase is therefore \(-\mu\pi/2\), with \(\mu\) the number of folds crossed, which is the Maslov index appearing in Equation (23.62). Rests on Theorem 9.72, Remark 9.79 and Proposition 23.55.
Derivation. Derives Proposition 23.56. Transversally to the fold the problem is one-dimensional: by Proposition 23.55 the two branches merge at \(\zeta=0\) and the amplitude carries the factor \(J^{-1/2}\), whose argument shifts by \(-\pi/2\) when \(J\) changes sign. Which branch is selected is not a matter of taste, and is fixed by the connection formulae Theorem 9.72: a solution decaying on the dark side matches, on the illuminated side, to \(2\abs{\vect{P}}^{-1/2}\sin\left[\hbar^{-1}\int\abs{\vect{P}}\dd\zeta +\pi/4\right]\), whose two exponential components differ in phase from the naive continuation by \(\mp\pi/2\); the outgoing one lags the incoming by \(\pi/2\). Adding one such contribution per crossing around a closed circuit gives \(-\mu\pi/2\), and demanding single-valuedness of the wave function around the circuit—\(\hbar^{-1}\oint p\,\dd q-\mu\pi/2\in2\pi\Z\)—is Equation (23.62). Remark 9.79 records the one-dimensional instance, \(\mu=2\) for a well with two simple turning points, and Keller's extension of the count to several degrees of freedom [Keller:1958].
∎Where the caustic is a classical turning surface, \(V=E\) with \(\vect{\nabla}V\neq\vect{0}\), the divergence Equation (23.89) is spurious: the exact solution is finite, and within a few multiples of the scale \(\ell\) of Equation (23.85) it is the Airy function \(\mathrm{Ai}\left(-\zeta/\ell\right)\) of Definition 9.68. The illuminated side carries fringes—the interference of the two branches of Proposition 23.55—whose maxima lie at the extrema of \(\mathrm{Ai}\), the brightest at \(\zeta=1.0188\,\ell\), and the dark side carries an exponential tail. For \(\zeta\gg\ell\) the Airy asymptotics Equation (9.79) reproduce Equation (23.89) exactly. Rests on Remark 23.51, Theorem 9.72 and Theorem 9.70.
Derivation. Derives Proposition 23.57. Near a smooth turning surface the potential may be replaced by its tangent plane, \(V\approx E+\abs{\vect{\nabla}V}\,\zeta\), and the stationary Schrödinger equation separates into motion along the surface, which is free at leading order, and motion normal to it, which is Equation (9.75)—the linearised turning-point equation of Ordinary Differential Equations and Sturm–Liouville Theory. Its solutions are the Airy functions (Proposition 9.69) on the scale Equation (9.74), which is Equation (23.85); the decaying combination on the forbidden side is \(\mathrm{Ai}\), and Theorem 9.72 is the statement that it matches the oscillatory semiclassical form on the allowed side. The extrema of \(\mathrm{Ai}(-z)\) are the zeros of \(\mathrm{Ai}'\), the first at \(z=1.0188\), and Equation (9.79) gives \(\mathrm{Ai}(-z)\sim\pi^{-1/2}z^{-1/4}\sin\left(\tfrac{2}{3}z^{3/2} +\tfrac{\pi}{4}\right)\), whose envelope is the \(z^{-1/4}\) of Equation (23.89).
∎Proposition 23.57 covers the caustic that is a turning surface, where the momentum itself vanishes. Most observed caustics are not of that kind: at the focus of a lens, at the cusp in a coffee cup, at the rainbow angle, the momentum is perfectly finite and it is the geometry of the family that folds. The fold Equation (23.88) fixes the transverse scale in that case by itself, at
with \(R\) the radius of curvature of the caustic and \(\bar{\lambda}\) of Equation (23.84), and the same Airy profile with \(\zeta/\ell_{ \text{f}}\) in place of \(\zeta/\ell\); the peak amplitude then exceeds the smooth geometrical value by a factor of order \(\left(R/\bar{\lambda}\right)^{1/6}\). This is Airy's result, obtained for the rainbow [Airy:1838] and set out in full by Born and Wolf [Born:1999]. Rests on Propositions 23.55 and 23.57.
Derivation. Derives Equation (23.90). Two inputs are used and both should be named. The first is the superposition hypothesis of physical optics: near the caustic the field is written as an integral over the members of the family,
\(S(a,\zeta)\) being the optical path along the member labelled \(a\) to the field point and \(A\) a slowly varying amplitude. The reduced wavelength Equation (23.84) is constant across the small region that matters, because at a caustic of this kind \(\abs{\vect{P}}\) is finite and smooth; the phase is therefore \(S/\bar{\lambda}\) with \(S\) an ordinary length. The geometrical field of Proposition 23.54 is what Equation (23.91) reduces to when the phase is stationary at isolated points, a stationary point in \(a\) being exactly a member that passes through the field point. The second input is the integral representation Equation (9.77) of the Airy function.
The phase is fixed by Equation (23.88) alone. Write \(s:=a-a_{c}\). The stationary points of \(S\) are the two members reaching \(\zeta\), which by Proposition 23.55 sit at \(s=\pm\sqrt{\zeta/c}\); so \(\pp S/\pp s\) vanishes exactly there and, being smooth with simple zeros, equals \(\kappa\left(s^{2}-\zeta/c\right)\) to leading order, \(\kappa\) being a constant of the family. Integrating,
The phase is cubic in the family parameter: that, and no property of the wave equation, is what a fold is.
Substituting Equation (23.92) into Equation (23.91), holding \(A\) fixed over the small range of \(s\) that contributes and rescaling by \(u=\left(\kappa/\bar{\lambda}\right)^{1/3}s\),
From Equations (23.88) and (23.92) (the fold makes the phase cubic, and the rescaling removes every constant from it). Put \(u=\ii t\). The integrand becomes the entire integrand of Equation (9.77) with \(z=-X\), and the ends of the contour, which run out along the imaginary \(t\) axis, may be rotated into the sectors of decay identified in Proposition 9.69 by Corollary 8.15. Hence \(\psi\propto\mathrm{Ai}\left(-X\right)\): oscillatory on the illuminated side \(X>0\), exponentially small on the dark side, with transverse scale
It remains to express Equation (23.94) through the radius of curvature. Take the caustic to be locally a circle of radius \(R\) and the family to be its tangent lines, labelled by the dimensionless angle \(s\) of the point of tangency, with \(\zeta\) the outward radial distance. The tangents through the point at radius \(R+\zeta\) touch where \(\cos s=R/\left(R+\zeta\right)\), so \(\zeta=\tfrac{1}{2}Rs^{2}\) and \(c=R/2\). The optical path from the tangency point to the field point is \(\sqrt{r^{2}-R^{2}}\) and the arc run along the caustic is \(R\arccos\left(R/r\right)\), so the path difference between the two tangents is, at \(r=R+\zeta\) and to order \(\zeta^{3/2}\),
whereas Equation (23.92) evaluated at its two stationary points gives \(\Delta S=\tfrac{4}{3}\kappa\left(\zeta/c\right)^{3/2}\). Comparing, and using \(c=R/2\), gives \(\kappa=R/2\) as well, so that Equation (23.94) is \(\ell_{\text{f}}=\left(R\bar{\lambda}^{2}/2\right)^{1/3}\), which is Equation (23.90).
The amplitude then follows without further work. The geometrical amplitude is \(\abs{\zeta}^{-1/4}\) by Equation (23.89), and the wave cuts that growth off at \(\zeta\sim\ell_{\text{f}}\); relative to the geometrical value a distance of order \(R\) from the caustic the peak is larger by \(\left(R/\ell_{\text{f}}\right)^{1/4} =2^{1/12}\left(R/\bar{\lambda}\right)^{1/6}\), which is the factor quoted above.
∎Caustics are the sharpest observable signature of the failure of the geometrical limit, and they are visible in every one of the three languages this chapter has used. In optics they are the bright cusped curve on the bottom of a cup, the twinkling of starlight through a turbulent atmosphere, and the supernumerary bows inside a rainbow, whose spacing is the fringe spacing of Proposition 23.57 and whose explanation was Airy's [Airy:1838]; the rainbow itself is a fold caustic of the family of rays refracted through a spherical drop, at a scattering angle of about \(138\,^\circ\), and it is treated in Section 61.8.2. In mechanics they are the safety parabola of Example 23.23—the envelope beyond which no projectile of given speed can reach—and the turning surfaces at which Equation (9.69) diverges. In gravity they are the caustics of gravitational lensing: the images of a background source multiply as it crosses one, and the sharp peak of a microlensing light curve is the crossing itself [Paczynski:1986], the first lensed system having been identified as a pair of quasar images [Walsh:1979]. What all three have in common is that the divergence is never observed: the wave theory replaces it with a finite peak of the width Equation (23.90), and measuring that width measures the wavelength. Rests on Propositions 23.54 and 23.57.