The Liouville–Arnold Theorem
This appendix proves Theorem 27.31 of Hamilton–Jacobi Theory and the Optical–Mechanical Analogy: a system of \(f\) degrees of freedom possessing \(f\) functionally independent integrals \(F_{1}=\Ham,F_{2},\ldots,F_{f}\) in involution has each compact connected component of a regular level set diffeomorphic to the torus \(T^{f}\), carries quasi-periodic motion on it, and admits action–angle variables in a neighbourhood of it. The last clause is what makes the theorem worth its length: it establishes Theorem 27.30 without the hypothesis of separability under which that theorem was proved, so that the action–angle description of a bounded integrable motion is a consequence of the involution relations alone and not of a lucky choice of coordinates.
Exactly one ingredient is imported, and Remark 27.32 already tells the reader which: that a discrete subgroup of \(\R^{f}\) with compact quotient is a lattice of full rank. It is stated below as Theorem A47.7 and used once, in The translation action and the torus. Everything else is proved here from the symplectic material of Symplectic Geometry of Phase Space and the manifold material of Differentiable Manifolds, Tensors, and Curvature; in particular Frobenius' theorem, which the argument needs twice, is Theorem 17.133 and is not reproved.
Setting and statement
Let \((M,\omega)\) be a symplectic manifold of dimension \(2f\) (Definition 28.8), let \(F_{1},\ldots,F_{f}\) be smooth real functions on \(M\) in involution,
which is Equation (27.54), and write \(\vect{F}=(F_{1},\ldots,F_{f}):M\longrightarrow\R^{f}\). Fix \(\vect{c}\in\R^{f}\) and suppose that \(\vect{c}\) is a regular value, so that the differentials \(\dd F_{1},\ldots,\dd F_{f}\) are linearly independent at every point of \(M_{\vect{c}}=\vect{F}^{-1}(\vect{c})\). Denote by \(X_{a}=X_{F_{a}}\) the Hamiltonian vector field of \(F_{a}\), defined by \(\iota_{X_{a}}\omega=\dd F_{a}\) as in Equation (28.5); by Equation (28.7) it acts on functions as \(X_{a}[g]=\pb{g}{F_{a}}\).
Nothing here fixes the SI dimension of \(F_{2},\ldots,F_{f}\): each is whatever conserved quantity it happens to be, and by Remark 28.11 the vector field \(X_{a}\) carries \([F_{a}]/\mathrm{J}\,\mathrm{s}\), so its flow parameter carries \(\mathrm{J}\,\mathrm{s}/[F_{a}]\) — seconds when \(F_{a}\) is the energy, radians when it is an angular momentum. The action variables of The action variables carry \(\mathrm{J}\,\mathrm{s}\) without exception, because \(\oint p_{a}\dd q^{a}\) does; the angles are pure numbers. That is the whole of the dimensional bookkeeping, and it is worth stating because the lattice \(\Gamma\) of The translation action and the torus lives in a copy of \(\R^{f}\) whose \(f\) axes carry \(f\) different units.
Under the hypotheses above, let \(N\) be a compact connected component of \(M_{\vect{c}}\). Then
-
\(N\) is an \(f\)-dimensional embedded submanifold of \(M\), diffeomorphic to the torus \(T^{f}=\R^{f}/\Z^{f}\), and it is Lagrangian: \(\omega\) vanishes on \(T_{P}N\) for every \(P\in N\);
-
in the angular coordinates supplied by that diffeomorphism the flow of \(\Ham=F_{1}\) is linear, so the motion is quasi-periodic with \(f\) frequencies;
-
on a neighbourhood of \(N\) there exist action–angle variables \((I_{a},\theta^{a})\), canonical, with \(\Ham=\Ham(I)\) and with the equations of motion Equation (27.51) and Equation (27.52).
Rests on Equation (27.54), Theorem 17.133 and Equation (28.8).
The commuting fields and the level set
On the open set \(U\subseteq M\) where \(\dd F_{1},\ldots,\dd F_{f}\) are linearly independent:
-
\(X_{a}[F_{b}]=0\) for all \(a,b\), so every \(X_{a}\) is tangent to every level set of \(\vect{F}\);
-
\(\comm{X_{a}}{X_{b}}=0\) as vector fields;
-
\(X_{1},\ldots,X_{f}\) are linearly independent at every point of \(U\), and \(\omega(X_{a},X_{b})=0\).
Derives Lemma A47.3. (1) By Equation (28.7), \(X_{a}[F_{b}]=\pb{F_{b}}{F_{a}}\), which vanishes by Equation (A47.1). A vector annihilating every \(\dd F_{b}\) is tangent to the level set through its base point.
(2) Equation (28.8) states that \(X_{\pb{u}{v}}=-\comm{X_{u}}{X_{v}}\), so \(\comm{X_{a}}{X_{b}}=-X_{\pb{F_{a}}{F_{b}}}=-X_{0}=0\). This is the one place where involution does its real work, and it is worth saying what it buys: the flows of the \(f\) integrals commute, which is what will make them into an action of an abelian group.
(3) Suppose \(\sum_{a}\lambda^{a}X_{a}(P)=0\) at some \(P\in U\). Contracting with \(\omega\) and using Equation (28.5) gives \(\sum_{a}\lambda^{a}\,\dd F_{a}(P)=0\), whence every \(\lambda^{a}\) vanishes by the independence of the differentials. Finally \(\omega(X_{a},X_{b})=\pb{F_{a}}{F_{b}}=0\) by Equation (28.7) and Equation (A47.1).
∎Let \(D_{P}=\operatorname{span}\set{X_{1}(P),\ldots,X_{f}(P)}\) for \(P\in U\). Then \(D\) is a smooth distribution of constant rank \(f\) (Definition 17.130), it is involutive, and it is integrable. Every connected component \(N\) of \(M_{\vect{c}}\cap U\) is an embedded \(f\)-dimensional submanifold with \(T_{P}N=D_{P}\), hence an integral manifold of \(D\), and it is Lagrangian. Rests on Lemma A47.3, Theorem 17.133 and Theorem 17.59.
Derives Proposition A47.4. Smoothness and constant rank are Lemma A47.3(3); involutivity is Lemma A47.3(2), since the bracket of two spanning fields vanishes and so lies in \(D\) trivially. Integrability then follows from Theorem 17.133, which also supplies, around each point of \(U\), a chart in which \(D\) is spanned by the first \(f\) coordinate fields Equation (17.281), so that \(U\) is foliated by \(f\)-dimensional integral manifolds.
Since \(\vect{c}\) is a regular value, \(\vect{F}\) restricted to \(U\) is a submersion onto a neighbourhood of \(\vect{c}\), and by the regular value theorem Theorem 17.59 the set \(M_{\vect{c}}\cap U\) is an embedded submanifold of dimension \(2f-f=f\) with \(T_{P}(M_{\vect{c}}\cap U)=\ker\dd\vect{F}_{P}\). By Lemma A47.3(1) every \(X_{a}(P)\) lies in that kernel, and by (3) the \(f\) of them are independent; a subspace of dimension \(f\) inside a space of dimension \(f\) is the whole of it, so \(T_{P}(M_{\vect{c}}\cap U)=D_{P}\). Each connected component is therefore a connected integral manifold of \(D\). That it is Lagrangian is Lemma A47.3(3) again: \(\omega\) evaluated on two vectors of \(T_{P}N=D_{P}\) is a combination of the numbers \(\omega(X_{a},X_{b})\), all zero.
∎The translation action and the torus
From here on \(N\) is a compact connected component of \(M_{\vect{c}}\).
Each \(X_{a}\) restricts to a complete vector field on \(N\), and
where \(\phi^{a}\) is the flow of \(X_{a}\), is independent of the order of the factors and defines a smooth action of the additive group \(\R^{f}\) on \(N\). Rests on Lemma A47.3, Proposition 17.131 and Theorem 17.125.
Derives Lemma A47.5. By Lemma A47.3(1) each \(X_{a}\) is tangent to \(N\), so its integral curves through points of \(N\) stay in \(N\) as long as they exist. A smooth vector field on a compact manifold is complete: by Theorem 17.125 each point has a neighbourhood on which the flow exists for a time at least \(\varepsilon\) depending on the neighbourhood, compactness extracts a finite subcover and hence a single \(\varepsilon>0\) that works everywhere on \(N\), and the flow is then extended to all times by composing it with itself. Commutativity of the flows is Proposition 17.131 applied to \(\comm{X_{a}}{X_{b}}=0\), which makes the composition Equation (A47.2) independent of the order and gives \(\Phi(\vect{t})\Phi(\vect{t}')=\Phi(\vect{t}+\vect{t}')\); \(\Phi(\vect{0})\) is the identity, and smoothness in \((\vect{t},P)\) jointly is that of each flow.
∎For every \(P\in N\) the orbit map \(\vect{t}\longmapsto\Phi(\vect{t})P\) is a local diffeomorphism of \(\R^{f}\) onto its image; the orbit is open in \(N\); and since \(N\) is connected there is exactly one orbit, so the action is transitive. Rests on Lemma A47.5, Corollary A28.7 and Proposition A47.4.
Derives Lemma A47.6. Differentiating Equation (A47.2) at \(\vect{t}=\vect{0}\) sends the \(a\)-th basis vector of \(\R^{f}\) to \(X_{a}(P)\), and by Lemma A47.3(3) together with Proposition A47.4 those \(f\) vectors are a basis of \(T_{P}N\). The differential of the orbit map at the origin is therefore an isomorphism onto \(T_{P}N\), and the inverse function theorem (Corollary A28.7, read in a chart of \(N\)) makes the orbit map a diffeomorphism of a neighbourhood of \(\vect{0}\) onto a neighbourhood of \(P\) in \(N\). Applying this at an arbitrary point \(\Phi(\vect{t}_{0})P\) of the orbit — which is legitimate because \(\Phi(\vect{t}_{0})\) is a diffeomorphism of \(N\) carrying the orbit map based at \(P\) to the one based at \(\Phi(\vect{t}_{0})P\) — shows that the orbit contains a neighbourhood of each of its points, i.e. is open. Distinct orbits are disjoint, so the orbits form a partition of \(N\) into open sets; connectedness leaves only one.
∎Let \(\Gamma\subseteq\R^{f}\) be a subgroup that is discrete, in the sense that some neighbourhood of \(\vect{0}\) meets \(\Gamma\) only in \(\vect{0}\). Then there are linearly independent vectors \(\vect{\lambda}_{1},\ldots,\vect{\lambda}_{k}\in\R^{f}\), \(k\leq f\), with
and the quotient group \(\R^{f}/\Gamma\) is compact if and only if \(k=f\), in which case it is diffeomorphic to the torus \(T^{f}\). Rests on Definitions 8.21 and 10.9.
Theorem A47.7 is the only statement in this section that is not proved, and Remark 27.32 of Hamilton–Jacobi Theory and the Optical–Mechanical Analogy names it as the debt in the same words. It is elementary — an induction on the dimension of the linear span of \(\Gamma\), at each step choosing a shortest non-zero element of the part of \(\Gamma\) lying in a new direction and showing that discreteness makes it generate that direction's contribution — but it is a theorem about subgroups of \(\R^{f}\), which is algebra and topology rather than mechanics, and it belongs with the material of Algebraic Structures and Topological and Metric Spaces, which do not carry it. The treatise's bibliography holds no source for it either, so the attribution is made in words: it is standard lattice theory, and the form used here is the one Arnold states in Mathematical Methods of Classical Mechanics (2nd edition, 1989), a work for which Symplectic Geometry of Phase Space also records that there is no key. Nothing else below is quoted.
Fix \(P\in N\) and let \(\Gamma=\set{\vect{t}\in\R^{f}\mid\Phi(\vect{t})P=P}\). Then \(\Gamma\) is a discrete subgroup of \(\R^{f}\) of rank exactly \(f\), the induced map
is a diffeomorphism, and \(N\) is diffeomorphic to \(T^{f}\). Rests on Lemma A47.6, Theorem A47.7 and Lemma A47.5.
Derives Proposition A47.9. \(\Gamma\) is a subgroup because \(\Phi\) is an action, and it is discrete because Lemma A47.6 makes the orbit map injective on some neighbourhood \(B\) of \(\vect{0}\), so \(\Gamma\cap B=\set{\vect{0}}\). The orbit map is constant on cosets of \(\Gamma\) and, by transitivity, descends to a bijection Equation (A47.4) of \(\R^{f}/\Gamma\) onto \(N\); it is a local diffeomorphism by the same lemma, hence a diffeomorphism. Since \(N\) is compact so is \(\R^{f}/\Gamma\), and Theorem A47.7 then gives \(\Gamma\) of rank \(f\) and \(\R^{f}/\Gamma\cong T^{f}\).
∎Under the identification Equation (A47.4), the Hamiltonian flow of \(\Ham=F_{1}\) starting at \(P\) is \(t\longmapsto t\,\vect{e}_{1}+\Gamma\): a straight line in \(\R^{f}\) read modulo the lattice. The motion is therefore quasi-periodic, with \(f\) frequencies fixed by the position of \(\vect{e}_{1}\) relative to the lattice basis, and it is periodic exactly when the line closes, that is when \(T\vect{e}_{1}\in\Gamma\) for some \(T>0\). Rests on Proposition A47.9 and Lemma A47.5.
Derives Corollary A47.10. The Hamiltonian vector field of \(F_{1}\) is \(X_{1}\), whose flow is \(\Phi(t\vect{e}_{1})\) by Equation (A47.2). The identification Equation (A47.4) carries it to translation by \(t\vect{e}_{1}\) on \(\R^{f}/\Gamma\), and a translation flow on a torus closes if and only if its generator is commensurable with the lattice. Remark 27.33 records what happens to these tori under perturbation, and Theorem 36.51 is the statement that most of them survive.
∎The action variables
Let \(\vect{\lambda}_{1}(\vect{c}),\ldots,\vect{\lambda}_{f}(\vect{c})\) be a basis of the lattice \(\Gamma\) supplied by Proposition A47.9, chosen to depend smoothly on \(\vect{c}\) — which is possible because a lattice basis is locally rigid: it is determined by finitely many periods, and a nearby torus has nearby periods, so the basis is continued uniquely. Write \(\gamma_{a}(\vect{c})\) for the closed curve \(t\in[0,1]\mapsto\Phi(t\vect{\lambda}_{a})P\) on the torus \(N(\vect{c})\); its homology class is independent of the base point \(P\), and \(\gamma_{1},\ldots,\gamma_{f}\) generate the first homology of \(N\) because they are the images of the lattice generators under Equation (A47.4).
With \(\theta\) the tautological one-form of Equation (28.3), that is \(\theta=p_{a}\dd q^{a}\) in any canonical chart, put
Rests on Proposition A47.9, Equation (28.3) and Equation (27.49).
The integral Equation (A47.5) depends only on the homology class of \(\gamma_{a}\) in \(N\), not on the representing curve, and it depends smoothly on \(\vect{c}\). Rests on Proposition A47.4, Definition A47.11 and Theorem A29.14.
Derives Lemma A47.12. \(N\) is Lagrangian by Proposition A47.4, so the restriction of \(\omega=-\dd\theta\) to \(N\) vanishes, that is the restriction of \(\theta\) to \(N\) is a closed one-form. Two homologous closed curves on \(N\) bound a two-chain there, and Stokes' theorem (Theorem A29.14) turns the difference of the two circuit integrals into the integral of \(\dd\theta=-\omega\) over that chain, which is zero. Smoothness in \(\vect{c}\) follows because the curves \(\gamma_{a}(\vect{c})\) were chosen to depend smoothly on \(\vect{c}\) and \(\theta\) is smooth.
∎The next statement is the one the chapter's own proof of Theorem 27.30 assumed without proving: that the actions may replace the constants \(\vect{c}\) as labels of the tori. Here it is a theorem, and its content is an exact identity between the Jacobian and the lattice.
With the conventions above,
so \(\vect{c}\longmapsto\vect{I}(\vect{c})\) is a local diffeomorphism and the \(I_{a}\) are independent functions on a neighbourhood of \(N\). Rests on Definition A47.11, Proposition A47.9 and Proposition 26.18.
Derives Proposition A47.13. Let \(S\) be the union of the tori \(N(\vect{c}')\) for \(\vect{c}'\) near \(\vect{c}\), a neighbourhood of \(N\) in \(M\), and on it define
the integral taken along a path from a fixed reference point. Since the restriction of \(\theta\) to each torus is closed (Lemma A47.12), \(W\) is well defined up to the periods, and by Equation (A47.5) carrying the path once around \(\gamma_{a}\) increases it by
Read \(W\) as a generating function of the second kind (Proposition 26.18) with the constants \(c_{b}\) as new momenta; the conjugate new coordinates are then \(\tau^{b}=\pp W/\pp c_{b}\), and by the second of Hamilton's equations in the new variables (Theorem 26.6) each \(\tau^{b}\) advances at unit rate along the flow of \(F_{b}\) and is constant along the flow of every other \(F_{a}\). In other words \(\tau^{b}\) is the \(b\)-th time coordinate of the action Equation (A47.2). Carrying the path once around \(\gamma_{a}\) means applying \(\Phi(\vect{\lambda}_{a})\), so
On the other hand \(\tau^{b}=\pp W/\pp c_{b}\), and the circuit \(\gamma_{a}\) lies inside a single torus, so the constants \(c_{b}\) are fixed along it and the two differentiations may be exchanged under the circuit integral, exactly as in Equation (27.53):
Comparing Equation (A47.9) with Equation (A47.10) gives Equation (A47.6). The determinant of the lattice generators is non-zero because they are linearly independent (Theorem A47.7).
∎For one degree of freedom and a bounded orbit of energy \(E\), \(f=1\) and \(\vect{\lambda}_{1}\) is the period \(T\) of the motion, so Equation (A47.6) reads \(\dd I/\dd E=T/2\pi=1/\omega\), which is exactly Equation (27.57), derived there from Hamilton's equations alone. For the harmonic oscillator, \(I=E/\omega\) with \(\omega\) independent of \(E\), and both sides are \(1/\omega\). Rests on Proposition A47.13 and Equation (27.57).
The angles, and the proof of the theorem
Proof of Theorem A47.2. Derives Theorem A47.2. Part (1) is Proposition A47.4 together with Proposition A47.9; part (2) is Corollary A47.10. It remains to build the action–angle variables.
By Proposition A47.13 the map \(\vect{c}\mapsto\vect{I}\) may be inverted on a neighbourhood, so \(W\) of Equation (A47.7) may be regarded as a function \(W(q,I)\) of the position coordinates and the actions. It is again a generating function of the second kind, now with the \(I_{a}\) as new momenta, so
which is Equation (27.50), and the transformation \((q,p)\mapsto(\theta,I)\) is canonical by Proposition 26.18. Repeating the exchange of differentiations that produced Equation (A47.10), now with \(I_{a}\) in place of \(c_{b}\),
so each \(\theta^{a}\) is well defined modulo \(2\pi\) and the \(f\) of them are angular coordinates on the torus. The \(I_{a}\) are functions of the \(c_{b}\) alone and are therefore constants of the motion; conversely \(\Ham=F_{1}=c_{1}\) is a function of the \(c_{b}\), hence of the \(I_{a}\), and contains no angle. Theorem 26.6 in the new variables then gives \(\dot{I}_{a}=-\pp\Ham/\pp\theta^{a}=0\), which is Equation (27.51), and \(\dot{\theta}^{a}=\pp\Ham/\pp I_{a}=\omega^{a}(I)\), constant along the motion, which integrates to Equation (27.52). That is part (3), and with it Theorem 27.30 freed of the separability hypothesis under which it was originally proved.
∎One step above deserves to be stated with its caveat rather than glossed. Writing \(W=W(q,I)\) presumes that the \(f\) position coordinates \(q^{a}\) are good coordinates on the torus, i.e. that \(N\to\R^{f}\), \(P\mapsto q(P)\), is a local diffeomorphism. A Lagrangian torus need not project that way everywhere: the projection degenerates on the caustics, which for the one-dimensional bounded orbit are the two turning points, where \(\pp p/\pp q\) blows up and the orbit folds back. The repair is local and is the one Proposition 26.18 already provides: near a point where some subset of the \(q^{a}\) fails, Legendre-transform in those variables and use a generating function of a different kind, whose arguments are the surviving \(q\)'s and the corresponding \(p\)'s. The circuit integrals Equation (A47.5) and Equation (A47.12) are unaffected, because \(\oint\theta\) is defined on \(N\) without reference to any projection. What the caustics do affect is the semiclassical use of \(W\): each fold crossed contributes a phase, and the count of them around \(\gamma_{a}\) is the Maslov index \(\mu_{a}\) of Equation (27.62). The classical theorem is indifferent to it; the quantization condition is not.
The Liouville–Arnold Theorem discharges the derivation owed at Theorem 27.31 of Hamilton–Jacobi Theory and the Optical–Mechanical Analogy, and Remark 27.32 there states the single imported ingredient, restated above as Theorem A47.7. Three connections back into the book are worth making explicit. First, the chapter observes that nothing in it depends on the theorem, because Proposition 27.20 establishes on its own that a separable system is integrable; what the appendix adds is the converse direction of usefulness — action–angle variables now exist for an integrable system whether or not anyone can separate its Hamilton–Jacobi equation, which is Theorem 27.30 with its hypothesis removed. Second, the tori built here are the objects whose fate under perturbation is the subject of Theorem 36.51, and Remark 27.33 is the honest statement that the hypotheses of Theorem 27.31 are met by almost no system. Third, the Lagrangian property established in Proposition A47.4 is what Remark 28.13 of Symplectic Geometry of Phase Space has in mind when it lists the Liouville–Arnold tori among the statements of symplectic geometry that carry physical content: they are global objects, and a symplectic manifold has nothing local to say.