Symplectic Geometry of Phase Space

Contents
  1. Symplectic vector spaces
  2. Symplectic manifolds
  3. Hamiltonian vector fields and the flow
  4. Invariants of the flow
  5. Poisson manifolds
  6. Symmetry, momentum maps, and reduction
  7. Geometric quantization
  8. What the symplectic view buys

Hamiltonian Mechanics assembled the coordinates and momenta into one object \(\eta^{c}\) and wrote Hamilton's equations as the single matrix equation Equation (22.57). That was a notational convenience. This chapter takes it seriously: the matrix \(J\) of Equation (22.56) is the component form of a geometric object living on phase space—a closed, nondegenerate two-form—and almost everything said about Hamiltonian systems is a statement about that form.

The payoff is that structure invisible in coordinates becomes visible. The Poisson bracket, the canonical transformations, Liouville's theorem, the Poincaré invariants of Section 22.3 and the conserved quantities attached to symmetries are not five facts but one fact seen five ways. And several quantities that a coordinate treatment has no reason to single out—the action variables, the adiabatic invariants, the geometric phases—are exactly the quantities that survive perturbation and that experiment therefore finds robust. That is the sense in which the symplectic view lets one see more in nature than the Newtonian one: it says which quantities are structural and which are accidents of the coordinates chosen to write them.

This chapter uses the manifold and differential-form machinery of Differentiable Manifolds, Tensors, and Curvature and the Lie-group material of Lie Groups, Lie Algebras, and Fibre Bundles; per the dimension rule of this treatise it is stated in general dimension, and the physical instantiation is \(2f\) with \(f\) the number of degrees of freedom.

Notation 24.1.

\(M\) denotes a smooth manifold of even dimension \(2f\), \(T_{x}M\) its tangent space at \(x\), and \(\Omega^{k}(M)\) the space of differential \(k\)-forms. The interior product of a vector field \(X\) with a form is written \(\iota_{X}\)—it is the operation written \(i_{\xi}\) in Equation (13.276)—the exterior derivative \(\dd\), and the Lie derivative along \(X\) is written \(\mathcal{L}_{X}\): note that in this chapter, and only here, the symbol \(\mathcal{L}\) is the Lie derivative and not the Lagrangian, which does not appear. Latin indices \(c,d=1,\ldots,2f\) label phase-space components as in Notation 22.1, and \(a,b=1,\ldots,f\) the degrees of freedom. Because \(\omega\) is reserved throughout for the symplectic form, an angular velocity is written \(\vect{\Omega}\) in Section 24.5 and an ordinary frequency \(\nu\) in Section 24.7.

Symplectic vector spaces

Definition 24.2 (Symplectic vector space).

A symplectic vector space is a pair \((V,\omega)\) with \(V\) a real vector space and \(\omega:V\times V\to\R\) a bilinear form that is

  1. antisymmetric: \(\omega(u,v)=-\omega(v,u)\); and

  2. nondegenerate: \(\omega(u,v)=0\) for all \(v\) implies \(u=0\).

Rests on Proposition 5.115.

Proposition 24.3 (Even dimension and the canonical basis).

A symplectic vector space has even dimension \(2f\), and admits a basis \(\left(e_{1},\ldots,e_{f},f^{1},\ldots,f^{f}\right)\) in which

\begin{equation}\tag{24.1} \omega(e_{a},e_{b})=0\ec\qquad \omega(f^{a},f^{b})=0\ec\qquad \omega(e_{a},f^{b})=\delta^{b}_{a}\ec \end{equation}

so that the matrix of \(\omega\) in that basis is the symplectic matrix Equation (22.56). All symplectic vector spaces of the same dimension are therefore isomorphic: unlike a metric, a symplectic form has no local invariants and no signature. Moreover \(\Sp(2f,\R)\) acts transitively on the set of such bases. Rests on Definition 24.2, Proposition 5.115 and Equation (22.56).

Proof.

Derives Proposition 24.3. The first two statements are Proposition 5.115 of Linear Algebra and Representation Theory at \(2n=2f\): that proposition proves, for a nondegenerate antisymmetric bilinear form on a real vector space, that the dimension is even and that a basis exists satisfying Equation (5.143), which are Equation (24.1) with the naming used here, and in which the matrix of the form is Equation (5.144)—the matrix \(J\) of Equation (22.56). Two symplectic vector spaces of dimension \(2f\) are then isomorphic by the linear map carrying one such basis to the other, which preserves \(\omega\) because it preserves the matrix.

Transitivity is the same observation read inside one space. Let \((e,f)\) and \((e',f')\) be two bases satisfying Equation (24.1) and let \(M\) be the matrix of the change of basis. Both bases give \(\omega\) the matrix \(J\), and a change of basis transforms the matrix of a bilinear form by \(\left[\omega\right]\mapsto{M}\transpose\left[\omega\right]M\) (Equation (5.139)); hence \({M}\transpose J\,M=J\) and \(M\in\Sp(2f,\R)\) by Equation (5.174). Conversely every \(M\in\Sp(2f,\R)\) carries a canonical basis to a canonical basis, so the action is transitive. Nothing beyond Proposition 5.115 is proved here; it is restated because the naming of the basis vectors is what fixes the block order of \(J\) used throughout this chapter.

Definition 24.4 (Symplectic group).

The symplectic group \(\Sp(2f,\R)\) is the group of linear maps \(M:V\to V\) preserving \(\omega\); in the canonical basis it is the set of \(2f\times2f\) real matrices satisfying

\begin{equation}\tag{24.2} M\transpose J\,M=J\ep \end{equation}

This is Definition 5.130 of Linear Algebra and Representation Theory, restated in the notation of this chapter; it is not a second definition. Rests on Definition 5.130 and Proposition 24.3.

Proposition 24.5 (Properties of the symplectic group).

\(\Sp(2f,\R)\) is a Lie group of dimension \(f(2f+1)\) whose Lie algebra \(\mathfrak{sp}(2f,\R)\) consists of the matrices \(A\) with \(A\transpose J+JA=0\), equivalently \(A=JS\) with \(S\) symmetric; every element has determinant \(+1\); and the group is noncompact for every \(f\geq1\). Rests on Definition 24.4, Proposition 5.134 and Theorem 5.133.

Proof.

Derives Proposition 24.5. The dimension, the Lie-algebra condition and the fact that \(\mathfrak{sp}(2f,\R)\) is a Lie algebra are Proposition 5.134 with Equation (5.178) at \(n=f\); the determinant is Theorem 5.133. That \(A\transpose J+JA=0\) is the same as \(A=JS\) with \(S\) symmetric follows by substituting \(A=JS\) and using \(J\transpose=-J\), \(J^{2}=-\identity_{2f}\) from Equation (22.56):

\[ \left(JS\right)\transpose J+J\left(JS\right) =S\transpose J\transpose J-S =S\transpose-S\ec \]

which vanishes exactly when \(S\) is symmetric; and every \(A\) arises this way with \(S=-JA\).

Noncompactness is proved by exhibiting an unbounded one-parameter family. For \(\lambda>0\) put

\[ M_{\lambda}= \begin{pmatrix} \lambda\identity_{f} & 0\\ 0 & \lambda^{-1}\identity_{f} \end{pmatrix}\ec \qquad J\,M_{\lambda}= \begin{pmatrix} 0 & \lambda^{-1}\identity_{f}\\ -\lambda\identity_{f} & 0 \end{pmatrix}\ec \]

so that \(M_{\lambda}\transpose\), which multiplies the first block row by \(\lambda\) and the second by \(\lambda^{-1}\), gives \(M_{\lambda}\transpose\left(J\,M_{\lambda}\right)=J\): each \(M_{\lambda}\) lies in \(\Sp(2f,\R)\). Its entries are unbounded as \(\lambda\to\infty\), so \(\Sp(2f,\R)\) is not a bounded subset of the \(4f^{2}\)-dimensional space of matrices and hence, by Theorem 6.12, not compact. Physically this family is the squeeze that trades position spread against momentum spread at constant phase area, and it is the reason a symplectic map can distort a region without limit while preserving its volume.

Remark 24.6 (Connectedness, and what it rests on).

\(\Sp(2f,\R)\) is also connected, and its maximal compact subgroup is \(\U(f)\); the standard argument is the polar decomposition \(M=OP\) of a real invertible matrix into an orthogonal and a positive-definite symmetric factor, followed by the observation that both factors of a symplectic matrix are themselves symplectic and that the positive factor can be contracted to the identity along \(P^{t}\), \(t\in[0,1]\). The polar decomposition is a statement of linear algebra and belongs in Linear Algebra and Representation Theory, which does not at present carry it; the claim is therefore recorded here as one this treatise states without proving. Nothing in this chapter uses it: connectedness enters only when one asks whether a symplectic matrix can be reached continuously from the identity, which is a question about generating canonical transformations by flows and is not needed for any result below. Rests on Proposition 24.5.

Remark 24.7 (Contrast with the orthogonal group).

A metric of signature \((p,q)\) gives the group \(\SO(p,q)\) of Minkowski Space and Its Symmetries, of dimension \(\tfrac{1}{2}n(n-1)\) with \(n=p+q\); a symplectic form in the same dimension \(n=2f\) gives a group of dimension \(f(2f+1)=\tfrac{1}{2}n(n+1)\), which is larger. Phase space has more symmetry than spacetime, not less, and that is why so much can be transformed away in it. Rests on Proposition 24.5.

Symplectic manifolds

Definition 24.8 (Symplectic manifold).

A symplectic manifold is a pair \((M,\omega)\) with \(M\) a smooth manifold of dimension \(2f\) and \(\omega\in\Omega^{2}(M)\) a two-form that is

  1. closed: \(\dd\omega=0\); and

  2. nondegenerate: \(\omega_{x}\) is a symplectic form on \(T_{x}M\) for every \(x\in M\).

This is Definition A.67, restated here because the whole chapter is about it. Rests on Definitions 13.98, 13.105 and 24.2.

Definition 24.9 (The canonical form on a cotangent bundle).

Let \(Q\) be the configuration manifold of a mechanical system and \(M=T^{*}Q\) its cotangent bundle—phase space. In coordinates \((q^{a},p_{a})\) adapted to the bundle, the canonical one-form (also called the tautological, Liouville, or symplectic potential) and the canonical two-form are

\begin{equation}\tag{24.3} \theta=p_{a}\,\dd q^{a}\ec\qquad \omega=-\dd\theta=\dd q^{a}\wedge\dd p_{a}\ep \end{equation}

Rests on Definitions 13.103 and 24.8.

Proposition 24.10 (The canonical form is symplectic and intrinsic).

The form \(\omega\) of Equation (24.3) is closed and nondegenerate, and \(\theta\) is defined without reference to coordinates by \(\theta_{\xi}(v)=\xi\!\left(\dd\pi(v)\right)\) for \(\xi\in T^{*}Q\) and \(v\in T_{\xi}(T^{*}Q)\), with \(\pi:T^{*}Q\to Q\) the projection. Phase space therefore carries a symplectic structure with no further input: no metric, no connection, and no choice of coordinates. Rests on Definition 24.9, Lemma 13.107 and Definition 13.52.

Proof.

Derives Proposition 24.10. Closed. \(\omega=-\dd\theta\) is exact, hence closed by Lemma 13.107; equivalently \(\dd\omega=-\dd^{2}\theta=0\) by Equation (13.245). The second equality in Equation (24.3) is the Leibniz rule for \(\dd\): \(\dd\left(p_{a}\dd q^{a}\right)=\dd p_{a}\wedge\dd q^{a} =-\dd q^{a}\wedge\dd p_{a}\).

Nondegenerate. In the coordinate basis \(\left(\pp/\pp q^{a},\pp/\pp p_{a}\right)\) of \(T_{\xi}(T^{*}Q)\) one reads off from \(\omega=\dd q^{a}\wedge\dd p_{a}\) that

\begin{align*} \omega\!\left(\pdv{}{q^{a}},\pdv{}{q^{b}}\right)&=0\ec\\ \omega\!\left(\pdv{}{p_{a}},\pdv{}{p_{b}}\right)&=0\ec\\ \omega\!\left(\pdv{}{q^{a}},\pdv{}{p_{b}}\right)&=\delta^{b}_{a}\ec \end{align*}

so the matrix of \(\omega_{\xi}\) is \(J\), which is invertible because \(J^{2}=-\identity_{2f}\) by Equation (22.56). A bilinear form with invertible matrix is nondegenerate, and the basis above is a canonical basis in the sense of Equation (24.1).

Intrinsic. Let \(\xi\in T^{*}Q\) lie over \(q=\pi(\xi)\) and choose an adapted chart, so that \(\xi=p_{a}\,\dd q^{a}\big|_{q}\) and a tangent vector \(v\in T_{\xi}(T^{*}Q)\) has components \(v=v_{q}^{a}\,\pp/\pp q^{a}+v_{p,a}\,\pp/\pp p_{a}\). The projection \(\pi\) forgets the momenta, so \(\dd\pi(v)=v_{q}^{a}\,\pp/\pp q^{a}\big|_{q}\), and therefore

\[ \xi\!\left(\dd\pi(v)\right)=p_{a}v_{q}^{a} =\left(p_{a}\,\dd q^{a}\right)(v)=\theta_{\xi}(v)\ep \]

The left-hand side is built from \(\xi\), \(v\) and \(\pi\) alone, none of which mentions a chart; the right-hand side is the coordinate expression Equation (24.3). The two agree in every adapted chart, so \(\theta\)—and with it \(\omega=-\dd\theta\)—is chart-independent. That the coordinate formula reproduces the definition by evaluating the covector on its own base point is what earns \(\theta\) the name tautological.

Remark 24.11 (The SI dimension of every object in this chapter).

Each product \(p_{a}q^{a}\) has the dimension of action, and every symplectic object is built from such products. So \(\theta\) and \(\omega\) carry the unit \(\mathrm{J}\,\mathrm{s}\); the Liouville form Equation (24.13) carries \(\left(\mathrm{J}\,\mathrm{s}\right)^{f}\); a symplectic capacity (Definition 24.33) carries \(\mathrm{J}\,\mathrm{s}\); and the Poisson bracket of two quantities \(u\) and \(v\) carries \([u][v]\) divided by \(\mathrm{J}\,\mathrm{s}\), which is why \(\pb{q^{a}}{p_{b}}\) in Equation (22.44) is a pure number. A Hamiltonian vector field \(X_{\Ham}\) built from an energy therefore carries \(/\mathrm{s}\): it is a rate, as its integral curves being trajectories requires. Two consequences are used later. First, \(\omega/\hbar\) is dimensionless, which is what allows it to be the curvature of a connection in Section 24.7; \(\hbar\) is the only constant of nature with the right unit to do that job, and Remark 5.116 already notes the coincidence. Second, the column \(\eta^{c}\) of Equation (22.55) mixes quantities of different dimension—the point of Remark 22.5—so no Euclidean length on phase space is dimensionally meaningful without a declared scale, a point that returns in Remark 24.27. Rests on Proposition 24.10 and Remark 22.5.

Theorem 24.12 (Darboux).

Every point of a symplectic manifold \((M,\omega)\) has a neighbourhood with coordinates \((q^{a},p_{a})\) in which

\begin{equation}\tag{24.4} \omega=\dd q^{a}\wedge\dd p_{a}\ep \end{equation}

Rests on Definition 24.8 and Theorem A.68.

Derives Theorem 24.12.

Theorem 24.12 is Theorem A.68, proved in Darboux's Theorem: Local Canonical Coordinates as part of Linear Algebra and Representation Theory and not reproved here. The argument there is Moser's deformation trick [Moser:1965]: connect \(\omega\) to its constant-coefficient value at the point by a path of closed nondegenerate forms, use nondegeneracy to solve for the vector field generating the isotopy, and integrate it. Its two hypotheses do exactly one job each—nondegeneracy supplies the vector field, closedness supplies the primitive through the local converse of the Poincaré lemma (Lemma A.77)—and the coordinates it produces are the canonical coordinates of Hamiltonian Mechanics. Corollary A.69 is the statement that any two symplectic manifolds of equal dimension are locally symplectomorphic.

Remark 24.13 (What Darboux's theorem forbids).

Darboux's theorem is the exact statement that a symplectic manifold has no local geometry. There is no symplectic analogue of curvature, no invariant distinguishing one point from another, and hence no local measurement that could reveal the symplectic structure of phase space. All symplectic invariants are global, which is why the theorems of this chapter that carry physical content—Liouville, Poincaré recurrence, the Liouville–Arnold tori of Theorem 23.31, the non-squeezing theorem—are statements about whole trajectories or whole regions and never about a neighbourhood of a point. The contrast with the Riemannian case of Geometric Formulation of Gravity, where curvature is a local observable and tidal forces measure it, is total. Rests on Theorem 24.12 and Corollary A.69.

Hamiltonian vector fields and the flow

Definition 24.14 (Hamiltonian vector field).

Let \(H\in C^{\infty}(M)\). Since \(\omega\) is nondegenerate, there is a unique vector field \(X_{H}\) with

\begin{equation}\tag{24.5} \iota_{X_{H}}\omega=\dd H\ep \end{equation}

It is the Hamiltonian vector field of \(H\). In Darboux coordinates

\begin{equation}\tag{24.6} X_{H}=\pdv{H}{p_{a}}\pdv{}{q^{a}} -\pdv{H}{q^{a}}\pdv{}{p_{a}}\ec \end{equation}

so that its integral curves are exactly the solutions of Hamilton's equations Equations (22.3) and (22.4). Rests on Definition 24.8, Equation (22.3) and Equation (22.4).

The coordinate form Equation (24.6) is a one-line check and is worth doing once, because every sign in this chapter descends from it. Write \(X=A^{a}\,\pp/\pp q^{a}+B_{a}\,\pp/\pp p_{a}\). Contracting \(\omega=\dd q^{a}\wedge\dd p_{a}\) with \(X\) and using the graded Leibniz rule for \(\iota_{X}\),

\[ \iota_{X}\omega =\left(\iota_{X}\dd q^{a}\right)\dd p_{a} -\dd q^{a}\left(\iota_{X}\dd p_{a}\right) =A^{a}\,\dd p_{a}-B_{a}\,\dd q^{a}\ec \]

while \(\dd H=\left(\pp H/\pp q^{a}\right)\dd q^{a} +\left(\pp H/\pp p_{a}\right)\dd p_{a}\). Matching the coefficients of \(\dd p_{a}\) gives \(A^{a}=\pp H/\pp p_{a}\) and matching those of \(\dd q^{a}\) gives \(B_{a}=-\pp H/\pp q^{a}\), which is Equation (24.6); these are Equations (22.3) and (22.4) read as the components of a vector field, and they are Equation (22.57) with the matrix \(J\) replaced by the form whose matrix it is.

Proposition 24.15 (Poisson bracket from the symplectic form).

With the conventions Equation (24.3) and Equation (24.5), the Poisson bracket of Definition 22.29 is

\begin{equation}\tag{24.7} \pb{u}{v}=\omega\left(X_{u},X_{v}\right)=X_{v}[u]\ec \end{equation}

and the assignment \(u\mapsto X_{u}\) satisfies

\begin{equation}\tag{24.8} X_{\pb{u}{v}}=-\left[X_{u},X_{v}\right]\ec \end{equation}

the bracket on the right being the commutator of vector fields. Rests on Definition 24.14, Definition 22.29 and Equation (22.46).

Proof.

Derives Proposition 24.15. The bracket. By Equation (24.5), \(\omega(X_{u},X_{v})=\left(\iota_{X_{u}}\omega\right)(X_{v}) =\dd u(X_{v})=X_{v}[u]\), which is the second equality in Equation (24.7) and holds on any symplectic manifold. For the first, work in Darboux coordinates and use Equation (24.6):

\begin{align*} X_{v}[u] &=\pdv{v}{p_{a}}\pdv{u}{q^{a}}-\pdv{v}{q^{a}}\pdv{u}{p_{a}}\\ &=\pdv{u}{q^{a}}\pdv{v}{p_{a}}-\pdv{u}{p_{a}}\pdv{v}{q^{a}}\ec \end{align*}

which is Equation (22.41). In particular \(\dot{u}=X_{\Ham}[u]=\pb{u}{\Ham}\) along the flow, so the geometric definition reproduces the evolution law of Hamiltonian Mechanics with no change of sign, and \(\pb{q^{a}}{p_{b}}=\delta^{a}_{b}\) as in Equation (22.44).

The homomorphism. Two vector fields are equal when they agree as derivations on every smooth function, so let \(w\in C^{\infty}(M)\) be arbitrary. Using \(X_{v}[w]=\pb{w}{v}\) twice,

\begin{align*} \left[X_{u},X_{v}\right][w] &=X_{u}\!\left[X_{v}[w]\right]-X_{v}\!\left[X_{u}[w]\right]\\ &=X_{u}\!\left[\pb{w}{v}\right]-X_{v}\!\left[\pb{w}{u}\right] =\pb{\pb{w}{v}}{u}-\pb{\pb{w}{u}}{v}\ep \end{align*}

On the other side, \(X_{\pb{u}{v}}[w]=\pb{w}{\pb{u}{v}}\), and the Jacobi identity Equation (22.46) in the form \(\pb{w}{\pb{u}{v}}+\pb{u}{\pb{v}{w}}+\pb{v}{\pb{w}{u}}=0\) gives, after using antisymmetry on the second term,

\begin{align*} X_{\pb{u}{v}}[w] &=\pb{u}{\pb{w}{v}}-\pb{v}{\pb{w}{u}}\\ &=-\pb{\pb{w}{v}}{u}+\pb{\pb{w}{u}}{v} =-\left[X_{u},X_{v}\right][w]\ep \end{align*}

Since \(w\) was arbitrary, Equation (24.8) follows. The minus sign is not removable: it is the statement that \(u\mapsto X_{u}\) is an antihomomorphism of Lie algebras with the conventions fixed above, and reversing it would require reversing the sign in either Equation (24.5) or Equation (22.41).

Theorem 24.16 (The Hamiltonian flow preserves the symplectic form).

Let \(\phi_{t}\) be the flow of \(X_{H}\). Then

\begin{equation}\tag{24.9} \mathcal{L}_{X_{H}}\omega=0\ec\qquad\text{equivalently}\qquad \phi_{t}^{*}\omega=\omega\quad\text{for all }t\ep \end{equation}

Rests on Definition 24.14, Proposition 13.128 and Definition 24.8.

Proof.

Derives Theorem 24.16. Cartan's magic formula Equation (13.276) applied to \(\omega\) and the field \(X_{H}\) reads

\[ \mathcal{L}_{X_{H}}\omega =\dd\left(\iota_{X_{H}}\omega\right) +\iota_{X_{H}}\left(\dd\omega\right)\ep \]

The first term is \(\dd\left(\dd H\right)=0\) by Equation (24.5) and the nilpotency Equation (13.245); the second is zero because \(\omega\) is closed. Each hypothesis of Definition 24.8 is used exactly once, which is why the definition demands both and no more.

For the second form, the Lie derivative is by Equation (13.267) the derivative of the dragged form, so that \(\dv{}{t}\phi_{t}^{*}\omega =\phi_{t}^{*}\left(\mathcal{L}_{X_{H}}\omega\right)=0\); hence \(\phi_{t}^{*}\omega\) is independent of \(t\) and equals its value \(\phi_{0}^{*}\omega=\omega\) at \(t=0\).

Definition 24.17 (Symplectomorphism).

A diffeomorphism \(\psi:M\to M\) with \(\psi^{*}\omega=\omega\) is a symplectomorphism, or canonical transformation. Rests on Definitions 24.8 and A.70.

Theorem 24.18 (Symplectic form of the transformation condition).

Let \(\eta\mapsto\zeta(\eta)\) be a transformation of phase space in the notation of Equation (22.55), with Jacobian matrix

\begin{equation}\tag{24.10} M^{c}{}_{d}=\pdv{\zeta^{c}}{\eta^{d}}\ep \end{equation}

The transformation is canonical if and only if

\begin{equation}\tag{24.11} M\transpose J\,M=J\ec \end{equation}

that is, if and only if its Jacobian lies in \(\Sp(2f,\R)\) at every point. This single matrix identity is equivalent to the \(4f^{2}\) direct conditions of Proposition 22.17, to the invariance of the fundamental Poisson brackets Equation (22.44), and to the invariance of the fundamental Lagrange brackets Equation (22.39). Rests on Definition 24.17, Equation (22.57) and Equation (22.42).

Proof.

Derives Theorem 24.18. The condition. Let the transformation be time-independent and invertible, write \(\Ham'(\zeta)=\Ham\!\left(\eta(\zeta)\right)\) for the same energy read in the new variables, and let \(\Ham\) be arbitrary. Differentiating \(\zeta=\zeta(\eta)\) along a trajectory and using Equation (22.57) for \(\eta\),

\[ \dot{\zeta}^{c}=M^{c}{}_{d}\,\dot{\eta}^{d} =M^{c}{}_{d}\,J^{d}{}_{e}\,\pdv{\Ham}{\eta^{e}}\ec \qquad \pdv{\Ham}{\eta^{e}}=M^{g}{}_{e}\,\pdv{\Ham'}{\zeta^{g}}\ec \]

the second being the chain rule. In matrix form \(\dot{\zeta}=M\,J\,M\transpose\,\pp\Ham'/\pp\zeta\); since \(\pp\Ham'/\pp\zeta\) may be prescribed arbitrarily at any one point by choosing \(\Ham\), the new variables obey Hamilton's equations with \(\Ham'\) for every \(\Ham\) if and only if

\begin{equation}\tag{24.12} M\,J\,M\transpose=J\ep \end{equation}

That Equation (24.12) and Equation (24.11) are equivalent is algebra. \(J\) is invertible, so either identity forces \(\det M\neq0\). Assume Equation (24.12); multiplying on the left by \(M^{-1}\) and on the right by \(\left(M\transpose\right)^{-1}\) gives \(J=M^{-1}J\left(M\transpose\right)^{-1}\), and inverting both sides with \(J^{-1}=-J\) gives \(-J=M\transpose\left(-J\right)M\), which is Equation (24.11). The same steps run backwards.

The three equivalent tests. All three are Equation (24.12) or Equation (24.11) read in a different bookkeeping. By Equation (22.42) the Poisson bracket of the new variables, computed in the old ones, is

\[ \pb{\zeta^{c}}{\zeta^{d}}_{q,p} =\left(\pdv{\zeta^{c}}{\eta}\right)\transpose J\left(\pdv{\zeta^{d}}{\eta}\right) =\left(M\,J\,M\transpose\right)^{cd}\ec \]

so the new variables satisfy the fundamental Poisson brackets Equation (22.44) exactly when Equation (24.12) holds. The Lagrange brackets of Equation (22.36) differentiate the other way round, \(\pp\eta/\pp\zeta=M^{-1}\), and give \(\left[\zeta^{c},\zeta^{d}\right]_{q,p} =\left(\left(M^{-1}\right)\transpose JM^{-1}\right)^{cd}\), which equals \(J\) exactly when Equation (24.11) holds for \(M^{-1}\), hence for \(M\). That the two families of brackets test the transposed identities is the duality recorded in Proposition 22.32. Finally, splitting \(M\) into the four \(f\times f\) blocks of Equation (22.55) and multiplying out Equation (24.11) block by block returns the four families Equations (22.16) and (22.17), which are \(4f^{2}\) scalar equations; this is Lemma 22.20, and it is what discharges the heading reserved without content in Section 22.5.1.

Coordinate-free reading. Nothing above is more than Definition 24.17 in components: \(\psi^{*}\omega=\omega\) says that the matrix of \(\omega\) is unchanged when transported by the Jacobian, and the matrix of \(\omega\) in canonical coordinates is \(J\).

Corollary 24.19 (Canonical transformations preserve phase volume).

\(\det M=1\) for every canonical transformation, so the phase-space volume element is invariant. Rests on Theorems 5.133 and 24.18.

Proof.

Derives Corollary 24.19. By Theorem 24.18 the Jacobian lies in \(\Sp(2f,\R)\) at every point, and every element of that group has determinant \(+1\) by Theorem 5.133. The volume element transforms as \(\dd^{2f}\zeta=\left(\det M\right)\dd^{2f}\eta=\dd^{2f}\eta\). Note that taking determinants directly in Equation (24.11) yields only \(\left(\det M\right)^{2}=1\); the sign is settled by the Pfaffian argument of Theorem 5.133 and not by the naive route.

Invariants of the flow

Definition 24.20 (Liouville volume).

The Liouville form of a symplectic manifold is the \(2f\)-form

\begin{align} \Omega&=\frac{1}{f!}\,\omega^{\wedge f} =\dd q^{1}\wedge\dd p_{1}\wedge\cdots \wedge\dd q^{f}\wedge\dd p_{f} \tag{24.13}\\ &=(-1)^{f(f-1)/2}\, \dd q^{1}\wedge\cdots\wedge\dd q^{f}\wedge \dd p_{1}\wedge\cdots\wedge\dd p_{f}\ec \tag{24.14} \end{align}

the last two expressions holding in Darboux coordinates. It is a volume form, and every symplectic manifold is therefore orientable. Rests on Definition 24.8, Definition 13.102 and Theorem 24.12.

The two coordinate expressions in Equation (24.13) differ by the sign of the permutation that separates the coordinates from the momenta, and the sign matters only for the choice of orientation, never for a volume. Writing \(\omega=\sum_{a}\omega_{a}\) with \(\omega_{a}=\dd q^{a}\wedge\dd p_{a}\), each \(\omega_{a}\) is a two-form and any two of them commute under \(\wedge\), while \(\omega_{a}\wedge\omega_{a}=0\); expanding the \(f\)-th power therefore leaves only the \(f!\) orderings of the distinct factors, so \(\omega^{\wedge f}=f!\,\omega_{1}\wedge\cdots\wedge\omega_{f}\), which is the first equality. Regrouping into all the \(\dd q\) followed by all the \(\dd p\) requires \(f(f-1)/2\) transpositions, which is the second.

Theorem 24.21 (Liouville).

The Hamiltonian flow preserves the Liouville volume:

\begin{equation}\tag{24.15} \mathcal{L}_{X_{H}}\Omega=0\ec\qquad\text{equivalently}\qquad \Div X_{H}=0\ep \end{equation}

The volume of any region of phase space is unchanged as the region is carried along by the motion, however violently the region is distorted. Rests on Definition 24.20, Theorem 24.16 and Proposition 7.73.

Proof.

Derives Theorem 24.21. Geometric route. The Lie derivative is a derivation of the exterior algebra, so it obeys the Leibniz rule on wedge products; applied to \(\Omega=\omega^{\wedge f}/f!\) this gives

\begin{align*} \mathcal{L}_{X_{H}}\Omega &=\frac{1}{f!}\sum_{k=1}^{f} \omega^{\wedge(k-1)}\wedge \left(\mathcal{L}_{X_{H}}\omega\right)\wedge \omega^{\wedge(f-k)}\\ &=\frac{1}{(f-1)!}\, \left(\mathcal{L}_{X_{H}}\omega\right)\wedge\omega^{\wedge(f-1)}\ec \end{align*}

using that two-forms commute under \(\wedge\), and this vanishes because \(\mathcal{L}_{X_{H}}\omega=0\) by Theorem 24.16. The Liouville form is built from \(\omega\) and nothing else, so it inherits every invariance \(\omega\) has.

Coordinate route. For a volume form \(\Omega\) the Lie derivative of a vector field satisfies \(\mathcal{L}_{X}\Omega=\left(\Div X\right)\Omega\), which in Darboux coordinates is the ordinary divergence. From Equation (24.6),

\begin{align*} \Div X_{H} &=\pdv{}{q^{a}}\!\left(\pdv{H}{p_{a}}\right) +\pdv{}{p_{a}}\!\left(-\pdv{H}{q^{a}}\right)\\ &=\frac{\pp^{2}H}{\pp q^{a}\pp p_{a}} -\frac{\pp^{2}H}{\pp p_{a}\pp q^{a}}=0\ec \end{align*}

the two terms cancelling term by term by the equality of mixed partial derivatives, Proposition 7.73. Note where the antisymmetry does its work: an arbitrary vector field on phase space has no reason to be divergence-free, and it is the opposite signs in Equation (24.6)—that is, the antisymmetry of \(J\)—that force the cancellation.

Remark 24.22 (Liouville's theorem is the foundation of statistical mechanics).

Theorem 24.21 is what makes the phase-space volume, and not some other measure, the natural counting measure for a mechanical system; the microcanonical ensemble of Statistical Mechanics is the uniform measure on an energy shell with respect to it, and it is distinguished by being invariant under the very dynamics whose long-time behaviour it is used to describe. The entropy so defined inherits the additive constant that only quantum mechanics fixes, through the cell size \(h^{f}\)—which is dimensionally the only thing it could be, by Remark 24.11. That this constant is measurable—in the Sackur–Tetrode equation for the absolute entropy of a monatomic gas [Sackur:1911] [Tetrode:1912]—is one of the places where phase-space geometry becomes an experimental number. Rests on Theorem 24.21 and Remark 24.11.

Integral invariants

Theorem 24.23 (Poincaré–Cartan integral invariant).

Let \(\gamma_{1}\) and \(\gamma_{2}\) be two closed curves encircling the same tube of trajectories in extended phase space \((q^{a},p_{a},t)\). Then

\begin{equation}\tag{24.16} \oint_{\gamma_{1}}\left(p_{a}\,\dd q^{a}-\Ham\,\dd t\right) =\oint_{\gamma_{2}}\left(p_{a}\,\dd q^{a}-\Ham\,\dd t\right)\ep \end{equation}

If both curves lie at a fixed instant, this reduces to the invariance of \(\oint p_{a}\,\dd q^{a}\), and hence, by Stokes' theorem, to Poincaré's theorem Theorem 22.22. Rests on Equation (24.3), Theorem A.311 and Equation (22.3).

Proof.

Derives Theorem 24.23. Work on extended phase space \(\widetilde{M}=M\times\R\) with coordinates \((q^{a},p_{a},t)\) and set \(\lambda=p_{a}\,\dd q^{a}-\Ham\,\dd t\), so that \(\dd\lambda=\dd p_{a}\wedge\dd q^{a}-\dd\Ham\wedge\dd t\), that is

\[ \dd\lambda =\dd p_{a}\wedge\dd q^{a} -\pdv{\Ham}{q^{a}}\,\dd q^{a}\wedge\dd t -\pdv{\Ham}{p_{a}}\,\dd p_{a}\wedge\dd t\ec \]

the term in \(\pp\Ham/\pp t\) dropping out because \(\dd t\wedge\dd t=0\). The trajectories of the system are the integral curves of

\[ Z=\pdv{}{t}+\pdv{\Ham}{p_{a}}\pdv{}{q^{a}} -\pdv{\Ham}{q^{a}}\pdv{}{p_{a}}\ec \]

that is, of \(X_{\Ham}\) carried along at unit rate in \(t\); this is Equation (24.6) together with \(\dot{t}=1\). Contract \(\dd\lambda\) with \(Z\), using \(\iota_{Z}\left(\alpha\wedge\beta\right) =\left(\iota_{Z}\alpha\right)\beta-\alpha\left(\iota_{Z}\beta\right)\) on one-forms \(\alpha,\beta\):

\begin{align*} \iota_{Z}\left(\dd p_{a}\wedge\dd q^{a}\right) &=-\pdv{\Ham}{q^{a}}\,\dd q^{a}-\pdv{\Ham}{p_{a}}\,\dd p_{a}\ec\\ -\pdv{\Ham}{q^{a}}\,\iota_{Z}\left(\dd q^{a}\wedge\dd t\right) &=-\pdv{\Ham}{q^{a}}\pdv{\Ham}{p_{a}}\,\dd t +\pdv{\Ham}{q^{a}}\,\dd q^{a}\ec\\ -\pdv{\Ham}{p_{a}}\,\iota_{Z}\left(\dd p_{a}\wedge\dd t\right) &=+\pdv{\Ham}{p_{a}}\pdv{\Ham}{q^{a}}\,\dd t +\pdv{\Ham}{p_{a}}\,\dd p_{a}\ep \end{align*}

Adding the three lines, the \(\dd q^{a}\) terms cancel, the \(\dd p_{a}\) terms cancel, and the two \(\dd t\) terms cancel against each other:

\begin{equation}\tag{24.17} \iota_{Z}\,\dd\lambda=0\ep \end{equation}

So the trajectory direction lies in the kernel of \(\dd\lambda\). Since \(\dd\lambda\) is a two-form on a space of odd dimension \(2f+1\) its kernel is at least one-dimensional; Equation (24.17) identifies it as the direction of the motion.

Now let \(T\) be the tube swept by the trajectories through a closed curve, a two-dimensional surface in \(\widetilde{M}\), and let \(\gamma_{1}\) and \(\gamma_{2}\) be two closed curves on \(T\) each encircling it once, with \(\Sigma\subset T\) the annulus they bound, so that \(\pp\Sigma=\gamma_{1}-\gamma_{2}\). At every point of \(\Sigma\) the tangent plane is two-dimensional and contains \(Z\); if \(v\) is any second tangent vector there, then \(\dd\lambda(Z,v)=0\) by Equation (24.17), so the pullback of \(\dd\lambda\) to \(\Sigma\) vanishes identically. The general Stokes theorem Theorem A.311 then gives

\[ \oint_{\gamma_{1}}\lambda-\oint_{\gamma_{2}}\lambda =\int_{\pp\Sigma}\lambda=\int_{\Sigma}\dd\lambda=0\ec \]

which is Equation (24.16).

If both curves lie in a surface of constant \(t\), then \(\dd t\) pulls back to zero on them and \(\lambda\) reduces to \(\theta=p_{a}\dd q^{a}\); the statement becomes the invariance of \(\oint p_{a}\,\dd q^{a}\) under the flow, and applying Theorem A.311 once more turns that into the invariance of \(\int\omega\) over a surface, which is Theorem 22.22. That the first Poincaré invariant is the only independent one is the content of Equation (24.13): the higher invariants are integrals of \(\omega^{\wedge k}\) and follow from the first.

Theorem 24.24 (Poincaré recurrence).

Let the Hamiltonian flow leave invariant a region \(D\subset M\) of finite Liouville volume—as it does, for instance, for bounded motion at fixed energy. Then for every neighbourhood \(U\subset D\), almost every point of \(U\) returns to \(U\) at arbitrarily late times. Rests on Theorem 24.21 and Definition 24.20.

Proof.

Derives Theorem 24.24. Write \(\mu\) for the Liouville volume of Definition 24.20, so that \(\mu(D)<\infty\), and fix a time step \(\tau>0\); let \(\phi=\phi_{\tau}\) be the time-\(\tau\) map of the flow. By Theorem 24.21 \(\phi\) preserves \(\mu\), and it maps \(D\) to itself.

Fix an integer \(k\geq1\) and let

\[ B_{k}=\set{x\in U\ :\ \phi^{nk}(x)\notin U \ \text{for every integer}\ n\geq1} \]

be the set of points of \(U\) that never return under the map \(\phi^{k}\). The sets \(\phi^{-nk}\left(B_{k}\right)\), \(n=0,1,2,\ldots\), are pairwise disjoint. Indeed, suppose \(x\) lay in both \(\phi^{-nk}\left(B_{k}\right)\) and \(\phi^{-mk}\left(B_{k}\right)\) with \(n<m\). Then \(y=\phi^{nk}(x)\in B_{k}\subset U\) while \(\phi^{(m-n)k}(y)=\phi^{mk}(x)\in B_{k}\subset U\), so \(y\) returns to \(U\) after \(m-n\geq1\) steps of \(\phi^{k}\), contradicting \(y\in B_{k}\).

Each of these sets lies in \(D\) and has the same volume \(\mu\left(B_{k}\right)\), because \(\phi\) preserves \(\mu\). If \(\mu\left(B_{k}\right)>0\) then the volume of the disjoint union of the first \(N\) of them is \(N\mu\left(B_{k}\right)\), which exceeds \(\mu(D)\) for \(N\) large enough—impossible, since they all sit inside \(D\). Hence \(\mu\left(B_{k}\right)=0\) for every \(k\).

Finally, a point of \(U\) that returns for arbitrarily large multiples of \(\tau\) is one lying outside every \(B_{k}\): if \(x\notin B_{k}\) then \(x\) returns to \(U\) at some time that is a positive multiple of \(k\tau\), and letting \(k\) run makes those return times unbounded. The exceptional set is \(\bigcup_{k\geq1}B_{k}\), a countable union of sets of volume zero and therefore of volume zero. So almost every point of \(U\) returns at arbitrarily late times [Poincare:1890].

Remark 24.25 (What the recurrence proof assumes about volume).

The argument just given treats the Liouville volume as a measure: it uses additivity over disjoint sets, monotonicity, invariance under a volume-preserving map, and—in the last step only—the fact that a countable union of null sets is null. Those are the axioms of a measure space, and this treatise does not carry measure theory: the probability chapter of Part II is built on densities and expectations and never defines a \(\sigma\)-algebra or a countably additive measure. The statement is therefore made here on that understanding, and the reader who wants the last step in full rigour needs the countable additivity that Part II owes. Nothing else in this chapter uses it; and the physically relevant half of the theorem—that a positive-volume set of initial conditions must return—needs only finite additivity, which is the elementary property of volume used everywhere else in this part. Rests on Theorem 24.24.

Remark 24.26 (Recurrence and irreversibility).

Theorem 24.24 says that a bounded mechanical system returns arbitrarily close to its initial state, which appears to contradict the second law of Classical Thermodynamics. It does not: the recurrence times for a macroscopic number of degrees of freedom exceed the age of the universe by an enormous factor, and the theorem is silent about them. The tension it does expose—between a time-reversible microscopic dynamics and an irreversible phenomenology—is Loschmidt's and Zermelo's objection, and it is treated where it belongs, in Statistical Mechanics. Rests on Theorem 24.24.

Symplectic capacity and non-squeezing

Liouville's theorem constrains the volume of a phase-space region and nothing else. It is natural to ask whether that is the whole story— whether any volume-preserving deformation of a region can be realised by a Hamiltonian flow. It cannot, and the obstruction is not a refinement of volume but an invariant of a different kind. This subsection states it. It is written as mathematics: this treatise records no experiment that measures a symplectic capacity, and the frequently drawn analogy with the uncertainty relation of The Postulates of Quantum Mechanics is suggestive rather than derivational, a point made precise in Remark 24.35.

Remark 24.27 (Fixing a scale, so that a ``ball'' means something).

A ball in phase space is not a dimensionally meaningful object: by Remark 24.11 a coordinate and its conjugate momentum carry different units, and no sum of their squares is a length. To speak of balls and cylinders one must first declare a scale. Fix a length \(\ell>0\) and an action \(S>0\) and pass to the dimensionless coordinates

\begin{equation}\tag{24.18} u^{a}=\frac{q^{a}}{\ell}\ec\qquad v_{a}=\frac{\ell\,p_{a}}{S}\ec\qquad\text{so that}\qquad \omega=S\,\dd u^{a}\wedge\dd v_{a}\ep \end{equation}

All statements below are made for the standard form \(\omega_{0}=\dd u^{a}\wedge\dd v_{a}\) on \(\R^{2f}\), and are converted back by the factor \(S\); every capacity then comes out as a pure number times \(S\), that is, in \(\mathrm{J}\,\mathrm{s}\). Different choices of \((\ell,S)\) give different balls, so the theorem below compares a ball and a cylinder built with the same choice. Write \(\langle\cdot,\cdot\rangle\) for the Euclidean inner product in these coordinates and \(\abs{\cdot}\) for its norm; then, with \(J\) the matrix Equation (22.56) and \(z,w\in\R^{2f}\) written as columns in the block order of Equation (22.55),

\begin{equation}\tag{24.19} \omega_{0}(z,w)=\gen{z,J\,w}\ec\qquad J\transpose J=\identity_{2f}\ec\qquad \gen{J\,z,z}=0\ep \end{equation}

The last two follow from \(J\transpose=-J\) and \(J^{2}=-\identity_{2f}\): \(J\) is an orthogonal matrix and \(J z\) is orthogonal to \(z\) for every \(z\). Rests on Remark 24.11 and Equation (22.56).

Definition 24.28 (Ball and cylinder).

In \(\left(\R^{2f},\omega_{0}\right)\) with the coordinates Equation (24.18), the ball of radius \(\rho\) and the cylinder of radius \(r\) over the first conjugate pair are

\begin{align} B(\rho)&=\set{z\in\R^{2f}\ :\ \textstyle\sum_{a}\left[\left(u^{a}\right)^{2} +\left(v_{a}\right)^{2}\right]<\rho^{2}}\ec \tag{24.20}\\ Z(r)&=\set{z\in\R^{2f}\ :\ \left(u^{1}\right)^{2}+\left(v_{1}\right)^{2}<r^{2}}\ep \tag{24.21} \end{align}

A symplectic embedding of one subset of \(\R^{2f}\) into another is a smooth embedding \(\varphi\) with \(\varphi^{*}\omega_{0}=\omega_{0}\). Rests on Remark 24.27 and Definition 24.17.

\(Z(r)\) has infinite volume for every \(r>0\), so no volume argument can prevent a ball of any radius from being squeezed into it. And volume preservation alone does not prevent it either:

Example 24.29 (A volume-preserving squeeze).

Take \(f=3\)—one particle in three-dimensional space—and, for \(\lambda>0\), the linear map

\begin{align*} \left(u^{1},v_{1}\right)&\mapsto \left(\lambda u^{1},\lambda v_{1}\right)\ec\\ \left(u^{2},v_{2}\right)&\mapsto \left(\lambda^{-1}u^{2},\lambda^{-1}v_{2}\right)\ec\\ \left(u^{3},v_{3}\right)&\mapsto \left(u^{3},v_{3}\right)\ep \end{align*}

Its determinant is \(\lambda^{2}\cdot\lambda^{-2}\cdot1=1\), so it preserves the Liouville volume exactly; and it carries \(B(\rho)\) into \(Z(\lambda\rho)\), which for small \(\lambda\) is as thin as one likes. It is not symplectic: it pulls \(\omega_{0}\) back to \(\lambda^{2}\,\dd u^{1}\wedge\dd v_{1} +\lambda^{-2}\,\dd u^{2}\wedge\dd v_{2} +\dd u^{3}\wedge\dd v_{3}\), which is not \(\omega_{0}\). The squeezing that volume preservation permits is thus achieved by trading area between two different conjugate pairs—exactly what a symplectic map may not do. Rests on Definition 24.28 and Theorem 24.21.

Proposition 24.30 (Linear non-squeezing).

Let \(A\in\Sp(2f,\R)\) and suppose \(A\left(B(\rho)\right)\subseteq Z(r)\). Then \(r\geq\rho\). Rests on Definition 24.28, Equation (24.19) and Definition 24.4.

Proof.

Derives Proposition 24.30. Let \(e_{1}\) and \(f^{1}\) be the first pair of the canonical basis, so that \(\omega_{0}(e_{1},f^{1})=1\) by Equation (24.1). For any \(w\in\R^{2f}\) the two coordinates in which the cylinder is defined are

\[ u^{1}(w)=\omega_{0}\!\left(w,f^{1}\right)\ec\qquad v_{1}(w)=-\omega_{0}\!\left(w,e_{1}\right)\ec \]

as one checks by evaluating \(\omega_{0}=\dd u^{a}\wedge\dd v_{a}\) on the basis. Because \(A\) preserves \(\omega_{0}\), \(\omega_{0}(Az,f^{1})=\omega_{0}\!\left(z,A^{-1}f^{1}\right)\) and likewise for \(e_{1}\). Put \(a=A^{-1}e_{1}\) and \(b=A^{-1}f^{1}\), so that

\begin{equation}\tag{24.22} \omega_{0}(a,b)=\omega_{0}\!\left(e_{1},f^{1}\right)=1\ec \end{equation}

and the composition of \(A\) with the projection onto the cylinder plane is the linear map \(L:\R^{2f}\to\R^{2}\),

\[ L(z)=\left(\omega_{0}(z,b),\,-\omega_{0}(z,a)\right) =\left(\gen{z,J\,b},\,-\gen{z,J\,a}\right)\ec \]

by Equation (24.19). The hypothesis \(A\left(B(\rho)\right)\subseteq Z(r)\) says that \(\abs{L(z)}<r\) for every \(z\) with \(\abs{z}<\rho\), hence by continuity \(\abs{L(z)}\leq r\) on the closed ball, that is \(\rho\norm{L}\leq r\) with \(\norm{L}\) the operator norm.

It remains to show \(\norm{L}\geq1\). The matrix of \(L\) has the two rows \(\left(J b\right)\transpose\) and \(-\left(J a\right)\transpose\), so

\[ L\,L\transpose= \begin{pmatrix} \abs{b}^{2} & -\gen{a,b}\\ -\gen{a,b} & \abs{a}^{2} \end{pmatrix}\ec \quad \det\left(L\,L\transpose\right) =\abs{a}^{2}\abs{b}^{2}-\gen{a,b}^{2}\ec \]

where \(\abs{Jb}=\abs{b}\), \(\abs{Ja}=\abs{a}\) and \(\gen{Jb,Ja}=\gen{b,a}\) because \(J\) is orthogonal by Equation (24.19). The determinant is the square of the Euclidean area \(\mathcal{A}\) of the parallelogram spanned by \(a\) and \(b\). Now \(Jb\) is orthogonal to \(b\) and has the same length, so \(Jb/\abs{b}\) is a unit vector perpendicular to \(b\), and the component of \(a\) along it is at most the distance from \(a\) to the line \(\R b\):

\[ 1=\abs{\omega_{0}(a,b)}=\abs{\gen{a,J b}} =\abs{b}\,\left|\gen{a,\tfrac{J b}{\abs{b}}}\right| \leq\abs{b}\cdot\operatorname{dist}\!\left(a,\R b\right) =\mathcal{A}\ec \]

using Equation (24.22). Hence \(\det\left(L L\transpose\right)=\mathcal{A}^{2}\geq1\). The matrix \(L L\transpose\) is symmetric positive semidefinite with two eigenvalues whose product is at least \(1\), so its largest eigenvalue is at least \(1\); and \(\norm{L}^{2}\) is that largest eigenvalue. Therefore \(\norm{L}\geq1\), and \(r\geq\rho\norm{L}\geq\rho\).

The proposition is elementary because a linear symplectic map sends the ball to an ellipsoid and the argument is one about areas of shadows. Gromov's theorem is the same statement without linearity, and it is not elementary at all: it was the first application of pseudoholomorphic curves, and no proof by the methods of this chapter is known.

Theorem 24.31 (Gromov's non-squeezing theorem).

Let \(f\geq2\) and \(\rho,r>0\). If there exists a symplectic embedding \(\varphi:B(\rho)\hookrightarrow Z(r)\) in the sense of Definition 24.28, then \(r\geq\rho\). Rests on Definition 24.28 and Proposition 24.30.

Remark 24.32 (Non-squeezing is quoted, and will not be proved here).

Theorem 24.31 carries no derivation in this book, and the omission is a decision rather than a debt. It is due to Gromov, in Pseudo holomorphic curves in symplectic manifolds (Inventiones Mathematicae 82, 1985); that paper has no key in this treatise's bibliography, so the attribution is made in words, on the same footing as Darboux's own memoir in Darboux's Theorem: Local Canonical Coordinates. The proof constructs a pseudoholomorphic disc through a prescribed point of the image and bounds its area by an elliptic-analysis argument—compactness of a moduli space of such discs, and a monotonicity estimate for their area—machinery that appears nowhere else in this treatise and that no other result here needs. Introducing it to serve one statement would unbalance the chapter, and this is the same treatment Part II gives Carleson's theorem in Remark 17.18.

No shortcut through the linear case exists. Proposition 24.30, which is proved in full by the methods of this chapter, constrains only linear symplectic maps and says nothing about a nonlinear embedding, and the obvious replacement argument—that a symplectomorphism preserves volume, so it cannot squeeze—is simply false, by Example 24.29. What the theorem is used for below is one statement, that the Gromov width of Equation (24.23) satisfies the third capacity axiom of Definition 24.33, and everything this chapter concludes from that is flagged in Remark 24.35 as mathematics rather than as measured physics.

Definition 24.33 (Symplectic capacity).

A symplectic capacity is an assignment \(c\) of a number in \([0,\infty]\) to every pair \(\left(U,\alpha\omega_{0}\right)\) with \(U\subseteq\R^{2f}\) and \(\alpha\neq0\) real, such that

  1. monotonicity: if \(\left(U,\alpha\omega_{0}\right)\hookrightarrow \left(V,\beta\omega_{0}\right)\) admits a symplectic embedding, then \(c\left(U,\alpha\omega_{0}\right)\leq c\left(V,\beta\omega_{0}\right)\);

  2. conformality: \(c\left(U,\alpha\omega_{0}\right) =\abs{\alpha}\,c\left(U,\omega_{0}\right)\);

  3. nontriviality: \(c\left(B(1),\omega_{0}\right)=c\left(Z(1),\omega_{0}\right)=\pi\).

The argument \(\omega_{0}\) is suppressed when it is the standard form. By Remark 24.27 a capacity of a region of physical phase space carries the unit \(\mathrm{J}\,\mathrm{s}\). Rests on Definition 24.28 and Remark 24.27.

The three axioms are those of Ekeland and Hofer. Only the third mentions a specific set, and it is the one that does all the work. Monotonicity and conformality alone are cheap: the \(f\)-th root of the Liouville volume satisfies both, since a symplectic embedding preserves volume and rescaling \(\omega_{0}\) by \(\alpha\) rescales the volume by \(\abs{\alpha}^{f}\). What that candidate cannot do is meet the third axiom, because \(Z(1)\) has infinite volume as soon as \(f\geq2\); in \(f=1\) the area is a capacity, and the whole content of the subject is that the case \(f\geq2\) is not vacuous. The following makes that precise—the existence of a capacity and the non-squeezing theorem are the same statement.

Proposition 24.34 (Capacities exist if and only if non-squeezing holds).

Fix \(f\geq1\). There exists a symplectic capacity on \(\left(\R^{2f},\omega_{0}\right)\) if and only if the non-squeezing statement of Theorem 24.31 holds in that dimension—for \(f=1\) the elementary statement that an area-preserving embedding of a disc into a disc cannot decrease the radius, for \(f\geq2\) Gromov's theorem. In that case the Gromov width

\begin{equation}\tag{24.23} c_{G}(U)=\sup\set{\pi\rho^{2}\ :\ B(\rho)\ \text{embeds symplectically in}\ U} \end{equation}

is a capacity, and it is the smallest one. Rests on Definition 24.33 and Theorem 24.31.

Proof.

Derives Proposition 24.34. First note that any capacity takes the value \(\pi\rho^{2}\) on \(B(\rho)\) and on \(Z(\rho)\). The dilation \(\delta_{\rho}(z)=\rho z\) carries \(B(1)\) onto \(B(\rho)\) and satisfies \(\delta_{\rho}^{*}\omega_{0}=\rho^{2}\omega_{0}\), so it is a symplectomorphism from \(\left(B(1),\rho^{2}\omega_{0}\right)\) to \(\left(B(\rho),\omega_{0}\right)\); monotonicity applied in both directions and then conformality with \(\alpha=\rho^{2}\) give

\[ c\left(B(\rho),\omega_{0}\right) =c\left(B(1),\rho^{2}\omega_{0}\right) =\rho^{2}c\left(B(1),\omega_{0}\right)=\pi\rho^{2}\ec \]

and the identical argument applies to \(Z(r)\), giving \(\pi r^{2}\).

A capacity implies non-squeezing. Let \(c\) be a capacity and let \(\varphi:B(\rho)\hookrightarrow Z(r)\) be a symplectic embedding. Then monotonicity gives \(\pi\rho^{2}=c\left(B(\rho)\right)\leq c\left(Z(r)\right)=\pi r^{2}\), hence \(r\geq\rho\).

Non-squeezing implies a capacity. Assume Theorem 24.31 and consider \(c_{G}\) of Equation (24.23). Monotonicity is immediate: a symplectic embedding \(U\hookrightarrow V\) composes with any \(B(\rho)\hookrightarrow U\), so every \(\rho\) admissible for \(U\) is admissible for \(V\). Conformality follows from the dilation computation above, since precomposing an embedding into \(\left(U,\alpha\omega_{0}\right)\) with \(\delta_{\sqrt{\alpha}}\) shows that \(\rho\) is admissible for \(\alpha\omega_{0}\) exactly when \(\rho/\sqrt{\alpha}\) is admissible for \(\omega_{0}\), so the supremum of \(\pi\rho^{2}\) is multiplied by \(\alpha>0\); the case \(\alpha<0\) reduces to it by the reflection \(v_{a}\mapsto-v_{a}\) in every pair, which carries \(\omega_{0}\) to \(-\omega_{0}\) and leaves every ball and every cylinder where it was. For nontriviality, the identity embeds \(B(1)\) in itself and in \(Z(1)\), so \(c_{G}\left(B(1)\right)\geq\pi\) and \(c_{G}\left(Z(1)\right)\geq\pi\); conversely a symplectic embedding \(B(\rho)\hookrightarrow Z(1)\) forces \(\rho\leq1\) by Theorem 24.31, so \(c_{G}\left(Z(1)\right)\leq\pi\), and \(B(1)\subset Z(1)\) with monotonicity gives \(c_{G}\left(B(1)\right)\leq\pi\) as well.

Minimality. If \(c\) is any capacity and \(B(\rho)\) embeds in \(U\), then \(\pi\rho^{2}=c\left(B(\rho)\right)\leq c(U)\); taking the supremum over such \(\rho\) gives \(c_{G}(U)\leq c(U)\).

Remark 24.35 (What non-squeezing does and does not say about nature).

A Hamiltonian flow is a symplectomorphism (Theorem 24.16), so the Gromov width of a phase-space region is conserved along the motion just as its volume is. That is a genuine constraint on classical dynamics, and it is strictly stronger than Liouville's: by Example 24.29 volume preservation permits a squeeze that Theorem 24.31 forbids. The shadow of a Hamiltonian-transported region on any one conjugate plane cannot be made arbitrarily small while the region stays large in the other directions.

Three cautions keep the statement honest. First, this treatise records no experiment that measures a symplectic capacity, and the section is therefore mathematics, not evidence-based physics. Second, the resemblance to the uncertainty relation of The Postulates of Quantum Mechanics is formal only: Definition 24.33 assigns a capacity to a region, and classical mechanics puts no floor under it—a single point, or a Lagrangian sheet of any extent, has capacity zero, and there is no classical reason for the number \(S\) of Remark 24.27 to be \(\hbar\) rather than anything else. The uncertainty relation is a theorem about noncommuting operators and states, not about embeddings, and it is the quantum postulate that supplies the scale. Third, what the analogy does establish is a dimensional coincidence worth keeping in view: a capacity carries \(\mathrm{J}\,\mathrm{s}\) by Remark 24.27, which is the unit of \(\hbar\), and Remark 24.11 traces that back to the pairing of a coordinate with its momentum. It is the same coincidence that makes Section 24.7 possible. Rests on Theorem 24.31, Theorem 24.16 and Remark 24.11.

Poisson manifolds

Not every space carrying a Poisson bracket is a symplectic manifold. The bracket, not the form, is the more general object, and it is the one mechanics actually needs—rigid-body dynamics lives on a space where the form degenerates.

Definition 24.36 (Poisson manifold).

A Poisson manifold is a manifold \(P\) with a bilinear operation \(\pb{\cdot}{\cdot}\) on \(C^{\infty}(P)\) that is antisymmetric, satisfies the Jacobi identity Equation (22.46), and obeys the Leibniz rule

\begin{equation}\tag{24.24} \pb{u}{vw}=\pb{u}{v}w+v\pb{u}{w}\ep \end{equation}

Equivalently, it carries a bivector field \(\pi\in\Gamma(\Lambda^{2}TP)\) with \(\pb{u}{v}=\pi(\dd u,\dd v)\) whose Schouten–Nijenhuis bracket with itself vanishes, \(\left[\pi,\pi\right]=0\); that vanishing is the Jacobi identity. Rests on Definition 22.29 and Equation (22.46).

Definition 24.37 (Casimir function).

A Casimir is a function \(C\) with \(\pb{C}{u}=0\) for every \(u\). A Casimir is conserved by every Hamiltonian, whatever it is; on a symplectic manifold only the constants are Casimirs, so a nonconstant Casimir is exactly a signature of degeneracy. Rests on Definition 24.36 and Proposition 24.15.

That the only Casimirs of a symplectic manifold are the constants is immediate from Proposition 24.15: \(\pb{C}{u}=-X_{C}[u]\) vanishes for every \(u\) only if \(X_{C}=0\), and then \(\dd C=\iota_{X_{C}}\omega=0\) by Equation (24.5), so \(C\) is locally constant. On a Poisson manifold the map \(\dd C\mapsto X_{C}\) has a kernel, and the Casimirs are what live in it.

Theorem 24.38 (Symplectic foliation, quoted).

A Poisson manifold is partitioned into immersed submanifolds—the symplectic leaves—each of which carries a symplectic form inducing the ambient bracket. The leaves are the level sets of the Casimirs where those are regular, and the dynamics of any Hamiltonian stays on the leaf it starts on. Rests on Definitions 24.8, 24.36 and 24.37.

Remark 24.39 (The foliation is quoted; its last clause is proved).

Only the existence of the leaves is quoted here, and it is worth being exact about which half that is. The bracket of Definition 24.36 assigns to each point the image of the bivector \(\pi\), a subspace of the tangent space spanned by the Hamiltonian vector fields; the Jacobi identity makes that family of subspaces involutive, since \(\left[X_{u},X_{v}\right]=-X_{\pb{u}{v}}\) is again a Hamiltonian vector field—the computation proving Equation (24.8) uses only the Jacobi identity and the derivation property, both of which Definition 24.36 supplies—and integrating an involutive family into submanifolds is the Frobenius theorem. The manifolds chapter Differentiable Manifolds, Tensors, and Curvature does not carry it, and even if it did it would not suffice: the rank of \(\pi\) varies from point to point—for the Lie–Poisson bracket Equation (24.26) of the rigid body it is zero at the origin of \(\mathfrak{so}(3)^{*}\) and two everywhere else—so the distribution is singular, and the tool required is the integrability theorem for singular distributions rather than the constant-rank Frobenius theorem. Weinstein's splitting theorem, the local normal form that generalizes Darboux (Theorem 24.12) to the Poisson case, belongs with the same missing half. All of it is differential topology, and this chapter states it rather than proving it for the reason given in Remark 17.18: the machinery has no other use in this book.

The last clause of Theorem 24.38 is the one the rigid body of Proposition 24.41 actually consumes, and it needs nothing beyond the definitions. If \(C\) is a Casimir then \(\dot{C}=\pb{C}{\Ham}=0\) along the flow of any \(\Ham\) whatever, so every Casimir is a constant of the motion and a trajectory stays in the level set of every Casimir it starts in. Where those level sets are regular they are the leaves, so no application below rests on the quoted half.

Definition 24.40 (Lie–Poisson bracket).

Let \(\mathfrak{g}\) be a Lie algebra and \(\mathfrak{g}^{*}\) its dual. The Lie–Poisson bracket on \(\mathfrak{g}^{*}\) is

\begin{equation}\tag{24.25} \pb{u}{v}_{\pm}(\mu)=\pm\left\langle\mu, \left[\frac{\delta u}{\delta\mu}, \frac{\delta v}{\delta\mu}\right]\right\rangle\ec \qquad\mu\in\mathfrak{g}^{*}\ec \end{equation}

the inner bracket being that of \(\mathfrak{g}\). Both signs give a Poisson structure in the sense of Definition 24.36—antisymmetry and the Leibniz rule are unaffected by an overall sign, and the Jacobi identity is quadratic in the bracket and therefore also unaffected—and which one a given mechanical system requires is fixed by the physics, not by convention. The rigid body referred to body axes takes the minus sign, Proposition 24.41. Rests on Definitions 5.123 and 24.36.

Proposition 24.41 (The free rigid body is a Lie–Poisson system).

Take \(\mathfrak{g}=\mathfrak{so}(3)\), identify \(\mathfrak{g}^{*}\) with \(\R^{3}\) so that \(\mu=\vect{L}\) is the angular momentum referred to body axes, and let \(\Ham=\tfrac{1}{2}L_{i}\left(I^{-1}\right)_{ij}L_{j}\) with \(I\) the inertia tensor. Then the minus Lie–Poisson bracket of Equation (24.25) is

\begin{equation}\tag{24.26} \pb{u}{v}_{-} =-\vect{L}\cdot\left(\nabla u\times\nabla v\right)\ec \qquad\text{so that}\qquad \pb{L_{i}}{L_{j}}_{-}=-\varepsilon_{ijk}L_{k}\ec \end{equation}

the gradients being taken with respect to \(\vect{L}\), and the equations of motion it generates are Euler's equations Theorem 29.25 for the torque-free body,

\begin{equation}\tag{24.27} \dv{\vect{L}}{t}=\vect{L}\times\vect{\Omega}\ec\qquad \vect{\Omega}=I^{-1}\vect{L}\ep \end{equation}

The function \(C=\abs{\vect{L}}^{2}\) is a Casimir, the leaves are the spheres \(\abs{\vect{L}}=\text{constant}\), and the trajectories are the intersections of those spheres with the energy ellipsoids \(\Ham=\text{constant}\). Rests on Definition 24.40, Theorem 29.25 and Definition 24.37.

Proof.

Derives Proposition 24.41. The bracket. Take the basis \(e_{i}\) of \(\mathfrak{so}(3)\) with \(\comm{e_{j}}{e_{k}}=\varepsilon_{jkl}e_{l}\), so that the structure constants are the Levi-Civita symbol (Definition 14.26 and Lemma 14.62), and identify \(\mathfrak{g}^{*}\cong\R^{3}\) by \(\left\langle\mu,e_{i}\right\rangle=L_{i}\). For a function \(u\) of \(\vect{L}\) the functional derivative is the ordinary gradient, \(\delta u/\delta\mu=\left(\pp u/\pp L_{j}\right)e_{j}\). Then

\begin{align*} \pb{u}{v}_{-} &=-\left\langle\mu, \comm{\pdv{u}{L_{j}}e_{j}}{\pdv{v}{L_{k}}e_{k}}\right\rangle =-\varepsilon_{jkl}L_{l}\,\pdv{u}{L_{j}}\pdv{v}{L_{k}}\\ &=-\vect{L}\cdot\left(\nabla u\times\nabla v\right)\ec \end{align*}

since \(\left(\nabla u\times\nabla v\right)_{l} =\varepsilon_{ljk}\left(\pp u/\pp L_{j}\right) \left(\pp v/\pp L_{k}\right)\) and \(\varepsilon_{jkl}=\varepsilon_{ljk}\). Putting \(u=L_{i}\), \(v=L_{j}\) gives \(\pb{L_{i}}{L_{j}}_{-}=-\vect{L}\cdot\left(\hat{e}_{i}\times \hat{e}_{j}\right)=-\varepsilon_{ijk}L_{k}\), which is the bracket declared in Remark 29.27.

The equations of motion. With \(\nabla\Ham=I^{-1}\vect{L}=\vect{\Omega}\)—the body-frame angular velocity, since \(\vect{L}=I\vect{\Omega}\)—and \(\nabla L_{i} =\hat{e}_{i}\),

\begin{align*} \dot{L}_{i}&=\pb{L_{i}}{\Ham}_{-} =-\vect{L}\cdot\left(\hat{e}_{i}\times\vect{\Omega}\right) =-\varepsilon_{mik}L_{m}\Omega_{k}\\ &=\varepsilon_{imk}L_{m}\Omega_{k} =\left(\vect{L}\times\vect{\Omega}\right)_{i}\ec \end{align*}

which is Equation (24.27), that is \(\dd\vect{L}/\dd t+\vect{\Omega}\times\vect{L}=0\)— Equation (29.21) with \(\vect{\tau}=\vect{0}\). In principal axes \(L_{i}=I_{i}\Omega_{i}\) and the components read \(I_{1}\dot{\Omega}_{1}=\left(I_{2}-I_{3}\right)\Omega_{2}\Omega_{3}\) and its cyclic partners, which are Equations (29.18), (29.19) and (29.20) with zero torque. Had the plus sign of Equation (24.25) been taken, the same computation would give \(\dd\vect{L}/\dd t=\vect{\Omega}\times\vect{L}\), the time reverse; that is why the sign is part of the physics.

The Casimir and the leaves. For \(C=\abs{\vect{L}}^{2}\) one has \(\nabla C=2\vect{L}\), so for every \(v\)

\[ \pb{C}{v}_{-} =-\vect{L}\cdot\left(2\vect{L}\times\nabla v\right)=0 \]

because the triple product has a repeated factor. Hence \(C\) is a Casimir in the sense of Definition 24.37, its regular level sets—the spheres of radius \(\abs{\vect{L}}>0\)—are the symplectic leaves, and every trajectory stays on one of them. Intersecting with the conserved \(\Ham\) gives the closed curves of Proposition 29.26. The odd dimension of \(\R^{3}\) is what forces this structure: no symplectic form exists there, and the two-dimensional leaves are where the symplectic geometry of this chapter actually lives, as Remark 29.27 anticipates.

Remark 24.42 (The intermediate axis).

The leaf picture settles the qualitative theory of free rotation in one figure and one line of algebra. Order the principal moments \(I_{1},I_{2},I_{3}\) and intersect the sphere \(\abs{\vect{L}}=L\) with the ellipsoids \(\Ham=E\): near the axes of largest and of smallest moment the intersections are small closed curves encircling the axis, and near the intermediate axis they cross. Linearising Equation (24.27) about steady rotation \(\vect{\Omega}=(\Omega,0,0)\) gives

\[ I_{2}\dot{\Omega}_{2}=\left(I_{3}-I_{1}\right)\Omega\,\Omega_{3}\ec \qquad I_{3}\dot{\Omega}_{3}=\left(I_{1}-I_{2}\right)\Omega\,\Omega_{2}\ec \]

and eliminating \(\Omega_{3}\),

\[ \ddot{\Omega}_{2} =\frac{\left(I_{3}-I_{1}\right)\left(I_{1}-I_{2}\right)}{I_{2}I_{3}} \,\Omega^{2}\,\Omega_{2}\ec \]

whose solutions grow exponentially exactly when \(\left(I_{3}-I_{1}\right)\left(I_{1}-I_{2}\right)>0\), that is when \(I_{1}\) lies strictly between the other two. Rotation about the intermediate axis is unstable and rotation about the other two is stable—the tennis-racket effect, and the reason a thrown book tumbles about one axis and not the other three. The rigid-body chapter treats it in full, Section 29.3. Rests on Propositions 24.41 and 29.26.

Symmetry, momentum maps, and reduction

Definition 24.43 (Momentum map).

Let a Lie group \(G\) with Lie algebra \(\mathfrak{g}\) act on \((M,\omega)\) by symplectomorphisms, and write \(\xi_{M}\) for the vector field generating the action of \(\xi\in\mathfrak{g}\). A momentum map is a map \(\vect{J}:M\to\mathfrak{g}^{*}\) such that, for every \(\xi\),

\begin{equation}\tag{24.28} \iota_{\xi_{M}}\omega =\dd\left\langle\vect{J},\xi\right\rangle\ec \end{equation}

that is, such that the component \(J_{\xi}\) is a Hamiltonian generating the action of \(\xi\). Rests on Definitions 14.2, 24.14 and 24.17.

Example 24.44 (Linear and angular momentum).

Take \(M=T^{*}\R^{3}\) with \(\omega=\dd q^{i}\wedge\dd p_{i}\), the phase space of one particle. For the translation group acting by \(\vect{q}\mapsto\vect{q}+s\vect{n}\) with \(\vect{n}\) a fixed unit vector, the generator is \(\xi_{M}=n^{i}\,\pp/\pp q^{i}\) and

\[ \iota_{\xi_{M}}\omega=n^{i}\,\dd p_{i} =\dd\left(\vect{n}\cdot\vect{p}\right)\ec \qquad\text{so}\qquad J_{\xi}=\vect{n}\cdot\vect{p}\ec \]

the component of linear momentum along \(\vect{n}\), in \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\). For the rotation group acting by \(\vect{q}\mapsto R\vect{q}\), \(\vect{p}\mapsto R\vect{p}\) about the axis \(\vect{n}\), the generator is \(\xi_{M}=\left(\vect{n}\times\vect{q}\right)^{i}\pp/\pp q^{i} +\left(\vect{n}\times\vect{p}\right)_{i}\pp/\pp p_{i}\) and

\[ \iota_{\xi_{M}}\omega =\left(\vect{n}\times\vect{q}\right)^{i}\dd p_{i} -\left(\vect{n}\times\vect{p}\right)_{i}\dd q^{i} =\dd\left(\vect{n}\cdot \left(\vect{q}\times\vect{p}\right)\right)\ec \]

as one checks by expanding \(\dd\left(\vect{n}\cdot(\vect{q}\times\vect{p})\right) =\left(\vect{p}\times\vect{n}\right)\cdot\dd\vect{q} +\left(\vect{n}\times\vect{q}\right)\cdot\dd\vect{p}\). So \(J_{\xi}=\vect{n}\cdot\vect{L}\), the component of angular momentum, in \(\mathrm{J}\,\mathrm{s}\). The momentum map of the full rotation group is \(\vect{J}=\vect{q}\times\vect{p}\). The name is not an analogy: the momentum map of a symmetry is the momentum conjugate to it. Rests on Definition 24.43 and Equation (24.3).

Theorem 24.45 (Noether, symplectic form).

If \(G\) acts by symplectomorphisms preserving \(H\), then the momentum map is conserved along the flow:

\begin{equation}\tag{24.29} \dv{}{t}\vect{J}=0\ep \end{equation}

For \(G\) the translation group \(\vect{J}\) is the linear momentum, for the rotation group the angular momentum, and for the one-parameter group generated by the flow itself the Hamiltonian. This is Noether's theorem of Lagrangian Mechanics, stated on phase space, where it is a single line. Rests on Definition 24.43, Proposition 24.15 and Equation (24.5).

Proof.

Derives Theorem 24.45. Fix \(\xi\in\mathfrak{g}\) and differentiate the component \(J_{\xi}\) along the motion. By Equation (24.7) the rate of change of any function is \(\dot{J}_{\xi}=X_{H}\!\left[J_{\xi}\right]=\dd J_{\xi}(X_{H})\), and by the defining property Equation (24.28) of the momentum map,

\begin{align*} \dd J_{\xi}\left(X_{H}\right) &=\left(\iota_{\xi_{M}}\omega\right)\left(X_{H}\right) =\omega\left(\xi_{M},X_{H}\right) =-\omega\left(X_{H},\xi_{M}\right)\\ &=-\left(\iota_{X_{H}}\omega\right)\left(\xi_{M}\right) =-\dd H\left(\xi_{M}\right)\ec \end{align*}

using Equation (24.5) in the last step. The final expression is \(-\xi_{M}[H]\), the rate of change of \(H\) along the group direction, and that vanishes because \(H\) is invariant under the action. So \(\dot{J}_{\xi}=0\) for every \(\xi\), which is Equation (24.29). The identifications are Example 24.44; for the one-parameter group generated by \(X_{H}\) itself the defining relation Equation (24.28) is Equation (24.5), so \(J=H\), and the conservation statement is that a time-independent Hamiltonian is conserved.

Notice how short the argument is compared with the Lagrangian one of Lagrangian Mechanics: there the conserved quantity has to be constructed from the variation of the action, here it is handed over by nondegeneracy—every symmetry generator is a Hamiltonian vector field, and its Hamiltonian is the conserved quantity [Noether:1918].

Lemma 24.46 (The level set of the momentum map is the symplectic orthogonal of the orbit).

Let \(\vect{J}\) be a momentum map for the action of \(G\) on \((M,\omega)\) and let \(x\in M\). Then

\begin{equation}\tag{24.30} \ker\dd\vect{J}_{x} =\left(T_{x}\left(G\cdot x\right)\right)^{\omega}\ec \end{equation}

where \(G\cdot x\) is the group orbit through \(x\) and, for a subspace \(W\subseteq T_{x}M\),

\[ W^{\omega} :=\set{w\in T_{x}M\ :\ \omega\left(w,z\right)=0\ \text{for all}\ z\in W} \]

is its symplectic orthogonal. In particular, if \(\mu\) is a regular value of \(\vect{J}\) then \(\dim\vect{J}^{-1}(\mu)=\dim M-\dim G\) when the action is locally free. Rests on Definition 24.43, Theorem 13.59 and Definition 24.8.

Proof.

Derives Lemma 24.46. The tangent space to the orbit at \(x\) is spanned by the values \(\xi_{M}(x)\), \(\xi\in\mathfrak{g}\), of the generators. For \(w\in T_{x}M\) and any \(\xi\), pairing the differential of \(\vect{J}\) with \(\xi\) and using Equation (24.28),

\[ \left\langle\dd\vect{J}_{x}(w),\xi\right\rangle =\dd J_{\xi}(w) =\left(\iota_{\xi_{M}}\omega\right)(w) =\omega\left(\xi_{M}(x),w\right)\ep \]

Hence \(\dd\vect{J}_{x}(w)=0\) if and only if \(\omega\left(\xi_{M}(x),w\right)=0\) for every \(\xi\), which by antisymmetry is exactly membership of \(\left(T_{x}(G\cdot x)\right)^{\omega}\). This is Equation (24.30).

For the dimension count, a locally free action has \(\dim T_{x}(G\cdot x)=\dim G\); nondegeneracy of \(\omega\) makes \(w\mapsto\omega(\cdot,w)\) an isomorphism \(T_{x}M\to T_{x}^{*}M\), so the symplectic orthogonal of a \(k\)-dimensional subspace has dimension \(2f-k\), and therefore \(\dim\ker\dd\vect{J}_{x}=\dim M-\dim G\). Since \(\mu\) is a regular value, \(\vect{J}^{-1}(\mu)\) is a submanifold with tangent space \(\ker\dd\vect{J}_{x}\) by the regular value theorem Theorem 13.59, of that dimension.

Theorem 24.47 (Marsden–Weinstein reduction).

Let \(G\) act freely and properly on \((M,\omega)\) with an equivariant momentum map \(\vect{J}\), and let \(\mu\in\mathfrak{g}^{*}\) be a regular value. Then the quotient

\begin{equation}\tag{24.31} M_{\mu}=\vect{J}^{-1}(\mu)/G_{\mu} \end{equation}

carries a unique symplectic form pulling back to the restriction of \(\omega\), and every \(G\)-invariant Hamiltonian on \(M\) descends to it. Its dimension is

\begin{equation}\tag{24.32} \dim M_{\mu}=\dim M-\dim G-\dim G_{\mu}\ep \end{equation}

Rests on Lemma 24.46, Theorem 13.59 and Definition 24.8.

Derives Theorem 24.47.

The geometric core of the argument is Lemma 24.46, proved above: the tangent space to the level set of the momentum map is the symplectic orthogonal of the group orbit, from which the degenerate directions of the restricted form turn out to be exactly the orbit directions of the isotropy group \(G_{\mu}\)—so that quotienting by that group removes the degeneracy, and removes nothing else. What makes the theorem long is everything around that core: that \(\vect{J}^{-1}(\mu)\) is a submanifold, that its quotient by \(G_{\mu}\) is again a smooth manifold with the quotient map a submersion, and that the descended form is well defined, closed and nondegenerate. Marsden–Weinstein Reduction carries it out.

One input there is quoted rather than proved, and is named as such in the appendix: the quotient-manifold theorem, that a free and proper action of a Lie group on a manifold has a smooth quotient with the projection a submersion. It is differential topology and belongs in Differentiable Manifolds, Tensors, and Curvature, which does not yet carry it. Its role is confined to producing the smooth structure on \(M_{\mu}\); the symplectic content of Theorem 24.47—the existence, uniqueness and nondegeneracy of the reduced form, and the dimension count Equation (24.32)—rests on Lemma 24.46 and on nothing quoted.

Remark 24.48 (Reduction is what physicists do without saying so).

Eliminating a cyclic coordinate by fixing its conjugate momentum and dropping the pair; passing to the centre-of-mass frame; reducing the two-body problem of Section 27.4 to the radial problem; using the effective potential Equation (27.32)—each of these is Theorem 24.47 applied to a particular group, and the effective potential is exactly the reduced Hamiltonian on \(M_{\mu}\) at the fixed value \(\mu=L\) of the angular momentum. The Lie–Poisson bracket of Definition 24.40 is itself a reduction: it is what the canonical bracket on \(T^{*}G\) becomes when the whole group is divided out, and that is why the rigid body of Proposition 24.41 lives on three variables rather than the six a cotangent bundle would demand. Rests on Theorem 24.47 and Proposition 24.41.

Geometric quantization

The programme of this section attempts to construct a quantum theory from a symplectic manifold intrinsically, using no coordinates and no prescription for ordering operators. It is placed here because it is the sharpest test of how much of quantum mechanics is already present in the geometry of phase space, and its answer is instructive in both directions: the linear structure of quantum mechanics is reproduced exactly, and the obstruction of Theorem 25.38 survives unchanged.

This section is mathematics, and is written as such. No experiment in this treatise tests geometric quantization: what is tested is the quantum mechanics of The Postulates of Quantum Mechanics, which the construction reproduces where it works and does not derive. The scope rule is respected by saying so plainly rather than by omitting the material, because the construction explains where two measured things come from—the Bohr–Sommerfeld rule Equation (23.61) and the half-integer shift in it—and those explanations are checkable.

Prequantization

Definition 24.49 (Prequantum datum).

A prequantum datum for \((M,\omega)\) is a triple \((L,h,\nabla)\) consisting of a complex line bundle \(L\to M\), a Hermitian metric \(h\) on it, and a connection \(\nabla\) compatible with \(h\), whose curvature

\begin{equation}\tag{24.33} F^{\nabla}(X,Y) =\nabla_{X}\nabla_{Y}-\nabla_{Y}\nabla_{X}-\nabla_{[X,Y]} \end{equation}

satisfies

\begin{equation}\tag{24.34} F^{\nabla}=\frac{\ii}{\hbar}\,\omega\ep \end{equation}

The right-hand side is dimensionless by Remark 24.11, as a curvature must be, and \(\hbar\) is the only constant of nature with the unit needed to make it so. Over a chart on which \(\omega=-\dd\theta\) one may take \(L\) trivial and

\begin{equation}\tag{24.35} \nabla_{X}s=X[s]-\frac{\ii}{\hbar}\,\theta(X)\,s\ec \end{equation}

whose curvature is \(-\left(\ii/\hbar\right)\dd\theta =\left(\ii/\hbar\right)\omega\) as required. Rests on Definitions 14.103, 14.105 and 24.8.

Definition 24.50 (Prequantum operator).

For \(u\in C^{\infty}(M)\) the prequantum operator acts on smooth sections of \(L\) by

\begin{equation}\tag{24.36} \widehat{u}\,s=-\ii\hbar\,\nabla_{X_{u}}s+u\,s\ec \end{equation}

with \(X_{u}\) the Hamiltonian vector field Equation (24.5). Both terms carry the unit of \(u\), by Remark 24.11. Rests on Definitions 24.14 and 24.49.

Proposition 24.51 (Prequantization is a Lie-algebra homomorphism).

For all \(u,v\in C^{\infty}(M)\),

\begin{equation}\tag{24.37} \comm{\widehat{u}}{\widehat{v}} =\ii\hbar\,\widehat{\pb{u}{v}}\ec \qquad\text{and}\qquad \widehat{1}=\identity\ep \end{equation}

That is, prequantization satisfies hypotheses (1) and (2) of Theorem 25.38 on all of \(C^{\infty}(M)\), with no restriction on the degree of the observable. Rests on Definition 24.50, Equation (24.8) and Equation (24.34).

Proof.

Derives Proposition 24.51. That \(\widehat{1}=\identity\) is immediate: \(X_{1}=0\) because \(\dd 1=0\), so Equation (24.36) leaves only multiplication by \(1\).

For the commutator, expand Equation (24.36) and treat the four terms separately. Multiplication operators commute, so \(\comm{u}{v}=0\). For the cross terms, the Leibniz rule for a connection gives \(\comm{\nabla_{X}}{v}s=\nabla_{X}(vs)-v\nabla_{X}s=X[v]\,s\), so

\[ -\ii\hbar\comm{\nabla_{X_{u}}}{v} +\ii\hbar\comm{\nabla_{X_{v}}}{u} =-\ii\hbar\,X_{u}[v]+\ii\hbar\,X_{v}[u] =2\ii\hbar\,\pb{u}{v}\ec \]

where the last step uses Equation (24.7) twice: \(X_{v}[u]=\pb{u}{v}\) and \(X_{u}[v]=\pb{v}{u}=-\pb{u}{v}\). For the remaining term, Equation (24.33) rearranged reads \(\comm{\nabla_{X_{u}}}{\nabla_{X_{v}}} =F^{\nabla}\left(X_{u},X_{v}\right) +\nabla_{\comm{X_{u}}{X_{v}}}\), so with Equation (24.34), Equation (24.7) and Equation (24.8),

\begin{align*} \left(-\ii\hbar\right)^{2} \comm{\nabla_{X_{u}}}{\nabla_{X_{v}}} &=-\hbar^{2}\left[\frac{\ii}{\hbar}\,\omega \left(X_{u},X_{v}\right) +\nabla_{-X_{\pb{u}{v}}}\right]\\ &=-\ii\hbar\,\pb{u}{v} +\hbar^{2}\nabla_{X_{\pb{u}{v}}}\ep \end{align*}

Adding the two displays,

\begin{align*} \comm{\widehat{u}}{\widehat{v}} &=\hbar^{2}\nabla_{X_{\pb{u}{v}}}+\ii\hbar\,\pb{u}{v}\\ &=\ii\hbar\left(-\ii\hbar\,\nabla_{X_{\pb{u}{v}}} +\pb{u}{v}\right) =\ii\hbar\,\widehat{\pb{u}{v}}\ec \end{align*}

which is Equation (24.37). The sign in Equation (24.34) is fixed by this calculation and is not free: with \(F^{\nabla}=-\left(\ii/\hbar\right)\omega\) the same steps would produce \(3\ii\hbar\pb{u}{v}\) in place of \(\ii\hbar\pb{u}{v}\), and the correspondence would fail on the very first pair of canonical variables.

Example 24.52 (The prequantum line and the Schrödinger representation).

Take one degree of freedom, \(M=T^{*}\R\) with \(\omega=\dd q\wedge\dd p\), \(L\) trivial, \(\theta=p\,\dd q\) and \(\nabla\) as in Equation (24.35). From Equation (24.6), \(X_{q}=-\pp/\pp p\) and \(X_{p}=\pp/\pp q\), so Equation (24.36) gives

\begin{equation}\tag{24.38} \widehat{q}\,s=\ii\hbar\,\pdv{s}{p}+q\,s\ec \qquad \widehat{p}\,s=-\ii\hbar\,\pdv{s}{q}\ec \end{equation}

the second because the connection term \(-\left(\ii/\hbar\right)p\,s\) in \(\nabla_{\pp/\pp q}s\) contributes \(-p\,s\) once multiplied by \(-\ii\hbar\), cancelling the \(+p\,s\) of Equation (24.36). One checks directly that \(\comm{\widehat{q}}{\widehat{p}}=\ii\hbar\), as Proposition 24.51 demands. But these operators act on functions of \(q\) and \(p\): the prequantum Hilbert space is far too large—twice as many variables as a wavefunction has, and, for the harmonic oscillator, a continuous spectrum where the observed one is discrete. Cutting it down is the job of the next step. Rests on Proposition 24.51 and Equation (24.6).

Polarization

Definition 24.53 (Real polarization and polarized sections).

A real polarization of \((M,\omega)\) is an integrable distribution \(P\subset TM\) of rank \(f\) on which \(\omega\) vanishes—a foliation of \(M\) by Lagrangian submanifolds. A section \(s\) of \(L\) is polarized if \(\nabla_{X}s=0\) for every \(X\) taking values in \(P\). Rests on Definitions 24.8 and 24.49.

Proposition 24.54 (The vertical polarization gives wave mechanics).

On \(M=T^{*}Q\) with \(\theta=p_{a}\dd q^{a}\), let \(P\) be spanned by the \(\pp/\pp p_{a}\). Then the polarized sections are the functions \(\psi(q)\) of the configuration variables alone, and on them

\begin{equation}\tag{24.39} \widehat{q}^{\,a}\psi=q^{a}\psi\ec\qquad \widehat{p}_{a}\psi=-\ii\hbar\,\pdv{\psi}{q^{a}}\ep \end{equation}

That is, the Schrödinger representation of The Postulates of Quantum Mechanics is what geometric quantization of a cotangent bundle in the vertical polarization produces. Rests on Definition 24.53, Definition 24.50 and Equation (24.35).

Proof.

Derives Proposition 24.54. \(P\) is integrable, being spanned by coordinate fields, and \(\omega\left(\pp/\pp p_{a},\pp/\pp p_{b}\right)=0\), so it is a real polarization; its leaves are the fibres of \(T^{*}Q\to Q\). Since \(\theta\left(\pp/\pp p_{a}\right)=0\), the connection Equation (24.35) reduces on \(P\) to the ordinary derivative, so a polarized section satisfies \(\pp s/\pp p_{a}=0\): it is a function \(\psi(q)\).

For the operators, \(X_{q^{a}}=-\pp/\pp p_{a}\) lies in \(P\), so \(\nabla_{X_{q^{a}}}\psi=0\) and Equation (24.36) leaves \(\widehat{q}^{\,a}\psi=q^{a}\psi\). For the momenta, \(X_{p_{a}}=\pp/\pp q^{a}\) and Equation (24.35) gives \(\nabla_{X_{p_{a}}}\psi=\pp\psi/\pp q^{a} -\left(\ii/\hbar\right)p_{a}\psi\), whence

\begin{align*} \widehat{p}_{a}\psi &=-\ii\hbar\left(\pdv{\psi}{q^{a}} -\frac{\ii}{\hbar}p_{a}\psi\right)+p_{a}\psi\\ &=-\ii\hbar\,\pdv{\psi}{q^{a}}-p_{a}\psi+p_{a}\psi =-\ii\hbar\,\pdv{\psi}{q^{a}}\ep \end{align*}

The cancellation is exact and is what makes the answer independent of \(p\), as it must be for an operator on \(\psi(q)\).

Remark 24.55 (What polarization costs).

Equation (24.36) defines \(\widehat{u}\) for every smooth \(u\), but \(\widehat{u}\) maps polarized sections to polarized sections only when the flow of \(X_{u}\) preserves the polarization. On \(T^{*}Q\) with the vertical polarization that restricts the directly quantizable observables to those at most quadratic in the momenta—for \(\Ham=p^{2}/2m\) one already finds \(\comm{\pp/\pp p}{\left(p/m\right)\pp/\pp q} =\left(1/m\right)\pp/\pp q\notin P\), so the free Hamiltonian moves the polarization and needs a separate treatment, in this case an easy one. Beyond the quadratic observables further devices are required, and none of them evades Theorem 25.38: there is no linear map from all the polynomials to operators satisfying Equation (24.37) irreducibly. Geometric quantization localizes the obstruction rather than removing it— prequantization satisfies the bracket condition on everything but is reducible, polarization restores irreducibility but shrinks the domain, and Theorem 25.38 is the theorem that no construction can have both. Rests on Proposition 24.54 and Theorem 25.38.

Bohr–Sommerfeld from holonomy

Proposition 24.56 (The Bohr–Sommerfeld condition is a triviality condition on the prequantum holonomy).

Let the motion be bounded and integrable, so that by Theorem 23.31 the level sets of the integrals are tori \(T\) carrying the basis cycles \(C_{a}\) of Definition 23.28, and take the polarization tangent to those tori. Then a nowhere-vanishing polarized section of \(L\) exists over \(T\) if and only if

\begin{equation}\tag{24.40} \oint_{C_{a}}p_{a}\,\dd q^{a}=n_{a}h\ec\qquad n_{a}\in\Z\ec \end{equation}

for every \(a\), which is the Bohr–Sommerfeld condition Equation (23.61). Rests on Definition 24.49, Theorem 23.31 and Definition 23.28.

Proof.

Derives Proposition 24.56. A Liouville–Arnold torus is Lagrangian: the actions are in involution, so \(\omega\) restricted to \(T\) vanishes. By Equation (24.34) the curvature of \(\nabla\) restricted to \(T\) therefore vanishes as well, so the connection is flat there and its parallel transport around a loop depends only on the homotopy class of the loop.

Compute that transport. Along a curve \(\gamma\) in \(T\) with tangent \(\dot{\gamma}\), the condition \(\nabla_{\dot{\gamma}}s=0\) reads, by Equation (24.35),

\[ \dv{s}{t}=\frac{\ii}{\hbar}\,\theta\!\left(\dot{\gamma}\right)s\ec \qquad\text{whence}\qquad s(t)=s(0)\,\ee^{\,\Phi(t)}\ec\quad \Phi(t)=\frac{\ii}{\hbar} \int_{0}^{t}\theta\!\left(\dot{\gamma}\right)\dd t'\ep \]

Around a closed cycle \(C_{a}\) the holonomy is therefore \(\exp\!\left(\left(\ii/\hbar\right)\oint_{C_{a}}\theta\right)\), and a globally defined nowhere-vanishing polarized section over \(T\) exists exactly when every such holonomy is \(1\)—parallel transport must return a section to itself. That is

\[ \frac{1}{\hbar}\oint_{C_{a}}p_{a}\,\dd q^{a}\in2\pi\Z\ec \qquad\text{that is}\qquad \oint_{C_{a}}p_{a}\,\dd q^{a}=2\pi\hbar\,n_{a}=n_{a}h\ec \]

using \(\theta=p_{a}\dd q^{a}\) from Equation (24.3) and \(h=2\pi\hbar\). This is Equation (24.40), and it is Equation (23.61) [Sommerfeld:1916].

Remark 24.57 (The half-integer, and the honest status of the construction).

Proposition 24.56 reproduces the old quantum rule with integers, and the old quantum rule with integers is wrong: the harmonic oscillator's ground state is not at zero energy. The missing piece is the metaplectic, or half-form, correction—one quantizes sections of \(L\otimes\sqrt{\Lambda}\) rather than of \(L\), where \(\sqrt{\Lambda}\) is a square root of the bundle of forms along the polarization—and its effect is to replace \(n_{a}\) by \(n_{a}+\mu_{a}/4\), with \(\mu_{a}\) the Maslov index of the cycle, counting the caustics it crosses. For one degree of freedom in a smooth well \(\mu=2\), so \(\oint p\,\dd q=\left(n+\tfrac{1}{2}\right)h\); for the oscillator of frequency \(\nu\) this integral is \(E/\nu\), giving \(E_{n}=\left(n+\tfrac{1}{2}\right)h\nu\), the observed spectrum with its zero-point energy. It is the same shift the semiclassical limit of wave mechanics produces, and the pending note at Proposition 23.38 records where that derivation belongs.

Two further statements are made here without proof, and are flagged as such. The existence of a prequantum datum at all constrains the cohomology class of \(\omega\): running the holonomy argument above over a closed two-surface \(\Sigma\subset M\) instead of a loop gives the necessary condition that \(\left(2\pi\hbar\right)^{-1}\int_{\Sigma}\omega\) be an integer, and that this condition is also sufficient is Weil's integrality theorem. Its proof needs the cohomology of a good cover of the manifold, which this treatise does not carry; and nothing above rests on it, since the Bohr–Sommerfeld statement just proved uses only the flatness of the connection on one torus. The construction as a whole is due to Kostant and to Souriau, independently, around 1970; neither is in this treatise's bibliography, so the attribution is made in words.

Finally, the scope rule. Geometric quantization is a construction, not a physical theory: it takes a classical system as input and manufactures a candidate quantum one, and Theorem 25.38 guarantees that it cannot do so for every observable. Where it succeeds it reproduces the quantum mechanics postulated in The Postulates of Quantum Mechanics, which is what experiment tests; nothing in this section is itself tested by an experiment recorded in this treatise, and the two checkable statements it touches—Equation (23.61) and the half-integer shift—are tested as spectroscopy, not as geometry. Rests on Proposition 24.56, Theorem 25.38 and Proposition 23.38.

What the symplectic view buys

The material above is geometry. This section records, as a checklist for the chapters that will use it, the places where the geometry makes a difference to what is measured. Each item is a pointer, not a derivation.