Differentiable Manifolds, Tensors, and Curvature
- Index conventions and the symmetry ladder
- Tensors under orthogonal transformations
- The theory of curves and surfaces in $\R^{3}$
- Differentiable manifolds
- Tensor analysis on manifolds
- Differential forms: $k$-forms and $k$-vectors
- The metric tensor
- Vielbein and local frames
- Flows, the Lie derivative, and Killing vectors
- Connections, parallel transport, and torsion
- Curvature
- The Cartan formalism
- Maximally symmetric spaces and the symmetry ladder
The differential geometry of curves and surfaces, and its generalisation to differentiable manifolds of arbitrary dimension, provides the language in which modern physics describes configuration spaces, spacetimes, and gauge fields. This chapter builds that language in the order in which the source manuscript develops it: tensors under orthogonal transformations in Euclidean space; the classical theory of curves and surfaces in \(\R^{3}\), where the metric, the Christoffel symbols, and the Riemann symbols first appear in concrete form; the definition of a differentiable manifold through charts and atlases; tensor analysis on manifolds and under general coordinate transformations; differential forms; the metric tensor; the covariant derivative and affine connections; and the curvature tensors.
The chapters on general relativity build directly on this material [Wald:1984] [Misner:1973].
Repeated indices are summed. In the sections on Cartesian tensors and on curve and surface theory all indices are written as subscripts and summation runs over repeated subscripts, as is customary for Cartesian tensors; from the section on differentiable manifolds onwards, coordinate indices are written as superscripts (\(x^{i}\)) and the summation convention pairs one upper with one lower index. Throughout this chapter, the type of a tensor is written \((r,s)\) (\(r\) contravariant and \(s\) covariant slots); the pair \((p,q)\) is reserved for the signature of a metric, with \(D = p + q\) the dimension. Where a derivation manipulates both sides of an equation, the operation applied is displayed after the equation being manipulated, e.g. \(\smash{A = B,\ \ \vect{x}_{l}\cdot(\;)}\) means “take the scalar product of both sides with \(\vect{x}_{l}\)”; this compact notation is inherited from the source manuscript.
Index conventions and the symmetry ladder
This chapter, and every chapter that builds on it, works in general dimension \(D = p + q\) with metric signature \((p,q)\) (project rule, Epistemology and the Scientific Method); the observed case is the instantiation \(p + q = 3 + 1\) of Part IV — Special Relativity and Part V — General Relativity and Cosmology. Two families of transformations act on everything we construct: diffeomorphisms (general coordinate transformations of the manifold) and local frame transformations (rotations of the orthonormal frames introduced in Section 13.8). Moreover, the maximally symmetric geometries of Section 13.13 are most transparently described by embedding the \((p+q)\)-dimensional space in flat spaces of one and two dimensions more. We therefore fix, once and for all, the covariant index notation for tensors and connections defined in \((p+q)\)-, \((p+q+1)\)- and \((p+q+2)\)-dimensional spaces. The following types of letters are used for indices of (pseudo-)tensors and connections under diffeomorphisms:
while for the indices of (pseudo-)tensors and connections under local Lorentz, (anti-)de Sitter and conformal transformations we use
respectively.
The groups these indices transform under form a ladder that recurs throughout the treatise; we fix its names and symbols in Table 13.1.
| Name | Symbol | Group |
|---|---|---|
| Poincaré | $\operatorname{P}_{p+q}$ | $\ISO(p,q)$ |
| Lorentz | $\operatorname{L}_{p+q}$ | $\SO(p,q)$ |
| translations | $\operatorname{T}_{p+q}$ | $\operatorname{T}(p,q)$ (abelian) |
| anti-de Sitter | $\operatorname{AdS}_{p+q}$ | $\SO(p,q+1)$ |
| de Sitter | $\operatorname{dS}_{p+q}$ | $\SO(p+1,q)$ |
| Conformal | $\operatorname{C}_{p+q}$ | $\SO(p+1,q+1)$ |
The frame metric of signature \((p,q)\) is fixed as
and likewise \(\eta_{AB}\) and \(\eta_{IJ}\) denote the flat metrics of the signatures indicated by the corresponding group in Table 13.1 — \((p,q+1)\) for \(\operatorname{AdS}_{p+q}\), \((p+1,q)\) for \(\operatorname{dS}_{p+q}\), and \((p+1,q+1)\) for \(\operatorname{C}_{p+q}\). In the relativistic chapters of the physical parts, where \(p+q = 3+1\) and the signature convention is \((-,+,+,+)\) (Part IV — Special Relativity), the dictionary is \((p,q) = (3,1)\) with the roles of the signs read from context; nothing in this chapter depends on the overall sign choice.
Tensors under orthogonal transformations
Before tensors are defined on a general manifold, it is instructive to meet them in their simplest habitat: Euclidean space referred to orthogonal coordinate systems, where all indices can be written as subscripts and the transformation matrices are constant.
Orthogonal coordinate transformations
Linear transformations
We say that a general coordinate transformation is linear if and only if it can be written as
with \(\Lambda_{ij}\) the transformation matrix, with constant coefficients.
Differentiating Equation (13.2) we see that
From Equations (13.2) and (13.3) we have
Renaming indices, we obtain the inverse transformation of the coordinates,
We denote
which is justified because
where \(\left(\Lambda^{-1}\right)_{ik}\) is the inverse transformation matrix, with constant coefficients. Therefore, from Equation (13.4),
Orthogonal coordinate systems
Consider an \(n\)-dimensional vector space generated by the basis \(\set{\hat{e}_{i}}^{n}_{i=1}\). We say that the given basis forms an orthogonal coordinate system if and only if
We denote the orthogonal coordinate system by \(K\).
Consider two orthogonal coordinate systems \(K\) and \(K'\) with respective bases \(\set{\hat{e}_{i}}^{n}_{i=1}\) and \(\set{\hat{e}'_{i}}^{n}_{i=1}\). We call orthogonal coordinate transformation the coordinate transformation between the two systems,
where \(a_{ij}\) is the transformation matrix, with constant coefficients. Rests on Definition 13.3.
Owing to Equation (13.8), the unit vectors of the new basis are related to the old ones by
and since the new coordinate system is also orthogonal,
Geometric interpretation of the transformation matrix
From Equations (13.7) and (13.9) we see that
On the other hand, let \(\theta_{ij}\) be the angles between the unit vectors \(\hat{e}_{j}\) and \(\hat{e}'_{i}\). Then
so that, by Equation (13.11),
The transformation matrix therefore contains the cosines of the angles formed between the unit vectors \(\hat{e}_{j}\) and \(\hat{e}'_{i}\).
Orthogonality condition
Consider two orthogonal coordinate systems \(K\) and \(K'\) related by an orthogonal transformation of the form Equation (13.8). Since \(K'\) is orthogonal, from Equations (13.7), (13.9) and (13.10),
Hence the orthogonality condition reads
In matrix notation this corresponds to
Taking the determinant of both sides,
Since \(\det(A)\neq 0\), the inverse \(A^{-1}\) exists, with
Comparing Equation (13.14) with Equation (13.16) we see that for orthogonal transformations
Since Equation (13.16) may equally be written \(A^{-1}A = I\), from Equation (13.17) we also have \(A\transpose A = I\); in index notation the orthogonality condition can therefore also be written
Inverse orthogonal transformation
From Equations (13.8) and (13.18),
Thus the inverse orthogonal coordinate transformation is given by
Finally, differentiating Equation (13.8) with respect to \(x_{j}\) and Equation (13.19) with respect to \(x'_{i}\), we see that
Cartesian tensors
We say that the \(n^{r}\) quantities \(T_{i_1\cdots i_r}\) are the components of a tensor of rank \(r\) if and only if under an orthogonal coordinate transformation their components satisfy
Rests on Definition 13.4.
In particular, for \(r = 2\),
To find the inverse relation, note from Equations (13.18) and (13.21) that
so that
Notable Cartesian tensors
We call vector a tensor of rank \(1\).
From the orthogonality condition Equation (13.13) we may write
Moreover, by its definition the Kronecker delta takes the same values in every orthogonal coordinate system, so we may write
Comparing with Equation (13.22), the Kronecker delta satisfies the transformation law of a rank-\(2\) tensor. Therefore the Kronecker delta is an invariant tensor under orthogonal transformations.
Algebra of Cartesian tensors
Addition
Consider the quantities \(A_{i_1\cdots i_r}\) and \(B_{i_1\cdots i_r}\), components of tensors under orthogonal transformations, both of rank \(r\). Define the quantities
Then, by Equation (13.21),
Therefore the addition of tensors forms the components of a new tensor under orthogonal transformations.
Multiplication by a scalar
Consider the quantities \(A_{i_1\cdots i_r}\), components of a rank-\(r\) tensor under orthogonal transformations, and a scalar \(\lambda\). Define
Then
Therefore the multiplication of a tensor by a scalar forms the components of a new tensor under orthogonal transformations.
Product
Consider the quantities \(A_{i_1\cdots i_r}\) and \(B_{i_1\cdots i_s}\), components of tensors under orthogonal transformations, of ranks \(r\) and \(s\) respectively. Define
Then, by Equation (13.21),
Therefore the product of tensors forms the components of a new tensor under orthogonal transformations.
Permutation of indices
Consider the quantities \(A_{i_1\cdots i_r}\), components of a rank-\(r\) tensor under orthogonal transformations. Define the symbol
in which the \(s\)-th and \(t\)-th indices of \(A_{i_1\cdots i_r}\) are permuted. Then, by Equation (13.21) and reordering the constant factors \(a_{ij}\),
Thus the permutation of indices of a tensor defines a new tensor under orthogonal transformations.
Contraction of indices
Consider the quantities \(A_{i_1\cdots i_r}\), components of a rank-\(r\) tensor under orthogonal transformations. Define the \(n^{r-2}\) quantities
Then, by Equations (13.18) and (13.21),
Thus the contraction of two indices of a tensor defines a new tensor, of rank \(r-2\), under orthogonal transformations.
Quotient law
Consider the quantities \(A_{i_1\cdots i_sj_1\cdots j_r}\), and consider further the quantities \(B_{j_1\cdots j_r}\), components of a rank-\(r\) tensor under orthogonal transformations. Let \(C_{i_1\cdots i_s}\) be given by
With respect to another orthogonal coordinate system,
Supposing that the \(C_{i_1\cdots i_s}\) are the components of a rank-\(s\) tensor under orthogonal transformations, from Equation (13.21) we have
Since the \(B_{k_1\cdots k_r}\) are arbitrary, applying \(a_{m_1k_1}\cdots a_{m_rk_r}\) to both sides of the resulting relation and using Equation (13.13),
Hence the quantities \(A_{i_1\cdots i_sj_1\cdots j_r}\) satisfying Equation (13.28) are the components of a tensor under orthogonal transformations. This result is called the quotient law.
The Levi-Civita symbol
Beside the Kronecker delta, three dimensions carry a second array with the same components in every frame. Define \(\varepsilon_{ijk}\), with \(i,j,k\in\set{1,2,3}\), to be totally antisymmetric with \(\varepsilon_{123}=+1\); equivalently \(\varepsilon_{ijk}=\sgn\sigma\) when \((i,j,k)=(\sigma(1),\sigma(2),\sigma(3))\) for a permutation \(\sigma\), and zero when two indices coincide. (The general \(n\)-index symbol reappears in Section 13.5.3 as a density of weight \(-1\) under arbitrary coordinate changes; under orthogonal changes the weight factor is \(\det a=\pm1\), which is the content of Equation (13.31) below.)
For \(n=3\), with \(\varepsilon_{ijk}\) as just defined,
and, for any \(3\times3\) array \(a_{ij}\),
In particular \(\varepsilon_{ijk}\) is an isotropic tensor under proper orthogonal transformations, and changes sign under improper ones. Rests on Definition 13.5 and Equation (5.19).
Derivation. Derives Proposition 13.6. Equation (13.29). Both sides vanish when \(j=k\) or when \(l=m\): on the left \(\varepsilon_{ijk}\) or \(\varepsilon_{ilm}\) is then zero, and on the right, putting \(j=k\), the two products coincide and cancel (similarly for \(l=m\)). So let \(j\neq k\) and \(l\neq m\). A nonzero \(\varepsilon_{ijk}\) forces \(\set{i,j,k}=\set{1,2,3}\), hence fixes \(i\); the same for \(\varepsilon_{ilm}\). The sum on the left is therefore nonzero only if the two constraints select the same \(i\), i.e. only if \(\set{l,m}=\set{j,k}\), and then exactly one term survives. If \((l,m)=(j,k)\) that term is \(\varepsilon_{ijk}^{2}=1\), while the right side is \(1-\delta_{jk}\delta_{kj}=1\); if \((l,m)=(k,j)\) it is \(\varepsilon_{ijk}\varepsilon_{ikj}=-1\), and the right side is \(\delta_{jk}\delta_{kj}-1=-1\). In every remaining case \(\set{l,m}\ne\set{j,k}\) and both sides are zero.
Equation (13.30). Contract Equation (13.29) with \(\delta_{jl}\): the right side gives \(\delta_{jj}\delta_{km}-\delta_{jm}\delta_{kj} = 3\delta_{km}-\delta_{km} = 2\delta_{km}\). Contracting once more with \(\delta_{km}\) gives \(2\delta_{kk}=6\).
Equation (13.31). The left side is totally antisymmetric in \((l,m,n)\): exchanging two of them exchanges two of the factors \(a\), which may be undone by exchanging the corresponding two summation indices of \(\varepsilon_{ijk}\), at the cost of one sign. A totally antisymmetric array of three three-valued indices is a multiple of \(\varepsilon_{lmn}\), the multiple being its value at \((l,m,n)=(1,2,3)\). That value is \(\varepsilon_{ijk}a_{i1}a_{j2}a_{k3}\), which is the Leibniz expansion Equation (5.19) of \(\det a\) read by columns: the nonzero terms are indexed by the permutations \(\sigma\) with \((i,j,k)=(\sigma(1),\sigma(2),\sigma(3))\), and \(\varepsilon_{\sigma(1)\sigma(2)\sigma(3)}=\sgn\sigma\).
The last sentence. Equation (13.31) holds for every \(3\times3\) array, so it may be applied to the transpose of \(a\), giving \(a_{li}a_{mj}a_{nk}\varepsilon_{ijk} =\varepsilon_{lmn}\det a\transpose=\varepsilon_{lmn}\det a\) by Equation (13.14). The left side is the rank-\(3\) transformation law Equation (13.21) applied to \(\varepsilon\); the right side is \(\varepsilon\) itself when \(\det a=+1\) and its negative when \(\det a=-1\), the two cases of Equation (13.15).
∎Isotropic tensors
The algebra above says how tensors combine. The question left open is which tensors are singled out by the group itself — those whose components are the same in every frame. That question is invariant theory for the rotation group rather than tensor algebra, and it is settled once and for all in Theorem 5.128. In three dimensions — and the dimension is essential, as Remark 13.7 records — a Cartesian tensor (Definition 13.5) whose components are unchanged by every proper rotation vanishes at rank one, is a multiple of \(\delta_{ij}\) at rank two, a multiple of \(\varepsilon_{ijk}\) at rank three, and a combination
of the three products of Kronecker deltas at rank four — three free numbers in place of \(3^{4}=81\). The invariance of \(\delta_{ij}\) and of \(\varepsilon_{ijk}\) under proper rotations, which is what makes the converse half true, is Equations (13.24) and (13.31) of this section.
A linear constitutive law between two rank-\(2\) Cartesian tensors, \(\sigma_{ij}=A_{ijkl}\,e_{kl}\), has \(81\) coefficients in three dimensions. If the medium is isotropic — no direction distinguished — then \(A_{ijkl}\) is an isotropic tensor, and Equation (13.32) reduces the \(81\) to three, or to two once \(e_{kl}\) is symmetric, since then only the combination \(b+c\) acts. That single step is what makes linear elasticity a two-constant theory (Lemma 30.26) and the Newtonian stress of a fluid a two-viscosity one. The classification is a statement about \(\SO(3)\) specifically and does not survive a change of dimension: it is not only the rank-\(3\) case that is three-dimensional. In two dimensions \(\varepsilon_{ij}\) is \(\SO(2)\)-invariant, so an isotropic rank-\(2\) tensor is \(a\,\delta_{ij}+b\,\varepsilon_{ij}\) rather than a multiple of \(\delta_{ij}\); in four, \(\varepsilon_{ijkl}\) is \(\SO(4)\)-invariant and is not a combination of Kronecker products, so a rank-\(4\) isotropic tensor carries a fourth independent coefficient. Every step of the proof in Linear Algebra and Representation Theory uses \(n=3\) — the parity argument needs \(\det\diag(1,-1,-1)=+1\), and the relabelling argument needs a \(3\)-cycle to be even. Rests on Theorem 5.128 and Definition 13.5.
The theory of curves and surfaces in $\R^{3}$
Theory of curves
Curves in $\R^{3}$
A curve is a differentiable map
with \(I_{\lambda}\) an open interval (so as to guarantee differentiability). The variable \(\lambda\) parametrises, or labels, the curve.
The tangent vector to the curve, with parameter \(t\), is defined as
Rests on Definition 13.8.
An admissible change of parameter is a function \(\alpha\) such that
The arc length of the curve from \(t_{0}\) to \(t\) is given by
From this we see that the inverse function of \(s\) is an admissible change of parameter for \(t\). Moreover, differentiating both sides of Equation (13.35) with respect to \(t\),
and therefore, from Equation (13.33),
We define the unit tangent vector as
From Equation (13.37) we see that \(\hat{T}\) is indeed a unit vector.
We define the normal curvature vector of the curve as
Rests on Definition 13.12.
The derivative of any unit vector \(\hat{e}\) lies in the plane orthogonal to \(\hat{e}\), for
and therefore the vector \(\vect{k}\) lies in the plane orthogonal to \(\hat{T}\).
We define the curvature function \(\kappa\) as
We define the principal unit normal vector as
Rests on Definition 13.13.
We thus find a relation between the derivative of the unit tangent vector and the unit normal,
We define the unit binormal vector as
We thereby obtain a system of three orthonormal vectors called the moving trihedron (the Frenet frame).
Torsion
Differentiating Equation (13.44) with respect to \(s\),
From Equation (13.43) the first term of this expression vanishes. From Equation (13.40) the vector \(\dv{\hat{N}}{s}\) lies in the plane spanned by \(\hat{T}\) and \(\hat{B}\), so we may write \(\dv{\hat{N}}{s}\) as a linear combination of \(\hat{T}\) and \(\hat{B}\). Thus
where the function \(\tau\) is called the torsion of the curve.
From Equation (13.44) we see that
and by Equations (13.43) and (13.46),
The Frenet–Serret equations
Equations Equations (13.43), (13.46) and (13.47) are the Frenet–Serret equations, which relate the vectors of the moving trihedron to their derivatives:
or, in matrix form,
The fundamental theorem of the theory of curves
The Frenet–Serret system determines the moving trihedron from the two functions \(\kappa\) and \(\tau\), and the trihedron determines the curve. The precise statement is the following.
Let \(I=\left]a,b\right[\subset\R\) and let \(\kappa,\tau: I\longrightarrow\R\) be continuous with \(\kappa(s)>0\) for every \(s\in I\). Then there exists a curve \(\vect{x}: I\longrightarrow\R^{3}\), parametrised by arc length, whose curvature function is \(\kappa\) and whose torsion is \(\tau\). It is unique up to a rigid motion of \(\R^{3}\): any two such curves are related by \(\vect{x}\longmapsto A\vect{x}+\vect{c}\) with \(A\in\SO(3)\) and \(\vect{c}\in\R^{3}\) constant. Rests on Equations (13.43), (13.46) and (13.47).
Derives Theorem 13.16. Existence. Read the Frenet–Serret equations as a linear system of ordinary differential equations for the nine components of the triple \((\hat{T},\hat{N},\hat{B})\),
The coefficients are continuous on \(I\), so by the existence and uniqueness theorem for linear systems (Ordinary Differential Equations and Sturm–Liouville Theory) the system has exactly one solution on the whole of \(I\) for each choice of initial data at some \(s_{0}\in I\). Choose as initial data any positively oriented orthonormal triple \((\hat{T}_{0},\hat{N}_{0},\hat{B}_{0})\).
The solution remains orthonormal. Write \(e_{1}=\hat{T}\), \(e_{2}=\hat{N}\), \(e_{3}=\hat{B}\) and consider the Gram matrix \(G_{AB}=e_{A}\cdot e_{B}\). From Equation (13.49),
a linear system for \(G\). Because \(\Omega\) is antisymmetric, the constant matrix \(G_{AB}=\delta_{AB}\) solves Equation (13.50); it agrees with the initial data at \(s_{0}\), and by uniqueness of solutions it is the solution. Hence the triple is orthonormal for every \(s\in I\), and \(\det(e_{1},e_{2},e_{3})\), a continuous function taking values in \(\set{+1,-1}\) on the interval \(I\), stays equal to its initial value \(+1\): the triple stays positively oriented.
Now define
Then \(\dd\vect{x}/\dd s=\hat{T}\) is a unit vector, so \(s\) is the arc-length parameter of \(\vect{x}\) by Equation (13.36); the first row of Equation (13.49) gives \(\dd\hat{T}/\dd s=\kappa\hat{N}\) with \(\hat{N}\) a unit vector and \(\kappa>0\), so the curvature function of \(\vect{x}\) is \(\kappa\) and its principal unit normal is \(\hat{N}\) (Equations (13.39) and (13.42)); orthonormality and orientation give \(\hat{B}=\hat{T}\times\hat{N}\), so \(\hat{B}\) is the binormal Equation (13.44); and the third row gives \(\dd\hat{B}/\dd s=-\tau\hat{N}\), so the torsion of \(\vect{x}\) is \(\tau\) by Equation (13.46).
Uniqueness. A rigid motion \(\vect{x}\longmapsto A\vect{x}+\vect{c}\) with \(A\in\SO(3)\) preserves arc length, scalar products and cross products, hence carries the trihedron of a curve to the trihedron of the image curve and leaves \(\kappa\) and \(\tau\) unchanged. Conversely, let \(\vect{x}\) and \(\tilde{\vect{x}}\) both realize the given \(\kappa\) and \(\tau\). Their trihedra at \(s_{0}\) are two positively oriented orthonormal bases of \(\R^{3}\), so there is a unique \(A\in\SO(3)\) carrying the second to the first; let \(\vect{c}=\vect{x}(s_{0})-A\tilde{\vect{x}}(s_{0})\). The curve \(A\tilde{\vect{x}}+\vect{c}\) then has the same \(\kappa\) and \(\tau\) as \(\vect{x}\) and the same trihedron at \(s_{0}\); since both trihedra solve Equation (13.49) with the same coefficients and the same data at \(s_{0}\), they coincide on all of \(I\). In particular the two unit tangents coincide, and the two curves, having the same value at \(s_{0}\), coincide by Equation (13.51).
∎The hypothesis \(\kappa>0\) cannot be dropped: where \(\kappa\) vanishes the principal normal Equation (13.42) is undefined, the trihedron is not determined by the curve, and a straight segment may be joined smoothly to curved pieces in inequivalent ways. The theorem is the geometric statement that \(\kappa\) and \(\tau\) are a complete set of invariants of a curve in \(\R^{3}\) under the group of rigid motions — the pattern that recurs throughout this chapter, where a geometric object is classified by the scalars that survive a group of transformations.
Theory of surfaces
Surfaces in $\R^{3}$
A surface is a differentiable map
with \(I_{u_1}\), \(I_{u_2}\), and \(A\) open sets.
We call curve of parameter \(u_1\) and curve of parameter \(u_2\) the curves
respectively, with \(a\) and \(b\) constants. Rests on Definitions 13.8 and 13.18.
From the definition we see that these curves lie on the surface.
The tangent vectors to the curves of parameter \(u_1\) and \(u_2\) are given by
respectively. These vectors lie in the tangent plane of the surface.
We define the unit normal vector of the surface as
Rests on Definition 13.19.
First fundamental form
Consider the differential
This differential is written as a linear combination of the vectors \(\vect{x}_{1}\) and \(\vect{x}_{2}\), and therefore also lies in the tangent plane of the surface. Consider a curve on the surface, given by
where \(s\) is the arc-length parameter of the curve. From Equation (13.36) we have
where \(\dd s\) is called the arc element. Thus
We call first fundamental form the square of the arc element of the curve given by Equation (13.56), that is, the quantity \(\dd s^{2}\). Rests on Definitions 13.11 and 13.18.
From Equations (13.55) and (13.58) we see that
We define the first fundamental coefficients as
The first fundamental form is therefore given by
From the definition we see that the \(g_{ij}\) are the components of a symmetric tensor of rank \(2\), which is called the metric tensor, or simply the metric. The determinant of the metric tensor is given by
We define the inverse metric tensor \(g^{ij}\) as the quantities that satisfy
Rests on Definition 13.22.
We can compute these quantities using matrix notation: writing
with
we see from Equation (13.64) that \([g^{ij}] = [g_{li}]^{-1}\), and therefore, from the algebra of matrices,
Length of a curve and area of a surface
Consider the curve on a surface \(\vect{x}\) given by Equation (13.56). Since the surface carries a first fundamental form, we can compute the length of the whole curve in terms of the first fundamental coefficients. As \(s = \int\dd s\), the length of the whole curve is given by
Now define an angle \(\gamma\) between the vectors \(\vect{x}_{1}\) and \(\vect{x}_{2}\). Since
from Equation (13.60) we have
On the other hand, the area of an infinitesimal parallelogram of the surface is given by
and from Equation (13.69),
Thus we see that
Integrating, the area of a surface is
Second fundamental form
The differential of any unit vector \(\hat{e}\) lies in the plane orthogonal to \(\hat{e}\), for
so the vector \(\dd\hat{n}\) lies in the plane orthogonal to \(\hat{n}\). By the chain rule, the differential \(\dd\hat{n}\) is given by
We define the second fundamental form as the quantity
Rests on Definition 13.20.
From Equations (13.55) and (13.74) we see that
We define the second fundamental coefficients as
Thus the second fundamental form is given by
From their definition, the second fundamental coefficients are the components of a symmetric tensor of rank \(2\). The determinant of the matrix formed by the second fundamental coefficients is given by
Alternative expression.
Note that the unit normal \(\hat{n}\) is orthogonal to each of the partial tangent vectors, \(\hat{n}\cdot\vect{x}_{i} = 0\). Hence
and therefore, from Equation (13.78), the second fundamental form is given by
Geometric interpretation.
Recall from linear algebra that the projection \(\vect{P}\) of a vector \(\vect{A}\) onto another vector \(\vect{B}\) is given by
Consider now the vector that, with respect to a tangent plane of the surface, indicates the small advance of the surface in a neighbourhood of the point of tangency. This vector is \(\vect{x}(u_1+\dd u_1,u_2+\dd u_2)-\vect{x}(u_1,u_2)\); thus the distance \(d\) from the surface to its tangent plane, in a neighbourhood of the point of tangency, equals the projection of this vector onto the unit normal \(\hat{n}\). From Equation (13.82),
Making a multivariable Taylor expansion of \(\vect{x}(u_1+\dd u_1,u_2+\dd u_2)\) about \((u_1,u_2)\),
Therefore the second fundamental form is, approximately, twice the distance from the surface to its tangent plane in a neighbourhood of the point of tangency.
The Gauss equations and the Christoffel symbols
Since at each point of the surface we have the three linearly independent vectors \(\vect{x}_{1}\), \(\vect{x}_{2}\), and \(\hat{n}\), we can write any vector of \(\R^{3}\) as a linear combination of them. In particular, we can express the partial derivative of \(\vect{x}_{i}\) with respect to \(u_{j}\) as
where \(\Gamma_{ij}^{\ k}\) and \(\alpha_{ij}\) are the coefficients of the linear combination. Taking the scalar product with \(\hat{n}\) and using Equation (13.80),
Thus we obtain the Gauss equations,
From Equation (13.83), taking the scalar product with \(\vect{x}_{l}\) and using Equations (13.54) and (13.60),
We define the Christoffel symbol of the first kind as the quantities
and the Christoffel symbol of the second kind as the quantities
Rests on Equation (13.83) and Definition 13.23.
Note from their definition that the Christoffel symbols are symmetric in their first two indices: considering that \(\vect{x}\) is a map of class \(C^{2}\), by Schwarz's theorem
Alternative expression for the Christoffel symbols.
From the definition of the first fundamental coefficients, Equation (13.60),
In the same way,
Then, since the Christoffel symbols are symmetric in their first two indices,
and therefore
From Equation (13.86), multiplying Equation (13.87) by \(g^{kl}\) and renaming indices,
Note from Equation (13.88) that setting \(j = k\) gives
where the last two terms cancel after renaming the summed indices; therefore
The Weingarten equations
By the same justification as for the Gauss equations, we can write the partial derivative of \(\hat{n}\) with respect to \(u_{i}\) as a linear combination of the vectors \(\vect{x}_{1}\), \(\vect{x}_{2}\), and \(\hat{n}\). Moreover, in analogy with Equations (13.40) and (13.73), the partial derivative of any unit vector \(\hat{e}\) with respect to \(u_{i}\) lies in the plane orthogonal to \(\hat{e}\):
Then \(\hat{n}_{i}\) lies in the plane orthogonal to \(\hat{n}\), so we may write
where the quantities \(\beta_{i}^{\ j}\) are the factors of the linear combination. To compute them, take the scalar product with \(\vect{x}_{k}\) and use Equations (13.60) and (13.77):
Finally, we obtain the Weingarten equations,
The Gauss–Codazzi equations
Starting from the Gauss equations,
Renaming indices, we have
By the compatibility conditions, \(\vect{x}_{ijk}-\vect{x}_{ikj}=\vect{0}\); subtracting the previous equations,
Since \(\vect{x}_{l}\) and \(\hat{n}\) are linearly independent, their factors vanish separately. From this we see that
which are the Gauss–Codazzi equations. These equations relate the metric to the second fundamental coefficients.
The Riemann symbols
We define the Riemann symbol of the first kind as the quantities
From Equation (13.92) we see that
We define the Riemann symbol of the second kind as the quantities
Rests on Definition 13.25, Definition 13.26 and Equation (13.92).
Curvatures of a surface
Normal curvature
Let \(\vect{x}=\vect{x}(u_1(t),u_2(t))\) be a curve whose points lie on the surface \(\vect{x}=\vect{x}(u_1,u_2)\), and consider the normal curvature vector \(\vect{k}\) of that curve (see Equation (13.39)).
We define the normal curvature vector of the surface as the projection (see Equation (13.82)) of the curve's normal curvature vector \(\vect{k}\) onto the unit normal \(\hat{n}\), that is,
We call normal curvature the quantity
Let us derive an alternative expression for the normal curvature. We have, using Equations (13.37) and (13.39),
Since the tangent vector \(\hat{T}\) of the curve is orthogonal to the surface normal \(\hat{n}\),
Therefore Equation (13.99) becomes, using Equations (13.37) and (13.38),
Finally, we arrive at
Principal curvatures and Gaussian curvature
From Equation (13.100) we see that
Introducing the auxiliary variable
we have
We now seek an extremal value of \(k_{n}\), denoted \(k_{np}\), which we call a principal curvature. That is,
This relation holds for a value of \(\lambda\) that makes \(k_{n}\) extremal. Substituting it back into Equation (13.103),
Note that Equation (13.104) can be rewritten in the form
and substituting this result into Equation (13.105) yields a second expression for the principal curvature,
Returning to the original variables, Equations (13.105) and (13.106) give the system of equations
Being a homogeneous system of two equations in the two variables \(\dd u_{1}\) and \(\dd u_{2}\), it has a non-trivial solution if and only if the determinant of the system vanishes:
We thus obtain a quadratic equation for \(k_{np}\) in terms of the first and second fundamental coefficients. Recall from the algebra of polynomials that for a quadratic equation \(ax^{2}+bx+c=0\) the product of its roots \(x_{1}\) and \(x_{2}\) is \(x_{1}x_{2}=c/a\).
We define the Gaussian curvature as the product of the two principal curvatures,
Geometric interpretation of the Gaussian curvature
The sign of \(K\) classifies the local shape of the surface at a point: where \(K>0\) the two principal curvatures share a sign and the surface curves away from its tangent plane on every side (elliptic point, as on a sphere or an ellipsoid); where \(K<0\) they have opposite signs and the surface crosses its tangent plane (a saddle, as on a hyperboloid or the inner rim of a torus); where \(K=0\) at least one principal curvature vanishes (a cylinder, a cone, or a plane). Figure 13.1 illustrates the flat and cylindrical cases, both with \(K=0\).
Theorema Egregium
The definition \(K=b/g\) (Equation (13.107)) is built from both fundamental forms: \(g_{ij}\), intrinsic to the surface, and \(b_{ij}\), which records how the surface bends inside the ambient space. Gauss's celebrated theorem is that the second form is, in the end, dispensable.
The Gaussian curvature \(K\) of a surface is determined by its first fundamental form \(g_{ij}\) and derivatives of \(g_{ij}\) alone; it does not depend on the second fundamental form \(b_{ij}\), i.e. not on how the surface is embedded in \(\R^{3}\). Rests on Definition 13.29, Definition 13.27, Equation (13.95) and Equation (13.88).
Derives Theorem 13.30. Specializing the Riemann symbol of the second kind (Definition 13.27) to \(D=2\) at \((l,i,j,k)=(1,2,1,2)\), and using \(b_{21}=b_{12}\),
so that \(K=b/g=R_{1212}/g\). Now trace the definition of \(R_{1212}\) backwards through what has already been established: by Equation (13.95), \(R^{l}_{\ ijk}\) — and hence, lowering with \(g_{lm}\), \(R_{lijk}\) — is built entirely from the Christoffel symbols \(\Gamma_{ij}^{\ k}\) and their first derivatives; and by Equation (13.88), \(\Gamma_{ij}^{\ k}\) is itself built entirely from \(g_{ij}\) and its first derivatives. No occurrence of \(b_{ij}\) survives this substitution: \(R_{1212}\), and therefore
is a function of \(g_{ij}\) and its derivatives up to second order alone. This proves the theorem; \(K\) is an intrinsic invariant of the surface.
∎If \(f:S\longrightarrow S'\) is a local isometry between two surfaces (Definition 13.120 specialized to \(D=2\)), then \(K_{S'}(f(P))=K_{S}(P)\) for every \(P\in S\). Rests on Theorem 13.30 and Definition 13.120.
Derives Corollary 13.31. An isometry pulls back the metric to itself, \(f^{\ast}g_{S'}=g_{S}\) (Equation (13.262)). Since \(K\) is, by Theorem 13.30, a fixed differential expression in the metric components alone, \(K_{S}(P)=K_{S'}(f(P))\) follows by evaluating that same expression on both sides of the pullback.
∎A flat sheet of paper bent, without stretching or tearing, into a cylinder realizes exactly this corollary: the bending is a local isometry (distances measured along the surface do not change), and indeed both surfaces have \(K=0\) (Figure 13.1). The same sheet cannot be bent into any piece of a sphere of radius \(R\), whose curvature \(K=1/R^{2}\neq0\) (Example 13.33 below) can never match the flat sheet's \(K=0\) — this is the everyday fact that maps of the round Earth cannot be flat without distortion, and the reason orange peel tears when flattened.
Isometric bending preserves \(K\) (flat sheet to cylinder, both \(K=0\)); no bending of a flat sheet reaches the sphere, whose \(K=1/R^{2}>0\) obstructs it by Theorem 13.30.
The proof above reuses the nineteenth-century route already built into Equations (13.88) and (13.95). The Cartan formalism of Section 13.12, developed later in this chapter, makes the same content more transparent and supplies, as a byproduct, an explicit formula for \(K\).
Fix any local orthogonal coordinates \((u_{1},u_{2})\) on the surface, so that \(g_{11}=E(u_{1},u_{2})\), \(g_{22}=G(u_{1},u_{2})\), \(g_{12}=0\) (such coordinates always exist locally — e.g. the curvature lines of Equation (13.104) away from umbilic points). An orthonormal coframe is then \(e^{1}=\sqrt{E}\,\dd u_{1}\), \(e^{2}=\sqrt{G}\,\dd u_{2}\), and the torsion-free structure equation Equation (13.311) in \(D=2\) reads \(\dd e^{1}=-\omega^{1}{}_{2}\wedge e^{2}\), \(\dd e^{2}=\omega^{1}{}_{2}\wedge e^{1}\), a linear system for the single connection one-form \(\omega^{1}{}_{2}\) that involves only \(\dd e^{1},\dd e^{2}\), hence only \(E,G\). Solving,
which is verified by substitution: with \(\dd e^{1}=-\dfrac{\pp_{2}E}{2\sqrt{E}}\,\dd u_{1}\wedge\dd u_{2}\) and \(\dd e^{2}=\dfrac{\pp_{1}G}{2\sqrt{G}}\,\dd u_{1}\wedge\dd u_{2}\), one checks \(-\omega^{1}{}_{2}\wedge e^{2}=\dd e^{1}\) and \(\omega^{1}{}_{2}\wedge e^{1}=\dd e^{2}\) directly. In \(D=2\) the Lorentz algebra \(\mathfrak{so}(2)\) is abelian, so the quadratic term of Equation (13.312) vanishes identically and \(R^{1}{}_{2}=\dd\omega^{1}{}_{2}\); defining \(K\) through \(R^{1}{}_{2}=K\,e^{1}\wedge e^{2}\) (consistent with Equation (13.107), checked below) and differentiating Equation (13.110) gives the classical Liouville formula,
manifestly a functional of \(E,G\) and their derivatives alone [DoCarmo:1976]. Not one component of \(\omega^{1}{}_{2}\), hence not one term of Equation (13.111), involves the embedding: the Cartan route makes the Theorema Egregium visible by construction rather than by cancellation, since \(b_{ij}\) never entered the computation.
In spherical coordinates \((u_{1},u_{2})=(\theta,\phi)\), \(E=R^{2}\), \(G=R^{2}\sin^{2}\theta\), so \(\pp_{2}E=0\) and \(\pp_{1}G=2R^{2}\sin\theta\cos\theta\). Then \(\pp_{1}G/\sqrt{EG}=2\cos\theta\), whose \(\theta\)-derivative is \(-2\sin\theta\), and Equation (13.111) gives
confirming \(K=1/R^{2}\) used above and matching the direct computation \(K=k_{np,1}k_{np,2}=(1/R)(1/R)\) from Definition 13.29, since \(b_{ij}=\pm g_{ij}/R\) gives \(b=\det b_{ij}=g/R^{2}\) regardless of the sign convention for the normal.
Definition 13.153 already recorded, for a general torsion-free connection, that \(R=2K\) in \(D=2\) (with \(R\) the Ricci scalar). The identity \(\Omega=\epsilon_{12}R^{12}+\epsilon_{21}R^{21} =2R^{1}{}_{2}=2K\,e^{1}\wedge e^{2}=2K\sqrt{g}\,\dd u_{1}\wedge\dd u_{2} =R\sqrt{g}\,\dd u_{1}\wedge\dd u_{2}\) (using \(R=2K\) in the last step), i.e. \(\epsilon_{ab}R^{ab}=R\sqrt{g}\,\dd^{2}u\) with no extra numerical factor, is exactly the \(D=2\) special case (\((D-2)!=0!=1\)) of the Einstein–Hilbert equivalence proven in Appendix A.8 — an independent check of the general-\(D\) combinatorial factor derived there.
Geodesic curvature
We define the projection curve as the curve projected onto the tangent plane at some point of a surface.
Under the same conditions used to define the normal curvature (see Equation (13.97)), we define the geodesic curvature vector as the normal curvature vector of the projection curve. Rests on Definitions 13.13, 13.28 and 13.35.
We may take the curve to pass through the origin of an orthogonal system; the curve projected onto the tangent plane, called \(\vect{x}_{g}\), is then given by
where \(\hat{T}'\) and \(\hat{N}'\) span the tangent plane at the point of tangency. Differentiating with respect to the arc-length parameter of the curve on the surface,
On the other hand, from Equation (13.39), the normal curvature vector of the projected curve is given by
where \(\hat{T}_{g}\) is the unit tangent vector defined in Equation (13.38) and \(s_{g}\) is the arc-length parameter defined in Equation (13.35), both for the projection curve. By the chain rule we may write
From Equation (13.37) we see that
so that
Furthermore, from the definition of the unit tangent vector (Equation (13.38)),
Again by Equation (13.37), substituting in Equation (13.116) and differentiating,
Finally, since in this case we are considering that at the point of tangency \(\hat{T}=\hat{T}'\), from Equations (13.113) and (13.114),
and from Equations (13.40) and (13.42),
Hence Equation (13.117) reduces to
Considering Equation (13.97), we see that
Alternative expression for the geodesic curvature vector.
Consider a curve on a surface \(\vect{x}\) parametrised by arc length (see Equation (13.56)). From the definition of the normal curvature vector of the curve (Equation (13.39)),
Comparing with Equation (13.119), we see that
Geodesic lines
Consider two points on a surface, and a curve on the surface containing them. Among all curves satisfying this, there must be one of least arc length, as given by Equation (13.67). Intuitively, such a curve possesses a projection curve with vanishing geodesic curvature.
We define the geodesic lines as the curves whose projection curve has vanishing geodesic curvature, that is, the curves satisfying
with the initial conditions
Rests on Definition 13.36, Equation (13.120) and Definition 13.26.
By the calculus of variations it can be proved that Equation (13.121) results from extremising the arc length given by Equation (13.67), as the following proposition shows.
Let \(u_{i}(s)\) be a regular curve on a surface, parametrised by arc length. It is a stationary point of the arc-length functional Equation (13.67) among curves with the same endpoints if and only if it satisfies the geodesic equation Equation (13.121). Rests on Equations (13.67), (13.87) and (13.121).
Derives Proposition 13.38. Write the functional Equation (13.67) for a curve \(u_{i}(t)\) joining two fixed points as
For a regular curve \(F>0\), so \(F\) is differentiable in its arguments and the Euler–Lagrange equations of Calculus of Variations,
are the condition for stationarity under variations vanishing at the endpoints. The functional is invariant under reparametrisation (Definition 13.10), so its stationary points come in reparametrisation classes and we may work with the representative parametrised by arc length, for which \(F\equiv 1\) and \(t=s\). With
setting \(F=1\) turns Equation (13.124) into
The product \(\dot{u}_{i}\dot{u}_{j}\) is symmetric in \(i\) and \(j\), so only the symmetric part of \(\pp_{i}g_{kj}\) in those indices contributes and we may replace it by \(\tfrac{1}{2}\left(\pp_{i}g_{kj}+\pp_{j}g_{ki}\right)\), giving
where the symmetry \(g_{kj}=g_{jk}\) was used. The bracket is exactly the Christoffel symbol of the first kind Equation (13.87), \(\Gamma_{ijk}\), so Equation (13.125) reads \(g_{kj}\ddot{u}_{j}+\Gamma_{ijk}\dot{u}_{i}\dot{u}_{j}=0\). Contracting with the inverse metric \(g^{kl}\) and using \(\Gamma_{ij}^{\ l}=\Gamma_{ijk}g^{kl}\) (Definition 13.26) gives
which is Equation (13.121). Every step is an equivalence, so the converse holds as well.
∎Stationarity is necessary but not sufficient for minimality. On the sphere the geodesics are the great circles, and the longer of the two arcs joining two non-antipodal points satisfies Equation (13.121) while plainly failing to minimise the length. What is true is local: a sufficiently short piece of a geodesic is the shortest curve between its endpoints. On a pseudo-Riemannian manifold Equation (13.121) becomes the autoparallel equation Equation (13.297) of the Levi-Civita connection, but the minimising statement does not carry over verbatim. The line element changes sign there, so the functional has to be written \(\int\sqrt{\abs{\dd s^{2}}}\) (Remark 13.151), and the conclusion then depends on the causal character of the curve: a short enough spacelike geodesic still minimises the arc length, whereas a short enough timelike geodesic maximises the proper time between its endpoints — the twin-paradox inequality of Part IV — Special Relativity — and on a null curve the functional vanishes identically, so every null curve is stationary for trivial reasons and the null geodesics are singled out by Equation (13.297) rather than by any extremal property of the length. See Remark 13.151, where the distinction between the shortest and the straightest curve is what torsion measures.
Geodesic parallelism
Consider a curve \(\vect{x}=\vect{x}(u_1(t),u_2(t))\) on the surface \(\vect{x}=\vect{x}(u_1,u_2)\). Assign to the curve a vector field \(\vect{A}\) of the form
At each point of the curve, \(\vect{A}\) is written as a linear combination of the partial tangent vectors of the surface, and therefore lies in the tangent plane. From vector analysis, the vectors \(\vect{A}\) are parallel for every value of \(t\) if and only if
Analogously, we say that the vectors \(\vect{A}\) are geodesically parallel for every value of \(t\) if and only if
From Equation (13.126),
Therefore the field \(\vect{A}\) is geodesically parallel if and only if
whence the condition of geodesic parallelism is
Holonomy and the local Gauss–Bonnet theorem
Transport a vector geodesically parallel all the way around a closed curve on a surface and it comes back to the point it started from — but not, in general, to the vector it started as. The failure is the holonomy of the curve, and the theorem of this subsection identifies it: it is the total Gaussian curvature enclosed. That is the local form of the Gauss–Bonnet theorem, and it is what makes curvature an observable rather than a bookkeeping device, since the angle it predicts can be measured by carrying a gyroscope, or a swinging pendulum, around a circuit.
If \(\vect{A}\) and \(\vect{B}\) are geodesically parallel along a curve on a surface, then \(\vect{A}\cdot\vect{B}\) is constant along it. In particular lengths and angles are preserved, so transport around a closed curve acts on the tangent plane as a rotation. Rests on Equation (13.129), Equation (13.126) and Definition 13.21.
Derivation. Derives Proposition 13.40. Both fields are tangent to the surface, and geodesic parallelism says that \(\dd\vect{A}/\dd t\) has vanishing tangential part, i.e.\ \(\vect{x}_{l}\cdot\dd\vect{A}/\dd t = 0\) for every \(l\); since \(\vect{B}=v_{i}\vect{x}_{i}\) is a combination of the \(\vect{x}_{l}\), \(\vect{B}\cdot\dd\vect{A}/\dd t = 0\), and likewise with the roles exchanged. Hence \(\dv{}{t}(\vect{A}\cdot\vect{B}) = \vect{B}\cdot\dv{\vect{A}}{t} + \vect{A}\cdot\dv{\vect{B}}{t} = 0\). A linear map of a two-dimensional Euclidean plane preserving the inner product is a rotation or a reflection, and transport is a rotation because it is reached continuously from the identity at \(t=0\).
∎Let \((u_{1},u_{2})\) be orthogonal coordinates on a surface, with \(g_{11}=E>0\), \(g_{22}=G>0\), \(g_{12}=0\), and let \(\gamma\) be a piecewise smooth simple closed curve, positively oriented in the \((u_{1},u_{2})\) plane, bounding a region \(D\) whose closure lies in the coordinate patch. Write \(\vartheta\) for the angle a geodesically parallel field makes with the orthonormal frame
Then the holonomy of \(\gamma\) — the total change of \(\vartheta\) on one circuit — is
with \(K\) the Gaussian curvature Equation (13.107). Reversing the orientation of \(\gamma\) reverses the sign. Rests on Proposition 13.40, Equation (13.111) and Theorem 7.99.
Derives Theorem 13.41. The angle obeys a first-order equation. By Proposition 13.40 a parallel field has constant length, so we may write \(\vect{A} = \abs{\vect{A}}\left(\cos\vartheta\,\hat{e}_{1} + \sin\vartheta\,\hat{e}_{2}\right)\). Put \(\hat{e}_{\perp} = -\sin\vartheta\,\hat{e}_{1} + \cos\vartheta\,\hat{e}_{2}\), a unit tangent vector orthogonal to \(\vect{A}\). Since \(\hat{e}_{1}\cdot\hat{e}_{1}=1\) the tangential part of \(\dd\hat{e}_{1}/\dd t\) is a multiple of \(\hat{e}_{2}\); call the coefficient \(\varpi=\hat{e}_{2}\cdot\dd\hat{e}_{1}/\dd t\), and then \(\hat{e}_{1}\cdot\dd\hat{e}_{2}/\dd t=-\varpi\) because \(\hat{e}_{1}\cdot\hat{e}_{2}=0\). Taking the scalar product of \(\dd\vect{A}/\dd t\) with \(\hat{e}_{\perp}\) and dividing by \(\abs{\vect{A}}\),
the last step using \(\hat{e}_{\perp}\cdot\dd\hat{e}_{1}/\dd t=\varpi\cos\vartheta\) and \(\hat{e}_{\perp}\cdot\dd\hat{e}_{2}/\dd t=\varpi\sin\vartheta\) together with \(\cos^{2}+\sin^{2}=1\). The left-hand side vanishes for a parallel field, because \(\hat{e}_{\perp}\) is tangent, so
\(\varpi\) is a one-form built from \(E\) and \(G\) alone. Using \(\vect{x}_{1}\cdot\vect{x}_{1}=E\), \(\vect{x}_{2}\cdot\vect{x}_{2}=G\), \(\vect{x}_{1}\cdot\vect{x}_{2}=0\),
so that, since the term coming from differentiating \(E^{-1/2}\) is proportional to \(\vect{x}_{1}\) and dies against \(\hat{e}_{2}\),
This is minus the connection one-form \(\omega^{1}{}_{2}\) of Equation (13.110), computed here again directly so that nothing is owed to the later Cartan formalism.
Green's theorem. Integrating Equation (13.132) once around \(\gamma\) and applying Theorem 7.99 in the \((u_{1},u_{2})\) plane,
the last step being the Liouville formula Equation (13.111). Substituting gives Equation (13.131). Reversing the orientation of \(\gamma\) reverses the sign of the line integral and hence of \(\Delta\vartheta\).
∎On the sphere of radius \(R\) in the coordinates \((u_{1},u_{2})=(\theta,\phi)\) of Example 13.33, \(E=R^{2}\) and \(G=R^{2}\sin^{2}\theta\), so \(\pp_{2}E=0\), \(\pp_{1}G=2R^{2}\sin\theta\cos\theta\) and \(\sqrt{EG}=R^{2}\sin\theta\). Equation (13.133) then gives the strikingly simple
so a parallel field carried once round the latitude circle \(\theta=\theta_{0}\), in the direction of increasing \(\phi\), turns relative to the frame Equation (13.130) by
The theorem itself is checked on a curve that does satisfy its hypotheses — the boundary of the coordinate rectangle \(\theta\in[\theta_{1},\theta_{2}]\), \(\phi\in[\phi_{1},\phi_{2}]\), whose closure lies in the patch. Only the two edges of constant \(\theta\) contribute to \(\oint\cos\theta\,\dd\phi\), and with the positive orientation they give \(\left(\phi_{2}-\phi_{1}\right) \left(\cos\theta_{2}-\cos\theta_{1}\right)\), so \(\Delta\vartheta = \left(\phi_{2}-\phi_{1}\right) \left(\cos\theta_{1}-\cos\theta_{2}\right)\); and since \(K=1/R^{2}\) with \(\dd A = R^{2}\sin\theta\,\dd\theta\,\dd\phi\), the right-hand side of Equation (13.131) is \(\left(\phi_{2}-\phi_{1}\right)\int_{\theta_{1}}^{\theta_{2}} \sin\theta\,\dd\theta\), the same number.
Comparing Equation (13.135) at \(\theta_{0}\) and at a small \(\varepsilon\) gives, by the same direct computation,
the total curvature of the annulus between the two latitudes.
The cap \(\theta\le\theta_{0}\) itself is not covered by this chart: the pole \(\theta=0\) is a coordinate singularity, where the frame Equation (13.130) is undefined, and the hypothesis of the theorem fails. The bookkeeping is exact and worth stating, because it is the whole content of the Foucault deficit. Letting \(\varepsilon\to0\) above, the total curvature of the cap is
which is precisely the solid angle the cap subtends at the centre, while Equation (13.135) gives \(\Delta\vartheta = -2\pi\cos\theta_{0}\): the two differ by exactly \(2\pi\), the limit of \(-\Delta\vartheta(\varepsilon)\), which is the turn the singular frame itself makes on one circuit of the pole. So it is the deficit from a full turn, \(2\pi-2\pi\cos\theta_{0}\), and not the frame-relative rotation, that is the geometric invariant of the region enclosed. Written in terms of the latitude \(\lambda=\pi/2-\theta_{0}\), that deficit is \(2\pi\left(1-\sin\lambda\right)\). Rests on Theorem 13.41 and Example 13.33.
Differentiable manifolds
The definitions given below serve to introduce the concept of a differentiable manifold, which results from generalising the idea of curves and surfaces. The definition is built up from the concept of a topological space (Topological and Metric Spaces), which we progressively equip with the pertinent structure so as to obtain an object rich in properties that can be linked with nature. One of the key concepts is that of a coordinate system, which is in intimate correspondence with reference frames: the tool of theoretical physics that lets us fix a point in the Universe with respect to which a physical system is being evaluated.
Charts, atlases, and differentiable structure
Coordinate systems
Consider a topological space. In the context of this chapter we refer to the elements of a topological space as points.
Let \(M\subset\R^{N}\) be a topological space and \(U\subset\R^{n}\) with \(n\leq N\). Consider a point \(P\in M\) and \(x=(x^{1},\ldots,x^{n})\in U\), and consider further the injective map \(\phi: x\longmapsto\phi(x)=P\).
A coordinate system: the map \(\phi\) carries a point \(x=(x^{1},\ldots,x^{n})\) of the parameter domain \(U\subset\R^{n}\) to a point \(P=\phi(x)\) of the topological space \(M\).
Intuitively, \(\phi\) must respect the topology of \(M\) in order to describe its points correctly through coordinates. To this end we must demand that points of \(M\) be close if and only if their coordinates are. The idea is to require
where \(x_{l}\) is a sequence of points in \(U\) and \(P_{l}=\phi(x_{l})\) a sequence of points in \(M\). Demanding “\(\Leftarrow\)” we meet the continuity of \(\phi\), and demanding “\(\Rightarrow\)” we meet the continuity of \(\phi^{-1}\). We further wish a one-to-one relation between the points of \(U\) and (at least some of, since in general more than one set like \(U\) is needed to describe all the points of \(M\)) the points of \(M\), which suggests that \(\phi\) must be bijective. In conclusion, the topology is respected if and only if \(\phi\) is a homeomorphism. By the definition of continuity, demanding “\(\Leftarrow\)” the set \(\phi^{-1}(\phi(U))=U\) must be open, and demanding “\(\Rightarrow\)” the set \((\phi^{-1})^{-1}(U)=\phi(U)\) must be open. The geometric interpretation is that (by the definition of an open set) we can place an open ball centred at \(P\) and one centred at \(x\),
Open balls centred at \(x\in U\) and at \(P=\phi(x)\in M\): the neighbours of \(x\) carry information about the neighbours of \(P\) precisely when \(\phi\) is a homeomorphism.
from which we see that the points neighbouring \(x\) should give us information about the points neighbouring \(P\) (by the definition of a neighbourhood), which happens if and only if the map \(\phi\) is smooth.
We say that the pair \((\phi,U)\) is an \(n\)-dimensional coordinate system (CS), or local chart, if and only if
-
the map \(\phi\) is a homeomorphism, which implies that \(U\) is open in \(\R^{n}\);
-
the map \(\phi\) is smooth.
Coordinates
As we have seen, given a coordinate system there is a one-to-one relation between the points of \(U\) and some points of \(M\). The variables \((x^{1},\ldots,x^{n})\) corresponding to a point \(P\in M\) receive the name of coordinates of the point \(P\) with respect to the coordinate system \((\phi,U)\). To symbolise (and not forget) the exact point of \(M\) that we wish to describe by means of \((\phi,U)\), we denote the coordinates of a point \(P\in M\) as \((x^{1},\ldots,x^{n})\stackrel{\text{not.}}{=}x^{i}(P)\) with \(i=1,\ldots,n\). Thus, when we want to refer to the point \(x\in U\) we write \(x\) or \((x^{1},\ldots,x^{n})\); when we want to refer explicitly to the point \(P\) through its coordinates, we write \(x^{i}(P)\); and for the coordinates of arbitrary points of \(M\) (within the codomain of \(\phi\)) we use \(x^{i}\). In summary,
Even though we have insisted on the difference between the quantities in \(U\) and the coordinates of the point, numerically they are the same, which is useful in practice; in the theory, however, it is necessary to be rigorous with the notation. It is very important to note that the coordinates of a point depend on the subset \(U\) of \(\R^{n}\) we have chosen. Note that \(\phi: U\longrightarrow\phi(U)\) is bijective, so this one-to-one relation characterises the points of \(\phi(U)\) by their coordinates in \(\R^{n}\). Note also that a correct notation for the point \(P\) in terms of its coordinates in \(U\) would be \(P=\phi(x)=\phi(x^{1},\ldots,x^{n})\), but not \(P=\phi(x^{i})\) nor \(P=\phi(x^{1}(P),\ldots,x^{n}(P))\).
Atlases
Consider a family of \(n\)-dimensional coordinate systems \(\set{(\phi_{a},U_{a})}_{a\in I}\) that cover the whole space \(M\), that is,
such that \(U_{a}\cap U_{b}\neq\varnothing\) for \(a\neq b\).
Let the overlap regions be
Two charts \((\phi_{a},U_{a})\) and \((\phi_{b},U_{b})\) covering overlapping portions of \(M\); the overlap region is \(\Omega_{ab}=\phi_{a}(U_{a})\cap\phi_{b}(U_{b})\).
The sets \(\phi^{-1}_{a}(\Omega_{ab})\) and \(\phi^{-1}_{b}(\Omega_{ab})\) are open because \(\phi_{a}\) and \(\phi_{b}\) are homeomorphisms. We further define the transition maps as
The transition map \(\phi_{ba}=\phi^{-1}_{b}\circ\phi_{a}\) between the coordinate domains of two overlapping charts.
The transition maps are continuous, being compositions of continuous maps, and since the sets \(\phi^{-1}_{a}(\Omega_{ab})\) and \(\phi^{-1}_{b}(\Omega_{ab})\) are open, the transition maps are moreover homeomorphisms. A point \(P\) may be described by \(P=\phi_{a}(x^{1}_{a},\ldots,x^{n}_{a})\) or by \(P=\phi_{b}(x^{1}_{b},\ldots,x^{n}_{b})\). This means we can perform a change of coordinate system of the form
which respects the topology of \(M\) because the \(\phi_{ba}\) are homeomorphisms. For the information about the points neighbouring \(P\) to be well defined after changing the coordinate system, we must demand that the transition maps be smooth.
We say that a family of coordinate systems \(\set{(\phi_{a},U_{a})}_{a\in I}\) is an \(n\)-dimensional atlas if and only if
-
the coordinate systems cover the whole space \(M\), that is,
\begin{equation*} M\subset\bigcup\limits_{a\in I}\phi_{a}(U_{a})\ec \end{equation*} -
the coordinate systems are smoothly related[These smoothly related changes of coordinate system are exclusively within the atlas. This concept must not be confused with the general coordinate transformations studied further below.].
Rests on Definition 13.43.
Differentiable maps on a topological space
Consider now a topological space \(M\) equipped with an atlas \(\set{(\phi_{a},U_{a})}_{a\in I}\), and let
be a map. A key question is: under what conditions is \(f\) differentiable? Since we know the usual definition of differentiability at a point \(x\in\R^{n}\), we wish to extend it to the points of a topological space by virtue of the given atlas. Of course, we know that a map that is not continuous is not differentiable, so at the very least we must demand that \(f\) be continuous. Choose the coordinate system \((\phi_{a},U_{a})\) and define the map \(\bar{f}_{a}=f\circ\phi_{a}\).
A map \(f: M\to\R\) read through a chart: the composition \(\bar{f}_{a}=f\circ\phi_{a}\) is an ordinary map between subsets of \(\R^{n}\) and \(\R\), on which differentiability makes sense.
The map \(\bar{f}_{a}\) is continuous, being a composition of continuous maps; moreover, the atlas allows us to define \(f\) on a neighbourhood of \(P\), so that \(\phi_{a}\) is well defined on a neighbourhood of \((x^{1}_{a},\ldots,x^{n}_{a})\), since \(\phi_{a}\) respects the topology (again a consequence of working with an atlas). Thus our map \(\bar{f}_{a}\) is in a position to be differentiable.
We say that \(f\) is differentiable at \(P=\phi_{a}(x)\) if and only if
-
\(f\) is continuous;
-
\(\bar{f}_{a}=f\circ\phi_{a}\) is differentiable.
But what would happen if we changed to a coordinate system \((\phi_{b},U_{b})\) and \(\bar{f}_{b}=f\circ\phi_{b}\) were not differentiable at \(P\)? This does not occur, since
and thus, since \(\phi_{ab}\) is smooth, it is at least of class \(C^{1}\) and hence differentiable; and since \(\bar{f}_{a}\) is differentiable, \(\bar{f}_{b}\) is differentiable, being a composition of differentiable maps. The atlas has again entered the game: it is the smoothness condition in the atlas that allows us to define a differentiable map on \(M\).
Differentiable structure
Consider a topological space now equipped with the atlases \(\set{(\phi_{a},U_{a})}_{a\in I}\) and \(\set{(\psi_{b},V_{b})}_{b\in J}\). The relation
is an equivalence relation.
On the set of \(n\)-dimensional atlases of a topological space \(M\), the relation
is an equivalence relation. Rests on Definitions 13.43 and 13.44.
Derives Proposition 13.46. First unpack Equation (13.138). The union of two atlases always satisfies the covering condition of Definition 13.44, since \(\mathcal{A}\) alone already covers \(M\); and the transition maps between two charts of \(\mathcal{A}\), or between two charts of \(\mathcal{B}\), are smooth because each family is an atlas. So \(\mathcal{A}\sim\mathcal{B}\) holds if and only if every crossed transition map \(\psi^{-1}\circ\phi\), between a chart \((\phi,U)\in\mathcal{A}\) and a chart \((\psi,V)\in\mathcal{B}\) with overlapping images, is smooth.
Reflexivity. \(\mathcal{A}\cup\mathcal{A}=\mathcal{A}\), which is an atlas by hypothesis.
Symmetry. The union is a symmetric operation, so the condition Equation (13.138) is symmetric as it stands. (Concretely: a crossed transition map in one direction is the inverse of the one in the other, \(\psi^{-1}\circ\phi=(\phi^{-1}\circ\psi)^{-1}\), and the atlas condition demands smoothness of both.)
Transitivity. Let \(\mathcal{A}\sim\mathcal{B}\) and \(\mathcal{B}\sim\mathcal{C}\), and take charts \((\phi,U)\in\mathcal{A}\) and \((\chi,W)\in\mathcal{C}\) whose images overlap. Let \(P\in\phi(U)\cap\chi(W)\). Since \(\mathcal{B}\) is an atlas it covers \(M\), so there is a chart \((\psi,V)\in\mathcal{B}\) with \(P\in\psi(V)\). The set
is open — it is the preimage under the homeomorphism \(\phi\) of an intersection of three sets open in \(M\) — and it contains \(\phi^{-1}(P)\). On it,
a composition of two smooth maps by the hypotheses \(\mathcal{A}\sim\mathcal{B}\) and \(\mathcal{B}\sim\mathcal{C}\), hence smooth. Smoothness is a local property and \(P\) was an arbitrary point of the overlap, so \(\chi^{-1}\circ\phi\) is smooth on the whole of \(\phi^{-1}(\phi(U)\cap\chi(W))\); the same argument with the roles of \(\phi\) and \(\chi\) exchanged gives the smoothness of \(\phi^{-1}\circ\chi\). Every crossed transition map of \(\mathcal{A}\cup\mathcal{C}\) is therefore smooth, and \(\mathcal{A}\cup\mathcal{C}\) is an atlas, that is \(\mathcal{A}\sim\mathcal{C}\).
Note where each hypothesis was used: reflexivity and symmetry are formal, while transitivity rests on the covering property of the intermediate atlas \(\mathcal{B}\) — without a chart of \(\mathcal{B}\) around \(P\) the factorisation Equation (13.139) would not exist.
∎We say that an equivalence class of atlases is a differentiable structure. Two atlases belonging to the same differentiable structure are called compatible. Rests on Definition 13.44.
Note the following: all the coordinate systems of two compatible atlases are smoothly related, because their union is an atlas (and the coordinate systems of an atlas are smoothly related).
Differentiable manifold
We say that a topological space \(M\) is an \(n\)-dimensional differentiable manifold if and only if \(M\) is equipped with an (\(n\)-dimensional) differentiable structure. When it is necessary to specify the dimension of the manifold, we denote it \(M_{n}\). Rests on Definition 13.47.
It is important to discuss the following. The dimension of the manifold refers to the number of coordinates we need to describe completely an arbitrary point of it. The definitions have been constructed in such a way that it is not possible to have an \(n\)-dimensional manifold in which some points are described by fewer than \(n\) coordinates.
Diffeomorphism
The notion of a differentiable map developed above compares a manifold with \(\R\); nothing prevents us from comparing two manifolds with each other.
Let \(M_{m}\) and \(N_{n}\) be differentiable manifolds. We say that a map \(f: M_{m}\longrightarrow N_{n}\) is smooth if and only if it is continuous and, for every coordinate system \((\phi_{a},U_{a})\) of \(M_{m}\) and every coordinate system \((\psi_{b},V_{b})\) of \(N_{n}\), the coordinate representation
which maps a subset of \(\R^{m}\) into \(\R^{n}\), is smooth wherever it is defined. Rests on Definitions 13.44, 13.45 and 13.48.
That the condition need be checked in one pair of charts only is the computation already performed for \(\bar{f}_{a}\) and \(\bar{f}_{b}\) above: changing either chart composes Equation (13.140) with a transition map, which is smooth because we are working inside atlases. Definition 13.45 is the special case \(N_{n}=E_{1}\) read in the Cartesian chart.
A diffeomorphism is a smooth homeomorphism \(f: M_{m}\longrightarrow N_{n}\) whose inverse \(f^{-1}\) is also smooth. Two manifolds related by a diffeomorphism are said to be diffeomorphic. Rests on Definitions 6.7 and 13.49.
Diffeomorphic manifolds are, for every purpose of this chapter, the same manifold. Applying the chain rule to \(f^{-1}\circ f=\id\) shows that the Jacobian of Equation (13.140) is invertible at every point, so \(m=n\); and then \((f\circ\phi_{a},U_{a})\) is a coordinate system of \(N_{n}\) whenever \((\phi_{a},U_{a})\) is one of \(M_{n}\), so \(f\) carries atlases to atlases and the differentiable structure of \(M_{n}\) (Definition 13.47) to that of \(N_{n}\). Every statement expressible through charts, smooth maps and tangent spaces alone therefore transfers. Diffeomorphism is to differential geometry what homeomorphism (Definition 6.7) is to topology.
Smoothness of the inverse must be demanded, for it does not follow: the map \(f:\R\longrightarrow\R\) with \(f(x)=x^{3}\) is a smooth homeomorphism, but \(f^{-1}(y)=y^{1/3}\) is not differentiable at the origin, so \(f\) is not a diffeomorphism — precisely as continuity of a bijection fails to give a homeomorphism (Remark 6.8).
The differential of a smooth map
Let \(f: M_{m}\longrightarrow N_{n}\) be smooth and let \(P\in M_{m}\). The differential of \(f\) at \(P\) is the map
that assigns to the vector tangent at \(P\) to a curve \(\lambda\longmapsto x(\lambda)\) (Definition 13.81) the vector tangent at \(f(P)\) to the image curve \(\lambda\longmapsto f(x(\lambda))\). Writing \(y^{\alpha}=f^{\alpha}(x^{1},\ldots,x^{m})\) for the coordinate representation Equation (13.140), the chain rule gives, in the coordinate bases \(\set{\pp/\pp x^{i}}\) and \(\set{\pp/\pp y^{\alpha}}\),
so \(f_{*P}\) is linear and is represented by the Jacobian matrix of \(f\) at \(P\). It is the pushforward used below in the discussion of pullbacks (Section 13.5.1). Rests on Definitions 13.49 and 13.81.
Immersion, embedding, and submanifold
Let \(f: M_{m}\longrightarrow N_{n}\) be smooth. We say that \(f\) is
-
an immersion if and only if \(f_{*P}\) is injective for every \(P\in M_{m}\), which requires \(m\leq n\);
-
a submersion if and only if \(f_{*P}\) is surjective for every \(P\in M_{m}\), which requires \(m\geq n\);
-
an embedding if and only if it is an injective immersion which is, in addition, a homeomorphism onto its image \(f(M_{m})\), that image carrying the topology it inherits as a subset of \(N_{n}\).
A subset \(S\subseteq N_{n}\) is an embedded submanifold of dimension \(m\) if and only if \(S\), equipped with the topology inherited from \(N_{n}\), admits an \(m\)-dimensional differentiable structure for which the inclusion \(\iota: S\longrightarrow N_{n}\) is an embedding. The integer \(n-m\) is the codimension of \(S\) in \(N_{n}\). Rests on Definitions 13.48 and 13.53.
The image of an embedding is an embedded submanifold: transporting the differentiable structure of \(M_{m}\) through the homeomorphism \(f: M_{m}\longrightarrow f(M_{m})\) produces an atlas on \(f(M_{m})\) whose charts are the \(f\circ\phi_{a}\), and with respect to it \(f\) becomes a diffeomorphism onto its image. The third condition in Definition 13.53 is not implied by the first two, as the following example shows.
Consider the figure-eight curve in \(E_{3}\),
It is an immersion: \(\dv{x^{i}}{\lambda}=(2\cos 2\lambda,\cos\lambda,0)\) never vanishes, because \(\cos\lambda=0\) forces \(\lambda=\pm\pi/2\) and there \(2\cos 2\lambda=-2\). It is injective: let \(\lambda_{1},\lambda_{2}\in\left]-\pi,\pi\right[\) have \(\gamma(\lambda_{1})=\gamma(\lambda_{2})\). Equality of the second components, \(\sin\lambda_{1}=\sin\lambda_{2}\), leaves besides \(\lambda_{2}=\lambda_{1}\) exactly two possibilities inside the interval, \(\lambda_{2}=\pi-\lambda_{1}\) (which lies in \(\left]-\pi,\pi\right[\) only for \(\lambda_{1}>0\)) and \(\lambda_{2}=-\pi-\lambda_{1}\) (only for \(\lambda_{1}<0\)); both branches must be checked, since the second is the one available on the negative half of the interval. On either of them \(\sin 2\lambda_{2}=-\sin 2\lambda_{1}\), because \(\sin\) has period \(2\pi\) and \(2\lambda_{2}=\pm 2\pi-2\lambda_{1}\). Equality of the first components then gives \(\sin 2\lambda_{1}=0\), that is \(\lambda_{1}\in\set{0,\pm\pi/2}\); and \(\lambda_{1}=0\) sends both branches out of the interval, while \(\lambda_{1}=\pi/2\) and \(\lambda_{1}=-\pi/2\) send the available branch back to \(\lambda_{2}=\lambda_{1}\). Hence \(\lambda_{1}=\lambda_{2}\) in every case.
It is not an embedding. As \(\lambda\to\pm\pi\) both components tend to \(0\), so the image accumulates at \(\gamma(0)=(0,0,0)\): every neighbourhood of \(\gamma(0)\) in the subspace topology of the image contains points \(\gamma(\lambda)\) with \(\lambda\) arbitrarily close to \(\pm\pi\), whose parameters are not close to \(0\). Hence \(\gamma^{-1}\) is discontinuous at \(\gamma(0)\), and the image, which is a self-touching figure eight, is not an embedded submanifold of \(E_{3}\): no neighbourhood of the origin in it is homeomorphic to an interval. Rests on Definitions 13.52 and 13.53.
Product of manifolds
Let \(M_{m}\) and \(N_{n}\) be differentiable manifolds with atlases \(\set{(\phi_{a},U_{a})}_{a\in I}\) and \(\set{(\psi_{b},V_{b})}_{b\in J}\). Then the topological product \(M_{m}\times N_{n}\), carrying the product topology, is a differentiable manifold with the product atlas
and
Derives Proposition 13.56. Each \(U_{a}\times V_{b}\) is open in \(\R^{m}\times\R^{n}=\R^{m+n}\), and \(\phi_{a}\times\psi_{b}\) is a homeomorphism onto its image for the product topology, being a product of homeomorphisms; it is smooth because each factor is, so Definition 13.43 is met and each pair is an \((m+n)\)-dimensional coordinate system. The images cover, since \(\phi_{a}(U_{a})\) cover \(M_{m}\) and \(\psi_{b}(V_{b})\) cover \(N_{n}\). Finally, the transition map between two such charts is
acting on the two blocks of coordinates separately; it is smooth because the transition maps of the two given atlases are. Both conditions of Definition 13.44 hold, so Equation (13.143) is an atlas, and the number of coordinates needed to locate a point, \((x^{1},\ldots,x^{m},y^{1},\ldots,y^{n})\), is \(m+n\), which is Equation (13.144).
∎Take \(M=S_{1}\), the circle of radius \(R\) with the two-chart atlas of Section 13.4.3, and \(N=\R\) with the single Cartesian chart. The index sets of Equation (13.143) then have two and one element, so the product atlas has \(2\times 1=2\) charts, both of the form \((\phi,z)\longmapsto(R\cos\phi,R\sin\phi,z)\) on the corresponding angular range, and it makes \(S_{1}\times\R\) a two-dimensional manifold, the cylinder of radius \(R\) in \(E_{3}\). Its dimension, \(1+1=2\), is Equation (13.144); it is the surface whose vanishing Gaussian curvature (Section 13.3.3) makes it locally isometric to the plane. Globally it is not: a flat strip is simply connected and the cylinder is not, so no isometry of the two exists — the cylinder is a strip with its two edges identified, not a subset of the plane. Rests on Proposition 13.56.
Hypersurfaces
A hypersurface of \(N_{n}\) is an embedded submanifold of codimension one, that is, of dimension \(n-1\). Rests on Definition 13.54.
Hypersurfaces are produced in practice by one scalar equation, and the theorem that licenses this is the following.
Let \(F: N_{n}\longrightarrow\R\) be smooth and let \(c\in\R\) be a regular value of \(F\): the differential \(F_{*P}\) is nonzero at every \(P\) with \(F(P)=c\), that is, in any chart around such a \(P\) at least one of the \(\pp F/\pp x^{i}\) is nonzero there. Then the level set
if non-empty, is a hypersurface of \(N_{n}\). Rests on Definitions 13.49, 13.52 and 13.58.
Derives Theorem 13.59. Let \(P\in S\) and choose a chart of \(N_{n}\) around \(P\) with coordinates \((x^{1},\ldots,x^{n})\). Relabelling the coordinates we may assume \(\pp F/\pp x^{n}\neq 0\) at \(P\). By the implicit function theorem of multivariable calculus there are then an open set \(W\subset\R^{n-1}\) containing \((x^{1},\ldots,x^{n-1})(P)\), an open interval \(I\) containing \(x^{n}(P)\), and a smooth function \(h: W\longrightarrow I\) such that, inside \(W\times I\),
Thus \(S\) meets the coordinate box \(W\times I\) in the graph of \(h\), and
is smooth and injective; its inverse is the restriction to \(S\) of the projection onto the first \(n-1\) coordinates, which is continuous, so \(\chi\) is a homeomorphism onto an open subset of \(S\) in the subspace topology. The family of all such \(\chi\), one around each point of \(S\), covers \(S\). Two of them are smoothly related, because the transition map is a composition of one such \(\chi\) with a projection and with a transition map of \(N_{n}\), all smooth. Hence Definition 13.44 is met and \(S\) is an \((n-1)\)-dimensional manifold whose inclusion into \(N_{n}\) reads, in these charts, \(u\longmapsto(u,h(u))\): an injective immersion — its Jacobian \((\identity,\pp h/\pp u)\) has rank \(n-1\) — and a homeomorphism onto its image by construction. By Definitions 13.53 and 13.58, \(S\) is a hypersurface.
∎The single input of the proof above is the implicit function theorem. Real Analysis develops differentiability in several variables but stops short of stating it, so it is stated and proved in the appendix — by an explicit contraction, in the manner of the Picard iteration of Theorem 9.8, so that no fixed-point theorem is imported. The form used here is the scalar one: where \(\pp F/\pp x^{n}\neq0\), the equation \(F=c\) determines \(x^{n}\) as a function of the remaining coordinates, of the same differentiability class as \(F\), which is exactly the assertion Equation (13.146) appealed to above.
Full derivation in Appendix A.
Derives Equation (13.146).
Take \(N_{3}=E_{3}\) with Cartesian coordinates \((x,y,z)\) and
On the level set the differential \(\pp F/\pp x^{i}=2(x,y,z)\) is nonzero, since the origin does not satisfy \(F=R^{2}\); hence \(R^{2}\) is a regular value and the sphere
is a hypersurface of \(E_{3}\), of dimension \(2\) and codimension \(1\). The charts furnished by Theorem 13.59 are here explicit: the six open hemispheres, on each of which one coordinate is a smooth function of the other two, for instance \(z=+\sqrt{R^{2}-x^{2}-y^{2}}\) on the northern one. No single chart suffices, by the compactness argument used above for the circle: \(S_{2}\) is closed and bounded, hence compact (Definition 6.9), while a chart domain must be open in \(\R^{2}\).
The induced metric (Definition 13.119) of \(S_{2}\) is the first fundamental form of Section 13.3.2; in the spherical parametrisation it is \(\dd s^{2}=R^{2}\dd\theta^{2} +R^{2}\sin^{2}\theta\,\dd\phi^{2}\), whose Gaussian curvature was computed to be \(K=1/R^{2}\) in Example 13.33. Rests on Theorem 13.59 and Definition 13.58.
Higher codimension and maps of constant rank
Theorem 13.59 cuts out a submanifold with one equation. Physics almost never presents one equation: a constraint surface in phase space, the level set of a momentum map, the shell of a conserved four-vector are all defined by several. The extension is not automatic — “at least one partial derivative is nonzero” has to become a rank condition — and the general statement is the following.
Let \(F: N_{n}\longrightarrow\R^{k}\) be smooth with \(1\le k\le n\), and let \(c\in\R^{k}\) be a regular value: at every \(P\) with \(F(P)=c\) the differential \(F_{*P}: T_{P}N_{n}\longrightarrow\R^{k}\) is surjective, equivalently the \(k\) differentials \(\dd F^{1},\ldots,\dd F^{k}\) are linearly independent at \(P\). Then
if non-empty, is an embedded submanifold of \(N_{n}\) of dimension \(n-k\), and at each of its points
For \(k=1\) this is Theorem 13.59. Rests on Theorem A.286, Definition 13.54 and Definition 13.52.
Derivation. Derives Theorem 13.62. Let \(P\in S\) and take a chart around \(P\) with coordinates \((x^{1},\ldots,x^{n})\), in which \(F\) becomes a smooth map of an open set of \(\R^{n}\) into \(\R^{k}\). Surjectivity of \(F_{*P}\) says that the \(k\times n\) Jacobian \(\pp F^{i}/\pp x^{j}\) has rank \(k\) at \(P\), so some \(k\) of its columns are independent; relabelling the coordinates, take them to be the last \(k\), and split \(x = (x',x'')\) with \(x'\in\R^{n-k}\) and \(x''\in\R^{k}\). Then \(\det\left[D_{x''}F(P)\right]\ne0\), which is the hypothesis Equation (A.523) of Theorem A.286 for the map \(F-c\). That theorem supplies open sets \(W\subseteq\R^{n-k}\) and \(V\subseteq\R^{k}\) around the two blocks of \(P\) and a smooth \(h:W\longrightarrow V\) with
From here the argument of Theorem 13.59 runs verbatim with \(n-1\) replaced by \(n-k\): \(S\) meets the coordinate box \(W\times V\) in the graph of \(h\); the map \(\chi(u) = (u,h(u))\) is smooth, injective, and a homeomorphism onto its image because its inverse is the restriction of the projection onto the first \(n-k\) coordinates; two such charts are smoothly related; and the inclusion has Jacobian \((\identity,\pp h/\pp u)\) of rank \(n-k\), so it is an injective immersion and a homeomorphism onto its image. By Definitions 13.53 and 13.54, \(S\) is an embedded submanifold of dimension \(n-k\).
For Equation (13.149): if \(\gamma\) is a curve in \(S\) through \(P\) then \(F\circ\gamma\equiv c\), so \(F_{*P}\dot\gamma = 0\) and \(T_{P}S\subseteq\ker F_{*P}\). Both spaces have dimension \(n-k\) — the first by the paragraph above, the second by the rank–nullity relation Equation (5.68) applied to the surjection \(F_{*P}\) — so they coincide.
∎The rank of \(F_{*}\) need not be maximal, and when it is merely constant the map still has a normal form. This is the statement that turns “the image of a degenerate Legendre map” or “the set swept out by a family of solutions” into a submanifold with a definite dimension, and it is proved from the same implicit function theorem.
Let \(U\subseteq\R^{N}\) be open and let \(f: U\longrightarrow\R^{N'}\) be of class \(C^{k}\), \(k\ge1\), with
Then about each \(p\in U\) and its image \(f(p)\) there are \(C^{k}\) changes of coordinates — diffeomorphisms of a neighbourhood of \(p\) in \(\R^{N}\) and of a neighbourhood of \(f(p)\) in \(\R^{N'}\) — in which \(f\) reads
Rests on Corollary A.292, Theorem A.286 and Definition 7.66.
Derivation. Derives Theorem 13.63. Translating source and target we may take \(p=0\) and \(f(0)=0\). Since \(Df(0)\) has rank \(\rho\), some \(\rho\times\rho\) minor of it is nonsingular; permuting the coordinates of \(\R^{N}\) and of \(\R^{N'}\) — themselves linear diffeomorphisms — we may assume it is the leading one. Split \(x=(x',x'')\in\R^{\rho}\times\R^{N-\rho}\) and \(y=(y',y'')\in\R^{\rho}\times\R^{N'-\rho}\), and write \(f=(f',f'')\) accordingly, so that \(\det\left[D_{x'}f'(0)\right]\ne0\).
Straightening the source. Put
a \(C^{k}\) map of \(U\) into \(\R^{N}\) whose Jacobian at the origin is the block matrix with rows \(\left(D_{x'}f',D_{x''}f'\right)\) and \(\left(0,\identity\right)\), of determinant \(\det\left[D_{x'}f'(0)\right]\ne0\). By the inverse function theorem (Corollary A.292) \(\varphi\) restricts to a \(C^{k}\) diffeomorphism of a neighbourhood of \(0\) onto an open set of \(\R^{N}\), inside which we choose a product of open balls \(W = W'\times W''\) containing \(0\); write \(\psi=\varphi^{-1}\) on \(W\). By construction the first \(\rho\) components of \(\varphi\) are \(f'\), so
with \(g = f''\circ\psi\) of class \(C^{k}\).
The rank forces \(g\) to forget \(u''\). Composing with a diffeomorphism does not change the rank of a differential, by the chain rule Equation (7.52) and the invertibility of \(D\psi\); so \(D(f\circ\psi)\) has rank \(\rho\) throughout \(W\). Its block form, read off Equation (13.154), has rows \(\left(\identity,0\right)\) and \(\left(D_{u'}g,D_{u''}g\right)\). The first \(\rho\) rows are already independent, so rank \(\rho\) forces every row of \(D_{u''}g\) to vanish: \(D_{u''}g\equiv0\) on \(W\). Since \(W''\) is a ball, hence connected, \(g(u',u'') = g(u')\) depends on \(u'\) alone.
Straightening the target. On the open set \(W'\times\R^{N'-\rho}\) define
a \(C^{k}\) diffeomorphism onto \(W'\times\R^{N'-\rho}\), its inverse being \((z',z'')\mapsto(z',z''+g(z'))\). It is defined on the image of Equation (13.154), whose first block lies in \(W'\), and
which is Equation (13.152).
∎Under the hypotheses of Theorem 13.63, every \(p\in U\) has a neighbourhood \(U_{0}\) such that \(f(U_{0})\) is an embedded submanifold of \(\R^{N'}\) of dimension \(\rho\), cut out near each of its points by exactly \(N'-\rho\) functions with linearly independent differentials. Rests on Theorem 13.63, Theorem 13.62 and Definition 13.54.
Derivation. Derives Corollary 13.64. In the coordinates of Theorem 13.63, take \(U_{0}=\psi(W)\). Its image is \(\set{(u',0)\mid u'\in W'}\), which is the level set of the last \(N'-\rho\) coordinate functions \(z^{\rho+1},\ldots,z^{N'}\) at the value \(0\); those are \(N'-\rho\) functions whose differentials \(\dd z^{\rho+1},\ldots,\dd z^{N'}\) are independent everywhere. By Theorem 13.62 the level set is an embedded submanifold of dimension \(N'-(N'-\rho)=\rho\), and transporting it back by the target diffeomorphism \(\chi^{-1}\), itself a diffeomorphism onto an open set, preserves both statements.
∎In this chapter rank names the number of indices of a tensor (Definition 13.5). In Theorems 13.62 and 13.63 it means the rank of a linear map (Definition 5.37), namely the dimension of the image of the differential. The two usages are standard and both are kept; which is meant is fixed by the object carrying the word. Note also that Theorem 13.63 is stated in \(\R^{N}\) because that is where the implicit function theorem lives; it transfers to a smooth map \(M_{m}\longrightarrow N_{n}\) of constant rank by reading it in charts, since a change of chart is exactly one of the coordinate changes the theorem already allows.
Group actions and quotient manifolds
The other standard way a manifold arises is not by cutting a bigger one with equations but by identifying its points: the configuration space of a rigid body is the set of frames modulo nothing, but the space of shapes is a set of configurations modulo rotations, and a reduced phase space is a level set modulo a symmetry group. The question is when the set of orbits is again a manifold, and the answer needs two hypotheses that are easy to state and easy to forget.
Let \(G\) be a Lie group (Lie Groups, Lie Algebras, and Fibre Bundles) acting on a manifold \(M_{n}\) by a smooth map \(G\times M_{n}\longrightarrow M_{n}\), \((g,P)\mapsto g\cdot P\), with \(e\cdot P = P\) and \(g\cdot(h\cdot P) = (gh)\cdot P\). The orbit of \(P\) is \(G\cdot P = \set{g\cdot P\mid g\in G}\), and the orbit space \(M_{n}/G\) is the set of orbits with the quotient topology. The action is
-
free iff \(g\cdot P = P\) for some \(P\) forces \(g=e\);
-
proper iff the map \(G\times M_{n}\longrightarrow M_{n}\times M_{n}\), \((g,P)\mapsto\left(g\cdot P,P\right)\), is proper, i.e. the preimage of every compact set is compact.
Let a Lie group \(G\) of dimension \(d\) act smoothly, freely and properly on \(M_{n}\). Then the orbit space \(M_{n}/G\) carries exactly one smooth structure, of dimension \(n-d\), for which the projection
is a smooth submersion (Definition 13.53); with it, \(M_{n}/G\) is Hausdorff and second countable. Every orbit is an embedded submanifold of \(M_{n}\) diffeomorphic to \(G\), and
A map \(f: M_{n}/G\longrightarrow N\) is smooth if and only if \(f\circ\pi\) is. Rests on Definition 13.66, Theorem 13.63 and Definition 13.54.
Full derivation in Appendix A.
Derives Theorem 13.67.
Let \(G=\R\) act on \(M=\R^{2}\) by \(t\cdot(x,y) = (x+t,\,y)\). The action is free, and it is proper because \((t,x,y)\mapsto\left((x+t,y),(x,y)\right)\) has continuous inverse \(\left((x',y'),(x,y)\right)\mapsto(x'-x,x,y)\) on its image, so the preimage of a compact set is a closed subset of a compact one. The orbits are the horizontal lines, and Equation (13.157) gives \(\dim(M/G) = 2-1 = 1\): the orbit space is the \(y\)-axis, as it should be.
Freeness is not decoration: let \(\SO(2)\) act on \(\R^{2}\) by rotation. The action is proper — the group is compact — but the origin is fixed, so it is not free, and the orbit space is the half-line \(\left[0,\infty\right[\), which is not a manifold at its endpoint. Properness is not decoration either: let \(\R\) act on the torus \(T^{2}=\R^{2}/\Z^{2}\) by \(t\cdot[x,y]=[x+t,\,y+\kappa t]\) with \(\kappa\) irrational. This action is free, but every orbit is dense in \(T^{2}\), because the successive returns of the line to a fixed meridian occur at the points \(\set{n\kappa\bmod1}\), which are distinct for distinct \(n\) and therefore accumulate. Hence the quotient topology on the orbit space is indiscrete: two distinct orbits have no disjoint neighbourhoods and the space is not Hausdorff, hence not a manifold. The failure is exactly the failure of properness, the group \(\R\) escaping to infinity while its points stay in a compact set. Rests on Theorem 13.67 and Definition 13.66.
Euclidean space
The Cartesian coordinate system
Consider the topological space \(M=\R^{n}\). Note that the pair \((\id_{\R^{n}},\R^{n})\) defines a coordinate system, for \(\id_{\R^{n}}\) is trivially a smooth homeomorphism: it is the bijective map that associates with each point \(x\in\R^{n}\) the point \(x\in\R^{n}\) itself.
The identity chart on \(\R^{n}\): each point is its own coordinate label.
We say that \((\id_{\R^{n}},\R^{n})\) is a Cartesian coordinate system of \(\R^{n}\). In this coordinate system the coordinates of a point \(P\in\R^{n}\) are \(x^{i}(P)\), which are called the Cartesian coordinates of the point \(P\) (with respect to the coordinate system \((\id_{\R^{n}},\R^{n})\)).
Euclidean space $E_{n}$
In the preceding context, a single Cartesian coordinate system suffices to cover all of \(\R^{n}\), so \(\set{(\id_{\R^{n}},\R^{n})}\) defines an \(n\)-dimensional atlas and, trivially, a differentiable structure. We conclude that \(\R^{n}\) is an \(n\)-dimensional differentiable manifold, which we denote \(E_{n}\) and call Euclidean space (or the Euclidean manifold); we also say that spaces of this type possess a Euclidean geometry. The fact that a single coordinate system describes the whole Euclidean space leads us to consider describing a manifold, such as a curve or a two-dimensional surface, as a submanifold of \(E_{n}\).
We say that the origin of \(E_{n}\) is the point \(O=\id_{\R^{n}}(0,\ldots,0)\) (with respect to the coordinate system \((\id_{\R^{n}},\R^{n})\)); therefore \(x^{i}(O)=0\) for all \(i\).
Pictorial representation
We represent pictorially a Cartesian coordinate system and a point \(P=\id_{\R^{n}}(x)\in E_{n}\) of coordinates \(x^{i}(P)\) as a diagram of perpendicular axes[That the axes are perpendicular has a mathematical background studied further below. This representation is due to René Descartes.]. For the case \(n=2\) it has been historically common to denote the coordinates of a generic point, \(x^{i}\), as \((x,y)\):
The plane \(E_{2}\) with perpendicular Cartesian axes \((x,y)\).
where \(E_{2}\) is called flat two-dimensional space. For the case \(n=3\) it has been historically common to denote the coordinates of a generic point, \(x^{i}\), as \((x,y,z)\):
The space \(E_{3}\) with Cartesian axes \((x,y,z)\).
where \(E_{3}\) receives the name flat three-dimensional space.
For the case \(n=2\), the Euclidean space \(E_{2}\) works very well for describing plane geometric figures, for example a square: given a square in the plane, we can specify each of its points by its coordinates. The same occurs with solid geometric figures in the space \(E_{3}\).
Vectors in $E_{n}$
Let \(P,Q\in E_{n}\) be two points. We define a vector in \(E_{n}\)[A vector in \(E_{n}\) must not be confused with a vector in the context of vector spaces.] as a straight-line segment joining \(P\) to \(Q\), endowed with the orientation from \(P\) towards \(Q\); we symbolise it \(\overrightarrow{PQ}\) and represent it pictorially as an arrow pointing from \(P\) to \(Q\).
A vector in \(E_{n}\): an oriented straight-line segment \(\overrightarrow{PQ}\).
We define the position vector of a point \(P\) with respect to the origin \(O\) as the vector \(\overrightarrow{OP}\). We introduce the notation \(\overrightarrow{OP}\stackrel{\text{not.}}{=}\vect{x}_{P}\).
The position vector \(\vect{x}_{P}=\overrightarrow{OP}\) of a point \(P\) with respect to the origin \(O\).
To refer to the position vector of arbitrary points of \(E_{n}\) we make no reference to the point, denoting it simply \(\vect{x}\). The position vector of the origin is denoted \(\vect{0}\).
Cartesian coordinates as components of the position vector
The fact that a single coordinate system describes the whole Euclidean space provides us with the bijective maps \(\vect{x}: P\longmapsto\vect{x}_{P}\) and \(x^{i}: P\longmapsto x^{i}(P)\); that is, the map assigning to each point of \(E_{n}\) its position vector (with respect to \(O\)), and the map assigning to each point of \(E_{n}\) its coordinates. A point of \(E_{n}\) is determined by its position vector or by its Cartesian coordinates, so we may establish a relation between the two concepts.
The Cartesian coordinates of \(P\) are the components of its position vector: \(\left(\vect{x}_{P}\right)^{i}=x^{i}(P)\).
We say that the Cartesian components of the position vector of a point \(P\) with respect to \(O\) are the Cartesian coordinates of \(P\) (with respect to the coordinate system \((\id_{\R^{n}},\R^{n})\)), which we symbolise as
(for a generic point we write \(\left(\vect{x}\right)^{i}=x^{i}\)). Note that this relation is one-to-one: given a coordinate system and an origin, the components of the position vector of a point are unique. With respect to the maps just mentioned we have, moreover, \(O\longmapsto\vect{0}\) and \(O\longmapsto(0,\ldots,0)\) respectively. In this sense, giving a coordinate system amounts to giving an origin (and vice versa) with respect to which we can locate points in \(E_{n}\).
$E_{n}$ as a vector space
As we know, the set of all points of \(\R^{n}\) forms an \(n\)-dimensional vector space. Moreover, each of these points gives the Cartesian components of the position vector of the corresponding point of Euclidean space \(E_{n}\). We may therefore picture the vector space \(\R^{n}\) as follows:
\(\R^{n}\) pictured as the space of position vectors of \(E_{n}\): each point doubles as the arrow from the origin to it.
Since the relation between the position vector of \(P\) (with respect to a given origin) and the coordinates of \(P\) (with respect to a given coordinate system) is one-to-one, from now on, given a Cartesian coordinate system (and hence an origin), we refer to the points of \(E_{n}\) by means of position vectors. This is only a notation to make the equations more transparent. Since we can now conceive the manifold \(E_{n}\) as a vector space — that is, position vectors as vectors of a vector space — we can define the Euclidean (or standard) inner product on \(E_{n}\) as
which clearly satisfies the axioms of an inner product. We further define the Euclidean (or standard) norm of \(E_{n}\) as the norm associated with the Euclidean inner product,
and, lastly, the Euclidean (or standard) metric of \(E_{n}\) as the metric associated with the Euclidean norm,
Displacements
Consider two points \(P\) and \(Q\) in Euclidean space, with position vectors \(\vect{x}_{P}\) and \(\vect{x}_{Q}\) with respect to a given origin. In this context it is common to symbolise the vector \(\overrightarrow{PQ}\) as
or, for two arbitrary points, simply as \(\delta\vect{x}\), which is called the displacement vector. It serves to denote the change of position between two points.
In the preceding context, one then analyses the case in which the Euclidean distance between these points is infinitesimally small.
Curves in Euclidean space
Differentiable curves
We say that a differentiable curve[More general, topologically founded definitions of a curve exist; they will not be needed in our studies.] is a one-dimensional differentiable submanifold of \(E_{n}\), which we denote \(\mathcal{C}\). Rests on Definition 13.48.
We therefore need only one coordinate to describe a curve locally — we say locally because in general more than one coordinate system is needed to describe it globally, though each of them involves a single coordinate. Thus the possible coordinate systems needed to describe the curve are of the form
The variable \(\lambda\) is a coordinate associated with the coordinate system \((\gamma,\left]a,b\right[\,)\), and \(\gamma(\lambda_{P})\), with \(\lambda_{P}\in\left]a,b\right[\), is a point of the curve. For a generic point of the curve we may identify \(\gamma(\lambda)\) with its position vector \(\vect{x}_{\gamma(\lambda)}\). The components of the position vector are therefore the quantities
which are the (Cartesian, since the curve lives in \(E_{n}\)) coordinates of the point \(\gamma(\lambda)\in\mathcal{C}\subset E_{n}\). The coordinate \(\lambda\) parametrises the curve: for each value of \(\lambda\in\R\) we have a position vector \(\vect{x}=\vect{x}(\lambda)\) characterising a point of the curve.
A curve in \(E_{n}\) described by a single chart \((\gamma,\left]a,b\right[\,)\): the coordinate \(\lambda\) labels the points \(\gamma(\lambda)\), whose position vectors are \(\vect{x}(\lambda)\).
It is necessary to clarify the difference between the two statements: \(\lambda\) is the coordinate of the point \(\gamma(\lambda)\in\mathcal{C}\), while the \(x^{i}(\gamma(\lambda))\) are the Cartesian coordinates of the point \(\gamma(\lambda)\in E_{n}\). Now, as mentioned before, in general it is not possible to describe a curve completely with a single coordinate system (we need an atlas). The problem of needing more than one coordinate system is not transcendental for calculations (though it does bring topological consequences to be studied in theoretical physics). Let us illustrate with a well-known and much-cited example: the circle of radius \(R\) as a hypersurface of \(E_{2}\), \(S_{1}=\set{(x,y)\in\R^{2}\mid x^{2}+y^{2}=R^{2}}\). We wish to describe by coordinate systems all points of \(E_{2}\) satisfying \(x^{2}+y^{2}=R^{2}\). A possible candidate is \((\gamma,\left]0,2\pi\right[\,)\), where
with the Cartesian coordinates of \(\gamma(\phi)\) given by
The circle \(S_{1}\subset E_{2}\) described by the angular chart \(\phi\mapsto(R\cos\phi,R\sin\phi)\) on \(\left]0,2\pi\right[\); the point \((R,0)\) escapes this single chart.
This coordinate system describes almost all of the desired region — almost, because the point \((R,0)\) is not included.
If we put the interval \(\left[0,2\pi\right[\) or \(\left]0,2\pi\right]\) in the domain of \(\gamma\) we would include the point, but in the definition of a differentiable manifold we worked only with open domains, to make sure we respect the topology (the maps of a coordinate system must be homeomorphisms). We could contemplate an interval \(\left]0,2\pi+\delta\right[\) so as to include the mentioned point, but then the map \(\gamma\) would no longer be bijective, which as we know is a crucial property of \(\gamma\) as a homeomorphism, meant to preserve the topology. How can we fix this problem? Is there a coordinate system that describes all the points of the circle? The answer is no, and as we shall see this is related to the fact that the circle is a compact subset of \(E_{2}\). If a single coordinate system \((\gamma,I)\) described every point of the circle, the homeomorphism \(\gamma: I\longrightarrow S_{1}\) would force \(I\) to be compact
— a contradiction, since \(I\) must necessarily be open.
We define the coordinate systems \((\gamma_{1},\left]-\delta,\pi+\delta\right[\,)\) and \((\gamma_{2},\left]\pi-\delta,2\pi+\delta\right[\,)\) with
where \(x^{i}(\gamma_{1}(\phi))=(R\cos\phi,R\sin\phi)\) with \(\phi\) varying in \(\left]-\delta,\pi+\delta\right[\), and \(x^{i}(\gamma_{2}(\phi))=(R\cos\phi,R\sin\phi)\) with \(\phi\) varying in \(\left]\pi-\delta,2\pi+\delta\right[\). We mention that we have used the same variable \(\phi\) for both coordinate systems because it is a dummy variable and there is no ambiguity (it merely happens that in both coordinate systems the curve behaves locally in the same way). The two coordinate systems cover the whole circle.
Thus \(\set{(\gamma_{1},\left]-\delta,\pi+\delta\right[\,), (\gamma_{2},\left]\pi-\delta,2\pi+\delta\right[\,)}\) defines an atlas on \(S_{1}\), which clearly possesses a one-dimensional differentiable structure. This means that \(S_{1}\) is a hypersurface of \(E_{2}\).
An example of a non-compact curve that can be described by a single coordinate system is the straight-line segment between two points \(P,Q\in E_{n}\), denoted \(L_{PQ}\). The straight line joining the two points is completely described by the coordinate system \((\gamma,\left]0,1\right[\,)\) with
where \(x^{i}(\gamma(\lambda))=(1-\lambda)x^{i}(P)+\lambda x^{i}(Q)\) with \(\lambda\) varying in \(\left]0,1\right[\). Visually, the curve “grows” from \(P\) to \(Q\) as \(\lambda\) grows in \(\left]0,1\right[\). Note further that \(x^{i}(\gamma(\lambda\to 0))=x^{i}(P)\) and \(x^{i}(\gamma(\lambda\to 1))=x^{i}(Q)\).
In the example of the circle of radius \(R\) we used for both coordinate systems the same parametrisation of the points of the curve, \((R\cos\phi,R\sin\phi)\); this does not mean we are describing the curve locally with only one coordinate system — it only means that the circle behaves locally in the same way in both coordinate systems, which is why we used the same variable. On the other hand, whenever we describe a curve locally from a Euclidean space by means of a coordinate system \((\gamma,\left]a,b\right[\,)\), we shall refer to the curve simply as the map \(x^{i}=x^{i}(\lambda)\), unless it is necessary to make explicit the interval over which \(\lambda\) varies together with the homeomorphism of that coordinate system. In this case, unlike for the circle, we give a parameter to specify that we are describing the curve locally. The stated notation is warranted by the fact that the coordinates of \(E_{n}\) used to describe the curve depend solely on the coordinate used to describe the curve locally.
Admissible change of parameter
Given a curve \(\mathcal{C}\), described by \(\vect{x}=\vect{x}(\lambda)\), we can perform coordinate changes in our coordinate systems of the form seen in Equation (13.163). A general coordinate transformation may take the form \(\bar{\lambda}=\bar{\lambda}(\lambda)\), and, as we know, the map
must be a smooth function of \(\lambda\) to guarantee that we possess a differentiable structure. The new coordinate system is then of the form \((\bar{\gamma},\left]c,d\right[\,)\) where
with \(\bar{\gamma}=\gamma\circ\bar{\lambda}\).
In the context of the theory of curves, we call the coordinate transformation \(\bar{\lambda}=\bar{\lambda}(\lambda)\) an admissible change of parameter if and only if
-
\(\bar{\lambda}\) is a function of class \(C^{1}\left(\left]a,b\right[\right)\), so as to guarantee the differentiable structure;
-
\(\forall\,\lambda\in\left]a,b\right[\), \(\dv{\bar{\lambda}}{\lambda}(\lambda)\neq 0\). This ensures the function has no extrema in \(\left]a,b\right[\), hence possesses an inverse on the given interval, so that \(\lambda=\lambda(\bar{\lambda})\) is also a coordinate transformation.
It is important to point out that an admissible change of parameter (a general coordinate transformation of the single coordinate describing the curve locally) must not be confused with a general coordinate transformation of the Euclidean space in which the curve is embedded. There is, however, no problem in performing a general coordinate transformation of the Euclidean space in order to describe the curve. As an example of an admissible change of parameter, let the straight-line curve \(L_{PQ}\) with \(P,Q\in E_{2}\) be described by
with \(\lambda\) varying in \(\left]0,1\right[\). We may perform the change of parameter \(\bar{\lambda}(\lambda)=1-\lambda\), which is clearly admissible, with \(\bar{\lambda}\) still varying in \(\left]0,1\right[\). The straight line is now described by
Visually, this change of parameter makes the curve “grow” from \(Q\) to \(P\) instead.
Tangent vector
Consider a differentiable curve in \(E_{n}\) described locally by \(x^{i}=x^{i}(\lambda)\), \(i=1,\ldots,n\), with respect to \((\gamma,\left]a,b\right[\,)\), or equivalently with respect to the basis \(\set{\hat{x}_{i}}\) of \(E_{n}\) as a vector space. We define the tangent vector to the curve at the point \(P=\gamma(\lambda_{P})\) as the vector
that is, as the variation of the coordinates of the curve as the parameter varies.
Arc element and length of a curve
Section 13.3.1 built the theory of curves extrinsically: the curve was a map into \(\R^{3}\), and its tangent, normal and binormal were vectors of the ambient Euclidean space, differentiated componentwise. The same content can be stated using nothing but the metric of the space the curve lives in and the covariant derivative it determines — the intrinsic formulation, which is the one that survives when the ambient space is removed. Throughout this stretch \((M,g)\) is a Riemannian manifold of dimension \(D\) (Definition 13.117 with \(q=0\)), with coordinates \(x^{\mu}\) and Levi-Civita connection \(\mathring{\Gamma}\) (Theorem 13.150); the reader who prefers a concrete setting may read \(M=E_{3}\) in arbitrary curvilinear coordinates, in which case \(g_{\mu\nu}\) is the Euclidean metric written in those coordinates.
The arc element is the line element Equation (13.259) of the displacement along the curve,
which for a surface in \(E_{3}\) is exactly the first fundamental form Equation (13.61), the metric \(g_{ij}\) there being the induced metric Equation (13.261).
Let \(x^{\mu}(\lambda)\), \(\lambda\in\left]a,b\right[\), be a differentiable curve in \((M,g)\). Its length is
the integrand being the metric norm of the tangent vector Equation (13.177). The curve is regular if that norm never vanishes. Rests on Definitions 13.81, 13.117 and 13.118.
The length Equation (13.168) is unchanged by a general coordinate transformation of \(M\) and by an admissible change of parameter (Definition 13.73). Rests on Definitions 13.73, 13.74 and 13.88.
Derives Proposition 13.75. Under a general coordinate transformation the integrand is a scalar: it is built by contracting \(g_{\mu\nu}\), a \((0,2)\) tensor (Definition 13.117), with two copies of the contravariant vector \(\dd x^{\mu}/\dd\lambda\) (Equation (13.199)), and the Jacobians cancel in pairs exactly as in Equation (13.206). Under a change of parameter \(\lambda=\lambda(\bar{\lambda})\) the chain rule gives \(\dd x^{\mu}/\dd\bar{\lambda} =(\dd\lambda/\dd\bar{\lambda})\,\dd x^{\mu}/\dd\lambda\), so the integrand is multiplied by \(\abs{\dd\lambda/\dd\bar{\lambda}}\), which is exactly the factor produced by \(\dd\lambda=(\dd\lambda/\dd\bar{\lambda})\dd\bar\lambda\) in the measure, the derivative having a fixed sign by Definition 13.73.
∎For a regular curve, the function
has \(\dd s/\dd\lambda>0\) everywhere and is therefore an admissible change of parameter. In the parameter \(s\) the tangent \(u^{\mu}=\dd x^{\mu}/\dd s\) is a unit vector,
and we call \(s\) the arclength parameter and \(u^{\mu}\) the unit tangent, the intrinsic counterpart of Definition 13.12. Rests on Definitions 13.73 and 13.74.
Curvature vector, binormal, and torsion
The curvature vector of a curve parametrised by arclength is the covariant derivative of its unit tangent along itself,
and the curvature function is \(\kappa=\sqrt{g_{\mu\nu}k^{\mu}k^{\nu}}\). Rests on Definition 13.76, Definition 13.145 and Theorem 13.150.
In \(E_{3}\) referred to Cartesian coordinates the Christoffel symbols vanish and Equation (13.171) collapses to \(\dd^{2}\vect{x}/\dd s^{2}\), which is the extrinsic curvature vector Equation (13.39); comparison with Equations (13.121) and (13.297) shows that the curves with \(k^{\mu}=0\) are precisely the geodesics.
\(g_{\mu\nu}u^{\mu}k^{\nu} = 0\). Rests on Definition 13.77, Definition 13.149 and Equation (13.170).
Derives Proposition 13.78. Differentiate Equation (13.170) along the curve. The Levi-Civita connection is metric compatible (Definition 13.149), so \(\mathring{\nabla}_{\lambda}g_{\mu\nu}=0\) and
This is the intrinsic form of Equation (13.40), the ambient differentiation there being replaced by covariant differentiation here.
∎Where \(\kappa>0\) we may therefore define the unit principal normal \(N^{\mu}=k^{\mu}/\kappa\), orthogonal to \(u^{\mu}\), in parallel with Definition 13.14. For \(D=3\) — the case treated extrinsically in Section 13.3.1, and the observed spatial dimension — the orthogonal complement of \(\gen{\set{u,N}}\) in each tangent space is one-dimensional, and we fix the unit binormal \(B^{\mu}\) as the unique unit vector completing \((u,N,B)\) to a positively oriented orthonormal frame; equivalently, with the Levi-Civita tensor Equation (13.257),
which reduces to the cross product Equation (13.44) in Cartesian coordinates.
Let \(D=3\) and let a curve in \((M,g)\) be parametrised by arclength with \(\kappa>0\). Then there is a function \(\tau\), the torsion, such that the frame \((u,N,B)\) obeys
In \(E_{3}\) referred to Cartesian coordinates these are Equations (13.43), (13.46) and (13.47). Rests on Proposition 13.78, Definition 13.149 and Equation (13.172).
Derives Theorem 13.79. Write \(e_{1}=u\), \(e_{2}=N\), \(e_{3}=B\) and label the frame with capital indices \(A,C\in\set{1,2,3}\), so that \(g_{\mu\nu}e_{A}^{\mu}e_{C}^{\nu}=\delta_{AC}\) along the whole curve. Expand the covariant derivative of each frame vector in the frame itself, which is legitimate because the frame is a basis of the tangent space at each point of the curve:
Differentiating \(g_{\mu\nu}e_{A}^{\mu}e_{C}^{\nu}=\delta_{AC}\) along the curve and using metric compatibility (Definition 13.149) gives \(\omega_{AC}+\omega_{CA}=0\): the matrix \(\omega\) is antisymmetric, so it has exactly three independent entries. They are fixed one by one. By Definition 13.77 and the definition of \(N\), \(\mathrm{D}u^{\mu}/\dd s=\kappa N^{\mu}\), which is Equation (13.173) and reads \(\omega_{21}=\kappa\), \(\omega_{11}=\omega_{31}=0\). Antisymmetry then gives \(\omega_{12}=-\kappa\) and \(\omega_{13}=0\). The single entry left undetermined is \(\omega_{32}\), which we define to be the torsion \(\tau\); antisymmetry gives \(\omega_{23}=-\tau\). Substituting into Equation (13.176) for \(C=2\) and \(C=3\) yields Equations (13.174) and (13.175). In Cartesian coordinates on \(E_{3}\) the Christoffel symbols vanish, \(\mathrm{D}/\dd s\) becomes \(\dd/\dd s\), and the three equations are those obtained extrinsically in Section 13.3.1.
∎The proof above is shorter than the extrinsic one of Section 13.3.1 because it uses no property of the ambient space: the antisymmetry of \(\omega\) is metric compatibility and nothing else, and it is what makes the coefficient matrix of the Frenet–Serret system antisymmetric in any Riemannian manifold. The same argument in dimension \(D\) produces \(D-1\) curvature functions and a frame of \(D\) vectors. What does not carry over unchanged is the fundamental theorem of the theory of curves — Theorem 13.16, that \(\kappa\) and \(\tau\) determine the curve up to a rigid motion — proved in Section 13.3.1 for \(E_{3}\). On a general Riemannian manifold prescribed \(\kappa\) and \(\tau\) still determine the curve uniquely once an initial point and an initial orthonormal frame are fixed, and for as far as the solution extends: the three Frenet equations above, together with \(\dd x^{\mu}/\dd s=u^{\mu}\), are a first-order system in which the frame enters linearly. Uniqueness up to isometry is a stronger statement, and it needs the isometries of Section 13.7.1 to act transitively on orthonormal frames, so that two solutions with different initial data can be carried onto one another. That is a genuine hypothesis and not a technicality: a generic metric admits no isometry but the identity, and then two curves with the same \(\kappa\) and \(\tau\) started at different points are simply two different curves. Enough isometries are available exactly on the maximally symmetric spaces of Theorem 13.160, whose curvature is constant, \(E_{3}\) being the case \(\epsilon=0\).
Tensor analysis on manifolds
Vectors, covectors, and tensors on a manifold
The definition of a vector on a manifold
The definition of a vector in \(\R^{N}\) is an oriented straight-line segment joining two points. This tells us that a vector as we know it is a bilocal object. Would this definition remain consistent if the space no longer had Euclidean geometry — that is, on a general manifold? Consider an \(n\)-dimensional manifold \(M\subset\R^{N}\). The Euclidean metric of \(\R^{N}\) allows us to define straight lines joining points on the manifold, but such straight lines are clearly not contained in \(M\).
Along which route, then, should the line between \(A\) and \(B\) be traced? There is no possible route unless we define beforehand a metric on the space \(M_{n}\), for without this concept we cannot have a concept of straight line. For this reason we cannot define a vector on the manifold as a bilocal object: we must develop a formalism to understand a vector as a local object.
Let \((\gamma,\left]a,b\right[\,)\) be a coordinate system describing a curve contained in \(M\). The curve is characterised by the position vectors \(\vect{x}_{\gamma(\lambda)}\) with \(\lambda\) varying in \(\left]a,b\right[\). Denote by \(\lambda_{0}\) the value of the parameter at which the point of the curve is \(P=\gamma(\lambda_{0})\in M\).
A curve through \(P=\gamma(\lambda_{0})\) on a manifold \(M\): its tangent vector at \(P\) does not lie in \(M\) but in a tangent plane touching \(M\) at \(P\).
We say that the vector \(V_{P}\in\R^{N}\) is a tangent vector to \(M\) at \(P\) if and only if it is tangent to the curve at \(P\), that is, if its components with respect to the coordinate system \((\gamma,\left]a,b\right[\,)\) are
The tangent vector \(V_{P}\) must not be viewed as a displacement lying in \(M\), but as lying in some kind of tangent surface coinciding with \(M\) at \(P\). Since \(M\) is equipped with an atlas, we can define a differentiable map of the form \(f: M\longrightarrow\R\). Then \(f\circ x:\R\longrightarrow\R\) is differentiable, being a composition of differentiable maps, and
But which map \(f\) do we choose? It does not matter. What matters is to see how the curve generates changes in \(f\). Indeed, \(f\circ x:\R\longrightarrow\R\) allows us to evaluate the changes in \(f\); these changes are given by the directional derivative (Equation (13.178)) in the direction of the curve, that is, in the direction of the vector \(V_{P}\).
We may conclude that the operator generating the changes of \(f\) in the direction of \(V_{P}\) is \(\left.\dv{}{\lambda}\right|_{\lambda_{0}}\). By the chain rule and Equation (13.177),
Therefore \(\left.\dv{}{\lambda}\right|_{\lambda_{0}}\) can be written as a linear combination of the basis \(\set{\pp_{i}}\), with components exactly \(V^{i}_{P}\).
We define a vector on \(M_{n}\) at \(P\) as the operator
We say that a vector field is an assignment rule that associates a vector with each point of a manifold. Rests on Definition 13.81.
We must then symbolise a vector field without reference to the point of the manifold, since we are considering all of them. In the context of the definition of a vector, we symbolise a vector field \(V\) as
When no reference to the point is made, it is understood that we are speaking of a field.
Tangent space and coordinate basis
Owing to the linearity of the derivative, and since it has the properties of addition and multiplication by a scalar, the set of all vectors tangent to \(M\) at \(P\) forms an \(n\)-dimensional vector space, which we call the tangent space and denote \(T_{M}(P)\) (for an arbitrary point, \(T_{M}\)). This may seem inconceivable, but so it is: at each point of the manifold we can define a vector space.
As discussed in the definition of a vector on a manifold, we can write a vector \(V\) on \(M\) as a linear combination of the vectors \(\pp_{i}\),
This means that the set \(\set{\pp_{i}}^{n}_{i=1}\) forms a basis of the tangent space, which we call the coordinate (or holonomic) basis. We therefore have
The vectors of the coordinate basis are tangent to the coordinate lines.
For example, the coordinate-basis vectors of \(E_{2}\) in Cartesian coordinates are \(\pp_{x}\) and \(\pp_{y}\), while in polar coordinates they are \(\pp_{r}\) and \(\pp_{\phi}\). The relation between the Cartesian and polar coordinates describing \(E_{2}\) is
so that
and likewise
whence
Covectors and the cotangent space
Let \(P\in M\). We define a covector at \(P\) as a linear functional
Rests on Definition 13.81.
From the theory of vector spaces we know that the set of all linear functionals defined on a vector space forms a vector space, called the dual. Thus the set of all covectors defined at a point \(P\in M\) forms the vector space
which we call the cotangent space at \(P\). As shown below (Equation (13.239)), the differentials \(\dd x^{i}\) furnish the basis dual to the coordinate basis, so that
Tensors
Let \(P\in M\), and consider \(T_{M}(P)=\gen{\set{e_{i}}}\) and \(T^{*}_{M}(P)=\gen{\set{\omega^{i}}}\). We define a tensor of type \((r,s)\) at \(P\) as a multilinear map
whose components are its values on the basis elements,
We say that a tensor is contravariant if it is of type \((r,0)\) with \(r\in\N\), and covariant if it is of type \((0,s)\) with \(s\in\N\). For example, a vector is a contravariant tensor of type \((1,0)\) and a covector is a covariant tensor of type \((0,1)\).
Pullback and pushforward
If we wished to compare two vectors assigned at distinct points of a differentiable manifold, we could not simply subtract them, since they belong to different vector spaces. We must provide a mechanism that maps one of the vectors into the space of the vector with which we wish to compare it.
Consider a diffeomorphism \(f: M\longrightarrow M\). For a point \(P\), the map \(f\) induces the pushforward \(f_{*}: T_{M}(P)\longrightarrow T_{M}(f(P))\). This amounts to taking a tangent vector at \(P\) and locating it at \(f(P)\).
One also has the pullback \(f^{*}: T^{*}_{M}(f(P))\longrightarrow T^{*}_{M}(P)\), which to each covector at \(f(P)\), say \(\omega_{f(P)}\), assigns by definition the covector \(f^{*}(\omega_{f(P)})=\omega_{f(P)}\circ f_{*}\) at \(P\). Covectors are linear functionals acting on vectors of the tangent space: for a tangent vector \(V_{P}\) at \(P\),
Vector fields on basic manifolds
The Cartesian chart of Section 13.4.2 is global, so the coordinate frame
consists of three vector fields defined at every point, and every vector field on \(E_{3}\) is \(V=V^{x}\pp_{x}+V^{y}\pp_{y}+V^{z}\pp_{z}\) with three functions as components. Spherical coordinates,
give a second chart, and the two frames are related by the chain rule, which is the transformation law of a coordinate basis, \(\bar{\pp}_{i}=(\pp x^{j}/\pp\bar{x}^{i})\pp_{j}\):
Two features are worth recording. First, the spherical chart is not global: it degenerates at \(r=0\) and along the polar axis \(\sin\theta=0\), where \(\pp_{\phi}\) vanishes and \(\phi\) is undefined — the same obstruction met for the circle in Section 13.4.3. Second, the spherical coordinate frame is not orthonormal: Equations (13.188), (13.189) and (13.190) give \(\abs{\pp_{r}}=1\), \(\abs{\pp_{\theta}}=r\) and \(\abs{\pp_{\phi}}=r\sin\theta\) in the Euclidean metric Equation (13.159). The normalized frame \(\hat{e}_{r}=\pp_{r}\), \(\hat{e}_{\theta}=\pp_{\theta}/r\), \(\hat{e}_{\phi}=\pp_{\phi}/(r\sin\theta)\) familiar from vector calculus is therefore not a coordinate frame: it is an anholonomic frame, with nonvanishing structure functions in the sense of Equation (13.266), and this is the elementary origin of the vielbein of Section 13.8. Rests on Definition 13.82 and Equation (13.199).
Let \(\mathcal{C}\) be a differentiable curve described locally by \((\gamma,\left]a,b\right[\,)\) as in Equation (13.163). Its tangent space at every point is one-dimensional, spanned by the single coordinate-basis vector \(\dd/\dd\lambda\), so every vector field on \(\mathcal{C}\) has the form
with one component function \(f\). Read in the ambient \(E_{n}\) through Equation (13.166), the field \(\dd/\dd\lambda\) is the tangent vector \(\left(\dd x^{i}/\dd\lambda\right)\hat{x}_{i}\) of the curve. Under an admissible change of parameter (Definition 13.73) the chain rule gives \(\dd/\dd\lambda=\left(\dd\bar{\lambda}/\dd\lambda\right) \dd/\dd\bar{\lambda}\), the one-dimensional instance of Equation (13.199). Since the coordinate field \(\dd/\dd\lambda\) never vanishes, a curve always carries a nowhere-vanishing vector field; the next example shows that this is a privilege of low dimension and not a general fact. Rests on Definitions 13.71 and 13.82.
Let \(S_{2}\subset E_{3}\) be the sphere of radius \(R\) (Example 13.61) and consider the three rotation generators of \(E_{3}\),
Each is tangent to \(S_{2}\), because each annihilates the defining function \(F=x^{2}+y^{2}+z^{2}\) of Equation (13.147): for instance \(L_{3}F=x(2y)-y(2x)=0\). A vector that annihilates \(F\) is tangent to the level set, since \(F\) is constant along any curve inside it; hence the \(L_{i}\) restrict to vector fields on \(S_{2}\). In the spherical chart Equation (13.187) with \(r=R\) held fixed, so that \((\theta,\phi)\) are coordinates on the sphere, comparison with Equations (13.189) and (13.190) gives
The last is immediate from Equation (13.190), which reads \(\pp_{\phi}=-y\,\pp_{x}+x\,\pp_{y}=L_{3}\); the other two follow by the same substitution. Computing the Lie brackets Equation (13.273) of the fields Equation (13.192) directly gives
so the three fields span a three-dimensional Lie algebra; it is isomorphic to \(\mathfrak{so}(3)\), the sign being absorbed by \(L_{i}\to-L_{i}\). These are the Killing vectors of the round sphere (Section 13.9.2), and they saturate the bound \(D(D+1)/2=3\) of Proposition 13.140 at \(D=2\): the round \(S_{2}\) is maximally symmetric.
Each of the three fields vanishes somewhere: \(L_{3}=\pp_{\phi}\) vanishes at the two poles \(\sin\theta=0\), and \(L_{1}\), \(L_{2}\) likewise vanish at the two points of \(S_{2}\) on their own axis. This is not an accident of the choice. No continuous vector field on \(S_{2}\) is nowhere vanishing — the “hairy ball” theorem. It follows from the Poincaré index theorem, which states that on a compact orientable surface the indices of the isolated zeros of a vector field sum to the Euler characteristic of the surface [DoCarmo:1976]; for \(S_{2}\) that characteristic is \(2\), so the sum of the indices cannot be zero and zeros cannot be absent. Two consequences are worth naming: \(S_{2}\) admits no global frame of the kind Equation (13.186) — it is not parallelizable, unlike \(E_{n}\) and unlike the curve of Example 13.86 — and consequently no vielbein (Section 13.8) on \(S_{2}\) can be defined by a single chart-independent formula valid everywhere. Charts are not a convenience here but a necessity. Rests on Definition 13.82, Example 13.61 and Definition 13.81.
Tensors under general coordinate transformations
General coordinate transformations
We call a coordinate system \(K\) a set of functions \(\set{x^{i}}^{n}_{i=1}\) such that the location of a point \(P\) is given by \(x^{i}(P)\).
We call general coordinate transformation (GCT) a transformation relating the coordinates \(x^{i}\) of a coordinate system \(K\) to the coordinates \(\bar{x}^{i}\) of a coordinate system \(\bar{K}\), that is, a transformation of the form
For practical purposes we assume that these \(n\) functions of \(n\) variables are of class \(C^{2}\).
In the context of coordinate systems it is necessary to extend the definition of the Kronecker delta to every coordinate system: we define the Kronecker delta symbol as the quantities
in every coordinate system.
A GCT is linear if and only if it can be written as \(\bar{x}^{i}=\Lambda^{i}{}_{j}x^{j}\) with a constant matrix \(\Lambda^{i}{}_{j}\) (compare Definition 13.2).
Transformation law of a vector
Since a vector field is the coordinate-independent object \(V = V^{i}\pp_{i}\), its components in two coordinate systems obey
where both sides have been applied to the coordinate function \(\bar{x}^{j}\); hence
Contravariant, covariant, and mixed tensors
We define the components of a contravariant tensor of rank \(r\) as quantities that change coordinate system under the transformation
placing the indices in the upper position to denote that these indices enumerate the components of a contravariant tensor. Analogously, we define the components of a covariant tensor of rank \(s\) as quantities that change coordinate system under the transformation
placing the indices in the lower position. Rests on Definition 13.88.
We define the components of a mixed tensor of order \(r+s\) as quantities that change coordinate system under the transformation
which we call a tensor of type \((r,s)\) and order \(r+s\). Rests on Definitions 13.88 and 13.89.
We call contravariant vector a tensor of type \((1,0)\); from Equation (13.200), its components with respect to another coordinate system are given by
We call covariant vector a tensor of type \((0,1)\); from Equation (13.201),
The displacement vector
For a GCT one cannot say that the coordinates themselves transform as a vector would: under a GCT the coordinates are neither contravariant nor covariant. Differentiating the relations Equation (13.197), however, we see that
so the differential of the coordinates transforms as a contravariant vector, which we call the displacement vector.
Scalars and invariance
The notion of the rank of a tensor under general transformations is the number of indices it carries, and the number of times the Jacobian matrix of the transformation must be applied to carry out a change of coordinates. If we apply this matrix zero times, we obtain the same quantities in both coordinate systems, since we are not transforming them at all. We may extend the definition to tensors of rank \(0\) by defining these invariant quantities: quantities that depend on position with respect to a coordinate system but take the same value in every coordinate system. That is,
More generally, we call invariant tensor a tensor satisfying
We call scalar an invariant tensor of rank \(0\).
Scalar product
Consider the components of a contravariant vector \(T^{i}\) and of a covariant vector \(U_{l}\). With respect to \(\bar{K}\),
Consider the quantity
from which we see that this type of product is a scalar under GCTs. We call this multiplication of vectors the scalar product.
The Kronecker delta as a mixed tensor
From its definition, the Kronecker delta is an invariant tensor. Observe that
Thus, from Equation (13.202), the Kronecker delta satisfies the definition of a tensor of type \((1,1)\).
Symmetry and antisymmetry
We say that a tensor with components \(T^{i_1\cdots i_r}{}_{l_1\cdots l_s}\) is a tensor symmetric with respect to the indices \(i_{p}\) and \(i_{q}\) if and only if
or a tensor symmetric with respect to the indices \(l_{p}\) and \(l_{q}\) if and only if
The quantities so defined are tensors under GCTs, by the permutation of indices. Analogously, the tensor is antisymmetric with respect to a pair of upper indices, or a pair of lower indices, if the interchange of the pair produces a minus sign; again the quantities so defined are tensors under GCTs.
Consider now a tensor of type \((1,1)\), and suppose that in some coordinate system
that is, quantities symmetric with respect to one contravariant and one covariant index. Then
from which we see that this type of symmetry does not generate tensors under GCTs. Symmetry or antisymmetry properties must therefore always be taken with respect to two contravariant or two covariant indices; for this reason we display them below only for covariant indices.
The symmetry or antisymmetry of a rank-\(2\) tensor reads
which in matrix notation is equivalent to
Symmetric and antisymmetric parts
We define the symmetric part of a rank-\(r\) tensor \(T_{i_1\cdots i_r}\) with respect to the indices \(i_{s}\) and \(i_{t}\) as the quantities
and analogously the antisymmetric part, with respect to the indices \(i_{s}\) and \(i_{t}\), as the quantities
Adding Equations (13.213) and (13.214) we see that
so a rank-\(r\) tensor can be decomposed into its symmetric and its antisymmetric part. For a tensor of order \(2\),
The vector-product tensor
We define the vector-product tensor as
From the algebra of tensors, the quantities \(C_{ij}\) are the components of a tensor of type \((0,2)\) under GCTs. Moreover, from Equation (13.217),
so that
that is, the vector-product tensor is an antisymmetric tensor of order \(2\).
Gradient of a scalar
Consider Equation (13.206) and apply the partial derivative with respect to the coordinate \(\bar{x}^{i}\):
from which we see that
that is, the gradient of a scalar is a covariant vector under GCTs.
A Jacobian identity
From the chain rule,
where \(\abs{\pp\bar{x}/\pp x}\) denotes the Jacobian determinant of the transformation.
Tensor densities
The quantities met so far transform under a GCT by factors of \(\pp\bar{x}/\pp x\) alone. A second family of objects, indispensable as soon as one integrates over a manifold, transforms by those factors and by a power of the Jacobian determinant.
Write
for the Jacobian determinant of the GCT Equation (13.197). Quantities \(\mathfrak{T}^{i_1\ldots i_r}{}_{j_1\ldots j_s}\) are the components of a tensor density of weight \(w \in \Z\), of type \((r,s)\), if under a GCT they transform as
A density of weight \(0\) is an ordinary tensor (Equations (13.200) and (13.201)); a density of weight \(0\) and type \((0,0)\) is a scalar Equation (13.206). A quantity transforming by \(\abs{J}^{-w}\) rather than \(J^{-w}\) is a relative tensor density of weight \(w\), and the two notions differ only for orientation-reversing transformations. Rests on Definitions 13.88 and 13.89.
Densities of the same weight and type form a vector space under pointwise addition; the product of a density of weight \(w_1\) with one of weight \(w_2\) is a density of weight \(w_1 + w_2\); contraction of an upper with a lower index preserves the weight; and the quotient of two densities of equal weight, where defined, is an ordinary tensor. Rests on Definition 13.92.
Derivation. Derives Proposition 13.93. Each clause is read off Equation (13.222). Linearity is immediate, since the same factor \(J^{-w}\) multiplies both summands. For the product, the Jacobian factors multiply as \(J^{-w_1}J^{-w_2} = J^{-(w_1+w_2)}\) while the partial-derivative factors are those of the tensor product. For a contraction, one factor \(\pp\bar{x}^{i}/\pp x^{k}\) meets one factor \(\pp x^{l}/\pp\bar{x}^{i}\) and the two collapse to \(\delta^{l}{}_{k}\) by Equation (13.198), leaving \(J^{-w}\) untouched. For the quotient, the factors \(J^{-w}\) cancel between numerator and denominator, so the result transforms with no Jacobian factor at all.
∎The Levi-Civita symbol \(\varepsilon_{i_1\ldots i_n}\), defined to take the values \(0,\pm1\) in every coordinate system, is a covariant density of weight \(-1\): expanding the determinant by the Leibniz formula gives
which is Equation (13.222) with \(w = -1\). Correspondingly \(\sqrt{\abs{\det g}}\) is a scalar density of weight \(+1\), because \(\bar{g}_{ij} = (\pp x^{k}/\pp\bar{x}^{i})(\pp x^{l}/\pp\bar{x}^{j}) g_{kl}\) gives \(\det\bar{g} = J^{-2}\det g\). Their product \(\sqrt{\abs{\det g}}\,\varepsilon_{i_1\ldots n}\) has weight \(0\) and is therefore an honest tensor — the volume form of Section 13.6.7 — and \(\sqrt{\abs{\det g}}\,\dd^{n}x\) is invariant, which is what makes an integral over a manifold well defined. Rests on Definition 13.92 and Proposition 13.93.
The sources develop none of this in general and treat only the following three-dimensional instances.
The pseudo-vector cross product
Consider the vector-product tensor (Equation (13.217)). Since this tensor is totally antisymmetric, we can define the associated dual pseudo-tensor, in three dimensions, as
which we call the pseudo-vector cross product, or cross product of two vectors. From Equation (13.217),
so that
The source calls Equation (13.225) a pseudo-vector and stops there. The name records a property, and the property has to be established; the three results below do that, and supply at the same time the elementary identities of the three-dimensional cross product, which the sources use without stating and which the physical parts of this treatise cref constantly.
For vectors \(\vect{A},\vect{B},\vect{C}\) of \(E_{3}\), with \((\vect{A}\times\vect{B})_{i}=\varepsilon_{ijk}A_{j}B_{k}\) as in Equation (13.225),
Rests on Proposition 13.6 and Equation (13.225).
Derivation. Derives Proposition 13.95. Equation (13.226): all three scalars equal \(\varepsilon_{ijk}A_{i}B_{j}C_{k}\) after renaming the summation indices, because \(\varepsilon_{ijk}=\varepsilon_{jki}=\varepsilon_{kij}\) — a cyclic permutation of three symbols is even.
Equation (13.227): by Equation (13.225) and Equation (13.29),
Equation (13.228): with the same contraction,
which is \(A_{j}A_{j}B_{k}B_{k}-\left(A_{j}B_{j}\right)^{2}\).
∎Under an orthogonal coordinate transformation Equation (13.8) with matrix \(a_{ij}\), the quantities Equation (13.225) obey
They therefore transform as the components of a Cartesian vector (Equation (13.21) with \(r=1\)) under a proper orthogonal transformation, \(\det a = +1\), and pick up an extra minus sign under an improper one, \(\det a = -1\): the cross product is a pseudo-vector, or axial vector, and not a vector. Rests on Proposition 13.6, Definition 13.5 and Equation (13.13).
Derivation. Derives Proposition 13.96. Write Equation (13.31) with the index names \((i,j,k)\) on the symbol and \((s,l,m)\) on the free slots, \(\varepsilon_{ijk}a_{is}a_{jl}a_{km}=\varepsilon_{slm}\det a\), and contract both sides with \(a_{ts}\), summing over \(s\). On the left the orthogonality condition Equation (13.13) gives \(a_{is}a_{ts}=\delta_{it}\), so
Now \(C_{jk}\) of Equation (13.217) is an ordinary rank-\(2\) tensor, so \(C'_{jk}=a_{jl}a_{km}C_{lm}\) by Equation (13.22), and Equation (13.224) with Equation (13.230) gives
which is Equation (13.229). The factor \(\det a=\pm1\) of Equation (13.15) is what distinguishes the two cases.
∎Stokes' theorem for an antisymmetric tensor field
For an antisymmetric tensor field \(H_{ij}\) with associated vector \(H_{i}\),
where \(\dd f^{ij}\) is the antisymmetric surface element of \(S\) and \(\pp S\) its boundary. The source states the relation without derivation; it is the flat three-dimensional instance of the general Stokes theorem for differential forms, and both are proved in the appendix.
Full derivation in Appendix A.
Derives Equation (13.231).
The relation between \(H_{ij}\) and \(H_{i}\) is left implicit in the source and is worth making explicit, because the identity is false under the other natural reading. It holds when \(H_{ij}\) is the rotor tensor of \(H_{i}\),
and not when \(H_{i}\) is the dual pseudo-vector \(\tfrac{1}{2}\varepsilon_{ijk}H_{jk}\) of Equation (13.224): with the latter reading the left-hand side of Equation (13.231) is the flux of \(H_{i}\) through \(S\), which has no reason to equal its circulation around \(\pp S\). With Equation (13.232), and writing the surface element as \(\dd f^{ij}=\varepsilon^{ijk}\dd f_{k}\), the left-hand side of Equation (13.231) becomes \(\tfrac{1}{2}\varepsilon^{ijk}(\pp_{i}H_{j}-\pp_{j}H_{i})\dd f_{k} =(\nabla\times\vect{H})\cdot\dd\vect{f}\), so in three dimensions Equation (13.231) is the classical Stokes theorem (Theorem 7.100) transcribed into tensor notation. In the language of Section 13.6 it is the statement \(\int_{S}\dd H=\oint_{\pp S}H\) for the one-form \(H=H_{i}\dd x^{i}\), whose exterior derivative has components Equation (13.232). In three dimensions nothing here is owed to the manifold machinery: the classical statement is derived outright in Theorem 7.100 from Green's theorem, with the scope of that derivation set out in Remark 7.102. The general form \(\int_{M}\dd\omega=\oint_{\pp M}\omega\) on a compact oriented manifold with boundary, of which Equation (13.231) is the flat three-dimensional instance, is proved in the appendix, and Equation (13.231) is deduced there from it under exactly the hypothesis Equation (13.232) made explicit above. The two derivations are independent — the appendix does not use Theorem 7.100 and Real Analysis does not use the appendix — so they check each other rather than one resting on the other.
Differential forms: $k$-forms and $k$-vectors
$k$-forms and $k$-vectors
We say that a \(k\)-form is a totally antisymmetric tensor of type \((0,k)\). In particular, a covector is a \(1\)-form. Rests on Definitions 13.83 and 13.84.
We say that a \(k\)-vector is a totally antisymmetric tensor of type \((k,0)\). In particular, a vector is a \(1\)-vector. Rests on Definitions 13.81 and 13.84.
The set of all \(k\)-forms at \(P\) forms a vector space, which we denote \(\Lambda^{k}_{M_n}(P)\). Of course,
Analogously, the set of all \(k\)-vectors forms a vector space, denoted \({}^{*}\Lambda^{k}_{M_n}(P)\), with
Both assertions, and the dimension of the two spaces, follow from a single count.
Let \(M_{n}\) be an \(n\)-dimensional differentiable manifold and \(P\in M_{n}\). The set \(\Lambda^{k}_{M_n}(P)\) of \(k\)-forms at \(P\) is a vector subspace of the space of tensors of type \((0,k)\) at \(P\), and
while \(\Lambda^{k}_{M_n}(P)=\set{0}\) for \(k>n\). The same holds for the space \({}^{*}\Lambda^{k}_{M_n}(P)\) of \(k\)-vectors, with upper indices throughout. Rests on Definitions 13.83, 13.84 and 13.98.
Derives Proposition 13.100. The tensors of type \((0,k)\) at \(P\) form a vector space under pointwise addition and scaling of their components (Definition 13.84). Total antisymmetry is the requirement
a family of linear homogeneous conditions on the components; its solution set is therefore a linear subspace, which is the first assertion. Equivalently: if \(\psi\) and \(\varphi\) are totally antisymmetric so is \(\alpha\psi+\beta\varphi\), because each condition Equation (13.236) is preserved by linear combination.
For the dimension, fix the coordinate cobasis \(\set{\dd x^{i}}\) of \(T^{*}_{M_n}(P)\) (Equation (13.239)) and read a \(k\)-form through its components \(\psi_{i_1\ldots i_k}\). Total antisymmetry has two consequences. First, any component with a repeated index vanishes: exchanging the two equal slots leaves the component unchanged and, by Equation (13.236), reverses its sign. Second, any component whose indices are distinct is determined by the one whose indices are arranged in increasing order, since a permutation \(\sigma\) of the slots multiplies the component by \(\sgn\sigma\):
Hence the linear map sending \(\psi\) to the family of its ordered components \(\psi_{i_1\ldots i_k}\) with \(i_1<\cdots<i_k\) is injective. It is also surjective: given arbitrary numbers assigned to the ordered index sets, Equation (13.237) extends them to a full family of components, and the extension is consistent because \(\sgn\) is a homomorphism, so the result is a totally antisymmetric tensor with the prescribed ordered components. The map is therefore an isomorphism, and \(\dim\Lambda^{k}_{M_n}(P)\) equals the number of strictly increasing \(k\)-tuples drawn from \(\set{1,\ldots,n}\), that is, the number of \(k\)-element subsets of an \(n\)-element set: \(\binom{n}{k}\), which is Equation (13.235). If \(k>n\) no such tuple exists — any \(k\) indices taken from \(n\) values must repeat one — so every component vanishes and the space is trivial. The argument used nothing but the antisymmetry of the components, so it applies verbatim to \(k\)-vectors.
∎Once the wedge product is available (Definition 13.102), the isomorphism just constructed is realized by an explicit basis: the \(\binom{n}{k}\) products \(\dd x^{i_1}\wedge\cdots\wedge\dd x^{i_k}\) with \(i_1<\cdots<i_k\). Equation (13.235) is the dimension count used in Section 13.6.7 to pair \(k\)-forms with \((n-k)\)-forms.
The basis of the cotangent space
Let \(P\in M_{n}\), with \(M_{n}\) a differentiable manifold, and let \(f: M_{n}\longrightarrow\R\) be a differentiable map. Consider the map
We claim that \(\dd f(V_{P}) = V_{P}(f)\). Indeed, \(\dd f\) is defined as
and, recalling the discussion of the definition of a vector, \(\set{\pp_{i}}\) is a basis of the tangent space at \(P\). Thus
From Equation (13.238), taking the vectors \(\pp_{i}\in T_{M_n}(P)\) and \(f = x^{j}\), we see that
so \(\set{\dd x^{i}}\) is the basis dual to \(\set{\pp_{i}}\), that is,
and in particular the \(\dd x^{i}\) are covectors, that is, \(1\)-forms. From this we see that for the set of all tensors of type \((0,k)\),
The wedge product
Wedge product and change of basis
Observe the following. The space of all \(2\)-forms, \(\Lambda^{2}_{M_n}(P)\), can be generated by the basis \(\set{\dd x^{i}\otimes\dd x^{j}}\), but we must be careful: the direct product of two \(2\)-forms is not necessarily a \(4\)-form, since the result is a tensor of type \((0,4)\) that is not necessarily antisymmetric; that is, it can happen that
For closure — referring to the antisymmetry property — we must define a new type of product. As we now show, this can be achieved by a change of basis. In the basis already discussed for the cotangent space,
where the antisymmetry of \(\psi_{ij}\) was used in the second step.
As we have seen, the \(\dd x^{i}\) are \(1\)-forms. We define the wedge product as the binary operation on the basis elements
Rests on Definition 13.98 and Equation (13.239).
Therefore
In short, we have performed the change of basis
Antisymmetry
From the definition of the wedge product it is immediate that
and clearly, taking \(j=i\) (no sum),
Powers of a $k$-form
We introduce the notation
which we call the \(m\)-th power of a \(k\)-form, with respect to the wedge product.
Commutation rule
Having defined the wedge product explicitly, we can analyse its commutativity. Let \(\psi\) be a \(k\)-form and \(\varphi\) an \(l\)-form. In the basis discussed above,
We thus obtain the commutation rule
Various properties follow from this relation by simple arithmetic. For example, consider the case \(\varphi=\psi\) with \(k\) odd. Since the square of an odd number is odd,
hence
The exterior derivative
We define the exterior derivative as the map
It is clear that the exterior derivative is a linear operator, owing to the linearity of the ordinary derivative \(\pp_{j}\).
Nilpotency
Let \(\psi\) be a \(k\)-form. Then
Now, if \(\psi_{i_1\ldots i_k}\) is of sufficiently high differentiability class, \(\pp_{i}\pp_{j}\) is symmetric in \(i\) and \(j\), while \(\dd x^{i}\wedge\dd x^{j}\) is antisymmetric in those indices; hence their contraction vanishes. Thus
Leibniz rule
Let \(\psi\) be a \(k\)-form and \(\varphi\) an \(l\)-form. Then
where in the last step the \(1\)-form \(\dd x^{j}\) was carried through the \(k\) factors \(\dd x^{i_1}\wedge\cdots\wedge\dd x^{i_k}\) at the cost of a factor \((-1)^{k}\), by Equation (13.241). We conclude the identity
known as the Leibniz rule for differential forms.
All the properties of the exterior derivative mentioned above are consequences of the properties of the ordinary derivative. This is because this special derivative is only a definition designed to deal with differential forms; the essential object is the derivative we all know.
Closed and exact forms
Let \(\psi\) be a \(k\)-form. We say that \(\psi\) is closed if and only if
Let \(\psi\in\Lambda^{k}_{M_n}(P)\) be a \(k\)-form. We say that \(\psi\) is exact if and only if there exists a \((k-1)\)-form \(\varphi\in\Lambda^{k-1}_{M_n}(P)\) such that
If a \(k\)-form is exact, then it is closed. Rests on Definition 13.105, Definition 13.106 and Equation (13.245).
Derivation. Derives Lemma 13.107. Immediate from the nilpotency of the exterior derivative, Equation (13.245): if \(\psi=\dd\varphi\) then \(\dd\psi=\dd(\dd\varphi)=0\).
∎The converse is not automatic, and its validity depends on the topology of the domain. Two statements settle the matter for everything this treatise needs: on a star-shaped domain the converse holds outright, and in general it fails, with a completely explicit counterexample.
An open set \(U\subseteq\R^{n}\) is star-shaped about \(x_{0}\in U\) if and only if the segment from \(x_{0}\) to \(x\) lies in \(U\) for every \(x\in U\). A ball is star-shaped about its centre, and every point of every open set has a star-shaped neighbourhood. Rests on Definition 6.2 and Equation (6.12).
Let \(U\subseteq\R^{n}\) be open and star-shaped about the origin, and let \(\psi\) be a closed \(k\)-form on \(U\) with \(k\ge1\). Then \(\psi\) is exact: \(\psi=\dd\varphi\) with \(\varphi\) the \((k-1)\)-form
Consequently a closed form is locally exact on any manifold, each point having a chart domain that may be taken star-shaped. Rests on Definitions 13.105, 13.106 and 13.108.
Derives Theorem 13.109. Write \(h\psi\) for the \((k-1)\)-form Equation (13.247); it is defined because \(tx\in U\) for \(t\in[0,1]\), and it is totally antisymmetric in \(i_{2},\ldots,i_{k}\) because \(\psi\) is. We prove the homotopy identity
from which the theorem follows by putting \(\dd\psi=0\).
In the convention of Definition 13.103 the components of the exterior derivative of an \(l\)-form \(\chi\) are \(\left(\dd\chi\right)_{i_{1}\ldots i_{l+1}} = (l+1)\,\pp_{[i_{1}}\chi_{i_{2}\ldots i_{l+1}]}\), the brackets denoting total antisymmetrization. Differentiating Equation (13.247) under the integral sign, the derivative falls either on the explicit factor \(x^{j}\) or on the argument \(tx\):
the first term arising from \(\delta^{j}_{[i_{1}} \psi_{|j|i_{2}\ldots i_{k}]}=\psi_{[i_{1}i_{2}\ldots i_{k}]} =\psi_{i_{1}\ldots i_{k}}\), and the vertical bars excluding \(j\) from the antisymmetrization. For the second term of Equation (13.248),
Separating the first slot from a total antisymmetrization over \(k+1\) indices gives the algebraic identity
which for \(k=1\) reads \(\pp_{j}\psi_{i}-\pp_{i}\psi_{j}=\pp_{j}\psi_{i}-\pp_{i}\psi_{j}\) and in general is the statement that antisymmetrizing over \(k+1\) slots may be performed by holding the first slot and antisymmetrizing the \(k\) transpositions that move \(j\) into it. Adding Equations (13.249) and (13.250) through Equation (13.251), the two terms carrying \(\pp_{[i_{1}}\psi_{|j|\ldots]}\) cancel and
By the chain rule the bracket is exactly \(\dv{}{t}\left[t^{k}\psi_{i_{1}\ldots i_{k}}(tx)\right]\), so the integral is \(\left[t^{k}\psi_{i_{1}\ldots i_{k}}(tx)\right]_{0}^{1} = \psi_{i_{1}\ldots i_{k}}(x)\), the lower limit vanishing because \(k\ge1\). That is Equation (13.248).
∎On \(U=\R^{2}\setminus\set{0}\), which is not star-shaped, put
It is closed: with \(\alpha_{1}=-y/(x^{2}+y^{2})\) and \(\alpha_{2}=x/(x^{2}+y^{2})\), both \(\pp_{1}\alpha_{2}\) and \(\pp_{2}\alpha_{1}\) equal \(\left(y^{2}-x^{2}\right)/ \left(x^{2}+y^{2}\right)^{2}\), so \(\dd\alpha=0\). It is not exact: on the unit circle \(x=\cos s\), \(y=\sin s\) one computes \(\oint\alpha = \int_{0}^{2\pi}\dd s = 2\pi\), whereas an exact form \(\dd f\) integrates to zero around any closed curve, since \(\oint\dd f = \int_{0}^{2\pi}\dv{}{s}f\left(\gamma(s)\right)\dd s = f(\gamma(2\pi))-f(\gamma(0)) = 0\) by the fundamental theorem of calculus (Theorem 7.43). So the converse of Lemma 13.107 genuinely fails, and what fails is topological: \(\alpha\) is exact on any star-shaped subset of \(U\) — there it is \(\dd\) of a branch of the polar angle — but no branch is single-valued on the whole punctured plane. Rests on Theorem 13.109 and Definition 13.106.
The symmetric-tensor analogue of Theorem 13.109 is the result that decides which strain fields of a continuum are realisable by an actual displacement, and it follows from the antisymmetric one by applying it twice.
For a \(C^{2}\) symmetric field \(e_{ij}=e_{ji}\) on an open set of \(\R^{n}\), put
It is antisymmetric under \(j\leftrightarrow l\) and under \(k\leftrightarrow m\) and symmetric under the exchange of the two pairs — the index symmetries of the curvature tensor (Proposition 13.154), which is what it is: the linearisation of Equation (13.95) about a flat metric in the perturbation \(e\). The equations \(\mathcal{R}(e)=0\) are the Saint-Venant compatibility conditions. Rests on Definition 13.91 and Proposition 13.154.
Let \(U\subseteq\R^{n}\) be open and star-shaped and let \(e_{ij}\) be a symmetric field of class \(C^{2}\) on \(U\). Then
The field \(u\) is unique up to \(u_{j}\mapsto u_{j}+a_{j} +\Omega_{jk}x^{k}\) with \(a\) constant and \(\Omega\) a constant antisymmetric array — an infinitesimal rigid motion. Rests on Theorem 13.109, Definition 13.111 and Proposition 7.73.
Derives Theorem 13.112. Necessity. Substituting \(e_{jm}=\tfrac{1}{2}\left(\pp_{j}u_{m}+\pp_{m}u_{j}\right)\) into Equation (13.253) and commuting partial derivatives freely (Proposition 7.73), the eight third derivatives cancel in four pairs: \(\pp_{l}\pp_{k}\pp_{j}u_{m}\) against \(\pp_{j}\pp_{k}\pp_{l}u_{m}\), \(\pp_{l}\pp_{k}\pp_{m}u_{j}\) against \(\pp_{l}\pp_{m}\pp_{k}u_{j}\), \(\pp_{j}\pp_{m}\pp_{l}u_{k}\) against \(\pp_{l}\pp_{m}\pp_{j}u_{k}\), and \(\pp_{j}\pp_{m}\pp_{k}u_{l}\) against \(\pp_{j}\pp_{k}\pp_{m}u_{l}\).
Sufficiency, first application. For each fixed pair \((j,l)\) consider the one-form on \(U\)
antisymmetric in \((j,l)\). Its exterior derivative has components \(\pp_{m}\left(\pp_{l}e_{jk}-\pp_{j}e_{lk}\right) -\pp_{k}\left(\pp_{l}e_{jm}-\pp_{j}e_{lm}\right) = -\mathcal{R}_{jlkm}(e)\), which vanishes by hypothesis. By Theorem 13.109 there are functions \(w_{jl}\) with \(\pp_{k}w_{jl}=\pp_{l}e_{jk}-\pp_{j}e_{lk}\); since the right-hand side changes sign under \(j\leftrightarrow l\), the function \(w_{jl}+w_{lj}\) has vanishing gradient and is therefore a constant on the connected set \(U\) (Corollary 7.36), which may be subtracted, so \(w\) may be taken antisymmetric.
Second application. For each fixed \(j\) consider \(\beta^{(j)}=\left(e_{jk}+w_{jk}\right)\dd x^{k}\). Its exterior derivative has components
by the symmetry of \(e\). So \(\beta^{(j)}=\dd u_{j}\) for some function \(u_{j}\), again by Theorem 13.109; that is \(\pp_{k}u_{j}=e_{jk}+w_{jk}\), and symmetrizing in \((j,k)\) kills \(w\) and returns \(e_{jk}=\tfrac{1}{2}(\pp_{j}u_{k}+\pp_{k}u_{j})\). The field \(u\) is \(C^{3}\) because its first derivatives are \(C^{2}\).
Uniqueness. If two solutions differ by \(u\) with \(\pp_{j}u_{k}+\pp_{k}u_{j}=0\), then
so all second derivatives vanish and \(u\) is affine, \(u_{j}=a_{j} +\Omega_{jk}x^{k}\); the condition then forces \(\Omega_{jk}+\Omega_{kj}=0\).
∎If \(\psi\) is a closed form and \(\varphi\) is an exact form, then \(\psi\wedge\varphi\) is a closed form. Rests on Definition 13.105, Definition 13.106, Lemma 13.107 and Equation (13.246).
Derivation. Derives Proposition 13.113. Let \(\psi\) be a closed \(k\)-form and \(\varphi\) an exact \(l\)-form. By the Poincaré lemma, \(\varphi\) is moreover closed; then by the Leibniz rule
vanishes.
∎The volume form
We define the volume form as
where \(g\) is the determinant of the metric tensor (Section 13.7). Rests on Definitions 13.102 and 13.117.
\(\eta\) is in fact an \(n\)-form, since
where \(\delta^{1\ldots n}_{i_1\ldots i_n}\) is the generalised Kronecker delta and \(\varepsilon_{i_1\ldots i_n}\) the Levi-Civita symbol. Hence \(\eta\) is an \(n\)-form with components
For a metric of signature \((p,q)\) with \(D=p+q\) the determinant \(g\) carries the sign \((-1)^{q}\), and the volume form is written with \(\sqrt{\abs{g}}\) in place of \(\sqrt{g}\); the source treats the Riemannian case \(q=0\).
The Hodge dual
Consider a differentiable manifold \(M_{n}\). Recall that at each point \(P\) of the manifold there exist a tangent space \(T_{M_n}(P)\) and a cotangent space \(T^{*}_{M_n}(P)\), which are \(n\)-dimensional vector spaces, and that the set of all \(k\)-forms is a vector space \(\Lambda^{k}_{M_n}(P)\) with (Proposition 13.100)
and likewise for the set of all \(k\)-vectors, \({}^{*}\Lambda^{k}_{M_n}(P)\). Owing to the fact that
we can associate four \(\binom{n}{k}\)-dimensional vector spaces with each point of the manifold: the space of \(k\)-forms, the space of \((n-k)\)-forms, the space of \(k\)-vectors, and the space of \((n-k)\)-vectors. Recalling that two vector spaces of the same dimension are isomorphic, the four spaces mentioned are isomorphic; that is, we can guarantee that isomorphisms exist among them.
Which are the isomorphisms that move us between these spaces? The answer to this question is not in itself essential: we define the maps and prove their quality as isomorphisms when necessary.
Let \(T\) be a \(k\)-vector with components \(T^{i_1\ldots i_k}\). We define the Hodge operator (duality operator, or star operator) as
We see that the change of components is given by
The metric tensor
A manifold, as defined so far, knows nearness but not distance: the charts of Section 13.4.1 let us differentiate, but nothing yet measures lengths, areas, volumes, or angles. As the source manuscript asks: the ambient \(\R^{N}\) of an embedded manifold carries the Euclidean metric, but can that law serve an arbitrary manifold? It cannot — a vector on a manifold lives in the tangent space at its point (Section 13.5.1), and each point carries its own. The remedy is a field that equips every tangent space with its own scalar product.
A metric on a \(D\)-dimensional manifold \(M\), \(D = p + q\), is a \((0,2)\) tensor field \(g = g_{\mu\nu}\,\dd x^{\mu}\otimes\dd x^{\nu}\) that is symmetric, \(g_{\mu\nu} = g_{\nu\mu}\), and nondegenerate, \(\det g_{\mu\nu} \neq 0\), and whose matrix at every point has \(p\) positive and \(q\) negative eigenvalues (signature \((p,q)\); the counts are point-independent on a connected manifold by continuity of the eigenvalues). The inverse metric \(g^{\mu\nu}\) is defined by \(g^{\mu\lambda}g_{\lambda\nu} = \delta^{\mu}_{\ \nu}\). The pair \((M,g)\) is a pseudo-Riemannian manifold (Riemannian iff \(q = 0\)). Rests on Definitions 13.48 and 13.84.
The invariant line element of a displacement \(\dd x^{\mu}\) is
It is a scalar under general coordinate transformations (Section 13.5.2), because the transformation Jacobians of \(g_{\mu\nu}\) and of the two \(\dd x\)'s cancel. For \(q \neq 0\) the line element is not positive definite: displacements split into those with \(\dd s^{2} > 0\), \(\dd s^{2} = 0\), and \(\dd s^{2} < 0\), the geometric germ of the causal structure of Part IV — Special Relativity. Rests on Definition 13.117.
The metric and its inverse convert between the tangent and cotangent spaces — they lower and raise indices, \(V_{\mu} = g_{\mu\nu}V^{\nu}\), \(\omega^{\mu} = g^{\mu\nu}\omega_{\nu}\) — and define the invariant volume element
whose invariance follows from the Jacobian transformation of \(\det g\) (Section 13.5.3: \(\sqrt{\abs{g}}\) is a scalar density of weight one).
Let \(\Sigma \subset M\) be a submanifold described by the embedding \(x^{\mu}(\xi^{i})\). The induced metric on \(\Sigma\) is the pullback of \(g\),
generalizing the first fundamental form of surface theory (Section 13.3.2), which is exactly Equation (13.261) for a surface in Euclidean \(\R^{3}\). Rests on Definition 13.117.
Isometries
A diffeomorphism \(f : M \longrightarrow M\) is an isometry iff the metric at \(P\) has the same value as the metric at \(f(P)\), compared by pulling back:
Rests on Definition 13.117.
The set of all isometries of \((M,g)\) forms a group under composition, denoted \(\operatorname{Isom}(M,g)\). Rests on Definition 13.120.
Derives Proposition 13.121. Closure: for isometries \(f\) and \(h\), the pullback of a composition is the composition of pullbacks in reverse order (Section 13.5.1), so
Associativity is inherited from composition of maps. The identity \(\id_{M}\) is an isometry, \(\id_{M}^{\ast}g_{P} = g_{P}\). Inverses: an isometry is a diffeomorphism, so \(f^{-1}\) exists, and applying \((f^{-1})^{\ast}\) to both sides of Equation (13.262) at the point \(f^{-1}(Q)\) gives \((f^{-1})^{\ast}g_{f^{-1}(Q)} = g_{Q}\), i.e. \(f^{-1}\) is an isometry.
∎The infinitesimal version of Equation (13.262) — the Killing equation — is derived in Section 13.9.2, and the manifolds whose isometry groups are as large as possible are classified in Section 13.13: they are precisely the flat, anti-de Sitter, and de Sitter geometries of Table 13.1.
Vielbein and local frames
The coordinate basis \(\pp_{\mu}\) diagonalizes differentiation, not the metric. For the physics of Part V — General Relativity and Cosmology — and for spinors, which transform under no representation of the diffeomorphism group — one needs at each point a basis in which the metric is the flat \(\eta_{ab}\) of Equation (13.1). Indices \(a,b,\ldots\) follow the conventions of Section 13.1.
A vielbein (frame field) is a set of \(D\) one-forms \(e^{a} = e^{a}{}_{\mu}\,\dd x^{\mu}\) such that
The dual frame vectors \(e_{a} = e_{a}{}^{\mu}\pp_{\mu}\) satisfy \(e^{a}{}_{\mu}e_{b}{}^{\mu} = \delta^{a}_{\ b}\) and \(e^{a}{}_{\mu}e_{a}{}^{\nu} = \delta_{\mu}^{\ \nu}\); frame indices are raised and lowered with \(\eta_{ab}\), coordinate indices with \(g_{\mu\nu}\), and the vielbein converts one type into the other, \(V^{a} = e^{a}{}_{\mu}V^{\mu}\). Taking determinants of Equation (13.264), \(e \equiv \det e^{a}{}_{\mu} = \sqrt{\abs{g}}\) (up to orientation), so the volume element Equation (13.260) is \(e\,\dd^{D}x\). Rests on Definition 13.117 and Equation (13.1).
The vielbein of a given metric is unique up to
a local (point-dependent) transformation with values in the Lorentz group \(\operatorname{L}_{p+q} = \SO(p,q)\) of Table 13.1. Rests on Definition 13.122.
Derives Proposition 13.123. If \(e\) and \(\tilde e\) both satisfy Equation (13.264), then \(\Lambda^{a}{}_{b} \equiv \tilde e^{a}{}_{\mu} e_{b}{}^{\mu}\) maps one to the other, and substituting into Equation (13.264) for both sides forces \(\eta_{ab}\Lambda^{a}{}_{c}\Lambda^{b}{}_{d} = \eta_{cd}\), which is the defining condition of \(\Ogrp(p,q)\); continuity and choice of orientation restrict to \(\SO(p,q)\). Conversely any such \(\Lambda(x)\) preserves Equation (13.264).
∎Geometry in the frame formulation therefore has two covariances, acting on different indices per Section 13.1: diffeomorphisms on \(\mu,\nu,\ldots\) and local Lorentz transformations on \(a,b,\ldots\) — the structure that makes gravity resemble a gauge theory (Section 13.12).
Unlike the coordinate basis, the frame \(e_{a}\) is in general anholonomic: its Lie brackets (Section 13.9) do not vanish,
and the \(C^{c}{}_{ab}\) measure the failure of the frame to come from coordinates.
Flows, the Lie derivative, and Killing vectors
A vector field \(\vect{\xi} = \xi^{\mu}\pp_{\mu}\) generates a flow: the one-parameter family of diffeomorphisms \(\phi_{t}\) obtained by sliding every point along the integral curves \(\dv{x^{\mu}}{t} = \xi^{\mu}(x)\). The Lie derivative measures the rate of change of a tensor field as it is dragged by this flow, comparing the dragged tensor with the actual one:
It requires no metric and no connection — only the differentiable structure.
That opening sentence asserts something, and the assertion is a theorem: that a smooth vector field has a flow, that the flow is unique, that it is smooth in the point as well as in the parameter, and that each \(\phi_{t}\) is a diffeomorphism. Everything in this section, and the Frobenius theorem of Section 13.9.1 below, rests on it.
An integral curve of a smooth vector field \(X\) on \(M_{n}\) is a smooth curve \(\gamma\) with \(\dot\gamma(t) = X\left(\gamma(t)\right)\) for every \(t\) in its domain. \(X\) is complete if and only if through every point there is an integral curve defined on all of \(\R\). Rests on Definitions 13.71 and 13.82.
Let \(X\) be a smooth vector field on \(M_{n}\) and \(P\in M_{n}\). Then there are an open set \(U\ni P\), a number \(\varepsilon>0\) and a smooth map
such that each \(\phi_{t}\) is a diffeomorphism of \(U\) onto an open subset of \(M_{n}\). Two integral curves of \(X\) that agree at one parameter value agree on the intersection of their domains; consequently
wherever both sides are defined, and \(\phi_{t}^{-1} = \phi_{-t}\). If \(X\) is complete, \(\phi\) is defined on all of \(\R\times M_{n}\) and \(t\mapsto\phi_{t}\) is a one-parameter group of diffeomorphisms of \(M_{n}\). Rests on Definition 13.124, Theorem 9.8 and Theorem A.74.
Derives Theorem 13.125. Work in a chart around \(P\), in which \(X\) becomes a smooth map \(X: V\longrightarrow\R^{n}\) on an open \(V\subseteq\R^{n}\) and \(P\) becomes \(x_{0}\). Smoothness makes \(X\) bounded and Lipschitz on any closed box inside \(V\) (Theorem 7.24 and Theorem 7.35 applied componentwise), so Theorem 9.8 gives a unique integral curve \(c\) with \(c(0)=x_{0}\) on some interval, and we fix \(T>0\) with \(c\left(\left[-T,T\right]\right)\subset V\).
Reduction to a field that vanishes at a point. For \(s\in[0,1]\) and \(z\) in the open set \(O = \set{z\in\R^{n}\mid z + c(Ts)\in V\ \text{for all}\ s\in[0,1]}\), which contains \(0\) because \(c([0,T])\) is compact and \(V\) open, put
This is smooth in \((s,z)\) and \(Z_{s}(0)=0\) for every \(s\), which is exactly the hypothesis of Theorem A.74. That theorem supplies a ball \(B\ni0\) and a smooth \(\zeta:[0,1]\times B\longrightarrow O\) with \(\pp_{s}\zeta_{s}(z) = Z_{s}\left(\zeta_{s}(z)\right)\), \(\zeta_{0}=\id\), each \(\zeta_{s}\) a diffeomorphism onto an open set. Define, for \(x\in U:=x_{0}+B\) and \(t\in[0,T]\),
Then \(\phi_{0}(x) = (x-x_{0}) + x_{0} = x\), and by Equation (13.270)
which is Equation (13.268). Each \(\phi_{t}\) is a diffeomorphism onto an open set, being \(\zeta_{t/T}\) composed with two translations, and \(\phi\) is smooth in \((t,x)\) jointly because \(\zeta\) and \(c\) are. Negative \(t\) is covered by running the same construction for \(-X\), whose integral curves are those of \(X\) traversed backwards; taking \(\varepsilon\) smaller than both times obtained gives Equation (13.268) on \(\left]-\varepsilon,\varepsilon\right[\,\times U\).
Uniqueness and the group law. Two integral curves agreeing at \(t_{1}\) satisfy the same initial value problem there, so they agree near \(t_{1}\) by Theorem 9.8; the set where they agree is therefore open, and it is closed by continuity, so it is the whole of the (connected) intersection of their domains. Now fix \(s\) and \(x\): both \(t\mapsto\phi_{t+s}(x)\) and \(t\mapsto\phi_{t}\left(\phi_{s}(x)\right)\) are integral curves of \(X\) taking the value \(\phi_{s}(x)\) at \(t=0\), whence Equation (13.269); putting \(s=-t\) gives \(\phi_{t}\circ\phi_{-t}=\id\). Completeness removes the restriction on the interval, and Equation (13.269) then says that \(t\mapsto\phi_{t}\) is a homomorphism of \((\R,+)\) into the diffeomorphisms of \(M_{n}\).
∎Only two things: the Picard–Lindelöf theorem of Ordinary Differential Equations and Sturm–Liouville Theory, for existence and uniqueness, and the flow theorem proved in the appendix for Moser's trick (Theorem A.74), for smooth dependence on the initial point. The latter is stated there for a time-dependent field vanishing at a point, which looks like a special case and is not: subtracting the value of \(X\) along the integral curve through \(P\), as Equation (13.270) does, turns any smooth field into one of that kind, at the cost of making it time-dependent — which the appendix theorem already allows. No fixed-point theorem and no compactness argument beyond Theorem 7.24 enters.
For a scalar, a vector, a covector, and a \((0,2)\) tensor,
with one drag term per index in general. All partial derivatives may be replaced by covariant derivatives of any torsion-free connection: the connection terms cancel pairwise. Rests on Equation (13.267), Definition 13.89 and Definition 13.90.
Derives Proposition 13.127. Expand \(\phi_{t}\) to first order, \(x'^{\mu} = x^{\mu} + t\,\xi^{\mu}\), in the transformation law of each tensor type (Section 13.5.2), and differentiate at \(t = 0\): the Jacobian \(\pp x'^{\mu}/\pp x^{\nu} = \delta^{\mu}_{\ \nu} + t\,\pp_{\nu}\xi^{\mu}\) contributes the drag terms with the signs shown (one \(+\) per lower index, one \(-\) per upper index), and the argument shift contributes the transport term \(\xi\cdot\pp\). The cancellation of connection terms is a direct check using the symmetry \(\Gamma^{\lambda}_{\ \mu\nu} = \Gamma^{\lambda}_{\ \nu\mu}\) of a torsion-free connection.
∎On differential forms (Section 13.6) the Lie derivative has a purely exterior expression:
On any \(k\)-form \(\alpha\),
where \(i_{\xi}\) is the interior product, \((i_{\xi}\alpha)_{\mu_{2}\ldots\mu_{k}} = \xi^{\mu_{1}}\alpha_{\mu_{1}\mu_{2}\ldots\mu_{k}}\). Rests on Proposition 13.127, Definition 13.103, Equation (13.245) and Equation (13.246).
Derives Proposition 13.128. On functions, \(i_{\xi}f \equiv 0\) and \(i_{\xi}\dd f = \xi^{\mu}\pp_{\mu}f = \mathcal{L}_{\xi}f\). On the coordinate differentials, \(\mathcal{L}_{\xi}\dd x^{\mu} = \dd\xi^{\mu} = \dd(i_{\xi}\dd x^{\mu}) + i_{\xi}\dd(\dd x^{\mu})\), using \(\dd^{2} = 0\) (Section 13.6.4). Both sides of Equation (13.276) are derivations of the exterior algebra (they obey the graded Leibniz rule — for the right side this follows from the graded Leibniz rules of \(\dd\) and \(i_{\xi}\)), and a derivation is fixed by its action on functions and their differentials, which generate all forms locally. The two sides agree there, hence everywhere.
∎The three operators \(\mathcal{L}\), \(\dd\) and \(i\) close on one another, and Equation (13.276) is only the first of the identities they satisfy. The second is the one that lets a computation move an interior product past a Lie derivative, which is what makes many arguments with forms one line long instead of a component calculation.
For any smooth vector fields \(X,Y\) and any \(k\)-form \(\alpha\),
Rests on Proposition 13.128, Proposition 13.127 and Definition 13.98.
Derives Proposition 13.129. Both sides are \((k-1)\)-forms; compare components. Write \(\beta = i_{Y}\alpha\), so \(\beta_{\mu_{2}\ldots\mu_{k}} = Y^{\nu}\alpha_{\nu\mu_{2}\ldots\mu_{k}}\) by the convention of Proposition 13.128, and use the general component formula of Proposition 13.127 — transport term plus one drag term per lower index:
where in each sum \(\lambda\) sits in the \(i\)-th slot. Subtracting, the transport terms of \(\alpha\) and all the drag terms carried by \(\mu_{2},\ldots,\mu_{k}\) cancel, leaving
by Equation (13.273), and that is \(\left(i_{\comm{X}{Y}}\alpha\right)_{\mu_{2}\ldots\mu_{k}}\).
∎Distributions and the Frobenius theorem
A single vector field, by Theorem 13.125, threads the manifold with curves: through every point passes one integral curve, and the curves fill the manifold. The question of this subsection is what happens when a \(k\)-dimensional field of tangent planes is prescribed instead of a field of tangent lines. Does a \(k\)-dimensional surface pass through each point with the prescribed tangent plane everywhere? For \(k=1\) the answer is always yes; for \(k\ge2\) it is yes precisely when an algebraic condition on brackets holds, and that is the Frobenius theorem. It is the theorem behind the symplectic foliation of a Poisson manifold, behind the counting of gauge degrees of freedom generated by first-class constraints, and behind the very meaning of the word nonholonomic: a constraint is nonholonomic exactly when the distribution it defines fails the condition below.
A distribution of rank \(k\) on \(M_{n}\) is an assignment \(P\mapsto D_{P}\subseteq T_{P}M_{n}\) of a \(k\)-dimensional subspace to each point which is smooth: every point has a neighbourhood on which there are smooth vector fields \(X_{1},\ldots,X_{k}\) spanning \(D_{Q}\) at each \(Q\) of it. A vector field \(X\) belongs to \(D\) iff \(X_{Q}\in D_{Q}\) everywhere. The distribution is
-
involutive iff \(\comm{X}{Y}\) belongs to \(D\) whenever \(X\) and \(Y\) do;
-
integrable iff through each \(P\in M_{n}\) there passes an integral manifold: an embedded \(k\)-dimensional submanifold \(S\ni P\) (Definition 13.54) with \(T_{Q}S = D_{Q}\) for every \(Q\in S\).
Rests on Definition 13.82, Definition 13.54 and Equation (13.273).
Let \(X\) and \(Y\) be smooth vector fields with flows \(\phi_{t}\) and \(\eta_{s}\). If \(\comm{X}{Y}=0\) then \(\left(\phi_{t}\right)_{*}Y = Y\) and
wherever both sides are defined. Rests on Theorem 13.125, Equation (13.267) and Equation (13.273).
Derives Proposition 13.131. Fix \(x\) and put \(W(t) = \left(\phi_{-t}\right)_{*}Y_{\phi_{t}(x)}\in T_{x}M_{n}\), a smooth curve in one fixed vector space. By the group law Equation (13.269), for any \(t_{0}\),
so differentiating at \(h=0\) and using the definition Equation (13.267) of the Lie derivative together with Equation (13.273), \(\dv{W}{t}(t_{0}) = \left(\phi_{-t_{0}}\right)_{*} \left(\mathcal{L}_{X}Y\right)_{y} = \left(\phi_{-t_{0}}\right)_{*}\comm{X}{Y}_{y} = 0\). Hence \(W(t)=W(0)=Y_{x}\), i.e. \(\left(\phi_{t}\right)_{*}Y_{x} = Y_{\phi_{t}(x)}\) for every \(t\) and \(x\).
A diffeomorphism that carries \(Y\) to \(Y\) carries integral curves of \(Y\) to integral curves of \(Y\): if \(\dot\gamma = Y\circ\gamma\) then \(\dv{}{s}\left(\phi_{t}\circ\gamma\right) = \left(\phi_{t}\right)_{*}Y_{\gamma} = Y_{\phi_{t}(\gamma)}\). So \(s\mapsto\phi_{t}(\eta_{s}(x))\) is the integral curve of \(Y\) through \(\phi_{t}(x)\), which is \(s\mapsto\eta_{s}(\phi_{t}(x))\) by the uniqueness clause of Theorem 13.125. That is Equation (13.278).
∎Let \(Y_{1},\ldots,Y_{k}\) be smooth vector fields on \(M_{n}\), pairwise commuting, \(\comm{Y_{a}}{Y_{b}}=0\), and linearly independent at \(P\). Then there is a chart \((u^{1},\ldots,u^{n})\) around \(P\) in which
Rests on Proposition 13.131, Theorem 13.125 and Corollary A.292.
Derives Proposition 13.132. Take any chart around \(P\) sending \(P\) to the origin of \(\R^{n}\), and follow it by the linear change of coordinates that sends \(Y_{1}(P),\ldots,Y_{k}(P)\) — independent by hypothesis, hence extendable to a basis — to the first \(k\) coordinate vectors. Write \(\phi^{a}\) for the flow of \(Y_{a}\) (Theorem 13.125) and, for \(u\) in a small neighbourhood of \(0\) in \(\R^{n}\), set
the argument on the right being the point of the slice \(u^{1}=\cdots=u^{k}=0\). This is defined and smooth near \(0\), each flow being smooth in both arguments.
Its differential at \(0\) is invertible: all parameters vanish there, so \(\pp\Psi/\pp u^{a}\big|_{0} = Y_{a}(P)\) for \(a\le k\), while \(\pp\Psi/\pp u^{j}\big|_{0} = \pp/\pp x^{j}\big|_{P}\) for \(j>k\), and these \(n\) vectors are a basis by the choice of the initial chart. By the inverse function theorem (Corollary A.292) \(\Psi\) is a diffeomorphism of a neighbourhood of \(0\) onto a neighbourhood of \(P\), and we take \(\Psi^{-1}\) as the chart.
Finally, fix \(a\). Because the flows commute (Proposition 13.131), the composition in Equation (13.280) may be reordered to put \(\phi^{a}\) outermost, and then \(\pp\Psi/\pp u^{a}\) is the velocity of the integral curve of \(Y_{a}\) through \(\Psi(u)\), that is \(Y_{a}(\Psi(u))\). Hence \(\Psi_{*}\left(\pp/\pp u^{a}\right) = Y_{a}\), which is Equation (13.279).
∎A smooth distribution \(D\) of constant rank \(k\) on \(M_{n}\) is integrable if and only if it is involutive. When it is, every point has a chart \((u^{1},\ldots,u^{n})\) in which
so that the slices \(u^{k+1}=\text{const},\ldots,u^{n}=\text{const}\) are integral manifolds: the chart domain is foliated by them, and every connected integral manifold inside it lies in one slice. Rests on Definition 13.130, Proposition 13.132 and Equation (13.273).
Derives Theorem 13.133. Integrable implies involutive. Let \(X,Y\) belong to \(D\) and let \(P\in M_{n}\) lie on an integral manifold \(S\). Both fields are tangent to \(S\), hence restrict to vector fields on \(S\); the bracket Equation (13.273) of two fields tangent to an embedded submanifold is again tangent to it, since in a chart adapted to \(S\) both fields have vanishing components transverse to \(S\) and Equation (13.273) differentiates those components only along directions inside \(S\). So \(\comm{X}{Y}_{P}\in T_{P}S = D_{P}\).
Involutive implies integrable. Fix \(P\) and a local spanning frame \(X_{1},\ldots,X_{k}\) of \(D\) (Definition 13.130). Choose a chart around \(P\) with coordinates \((x^{1},\ldots,x^{n})\) such that \(D_{P}\) is transverse to the span of \(\pp/\pp x^{k+1},\ldots,\pp/\pp x^{n}\); equivalently, writing \(X_{a} = X_{a}^{\ i}\pp/\pp x^{i}\), the \(k\times k\) block \(A^{\ b}_{a} = X^{\ b}_{a}\) with \(b\le k\) is invertible at \(P\), hence on a neighbourhood. Replace the frame by
still a frame for \(D\), since the change is by an invertible matrix of functions.
Now compute \(\comm{Y_{a}}{Y_{b}}\) from Equation (13.273). The components of \(Y_{a}\) along \(\pp/\pp x^{c}\) with \(c\le k\) are the constants \(\delta^{c}_{a}\), so every derivative of them vanishes and the bracket has no component along \(\pp/\pp x^{1},\ldots,\pp/\pp x^{k}\):
By involutivity the left side belongs to \(D\), so it is a combination \(\lambda^{a}Y_{a}\); comparing the components along \(\pp/\pp x^{c}\), \(c\le k\), which are \(\lambda^{c}\) on the right and zero on the left by Equation (13.283), gives \(\lambda^{c}=0\) and hence
The frame \(Y_{1},\ldots,Y_{k}\) is therefore commuting and pointwise independent, and Proposition 13.132 supplies a chart in which \(Y_{a}=\pp/\pp u^{a}\). Since the \(Y_{a}\) span \(D\), this is Equation (13.281), and the slices \(u^{k+1},\ldots,u^{n}=\text{const}\) are embedded \(k\)-dimensional submanifolds whose tangent space at each point is exactly \(D\): integral manifolds through every point of the chart domain.
For the last clause, let \(S\) be a connected integral manifold inside the chart domain. On \(S\) each function \(u^{j}\) with \(j>k\) has vanishing differential, because \(T_{Q}S=D_{Q}\) is spanned by the \(\pp/\pp u^{a}\) with \(a\le k\); a function with vanishing differential on a connected manifold is constant (Corollary 7.36 applied along paths in a chart). So \(S\) lies in one slice.
∎The dual formulation is the one that appears whenever the planes are prescribed by equations rather than by spanning fields — a mechanical constraint written as a set of one-forms annihilating the admissible velocities is exactly that situation.
Let \(D\) have rank \(k\) and let \(\theta^{1},\ldots,\theta^{n-k}\) be pointwise independent one-forms whose common kernel is \(D\), so that a vector \(v\) lies in \(D\) if and only if \(\theta^{\alpha}(v)=0\) for every \(\alpha\). Then \(D\) is integrable if and only if the ideal generated by the \(\theta^{\alpha}\) is closed under exterior differentiation, i.e.\ iff there are one-forms \(\eta^{\alpha}{}_{\beta}\) with
equivalently iff \(\dd\theta^{\alpha}\wedge\theta^{1}\wedge\cdots\wedge\theta^{n-k}=0\) for every \(\alpha\). Rests on Theorem 13.133, Definition 13.103 and Definition 13.130.
Derives Corollary 13.134. For a one-form \(\theta=\theta_{\mu}\dd x^{\mu}\), the components of \(\dd\theta\) are \(\pp_{\mu}\theta_{\nu}-\pp_{\nu}\theta_{\mu}\) (Definition 13.103), and for any vector fields \(X,Y\)
as expanding the three terms on the right with Equation (13.273) shows: the terms carrying a derivative of \(X\) or of \(Y\) cancel in pairs and only the derivatives of \(\theta\) survive. If \(X\) and \(Y\) belong to \(D\), the first two terms on the right of Equation (13.286) vanish, so
Involutivity says the right-hand side vanishes for every \(\alpha\) and every \(X,Y\) in \(D\) — because the \(\theta^{\alpha}\) detect exactly membership of \(D\) — and that is the statement that each \(\dd\theta^{\alpha}\) annihilates \(D\times D\). Completing the \(\theta^{\alpha}\) to a coframe by one-forms \(\sigma^{a}\) dual to a frame of \(D\), a two-form annihilating \(D\times D\) has no \(\sigma^{a}\wedge\sigma^{b}\) term, i.e. every term carries at least one factor \(\theta^{\beta}\), which is Equation (13.285); and a two-form of that shape wedges to zero against \(\theta^{1}\wedge\cdots\wedge\theta^{n-k}\), while conversely a \(\sigma^{a}\wedge\sigma^{b}\) term would survive that wedge. The three conditions are therefore equivalent, and by Theorem 13.133 each is equivalent to integrability.
∎Theorem 13.133 assumes the rank of \(D\) is the same at every point, and the hypothesis is doing work: Equation (13.282) needs a local frame of fixed size, and Equation (13.281) asserts a fixed number of slice coordinates. Distributions whose rank jumps do occur — a bivector field may be invertible on the complement of a point and vanish at it, so that the distribution it spans has rank \(2\) almost everywhere and rank \(0\) there — and at such a point the conclusion above says nothing whatever, not even that a leaf of some smaller dimension passes through. This treatise establishes no generalisation covering that case, and nothing in it rests on one; the limitation is recorded here so that an appeal to Theorem 13.133 elsewhere in the book can be checked against its hypothesis rather than against its name.
Let \(\Gamma\) be a subgroup of \(\left(\R^{f},+\right)\) which is discrete, i.e. \(\Gamma\cap B_{r}(0)=\set{0}\) for some \(r>0\). Then there are \(m\le f\) linearly independent vectors \(w_{1},\ldots,w_{m}\in\Gamma\) with
and \(m\) is the dimension of the linear span of \(\Gamma\). Rests on Definitions 5.13 and 6.9.
Derives Lemma 13.136. First, a bounded set meets \(\Gamma\) in finitely many points: two distinct elements of \(\Gamma\) differ by a nonzero element of \(\Gamma\), hence by at least \(r\), so a ball of radius \(R\) can contain at most finitely many of them.
Induct on \(m=\dim\operatorname{span}\Gamma\). For \(m=0\), \(\Gamma=\set{0}\). Let \(m\ge1\) and choose linearly independent \(v_{1},\ldots,v_{m}\in\Gamma\) spanning that span. Put \(W=\operatorname{span}\set{v_{1},\ldots,v_{m-1}}\). The subgroup \(\Gamma'=\Gamma\cap W\) is discrete with span exactly \(W\), so by the induction hypothesis \(\Gamma'=\Z w_{1}\oplus\cdots\oplus\Z w_{m-1}\) with the \(w_{i}\) linearly independent, hence a basis of \(W\).
Every \(x\in\Gamma\) decomposes uniquely as \(x=\sum_{i<m}a_{i}w_{i}+t\,v_{m}\) with \(a_{i},t\) real. Let
a bounded set, hence finite, and non-empty: subtracting from \(v_{m}\) suitable integer multiples of the \(w_{i}\) — which are in \(\Gamma\) — brings its \(a_{i}\) into \([0,1[\) without changing \(t=1\). Choose \(w_{m}\in S\) with \(t=t_{0}\) minimal.
Let \(x\in\Gamma\) have \(v_{m}\)-coefficient \(t\), put \(n=\lfloor t/t_{0} \rfloor\) and consider \(x-n\,w_{m}\in\Gamma\), whose \(v_{m}\)-coefficient is \(t-nt_{0}\in\left[0,t_{0}\right[\). If it were positive, subtracting integer multiples of the \(w_{i}\) would produce an element of \(S\) with a \(v_{m}\)-coefficient smaller than \(t_{0}\), contradicting minimality. So \(t=nt_{0}\) and \(x-n\,w_{m}\in\Gamma\cap W=\Gamma'\). Hence \(\Gamma=\Gamma'+\Z w_{m}\), the sum is direct because \(w_{m}\notin W\), and \(w_{1},\ldots,w_{m}\) are linearly independent. Finally \(m\le f\), being the number of independent vectors in \(\R^{f}\).
∎Let \(M_{f}\) be a compact connected manifold of dimension \(f\) carrying \(f\) complete, pairwise commuting vector fields \(Y_{1},\ldots,Y_{f}\) that are linearly independent at every point. Then the map
is a smooth action of the abelian group \(\left(\R^{f},+\right)\) which is transitive, and \(M_{f}\) is diffeomorphic to the torus \(T^{f}=\R^{f}/\Z^{f}\). Rests on Proposition 13.131, Proposition 13.132 and Theorem 13.125.
Derives Theorem 13.137. Each \(\phi^{a}\) is defined for all parameters by completeness (Definition 13.124 and Theorem 13.125) and the \(f\) flows commute in pairs (Proposition 13.131), so Equation (13.289) does not depend on the order of composition and \(\Theta(t,\Theta(s,x))=\Theta(t+s,x)\): it is a smooth action of \(\R^{f}\).
The orbits are open, hence there is one. Fix \(x\) and apply Proposition 13.132 with \(k=f=n\): near \(x\) there is a chart in which \(Y_{a}=\pp/\pp u^{a}\), and in it \(\Theta(\cdot,x)\) is \(t\mapsto u = t\) near \(t=0\). So the orbit of \(x\) contains a neighbourhood of \(x\); being a union of such, every orbit is open. The orbits partition \(M_{f}\), and \(M_{f}\) is connected, so there is exactly one: the action is transitive. The same computation shows that \(\Theta(\cdot,x)\) is a local diffeomorphism at every \(t\), by the group law.
The stabiliser is a lattice. Let \(\Gamma=\set{t\in\R^{f}\mid\Theta(t,x)=x}\), a subgroup of \(\R^{f}\); since \(\Theta(\cdot,x)\) is injective near \(0\), \(\Gamma\) meets some neighbourhood of \(0\) only in \(0\), so \(\Gamma\) is discrete, and by transitivity \(\Theta(\cdot,x)\) descends to a bijection \(\R^{f}/\Gamma\longrightarrow M_{f}\), a diffeomorphism because \(\Theta(\cdot,x)\) is a local diffeomorphism everywhere. By Lemma 13.136, \(\Gamma=\Z w_{1}\oplus\cdots\oplus\Z w_{m}\) with \(w_{1},\ldots,w_{m}\) linearly independent; completing them to a basis of \(\R^{f}\) and using that basis as coordinates identifies \(\R^{f}/\Gamma\) with \(T^{m}\times\R^{f-m}\). Compactness of \(M_{f}\) forces \(m=f\), since \(\R^{f-m}\) is non-compact for \(m<f\), and the same change of basis carries \(\Gamma\) to \(\Z^{f}\). Hence \(M_{f}\cong\R^{f}/\Z^{f}=T^{f}\).
∎Theorem 13.137 is the differential-topological half of the Liouville–Arnold theorem of Hamiltonian mechanics: an integrable system with \(f\) degrees of freedom supplies, on a compact connected level set of its \(f\) commuting conserved quantities, exactly \(f\) commuting independent vector fields, and the theorem then says the level set is a torus. The symplectic half — that the resulting angle coordinates are canonical — belongs to the mechanics part and is not a statement about manifolds alone.
Killing vectors
\(\vect{\xi}\) is a Killing vector of \((M,g)\) iff its flow consists of isometries, i.e. iff \(\mathcal{L}_{\xi}g = 0\). Rests on Definition 13.120 and Equation (13.267).
With \(\nabla\) the Levi-Civita connection of \(g\) (Theorem 13.150),
A \(D\)-dimensional manifold admits at most \(D(D+1)/2\) linearly independent Killing vectors. Rests on Definition 13.139, Proposition 13.127, Theorem 13.150, Definition 13.149 and Theorem 13.152.
Derives Proposition 13.140. Replace \(\pp\) by \(\nabla\) in Equation (13.275) (Proposition 13.127) and use metricity \(\nabla_{\lambda}g_{\mu\nu} = 0\): the transport term drops and the drag terms become \(\nabla_{\mu}\xi_{\nu} + \nabla_{\nu}\xi_{\mu}\). For the count: differentiating Equation (13.290) and commuting derivatives (Section 13.11) yields \(\nabla_{\mu}\nabla_{\nu}\xi_{\lambda} = R^{\rho}{}_{\mu\nu\lambda}\,\xi_{\rho}\) (torsion-free), so a Killing vector is determined everywhere by the \(D\) values \(\xi_{\mu}(P)\) and the \(D(D-1)/2\) values \(\nabla_{[\mu}\xi_{\nu]}(P)\) at a single point: \(D + D(D-1)/2 = D(D+1)/2\).
∎A Killing vector is worth having because it gives a conserved quantity, and the conservation law — not the symmetry — is what a physical chapter actually uses. It is one line from Killing's equation.
Let \(\vect{\xi}\) be a Killing vector of \((M,g)\) and let \(\gamma\) be an affinely parametrized geodesic of the Levi-Civita connection, with tangent \(u^{\mu}=\dv{x^{\mu}}{t}\) and \(u^{\nu}\nabla_{\nu}u^{\mu}=0\). Then
that is, \(\xi_{\mu}u^{\mu}\) is constant along \(\gamma\). Rests on Proposition 13.140, Definition 13.146 and Definition 13.149.
Derives Proposition 13.141. The derivative along the curve of a scalar is \(u^{\nu}\nabla_{\nu}\) of it, so
The second term vanishes by the geodesic equation (Equation (13.297) with \(\Gamma=\mathring{\Gamma}\)). In the first, \(u^{\nu}u^{\mu}\) is symmetric under \(\mu\leftrightarrow\nu\), so only the symmetric part of \(\nabla_{\nu}\xi_{\mu}\) contributes:
and the bracket vanishes by Killing's equation Equation (13.290).
∎Not every geodesic invariant comes from an isometry. The object that produces the rest is the symmetric-tensor generalisation of a Killing vector.
A Killing tensor of rank \(r\) on \((M,g)\) is a totally symmetric tensor field \(K_{\mu_{1}\ldots\mu_{r}}\) satisfying
the round brackets denoting here the total symmetrization over the \(r+1\) indices they enclose — the average over all \((r+1)!\) orderings of them — which extends to any number of indices the two-index convention Equation (13.213) of Definition 13.91. For \(r=1\), Equation (13.292) is Killing's equation Equation (13.290). Rests on Definitions 13.91, 13.139 and 13.149.
Let \(K\) be a Killing tensor of rank \(r\) and \(\gamma\) an affinely parametrized geodesic with tangent \(u^{\mu}\). Then
Conversely, a function on the tangent bundle that is a homogeneous polynomial of degree \(r\) in \(u\) and constant along every geodesic has totally symmetric coefficients obeying Equation (13.292). Rests on Definition 13.142, Definition 13.146 and Proposition 13.141.
Derives Theorem 13.143. Differentiating along the curve and using \(u^{\nu}\nabla_{\nu}u^{\mu}=0\) for each of the \(r\) factors,
The product \(u^{\lambda}u^{\mu_{1}}\cdots u^{\mu_{r}}\) is totally symmetric in its \(r+1\) indices, so contracting it with any tensor picks out that tensor's totally symmetric part; the right-hand side is therefore \(\left(\nabla_{(\lambda}K_{\mu_{1}\ldots\mu_{r})}\right) u^{\lambda}u^{\mu_{1}}\cdots u^{\mu_{r}}\), which vanishes by Equation (13.292).
For the converse, the same computation shows that the derivative along every geodesic is \(\left(\nabla_{(\lambda}K_{\mu_{1}\ldots\mu_{r})} \right)u^{\lambda}u^{\mu_{1}}\cdots u^{\mu_{r}}\), and through a given point a geodesic may be started with any tangent \(u\) (Equation (13.297) is a second-order equation with arbitrary initial position and velocity, solvable by Theorem 9.8). A totally symmetric tensor whose contraction with \(u^{\otimes(r+1)}\) vanishes for every \(u\) is zero, by polarization: the contraction is a homogeneous polynomial in the components of \(u\) whose coefficients are exactly the independent components of the symmetric tensor.
∎The metric itself is a Killing tensor of rank \(2\): metricity (Definition 13.149) gives \(\nabla_{\lambda}g_{\mu\nu}=0\), so Equation (13.292) holds trivially, and the invariant Equation (13.293) is \(g_{\mu\nu}u^{\mu}u^{\nu}\) — the statement that a geodesic keeps the norm of its tangent, i.e. that timelike, null and spacelike geodesics stay so. Symmetrized products of Killing vectors are Killing tensors too: if \(\xi\) and \(\chi\) are Killing vectors then \(K_{\mu\nu}=\xi_{(\mu}\chi_{\nu)}\) obeys Equation (13.292), since Equation (13.290) makes each \(\nabla_{(\lambda}\xi_{\mu)}\) vanish. A Killing tensor is called irreducible when it is not a combination of the metric and such products; an irreducible one gives a constant of the motion that no isometry accounts for, which is why the bound \(D(D+1)/2\) of Proposition 13.140 does not bound the number of geodesic invariants. Rests on Definition 13.142, Definition 13.149 and Proposition 13.140.
Metrics that saturate this bound are maximally symmetric (Section 13.13). Relaxing isometry to angle preservation defines the conformal Killing vectors, \(\mathcal{L}_{\xi}g = 2\sigma g\) with \(\sigma = \nabla_{\mu}\xi^{\mu}/D\); for flat space of signature \((p,q)\) with \(D > 2\) their maximal number is \((D+1)(D+2)/2\), the dimension of the conformal group \(\operatorname{C}_{p+q} = \SO(p+1,q+1)\) (Section 13.13.1).
Connections, parallel transport, and torsion
In Euclidean geometry a homogeneous (constant) vector field is one whose derivative with respect to position vanishes; its vectors are parallel and of equal length. How does this notion extend to an arbitrary manifold? The partial derivative fails: differentiating the vector transformation law \(V'^{\mu} = (\pp x'^{\mu}/\pp x^{\nu})V^{\nu}\) gives
whose second, inhomogeneous term spoils tensoriality whenever the coordinate change is nonlinear. The covariant derivative is the repair, and it is the concept on which the whole of theoretical physics leans.
An affine connection \(\nabla\) assigns to vector fields the \((1,1)\)-tensor-valued derivative
extended to covectors by \(\nabla_{\mu}\omega_{\nu} = \pp_{\mu}\omega_{\nu} - \Gamma^{\lambda}{}_{\mu\nu}\omega_{\lambda}\) (so that \(\nabla(V^{\mu}\omega_{\mu})\) obeys the Leibniz rule and reduces to \(\pp\)), and to arbitrary \((r,s)\) tensors with one \(+\Gamma\) term per upper and one \(-\Gamma\) term per lower index. The connection coefficients are not tensor components: requiring Equation (13.295) to transform as a tensor forces
the inhomogeneous term cancelling that of Equation (13.294). Rests on Definitions 13.82 and 13.90.
A tensor \(T\) is parallelly transported along a curve with tangent \(u^{\mu} = \dv{x^{\mu}}{t}\) iff \(u^{\mu}\nabla_{\mu}T = 0\). A curve is an autoparallel iff it parallelly transports its own tangent,
generalizing the geodesics of surface theory (Section 13.3.3). Only the symmetric part \(\Gamma^{\lambda}{}_{(\mu\nu)}\) enters Equation (13.297). Rests on Definition 13.145.
The torsion of a connection is
Rests on Definition 13.145.
\(T^{\lambda}{}_{\mu\nu}\) is a tensor, although \(\Gamma^{\lambda}{}_{\mu\nu}\) is not. Rests on Definition 13.147, Equation (13.296), Definition 13.90 and Proposition 7.73.
Derives Proposition 13.148. The inhomogeneous term of Equation (13.296) is symmetric in \(\mu\nu\) (mixed second partials commute, Proposition 7.73), so it cancels in the antisymmetrization Equation (13.298), leaving the homogeneous tensor law.
∎Geometrically, torsion measures the failure of infinitesimal parallelograms to close: transporting \(\dd x_{(1)}\) along \(\dd x_{(2)}\) and vice versa, the two endpoints differ by \(T^{\lambda}{}_{\mu\nu}\,\dd x_{(1)}^{\mu}\dd x_{(2)}^{\nu}\). In Section 13.12 it becomes one of the two field strengths of the geometry.
A connection is metric compatible iff \(\nabla_{\lambda}g_{\mu\nu} = 0\): parallel transport preserves lengths and angles. Rests on Definitions 13.117, 13.145 and 13.146.
On a pseudo-Riemannian manifold there is exactly one metric-compatible connection for each prescribed torsion. Lowering the first index, \(\Gamma_{\lambda\mu\nu} = g_{\lambda\rho}\Gamma^{\rho}{}_{\mu\nu}\),
with \(T_{\lambda\mu\nu} = g_{\lambda\rho}T^{\rho}{}_{\mu\nu}\). The torsion-free case \(K = 0\) is the Levi-Civita connection \(\mathring{\Gamma}\) (the Christoffel symbols, generalizing Section 13.3.2); the contorsion \(K_{\lambda\mu\nu}\) is a tensor, antisymmetric in its outer pair, \(K_{\lambda\mu\nu} = -K_{\nu\mu\lambda}\). Rests on Definitions 13.117, 13.145, 13.147 and 13.149.
Derives Theorem 13.150. Write metricity three times with cyclically permuted indices,
and form the combination (second) \(+\) (third) \(-\) (first). Grouping with the torsion Equation (13.298), \(\Gamma_{\lambda\mu\nu} + \Gamma_{\lambda\nu\mu} = 2\Gamma_{\lambda\mu\nu} - T_{\lambda\mu\nu}\), \(\Gamma_{\nu\mu\lambda} - \Gamma_{\nu\lambda\mu} = T_{\nu\mu\lambda}\), and \(\Gamma_{\mu\nu\lambda} - \Gamma_{\mu\lambda\nu} = T_{\mu\nu\lambda}\), so that
Solving for \(\Gamma_{\lambda\mu\nu}\) and using the antisymmetry \(T_{\nu\mu\lambda} = -T_{\nu\lambda\mu}\), \(T_{\mu\nu\lambda} = -T_{\mu\lambda\nu}\) to reorder gives Equation (13.299). Uniqueness is manifest: the metric determines \(\mathring{\Gamma}\), the prescribed torsion determines \(K\). The outer antisymmetry of \(K\) is a two-line check using \(T_{\lambda\mu\nu} = -T_{\lambda\nu\mu}\), and it is precisely the condition that the contorsion contribution preserve metricity.
∎Curves that extremize the length functional \(\int\sqrt{\abs{\dd s^{2}}}\) obey Equation (13.297) with \(\mathring{\Gamma}\) (the variational computation of Lagrangian Mechanics applied to the line element). Since \(K^{\lambda}{}_{(\mu\nu)} \neq 0\) in general, autoparallels of a torsionful connection differ from geodesics of its metric: straightest and shortest part ways.
The spin connection
Frame-indexed fields need their own connection: local Lorentz transformations Equation (13.265) are point-dependent, so \(\pp_{\mu}V^{a}\) is not covariant. The spin connection \(\omega_{\mu}{}^{a}{}_{b}\) enters as
and the requirement that frame and coordinate descriptions agree — the vielbein postulate, total covariant constancy of the vielbein — ties it to \(\Gamma\):
Metricity of \(\Gamma\) is then equivalent to \(\omega_{\mu ab} = -\omega_{\mu ba}\): the spin connection takes values in the Lorentz algebra \(\mathfrak{so}(p,q)\) (Lie Groups, Lie Algebras, and Fibre Bundles). Under a local Lorentz transformation it transforms inhomogeneously,
exactly like a gauge potential — the observation the Cartan formalism turns into machinery (Section 13.12).
Curvature
In the theory of curves a curve is characterized by its curvatures, built from first and second derivatives (Section 13.3.1); the curvature of a manifold must likewise live in the first and second derivatives of the metric, and vanish for Euclidean space. The correct object measures the failure of parallel transport around a closed loop, i.e. the non-commutativity of covariant derivatives.
For any affine connection and any vector field,
with
\(R^{\lambda}{}_{\rho\mu\nu}\) is a tensor: the Riemann curvature tensor. Rests on Definition 13.145, Definition 13.147 and Proposition 7.73.
Derives Theorem 13.152. Expand \(\nabla_{\mu}\nabla_{\nu}V^{\lambda} = \pp_{\mu}(\nabla_{\nu}V^{\lambda}) + \Gamma^{\lambda}{}_{\mu\sigma}\nabla_{\nu}V^{\sigma} - \Gamma^{\rho}{}_{\mu\nu}\nabla_{\rho}V^{\lambda}\) (the middle slot of \(\nabla_{\nu}V^{\lambda}\) is a covector index) and antisymmetrize in \(\mu\nu\). The \(\pp\pp V\) terms cancel by equality of mixed partials; the terms with one \(\Gamma\) and one \(\pp V\) cancel except for those coming from the last piece, which assemble into \(-T^{\rho}{}_{\mu\nu}\nabla_{\rho}V^{\lambda}\) by Equation (13.298); the remaining terms contain no derivatives of \(V\) and define Equation (13.306). Tensoriality follows because the left side of Equation (13.305) and the torsion term are tensors for arbitrary \(V\).
∎The Ricci tensor and curvature scalar are
For a surface (\(D = 2\), torsion-free) \(R = 2K\) with \(K\) the Gaussian curvature of Section 13.3.3, closing the circle with the classical theory. Rests on Theorem 13.152 and Definition 13.117.
Always: \(R^{\lambda}{}_{\rho\mu\nu} = -R^{\lambda}{}_{\rho\nu\mu}\). With metricity: \(R_{\lambda\rho\mu\nu} = -R_{\rho\lambda\mu\nu}\). For the Levi-Civita connection additionally
(first and second Bianchi identities). With torsion the last two acquire torsion terms, displayed in form language in Proposition 13.157. Rests on Theorem 13.152, Definition 13.149, Theorem 13.150 and Proposition 13.157.
Derives Proposition 13.154. The first is manifest from Equation (13.306). The second: apply Equation (13.305) to the scalar \(g_{\lambda\rho}V^{\lambda}W^{\rho}\) and use metricity. The torsion-free identities follow from Equation (13.306) with \(\Gamma^{\lambda}{}_{[\mu\nu]} = 0\) by direct cyclic summation — carried out most efficiently in the Cartan formalism below (Proposition 13.157 with \(T^{a} = 0\)); the pair-exchange symmetry is an algebraic consequence of the other three.
∎Let \(x^{\lambda}(\tau,s)\) be a smooth one-parameter family of geodesics of a torsion-free connection, with tangent \(u^{\lambda} = \pp x^{\lambda}/\pp\tau\) and connecting field \(\xi^{\lambda} = \pp x^{\lambda}/\pp s\). Then \(\xi\) obeys the Jacobi equation
with \(\mathrm{D}/\dd\tau = u^{\mu}\nabla_{\mu}\) the covariant derivative along the curve. Rests on Theorem 13.152 and Definition 13.146.
Derives Proposition 13.155. The coordinate fields commute: \([u,\xi]^{\lambda} = u^{\mu}\pp_{\mu}\xi^{\lambda} - \xi^{\mu}\pp_{\mu}u^{\lambda} = \pp_{\tau}\pp_{s}x^{\lambda} - \pp_{s}\pp_{\tau}x^{\lambda} = 0\). For any connection, \(\nabla_{u}\xi^{\lambda} - \nabla_{\xi}u^{\lambda} = [u,\xi]^{\lambda} + T^{\lambda}{}_{\mu\nu}u^{\mu}\xi^{\nu}\) by Equation (13.298), so with \(T = 0\) the two mixed derivatives agree, \(\nabla_{u}\xi = \nabla_{\xi}u\). Hence
the exchange being the Ricci identity Equation (13.305) applied to the pair \((u,\xi)\): the commutator term \(\nabla_{[u,\xi]}\) vanishes because \([u,\xi] = 0\), and the torsion term because \(T = 0\). The first term on the right vanishes because every member of the family is a geodesic — \(\nabla_{u}u = 0\) identically in \(s\), and \(\nabla_{\xi}\) differentiates that identity.
∎The Cartan formalism
The results of Sections 13.8, 13.10 and 13.11 compress into two lines when written in the language of differential forms (Section 13.6), with all coordinate indices hidden. The dynamical variables are the vielbein one-forms \(e^{a}\) and the \(\mathfrak{so}(p,q)\)-valued connection one-form
and the Lorentz-covariant exterior derivative acts on frame-indexed forms as \(\mathrm{D}\alpha^{a} = \dd\alpha^{a} + \omega^{a}{}_{b}\wedge\alpha^{b}\) (one \(\omega\) term per frame index).
The torsion and curvature of Definition 13.147 and Theorem 13.152 are the two-forms
with \(T^{a}{}_{\mu\nu} = e^{a}{}_{\lambda}T^{\lambda}{}_{\mu\nu}\) and \(R^{a}{}_{b\mu\nu} = e^{a}{}_{\lambda}\,e_{b}{}^{\rho}\,R^{\lambda}{}_{\rho\mu\nu}\). Rests on Definition 13.122, Equation (13.303), Definition 13.147, Theorem 13.152, Definition 13.103 and Definition 13.102.
Derives Theorem 13.156. For Equation (13.311): by the vielbein postulate Equation (13.303), \(\pp_{\mu}e^{a}{}_{\nu} + \omega_{\mu}{}^{a}{}_{b}e^{b}{}_{\nu} = \Gamma^{\lambda}{}_{\mu\nu}e^{a}{}_{\lambda}\); antisymmetrizing in \(\mu\nu\), the left side gives the components of \(\dd e^{a} + \omega^{a}{}_{b}\wedge e^{b}\) and the right side gives \(\tfrac12(\Gamma^{\lambda}{}_{\mu\nu} - \Gamma^{\lambda}{}_{\nu\mu}) e^{a}{}_{\lambda} = \tfrac12 T^{\lambda}{}_{\mu\nu}e^{a}{}_{\lambda}\). For Equation (13.312): solve Equation (13.303) for \(\omega_{\mu}{}^{a}{}_{b}\) in terms of \(\Gamma\) and \(e\), insert into \(\dd\omega + \omega\wedge\omega\), and the inhomogeneous \(\pp e\) contributions cancel, leaving exactly the components Equation (13.306) converted with two vielbeins. (Equivalently: \(\mathrm{D}^{2}V^{a} = R^{a}{}_{b}V^{b}\) for any zero-form \(V^{a}\), by a two-line computation, which identifies Equation (13.312) as the frame version of the Ricci identity Equation (13.305).)
∎Applying \(\dd\) to the structure equations,
In the torsion-free case, Equation (13.313) reduces to the first Bianchi identity \(R^{a}{}_{b}\wedge e^{b} = 0\), i.e.\ \(R^{\lambda}{}_{[\rho\mu\nu]} = 0\), and Equation (13.314) is the second identity of Equation (13.308). Rests on Theorem 13.156, Equation (13.245) and Equation (13.246).
Derives Proposition 13.157. Differentiate Equation (13.311) with \(\dd^{2} = 0\) (Section 13.6.4):
which is Equation (13.313) after moving the \(\omega\wedge T\) term to the left as part of \(\mathrm{D}T^{a}\). Differentiating Equation (13.312) likewise, \(\dd R^{a}{}_{b} = \dd\omega^{a}{}_{c}\wedge\omega^{c}{}_{b} - \omega^{a}{}_{c}\wedge\dd\omega^{c}{}_{b} = R^{a}{}_{c}\wedge\omega^{c}{}_{b} - \omega^{a}{}_{c}\wedge R^{c}{}_{b}\) (the \(\omega\wedge\omega\wedge\omega\) terms cancel pairwise), which is \(\mathrm{D}R^{a}{}_{b} = 0\).
∎Given \(e^{a}\) and a prescribed torsion two-form \(T^{a}\), the first structure equation Equation (13.311) determines \(\omega^{ab} = -\omega^{ba}\) uniquely; for \(T^{a} = 0\), \(\omega = \omega(e)\) is the frame form of the Levi-Civita connection, with the contorsion one-form \(K^{a}{}_{b}\) (\(T^{a} = K^{a}{}_{b}\wedge e^{b}\)) supplying the general case, in exact parallel with Equation (13.299). Rests on Theorem 13.156, Theorem 13.150 and Equation (13.266).
Derives Proposition 13.158. Counting: \(\omega_{\mu}{}^{ab}\) has \(D\cdot D(D-1)/2\) independent components, exactly the number of components of the two-form system Equation (13.311); the explicit solution repeats the cyclic permutation trick of Theorem 13.150 on the anholonomy coefficients Equation (13.266),
whose verification is the same computation as in Theorem 13.150 transcribed to frame indices.
∎Equations Equations (13.304), (13.311) and (13.312) exhibit the frame bundle of \((M,g)\) as a gauge structure with group \(\operatorname{L}_{p+q} = \SO(p,q)\): \(\omega\) is the gauge potential, \(R\) its field strength, and \(e\) a Lorentz-vector-valued one-form whose field strength is the torsion. The fibre-bundle formulation is taken up in Lie Groups, Lie Algebras, and Fibre Bundles; the physical instantiation is general relativity (Part V — General Relativity and Cosmology), where the Einstein–Hilbert action becomes
the equality of the two integrands being proven in Appendix A.8 and used in Geometric Formulation of Gravity.
Maximally symmetric spaces and the symmetry ladder
The spaces admitting the full complement of \(D(D+1)/2\) Killing vectors (Proposition 13.140) are characterized by a curvature built from the metric alone, with one constant:
with \(T^{a} = 0\) and \(\ell\) a length (for \(\epsilon = 0\) the curvature vanishes and \(\ell\) drops out). In form language, \(R^{ab} = (\epsilon/\ell^{2})\,e^{a}\wedge e^{b}\).
Let \((M,g)\) be a connected pseudo-Riemannian manifold of dimension \(D\geq 2\) with the Levi-Civita connection, and suppose it admits the maximal number \(D(D+1)/2\) of linearly independent Killing vectors. Then its curvature has the form Equation (13.318), with
\(R\) being the constant curvature scalar of Equation (13.307). Rests on Proposition 13.140, Proposition 13.154 and Definition 13.153.
Derives Theorem 13.160. Step 1: the data a Killing vector may carry at a point. The proof of Proposition 13.140 showed that a Killing vector is determined throughout \(M\) by the \(D\) numbers \(\xi_{\mu}(P)\) and the \(D(D-1)/2\) numbers \(\nabla_{[\mu}\xi_{\nu]}(P)\) at one point \(P\); that is, the linear map sending \(\xi\) to this data is injective. By hypothesis its domain has dimension \(D(D+1)/2\), which is exactly the dimension of its target, so the map is a bijection: for every vector \(v_{\mu}\) and every antisymmetric \(A_{\mu\nu}\) there is a Killing vector with \(\xi_{\mu}(P)=v_{\mu}\) and \(\nabla_{\mu}\xi_{\nu}(P)=A_{\mu\nu}\), the latter being antisymmetric already by Killing's equation Equation (13.290).
Step 2: isotropy of the curvature. The flow of a Killing vector consists of isometries (Definition 13.139), and the Riemann tensor is built from the metric by a rule that commutes with diffeomorphisms, so \(\phi_{t}^{\ast}R[g]=R[\phi_{t}^{\ast}g]=R[g]\); differentiating at \(t=0\) gives \(\mathcal{L}_{\xi}R=0\). Writing the Lie derivative of a \((0,4)\) tensor with covariant derivatives, which is legitimate for a torsion-free connection (Proposition 13.127),
Evaluate at \(P\) on a Killing vector with \(\xi(P)=0\) and \(\nabla_{\mu}\xi_{\nu}(P)=A_{\mu\nu}\), which Step 1 provides for arbitrary antisymmetric \(A\). The transport term drops and
holds at \(P\) for every antisymmetric \(A\).
Step 3: the algebra. Fix two indices \(\alpha,\beta\) and take \(A_{\sigma\tau}=g_{\sigma\alpha}g_{\tau\beta}-g_{\sigma\beta} g_{\tau\alpha}\), so that \(A_{\lambda}{}^{\sigma}=g_{\lambda\alpha}\delta^{\sigma}_{\ \beta} -g_{\lambda\beta}\delta^{\sigma}_{\ \alpha}\). Then Equation (13.321) becomes
Contract Equation (13.322) with \(g^{\lambda\alpha}\). The first term gives \(D\,R_{\beta\rho\mu\nu}\); the second and the third give \(-R_{\beta\rho\mu\nu}\) each, the third after using \(R_{\rho\beta\mu\nu}=-R_{\beta\rho\mu\nu}\) (Proposition 13.154); the fourth vanishes by the same antisymmetry; the sixth and eighth produce Ricci tensors (Definition 13.153), \(-g_{\mu\beta}R_{\rho\nu}\) and \(+g_{\nu\beta}R_{\rho\mu}\); the fifth and seventh survive as \(R_{\mu\rho\beta\nu}\) and \(R_{\nu\rho\mu\beta}\). Hence
Use the pair-exchange symmetry of Equation (13.308) to write \(R_{\mu\rho\beta\nu}=R_{\beta\nu\mu\rho}\) and \(R_{\nu\rho\mu\beta}=-R_{\beta\mu\nu\rho}\), both with \(\beta\) in the first slot, and then the first Bianchi identity \(R_{\beta[\rho\mu\nu]}=0\) in the form \(R_{\beta\nu\mu\rho}=R_{\beta\rho\mu\nu}+R_{\beta\mu\nu\rho}\). The last two terms of the left side of Equation (13.323) collapse to \(R_{\beta\rho\mu\nu}\), leaving
after renaming \(\beta\) to \(\lambda\).
The left side of Equation (13.324) is antisymmetric under \(\lambda\leftrightarrow\rho\), so the right side must be too:
Contracting this with \(g^{\lambda\mu}\) gives \((D-1)R_{\rho\nu} = -R_{\rho\nu}+g_{\rho\nu}R\), that is
Substituting Equation (13.325) into Equation (13.324),
which is the tensor structure of Equation (13.318) at the point \(P\), and \(P\) was arbitrary.
Step 4: the coefficient is constant. By Step 1 there is, at each point \(P\) and for each vector \(v\), a Killing vector with \(\xi(P)=v\). Since \(R\) is a scalar built from the metric and the flow of \(\xi\) consists of isometries, \(\mathcal{L}_{\xi}R=\xi^{\sigma}\pp_{\sigma}R=0\) (Equation (13.272)), so \(v^{\sigma}\pp_{\sigma}R=0\) at \(P\) for every \(v\), whence \(\pp_{\sigma}R=0\) there. As \(P\) was arbitrary and \(M\) is connected, \(R\) is constant. Writing that constant as Equation (13.319) — with \(\epsilon=0\) when \(R=0\), and otherwise \(\epsilon=\pm 1\) the sign of \(R\) and \(\ell=\left(\abs{R}/D(D-1)\right)^{-1/2}\) the associated length — turns Equation (13.326) into Equation (13.318).
∎Step 4 did not need the second Bianchi identity, only homogeneity. Where a space is isotropic about every point but the full Killing count is not assumed, Equation (13.326) still holds with a function \(K(x)\) in place of the constant, and inserting it into the second Bianchi identity of Equation (13.308) and contracting twice yields \((D-1)(D-2)\nabla_{\sigma}K=0\): for \(D>2\) pointwise isotropy alone already forces \(K\) to be constant. This is Schur's theorem, and it is the reason the three cases below exhaust the possibilities.
The three cases realize the groups of Table 13.1:
-
[Flat space (\(\epsilon = 0\)).] The manifold is \(\R^{p,q}\) with metric \(\eta_{\mu\nu}\); the Killing vectors are the \(D\) translations \(\pp_{\mu}\) (group \(\operatorname{T}_{p+q}\), abelian) and the \(D(D-1)/2\) pseudo-rotations \(x_{\mu}\pp_{\nu} - x_{\nu}\pp_{\mu}\) (group \(\operatorname{L}_{p+q} = \SO(p,q)\)), closing into the Poincaré group \(\operatorname{P}_{p+q} = \ISO(p,q)\), the semidirect product of the two.
-
[Anti-de Sitter space (\(\epsilon = -1\)).] \(\operatorname{AdS}_{p+q}\) is the quadric
\begin{equation}\tag{13.327} \eta_{AB}\,X^{A}X^{B} = -\ell^{2} \end{equation}in the flat \((p+q+1)\)-dimensional space of signature \((p,q+1)\), with embedding coordinates \(X^{m}\) and frame metric \(\eta_{AB}\) per the conventions of Section 13.1. The induced metric (Definition 13.119) satisfies Equation (13.318) with \(\epsilon = -1\), by Proposition 13.162 below. Every linear transformation of \(X^{m}\) preserving \(\eta_{AB}\) preserves the quadric and its induced metric, so \(\operatorname{Isom}(\operatorname{AdS}_{p+q}) \supseteq \SO(p,q+1)\); its dimension \((D+1)D/2\) saturates the Killing bound, so equality holds (connected component).
-
[de Sitter space (\(\epsilon = +1\)).] Identically, \(\operatorname{dS}_{p+q}\) is the quadric \(\eta_{AB}X^{A}X^{B} = +\ell^{2}\) in signature \((p+1,q)\), with isometry group \(\SO(p+1,q)\).
The curvature of the two quadrics is computed once, for both signs at once, by the Gauss equation of Section 13.3.2 transcribed to a pseudo-Riemannian hypersurface of a flat space.
Let \(\eta_{AB}\) be the flat metric of a \((D+1)\)-dimensional space with Cartesian coordinates \(X^{A}\), let \(\sigma=\pm 1\), and let \(\Sigma\) be the quadric
described by an embedding \(X^{A}(\xi^{i})\), \(i=1,\ldots,D\). Then the metric \(h_{ij}\) induced on \(\Sigma\) (Definition 13.119) has curvature
that is, Equation (13.318) with \(\epsilon = \sigma\). Rests on Definition 13.119, Equation (13.83) and Theorem 13.152.
Derives Proposition 13.162. Write \(e_{i}^{A}=\pp X^{A}/\pp\xi^{i}\) for the tangent vectors, so that \(h_{ij}=\eta_{AB}e_{i}^{A}e_{j}^{B}\) by Equation (13.261). Differentiating Equation (13.328) along \(\Sigma\) gives \(\eta_{AB}X^{A}e_{i}^{B}=0\): the position vector is normal to the quadric. Hence
the last relation because \(X^{A}\) is itself the embedding.
At each point the \(D+1\) vectors \(\set{e_{i}^{A},N^{A}}\) are a basis of the ambient space, so we may expand, exactly as in the Gauss equations Equation (13.83),
the factor \(\sigma\) being what makes the stated formula for \(b_{ij}\) follow from Equation (13.330) on contracting with \(\eta_{AB}N^{B}\); contracting instead with \(\eta_{AB}e_{k}^{B}\) identifies \(\Gamma^{k}{}_{ij}\) as the Levi-Civita connection of \(h_{ij}\), by the same computation that produced Equation (13.87) from Equation (13.83). Differentiating \(\eta_{AB}N^{A}e_{j}^{B}=0\) with respect to \(\xi^{i}\) and using Equation (13.330),
The second fundamental form of the quadric is thus proportional to its first: every direction is a curvature direction, which is the analytic form of the statement that the quadric looks the same from every side.
The ambient space is flat and referred to Cartesian coordinates, so \(\pp_{i}\pp_{j}e_{k}^{A}\) is symmetric in \(i\) and \(j\). Differentiating Equation (13.331) once more and antisymmetrizing in \(i,j\), the component along \(e_{m}^{A}\) must vanish:
the last bracket arising from the terms \(\sigma b_{jk}\pp_{i}N^{A}\) through Equation (13.330). The first four terms are precisely \(R^{m}{}_{kij}\) of the induced metric, in the convention Equation (13.306). Inserting Equation (13.332),
and lowering the first index with \(h_{lm}\) and renaming gives Equation (13.329).
For anti-de Sitter space Equation (13.327) has \(\sigma=-1\), so \(\epsilon=-1\); for de Sitter space \(\sigma=+1\), so \(\epsilon=+1\). In both cases the induced metric is nondegenerate of the signature stated, because the normal direction Equation (13.330) carries the extra sign \(\sigma\) of the ambient metric and is removed from it.
∎In the limit \(\ell \to \infty\) both quadrics flatten and their isometry groups contract to \(\operatorname{P}_{p+q}\) — the İnönü–Wigner contraction, treated algebraically in Lie Groups, Lie Algebras, and Fibre Bundles.
The conformal group
Relaxing isometry to angle preservation, \(f^{\ast}g = \Omega^{2}(x)\,g\), enlarges the symmetry. For flat space of signature \((p,q)\), \(D = p + q > 2\), the conformal Killing vectors are
for a total of \((D+1)(D+2)/2\) generators. That this list is complete is the content of the following proposition.
Let \(\eta_{\mu\nu}\) be the flat metric of signature \((p,q)\) in Cartesian coordinates, \(D=p+q>2\). Every solution of the conformal Killing equation
is a polynomial of degree at most two in \(x\), and the solution space is spanned by the vector fields Equation (13.333); its dimension is \((D+1)(D+2)/2\). Rests on Equation (13.275), Definition 13.139 and Equation (13.333).
Derives Proposition 13.163. Equation (13.334) is \(\mathcal{L}_{\xi}\eta=2\sigma\eta\) written out with Equation (13.275), the trace of which fixes \(\sigma\) as stated.
Step 1: second derivatives in terms of \(\sigma\). Put \(E_{\mu\nu}=\pp_{\mu}\xi_{\nu}+\pp_{\nu}\xi_{\mu}-2\sigma\eta_{\mu\nu}\), so that \(E_{\mu\nu}=0\). Form the combination \(\pp_{\rho}E_{\mu\nu}+\pp_{\mu}E_{\nu\rho}-\pp_{\nu}E_{\rho\mu}=0\). Six of the eight derivative terms cancel in pairs because partial derivatives commute, and what survives is
Step 2: \(\sigma\) is affine. Apply \(\pp^{\nu}\) to Equation (13.335). On the left this gives \(\pp_{\mu}\pp_{\rho}\left(\pp^{\nu}\xi_{\nu}\right) =D\,\pp_{\mu}\pp_{\rho}\sigma\), and on the right \(2\pp_{\mu}\pp_{\rho}\sigma-\eta_{\rho\mu}\Box\sigma\) with \(\Box=\pp^{\nu}\pp_{\nu}\), whence
Tracing Equation (13.336) gives \((D-2)\Box\sigma=-D\,\Box\sigma\), that is \(\Box\sigma=0\) for \(D\neq 1\); and then Equation (13.336) itself gives \(\pp_{\mu}\pp_{\rho}\sigma=0\), since \(D>2\). Hence \(\sigma\) is an affine function, which we write
with \(\lambda\) and \(b_{\mu}\) constant, the factor \(2\) chosen for later convenience. This is the only place where \(D>2\) is used, and it is where the argument genuinely fails in two dimensions, whose conformal algebra is infinite-dimensional.
Step 3: \(\xi\) is quadratic. With \(\pp\pp\sigma=0\) the right-hand side of Equation (13.335) is constant, so all third derivatives of \(\xi\) vanish and its Taylor expansion about the origin terminates:
Step 4: the three orders. Each order of Equation (13.338) is constrained separately, because Equation (13.334) must hold identically in \(x\).
-
Constant part. \(a_{\nu}\) is unconstrained: \(D\) parameters, the translations \(\xi^{\mu}=a^{\mu}\).
-
Linear part. Here Equation (13.334) reads \(b_{\nu\mu}+b_{\mu\nu}=2\lambda\eta_{\mu\nu}\), so \(b_{\nu\mu}=\lambda\eta_{\nu\mu}+\lambda_{\nu\mu}\) with \(\lambda_{\nu\mu}=-\lambda_{\mu\nu}\) arbitrary: one parameter for the dilatation \(\xi^{\mu}=\lambda\,x^{\mu}\) and \(D(D-1)/2\) for the pseudo-rotations \(\xi^{\mu}=\lambda^{\mu}{}_{\nu}x^{\nu}\).
-
Quadratic part. By Equations (13.335) and (13.337), \(\pp_{\mu}\pp_{\rho}\xi_{\nu} = 2\left(\eta_{\mu\nu}b_{\rho}+\eta_{\nu\rho}b_{\mu} -\eta_{\rho\mu}b_{\nu}\right)\), so the last term of Equation (13.338) is
\begin{equation*} \left(\eta_{\mu\nu}b_{\rho}+\eta_{\nu\rho}b_{\mu} -\eta_{\rho\mu}b_{\nu}\right)x^{\mu}x^{\rho} = 2\left(b\cdot x\right)x_{\nu} - x^{2}b_{\nu}\ec \end{equation*}which is the special conformal transformation of Equation (13.333): \(D\) parameters \(b_{\mu}\).
Conversely each of the four families satisfies Equation (13.334) by direct substitution — for the last, with \(\sigma=2\,b\cdot x\) — so the solution space is exactly their span. The parameters are independent, being the coefficients of distinct powers of \(x\) together with the irreducible pieces of \(b_{\nu\mu}\), and they number
which is the dimension of \(\SO(p+1,q+1)\) quoted below.
∎These close into the conformal group \(\operatorname{C}_{p+q} = \SO(p+1,q+1)\) of Table 13.1, realized linearly on the \((p+q+2)\)-dimensional flat space of signature \((p+1,q+1)\): flat \((p,q)\)-space is identified with rays of the null cone
(projective light-cone construction), on which \(\SO(p+1,q+1)\) acts by its defining representation, indices \(M,N\) and \(I,J\) per Section 13.1. The dimension matches: \(\dim\SO(p+1,q+1) = (D+1)(D+2)/2\).
The observed case
Physics instantiates this chapter at \(p + q = 3 + 1\). The kinematics of Part IV — Special Relativity is the flat row of Table 13.1; general relativity (Part V — General Relativity and Cosmology) is the torsion-free, Levi-Civita case of Sections 13.10, 13.11 and 13.12, tested by the experiments recorded there; and the accelerating expansion of the Universe (Evidence-Based Cosmology) makes the de Sitter row observationally relevant. Torsion, by contrast, is retained in this chapter as mathematics: no experiment to date requires \(T^{a} \neq 0\), and the physical parts of this treatise accordingly set it to zero — the scope rule of Epistemology and the Scientific Method applied to geometry itself.