Lorentz Transformations

Contents
  1. The postulates
  2. Lorentz boosts
  3. Spacetime diagrams
  4. Consequences of the postulates
  5. Composition of velocities
  6. General Lorentz transformations

A Lorentz transformation relates the space-time coordinates that two inertial frames assign to one and the same event. This chapter derives them from the two postulates of Experiments: Light, the Aether, and Time, reads their consequences off the geometry of Minkowski space, and characterizes the most general such transformation as a boost composed with a spatial rotation. The founding paper is Einstein's [Einstein:1905a]; the geometric reading followed from Minkowski [Minkowski:1909], and the group structure was identified by Poincaré [Poincare:1906]. Standard modern treatments are [Jackson:1999] [Landau:1976].

The postulates

Everything in this part follows from two statements, both of them experimental claims rather than mathematical ones.

Postulate 38.1 (Principle of relativity).

The laws of physics take the same form in every inertial frame. No experiment performed within a freely moving laboratory can determine its velocity.

Postulate 38.2 (Invariance of the speed of light).

Light propagates in vacuum with the same speed \(c\) in every inertial frame, independently of the motion of its source.

Postulate 38.1 is old — it is Galileo's, and Newtonian mechanics obeys it. Postulate 38.2 is the new one, and it is what the aether experiments of Experiments: Light, the Aether, and Time force upon us: the Michelson–Morley null result (Section 37.3) says that the speed of light does not depend on the motion of the apparatus, and the Fizeau measurement (Section 37.4) says that it does not add to the speed of its medium in the Galilean way. The two postulates are incompatible with Galilean kinematics, in which speeds simply add, and the whole of this chapter is the working out of what must replace it.

Since 1983 the second postulate has been built into the system of units: \(c\) is a defined constant, exactly \(299792458\,\mathrm{m}/\mathrm{s}\), and the metre is the distance light travels in \(1/c\) seconds (Measurement, SI Units, and the Theory of Errors). Measuring \(c\) is therefore no longer possible; what the experiments now test is the isotropy and source-independence asserted above, and those are still measured, to the precision reported in Section 37.5.

Notation 38.3 (Standard configuration).

Throughout, \(K\) and \(\bar{K}\) are inertial frames whose origins coincide at \(t=\bar{t}=0\), whose axes are parallel, and whose relative velocity is \(v\) along the common \(x\)-axis. We write

\begin{equation}\tag{38.1} \beta:=\frac{v}{c}\ec\qquad \gamma:=\frac{1}{\sqrt{1-\beta^{2}}}\ec \end{equation}

so that \(\gamma\geq1\), with \(\gamma=1\) only for \(\beta=0\), and \(\gamma\to\infty\) as \(\beta\to1\). Coordinates are collected as \(x^{\mu}=(x^{0},x^{i})=(ct,\vect{x})\), with Greek indices running over \(0,1,2,3\) and Latin over \(1,2,3\).

Lorentz boosts

The one-dimensional boost

Theorem 38.4 (Lorentz boost).

In the standard configuration of Notation 38.3, the two postulates force the coordinate transformation

\begin{align} \tag{38.2} \bar{x} &= \gamma\left(x-vt\right)\ec\\ \tag{38.3} \bar{t} &= \gamma\left(t-\frac{v}{c^{2}}x\right)\ec\\ \tag{38.4} \bar{y} &= y\ec\qquad \bar{z}=z\ep \end{align}

Rests on Postulate 38.1, Postulate 38.2 and Notation 38.3.

Proof.

Derives Theorem 38.4. Step 1: the transformation is linear. Space and time are homogeneous — no event and no place is privileged — so a uniform velocity in \(K\) must appear as a uniform velocity in \(\bar{K}\): straight worldlines go to straight worldlines. A map carrying straight lines to straight lines and fixing the coincident origins is linear. Write, for the \(x\) and \(t\) pair,

\begin{equation}\tag{38.5} \bar{x}=Ax+Bt\ec\qquad \bar{t}=Dx+Et\ec \end{equation}

with coefficients depending on \(v\) alone.

Step 2: the origin of \(\bar{K}\). The point \(\bar{x}=0\) is by construction the origin of \(\bar{K}\), which moves in \(K\) along \(x=vt\). Substituting into Equation (38.5) gives \(Avt+Bt=0\) for all \(t\), whence \(B=-Av\) and

\begin{equation}\tag{38.6} \bar{x}=A\left(x-vt\right)\ep \end{equation}

From Equation (38.5) (imposing that \(\bar{x}=0\) be the worldline \(x=vt\)).

Step 3: the invariance of \(c\). Emit a light pulse from the common origin at \(t=\bar{t}=0\) in the \(+x\) direction. In \(K\) it obeys \(x=ct\), in \(\bar{K}\) it must obey \(\bar{x}=c\bar{t}\) by Postulate 38.2. From Equations (38.5) and (38.6),

\[ \bar{x}=A\left(c-v\right)t\ec\qquad \bar{t}=\left(Dc+E\right)t\ec \]

so \(\bar{x}=c\bar{t}\) requires

\begin{equation}\tag{38.7} A\left(c-v\right)=c\left(Dc+E\right)\ep \end{equation}

A pulse sent in the \(-x\) direction obeys \(x=-ct\) and \(\bar{x}=-c\bar{t}\), and the same substitution gives

\begin{equation}\tag{38.8} A\left(c+v\right)=c\left(E-Dc\right)\ep \end{equation}

Adding Equations (38.7) and (38.8) yields \(2Ac=2cE\), hence \(E=A\); subtracting them yields \(2Av=-2Dc^{2}\), hence \(D=-Av/c^{2}\). So far

\begin{equation}\tag{38.9} \bar{x}=A\left(x-vt\right)\ec\qquad \bar{t}=A\left(t-\frac{v}{c^{2}}x\right)\ep \end{equation}

From Equations (38.7) and (38.8) (by addition and subtraction).

Step 4: the principle of relativity fixes \(A\). By Postulate 38.1 neither frame is privileged, so expressing \((x,t)\) in terms of \((\bar{x},\bar{t})\) must reproduce Equation (38.9) with \(v\) replaced by \(-v\) — from \(\bar{K}\), the frame \(K\) recedes at \(-v\). Moreover \(A\) can depend only on \(\abs{v}\): space is isotropic, so reflecting the \(x\)-axis, which sends \(v\to-v\), cannot change a scale factor. Hence \(A(-v)=A(v)\), and

\[ x=A\left(\bar{x}+v\bar{t}\right)\ec\qquad t=A\left(\bar{t}+\frac{v}{c^{2}}\bar{x}\right)\ep \]

Substituting Equation (38.9) into the first of these,

\begin{align*} x &= A\left[A\left(x-vt\right) +vA\left(t-\frac{v}{c^{2}}x\right)\right]\\ &= A^{2}\left[x-vt+vt-\frac{v^{2}}{c^{2}}x\right] = A^{2}\left(1-\beta^{2}\right)x\ec \end{align*}

so \(A^{2}(1-\beta^{2})=1\). The root is fixed by requiring the transformation to reduce to the identity at \(v=0\) rather than to a reflection, so \(A=+\left(1-\beta^{2}\right)^{-1/2}=\gamma\). This is Equations (38.2) and (38.3).

Step 5: the transverse coordinates. Consider a rod of proper length \(\ell\) lying along the \(y\)-axis, at rest in \(\bar{K}\), and suppose \(\bar{y}=\kappa(v)y\) with \(\kappa\neq1\). By isotropy \(\kappa\) depends only on \(\abs{v}\), and by the principle of relativity the inverse relation is \(y=\kappa(v)\bar{y}\); composing gives \(\kappa^{2}=1\), so \(\kappa=1\) (again excluding the reflection). Lengths perpendicular to the motion are unchanged, which is Equation (38.4).

Setting \(c\to\infty\) in Equations (38.2) and (38.3) returns \(\bar{x}=x-vt\) and \(\bar{t}=t\): the Galilean transformation is the limit in which light is infinitely fast. Every departure from Newtonian kinematics in this part is controlled by the single dimensionless number \(\beta\), and at \(\beta=10^{-6}\) — a satellite in low orbit — \(\gamma-1\) is about \(5\times10^{-13}\), which is why the Galilean form survived two centuries of testing.

In matrix form, with \(x^{\mu}=(ct,x,y,z)\),

\begin{equation}\tag{38.10} \bar{x}^{\mu}=\Lambda^{\mu}{}_{\nu}x^{\nu}\ec\qquad \left[\Lambda\right]= \begin{pmatrix} \gamma & -\gamma\beta & 0 & 0\\ -\gamma\beta & \gamma & 0 & 0\\ 0 & 0 & 1 & 0\\ 0 & 0 & 0 & 1 \end{pmatrix}\ep \end{equation}

From Equations (38.2), (38.3) and (38.4) (collecting the three coordinate relations into one matrix).

Minkowski space

In a Euclidean space, rotations — linear homogeneous transformations — leave the norm invariant, and with it the Euclidean metric. A boost is also linear and homogeneous and resembles a rotation formally, though not physically.[As Section 39.2.4 shows, a Euclidean rotation and a Lorentz boost are even topologically distinct: the rotation group is compact and the Lorentz group is not.] It is natural to ask which metric a boost preserves — equivalently, since the metric defines the space, in which space a boost is a rigid motion.

Proposition 38.5 (The invariant interval).

Under the boost Equations (38.2), (38.3) and (38.4),

\begin{equation}\tag{38.11} c^{2}\bar{t}^{\,2}-\bar{x}^{2}-\bar{y}^{2}-\bar{z}^{2} =c^{2}t^{2}-x^{2}-y^{2}-z^{2}\ep \end{equation}

Rests on Theorem 38.4 and Equation (38.1).

Proof.

Derives Proposition 38.5. The transverse terms are unchanged. For the other two,

\begin{align*} c^{2}\bar{t}^{\,2}-\bar{x}^{2} &= \gamma^{2}\left(ct-\beta x\right)^{2} -\gamma^{2}\left(x-\beta ct\right)^{2}\\ &= \gamma^{2}\left[c^{2}t^{2}-2\beta ctx+\beta^{2}x^{2} -x^{2}+2\beta ctx-\beta^{2}c^{2}t^{2}\right]\\ &= \gamma^{2}\left(1-\beta^{2}\right) \left(c^{2}t^{2}-x^{2}\right) = c^{2}t^{2}-x^{2}\ec \end{align*}

the last step by Equation (38.1).

Definition 38.6 (Minkowski metric and interval).

Minkowski space is \(\R^{4}\) equipped with the symmetric bilinear form of signature \((1,3)\)

\begin{equation}\tag{38.12} \left[\eta_{\mu\nu}\right]=\diag\left(1,-1,-1,-1\right)\ec \end{equation}

and the invariant interval between two events separated by \(\Delta x^{\mu}\) is

\begin{equation}\tag{38.13} \Delta s^{2}:=\eta_{\mu\nu}\Delta x^{\mu}\Delta x^{\nu} =c^{2}\Delta t^{2}-\abs{\Delta\vect{x}}^{2}\ep \end{equation}

Rests on Proposition 38.5.

Proposition 38.5 says precisely that a boost preserves Equation (38.13). Written with \(\Lambda\), the condition that a linear transformation preserve \(\eta\) is

\begin{equation}\tag{38.14} \eta_{\mu\nu}\Lambda^{\mu}{}_{\lambda}\Lambda^{\nu}{}_{\rho} =\eta_{\lambda\rho}\ec \qquad\text{that is}\qquad \Lambda\transpose\eta\,\Lambda=\eta\ec \end{equation}

the exact analogue of the orthogonality condition \(O\transpose O=\identity\) that defines a Euclidean rotation, with \(\identity\) replaced by \(\eta\). Equation (38.14) is taken as the definition of a Lorentz transformation in Section 39.2.1, and the whole of Minkowski Space and Its Symmetries is the study of the group it defines.

Remark 38.7 (Signature conventions).

This treatise uses \((+,-,-,-)\), so that \(\Delta s^{2}>0\) for timelike separations and proper time is real. The opposite convention \((-,+,+,+)\) is equally common — Appendix A.1 uses it, following the gravitational-wave literature — and flips the sign of every interval. Nothing physical depends on the choice; every observable is a ratio or a sign comparison in which the convention cancels.

The light cone

The sign of \(\Delta s^{2}\) is invariant, so all observers agree on it. It partitions the events of Minkowski space, relative to any chosen event \(O\), into the regions of Figure 38.1.

[figure: st-lightcone.pdf]

The causal structure of Minkowski space, drawn with one spatial dimension suppressed. Light through the event \(O\) travels on the \(45\,^\circ\) lines \(x=\pm ct\), where \(\Delta s^{2}=0\). Events with \(\Delta s^{2}>0\) are timelike separated from \(O\) and lie inside the cone, divided into a future and a past that every observer agrees upon; events with \(\Delta s^{2}<0\) are spacelike separated and lie outside it. A massive particle's worldline stays everywhere inside the cone, since its speed never reaches \(c\).

Definition 38.8 (Causal classification).

A separation \(\Delta x^{\mu}\) between two events is timelike if \(\Delta s^{2}>0\), null or lightlike if \(\Delta s^{2}=0\), and spacelike if \(\Delta s^{2}<0\). Rests on Definition 38.6.

Proposition 38.9 (Causal order is absolute).

If two events are timelike or null separated, every inertial observer agrees which came first. If they are spacelike separated, their time order depends on the frame, and there are frames in which they are simultaneous. Rests on Equation (38.3) and Definition 38.8.

Proof.

Derives Proposition 38.9. Take the events at \(\Delta x^{\mu}=(c\Delta t,\Delta x,0,0)\) with \(\Delta t>0\). By Equation (38.3), \(c\Delta\bar{t}=\gamma\left(c\Delta t-\beta\Delta x\right)\), which changes sign when \(\beta\) passes through \(c\Delta t/\Delta x\). This value is an admissible velocity, \(\abs{\beta}<1\), exactly when \(\abs{\Delta x}>c\abs{\Delta t}\), that is when \(\Delta s^{2}<0\). For timelike and null separations no admissible boost reverses the order; for spacelike separations the boost with \(\beta=c\Delta t/\Delta x\) makes \(\Delta\bar{t}=0\) and any larger \(\beta\) reverses the order.

This is the reason the invariant speed also bounds causal influence. If a signal could travel faster than \(c\), it would connect spacelike-separated events; by Proposition 38.9 some observer would see it arrive before it was sent, and the notion of cause would not survive a change of frame. Nothing in the postulates forbids superluminal coordinate speeds that carry no signal — the intersection point of a pair of closing scissor blades is a standard example — and nothing in them permits superluminal signalling.

The general boost

For a relative velocity \(\vect{v}=c\vect{\beta}\) in an arbitrary direction, resolve the position into components along and across the motion,

\begin{equation}\tag{38.15} \vect{x}_{\parallel} =\frac{\left(\vect{x}\cdot\vect{\beta}\right)} {\abs{\vect{\beta}}^{2}}\vect{\beta}\ec \qquad \vect{x}_{\perp}=\vect{x}-\vect{x}_{\parallel}\ep \end{equation}

Theorem 38.4 applies to the longitudinal pair and leaves the transverse part alone, so

\begin{align} \tag{38.16} c\bar{t} &= \gamma\left(ct-\vect{\beta}\cdot\vect{x}\right)\ec\\ \tag{38.17} \bar{\vect{x}}_{\parallel} &= \gamma\left(\vect{x}_{\parallel}-\vect{\beta}\,ct\right)\ec\\ \tag{38.18} \bar{\vect{x}}_{\perp} &= \vect{x}_{\perp}\ep \end{align}

Recombining \(\bar{\vect{x}}=\bar{\vect{x}}_{\parallel} +\bar{\vect{x}}_{\perp}\) and using Equation (38.15),

\begin{equation}\tag{38.19} \bar{\vect{x}}=\vect{x} +\frac{\gamma-1}{\abs{\vect{\beta}}^{2}} \left(\vect{\beta}\cdot\vect{x}\right)\vect{\beta} -\gamma\vect{\beta}\,ct\ep \end{equation}

From Equations (38.15), (38.17) and (38.18) (recombining the longitudinal and transverse parts).

Reading off the components of \(\bar{x}^{\mu}=\Lambda^{\mu}{}_{\nu}(\beta)x^{\nu}\) gives

\begin{align} \Lambda^0_{\ 0}(\beta)&=\gamma\ec\tag{38.20}\\ \Lambda^0_{\ i}(\beta)&=-\gamma\beta_i\ec\tag{38.21}\\ \Lambda^i_{\ 0}(\beta)&=-\gamma\beta^i\ec\tag{38.22}\\ \Lambda^i_{\ j}(\beta)&=\delta^i_{\ j} +\frac{\gamma-1}{\abs{\vect{\beta}}^2}\beta^i\beta_j\ec \tag{38.23} \end{align}

or, in matrix form,

\begin{equation}\tag{38.24} \left[\Lambda(\beta)\right]= \begin{pmatrix} \gamma & -\gamma\beta_1 & -\gamma\beta_2 & -\gamma\beta_3\\[2pt] -\gamma\beta^1 & 1+\frac{\gamma-1}{\abs{\vect{\beta}}^2}\beta^1\beta_1 & \frac{\gamma-1}{\abs{\vect{\beta}}^2}\beta^1\beta_2 & \frac{\gamma-1}{\abs{\vect{\beta}}^2}\beta^1\beta_3\\[2pt] -\gamma\beta^2 & \frac{\gamma-1}{\abs{\vect{\beta}}^2}\beta^2\beta_1 & 1+\frac{\gamma-1}{\abs{\vect{\beta}}^2}\beta^2\beta_2 & \frac{\gamma-1}{\abs{\vect{\beta}}^2}\beta^2\beta_3\\[2pt] -\gamma\beta^3 & \frac{\gamma-1}{\abs{\vect{\beta}}^2}\beta^3\beta_1 & \frac{\gamma-1}{\abs{\vect{\beta}}^2}\beta^3\beta_2 & 1+\frac{\gamma-1}{\abs{\vect{\beta}}^2}\beta^3\beta_3 \end{pmatrix}\ep \end{equation}

From Equations (38.20), (38.21), (38.22) and (38.23) (assembling the components into a matrix).

Taking \(\vect{\beta}=(\beta,0,0)\) recovers Equation (38.10), since then \(\gamma-1\) multiplies only the \(11\) entry and \(1+(\gamma-1)=\gamma\).

Spacetime diagrams

The consequences of Theorem 38.4 are algebraically trivial and conceptually treacherous; they become obvious when drawn. A spacetime diagram plots \(ct\) upward against \(x\) rightward, so that light travels on the \(45\,^\circ\) lines and a particle's history is a curve — its worldline — everywhere steeper than \(45\,^\circ\).

The boosted frame is drawn in the same picture rather than on a separate sheet. Its time axis is the set \(\bar{x}=0\), which by Equation (38.2) is the line \(x=vt\); its space axis is the set \(\bar{t}=0\), which by Equation (38.3) is the line \(ct=\beta x\). Both therefore make the same angle \(\alpha=\arctan\beta\) with the corresponding unprimed axis, but on opposite sides, as in Figure 38.2: the axes close like a pair of scissors on the light line, and reach it together only as \(\beta\to1\).

[figure: st-boosted-axes.pdf]

The axes of a boosted frame drawn inside the diagram of the unboosted one, for \(\beta=0.5\). The \(c\bar{t}\)-axis is the worldline of the moving origin and the \(\bar{x}\)-axis is the locus of events the moving observer calls simultaneous; both make the angle \(\alpha=\arctan\beta\) with their unprimed counterparts, on opposite sides of the light line \(x=ct\). Because the two axes tilt towards the light line rather than rotating rigidly, the surfaces of constant time in the two frames are genuinely different families.

One warning must be issued before the diagram is used to read off numbers. The page is Euclidean and the geometry it depicts is not, so distances measured with a ruler on the paper are meaningless. What is meaningful is the invariant Equation (38.13), whose level sets are the hyperbolae of Figure 38.3. Those hyperbolae are the orbits of the boost — a boost slides events along them exactly as a Euclidean rotation slides points along circles — and they, not the ruler, calibrate the axes.

[figure: st-hyperbolae.pdf]

Calibration. The hyperbolae \(\left(ct\right)^{2}-x^{2}=\pm1\) are invariant under boosts, so a boost slides events along them, and the unit tick on any frame's time axis is where that axis cuts the timelike hyperbola. The primed ticks are farther from the origin on the page than the unprimed ones, and represent the same invariant interval: this is why tick spacings in the two frames may never be compared by eye.

Consequences of the postulates

Relativity of simultaneity

Phenomenon 38.10 (Relativity of simultaneity).

Two spatially separated events that are simultaneous in one inertial frame are not simultaneous in another moving relative to it. For events separated by \(\Delta x\) and simultaneous in \(K\),

\begin{equation}\tag{38.25} \Delta\bar{t}=-\gamma\frac{v}{c^{2}}\Delta x\ep \end{equation}

Rests on Theorem 38.4 and Equation (38.3).

Derivation. Derives Phenomenon 38.10. Apply Equation (38.3) to the separation of the two events. With \(\Delta t=0\),

\[ \Delta\bar{t}=\gamma\left(\Delta t -\frac{v}{c^{2}}\Delta x\right) =-\gamma\frac{v}{c^{2}}\Delta x\ec \]

which vanishes for all \(\Delta x\) only if \(v=0\).

Figure 38.4 shows why this is not a paradox. “The set of events happening now” is a slice through spacetime, and the two frames slice at different angles. Neither slice is correct; simultaneity is simply not a property of a pair of events, but of a pair of events and a frame. Note that Equation (38.25) grows with the separation \(\Delta x\): disagreement about simultaneity is unmeasurable in a laboratory and enormous over astronomical distances.

[figure: st-simultaneity.pdf]

Relativity of simultaneity. The events \(A\) and \(B\) lie on one horizontal line, so \(K\) assigns them the same \(t\). The surfaces of constant \(\bar{t}\) are tilted, so \(A\) and \(B\) lie on different ones and \(\bar{K}\) assigns them different times — here \(B\) happens first. Nothing has moved between the two readings; only the family of slices called “the same instant” has been re-chosen.

Remark 38.11 (Simultaneity is a convention, within limits).

That distant simultaneity must be defined rather than discovered was Einstein's opening move [Einstein:1905a]: he stipulated that light takes equal times on the outward and return legs of a round trip, which is what synchronizes separated clocks. The one-way speed of light is not independently measurable — measuring it presupposes synchronized clocks at the two ends — so this is a coordinative definition in the sense of Section 1.3.2, not an experimental result. What is measured, and what Experiments: Light, the Aether, and Time reports, is the two-way speed and the isotropy of the round trip.

Time dilation

Definition 38.12 (Proper time).

The proper time \(\tau\) along a worldline is the time read by a clock carried along it, defined through the invariant interval by

\begin{equation}\tag{38.26} c^{2}\dd\tau^{2}:=\dd s^{2} =c^{2}\dd t^{2}-\abs{\dd\vect{x}}^{2} \ec\qquad\text{so}\qquad \dd\tau=\frac{\dd t}{\gamma}\ep \end{equation}

Being built from \(\dd s^{2}\), proper time is invariant: every frame agrees on what a given clock reads. Rests on Definition 38.6 and Proposition 38.5.

Phenomenon 38.13 (Time dilation).

A clock moving uniformly with speed \(v\) relative to a frame \(K\) runs slow as judged in \(K\): an interval \(\Delta\tau\) of its proper time corresponds to

\begin{equation}\tag{38.27} \Delta t=\gamma\Delta\tau>\Delta\tau \end{equation}

of \(K\)-coordinate time [Ives:1938] [Rossi:1941]. Rests on Theorem 38.4, Equation (38.3) and Definition 38.12.

Derivation. Derives Phenomenon 38.13. The clock is at rest in \(\bar{K}\), so the two ticks occur at the same \(\bar{x}\) and are separated by \(\Delta\bar{t}=\Delta\tau\). Inverting Equation (38.3) — equivalently, applying the boost with \(-v\) —

\[ \Delta t=\gamma\left(\Delta\bar{t} +\frac{v}{c^{2}}\Delta\bar{x}\right) =\gamma\Delta\bar{t}=\gamma\Delta\tau\ec \]

since \(\Delta\bar{x}=0\). The same result follows from Equation (38.26) directly.

[figure: st-time-dilation.pdf]

Time dilation. The moving clock sits at \(\bar{x}=0\), so its worldline is the \(c\bar{t}\)-axis; at the event \(E\) it has ticked off a proper time \(\tau\). The calibration hyperbola through \(E\) marks every event lying a proper time \(\tau\) from \(O\), in any frame, and it is what makes the comparison meaningful. The height of \(E\) on the \(ct\)-axis is \(\gamma c\tau\), larger than \(c\tau\).

Remark 38.14 (The effect is reciprocal, and that is consistent).

\(\bar{K}\) judges \(K\)'s clocks to run slow by the same factor. The two statements do not contradict each other because each concerns a different pair of events: comparing a single moving clock with two stationary ones requires those two to be synchronized, and Phenomenon 38.10 says the frames disagree about that synchronization by exactly the amount needed. An asymmetry appears only when one clock changes frames, which is not a boost; that is the twin paradox, resolved in Section 40.6. Rests on Phenomena 38.10 and 38.13.

Length contraction

Definition 38.15 (Proper length).

The proper length \(L_{0}\) of a body is its length in the frame in which it is at rest. Its length in any other frame is defined as the distance between its two ends taken at the same time in that frame; the qualification is essential, because by Phenomenon 38.10 the frames disagree about what it means. Rests on Phenomenon 38.10.

Phenomenon 38.16 (Length contraction).

A body of proper length \(L_{0}\) moving with speed \(v\) along its own length has, in the frame through which it moves, the length

\begin{equation}\tag{38.28} L=\frac{L_{0}}{\gamma}<L_{0}\ep \end{equation}

Dimensions perpendicular to the motion are unaffected. Rests on Definition 38.15, Equation (38.2) and Equation (38.4).

Derivation. Derives Phenomenon 38.16. Let the body be at rest in \(\bar{K}\) with its ends at \(\bar{x}_{1}\) and \(\bar{x}_{2}\), so \(L_{0}=\bar{x}_{2}-\bar{x}_{1}\). In \(K\) the ends are located simultaneously, at some common \(t\). Applying Equation (38.2) to each end and subtracting,

\[ L_{0}=\bar{x}_{2}-\bar{x}_{1} =\gamma\left[\left(x_{2}-x_{1}\right) -v\left(t-t\right)\right] =\gamma L\ec \]

which is Equation (38.28). The transverse statement is Equation (38.4).

Figure 38.6 makes the asymmetry with time dilation plain. The body sweeps out a strip in spacetime; the two frames measure its width along different directions, and the strip is what it is. The factor appears in the numerator for lengths and the denominator for times because one measurement is taken along a line of constant \(t\) and the other along a worldline.

[figure: st-length-contraction.pdf]

Length contraction. A rod at rest in \(\bar{K}\) sweeps out the shaded strip between the worldlines of its ends. Its rest length \(L_{0}\) is the interval \(PQ\) measured along the \(\bar{x}\)-axis, between events \(\bar{K}\) calls simultaneous. \(K\) instead measures \(PR\) along its own line of constant \(t\) and obtains \(L_{0}/\gamma\). The rod did not shrink; the two observers measured the width of the same strip in different directions.

Remark 38.17 (Contraction is not what a camera records).

Phenomenon 38.16 concerns simultaneous positions, not appearances. Light reaching an eye or a plate at one instant left different parts of the body at different times, and when that is accounted for a rapidly moving object looks rotated rather than flattened — the Terrell–Penrose effect. Nothing in Equation (38.28) changes; the difference lies in which set of events the measurement collects.

Composition of velocities

Phenomenon 38.18 (Relativistic composition of velocities).

Let a particle have velocity components \(u^{i}=\dd x^{i}/\dd t\) in \(K\), and let \(\bar{K}\) move with speed \(v\) along \(x\). Then in \(\bar{K}\) [Fizeau:1851]

\begin{align} \tag{38.29} \bar{u}^{x} &= \frac{u^{x}-v}{1-u^{x}v/c^{2}}\ec\\ \tag{38.30} \bar{u}^{y} &= \frac{u^{y}}{\gamma\left(1-u^{x}v/c^{2}\right)}\ec \qquad \bar{u}^{z}=\frac{u^{z}}{\gamma\left(1-u^{x}v/c^{2}\right)}\ep \end{align}

Rests on Theorem 38.4, Equation (38.2), Equation (38.3) and Equation (38.4).

Derivation. Derives Phenomenon 38.18. Differentiate Equations (38.2), (38.3) and (38.4):

\[ \dd\bar{x}=\gamma\left(\dd x-v\dd t\right)\ec\qquad \dd\bar{t}=\gamma\left(\dd t-\frac{v}{c^{2}}\dd x\right)\ec\qquad \dd\bar{y}=\dd y\ep \]

Dividing the first by the second, and cancelling \(\gamma\) and \(\dd t\),

\[ \bar{u}^{x}=\frac{\dd\bar{x}}{\dd\bar{t}} =\frac{\dd x-v\dd t}{\dd t-\left(v/c^{2}\right)\dd x} =\frac{u^{x}-v}{1-u^{x}v/c^{2}}\ec \]

which is Equation (38.29). Dividing the third by the second keeps one factor of \(\gamma\) and gives Equation (38.30); note that a transverse velocity changes even though a transverse length does not, because the time interval is transformed.

Corollary 38.19 (The speed of light is a fixed point).

If \(u^{x}=c\) then \(\bar{u}^{x}=c\) for every \(\abs{v}<c\); and if \(\abs{u^{x}}<c\) and \(\abs{v}<c\) then \(\abs{\bar{u}^{x}}<c\). Rests on Equation (38.29).

Proof.

Derives Corollary 38.19. For the first, substitute \(u^{x}=c\) into Equation (38.29):

\[ \bar{u}^{x}=\frac{c-v}{1-v/c} =\frac{c\left(1-v/c\right)}{1-v/c}=c\ep \]

For the second, put \(u^{x}=c\beta_{1}\) and \(v=c\beta_{2}\) with \(\abs{\beta_{i}}<1\). Then

\[ 1-\frac{\bar{u}^{x}}{c} =1-\frac{\beta_{1}-\beta_{2}}{1-\beta_{1}\beta_{2}} =\frac{\left(1-\beta_{1}\right)\left(1+\beta_{2}\right)} {1-\beta_{1}\beta_{2}}>0\ec \]

and symmetrically \(1+\bar{u}^{x}/c>0\), so \(\abs{\bar{u}^{x}}<c\).

Corollary 38.19 is the resolution of the puzzle that motivated the whole construction: no matter how fast a source or an observer moves, light still passes them at \(c\), and no composition of subluminal velocities ever reaches \(c\). In the Galilean limit \(c\to\infty\) the denominators become \(1\) and Equation (38.29) returns \(\bar{u}^{x}=u^{x}-v\).

Example 38.20 (Fizeau's experiment as a first-order effect).

Light travelling through water of refractive index \(n\) moving at speed \(v\) has, in the water's rest frame, the speed \(u'=c/n\). In the laboratory, Equation (38.29) gives

\[ u=\frac{c/n+v}{1+v/(nc)} \simeq\frac{c}{n}+v\left(1-\frac{1}{n^{2}}\right) +O\!\left(\frac{v^{2}}{c}\right)\ec \]

expanding to first order in \(v/c\). The coefficient \(1-1/n^{2}\) is precisely Fresnel's drag coefficient, which Fizeau measured in 1851 (Section 37.4) and which had stood for half a century as an unexplained empirical rule. It is a first-order consequence of Equation (38.29), requiring no aether and no partial dragging of one. Rests on Equation (38.29).

Rapidity

The composition law Equation (38.29) is not addition, which makes repeated boosts awkward. One change of variable repairs this.

Definition 38.21 (Rapidity).

The rapidity \(\varphi\) of a boost is defined by

\begin{equation}\tag{38.31} \beta=\tanh\varphi\ec\qquad\text{whence}\qquad \gamma=\cosh\varphi\ec\qquad \gamma\beta=\sinh\varphi\ep \end{equation}

Rests on Equation (38.1).

The identity \(\cosh^{2}\varphi-\sinh^{2}\varphi=1\) is exactly \(\gamma^{2}-\gamma^{2}\beta^{2}=1\), so the parametrization is consistent, and \(\varphi\) runs over all of \(\R\) as \(\beta\) runs over \((-1,1)\). Substituting into Equation (38.10), the boost becomes

\begin{equation}\tag{38.32} \left[\Lambda\right]= \begin{pmatrix} \cosh\varphi & -\sinh\varphi\\ -\sinh\varphi & \cosh\varphi \end{pmatrix} \end{equation}

in the \((ct,x)\) block — a hyperbolic rotation. It is the exact analogue of the circular rotation

\[ \left[R(\theta)\right]= \begin{pmatrix} \cos\theta & -\sin\theta\\ \sin\theta & \cos\theta \end{pmatrix}\ec \]

with the trigonometric functions replaced by hyperbolic ones. This is the precise sense in which a boost is a rotation in Minkowski space, and it is why the orbits in Figure 38.3 are hyperbolae rather than circles.

Proposition 38.22 (Rapidities add).

Two collinear boosts of rapidities \(\varphi_{1}\) and \(\varphi_{2}\) compose into a single boost of rapidity \(\varphi_{1}+\varphi_{2}\). Rests on Definition 38.21 and Equation (38.32).

Proof.

Derives Proposition 38.22. Multiplying two matrices of the form Equation (38.32) and using the hyperbolic addition formulas,

\[ \begin{pmatrix} \cosh\varphi_{1} & -\sinh\varphi_{1}\\ -\sinh\varphi_{1} & \cosh\varphi_{1} \end{pmatrix} \begin{pmatrix} \cosh\varphi_{2} & -\sinh\varphi_{2}\\ -\sinh\varphi_{2} & \cosh\varphi_{2} \end{pmatrix} =\begin{pmatrix} \cosh\left(\varphi_{1}+\varphi_{2}\right) & -\sinh\left(\varphi_{1}+\varphi_{2}\right)\\ -\sinh\left(\varphi_{1}+\varphi_{2}\right) & \cosh\left(\varphi_{1}+\varphi_{2}\right) \end{pmatrix}\ep \]

Applying \(\tanh\) to the sum reproduces Equation (38.29), since

\[ \tanh\left(\varphi_{1}+\varphi_{2}\right) =\frac{\tanh\varphi_{1}+\tanh\varphi_{2}} {1+\tanh\varphi_{1}\tanh\varphi_{2}}\ep \]

Rapidity is therefore the natural parameter: it is additive, unbounded, and identifies the collinear boosts as a one-parameter group isomorphic to \((\R,+)\). That \(\varphi\) ranges over an unbounded set while \(\beta\) is confined to \((-1,1)\) is the analytic face of the non-compactness proved in Section 39.2.1.

General Lorentz transformations

Boosts composed with rotations

In the most general case two coordinate systems have a constant relative velocity and axes that are not parallel. The coordinates with respect to \(K\) and \(\bar{K}\) are then related by a transformation \(\Lambda\) containing not only the boosts \(\bar{x}^{\mu}=\Lambda^{\mu}{}_{\nu} (\beta)x^{\nu}\) but also three-dimensional rotations.

For a spatial rotation the coordinate transformation is linear and homogeneous, \(\bar{x}^i=R^i_{\ j}(\theta)x^j\) with \(R(\theta)\in\SO(3)\). In the general case we must therefore consider

\begin{equation}\tag{38.33} \bar{x}^{\mu}=\Lambda^{\mu}_{\ \nu}(\beta,\theta)x^{\nu}\ec \end{equation}

where \(\Lambda(0,\theta)\) is a spatial rotation and \(\Lambda(\beta,0)\) is the boost Equation (38.24).

From Equation (38.33),

\begin{align} \bar{x}^0&=\Lambda^0_{\ 0}(0,\theta)x^0 +\Lambda^0_{\ i}(0,\theta)x^i\ec\\ \bar{x}^i&=\Lambda^i_{\ 0}(0,\theta)x^0 +\Lambda^i_{\ j}(0,\theta)x^j\ec \end{align}

and in a spatial rotation the time coordinate is left invariant, \(\bar{x}^0=x^0\). Hence \(\Lambda^0_{\ 0}(0,\theta)=1\) and \(\Lambda^0_{\ i}(0,\theta)=0\); and for the spatial part to be a rotation, \(\Lambda^i_{\ 0}(0,\theta)=0\) and \(\Lambda^i_{\ j}(0,\theta)=R^i_{\ j}(\theta)\). In summary the general case is characterized by

\begin{equation*} \begin{array}{ll} \Lambda^0_{\ 0}(\beta,0)=\gamma & \Lambda^0_{\ 0}(0,\theta)=1\\[2pt] \Lambda^0_{\ i}(\beta,0)=-\gamma\beta_i & \Lambda^0_{\ i}(0,\theta)=0\\[2pt] \Lambda^i_{\ 0}(\beta,0)=-\gamma\beta^i & \Lambda^i_{\ 0}(0,\theta)=0\\[2pt] \Lambda^i_{\ j}(\beta,0)=\delta^i_{\ j} +\frac{\gamma-1}{\abs{\vect{\beta}}^2}\beta^i\beta_j & \Lambda^i_{\ j}(0,\theta)=R^i_{\ j}(\theta)\ep \end{array} \end{equation*}
Theorem 38.23 (Polar decomposition of a Lorentz transformation).

Every proper orthochronous Lorentz transformation (Definition 39.5) factors uniquely as

\begin{equation}\tag{38.36} \Lambda=\Lambda\left(\vect{\beta},0\right)R\left(\theta\right)\ec \end{equation}

a pure boost followed by a spatial rotation. Rests on Definition 39.5, Equation (38.14), Equation (38.20), Equation (38.22) and Proposition 39.6.

Proof.

Derives Theorem 38.23. Existence. The first column \(\Lambda^{\mu}{}_{0}\) is the image of the time direction. Taking \(\mu=\nu=0\) in Equation (38.14) gives \(\left(\Lambda^{0}{}_{0}\right)^{2} -\sum_{i}\left(\Lambda^{i}{}_{0}\right)^{2}=1\), so that column is a unit timelike vector, and it is future-directed because \(\Lambda^{0}{}_{0}\geq1\) for an orthochronous transformation.

Let \(B:=\Lambda(\vect{\beta},0)\) be the pure boost with

\begin{equation}\tag{38.37} \gamma:=\Lambda^{0}{}_{0}\ec\qquad \beta^{i}:=-\frac{\Lambda^{i}{}_{0}}{\Lambda^{0}{}_{0}}\ec \end{equation}

which is admissible since \(\abs{\vect{\beta}}<1\) follows from the unit condition, and which by Equations (38.20) and (38.22) has exactly the same first column as \(\Lambda\).

Put \(R:=B^{-1}\Lambda\), a product of Lorentz transformations and hence Lorentz. Its first column is \(R^{\mu}{}_{0}=\left(B^{-1}\right)^{\mu}{}_{\nu}B^{\nu}{}_{0} =\delta^{\mu}{}_{0}\), so \(R^{0}{}_{0}=1\) and \(R^{i}{}_{0}=0\). Applying the row identity of Proposition 39.6, \(\left(R^{0}{}_{0}\right)^{2}=1+\sum_{i}\left(R^{0}{}_{i}\right)^{2}\), gives \(\sum_{i}\left(R^{0}{}_{i}\right)^{2}=0\) and hence \(R^{0}{}_{i}=0\). So \(R=\diag\left(1,\tilde{R}\right)\) with \(\tilde{R}\tilde{R}\transpose=\identity\), that is \(\tilde{R}\in\Ogrp(3)\); and since \(\det B=1\) — directly from Equation (38.10), \(\gamma^{2}-\gamma^{2}\beta^{2}=1\) — while \(\det\Lambda=1\), also \(\det\tilde{R}=1\) and \(\tilde{R}\in\SO(3)\). Then \(\Lambda=BR\) is the required factorization.

Uniqueness. Suppose \(\Lambda=B'R'\) with \(B'\) a pure boost and \(R'\) a rotation. A rotation fixes the time direction, \(R'^{\nu}{}_{0} =\delta^{\nu}{}_{0}\), so \(\Lambda^{\mu}{}_{0}=B'^{\mu}{}_{\nu}R'^{\nu}{}_{0}=B'^{\mu}{}_{0}\): the boost \(B'\) has the same first column as \(\Lambda\), and a pure boost is determined by that column through Equation (38.37). Hence \(B'=B\) and therefore \(R'=B^{-1}\Lambda=R\).

The counting is \(6=3+3\): three parameters for the boost velocity and three for the rotation, matching the dimension of the Lorentz group computed in Section 39.2.7.

Remark 38.24 (Boosts alone do not form a group).

The decomposition Equation (38.36) is not a formality. The product of two boosts in different directions is not a boost: it is a boost composed with a nontrivial rotation, the Thomas rotation. Boosts are therefore not closed under composition and do not form a subgroup, whereas rotations do. The physical consequence — the Thomas precession of a spin carried around a closed path — is measured in atomic fine structure, where it supplies the factor of \(\tfrac{1}{2}\) in the spin–orbit coupling (The Hydrogen Atom). Rests on Theorem 38.23.

Directions normal to the relative velocity

Equation (38.4) states that lengths perpendicular to the boost are unchanged, but it does not follow that directions are unchanged: an angle mixes a transverse with a longitudinal component, and only one of them transforms.

Proposition 38.25 (Aberration of light).

Let a light ray travel in \(K\) at an angle \(\vartheta\) to the \(x\)-axis. In \(\bar{K}\) its angle \(\bar{\vartheta}\) satisfies

\begin{equation}\tag{38.38} \cos\bar{\vartheta} =\frac{\cos\vartheta-\beta}{1-\beta\cos\vartheta}\ep \end{equation}

Rests on Equation (38.29) and Corollary 38.19.

Proof.

Derives Proposition 38.25. The ray has \(u^{x}=c\cos\vartheta\) and \(u^{y}=c\sin\vartheta\). Substituting into Equation (38.29) and dividing by \(c\),

\[ \cos\bar{\vartheta}=\frac{\bar{u}^{x}}{c} =\frac{c\cos\vartheta-v}{c\left(1-\beta\cos\vartheta\right)} =\frac{\cos\vartheta-\beta}{1-\beta\cos\vartheta}\ec \]

where the speed in \(\bar{K}\) is again \(c\) by Corollary 38.19.

Aberration is why the apparent positions of the fixed stars trace small ellipses over a year as the Earth's velocity swings through \(\pm29.8\,\mathrm{km}/\mathrm{s}\): the effect was discovered by Bradley in 1728, long before its explanation, and its first-order part \(\Delta\vartheta\simeq\beta\sin\vartheta\) is Galilean. The relativistic content of Equation (38.38) lies in the denominator, at second order. At \(\beta\to1\) the formula concentrates almost all directions into a narrow forward cone of half-angle \(\sim1/\gamma\) — the relativistic beaming that makes synchrotron radiation a forward searchlight (Radiation and Scattering of Electromagnetic Waves).