Part II · Special Relativity — Chapter 2.3

Minkowski Geometry

Special relativity is Euclidean geometry with one sign flipped. Everything strange about it is that sign, working.

Where we are

Chapter 2.2 produced a transformation and then a pile of consequences. Simultaneity fails. Moving clocks run slow. Moving rods are short. Velocities do not add, and rapidities do. Each of those was derived, each is correct, and taken together they are a list. Lists are how a subject looks before anyone has found its organising idea.

This chapter finds the organising idea. It is already sitting in 2.2 §2.4, noticed in passing and then left alone: every one of those transformations preserves the combination

Δs2=c2Δt2Δx2Δy2Δz2\Delta s^{2}=c^{2}\Delta t^{2}-\Delta x^{2}-\Delta y^{2}-\Delta z^{2}.

One number, agreed on by everybody. We are going to take that seriously and see what it forces. What it forces is a geometry. By that we mean a four-dimensional space with a notion of distance, in which the Lorentz transformations are the rigid motions, boosts are rotations, worldlines have lengths, and the length of a worldline is the time a clock carried along it reads. After that, the list stops being a list. Time dilation is a statement about lengths. The twin paradox is a statement about triangles. The light barrier is a statement about which directions exist.

Here is the thesis in one sentence, and it is worth reading twice: special relativity is Euclidean geometry with one sign flipped. Nearly every difference between the two subjects traces back to a single minus sign in the distance formula. Why boosts are hyperbolic rather than circular. Why the straight path is the longest and not the shortest. Why some pairs of distinct points are at zero distance. Why there is a speed limit at all. All of those come from that one sign. We will flip it, and then follow it everywhere it goes.

One thing this chapter does not do honestly. It writes xμx^{\mu}, ημν\eta_{\mu\nu} and Δs2=ημνΔxμΔxν\Delta s^{2}=\eta_{\mu\nu}\Delta x^{\mu}\Delta x^{\nu}, and uses index notation informally, without defining what an index means. That debt is real, it is flagged where it is incurred, and Chapter 2.4 pays it in full.

Tools you'll need  — Chapter 2.2: the boost ct=γ(ctβx)ct'=\gamma(ct-\beta x), x=γ(xβct)x'=\gamma(x-\beta ct), the matrix form, and the rapidity ϕ\phi with β=tanhϕ\beta=\tanh\phi. Nothing from 2.2 is re-derived here. Chapter 0.4 §3 and §5: linear maps as matrices, composition as matrix multiplication, and the determinant. Section 1 asks which linear maps preserve a given quadratic expression, which is exactly a 0.4 question. Chapter 0.5 §1: the inner-product axioms, and the observation that once you have a bilinear form you have lengths, angles and orthogonality. This chapter removes one of those axioms on purpose and watches which conclusions survive. Chapter 1.2 §9, worked example 2: extremising ds\int\dd s gives straight lines in the plane and great circles on a sphere. Section 6 runs that machine one more time, in spacetime.

1 · The invariant interval

Everything in this chapter is about two events. Not two objects, and not two particles, but two events, each one a point of spacetime with coordinates (t,x,y,z)(t,x,y,z) in some inertial frame. The single combination of their coordinates that the whole subject turns on is this one:

Δs2    c2Δt2Δx2Δy2Δz2, \Delta s^{2} \;\equiv\; c^{2}\Delta t^{2} - \Delta x^{2} - \Delta y^{2} - \Delta z^{2}, (2.3.1)

Here Δt=t2t1\Delta t = t_{2}-t_{1} and so on. The quantity is called the interval between the two events, or sometimes the spacetime separation.

Three warnings come with it, and all three are paid off before the chapter ends.

  • The notation Δs2\Delta s^{2} is a single symbol. It is not the square of something called Δs\Delta s, and it can come out negative.
  • It is not a distance.
  • The minus signs are not a typographical accident. They are the entire subject.

1.1 · It is the same in every frame

The claim to be proved is that two observers in relative motion, handed the same two events, compute the same Δs2\Delta s^{2}. The proof is direct substitution and it is short. Start from the boost of Chapter 2.2,

ct=γ(ctβx),x=γ(xβct),y=y,z=z, ct' = \gamma\big(ct - \beta x\big), \qquad x' = \gamma\big(x - \beta\,ct\big), \qquad y'=y, \qquad z'=z, (2.3.2)

and apply it to a separation rather than to a single event. That step is legitimate for exactly one reason, and 2.2 §1.2 worked hard to earn it. The transformation is linear, so it acts on differences the same way it acts on coordinates. Our goal is to see what happens to the two-dimensional part c2Δt2Δx2c^{2}\Delta t'^{2}-\Delta x'^{2}, so substitute the boost into it and grind:

c2Δt2Δx2=γ2(cΔtβΔx)2γ2(ΔxβcΔt)2=γ2[c2Δt22βcΔtΔx+β2Δx2]γ2[Δx22βcΔtΔx+β2c2Δt2]=γ2(1β2)c2Δt2    γ2(1β2)Δx2=c2Δt2Δx2, \begin{aligned} c^{2}\Delta t'^{2}-\Delta x'^{2} &= \gamma^{2}\big(c\Delta t-\beta\Delta x\big)^{2} - \gamma^{2}\big(\Delta x-\beta\,c\Delta t\big)^{2}\\[4pt] &= \gamma^{2}\Big[c^{2}\Delta t^{2} - 2\beta\,c\Delta t\,\Delta x + \beta^{2}\Delta x^{2}\Big]\\[2pt] &\qquad - \gamma^{2}\Big[\Delta x^{2} - 2\beta\,c\Delta t\,\Delta x + \beta^{2}c^{2}\Delta t^{2}\Big]\\[4pt] &= \gamma^{2}\big(1-\beta^{2}\big)\,c^{2}\Delta t^{2} \;-\; \gamma^{2}\big(1-\beta^{2}\big)\,\Delta x^{2}\\[4pt] &= c^{2}\Delta t^{2}-\Delta x^{2}, \end{aligned} (2.3.3)

The last line used γ2(1β2)=1\gamma^{2}(1-\beta^{2})=1, which is the definition of γ\gamma rearranged.

Look at what became of the cross terms. They cancelled identically, and they had to. They enter the two brackets with the same coefficient and with opposite overall sign, so nothing else was available to them.

The other two directions cost us nothing, because Δy=Δy\Delta y'=\Delta y and Δz=Δz\Delta z'=\Delta z. We may therefore subtract those two squares from each side without disturbing the equality, and that gives the full four-dimensional statement:

  c2Δt2Δx2Δy2Δz2  =  c2Δt2Δx2Δy2Δz2.   \boxed{\;c^{2}\Delta t'^{2}-\Delta x'^{2}-\Delta y'^{2}-\Delta z'^{2} \;=\; c^{2}\Delta t^{2}-\Delta x^{2}-\Delta y^{2}-\Delta z^{2}.\;} (2.3.4)

Every inertial observer, handed the same two events, computes the same Δs2\Delta s^{2}. They will disagree about Δt\Delta t. They will disagree about Δx\Delta x. They will disagree about which event came first, sometimes. They agree about this.

Chapter 2.2 reached this same equation in its §2.4 and treated it as a consistency check. We had demanded that light go at cc, and the check confirmed that a spherical light pulse stays spherical. That reading undersells the result badly, and it is worth being precise about why.

Postulate 2 constrains only the pairs of events with Δs2=0\Delta s^{2}=0, which is one three-dimensional family out of all the pairs there are. What came back out is a statement about every pair of events, at every value of Δs2\Delta s^{2}. The theory handed us more invariance than we asked for. A quantity that turns out to be invariant for no reason you demanded is exactly the kind of quantity worth building a subject on.

1.2 · The reframe: this is what "length" means

Now comes the move that makes the rest of the book easier. Let's ask what a rotation is.

You probably answer "a rigid turn about an axis". That is a picture rather than a definition. The definition used in mathematics is algebraic, and it is the one Chapter 0.4 §5 was quietly preparing you for:

rotation of R3 is a linear map that preserves x2+y2+z2. \text{a \textbf{rotation} of }\R^{3}\ \text{is a linear map that preserves}\ x^{2}+y^{2}+z^{2}. (2.3.5)

That is the whole content. You do not need to mention turning, or axes, or angles. Demand that a linear map leave x2+y2+z2x^{2}+y^{2}+z^{2} unchanged for every vector, and rotations together with reflections are what you get. They form the group O(3)\mathrm{O}(3), and requiring the determinant to be +1+1 throws away the reflections.

Notice which way the logic runs. It runs from the quadratic form to the transformations, and not the other way round.

Now notice what the preserved quantity is. It is what we call length. That is not a separate fact about the world sitting alongside the definition. It is the definition of length: the thing on which all observers with differently-oriented axes agree.

Read (2.3.4) again with that in mind. It says:

The whole chapter, in one substitution

Lorentz transformations are the linear maps that preserve Δs2\Delta s^{2}.

So Δs2\Delta s^{2} plays the role that squared length plays in Euclidean geometry, and the Lorentz transformations play the role that rotations play. Change the quadratic form from x2+y2+z2x^{2}+y^{2}+z^{2} to c2t2x2y2z2c^{2}t^{2}-x^{2}-y^{2}-z^{2}, and the group of "rigid motions" changes from rotations to boosts-and-rotations. Nothing else about the logic changes at all.

That is not an analogy. It is the same definition with a different quadratic form in it.

Two remarks on how much that box is claiming.

The first is that a converse is involved. What we proved above is that Lorentz transformations preserve Δs2\Delta s^{2}. The statement that the linear maps preserving Δs2\Delta s^{2} are exactly the Lorentz transformations runs the other way, and converses need proof. The grind box below proves it in the case that matters, which is one space dimension. It also finds three extra maps hiding in the answer, which turn out to be parity, time reversal, and the two applied together.

The second is that the word "geometry" is a promise. Calling this a geometry commits us to lengths, angles, straight lines and triangles behaving like a geometry. Sections 3 to 6 make good on each of those. Section 6 finds that one of them behaves like a geometry with the inequality reversed. That is the sign flip, arriving where it hurts.

Grind box — the converse: every linear map preserving c2t2x2c^{2}t^{2}-x^{2}, classified

Work in the (ct,x)(ct,x) plane, which is where all the content lives. Write the unknown map as a general 2×22\times2 matrix and let the invariance condition tell us what its entries have to be. So let

M=(abde),(ctx)=M(ctx), M = \begin{pmatrix} a & b\\ d & e\end{pmatrix}, \qquad \begin{pmatrix}ct'\\ x'\end{pmatrix} = M\begin{pmatrix}ct\\ x\end{pmatrix},

and demand c2t2x2=c2t2x2c^{2}t'^{2}-x'^{2}=c^{2}t^{2}-x^{2} for all (ct,x)(ct,x). The word "all" is what makes this usable, because two quadratics that agree everywhere must agree coefficient by coefficient. So expand the left side and see what the coefficients are:

(act+bx)2(dct+ex)2=(a2d2)c2t2+2(abde)ctx+(b2e2)x2. \big(a\,ct+b\,x\big)^{2}-\big(d\,ct+e\,x\big)^{2} = \big(a^{2}-d^{2}\big)c^{2}t^{2} + 2\big(ab-de\big)ct\,x + \big(b^{2}-e^{2}\big)x^{2}.

Matching coefficients of the three independent monomials gives three equations:

a2d2=1,abde=0,b2e2=1. a^{2}-d^{2}=1, \qquad ab-de=0, \qquad b^{2}-e^{2}=-1.

Solve the first. We have a2=1+d21a^{2}=1+d^{2}\geq1, so a1\abs a\geq1, and any number of modulus at least one can be written as a cosh\cosh. Put a=σcoshϕa=\sigma\cosh\phi with σ=±1\sigma=\pm1 and ϕ0\phi\geq0. Then d2=sinh2ϕd^{2}=\sinh^{2}\phi, and by allowing ϕ\phi to take either sign we can write d=σsinhϕd=-\sigma\sinh\phi without loss of generality.

Solve the third. The same argument applies, since e2=1+b21e^{2}=1+b^{2}\geq1 likewise, so e=τcoshψe=\tau\cosh\psi and b=τsinhψb=-\tau\sinh\psi with τ=±1\tau=\pm1. Notice that we have had to allow a second angle ψ\psi here. The middle equation is what will force it to equal the first.

Use the second. Substituting the four entries into abde=0ab-de=0,

abde=στcoshϕsinhψ+στsinhϕcoshψ=στsinh(ϕψ)=0, ab-de = -\sigma\tau\cosh\phi\sinh\psi + \sigma\tau\sinh\phi\cosh\psi = \sigma\tau\,\sinh(\phi-\psi) = 0,

so ψ=ϕ\psi=\phi, since sinh\sinh vanishes only at zero. Therefore

M=(σcoshϕτsinhϕσsinhϕτcoshϕ)=(σ00τ)(coshϕsinhϕsinhϕcoshϕ). M = \begin{pmatrix} \sigma\cosh\phi & -\tau\sinh\phi\\ -\sigma\sinh\phi & \tau\cosh\phi\end{pmatrix} = \begin{pmatrix}\sigma&0\\0&\tau\end{pmatrix}\begin{pmatrix}\cosh\phi & -\sinh\phi\\ -\sinh\phi & \cosh\phi\end{pmatrix}.

Read the factorisation on the right. The second factor is exactly Chapter 2.2's boost in rapidity form. The first factor is one of four sign choices: σ=τ=+1\sigma=\tau=+1 is the identity, σ=+1,τ=1\sigma=+1,\tau=-1 is parity xxx\to-x, σ=1,τ=+1\sigma=-1,\tau=+1 is time reversal ttt\to-t, and σ=τ=1\sigma=\tau=-1 is both. Note detM=στ\det M=\sigma\tau.

The conclusion. The linear maps preserving c2t2x2c^{2}t^{2}-x^{2} are precisely the boosts, possibly composed with parity and/or time reversal. So the converse holds, up to those discrete factors. That is the same structure as in Euclidean geometry, where preserving x2+y2x^{2}+y^{2} gives rotations possibly composed with a reflection.

The group therefore falls into four disconnected pieces. The piece containing the identity is the one with σ=τ=+1\sigma=\tau=+1, the pure boosts, and it is the piece physics uses without comment. It has a name, the proper orthochronous Lorentz group. Proper means det=+1\det=+1, and orthochronous means it does not reverse the direction of time. ⚑ That the same four-component structure survives in the full four-dimensional group, with rotations included, is quoted here and unpacked in Chapter 6.1.

A detail worth noticing. Nothing in this derivation used a postulate, a light ray, or a physical assumption. Given the quadratic form, the transformations are pure algebra. All the physics went into choosing c2t2x2c^{2}t^{2}-x^{2} rather than c2t2+x2c^{2}t^{2}+x^{2}, and Chapter 2.2 is where that choice was forced.

1.3 · The sign convention, and the one half the world uses instead

Nothing above depended on whether we write c2Δt2Δx2c^{2}\Delta t^{2}-\Delta x^{2} or Δx2c2Δt2\Delta x^{2}-c^{2}\Delta t^{2}. The two differ by an overall factor of 1-1, they are preserved by exactly the same transformations, and every physical statement can be phrased in either. Books nevertheless choose, and then never mention it again, which is how readers get ambushed.

This book uses the timelike-positive convention, often written (+,,,)(+,-,-,-) and sometimes called the "mostly minus" or "particle physics" signature:

Δs2=c2Δt2Δx2Δy2Δz2,so a clock’s own path has Δs2>0. \Delta s^{2} = c^{2}\Delta t^{2}-\Delta x^{2}-\Delta y^{2}-\Delta z^{2}, \qquad\text{so a clock's own path has } \Delta s^{2}\gt0. (2.3.6)

The reason is Parts V to VII. In particle physics the object you write down forty times a day is the four-momentum, and its invariant square in this convention is pp=m2c2p\cdot p = m^{2}c^{2}. That is mass squared: positive, with no sign to remember. In the other convention it is m2c2-m^{2}c^{2}, and every mass-shell condition in the book would carry a minus sign for no benefit. Since Parts V–VII are much longer than Part III, we optimise for them.

The other convention, (,+,+,+)(-,+,+,+) or "mostly plus", is standard in general relativity, where the object you write forty times a day is a spatial line element and you would rather it came out positive. If you open a relativity textbook and find ds2=c2dt2+dx2\dd s^{2}=-c^{2}\dd t^{2}+\dd x^{2}, nothing is wrong. Multiply by 1-1 and carry on.

Here is what the choice does and does not touch. What changes: the sign of Δs2\Delta s^{2} classifying timelike versus spacelike, the sign in ppp\cdot p, and the sign of ημν\eta_{\mu\nu} itself. What does not change: a single physical prediction. It is a units-style choice, like measuring angles in degrees, and it carries the same amount of physics.

In plain terms 2.3.1

Out of a demand made about one narrow family of events came an equality holding for every pair of them, and that overshoot is what makes a geometry available. The combination of time separation and space separation that all observers compute alike has a name, the interval, and the way to see what it is doing is to ask first what a rotation is.

The answer taught first is a rigid turn about an axis, which is a picture rather than a definition. The definition is algebraic: a rotation is a linear map leaving the sum of the squared coordinates alone. Nothing about turning is needed, and the logic runs from the quadratic expression to the transformations rather than the other way. The quantity those transformations preserve is what the word length means, not a separate fact about the world but the name for whatever observers with differently oriented axes agree on.

Read the invariance again with that in mind, and the whole chapter is one substitution. Flip the sign of the three spatial terms in the quadratic expression, and the rigid motions change from rotations into the transformations of relativity. That is not an analogy; it is the same definition with a different expression inside it. The converse holds too, since the maps preserving the flipped expression are exactly the boosts, up to a reflection of space or of time.

2 · Coordinates, the metric, and one line of index notation

We are about to need a compact way to write (2.3.1) that does not degenerate into four-fold bookkeeping every time. The notation that does this is one of the great labour-saving devices in physics. We introduce it here informally, and label it honestly as informal, because §3 onward reads a great deal better in it.

2.1 · One symbol for four coordinates

Group the four coordinates of an event into a single object with an index:

xμ  =  (x0,x1,x2,x3)  =  (ct,x,y,z). x^{\mu} \;=\; \big(x^{0},\,x^{1},\,x^{2},\,x^{3}\big) \;=\; \big(ct,\,x,\,y,\,z\big). (2.3.7)

Greek indices μ,ν,ρ,σ\mu,\nu,\rho,\sigma run over 0,1,2,30,1,2,3; Latin indices i,j,ki,j,k over the spatial 1,2,31,2,3 only. The time coordinate is x0=ctx^{0}=ct and not tt, so that all four entries have the dimensions of length. That is not cosmetic: Chapter 2.2 §2.2 already showed that ctct and xx enter the boost symmetrically, and writing tt instead would hide the symmetry behind a units conversion.

One warning about the superscript. It is an index, not a power. So x2x^{2} is the yy-coordinate. Where we mean the square of something, we will write it so that the context is unmissable. This collision of notation is genuinely annoying, and it is also universal, and everyone lives with it.

2.2 · The metric, and the summation convention

Now define a 4×44\times4 array of numbers

ημν  =  diag(1,1,1,1)  =  (1000010000100001), \eta_{\mu\nu} \;=\; \mathrm{diag}(1,-1,-1,-1) \;=\; \begin{pmatrix}1&0&0&0\\ 0&-1&0&0\\ 0&0&-1&0\\ 0&0&0&-1\end{pmatrix}, (2.3.8)

called the Minkowski metric. It carries the minus signs of (2.3.1) and does nothing else. With it, the interval is

Δs2  =  μ=03ν=03ημνΔxμΔxν    ημνΔxμΔxν. \Delta s^{2} \;=\; \sum_{\mu=0}^{3}\sum_{\nu=0}^{3}\eta_{\mu\nu}\,\Delta x^{\mu}\,\Delta x^{\nu} \;\equiv\; \eta_{\mu\nu}\,\Delta x^{\mu}\Delta x^{\nu}. (2.3.9)

The second form uses the Einstein summation convention: an index appearing twice in a single term, once up and once down, is summed over its whole range, and the summation sign is not written. It is a convention about typography and nothing more. The saving is nonetheless substantial. The double sum has sixteen terms, of which twelve vanish because η\eta is diagonal, and you never want to write that out twice.

Before leaning on the shorthand, let's check that (2.3.9) really does say what it should. Since η\eta is diagonal, the only surviving terms are the ones with μ=ν\mu=\nu, so write those four out:

ημνΔxμΔxν=(+1)(Δx0)2+(1)(Δx1)2+(1)(Δx2)2+(1)(Δx3)2, \eta_{\mu\nu}\Delta x^{\mu}\Delta x^{\nu} = (+1)\big(\Delta x^{0}\big)^{2} + (-1)\big(\Delta x^{1}\big)^{2}+(-1)\big(\Delta x^{2}\big)^{2}+(-1)\big(\Delta x^{3}\big)^{2}, (2.3.10)

That is c2Δt2Δx2Δy2Δz2c^{2}\Delta t^{2}-\Delta x^{2}-\Delta y^{2}-\Delta z^{2} ✓. So the metric is doing exactly one job at this stage: supplying a sign for each term.

In ordinary three-dimensional space the corresponding array is δij=diag(1,1,1)\delta_{ij}=\mathrm{diag}(1,1,1), which supplies all plus signs. That is why nobody ever writes it, and why the metric is invisible in first-year physics. The whole of the difference between Euclidean geometry and Minkowski geometry is the difference between δij\delta_{ij} and ημν\eta_{\mu\nu}.

Grind box — the summation convention, written out once in full, and the rules for using it

The unabbreviated version of (2.3.9), all sixteen terms, with Δxμ=(a,b,c,d)\Delta x^{\mu}=(a,b,c,d) to avoid symbol collisions:

ημνΔxμΔxν=  η00aa+η01ab+η02ac+η03ad+  η10ba+η11bb+η12bc+η13bd+  η20ca+η21cb+η22cc+η23cd+  η30da+η31db+η32dc+η33dd=  a2b2c2d2. \begin{aligned} \eta_{\mu\nu}\Delta x^{\mu}\Delta x^{\nu} =\;& \eta_{00}aa + \eta_{01}ab + \eta_{02}ac + \eta_{03}ad\\ +\;& \eta_{10}ba + \eta_{11}bb + \eta_{12}bc + \eta_{13}bd\\ +\;& \eta_{20}ca + \eta_{21}cb + \eta_{22}cc + \eta_{23}cd\\ +\;& \eta_{30}da + \eta_{31}db + \eta_{32}dc + \eta_{33}dd\\[4pt] =\;& a^{2}-b^{2}-c^{2}-d^{2}. \end{aligned}

Three rules, all of which will be justified properly in Chapter 2.4 and all of which you should obey immediately.

(1) A repeated index is dead. In ημνΔxμΔxν\eta_{\mu\nu}\Delta x^{\mu}\Delta x^{\nu} the labels μ\mu and ν\nu have been summed away, so the result carries no index and does not depend on what you called them. Renaming μα\mu\to\alpha throughout changes nothing. Such indices are called dummy indices, and they behave exactly like the kk in kak\sum_{k}a_{k}.

(2) An unrepeated index is alive, and must match on both sides. Vμ=ΛμνWνV^{\mu}=\Lambda^{\mu}{}_{\nu}W^{\nu} is a legal equation: ν\nu is summed, μ\mu is free and appears once on each side, so the equation is four equations. By contrast Vμ=WνV^{\mu}=W^{\nu} is not an equation at all, because it does not say which component equals which.

(3) Never use the same letter three times. ημμΔxμ\eta_{\mu\mu}\Delta x^{\mu} is ill-formed: the convention cannot tell which pair you meant to sum. If you need two separate sums, use two separate letters. This rule catches more algebra errors than any other.

What is not being claimed here. We have written some indices up (xμx^{\mu}) and some down (ημν\eta_{\mu\nu}), and given no reason. There is a reason, it is important, and it is not that one is a row and the other a column. Chapter 2.4 §3 gives it. Until then, treat the placement as a spelling rule you are obeying on trust, and note that the trust is being tracked: every expression in this chapter has each summed index appearing once up and once down, and you may check that it does.

2.3 · What a four-vector is, provisionally

We need a working notion of "vector" in spacetime, and there is one object whose behaviour under a boost we already know completely. The displacement between two nearby events, dxμ=(cdt,dx,dy,dz)\dd x^{\mu}=(c\,\dd t,\dd x,\dd y,\dd z), transforms in exactly the way (2.3.2) says it does. We take that as our template:

Working definition — four-vector

A four-vector is a set of four quantities VμV^{\mu}, one per frame, that transform between frames the same way dxμ\dd x^{\mu} does:

Vμ  =  ΛμνVν, V'^{\mu} \;=\; \Lambda^{\mu}{}_{\nu}\,V^{\nu},

with Λμν\Lambda^{\mu}{}_{\nu} the same matrix that transforms coordinate displacements.

This is deliberately informal. It says "these ones, like that one", which is a definition by example, and it leaves at least three questions unanswered. Why are some indices up and others down? What is the object VμV^{\mu}, as opposed to what its components do? Does the definition depend on the transformation being a Lorentz transformation? Chapter 2.4 answers all three and turns this paragraph into a theorem. Nothing in this chapter needs the answers. Everything in this chapter needs the template.

Why bother now? Because the payoff arrives immediately and repeatedly. If VμV^{\mu} and WμW^{\mu} are both four-vectors, then the combination

VW    ημνVμWν  =  V0W0V1W1V2W2V3W3 V\cdot W \;\equiv\; \eta_{\mu\nu}V^{\mu}W^{\nu} \;=\; V^{0}W^{0}-V^{1}W^{1}-V^{2}W^{2}-V^{3}W^{3} (2.3.11)

is a single number that every observer agrees on. The reason is exactly the reason Δs2\Delta s^{2} is agreed on, since Δs2\Delta s^{2} is the special case V=W=ΔxV=W=\Delta x.

What we have built is a machine for manufacturing invariants, and we will keep feeding it: Δxμ\Delta x^{\mu} in this chapter, the four-velocity uμu^{\mu} in §7, the four-momentum pμp^{\mu} in Chapter 2.5 (where pp=m2c2p\cdot p=m^{2}c^{2} turns out to be the whole of relativistic kinematics), and the four-current in Chapter 2.6.

2.4 · The condition on Λ\Lambda

One more line before we start using any of this, and it is the line Chapter 2.4 and Chapter 6.1 both build on. What we want is the condition a matrix has to satisfy in order to be a Lorentz transformation at all. So write the transformation in matrix form, as Chapter 2.2 §3 did, and demand invariance of the interval. In index notation:

ημνΔxμΔxν  =  ημνΛμρΛνσΔxρΔxσ  =!  ηρσΔxρΔxσ. \eta_{\mu\nu}\,\Delta x'^{\mu}\Delta x'^{\nu} \;=\; \eta_{\mu\nu}\,\Lambda^{\mu}{}_{\rho}\Lambda^{\nu}{}_{\sigma}\,\Delta x^{\rho}\Delta x^{\sigma} \;\overset{!}{=}\; \eta_{\rho\sigma}\,\Delta x^{\rho}\Delta x^{\sigma}. (2.3.12)

Since the separation Δxρ\Delta x^{\rho} is arbitrary and both sides are symmetric in ρσ\rho\leftrightarrow\sigma, the coefficients must match term by term:

  ημνΛμρΛνσ  =  ηρσ,that isΛTηΛ=η.   \boxed{\;\eta_{\mu\nu}\,\Lambda^{\mu}{}_{\rho}\,\Lambda^{\nu}{}_{\sigma} \;=\; \eta_{\rho\sigma}\,,\qquad\text{that is}\qquad \Lambda^{\mathsf T}\eta\,\Lambda = \eta.\;} (2.3.13)

Let's look at what that boxed line is saying. Compare ΛTηΛ=η\Lambda^{\mathsf T}\eta\Lambda=\eta with the condition defining a rotation matrix, RT1R=1R^{\mathsf T}\,\mathbb{1}\,R=\mathbb{1}, that is, RTR=1R^{\mathsf T}R=\mathbb{1}. Same equation, one symbol different.

(2.3.13) is the defining equation of the Lorentz group, written O(1,3)\mathrm{O}(1,3), where the notation records that the metric has one plus and three minuses. Chapter 6.1 will take it as the starting point rather than as the conclusion, differentiate it near the identity, and read off the six generators. Problem 4 does the counting by hand and gets six.

Two consequences drop out of that condition without any further work. Here is the first. We want to know what the condition says about the overall scale of Λ\Lambda, so take determinants of (2.3.13) and use det(AB)=detAdetB\det(AB)=\det A\det B from Chapter 0.4 §5:

(detΛ)2detη=detηdetΛ=±1, \big(\det\Lambda\big)^{2}\det\eta = \det\eta \qquad\Longrightarrow\qquad \det\Lambda = \pm1, (2.3.14)

So Lorentz transformations preserve four-dimensional volume in spacetime, up to sign. Chapter 2.2 §3 checked det=+1\det=+1 for the boost explicitly.

The second consequence is the group property. The set of matrices satisfying (2.3.13) is closed under multiplication and inversion. Chapter 2.2 §3.1 verified that by brute force for boosts along one axis, and here it takes one line: (Λ2Λ1)Tη(Λ2Λ1)=Λ1T(Λ2TηΛ2)Λ1=Λ1TηΛ1=η(\Lambda_{2}\Lambda_{1})^{\mathsf T}\eta(\Lambda_{2}\Lambda_{1}) = \Lambda_{1}^{\mathsf T}(\Lambda_{2}^{\mathsf T}\eta\Lambda_{2})\Lambda_{1}=\Lambda_{1}^{\mathsf T}\eta\Lambda_{1}=\eta.

Grind box — checking ΛTηΛ=η\Lambda^{\mathsf T}\eta\Lambda=\eta on the boost, entry by entry

Only the upper-left 2×22\times2 block does anything, so drop yy and zz and take η=diag(1,1)\eta=\mathrm{diag}(1,-1) with

Λ=(γγβγβγ)=(coshϕsinhϕsinhϕcoshϕ). \Lambda = \begin{pmatrix}\gamma & -\gamma\beta\\ -\gamma\beta & \gamma\end{pmatrix} = \begin{pmatrix}\cosh\phi & -\sinh\phi\\ -\sinh\phi & \cosh\phi\end{pmatrix}.

The boost matrix is symmetric, so ΛT=Λ\Lambda^{\mathsf T}=\Lambda. Take the triple product in two stages and compute ηΛ\eta\Lambda first. Multiplying on the left by diag(1,1)\mathrm{diag}(1,-1) flips the sign of the second row:

ηΛ=(coshϕsinhϕsinhϕcoshϕ). \eta\Lambda = \begin{pmatrix}\cosh\phi & -\sinh\phi\\ \sinh\phi & -\cosh\phi\end{pmatrix}.

Now multiply that result on the left by ΛT\Lambda^{\mathsf T}, which is Λ\Lambda itself:

ΛTηΛ=(coshϕsinhϕsinhϕcoshϕ)(coshϕsinhϕsinhϕcoshϕ)=(cosh2ϕsinh2ϕcoshϕsinhϕ+sinhϕcoshϕsinhϕcoshϕ+coshϕsinhϕsinh2ϕcosh2ϕ)=(1001)  =  η \begin{aligned} \Lambda^{\mathsf T}\eta\Lambda &= \begin{pmatrix}\cosh\phi & -\sinh\phi\\ -\sinh\phi & \cosh\phi\end{pmatrix}\begin{pmatrix}\cosh\phi & -\sinh\phi\\ \sinh\phi & -\cosh\phi\end{pmatrix}\\[6pt] &= \begin{pmatrix}\cosh^{2}\phi-\sinh^{2}\phi & -\cosh\phi\sinh\phi+\sinh\phi\cosh\phi\\ -\sinh\phi\cosh\phi+\cosh\phi\sinh\phi & \sinh^{2}\phi-\cosh^{2}\phi\end{pmatrix}\\[6pt] &= \begin{pmatrix}1&0\\0&-1\end{pmatrix} \;=\; \eta \quad\checkmark \end{aligned}

The diagonal entries are the hyperbolic Pythagorean identity. The off-diagonal ones cancel because the matrix is symmetric. Note which identity did the work: cosh2sinh2=1\cosh^{2}-\sinh^{2}=1. In the Euclidean case the same computation with η1\eta\to\mathbb{1} and ΛR(θ)\Lambda\to R(\theta) runs on cos2+sin2=1\cos^{2}+\sin^{2}=1 instead. One sign, two geometries, and that is what §3 is about.

Familiar ground — the interval is a Mahalanobis distance with one sign changed

You already compute quadratic forms with a matrix in the middle, and you compute them for exactly the reason (2.3.9) exists. Given a panel of measurements with covariance Σ\Sigma, the Mahalanobis distance of an observation from the mean is

d2  =  (xμ) ⁣Σ1(xμ)  =  (Σ1)ijΔxiΔxj, d^{2} \;=\; (x-\mu)^{\!\top}\Sigma^{-1}(x-\mu) \;=\; \big(\Sigma^{-1}\big)_{ij}\,\Delta x^{i}\,\Delta x^{j},

written on the right in this section's notation. Its entire purpose is to be a number that does not depend on the units the individual features were recorded in, nor on how strongly they happen to be correlated: one agreed answer, assembled from components every laboratory writes down differently. That is the paragraph above (2.3.9), word for word, with Σ1\Sigma^{-1} standing where ημν\eta_{\mu\nu} stands.

The parallel runs past the formula. Chapter 0.5 §6 guarantees that Σ1\Sigma^{-1}, being real and symmetric, is diagonalised by an orthogonal change of basis. That change of basis is the whitening transform, after which d2d^{2} is a plain sum of squares and the maps preserving it are ordinary rotations. Here the same manoeuvre stops one step short. (2.3.8) is already diagonal, and no real change of basis can turn its diagonal into four plus signs, because the number of minus signs is an invariant of the form and not an artefact of the basis.

What that one sign buys is the whole of §4. Σ1\Sigma^{-1} is positive definite, so d20d^{2}\ge0 always, with equality only for an observation sitting exactly at the mean. There is nothing there to classify and no structure to find. η\eta is not positive definite, so Δs2\Delta s^{2} comes out positive, negative or exactly zero, its zero set is a whole cone rather than a single point, and which of the three a given pair of events yields is the causal structure of the world.

The same fact can be stated through the transformations. Those preserving d2d^{2} form a compact family, so no amount of whitened rotating takes you far. Those preserving Δs2\Delta s^{2} do not, which is why γ\gamma has no upper bound while a correlation coefficient has two.

Two honest differences. Σ\Sigma is estimated from a sample and is a property of a population, whereas η\eta is neither estimated nor a property of anything anybody can vary. And a Mahalanobis distance runs from a point to a distribution, while the interval runs between two events, with no distribution anywhere in it.

In plain terms 2.3.2

Notation deserves attention when it makes a structure visible rather than merely shorter, and two decisions do that here. The first is to measure time in the same units as distance, so that two coordinates entering the transformation on an equal footing are seen to do so instead of being held apart by a conversion factor. The second is to collect the four coordinates of an event under one symbol carrying a label.

The signs then need somewhere to live, and they live in a small square array with one plus and three minuses down its diagonal, called the metric. Its entire job at this stage is to supply a sign to each term. In ordinary space the corresponding array is all plus signs, which is why nobody ever writes it and why the metric is invisible in first-year physics. The whole difference between the two geometries is the difference between those two arrays.

What the arrangement buys is a machine for manufacturing agreement. Take any two collections of four numbers that transform between frames the way a displacement between events does, combine them through the metric, and the number resulting is one every observer computes alike. The interval is the special case where both collections are the same displacement, and the machine will be fed repeatedly: a velocity shortly, a momentum in the next chapter, a current after that.

3 · Boosts are hyperbolic rotations

This is the section the chapter exists for. Everything before it was setup and everything after it is consequence.

3.1 · Two matrices, side by side

Put the Euclidean rotation and the Lorentz boost next to each other and read across.

A rotation in the xyxy plane preserves x2+y2x^{2}+y^{2} and is built from a pair of functions satisfying cos2θ+sin2θ=1\cos^{2}\theta+\sin^{2}\theta=1:

(xy)=(cosθsinθsinθcosθ)(xy),x2+y2=x2+y2. \begin{pmatrix}x'\\ y'\end{pmatrix} = \begin{pmatrix}\cos\theta & \sin\theta\\ -\sin\theta & \cos\theta\end{pmatrix}\begin{pmatrix}x\\ y\end{pmatrix}, \qquad x'^{2}+y'^{2}=x^{2}+y^{2}. (2.3.15)

Now the boost, in the form Chapter 2.2 §6.1 put it into. A boost in the (ct,x)(ct,x) plane preserves c2t2x2c^{2}t^{2}-x^{2}, and it is built from a pair of functions satisfying cosh2ϕsinh2ϕ=1\cosh^{2}\phi-\sinh^{2}\phi=1:

(ctx)=(coshϕsinhϕsinhϕcoshϕ)(ctx),c2t2x2=c2t2x2. \begin{pmatrix}ct'\\ x'\end{pmatrix} = \begin{pmatrix}\cosh\phi & -\sinh\phi\\ -\sinh\phi & \cosh\phi\end{pmatrix}\begin{pmatrix}ct\\ x\end{pmatrix}, \qquad c^{2}t'^{2}-x'^{2}=c^{2}t^{2}-x^{2}. (2.3.16)

The parallel is not loose, so let's set out the three points of contact. Each matrix is built from the two functions whose squares obey the identity matching its own quadratic form. Each preserves that form. Each is parametrised by one real number. And in both cases the parameter is additive:

R(θ2)R(θ1)=R(θ1+θ2),Λ(ϕ2)Λ(ϕ1)=Λ(ϕ1+ϕ2), R(\theta_{2})R(\theta_{1})=R(\theta_{1}+\theta_{2}), \qquad\qquad \Lambda(\phi_{2})\Lambda(\phi_{1})=\Lambda(\phi_{1}+\phi_{2}), (2.3.17)

The first of those follows from the circular addition formulas and the second from the hyperbolic ones, and Chapter 2.2 §6.2 multiplied the matrices out to check. That additivity deserves a name.

The rapidity is the angle

The rapidity ϕ\phi of Chapter 2.2 is not merely "the variable in which velocities add". It is the hyperbolic angle of the rotation, and it adds for exactly the reason Euclidean angles add: composing two rotations of the same kind about the same axis adds their angles, because that is what a one-parameter group of rotations does.

Chapter 2.2 arrived at ϕ\phi by noticing that β3=(β1+β2)/(1+β1β2)\beta_{3}=(\beta_{1}+\beta_{2})/(1+\beta_{1}\beta_{2}) is the tanh\tanh addition theorem and taking the hint. Here the hint is explained. Velocity is β=tanhϕ\beta=\tanh\phi, and tanh\tanh of an angle is a slope. Slopes never add. Angles do. Three centuries of physics added the slopes.

Two differences are worth flagging before we exploit the parallel, because they are the places where the two geometries part company.

The signs in the off-diagonal entries. The rotation matrix has +sin+\sin above the diagonal and sin-\sin below. The boost has sinh-\sinh in both places. So the boost matrix is symmetric and the rotation is not. That is not cosmetic. Chapter 2.2's Problem 3 used precisely this fact to prove that two boosts in different directions leave a rotation behind.

The range of the parameter. The angle θ\theta is periodic. Rotate by 2π2\pi and you are back where you started, and the orbit of a point closes. The rapidity ϕ\phi is not periodic and runs over all of R\R, so you may boost forever and never return. In the same way cos\cos and sin\sin are bounded, while cosh\cosh and sinh\sinh are not.

That last contrast has a name. The rotation group is compact and the boosts are not, and this is the deepest structural difference between them. It is responsible for a great deal in Part VI. It is also the reason γ=coshϕ\gamma=\cosh\phi has no upper bound while a rotation never stretches anything.

3.2 · The orbits: circles become hyperbolae

Now ask the question that turns algebra into a picture. Take one point and apply every transformation in the family. What curve does it sweep out?

Euclidean. Start at (x0,y0)(x_{0},y_{0}) and rotate through every θ\theta. Since x2+y2x^{2}+y^{2} is preserved, every image point satisfies x2+y2=r2x^{2}+y^{2}=r^{2} with r2=x02+y02r^{2}=x_{0}^{2}+y_{0}^{2} fixed. The orbit is a circle. And every point of that circle is reached, since the angle covers the full range. So: the orbits of the rotation group are the circles centred on the origin, and a circle is precisely a locus of constant distance from the origin. This is so familiar that it takes an effort to notice it is a theorem.

Minkowski. Start at (ct0,x0)(ct_{0},x_{0}) and boost through every ϕ\phi. Since c2t2x2c^{2}t^{2}-x^{2} is preserved, every image satisfies

c2t2x2  =  kwithk=c2t02x02  fixed. c^{2}t^{2}-x^{2} \;=\; k \qquad\text{with}\qquad k = c^{2}t_{0}^{2}-x_{0}^{2}\ \text{ fixed}. (2.3.18)

That is the equation of a hyperbola, a rectangular one with asymptotes ct=±xct=\pm x. So much for the curve the image stays on. We also want to know how the point moves along it, so write the starting point in the form (ct0,x0)=(λcoshϕ0,λsinhϕ0)(ct_{0},x_{0}) = \big(\lambda\cosh\phi_{0},\,\lambda\sinh\phi_{0}\big) for k=λ2>0k=\lambda^{2}\gt0 and apply (2.3.16) to it:

ct=λ(coshϕ0coshϕsinhϕ0sinhϕ)=λcosh(ϕ0ϕ),x=λ(sinhϕ0coshϕcoshϕ0sinhϕ)=λsinh(ϕ0ϕ), \begin{aligned} ct' &= \lambda\big(\cosh\phi_{0}\cosh\phi-\sinh\phi_{0}\sinh\phi\big)=\lambda\cosh(\phi_{0}-\phi),\\[3pt] x' &= \lambda\big(\sinh\phi_{0}\cosh\phi-\cosh\phi_{0}\sinh\phi\big)=\lambda\sinh(\phi_{0}-\phi), \end{aligned} (2.3.19)

So all the boost does is slide the hyperbolic angle: ϕ0ϕ0ϕ\phi_{0}\to\phi_{0}-\phi. Compare the Euclidean statement that a rotation slides the polar angle, θ0θ0θ\theta_{0}\to\theta_{0}-\theta. The structure is identical, and the parametrisation (cosh,sinh)(\cosh,\sinh) is doing exactly the job that (cos,sin)(\cos,\sin) does.

The invariant hyperbolae are the circles of this geometry

The curves c2t2x2=kc^{2}t^{2}-x^{2}=k are the loci of events at fixed interval from the origin. They are what "a circle of radius rr" means when the distance function has a minus sign in it, and the three families they fall into are the whole of §4's causal structure, drawn:

  • k>0k \gt 0: two branches opening upward and downward, entirely inside the light cone. The upper branch is the set of events at proper time k/c\sqrt{k}/c into the future of the origin.
  • k<0k \lt 0: two branches opening left and right, entirely outside the light cone, at proper distance k\sqrt{-k} from the origin.
  • k=0k = 0: the hyperbola degenerates into its own asymptotes, the two lines ct=±xct=\pm x, which are the light cone. This is the one case with no Euclidean counterpart, because x2+y2=0x^{2}+y^{2}=0 has only the single point x=y=0x=y=0 as a solution while c2t2x2=0c^{2}t^{2}-x^{2}=0 has a whole pair of lines.

A boost moves points along these curves and never across them. Which is to say: a boost can change an event's time coordinate by a factor of a thousand and its position coordinate likewise, and cannot change its interval from the origin by one part in 101510^{15}. The figure below is built to make you watch exactly that happen.

One consequence deserves saying out loud, because it kills a persistent confusion. The light cone k=0k=0 is an orbit too. A boost slides a null point along the ray ct=xct=x, multiplying both coordinates by eϕ\ee^{-\phi}. (Put k=0k=0 into (2.3.19), or just note that coshϕsinhϕ=eϕ\cosh\phi-\sinh\phi=\ee^{-\phi}.) It can shrink the coordinates toward zero or blow them up without bound, and it can never move the point off the line.

That is Postulate 2 in geometric dress. The light cone is not merely a surface that happens to be invariant. It is the degenerate member of the family of invariant hyperbolae, and no continuous boost can carry a point across it.

3.3 · Why the boosted axes look wrong, and why the ticks are wrong too

Chapter 2.2's figure showed the primed axes scissoring symmetrically toward the light line, and readers invariably ask two questions about it. First: why are the ctct' and xx' axes not perpendicular, if SS' is a perfectly good frame with perpendicular axes? Second: where along those axes is "one second", and why is it not where I would put it? Both have the same answer, and we can now give it.

Perpendicularity. In Euclidean geometry, two directions are perpendicular when their dot product vanishes. In this geometry the dot product is (2.3.11), with the minus signs in it. The right word here is orthogonal, and it means ημνVμWν=0\eta_{\mu\nu}V^{\mu}W^{\nu}=0.

So let's test the two primed axes against that condition. The ctct' axis is the set of events with x=0x'=0, so it points along Tμ=(coshϕ,sinhϕ,0,0)T^{\mu}=(\cosh\phi,\sinh\phi,0,0). The xx' axis is the set with t=0t'=0, so it points along Xμ=(sinhϕ,coshϕ,0,0)X^{\mu}=(\sinh\phi,\cosh\phi,0,0). Their Minkowski dot product is

TX=ημνTμXν=coshϕsinhϕsinhϕcoshϕ=0. T\cdot X = \eta_{\mu\nu}T^{\mu}X^{\nu} = \cosh\phi\sinh\phi - \sinh\phi\cosh\phi = 0. (2.3.20)

They are orthogonal, exactly, at every ϕ\phi. They do not look it, because your eye is applying the Euclidean dot product of the page, in which their product is coshϕsinhϕ+sinhϕcoshϕ=sinh2ϕ0\cosh\phi\sinh\phi+\sinh\phi\cosh\phi=\sinh2\phi\neq0. The axes are perpendicular in the geometry and not on the paper, and it is the paper that is wrong.

There is a pleasing detail here. The two axes are mirror images in the line ct=xct=x, since swapping coshsinh\cosh\leftrightarrow\sinh maps one to the other. That reflection is what "scissoring symmetrically about the light line" means, and it is forced by the fact that light itself is orthogonal to itself. A null vector N=(1,1,0,0)N=(1,1,0,0) has NN=0N\cdot N=0.

Calibration. Where is the event "one second of SS' time, at SS''s origin"? By definition it is at x=0x'=0, ct=1ct'=1, so its interval from the origin is c2t2x2=1c^{2}t'^{2}-x'^{2}=1. Intervals are invariant, so in the unprimed frame it is the point where the ctct' axis meets the invariant hyperbola c2t2x2=1c^{2}t^{2}-x^{2}=1. From the parametrisation, that point is

(ct,x)  =  (coshϕ, sinhϕ)  =  (γ, γβ). \big(ct,\,x\big) \;=\; \big(\cosh\phi,\ \sinh\phi\big) \;=\; \big(\gamma,\ \gamma\beta\big). (2.3.21)

So the invariant hyperbolae calibrate the axes. To mark off unit ticks along any observer's time axis, intersect that axis with the hyperbolae c2t2x2=1,4,9,c^{2}t^{2}-x^{2}=1,4,9,\dots. To mark off unit ticks along the space axis, intersect it with x2c2t2=1,4,9,x^{2}-c^{2}t^{2}=1,4,9,\dots. The recipe works for every observer, because the hyperbolae belong to no observer.

How far off is the eye? The unit tick on the unprimed ctct axis sits at Euclidean distance 11 from the origin on the page. The unit tick on the ctct' axis sits at page-distance

cosh2ϕ+sinh2ϕ  =  cosh2ϕ  =  γ2+γ2β2  =  γ1+β2. \sqrt{\cosh^{2}\phi+\sinh^{2}\phi} \;=\; \sqrt{\cosh2\phi} \;=\; \sqrt{\gamma^{2}+\gamma^{2}\beta^{2}} \;=\; \gamma\sqrt{1+\beta^{2}}. (2.3.22)

Put β=0.6\beta=0.6 into that, where ϕ=artanh0.6=ln2\phi=\operatorname{artanh}0.6=\ln 2 exactly, which is a small gift. The page-distance comes out as 1.252+0.752=2.125=1.4577\sqrt{1.25^{2}+0.75^{2}}=\sqrt{2.125}=1.4577. So the primed observer's second is drawn 46% longer than yours. It is not longer. The page is lying, uniformly and predictably, by the factor γ1+β2\gamma\sqrt{1+\beta^{2}}:

β\beta0.20.20.40.40.50.50.60.60.80.80.90.9
ϕ\phi0.20270.20270.42360.42360.54930.54930.69310.69311.09861.09861.47221.4722
page stretch1.0411.0411.1751.1751.2911.2911.4581.4582.1342.1343.0873.087

Once you know that, length contraction and time dilation become visible on a diagram rather than mysterious. Chapter 2.2 §4.3 said that two observers "slice the same four-dimensional object at different angles". The slices are the lines parallel to each observer's xx axis. The calibration is why the same physical rod, sliced two ways, is assigned two different lengths without anything being squeezed.

⚠ Why this isn't obvious — the page is Euclidean and the geometry is not

This is the single most common source of confusion when reading spacetime diagrams, so it is worth stating in its bluntest form: on a spacetime diagram, longer on the page means shorter in proper time. Your visual system measures distances with Δx2+Δy2\sqrt{\Delta x^{2}+\Delta y^{2}} because that is the geometry of the paper. The diagram's content is c2Δt2Δx2\sqrt{c^{2}\Delta t^{2}-\Delta x^{2}}. These do not merely differ by a scale factor. They run in opposite directions.

Numbers make it concrete. Take Chapter 2.2's twins, in years and light-years, from (ct,x)=(0,0)(ct,x)=(0,0) to (10,0)(10,0):

WorldlineLength on the pageProper time
straight, at rest10.0010.0010.0010.00 yr
out to x=4x=4 at β=0.8\beta=0.8 and back241=12.812\sqrt{41}=12.816.006.00 yr
constant proper acceleration, same turning point12.9712.974.944.94 yr

Perfectly monotone, and backwards. The path that looks longest is the one along which least time elapses. Every drawn segment that tilts away from vertical is buying page-length with proper time, at a rate set by the minus sign. So train yourself to stop trusting the picture for lengths. Keep trusting it for everything else, meaning incidence, ordering, and which side of a light ray something is on, because those it gets right.

3.4 · Seeing it

MINKOWSKI invariant (ct)² − x² rapidity φ = 0.0000 β = tanh φ = 0.00000 γ = cosh φ = 1.00000
in S: ct = +1.50000 x = +0.50000 in S′: ct′ = +1.50000 x′ = +0.50000
Δs² in S = +2.0000000000 | Δs² in S′ = +2.0000000000 | difference = 0.00e+0
TIMELIKE, absolute future — proper time from the origin τ = 1.414214 yr
One sign, two geometries. Inside this figure only, units are chosen with c=1c=1 — years for time, light-years for distance — so the vertical axis is ctct and light travels at exactly 4545^{\circ}. Everything is computed numerically from cosh\cosh, sinh\sinh, cos\cos, sin\sin; no result is drawn from a formula stated elsewhere. (i) Minkowski mode. The faint curves are the invariant hyperbolae c2t2x2=±k2c^{2}t^{2}-x^{2}=\pm k^{2}; the orange lines are the light cone, the degenerate member k=0k=0. The hollow marker is an event as you placed it — click or drag anywhere in the frame to move it — and the solid marker is the same event as labelled by an observer boosted by rapidity ϕ\phi. Slide ϕ\phi and watch: the solid marker travels along the bold hyperbola through the event and never leaves it, while ctct and xx separately change by large factors. The readout gives Δs2\Delta s^{2} computed independently in both frames; the two agree to the last digit shown. The small ticks along the bold curve are at rapidity intervals of 0.50.5, so you can see that the boost slides the point uniformly in ϕ\phi — the hyperbola is the orbit, and ϕ\phi is its natural parameter. (ii) Euclidean mode. The identical figure with one sign changed: the invariant is now x2+y2x^{2}+y^{2}, the level curves are circles, the transformation is an ordinary rotation, and the point slides along its circle instead. Ticks are at angle intervals of 0.50.5 rad. Flip between the two modes and you are watching the entire difference between Euclidean geometry and special relativity. (iii) Background shading in Minkowski mode marks the causal regions of §4 — absolute future, absolute past, and elsewhere — and the last readout line classifies the current event and reports its proper time from the origin (if timelike) or proper distance (if spacelike). Note that no amount of boosting ever moves the event out of the region it started in.
Grind box — why ϕ\phi deserves to be called an angle: it is twice a sector area, exactly as θ\theta is

Calling ϕ\phi an "angle" is so far only an analogy of algebraic form. Here is the fact that makes it more than that, and the computation is one line in each geometry.

Chapter 0.7 §5 derived Green's theorem, one corollary of which is that the area swept by the radius vector from the origin along a curve (u(s),v(s))\big(u(s),v(s)\big) is

A=12(udvvdu)=12(uv˙vu˙)ds. A = \frac12\int\big(u\,\dd v - v\,\dd u\big) = \frac12\int\big(u\dot v - v\dot u\big)\,\dd s.

Circle. Parametrise the unit circle as u=cossu=\cos s, v=sinsv=\sin s. Then

uv˙vu˙=cos2s+sin2s=1A=θ2. u\dot v - v\dot u = \cos^{2}s + \sin^{2}s = 1 \qquad\Longrightarrow\qquad A = \frac{\theta}{2}.

Hyperbola. Parametrise the unit hyperbola u2v2=1u^{2}-v^{2}=1 as u=coshsu=\cosh s, v=sinhsv=\sinh s. Then

uv˙vu˙=cosh2ssinh2s=1A=ϕ2. u\dot v - v\dot u = \cosh^{2}s - \sinh^{2}s = 1 \qquad\Longrightarrow\qquad A = \frac{\phi}{2}.

The same integrand, the same value, the same conclusion: the parameter is twice the area of the sector it sweeps out, in both geometries. That is the honest sense in which ϕ\phi is an angle. And it explains additivity without any trigonometry: applying two transformations sweeps two adjacent sectors, and areas add. Angles add because areas add, in the circular case and the hyperbolic case alike.

It also explains why the analogy could not have been avoided. "Angle" was never fundamentally about turning. It was about the invariant measure along the orbit of a one-parameter group, and both groups have one. Chapter 6.1 will call this measure the group parameter and will not care which geometry it came from.

A numerical check, since the claim is easy to test. Take ϕ=1\phi=1. The sector bounded by the uu-axis, the hyperbola, and the ray to (cosh1,sinh1)=(1.5431,1.1752)(\cosh1,\sinh1)=(1.5431,1.1752) has area 12(1.5431)(1.1752)11.5431u21du=0.90680.4068=0.5000\tfrac12(1.5431)(1.1752) - \int_{1}^{1.5431}\sqrt{u^{2}-1}\,\dd u = 0.9068 - 0.4068 = 0.5000, which is ϕ/2\phi/2 ✓. (The integral evaluates to 12uu2112arcoshu\tfrac12 u\sqrt{u^{2}-1}-\tfrac12\operatorname{arcosh}u between the limits.)

In plain terms 2.3.3

Take one event and apply every transformation in the family to it, then ask what curve the image traces. In ordinary space the answer is a circle, since the sum of squares is preserved and the angle sweeps through everything, and a circle is a locus of constant distance from the centre. Flip the sign and the same question returns a hyperbola, with the light rays through the starting point as its asymptotes.

So hyperbolas are the circles of this geometry, the loci of events at fixed interval from a given one. Boosting slides a point along the hyperbola it started on and never carries it to a neighbouring one, which shows up as large changes in both coordinates alongside no change in the interval. The light rays are the degenerate member, where the curve collapses onto its own asymptotes, and that member has no Euclidean counterpart, because a sum of squares vanishes at one point while a difference vanishes along two lines.

This disposes of the complaint that boosted axes look wrong. They are genuinely at right angles in this geometry and fail to look it because the eye applies the geometry of the paper; the invariant hyperbolas supply the tick marks, since they belong to no observer. Longer on the paper means less elapsed time, so the two measures do not merely differ by a scale factor, they run in opposite directions.

4 · Causal structure

In Euclidean space, x2+y2+z2x^{2}+y^{2}+z^{2} is positive unless both points coincide. Every pair of distinct points is at a positive distance, and there is nothing more to say. Flip the sign and a new phenomenon appears immediately: Δs2\Delta s^{2} can be positive, negative, or zero, and which it is turns out to be the most important thing about a pair of events.

4.1 · Three kinds of separation

The classification

For two events with interval Δs2=c2Δt2Δx2\Delta s^{2}=c^{2}\Delta t^{2}-\abs{\Delta\vv x}^{2}:

  • Timelike if Δs2>0\Delta s^{2}\gt0: cΔt>Δxc\abs{\Delta t}\gt\abs{\Delta \vv x}. Light has time to spare. A material particle can be present at both.
  • Null (or lightlike) if Δs2=0\Delta s^{2}=0: cΔt=Δxc\abs{\Delta t}=\abs{\Delta\vv x}. Only light connects them, and only just.
  • Spacelike if Δs2<0\Delta s^{2}\lt0: cΔt<Δxc\abs{\Delta t}\lt\abs{\Delta \vv x}. Not even light can get from one to the other.

The classification is Lorentz-invariant, and the proof is two words: Δs2\Delta s^{2} is invariant, so the sign of Δs2\Delta s^{2} is invariant. That is the entire argument, and its brevity is the point. Everything else in this section is unpacking what that one-line theorem means.

It is worth being clear about how strong that is. Whether two events are timelike separated is not a matter of who is looking, how fast they are moving, or which coordinates they chose. It is a fact about the pair, of exactly the same standing as "these two points are three metres apart" in Euclidean geometry. It is a frame-independent property, and it is the first one we have met that is not just a number but a structure.

4.2 · The light cone

Fix an event OO and ask which events are null-separated from it. With OO at the origin, the condition Δs2=0\Delta s^{2}=0 reads

c2t2=x2+y2+z2,that isx=±ct. c^{2}t^{2} = x^{2}+y^{2}+z^{2}, \qquad\text{that is}\qquad \abs{\vv x} = \pm\,ct. (2.3.23)

At each time t>0t\gt0 this is a sphere of radius ctct, which is the expanding flash of light emitted at OO. Stack those spheres up the time axis and you get a cone. Do the same for t<0t\lt0 and you get a second cone opening downward, made of the light that could have converged on OO. Together they are the light cone of OO.

The name comes from what it looks like drawn. Suppress yy and zz and it is the pair of 4545^{\circ} lines in the figure of §3.4. Suppress only zz and it is an honest cone.

The cone partitions spacetime into three regions, and every one of them is invariant, because the cone is invariant and the sign of Δt\Delta t inside it is invariant (§4.3):

RegionConditionMeaning
Absolute futureΔs20\Delta s^{2}\geq0 and Δt>0\Delta t\gt0Events OO can influence. Every observer agrees they happen after OO.
Absolute pastΔs20\Delta s^{2}\geq0 and Δt<0\Delta t\lt0Events that can influence OO. Every observer agrees they happen before OO.
ElsewhereΔs2<0\Delta s^{2}\lt0No causal contact in either direction. Time order is frame-dependent.

The word absolute is doing real work there and is not decoration. It marks the parts of "before" and "after" that survive the loss of absolute simultaneity. Chapter 2.2 destroyed the universal "now", and it is easy to conclude from that wreckage that time order is entirely conventional. It is not. Precisely the causally connectable part of it is absolute. That is the minimum required for physics to make sense, and, as we are about to see, it is also exactly the maximum available.

"Elsewhere" is the genuinely new region, and it has no Newtonian counterpart. In Galilean physics every event is either before, after, or simultaneous with OO, and the three cases exhaust everything. Here the third of those has swollen from a three-dimensional slice into a four-dimensional region, and it contains, at this instant, everything in the universe more than a few light-nanoseconds away from you. The Andromeda galaxy has no "now". It has a two-million-year-thick band of events about which your question "what is happening there right now?" has no frame-independent answer.

4.3 · The causality theorem

Claim. If two events are timelike or null separated, every inertial observer agrees on which one happened first.

Proof. Suppose Δs20\Delta s^{2}\geq0. The first thing to extract from that hypothesis is a bound on how far apart in space the two events are, relative to how far apart in time, so write the condition out and then drop the two transverse terms:

c2Δt2    Δx2+Δy2+Δz2    Δx2ΔxcΔt    1, c^{2}\Delta t^{2} \;\geq\; \Delta x^{2}+\Delta y^{2}+\Delta z^{2} \;\geq\; \Delta x^{2} \qquad\Longrightarrow\qquad \left|\frac{\Delta x}{c\,\Delta t}\right| \;\leq\; 1, (2.3.24)

The division there is legitimate because Δt0\Delta t\neq0. (If Δt=0\Delta t=0 then Δs20\Delta s^{2}\leq0, with equality only for coincident events.) Now choose axes so that the boost is along xx. That costs nothing, since a rotation changes neither Δt\Delta t nor the classification. What we want is a comparison between Δt\Delta t' and Δt\Delta t, so take the time equation of (2.3.2) and factor cΔtc\Delta t out of it:

cΔt=γ(cΔtβΔx)=γcΔt[1βΔxcΔt]. c\,\Delta t' = \gamma\big(c\,\Delta t-\beta\,\Delta x\big) = \gamma\,c\,\Delta t\left[1 - \beta\,\frac{\Delta x}{c\,\Delta t}\right]. (2.3.25)

The sign of Δt\Delta t' now rests entirely on the sign of that bracket, because γ\gamma is positive and cΔtc\Delta t carries the sign we are testing. So bound the bracket below, using the inequality we have just derived:

1βΔxcΔt    1βΔxcΔt    1β  >  0, 1 - \beta\,\frac{\Delta x}{c\,\Delta t} \;\geq\; 1 - \abs\beta\left|\frac{\Delta x}{c\,\Delta t}\right| \;\geq\; 1-\abs\beta \;\gt\; 0, (2.3.26)

The last step used β<1\abs\beta\lt1, which is not an extra assumption but a property of every Lorentz transformation. Since γ>0\gamma\gt0 and the bracket is strictly positive, Δt\Delta t' has the same sign as Δt\Delta t in every frame. \blacksquare

Notice exactly where the proof lives: in the collision of two inequalities, Δx/cΔt1\abs{\Delta x/c\Delta t}\leq1 from the separation being timelike, and β<1\abs\beta\lt1 from the observer being an observer. Neither alone is enough. If observers could exceed cc, or if influences could be spacelike, the bracket could reach zero and change sign. Causality is protected by the product of two things being less than one.

A sharper version comes free. We would like a lower bound on the size of Δt\Delta t' and not just on its sign, so rearrange the invariance relation into c2Δt2=Δs2+Δx2Δs2c^{2}\Delta t'^{2}=\Delta s^{2}+\abs{\Delta\vv x'}^{2}\geq\Delta s^{2} and take square roots:

cΔt    Δs2for every frame, c\,\abs{\Delta t'} \;\geq\; \sqrt{\Delta s^{2}} \qquad\text{for every frame,} (2.3.27)

with equality exactly in the frame where Δx=0\Delta\vv x'=0. So for a timelike pair, Δt\Delta t' does not merely fail to vanish. It is bounded away from zero by a frame-independent amount, and no observer can even make the two events nearly simultaneous. Worked example 1 puts numbers on this.

4.4 · Spacelike: the order is genuinely up for grabs

Now the other case, and it is not a near miss. Claim. If two events are spacelike separated, there is a frame in which they are simultaneous, a frame in which AA precedes BB, and a frame in which BB precedes AA.

Proof. Take Δs2<0\Delta s^{2}\lt0 and rotate the axes so that the spatial separation lies along xx. Then Δy=Δz=0\Delta y=\Delta z=0 and c2Δt2<Δx2c^{2}\Delta t^{2}\lt\Delta x^{2}, so Δx0\Delta x\neq0 and we may divide by it:

cΔtΔx  <  1. \left|\frac{c\,\Delta t}{\Delta x}\right| \;\lt\; 1. (2.3.28)

Now go looking for the frame in which the two events are simultaneous. Reading (2.3.25), the bracket there vanishes, and with it Δt\Delta t', at the single velocity

β0  =  cΔtΔx, \beta_{0} \;=\; \frac{c\,\Delta t}{\Delta x}, (2.3.29)

By (2.3.28) that value satisfies β0<1\abs{\beta_{0}}\lt1, so it is an allowed velocity and the frame exists. Now for the two orders. The quantity cΔt=γ(cΔtβΔx)c\Delta t'=\gamma(c\Delta t-\beta\Delta x) is a continuous, strictly monotonic function of β\beta that passes through zero at β0\beta_{0}, its derivative in β\beta never vanishing for spacelike separations. So it is positive on one side of β0\beta_{0} and negative on the other. Boosts with β\beta slightly less than β0\beta_{0} give one order, and boosts with β\beta slightly more give the other. \blacksquare

In the simultaneous frame the invariance of the interval gives Δx2=Δs2-\Delta x'^{2}=\Delta s^{2}, so the two events are separated by

Δσ    Δs2  =  Δx, \Delta\sigma \;\equiv\; \sqrt{-\Delta s^{2}} \;=\; \abs{\Delta x'}, (2.3.30)

That quantity is the proper distance. It is the spacelike counterpart of proper time, and it is the distance measured by the one observer for whom the two events are "at the same moment". The two cases mirror each other exactly. Just as Δs2\sqrt{\Delta s^{2}} is the smallest cΔtc\abs{\Delta t'} any observer can report for a timelike pair, Δs2\sqrt{-\Delta s^{2}} is the smallest Δx\abs{\Delta\vv x'} any observer can report for a spacelike pair, by the identical algebra.

Two events being reversible in order sounds alarming until you notice what it costs to arrange, and what it does not permit. Reversal requires β>β0=cΔt/Δx\abs\beta\gt\abs{\beta_{0}}=c\abs{\Delta t}/\abs{\Delta x}, and that is possible exactly when the separation is spacelike. The condition for the reordering to be achievable and the condition for the events to be causally disconnected are the same condition. That is not luck. It is the theorem below.

4.5 · Why cc is a causal speed limit and not merely light's speed

Assemble the two results. Time order is absolute for causally connectable pairs, and reversible for causally disconnected ones. Nothing so far forbids a faster-than-light influence. We have just not allowed one yet. Let's allow one and watch what happens.

Suppose some influence travels from event AA to event BB at speed u>cu\gt c in some frame. It could be a signal, a particle, or a "quantum of whatever". Then Δx=uΔt>cΔt\abs{\Delta \vv x}=u\abs{\Delta t}\gt c\abs{\Delta t}, so AA and BB are spacelike separated. By §4.4 there is then an inertial frame in which BB precedes AA. But AA was the cause and BB the effect. So:

The theorem

If any influence could travel faster than cc, then some inertial observer would see the effect happen before the cause.

That does not mean "would see it that way as an optical illusion". It means the observer would find, in coordinates constructed with their own synchronised clocks, that the effect is recorded at an earlier time than the cause. And by Postulate 1 that observer's description is as valid as anybody's.

An effect preceding its cause is uncomfortable, but it is not yet a contradiction. The contradiction takes one more step, and that step is constructive. With two superluminal signals you can send a message into your own past, and here is how.

The construction

Work with c=1c=1, in seconds and light-seconds, and take the idealised case of a signal that is instantaneous in the frame of whoever sends it. (The finite-uu version is in the grind box. It needs β>2u/(u2+1)\beta\gt2u/(u^{2}+1) and is otherwise identical.)

Step 1. Alice is at rest in frame SS at x=0x=0. At t=0t=0 she sends Bob a signal that is instantaneous in SS, meaning that it travels along the line t=0t=0. Bob sits at x=L=1x=L=1 light-second, so the reception event is

P:(ct,x)=(0, 1). P: \quad (ct,\,x) = (0,\ 1). (2.3.31)

Step 2. Bob is not at rest in SS. He moves at β=0.6\beta=0.6, so he is at rest in SS'. We are going to need the reception event in his coordinates, so transform PP with the boost:

ctP=γ(ctPβxP)=1.25(00.6×1)=0.75,xP=γ(xPβctP)=1.25(10)=1.25. \begin{aligned} ct'_{P} &= \gamma\big(ct_{P}-\beta x_{P}\big) = 1.25\,(0-0.6\times1) = -0.75,\\[3pt] x'_{P} &= \gamma\big(x_{P}-\beta\,ct_{P}\big) = 1.25\,(1-0) = 1.25. \end{aligned} (2.3.32)

Step 3. On receiving the message, Bob immediately replies with a signal that is instantaneous in his own frame, so it travels along the line t=tP=0.75t'=t'_{P}=-0.75. The question is where that line crosses Alice's worldline x=0x=0. On x=0x=0 we have ct=γctct'=\gamma\,ct, so

γct=0.75ct=0.6. \gamma\,ct = -0.75 \qquad\Longrightarrow\qquad ct = -0.6. (2.3.33)

Alice receives Bob's reply at t=0.6 st=-0.6\ \mathrm{s}: six tenths of a second before she sent the original message. In general the return time is t=βL/ct=-\beta L/c, so the effect is available at will and can be made as large as you like by increasing LL.

Step 4, the contradiction. Alice now writes a program: if a reply arrives before t=0t=0, do not send the message. The premise of the loop is that the message was sent. The conclusion is that it was not. This is not a paradox of interpretation, a puzzle about free will, or a subtlety about what "cause" means. It is a straightforward logical contradiction, and it is produced by three ingredients: the existence of superluminal signalling, the relativity of simultaneity, and Postulate 1, which grants Bob the same right as Alice to call his own frame's instantaneous signal instantaneous. Since the last two are established, the first must go.

What cc actually is

This is why the constant cc is a causal speed limit and not merely "the speed of light". Nothing in the argument above mentioned light. The limit is a property of the causal structure of spacetime, of the light cone as the boundary between "can influence" and "cannot". Light travels at cc because light is massless, and not the other way round.

The distinction is not pedantry, and it pays off twice later. Suppose the photon turned out to have a tiny mass. Light would then travel slightly slower than cc, and not one line of this chapter would change. The constant cc would remain the limit, and light would merely stop saturating it.

The second payoff is in general relativity, where there is no global speed to speak of. What survives there is exactly this: the light cone at each point, the local causal structure, which Part III keeps while discarding almost everything else. That is why Chapter 3.1 can throw away flat spacetime and still have a theory.

Grind box — the loop with a finite superluminal speed, and the exact threshold

The instantaneous case above is clean, but it looks like a limiting trick. It is not one. Any u>cu\gt c works, provided the relative velocity of the two parties is large enough, and here is that condition derived.

Units c=1c=1. Alice at x=0x=0 sends at t=0t=0 a signal of speed u>1u\gt1 in SS toward +x+x, and Bob receives it at

P=(tP,xP)=(L/u, L). P = \big(t_{P},\,x_{P}\big) = \big(L/u,\ L\big).

Bob is at rest in SS' with velocity β\beta, so transform that reception event into his coordinates:

tP=γL(1uβ),xP=γL(1βu). t'_{P} = \gamma L\left(\frac1u-\beta\right), \qquad x'_{P} = \gamma L\left(1-\frac{\beta}{u}\right).

Bob replies with a signal of speed uu in his own frame, travelling in the x-x' direction: x(t)=xPu(ttP)x'(t') = x'_{P} - u\big(t'-t'_{P}\big). Alice's worldline x=0x=0 is, in SS', the line x=βtx'=-\beta t'. Setting the two equal and solving,

t(uβ)=xP+utPt=xP+utPuβ. t'\big(u-\beta\big) = x'_{P}+u\,t'_{P} \qquad\Longrightarrow\qquad t' = \frac{x'_{P}+u\,t'_{P}}{u-\beta}.

Substitute the two transformed coordinates and factor out γL\gamma L:

xP+utP=γL[(1βu)+(1βu)]=γL[2β(u+1u)]. x'_{P}+u\,t'_{P} = \gamma L\left[\left(1-\frac\beta u\right)+\left(1-\beta u\right)\right] = \gamma L\left[2-\beta\left(u+\frac1u\right)\right].

Alice's own clock reads t=t/γt=t'/\gamma on her worldline, so the reply arrives at

t  =  Luβ[2βu2+1u]. t \;=\; \frac{L}{u-\beta}\left[2-\beta\,\frac{u^{2}+1}{u}\right].

Since u>βu\gt\beta, the sign is the sign of the bracket, and the reply beats the message (t<0t\lt0) exactly when

  β  >  2uu2+1.   \boxed{\;\beta \;\gt\; \frac{2u}{u^{2}+1}.\;}

Sanity checks. The threshold is less than 11 for every u>1u\gt1, since 2u/(u2+1)<1(u1)2>02u/(u^{2}+1)\lt1 \Leftrightarrow (u-1)^{2}\gt0 ✓, so an allowed β\beta always exists. As uu\to\infty the threshold goes to 00: with truly instantaneous signals any relative motion suffices, matching the main text. As u1+u\to1^{+} the threshold goes to 11: at exactly light speed you would need an observer at exactly light speed, of which there are none, and causality is safe. The barrier is not a convention. It is precisely where the construction stops working.

Numbers. Take u=2cu=2c, so the threshold is 22/5=0.82\cdot2/5=0.8. With L=1L=1 light-second:

β\beta0.800.800.850.850.900.900.950.95
reply arrives at tt (s)0.0000.0000.109-0.1090.227-0.2270.357-0.357

At threshold the reply arrives exactly as the message departs; beyond it, before.

One honest caveat. What the argument rules out is superluminal signalling, meaning the transmission of information or influence, which is inconsistent with relativity plus logic. It does not forbid superluminal coordinate speeds, which Chapter 2.2 §5 already showed are harmless (the spotlight, the scissors), and it does not by itself forbid a particle that always moves faster than cc and cannot be used to signal. ⚑ Whether such a particle is consistent quantum-mechanically is a separate question with a well-known answer, which is that it is not, and the argument there is Chapter 5.1's rather than ours.

⚠ Why this isn't obvious — Δs2\Delta s^{2} is not a distance

Chapter 0.5 §1 laid down the axioms an inner product must satisfy, and one of them was positive-definiteness: v,v>0\avg{v,v}\gt0 for every v0v\neq0. Everything downstream in that chapter leaned on it. The Cauchy–Schwarz inequality needs it. So does the triangle inequality, the definition of v=v,v\norm{v}=\sqrt{\avg{v,v}} as a real number, the interpretation of u,v/uv\avg{u,v}/\norm u\norm v as a cosine, and the whole apparatus of angles.

ημν\eta_{\mu\nu} does not satisfy it, and the violation is deliberate. The vector (1,1,0,0)(1,1,0,0) is not zero, and ημνVμVν=11=0\eta_{\mu\nu}V^{\mu}V^{\nu}=1-1=0. So we have a nonzero vector of zero length, and vectors of negative length squared, and no way to define V\norm{V} as a real number for all VV. Calling η\eta a "metric" is standard and is a slight abuse. The precise term is a non-degenerate symmetric bilinear form of signature (1,3)(1,3), and the word non-degenerate marks the axiom we did keep.

What survives: bilinearity, symmetry, non-degeneracy (if VW=0V\cdot W=0 for every WW then V=0V=0), the invariance of VWV\cdot W under the transformations, the notion of orthogonality VW=0V\cdot W=0, and, most important of all, the ability to classify vectors into three invariant types. That is enough to do geometry with, and §5 and §6 go on to do exactly that.

What does not: positivity, the triangle inequality (§6 reverses it), Cauchy–Schwarz (§6 reverses that too, for timelike vectors), and the idea that "distance zero" means "same point". Vectors orthogonal to themselves exist, and they are exactly the null vectors. That sentence would be flatly false in Chapter 0.5, and here it is the defining feature of the light cone.

So the honest summary is: we removed one axiom from the definition of an inner product, and in exchange the geometry acquired a causal structure. That trade is what the rest of physics is built on.

In plain terms 2.3.4

In ordinary space the distance between two distinct points is positive and there is nothing further to say. Reverse one sign and the same quantity comes out positive, negative or exactly zero, and which of the three it is turns out to be the most important thing about the pair. The proof that the classification belongs to the pair rather than the observer is two words long: the quantity is invariant, so its sign is.

What has appeared is not another agreed number but an agreed structure: what survives of before and after once the universal present is gone. Events light or matter could reach keep their order in every frame, while events too far apart for anything to cross between have an order different observers reverse without anybody being wrong, since nothing was ever at stake in it. Time order is absolute where it could matter and negotiable where it could not.

The speed limit follows, and it belongs to the geometry rather than to light. Let some influence outrun it and the two events it joins are of the negotiable kind, so some observer records the effect before the cause. Two such influences are worse: one person signals another, the second replies with an influence instantaneous in his own frame, the reply arrives before the first was sent, and she declines to send it. Light travels at the limit because light is massless.

5 · Proper time is the length of a worldline

So far the interval has connected two events by a straight line. Now let's bend it.

A particle's history in spacetime is a curve, called its worldline. Along that curve, take two neighbouring events separated by dxμ=(cdt,dx,dy,dz)\dd x^{\mu}=(c\,\dd t,\dd x,\dd y,\dd z). Suppose the particle moves slower than light, which by §4 is exactly the statement that its worldline is timelike at every point. Then ds2>0\dd s^{2}\gt0, and we may define

c2dτ2    ds2  =  c2dt2dx2dy2dz2. c^{2}\,\dd\tau^{2} \;\equiv\; \dd s^{2} \;=\; c^{2}\dd t^{2}-\dd x^{2}-\dd y^{2}-\dd z^{2}. (2.3.34)

What we want is dτ\dd\tau expressed through the coordinate time dt\dd t and the particle's ordinary speed, since those are the things an observer measures. So divide the right-hand side by c2dt2c^{2}\dd t^{2} and pull dt\dd t back out in front:

dτ=dt11c2[(dxdt)2+(dydt)2+(dzdt)2]=1β2(t) dt=dtγ(t), \dd\tau = \dd t\sqrt{1 - \frac{1}{c^{2}}\left[\left(\dv{x}{t}\right)^{2}+\left(\dv{y}{t}\right)^{2}+\left(\dv{z}{t}\right)^{2}\right]} = \sqrt{1-\beta^{2}(t)}\ \dd t = \frac{\dd t}{\gamma(t)}, (2.3.35)

Here β(t)\beta(t) is the particle's instantaneous speed in units of cc. It is a function of tt, because the particle may do whatever it likes along the way. Now add the contributions up by integrating along the path:

  τ[path]  =  t1t21β2(t) dt  =  1cds.   \boxed{\;\tau[\,\text{path}\,] \;=\; \int_{t_{1}}^{t_{2}}\sqrt{1-\beta^{2}(t)}\ \dd t \;=\; \frac1c\int \dd s.\;} (2.3.36)

This quantity is the proper time along the worldline. Three things about it, in increasing order of importance.

It is invariant. Each ds\dd s is invariant, and a sum of invariants is invariant. Two observers computing (2.3.36) for the same path will use different β(t)\beta(t), different t1,t2t_{1},t_{2}, and different intermediate values, and will get the same number.

It is a functional of the path, not a function of the endpoints. This is the language of Chapter 1.2. Here τ\tau takes an entire curve as input and returns one number, in precisely the way that the arc-length functional A[y]=1+y2dx\mathcal{A}[y]=\int\sqrt{1+y'^{2}}\,\dd x of 1.2 §1 does. Set the two of them side by side:

A[y]=dx2+dy2Euclidean length,cτ[path]=c2dt2dx2Minkowski length \underbrace{\mathcal{A}[y]=\int\sqrt{\dd x^{2}+\dd y^{2}}}_{\text{Euclidean length}}, \qquad\qquad \underbrace{c\,\tau[\,\text{path}\,]=\int\sqrt{c^{2}\dd t^{2}-\dd x^{2}}}_{\text{Minkowski length}} (2.3.37)

There is the sign again, in the only place it could possibly be. Proper time is arc length. It is the length of the worldline, measured with the geometry that spacetime actually has. Everything §6 does follows from taking that sentence at face value.

It is what a clock reads. This is the part that keeps proper time from being an abstraction. Take the interval between two neighbouring events on a clock's own worldline. In the clock's instantaneous rest frame it is not moving, so dx=dy=dz=0\dd x'=\dd y'=\dd z'=0 and ds2=c2dt2\dd s^{2}=c^{2}\dd t'^{2}. That is to say, the interval is the time elapsed on that clock. Since ds2\dd s^{2} is invariant, the same holds computed in any frame at all.

So τ\tau is the accumulated reading of a clock carried along the path. And (2.3.36) says that a moving clock accumulates less of it than the coordinate time, by exactly γ\gamma at each instant. That is Chapter 2.2's time dilation, now derived as a statement about the length of a curve rather than about clocks being defective.

Grind box — the assumption hidden in "a clock measures proper time"

The argument above computed dτ\dd\tau in the clock's instantaneous rest frame, and then added the pieces up. There is an assumption buried in "instantaneous", and it deserves naming because it is a physical input rather than a mathematical one.

The clock hypothesis. An ideal clock's rate depends only on its instantaneous velocity, not on its acceleration. Equivalently: an accelerating clock reads the same as a momentarily co-moving inertial clock, at every moment.

This is not a theorem of special relativity. It is an extra assumption, and it is easy to see that it could fail: a pendulum clock in an accelerating rocket manifestly does not obey it, and neither does any device whose mechanism is disturbed by being shaken. What the hypothesis really asserts is that some physical processes are ideal in this sense, so that "proper time" is realisable and not merely definable.

⚑ Experimentally it holds to an astonishing degree. Muons circulating in a storage ring at γ29\gamma\approx29 experience a proper acceleration of order 1018g10^{18}g, and their decay rate is slowed by exactly the factor γ\gamma predicted from their speed alone, with no acceleration-dependent correction at the 10310^{-3} level. Atomic clocks flown on aircraft and the clocks aboard GPS satellites agree with (2.3.36) integrated along their actual trajectories. We quote those results. Deriving them would require the internal dynamics of the clock.

Why it matters twice. First, without it, §6's twin calculation would be incomplete, because you could always object that the turnaround does something unmodelled to the traveller's clock. Second, in general relativity the clock hypothesis is promoted to a postulate. There, proper time along a worldline is defined to be what a clock carried along it reads, and that definition is the bridge between the geometry and any measurement at all. Chapter 3.3 leans on it in its first paragraph.

In plain terms 2.3.5

Straight separations have carried everything so far, and no real history is straight. Bend the path, chop it into pieces short enough to be straight, take the interval along each and add them, and what accumulates is the proper time. It is a functional of the whole path rather than a function of its endpoints, in the sense the action was, and it is the arc length of the worldline in the geometry spacetime actually has.

Set it beside the arc length of a curve in the plane and the two expressions differ in one place only, the sign under the root. That single difference does all the work of the argument to come. The quantity is invariant besides, because each piece is, so two observers who disagree about every intermediate number agree about the total.

It is also what a clock reads, which keeps it from being an abstraction: in a clock's own momentary rest frame the interval between neighbouring events on its path is the time it displays. Dilation stops being a statement about defective clocks and becomes one about the length of a curve. One physical assumption hides in the word momentary and deserves naming: an ideal clock's rate depends on its speed and not on its acceleration. That is no theorem, a pendulum in a launching rocket violates it, and it holds experimentally to a degree hard to credit.

6 · The twin paradox is a theorem about triangles

Now the payoff. In the plane, the straight line is the shortest path between two points, and everyone has known that since childhood. In spacetime the corresponding theorem is true with the inequality turned around, and the reversal is the minus sign, doing its most consequential work.

6.1 · The theorem

Straight is longest

Among all timelike worldlines connecting two fixed timelike-separated events, the straight one has the greatest proper time. Every other path takes strictly less, and the more it deviates, the less it takes.

Proof. Let the two events be E1E_{1} and E2E_{2}, timelike separated. The whole proof rests on picking the right frame first, so let's pick it. By §4.3 and the invariance of the interval there is a frame in which the two events occur at the same place. Take βrest=Δx/(cΔt)\beta_{\text{rest}}=\Delta x/(c\Delta t), which satisfies βrest<1\abs{\beta_{\text{rest}}}\lt1 precisely because the separation is timelike, exactly mirroring the construction in §4.4.

Work in that frame from here on. There Δx=0\Delta\vv x=0, and the coordinate time between the events is T=ΔtT=\Delta t, which by (2.3.27) is Δs2/c\sqrt{\Delta s^{2}}/c.

The straight worldline joining them is the one at rest in this frame. Along it β0\beta\equiv0, so (2.3.36) gives τstraight=T\tau_{\text{straight}}=T.

Now take any timelike path between the same two events. It has the same endpoints, so it runs over the same coordinate-time interval, and that lets us compare the two proper times integral by integral:

τ[path]  =  0T1β2(t) dt    0T1 dt  =  T  =  τstraight, \tau[\text{path}] \;=\; \int_{0}^{T}\sqrt{1-\beta^{2}(t)}\ \dd t \;\leq\; \int_{0}^{T}1\ \dd t \;=\; T \;=\; \tau_{\text{straight}}, (2.3.38)

The inequality holds because 1β21\sqrt{1-\beta^{2}}\leq1 for every β\beta, with equality only at β=0\beta=0. So every path loses, and it loses strictly unless it is at rest for the entire journey, which is to say unless it is the straight one. \blacksquare

That is the whole proof. It is four lines because we chose the right frame first, and we were allowed to choose it because the answer is frame-independent. The integrand 1β2\sqrt{1-\beta^{2}} is a penalty for moving, charged continuously, and the straight path is the one that never pays it.

6.2 · The same result from Chapter 1.2's machinery

The direct argument is airtight but it uses a special frame. The variational argument does not, and it is the one that survives into Part III, so it is worth doing even though we already have the answer.

Extremise the functional (2.3.36). In one space dimension, with x˙=dx/dt\dot x=\dd x/\dd t,

τ[x]=t1t2F(x˙)dt,F(x˙)=1x˙2c2. \tau[x] = \int_{t_{1}}^{t_{2}} F(\dot x)\,\dd t, \qquad F(\dot x)=\sqrt{1-\frac{\dot x^{2}}{c^{2}}}. (2.3.39)

This is a Chapter 1.2 problem of exactly the type solved in its worked example 2. The Euler–Lagrange equation is

ddt(Fx˙)Fx=0, \dv{}{t}\left(\pdv{F}{\dot x}\right) - \pdv{F}{x} = 0, (2.3.40)

Now notice that FF contains no xx at all, which is the statement that spacetime is homogeneous and has no preferred place in it. So the second term vanishes, and what is left says that the first bracket is a constant of the motion:

Fx˙=x˙/c21x˙2/c2=const    Kc2. \pdv{F}{\dot x} = \frac{-\dot x/c^{2}}{\sqrt{1-\dot x^{2}/c^{2}}} = \text{const} \;\equiv\; -\frac{K}{c^{2}}. (2.3.41)

Solve for x˙\dot x exactly as 1.2 did for the plane geodesic: square, clear the denominator, and collect,

x˙2=K2(1x˙2c2)    x˙2(1+K2c2)=K2  x˙=K1+K2/c2=const. \begin{aligned} \dot x^{2} = K^{2}\left(1-\frac{\dot x^{2}}{c^{2}}\right) \;&\Longrightarrow\; \dot x^{2}\left(1+\frac{K^{2}}{c^{2}}\right)=K^{2}\\[4pt] &\Longrightarrow\; \dot x = \frac{K}{\sqrt{1+K^{2}/c^{2}}} = \text{const}. \end{aligned} (2.3.42)

Constant velocity, which is to say x=vt+x0x=vt+x_{0}, a straight worldline. We assumed nothing about the answer, only stationarity.

There is a free bonus sitting in that last expression. Whatever real constant KK you pick, the resulting x˙\abs{\dot x} comes out strictly less than cc. The extremal paths are automatically timelike, and the variational principle cannot produce a superluminal answer even if you ask it to.

The variational calculation identifies the straight line as stationary. Which kind of stationary point it is has to come from somewhere else, and §6.1 supplies the answer: a maximum. Chapter 1.2 §5.1 made the same distinction for the action and warned that stationary need not mean minimal. Here is the cleanest possible example of that warning being necessary.

Where this goes, and it is the most important forward pointer in Part II

Look at what the calculation actually needed. It needed a functional of the form ds\int\dd s with ds2=ημνdxμdxν\dd s^{2}=\eta_{\mu\nu}\dd x^{\mu}\dd x^{\nu}, and the Euler–Lagrange machinery of Chapter 1.2. It did not need flatness, it did not need η\eta to be constant, and it did not need Lorentz transformations.

Chapter 3.3 will keep this variational principle word for word and make exactly one change. It replaces the constant array ημν\eta_{\mu\nu} by a position-dependent one, gμν(x)g_{\mu\nu}(x), which is ten functions of where you are. The same functional, the same Euler–Lagrange equation, the same request for a stationary path. The resulting equation is called the geodesic equation, and its solutions are called free fall.

That is the whole of general relativity's kinematics, and this section is where it starts to be earned. Chapter 1.2 already showed you the pattern once: the same ds\int\dd s gave a straight line in the plane and a great circle on a sphere, and the only thing that changed was ds\dd s. Gravity is that substitution performed on spacetime. A planet orbits the Sun for the same reason a great-circle route from London to Tokyo goes over the Arctic. It is going as straight as the geometry allows.

6.3 · The reverse triangle inequality

The theorem of §6.1 has a clean algebraic form worth stating on its own, because it is the exact mirror image of the most familiar inequality in geometry.

First the vocabulary. Call a four-vector VμV^{\mu} timelike if VV>0V\cdot V\gt0, and future-pointing if V0>0V^{0}\gt0. For such vectors write VVV\abs V\equiv\sqrt{V\cdot V}, which is a real number. With those two words in place the theorem reads:

For future-pointing timelike U,V:U+V    U+V, \text{For future-pointing timelike } U,V: \qquad \abs{U+V} \;\geq\; \abs U + \abs V, (2.3.43)

with equality only when UU and VV are parallel. Compare the Euclidean statement u+vu+v\norm{u+v}\leq\norm u+\norm v, and notice that the inequality has flipped. In the plane, going via a third point makes the journey longer. In spacetime it makes the elapsed time shorter, and the two-leg twin path is precisely U+VU+V, with UU the outbound leg and VV the inbound one. Problem 1 asks you to prove (2.3.43). The proof turns on a reversed Cauchy–Schwarz inequality UVUVU\cdot V\geq\abs U\abs V, which is also flipped, and flipped for the same reason.

6.4 · Closing the loop with Chapter 2.2

Chapter 2.2 §7.1 worked the twin paradox in detail. The traveller goes 44 light-years at β=0.8\beta=0.8, turns, and returns. Earth ages 1010 years and the traveller ages 66. The resolution offered there was that the traveller changes inertial frames while the stay-at-home does not, followed by a careful accounting of the 6.46.4 years that the traveller's simultaneity slice sweeps across Earth's worldline during the turnaround. All correct, and all now unnecessary.

Here is the same result in one sentence. The travelling twin ages less because a bent timelike path is shorter, and shorter means less proper time.

τstraight=10202=10 yr,τbent=25242=2×3=6 yr. \tau_{\text{straight}} = \sqrt{10^{2}-0^{2}} = 10\ \text{yr}, \qquad \tau_{\text{bent}} = 2\sqrt{5^{2}-4^{2}} = 2\times3 = 6\ \text{yr}. (2.3.44)

Two paths, same endpoints, different lengths. Nobody is scandalised that a detour adds mileage. The only unfamiliar thing is that in this geometry a detour subtracts time.

Notice what the geometric statement does not require. It never mentions acceleration, it never uses a non-inertial frame, and it never needs the traveller's point of view at all. Acceleration enters only as the answer to a different question, which is why is one path bent? And a kink in a worldline is detectable from the inside, by an accelerometer, which is why there was never a symmetry between the twins to appeal to.

The often-heard slogan "the twin paradox is really about acceleration" is therefore half wrong. Acceleration is what makes a path bent. Bentness is what costs proper time. And the amount of time lost is not a function of the acceleration but of the whole shape of the path. Worked example 2 makes that concrete by comparing three paths between the same two events, one of them smoothly accelerated throughout with no kink anywhere.

Grind box — the second variation, and why the extremum is a maximum rather than a minimum

Section 6.1 proved the maximum by a global argument. Here is the local version, which is what Chapter 1.2 §5.1's second-variation machinery would say, and it exposes the sign in a different place.

Work in the frame where the endpoints coincide spatially, and perturb the straight path x(t)=0x(t)=0 by a small εξ(t)\varepsilon\,\xi(t) with ξ(t1)=ξ(t2)=0\xi(t_{1})=\xi(t_{2})=0. Then β=εξ˙/c\beta = \varepsilon\dot\xi/c and

τ[ε]=1ε2ξ˙2c2 dt=[1ε2ξ˙22c2ε4ξ˙48c4]dt, \tau[\varepsilon] = \int\sqrt{1-\frac{\varepsilon^{2}\dot\xi^{2}}{c^{2}}}\ \dd t = \int\left[1-\frac{\varepsilon^{2}\dot\xi^{2}}{2c^{2}} - \frac{\varepsilon^{4}\dot\xi^{4}}{8c^{4}}-\cdots\right]\dd t,

by the binomial series of Chapter 0.3. Read off the two variations:

δτ=0(no term linear in ε),δ2τ=ε2c2ξ˙2dt  <  0. \delta\tau = 0 \quad\text{(no term linear in }\varepsilon\text{)}, \qquad \delta^{2}\tau = -\frac{\varepsilon^{2}}{c^{2}}\int\dot\xi^{2}\,\dd t \;\lt\; 0 .

The first variation vanishes for every perturbation, confirming stationarity without solving anything. The second variation is negative definite, meaning strictly negative unless ξ˙0\dot\xi\equiv0, which is to say unless the perturbation is nothing at all. Hence a maximum, and a genuine one rather than a saddle.

Where the sign came from. Compare the Euclidean calculation for a straight line in the plane, perturbed the same way: 1+ε2ξ2dx=[1+12ε2ξ2+]\int\sqrt{1+\varepsilon^{2}\xi'^{2}}\,\dd x = \int[1+\tfrac12\varepsilon^{2}\xi'^{2}+\cdots], giving δ2A=+ε2ξ2>0\delta^{2}\mathcal{A}=+\varepsilon^{2}\int\xi'^{2}\gt0: a minimum. The two calculations are identical except for the sign inside the square root, which is the sign in the metric, which flips the sign of the second variation, which flips minimum into maximum. One symbol, all the way through.

And a warning that will matter in Part III. This is a local statement. It says the straight path beats all nearby paths. In flat spacetime the global statement is true as well, and that is what §6.1 proved. In curved spacetime it can fail. There can be several geodesics between the same two events, and only one of them maximises. That is the conjugate-point phenomenon of Chapter 1.2 §5.1, and in Chapter 3.8 the several geodesics become the several images of a gravitationally lensed quasar.

In plain terms 2.3.6

The most familiar sentence in geometry is that a straight line is the shortest route between two points, and here it fails in the most interesting available way. Among all the paths a material object could take between two given events, the unaccelerated one accumulates the greatest elapsed time, and every other loses strictly.

The proof is four lines because the frame may be chosen first, and it may be chosen because the answer does not depend on the choice. Work where the two events happen at the same place. The path that stays put accumulates the full coordinate time, while every other pays a penalty at each instant for moving, a penalty never zero and never negative. So the travelling twin ages less for the same reason a detour through the next town adds mileage, except that in this geometry the detour subtracts, and the triangle inequality has turned around.

Two debts are settled by that. The warning that a stationary path need not be a minimum, issued when the action principle arrived, gets its cleanest illustration, since this stationary path is a maximum. And the calculation needed nothing but a functional of that shape and the machinery of the last part, not flatness and not constancy of the array of signs. Replace that array by one whose entries vary from place to place and the identical calculation returns the paths of free fall.

7 · Four-velocity, briefly

One construction before we stop. It is short, it is forced, and Chapter 2.5 is built entirely on it.

Ordinary velocity is dx/dt\dd\vv x/\dd t, and Chapter 2.2's Problem 1 exposed its defect. The numerator is a piece of a four-vector, but the denominator is a coordinate, so the quotient transforms in a mess. Both the top and the bottom change when you change frames, and they do not change compatibly.

The repair is to divide by something invariant instead, and §5 has just supplied the only natural candidate. Define the four-velocity

uμ    dxμdτ. u^{\mu} \;\equiv\; \dv{x^{\mu}}{\tau}. (2.3.45)

Since dxμ\dd x^{\mu} transforms as a four-vector and dτ\dd\tau is a number every observer agrees on, uμu^{\mu} is a four-vector. The numerator picks up a factor of Λ\Lambda and the denominator picks up nothing. That is the entire reason for the definition.

We will want its components in terms of ordinary velocity, and they follow at once from dτ=dt/γ\dd\tau=\dd t/\gamma together with the chain rule:

uμ=dtdτdxμdt=γ(c, v)=(γc, γv). u^{\mu} = \dv{t}{\tau}\,\dv{x^{\mu}}{t} = \gamma\,\big(c,\ \vv v\big) = \big(\gamma c,\ \gamma\vv v\big). (2.3.46)

And now the property that makes this the right object to have defined. Feed uμu^{\mu} into the invariant-manufacturing machine of §2.3 and compute its own square:

uu  =  ημνdxμdτdxνdτ  =  ημνdxμdxνdτ2  =  ds2dτ2  =  c2. u\cdot u \;=\; \eta_{\mu\nu}\dv{x^{\mu}}{\tau}\dv{x^{\nu}}{\tau} \;=\; \frac{\eta_{\mu\nu}\,\dd x^{\mu}\dd x^{\nu}}{\dd\tau^{2}} \;=\; \frac{\dd s^{2}}{\dd\tau^{2}} \;=\; c^{2}. (2.3.47)

That holds identically, for every particle, at every moment, in every frame, whatever it is doing. Check it against the components: γ2c2γ2v2=γ2c2(1β2)=c2\gamma^{2}c^{2}-\gamma^{2}v^{2}=\gamma^{2}c^{2}(1-\beta^{2})=c^{2} ✓.

An object of fixed length

The four-velocity is a four-vector of constant Minkowski length cc. It can point in any timelike future direction, and that is all the freedom it has. It lives on the invariant hyperbola of §3, which in four dimensions is a three-dimensional hyperboloid, and boosting a particle slides its four-velocity along that surface exactly as the figure in §3.4 slides an event along a hyperbola.

This is why uμu^{\mu} and not v\vv v is the right generalisation of velocity. The constraint uu=c2u\cdot u=c^{2} is frame-independent, whereas "speed less than cc" is a statement about components. Chapter 2.5 multiplies by the rest mass to get pμ=muμp^{\mu}=mu^{\mu}, whose invariant square is then pp=m2c2p\cdot p=m^{2}c^{2} with no work at all, and that single equation contains E=γmc2E=\gamma mc^{2}, E2=p2c2+m2c4E^{2}=p^{2}c^{2}+m^{2}c^{4}, and the massless case E=pcE=pc as its three readings.

⚠ The slogan, and what is actually true

You will hear it said that everything moves through spacetime at speed cc, and that time dilation is what happens when motion through space is "bought" out of motion through time. As a mnemonic it is not bad. As a derivation it is not one, and it is worth separating the two.

True: uu=c2u\cdot u=c^{2} for every particle, by (2.3.47). That is a real theorem and it is what the slogan is gesturing at.

Careful: the components of uμu^{\mu} are (γc,γv)(\gamma c,\gamma\vv v), and these do not trade off the way a budget does. Writing the invariant out,

(u0)2=c2+u2, \big(u^{0}\big)^{2} = c^{2} + \abs{\vv u}^{2},

So as spatial motion grows, the "motion through time" u0=γcu^{0}=\gamma c grows too. It does not shrink. There is no fixed pot of cc being divided up. The two pieces combine with a minus sign, which is precisely why the total can stay at cc while both pieces run off to infinity. A picture built on "you only have so much speed to share out" will give you the wrong sign on the first question that tests it.

The honest version of the slogan: a particle's four-velocity is a unit timelike vector (scaled by cc), and what encodes the velocity is its direction rather than its magnitude. Time dilation is the statement that a unit vector tilted away from your time axis has a smaller projection onto it, and that projection is dτ/dt=1/γ\dd\tau/\dd t=1/\gamma. That is a statement about a fixed-length vector being tilted, which is exactly §3's picture, and it is derivable. The budget metaphor is not.

In plain terms 2.3.7

One construction remains, and it is forced rather than chosen. Ordinary velocity has a defect that surfaces once frames matter: its numerator and denominator belong to different worlds, the displacement being part of a four-dimensional object while the elapsed time is one observer's coordinate. The repair is to divide by something nobody disagrees about, and the worldline's length is the one candidate.

What comes out has a fixed length, the same for every particle at every moment whatever it does, so all its freedom is in its direction. Dilation then reads as a fixed-length object tilted from your time axis having a smaller projection along it. The slogan about everything moving through spacetime at one speed gestures at this and is a mnemonic rather than a derivation, since the two pieces combine with a minus sign and both run to infinity rather than trading against each other.

That is as far as three chapters carry it. The speed limit is built into the geometry rather than into any material; observers slice one fabric at different angles and agree on the interval; the straightest history carries the most time. But the word used throughout for a collection of four numbers has been a definition by resemblance, and the up-and-down placement of its labels has been a spelling rule obeyed on trust. Saying what makes such a collection an object rather than a list comes next.

8 · Worked examples

Worked example 1 — a full causal-structure computation

Four events are given in a frame SS, in units where c=1c=1: time in years, distance in light-years, so coordinates are written (ct,x)(ct,x) with y=z=0y=z=0 throughout.

EventP1P_{1}P2P_{2}P3P_{3}P4P_{4}
(ct,x)(ct,\,x)(0,0)(0,\,0)(3,1)(3,\,1)(1,4)(1,\,4)(5,5)(5,\,5)

(a) Classify all six pairs. (b) For the spacelike pair P2P3P_{2}P_{3}, find the frame in which they are simultaneous and a frame in which their order is reversed, and verify the interval is unchanged. (c) Show that no boost reverses the timelike pair P1P2P_{1}P_{2}, and find the frame that comes closest.

(a) The classification. Compute Δs2=c2Δt2Δx2\Delta s^{2}=c^{2}\Delta t^{2}-\Delta x^{2} for each pair. Six pairs, six subtractions:

PaircΔtc\Delta tΔx\Delta xΔs2\Delta s^{2}TypeInvariant meaning
P1P2P_{1}\to P_{2}331191=+89-1=+8timelikeτ=8=2.828\tau=\sqrt8=2.828 yr
P1P3P_{1}\to P_{3}1144116=151-16=-15spacelikeσ=15=3.873\sigma=\sqrt{15}=3.873 ly
P1P4P_{1}\to P_{4}55552525=025-25=0nullon the light cone
P2P3P_{2}\to P_{3}2-23349=54-9=-5spacelikeσ=5=2.236\sigma=\sqrt5=2.236 ly
P2P4P_{2}\to P_{4}2244416=124-16=-12spacelikeσ=12=3.464\sigma=\sqrt{12}=3.464 ly
P3P4P_{3}\to P_{4}4411161=+1516-1=+15timelikeτ=15=3.873\tau=\sqrt{15}=3.873 yr

Now read the structure off. P2P_{2} and P4P_{4} both lie in the absolute future of P1P_{1}'s light cone or on it, with P2P_{2} strictly inside and P4P_{4} exactly on it, which means a light signal emitted at P1P_{1} arrives precisely at P4P_{4}. P3P_{3} is in P1P_{1}'s elsewhere, so nothing that happens at P1P_{1} can affect P3P_{3}, and nothing at P3P_{3} can affect P1P_{1}. Yet P3P_{3} can affect P4P_{4} (Δs2=+15\Delta s^{2}=+15), and so can P1P_{1}. Causal influence is not a transitive-and-total ordering. It is a partial order, and the light cones are its structure.

Note also that the classification survives translation: only differences entered, so moving the origin changes nothing. That is Chapter 2.2's homogeneity, still holding.

(b) The spacelike pair P2P3P_{2}P_{3}. In SS, P3P_{3} happens 22 years before P2P_{2} (cΔt=2c\Delta t=-2 going from P2P_{2} to P3P_{3}) and 33 light-years to the right.

Simultaneous frame. By (2.3.29),

β0=cΔtΔx=23=0.6667, \beta_{0} = \frac{c\,\Delta t}{\Delta x} = \frac{-2}{3} = -0.6667,

which has modulus less than 11 ✓, as guaranteed for a spacelike pair. That frame moves in the x-x direction at two thirds of light speed. In it, γ0=(14/9)1/2=1.3416\gamma_{0}=(1-4/9)^{-1/2}=1.3416 and

Δx=γ0(Δxβ0cΔt)=1.3416(3(0.6667)(2))=1.3416×1.6667=2.2361, \begin{aligned} \Delta x' &= \gamma_{0}\big(\Delta x-\beta_{0}\,c\Delta t\big)\\[3pt] &= 1.3416\big(3-(-0.6667)(-2)\big) = 1.3416\times1.6667 = 2.2361, \end{aligned}

and 2.2361=52.2361=\sqrt5 ✓, which is the proper distance, exactly as (2.3.30) promised. The two events are 5\sqrt5 light-years apart and simultaneous, for that observer.

Reversed frame. Any β\beta beyond β0\beta_{0} in the same direction does it; take β=0.8\beta=-0.8, γ=5/3\gamma=5/3:

cΔt=53(2(0.8)(3))=53(0.4)=+0.6667,Δx=53(3(0.8)(2))=53(1.4)=2.3333. \begin{aligned} c\,\Delta t' &= \tfrac53\big(-2-(-0.8)(3)\big) = \tfrac53(0.4) = +0.6667,\\[3pt] \Delta x' &= \tfrac53\big(3-(-0.8)(-2)\big) = \tfrac53(1.4) = 2.3333. \end{aligned}

The sign of Δt\Delta t' has flipped. In SS it was P3P_{3} that came first, and in this frame P2P_{2} does. And the interval is untouched:

Δs2=(0.6667)2(2.3333)2=0.44445.4444=5.0000 \Delta s'^{2} = (0.6667)^{2}-(2.3333)^{2} = 0.4444-5.4444 = -5.0000 \quad\checkmark

Both frames are correct. Neither event caused the other, and neither could, being spacelike separated, so nothing whatever is at stake in the disagreement.

(c) The timelike pair P1P2P_{1}P_{2}, which no boost can reorder. Here cΔt=3c\Delta t=3, Δx=1\Delta x=1. From (2.3.25),

cΔt=γ(3β), c\,\Delta t' = \gamma\big(3-\beta\big),

Since β<1<3\abs\beta\lt1\lt3, the bracket is positive for every admissible β\beta. At worst β1\beta\to1^{-} gives 3β23-\beta\to2, still positive, while γ\gamma\to\infty. So Δt>0\Delta t'\gt0 always, and in fact Δt\Delta t' can be made arbitrarily large but never small. Tabulate:

β\beta0.9-0.900+1/3+1/3+0.8+0.8+0.99+0.99
cΔtc\Delta t' (yr)8.9478.9473.0003.0002.8282.8283.6673.66714.2514.25
Δx\Delta x' (ly)8.4888.4881.0001.000002.333-2.33313.96-13.96
Δs2\Delta s'^{2}8.0008.0008.0008.0008.0008.0008.0008.0008.0008.000

The last row is the point of the whole exercise: two coordinates swinging over an order of magnitude and changing sign, and one number sitting perfectly still.

The closest approach. The minimum of cΔtc\Delta t' is at β=Δx/(cΔt)=1/3\beta=\Delta x/(c\Delta t)=1/3, the frame in which the two events happen at the same place, and there

cΔtmin=Δs2=8=2.8284 yr,Δx=0, c\,\Delta t'_{\min} = \sqrt{\Delta s^{2}} = \sqrt8 = 2.8284\ \text{yr}, \qquad \Delta x' = 0,

confirming (2.3.27). A clock can be present at both events, since they are timelike separated. It needs to travel 11 light-year in 33 years, which is β=1/3\beta=1/3. Such a clock reads 2.8282.828 years between them, and that is the proper time of the pair. No observer can shrink the gap below it, and none can make the two events simultaneous, let alone reversed.

The moral. Time order is not "relative". It is absolute exactly where it could matter and relative exactly where it could not, and the boundary between the two regimes is the light cone. That is a far more disciplined statement than "everything is relative", and it is the one the theory actually makes.

Worked example 2 — three worldlines between the same two events

Two events: departure E1=(ct,x)=(0,0)E_{1}=(ct,x)=(0,0) and reunion E2=(10,0)E_{2}=(10,0), again in years and light-years with c=1c=1. Compute the proper time along three timelike paths joining them: (a) the straight one; (b) Chapter 2.2's two-leg trip out to x=4x=4 at β=0.8\beta=0.8 and back; (c) a smooth worldline of constant proper acceleration with the same turning point. Verify that the straight one wins.

(a) Straight. The path x(t)=0x(t)=0 has β0\beta\equiv0, so (2.3.36) gives

τa=01010 dt=10.000 yr. \tau_{a} = \int_{0}^{10}\sqrt{1-0}\ \dd t = 10.000\ \text{yr}.

Equivalently, straight from the interval: τ=Δs2/c=10202=10\tau=\sqrt{\Delta s^{2}}/c=\sqrt{10^{2}-0^{2}}=10.

(b) Two straight legs. Outbound from (0,0)(0,0) to (5,4)(5,4), inbound from (5,4)(5,4) to (10,0)(10,0). Each leg is straight, so its proper time is the interval along it:

τb=5242+5242=3+3=6.000 yr. \tau_{b} = \sqrt{5^{2}-4^{2}} + \sqrt{5^{2}-4^{2}} = 3+3 = 6.000\ \text{yr}.

Speed on each leg: β=4/5=0.8\beta=4/5=0.8, so 1β2=0.6\sqrt{1-\beta^{2}}=0.6 and 5×0.6=35\times0.6=3 ✓ by the other route. This is Chapter 2.2 §7.1's traveller, and the deficit is 106=410-6=4 years.

(c) Constant proper acceleration. Chapter 2.2's Problem 4 showed that a worldline of constant proper acceleration is a hyperbola (xxc)2c2(ttc)2=k2\big(x-x_{c}\big)^{2}-c^{2}\big(t-t_{c}\big)^{2}=k^{2}, which §3 now identifies as an invariant hyperbola, translated. Take the branch that leaves x=0x=0, turns around, and comes back, symmetric about ct=5ct=5:

x(t)=xck2+c2(t5)2,xc=k2+25, x(t) = x_{c}-\sqrt{k^{2}+c^{2}(t-5)^{2}}, \qquad x_{c}=\sqrt{k^{2}+25},

where xcx_{c} was fixed by demanding x(0)=0x(0)=0, and x(10)=0x(10)=0 then follows by symmetry. Choose kk so that the turning point matches path (b)'s, namely x(5)=xck=4x(5)=x_{c}-k=4:

k2+25k=4    k2+25=k2+8k+16    k=98=1.125 ly. \sqrt{k^{2}+25}-k = 4 \;\Longrightarrow\; k^{2}+25 = k^{2}+8k+16 \;\Longrightarrow\; k = \tfrac98 = 1.125\ \text{ly}.

Is it timelike everywhere? Write u=c(t5)u=c(t-5). Then dx/du=u/k2+u2\dd x/\dd u = -u/\sqrt{k^{2}+u^{2}}, whose modulus is less than 11 for every uu ✓. The fastest it ever goes is at the endpoints, βmax=5/1.1252+25=0.9756\abs\beta_{\max}=5/\sqrt{1.125^{2}+25}=0.9756.

Proper time. Along the path,

c2dτ2=du2dx2=du2[1u2k2+u2]=k2du2k2+u2, c^{2}\dd\tau^{2} = \dd u^{2}-\dd x^{2} = \dd u^{2}\left[1-\frac{u^{2}}{k^{2}+u^{2}}\right] = \frac{k^{2}\,\dd u^{2}}{k^{2}+u^{2}},

so cdτ=kdu/k2+u2c\,\dd\tau = k\,\dd u/\sqrt{k^{2}+u^{2}}, which integrates to an inverse hyperbolic sine:

cτc=k[arsinhuk]5+5=2karsinh5k=2(1.125)arsinh(4.4444). c\,\tau_{c} = k\Big[\operatorname{arsinh}\frac{u}{k}\Big]_{-5}^{+5} = 2k\operatorname{arsinh}\frac{5}{k} = 2(1.125)\operatorname{arsinh}(4.4444).

Since arsinhz=ln ⁣(z+1+z2)\operatorname{arsinh}z=\ln\!\big(z+\sqrt{1+z^{2}}\big), we get arsinh(4.4444)=ln9=2.19722\operatorname{arsinh}(4.4444)=\ln 9=2.19722 (the square root comes out to exactly 4.55564.5556, so the argument is exactly 99), hence

τc=2.25×2.19722=4.9438 yr. \tau_{c} = 2.25\times2.19722 = 4.9438\ \text{yr}.

(Checked by direct numerical summation of dt2dx2\sqrt{\dd t^{2}-\dd x^{2}} over two million steps along the path: 4.943755304.94375530, against the closed form 4.943755304.94375530.)

The comparison.

PathProper timeDeficitLength on the page
(a) straight10.000010.0000 yr10.0010.00
(b) two legs, β=0.8\beta=0.86.00006.0000 yr4.004.00 yr241=12.812\sqrt{41}=12.81
(c) constant proper acceleration4.94384.9438 yr5.065.06 yr12.9712.97

The straight path wins, as §6.1 guarantees it must. And the ordering of the last column is exactly the reverse of the second, with more page-length going together with less proper time, monotonically. That is the ⚠ callout of §3.3 in numerical form.

Two things worth extracting.

Acceleration is not the mechanism. Path (c) has no kink anywhere. It is smooth, its proper acceleration is constant, and it never changes inertial frame abruptly. It nonetheless loses more proper time than path (b), which does have a kink. So "the twin who accelerates ages less" is not a law. The amount lost depends on the whole shape of the path, not on where or how hard it accelerated. The correct statement remains the geometric one: bent is shorter, and (c) is bent more.

The numbers are not exotic. The proper acceleration of path (c) is a=c2/ka=c^{2}/k, and Chapter 2.2's Problem 4 gives c2/g=0.9687c^{2}/g=0.9687 light-years, so

ag=c2/gk=0.96871.125=0.861. \frac{a}{g} = \frac{c^{2}/g}{k} = \frac{0.9687}{1.125} = 0.861.

A crew accelerating at 0.86g0.86g, slightly less than standing on Earth, makes a round trip that takes ten years by Earth's clocks and four years eleven months by theirs, reaching 44 light-years out and 97.6%97.6\% of light speed at the extremes. This is not a thought experiment about impossible machines. It is a comfortable ride, and the only obstacle is fuel, which is Chapter 2.5's problem.

9 · Your turn

Problem 1 — the reverse triangle inequality

Let UμU^{\mu} and VμV^{\mu} be future-pointing timelike four-vectors: UU>0U\cdot U\gt0, VV>0V\cdot V\gt0, U0>0U^{0}\gt0, V0>0V^{0}\gt0. Write U=UU\abs U=\sqrt{U\cdot U}. (a) Prove the reversed Cauchy–Schwarz inequality UVUVU\cdot V\geq\abs U\,\abs V, with equality only when UU and VV are parallel. (b) Deduce U+VU+V\abs{U+V}\geq\abs U+\abs V. (c) Explain in one sentence which Euclidean fact each of these is the mirror image of, and where the sign flip entered.

Solution

(a) Reversed Cauchy–Schwarz. Both sides of the claimed inequality are invariant, so we may evaluate them in whatever frame is convenient. That is the standard and enormously useful move here, and it is licensed by (2.3.4). Choose the rest frame of UU. Since UU is timelike and future-pointing, §6.1's construction supplies a boost making its spatial part vanish, so

Uμ=(U, 0), U^{\mu} = \big(\abs U,\ \vv 0\big),

because then UU=(U0)2=U2U\cdot U=(U^{0})^{2}=\abs U^{2} and U0>0U^{0}\gt0. In that frame write Vμ=(V0,V)V^{\mu}=(V^{0},\vv V). The dot product is easy:

UV=ημνUμVν=UV0. U\cdot V = \eta_{\mu\nu}U^{\mu}V^{\nu} = \abs U\,V^{0}.

Now use VV's own invariant, VV=(V0)2V2=V2V\cdot V=(V^{0})^{2}-\abs{\vv V}^{2}=\abs V^{2}, to write V0=V2+V2V^{0}=\sqrt{\abs V^{2}+\abs{\vv V}^{2}}. We take the positive root, since VV is future-pointing, and by the causality theorem of §4.3 that property is frame-independent. Hence

UV=UV2+V2    UV, U\cdot V = \abs U\sqrt{\abs V^{2}+\abs{\vv V}^{2}} \;\geq\; \abs U\,\abs V,

with equality if and only if V=0\vv V=\vv 0, i.e. VV is at rest in UU's frame, i.e. VUV\parallel U. \blacksquare

The whole reversal happened in one place: V2\abs{\vv V}^{2} enters V0V^{0} with a plus sign after the minus in the metric has been moved to the other side. In Euclidean space the corresponding step subtracts and you get \leq.

(b) The triangle inequality. U+VU+V is future-pointing (its time component is a sum of positives) and timelike (by (a): (U+V)(U+V)=U2+2UV+V2>0(U+V)\cdot(U+V)=\abs U^{2}+2U\cdot V+\abs V^{2}\gt0), so U+V\abs{U+V} is defined. Expand and apply (a):

U+V2=U2+2UV+V2    U2+2UV+V2=(U+V)2. \abs{U+V}^{2} = \abs U^{2} + 2\,U\cdot V + \abs V^{2} \;\geq\; \abs U^{2}+2\abs U\abs V+\abs V^{2} = \big(\abs U+\abs V\big)^{2}.

Take square roots, which is legitimate because both sides are positive, and the result is U+VU+V\abs{U+V}\geq\abs U+\abs V, with equality only for parallel vectors. \blacksquare

(c) The mirrors. (a) mirrors u,vuv\abs{\avg{u,v}}\leq\norm u\norm v of Chapter 0.5 §1, and (b) mirrors u+vu+v\norm{u+v}\leq\norm u+\norm v. Both flip because the metric is not positive-definite: in Chapter 0.5 the proof of Cauchy–Schwarz went by demanding uλv,uλv0\avg{u-\lambda v,u-\lambda v}\geq0 for all λ\lambda, and that step is exactly the axiom §4's ⚠ callout removed.

What it says physically. Take UU and VV to be the two legs of the twin's journey, as four-vectors from departure to turnaround and turnaround to reunion. Then U+VU+V is the straight path, U+V=c×\abs{U+V}=c\times (stay-at-home's proper time) and U+V=c×\abs U+\abs V=c\times (traveller's). The inequality is the twin paradox, and it is a two-line theorem about vectors. Numerically, with U=(5,4)U=(5,4) and V=(5,4)V=(5,-4) in the units of §6.4: U=V=3\abs U=\abs V=3, U+V=(10,0)U+V=(10,0), U+V=106\abs{U+V}=10\geq6 ✓, and UV=25+16=413×3=9U\cdot V=25+16=41\geq3\times3=9 ✓, with plenty to spare, because the two legs are far from parallel.

Problem 2 — simultaneity and reversal for a spacelike pair

Two events in a frame SS, in years and light-years with c=1c=1: E1=(ct,x)=(0,0)E_{1}=(ct,x)=(0,0) and E2=(2,5)E_{2}=(2,5). (a) Classify the pair. (b) Find the frame in which they are simultaneous, and the distance between them in that frame. (c) Find a frame in which E2E_{2} precedes E1E_{1}, and verify the interval. (d) Show that no frame can make them occur at the same place, and say why that is the mirror image of the timelike case.

Solution

(a) Δs2=2252=425=21<0\Delta s^{2}=2^{2}-5^{2}=4-25=-21\lt0, so the pair is spacelike. Nothing at E1E_{1} can influence E2E_{2}, because a signal would need speed 5/2=2.5c5/2=2.5c.

(b) Simultaneous frame. Set Δt=0\Delta t'=0 in cΔt=γ(cΔtβΔx)c\Delta t'=\gamma(c\Delta t-\beta\Delta x):

β0=cΔtΔx=25=0.4,γ0=110.16=1.09109. \beta_{0} = \frac{c\Delta t}{\Delta x} = \frac{2}{5}=0.4, \qquad \gamma_{0}=\frac{1}{\sqrt{1-0.16}}=1.09109.

Then

Δx=γ0(Δxβ0cΔt)=1.09109(50.8)=1.09109×4.2=4.58258, \Delta x' = \gamma_{0}\big(\Delta x-\beta_{0}c\Delta t\big) = 1.09109\,(5-0.8) = 1.09109\times4.2 = 4.58258,

and 4.58258=214.58258=\sqrt{21} ✓, the proper distance (2.3.30). Check the interval: 02(21)2=210^{2}-(\sqrt{21})^{2}=-21 ✓. An observer moving at 0.4c0.4c says these two events happen at the same moment, 4.5834.583 light-years apart.

(c) Reversed frame. Any β>0.4\beta\gt0.4 works. Take β=0.8\beta=0.8, γ=5/3\gamma=5/3:

cΔt=53(20.8×5)=53(2)=3.3333,Δx=53(50.8×2)=53(3.4)=5.6667. c\Delta t' = \tfrac53\,(2-0.8\times5) = \tfrac53(-2) = -3.3333, \qquad \Delta x' = \tfrac53\,(5-0.8\times2) = \tfrac53(3.4)=5.6667. Δs2=(3.3333)2(5.6667)2=11.111132.1111=21.0000 \Delta s'^{2} = (-3.3333)^{2}-(5.6667)^{2} = 11.1111-32.1111 = -21.0000 \quad\checkmark

In SS, E1E_{1} came first by 22 years. In this frame E2E_{2} comes first by 3.333.33 years. Both are correct, and nothing is at stake.

(d) No common-place frame. Setting Δx=0\Delta x'=0 would need β=Δx/(cΔt)=5/2=2.5\beta=\Delta x/(c\Delta t)=5/2=2.5, which exceeds 11 and is not a velocity. More robustly, from invariance Δx2=c2Δt2+2121>0\abs{\Delta\vv x'}^{2}=c^{2}\Delta t'^{2}+21\geq21\gt0 in every frame, so Δx\Delta\vv x' can never vanish and in fact can never fall below 21\sqrt{21}.

That is the exact mirror of the timelike result (2.3.27), where cΔtΔs2c\abs{\Delta t'}\geq\sqrt{\Delta s^{2}} and no frame can make the events simultaneous. Swap the roles of space and time and the two statements map onto each other, which is what you should expect of a geometry whose only asymmetry between them is a sign. In slogan form: a timelike pair has an invariant time and a negotiable separation. A spacelike pair has an invariant separation and a negotiable time.

Problem 3 — the accelerated worldline is a hyperbola, and it has a horizon

A rocket has constant proper acceleration aa. Chapter 2.2's Problem 4 gave its rapidity as ϕ=aτ/c\phi=a\tau/c and its worldline, in the frame where it is momentarily at rest at τ=0\tau=0, as

ct(τ)=c2asinhaτc,x(τ)=c2acoshaτc. ct(\tau)=\frac{c^{2}}{a}\sinh\frac{a\tau}{c}, \qquad x(\tau)=\frac{c^{2}}{a}\cosh\frac{a\tau}{c}.

(a) Show that this is an invariant hyperbola of §3 and identify which one. (b) Find its asymptotes. (c) A beacon sits at the spatial origin x=0x=0 and flashes at times t0t_{0}. Show that flashes with t00t_{0}\geq0 never reach the rocket, and that flashes with t0<0t_{0}\lt0 always do. (d) Identify the resulting Rindler horizon and say precisely what it is a horizon for.

Solution

(a) It is an invariant hyperbola. Use cosh2sinh2=1\cosh^{2}-\sinh^{2}=1:

x2c2t2=c4a2[cosh2aτcsinh2aτc]=(c2a)2, x^{2}-c^{2}t^{2} = \frac{c^{4}}{a^{2}}\left[\cosh^{2}\frac{a\tau}{c}-\sinh^{2}\frac{a\tau}{c}\right] = \left(\frac{c^{2}}{a}\right)^{2},

a constant. So the worldline lies on c2t2x2=(c2/a)2c^{2}t^{2}-x^{2}=-\big(c^{2}/a\big)^{2}, which is one of §3's spacelike hyperbolae, the right-hand branch, with k=(c2/a)2k=-(c^{2}/a)^{2}, so that the proper distance k\sqrt{-k} from the origin is c2/ac^{2}/a. Write bc2/ab\equiv c^{2}/a for brevity. It is a length, and it is the rocket's distance from the origin at t=0t=0.

This is a striking fact worth pausing on. A boost slides points along these curves ((2.3.19)), so the uniformly accelerated worldline is an orbit of the boost group. Boosting the rocket by ϕ\phi gives you the same worldline, reparametrised. That is the exact analogue of "a circle is carried to itself by rotation", and it is why this trajectory is the spacetime counterpart of uniform circular motion: the curve of constant "curvature", generated by a one-parameter symmetry. It is also why Chapter 2.2's crew feel a constant gg forever. Their situation is literally unchanged by the passage of proper time, up to a boost.

(b) Asymptotes. As t\abs t\to\infty, x=b2+c2t2ctx=\sqrt{b^{2}+c^{2}t^{2}}\to c\abs t, so the asymptotes are x=±ctx=\pm ct: the light cone of the origin. The rocket approaches the speed of light without reaching it, forever, hugging the cone from outside.

(c) Which flashes arrive. A flash emitted at (t0,0)(t_{0},0) travels along x=c(tt0)x=c(t-t_{0}). It reaches the rocket when

c(tt0)=b2+c2t2. c\big(t-t_{0}\big) = \sqrt{b^{2}+c^{2}t^{2}}.

Square both sides, which is permissible provided t>t0t\gt t_{0}, a condition we check afterwards:

c2t22c2tt0+c2t02=b2+c2t2    t=c2t02b22c2t0. c^{2}t^{2}-2c^{2}t\,t_{0}+c^{2}t_{0}^{2} = b^{2}+c^{2}t^{2} \;\Longrightarrow\; t = \frac{c^{2}t_{0}^{2}-b^{2}}{2c^{2}t_{0}}.

Note that c2t2c^{2}t^{2} cancels, so the equation is linear in tt and there is exactly one candidate meeting time.

Case t0>0t_{0}\gt0. Require t>t0t\gt t_{0}:

c2t02b22c2t0>t0    c2t02b2>2c2t02    b2>c2t02, \frac{c^{2}t_{0}^{2}-b^{2}}{2c^{2}t_{0}} \gt t_{0} \;\Longleftrightarrow\; c^{2}t_{0}^{2}-b^{2} \gt 2c^{2}t_{0}^{2} \;\Longleftrightarrow\; -b^{2}\gt c^{2}t_{0}^{2},

which is impossible. Case t0=0t_{0}=0. The formula has zero in the denominator, so there is no finite solution. Case t0<0t_{0}\lt0. Dividing by 2c2t02c^{2}t_{0} now reverses the inequality, and the condition becomes b2<c2t02-b^{2}\lt c^{2}t_{0}^{2}, which is always true. So the flash does arrive, after a delay that grows without bound as t00t_{0}\to0^{-}.

Numbers, with a=ga=g and Chapter 2.2's b=c2/g=0.9687 lyb=c^{2}/g=0.9687\ \mathrm{ly}: a flash at t0=1 yrt_{0}=-1\ \mathrm{yr} is caught at t=(10.9384)/(2)=0.031 yrt=(1-0.9384)/(-2)=-0.031\ \mathrm{yr}; one at t0=0.1 yrt_{0}=-0.1\ \mathrm{yr} is caught at t=+4.64 yrt=+4.64\ \mathrm{yr}; one at t0=0.01 yrt_{0}=-0.01\ \mathrm{yr} at t=+46.9 yrt=+46.9\ \mathrm{yr}. The last light to make it takes forever to arrive, arbitrarily redshifted.

(d) The horizon. The dividing surface is the null line x=ctx=ct, which is the future light cone of the origin and simultaneously the asymptote from (b). Generalise the calculation to a flash from any event (te,xe)(t_{e},x_{e}). The ray x=xe+c(tte)x=x_{e}+c(t-t_{e}) eventually overtakes the asymptote x=ctx=ct if and only if xecte>0x_{e}-ct_{e}\gt0. So

  the rocket can receive a signal from event E  xE>ctE.   \boxed{\;\text{the rocket can receive a signal from event }E\ \Longleftrightarrow\ x_{E} \gt c\,t_{E}.\;}

Events on the far side of x=ctx=ct are permanently invisible to it. That surface is the Rindler horizon.

What it is a horizon for, stated carefully. It is not a property of spacetime. This is flat Minkowski spacetime, with no matter, no curvature, and no singularity anywhere. It is a property of the observer. It is the boundary of the region from which that particular eternally accelerating worldline can ever receive news. An inertial observer sails across it without noticing. Different accelerations aa give different horizons, and if you stop accelerating then yours vanishes and all the delayed signals arrive.

⚑ Three forward pointers, all quoted. (i) The region x>ctx\gt c\abs t that the rocket can explore is called the Rindler wedge, and the coordinates adapted to the family of such observers are Rindler coordinates. (ii) Chapter 3.8 finds a horizon at the Schwarzschild radius of a black hole with exactly the same local character, a one-way surface with no local marker on it, and the argument that it is not a singularity is essentially the one above. (iii) Chapter 3.1 uses the equivalence principle to argue that a uniformly accelerated observer is locally indistinguishable from one at rest in a gravitational field, at which point the fact that acceleration alone can manufacture a horizon in empty flat space stops being a curiosity and becomes a clue.

Problem 4 — counting the Lorentz group

(a) Verify that a spatial rotation Λ=diag(1,R)\Lambda=\mathrm{diag}(1,R) with RTR=1R^{\mathsf T}R=\mathbb 1 satisfies (2.3.13), so rotations are Lorentz transformations. (b) By counting independent equations in ΛTηΛ=η\Lambda^{\mathsf T}\eta\Lambda=\eta, show that the solutions form a six-parameter family. (c) Identify the six as three boosts and three rotations by writing Λ=1+ω\Lambda=\mathbb 1+\omega for infinitesimal ω\omega and showing that ωμνημρωρν\omega_{\mu\nu}\equiv\eta_{\mu\rho}\omega^{\rho}{}_{\nu} is antisymmetric. (d) Connect the count to Chapter 2.4's arithmetic for antisymmetric tensors.

Solution

(a) Rotations qualify. In block form with η=diag(1,13)\eta=\mathrm{diag}(1,-\mathbb 1_{3}),

ΛTηΛ=(100RT)(1001)(100R)=(100RTR)=(1001)=η  \Lambda^{\mathsf T}\eta\Lambda = \begin{pmatrix}1&0\\0&R^{\mathsf T}\end{pmatrix}\begin{pmatrix}1&0\\0&-\mathbb 1\end{pmatrix}\begin{pmatrix}1&0\\0&R\end{pmatrix} = \begin{pmatrix}1&0\\0&-R^{\mathsf T}R\end{pmatrix} = \begin{pmatrix}1&0\\0&-\mathbb 1\end{pmatrix} = \eta \ \checkmark

That makes sense in hindsight. A rotation leaves tt alone and preserves x2+y2+z2x^{2}+y^{2}+z^{2}, so it preserves c2t2(x2+y2+z2)c^{2}t^{2}-(x^{2}+y^{2}+z^{2}). The Lorentz group therefore contains the rotation group as a subgroup, namely the transformations relating two observers who are at rest with respect to each other but have turned their axes.

(b) Six parameters. Λ\Lambda is a 4×44\times4 matrix, so it has 1616 unknown entries. The condition ΛTηΛ=η\Lambda^{\mathsf T}\eta\Lambda=\eta is an equation between two 4×44\times4 matrices, so naively that is 1616 equations. But both sides are symmetric. The left side is symmetric because (ΛTηΛ)T=ΛTηTΛ=ΛTηΛ\big(\Lambda^{\mathsf T}\eta\Lambda\big)^{\mathsf T}=\Lambda^{\mathsf T}\eta^{\mathsf T}\Lambda=\Lambda^{\mathsf T}\eta\Lambda, using ηT=η\eta^{\mathsf T}=\eta, and the right side is η\eta. A symmetric 4×44\times4 matrix equation carries only

n(n+1)2n=4=4×52=10 \frac{n(n+1)}{2}\bigg|_{n=4} = \frac{4\times5}{2} = 10

independent equations. Hence

1610=6  free parameters. 16 - 10 = 6 \ \text{ free parameters.}

(The counting assumes the ten constraints are independent, which they are; (c) confirms it by exhibiting a six-dimensional solution space explicitly at the linear level.)

(c) Three and three. Put Λ=1+ω\Lambda=\mathbb 1+\omega with ω\omega small and keep first order:

(1+ωT)η(1+ω)=η+ωTη+ηω+O(ω2)=η    ωTη+ηω=0. \big(\mathbb 1+\omega^{\mathsf T}\big)\eta\big(\mathbb 1+\omega\big) = \eta + \omega^{\mathsf T}\eta + \eta\omega + O(\omega^{2}) = \eta \;\Longrightarrow\; \omega^{\mathsf T}\eta + \eta\omega = 0.

Define the matrix ω^ηω\hat\omega\equiv\eta\omega, whose entries are ωμν=ημρωρν\omega_{\mu\nu}=\eta_{\mu\rho}\omega^{\rho}{}_{\nu}, so the index has been lowered, in the sense Chapter 2.4 §4.1 makes precise. Then

ω^T=ωTηT=ωTη=ηω=ω^. \hat\omega^{\mathsf T} = \omega^{\mathsf T}\eta^{\mathsf T} = \omega^{\mathsf T}\eta = -\eta\omega = -\hat\omega.

So ωμν=ωνμ\omega_{\mu\nu}=-\omega_{\nu\mu}: antisymmetric. Its independent entries are the strictly-upper-triangular ones, and they split naturally:

  • ω0i\omega_{0i} for i=1,2,3i=1,2,3, mixing time with a space direction. Three boosts, one per axis.
  • ωij\omega_{ij} for 1i<j31\leq i\lt j\leq3, mixing two space directions. Three rotations, one per plane xyxy, yzyz, zxzx.

Total 3+3=63+3=6, matching (b) ✓. Now check the shape against Chapter 2.2's boost. Expanding Λ(ϕ)\Lambda(\phi) to first order in ϕ\phi gives ω01=ω10=ϕ\omega^{0}{}_{1}=\omega^{1}{}_{0}=-\phi, and lowering the first index with η\eta turns the second of these into ω10=η11ϕ=+ϕ\omega_{10}=-\eta_{11}\phi=+\phi while ω01=ϕ\omega_{01}=-\phi, which is antisymmetric ✓. Note that ωμν\omega^{\mu}{}_{\nu} itself is not antisymmetric. The property appears only after lowering, which is a first hint that index position carries meaning.

(d) The connection to Chapter 2.4. Chapter 2.4 §7.2 counts the independent components of an antisymmetric rank-2 tensor in nn dimensions as 12n(n1)\tfrac12 n(n-1), giving 1243=6\tfrac12\cdot4\cdot3=6 in spacetime. It remarks there that six is not four, so such an object cannot masquerade as a four-vector the way B\vv B masquerades in three dimensions. Here is the same six, arrived at completely independently: not by counting slots in a tensor, but by counting how many ways there are to move while preserving η\eta.

They are the same six for a reason. The generators of the Lorentz group are an antisymmetric rank-2 object, and Chapter 2.6 will find that the electromagnetic field tensor FμνF^{\mu\nu} is another one with the same six slots: three of E\vv E and three of B\vv B. ⚑ That the match is structural rather than coincidental is Chapter 6.1's business. The field strength of a gauge theory lives in the Lie algebra of its symmetry group, and for Lorentz symmetry that algebra is exactly the antisymmetric matrices counted above.

The brick you just laid

You turned Chapter 2.2's algebra into a geometry. The interval Δs2=c2Δt2Δx2Δy2Δz2\Delta s^{2}=c^{2}\Delta t^{2}-\Delta x^{2}-\Delta y^{2}-\Delta z^{2} is invariant, proved by direct substitution. The Lorentz transformations are defined as the linear maps preserving it, exactly as rotations are the maps preserving x2+y2+z2x^{2}+y^{2}+z^{2}. And the whole subject is Euclidean geometry with one sign changed.

You met xμx^{\mu}, ημν\eta_{\mu\nu}, the summation convention, and ΛTηΛ=η\Lambda^{\mathsf T}\eta\Lambda=\eta, all of it informally, with the debt to Chapter 2.4 recorded. You saw that boosts are hyperbolic rotations, that the rapidity is the hyperbolic angle and adds because areas add, that the orbits are the invariant hyperbolae, and that those hyperbolae calibrate every observer's axes and explain why the page misleads.

You classified separations as timelike, null or spacelike, proved the classification and the causal order are invariant, proved that spacelike order is always reversible, and built a closed causal loop from two superluminal signals. That is why cc is a limit on causation and not a fact about light.

You defined proper time as the length of a worldline, proved that the straight worldline maximises it, both directly and variationally, and identified the twin paradox as the reverse triangle inequality. And you met the four-velocity, an object of permanently fixed length cc.

Where this gets spent. Chapter 2.4 takes every piece of notation introduced here on trust, meaning xμx^{\mu}, ημν\eta_{\mu\nu}, upper and lower indices, and "transforms like dxμ\dd x^{\mu}", and turns all of it into definitions, with the transformation law as the definition of a tensor. Its opening sentence is about the debt this chapter incurred. Chapter 2.5 feeds the four-velocity of §7 into pμ=muμp^{\mu}=mu^{\mu}, gets pp=m2c2p\cdot p=m^{2}c^{2} from (2.3.47) in one line, and reads E=mc2E=mc^{2} off it. Chapter 2.6 finds the six components counted in Problem 4 sitting inside the electromagnetic field tensor. Chapter 3.1 keeps the light cone and throws away everything else, which is what "spacetime is locally Minkowski" means. Chapter 3.2 builds the manifold on which that statement can be made. Chapter 3.3 takes (2.3.36) unchanged, replaces ημν\eta_{\mu\nu} by gμν(x)g_{\mu\nu}(x), and calls the extremal paths gravity. That is the single most important thing this chapter sets up. Chapter 3.8 meets Problem 3's horizon again, this time around a black hole. And Chapter 7.3 returns to the observation that the invariance group of a metric is where the physics is, in two dimensions, where that group turns out to be infinite-dimensional and the string becomes solvable because of it.