Part II · Special Relativity — Chapter 2.5

Relativistic Dynamics

Conservation laws have to be statements about four-vectors, or they are not laws. Everything else in this chapter is the bill for that sentence.

Where we are

Chapters 2.2 to 2.4 rebuilt kinematics. We know how coordinates transform, we know what is invariant, and we know what a tensor is. We also know why writing a law as an equation between tensors makes it frame-independent by inspection.

What we do not yet have is any mechanics. Momentum is still mvm\vv v, energy is still 12mv2\half mv^{2}, and force is still mam\vv a. All three came over unchanged from Chapter 1.1, and all three rest on a Galilean picture that Chapter 2.2 destroyed.

This chapter rebuilds them. It has exactly one organising idea, and it is worth stating before any algebra so you can watch it do all the work:

Momentum conservation, as you were taught it, is a statement about three numbers. Three numbers are not a frame-independent object. So the law as stated cannot be fundamental.

Repairing that is not optional, and the repair is not free. There is essentially one four-vector you can build out of a particle's mass and motion. Once you insist that it is the conserved thing, a fourth conserved quantity arrives along with the three you wanted. You did not ask for it, no experiment forced it on you, and you have no way to refuse it.

That fourth quantity has a low-speed limit of 12mv2\half mv^{2} plus a constant, and the constant is mc2mc^{2}. Everything famous in this chapter follows from that one structural demand. By §3 you will have watched E=mc2E=mc^{2} arrive as an unwanted guest rather than as a revelation.

Tools you'll need  — Chapter 0.3 §2: the binomial series (1+x)α(1+x)^{\alpha} for arbitrary real α\alpha, and its worked example 2, which already expanded γmc2\gamma mc^{2} without saying what it was. Chapter 0.6 §5: the multivariable chain rule, used constantly below to convert d/dτ\dd/\dd\tau into γd/dt\gamma\,\dd/\dd t. Chapter 1.2 §3.5 and §8.1: the Euler–Lagrange equation, and the table of actions this book will write down. Section 5 here supplies the second entry in that table. Chapter 1.4 §3: Noether's theorem, in particular that invariance under time translation gives a conserved energy. We use it in §5 as an independent check. Chapter 2.3 §5 (proper time as the length of a worldline), §6 (straight worldlines maximise it), §4 (the causal classification of four-vectors), and §7 (four-velocity, sketched there and built properly here). Chapter 2.4 §4 (raising and lowering), §5 (the operations preserving tensor character, and the quotient theorem), and above all §6, the invariance theorem, which is the load-bearing wall of this entire chapter.

1 · Velocity is not a four-vector, and the repair

Start with the object we would most like to keep. A particle traces a worldline through spacetime, and the obvious thing to call its velocity is the rate at which its coordinates change with time:

Vμ    dxμdt  =  (c, v),v=dxdt. V^{\mu} \;\equiv\; \dv{x^{\mu}}{t} \;=\; \big(c,\ \vv v\big), \qquad \vv v = \dv{\vv x}{t}. (2.5.1)

That is four numbers, built from a four-vector, containing exactly the information you want. It is also not a four-vector. The reason is worth stating precisely rather than gesturing at, so let's do that first.

1.1 · Exactly what goes wrong

Being a four-vector means one thing and one thing only, by Chapter 2.4 §2: the four components must transform as Aμ=ΛμνAνA'^{\mu}=\Lambda^{\mu}{}_{\nu}A^{\nu}. So let's test (2.5.1) against that requirement.

The numerator is fine. dxμ\dd x^{\mu} is the prototype four-vector, so dxμ=Λμνdxν\dd x'^{\mu}=\Lambda^{\mu}{}_{\nu}\dd x^{\nu}. The denominator is the problem. tt is not a scalar. It is the zeroth component of a four-vector, and under a boost along xx it becomes

dt  =  γ(dtβcdx)  =  γdt(1βvxc), \dd t' \;=\; \gamma\left(\dd t - \frac{\beta}{c}\,\dd x\right) \;=\; \gamma\,\dd t\left(1-\frac{\beta v_{x}}{c}\right), (2.5.2)

where the second form divides out dt\dd t and uses dx/dt=vx\dd x/\dd t=v_{x}. Now put the transformed numerator over that transformed denominator and see what we are left with:

Vμ  =  dxμdt  =  Λμνdxνγdt(1βvx/c)  =  1γ(1βvx/c)  ΛμνVν. V'^{\mu} \;=\; \frac{\dd x'^{\mu}}{\dd t'} \;=\; \frac{\Lambda^{\mu}{}_{\nu}\,\dd x^{\nu}}{\gamma\,\dd t\,(1-\beta v_{x}/c)} \;=\; \frac{1}{\gamma\left(1-\beta v_{x}/c\right)}\;\Lambda^{\mu}{}_{\nu}V^{\nu}. (2.5.3)

Let's look at what that line is actually saying. VμV^{\mu} transforms almost correctly. It picks up the right matrix Λ\Lambda, and then it is spoiled by an extra scalar factor out front.

Worse, that factor depends on the particle's own velocity vxv_{x}. Two different particles observed in the same boost are therefore multiplied by two different numbers, and no rescaling of the definition can repair that.

This single line is also the whole reason Chapter 2.2's velocity-addition law came out as an awkward quotient rather than a matrix acting on a column. The awkward denominator in u=(u+v)/(1+uv/c2)u'=(u+v)/(1+uv/c^{2}) is the factor in (2.5.3).

Now state the diagnosis in a form that tells you the cure. We divided a four-vector by a frame-dependent quantity. So divide by a frame-independent one instead. Chapter 2.3 §5 supplied exactly one natural candidate along a worldline: the proper time τ\tau, which every observer computes and every observer agrees on.

1.2 · The bridge: dt/dτ=γ\dd t/\dd\tau=\gamma

Before we can use τ\tau we need to know how it relates to coordinate time. This is a two-line recomputation of a result from Chapter 2.3 §5. It is worth having in front of you, because we will use it perhaps thirty times.

Proper time along a worldline is defined by c2dτ2=ds2c^{2}\dd\tau^{2}=\dd s^{2}. So our first job is to evaluate ds2\dd s^{2} between two neighbouring events on the worldline, and then to pull a factor of dt2\dd t^{2} out to the front:

ds2  =  c2dt2dx2dy2dz2  =  c2dt2[11c2((dxdt)2+(dydt)2+(dzdt)2)]. \begin{aligned} \dd s^{2} \;&=\; c^{2}\dd t^{2}-\dd x^{2}-\dd y^{2}-\dd z^{2}\\[4pt] \;&=\; c^{2}\dd t^{2}\left[1-\frac{1}{c^{2}}\left(\left(\dv{x}{t}\right)^{2}+\left(\dv{y}{t}\right)^{2}+\left(\dv{z}{t}\right)^{2}\right)\right]. \end{aligned} (2.5.4)

The bracket is 1v2/c2=1/γ21-v^{2}/c^{2}=1/\gamma^{2}, so c2dτ2=c2dt2/γ2c^{2}\dd\tau^{2}=c^{2}\dd t^{2}/\gamma^{2}. Take the positive root, since both τ\tau and tt increase toward the future, and we have the relation we came for:

  dτ=dtγdtdτ=γ.   \boxed{\;\dd\tau = \frac{\dd t}{\gamma} \qquad\Longleftrightarrow\qquad \dv{t}{\tau} = \gamma.\;} (2.5.5)

Two remarks here will save confusion later.

First, γ\gamma in this formula is the particle's own γ\gamma, evaluated with its instantaneous speed in whatever frame you are using. It is a function of tt, and it has nothing to do with the γ\gamma of any boost between frames.

Second, (2.5.5) is the reason d/dτ\dd/\dd\tau is computable at all. Whenever a τ\tau-derivative turns up, you convert it into the tt-derivative you already know how to take:

ddτ  =  dtdτddt  =  γddt, \dv{}{\tau} \;=\; \dv{t}{\tau}\,\dv{}{t} \;=\; \gamma\,\dv{}{t}, (2.5.6)

which is nothing but the chain rule of Chapter 0.6 §5 applied to a one-parameter reparametrisation.

1.3 · The four-velocity

Now define what we were forced to define:

uμ    dxμdτ. u^{\mu} \;\equiv\; \dv{x^{\mu}}{\tau}. (2.5.7)

This is a four-vector, and the proof is one sentence. The numerator transforms with Λ\Lambda and the denominator is a scalar, so uμ=Λμνuνu'^{\mu}=\Lambda^{\mu}{}_{\nu}u^{\nu} with no leftover factor. That is the entire content of the definition, and it is why relativity keeps replacing d/dt\dd/\dd t by d/dτ\dd/\dd\tau everywhere.

Next we want the components of uμu^{\mu} in terms of quantities a laboratory can measure, so apply the conversion rule (2.5.6) to the definition:

uμ  =  γdxμdt  =  γ(c, v)  =  (γc, γvx, γvy, γvz). u^{\mu} \;=\; \gamma\,\dv{x^{\mu}}{t} \;=\; \gamma\,\big(c,\ \vv v\big) \;=\; \big(\gamma c,\ \gamma v_{x},\ \gamma v_{y},\ \gamma v_{z}\big). (2.5.8)

Note what happened, because it is the pattern of the whole chapter. The repaired object is the old object times γ\gamma.

At low speed γ1\gamma\to1 and uμ(c,v)u^{\mu}\to(c,\vv v), so nothing familiar is lost. At high speed the components run off to infinity. That is exactly the behaviour that lets a fixed-length object point in ever more extreme directions without ever exceeding the speed limit.

Which brings us to the property that makes uμu^{\mu} special.

1.4 · uu=c2u\cdot u=c^{2}, by explicit contraction

We want the Minkowski length of uμu^{\mu}, so contract it with itself using the metric. Chapter 2.4 §4.3 tells us how the bookkeeping goes. The time component keeps its sign and the spatial components flip:

uu    ημνuμuν  =  (u0)2(u1)2(u2)2(u3)2=  γ2c2γ2(vx2+vy2+vz2)=  γ2c2(1v2c2)=  γ2c21γ2  =  c2. \begin{aligned} u\cdot u \;\equiv\; \eta_{\mu\nu}u^{\mu}u^{\nu} \;&=\; \big(u^{0}\big)^{2}-\big(u^{1}\big)^{2}-\big(u^{2}\big)^{2}-\big(u^{3}\big)^{2}\\[3pt] &=\; \gamma^{2}c^{2} - \gamma^{2}\big(v_{x}^{2}+v_{y}^{2}+v_{z}^{2}\big)\\[3pt] &=\; \gamma^{2}c^{2}\left(1-\frac{v^{2}}{c^{2}}\right)\\[3pt] &=\; \gamma^{2}c^{2}\cdot\frac{1}{\gamma^{2}} \;=\; c^{2}. \end{aligned} (2.5.9)

The answer is c2c^{2}, identically. Not for special particles, and not in special frames. It holds for every particle, at every instant, in every frame, whatever that particle happens to be doing. The four-velocity is a four-vector of fixed Minkowski length cc, and all it can do is point in different future-timelike directions.

There is a second proof, shorter, which explains why rather than merely confirming. It is the one to remember. Write the two τ\tau-derivatives as a single ratio of differentials and the answer falls out of the definition of τ\tau:

uu  =  ημνdxμdτdxνdτ  =  ημνdxμdxνdτ2  =  ds2dτ2  =  c2dτ2dτ2  =  c2. u\cdot u \;=\; \eta_{\mu\nu}\dv{x^{\mu}}{\tau}\dv{x^{\nu}}{\tau} \;=\; \frac{\eta_{\mu\nu}\,\dd x^{\mu}\dd x^{\nu}}{\dd\tau^{2}} \;=\; \frac{\dd s^{2}}{\dd\tau^{2}} \;=\; \frac{c^{2}\dd\tau^{2}}{\dd\tau^{2}} \;=\; c^{2}. (2.5.10)

So the constraint is not a fact about particles at all. It is the definition of τ\tau, rearranged.

We chose to parametrise the worldline by its own arc length, and a curve parametrised by arc length has a unit tangent vector. Here "unit" means length cc, because we measured arc length in seconds and tangent vectors in metres per second. Everything in this chapter that looks like a miracle is this sentence in disguise.

1.5 · Four-acceleration, and an orthogonality you get for free

Differentiate again with respect to the invariant:

aμ    duμdτ  =  d2xμdτ2. a^{\mu} \;\equiv\; \dv{u^{\mu}}{\tau} \;=\; \frac{\dd^{2}x^{\mu}}{\dd\tau^{2}}. (2.5.11)

Same argument as before, run twice over: aμa^{\mu} is a four-vector.

Now do something that costs one line and pays for the rest of the chapter. We want to know what the fixed length of uμu^{\mu} forces on aμa^{\mu}, so differentiate the identity (2.5.9) with respect to τ\tau. Its right-hand side is the constant c2c^{2}, so whatever the left-hand side works out to has to vanish:

0  =  ddτ(ημνuμuν)  =  ημνduμdτuν+ημνuμduνdτ  =  2ημνaμuν, 0 \;=\; \dv{}{\tau}\big(\eta_{\mu\nu}u^{\mu}u^{\nu}\big) \;=\; \eta_{\mu\nu}\dv{u^{\mu}}{\tau}u^{\nu} + \eta_{\mu\nu}u^{\mu}\dv{u^{\nu}}{\tau} \;=\; 2\,\eta_{\mu\nu}a^{\mu}u^{\nu}, (2.5.12)

where the two terms merged because ημν\eta_{\mu\nu} is a constant array and is symmetric, so the second sum is the first with the dummy names swapped. Divide by the 2 and what remains is

  ua  =  0   \boxed{\;u\cdot a \;=\; 0\;} (2.5.13)

and it holds always, for every worldline. In Minkowski language, four-acceleration is orthogonal to four-velocity. The word "orthogonal" is doing honest work here. It is the same bilinear form, and the same definition of perpendicularity, that Chapter 0.5 gave for Euclidean space, with one sign changed.

What the orthogonality means

In Euclidean geometry, if a point moves on a sphere of fixed radius, its velocity is tangent to the sphere, which is to say perpendicular to the radius vector. That is the same statement as (2.5.13), one dimension up and one sign over. The four-velocity lives on the hyperboloid uu=c2u\cdot u=c^{2}, and acceleration can only move it along that surface, never off it.

So there is no such thing as accelerating "faster through spacetime". A force can turn your four-velocity, tilting it further from your original time axis and thereby increasing γ\gamma without limit. What it cannot do is lengthen it.

The speed limit is therefore not a barrier the particle runs into. It is the statement that the only motion on offer is rotation of a fixed-length vector, and a rotation never gets you off the surface you started on. Chapter 2.3 §3.2 drew this surface as the invariant hyperbolae, and here it is again, doing dynamics.

It also does immediate accounting work. ua=0u\cdot a=0 says the four-acceleration has three independent components rather than four, since one of them is always fixed by the other three. In §6 that single constraint turns into the work–energy theorem, which you will therefore not have to assume.

Grind box — components of aμa^{\mu}, and ua=0u\cdot a=0 checked the hard way

We will need the components in §6, and computing them is a good drill in (2.5.6). First we need the derivative of γ\gamma. With γ=(1v2/c2)1/2\gamma=(1-v^{2}/c^{2})^{-1/2} and v2=vvv^{2}=\vv v\cdot\vv v, the chain rule gives

dγdt=12(1v2c2)3/2(1c2d(vv)dt)=γ3c2va, \dv{\gamma}{t} = -\frac12\left(1-\frac{v^{2}}{c^{2}}\right)^{-3/2}\cdot\left(-\frac{1}{c^{2}}\dv{(\vv v\cdot\vv v)}{t}\right) = \frac{\gamma^{3}}{c^{2}}\,\vv v\cdot\vv a,

using d(vv)/dt=2va\dd(\vv v\cdot\vv v)/\dd t = 2\vv v\cdot\vv a, with adv/dt\vv a\equiv\dd\vv v/\dd t the ordinary three-acceleration. Write γ˙\dot\gamma for this. Now apply (2.5.6) to (2.5.8):

aμ=γddt(γc, γv)=(γγ˙c,  γγ˙v+γ2a). a^{\mu} = \gamma\dv{}{t}\big(\gamma c,\ \gamma\vv v\big) = \Big(\gamma\dot\gamma c,\ \ \gamma\dot\gamma\,\vv v + \gamma^{2}\vv a\Big).

Now contract, carefully, keeping every term:

ua=(γc)(γγ˙c)(γv)(γγ˙v+γ2a)=γ2γ˙c2γ2γ˙v2γ3va=γ2γ˙c2(1v2c2)γ3va=γ˙c2γ3va. \begin{aligned} u\cdot a &= (\gamma c)(\gamma\dot\gamma c) - (\gamma\vv v)\cdot\big(\gamma\dot\gamma\vv v+\gamma^{2}\vv a\big)\\[3pt] &= \gamma^{2}\dot\gamma c^{2} - \gamma^{2}\dot\gamma v^{2} - \gamma^{3}\,\vv v\cdot\vv a\\[3pt] &= \gamma^{2}\dot\gamma c^{2}\left(1-\frac{v^{2}}{c^{2}}\right) - \gamma^{3}\,\vv v\cdot\vv a\\[3pt] &= \dot\gamma c^{2} - \gamma^{3}\,\vv v\cdot\vv a. \end{aligned}

Substituting γ˙=γ3(va)/c2\dot\gamma=\gamma^{3}(\vv v\cdot\vv a)/c^{2} from the first line makes the two terms cancel identically. ✓ The abstract one-liner (2.5.12) and the brute-force component calculation agree, as they must. The point of doing both once is to see how much labour the index notation is saving.

One consequence worth extracting now. Go to the particle's instantaneous rest frame, where v=0\vv v=0 and uμ=(c,0)u^{\mu}=(c,\vv 0). Then ua=ca0=0u\cdot a = c\,a^{0}=0 forces a0=0a^{0}=0, so in that frame aμ=(0,a0)a^{\mu}=(0,\vv a_{0}), where a0\vv a_{0} is the acceleration the particle itself feels. Call it the proper acceleration. Its invariant square is

aa=0a02=α2,αa0, a\cdot a = 0 - \abs{\vv a_{0}}^{2} = -\,\alpha^{2}, \qquad \alpha \equiv \abs{\vv a_{0}},

so four-acceleration is always a spacelike vector, and its invariant magnitude is the proper acceleration, meaning what an accelerometer bolted to the particle reads. That number is frame-independent. This is why "the rocket burns at 1g1g" is a meaningful statement, while "the rocket's acceleration is 9.8ms29.8\,\mathrm{m\,s^{-2}}" is a statement about somebody's coordinates. Problem 2 uses this.

In plain terms 2.5.1

The object sketched at the close of the last chapter now gets built, and the diagnosis is sharper than the sketch allowed. Divide a displacement between two events by the time one observer's clock assigns to it: the numerator behaves impeccably, being the prototype everything else is measured against, while the denominator belongs to whoever holds the clock. What results is spoiled by a stray factor depending on the particle's own speed, so two particles watched through one change of frame are spoiled by different numbers and no rescaling rescues either.

Divide instead by the reading of the clock the particle carries, which everybody computes and everybody agrees about, and the defect has nowhere to come from. What comes out has the same length for every particle at every instant, and that constraint is not a discovery about matter but the definition of the carried clock rearranged: a curve measured by its own length has a tangent of fixed size.

Differentiate once more and something arrives for nothing. Since the length never changes, its rate of change stands perpendicular to it in the interval's geometry, so a force can turn a fixed-length arrow and can never stretch it. The speed limit stops being a barrier a particle runs into and becomes the observation that turning is the only motion on offer, and turning keeps you on the surface you began on.

2 · Four-momentum, and the argument this chapter turns on

We have two ingredients on the table. One is a four-vector that describes a particle's motion. The other is a scalar that describes the particle itself, namely its rest mass mm, which by the conventions of this Part always means the rest mass and never anything else.

Multiplying a four-vector by a scalar gives a four-vector (2.4 §5.1, operation 2), so we are entitled to define the four-momentum

pμ    muμ  =  γm(c, v)  =  (γmc,  γmv). p^{\mu} \;\equiv\; m\,u^{\mu} \;=\; \gamma m\,\big(c,\ \vv v\big) \;=\; \big(\gamma mc,\ \ \gamma m\vv v\big). (2.5.14)

Its spatial part is γmv\gamma m\vv v, which reduces to the Newtonian mvm\vv v when γ1\gamma\to1. So we have at least a candidate for "momentum", and we will call it pγmv\vv p\equiv\gamma m\vv v.

The time component p0=γmcp^{0}=\gamma mc is, at this stage, an unnamed number that came along for the ride. We are not going to name it until §3, and we are not going to name it by guessing. The name will be forced on us.

Everything now rests on a single question: why should this be the conserved quantity? The answer is the intellectual centre of the chapter, so we take it slowly.

2.1 · What a conservation law has to be

A conservation law is a claim about a physical process. Some quantity computed before the process equals the same quantity computed after it. Consider a collision, with particles coming in and particles going out, and the assertion

inQi  =  outQf. \sum_{\text{in}} Q_{i} \;=\; \sum_{\text{out}} Q_{f}. (2.5.15)

Now ask the question that Chapter 2.4 taught us to ask about every equation. Who is asserting this, and does the assertion survive translation into another observer's language?

The question is not pedantry. If the assertion does not survive, then (2.5.15) is not a law of nature. It is a report from one laboratory, of no more fundamental standing than "the train is moving".

Chapter 2.4 §6 gave the tool for deciding. If the two sides of an equation are tensors of the same type, then the equation holds in every frame as soon as it holds in one. If they are not, all bets are off.

Chapter 2.4's closing warning was blunt about which side of that line statements like p=0\vv p=0 fall on. They are statements about components, and therefore about a frame, and they must never be carried across a boost unexamined.

So the demand is clear. Whatever is conserved must be a tensor, and the conservation statement must be an equation between tensors of the same type. Three numbers are not a tensor under the Lorentz group. Four numbers that transform with Λ\Lambda are.

2.2 · Newtonian momentum conservation does not survive a boost

That is an argument from principle. Now let's see the failure in arithmetic, in a collision you can hold in your head.

Two identical lumps of putty, each of rest mass mm, approach each other along the xx-axis in frame SS with speeds +u+u and u-u. They collide and stick. By symmetry the composite is at rest in SS, and we will call its rest mass MM, whatever that turns out to be.

Start by checking Newtonian momentum conservation in SS, where the arithmetic is easiest:

mu+m(u)before  =  0  =  M0after. \underbrace{mu + m(-u)}_{\text{before}} \;=\; 0 \;=\; \underbrace{M\cdot 0}_{\text{after}}. \qquad\checkmark (2.5.16)

Perfect. But note how little that check actually tested. By symmetry it would have passed for any definition of momentum of the form p=mf(v)v\vv p = m f(v)\,\vv v, with ff any function whatever of the speed, because the two terms cancel no matter what ff is.

A symmetric collision in a single frame is almost completely uninformative. The information is in the boost, so let's boost.

Move to a frame SS' travelling at v-v along xx relative to SS, so that in SS' everything is drifting in the +x+x direction. Chapter 2.2's velocity-addition law gives the new velocities of the two lumps:

u1  =  u+v1+uv/c2,u2  =  u+v1uv/c2, u_{1}' \;=\; \frac{u+v}{1+uv/c^{2}}, \qquad\qquad u_{2}' \;=\; \frac{-u+v}{1-uv/c^{2}}, (2.5.17)

and the composite, at rest in SS, moves at vv in SS'. Our next goal is the incoming Newtonian momentum in SS', so add those two velocities. Write auv/c2a\equiv uv/c^{2} for compactness and put the two fractions over the common denominator 1a21-a^{2}:

u1+u2=(u+v)(1a)+(vu)(1+a)(1+a)(1a)=(u+vauav)+(vu+avau)1a2=2v2au1a2  =  2v(1u2/c2)1u2v2/c4, \begin{aligned} u_{1}'+u_{2}' &= \frac{(u+v)(1-a) + (v-u)(1+a)}{(1+a)(1-a)}\\[4pt] &= \frac{\big(u+v-au-av\big) + \big(v-u+av-au\big)}{1-a^{2}}\\[4pt] &= \frac{2v-2au}{1-a^{2}} \;=\; \frac{2v\left(1-u^{2}/c^{2}\right)}{1-u^{2}v^{2}/c^{4}}, \end{aligned} (2.5.18)

where the last step used au=u2v/c2au = u^{2}v/c^{2}, so that 2v2au=2v(1u2/c2)2v-2au = 2v(1-u^{2}/c^{2}).

Meanwhile the Newtonian momentum after the collision is MvMv, and Newtonian mechanics also insists that mass is additive, M=2mM=2m. So here are the two sides of the conservation law as SS' reads them:

2mv1u2/c21u2v2/c4beforeversus2mvafter. \underbrace{2mv\,\frac{1-u^{2}/c^{2}}{1-u^{2}v^{2}/c^{4}}}_{\text{before}} \qquad\text{versus}\qquad \underbrace{2mv}_{\text{after}}. (2.5.19)

These are equal only if u=0u=0 or v=0v=0, which is to say only if there was no collision or no boost. Put in numbers. For u=v=0.6cu=v=0.6c the fraction is (10.36)/(10.1296)=0.64/0.8704=0.7353(1-0.36)/(1-0.1296)=0.64/0.8704=0.7353, so an observer in SS' finds the outgoing momentum 36% larger than the incoming momentum. The books balance in SS and fail badly in SS'.

⚠ What exactly failed

It is tempting to say "Newtonian momentum isn't conserved at high speed". That is not what the calculation showed. In frame SS it was conserved exactly, to all orders, as (2.5.16) shows. What failed is agreement between observers about whether it was conserved. One frame says the law holds, and another says it is violated by 36%.

That is a much worse disease than being approximately wrong, and no small correction term can cure it. A quantity whose conservation is frame-dependent is not describing anything about the collision. It is describing the laboratory.

This is precisely the defect Chapter 2.4 §6.2 diagnosed in F=ma\vv F=m\vv a. That chapter said its content changes when you change frames, which disqualifies it as a statement about nature rather than about a laboratory. Here the same defect has been shown rather than asserted.

2.3 · The four-vector version cannot fail

Now run the same collision with pμp^{\mu} and watch the disease vanish for structural reasons.

Suppose that in frame SS the total four-momentum is conserved, meaning inpμ=outpμ\sum_{\text{in}}p^{\mu} = \sum_{\text{out}}p^{\mu} for all four components. To test that claim in another frame we want a single object to transform, so define the shortfall

ΔPμ    outpμ    inpμ. \Delta P^{\mu} \;\equiv\; \sum_{\text{out}}p^{\mu} \;-\; \sum_{\text{in}}p^{\mu}. (2.5.20)

Each pμp^{\mu} is a four-vector, and sums and differences of tensors of the same type are tensors of that type (2.4 §5.1, operation 1). So ΔPμ\Delta P^{\mu} is itself a four-vector.

Our hypothesis is now the tensor equation ΔPμ=0\Delta P^{\mu}=0, and 2.4 §6 says a tensor equation true in one frame is true in all of them. Written out explicitly, that reads

ΔPμ  =  ΛμνΔPν  =  Λμν0  =  0. \Delta P'^{\mu} \;=\; \Lambda^{\mu}{}_{\nu}\,\Delta P^{\nu} \;=\; \Lambda^{\mu}{}_{\nu}\cdot 0 \;=\; 0. (2.5.21)

That is the whole proof, and it is deliberately anticlimactic. The transformation law is linear and homogeneous, so it maps zero to zero, and there is nowhere for a violation to come from. If four-momentum is conserved for one observer, it is conserved for every observer, automatically, with no further conditions.

Grind box — the same collision, done relativistically, term by term

The abstract argument is airtight, but you should see the numbers come out once. We need one identity, and it is worth having permanently.

The velocity-addition identity for γ\gamma. If u=(u+v)/(1+uv/c2)u'=(u+v)/(1+uv/c^{2}), then

1u2c2=(1+uvc2)2(u+v)2c2(1+uvc2)2. 1-\frac{u'^{2}}{c^{2}} = \frac{\left(1+\dfrac{uv}{c^{2}}\right)^{2}-\dfrac{(u+v)^{2}}{c^{2}}}{\left(1+\dfrac{uv}{c^{2}}\right)^{2}}.

Expand the numerator, being careful with every term:

num=1+2uvc2+u2v2c4u2c22uvc2v2c2=1u2c2v2c2+u2v2c4  =  (1u2c2)(1v2c2). \begin{aligned} \text{num} &= 1 + \frac{2uv}{c^{2}} + \frac{u^{2}v^{2}}{c^{4}} - \frac{u^{2}}{c^{2}} - \frac{2uv}{c^{2}} - \frac{v^{2}}{c^{2}}\\[3pt] &= 1 - \frac{u^{2}}{c^{2}} - \frac{v^{2}}{c^{2}} + \frac{u^{2}v^{2}}{c^{4}} \;=\; \left(1-\frac{u^{2}}{c^{2}}\right)\left(1-\frac{v^{2}}{c^{2}}\right). \end{aligned}

The 2uv/c22uv/c^{2} terms cancel and what remains factorises. Taking reciprocals and square roots,

  γ(u)=γ(u)γ(v)(1+uvc2).   \boxed{\;\gamma(u') = \gamma(u)\,\gamma(v)\left(1+\frac{uv}{c^{2}}\right).\;}

This is just the statement that rapidities add, cosh(ϕu+ϕv)=coshϕucoshϕv+sinhϕusinhϕv\cosh(\phi_{u}+\phi_{v}) = \cosh\phi_{u}\cosh\phi_{v} +\sinh\phi_{u}\sinh\phi_{v}, written in velocities. Chapter 2.2 §5 proved the rapidity version.

Momentum before, in SS'. Using the identity with uuu\to u and then uuu\to-u:

γ(u1)u1=γ(u)γ(v)(1+uvc2)u+v1+uv/c2=γ(u)γ(v)(u+v), \gamma(u_{1}')\,u_{1}' = \gamma(u)\gamma(v)\left(1+\frac{uv}{c^{2}}\right)\cdot\frac{u+v}{1+uv/c^{2}} = \gamma(u)\gamma(v)\,(u+v), γ(u2)u2=γ(u)γ(v)(1uvc2)vu1uv/c2=γ(u)γ(v)(vu). \gamma(u_{2}')\,u_{2}' = \gamma(u)\gamma(v)\left(1-\frac{uv}{c^{2}}\right)\cdot\frac{v-u}{1-uv/c^{2}} = \gamma(u)\gamma(v)\,(v-u).

The messy denominators cancel exactly, which is the identity earning its keep. Adding the two and multiplying by mm:

beforeγmu  =  mγ(u)γ(v)[(u+v)+(vu)]  =  2mγ(u)γ(v)v. \sum_{\text{before}} \gamma m\,u' \;=\; m\gamma(u)\gamma(v)\big[(u+v)+(v-u)\big] \;=\; 2m\,\gamma(u)\gamma(v)\,v.

Momentum after, in SS'. The composite has rest mass MM and moves at vv, so its momentum is γ(v)Mv\gamma(v)Mv. Equality demands

γ(v)Mv=2mγ(u)γ(v)vM=2γ(u)m. \gamma(v)\,M\,v = 2m\,\gamma(u)\gamma(v)\,v \qquad\Longrightarrow\qquad M = 2\gamma(u)\,m.

Read that twice. Relativistic momentum conservation in the boosted frame does not merely permit the composite's rest mass to exceed 2m2m. It requires it, and fixes the excess exactly. The lump of putty is heavier than its parts, by a factor γ(u)\gamma(u), and nothing in the calculation gave us a choice. Section 9 is about that number.

The time components, for free. Do the same with p0=γmcp^{0}=\gamma mc:

beforeγmc=mcγ(u)γ(v)[(1+uvc2)+(1uvc2)]=2mcγ(u)γ(v), \sum_{\text{before}}\gamma mc = mc\,\gamma(u)\gamma(v)\left[\left(1+\frac{uv}{c^{2}}\right)+\left(1-\frac{uv}{c^{2}}\right)\right] = 2mc\,\gamma(u)\gamma(v),

while after the collision p0=γ(v)Mc=2γ(u)γ(v)mcp^{0}=\gamma(v)Mc = 2\gamma(u)\gamma(v)mc, using the MM just derived. Equal ✓, and note that we did not impose this. It came out. That is (2.5.21) being true in coordinates.

And in SS itself. Before the collision p0=2γ(u)mcp^{0}=2\gamma(u)mc, and after it p0=Mc=2γ(u)mcp^{0}=Mc=2\gamma(u)mc ✓, while the spatial parts vanish on both sides by symmetry. So the four-vector statement holds in SS, holds in SS', and by (2.5.21) holds in every frame there is.

2.4 · The part you cannot refuse: momentum forces energy

So far we have shown that four-momentum conservation is self-consistent across frames. Now comes the sharper statement, and it is the one that makes this chapter inevitable.

Suppose you are a minimalist. You accept that the conserved object must be built from the four-vector pμp^{\mu}, since you have seen §2.2 and you are not going back. But you want to conserve only the three spatial components, because those are the ones you have experimental evidence for. You want the law "ΔP=0\Delta\vv P=0 in every frame", and you decline to say anything at all about ΔP0\Delta P^{0}.

You cannot have it. Here is why, in three lines.

Theorem — conserving three components forces the fourth

Let ΔPμ\Delta P^{\mu} be a four-vector. If its three spatial components vanish in every inertial frame, then its time component vanishes too.

Proof. By hypothesis ΔPμ=(ΔP0,0,0,0)\Delta P^{\mu}=(\Delta P^{0},\,0,\,0,\,0) in the original frame. Boost along xx with speed βc\beta c. By Chapter 2.4's transformation law (2.5.21) the new first spatial component is

ΔP1  =  γ(ΔP1βΔP0)  =  γβΔP0. \Delta P'^{1} \;=\; \gamma\big(\Delta P^{1} - \beta\,\Delta P^{0}\big) \;=\; -\,\gamma\beta\,\Delta P^{0}.

The hypothesis says this must vanish for every β\beta. Since γβ0\gamma\beta\neq0 for β0\beta\neq0, we need ΔP0=0\Delta P^{0}=0. \blacksquare

Look at what that argument did. It used no dynamics, no experiment, and no assumption about forces or collisions. It used only the transformation law.

And what it says is this. A universe in which momentum is conserved for every observer is a universe in which p0p^{0} is conserved too. You are not permitted to conserve three components of a four-vector. The boost mixes the fourth into the first three and drags it into the law whether you invited it or not.

That is the whole chapter in one paragraph. The demand that conservation laws be frame-independent selects pμp^{\mu} out of all the candidate momenta. Having selected it, you are stuck with its time component, a fourth conserved quantity you did not ask for, did not derive from any experiment, and cannot discard. Section 3 asks what on earth it is.

⚠ Two honest caveats about what has been proved

This is not a proof that anything is conserved. Nothing so far says nature conserves pμp^{\mu}. That is a physical claim, and Chapter 1.4 gave its real origin: Noether's theorem, applied to the invariance of the laws under translations in space, which gives momentum, and in time, which gives energy.

What we have proved here is conditional, and it comes in two parts. If some four-vector quantity is conserved in one frame, it is conserved in all of them. And if a frame-independent momentum conservation law exists at all, then the conserved object must be a four-vector and its time component comes along with it.

Relativity narrows the field of candidates to essentially one. Experiment then confirms that this one is right, and it does so every day in every particle detector on Earth.

The choice of mm as the scalar was not forced by tensor algebra alone. pμ=muμp^{\mu}=mu^{\mu} is a four-vector, but so is f(m)uμf(m)u^{\mu} for any function ff, and so is mg(m/M0)uμm\,g(m/M_{0})\,u^{\mu} for any dimensionless combination you can invent. What fixes the choice is the correspondence limit. As γ1\gamma\to1 the spatial part must become the Newtonian mvm\vv v that three centuries of experiment supports, and that pins the scalar to mm itself. Covariance narrows the field and correspondence chooses within it. That pair of moves is the method of the entire second half of this book.

In plain terms 2.5.2

Every conservation law is somebody's claim about a process, and what matters is whether the claim survives translation into another observer's language, since one that does not is a report about a laboratory rather than a fact about the world. Three numbers cannot survive, and the failure is not gentle. Two identical lumps of putty approaching at equal speeds and sticking conserve the old momentum exactly in the symmetric frame, while an observer drifting past finds the outgoing total larger than the incoming one by more than a third.

Notice precisely what that is. It is not that the old momentum is slightly wrong at high speed, since in that frame it was right to every order. What fails is agreement about whether the law holds at all, and a disease of that kind admits no correction term. Conserve the four-part object instead and the disease has nowhere to live, since the transformation between observers is linear and carries zero to zero.

Then comes the part that cannot be declined. Suppose you accept the object but wish to conserve only its three familiar entries. A change of frame mixes the fourth into the first three, so demanding that the three vanish for every observer forces the fourth to vanish as well. Frame-independence has selected the object out of all the candidates, and having taken three of its components you are stuck with the fourth.

3 · What the time component is

We have an unnamed conserved quantity cp0=γmc2c\,p^{0}=\gamma mc^{2}. Its dimensions are those of energy, since it is a mass times a velocity squared. Dimensions are suggestive and prove nothing, so let's find out what it actually is by the only honest route available. We expand it at low speed and see what it turns into.

3.1 · The expansion

Apply the binomial series of Chapter 0.3 §2, which for arbitrary real exponent α\alpha reads (1+x)α=k0(αk)xk(1+x)^{\alpha}=\sum_{k\ge0}\binom{\alpha}{k}x^{k} with (αk)=α(α1)(αk+1)/k!\binom{\alpha}{k}=\alpha(\alpha-1)\cdots(\alpha-k+1)/k!. Take α=12\alpha=-\half and x=β2x=-\beta^{2}, where β=v/c\beta=v/c, and the first few coefficients come out as

(1/21)=12,(1/22)=(12)(32)2=38,(1/23)=(12)(32)(52)6=516. \begin{aligned} \binom{-1/2}{1} &= -\frac12, \qquad\qquad \binom{-1/2}{2} = \frac{(-\frac12)(-\frac32)}{2} = \frac38,\\[5pt] \binom{-1/2}{3} &= \frac{(-\frac12)(-\frac32)(-\frac52)}{6} = -\frac{5}{16}. \end{aligned} (2.5.22)

Substituting x=β2x=-\beta^{2} flips the sign of every odd power, so all the signs become positive:

γ  =  (1β2)1/2  =  1+12β2+38β4+516β6+ \gamma \;=\; \big(1-\beta^{2}\big)^{-1/2} \;=\; 1 + \frac12\beta^{2} + \frac38\beta^{4} + \frac{5}{16}\beta^{6} + \cdots (2.5.23)

That is the expansion of γ\gamma alone. The quantity we actually want is γmc2\gamma mc^{2}, so multiply through by mc2mc^{2} and restore β=v/c\beta=v/c, keeping an eye on what each term becomes:

γmc2  =  mc2constant  +  12mv2Newtonian  +  38mv4c2first correction  +  516mv6c4+ \gamma mc^{2} \;=\; \underbrace{mc^{2}}_{\text{constant}} \;+\; \underbrace{\half mv^{2}}_{\text{Newtonian}} \;+\; \underbrace{\frac38\frac{mv^{4}}{c^{2}}}_{\text{first correction}} \;+\; \frac{5}{16}\frac{mv^{6}}{c^{4}} + \cdots (2.5.24)

3.2 · The identification, made carefully

Now the part that deserves care, because it is usually skated over. What (2.5.24) shows is a mathematical fact about a function, not a physical identification. It says that the conserved quantity γmc2\gamma mc^{2} equals a constant, plus the Newtonian kinetic energy, plus terms that vanish as v/c0v/c\to0.

Getting from there to "γmc2\gamma mc^{2} is the energy" takes one further argument, and here it is.

Newtonian mechanics has its own conserved quantity for elastic collisions, i12mivi2\sum_{i}\half m_{i} v_{i}^{2}, verified to exhaustion at low speeds. Our new law says instead that iγimic2\sum_{i}\gamma_{i}m_{i}c^{2} is conserved. We want to see how the two claims are related, so expand the new one particle by particle at low speed:

iγimic2  =  imic2()  +  i12mivi2  +  O ⁣(v4c2). \sum_{i}\gamma_{i}m_{i}c^{2} \;=\; \underbrace{\sum_{i}m_{i}c^{2}}_{(\ast)} \;+\; \sum_{i}\half m_{i}v_{i}^{2} \;+\; O\!\left(\frac{v^{4}}{c^{2}}\right). (2.5.25)

Now suppose the collision does not change what the particles are, so the same species go in as come out. That is the only kind of collision Newtonian mechanics ever contemplated.

Then the sum ()(\ast) is the same before and after, and it cancels out of the conservation statement entirely. What is left is exactly Newtonian kinetic-energy conservation.

So the new law contains the old one, and the quantity it conserves must be what the old theory called energy. That, and only that, is what licenses the name:

  E    γmc2,pμ  =  (Ec, p),p=γmv.   \boxed{\;E \;\equiv\; \gamma mc^{2}, \qquad p^{\mu} \;=\; \left(\frac{E}{c},\ \vv p\right), \qquad \vv p = \gamma m\vv v.\;} (2.5.26)

The four-momentum is the energy and the momentum, welded into one object by the boost that mixes them. From here on pμp^{\mu} is called the energy–momentum four-vector, and the two conservation laws you learned separately are one law with four components.

The constant that refuses to be a constant

In Newtonian mechanics the zero of energy is arbitrary. Adding a constant to EE changes no equation of motion, no conservation law, and no prediction. That freedom is why nobody ever asked what the "true" energy of a stationary object was. The question had no content.

Relativity removes the freedom, and the mechanism is visible in (2.5.25). The would-be constant is imic2\sum_{i}m_{i}c^{2}, and it cancels only if the rest masses do not change.

Processes exist that change them. A nucleus binds. A particle decays. Two lumps of putty stick together and end up with rest mass 2γm2\gamma m rather than 2m2m, as §2.3's grind box proved. In any such process the sum ()(\ast) is different before and after, and the difference has to be paid for out of kinetic energy.

So mc2mc^{2} is not an additive constant one may drop. It is a reservoir, and processes exist that draw on it. Setting v=0v=0 in (2.5.26) names it:

E0  =  mc2. E_{0} \;=\; mc^{2}.

That is the rest energy. Note carefully what it is not. It is not a formula for the energy of a moving body, which is γmc2\gamma mc^{2}. It is not a claim that mass "is" energy in some substantial sense. It is the value of the conserved quantity EE for a particle that happens to be at rest, and its only physical content is the one just derived: changes in rest mass show up as changes in kinetic energy, at an exchange rate of c2c^{2}.

3.3 · The first correction, and where you can see it

The 38mv4/c2\tfrac38 mv^{4}/c^{2} term in (2.5.24) is not decoration. To judge its size we need something to judge it against, so compare it with the Newtonian term:

38mv4/c212mv2  =  34v2c2  =  34β2. \frac{\tfrac38 mv^{4}/c^{2}}{\tfrac12 mv^{2}} \;=\; \frac34\,\frac{v^{2}}{c^{2}} \;=\; \frac34\beta^{2}. (2.5.27)

The correction is quadratically small in β\beta. That is why Newtonian mechanics survived so long, and it is also why the correction becomes visible the moment you have a system with a characteristic speed and a spectroscope.

The electron in a hydrogen atom has βα=1/137\beta\approx\alpha=1/137. Chapter 0.3 §5 got that from dimensional analysis alone, since the natural speed ke/k_{e}/\hbar divided by cc is exactly α\alpha. Putting that number into the ratio above gives

34β2    34α2  =  34×5.325×105  =  3.99×105, \frac34\beta^{2} \;\approx\; \frac34\alpha^{2} \;=\; \frac34\times5.325\times10^{-5} \;=\; 3.99\times10^{-5}, (2.5.28)

a relative shift of four parts in 10510^{5} in the atom's energy levels. On a 13.6eV13.6\,\mathrm{eV} scale that is of order α2×13.6eV7×104eV\alpha^{2}\times13.6\,\mathrm{eV}\approx7\times10^{-4}\,\mathrm{eV}, which is comfortably resolvable.

It is one of the three contributions to the fine structure of hydrogen that Chapter 0.3's worked example 2 flagged and could not yet explain. Chapter 4.16 computes the full splitting. The point here is only that the leading piece of it is a term in (2.5.24) and nothing more exotic.

Grind box — the same correction in momentum, and why the coefficient changes sign

Open a quantum mechanics text and you will find the relativistic correction to the hydrogen Hamiltonian written as p^4/8m3c2-\hat p^{4}/8m^{3}c^{2}, a negative term with coefficient 18\tfrac18. We just derived a positive term with coefficient 38\tfrac38. Both are correct. Chapter 0.3's worked example 2 noticed the clash and named the culprit. Now that we have actually derived p=γmv\vv p=\gamma m\vv v and the mass shell, we can close it arithmetically.

The kinetic energy in terms of pp. Anticipating §4's mass shell, E=p2c2+m2c4E=\sqrt{p^{2}c^{2}+m^{2}c^{4}}. Factor out mc2mc^{2} and expand with the binomial series again, now with x=p2/m2c2x=p^{2}/m^{2}c^{2} and α=+12\alpha=+\tfrac12, for which (1/21)=12\binom{1/2}{1}=\tfrac12 and (1/22)=(12)(12)2=18\binom{1/2}{2}=\frac{(\frac12)(-\frac12)}{2}=-\tfrac18:

E=mc2(1+p2m2c2)1/2=mc2[1+p22m2c2p48m4c4+]=mc2+p22mp48m3c2+ E = mc^{2}\left(1+\frac{p^{2}}{m^{2}c^{2}}\right)^{1/2} = mc^{2}\left[1+\frac{p^{2}}{2m^{2}c^{2}} - \frac{p^{4}}{8m^{4}c^{4}}+\cdots\right] = mc^{2}+\frac{p^{2}}{2m}-\frac{p^{4}}{8m^{3}c^{2}}+\cdots

There it is, sign and coefficient. Quantum mechanics uses this form because p^\hat p is the natural operator, not vv.

Why the two agree. The variable is different, because p=γmvp=\gamma mv rather than mvmv. So expand γ\gamma to the order needed,

p2=γ2m2v2=m2v2(1v2c2)1=m2v2(1+v2c2+), p^{2} = \gamma^{2}m^{2}v^{2} = m^{2}v^{2}\left(1-\frac{v^{2}}{c^{2}}\right)^{-1} = m^{2}v^{2}\left(1+\frac{v^{2}}{c^{2}}+\cdots\right),

so the two terms become

p22m=mv22+mv42c2+,p48m3c2=m4v48m3c2+=mv48c2+ \frac{p^{2}}{2m} = \frac{mv^{2}}{2}+\frac{mv^{4}}{2c^{2}}+\cdots, \qquad -\frac{p^{4}}{8m^{3}c^{2}} = -\frac{m^{4}v^{4}}{8m^{3}c^{2}}+\cdots = -\frac{mv^{4}}{8c^{2}}+\cdots

Add them: 12mv2+(1218)mv4/c2=12mv2+38mv4/c2\tfrac12 mv^{2} + \left(\tfrac12-\tfrac18\right)mv^{4}/c^{2} = \tfrac12 mv^{2}+\tfrac38 mv^{4}/c^{2} ✓, which is (2.5.24). The apparent discrepancy was bookkeeping. A 12\tfrac12 of the quartic term comes from pp not being mvmv, and a 18-\tfrac18 comes from the square root, and 1218=38\tfrac12-\tfrac18=\tfrac38.

The moral is one you will need again in Part V: a perturbative coefficient is meaningless until you say which variable is being held fixed. Terms move between orders when you change variables. Only the physical prediction is invariant, which here means the size of the fine-structure shift, and that is α2\sim\alpha^{2} either way.

In plain terms 2.5.3

A fourth conserved quantity stands about unnamed, and the only honest route to its identity is to expand it for a slow particle and watch what it becomes. Out comes a constant belonging to the particle, then the Newtonian kinetic energy, then corrections falling with the square of the speed ratio. Nothing there is a concession: nearly everything is an approximation, and the controlled kind names its next term. What licenses the name is that when a collision leaves the participants the species it found them, the constant cancels and kinetic energy is conserved as before.

The constant is where the interest lies. Newtonian mechanics let you shift the zero of energy by any amount, which is why nobody asked a stationary object's energy; the question had no content. That freedom is withdrawn, since the constant cancels only while the rest masses hold, and processes change them: a nucleus binds, a particle decays, two lumps of putty stick and come out heavier than their parts.

Rest energy is a reservoir rather than an offset, at a fixed exchange rate against motion. Notice how little was requested and how much arrived. Nobody went looking for it, no experiment produced it, no argument about matter suggested it. It was conscripted: the demand that a law read alike for everybody left a fourth component lying about, and what it turned out to be is the equation everyone recites.

4 · The mass shell

Every four-vector has an invariant square, and the invariant square of pμp^{\mu} turns out to contain everything.

4.1 · pp=m2c2p\cdot p=m^{2}c^{2}, and what it says

Contract pμ=muμp^{\mu}=mu^{\mu} with itself. The mass is a scalar, so it comes straight out of the contraction, and (2.5.9) finishes the job:

pp  =  ημν(muμ)(muν)  =  m2(uu)  =  m2c2. p\cdot p \;=\; \eta_{\mu\nu}\,(mu^{\mu})(mu^{\nu}) \;=\; m^{2}\,\big(u\cdot u\big) \;=\; m^{2}c^{2}. (2.5.29)

That is the answer in terms of the mass. We also want it in terms of the things a detector actually reports, so compute the same quantity from the components in (2.5.26), remembering that lowering flips the sign of the spatial parts:

pp  =  (p0)2p2  =  E2c2p2,pp. p\cdot p \;=\; \big(p^{0}\big)^{2} - \abs{\vv p}^{2} \;=\; \frac{E^{2}}{c^{2}} - p^{2}, \qquad p\equiv\abs{\vv p}. (2.5.30)

Two expressions for one quantity. Set them equal and multiply through by c2c^{2}:

  E2  =  p2c2+m2c4.   \boxed{\;E^{2} \;=\; p^{2}c^{2} + m^{2}c^{4}.\;} (2.5.31)

This is the mass-shell relation, and it is the most useful equation in particle physics. Three observations before we use it.

It is an invariant statement. The left side involves EE and the right involves pp. Both change under a boost, and the particular combination E2p2c2E^{2}-p^{2}c^{2} does not. Different observers assign different energies and momenta to the same electron, and all of them compute the same mm.

So mass is a Lorentz invariant. It is not a quantity that grows with speed. It is the fixed label attached to the four-vector, exactly as the interval is the fixed label attached to a pair of events.

It is Chapter 2.3's geometry, in different variables. Plot EE against pp and (2.5.31) is a hyperbola with asymptote E=pcE=pc. Compare 2.3 §3.2, where the orbits of a boost in the (ct,x)(ct,x) plane were the hyperbolae c2t2x2=constc^{2}t^{2}-x^{2}=\text{const} with asymptotes at 4545^{\circ}.

Same curves, and for the same reason. A boost is a hyperbolic rotation, and it slides any four-vector along the level surface of its own invariant square. Momentum space is spacetime with the axes relabelled.

It has no reference to τ\tau. That matters more than it looks, and §4.3 collects the debt.

4.2 · Velocity from the four-momentum

Before we take the massless case we need one more result, and the order matters, because this result is what makes the massless case work at all. We want the particle's velocity written in terms of EE and p\vv p. Both components of pμp^{\mu} carry the same factor γm\gamma m, so dividing one by the other kills it:

pE  =  γmvγmc2  =  vc2  v  =  pc2E.   \frac{\vv p}{E} \;=\; \frac{\gamma m\vv v}{\gamma mc^{2}} \;=\; \frac{\vv v}{c^{2}} \qquad\Longrightarrow\qquad \boxed{\;\vv v \;=\; \frac{\vv p\,c^{2}}{E}.\;} (2.5.32)

Read it as a recipe. Given a particle's energy and momentum, its velocity is the ratio. No γ\gamma, no mm, no square roots.

Two checks. At low speed, Emc2E\approx mc^{2} and pmv\vv p\approx m\vv v, so the ratio returns v\vv v ✓. And in general EpcE\ge pc by (2.5.31) whenever m20m^{2}\ge0, so the formula automatically gives vcv\le c. The speed limit is not imposed on (2.5.32). It falls out of it.

4.3 · Massless particles

Set m=0m=0 in (2.5.31). Taking the positive root, since energies of real particles are positive:

E  =  pc. E \;=\; pc. (2.5.33)

That fixes the relation between energy and momentum. To get the speed, feed it into the recipe (2.5.32) and watch the momentum cancel:

v  =  pc2E  =  pc2pc  =  c. v \;=\; \frac{pc^{2}}{E} \;=\; \frac{pc^{2}}{pc} \;=\; c. (2.5.34)

Exactly cc, for every energy, with no limiting process and no approximation. A massless particle does not travel near the speed of light. It travels at it, always, and it has no rest frame in which to be at rest.

⚠ Why this isn't obvious — "massless" is a statement about a four-vector

It is natural to read m=0m=0 as the end of a sequence: particles get lighter and lighter, and in the limit you have a photon. That reading is wrong, and it is wrong in a way that will cost you.

The construction pμ=muμp^{\mu}=mu^{\mu} does not survive the limit. As m0m\to0 at fixed energy, γ\gamma\to\infty, so pμ=muμp^{\mu}=mu^{\mu} becomes 0×0\times\infty. That is an indeterminate form, not a limit.

Worse, uμ=dxμ/dτu^{\mu}=\dd x^{\mu}/\dd\tau needs τ\tau, and a null worldline has ds2=0\dd s^{2}=0 at every point, so dτ=0\dd\tau=0 identically and there is no proper time to divide by. Photons do not have worldlines parametrised by proper time. A photon has no clock, and "the time experienced by a photon" is not a small number. It is a meaningless phrase.

The correct definition takes pμp^{\mu} as primary. A massless particle is one whose four-momentum is a null vector, pp=0p\cdot p=0. That is a perfectly good statement with no τ\tau, no γ\gamma, and no mm multiplying anything.

Everything else follows from it. (2.5.31) gives E=pcE=pc, and (2.5.32) gives v=cv=c, and Chapter 2.3 §4's causal classification already told you where such a vector points, namely along the light cone. Nothing new had to be assumed. The massless case is the null case of a classification we already had.

This inversion is worth naming, because it is the version that survives into quantum field theory. The four-momentum is fundamental, and mass is the label ppp\cdot p carries. In field theory particles are excitations of fields and mm appears as a parameter in a Lagrangian long before anything is moving.

m = 1.000 p = 0.620 E = √(p²+m²) = 1.1766
v/c = pc²/E = 0.5269 γ = E/mc² = 1.1766
Newtonian m + p²/2m = 1.1922 error = 1.325 %
Newtonian error passes 1% at p = 0.571 (v/c = 0.496, γ = 1.152)
The mass shell. Inside this figure only, units are chosen with c=1c=1 and everything measured in GeV, so EE, pcpc and mc2mc^{2} are all just numbers on the same scale and the asymptote E=pcE=pc is the line at 4545^{\circ}. Every curve and every readout is computed numerically from (2.5.31), (2.5.32) and ENewt=mc2+p2/2mE_{\text{Newt}}=mc^{2}+p^{2}/2m; nothing is drawn from a stored result. (i) The bold blue curve is the exact hyperbola E=p2c2+m2c4E=\sqrt{p^{2}c^{2}+m^{2}c^{4}}; the faint curves behind it are the same relation at other masses — the invariant hyperbolae of Chapter 2.3 §3.2, now in momentum space, one for each value of ppp\cdot p. (ii) The orange dashed line is the asymptote E=pcE=pc, the degenerate member m=0m=0 of that family — the light cone of momentum space. Drag the mass slider down and watch the hyperbola flatten onto it; at m=0m=0 the curve is the line, and the readout shows v/cv/c locked at exactly 11. (iii) The purple curve is the Newtonian energy mc2+p2/2mmc^{2}+p^{2}/2m. It leaves the hyperbola with the same value and the same slope at p=0p=0 — the two agree to second order, which is (2.5.24) — and then diverges upward without limit, because Newtonian mechanics believes any energy buys any speed. Move the momentum slider and watch the error readout: it stays under 1% until the particle is moving at half the speed of light, then climbs through 6% at v=0.71cv=0.71c and 34% at v=0.89cv=0.89c. That quartic-then-catastrophic behaviour is why nobody noticed relativity until they built accelerators. (iv) The marker sits on the exact curve at the momentum you choose; drag it, or click anywhere in the frame. Its readouts are EE, γ=E/mc2\gamma=E/mc^{2}, and v/c=pc/Ev/c=pc/E — the last computed from (2.5.32) and not from any γ\gamma.
Grind box — why the Newtonian curve is so good for so long

The figure claims the Newtonian energy is within 1% of the truth up to v0.5cv\approx0.5c, which is surprising if you expect errors of order β2\beta^{2}. Here is the reason, and it is a nice piece of Chapter 0.3 bookkeeping.

Work in units of mcmc and mc2mc^{2}. Write x=p/mcx=p/mc, so the exact energy is E/mc2=1+x2E/mc^{2}=\sqrt{1+x^{2}} and the Newtonian one is 1+x2/21+x^{2}/2. The fractional error is

ε(x)  =  1+x2/21+x21. \varepsilon(x) \;=\; \frac{1+x^{2}/2}{\sqrt{1+x^{2}}} - 1.

Expand with the binomial series, (1+x2)1/2=112x2+38x4(1+x^{2})^{-1/2}=1-\tfrac12x^{2}+\tfrac38x^{4}-\cdots:

(1+x22)(1x22+3x48)=1x22+3x48+x22x44+O(x6)=1+x48+O(x6). \left(1+\frac{x^{2}}{2}\right)\left(1-\frac{x^{2}}{2}+\frac{3x^{4}}{8}\right) = 1 - \frac{x^{2}}{2} + \frac{3x^{4}}{8} + \frac{x^{2}}{2} - \frac{x^{4}}{4} + O(x^{6}) = 1 + \frac{x^{4}}{8}+O(x^{6}).

The x2x^{2} terms cancel exactly. That cancellation is not luck. The Newtonian formula was constructed to reproduce the first two terms of (2.5.24), so the leading error is necessarily the third. Hence εx4/8\varepsilon\approx x^{4}/8, which is quartic rather than quadratic, and a quartic is very small for a while and then is not.

Numbers. Setting ε=0.01\varepsilon=0.01 exactly, numerically rather than from the expansion, gives x=0.5715x=0.5715, whence v/c=x/1+x2=0.4962v/c=x/\sqrt{1+x^{2}}=0.4962 and γ=1.152\gamma=1.152. Some more, exact:

p/mcp/mcv/cv/cE/mc2E/mc^{2}Newtonianerror
0.50.50.4470.4471.11801.11801.12501.12500.62%0.62\%
110.7070.7071.41421.41421.50001.50006.07%6.07\%
220.8940.8942.23612.23613.00003.000034.2%34.2\%
330.9490.9493.16233.16235.50005.500073.9%73.9\%
550.9810.9815.09905.099013.50013.500165%165\%

This is the quantitative answer to "why didn't anyone notice?". Everything in nineteenth-century mechanics ran at β<104\beta\lt10^{-4}, where ε<1017\varepsilon\lt10^{-17}. You do not find relativity by doing careful mechanics. You find it by doing electromagnetism, which is Chapter 2.1's story, or by building something that goes fast.

⚠ Why this isn't obvious — nothing here forbids v>cv\gt c by decree

Look back over §§1–4 and notice what is not in them. There is no axiom saying "speeds cannot exceed cc". The prohibition is a consequence, and it arrives twice, by different routes.

Causally (Chapter 2.3 §4.5). A signal travelling faster than cc can be boosted into a frame where it arrives before it left, and combining two such signals builds a closed causal loop. That argument came from the geometry alone and applies to influences, not just to objects.

Energetically (this chapter). From (2.5.26), pushing a massive particle to vcv\to c requires γ\gamma\to\infty and hence EE\to\infty. The barrier is not a wall. It is a bill. And (2.5.32) shows the same thing from the other side: no finite EE and pp satisfying (2.5.31) with m>0m\gt0 can give pc2/E=cpc^{2}/E=c.

The two arguments are independent and they agree, which is the sort of over-determination that makes a structure believable.

Note also what neither of them forbids. A hypothetical particle with pp<0p\cdot p\lt0, meaning a spacelike four-momentum and an "imaginary mass", would satisfy v>cv\gt c happily. Nothing in the algebra excludes it. What excludes it is causality, together with the fact that in quantum field theory such a field signals an unstable vacuum rather than a fast particle. Chapter 6.6 makes that precise, because the Higgs field before symmetry breaking is exactly such a case, and the resolution is not that anything travels faster than light.

In plain terms 2.5.4

Contract the welded energy and momentum with itself and the mass falls out as a number nobody disputes. Observers assign the same electron different energies and different momenta and all compute the same mass, which makes mass a label the object carries rather than something growing with speed, exactly as the interval is the label a pair of events carries. Plot energy against momentum and the hyperbolae drawn two chapters ago for time and position reappear, since a change of frame slides any such object along its own invariant square.

One consequence deserves extracting first, because it makes the massless case work: the velocity is the ratio of momentum to energy, with no dilation factor and no mass in it anywhere. That recipe does not care whether a mass exists. Set the mass to zero and it returns the speed limit exactly, at every energy, with no approximation.

Reading the massless case as the end of a sequence of ever lighter particles is the expensive mistake to avoid. The construction that built momentum from a mass and a carried clock does not survive that limit, since a light ray has no carried clock to divide by, and the phrase naming the time a light ray experiences names nothing. The repair is to invert the order: take the four-part momentum as primary and let mass be the label it carries.

5 · The action is proper time

We have the repaired momentum p=γmv\vv p=\gamma m\vv v, and we expect the equation of motion to read dp/dt=F\dd\vv p/\dd t=\vv F. That is Newton's second law in the form he actually wrote it, with the repaired momentum substituted for the old one. (Section 6 shows this is the spatial part of a proper tensor equation. For now, take it as the target.)

Chapter 1.2 established that respectable dynamics comes from an action principle. So the question in front of us is this: what Lagrangian produces this?

We are not going to guess. We impose the answer and solve for LL. That is the honest way round, and it is exactly how Chapter 1.2 §4 checked that L=TVL=T-V reproduces Newton, run backwards.

5.1 · Solving for LL

Take a particle in a potential, L=L0(v)V(x)L=L_{0}(\vv v)-V(\vv x), with the free part depending only on the velocity. We need the Euler–Lagrange equation for the coordinate xix^{i}, which gives one equation per degree of freedom, ddt(L/q˙i)=L/qi\dv{}{t}\big(\partial L/\partial\dot q^{i}\big)=\partial L/\partial q^{i}, from Chapter 1.2 §3.5. Here it reads

ddt(L0vi)  =  Lxi  =  Vxi  =  Fi. \dv{}{t}\left(\pdv{L_{0}}{v^{i}}\right) \;=\; \pdv{L}{x^{i}} \;=\; -\pdv{V}{x^{i}} \;=\; F_{i}. (2.5.35)

We want the left-hand side to be dpi/dt\dd p_{i}/\dd t with pi=γmvip_{i}=\gamma m v^{i}. Matching the two demands

L0vi  =  γmvi  =  mvi1v2/c2. \pdv{L_{0}}{v^{i}} \;=\; \gamma m\,v^{i} \;=\; \frac{m\,v^{i}}{\sqrt{1-v^{2}/c^{2}}}. (2.5.36)

Now solve this for L0L_{0}. The right-hand side is viv^{i} times a function of the speed alone, so L0L_{0} can depend on v\vv v only through v=vv=\abs{\vv v}. The chain rule then gives L0/vi=L0(v)v/vi=L0(v)vi/v\partial L_{0}/\partial v^{i} = L_{0}'(v)\,\partial v/\partial v^{i} = L_{0}'(v)\,v^{i}/v. Comparing with (2.5.36), the direction viv^{i} cancels and we are left with a single ordinary differential equation:

L0(v)  =  mv1v2/c2. L_{0}'(v) \;=\; \frac{m\,v}{\sqrt{1-v^{2}/c^{2}}}. (2.5.37)

All that remains is to integrate it. The substitution that clears the square root is s=1v2/c2s=1-v^{2}/c^{2}, for which ds=2vdv/c2\dd s = -2v\,\dd v/c^{2} and therefore vdv=c22dsv\,\dd v = -\tfrac{c^{2}}{2}\dd s:

L0  =  mvdvs  =  mc22s1/2ds=  mc222s+const  =  mc21v2c2+const. \begin{aligned} L_{0} \;&=\; \int\frac{m\,v\,\dd v}{\sqrt{s}} \;=\; -\frac{mc^{2}}{2}\int s^{-1/2}\,\dd s\\[4pt] &=\; -\frac{mc^{2}}{2}\cdot 2\sqrt{s} + \text{const} \;=\; -mc^{2}\sqrt{1-\frac{v^{2}}{c^{2}}} + \text{const}. \end{aligned} (2.5.38)

The constant is genuinely free. Chapter 1.2's Problem 4 showed that adding a constant to LL adds a term linear in tt to SS, which is a total derivative and changes no equation of motion. Set it to zero, for a reason that will be apparent in ten lines:

  L  =  mc21β2    V(x)  =  mc2γV.   \boxed{\;L \;=\; -mc^{2}\sqrt{1-\beta^{2}} \;-\; V(\vv x) \;=\; -\frac{mc^{2}}{\gamma} - V.\;} (2.5.39)

Check the limit before going on. Expanding the square root with the binomial series, 1β2=112β218β4\sqrt{1-\beta^{2}} = 1-\tfrac12\beta^{2}-\tfrac18\beta^{4}-\cdots, so

L  =  mc2+12mv2+mv48c2+V, L \;=\; -mc^{2} + \half mv^{2} + \frac{mv^{4}}{8c^{2}} + \cdots - V, (2.5.40)

which is Chapter 1.2's TVT-V plus the irrelevant constant mc2-mc^{2} plus small corrections ✓.

Note in passing that LL is emphatically not TVT-V any more. The quantity mc2/γ-mc^{2}/\gamma is not the kinetic energy of anything. Chapter 1.2 §5 warned that "TVT-V" is a special case rather than a definition, and here is the first place that warning cashes out.

5.2 · The rewrite that makes it obvious

Now form the action and use (2.5.5) in the direction we have not yet used it. Until now we have converted τ\tau-derivatives into tt-derivatives. This time we go the other way and turn dt\dd t into dτ\dd\tau. For a free particle (V=0V=0),

S  =  t1t2Ldt  =  mc2 ⁣t1t21β2 dt  =    mc2 ⁣dτ    =  mc2τ[path]. S \;=\; \int_{t_{1}}^{t_{2}} L\,\dd t \;=\; -mc^{2}\!\int_{t_{1}}^{t_{2}}\sqrt{1-\beta^{2}}\ \dd t \;=\; \boxed{\;-mc^{2}\!\int\dd\tau\;} \;=\; -mc^{2}\,\tau[\text{path}]. (2.5.41)

Stop here. This is the most important equation in the chapter after (2.5.31), and it deserves to be read as a sentence rather than a formula.

The free relativistic action is the proper time

Up to a constant mc2-mc^{2}, the action of a free particle is the reading of the clock it carries. Not a quantity resembling it, and not a quantity proportional to it in some limit. It is that number.

So the two principles this book has stated separately are one principle. Chapter 1.2 said that nature makes SS stationary. Chapter 2.3 §6 proved that among all worldlines joining two events, the straight one has the greatest proper time.

Put those side by side and the minus sign does its job. Maximising τ\tau is minimising mc2τ-mc^{2}\tau, because mc2>0mc^{2}\gt0. The minus sign in front of mc2mc^{2} is not a convention chosen for tidiness. It is what converts 2.3's maximum into 1.2's stationary point.

Chapter 1.2's grind box on the second variation already checked the Legendre condition for exactly this Lagrangian and found 2L/x˙2=mγ3>0\partial^{2}L/\partial\dot x^{2} = m\gamma^{3} \gt0, confirming that the stationary point is a genuine minimum of SS.

Two loops close at once. First, the free particle's trajectory is the longest worldline, which is why the travelling twin comes back younger. She took a shorter path in the only sense of "length" spacetime recognises. Second, this is entry two in Chapter 1.2 §8.1's table of actions, quoted forward there and derived here. Whatever it is worth, the derivation cost eleven lines.

And the forward view. Chapter 3.3 keeps (2.5.41) verbatim and changes only how dτ\dd\tau is computed, replacing ημν\eta_{\mu\nu} by a position-dependent metric gμν(x)g_{\mu\nu}(x). Extremising the same functional then gives the geodesic equation, and general relativity's statement that free particles fall along geodesics is this statement, with a different ruler. Nothing about the principle changes at all.

Grind box — three independent checks on LL, including Noether's

Check 1: the canonical momentum. Chapter 1.2 §6 defines the momentum conjugate to xix^{i} as L/x˙i\partial L/\partial\dot x^{i}. From (2.5.39), with x˙i=vi\dot x^{i}=v^{i},

Lx˙i=mc212(1β2)1/2(2vic2)=mvi1β2=γmvi=pi.  \pdv{L}{\dot x^{i}} = -mc^{2}\cdot\frac{1}{2}\big(1-\beta^{2}\big)^{-1/2}\cdot\left(-\frac{2v^{i}}{c^{2}}\right) = \frac{m v^{i}}{\sqrt{1-\beta^{2}}} = \gamma m v^{i} = p_{i}. \ \checkmark

Which it had better, since (2.5.36) is what we solved. But it is reassuring that the canonical momentum of Lagrangian mechanics and the spatial part of the four-momentum are literally the same object. They did not have to be.

Check 2: Noether's energy. Chapter 1.4 §3 showed that if LL has no explicit time dependence, the conserved Noether charge is H=ix˙iL/x˙iLH=\sum_{i}\dot x^{i}\,\partial L/\partial\dot x^{i} - L. Compute it for the free case:

H=v(γmv)+mc21β2=γmv2+mc2γ=γm(v2+c2γ2)  =  γm(v2+c2v2)  =  γmc2. \begin{aligned} H &= \vv v\cdot(\gamma m\vv v) + mc^{2}\sqrt{1-\beta^{2}}\\[3pt] &= \gamma m v^{2} + \frac{mc^{2}}{\gamma}\\[3pt] &= \gamma m\left(v^{2} + \frac{c^{2}}{\gamma^{2}}\right) \;=\; \gamma m\big(v^{2} + c^{2}-v^{2}\big) \;=\; \gamma mc^{2}. \end{aligned}

where the third line used c2/γ2=c2(1β2)=c2v2c^{2}/\gamma^{2}=c^{2}(1-\beta^{2})=c^{2}-v^{2}. So the Noether charge of time-translation invariance is exactly E=γmc2E=\gamma mc^{2}, the quantity §3 identified by an entirely different argument, namely matching a Taylor expansion to Newtonian kinetic energy. Two independent derivations, one answer. ✓

Check 3: the equation of motion, in full. Confirm that (2.5.39) really does give the promised law and not something that merely resembles it. With V=0V=0, Euler–Lagrange says d(γmvi)/dt=0\dd(\gamma mv^{i})/\dd t = 0, so γmv\gamma m\vv v is constant, so v\vv v is constant, since γ\gamma is a monotonic function of vv and fixing γmv\gamma m v therefore fixes vv. The path is a straight line traversed uniformly. Free particles move in straight lines at constant speed, which is Newton's first law, recovered from an action whose Lagrangian is nothing like 12mv2\half mv^{2}. ✓

An aside on manifest covariance. (2.5.41) is manifestly a scalar, since mm is a scalar and cc is a scalar and τ\tau is a scalar. So the action is frame-independent by inspection, in the sense of Chapter 2.4 §6. The Lagrangian (2.5.39) is not, and neither is dt\dd t. Only their product is. This is the standard situation in relativistic mechanics, and it is why the action formulation is preferred. The object with the good transformation properties is the integral, not the integrand-in-tt.

In plain terms 2.5.5

Guessing which number to attach to each history would be a poor way to proceed, so the calculation runs backwards: demand that the machinery of the first part return the repaired equation of motion, and solve for it. What comes back for a free particle, once a harmless constant is set aside, is startling in its plainness. The number attached to a history is the reading of the clock carried along it, multiplied by the mass and the square of the speed limit, with a minus sign in front.

Two statements this book has made separately are therefore one statement. The first part said nature selects the history at which the number stops changing; the chapter before last proved that among all histories joining two events the unaccelerated one carries the most time. The minus sign reconciles them, since making a quantity largest is making its negative smallest.

Something else is settled on the way past. The number being attached is no longer kinetic energy less potential energy, resembling that combination in nothing but its low-speed limit, so the warning issued when the action principle first arrived, that the familiar difference is a special case and not a definition, collects here. Notice too which object is well behaved: neither the integrand nor the element of time multiplying it is the same for every observer, and only their product is.

6 · Force, and the death of relativistic mass

Newton's second law is the one piece of the old mechanics we have not yet replaced. Chapter 2.4 §6.2 already announced the verdict, which is that F=ma\vv F=m\vv a is not a tensor equation while dpμ/dτ=fμ\dd p^{\mu}/\dd\tau=f^{\mu} is. Announcing is not deriving. Here is the object, and here is what it costs.

6.1 · The covariant force

Define the four-force by the only construction available: differentiate the four-momentum by the invariant.

fμ    dpμdτ  =  mduμdτ  =  maμ, f^{\mu} \;\equiv\; \dv{p^{\mu}}{\tau} \;=\; m\,\dv{u^{\mu}}{\tau} \;=\; m\,a^{\mu}, (2.5.42)

where the last step holds for particles of constant rest mass. Both sides are (1,0)(1,0) tensors, so by Chapter 2.4 §6 this is a legitimate law. Written once, it holds in every frame.

It also inherits a constraint at no cost. We know ua=0u\cdot a=0 from (2.5.13), so contract the definition with uu and let the mass ride along:

fu  =  m(au)  =  0. f\cdot u \;=\; m\,(a\cdot u) \;=\; 0. (2.5.43)

So a four-force has only three independent components. Whatever three you specify, the fourth is determined. That is not a defect. It is a theorem with a name you already know, as we are about to see.

6.2 · The three-force, and the work–energy theorem for free

Connect to the laboratory. Define the ordinary three-force as the rate of change of the ordinary momentum with respect to ordinary time, Fdp/dt\vv F\equiv\dd\vv p/\dd t. That is Newton's second law in its correct form, the form he actually wrote. Using (2.5.6) on (2.5.42), component by component:

fμ  =  γddt(Ec, p)  =  γ(1cdEdt,  F). f^{\mu} \;=\; \gamma\,\dv{}{t}\left(\frac{E}{c},\ \vv p\right) \;=\; \gamma\left(\frac{1}{c}\dv{E}{t},\ \ \vv F\right). (2.5.44)

Now we cash in the constraint. We have fμf^{\mu} in laboratory quantities, and (2.5.43) says its contraction with uμu^{\mu} vanishes. So write out that contraction, using uμ=γ(c,v)u^{\mu}=\gamma(c,\vv v), and see what it forces:

0  =  fu  =  γ(1cdEdt)(γc)γFγv  =  γ2(dEdtFv), 0 \;=\; f\cdot u \;=\; \gamma\left(\frac1c\dv{E}{t}\right)(\gamma c) - \gamma\vv F\cdot\gamma\vv v \;=\; \gamma^{2}\left(\dv{E}{t} - \vv F\cdot\vv v\right), (2.5.45)

and since γ20\gamma^{2}\neq0, the bracket itself must vanish, which leaves

  dEdt  =  Fv.   \boxed{\;\dv{E}{t} \;=\; \vv F\cdot\vv v.\;} (2.5.46)

That is the work–energy theorem. The rate of change of energy is the power delivered by the force. We did not assume it, define it, or import it. It is the time component of a four-vector equation, forced by the geometric identity ua=0u\cdot a=0.

In Newtonian mechanics the work–energy theorem is a separate derivation. Here it is the statement that the four-force is orthogonal to the four-velocity, which is the statement that the four-velocity has fixed length, which is the definition of proper time. Everything is the same thing.

6.3 · F\vv F and a\vv a are not parallel

Now the damage. Expand F=d(γmv)/dt\vv F=\dd(\gamma m\vv v)/\dd t with the product rule, using γ˙=γ3(va)/c2\dot\gamma=\gamma^{3}(\vv v\cdot\vv a)/c^{2} from §1.5's grind box:

F  =  ddt(γmv)  =  γma+γ˙mv  =  γm[a+γ2c2(va)v]. \vv F \;=\; \dv{}{t}\big(\gamma m\vv v\big) \;=\; \gamma m\,\vv a + \dot\gamma\,m\vv v \;=\; \gamma m\left[\vv a + \frac{\gamma^{2}}{c^{2}}\big(\vv v\cdot\vv a\big)\,\vv v\right]. (2.5.47)

The bracket contains a piece along a\vv a and a piece along v\vv v. Unless those two directions coincide, or unless the second term vanishes, the force and the acceleration point in different directions.

Push a fast particle sideways and it does not accelerate sideways. It accelerates partly forward, or partly backward, depending on the sign of va\vv v\cdot\vv a. Nothing in Newtonian mechanics prepares you for this, and it is not a small effect.

The two special cases where they are parallel are worth having explicitly.

Longitudinal, meaning a push along the motion. Then av\vv a\parallel\vv v, so va=va\vv v\cdot\vv a=va and the bracket is a(1+γ2β2)v^a\big(1+\gamma^{2}\beta^{2}\big)\hat v. Use the identity γ2γ2β2=γ2(1β2)=1\gamma^{2}-\gamma^{2}\beta^{2}=\gamma^{2}(1-\beta^{2})=1, that is 1+γ2β2=γ21+\gamma^{2}\beta^{2}=\gamma^{2}, and the whole factor collapses:

F  =  γmaγ2  =  γ3ma. F_{\parallel} \;=\; \gamma m\,a_{\parallel}\cdot\gamma^{2} \;=\; \gamma^{3}m\,a_{\parallel}. (2.5.48)

Transverse, meaning a push across the motion. Then va=0\vv v\cdot\vv a=0, the second term dies, and what is left is

F  =  γma. F_{\perp} \;=\; \gamma m\,a_{\perp}. (2.5.49)

Two different laws, differing by γ2\gamma^{2}, for the same particle at the same instant. At γ=10\gamma=10 a longitudinal push is a hundred times less effective at producing acceleration than a transverse one of the same magnitude.

This is not an exotic regime. It is routine at any electron accelerator, and it is why the transverse focusing of a beam and its longitudinal acceleration are engineered as separate problems with separate hardware.

Grind box — inverting the relation: a\vv a in terms of F\vv F

(2.5.47) gives F\vv F from a\vv a. For solving actual problems you want the other direction, and inverting it is a small exercise in projection that is worth doing once, because the same trick appears in Chapter 3.3.

Step 1, get the component along v\vv v. Dot (2.5.47) with v\vv v:

Fv=γm[va+γ2c2(va)v2]=γm(va)[1+γ2β2]=γ3m(va), \vv F\cdot\vv v = \gamma m\left[\vv v\cdot\vv a + \frac{\gamma^{2}}{c^{2}}(\vv v\cdot\vv a)v^{2}\right] = \gamma m\,(\vv v\cdot\vv a)\left[1+\gamma^{2}\beta^{2}\right] = \gamma^{3}m\,(\vv v\cdot\vv a),

using 1+γ2β2=γ21+\gamma^{2}\beta^{2}=\gamma^{2} again. So

va=Fvγ3m. \vv v\cdot\vv a = \frac{\vv F\cdot\vv v}{\gamma^{3}m}.

Step 2, substitute back. Put that into (2.5.47) and solve for a\vv a:

F=γma+γmγ2c2Fvγ3mv=γma+(Fv)c2v, \vv F = \gamma m\,\vv a + \gamma m\cdot\frac{\gamma^{2}}{c^{2}}\cdot\frac{\vv F\cdot\vv v}{\gamma^{3}m}\,\vv v = \gamma m\,\vv a + \frac{(\vv F\cdot\vv v)}{c^{2}}\,\vv v,

hence

  a=1γm[F(Fv)c2v].   \boxed{\;\vv a = \frac{1}{\gamma m}\left[\vv F - \frac{(\vv F\cdot\vv v)}{c^{2}}\,\vv v\right].\;}

Check it. For Fv\vv F\parallel\vv v the bracket is F(1β2)=F/γ2\vv F(1-\beta^{2})=\vv F/\gamma^{2}, so a=F/γ3ma=F/\gamma^{3}m ✓ matching (2.5.48). For Fv\vv F\perp\vv v the second term vanishes, so a=F/γma=F/\gamma m ✓ matching (2.5.49).

What the formula says. The bracket removes from F\vv F the piece (Fv^)β2v^(\vv F\cdot\hat v)\beta^{2}\hat v, a projection along the motion weighted by β2\beta^{2}. So the acceleration is a distorted image of the force, squashed along the direction of travel.

At β1\beta\to1 the longitudinal response is suppressed entirely. You can push a nearly-light-speed particle as hard as you like along its motion and it barely speeds up, even though its energy rises steadily at the rate (2.5.46) says it must. The energy has to go somewhere, and it goes into γ\gamma rather than into vv.

One consequence worth naming. Set F\vv F constant, along xx, starting from rest. Then d(γmv)/dt=F\dd(\gamma mv)/\dd t=F integrates immediately to γmv=Ft\gamma mv=Ft, so

v(t)=Ft/m1+(Ft/mc)2    cas t, v(t) = \frac{Ft/m}{\sqrt{1+(Ft/mc)^{2}}} \;\longrightarrow\; c \quad\text{as } t\to\infty,

approaching but never reaching cc, whereas the Newtonian v=Ft/mv=Ft/m sails past it at t=mc/Ft=mc/F. Same constant force, same duration, entirely different destination. And the relativistic answer required no new physical assumption beyond p=γmv\vv p=\gamma m\vv v.

6.4 · Why this book will not say "relativistic mass"

There is a tempting way to make (2.5.48) and (2.5.49) look like F=ma\vv F=m\vv a, which is to absorb the γ\gamma's into the mass. Define mrel=γmm_{\text{rel}}=\gamma m and you can write p=mrelv\vv p=m_{\text{rel}}\vv v and E=mrelc2E=m_{\text{rel}}c^{2}, and the formulas look Newtonian again. Textbooks did this for sixty years. It is a mistake, and the calculation above is the cleanest way to see why.

It would have to be two different numbers at once. To keep F=ma\vv F=m\vv a you need m=γ3mm_{\parallel}=\gamma^{3}m for a longitudinal push and m=γmm_{\perp}=\gamma m for a transverse one. A "mass" that depends on which way you shove is not a property of the particle. It is a badly chosen name for the components of a tensor relation.

And for a general push, by (2.5.47), no scalar whatsoever works, because F\vv F and a\vv a are not even parallel. You would need a matrix, and calling a matrix "the mass" abandons the only thing the word was supposed to mean.

It makes E=mc2E=mc^{2} empty. With mrelE/c2m_{\text{rel}}\equiv E/c^{2}, the celebrated equation E=mrelc2E=m_{\text{rel}}c^{2} says E=EE=E. All the physical content lives in the rest mass and is invisible in the relativistic one: that a particle at rest still has energy mc2mc^{2}, that binding changes rest mass, and that mass is not additive.

It hides the invariant. The real structure is (2.5.31). Here mm is what you get by contracting pμp^{\mu} with itself, a Lorentz scalar, the same for all observers, in exactly the way the interval is the same for all observers.

Redefining "mass" to mean E/c2E/c^{2} replaces an invariant with a component and throws away the geometry. It is the momentum-space version of saying that a metre stick "really" gets shorter when it moves.

So mm means rest mass, always, everywhere in this book, exactly as the conventions for Part II state. Energy is E=γmc2E=\gamma mc^{2} and momentum is p=γmv\vv p=\gamma m\vv v, and the γ\gamma's stay where they are, visible, attached to the motion rather than smuggled into the particle. The phrase "relativistic mass" does not appear again.

In plain terms 2.5.6

The second law costs more to replace than it appears to. Differentiate the four-part momentum by the carried clock and you have a force with the right credentials, carrying a constraint nobody imposed: since the four-velocity has fixed length, the force stands perpendicular to it, so only three of its entries are free. Write out the entry that is not free and it says energy changes at the rate the ordinary force does work, which in Newtonian mechanics was a separate derivation and here is the fixed length restated.

Then the damage. Force and acceleration stop pointing the same way. Push a fast particle along its motion and it responds far more sluggishly than to an equal push across it: at a dilation factor of ten, a hundred times less so. That is why transverse focusing of a beam and forward acceleration are engineered as separate problems with separate hardware.

A habit sixty years of textbooks kept dies with it. Absorbing the dilation factors into the mass makes the formulas look Newtonian again, at the price of the mass being two numbers at once, one for a push along the motion and one across. For any other direction no single number works, the two vectors not being parallel, so a matrix is needed, and calling a matrix the mass abandons what the word was for. The phrase will not appear again.

7 · Light: the wave four-vector, and Doppler

Everything so far has been about particles. Waves need one more four-vector, and it comes from an argument so simple it is easy to undervalue.

7.1 · Why kμk^{\mu} has to be a four-vector

A plane wave has a phase

Φ(t,x)  =  kxωt, \Phi(t,\vv x) \;=\; \vv k\cdot\vv x - \omega t, (2.5.50)

with k\vv k the wave vector (k=2π/λ\abs{\vv k}=2\pi/\lambda, pointing along propagation) and ω\omega the angular frequency. Here is the physical claim, and it is the whole argument:

Counting crests is frame-independent

A crest is not a thing that moves. It is a set of events at which the field is momentarily maximal. So "is this event a crest?" is a yes-or-no question about a single point of spacetime, and both observers are being asked about the same point. They must give the same answer.

Now count. Put a detector at a fixed place and let it click once per crest between two events AA and BB on its own worldline. The number of clicks is an integer.

An observer flying past at 0.9c0.9c watches the same detector and counts the same integer. Each click is an event on one worldline, and 2.3 §4.3 proved that the ordering of events on a single timelike worldline is invariant. Nobody can gain or lose a click by changing frames.

That count is ΔΦ/2π\Delta\Phi/2\pi. An integer cannot transform, and AA and BB were arbitrary, so Φ\Phi is a Lorentz scalar.

Now use the scalar to manufacture the four-vector. We want (2.5.50) written as a contraction, since a contraction is the shape the quotient theorem knows how to act on. With xμ=(ct,x,y,z)x^{\mu}=(ct,x,y,z) and the metric lowering the spatial indices,

define  kμ(ωc, k);thenkμxμ=k0(ct)kx  =  ωtkx  =  Φ. \begin{aligned} \text{define }\ k^{\mu} &\equiv \left(\frac{\omega}{c},\ \vv k\right); \qquad\text{then}\\[4pt] k_{\mu}x^{\mu} &= k^{0}(ct) - \vv k\cdot\vv x \;=\; \omega t - \vv k\cdot\vv x \;=\; -\Phi. \end{aligned} (2.5.51)

So Φ=kμxμ\Phi=-k_{\mu}x^{\mu}. We know xμx^{\mu} is a four-vector, and we have just argued that Φ\Phi is a scalar for every event xμx^{\mu}, not merely for one.

That is precisely the hypothesis of the quotient theorem (Chapter 2.4 §5.1's grind box). If kμAμk_{\mu}A^{\mu} is a scalar for every four-vector AμA^{\mu}, then kμk_{\mu} is a covector, and raising with η\eta makes kμk^{\mu} a four-vector. Hence

  kμ  =  (ωc, k)  is a four-vector.   \boxed{\;k^{\mu} \;=\; \left(\frac{\omega}{c},\ \vv k\right)\ \text{ is a four-vector.}\;} (2.5.52)

For light in vacuum the dispersion relation is ω=ck\omega=c\abs{\vv k}, which Chapter 2.1 §2 derived from Maxwell. That says precisely kk=ω2/c2k2=0k\cdot k = \omega^{2}/c^{2}-\abs{\vv k}^{2}=0, so the wave four-vector of light is null. Set that beside pp=0p\cdot p=0 from §4.3 and the shape of Chapter 4.1's punchline is already visible. We come back to it below.

7.2 · Longitudinal Doppler

Everything about the Doppler effect is now one matrix multiplication. Let a light wave travel in the +x+x direction in frame SS, so

kμ  =  ωc(1, 1, 0, 0), k^{\mu} \;=\; \frac{\omega}{c}\big(1,\ 1,\ 0,\ 0\big), (2.5.53)

the spatial part having magnitude ω/c\omega/c as the null condition requires. Now boost to SS' moving at +βc+\beta c along xx, which is an observer running away from the source, in the same direction the light is going. Chapter 2.4 §2.1's boost matrix acts on any four-vector, so it acts on this one:

k0=γ(k0βk1)=γωc(1β),k1=γ(k1βk0)=γωc(1β). \begin{aligned} k'^{0} &= \gamma\big(k^{0} - \beta k^{1}\big) = \gamma\frac{\omega}{c}\big(1-\beta\big),\\[3pt] k'^{1} &= \gamma\big(k^{1} - \beta k^{0}\big) = \gamma\frac{\omega}{c}\big(1-\beta\big). \end{aligned} (2.5.54)

Both components pick up the same factor, as they must, since kμk'^{\mu} has to be null too. The frequency we want is ω=ck0\omega'=ck'^{0}, so read that off and simplify with γ=1/(1β)(1+β)\gamma=1/\sqrt{(1-\beta)(1+\beta)}:

ω  =  1β(1β)(1+β)  ω  =    1β1+β ω   \omega' \;=\; \frac{1-\beta}{\sqrt{(1-\beta)(1+\beta)}}\;\omega \;=\; \boxed{\;\sqrt{\frac{1-\beta}{1+\beta}}\ \omega\;} (2.5.55)

For a receding observer, β>0\beta\gt0 and ω<ω\omega'\lt\omega, which is a redshift. For an approaching one, replace ββ\beta\to-\beta and get a blueshift. At β1\beta\ll1 this reduces to ωω(1β)\omega'\approx\omega(1-\beta), the classical result, so nothing familiar is lost. What is new is the exactness, and the second-order term, which is where the physics is.

7.3 · Transverse Doppler: time dilation, seen directly

Now a case with no classical counterpart at all. Let a source move at +βc+\beta c along xx past a detector, and consider the light that reaches the detector travelling in the +y+y direction in the detector's frame. That is the light emitted, as the lab sees it, at the moment of closest approach, so that the source's velocity is entirely perpendicular to the line of sight. Classically there is no Doppler shift whatsoever in that configuration, because the distance is momentarily not changing.

In the lab frame SS, the received wave has

kμ  =  ωc(1, 0, 1, 0). k^{\mu} \;=\; \frac{\omega}{c}\big(1,\ 0,\ 1,\ 0\big). (2.5.56)

What we want to compare it with is the frequency the source emits, so transform to the source's rest frame SS', which moves at +βc+\beta c along xx. Only the time component is needed:

k0  =  γ(k0βk1)  =  γωc0  =  γωc. k'^{0} \;=\; \gamma\big(k^{0}-\beta k^{1}\big) \;=\; \gamma\frac{\omega}{c} - 0 \;=\; \gamma\,\frac{\omega}{c}. (2.5.57)

But ck0ck'^{0} is the frequency in the source's own frame, which is the frequency the source actually emits. Call it ω0\omega_{0}, a property of the atom. Then ω0=γω\omega_{0}=\gamma\omega, which rearranges to

  ω  =  ω0γ.   \boxed{\;\omega \;=\; \frac{\omega_{0}}{\gamma}.\;} (2.5.58)

A redshift, of exactly the factor γ\gamma, in a geometry where the classical prediction is precisely nothing. And its content is unmistakable. The factor 1/γ1/\gamma is the time-dilation factor of Chapter 2.2. The source is a clock, the clock is moving, so it runs slow, and you see its emission slowed.

The transverse Doppler shift is time dilation with no admixture of anything else, which is exactly why it was worth measuring.

Grind box — the general Doppler formula, aberration, and what Ives and Stilwell actually measured

Arbitrary angle. Let the light travel at angle θ\theta to the xx-axis in SS, so kμ=(ω/c)(1,cosθ,sinθ,0)k^{\mu}=(\omega/c)(1,\cos\theta,\sin\theta,0). Boosting along xx by β\beta:

ω=γω(1βcosθ),k1=γωc(cosθβ),k2=ωcsinθ. \begin{aligned} \omega' &= \gamma\omega\big(1-\beta\cos\theta\big),\\[2pt] k'^{1} &= \gamma\frac{\omega}{c}\big(\cos\theta-\beta\big), \qquad k'^{2} = \frac{\omega}{c}\sin\theta. \end{aligned}

The first line is the general Doppler formula. Check the special cases. Setting θ=0\theta=0 gives ω=γω(1β)\omega'=\gamma\omega(1-\beta), which is (2.5.55) ✓. Setting θ=π/2\theta=\pi/2 gives ω=γω\omega'=\gamma\omega, which is (2.5.58) ✓, with ω=ω0\omega'=\omega_{0}.

Aberration. Divide the second line by the first to get the direction in SS':

cosθ=k1k0=cosθβ1βcosθ, \cos\theta' = \frac{k'^{1}}{k'^{0}} = \frac{\cos\theta-\beta}{1-\beta\cos\theta},

which is the relativistic aberration of light. It is the same formula that turns the rain on your windscreen into a forward-slanting streak, but exact, and with the crucial difference that the aberration of light does not depend on any medium. This is Problem 3, where it becomes the headlight effect.

⚑ Ives and Stilwell, 1938: quoted experiment, derived analysis. The transverse configuration is hard to arrange, because "transverse" must be specified in a definite frame and a tiny angular error contaminates the result with a first-order longitudinal shift, which is γ1β2/2\gamma-1\approx\beta^{2}/2 times larger in the relevant sense. Ives and Stilwell got around this by measuring longitudinally in both directions at once. Let a source recede with β\beta and approach with β\beta, so that in wavelengths λ±=λ0(1±β)/(1β)\lambda_{\pm}=\lambda_{0}\sqrt{(1\pm\beta)/(1\mp\beta)}. Average them:

λ++λ2=λ02[1+β1β+1β1+β]=λ02(1+β)+(1β)(1β)(1+β)=γλ0. \frac{\lambda_{+}+\lambda_{-}}{2} = \frac{\lambda_{0}}{2}\left[\sqrt{\frac{1+\beta}{1-\beta}}+\sqrt{\frac{1-\beta}{1+\beta}}\right] = \frac{\lambda_{0}}{2}\cdot\frac{(1+\beta)+(1-\beta)}{\sqrt{(1-\beta)(1+\beta)}} = \gamma\,\lambda_{0}.

The first-order shifts cancel exactly and what survives is pure γ\gamma, which is the transverse effect extracted from a longitudinal measurement. Ives and Stilwell ran hydrogen canal rays at β5×103\beta\approx5\times10^{-3}, where γ11.25×105\gamma-1\approx1.25\times10^{-5}, and found the predicted second-order displacement of the mean. It remains one of the cleanest direct tests of time dilation, and the derivation above is the entire theory of the experiment. We quote the experiment, and the algebra is ours.

A four-vector waiting to be identified

Two null four-vectors have now appeared in this chapter for entirely different reasons: pμp^{\mu} for a massless particle (§4.3), and kμk^{\mu} for a light wave (§7.1). Nothing so far connects them. But notice that they have the same transformation law and the same null condition. Compare (2.5.33) with ω=ck\omega=c\abs{\vv k} and they also have the same relation between time and space components.

Chapter 4.1 supplies the missing constant. The relation is

pμ=kμ,i.e.E=ω,p=k, p^{\mu} = \hbar\,k^{\mu}, \qquad\text{i.e.}\qquad E=\hbar\omega,\quad \vv p = \hbar\vv k,

and the reason it must be exactly this, with a single constant and no angle-dependent fudge, is that both sides are four-vectors. One scalar relates them, or none does.

So the relativity of this chapter forces E=ωE=\hbar\omega and p=k\vv p=\hbar\vv k to stand or fall together. You cannot have the Planck relation for energy without the de Broglie relation for momentum. That is a strong structural statement, and it was available years before anyone believed either half of it.

In plain terms 2.5.7

Waves need one more four-part object, and the argument producing it is counting rather than calculation. A crest is not a thing that travels but a set of events at which the field is momentarily largest, so whether a given event is a crest is a question about one point of spacetime that both observers are asked about. Let a detector click once per crest between two events on its own history: the count is a whole number, nobody gains or loses one by moving, and whole numbers cannot transform. So the phase is agreed, and whatever pairs with position to produce it must be four-part too.

Frequency and wavelength are thereby welded together as energy and momentum are, and the whole Doppler effect becomes one multiplication. Light from a source passing at closest approach, where the separation is momentarily unchanging and the old theory predicts no shift, arrives reddened by exactly the dilation factor.

Two null objects have now appeared for unrelated reasons, one for a massless particle and one for a light wave, sharing a transformation law, a null condition and the same relation between time and space parts. One constant relating them relates both halves at once, since a single scalar between two such objects covers everything or nothing. Planck's relation for energy and de Broglie's for momentum stand or fall together, which was available years before anybody believed either.

8 · Collisions: invariant mass, and why colliders exist

The whole practical value of §2 is that it turns collision problems into linear algebra. You write down the four-momenta, you add them, you square whatever combination kills the unknowns, and you are done. This section builds the three tools and then uses them to say something quantitative and slightly outrageous about accelerator design.

8.1 · The invariant mass of a system

Take any collection of particles, whatever they are doing, and add up their four-momenta:

Pμ    ipiμ  =  (iEic,  ipi). P^{\mu} \;\equiv\; \sum_{i} p_{i}^{\mu} \;=\; \left(\frac{\sum_{i}E_{i}}{c},\ \ \sum_{i}\vv p_{i}\right). (2.5.59)

A sum of four-vectors is a four-vector, so PμP^{\mu} transforms properly and has an invariant square. That square is the object we want, so give it a name. Define the system's invariant mass MM by

M2c2    PP  =  (iEi)2c2ipi2. M^{2}c^{2} \;\equiv\; P\cdot P \;=\; \frac{\big(\sum_{i}E_{i}\big)^{2}}{c^{2}} - \abs{\textstyle\sum_{i}\vv p_{i}}^{2}. (2.5.60)

Everyone agrees on MM. And the system's total four-momentum is conserved, by §2, so MM is conserved as well. That is a strong constraint on what a system may turn into.

8.2 · The centre-of-momentum frame

Suppose PμP^{\mu} is timelike, PP>0P\cdot P\gt0. Then by exactly the construction of Chapter 2.3 §6.1 there is a frame in which its spatial part vanishes. Boost with βcm=P/P0=cipi/iEi\vv\beta_{\text{cm}}=\vv P\big/P^{0} = c\sum_{i}\vv p_{i}\big/\sum_{i}E_{i}, which has magnitude less than 11 precisely because PμP^{\mu} is timelike. In that frame, called the centre-of-momentum frame or CM frame,

PμCM  =  (ECMc, 0),soPP=ECM2c2ECM=Mc2. P^{\mu}\Big|_{\text{CM}} \;=\; \left(\frac{E_{\text{CM}}}{c},\ \vv 0\right), \qquad\text{so}\qquad P\cdot P = \frac{E_{\text{CM}}^{2}}{c^{2}} \quad\Longrightarrow\quad E_{\text{CM}} = Mc^{2}. (2.5.61)

So the invariant mass of a system is its total energy in the frame where it is collectively at rest, divided by c2c^{2}. Notice that this is the direct generalisation of E0=mc2E_{0}=mc^{2} from one particle to many.

Notice also that βcm\vv\beta_{\text{cm}} is exactly (2.5.32) applied to the system's four-momentum, which is a small piece of evidence that we defined things sensibly.

⚠ Why this isn't obvious — mass is not a substance

Take two photons of energy EE flying in opposite directions along xx. Each is massless. Add their four-momenta:

Pμ=Ec(1,1,0,0)+Ec(1,1,0,0)=(2Ec,0,0,0), P^{\mu} = \frac{E}{c}(1,1,0,0) + \frac{E}{c}(1,-1,0,0) = \left(\frac{2E}{c},\,0,\,0,\,0\right),

so PP=4E2/c2>0P\cdot P = 4E^{2}/c^{2} \gt 0 and the pair has invariant mass M=2E/c2M=2E/c^{2}. That is nonzero, and it is built entirely out of massless constituents. There is even a CM frame, and it is the one you are already in, since P=0\vv P=0.

Now aim the same two photons in the same direction. Then Pμ=(2E/c)(1,1,0,0)P^{\mu}=(2E/c)(1,1,0,0), PP=0P\cdot P=0, and M=0M=0. Same two objects, same energies, different invariant mass.

So mass is not stuff that the parts have and the whole inherits. It is a property of the total four-momentum, and specifically of how much the constituent four-momenta point in different directions in spacetime. Randomly directed momentum inside a box shows up as mass of the box.

This is not an analogy and not a special case. It is where most of your own mass comes from, as §9 explains, and it is why "matter is made of massive things" is the wrong picture at the bottom.

8.3 · Mandelstam ss

For a two-body collision the invariant mass of the initial state gets its own name. Define

s    (p1+p2)(p1+p2)  c2,sos  =  ECM. s \;\equiv\; \big(p_{1}+p_{2}\big)\cdot\big(p_{1}+p_{2}\big)\;c^{2}, \qquad\text{so}\qquad \sqrt{s} \;=\; E_{\text{CM}}. (2.5.62)

(The factor of c2c^{2} is bookkeeping so that s\sqrt{s} is an energy in SI units. In the natural units of Part V, where c=1c=1, the definition is simply s=(p1+p2)2s=(p_{1}+p_{2})^{2} and this parenthesis is unnecessary. Chapter 5.9 introduces the companions tt and uu. Here we need only ss.)

Our goal is to see what ss depends on, so expand the product once, using pipi=mi2c2p_{i}\cdot p_{i}=m_{i}^{2}c^{2} on the two square terms:

sc2  =  p1p1+2p1p2+p2p2  =  (m12+m22)c2+2p1p2. \frac{s}{c^{2}} \;=\; p_{1}\cdot p_{1} + 2\,p_{1}\cdot p_{2} + p_{2}\cdot p_{2} \;=\; \big(m_{1}^{2}+m_{2}^{2}\big)c^{2} + 2\,p_{1}\cdot p_{2}. (2.5.63)

Everything now depends on the single cross term, and p1p2p_{1}\cdot p_{2} is an invariant we can evaluate in whichever frame is convenient. Two configurations matter.

Fixed target. Particle 1 has energy EE and hits particle 2 at rest, p2μ=(m2c,0)p_{2}^{\mu}=(m_{2}c,\vv 0). Then p1p2=(E/c)(m2c)p10=m2Ep_{1}\cdot p_{2}=(E/c)(m_{2}c) - \vv p_{1}\cdot\vv 0 = m_{2}E, so

sfixed  =  (m12+m22)c4  +  2m2c2E E  mc2 s    2m2c2E. s_{\text{fixed}} \;=\; \big(m_{1}^{2}+m_{2}^{2}\big)c^{4} \;+\; 2\,m_{2}c^{2}\,E \qquad\xrightarrow[\ E\ \gg\ mc^{2}\ ]{}\qquad \sqrt{s} \;\approx\; \sqrt{2m_{2}c^{2}E}. (2.5.64)

Collider. Two beams of energy EE meet head-on with equal and opposite momenta. Then pi=0\sum\vv p_{i}=0, so you are already in the CM frame and

scollider  =  ECM  =  2E. \sqrt{s_{\text{collider}}} \;=\; E_{\text{CM}} \;=\; 2E. (2.5.65)
The whole argument for building colliders, in one line

Compare the two scalings. Doubling the beam energy of a collider doubles s\sqrt{s}. Doubling the beam energy of a fixed-target machine multiplies s\sqrt{s} by 2\sqrt{2}.

scolliderE,sfixedE. \sqrt{s_{\text{collider}}} \propto E, \qquad\qquad \sqrt{s_{\text{fixed}}} \propto \sqrt{E}.

Useful energy is s\sqrt{s}, not EE. In the fixed-target case the rest of it is spent dragging the centre of mass forward, and is unavailable for making anything. The consequence compounds viciously with energy, and Worked example 2 puts a number on it that is hard to believe until you check it.

8.4 · Threshold energies, and the antiproton

Here is the technique that makes ss worth defining. A reaction can proceed only if the initial state carries enough invariant mass to build the final state.

The minimum case is unambiguous. At threshold the products have no kinetic energy left over in the CM frame, so they are all at rest there and moving together as one. The final invariant mass is then just the sum of the final rest masses. Since ss is conserved and invariant,

s threshold  =  (finalmf)c2. \sqrt{s}\ \Big|_{\text{threshold}} \;=\; \left(\sum_{\text{final}} m_{f}\right)c^{2}. (2.5.66)

Take the reaction that mattered historically, which is producing an antiproton by smashing a proton beam into a hydrogen target,

p+p    p+p+p+pˉ. p + p \;\longrightarrow\; p + p + p + \bar p. (2.5.67)

Charge and baryon number both force the extra proton to accompany the antiproton, since you cannot make pˉ\bar p alone, so four particles of mass mpm_{p} must emerge. (Antiparticles have the same rest mass as their partners. ⚑ We quote that here, and Chapter 5.5 derives it, where it falls out of the Dirac equation.) Now set s=4mpc2\sqrt{s}=4m_{p}c^{2} in (2.5.64) with m1=m2=mpm_{1}=m_{2}=m_{p} and solve for the beam energy:

16mp2c4=2mp2c4+2mpc2EE=14mp2c42mpc2  =  7mpc2, \begin{aligned} 16\,m_{p}^{2}c^{4} &= 2m_{p}^{2}c^{4} + 2m_{p}c^{2}E\\[3pt] \Longrightarrow\qquad E &= \frac{14\,m_{p}^{2}c^{4}}{2m_{p}c^{2}} \;=\; 7\,m_{p}c^{2}, \end{aligned} (2.5.68)

That EE is the projectile's total energy, and what an accelerator is rated by is the kinetic energy it delivers. So subtract the rest energy the proton had before anyone switched the machine on:

T=Empc2=6mpc2=6×938.27 MeV=5.63 GeV. T = E - m_{p}c^{2} = 6\,m_{p}c^{2} = 6\times938.27\ \mathrm{MeV} = 5.63\ \mathrm{GeV}. (2.5.69)

Six proton rest energies to make two particles' worth of new rest mass. The factor of three overhead is exactly the energy locked up in the forward motion of the centre of mass, which (2.5.64) says you cannot spend. ⚑ Historically: the Bevatron at Berkeley was designed to reach 6.2 GeV6.2\ \mathrm{GeV}, comfortably above this threshold and for this reason, and the antiproton was found there in 1955. The number in (2.5.69) is a piece of accelerator engineering that came out of four lines of four-vector algebra.

Grind box — the general threshold formula, and a second worked case

Do the algebra once in general so you never have to repeat it. A projectile of mass m1m_{1} and total energy EE strikes a stationary target of mass m2m_{2} and produces final particles of total rest mass Mfmf\mathcal{M}\equiv\sum_{f}m_{f}. Equate (2.5.64) with (2.5.66):

(m12+m22)c4+2m2c2E  =  M2c4, \big(m_{1}^{2}+m_{2}^{2}\big)c^{4} + 2m_{2}c^{2}E \;=\; \mathcal{M}^{2}c^{4},   Ethreshold  =  [M2m12m22]c22m2.   \boxed{\;E_{\text{threshold}} \;=\; \frac{\big[\mathcal{M}^{2}-m_{1}^{2}-m_{2}^{2}\big]c^{2}}{2m_{2}}.\;}

Check against (2.5.68): M=4mp\mathcal{M}=4m_{p}, m1=m2=mpm_{1}=m_{2}=m_{p} gives E=(1611)mpc2/2=7mpc2E=(16-1-1)m_{p}c^{2}/2=7m_{p}c^{2} ✓.

A second case, pion photoproduction, γ+pp+π0\gamma+p\to p+\pi^{0}. Now m1=0m_{1}=0 for the photon, m2=mp=938.27 MeV/c2m_{2}=m_{p}=938.27\ \mathrm{MeV}/c^{2}, and M=mp+mπ0\mathcal{M}=m_{p}+m_{\pi^{0}} with mπ0c2=134.98 MeVm_{\pi^{0}}c^{2}=134.98\ \mathrm{MeV} (⚑ quoted mass). Then

Eγ=(mp+mπ)2mp22mpc2=2mpmπ+mπ22mpc2=mπc2(1+mπ2mp). E_{\gamma} = \frac{(m_{p}+m_{\pi})^{2}-m_{p}^{2}}{2m_{p}}c^{2} = \frac{2m_{p}m_{\pi}+m_{\pi}^{2}}{2m_{p}}c^{2} = m_{\pi}c^{2}\left(1+\frac{m_{\pi}}{2m_{p}}\right).

Numerically 134.98×(1+0.0719)=144.7 MeV134.98\times(1+0.0719)=144.7\ \mathrm{MeV}. The photon must supply about 7% more than the pion's rest energy, the surplus being the recoil the proton is obliged to take. Note how the formula degrades gracefully. For a very heavy target, m2m_{2}\to\infty and the overhead vanishes, which is why a nucleus makes a better anvil than a single proton.

The same physics, from the CM frame. There is a shortcut worth knowing. The overhead in fixed-target work is the kinetic energy of the centre of mass, which by (2.5.32) moves at βcm=p1c/(E+m2c2)\beta_{\text{cm}}=p_{1}c/(E+m_{2}c^{2}). Everything above can be re-derived by boosting to that frame and demanding the products be at rest, but the invariant route is shorter precisely because it never needs βcm\beta_{\text{cm}} at all. That is the habit to acquire: find the invariant, evaluate it in the easy frame, use it in the hard one.

In plain terms 2.5.8

Add the four-part momenta of any collection of particles, whatever they are doing, and the total is another object of the same kind with an invariant square of its own. That square defines the collection's mass; everybody agrees on it, and since the total is conserved it is conserved too. In the frame where the collection is collectively at rest it is the total energy over the square of the speed limit, which is the single-particle statement about rest energy carried over to many.

What this does to the word mass repays dwelling on. Two pulses of light of equal energy flying apart possess a mass, built entirely from constituents having none, while the same two aimed the same way possess none. Nothing changed but where they pointed. Mass is a property of the total, measuring how far the constituent momenta point in different directions, so randomly directed motion inside a box shows up as mass of the box.

The practical consequence is a piece of accelerator engineering. Only that invariant mass is available for making anything new, and energy tied up in the forward motion of the whole cannot be spent. Doubling the beam energy of a machine whose beams meet head-on doubles what is available; doubling it in one firing into a stationary target multiplies it by the square root of two. That difference in scaling is the whole argument for building colliders.

9 · Mass is not additive

Section 2.3's grind box already produced the result that makes this section necessary. Two lumps of putty of rest mass mm each, colliding and sticking, form a body of rest mass 2γm2\gamma m, not 2m2m. The kinetic energy did not disappear. It became rest mass. Run that backwards and you have the most consequential fact in applied physics.

9.1 · Binding energy and the mass defect

Consider a bound system: a nucleus, an atom, a planet. Assemble it from constituents that start far apart and at rest, and let it settle by radiating away the excess. The total four-momentum is conserved throughout, so in the frame where the final object sits at rest,

(imi)c2  =  Mboundc2  +  B,i.e.  Mbound=imiBc2   \Big(\sum_{i}m_{i}\Big)c^{2} \;=\; M_{\text{bound}}\,c^{2} \;+\; B, \qquad\text{i.e.}\qquad \boxed{\;M_{\text{bound}} = \sum_{i}m_{i} - \frac{B}{c^{2}}\;} (2.5.70)

where B>0B\gt0 is the energy carried away, called the binding energy. The bound system is lighter than its parts, by exactly the energy you would have to put back to take it apart, over c2c^{2}.

That difference is the mass defect. It is not a small correction to a picture in which mass is additive. It is a demonstration that mass was never additive.

Real numbers, since the whole point is that this is measurable. ⚑ The masses below are quoted from standard tables, and everything done with them is ours.

SystemConstituents (MeV/c2\mathrm{MeV}/c^{2})Bound massBBB/mc2B/\sum m c^{2}
Deuteron  2\ ^{2}H938.272+939.565=1877.837938.272 + 939.565 = 1877.8371875.6131875.6132.224 MeV2.224\ \mathrm{MeV}0.118%0.118\%
Helium-42(938.272)+2(939.565)=3755.6752(938.272)+2(939.565)=3755.6753727.3793727.37928.30 MeV28.30\ \mathrm{MeV}0.753%0.753\%
Hydrogen atom938.272+0.511=938.783938.272+0.511=938.783938.783938.78313.6 eV13.6\ \mathrm{eV}1.4×1081.4\times10^{-8}

Read the last column. Chemistry runs at 10810^{-8} or 101010^{-10} of the rest energy. That is why Lavoisier could weigh reactants and products and conclude that mass is conserved. His balance was many orders of magnitude too coarse to see the defect, and he was right to the precision available.

Nuclear binding runs at 10310^{-3}, which is five to seven orders of magnitude larger. That gap is the entire reason nuclear energy is a different kind of thing rather than a better kind of chemistry.

Grind box — the arithmetic, and a fusion reaction end to end

Deuteron. A proton and a neutron bound together. The constituent rest energies are 938.272+939.565=1877.837 MeV938.272+939.565=1877.837\ \mathrm{MeV}, and the measured deuteron rest energy is 1875.613 MeV1875.613\ \mathrm{MeV}. Difference:

B=1877.8371875.613=2.224 MeV,Bmc2=2.2241877.837=1.184×103. B = 1877.837 - 1875.613 = 2.224\ \mathrm{MeV}, \qquad \frac{B}{\sum m c^{2}} = \frac{2.224}{1877.837} = 1.184\times10^{-3}.

So a deuteron weighs about 0.12%0.12\% less than the parts you made it from. That is a fractional mass difference roughly 10510^{5} times larger than any chemical bond, and comfortably measurable with a mass spectrometer.

Helium-4. The constituents come to 2×938.272+2×939.565=3755.6752\times938.272 + 2\times939.565 = 3755.675 and the measured mass to 3727.3793727.379, so B=28.296 MeVB=28.296\ \mathrm{MeV}. That is 7.07 MeV7.07\ \mathrm{MeV} per nucleon and 0.753%0.753\% of the total. Helium-4 is unusually tightly bound for its size, which is why α\alpha particles exist as a distinct thing and why stellar nucleosynthesis piles up at helium.

A complete reaction: D–T fusion.  2H+3H4He+n\ ^{2}\mathrm{H} + {}^{3}\mathrm{H} \to {}^{4}\mathrm{He} + n. Using nuclear rest energies in MeV:

in:1875.613+2808.921=4684.534out:3727.379+939.565=4666.944released:Q=4684.5344666.944=17.59 MeV. \begin{aligned} \text{in:}\quad & 1875.613 + 2808.921 = 4684.534\\ \text{out:}\quad & 3727.379 + 939.565 = 4666.944\\[3pt] \text{released:}\quad & Q = 4684.534 - 4666.944 = 17.59\ \mathrm{MeV}. \end{aligned}

As a fraction of the input rest energy, 17.59/4684.534=3.755×10317.59/4684.534 = 3.755\times10^{-3}, which is 0.375%0.375\%. Per kilogram of fuel that is 0.00375c2=3.4×1014 Jkg10.00375\,c^{2}=3.4\times10^{14}\ \mathrm{J\,kg^{-1}}, against roughly 5×107 Jkg15\times10^{7}\ \mathrm{J\,kg^{-1}} for burning hydrocarbons. The ratio is about seven million.

The lesson to extract. Every one of these is the same calculation: add the rest energies before, add them after, and the difference is kinetic energy. Nothing was "converted into energy" in a way that requires new physics. The bookkeeping quantity iEi\sum_{i}E_{i} was conserved throughout, and all that changed was how much of it was sitting in the rest-mass column.

⚠ Why this isn't obvious — "mass converts to energy" is the wrong sentence

The standard telling is that in a nuclear reaction "some mass is converted into energy, according to E=mc2E=mc^{2}". Nothing in this chapter supports that reading, and it causes real confusion.

What is actually true. The conserved quantity is E=iγimic2E=\sum_{i}\gamma_{i}m_{i}c^{2}, and it never changes. In a fusion reaction the rest masses of the products are smaller than those of the reactants, so the rest-energy share of the total falls, and the difference reappears in the kinetic-energy share. Nothing was created or destroyed and nothing turned into anything. Energy was reallocated between two columns of the same ledger.

And the system's mass? Do the reaction inside a sealed, perfectly reflecting box and weigh the box. Its invariant mass is unchanged, because the kinetic energy and radiation are still inside, and by §8.2 the box's invariant mass is its total energy in its rest frame over c2c^{2}. The mass only drops once you let the heat out.

So the honest sentence is this. Rest energy is released as kinetic energy, and if the kinetic energy escapes, the system's rest mass drops by exactly the energy that left, over c2c^{2}.

A related trap. "E=mc2E=mc^{2}" is about rest energy. The energy of a moving body is E=γmc2E=\gamma mc^{2}, and the two differ by every factor that matters at an accelerator. When a physicist writes E=mc2E=mc^{2} they mean the v=0v=0 case of (2.5.26), and the famous equation is famous partly because that qualification is usually dropped.

9.2 · Where your mass actually comes from

Push the argument one step further than nuclear physics does and it stops being a correction and becomes the main effect.

A proton has rest energy 938 MeV938\ \mathrm{MeV}. It is made of three light quarks whose rest energies sum to roughly 9 MeV9\ \mathrm{MeV}. ⚑ That figure is quoted from Chapter 6.5's tables, and note that even defining a quark mass takes care, since quarks are never found alone. It accounts for about 1%1\% of the proton.

The other 99%99\% is exactly what §8.2's warning described. It is the energy of quarks moving relativistically inside a small volume, plus the energy of the gluon field binding them, all of it appearing as invariant mass because the constituent four-momenta point in many different directions and their spatial parts cancel while their energies add.

So the mass of ordinary matter is, to better than 99%99\%, not a property of its ingredients. It is confined kinetic and field energy, measured in the frame where the total momentum vanishes. That is (2.5.60), applied to a bound state of a strongly coupled field theory.

Chapter 6.5 does that calculation, or rather explains why it can only be done numerically. The structural statement is available now, and it belongs to this chapter rather than to quantum chromodynamics.

Familiar ground — mass adds the way a variance does, and then stops

Everybody has had to explain that variances add and standard deviations do not. The reason is a norm identity, a+b2=a2+b2+2ab\norm{a+b}^{2}=\norm{a}^{2}+\norm{b}^{2}+2\,a\cdot b, whose cross term is the covariance. When the parts are uncorrelated the cross term vanishes and the squares add. Section 8 is that identity in a space with a different inner product.

Write (2.5.60) out for two particles, using (2.5.59):

M2c4  =  m12c4+m22c4+2(E1E2p1p2c2). M^{2}c^{4} \;=\; m_{1}^{2}c^{4} + m_{2}^{2}c^{4} + 2\big(E_{1}E_{2} - \vv p_{1}\cdot\vv p_{2}\,c^{2}\big).

The squares add and a cross term measures alignment, exactly as before. Two photons flying apart have E1E2p1p2c2=2E1E2E_{1}E_{2}-\vv p_{1}\cdot\vv p_{2}c^{2}=2E_{1}E_{2} and therefore a mass built entirely out of massless parts. The same two aimed the same way have a cross term of zero and no mass at all.

Nothing changed but the correlation between the directions, which is why §9.2's proton weighs what it does. Its constituents' momenta point every way, the cross terms survive in full, and the mass of the whole is mostly the misalignment of the parts.

Now the place the analogy stops, which is worth more than the place it holds. A covariance may be negative, so Var(X+Y)\mathrm{Var}(X+Y) can fall below VarX+VarY\mathrm{Var}\,X+\mathrm{Var}\,Y. The Minkowski cross term cannot do the corresponding thing. For two future-pointing momenta,

E1E2p1p2c2    m1m2c4, E_{1}E_{2} - \vv p_{1}\cdot\vv p_{2}\,c^{2} \;\ge\; m_{1}m_{2}c^{4},

with equality only when the two move together, so Mm1+m2M\ge m_{1}+m_{2} for any collection of free particles. There is no anti-correlated arrangement that lightens a system. A bound system does weigh less than its pieces, and §9.1 is careful about why: energy physically left, carried away as radiation, which is not this identity at all.

The algebra of the norm of a sum is shared. The sign structure that makes this one a one-sided bound is not, and it comes from the same minus sign that produced the causal structure two chapters ago.

In plain terms 2.5.9

A bound system, assembled from constituents that began far apart and settled by radiating the excess away, weighs less than its parts by exactly what left. Chemistry does this at one part in a hundred million, far under what Lavoisier's balance could resolve, which is why he pronounced mass conserved. Nuclear binding does it at one part in a thousand, and that gap of five orders is why nuclear energy differs from chemistry in kind.

The usual telling, in which mass converts into energy, is not what happens. The conserved total never changes; only which column it sits in, rest mass or motion. Run the reaction in a sealed reflecting box and weigh it: the mass is unchanged, everything released still inside. It falls only when the heat is let out.

One step further and the correction becomes the main effect. The three quarks in a proton account for about one hundredth of its mass; the rest is their confined motion and the binding field, appearing as mass because their momenta point every way and cancel while the energies add. That number is the one the chapter on expansion said no series would find, lying so flat near the origin that every term reports it as zero. Mass entered as the amount of stuff in a body and leaves as a label on a four-part object, better than ninety-nine per cent of yours borrowed motion.

10 · Worked examples

Worked example 1 — Compton scattering, in full

A photon of wavelength λ\lambda strikes a free electron at rest. The photon scatters through an angle θ\theta and emerges with wavelength λ\lambda', while the electron recoils in some unknown direction with some unknown speed. Find λλ\lambda'-\lambda.

Why this is the model calculation. There are four unknowns after the collision, namely the photon's new wavelength and the electron's three momentum components, and there are only four conservation equations. But we are not asked about the electron at all.

So here is the technique. Isolate the four-momentum you do not care about on one side, and square it. Squaring converts an unknown four-vector into the one number you do know about it, namely m2c2m^{2}c^{2}. That single move removes the recoil direction from the problem without ever computing it.

Setup. ⚑ One quoted input, and only one: a photon of wavelength λ\lambda has energy E=hc/λE=hc/\lambda. That is Planck and Einstein, and Chapter 4.1 is where it is argued for. Everything else below is this chapter's. Given it, §4.3's E=pcE=pc gives the photon's momentum p=h/λ\abs{\vv p}=h/\lambda, so with n^\hat n a unit vector along its travel,

pγμ=hλ(1, n^),pγμ=hλ(1, n^),n^n^=cosθ, p_{\gamma}^{\mu} = \frac{h}{\lambda}\big(1,\ \hat n\big), \qquad p_{\gamma}'^{\mu} = \frac{h}{\lambda'}\big(1,\ \hat n'\big), \qquad \hat n\cdot\hat n' = \cos\theta,

and the electron starts at rest, peμ=(mec, 0)p_{e}^{\mu}=(m_{e}c,\ \vv 0), with unknown final peμp_{e}'^{\mu}.

Step 1. Conservation, rearranged to isolate the unknown.

pγμ+peμ=pγμ+peμpeμ=pγμpγμ+peμ. p_{\gamma}^{\mu} + p_{e}^{\mu} = p_{\gamma}'^{\mu} + p_{e}'^{\mu} \qquad\Longrightarrow\qquad p_{e}'^{\mu} = p_{\gamma}^{\mu} - p_{\gamma}'^{\mu} + p_{e}^{\mu}.

Step 2. Square both sides. The left side is known, since pepe=me2c2p_{e}'\cdot p_{e}' = m_{e}^{2}c^{2} by (2.5.29), whatever the electron ended up doing. Expanding the right side needs six terms, so take them one at a time:

pepe  =  pγpγ=0+pγpγ=0+pepe=me2c22pγpγ+2pγpe2pγpe. \begin{aligned} p_{e}'\cdot p_{e}' \;=\;& \underbrace{p_{\gamma}\cdot p_{\gamma}}_{=\,0} + \underbrace{p_{\gamma}'\cdot p_{\gamma}'}_{=\,0} + \underbrace{p_{e}\cdot p_{e}}_{=\,m_{e}^{2}c^{2}}\\[3pt] &- 2\,p_{\gamma}\cdot p_{\gamma}' + 2\,p_{\gamma}\cdot p_{e} - 2\,p_{\gamma}'\cdot p_{e}. \end{aligned}

The first two vanish because photons are massless. Setting the whole thing equal to me2c2m_{e}^{2}c^{2}, the me2c2m_{e}^{2}c^{2} terms cancel and we are left with a relation among three dot products:

pγpγ  =  pγpe    pγpe. p_{\gamma}\cdot p_{\gamma}' \;=\; p_{\gamma}\cdot p_{e} \;-\; p_{\gamma}'\cdot p_{e}.

The electron's final state has vanished entirely. That is the whole trick.

Step 3. Evaluate the three dot products. Each is elementary. Keep the metric signs straight.

pγpγ=hλhλ(11n^n^)=h2λλ(1cosθ), p_{\gamma}\cdot p_{\gamma}' = \frac{h}{\lambda}\frac{h}{\lambda'}\big(1\cdot1 - \hat n\cdot\hat n'\big) = \frac{h^{2}}{\lambda\lambda'}\big(1-\cos\theta\big), pγpe=hλmec(1)hλn^0=hmecλ,pγpe=hmecλ. p_{\gamma}\cdot p_{e} = \frac{h}{\lambda}\,m_{e}c\,(1) - \frac{h}{\lambda}\hat n\cdot\vv 0 = \frac{h\,m_{e}c}{\lambda}, \qquad p_{\gamma}'\cdot p_{e} = \frac{h\,m_{e}c}{\lambda'}.

Step 4. Assemble.

h2λλ(1cosθ)  =  hmec(1λ1λ)  =  hmecλλλλ. \frac{h^{2}}{\lambda\lambda'}\big(1-\cos\theta\big) \;=\; h\,m_{e}c\left(\frac{1}{\lambda}-\frac{1}{\lambda'}\right) \;=\; h\,m_{e}c\,\frac{\lambda'-\lambda}{\lambda\lambda'}.

Now multiply through by λλ/(hmec)\lambda\lambda'/(h\,m_{e}c). The product λλ\lambda\lambda' was the only place the two wavelengths appeared together, and it cancels completely:

  λλ  =  hmec(1cosθ).   \boxed{\;\lambda' - \lambda \;=\; \frac{h}{m_{e}c}\big(1-\cos\theta\big).\;}

What the answer says. The shift does not depend on λ\lambda. It is bounded, between 00 at θ=0\theta=0 and 2h/mec2h/m_{e}c at θ=180\theta=180^{\circ}. And it is a pure length times a pure number.

The Compton wavelength, and a payoff from Chapter 0.3. The combination

λChmec=6.626×1034(9.109×1031)(2.998×108)=2.426×1012 m=2.426 pm \lambda_{C} \equiv \frac{h}{m_{e}c} = \frac{6.626\times10^{-34}}{(9.109\times10^{-31})(2.998\times10^{8})} = 2.426\times10^{-12}\ \mathrm{m} = 2.426\ \mathrm{pm}

is the only length you can build from hh, mem_{e} and cc. Check that against Chapter 0.3 §5's method. With [h]=ML2T1[h]=\mathsf{ML}^{2}\mathsf{T}^{-1}, [me]=M[m_{e}]=\mathsf{M} and [c]=LT1[c]=\mathsf{LT}^{-1}, seeking hAmeBcCh^{A}m_{e}^{B}c^{C} of dimension L\mathsf{L} gives A+B=0A+B=0, 2A+C=12A+C=1, AC=0-A-C=0, whence A=1A=1, C=1C=-1, B=1B=-1, uniquely.

So dimensional analysis alone guaranteed that if the answer is a wavelength shift built from these three constants, it must be λC\lambda_{C} times a dimensionless function of θ\theta. The four-vector calculation supplied the function, and it is 1cosθ1-\cos\theta.

Numbers, and why anyone believed it. Compton used molybdenum KαK_{\alpha} X-rays, λ=71 pm\lambda=71\ \mathrm{pm}. At θ=90\theta=90^{\circ} the shift is 2.426 pm2.426\ \mathrm{pm}, a 3.4%3.4\% change, easily resolved. At θ=180\theta=180^{\circ} it is 4.853 pm4.853\ \mathrm{pm}, or 6.8%6.8\%.

Crucially, the shift is independent of the incident wavelength and of the target material, which is impossible for any classical scattering mechanism. A classical wave shakes an electron at the driving frequency and the electron re-radiates at that same frequency, giving no shift at all.

Chapter 4.1 uses exactly this calculation as evidence that light carries momentum in discrete parcels obeying (2.5.33). The collision is between two particles, and the bookkeeping that works is this chapter's.

Worked example 2 — what a fixed-target LHC would cost

The LHC collides protons head-on at 6.8 TeV6.8\ \mathrm{TeV} per beam, giving s=13.6 TeV\sqrt{s}=13.6\ \mathrm{TeV}. What beam energy would a fixed-target machine need to reach the same s\sqrt{s} against a stationary proton? And conversely, what does the LHC's own beam achieve if you point it at a stationary target instead? (⚑ The machine parameters are quoted, and the physics is ours.)

Part (a), the fixed-target equivalent. Rearrange (2.5.64) with m1=m2=mpm_{1}=m_{2}=m_{p}:

E=s2mp2c42mpc2. E = \frac{s - 2m_{p}^{2}c^{4}}{2m_{p}c^{2}}.

With s=13.6 TeV=1.36×104 GeV\sqrt{s}=13.6\ \mathrm{TeV}=1.36\times10^{4}\ \mathrm{GeV} and mpc2=0.9383 GeVm_{p}c^{2}=0.9383\ \mathrm{GeV}, the 2mp2c4=1.76 GeV22m_{p}^{2}c^{4}=1.76\ \mathrm{GeV}^{2} is utterly negligible against s=1.85×108 GeV2s=1.85\times10^{8}\ \mathrm{GeV}^{2}, so

E1.850×1082×0.9383 GeV=9.86×107 GeV=9.86×104 TeV99 PeV. E \approx \frac{1.850\times10^{8}}{2\times0.9383}\ \mathrm{GeV} = 9.86\times10^{7}\ \mathrm{GeV} = 9.86\times10^{4}\ \mathrm{TeV} \approx 99\ \mathrm{PeV}.

Compare that with the 6.8 TeV6.8\ \mathrm{TeV} each LHC beam actually carries. The fixed-target machine would need

9.86×1076.8×1031.45×104 \frac{9.86\times10^{7}}{6.8\times10^{3}} \approx 1.45\times10^{4}

times more energy per particle, which is about fourteen and a half thousand. The LHC's ring is 27 km27\ \mathrm{km} around and its bending magnets are near the limit of what superconductors will do. Scaling the same technology by 1.45×1041.45\times10^{4} means a ring on the order of 4×105 km4\times10^{5}\ \mathrm{km}, which is about the distance to the Moon. That is the E\sqrt{E} scaling, made visceral.

Part (b), the LHC beam on a fixed target. Same formula, other direction, with E=6.8×103 GeVE=6.8\times10^{3}\ \mathrm{GeV}:

s=2mpc2E+2mp2c4=2(0.9383)(6800)+1.76  GeV=113 GeV. \sqrt{s} = \sqrt{2m_{p}c^{2}E + 2m_{p}^{2}c^{4}} = \sqrt{2(0.9383)(6800)+1.76}\ \ \mathrm{GeV} = 113\ \mathrm{GeV}.

So a 6.8 TeV6.8\ \mathrm{TeV} proton hitting a stationary proton makes only 113 GeV113\ \mathrm{GeV} of invariant mass available, which is 1.7% of the beam energy. The other 98.3% is spent dragging the centre of mass down the beam pipe and is gone.

The overhead is not a constant. Measure the waste by the ratio of beam energy to available energy, E/sE/\sqrt{s}. Rearranging (2.5.64) for equal masses gives

Es  =  s2mpc2mpc2s   smpc2   s2mpc2, \frac{E}{\sqrt{s}} \;=\; \frac{\sqrt{s}}{2m_{p}c^{2}} - \frac{m_{p}c^{2}}{\sqrt{s}} \;\xrightarrow[\ \sqrt{s}\gg m_{p}c^{2}\ ]{}\; \frac{\sqrt{s}}{2m_{p}c^{2}},

which grows linearly with the energy you are trying to reach. At the antiproton threshold of §8.4, s=4mpc2=3.75 GeV\sqrt{s}=4m_{p}c^{2}=3.75\ \mathrm{GeV} and the ratio is 214=1.752-\tfrac14=1.75. You buy 3.75 GeV3.75\ \mathrm{GeV} of invariant mass with a 6.57 GeV6.57\ \mathrm{GeV} beam, which is tolerable. At the LHC's beam energy the same ratio is 6800/113=606800/113=60.

So there is no fixed inefficiency to design around. The waste compounds. That is why every energy-frontier machine built since the 1970s has been a collider, and why fixed-target experiments now go looking for rare processes at modest s\sqrt{s} rather than for new mass.

11 · Your turn

Problem 1 — a photon cannot decay into an electron–positron pair

In empty space, consider the proposed process γe+e+\gamma \to e^{-} + e^{+}. Show, using invariant mass alone, that it is impossible, no matter how energetic the photon is. Then explain in one sentence why the same process does occur routinely near a heavy nucleus.

Solution

Four-momentum conservation would require pγμ=pμ+p+μp_{\gamma}^{\mu} = p_{-}^{\mu}+p_{+}^{\mu}. Squaring both sides is legitimate because both are four-vectors, and the square is an invariant, so if the two sides are equal their squares are equal in every frame.

Left side. The photon is massless, so pγpγ=0p_{\gamma}\cdot p_{\gamma}=0.

Right side. Let Pμ=pμ+p+μP^{\mu}=p_{-}^{\mu}+p_{+}^{\mu}. Each electron four-momentum is timelike and future-pointing, since energies are positive, so their sum is too, and by §8.2 there is a frame in which P=0\vv P=0. Evaluate the invariant there, which is the whole art:

PPCM=(E+E+)2c20. P\cdot P \Big|_{\text{CM}} = \frac{(E_{-}+E_{+})^{2}}{c^{2}} - 0.

In that frame the two particles have equal and opposite momenta, so equal energies, and each satisfies E±mec2E_{\pm}\ge m_{e}c^{2} by (2.5.31). Hence

PP    (2mec2)2c2  =  4me2c2  >  0. P\cdot P \;\ge\; \frac{(2m_{e}c^{2})^{2}}{c^{2}} \;=\; 4m_{e}^{2}c^{2} \;\gt\; 0.

So the process demands 0=PP4me2c20 = P\cdot P \ge 4m_{e}^{2}c^{2}, a contradiction. The photon's energy never entered the argument, because PPP\cdot P is an invariant and we were free to compute it in the pair's own rest frame. \blacksquare

Restated geometrically, which is the version to remember: a null vector cannot be the sum of two timelike future-pointing vectors. The sum of future-timelike vectors is future-timelike, strictly. The light cone is the boundary, and you cannot get back onto it by adding things from the inside.

Near a nucleus. The nucleus absorbs four-momentum, so the reaction is really γ+Ne+e++N\gamma + N \to e^{-}+e^{+}+N, and the initial invariant mass is no longer zero. By (2.5.63) it is s=mN2c4+2mNc2Eγs=m_{N}^{2}c^{4}+2m_{N}c^{2}E_{\gamma}, which exceeds (mN+2me)2c4(m_{N}+2m_{e})^{2}c^{4} once EγE_{\gamma} is large enough. Applying the grind box's threshold formula with m1=0m_{1}=0, m2=mNm_{2}=m_{N}, M=mN+2me\mathcal M = m_{N}+2m_{e}:

Eγthr=(mN+2me)2mN22mNc2=2mec2(1+memN), E_{\gamma}^{\text{thr}} = \frac{(m_{N}+2m_{e})^{2}-m_{N}^{2}}{2m_{N}}c^{2} = 2m_{e}c^{2}\left(1+\frac{m_{e}}{m_{N}}\right),

which for a heavy nucleus is barely above 2mec2=1.022 MeV2m_{e}c^{2}=1.022\ \mathrm{MeV}. The nucleus takes almost no energy but supplies the momentum balance. It is there to break the kinematics, not to pay for it. This is why 1.022 MeV1.022\ \mathrm{MeV} is the number quoted for pair production, and why positron emission tomography works at 511 keV511\ \mathrm{keV} per photon in the reverse process.

Problem 2 — the relativistic rocket, and its energy budget

A ship accelerates with constant proper acceleration α\alpha, meaning the accelerometer bolted to its floor always reads α\alpha, starting from rest at τ=0\tau=0.

(a) Using uu=c2u\cdot u=c^{2}, ua=0u\cdot a=0 and aa=α2a\cdot a=-\alpha^{2}, show that uμ=c(coshϕ,sinhϕ,0,0)u^{\mu}=c(\cosh\phi,\sinh\phi,0,0) with ϕ=ατ/c\phi = \alpha\tau/c, and hence that v/c=tanh(ατ/c)v/c=\tanh(\alpha\tau/c) and γ=cosh(ατ/c)\gamma=\cosh(\alpha\tau/c).

(b) With α=g=9.81 ms2\alpha=g=9.81\ \mathrm{m\,s^{-2}}, find v/cv/c and γ\gamma after one year and after ten years of ship-time, and the energy per kilogram of payload required in each case.

Solution

(a) The constraint uu=c2u\cdot u=c^{2} says the four-velocity is confined to a hyperbola in the (u0,u1)(u^{0},u^{1}) plane, and the general parametrisation of X2Y2=c2X^{2}-Y^{2}=c^{2} is X=ccoshϕX=c\cosh\phi, Y=csinhϕY=c\sinh\phi for some function ϕ(τ)\phi(\tau). That is precisely Chapter 2.3 §3.2's hyperbolic parametrisation, which is why rapidity was worth defining. So uμ=c(coshϕ,sinhϕ,0,0)u^{\mu}=c(\cosh\phi,\sinh\phi,0,0) with no loss of generality. Differentiate:

aμ=duμdτ=cdϕdτ(sinhϕ, coshϕ, 0, 0). a^{\mu} = \dv{u^{\mu}}{\tau} = c\,\dv{\phi}{\tau}\big(\sinh\phi,\ \cosh\phi,\ 0,\ 0\big).

Check the orthogonality first, as a sanity test: ua=c2ϕ˙(coshϕsinhϕsinhϕcoshϕ)=0u\cdot a = c^{2}\dot\phi(\cosh\phi\sinh\phi - \sinh\phi\cosh\phi)=0 ✓, automatically. The constraint (2.5.13) is built into the parametrisation. Now the magnitude:

aa=c2ϕ˙2(sinh2ϕcosh2ϕ)=c2ϕ˙2. a\cdot a = c^{2}\dot\phi^{2}\big(\sinh^{2}\phi-\cosh^{2}\phi\big) = -c^{2}\dot\phi^{2}.

Setting this equal to α2-\alpha^{2} gives ϕ˙=α/c\dot\phi = \alpha/c, so ϕ=ατ/c\phi=\alpha\tau/c with ϕ(0)=0\phi(0)=0 for a start from rest. Finally, from uμ=γ(c,v)u^{\mu}=\gamma(c,\vv v) we read off γ=coshϕ\gamma=\cosh\phi and γv=csinhϕ\gamma v = c\sinh\phi, hence

vc=tanh ⁣(ατc),γ=cosh ⁣(ατc). \frac{v}{c} = \tanh\!\left(\frac{\alpha\tau}{c}\right), \qquad \gamma = \cosh\!\left(\frac{\alpha\tau}{c}\right).

Constant proper acceleration is uniform growth of rapidity. That is the clean statement, and it is why tanh\tanh appears. Velocities do not add, rapidities do, and a constant push adds rapidity at a constant rate. It also settles the speed limit question for good, since tanh\tanh approaches 11 and never reaches it, however long you burn.

(b) The natural timescale is c/g=(2.998×108)/9.81=3.056×107 s=0.968 yrc/g = (2.998\times10^{8})/9.81 = 3.056\times10^{7}\ \mathrm{s} = 0.968\ \mathrm{yr}. So one year of ship-time is ϕ=1/0.968=1.033\phi = 1/0.968 = 1.033.

ship-time τ\tauϕ=gτ/c\phi=g\tau/cv/c=tanhϕv/c=\tanh\phiγ=coshϕ\gamma=\cosh\phikinetic energy per kg
1 yr1\ \mathrm{yr}1.0331.0330.77500.77501.5821.5825.23×1016 J5.23\times10^{16}\ \mathrm{J}
10 yr10\ \mathrm{yr}10.3310.3312×1091-2\times10^{-9}1.53×1041.53\times10^{4}1.37×1021 J1.37\times10^{21}\ \mathrm{J}

The kinetic energy is Emc2=(γ1)mc2E-mc^{2}=(\gamma-1)mc^{2}, with c2=8.988×1016 Jkg1c^{2}=8.988\times10^{16}\ \mathrm{J\,kg^{-1}}. After one year it is 0.582mc20.582\,mc^{2} per kilogram, which is more than half the payload's rest energy, and that is only the payload, ignoring the fuel needed to carry the fuel. After ten years it is 1.5×1041.5\times10^{4} times the rest energy. To deliver one kilogram you must supply the total annihilation energy of about fifteen tonnes, and no engine is perfect.

The point. Relativity does not forbid interstellar travel. It prices it, and it prices it exponentially, because γ=coshϕ12eϕ\gamma=\cosh\phi\sim\tfrac12\ee^{\phi} grows exponentially in ship-time, and so does the distance covered, x=(c2/g)(coshϕ1)x=(c^{2}/g)(\cosh\phi-1). Set x=26000x=26\,000 light-years, the distance to the galactic centre, and you get coshϕ=2.68×104\cosh\phi=2.68\times10^{4}, hence ϕ=10.89\phi=10.89 and τ=10.5\tau=10.5 years of ship-time. That is the romance, and it is real. The bill is (γ1)mc22.7×104mc2(\gamma-1)mc^{2}\approx2.7\times10^{4}\,mc^{2} per kilogram delivered. And since the ship never quite reaches cc, some 2.6×1042.6\times10^{4} years will have passed on Earth by the time it arrives.

Problem 3 — the headlight effect

A source at rest in frame SS' emits photons uniformly in all directions. SS' moves at βc\beta c along xx relative to the lab. Using the aberration formula from §7's grind box, show that the photons emitted into the forward hemisphere in SS' (those with θ<90\theta'\lt90^{\circ}) are all compressed, in the lab, into a cone about the xx-axis of half-angle θ1/2=arccosβ\theta_{1/2}=\arccos\beta, and that for γ1\gamma\gg1 this is approximately 1/γ1/\gamma. Evaluate for γ=10\gamma=10 and γ=1000\gamma=1000, and comment on what this does to the apparent brightness of a relativistic jet pointed at you.

Solution

The grind box in §7 derived cosθ=(cosθβ)/(1βcosθ)\cos\theta' = (\cos\theta-\beta)/(1-\beta\cos\theta), mapping lab angle to source angle. We want the inverse, which is the same formula with ββ\beta\to-\beta, since that swaps the roles of the frames:

cosθ=cosθ+β1+βcosθ. \cos\theta = \frac{\cos\theta'+\beta}{1+\beta\cos\theta'}.

The boundary ray. Put θ=90\theta'=90^{\circ}, so that cosθ=0\cos\theta'=0. Then cosθ=β\cos\theta = \beta. So the photon emitted exactly sideways in the source frame arrives in the lab at angle θ1/2=arccosβ\theta_{1/2}=\arccos\beta to the direction of motion. Since the map is monotonic in cosθ\cos\theta', every ray with θ<90\theta'\lt90^{\circ} lands inside that cone.

The small-angle form. For θ1/2\theta_{1/2} small, sinθ1/2=1cos2θ1/2=1β2=1/γ\sin\theta_{1/2}=\sqrt{1-\cos^{2}\theta_{1/2}} = \sqrt{1-\beta^{2}} = 1/\gamma, so θ1/21/γ\theta_{1/2}\approx 1/\gamma radians. Neat, and worth remembering as a rule of thumb.

γ\gammaβ\betaθ1/2=arccosβ\theta_{1/2}=\arccos\betasolid angle fraction
10100.994990.994995.73=0.100 rad5.73^{\circ}=0.100\ \mathrm{rad}2.5×103\approx2.5\times10^{-3}
100010000.99999950.99999950.0573=1.0×103 rad0.0573^{\circ}=1.0\times10^{-3}\ \mathrm{rad}2.5×107\approx2.5\times10^{-7}

(The solid-angle fraction of a cone of half-angle θ\theta is (1cosθ)/2θ2/4=1/4γ2(1-\cos\theta)/2\approx\theta^{2}/4 = 1/4\gamma^{2}.)

Brightness. Half of all the emitted photons, meaning the forward hemisphere, which is 2π2\pi steradians in the source frame, end up inside a lab cone of solid angle π/γ2\approx\pi/\gamma^{2}. The photon flux per unit solid angle is therefore boosted by roughly γ2\gamma^{2}.

Each of those photons is also blueshifted by the longitudinal Doppler factor (1+β)/(1β)2γ\sqrt{(1+\beta)/(1-\beta)}\approx2\gamma from (2.5.55), and they arrive at a compressed rate for the same reason. Multiplying the effects gives an apparent brightness enhanced by a large power of γ\gamma. The standard estimate for a continuously emitting jet is γ4\gamma^{4} in the observed flux, though the exact exponent depends on the source's spectral shape.

Consequence. A jet pointed within 1/γ1/\gamma of your line of sight looks enormously brighter than the identical jet pointed away. That is why blazars, which are active galaxies whose jets happen to aim at Earth, dominate the catalogues of the brightest gamma-ray sources despite being a tiny fraction of active galaxies. It is also why radio astronomers see one-sided jets from manifestly two-sided sources, the receding jet being de-beamed by the same large factor. Neither observation requires any asymmetry in the source. Both are (2.5.52), boosted.

Problem 4 — mass defect and energy release in D–D fusion

Consider 2H+2H3He+n^{2}\mathrm{H} + {}^{2}\mathrm{H} \to {}^{3}\mathrm{He} + n. Using nuclear rest energies mdc2=1875.613m_{d}c^{2}=1875.613, m3Hec2=2808.391m_{^{3}\mathrm{He}}c^{2}=2808.391, mnc2=939.565m_{n}c^{2}=939.565 (all in MeV), compute (a) the energy released QQ; (b) QQ as a fraction of the input rest energy; (c) the energy released per kilogram of deuterium fuel, and compare with burning an equal mass of methane (5.6×107 Jkg1\approx5.6\times10^{7}\ \mathrm{J\,kg^{-1}}). (d) Finally, the neutron and the 3^{3}He share QQ as kinetic energy; use momentum conservation in the CM frame to find how it splits, and check that the non-relativistic treatment is justified.

Solution

(a) Total four-momentum conservation, evaluated in the CM frame where the initial nuclei are brought together with negligible kinetic energy, gives inmc2=outmc2+Q\sum_{\text{in}} mc^{2} = \sum_{\text{out}}mc^{2} + Q:

in:2×1875.613=3751.226 MeVout:2808.391+939.565=3747.956 MeVQ=3751.2263747.956=3.270 MeV. \begin{aligned} \text{in:}\quad & 2\times1875.613 = 3751.226\ \mathrm{MeV}\\ \text{out:}\quad & 2808.391 + 939.565 = 3747.956\ \mathrm{MeV}\\[3pt] Q &= 3751.226-3747.956 = 3.270\ \mathrm{MeV}. \end{aligned}

(b) Q/inmc2=3.270/3751.226=8.72×104Q/\sum_{\text{in}}mc^{2} = 3.270/3751.226 = 8.72\times10^{-4}, which is 0.0872%0.0872\%. That is under a tenth of a percent of the rest energy, and yet:

(c) Energy per kilogram is that fraction times c2c^{2}:

8.72×104×8.988×1016 Jkg1=7.8×1013 Jkg1. 8.72\times10^{-4} \times 8.988\times10^{16}\ \mathrm{J\,kg^{-1}} = 7.8\times10^{13}\ \mathrm{J\,kg^{-1}}.

Against methane's 5.6×1075.6\times10^{7}, that is a factor of 1.4×1061.4\times10^{6}, a million and a half. One kilogram of deuterium releases what about 1400 tonnes of methane would.

The entire difference between chemistry and nuclear physics is the difference between the last column of §9.1's table for a hydrogen atom (10810^{-8}) and for a nucleus (10310^{-3}). The mechanism is identical and only the scale of the binding differs.

(d) In the CM frame the products have equal and opposite momenta, pp. Non-relativistically the kinetic energies are Ti=p2/2miT_{i}=p^{2}/2m_{i}, so they are inversely proportional to the masses:

TnT3He=m3Hemn=2808.391939.565=2.989. \frac{T_{n}}{T_{^{3}\mathrm{He}}} = \frac{m_{^{3}\mathrm{He}}}{m_{n}} = \frac{2808.391}{939.565} = 2.989.

With Tn+T3He=Q=3.270 MeVT_{n}+T_{^{3}\mathrm{He}}=Q=3.270\ \mathrm{MeV}, this gives

Tn=Q2.9893.989=2.451 MeV,T3He=0.819 MeV. T_{n} = Q\,\frac{2.989}{3.989} = 2.451\ \mathrm{MeV}, \qquad T_{^{3}\mathrm{He}} = 0.819\ \mathrm{MeV}.

The light one takes most of the energy. That is a general feature of two-body decay, and it is the reason fusion reactors must cope with fast neutrons.

Justifying the approximation. The neutron's kinetic energy is 2.451/939.565=2.6×1032.451/939.565 = 2.6\times10^{-3} of its rest energy, so γ=1.0026\gamma=1.0026 and β=0.072\beta=0.072. By the grind box in §4, the fractional error in using T=p2/2mT=p^{2}/2m at this speed is about x4/8x^{4}/8 with x=γβ=0.0723x=\gamma\beta=0.0723, which comes to 3×1063\times10^{-6}. Three parts per million is far below the precision of the input masses, so the non-relativistic split is fine. Note that we needed relativity to compute QQ at all, and then did not need it to divide QQ up. That combination is typical of nuclear physics.

The brick you just laid

You have mechanics that survives a boost. The route was one demand, that a conservation law must be an equation between tensors, and everything else was consequence.

Differentiate by the invariant τ\tau to get uμu^{\mu} with uu=c2u\cdot u=c^{2}. Multiply by mm to get pμp^{\mu}. Observe that conserving three components of a four-vector in every frame forces the fourth, and that the fourth expands to mc2+12mv2+38mv4/c2+mc^{2}+\half mv^{2}+\tfrac38 mv^{4}/c^{2}+\cdots. From the invariant square came E2=p2c2+m2c4E^{2}=p^{2}c^{2}+m^{2}c^{4}, from that the massless case as the null case, and from v=pc2/E\vv v=\vv p c^{2}/E the fact that null four-momentum means exactly cc.

You also have the action. S=mc2 ⁣dτS=-mc^{2}\!\int\dd\tau was solved for, not guessed, and it closed Chapter 2.3's maximisation theorem. Extremising the action and maximising proper time are the same instruction, differing by the minus sign that converts one into the other. That is entry two in Chapter 1.2 §8.1's table, now derived.

And you have the tools that make relativistic problems tractable: contract to kill unknowns, boost to the frame where the answer is easy, and remember that invariant mass is a property of a four-momentum rather than a substance carried by parts.

Where this gets spent. Chapter 2.6 takes dpμ/dτ=fμ\dd p^{\mu}/\dd\tau=f^{\mu} and asks what four-vector the electromagnetic field puts on the right. The answer is qFμνuνqF^{\mu\nu}u_{\nu}, and its time component is the power (2.5.46). That chapter also finds the missing momentum that Chapter 1.1 left unaccounted for, and it is in the field.

Chapter 3.6 builds TμνT^{\mu\nu}, the object that sources gravity, out of exactly the energy and momentum densities defined here. Energy gravitates, not mass, which is why light bends.

Chapter 4.1 reuses Worked example 1 as evidence that photons carry momentum, and promotes pμ=kμp^{\mu}=\hbar k^{\mu} from a structural suspicion to a law. Chapter 5.1 shows that E=γmc2E=\gamma mc^{2} plus quantum mechanics makes particle number unfixable, since you can always find 2mc22mc^{2} somewhere, which is why fields replace particles. Chapter 5.9 computes a real cross-section in the Mandelstam variables §8.3 introduced. And Chapter 6.5 explains the 99%99\% of the proton's mass that §9.2 could locate but not account for.