Part II · Special Relativity — Chapter 2.6

Electromagnetism Is Relativity

Maxwell's equations never needed fixing. They were relativistic before anyone knew what that meant. That is exactly why they produced a frame-independent cc and started the crisis in the first place.

Where we are

Chapter 2.1 left you at a fork. Three things were believed in 1900: the principle of relativity, Maxwell's equations, and the Galilean transformation. Any two of them sat together comfortably. All three at once were impossible.

Three branches led away from that fork.

  • Branch (A) said relativity fails for optics, and there is an ether frame after all.
  • Branch (B) said Maxwell is wrong.
  • Branch (C) said the Galilean transformation is wrong.

Chapters 2.2 to 2.5 took branch (C) and rebuilt kinematics and dynamics from the wreckage.

This chapter collects the payment. Rewritten in the language of Chapter 2.4, Maxwell's four equations become two, and one of those two is not physics at all but bookkeeping.

The paradox of 2.1 §3 evaporates as well. That paradox was the question cc with respect to what? It goes away not because we answer it, but because we finally see that it was malformed. Maxwell's equations single out a speed. They never singled out a frame. Those are two different things. Only after 2.2 do you have a geometry in which a speed can come out the same for everyone without there being anything for it to be measured relative to.

Two promises also come due here.

The first is from Chapter 2.4 §7.2. That section counted six independent components in an antisymmetric 4×44\times4 tensor and told you they would turn out to be E\vv E and B\vv B. Sections 2 and 3 below deliver them, by construction, with every component checked.

The second is from Chapter 1.1 §3.3. That section showed Newton's third law failing for two moving charges, left an explicit IOU marked "the field carries momentum", and told you Chapter 2.6 would pay it. Section 10 pays it, down to the last factor of 32\tfrac32.

Tools you'll need  — Chapter 0.7: divergence and curl in Cartesian components (§3.3, §4.2), the continuity equation (§6), and the two second-derivative identities ×ϕ=0\nabla\times\nabla\phi=\vv 0 and (×A)=0\nabla\cdot(\nabla\times\vv A)=0 (§7), together with §7.3 on why the vector potential exists and is not unique. Chapter 1.1 §3.3: the two moving charges that break the third law. Section 10 here is the resolution, so it is worth rereading the setup. Chapter 1.2 §8: the field Euler–Lagrange equation, and §8.1's table of actions. Section 9 here supplies the third entry. Chapter 1.4 §6: Noether's theorem for fields, and its grind box on improvement terms, which §10 uses without apology. Chapter 2.1 §2 (Maxwell produces cc), §4 (the wave equation is not Galilean invariant) and the fork of §4.3. Chapter 2.4 §2.1 (the boost matrix Λμν\Lambda^{\mu}{}_{\nu}), §4 (raising and lowering), §5 (the tensor transformation law and the quotient theorem), §6 (the invariance theorem), §7.2 (the count of six), §8.1 (the Levi-Civita symbol). Chapter 2.5 §1.3 (four-velocity), §2 (four-momentum), §6.1 (the covariant force dpμ/dτ=fμ\dd p^{\mu}/\dd\tau=f^{\mu} with fu=0f\cdot u=0).

⚑ What is being assumed, stated once

This chapter does not derive electromagnetism from nothing. Two things are taken as given.

(i) Maxwell's four equations, quoted in Chapter 2.1 §2 as the empirical input of that chapter, in SI units and with sources present:

E=ρϵ0,×E=Bt,B=0,×B=μ0J+μ0ϵ0Et. \begin{aligned} \nabla\cdot\vv E &= \frac{\rho}{\epsilon_{0}}, &\qquad \nabla\times\vv E &= -\pdv{\vv B}{t},\\[4pt] \nabla\cdot\vv B &= 0, &\qquad \nabla\times\vv B &= \mu_{0}\vv J + \mu_{0}\epsilon_{0}\pdv{\vv E}{t}. \end{aligned}

(ii) Charge invariance. The electric charge of a body is the same number in every inertial frame. This is an experimental fact rather than a theorem, and it has been tested very sharply. Atoms and molecules come out neutral to better than one part in 102110^{21}, even though the electrons inside them move at wildly different speeds from the nuclei. Chapter 6.3 will make charge invariance a consequence of gauge symmetry. Here it is an input, and §6 is careful about exactly how much work it does.

Everything else below is derived. What the chapter produces is three observations. The first is that (i) is already a Lorentz-tensor equation in disguise. The second is that the disguise is the only thing that ever made it look incompatible with relativity. The third is that Chapter 6.3 can then run the whole construction backwards and obtain (i) from a symmetry principle.

1 · The four-current

We start with the simplest object in the theory. We are going to build it by requirement rather than by decree, which means we will state what we need of it and let that fix it completely.

Chapter 0.7 §6 derived, from nothing but "charge is neither created nor destroyed", the continuity equation

ρt  +  J  =  0, \pdv{\rho}{t} \;+\; \nabla\cdot\vv J \;=\; 0, (2.6.1)

Here ρ\rho is charge per unit volume and J\vv J is current density. Now look at (2.6.1) with the eyes of Chapter 2.4. It is a sum of four terms: one time derivative of one function, and three space derivatives of three functions. That is exactly the shape of a four-divergence μVμ\partial_{\mu}V^{\mu}. The catch is the word if: it is that shape only if the four functions can be arranged into a four-vector.

So let us not guess. Let us demand. Suppose there exists a four-vector jμj^{\mu} whose four-divergence reproduces (2.6.1), and ask what its components would have to be. Before we can compare anything we need the derivative operator in the right form. Recall from Chapter 2.4 §8 that μ=xμ\partial_{\mu}=\pdv{}{x^{\mu}} carries a lower index, and that x0=ctx^{0}=ct, so

0  =  x0  =  (ct)  =  1ct,i  =  xi. \partial_{0} \;=\; \pdv{}{x^{0}} \;=\; \pdv{}{(ct)} \;=\; \frac1c\pdv{}{t}, \qquad \partial_{i} \;=\; \pdv{}{x^{i}}. (2.6.2)

Our goal now is to write out the four-divergence in ordinary three-dimensional language, so that we can set it side by side with the continuity equation. Write the unknown four-vector as jμ=(j0,j)j^{\mu}=(j^{0},\vv j), split the sum over μ\mu into its time part and its three space parts, and use the two derivatives we just wrote down:

μjμ  =  0j0+iji  =  1cj0t  +  j. \partial_{\mu}j^{\mu} \;=\; \partial_{0}j^{0} + \partial_{i}j^{i} \;=\; \frac1c\pdv{j^{0}}{t} \;+\; \nabla\cdot\vv j. (2.6.3)

Now compare term by term with (2.6.1). The spatial part matches immediately if j=J\vv j=\vv J. The temporal part matches if j0/c=ρj^{0}/c=\rho, that is, if j0=cρj^{0}=c\rho. There is no freedom left anywhere. The two expressions agree as functions, for arbitrary ρ\rho and J\vv J, only for that one choice. So the four-vector we demanded exists, and it is this one:

  jμ  =  (cρ, J),μjμ  =  0.   \boxed{\;j^{\mu} \;=\; \big(c\rho,\ \vv J\big), \qquad \partial_{\mu}j^{\mu} \;=\; 0.\;} (2.6.4)

The factor of cc is not cosmetic. It is dimensional bookkeeping forced on us by x0=ctx^{0}=ct, exactly as the cc in xμ=(ct,x)x^{\mu}=(ct,\vv x) was.

Now let's look at what has happened to charge conservation. It arrived as a four-term partial differential equation whose behaviour under a change of frame was completely obscure. It now reads μjμ=0\partial_{\mu}j^{\mu}=0. That is a contraction of a (0,1)(0,1) index against a (1,0)(1,0) index, so it is a scalar by Chapter 2.4 §5.1, so it is the same statement in every frame by the invariance theorem of §6. Charge conservation is manifestly relativistic, and it takes one line to say so.

Why ρ\rho and J\vv J had to be one object

You should find this unsurprising. If you do not, here is the physical version. A line of static charges has ρ0\rho\neq0 and J=0\vv J=\vv 0. Run past it and the charges are moving, so now J0\vv J\neq\vv 0.

Charge density and current density are the same thing seen from different frames. That is exactly what happened to tt and x\vv x in Chapter 2.2, and to EE and p\vv p in Chapter 2.5. Section 6 of this chapter turns that one observation into the entire explanation of magnetism.

⚑ Quoted forward — where this current will come from

Chapter 1.4 §6 proved that a continuous symmetry of a field theory produces a conserved current satisfying μjμ=0\partial_{\mu}j^{\mu}=0, and §6.2 of that chapter ran the machinery on a global phase rotation of a complex field. That current, once the phase is allowed to vary from point to point, is (2.6.4). We are not in a position to show that yet, since it needs the gauge principle of Chapter 6.3. So for now jμj^{\mu} is built out of measured densities.

The logical order is still worth knowing in advance. Conservation of charge is not an extra postulate bolted onto electromagnetism. It is a theorem about a symmetry.

Grind box — why ρ\rho alone is not a scalar, and how the γ\gamma's work out

A common first guess is that charge density should be an invariant, since charge is. It is not, and the reason is instructive.

Take NN charges qq sitting at rest in a box of volume V0V_{0}, so ρ0=Nq/V0\rho_{0}=Nq/V_{0} in their rest frame. View them from a frame in which the box moves at speed vv along xx. The charge NqNq is unchanged, since that is the invariance assumption (ii). The volume is not unchanged. The box is contracted along xx by γ\gamma, so V=V0/γV=V_{0}/\gamma, and therefore

ρ  =  NqV  =  γρ0,J  =  ρv  =  γρ0v. \rho \;=\; \frac{Nq}{V} \;=\; \gamma\,\rho_{0}, \qquad \vv J \;=\; \rho\vv v \;=\; \gamma\rho_{0}\vv v.

Now assemble the four-vector and compare with the four-velocity uμ=γ(c,v)u^{\mu}=\gamma(c,\vv v) of Chapter 2.5 §1.3:

jμ  =  (cγρ0, γρ0v)  =  ρ0γ(c,v)  =  ρ0uμ. j^{\mu} \;=\; \big(c\gamma\rho_{0},\ \gamma\rho_{0}\vv v\big) \;=\; \rho_{0}\,\gamma\big(c,\vv v\big) \;=\; \rho_{0}\,u^{\mu}.

So the four-current of a moving cloud of charge is the invariant rest-frame density times the four-velocity. That is manifestly a four-vector, being a scalar multiplying a (1,0)(1,0) tensor.

Notice how the single factor of γ\gamma from length contraction is exactly the factor of γ\gamma that uμu^{\mu} already carries. It is the same γ\gamma, arriving twice for the same reason. Had charge density been invariant instead, jμj^{\mu} would not have been a four-vector at all, and (2.6.4) would be false.

One caution. The relation jμ=ρ0uμj^{\mu}=\rho_{0}u^{\mu} holds for a single species of charge all moving together. A copper wire has two species, a stationary lattice and a drifting electron gas, and you add their four-currents. That addition is the whole of §6, and the fact that the two species carry different factors of γ\gamma is the whole of magnetism.

In plain terms 2.6.1

Charge conservation arrived long ago as a four-term equation whose behaviour under a change of frame was completely opaque: one derivative in time of one function and three in space of three others. Set beside the machinery of the last two chapters, that is the shape of a four-dimensional divergence, provided the four functions assemble into one object. Rather than guess at the arrangement, demand it, and no freedom is left: the charge density, carrying one factor of the speed limit so the entries are commensurable, sits above the three of current.

Charge density is not itself an agreed number, although the charge is. The same charge occupies a contracted volume when you run past it, so the density picks up the dilation factor the four-velocity already carried. Density of charge and density of current are one thing seen from different states of motion, exactly as elapsed time and displacement were, and as energy and momentum were a chapter ago.

The machine assembled three chapters back for manufacturing agreement between observers was promised a velocity, then a momentum, then a current, and it now has all three and wants nothing further. What it returns on this last feeding is charge conservation as a single contraction, the same statement for everybody at a glance, where before it was four terms whose fate under a change of frame nobody could see.

2 · The four-potential and the field tensor

Now for the six components that Chapter 2.4 promised. The plan of this section is short. We recall the two potentials that Maxwell's equations force into existence, we stack them into one four-component array, and we then find that E\vv E and B\vv B are what you get by differentiating that array.

2.1 · Recalling the potentials

Chapter 0.7 §7.3 established the first half. Since (×A)=0\nabla\cdot(\nabla\times\vv A)=0 identically, writing

B  =  ×A \vv B \;=\; \nabla\times\vv A (2.6.5)

makes B=0\nabla\cdot\vv B=0 automatic. Feed that into Faraday's law:

×E  =  Bt  =  t(×A)  =  ×At×(E+At)  =  0, \begin{aligned} \nabla\times\vv E \;&=\; -\pdv{\vv B}{t} \;=\; -\pdv{}{t}\big(\nabla\times\vv A\big) \;=\; -\nabla\times\pdv{\vv A}{t}\\[5pt] &\Longrightarrow\qquad \nabla\times\left(\vv E + \pdv{\vv A}{t}\right) \;=\; \vv 0, \end{aligned} (2.6.6)

The middle step there is the commutation of t\partial_{t} with the spatial derivatives inside the curl, which is Clairaut's theorem again, as in Chapter 2.1 §2.2.

Now use what that last line tells us. Chapter 0.7 §2.3 showed that a curl-free field on a simply connected region is a gradient. So E+tA=ϕ\vv E+\partial_{t}\vv A=-\nabla\phi for some scalar ϕ\phi. The minus sign is the standard convention, chosen so that ϕ\phi reduces to the electrostatic potential in the static case. Rearranging that for E\vv E, and carrying the curl formula for B\vv B along beside it, we have both fields written in terms of potentials:

E  =  ϕ    At,B  =  ×A. \vv E \;=\; -\nabla\phi \;-\; \pdv{\vv A}{t}, \qquad\qquad \vv B \;=\; \nabla\times\vv A. (2.6.7)

Two of Maxwell's four equations have now been used up entirely. They are the reason the potentials exist, and once you work with ϕ\phi and A\vv A they hold identically, with nothing left to check. Hold on to that thought. Section 3.4 makes it the punchline of the whole section.

2.2 · Assembling AμA^{\mu}

We now have four functions in hand: one scalar ϕ\phi and the three components of A\vv A. The pattern of §1 suggests trying Aμ=(αϕ,A)A^{\mu}=(\alpha\phi,\vv A) for some constant α\alpha that dimensions will fix. It also suggests checking afterwards that the result transforms as a four-vector, rather than assuming that it does.

Dimensions first. In SI, A\vv A has units of Tm=Vsm1\mathrm{T\,m}=\mathrm{V\,s\,m^{-1}} and ϕ\phi has units of V\mathrm{V}. So ϕ/c\phi/c has the units of A\vv A, and α=1/c\alpha=1/c is the only choice that makes the four components commensurable. It is the same cc, appearing for the same reason, as in xμ=(ct,x)x^{\mu}=(ct,\vv x) and jμ=(cρ,J)j^{\mu}=(c\rho,\vv J). With that settled, define

  Aμ    (ϕc, A).   \boxed{\;A^{\mu} \;\equiv\; \left(\frac{\phi}{c},\ \vv A\right).\;} (2.6.8)

One warning about the letter. Chapters 2.2 to 2.5 used ϕ\phi for rapidity, and from here to the end of the book it means the electric scalar potential. That is an unfortunate clash of symbols, and the whole subject lives with it. Rapidity does not appear again in this chapter, so nothing is ambiguous on this page. When Part V has to write both in one line, it writes the rapidity as φ\varphi.

Is AμA^{\mu} a four-vector? Not by fiat. The check comes in §4. There we will find that in the Lorenz gauge AμA^{\mu} satisfies Aμ=μ0jμ\Box A^{\mu}=\mu_{0}j^{\mu}, where \Box is a scalar operator and jμj^{\mu} is a four-vector by §1. That forces AμA^{\mu} to be a four-vector too, since otherwise the equation could not hold in every frame. Until we have that argument in hand, treat (2.6.8) as the definition of an array of four numbers and nothing more.

2.3 · The field tensor

Look at (2.6.7) again. Every entry is one derivative of one potential component, minus another derivative of another potential component. In four-dimensional language there is exactly one object with that shape:

  Fμν    μAν    νAμ.   \boxed{\;F^{\mu\nu} \;\equiv\; \partial^{\mu}A^{\nu} \;-\; \partial^{\nu}A^{\mu}.\;} (2.6.9)

It is antisymmetric, Fνμ=FμνF^{\nu\mu}=-F^{\mu\nu}, and you can see that by inspection. Swapping μ\mu and ν\nu swaps the two terms. That is not a design choice anyone made. It is forced by the construction, and the callout at the end of §2 explains why the construction itself is forced.

The upper-index derivative is μ=ημνν\partial^{\mu}=\eta^{\mu\nu}\partial_{\nu}, and since ημν=diag(1,1,1,1)\eta^{\mu\nu}=\mathrm{diag}(1,-1,-1,-1) this flips the sign of the spatial components:

μ  =  (1ct, ). \partial^{\mu} \;=\; \left(\frac1c\pdv{}{t},\ -\nabla\right). (2.6.10)

That minus sign is where readers get hurt. Write it out once and keep it.

2.4 · The six components, one at a time

The three F0iF^{0i}. Put μ=0\mu=0, ν=i\nu=i into (2.6.9):

F0i  =  0AiiA0=  1cAit    (i) ⁣(ϕc)=  1c(Ait+iϕ)  =  1c(iϕAit)  =  Eic, \begin{aligned} F^{0i} \;&=\; \partial^{0}A^{i} - \partial^{i}A^{0}\\[3pt] &=\; \frac1c\pdv{A^{i}}{t} \;-\; \left(-\partial_{i}\right)\!\left(\frac{\phi}{c}\right)\\[3pt] &=\; \frac1c\left(\pdv{A^{i}}{t} + \partial_{i}\phi\right) \;=\; -\frac1c\left(-\partial_{i}\phi - \pdv{A^{i}}{t}\right) \;=\; -\frac{E^{i}}{c}, \end{aligned} (2.6.11)

the last step being (2.6.7) read backwards. So the first row of the matrix is E/c-\vv E/c, and the first column is +E/c+\vv E/c by antisymmetry.

The three FijF^{ij}. Both indices spatial, so both derivatives pick up the minus sign from (2.6.10):

Fij  =  iAjjAi  =  (iAjjAi). F^{ij} \;=\; \partial^{i}A^{j} - \partial^{j}A^{i} \;=\; -\left(\partial_{i}A^{j} - \partial_{j}A^{i}\right). (2.6.12)

Our goal now is to recognise that bracket as the curl, so that B\vv B appears. Chapter 0.7 §4.2 wrote the curl in components. In Levi-Civita notation, which is Chapter 2.4 §8.1 restricted to three indices, it reads Bk=ϵklmlAmB^{k}=\epsilon^{klm}\partial_{l}A^{m}. To get at the bracket we contract that with another epsilon and use the standard identity ϵijkϵklm=δilδjmδimδjl\epsilon^{ijk}\epsilon^{klm}=\delta^{il}\delta^{jm}-\delta^{im}\delta^{jl}:

ϵijkBk  =  ϵijkϵklmlAm  =  (δilδjmδimδjl)lAm  =  iAjjAi. \epsilon^{ijk}B^{k} \;=\; \epsilon^{ijk}\epsilon^{klm}\,\partial_{l}A^{m} \;=\; \big(\delta^{il}\delta^{jm}-\delta^{im}\delta^{jl}\big)\partial_{l}A^{m} \;=\; \partial_{i}A^{j}-\partial_{j}A^{i}. (2.6.13)

The right-hand side there is precisely the bracket sitting inside (2.6.12), so we can substitute it and read off the purely spatial part of the field tensor:

Fij  =  ϵijkBk. F^{ij} \;=\; -\,\epsilon^{ijk}B^{k}. (2.6.14)

Three entries, since the diagonal vanishes and the array is antisymmetric. Six in total. Assemble them:

Fμν  =  (0Ex/cEy/cEz/cEx/c0BzByEy/cBz0BxEz/cByBx0) F^{\mu\nu} \;=\; \begin{pmatrix} 0 & -E_{x}/c & -E_{y}/c & -E_{z}/c\\[2pt] E_{x}/c & 0 & -B_{z} & B_{y}\\[2pt] E_{y}/c & B_{z} & 0 & -B_{x}\\[2pt] E_{z}/c & -B_{y} & B_{x} & 0 \end{pmatrix} (2.6.15)

Check one entry against (2.6.14) to be sure the epsilon signs landed where they should: F12=ϵ12kBk=ϵ123B3=BzF^{12}=-\epsilon^{12k}B^{k}=-\epsilon^{123}B^{3}=-B_{z} ✓, and F13=ϵ132B2=+ByF^{13}=-\epsilon^{132}B^{2}=+B_{y} ✓ since ϵ132=1\epsilon^{132}=-1.

The promise of Chapter 2.4 §7.2, delivered

That chapter counted the components. An antisymmetric rank-2 tensor in four dimensions has 1243=6\tfrac12\cdot4\cdot3=\mathbf{6} of them. Six is too many for a four-vector, so such a tensor has to be its own kind of object. Chapter 2.4 then told you the six would turn out to be the three components of E\vv E and the three of B\vv B, and said nothing more.

Here they are. This is not an analogy, and it is not a coincidence of counting. The object μAννAμ\partial^{\mu}A^{\nu}-\partial^{\nu}A^{\mu} was built out of the potentials that Maxwell's own equations forced into existence, and its six independent entries came out as E/c\vv E/c and B\vv B with no room to choose otherwise.

E\vv E and B\vv B are not two fields. They are one tensor, sliced up by an observer. Which slices you call "electric" and which "magnetic" depends on your state of motion, exactly as which part of xμx^{\mu} you call "time" does. Section 5 computes the slicing rule and §6 shows you a wire where the whole of magnetism is nothing but a re-slicing.

2.5 · Both indices down, and the signs that flip

We will need FμνF_{\mu\nu} constantly, and this is exactly the step at which sign errors are manufactured. Lower with two metrics, per Chapter 2.4 §4.1:

Fμν  =  ημαηνβFαβ. F_{\mu\nu} \;=\; \eta_{\mu\alpha}\,\eta_{\nu\beta}\,F^{\alpha\beta}. (2.6.16)

Since η\eta is diagonal, each entry simply picks up the factor ημμηνν\eta_{\mu\mu}\eta_{\nu\nu}, with no sum implied. For a 0i0i entry that factor is (+1)(1)=1(+1)(-1)=-1. For an ijij entry it is (1)(1)=+1(-1)(-1)=+1. Applying those two factors to the entries we found above gives the rule we will be using for the rest of the chapter:

F0i  =  F0i  =  +Eic,Fij  =  +Fij  =  ϵijkBk. F_{0i} \;=\; -F^{0i} \;=\; +\frac{E^{i}}{c}, \qquad\qquad F_{ij} \;=\; +F^{ij} \;=\; -\epsilon^{ijk}B^{k}. (2.6.17)
⚠ Only the electric entries flip

Lowering both indices of FF reverses the sign of the E\vv E entries and leaves the B\vv B entries alone. One index is spatial in the first case and both are spatial in the second, so one minus sign survives in the first and two cancel in the second. If a calculation below ever comes out with E\vv E and B\vv B carrying the wrong relative sign, this is where to look first.

Fμν  =  (0Ex/cEy/cEz/cEx/c0BzByEy/cBz0BxEz/cByBx0) F_{\mu\nu} \;=\; \begin{pmatrix} 0 & E_{x}/c & E_{y}/c & E_{z}/c\\[2pt] -E_{x}/c & 0 & -B_{z} & B_{y}\\[2pt] -E_{y}/c & B_{z} & 0 & -B_{x}\\[2pt] -E_{z}/c & -B_{y} & B_{x} & 0 \end{pmatrix} (2.6.18)
⚠ Why this isn't obvious — antisymmetry is not a choice

It is tempting to read (2.6.9) as a clever guess. The tempting story goes like this. Somebody noticed that an antisymmetric tensor has six slots, and that E\vv E and B\vv B have six components between them, so they packed them in.

That reading gets the logic backwards. The correct order matters, because Chapter 6.3 runs on it.

Here is the chain, in order.

  • (1) B=0\nabla\cdot\vv B=0 and Faraday's law force the existence of potentials ϕ,A\phi,\vv A (§2.1).
  • (2) Those potentials are not unique, since you may add μχ\partial^{\mu}\chi to AμA^{\mu} without changing a field. So the physical fields have to be built from AμA^{\mu} in a way that kills that freedom.
  • (3) The only combination of one derivative and one potential that is annihilated by AμAμ+μχA^{\mu}\to A^{\mu}+\partial^{\mu}\chi is the antisymmetric one. The reason is that μνχ\partial^{\mu}\partial^{\nu}\chi is symmetric, so only the antisymmetric part of μAν\partial^{\mu}A^{\nu} can avoid it.
  • (4) Therefore FμνF^{\mu\nu} is antisymmetric, therefore it has six components, and therefore those components are E\vv E and B\vv B.

Gauge invariance is upstream of everything. The six-component count is a consequence of it, which makes Chapter 2.4's promise a prediction rather than a coincidence noticed after the fact.

Notice also what happens to the other half of μAν\partial^{\mu}A^{\nu}. The symmetric part μAν+νAμ\partial^{\mu}A^{\nu}+\partial^{\nu}A^{\mu}, with its ten components, is not gauge invariant and carries no physics here. It is nonetheless exactly the object that does carry physics in Chapter 3.6, where the potential is a symmetric hμνh_{\mu\nu} and the field is the linearised Riemann tensor. Same construction, different symmetry, and that difference is the difference between a spin-1 force and a spin-2 one.

Grind box — every entry of FμνF^{\mu\nu} from F=AAF=\partial A-\partial A, by hand

Nothing here is subtle. The point is to have watched all sixteen numbers appear once, so that the matrix above is a fact rather than a memory. Write Aμ=(ϕ/c,Ax,Ay,Az)A^{\mu}=(\phi/c,A_{x},A_{y},A_{z}) and μ=(1ct,x,y,z)\partial^{\mu}=(\tfrac1c\partial_{t},-\partial_{x},-\partial_{y},-\partial_{z}).

Diagonal. Fμμ=μAμμAμ=0F^{\mu\mu}=\partial^{\mu}A^{\mu}-\partial^{\mu}A^{\mu}=0 for each fixed μ\mu (no sum). Four zeros.

F01F^{01}.

F01=1cAxt(x)ϕc=1c(Axt+ϕx)=1cEx, F^{01} = \frac1c\pdv{A_{x}}{t} - \left(-\pdv{}{x}\right)\frac{\phi}{c} = \frac1c\left(\pdv{A_{x}}{t}+\pdv{\phi}{x}\right) = -\frac1c E_{x},

using Ex=xϕtAxE_{x}=-\partial_{x}\phi-\partial_{t}A_{x}. Identically F02=Ey/cF^{02}=-E_{y}/c and F03=Ez/cF^{03}=-E_{z}/c.

F12F^{12}. Both spatial, so both derivatives carry the minus:

F12=(x)Ay(y)Ax=(AyxAxy)=Bz, F^{12} = \left(-\pdv{}{x}\right)A_{y} - \left(-\pdv{}{y}\right)A_{x} = -\left(\pdv{A_{y}}{x}-\pdv{A_{x}}{y}\right) = -B_{z},

since Bz=xAyyAxB_{z}=\partial_{x}A_{y}-\partial_{y}A_{x} from Chapter 0.7 §4.2. Cycling xyzxx\to y\to z\to x gives F23=BxF^{23}=-B_{x} and F31=ByF^{31}=-B_{y}, and the last of these is F13=+ByF^{13}=+B_{y}.

Lower half. By antisymmetry, F10=+Ex/cF^{10}=+E_{x}/c, F21=+BzF^{21}=+B_{z}, and so on. Sixteen numbers, six of them independent, exactly as (2.6.15) says.

Sanity check on units. Every entry of FμνF^{\mu\nu} must have the same dimensions, or the object is not a tensor. BB is in tesla. And E/cE/c is Vm1/(ms1)=Vsm2=T\mathrm{V\,m^{-1}}/(\mathrm{m\,s^{-1}}) = \mathrm{V\,s\,m^{-2}} = \mathrm{T} ✓. So the factor of cc in E/cE/c is not decoration. Without it the six components could not be six components of one thing.

In plain terms 2.6.2

Two of the four equations are spent before the real work starts, and what they buy is the existence of the potentials: one number and one three-part object, out of which both fields are then built. Look at how they are built and every entry is one derivative of one potential component minus a different derivative of a different one. In four dimensions exactly one object has that shape, and it reverses its sign when its two labels are exchanged.

The antisymmetry is forced rather than noticed. Potentials are not unique, since a whole function's worth of freedom may be added without altering any field, and the only way to build something blind to that freedom from one derivative and one potential is to keep the part reversing sign, the leftover being symmetric and impossible to dodge otherwise. Six independent entries follow, and that count was performed a chapter ago with the answer promised and withheld.

Here is the answer. The six are the three electric components and the three magnetic ones, arriving by construction with no room to choose otherwise. They are not two fields that happen to fit inside one container. They are one object sliced by an observer, and which slices somebody calls electric depends on that observer's motion in exactly the way that which part of a separation between events somebody calls time depends on it.

3 · Four equations become two

This is the centrepiece of the chapter, and nothing in it will be waved through. We are going to write down two tensor equations and then expand every single component of both, so that you can see Maxwell's familiar four emerge one at a time.

3.1 · The inhomogeneous half

What can we build from FμνF^{\mu\nu}, one derivative, and the current? The index structure almost writes the answer for us. Contracting the derivative into the first slot gives μFμν\partial_{\mu}F^{\mu\nu}, which has one free upper index. So does jνj^{\nu}. Two objects with matching index structure can be set equal, and there is only one way to do it, with a single dimensionful constant left undetermined:

  μFμν  =  μ0jν.   \boxed{\;\partial_{\mu}F^{\mu\nu} \;=\; \mu_{0}\,j^{\nu}.\;} (2.6.19)

We now show that (2.6.19) is Gauss's law together with the Ampère–Maxwell law, and that the constant is μ0\mu_{0}. There are four components in it, and we will do all four.

Component ν=0\nu=0. The sum over μ\mu runs over 0,1,2,30,1,2,3, but F00=0F^{00}=0, so only the spatial terms survive:

μFμ0  =  0F00=0  +  iFi0  =  i ⁣(Eic)  =  1cE, \partial_{\mu}F^{\mu 0} \;=\; \underbrace{\partial_{0}F^{00}}_{=\,0} \;+\; \partial_{i}F^{i0} \;=\; \partial_{i}\!\left(\frac{E^{i}}{c}\right) \;=\; \frac{1}{c}\,\nabla\cdot\vv E, (2.6.20)

reading Fi0=+Ei/cF^{i0}=+E^{i}/c off the first column of (2.6.15). The right-hand side is μ0j0=μ0cρ\mu_{0}j^{0}=\mu_{0}c\rho. Equate and multiply by cc:

E  =  μ0c2ρ  =  ρϵ0, \nabla\cdot\vv E \;=\; \mu_{0}c^{2}\rho \;=\; \frac{\rho}{\epsilon_{0}}, (2.6.21)

The last step there uses μ0ϵ0=1/c2\mu_{0}\epsilon_{0}=1/c^{2}. That is not an extra input. It is Chapter 2.1's result c=1/μ0ϵ0c=1/\sqrt{\mu_{0}\epsilon_{0}} rearranged.

That is Gauss's law. It also fixes the undetermined constant in (2.6.19) to be μ0\mu_{0}, because any other choice would give the wrong Coulomb force.

Component ν=j\nu=j (spatial). Now both the time term and the space terms contribute:

μFμj  =  0F0j  +  iFij  =  1ct ⁣(Ejc)  +  i(ϵijkBk). \partial_{\mu}F^{\mu j} \;=\; \partial_{0}F^{0j} \;+\; \partial_{i}F^{ij} \;=\; \frac1c\pdv{}{t}\!\left(-\frac{E^{j}}{c}\right) \;+\; \partial_{i}\big(-\epsilon^{ijk}B^{k}\big). (2.6.22)

The first piece is already in ordinary language. The second piece is not, so let's handle it. We want to see a curl there, so swap the first two indices of the epsilon, which costs one sign, and compare with the component form of the curl:

ϵijkiBk  =  +ϵjikiBk  =  (×B)j. -\epsilon^{ijk}\partial_{i}B^{k} \;=\; +\epsilon^{jik}\partial_{i}B^{k} \;=\; \big(\nabla\times\vv B\big)^{j}. (2.6.23)

Putting that back into (2.6.22), and setting the whole thing equal to μ0jj=μ0Jj\mu_{0}j^{j}=\mu_{0}J^{j}, the component equation reads 1c2tEj+(×B)j=μ0Jj-\frac{1}{c^{2}}\partial_{t}E^{j}+(\nabla\times\vv B)^{j}=\mu_{0}J^{j}. Collecting the three values of jj back into vectors and moving the time derivative to the right,

×B  =  μ0J  +  1c2Et  =  μ0J+μ0ϵ0Et. \nabla\times\vv B \;=\; \mu_{0}\vv J \;+\; \frac{1}{c^{2}}\pdv{\vv E}{t} \;=\; \mu_{0}\vv J + \mu_{0}\epsilon_{0}\pdv{\vv E}{t}. (2.6.24)

That is the Ampère–Maxwell law, displacement current and all. One tensor equation, four components, two of Maxwell's four laws.

Look at where the displacement term came from. Maxwell had to add it by hand, and Chapter 0.7's Problem 4 showed that charge conservation forces it. Here it is not added at all. It is the μ=0\mu=0 term of a sum that has to run over all four values of μ\mu, because the index is contracted and a contracted index sums over everything.

3.2 · The homogeneous half

Two of Maxwell's equations are left: B=0\nabla\cdot\vv B=0 and Faraday's law. Neither contains a source, so the covariant version must have zero on the right. The object that works is the totally antisymmetrised derivative of FF, written out as a cyclic sum:

  λFμν  +  μFνλ  +  νFλμ  =  0.   \boxed{\;\partial_{\lambda}F_{\mu\nu} \;+\; \partial_{\mu}F_{\nu\lambda} \;+\; \partial_{\nu}F_{\lambda\mu} \;=\; 0.\;} (2.6.25)

Before expanding, notice two structural facts, because they save all the work.

It is totally antisymmetric in λμν\lambda\mu\nu. Swap any two of the three indices and the three terms permute into one another with an overall sign. Check λμ\lambda\leftrightarrow\mu for yourself: the first term becomes μFλν=μFνλ\partial_{\mu}F_{\lambda\nu}=-\partial_{\mu}F_{\nu\lambda}, which is minus the second, and so on around. So the left-hand side is an antisymmetric rank-3 object.

Therefore there are only four independent components. By the count of Chapter 2.4 §7.2, generalised, a totally antisymmetric rank-3 array in four dimensions has (43)=4\binom{4}{3}=4 independent entries. There is one entry for each way of choosing three distinct indices from {0,1,2,3}\{0,1,2,3\}, namely (1,2,3)(1,2,3), (0,1,2)(0,1,2), (0,2,3)(0,2,3) and (0,3,1)(0,3,1). Any triple with a repeated index vanishes identically.

Four components, and four is exactly the number of scalar equations we need. B=0\nabla\cdot\vv B=0 is one of them, and Faraday's law is the other three. The bookkeeping fits before we compute anything at all.

Triple (λμν)=(1,2,3)(\lambda\mu\nu)=(1,2,3). All three indices spatial. Using Fij=ϵijkBkF_{ij}=-\epsilon^{ijk}B^{k} from (2.6.17), so that F23=BxF_{23}=-B_{x}, F31=ByF_{31}=-B_{y}, F12=BzF_{12}=-B_{z}:

1F23+2F31+3F12  =  (xBx+yBy+zBz)  =  B  =  0. \partial_{1}F_{23} + \partial_{2}F_{31} + \partial_{3}F_{12} \;=\; -\big(\partial_{x}B_{x}+\partial_{y}B_{y}+\partial_{z}B_{z}\big) \;=\; -\nabla\cdot\vv B \;=\; 0. (2.6.26)

No magnetic monopoles.

Triple (λμν)=(0,1,2)(\lambda\mu\nu)=(0,1,2). One time index. Now F12=BzF_{12}=-B_{z}, F20=F02=Ey/cF_{20}=-F_{02}=-E_{y}/c, F01=+Ex/cF_{01}=+E_{x}/c, and 0=1ct\partial_{0}=\tfrac1c\partial_{t}:

0F12+1F20+2F01=1ct(Bz)+x(Eyc)+y(Exc)=1c[Bzt+(EyxExy)(×E)z]  =  0. \begin{aligned} \partial_{0}F_{12} + \partial_{1}F_{20} + \partial_{2}F_{01} &= \frac1c\pdv{}{t}\big({-B_{z}}\big) + \pdv{}{x}\left(-\frac{E_{y}}{c}\right) + \pdv{}{y}\left(\frac{E_{x}}{c}\right)\\[4pt] &= -\frac1c\left[\pdv{B_{z}}{t} + \underbrace{\left(\pdv{E_{y}}{x}-\pdv{E_{x}}{y}\right)}_{\textstyle (\nabla\times\vv E)_{z}}\right] \;=\; 0. \end{aligned} (2.6.27)

Multiply by c-c and you have the zz-component of Faraday's law, (×E)z=tBz(\nabla\times\vv E)_{z}=-\partial_{t}B_{z}. The triples (0,2,3)(0,2,3) and (0,3,1)(0,3,1) give the xx- and yy-components by the identical computation with the labels cycled. The grind box does them explicitly, so that "by symmetry" is not doing any hidden work here.

Grind box — the other two Faraday components, written out

Triple (0,2,3)(0,2,3). The entries needed are F23=BxF_{23}=-B_{x}, F30=F03=Ez/cF_{30}=-F_{03}=-E_{z}/c, F02=+Ey/cF_{02}=+E_{y}/c:

0F23+2F30+3F02=1ct(Bx)+y(Ezc)+z(Eyc)=1c[Bxt+(EzyEyz)]=1c[Bxt+(×E)x]. \begin{aligned} \partial_{0}F_{23}+\partial_{2}F_{30}+\partial_{3}F_{02} &= \frac1c\pdv{}{t}(-B_{x}) + \pdv{}{y}\left(-\frac{E_{z}}{c}\right) + \pdv{}{z}\left(\frac{E_{y}}{c}\right)\\[3pt] &= -\frac1c\left[\pdv{B_{x}}{t} + \left(\pdv{E_{z}}{y}-\pdv{E_{y}}{z}\right)\right] = -\frac1c\left[\pdv{B_{x}}{t}+(\nabla\times\vv E)_{x}\right]. \end{aligned}

Triple (0,3,1)(0,3,1). Entries F31=ByF_{31}=-B_{y}, F10=Ex/cF_{10}=-E_{x}/c, F03=+Ez/cF_{03}=+E_{z}/c:

0F31+3F10+1F03=1ct(By)+z(Exc)+x(Ezc)=1c[Byt+(ExzEzx)]=1c[Byt+(×E)y]. \begin{aligned} \partial_{0}F_{31}+\partial_{3}F_{10}+\partial_{1}F_{03} &= \frac1c\pdv{}{t}(-B_{y}) + \pdv{}{z}\left(-\frac{E_{x}}{c}\right) + \pdv{}{x}\left(\frac{E_{z}}{c}\right)\\[3pt] &= -\frac1c\left[\pdv{B_{y}}{t} + \left(\pdv{E_{x}}{z}-\pdv{E_{z}}{x}\right)\right] = -\frac1c\left[\pdv{B_{y}}{t}+(\nabla\times\vv E)_{y}\right]. \end{aligned}

Setting each to zero gives the xx- and yy-components of Faraday's law. Together with (2.6.27) and (2.6.26) that is all four independent components of (2.6.25), and all four of Maxwell's homogeneous scalar equations. Nothing is left over and nothing is missing.

Why the cyclic pattern (0,1,2),(0,2,3),(0,3,1)(0,1,2),(0,2,3),(0,3,1) and not (0,1,3)(0,1,3)? Because (0,3,1)(0,3,1) and (0,1,3)(0,1,3) are the same component up to a sign, and picking the cyclic representative keeps every equation's sign uniform. Choosing the other one is not an error. It gives the same equation multiplied by 1-1.

3.3 · The scoreboard

Covariant statementComponentTraditional name
μFμν=μ0jν\partial_{\mu}F^{\mu\nu}=\mu_{0}j^{\nu}ν=0\nu=0Gauss's law
ν=1,2,3\nu=1,2,3Ampère–Maxwell
[λFμν]=0\partial_{[\lambda}F_{\mu\nu]}=0(1,2,3)(1,2,3)B=0\nabla\cdot\vv B=0
(0,ij)(0,ij)Faraday's law

Four equations, eight scalar equations in the traditional bookkeeping, compressed into two lines.

There is a second payoff, and it is the larger one. Both lines are equations between tensors of the same type, so the invariance theorem of Chapter 2.4 §6 applies to each of them. Each holds in every inertial frame the moment it holds in one. No chain-rule computation, no cross terms, none of the carnage of Chapter 2.1 §4.2. That is the whole purpose of the tensor apparatus, and this is the first place the book collects on it.

3.4 · Two of Maxwell's equations are not physics

Now for the structural point of the section. It is easy to prove and it deserves emphasis out of all proportion to that ease.

Our goal is to see what the homogeneous half says once the potentials are put back in. So substitute the definition (2.6.9) into (2.6.25), with all indices down so that Fμν=μAννAμF_{\mu\nu}=\partial_{\mu}A_{\nu}-\partial_{\nu}A_{\mu}:

λFμν+μFνλ+νFλμ  =  λμAνλνAμ+  μνAλμλAν+  νλAμνμAλ. \begin{aligned} \partial_{\lambda}F_{\mu\nu} + \partial_{\mu}F_{\nu\lambda} + \partial_{\nu}F_{\lambda\mu} \;=\;& \partial_{\lambda}\partial_{\mu}A_{\nu} - \partial_{\lambda}\partial_{\nu}A_{\mu}\\ +\;& \partial_{\mu}\partial_{\nu}A_{\lambda} - \partial_{\mu}\partial_{\lambda}A_{\nu}\\ +\;& \partial_{\nu}\partial_{\lambda}A_{\mu} - \partial_{\nu}\partial_{\mu}A_{\lambda}. \end{aligned} (2.6.28)

Six terms. Let's pair them off. The first cancels the fourth, the second cancels the fifth, and the third cancels the sixth. Each pair differs only by the order of two partial derivatives, and those commute by Clairaut's theorem. So the sum is identically zero.

Half of Maxwell's equations are bookkeeping

Given that E\vv E and B\vv B come from potentials, B=0\nabla\cdot\vv B=0 and Faraday's law cannot fail. They carry no information about how electromagnetism works. They are the statement that mixed partial derivatives commute, wearing a hat.

You have met both of them before, in Chapter 0.7 §7. There, (×A)=0\nabla\cdot(\nabla\times\vv A)=0 was six terms cancelling in pairs by Clairaut, and that identity is exactly (2.6.26). Alongside it, ×ϕ=0\nabla\times\nabla\phi=\vv 0 was the same cancellation with fewer terms, and that identity is exactly the content of Faraday's law once E\vv E is written in terms of ϕ\phi and A\vv A.

So the six-term cancellation in (2.6.28) is both of Chapter 0.7's identities at once. In four dimensions they are one identity.

The physical content of electromagnetism therefore sits entirely in (2.6.19), which says that charges and currents source the field. The other two equations are the price of using potentials, and they are free.

⚑ Quoted forward — this is d2=0\dd^{2}=0

Chapter 0.7 §7.4 lined up gradient, curl and divergence and observed that composing two consecutive arrows gives zero, both times, and promised that in the language of differential forms both statements are the single equation dd=0\dd\circ\dd=0. Chapter 3.5 supplies that language. In it, AμA_{\mu} is a one-form AA, the definition F=dAF=\dd A is (2.6.9), and (2.6.25) collapses to the single symbol dF=ddA=0\dd F=\dd\dd A=0. The inhomogeneous half becomes d ⁣ ⁣F=μ0j\dd\!\star\!F=\mu_{0}\star j.

The asymmetry you have just noticed between the two halves, that one is an identity and the other is a field equation, then becomes the visible asymmetry between dF\dd F and d ⁣ ⁣F\dd\!\star\!F. We flag this as a promise rather than a derivation. You now know what it is a promise about.

In plain terms 2.6.3

Written in the new language four equations become two, and the compression is not typographical. One of them carries a source and admits a single sensible arrangement of labels; expanding its four entries returns the law relating field to charge and the one relating circulating magnetic field to current. The term Maxwell inserted by hand, whose absence made the set inconsistent with charge conservation, is not inserted at all: it is one entry of a sum obliged to run over all four values because the label is contracted.

The other equation is not physics. Write the field tensor in terms of the potentials and six terms cancel in pairs, each pair differing only in the order of two derivatives. Given potentials, the absence of magnetic sources and the law of induction cannot fail, and carry no information about how electromagnetism works. They are the toolkit's two identities, the vanishing swirl of a gradient and the vanishing outflow of a curl, which in four dimensions are one.

Both lines relate objects of the same type, so each holds in every frame the moment it holds in one, with no chain rule, no cross terms and none of the carnage that transforming the wave equation by hand produced earlier. The content of the subject is thereby one sentence: charges and currents make fields. The rest is the price of using potentials, and the price is nothing.

4 · Gauge invariance

4.1 · The freedom

The potentials are not unique, and Chapter 0.7 §7.3 already said why. Adding a gradient to A\vv A leaves ×A\nabla\times\vv A alone, because the curl of a gradient vanishes. In four-dimensional language the statement is cleaner, and it covers ϕ\phi as well. Let χ(x)\chi(x) be any smooth scalar function of spacetime, and transform the four-potential by

Aμ    A~μ  =  Aμ+μχ. A^{\mu} \;\longrightarrow\; \tilde A^{\mu} \;=\; A^{\mu} + \partial^{\mu}\chi. (2.6.29)

The question is what that does to the fields, so our next step is to build the field tensor out of the new potential and see what survives. Feed A~μ\tilde A^{\mu} straight into the definition (2.6.9):

F~μν  =  μA~ννA~μ  =  Fμν+μνχνμχ=0  =  Fμν. \tilde F^{\mu\nu} \;=\; \partial^{\mu}\tilde A^{\nu} - \partial^{\nu}\tilde A^{\mu} \;=\; F^{\mu\nu} + \underbrace{\partial^{\mu}\partial^{\nu}\chi - \partial^{\nu}\partial^{\mu}\chi}_{=\,0} \;=\; F^{\mu\nu}. (2.6.30)

That is Clairaut again, the third appearance in as many sections, and the last. The fields are untouched. Since the fields are what exert forces, no experiment can distinguish AμA^{\mu} from A~μ\tilde A^{\mu}.

It is worth seeing what this looks like in the older notation, so let's unpack (2.6.29) into three-vector language. The one thing to watch is the minus sign in μ=(1ct,)\partial^{\mu}=(\tfrac1c\partial_{t},-\nabla):

ϕ~c=ϕc+1cχt,A~=Aχϕ~=ϕ+χt,A~=Aχ. \frac{\tilde\phi}{c} = \frac{\phi}{c} + \frac1c\pdv{\chi}{t}, \qquad \tilde{\vv A} = \vv A - \nabla\chi \qquad\Longrightarrow\qquad \tilde\phi = \phi + \pdv{\chi}{t}, \quad \tilde{\vv A} = \vv A - \nabla\chi. (2.6.31)

You may have seen the familiar textbook form, which has A+χ\vv A+\nabla\chi and ϕtχ\phi-\partial_{t}\chi. That is the same transformation with χχ\chi\to-\chi. Nothing whatever depends on which sign convention you adopt.

You have seen this freedom before

Chapter 1.2's Problem 4 established that adding a total time derivative to a Lagrangian, LL+dFdtL\to L+\dv{F}{t}, changes the action only by a boundary term and therefore leaves the equations of motion untouched. Chapter 1.4 §1.2 then built the definition of a symmetry around exactly that loophole. Gauge freedom is that freedom.

Section 9 makes the correspondence exact. The interaction term in the electromagnetic Lagrangian is jμAμ-j_{\mu}A^{\mu}, and under (2.6.29) it changes by

jμμχ  =  μ(jμχ)+χμjμ=0  =  μ(jμχ), -j_{\mu}\partial^{\mu}\chi \;=\; -\partial^{\mu}\big(j_{\mu}\chi\big) + \chi\,\underbrace{\partial^{\mu}j_{\mu}}_{=\,0} \;=\; -\partial^{\mu}\big(j_{\mu}\chi\big),

That is a pure four-divergence, which is the field-theory version of dFdt\dv{F}{t}. It integrates to a boundary term by the divergence theorem, so it changes nothing.

Note what had to be true for that to work: charge conservation. Gauge invariance of the action and conservation of charge are the same statement seen from two sides. That is the first hint of the structure Chapter 6.3 turns into a machine.

4.2 · Choosing a gauge, and recovering 2.1's wave equation

Freedom is a nuisance if you want to solve for something, so we spend it. Impose the Lorenz condition

μAμ  =  01c2ϕt+A  =  0, \partial_{\mu}A^{\mu} \;=\; 0 \qquad\Longleftrightarrow\qquad \frac{1}{c^{2}}\pdv{\phi}{t} + \nabla\cdot\vv A \;=\; 0, (2.6.32)

The equivalence there follows from μAμ=1ct(ϕ/c)+A\partial_{\mu}A^{\mu}=\tfrac1c\partial_{t}(\phi/c)+\nabla\cdot\vv A.

Notice that (2.6.32) is itself a scalar equation, being a contraction, so imposing it in one frame imposes it in all of them. That is not true of the other common choice, the Coulomb gauge A=0\nabla\cdot\vv A=0. That one singles out a frame, which makes it useless for our purposes here. It is convenient elsewhere, and §10 uses it once, deliberately.

It is always attainable. Suppose you are handed potentials with μAμ=f(x)0\partial_{\mu}A^{\mu}=f(x)\neq0. Gauge-transform by some χ\chi and compute the new divergence:

μA~μ  =  μAμ+μμχ  =  f+χ,μμ=1c22t22. \partial_{\mu}\tilde A^{\mu} \;=\; \partial_{\mu}A^{\mu} + \partial_{\mu}\partial^{\mu}\chi \;=\; f + \Box\chi, \qquad \Box\equiv\partial_{\mu}\partial^{\mu} = \frac{1}{c^{2}}\pdv{^{2}}{t^{2}} - \nabla^{2}. (2.6.33)

So the Lorenz condition holds for the new potentials precisely when χ=f\Box\chi=-f. That is an inhomogeneous wave equation for χ\chi with a known source.

⚑ Quoted — solvability of the wave equation

For any reasonable source ff the equation χ=f\Box\chi=-f has solutions. This is the standard existence theorem for the wave operator (the solution is built from the retarded Green's function, χ(x,t)=14π ⁣ ⁣f(x,txx/c)/xxd3x\chi(\vv x,t)=\tfrac{1}{4\pi}\!\int\! f(\vv x',t-\abs{\vv x-\vv x'}/c)/\abs{\vv x-\vv x'}\,\dd^{3}x'), and we quote it rather than prove it. The proof is a chapter of analysis rather than of physics, and Chapter 5.4 constructs the Green's functions properly when it needs the propagator. What matters here is only that a solution exists, so the Lorenz gauge is always available.

Note also that the solution is not unique. Any χ0\chi_{0} with χ0=0\Box\chi_{0}=0 may be added to it, and that leftover freedom is called the residual gauge. It is precisely why the photon has two polarisation states rather than four, and Chapter 5.8 spends real effort on it.

Now the payoff. Substitute the definition of FF into the field equation (2.6.19) and impose (2.6.32):

μ0jν  =  μFμν  =  μ(μAννAμ)  =  (μμ)Aν    ν(μAμ=0)  =  Aν. \begin{aligned} \mu_{0}j^{\nu} \;=\; \partial_{\mu}F^{\mu\nu} &\;=\; \partial_{\mu}\big(\partial^{\mu}A^{\nu} - \partial^{\nu}A^{\mu}\big)\\[3pt] &\;=\; \big(\partial_{\mu}\partial^{\mu}\big)A^{\nu} \;-\; \partial^{\nu}\big(\underbrace{\partial_{\mu}A^{\mu}}_{=\,0}\big) \;=\; \Box A^{\nu}. \end{aligned} (2.6.34)

The middle step moved μ\partial_{\mu} past ν\partial^{\nu}, which is legitimate because partial derivatives commute. That let the Lorenz condition kill the second term outright. So all of electromagnetism, in the Lorenz gauge, has come down to one line:

  Aμ  =  μ0jμ.   \boxed{\;\Box A^{\mu} \;=\; \mu_{0}\,j^{\mu}.\;} (2.6.35)

Four uncoupled wave equations, one per component, each with its own source. In vacuum (jμ=0j^{\mu}=0) this is Aμ=0\Box A^{\mu}=0, that is,

1c22Aμt2  =  2Aμ, \frac{1}{c^{2}}\pdv{^{2}A^{\mu}}{t^{2}} \;=\; \nabla^{2}A^{\mu}, (2.6.36)

That is the wave equation of Chapter 2.1 §2. It was derived there by taking the curl of Faraday's law and grinding through ×(×E)=(E)2E\nabla\times(\nabla\times\vv E)=\nabla(\nabla\cdot\vv E)-\nabla^{2}\vv E. Same equation, same cc, obtained here in three lines because the bookkeeping was done first.

The crisis of 2.1, dissolved

Chapter 2.1 §3 asked the question that broke nineteenth-century physics. The wave equation contains a speed cc. Every other wave equation's speed is measured relative to a medium. So what is cc measured relative to?

Look at (2.6.35) and you can see why the question has no answer. The operator =μμ\Box=\partial_{\mu}\partial^{\mu} is a scalar, with two indices contracted, so by Chapter 2.4 §5.1 it takes the same form in every inertial frame. And cc enters it only through ημν\eta_{\mu\nu}, which is to say only through the metric of spacetime itself.

So the constant cc in Maxwell's equations is a property of the geometry in which the fields live, rather than of any substance they live in. Asking what it is measured relative to is like asking what the number π\pi is measured relative to.

Branch (A) of 2.1's fork is now worse off than merely unsupported by Michelson and Morley. It is structurally unavailable. Maxwell's equations do pick out a preferred speed. They never picked out a preferred frame, and only the assumption that t=tt'=t made those two look like the same claim.

Electromagnetism was never the theory that needed fixing. It was relativistic from birth. What had to give way was mechanics, with its Galilean addition of velocities, and Chapter 2.5 has already paid that bill.

⚠ Where gauge freedom is going

In this chapter gauge invariance is a convenience. It lets you choose (2.6.32) and turn a coupled mess into four wave equations. That is a wild understatement of its importance.

Chapter 6.3 reverses the logic. Instead of noticing that electromagnetism happens to have this redundancy, it demands that the phase of a charged quantum field be adjustable independently at every point of spacetime. That is a symmetry with a whole function's worth of parameters, and it is a heavy demand. Chapter 6.3 then finds that the demand cannot be met unless a vector field AμA^{\mu} exists, transforming exactly as (2.6.29) and coupling exactly as jμAμ-j_{\mu}A^{\mu}. The entire content of this chapter comes back out as a consequence.

Repeat the trick with a non-commuting symmetry group and out come the weak and strong interactions (Chapter 6.4). Gauge freedom stops being a convenience and becomes the generating principle of every known force except gravity. And gravity turns out to be the same trick applied to the Lorentz group.

⚠ Why this isn't obvious — the potentials know more than the fields

The natural reading of §4.1 is that AμA^{\mu} is a convenient fiction and FμνF^{\mu\nu} is the real thing. The argument runs like this. The fields are gauge invariant, the potentials are not, so only the fields can be physical. That reading is almost right, and the exception to it is one of the most instructive facts in physics.

Chapter 0.7 §2.4 built a vector field on the punctured plane that is curl-free everywhere and yet has circulation 2π2\pi around every loop enclosing the puncture. Its potential exists locally and is the polar angle, which is multivalued. Now read that as electromagnetism. Take a long solenoid with B0\vv B\neq\vv 0 inside and B=0\vv B=\vv 0 everywhere outside. Outside, A\vv A is curl-free. It still cannot be set to zero, because Adr\oint\vv A\cdot\dd\vv r around a loop enclosing the solenoid equals the enclosed flux, by Stokes' theorem, and that is not zero. The region outside the solenoid is not simply connected, and Chapter 0.7's counterexample is exactly this situation.

Classically nothing follows, since no charge outside ever feels a force. Quantum mechanically something does follow. Chapter 5.6 shows that a charged particle's amplitude picks up a phase exp ⁣(iq ⁣ ⁣Adr)\exp\!\big(\tfrac{\ii q}{\hbar}\!\int\!\vv A\cdot\dd\vv r\big) along its path. So two paths passing on opposite sides of the solenoid differ in phase by qΦ/q\Phi/\hbar, and the interference pattern shifts, even though the particle never enters a region where the field is nonzero.

This is the Aharonov–Bohm effect. It was measured in 1960, and it settles the question. The potential carries information that the field does not, and that information is topological. What is physical is neither AμA^{\mu}, which is gauge dependent, nor FμνF^{\mu\nu}, which is too little. It is the gauge-invariant loop integral Aμdxμ\oint A_{\mu}\dd x^{\mu}, called the holonomy. Chapter 6.3 builds the whole of gauge theory on that object, and Chapter 6.5 finds that in the strong interaction it is essentially all there is.

Familiar ground — gauge freedom is the reference category

Fit a model with a categorical predictor of several levels and you have already met (2.6.29). Write the fitted value for group gg as y^=μ+αg\hat y = \mu + \alpha_{g}. The parameters are not identified, because replacing

αg    αg+λ,μ    μλ \alpha_{g} \;\longrightarrow\; \alpha_{g} + \lambda, \qquad \mu \;\longrightarrow\; \mu - \lambda

for any constant λ\lambda leaves every fitted value, every residual, every contrast and the entire likelihood exactly as they were. Software does not announce this. It silently imposes a constraint of its own, either dropping a reference level so that α1=0\alpha_{1}=0 or using sum-to-zero coding so that gαg=0\sum_{g}\alpha_{g}=0, and then prints whichever coefficients that constraint produces. A coefficient reported without its coding scheme is not a quantity.

Every clause of that paragraph is this section. The potential AμA^{\mu} is the parameter, (2.6.29) is the reparametrisation, and (2.6.30) is the statement that nothing measurable moves. The Lorenz condition (2.6.32) is a coding scheme, adopted in §4.2 because it makes the equations uncouple, exactly as sum-to-zero coding is adopted because it makes a table symmetric.

The test for what is real is the same test in both subjects. A quantity is physical if and only if the reparametrisation leaves it alone, which is estimability under another name. In both cases what survives are the contrasts.

Three differences, and the third is why Part VI exists.

  • First, the redundancy here is a whole function's worth. Here χ\chi is arbitrary at every point of spacetime, rather than being one number for the whole model, and §4.2's counting is what that costs.
  • Second, the potentials are not pure bookkeeping. Chapter 0.7's solenoid already showed that AμA^{\mu} carries something FμνF^{\mu\nu} does not, and a design matrix has no counterpart to that.
  • Third, and this is the one that matters, in statistics the redundancy is a nuisance and nothing whatever is lost by constraining it away. Here it turns out to be generative.

Take that third difference seriously for a moment. Demand that the freedom hold separately at each point rather than once for the whole universe, and a field is forced into existence to enforce it. Every force in the Standard Model is that demand made of a different group, and Chapter 6.3 is the one place in this book where a choice of reference category writes down a law of nature.

In plain terms 2.6.4

A whole function's worth of freedom sits inside the potentials and no experiment can see it. Add the four-dimensional gradient of anything smooth and both fields come out unaltered, for the reason that has now done this work three times, that two derivatives are indifferent to their order. Since the fields are what push charges about, potentials differing that way describe the same world. Freedom is an obstacle when you want to solve for something, so it gets spent, and one condition spends enough.

The condition chosen is itself a contraction, so imposing it in one frame imposes it in all, which the other common choice cannot claim. With it in force the four components stop talking to one another and each satisfies a wave equation with its own source. Remove the sources and what remains is the equation that opened this part, obtained in three lines because the bookkeeping was done first instead of last.

That dissolves the question which broke nineteenth-century physics. The operator in that equation is a contraction, so it reads the same for everybody, and the speed enters it only through the array of signs, which is to say through the geometry rather than through any substance filling space. Maxwell's equations do single out a speed. They never singled out a frame, and only the assumption of one universal clock made those two look like the same claim.

5 · E\vv E and B\vv B mix

If E\vv E and B\vv B really are the components of one tensor, then a boost must shuffle them into each other, and Chapter 2.4 tells us exactly how. For a (2,0)(2,0) tensor the law is

Fμν  =  ΛμαΛνβFαβ, F'^{\mu\nu} \;=\; \Lambda^{\mu}{}_{\alpha}\,\Lambda^{\nu}{}_{\beta}\,F^{\alpha\beta}, (2.6.37)

In matrix notation, with FF the array of (2.6.15), that reads F=ΛFΛTF'=\Lambda F\Lambda^{\mathsf T}. There is one Λ\Lambda acting on the rows and one acting on the columns, which is what two upper indices buys you. To make it concrete we take the standard boost of Chapter 2.4 §2.1, with SS' moving at +v+v along xx:

Λμν  =  (γγβ00γβγ0000100001),β=vc. \Lambda^{\mu}{}_{\nu} \;=\; \begin{pmatrix} \gamma & -\gamma\beta & 0 & 0\\ -\gamma\beta & \gamma & 0 & 0\\ 0&0&1&0\\ 0&0&0&1\end{pmatrix}, \qquad \beta=\frac{v}{c}. (2.6.38)

Carrying out the two matrix multiplications, which the grind box does entry by entry, and then reading the primed fields off the result, gives the transformation rules in full.

How E\vv E and B\vv B transform under a boost along xx

Splitting each field into the component along the boost (\parallel) and the two perpendicular to it (\perp):

E=E,B=B,Ey=γ(EyvBz),By=γ(By+vc2Ez),Ez=γ(Ez+vBy),Bz=γ(Bzvc2Ey). \begin{aligned} E'_{\parallel} &= E_{\parallel}, &\qquad B'_{\parallel} &= B_{\parallel},\\[6pt] E'_{y} &= \gamma\big(E_{y}-vB_{z}\big), &\qquad B'_{y} &= \gamma\Big(B_{y}+\frac{v}{c^{2}}E_{z}\Big),\\[4pt] E'_{z} &= \gamma\big(E_{z}+vB_{y}\big), &\qquad B'_{z} &= \gamma\Big(B_{z}-\frac{v}{c^{2}}E_{y}\Big). \end{aligned}

Or, compactly, with v\vv v the boost velocity:

E=γ(E+v×B),B=γ(B1c2v×E). \vv E'_{\perp} = \gamma\big(\vv E + \vv v\times\vv B\big)_{\perp}, \qquad \vv B'_{\perp} = \gamma\Big(\vv B - \frac{1}{c^{2}}\vv v\times\vv E\Big)_{\perp}.

Three things are worth noticing immediately.

The components along the boost do not change. That is the opposite of what happens to a four-vector, whose parallel component is the one that mixes with time. The reason is visible in (2.6.15). Here ExE_{x} sits in the 0101 slot, so both of its indices lie in the block the boost acts on, and it acquires two factors of Λ\Lambda that undo each other. Meanwhile BxB_{x} sits in the 2323 slot, which the boost does not touch at all.

The combination E+v×B\vv E+\vv v\times\vv B is already familiar. It is the Lorentz force per unit charge. Section 8 shows that is not a coincidence.

If B=0\vv B=\vv 0 in one frame, B0\vv B'\neq\vv 0 in another. The one exception is when E\vv E happens to be parallel to v\vv v. So a pure electric field is not a frame-independent notion at all. Magnetism is what an electric field looks like from a moving frame, and §6 makes that quantitative in the one case where you can check it against a laboratory measurement.

Grind box — the boost of FμνF^{\mu\nu}, entry by entry

The transformation is F=ΛFΛTF'=\Lambda F\Lambda^{\mathsf T}. Since Λ\Lambda is symmetric here, ΛT=Λ\Lambda^{\mathsf T}=\Lambda, and it acts nontrivially only on the 0011 block. Write Λ\Lambda as the 2×22\times2 block (γγβγβγ)\begin{pmatrix}\gamma&-\gamma\beta\\-\gamma\beta&\gamma\end{pmatrix} on indices (0,1)(0,1) and the identity on (2,3)(2,3).

Entries with both indices in {0,1}\{0,1\}. Only F01F^{01} is nonzero there. The 2×22\times2 block of FF is (0Ex/cEx/c0)\begin{pmatrix}0&-E_{x}/c\\ E_{x}/c&0\end{pmatrix}, which is (Ex/c)(-E_{x}/c) times the antisymmetric symbol. For any 2×22\times2 matrix MM and the antisymmetric ε\varepsilon we have MεMT=det(M)εM\varepsilon M^{\mathsf T}=\det(M)\,\varepsilon, a two-dimensional special case of the determinant identity in Chapter 2.4 §8.1. Here det=γ2γ2β2=γ2(1β2)=1\det=\gamma^{2}-\gamma^{2}\beta^{2}=\gamma^{2}(1-\beta^{2})=1. Hence

F01=F01Ex=Ex. F'^{01}=F^{01} \qquad\Longrightarrow\qquad E'_{x}=E_{x}.

Entries with both indices in {2,3}\{2,3\}. Λ\Lambda is the identity there, so F23=F23F'^{23}=F^{23}, i.e. Bx=BxB'_{x}=B_{x}.

Mixed entries. These carry one factor of the block and one identity. For F02F'^{02}:

F02=Λ0αΛ2βFαβ=Λ0αFα2=γF02γβF12. F'^{02} = \Lambda^{0}{}_{\alpha}\Lambda^{2}{}_{\beta}F^{\alpha\beta} = \Lambda^{0}{}_{\alpha}F^{\alpha 2} = \gamma F^{02} - \gamma\beta F^{12}.

Substitute F02=Ey/cF^{02}=-E_{y}/c and F12=BzF^{12}=-B_{z}:

Eyc=γEyc+γβBzEy=γ(EyβcBz)=γ(EyvBz). -\frac{E'_{y}}{c} = -\gamma\frac{E_{y}}{c} + \gamma\beta B_{z} \qquad\Longrightarrow\qquad E'_{y} = \gamma\big(E_{y} - \beta c B_{z}\big) = \gamma\big(E_{y}-vB_{z}\big).

For F03F'^{03}, with F03=Ez/cF^{03}=-E_{z}/c and F13=+ByF^{13}=+B_{y}:

Ezc=γF03γβF13=γEzcγβBy    Ez=γ(Ez+vBy). -\frac{E'_{z}}{c} = \gamma F^{03} - \gamma\beta F^{13} = -\gamma\frac{E_{z}}{c} - \gamma\beta B_{y} \;\Longrightarrow\; E'_{z} = \gamma\big(E_{z}+vB_{y}\big).

For F12F'^{12}, now the first index sits in the block as a row:

F12=Λ1αFα2=γβF02+γF12=γβEycγBz, F'^{12} = \Lambda^{1}{}_{\alpha}F^{\alpha2} = -\gamma\beta F^{02} + \gamma F^{12} = \gamma\beta\frac{E_{y}}{c} - \gamma B_{z},

and since F12=BzF'^{12}=-B'_{z},   Bz=γ(BzvEy/c2)\;B'_{z}=\gamma\big(B_{z}-vE_{y}/c^{2}\big). Identically, F13=γβF03+γF13F'^{13}=-\gamma\beta F^{03}+\gamma F^{13} gives By=γ(By+vEz/c2)B'_{y}=\gamma\big(B_{y}+vE_{z}/c^{2}\big).

Six components, six rules, no others. Every one of them was checked symbolically against ΛFΛT\Lambda F\Lambda^{\mathsf T} before this box was written.

A consistency check worth doing. Apply the rules twice, once with vv and once with v-v, and you must get back what you started with. For EyE_{y}: γ(Ey+vBz)=γ2[(EyvBz)+v(BzvEy/c2)]=γ2Ey(1v2/c2)=Ey\gamma(E'_{y}+vB'_{z}) = \gamma^{2}\big[(E_{y}-vB_{z}) + v(B_{z}-vE_{y}/c^{2})\big] = \gamma^{2}E_{y}(1-v^{2}/c^{2}) = E_{y} ✓. The γ2(1β2)=1\gamma^{2}(1-\beta^{2})=1 that makes it work is the same identity that made detΛ=1\det\Lambda=1 above.

⚠ "E\vv E and B\vv B are the same thing" — what that does and does not mean

It does not mean you can turn an electric field into a magnetic field by running. The invariants of §7 forbid that in general. If EB0\vv E\cdot\vv B\neq0 in one frame it is nonzero in every frame, so a field configuration with both fields present and non-orthogonal has no frame at all in which either one vanishes.

What it does mean is more precise and weaker. E\vv E and B\vv B are components of a single object, and the split between them is observer-dependent in exactly the way the split of a four-vector into time and space parts is. Two observers disagree about how much of FμνF^{\mu\nu} is electric in the same sense that they disagree about how much of pμp^{\mu} is energy. Neither of them is confused. The word "electric" names a slice, and they are slicing differently. What they do agree about is FμνF^{\mu\nu} itself, and the two scalars §7 builds from it.

In plain terms 2.6.5

If the two fields are entries of one object, changing frames must shuffle them into each other, and the rule needs nothing beyond the transformation law in hand. The entries along the direction of motion come through untouched, the reverse of what happens to a four-part vector, whose parallel component is the one mixing with time; a boost acting on both labels of one entry undoes itself. Across the motion the fields trade, and the combination appearing in the trade is the familiar force per unit charge.

The immediate consequence is that a purely electric field is not a notion anybody can defend as absolute. Let one observer find no magnetism anywhere and another moving past will find some. Magnetism is what an electric field looks like from a moving frame, a slogan about to be turned into arithmetic checkable against a laboratory bench.

What the slogan does not license is the belief that either field can always be transformed away. Some configurations refuse, and two agreed numbers built by contracting every label decide which. The exact statement is weaker and better: the split between electric and magnetic is observer-dependent in precisely the way the split of a displacement into time and space is. Two observers disagree about how much of the object is electric in the same sense that they disagree about how much of a momentum is energy.

6 · The wire — magnetism as electrostatics in disguise

This is the argument that makes the chapter's title literal, and it can be checked against a current balance on a laboratory bench.

6.1 · The setup, in the lab

Take a long straight wire along the xx-axis. We model it as Chapter 1.1's grind boxes model everything, with the crudest structure that still has the right physics. That means a rigid lattice of positive ions at rest, and a gas of conduction electrons drifting through it.

  • Lattice: at rest in the lab, linear charge density +λ0+\lambda_{0}.
  • Electrons: drifting with velocity vdx^-v_{d}\hat{\vv x}, linear charge density λ0-\lambda_{0} as measured in the lab.

The two densities are equal and opposite in the lab, so the wire is electrically neutral there. That is an experimental fact about wires rather than an assumption we are making. A current-carrying wire does not attract a stationary pith ball. With the densities set, the conventional current is

I  =  (λ0)×(vd)  =  +λ0vd, I \;=\; (-\lambda_{0})\times(-v_{d}) \;=\; +\lambda_{0}v_{d}, (2.6.39)

flowing in the +x^+\hat{\vv x} direction, because negative charge moving one way is positive current the other.

Now put a test charge q>0q\gt0 at perpendicular distance rr from the wire, at position ry^r\hat{\vv y}, moving parallel to the wire with velocity ux^u\hat{\vv x}.

In the lab the wire is neutral, so E=0\vv E=\vv 0 and there is no electric force whatever. That leaves the magnetic force to account for everything. The magnetic field of a long straight wire at distance rr is B=μ0I2πrx^×r^\vv B=\dfrac{\mu_{0}I}{2\pi r}\,\hat{\vv x}\times\hat{\vv r}. In our geometry r^=y^\hat{\vv r}=\hat{\vv y} and x^×y^=z^\hat{\vv x}\times\hat{\vv y}=\hat{\vv z}, so the field and the force on the test charge come out as

B  =  μ0I2πrz^,Flab  =  qux^×B  =  μ0quλ0vd2πry^. \vv B \;=\; \frac{\mu_{0}I}{2\pi r}\,\hat{\vv z}, \qquad\qquad \vv F_{\text{lab}} \;=\; q\,u\hat{\vv x}\times\vv B \;=\; -\,\frac{\mu_{0}q\,u\,\lambda_{0}v_{d}}{2\pi r}\,\hat{\vv y}. (2.6.40)

The force is attractive. It points from the charge toward the wire, since x^×z^=y^\hat{\vv x}\times\hat{\vv z}=-\hat{\vv y}. The force is purely magnetic and there is no ambiguity about that.

6.2 · The same situation, from the test charge's frame

Now board the test charge. Transform to the frame SS' moving at +ux^+u\hat{\vv x}. In SS' the test charge is at rest, so it cannot feel a magnetic force at all. The Lorentz force on a stationary charge has no magnetic term, whatever B\vv B' may be.

And yet the charge is certainly still accelerating toward the wire. Whether it hits the wire is not a matter of opinion. So some other force must be doing the work, and there is only one candidate.

Our next step is to compute the two charge densities in SS'. This is the crux of the whole argument, so we go slowly.

The lattice. It is at rest in the lab with density λ0\lambda_{0}, so its proper density, meaning the density in its own rest frame, is λ0\lambda_{0}. In SS' it moves at speed uu, so the lattice is length-contracted by γu\gamma_{u} and the same charge occupies less length:

λ+  =  γuλ0,γu=(1u2/c2)1/2. \lambda'_{+} \;=\; \gamma_{u}\,\lambda_{0}, \qquad \gamma_{u}=\big(1-u^{2}/c^{2}\big)^{-1/2}. (2.6.41)

The electrons. These were already moving in the lab, so we must first back out their proper density. In the lab they move at vdv_{d} and have density λ0-\lambda_{0}. Running the logic of (2.6.41) backwards, their proper density is λ0/γd-\lambda_{0}/\gamma_{d} with γd=(1vd2/c2)1/2\gamma_{d}=(1-v_{d}^{2}/c^{2})^{-1/2}. To contract that proper density into SS' we need their speed in SS', which the velocity-addition law of Chapter 2.2 §5 supplies:

v  =  vdu1+vdu/c2, v'_{-} \;=\; \frac{-v_{d}-u}{1+v_{d}u/c^{2}}, (2.6.42)

What the contraction actually needs is not that velocity but the Lorentz factor built from it, so substitute the last line into γ=(1v2/c2)1/2\gamma=(1-v^{2}/c^{2})^{-1/2} and simplify. The grind box does the algebra. The result is remarkably tidy:

γ  =  γuγd(1+uvdc2). \gamma'_{-} \;=\; \gamma_{u}\,\gamma_{d}\left(1+\frac{u v_{d}}{c^{2}}\right). (2.6.43)

Now we have everything the electrons need. Their density in SS' is their proper density λ0/γd-\lambda_{0}/\gamma_{d} multiplied by the contraction factor we just computed:

λ  =  λ0γdγ  =  γuλ0(1+uvdc2). \lambda'_{-} \;=\; -\frac{\lambda_{0}}{\gamma_{d}}\cdot\gamma'_{-} \;=\; -\gamma_{u}\lambda_{0}\left(1+\frac{uv_{d}}{c^{2}}\right). (2.6.44)

The two densities have been contracted by different factors. Both of them picked up γu\gamma_{u}. The electrons picked up an extra (1+uvd/c2)(1+uv_{d}/c^{2}) on top of that, because they were already moving before we boosted. Adding the two densities gives the net charge on the wire as SS' sees it:

  λ  =  λ++λ  =  γuλ0[1(1+uvdc2)]  =  γuλ0uvdc2.   \boxed{\;\lambda' \;=\; \lambda'_{+} + \lambda'_{-} \;=\; \gamma_{u}\lambda_{0}\left[1 - \left(1+\frac{uv_{d}}{c^{2}}\right)\right] \;=\; -\,\gamma_{u}\lambda_{0}\,\frac{u v_{d}}{c^{2}}.\;} (2.6.45)

The wire is not neutral in this frame. It carries a net negative charge per unit length. The test charge is positive, so it is attracted, and that is the direction we already know the force to have. To leading order in uvd/c2uv_{d}/c^{2} we may set γu1\gamma_{u}\to1, which gives λλ0uvd/c2\lambda'\approx-\lambda_{0}uv_{d}/c^{2}.

6.3 · The force, and the comparison

The field of an infinite line charge λ\lambda' at perpendicular distance rr' is radial, with magnitude E=λ/(2πϵ0r)E'=\lambda'/(2\pi\epsilon_{0}r'). Distances perpendicular to the boost are unchanged, so r=rr'=r. The test charge is stationary in SS', so the force on it is purely electric:

F  =  qE  =  qλ2πϵ0ry^  =  γuqλ0uvd2πϵ0c2ry^  =  μ0γuquλ0vd2πry^, \vv F' \;=\; q\vv E' \;=\; \frac{q\lambda'}{2\pi\epsilon_{0}r}\,\hat{\vv y} \;=\; -\,\frac{\gamma_{u}\,q\lambda_{0}uv_{d}}{2\pi\epsilon_{0}c^{2}r}\,\hat{\vv y} \;=\; -\,\frac{\mu_{0}\,\gamma_{u}\,q\,u\,\lambda_{0}v_{d}}{2\pi r}\,\hat{\vv y}, (2.6.46)

The last step there used 1/(ϵ0c2)=μ01/(\epsilon_{0}c^{2})=\mu_{0}. Now set that beside the lab result (2.6.40), which is what we came here to compare it with:

  F  =  γuFlab.   \boxed{\;\vv F' \;=\; \gamma_{u}\,\vv F_{\text{lab}}.\;} (2.6.47)

That factor γu\gamma_{u} is exactly the factor by which a transverse force is expected to change between the two frames. Here is why, in three lines.

Momentum transverse to the boost is unchanged, py=pyp'_{y}=p_{y}, because pμp^{\mu} is a four-vector and the boost matrix of Chapter 2.4 §2.1 acts as the identity on the yy and zz rows. Time is not unchanged. For a particle instantaneously at rest in SS' we have dx=0\dd x'=0, so the time transformation gives dt=γu(dt+udx/c2)=γudt\dd t=\gamma_{u}(\dd t'+u\,\dd x'/c^{2})=\gamma_{u}\dd t'. Dividing the first fact by the second,

Fy  =  dpydt  =  dpyγudt  =  Fyγu, F_{y} \;=\; \dv{p_{y}}{t} \;=\; \frac{\dd p'_{y}}{\gamma_{u}\,\dd t'} \;=\; \frac{F'_{y}}{\gamma_{u}}, (2.6.48)

which is (2.6.47) rearranged. The two calculations agree identically, not approximately.

What just happened

One observer says this. The wire is neutral, the charge is moving, there is a magnetic field, and the force is qu×Bq\vv u\times\vv B.

The other says this. There is no magnetic force here, because nothing is moving. The wire carries a net negative charge, and what you are watching is Coulomb attraction.

They are both right, they compute the same trajectory, and neither description is more fundamental than the other. What one observer calls magnetism, another calls electrostatics acting on a wire that is not neutral.

So magnetism is a relativistic correction to Coulomb's law, the v2/c2v^{2}/c^{2} piece of it. The only reason it is not a fantastically small effect is that the zeroth-order term has been cancelled to fantastic precision by the neutrality of matter.

6.4 · How small is vd/cv_{d}/c, and why you can still pick up a nail

Let's put numbers in, because they are startling. Copper has one conduction electron per atom, density 8960 kgm38960\ \mathrm{kg\,m^{-3}} and molar mass 63.55 gmol163.55\ \mathrm{g\,mol^{-1}}, so the number density of mobile electrons is

n  =  896063.55×103×6.022×1023  =  8.49×1028 m3. n \;=\; \frac{8960}{63.55\times10^{-3}}\times6.022\times10^{23} \;=\; 8.49\times10^{28}\ \mathrm{m^{-3}}. (2.6.49)

In a wire of cross-section 1 mm21\ \mathrm{mm^{2}} carrying 1 A1\ \mathrm{A}, the drift speed follows from I=neAvdI=neAv_{d}:

vd  =  IneA  =  1(8.49×1028)(1.602×1019)(106)  =  7.35×105 ms1, v_{d} \;=\; \frac{I}{neA} \;=\; \frac{1}{(8.49\times10^{28})(1.602\times10^{-19})(10^{-6})} \;=\; 7.35\times10^{-5}\ \mathrm{m\,s^{-1}}, (2.6.50)

That is 0.074 mms10.074\ \mathrm{mm\,s^{-1}}, slower than a growing fingernail. Expressed as a fraction of the speed of light, which is the form the physics above cares about, it is

vdc  =  2.45×1013. \frac{v_{d}}{c} \;=\; 2.45\times10^{-13}. (2.6.51)

So magnetism is an effect of relative order 101310^{-13}, and yet electromagnets lift cars. How can both of those be true?

The resolution is in (2.6.45). The surviving effect is not λ0(vd/c)\lambda_{0}(v_{d}/c) on its own. It is the difference between two enormous numbers, and the enormous number is λ0\lambda_{0} itself. In that same wire,

λ0  =  neA  =  1.36×104 Cm1, \lambda_{0} \;=\; neA \;=\; 1.36\times10^{4}\ \mathrm{C\,m^{-1}}, (2.6.52)

thirteen thousand coulombs of mobile charge per metre, exactly cancelled by thirteen thousand coulombs of lattice charge. A tiny fractional imbalance in a colossal cancellation is still a substantial charge. Worked example 1 does the arithmetic all the way to a force.

w = 0.000 c
0.500
frame speed w = 0.000 c, gamma_w = 1.000 | test charge in this frame: u = 0.700 c
λ₊/λ₀ = +1.000 λ₋/λ₀ = −1.000 → net 0.000 (formula −gamma_w·w·v_d/c² = 0.000)
spacings: a₊ = 1.000 a, a₋ = 1.000 a, ratio a₋/a₊ = 1.000 (1+w v_d/c²)⁻¹ = 1.000
F_electric = 0.000 F₀, F_magnetic = 1.000 F₀, total = 1.000 F₀ | INVARIANT gamma·F⊥ = 1.400 F₀ (gamma_u = 1.400)
a real copper wire (1 mm², 1 A): v_d/c = 2.45e-13, λ₀ = 1.360e+4 C/m, net λ′ = -1.113e-12 C/m at u = 10⁵ m/s ⇒ F = 3.204e-19 N
One wire, two frames. The test charge qq moves at a fixed u=0.7cu=0.7c in the lab; the slider carries you continuously from the lab frame to qq's rest frame. The drift speed is exaggerated by about twelve orders of magnitude so that the contraction is visible at all — in a real copper wire vd/c2.5×1013v_{d}/c\approx2.5\times10^{-13}, as (2.6.51) computes and the third button reports. Everything drawn and every number displayed is evaluated live from (2.6.41)(2.6.45) and the field transformation rules of §5; nothing is a stored result. (i) Watch the two rows of charges. Both compress as you boost, but by different factors: the lattice by γw\gamma_{w}, the electrons by γw(1+wvd/c2)\gamma_{w}(1+wv_{d}/c^{2}). The spacing readout shows the ratio, which is (1+wvd/c2)1(1+wv_{d}/c^{2})^{-1} and has nothing to do with how the picture was drawn. (ii) The net line charge appears as soon as you leave the lab frame, growing from zero to γuλ0uvd/c2-\gamma_{u}\lambda_{0}uv_{d}/c^{2}. (iii) The force budget bar splits the transverse force into its electric (orange) and magnetic (blue) parts. In the lab it is all magnetic; in qq's frame it is all electric. The total is not constant, and should not be: it grows by γ\gamma, because forces are not invariants. What is constant is the lower bar, γqF\gamma_{q}F_{\perp} — the transverse component of the four-force of Chapter 2.5 §6.1, which is a component of a four-vector and therefore untouched by a boost along xx. That bar does not move by a pixel as you drag, and it is the honest statement of "the physics is the same".
Grind box — the composition identity γ=γuγd(1+uvd/c2)\gamma'_{-}=\gamma_{u}\gamma_{d}(1+uv_{d}/c^{2})

Let a particle move at velocity ww along xx in frame SS, and let SS' move at +u+u. The velocity in SS' is w=(wu)/(1uw/c2)w'=(w-u)/(1-uw/c^{2}). Compute 1w2/c21-w'^{2}/c^{2}:

1w2c2=1(wu)2c2(1uwc2)2=c2(1uwc2)2(wu)2c2(1uwc2)2. \begin{aligned} 1 - \frac{w'^{2}}{c^{2}} &= 1 - \frac{(w-u)^{2}}{c^{2}\left(1-\frac{uw}{c^{2}}\right)^{2}} = \frac{c^{2}\left(1-\frac{uw}{c^{2}}\right)^{2} - (w-u)^{2}}{c^{2}\left(1-\frac{uw}{c^{2}}\right)^{2}}. \end{aligned}

Expand the numerator, writing a=u/ca=u/c, b=w/cb=w/c and factoring out c2c^{2}:

(1ab)2(ba)2=12ab+a2b2b2+2aba2=(1a2)(1b2). (1-ab)^{2} - (b-a)^{2} = 1 - 2ab + a^{2}b^{2} - b^{2} + 2ab - a^{2} = (1-a^{2})(1-b^{2}).

The cross terms cancel exactly, and that is the only step in this box with any content. Putting that factorised numerator back and taking the reciprocal square root,

1w2c2=(1u2/c2)(1w2/c2)(1uw/c2)2γw=γuγw(1uwc2). 1-\frac{w'^{2}}{c^{2}} = \frac{(1-u^{2}/c^{2})(1-w^{2}/c^{2})}{\left(1-uw/c^{2}\right)^{2}} \qquad\Longrightarrow\qquad \gamma_{w'} = \gamma_{u}\gamma_{w}\left(1-\frac{uw}{c^{2}}\right).

For the electrons, w=vdw=-v_{d}, and the minus sign in the bracket becomes a plus: γ=γuγd(1+uvd/c2)\gamma'_{-}=\gamma_{u}\gamma_{d}(1+uv_{d}/c^{2}), which is (2.6.43). ✓

Bonus. Read the same identity in reverse and it says that γw=γuγw(1uw/c2)\gamma_{w'}=\gamma_{u}\gamma_{w}(1-uw/c^{2}) is the 00-component of the boosted four-velocity. That is to say, it is u0=Λ0μuμu'^{0}=\Lambda^{0}{}_{\mu}u^{\mu} written out. Nothing new was needed here. Velocity addition is the transformation of uμu^{\mu}, as Chapter 2.5 remarked.

And the same identity again. Notice that the spacings drawn in the figure obey a/a+=(1+wvd/c2)1a'_{-}/a'_{+}=(1+wv_{d}/c^{2})^{-1}, which is this identity divided by itself. The two rows of charges compress by different amounts because they enter the boost with different γ\gamma's, and for no other reason.

⚠ Why this isn't obvious — what the wire argument does not prove

"Magnetism is just relativity applied to electrostatics" is a true and beautiful slogan, and it is routinely overstated. Here is precisely what the argument above used.

  1. Coulomb's law, for the field of a static line charge. That is an input.
  2. Charge invariance. Without it, λ+=γuλ0\lambda'_{+}=\gamma_{u}\lambda_{0} is false and the whole calculation collapses. This is an experimental fact, flagged in the ⚑ callout at the head of the chapter. It is not a consequence of relativity.
  3. The field concept, meaning the idea that the force is mediated by something local, so that "the field at the test charge" is a meaningful quantity to transform. A pure action-at-a-distance theory would need a different argument.
  4. The Lorentz force law in the lab, to have something to compare against.

What relativity supplies is the relation between the two descriptions, and it supplies it with no freedom at all. Given Coulomb, charge invariance and the field concept, the magnetic force is not optional and its magnitude is fixed.

That is a very strong statement. It means the constant μ0\mu_{0} is not independent of ϵ0\epsilon_{0}, which is Chapter 2.1's c=1/μ0ϵ0c=1/\sqrt{\mu_{0}\epsilon_{0}} arriving from the other direction. It is still not the same as saying "magnetism can be derived from electrostatics alone", which is false.

In plain terms 2.6.6

The argument making the claim literal is checkable against a current balance. A current-carrying wire is electrically neutral in the laboratory, which is experiment and not assumption, since a live wire does not attract a pith ball, so a charge moving alongside feels a purely magnetic force. Now board that charge. In its own frame nothing moves, so no magnetic force is available, yet it still accelerates towards the wire, whether it strikes not being a matter of opinion.

Something else is doing the work, and there is one candidate. The wire holds two populations of charge, a lattice at rest and a gas of electrons drifting through it, contracting by different factors when frames change, the electrons having been moving already and the lattice not. The cancellation that made the wire neutral is spoiled, the wire carries a net charge, and the attraction is ordinary electrostatics. Both compute the same trajectory and neither description is deeper.

The numbers make it startling. Electrons drift along a copper wire at under a tenth of a millimetre per second, two and a half parts in ten million million of light speed, so magnetism is an effect of that size. Electromagnets lift cars because the unbalanced quantity is colossal: thirteen thousand coulombs of mobile charge per metre, cancelled to that precision by the lattice. A minute fractional imbalance in an enormous cancellation is still a substantial charge.

7 · The invariants

If E\vv E and B\vv B separately are observer-dependent, what do all observers agree about? Chapter 2.4 §5.1 answers that in general terms. What everybody agrees about is whatever you can build by contracting every index away. From a single antisymmetric rank-2 tensor there are exactly two such quantities, and this section constructs both of them.

7.1 · The first invariant

The simplest thing you can do with two copies of FF is contract every index of one against every index of the other, which leaves nothing free:

FμνFμν  =  μ,νFμνFμν. F_{\mu\nu}F^{\mu\nu} \;=\; \sum_{\mu,\nu}F_{\mu\nu}F^{\mu\nu}. (2.6.53)

To see what that is in terms of E\vv E and B\vv B, split the double sum by index type. The diagonal contributes nothing, since the diagonal of FF is zero. The 0i0i and i0i0 entries contribute twice over, once each way round, and by (2.6.17) we have F0iF0i=(Ei/c)(Ei/c)=(Ei)2/c2F_{0i}F^{0i}=(E^{i}/c)(-E^{i}/c)=-(E^{i})^{2}/c^{2}. So those terms give

2iF0iF0i  =  2c2(Ex2+Ey2+Ez2)  =  2E2c2. 2\sum_{i}F_{0i}F^{0i} \;=\; -\frac{2}{c^{2}}\big(E_{x}^{2}+E_{y}^{2}+E_{z}^{2}\big) \;=\; -\frac{2E^{2}}{c^{2}}. (2.6.54)

That accounts for the electric entries. Now for the purely spatial ones. Lowering both indices leaves them alone, so Fij=Fij=ϵijkBkF_{ij}=F^{ij}=-\epsilon^{ijk}B^{k}, and their contribution is a sum over two epsilons:

i,jFijFij  =  i,jϵijkϵijlBkBl  =  2δklBkBl  =  2B2, \sum_{i,j}F_{ij}F^{ij} \;=\; \sum_{i,j}\epsilon^{ijk}\epsilon^{ijl}B^{k}B^{l} \;=\; 2\,\delta^{kl}B^{k}B^{l} \;=\; 2B^{2}, (2.6.55)

using the contraction identity ϵijkϵijl=2δkl\epsilon^{ijk}\epsilon^{ijl}=2\delta^{kl} (sum on i,ji,j), which is the previous identity ϵijkϵklm=δilδjmδimδjl\epsilon^{ijk}\epsilon^{klm}=\delta^{il}\delta^{jm}-\delta^{im}\delta^{jl} with one more index contracted. Adding:

  FμνFμν  =  2(B2E2c2).   \boxed{\;F_{\mu\nu}F^{\mu\nu} \;=\; 2\left(B^{2} - \frac{E^{2}}{c^{2}}\right).\;} (2.6.56)

7.2 · The second invariant

The other contraction available uses the Levi-Civita symbol of Chapter 2.4 §8.1:

ϵμνρσFμνFρσ. \epsilon_{\mu\nu\rho\sigma}F^{\mu\nu}F^{\rho\sigma}. (2.6.57)

Total antisymmetry means the only surviving terms are those in which {μν}\{\mu\nu\} and {ρσ}\{\rho\sigma\} are complementary pairs. There are three such pair-splittings, namely {0123}\{01|23\}, {0213}\{02|13\} and {0312}\{03|12\}. Each of the three accounts for eight of the 2424 nonzero permutations: two orderings inside the first pair, two inside the second, and two choices of which pair comes first.

Flipping any one of those three things costs a sign in ϵ\epsilon and a sign in one factor of FF, so the two signs cancel and all eight terms are equal. That reduces the whole sum to three representative terms with a factor of eight in front:

ϵμνρσFμνFρσ=8[ϵ0123F01F23+ϵ0213F02F13+ϵ0312F03F12]=8[(Exc)(Bx)(Eyc)(By)+(Ezc)(Bz)], \begin{aligned} \epsilon_{\mu\nu\rho\sigma}F^{\mu\nu}F^{\rho\sigma} &= 8\Big[\epsilon_{0123}F^{01}F^{23} + \epsilon_{0213}F^{02}F^{13} + \epsilon_{0312}F^{03}F^{12}\Big]\\[3pt] &= 8\left[\left(-\frac{E_{x}}{c}\right)(-B_{x}) - \left(-\frac{E_{y}}{c}\right)(B_{y}) + \left(-\frac{E_{z}}{c}\right)(-B_{z})\right], \end{aligned} (2.6.58)

using ϵ0123=+1\epsilon_{0123}=+1, ϵ0213=1\epsilon_{0213}=-1 (one transposition) and ϵ0312=+1\epsilon_{0312}=+1 (two transpositions), together with the entries of (2.6.15). All three terms come out positive:

  ϵμνρσFμνFρσ  =  8c  EB.   \boxed{\;\epsilon_{\mu\nu\rho\sigma}F^{\mu\nu}F^{\rho\sigma} \;=\; \frac{8}{c}\;\vv E\cdot\vv B.\;} (2.6.59)
⚑ It is a pseudoscalar, per Chapter 2.4 §8.1

That chapter proved ϵμνρσ=(detΛ)1ϵμνρσ\epsilon'_{\mu\nu\rho\sigma}=(\det\Lambda)^{-1}\epsilon_{\mu\nu\rho\sigma} and noted that detΛ=±1\det\Lambda=\pm1 for Lorentz transformations. So (2.6.59) is genuinely invariant under boosts and rotations (detΛ=+1\det\Lambda=+1) but changes sign under a spatial reflection. It is a pseudoscalar. That is not a defect. It is a fact about EB\vv E\cdot\vv B, which is the dot product of a vector with a pseudovector.

Chapter 2.4 also flagged, and we re-flag here, that for general coordinate changes ϵ\epsilon is a tensor density and needs a factor of g\sqrt{-g}. Within Part II, where detΛ=±1\det\Lambda=\pm1, that never bites.

One thing worth knowing in advance. A term proportional to (2.6.59) may legally be added to the Lagrangian of §9, and it is the notorious θ\theta-term. Chapter 6.5 explains why the strong interaction's version of it is measured to be astonishingly close to zero.

7.3 · What the invariants forbid

We now have two numbers that everyone agrees on. They are strong constraints, and it is worth being exact about what each of them rules out.

A plane light wave stays a plane light wave. Chapter 2.1 §2 found that in an electromagnetic wave EB\vv E\perp\vv B and E=cBE=cB. Then both invariants vanish:

B2E2c2  =  B2B2  =  0,EB  =  0. B^{2}-\frac{E^{2}}{c^{2}} \;=\; B^{2}-B^{2} \;=\; 0, \qquad\qquad \vv E\cdot\vv B \;=\; 0. (2.6.60)

Both are zero in every frame, so E=cBE'=cB' and EB\vv E'\perp\vv B' for every observer, no matter how fast. You may Doppler-shift a light wave to any frequency you like and change its amplitude by any factor, and Chapter 2.5 §7.2 computed exactly how. What you can never do is make it into something other than a light wave.

In particular, you cannot bring light to rest. A stationary field configuration would have to have E\vv E and B\vv B independent of each other, and there is no frame in which the ratio E/cBE/cB is anything but 11.

When can you kill the magnetic field? Suppose there exists a frame with B=0\vv B'=\vv 0. Then in that frame the invariants read E2/c2-E'^{2}/c^{2} and 00. Since invariants are invariants, that requires

B2E2c2  <  0(i.e. E>cB),andEB  =  0, B^{2}-\frac{E^{2}}{c^{2}} \;\lt\; 0 \quad\text{(i.e. } E\gt cB\text{)}, \qquad\text{and}\qquad \vv E\cdot\vv B \;=\; 0, (2.6.61)

Those two conditions must hold in every frame, and in particular in yours. So the condition is checkable without leaving home. The fields must be perpendicular, and the electric one must dominate. These conditions are also sufficient, and Worked example 2 constructs the frame explicitly in the mirror-image case.

When can you kill the electric field? Run the identical argument with the roles swapped and the conditions come out as

B2E2c2  >  0(i.e. E<cB),andEB  =  0. B^{2}-\frac{E^{2}}{c^{2}} \;\gt\; 0 \quad\text{(i.e. } E\lt cB\text{)}, \qquad\text{and}\qquad \vv E\cdot\vv B \;=\; 0. (2.6.62)

Worked example 2 does this case in full and finds the frame velocity, which turns out to be the drift velocity every mass spectrometer relies on.

And if EB0\vv E\cdot\vv B\neq0, neither field can be removed in any frame, ever. The two fields are then irreducibly both present, and the best you can do is find a frame where they are parallel.

Reading the invariants as a classification

The pair (FμνFμν, ϵFF)\big(F_{\mu\nu}F^{\mu\nu},\ \epsilon F F\big) classifies field configurations into frame-independent types, exactly as the sign of Δs2\Delta s^{2} classified pairs of events in Chapter 2.3 into timelike, spacelike and null.

InvariantsNameSimplest frame
E>cBE\gt cB, EB=0\vv E\cdot\vv B=0electrica frame with B=0\vv B=\vv 0: pure electrostatics
E<cBE\lt cB, EB=0\vv E\cdot\vv B=0magnetica frame with E=0\vv E=\vv 0: pure magnetostatics
E=cBE=cB, EB=0\vv E\cdot\vv B=0nullnone simpler, a radiation field
EB0\vv E\cdot\vv B\neq0generica frame with EB\vv E\parallel\vv B

The null case is the boundary between the other two, and it is exactly the case that cannot be simplified. That happens in the same way that a null interval sits on the boundary of the light cone and cannot be transformed into either a pure time separation or a pure space separation. Light is the null case of electromagnetism in precisely the sense that a light ray is the null case of a worldline.

That table has a shape, and it is worth seeing it. Take E=(E,E,0)\vv E=(E_{\parallel},E_{\perp},0) and B=(B,0,B)\vv B=(B_{\parallel},0,B_{\perp}). That is a configuration on which a boost along x^\hat{\vv x} closes in two dimensions, so the picture stays flat, and the four rows of the table become four regions of one plane.

φ = +0.000 β = tanh φ = +0.000000
0.80
0.35
0.30
0.50
E′ = ( 0.300000, 0.800000, 0.000000) B′ = ( 0.500000, 0.000000, 0.350000)
F_μν F^μν by contraction of F′: -0.7150000000 2(B′²−E′²): -0.7150000000 Δ 1.1e-16
ε F F by the 24-term sum: +1.2000000000 8 E′·B′: +1.2000000000 Δ 0.0e+0
type, from the invariants alone: GENERIC an x-boost at φ* = +0.469135 (β* = 0.437500) drives B⊥ to zero, but B∥ = 0.50 survives it.
over the whole range |φ| ≤ 3: E∥ moves 9.9e-15, B∥ moves 0.0e+0, max(|E′_z|,|B′_y|) = 0.0e+0, B⊥²−E⊥² drifts 4.1e-14
drag the rapidity: the marker moves along its curve and never off it.
The orbit a field cannot leave. Inside this figure only, units are chosen with c=1c=1, so E\vv E and B\vv B are numbers on one scale and the lines B=±EB_{\perp}=\pm E_{\perp} sit at 4545^{\circ}. Everything drawn and every number printed comes from assembling FμνF^{\mu\nu} as in (2.6.15) and multiplying: F=ΛFΛTF'=\Lambda F\Lambda^{\mathsf T} with Λ\Lambda from (2.6.38), two 4×44\times4 products at every point of every curve. None of §5's transformation rules is used, and both invariants are contracted straight out of FF' — the first by lowering both indices, the second by a genuine twenty-four-term sum over the permutations of ϵμνρσ\epsilon_{\mu\nu\rho\sigma} — never from (2.6.56) or (2.6.59). So the two numbers on each of those readout lines are two independent computations of one quantity, and they never differ by more than 101510^{-15}. (i) The upper panel is the transverse plane, EE_{\perp} across and BB_{\perp} up. Every curve in it is an orbit — a seed configuration pushed through the rapidities and drawn wherever the matrix multiplication puts it, not a conic evaluated from a formula. The faint grey curves are the family, the two orange lines are the seeds with E=B|E_{\perp}|=|B_{\perp}|, and the bold blue curve is the orbit of the configuration you have set. The hollow marker is the lab frame, the filled one is the frame you are in, and the dots along the bold curve are spaced by Δϕ=0.5\Delta\phi=0.5 — uniform in rapidity, the device Chapter 2.3 used on the light cone and for the same reason. These are Chapter 2.3's invariant hyperbolae and Chapter 2.5's mass shell for the third time, now in field space; which curve you sit on is §7.3's classification, and the type readout computes it from the invariants alone. (ii) The lower panel is the longitudinal plane, and its whole content is that nothing happens in it. EE_{\parallel} sits in the 0101 slot and BB_{\parallel} in the 2323 slot, so a boost along x^\hat{\vv x} acts on both of the first one's indices and neither of the second's; and in this configuration their product is EB\vv E\cdot\vv B. The last readout sweeps the whole rapidity range and reports the largest excursion either component ever makes — zero to machine precision — beside the largest value EzE'_{z} and ByB'_{y} ever reach, which is exactly zero, and the drift of B2E2B_{\perp}^{2}-E_{\perp}^{2}, which never exceeds 101310^{-13}. The two-dimensional closure is measured, not assumed. (iii) Press a light wave and drag the rapidity as hard as it will go. Both invariants are zero, the type reads null, and the marker slides down the 4545^{\circ} line toward the origin with the two fields shrinking together in lockstep. It never leaves the line and it never arrives: you have Doppler-shifted a light wave through every frame there is and failed either to change what it is or to bring it to rest. (iv) Then press EB0\vv E\cdot\vv B\neq0, watch the type flip to generic, and drag again with your eye on the lower panel. The marker does not move by a pixel. Neither field can be removed by anybody, ever — and the reason is on the screen rather than in a table, because the parts that would have to go are exactly the parts a boost along the motion never touches.
In plain terms 2.6.7

Since each field separately depends on who is looking, the quantities worth having are those nobody can dispute, which means whatever survives contracting every label away. From one antisymmetric object there are exactly two such numbers: the difference between the squared magnitudes of the fields, and the amount by which they overlap. Everybody computes the same pair, whatever their motion.

The pair sorts field configurations into types not open to argument, as the sign of the interval sorted pairs of events into three kinds. Where the fields stand perpendicular and the electric one dominates, some observer finds no magnetism at all; where the magnetic dominates, some observer finds no electricity; and where the two are not perpendicular, neither can be removed by anybody ever, the best available simplification being a frame in which they lie parallel. The test is run without leaving home, because the quantities tested are agreed.

Light is the case where both numbers vanish, and since they vanish for everybody, no change of frame turns a light wave into anything else. Its frequency may be shifted to whatever you like and its amplitude scaled by any factor, and it can never be brought to rest, because no frame exists in which the two fields become independent. Light sits at the boundary between the electric and magnetic types exactly as a light ray sits on the boundary of the light cone.

8 · The Lorentz force, covariantly

Chapter 2.5 §6.1 established the shape any relativistic force law must have:

dpμdτ  =  fμ,with the constraintfu  =  0. \dv{p^{\mu}}{\tau} \;=\; f^{\mu}, \qquad\text{with the constraint}\qquad f\cdot u \;=\; 0. (2.6.63)

What can fμf^{\mu} be for a charge in an electromagnetic field? Three requirements narrow it almost to nothing. It must be a four-vector, it must be linear in the charge, and it must be built from the only field object we have. With uμu^{\mu} the particle's four-velocity, there is essentially one candidate:

  dpμdτ  =  qFμνuν.   \boxed{\;\dv{p^{\mu}}{\tau} \;=\; q\,F^{\mu\nu}u_{\nu}.\;} (2.6.64)

Chapter 2.4's grind box on the quotient theorem already flagged this construction. Requiring FμνuνF^{\mu\nu}u_{\nu} to be a four-vector for every four-velocity is what certifies FμνF^{\mu\nu} as a tensor in the first place. Our job now is to verify that (2.6.64) is the force law you already know.

8.1 · The spatial components

Lower the index on the four-velocity uμ=γ(c,v)u^{\mu}=\gamma(c,\vv v) of Chapter 2.5 §1.3, which flips the sign of the spatial part:

uν  =  ηναuα  =  γ(c,v). u_{\nu} \;=\; \eta_{\nu\alpha}u^{\alpha} \;=\; \gamma\big(c,\,-\vv v\big). (2.6.65)

Take μ=i\mu=i and split the sum over ν\nu into its time and space parts:

qFiνuν  =  q[Fi0u0  +  Fijuj]=  q[Eic(γc)  +  (ϵijkBk)(γvj)]=  γq[Ei+ϵijkvjBk]  =  γq[E+v×B]i. \begin{aligned} qF^{i\nu}u_{\nu} \;&=\; q\Big[F^{i0}u_{0} \;+\; F^{ij}u_{j}\Big]\\[3pt] &=\; q\left[\frac{E^{i}}{c}\,(\gamma c) \;+\; \big(-\epsilon^{ijk}B^{k}\big)\big(-\gamma v^{j}\big)\right]\\[3pt] &=\; \gamma q\left[E^{i} + \epsilon^{ijk}v^{j}B^{k}\right] \;=\; \gamma q\left[\vv E + \vv v\times\vv B\right]^{i}. \end{aligned} (2.6.66)

The left-hand side is dpi/dτ=γdpi/dt\dd p^{i}/\dd\tau=\gamma\,\dd p^{i}/\dd t, since dt/dτ=γ\dd t/\dd\tau=\gamma (Chapter 2.5 §1.2). Cancel the common γ\gamma:

  dpdt  =  q(E+v×B).   \boxed{\;\dv{\vv p}{t} \;=\; q\big(\vv E + \vv v\times\vv B\big).\;} (2.6.67)

That is the Lorentz force law, the one Chapter 1.1 §3.3 had to quote without justification. Here it arrives as the spatial part of a four-vector equation. Notice also that p=γmv\vv p=\gamma m\vv v here rather than mvm\vv v, so the relativistic correction rides along for free.

8.2 · The time component is the power

Take μ=0\mu=0. The ν=0\nu=0 term dies because F00=0F^{00}=0:

qF0νuν  =  qF0juj  =  q(Ejc)(γvj)  =  γqcEv. qF^{0\nu}u_{\nu} \;=\; qF^{0j}u_{j} \;=\; q\left(-\frac{E^{j}}{c}\right)\big(-\gamma v^{j}\big) \;=\; \frac{\gamma q}{c}\,\vv E\cdot\vv v. (2.6.68)

And the left-hand side is dp0/dτ=γd(E/c)/dt\dd p^{0}/\dd\tau=\gamma\,\dd(E/c)/\dd t, where EE is now the particle's energy of Chapter 2.5 §3 (an unfortunate clash of symbols that the whole subject lives with). Cancelling γ/c\gamma/c:

dEdt  =  qEv. \dv{E}{t} \;=\; q\,\vv E\cdot\vv v. (2.6.69)

The magnetic field does no work. That familiar fact is not a separate observation here. It is a consequence of where B\vv B sits in the matrix. The F0jF^{0j} entries are purely electric, so only E\vv E can appear in the energy equation at all.

And (2.6.69) is Chapter 2.5's work–energy theorem dE/dt=Fv\dd E/\dd t=\vv F\cdot\vv v with F\vv F the Lorentz force, since (v×B)v=0(\vv v\times\vv B)\cdot\vv v=0 automatically.

8.3 · The constraint is automatic, and why

Chapter 2.5 §6.1 proved that any four-force must satisfy fu=0f\cdot u=0, because uu=c2u\cdot u=c^{2} is constant. That is a strong condition, and it is not obvious that an arbitrary guess for fμf^{\mu} would satisfy it. Check (2.6.64):

fu  =  fμuμ  =  qFμνuνuμ  =  0. f\cdot u \;=\; f^{\mu}u_{\mu} \;=\; q\,F^{\mu\nu}u_{\nu}u_{\mu} \;=\; 0. (2.6.70)

Why zero? Because FμνF^{\mu\nu} is antisymmetric under μν\mu\leftrightarrow\nu while uμuνu_{\mu}u_{\nu} is symmetric, and a symmetric object fully contracted with an antisymmetric one vanishes.

Here is that argument once, in case it has not been made explicit before. Relabel the dummy indices μν\mu\leftrightarrow\nu in the sum, which changes nothing, and then use the two symmetry properties to put the labels back the way they were:

S    Fμνuμuν  =  Fνμuνuμ  =  Fμνuμuν  =  SS=0. S \;\equiv\; F^{\mu\nu}u_{\mu}u_{\nu} \;=\; F^{\nu\mu}u_{\nu}u_{\mu} \;=\; -F^{\mu\nu}u_{\mu}u_{\nu} \;=\; -S \qquad\Longrightarrow\qquad S=0. (2.6.71)
The structural check that had to work

Let's look at what just happened. The constraint fu=0f\cdot u=0 is a statement about the geometry of spacetime. It says a four-force can only push a four-velocity sideways, because four-velocities all have the same length. It came from Chapter 2.5 with no reference to electromagnetism whatsoever.

And it is satisfied by (2.6.64) because FμνF^{\mu\nu} is antisymmetric. That fact came from F=AAF=\partial A-\partial A, with no reference to spacetime geometry whatsoever. Two independent constraints, arriving from two unrelated directions, match exactly.

The physical content of the match is (2.6.69). The reason the magnetic force does no work is the reason a four-force is orthogonal to a four-velocity. Had FμνF^{\mu\nu} carried any symmetric part, charges would gain energy from nothing.

In plain terms 2.6.8

Three requirements and a shortage of materials pin down the force on a charge. It must be a four-part object, proportional to the charge, and assembled out of the particle's four-velocity and the one field object available; essentially one candidate meets all three. Expanding it returns the force law the first part of this book had to quote without justification, with the repaired momentum riding along at no extra cost, while its time entry says the particle's energy changes at the rate the electric field does work.

That the magnetic field does no work is therefore not a separate fact to be remembered. It is a remark about where the magnetic entries sit, since they occupy the purely spatial part of the array and so cannot appear in the equation governing energy. The familiar statement and the arrangement of the object are one thing.

Then a check that had to work, and does. Any four-force must stand perpendicular to the four-velocity, a constraint out of the geometry of the interval with no reference to electricity. The candidate satisfies it because the field tensor reverses sign when its labels are swapped, a property out of the freedom in the potentials with no reference to geometry. Two constraints from unrelated directions, matching exactly. Had the field tensor carried any part that did not reverse sign, charges would gain energy from nothing.

9 · The action

Chapter 1.2 §8.1 listed eight actions and promised each would be constructed in its own chapter. Line three said electromagnetic field, Chapter 2.6. Here it is, with the source term and the SI constants restored:

  LEM  =  14μ0FμνFμν    jμAμ.   \boxed{\;\mathcal{L}_{\text{EM}} \;=\; -\frac{1}{4\mu_{0}}\,F_{\mu\nu}F^{\mu\nu} \;-\; j_{\mu}A^{\mu}.\;} (2.6.72)

Two terms, one line, and it contains every result of §3. The claim we have to verify is that feeding (2.6.72) to the field Euler–Lagrange equation of Chapter 1.2 §8 returns μFμν=μ0jν\partial_{\mu}F^{\mu\nu}=\mu_{0}j^{\nu}.

9.1 · Varying it

The dynamical variable here is AνA_{\nu}, which is four fields, one for each value of ν\nu. So the Euler–Lagrange equation carries a free index:

μ ⁣(L(μAν))    LAν  =  0. \partial_{\mu}\!\left(\pdv{\mathcal{L}}{(\partial_{\mu}A_{\nu})}\right) \;-\; \pdv{\mathcal{L}}{A_{\nu}} \;=\; 0. (2.6.73)

The second term is immediate. Only jμAμ=jνAν-j_{\mu}A^{\mu}=-j^{\nu}A_{\nu} contains AνA_{\nu} undifferentiated, so L/Aν=jν\partial\mathcal{L}/\partial A_{\nu}=-j^{\nu}. The first term is where the work is, and the grind box does it slowly. The answer is

L(μAν)  =  1μ0Fμν. \pdv{\mathcal{L}}{(\partial_{\mu}A_{\nu})} \;=\; -\frac{1}{\mu_{0}}F^{\mu\nu}. (2.6.74)

We now have both pieces the Euler–Lagrange equation asks for, so put them into (2.6.73) and see what field equation comes out:

μ ⁣(1μ0Fμν)+jν  =  0μFμν  =  μ0jν. \partial_{\mu}\!\left(-\frac{1}{\mu_{0}}F^{\mu\nu}\right) + j^{\nu} \;=\; 0 \qquad\Longrightarrow\qquad \partial_{\mu}F^{\mu\nu} \;=\; \mu_{0}j^{\nu}. (2.6.75)

That is (2.6.19). Gauss's law and Ampère–Maxwell, both of them, out of one scalar. The homogeneous pair needs no varying at all. It holds identically the moment you write FF in terms of AA, which was the point of §3.4.

Grind box — differentiating FαβFαβF_{\alpha\beta}F^{\alpha\beta} with respect to μAν\partial_{\mu}A_{\nu}

This is pure index gymnastics and it is worth doing once with every step visible, because the same manipulation reappears in Chapters 5.2, 6.3 and 6.4 with more indices.

Step 1. The chain rule on a product. Write LF=14μ0FαβFαβ\mathcal{L}_{F}=-\tfrac{1}{4\mu_{0}}F_{\alpha\beta}F^{\alpha\beta} and note FαβFαβ=ηαγηβδFαβFγδF_{\alpha\beta}F^{\alpha\beta}=\eta^{\alpha\gamma}\eta^{\beta\delta}F_{\alpha\beta}F_{\gamma\delta}, so it is quadratic in the FαβF_{\alpha\beta} with constant coefficients. Hence

(μAν)(FαβFαβ)  =  2FαβFαβ(μAν). \pdv{}{(\partial_{\mu}A_{\nu})}\big(F_{\alpha\beta}F^{\alpha\beta}\big) \;=\; 2F^{\alpha\beta}\,\pdv{F_{\alpha\beta}}{(\partial_{\mu}A_{\nu})}.

Step 2. The inner derivative. Since Fαβ=αAββAαF_{\alpha\beta}=\partial_{\alpha}A_{\beta}-\partial_{\beta}A_{\alpha} and the sixteen quantities μAν\partial_{\mu}A_{\nu} are treated as independent variables,

(αAβ)(μAν)  =  δμαδνβFαβ(μAν)  =  δμαδνβδμβδνα. \pdv{(\partial_{\alpha}A_{\beta})}{(\partial_{\mu}A_{\nu})} \;=\; \delta^{\mu}{}_{\alpha}\,\delta^{\nu}{}_{\beta} \qquad\Longrightarrow\qquad \pdv{F_{\alpha\beta}}{(\partial_{\mu}A_{\nu})} \;=\; \delta^{\mu}{}_{\alpha}\delta^{\nu}{}_{\beta} - \delta^{\mu}{}_{\beta}\delta^{\nu}{}_{\alpha}.

The two deltas in each product are what turn "differentiate with respect to a particular component" into "select that component". The two terms are the two places A\partial A appears in FF.

Step 3. Contract. Now feed Step 2 back into Step 1. The deltas rename indices:

2Fαβ(δμαδνβδμβδνα)  =  2(FμνFνμ)  =  4Fμν, 2F^{\alpha\beta}\big(\delta^{\mu}{}_{\alpha}\delta^{\nu}{}_{\beta} - \delta^{\mu}{}_{\beta}\delta^{\nu}{}_{\alpha}\big) \;=\; 2\big(F^{\mu\nu} - F^{\nu\mu}\big) \;=\; 4F^{\mu\nu},

where the last step used antisymmetry. There is the factor of 44 that the 14\tfrac14 in (2.6.72) was put there to cancel. Carrying the prefactor along,

LF(μAν)  =  14μ04Fμν  =  1μ0Fμν, \pdv{\mathcal{L}_{F}}{(\partial_{\mu}A_{\nu})} \;=\; -\frac{1}{4\mu_{0}}\cdot 4F^{\mu\nu} \;=\; -\frac{1}{\mu_{0}}F^{\mu\nu},

which is (2.6.74). The source term contains no derivatives of AA and contributes nothing here. ✓

Why the sign of the whole Lagrangian is what it is. Expand (2.6.72) into three-vector language using (2.6.56):

14μ0FμνFμν  =  12μ0(B2E2c2)  =  ϵ0E22B22μ0. -\frac{1}{4\mu_{0}}F_{\mu\nu}F^{\mu\nu} \;=\; -\frac{1}{2\mu_{0}}\left(B^{2}-\frac{E^{2}}{c^{2}}\right) \;=\; \frac{\epsilon_{0}E^{2}}{2} - \frac{B^{2}}{2\mu_{0}}.

That is kinetic minus potential. The electric term carries the time derivatives of A\vv A and plays the role of TT, while the magnetic term carries the spatial derivatives and plays the role of VV. The overall sign of (2.6.72) is fixed by demanding that TT come out positive, which is Chapter 1.2 §5's discussion transplanted verbatim to a field. Get it wrong and the field has negative kinetic energy, which Chapter 5.2 shows is a catastrophe rather than a sign convention.

9.2 · Why this Lagrangian and not another

The variation above is a verification rather than a derivation. We wrote down (2.6.72) and checked it. So where did it come from? The honest answer is a short list of demands.

  1. Lorentz invariance. L\mathcal{L} must be a scalar, or the action is not the same number for every observer and the theory has a preferred frame. That immediately restricts us to contractions: FμνFμνF_{\mu\nu}F^{\mu\nu}, ϵFF\epsilon FF, AμAμA_{\mu}A^{\mu}, jμAμj_{\mu}A^{\mu}, and products and derivatives of these.
  2. Gauge invariance. L\mathcal{L} must be unchanged, up to a four-divergence, under (2.6.29). This kills AμAμA_{\mu}A^{\mu} outright, since that term is not gauge invariant. It would have been a photon mass term. The photon is massless because a mass term is not gauge invariant. The same demand permits jμAμj_{\mu}A^{\mu}, and only because charge is conserved, as §4.1 showed.
  3. Second-order equations of motion. L\mathcal{L} should contain at most first derivatives of AA, so that Euler–Lagrange returns something second order and the initial-value problem is the usual one.
  4. Parity, if you want it. ϵμνρσFμνFρσ\epsilon_{\mu\nu\rho\sigma}F^{\mu\nu}F^{\rho\sigma} passes tests 1–3 but is a pseudoscalar, and moreover is a total derivative, so it does not affect the classical field equations at all.

What survives at lowest order is FμνFμνF_{\mu\nu}F^{\mu\nu} and jμAμj_{\mu}A^{\mu}, with two constants in front. One of those constants is fixed by matching Coulomb's law, and the other is absorbed into the normalisation of AA. The form is very nearly forced.

Three doors this opens

Chapter 6.4, Yang–Mills. Take (2.6.72), let AμA^{\mu} carry an internal index and become matrix-valued, and define Fμν=μAννAμig[Aμ,Aν]F_{\mu\nu}=\partial_{\mu}A_{\nu}-\partial_{\nu}A_{\mu}-\ii g[A_{\mu},A_{\nu}]. Then write down 12trFμνFμν-\tfrac12\operatorname{tr}F_{\mu\nu}F^{\mu\nu}, which is the same expression, entry seven in Chapter 1.2's table. The commutator is the only new ingredient, and it is what makes gluons carry colour charge and interact with each other. The strong and weak interactions are this chapter with a non-commuting symmetry group.

Chapter 5.11, effective field theory. The argument of §9.2 is the beginning of a general method. List the fields, list the symmetries, write down every term allowed, and order them by how many derivatives they have. The leading term is the theory, and the rest are corrections suppressed by powers of energy. Electromagnetism looks fundamental partly because the leading allowed term is so simple. Higher terms do exist. One of them, (FμνFμν)2\big(F_{\mu\nu}F^{\mu\nu}\big)^{2}, is the Euler–Heisenberg term, which makes light scatter off light, and it is suppressed by four powers of the electron mass. That is why nobody noticed it for a century.

Chapter 5.2, quantisation. A Lagrangian is the input to a path integral. Once you have (2.6.72), quantising the electromagnetic field is a well-posed problem, and its answer is the photon. The vertex of quantum electrodynamics is the term jμAμ-j_{\mu}A^{\mu} (Chapter 5.8), which is the entire interaction between light and matter and which you have just written down.

⚠ A note on units, before Part V

The factors of μ0\mu_{0}, ϵ0\epsilon_{0} and cc cluttering this chapter are SI, and they carry real information here. They are what let you compare a derived force against a laboratory measurement in newtons, which §6 and the worked examples do.

From Part V onward the book switches to natural units, =c=1\hbar=c=1, in the Heaviside–Lorentz convention where the 4π4\pi is moved out of Coulomb's law and ϵ0=μ0=1\epsilon_{0}=\mu_{0}=1. In those units (2.6.72) reads L=14FμνFμνjμAμ\mathcal{L}=-\tfrac14F_{\mu\nu}F^{\mu\nu}-j_{\mu}A^{\mu}, which is how you will see it written everywhere else, and every equation in this chapter loses its decoration.

Nothing physical changes when you do that, because the constants can always be restored by dimensional analysis. We keep them here because Part II is about measurable consequences, and dropping cc while still learning what cc means would be perverse.

In plain terms 2.6.9

Compress everything so far and one line is left: one number formed by contracting the field tensor with itself, and one term coupling the current to the potential. Hand that line to the machinery of the first part and out comes the equation saying charges and currents make fields, constants and all. The other half needs no varying, holding the instant the field is written in terms of potentials.

How nearly the form is forced is what to carry away. Demand a number every observer computes alike and only contractions are admitted. Demand that the redundancy in the potentials stay invisible and one candidate is killed outright, the one that would have given the force's carrier a mass, which is the whole reason light has none. Demand no more than first derivatives and almost nothing is left. Two terms survive, their constants fixed by matching the static force and by a choice of normalisation.

The method is worth more than the result. List the ingredients, list the symmetries, write down every term they permit and order these by how many derivatives they carry, whereupon the leading one is the theory and the rest are corrections. Let the potential carry an internal label and become a matrix, add the one ingredient non-commuting labels supply, and the same line written again is the strong interaction, whose carriers act on one another because that ingredient says they must.

10 · Field momentum, and Newton's third law repaired

Chapter 1.1 §3.3 put two moving charges alone in the universe and showed that the total force on them does not vanish. Momentum appeared from nowhere at a computable rate.

That chapter offered two escapes. Either momentum conservation is false, or momentum is stored somewhere that is not a particle. It took the second escape, named the electromagnetic field as the only candidate, and left an IOU. This section pays it, and the payment is exact.

10.1 · The energy–momentum tensor, from Noether

Chapter 1.4 proved that every continuous symmetry yields a conserved current. The symmetry we have not yet used is the most basic one there is. The laws are the same here as there, and now as then, which is invariance under spacetime translations xμxμ+aμx^{\mu}\to x^{\mu}+a^{\mu}. Since aμa^{\mu} has four components, we get four conserved currents, one per direction, and those four assemble into a rank-2 object.

Run Chapter 1.4 §6's machinery. Under a translation the field changes by δAλ=aννAλ\delta A_{\lambda}=-a^{\nu}\partial_{\nu}A_{\lambda} and the Lagrangian density by δL=aννL=μ(aνδμνL)\delta\mathcal{L}=-a^{\nu}\partial_{\nu}\mathcal{L}=\partial_{\mu}\big(-a^{\nu}\delta^{\mu}{}_{\nu}\mathcal{L}\big) That last expression is a four-divergence, which is exactly the loosened invariance condition of Chapter 1.4 §1.2. So the symmetry qualifies, and the Noether current for the ν\nu-th translation is

Tcanμν  =  L(μAλ)νAλ    ημνL  =  1μ0FμλνAλ  +  ημν4μ0FαβFαβ, T^{\mu\nu}_{\text{can}} \;=\; \pdv{\mathcal{L}}{(\partial_{\mu}A_{\lambda})}\,\partial^{\nu}A_{\lambda} \;-\; \eta^{\mu\nu}\mathcal{L} \;=\; -\frac{1}{\mu_{0}}F^{\mu\lambda}\,\partial^{\nu}A_{\lambda} \;+\; \frac{\eta^{\mu\nu}}{4\mu_{0}}F_{\alpha\beta}F^{\alpha\beta}, (2.6.76)

using (2.6.74) and setting jμ=0j^{\mu}=0 (we are asking what the free field carries). It satisfies μTcanμν=0\partial_{\mu}T^{\mu\nu}_{\text{can}}=0 by Noether's theorem.

And it is wrong in two ways. It is not symmetric in μν\mu\nu. It is also not gauge invariant, because it contains AλA_{\lambda} bare, so different gauges would assign different energy densities to the same physical field.

Both defects are fixed by the freedom that Chapter 1.4's grind box identified. Adding λΣλμν\partial_{\lambda}\Sigma^{\lambda\mu\nu}, with Σ\Sigma antisymmetric in its first two indices, changes no conserved charge. So we are free to choose a Σ\Sigma that cleans both problems up at once. Take

Σλμν  =  1μ0FμλAν,antisymmetric in λμ because F is. \Sigma^{\lambda\mu\nu} \;=\; \frac{1}{\mu_{0}}F^{\mu\lambda}A^{\nu}, \qquad\text{antisymmetric in }\lambda\mu\text{ because }F\text{ is}. (2.6.77)

Its divergence is, using λFμλ=λFλμ=μ0jμ=0\partial_{\lambda}F^{\mu\lambda}=-\partial_{\lambda}F^{\lambda\mu}=-\mu_{0}j^{\mu}=0 in the source-free region,

λΣλμν  =  1μ0[(λFμλ)=0Aν+FμλλAν]  =  1μ0FμλλAν. \partial_{\lambda}\Sigma^{\lambda\mu\nu} \;=\; \frac{1}{\mu_{0}}\Big[\underbrace{\big(\partial_{\lambda}F^{\mu\lambda}\big)}_{=\,0}A^{\nu} + F^{\mu\lambda}\partial_{\lambda}A^{\nu}\Big] \;=\; \frac{1}{\mu_{0}}F^{\mu\lambda}\,\partial_{\lambda}A^{\nu}. (2.6.78)

Add it to (2.6.76) and the two A\partial A terms combine into an FF:

Tμν  =  1μ0Fμλ(νAλλAν)=Fνλ  +  ημν4μ0FαβFαβ. T^{\mu\nu} \;=\; -\frac{1}{\mu_{0}}F^{\mu\lambda}\underbrace{\big(\partial^{\nu}A_{\lambda} - \partial_{\lambda}A^{\nu}\big)}_{\textstyle =\,F^{\nu}{}_{\lambda}} \;+\; \frac{\eta^{\mu\nu}}{4\mu_{0}}F_{\alpha\beta}F^{\alpha\beta}. (2.6.79)

One last cosmetic step puts the indices in their conventional places. Since Fνλ=FλνF^{\nu}{}_{\lambda}=-F_{\lambda}{}^{\nu}, the minus sign in front can be absorbed, and we are left with the object we were after:

  Tμν  =  1μ0(FμλFλν  +  14ημνFαβFαβ).   \boxed{\;T^{\mu\nu} \;=\; \frac{1}{\mu_{0}}\left(F^{\mu\lambda}F_{\lambda}{}^{\nu} \;+\; \frac{1}{4}\eta^{\mu\nu}F_{\alpha\beta}F^{\alpha\beta}\right).\;} (2.6.80)

Now it is manifestly gauge invariant, since only FF appears in it. It is also symmetric, and that is worth checking rather than believing. The first term is ηλρFμλFρν\eta_{\lambda\rho}F^{\mu\lambda}F^{\rho\nu}. Swap μν\mu\leftrightarrow\nu and relabel the dummies λρ\lambda\leftrightarrow\rho, and you get it back, with the two minus signs from antisymmetry cancelling each other.

That symmetry is not cosmetic. Chapter 3.6 puts TμνT^{\mu\nu} on the right-hand side of Einstein's equation, and the left-hand side of that equation is symmetric by construction.

10.2 · What its components are

T00T^{00}. Only λ\lambda spatial contributes to F0λFλ0F^{0\lambda}F_{\lambda}{}^{0}, and lowering the index costs a minus:

F0λFλ0  =  F0iηijFj0  =  F0iFi0  =  +i(F0i)2  =  E2c2, F^{0\lambda}F_{\lambda}{}^{0} \;=\; F^{0i}\eta_{ij}F^{j0} \;=\; -F^{0i}F^{i0} \;=\; +\sum_{i}\big(F^{0i}\big)^{2} \;=\; \frac{E^{2}}{c^{2}}, (2.6.81)

using Fi0=F0iF^{i0}=-F^{0i} and F0i=Ei/cF^{0i}=-E^{i}/c. With η00=1\eta^{00}=1 and (2.6.56) for the second term,

T00  =  1μ0[E2c2+142(B2E2c2)]  =  E22μ0c2+B22μ0  =    ϵ0E22+B22μ0   T^{00} \;=\; \frac{1}{\mu_{0}}\left[\frac{E^{2}}{c^{2}} + \frac{1}{4}\cdot2\left(B^{2}-\frac{E^{2}}{c^{2}}\right)\right] \;=\; \frac{E^{2}}{2\mu_{0}c^{2}} + \frac{B^{2}}{2\mu_{0}} \;=\; \boxed{\;\frac{\epsilon_{0}E^{2}}{2} + \frac{B^{2}}{2\mu_{0}}\;} (2.6.82)

The last step used 1/(μ0c2)=ϵ01/(\mu_{0}c^{2})=\epsilon_{0}. That is the energy density of the electromagnetic field. You have probably met that expression as a stated fact. Here it is derived, as the conserved Noether charge density of time translation, which is what the word "energy" has meant since Chapter 1.4 §3.1.

T0iT^{0i}. Now η0i=0\eta^{0i}=0, so only the first term survives:

F0λFλi  =  F0jηjkFki  =  F0jFji  =  +F0jFij=  (Ejc)(ϵijkBk)  =  1cϵijkEjBk. \begin{aligned} F^{0\lambda}F_{\lambda}{}^{i} \;&=\; F^{0j}\,\eta_{jk}\,F^{ki} \;=\; -\,F^{0j}F^{ji} \;=\; +\,F^{0j}F^{ij}\\[3pt] &=\; \left(-\frac{E^{j}}{c}\right)\left(-\epsilon^{ijk}B^{k}\right) \;=\; \frac{1}{c}\,\epsilon^{ijk}E^{j}B^{k}. \end{aligned} (2.6.83)

Only λ=j\lambda=j contributes there, since F00=0F^{00}=0. The metric ηjk\eta_{jk} supplies one minus sign and the antisymmetry Fji=FijF^{ji}=-F^{ij} supplies the other, so the two cancel. Putting that into the definition of TμνT^{\mu\nu},

T0i  =  1μ0cϵijkEjBk  =  1c[E×Bμ0]i  =  Sic,S    E×Bμ0. T^{0i} \;=\; \frac{1}{\mu_{0}c}\,\epsilon^{ijk}E^{j}B^{k} \;=\; \frac{1}{c}\left[\frac{\vv E\times\vv B}{\mu_{0}}\right]^{i} \;=\; \frac{S^{i}}{c}, \qquad \vv S \;\equiv\; \frac{\vv E\times\vv B}{\mu_{0}}. (2.6.84)

S\vv S is the Poynting vector. Now recall what T0μ/cT^{0\mu}/c is by construction. It is the density of pμp^{\mu}, with the 00-index labelling the conserved density and the μ\mu labelling which component of momentum we mean. Reading the spatial μ\mu off, the momentum density of the field is

  g  =  T0ice^i  =  Sc2  =  ϵ0E×B.   \boxed{\;\vv g \;=\; \frac{T^{0i}}{c}\hat{\vv e}_{i} \;=\; \frac{\vv S}{c^{2}} \;=\; \epsilon_{0}\,\vv E\times\vv B.\;} (2.6.85)

Which is precisely the expression Chapter 1.1 promised, arriving here as a component of a Noether current rather than as an assertion.

10.3 · The conservation law, with sources

Now restore the charges, because it is the exchange between field and matter that we care about. Take the divergence of (2.6.80):

μ0μTμν  =  (μFμλ)Fλν(1)  +  FμλμFλν(2)  +  12FαβνFαβ(3). \mu_{0}\,\partial_{\mu}T^{\mu\nu} \;=\; \underbrace{\big(\partial_{\mu}F^{\mu\lambda}\big)F_{\lambda}{}^{\nu}}_{(1)} \;+\; \underbrace{F^{\mu\lambda}\,\partial_{\mu}F_{\lambda}{}^{\nu}}_{(2)} \;+\; \underbrace{\tfrac12 F^{\alpha\beta}\,\partial^{\nu}F_{\alpha\beta}}_{(3)}. (2.6.86)

Term (1) is the field equation itself, μFμλ=μ0jλ\partial_{\mu}F^{\mu\lambda}=\mu_{0}j^{\lambda}, so it is μ0jλFλν\mu_{0}j^{\lambda}F_{\lambda}{}^{\nu}. Terms (2) and (3) cancel identically, by the Bianchi identity, and the grind box does that cancellation in full. Dividing through by μ0\mu_{0},

  μTμν  =  jλFλν  =  Fνλjλ.   \boxed{\;\partial_{\mu}T^{\mu\nu} \;=\; j_{\lambda}F^{\lambda\nu} \;=\; -\,F^{\nu\lambda}j_{\lambda}.\;} (2.6.87)

The right-hand side is not zero, and it should not be. The field is exchanging energy and momentum with the charges. Let's see what that exchange term actually is. For ν=i\nu=i, using j0=cρj_{0}=c\rho and jj=Jjj_{j}=-J^{j},

Fiλjλ  =  Eic(cρ)+(ϵijkBk)(Jj)  =  [ρE+J×B]i, F^{i\lambda}j_{\lambda} \;=\; \frac{E^{i}}{c}(c\rho) + \big(-\epsilon^{ijk}B^{k}\big)\big(-J^{j}\big) \;=\; \big[\rho\vv E + \vv J\times\vv B\big]^{i}, (2.6.88)

That is the Lorentz force per unit volume, which is (2.6.67) written for a continuous distribution. Meanwhile the left-hand side, with 0=1ct\partial_{0}=\tfrac1c\partial_{t} and T0i=cgiT^{0i}=cg^{i}, is tgi+jTji\partial_{t}g^{i}+\partial_{j}T^{ji}. Putting the two sides together,

gt  +  T  =  (ρE+J×B)  =  fmech, \pdv{\vv g}{t} \;+\; \nabla\cdot\mathsf{T} \;=\; -\big(\rho\vv E + \vv J\times\vv B\big) \;=\; -\,\vv f_{\text{mech}}, (2.6.89)

Here T\mathsf{T} is the spatial block TijT^{ij}, called the Maxwell stress tensor.

We want a statement about totals rather than densities, so integrate that over all space and take the three terms one at a time. The middle term becomes a surface integral at infinity by the divergence theorem. For a system of charges confined to a finite region the fields fall off at least as fast as 1/r21/r^{2}, so Tij1/r4T^{ij}\sim1/r^{4} while the area grows only as r2r^{2}, and the surface term vanishes. The last term integrates to the total force on the matter, which is dPmech/dt\dd\vv P_{\text{mech}}/\dd t. What is left is

  ddt(Pmech+Pfield)  =  0,Pfield  =  ϵ0 ⁣ ⁣E×B  d3x.   \boxed{\;\dv{}{t}\Big(\vv P_{\text{mech}} + \vv P_{\text{field}}\Big) \;=\; \vv 0, \qquad \vv P_{\text{field}} \;=\; \epsilon_{0}\!\int\!\vv E\times\vv B\;\dd^{3}x.\;} (2.6.90)

The ν=0\nu=0 component of (2.6.87) gives, by the identical route, tu+S=JE\partial_{t}u+\nabla\cdot\vv S=-\vv J\cdot\vv E. That is Poynting's theorem. It says the field energy in a region falls by the flux of S\vv S out of it, plus the work done on the charges inside. The two statements are one four-vector equation.

Grind box — why terms (2) and (3) cancel

Write term (2) with cleaner dummy names, μα\mu\to\alpha and λβ\lambda\to\beta:

(2)  =  FαβαFβν. (2) \;=\; F^{\alpha\beta}\,\partial_{\alpha}F_{\beta}{}^{\nu}.

Because FαβF^{\alpha\beta} is antisymmetric, only the antisymmetric part of whatever it multiplies survives. So we may replace the second factor by half of it minus its αβ\alpha\leftrightarrow\beta swap:

(2)  =  12Fαβ(αFβνβFαν). (2) \;=\; \tfrac12 F^{\alpha\beta}\big(\partial_{\alpha}F_{\beta}{}^{\nu} - \partial_{\beta}F_{\alpha}{}^{\nu}\big).

Now bring in the Bianchi identity (2.6.25), with the third index raised. Raising is legitimate here because we raise it on every term with the same η\eta:

αFβν+βFνα+νFαβ  =  0. \partial_{\alpha}F_{\beta}{}^{\nu} + \partial_{\beta}F^{\nu}{}_{\alpha} + \partial^{\nu}F_{\alpha\beta} \;=\; 0.

Use Fνα=FανF^{\nu}{}_{\alpha}=-F_{\alpha}{}^{\nu} on the middle term and rearrange:

αFβνβFαν  =  νFαβ. \partial_{\alpha}F_{\beta}{}^{\nu} - \partial_{\beta}F_{\alpha}{}^{\nu} \;=\; -\,\partial^{\nu}F_{\alpha\beta}.

Substitute:

(2)  =  12FαβνFαβ  =  (3). (2) \;=\; -\tfrac12 F^{\alpha\beta}\,\partial^{\nu}F_{\alpha\beta} \;=\; -\,(3). \qquad\blacksquare

Note what did the work there. It was the homogeneous half of Maxwell, the half §3.4 called bookkeeping. It is not decorative after all. Without it TμνT^{\mu\nu} would not be conserved, and energy and momentum would not balance.

10.4 · Back to Chapter 1.1

Here is the configuration again, unchanged. Charge q1q_{1} sits at the origin moving with v1=vx^\vv v_{1}=v\hat{\vv x}. Charge q2q_{2} sits at r2=dx^\vv r_{2}=d\hat{\vv x}, directly ahead of it, moving with v2=vy^\vv v_{2}=v\hat{\vv y}. Both are in uniform motion at the instant considered.

What 1.1 found. Charge 1 produces no magnetic field directly ahead of itself, so charge 2 feels no magnetic force. Charge 2 does produce one at charge 1's location, so charge 1 does feel a magnetic force. The magnetic residue was

(F12+F21)mag  =  μ0q1q2v24πd2y^. \big(\vv F_{12}+\vv F_{21}\big)_{\text{mag}} \;=\; -\,\frac{\mu_{0}q_{1}q_{2}v^{2}}{4\pi d^{2}}\,\hat{\vv y}. (2.6.91)

What 1.1 was entitled to drop, and we are not. That chapter said the electric forces "are equal and opposite along the line joining them, so they cancel". That is true to zeroth order in v/cv/c and false at order v2/c2v^{2}/c^{2}, which is the very order the magnetic residue lives at.

Chapter 1.1 was making a qualitative point, so the omission cost it nothing. We are about to make a quantitative one, so we need the correction. It is three lines with §5's transformation rules.

Charge 1's field at charge 2. Boost to charge 1's rest frame SS'. The field point (t=0, x=d)(t=0,\ x=d) has x=γ(dv0)=γdx'=\gamma(d-v\cdot0)=\gamma d, so in SS' the charge sits at a Coulomb distance γd\gamma d and the field there is q1/(4πϵ0γ2d2)q_{1}/(4\pi\epsilon_{0}\gamma^{2}d^{2}), pointing along x^\hat{\vv x}. That direction is parallel to the boost, and by §5 the parallel component EE_{\parallel} is unchanged on transforming back. So

E1(r2)  =  q14πϵ0d21γ2  x^. \vv E_{1}(\vv r_{2}) \;=\; \frac{q_{1}}{4\pi\epsilon_{0}d^{2}}\,\frac{1}{\gamma^{2}}\;\hat{\vv x}. (2.6.92)

Charge 2's field at charge 1. Boost to charge 2's rest frame, which moves along y^\hat{\vv y}. The displacement from charge 2 to the origin, dx^-d\hat{\vv x}, is perpendicular to that boost, so it is unchanged and the Coulomb distance is dd. The field there is q2/(4πϵ0d2)x^-q_{2}/(4\pi\epsilon_{0}d^{2})\,\hat{\vv x}. That direction is also perpendicular to the boost, so by §5 the field gets multiplied by γ\gamma on transforming to the lab:

E2(r1)  =  q2γ4πϵ0d2  x^. \vv E_{2}(\vv r_{1}) \;=\; -\,\frac{q_{2}\gamma}{4\pi\epsilon_{0}d^{2}}\;\hat{\vv x}. (2.6.93)

The two electric forces are q2E1(r2)q_{2}\vv E_{1}(\vv r_{2}) and q1E2(r1)q_{1}\vv E_{2}(\vv r_{1}), and they no longer cancel. One is weakened by γ2\gamma^{-2} and the other is strengthened by γ\gamma, because charge 1 is moving along the line joining them while charge 2 is moving across it. Adding them, and expanding to order β2\beta^{2} with γ2=1β2\gamma^{-2}=1-\beta^{2} and γ1+12β2\gamma\approx1+\tfrac12\beta^{2}, gives the electric residue:

(F12+F21)elec=q1q24πϵ0d2(1γ2γ)x^q1q24πϵ0d2(32β2)x^  =  32μ0q1q2v24πd2x^, \begin{aligned} \big(\vv F_{12}+\vv F_{21}\big)_{\text{elec}} &= \frac{q_{1}q_{2}}{4\pi\epsilon_{0}d^{2}}\Big(\frac{1}{\gamma^{2}} - \gamma\Big)\hat{\vv x}\\[3pt] &\approx \frac{q_{1}q_{2}}{4\pi\epsilon_{0}d^{2}}\Big(-\tfrac32\beta^{2}\Big)\hat{\vv x} \;=\; -\,\frac{3}{2}\,\frac{\mu_{0}q_{1}q_{2}v^{2}}{4\pi d^{2}}\,\hat{\vv x}, \end{aligned} (2.6.94)

The last step used β2/ϵ0=v2/(c2ϵ0)=μ0v2\beta^{2}/\epsilon_{0}=v^{2}/(c^{2}\epsilon_{0})=\mu_{0}v^{2}. So the third law fails in the x^\hat{\vv x} direction too, and by a larger amount than it fails in y^\hat{\vv y}. Adding the two residues, the total rate at which mechanical momentum is being created is

dPmechdt  =  μ0q1q2v24πd2(32x^+y^). \dv{\vv P_{\text{mech}}}{t} \;=\; -\,\frac{\mu_{0}q_{1}q_{2}v^{2}}{4\pi d^{2}}\left(\frac{3}{2}\hat{\vv x} + \hat{\vv y}\right). (2.6.95)

10.5 · And where it went

Now we compute the field momentum, to see whether it is disappearing at the same rate. Two simplifications are available at the order we need. Each charge's magnetic field is Ba=va×Ea/c2\vv B_{a}=\vv v_{a}\times\vv E_{a}/c^{2}, which §5 gives exactly, by boosting a rest frame where B=0\vv B=\vv 0. And each Ea\vv E_{a} may be taken as Coulomb, since E×B\vv E\times\vv B is already first order in v/cv/c while corrections to E\vv E are second. With the fields written as sums of the two charges' contributions, (2.6.90) expands to

Pfield  =  ϵ0 ⁣ ⁣(E1+E2)×(B1+B2)d3x. \vv P_{\text{field}} \;=\; \epsilon_{0}\!\int\!\big(\vv E_{1}+\vv E_{2}\big)\times\big(\vv B_{1}+\vv B_{2}\big)\,\dd^{3}x. (2.6.96)

The two self terms Ea×Ba\vv E_{a}\times\vv B_{a} depend only on va\vv v_{a}, which is constant. So however large they are, and they are in fact infinite, as the caution below explains, they contribute nothing to dP/dt\dd\vv P/\dd t. That leaves the cross terms. Expanding each of those with a×(b×c)=b(ac)c(ab)\vv a\times(\vv b\times\vv c)=\vv b(\vv a\cdot\vv c)-\vv c(\vv a\cdot\vv b),

Pfieldint  =  ϵ0c2 ⁣ ⁣[(v1+v2)(E1 ⁣ ⁣E2)    E2(E1 ⁣ ⁣v2)    E1(E2 ⁣ ⁣v1)]d3x. \vv P_{\text{field}}^{\text{int}} \;=\; \frac{\epsilon_{0}}{c^{2}}\!\int\!\Big[\big(\vv v_{1}+\vv v_{2}\big)\big(\vv E_{1}\!\cdot\!\vv E_{2}\big) \;-\; \vv E_{2}\big(\vv E_{1}\!\cdot\!\vv v_{2}\big) \;-\; \vv E_{1}\big(\vv E_{2}\!\cdot\!\vv v_{1}\big)\Big]\dd^{3}x. (2.6.97)

Two integrals are needed to finish this, and both are elementary. The grind box evaluates them. Writing d=r2r1\vv d=\vv r_{2}-\vv r_{1}, d=dd=\abs{\vv d} and n^=d/d\hat{\vv n}=\vv d/d, they are

 ⁣E1 ⁣ ⁣E2  d3x  =  q1q24πϵ02d, ⁣E1iE2j  d3x  =  q1q28πϵ02d(δijninj). \int\!\vv E_{1}\!\cdot\!\vv E_{2}\;\dd^{3}x \;=\; \frac{q_{1}q_{2}}{4\pi\epsilon_{0}^{2}\,d}, \qquad\quad \int\! E_{1}^{i}E_{2}^{j}\;\dd^{3}x \;=\; \frac{q_{1}q_{2}}{8\pi\epsilon_{0}^{2}\,d}\Big(\delta^{ij}-n^{i}n^{j}\Big). (2.6.98)

Now substitute those two results back. Write Vv1+v2\vv V\equiv\vv v_{1}+\vv v_{2}. The two tensor terms combine into MijVjM^{ij}V^{j}, using the symmetry of MijE1iE2jd3xM^{ij}\equiv\int E_{1}^{i}E_{2}^{j}\,\dd^{3}x, and the prefactors collapse with 1/(ϵ0c2)=μ01/(\epsilon_{0}c^{2})=\mu_{0}:

Pfieldint=ϵ0c2q1q24πϵ02d[V12(Vn^(n^ ⁣ ⁣V))]=  μ0q1q28πd[V+n^(n^ ⁣ ⁣V)].   \begin{aligned} \vv P_{\text{field}}^{\text{int}} &= \frac{\epsilon_{0}}{c^{2}}\cdot\frac{q_{1}q_{2}}{4\pi\epsilon_{0}^{2}d}\left[\vv V - \tfrac12\Big(\vv V - \hat{\vv n}\big(\hat{\vv n}\!\cdot\!\vv V\big)\Big)\right]\\[4pt] &= \boxed{\;\frac{\mu_{0}q_{1}q_{2}}{8\pi d}\Big[\vv V + \hat{\vv n}\big(\hat{\vv n}\!\cdot\!\vv V\big)\Big].\;} \end{aligned} (2.6.99)

That is the general answer. For the configuration at hand, at the instant t=0t=0, we have n^=x^\hat{\vv n}=\hat{\vv x}, V=v(x^+y^)\vv V=v(\hat{\vv x}+\hat{\vv y}) and n^V=v\hat{\vv n}\cdot\vv V=v, so

Pfield  =  μ0q1q2v8πd(2x^+y^)    0. \vv P_{\text{field}} \;=\; \frac{\mu_{0}q_{1}q_{2}v}{8\pi d}\big(2\hat{\vv x}+\hat{\vv y}\big) \;\neq\; \vv 0. (2.6.100)

There is momentum in the empty space between the charges.

What we actually want is the rate at which that momentum changes, so differentiate. The separation vector evolves as d(t)=(dvt)x^+vty^\vv d(t)=(d-vt)\hat{\vv x}+vt\,\hat{\vv y}, since charge 1 moves in +x+x and charge 2 moves in +y+y. At t=0t=0 that gives d˙=v\dot d=-v and n^˙=(v/d)y^\dot{\hat{\vv n}}=(v/d)\hat{\vv y}. Grinding through the product rule, as the grind box does,

dPfielddtt=0  =  μ0q1q2v24πd2(32x^+y^). \dv{\vv P_{\text{field}}}{t}\bigg|_{t=0} \;=\; \frac{\mu_{0}q_{1}q_{2}v^{2}}{4\pi d^{2}}\left(\frac{3}{2}\hat{\vv x} + \hat{\vv y}\right). (2.6.101)

Now set that beside (2.6.95), which is the rate at which the particles were gaining momentum from nowhere. The two match term by term, including the awkward 32\tfrac32 that Chapter 1.1 never saw. So the books balance exactly:

  dPmechdt  +  dPfielddt  =  0.   \boxed{\;\dv{\vv P_{\text{mech}}}{t} \;+\; \dv{\vv P_{\text{field}}}{t} \;=\; \vv 0.\;} (2.6.102)
The longest-running promise in the book, closed

Chapter 1.1 §3.3 wrote down a momentum non-conservation of μ0q1q2v2y^/4πd2-\mu_{0}q_{1}q_{2}v^{2}\hat{\vv y}/4\pi d^{2} and offered you a choice. Either abandon momentum conservation, or accept that the field is a physical object with momentum of its own.

The second was right, and now it is computed. The missing momentum is (2.6.100), it lives in the field, and its rate of change is exactly minus the rate at which the particles are gaining momentum from nowhere. That holds in both components, including the 32\tfrac32 that only shows up if you are honest about the electric forces at order v2/c2v^{2}/c^{2}.

Notice how the repair works. Newton's third law presumed two things: that all the momentum in the universe is carried by particles, and that particles exchange it instantaneously. Relativity forbids the second of those, and once it goes the first cannot survive either. If charge 1 moves now and charge 2 will not know for d/cd/c seconds, the momentum has to be somewhere during the interval.

So the third law is not a fundamental principle that electromagnetism violates. It is a low-velocity approximation to a four-vector conservation law, and it fails exactly where the approximation does.

The conceptual bill is large and worth paying explicitly. A field can no longer be read as a bookkeeping device, a table of where the force would be if you put a test charge there. Something that stores energy, stores momentum, exerts stresses, and can carry momentum across a room during the interval when neither particle has it, is not a table. It is a physical system with its own degrees of freedom and its own dynamics.

That is why Part V quantises it, and why the resulting quanta, photons, are as real as electrons rather than being a way of talking about electrons.

⚠ Two honest cautions

The self terms are infinite. ϵ0Ea×Bad3x\epsilon_{0}\int\vv E_{a}\times\vv B_{a}\,\dd^{3}x diverges at the location of a point charge, as does the self-energy ϵ02Ea2\tfrac{\epsilon_{0}}{2}\int E_{a}^{2}. We dodged it above by noting that these terms are constant when va\vv v_{a} is, so they drop out of the rate.

That dodge is legitimate here and evasive in general. The divergence is the classical electron self-energy problem. It is genuinely unsolved in classical physics, and it is the ancestor of the renormalisation programme of Chapter 5.11. Point charges are an idealisation, and they bite.

What we computed and what we did not. The general theorem (2.6.90) is exact, derived from μTμν\partial_{\mu}T^{\mu\nu} with no approximation. The two-charge check is narrower. It keeps the velocity-dependent terms to order v2/c2v^{2}/c^{2} and drops the acceleration fields, the radiation ones. Their contribution is balanced separately against the v˙\dot{\vv v} terms in (2.6.99), which is a second and independent piece of bookkeeping, controlled by a different small parameter: the ratio of the charges' Coulomb energy to their kinetic energy.

Nothing was assumed in order to make (2.6.102) come out. Every coefficient was computed first and compared afterwards.

Grind box — the two field integrals of (2.6.98)

Both are done with the same two moves. Integrate by parts, then use Gauss's law Ea=ρa/ϵ0\nabla\cdot\vv E_{a}=\rho_{a}/\epsilon_{0} with ρa=qaδ3(xra)\rho_{a}=q_{a}\delta^{3}(\vv x-\vv r_{a}). Surface terms vanish because the fields fall off as 1/r21/r^{2} and the integrands as 1/r31/r^{3} or faster.

The scalar one. Write E1=ϕ1\vv E_{1}=-\nabla\phi_{1} and integrate by parts:

 ⁣E1 ⁣ ⁣E2d3x= ⁣ ⁣ϕ1E2d3x= ⁣ϕ1( ⁣ ⁣E2)d3x=q2ϵ0ϕ1(r2)=q1q24πϵ02d. \int\!\vv E_{1}\!\cdot\!\vv E_{2}\,\dd^{3}x = -\!\int\!\nabla\phi_{1}\cdot\vv E_{2}\,\dd^{3}x = \int\!\phi_{1}\,\big(\nabla\!\cdot\!\vv E_{2}\big)\,\dd^{3}x = \frac{q_{2}}{\epsilon_{0}}\,\phi_{1}(\vv r_{2}) = \frac{q_{1}q_{2}}{4\pi\epsilon_{0}^{2}d}.

Multiplied by ϵ0\epsilon_{0}, that is the familiar interaction energy q1q2/4πϵ0dq_{1}q_{2}/4\pi\epsilon_{0}d, which is a useful check that the normalisation is right.

The tensor one. Call it Mij(d)M^{ij}(\vv d). Three facts pin it down.

(i) It is symmetric. Invert through the midpoint, xr1+r2x\vv x\to\vv r_{1}+\vv r_{2}-\vv x. Under that map xr1(xr2)\vv x-\vv r_{1}\mapsto-(\vv x-\vv r_{2}), so E1(q1/q2)E2\vv E_{1}\mapsto-(q_{1}/q_{2})\vv E_{2} and E2(q2/q1)E1\vv E_{2}\mapsto-(q_{2}/q_{1})\vv E_{1}, and the Jacobian is 11. Hence Mij=MjiM^{ij}=M^{ji}.

(ii) Its form. The only vectors in the problem are n^\hat{\vv n} and the coordinate axes, so by rotational symmetry about n^\hat{\vv n} and (i),

Mij  =  Adδij  +  Bdninj, M^{ij} \;=\; \frac{A}{d}\,\delta^{ij} \;+\; \frac{B}{d}\,n^{i}n^{j},

with AA and BB constants. The 1/d1/d is there because MM has the dimensions of the scalar integral, which scales as 1/d1/d.

(iii) Two equations. The trace is the scalar integral just computed:

3A+B  =  q1q24πϵ02. 3A + B \;=\; \frac{q_{1}q_{2}}{4\pi\epsilon_{0}^{2}}.

For the second, differentiate with respect to d\vv d and contract. Since E2\vv E_{2} depends on xr2\vv x-\vv r_{2} and r2=r1+d\vv r_{2}=\vv r_{1}+\vv d, we have E2j/dk=kE2j\partial E_{2}^{j}/\partial d^{k}=-\partial_{k}E_{2}^{j}, so

Mijdj= ⁣ ⁣E1ijE2jd3x=q2ϵ0E1i(r2)=q1q24πϵ02nid2. \pdv{M^{ij}}{d^{j}} = -\!\int\! E_{1}^{i}\,\partial_{j}E_{2}^{j}\,\dd^{3}x = -\frac{q_{2}}{\epsilon_{0}}E_{1}^{i}(\vv r_{2}) = -\frac{q_{1}q_{2}}{4\pi\epsilon_{0}^{2}}\frac{n^{i}}{d^{2}}.

Evaluate the same derivative on the parametrised form, using j(1/d)=nj/d2\partial_{j}(1/d)=-n^{j}/d^{2} and j(didj/d3)=ni/d2\partial_{j}\big(d^{i}d^{j}/d^{3}\big)=n^{i}/d^{2}:

dj[Aδij+Bninjd]=Anid2+Bnid2=(BA)nid2. \pdv{}{d^{j}}\left[\frac{A\delta^{ij}+Bn^{i}n^{j}}{d}\right] = -\frac{A\,n^{i}}{d^{2}} + \frac{B\,n^{i}}{d^{2}} = \frac{(B-A)\,n^{i}}{d^{2}}.

Comparing, BA=q1q2/4πϵ02B-A=-q_{1}q_{2}/4\pi\epsilon_{0}^{2}. Solve the two equations: A=q1q2/8πϵ02A=q_{1}q_{2}/8\pi\epsilon_{0}^{2} and B=AB=-A. Hence

Mij  =  q1q28πϵ02d(δijninj), M^{ij} \;=\; \frac{q_{1}q_{2}}{8\pi\epsilon_{0}^{2}d}\big(\delta^{ij}-n^{i}n^{j}\big),

which is (2.6.98). Note the structure. Here δnn\delta-nn is the projector onto directions perpendicular to the line joining the charges, so Mijnj=0M^{ij}n^{j}=0, and the longitudinal part of the integral vanishes identically. This was also checked by direct numerical integration before being written down.

Grind box — differentiating (2.6.99)

Write C=μ0q1q2/8πC=\mu_{0}q_{1}q_{2}/8\pi and V=v(x^+y^)\vv V=v(\hat{\vv x}+\hat{\vv y}), a constant. Then

Pfield  =  C[Vd+n^(n^V)d]. \vv P_{\text{field}} \;=\; C\left[\frac{\vv V}{d} + \frac{\hat{\vv n}\,(\hat{\vv n}\cdot\vv V)}{d}\right].

The kinematics at t=0t=0. d(t)=r2r1=(dvt)x^+vty^\vv d(t)=\vv r_{2}-\vv r_{1}=(d-vt)\hat{\vv x}+vt\,\hat{\vv y}, so d˙=v(y^x^)\dot{\vv d}=v(\hat{\vv y}-\hat{\vv x}) and

d˙=n^d˙=x^v(y^x^)=v,n^˙=d˙ddd˙d2=v(y^x^)d+vx^d=vdy^. \dot d = \hat{\vv n}\cdot\dot{\vv d} = \hat{\vv x}\cdot v(\hat{\vv y}-\hat{\vv x}) = -v, \qquad \dot{\hat{\vv n}} = \frac{\dot{\vv d}}{d} - \frac{\vv d\,\dot d}{d^{2}} = \frac{v(\hat{\vv y}-\hat{\vv x})}{d} + \frac{v\hat{\vv x}}{d} = \frac{v}{d}\hat{\vv y}.

Note n^˙n^\dot{\hat{\vv n}}\perp\hat{\vv n}, as it must be for a unit vector.

First term.

ddtVd=d˙d2V=vd2v(x^+y^)=v2d2(x^+y^). \dv{}{t}\frac{\vv V}{d} = -\frac{\dot d}{d^{2}}\vv V = \frac{v}{d^{2}}\,v(\hat{\vv x}+\hat{\vv y}) = \frac{v^{2}}{d^{2}}\big(\hat{\vv x}+\hat{\vv y}\big).

Second term. Three pieces by the product rule:

ddt[n^(n^V)d]=d˙d2n^(n^V)+n^˙(n^V)d+n^(n^˙V)d=vd2x^(v)  +  (v/d)y^(v)d  +  x^((v/d)y^v(x^+y^))d=v2d2x^+v2d2y^+v2d2x^  =  v2d2(2x^+y^). \begin{aligned} \dv{}{t}\left[\frac{\hat{\vv n}(\hat{\vv n}\cdot\vv V)}{d}\right] &= -\frac{\dot d}{d^{2}}\hat{\vv n}(\hat{\vv n}\cdot\vv V) + \frac{\dot{\hat{\vv n}}(\hat{\vv n}\cdot\vv V)}{d} + \frac{\hat{\vv n}(\dot{\hat{\vv n}}\cdot\vv V)}{d}\\[4pt] &= \frac{v}{d^{2}}\,\hat{\vv x}\,(v) \;+\; \frac{(v/d)\hat{\vv y}\,(v)}{d} \;+\; \frac{\hat{\vv x}\,\big((v/d)\hat{\vv y}\cdot v(\hat{\vv x}+\hat{\vv y})\big)}{d}\\[4pt] &= \frac{v^{2}}{d^{2}}\hat{\vv x} + \frac{v^{2}}{d^{2}}\hat{\vv y} + \frac{v^{2}}{d^{2}}\hat{\vv x} \;=\; \frac{v^{2}}{d^{2}}\big(2\hat{\vv x}+\hat{\vv y}\big). \end{aligned}

Total.

dPfielddt=Cv2d2[(x^+y^)+(2x^+y^)]=μ0q1q2v28πd2(3x^+2y^), \dv{\vv P_{\text{field}}}{t} = C\frac{v^{2}}{d^{2}}\Big[(\hat{\vv x}+\hat{\vv y}) + (2\hat{\vv x}+\hat{\vv y})\Big] = \frac{\mu_{0}q_{1}q_{2}v^{2}}{8\pi d^{2}}\big(3\hat{\vv x}+2\hat{\vv y}\big),

which is (2.6.101) after pulling out a factor of 22: 18π(3x^+2y^)=14π(32x^+y^)\tfrac{1}{8\pi}(3\hat{\vv x}+2\hat{\vv y}) = \tfrac{1}{4\pi}\big(\tfrac32\hat{\vv x}+\hat{\vv y}\big). ✓ Every term of this box, and of (2.6.94), was verified symbolically. The sum (2.6.102) is zero identically rather than numerically.

Where TμνT^{\mu\nu} goes next

You have built a symmetric, conserved, traceless rank-2 tensor whose 0000 component is energy density and whose 0i0i components are momentum density. Chapter 3.6 will take Einstein's field equation

Gμν  =  8πGc4Tμν G^{\mu\nu} \;=\; \frac{8\pi G}{c^{4}}\,T^{\mu\nu}

and put this object on the right-hand side. That is what "energy gravitates" means technically. It is not mass but TμνT^{\mu\nu} that sources spacetime curvature, so the electromagnetic field itself gravitates, with a strength you can now compute. That includes light, a magnetic field, and the energy stored in a capacitor.

It is also why the equation had to be symmetric in μν\mu\nu, which is why §10.1's improvement term was not optional.

And the tracelessness ημνTμν=0\eta_{\mu\nu}T^{\mu\nu}=0, which you can verify in one line from (2.6.80), is the statement that the photon is massless. It reappears in Chapter 5.11 as the conformal symmetry of classical electromagnetism, a symmetry that quantum corrections break.

In plain terms 2.6.10

The oldest debt in the book is settled by computation rather than assertion. Two charges in uniform motion, alone in the universe, were seen in the first part to push on each other unequally, momentum appearing from nowhere. Work out what the empty space between them holds and differentiate: the field loses momentum at exactly the rate the particles gain it, awkward factor of three halves and all.

The diagnosis matters more than the arithmetic. The third law assumed all momentum belongs to particles, handed over the instant either moves. Relativity forbids the second clause: once one charge moves and the other cannot learn of it until light crosses the gap, the momentum must be somewhere meanwhile. What holds it there is no table of where a force would be but a physical system with its own degrees of freedom, which is why a later part must quantise it.

Part II has done its work. You hold a geometry where one speed is the same for everybody with nothing to measure it against, a language making frame-independence visible in an equation's shape, a mechanics of four-entry objects, which is where energy came from, and one field with electricity and magnetism as its slices. One thing has not moved: gravity is still a force reaching across empty space, arriving the moment it is sent, and nothing in this part permits an influence with no delay.

11 · Worked examples

Worked example 1 — the wire, in full numbers

A copper wire of cross-section 1 mm21\ \mathrm{mm^{2}} carries 1 A1\ \mathrm{A}. A proton (q=+eq=+e) travels parallel to it at u=105 ms1u=10^{5}\ \mathrm{m\,s^{-1}}, at a perpendicular distance r=1 cmr=1\ \mathrm{cm}. Compute the force on it twice, first magnetically in the lab and then electrostatically in its own rest frame. Then say how big the charge imbalance is.

The lab. The field of the wire at 1 cm1\ \mathrm{cm} is

B=μ0I2πr=(4π×107)(1)2π(0.01)=2.000×105 T, B = \frac{\mu_{0}I}{2\pi r} = \frac{(4\pi\times10^{-7})(1)}{2\pi(0.01)} = 2.000\times10^{-5}\ \mathrm{T},

about a third of the Earth's field. The magnetic force is

Flab=quB=(1.602×1019)(105)(2.000×105)=3.204×1019 N, F_{\text{lab}} = quB = (1.602\times10^{-19})(10^{5})(2.000\times10^{-5}) = 3.204\times10^{-19}\ \mathrm{N},

directed toward the wire. There is no electric force, because the wire is neutral in this frame.

The wire's insides. From (2.6.49), n=8.49×1028 m3n=8.49\times10^{28}\ \mathrm{m^{-3}}, so the mobile charge per metre is

λ0=neA=(8.49×1028)(1.602×1019)(106)=1.360×104 Cm1, \lambda_{0} = neA = (8.49\times10^{28})(1.602\times10^{-19})(10^{-6}) = 1.360\times10^{4}\ \mathrm{C\,m^{-1}},

and the drift speed is vd=I/λ0=7.351×105 ms1v_{d}=I/\lambda_{0}=7.351\times10^{-5}\ \mathrm{m\,s^{-1}}, i.e. 0.0735 mms10.0735\ \mathrm{mm\,s^{-1}}.

The proton's frame. Here u/c=3.34×104u/c=3.34\times10^{-4}, so γu=1.000000056\gamma_{u}=1.000000056. That is utterly negligible, and we keep it anyway, because the whole effect we are chasing is of this size. From (2.6.45),

λ  =  γuλ0uvdc2  =  (1.360×104)(105)(7.351×105)8.988×1016=  1.113×1012 Cm1. \begin{aligned} \lambda' \;&=\; -\gamma_{u}\lambda_{0}\frac{uv_{d}}{c^{2}} \;=\; -\frac{(1.360\times10^{4})(10^{5})(7.351\times10^{-5})}{8.988\times10^{16}}\\[3pt] &=\; -1.113\times10^{-12}\ \mathrm{C\,m^{-1}}. \end{aligned}

The electric field of that line charge at r=1 cmr=1\ \mathrm{cm} is

E=λ2πϵ0r=1.113×10122π(8.854×1012)(0.01)=2.000 Vm1, E' = \frac{\lambda'}{2\pi\epsilon_{0}r} = \frac{-1.113\times10^{-12}}{2\pi(8.854\times10^{-12})(0.01)} = -2.000\ \mathrm{V\,m^{-1}},

pointing toward the wire, and the force on the proton is

F=qE=(1.602×1019)(2.000)=3.204×1019 N. F' = qE' = (1.602\times10^{-19})(2.000) = 3.204\times10^{-19}\ \mathrm{N}.

Compare. F/Flab=1.000000056=γuF'/F_{\text{lab}} = 1.000000056 = \gamma_{u} ✓, exactly as (2.6.47) requires. Notice also the clean intermediate result E=uBE'=-uB to this accuracy. That is Ey=γ(EyuBz)E'_{y}=\gamma(E_{y}-uB_{z}) with Ey=0E_{y}=0, which is the field transformation rule of §5 arriving by a completely independent route.

How big is the imbalance? Count it in electrons per metre:

λe=6.94×106 m1,againstλ0e=8.49×1022 m1. \frac{\abs{\lambda'}}{e} = 6.94\times10^{6}\ \mathrm{m^{-1}}, \qquad\text{against}\qquad \frac{\lambda_{0}}{e} = 8.49\times10^{22}\ \mathrm{m^{-1}}.

That is a fractional imbalance of 8.2×10178.2\times10^{-17}, or one part in twelve million billion. And that is the entire magnetic force on this proton, a discrepancy in the seventeenth significant figure of a cancellation. If the positive and negative charge in a wire did not cancel to far better than that, the residual electrostatic force would swamp every magnetic effect ever measured, and nobody would have discovered magnetism in a laboratory at all.

Worked example 2 — a boost that eliminates the electric field

In the lab, E=E0y^\vv E=E_{0}\hat{\vv y} and B=B0z^\vv B=B_{0}\hat{\vv z} are uniform, crossed, and satisfy E0<cB0E_{0}\lt cB_{0}. Find a frame in which the electric field vanishes, describe the motion of a charge in that frame, and identify the frame velocity in a form that does not refer to the axes.

Check first that it is possible. The invariants of §7 are B02E02/c2>0B_{0}^{2}-E_{0}^{2}/c^{2}\gt0 and EB=0\vv E\cdot\vv B=0, which is precisely condition (2.6.62). So a frame with E=0\vv E'=\vv 0 is permitted. Now construct it.

The construction. Boost along x^\hat{\vv x} at speed vv. From §5, Ex=Ex=0E'_{x}=E_{x}=0 and Ez=γ(Ez+vBy)=0E'_{z}=\gamma(E_{z}+vB_{y})=0 automatically. The only surviving component is

Ey=γ(E0vB0), E'_{y} = \gamma\big(E_{0} - vB_{0}\big),

and since γ\gamma is never zero, E=0\vv E'=\vv 0 requires

  v=E0B0.   \boxed{\;v = \frac{E_{0}}{B_{0}}.\;}

That speed is subluminal precisely when E0<cB0E_{0}\lt cB_{0}. There is the invariant condition again, arriving this time as a kinematic constraint rather than as a separate assumption. Note in passing what happens if E0>cB0E_{0}\gt cB_{0}. The required vv then exceeds cc, so there is no such frame, and the invariant told you so in advance without any construction at all.

Axis-free form. With E×B=E0B0(y^×z^)=E0B0x^\vv E\times\vv B = E_{0}B_{0}(\hat{\vv y}\times\hat{\vv z}) = E_{0}B_{0}\hat{\vv x} and B2=B02B^{2}=B_{0}^{2},

E×BB2=E0B0x^=v. \frac{\vv E\times\vv B}{B^{2}} = \frac{E_{0}}{B_{0}}\hat{\vv x} = \vv v.

So the frame in question moves with the drift velocity

  vdrift  =  E×BB2   \boxed{\;\vv v_{\text{drift}} \;=\; \frac{\vv E\times\vv B}{B^{2}}\;}

That formula makes no reference to any choice of axes, and it is valid whenever EB\vv E\perp\vv B and E<cBE\lt cB.

What the motion looks like. In the primed frame there is only a magnetic field, of magnitude

B=γ(B0vc2E0)=γB0(1E02B02c2)=B0γ=B02E02c2, B' = \gamma\left(B_{0} - \frac{v}{c^{2}}E_{0}\right) = \gamma B_{0}\left(1-\frac{E_{0}^{2}}{B_{0}^{2}c^{2}}\right) = \frac{B_{0}}{\gamma} = \sqrt{B_{0}^{2}-\frac{E_{0}^{2}}{c^{2}}},

using v=E0/B0v=E_{0}/B_{0} and γ2(1v2/c2)=1\gamma^{2}(1-v^{2}/c^{2})=1. The last expression is the square root of the first invariant divided by two, as it has to be. With E=0\vv E'=\vv 0 the invariant (2.6.56) reads 2B22B'^{2}, and that must agree with 2(B02E02/c2)2(B_{0}^{2}-E_{0}^{2}/c^{2}).

A charge in a pure magnetic field feels a force always perpendicular to its velocity and does no work, by (2.6.69). So it moves in a circle at constant speed, which is the cyclotron gyration.

Back in the lab, that circle is being carried along at vdrift\vv v_{\text{drift}}, so the trajectory is a cycloid-like drift. Here is the point to take away. The drift velocity does not depend on the charge, the sign of the charge, or the mass. Every particle drifts at E×B/B2\vv E\times\vv B/B^{2}.

Why every mass spectrometer contains one. Turn the argument around. A particle that enters crossed fields moving at exactly vdrift\vv v_{\text{drift}} is at rest in the primed frame, and a charge at rest in a pure magnetic field feels no force at all. So it passes through undeflected, while anything faster or slower is bent aside.

That is a velocity selector: a slit, two plates, a magnet, and one equation. It selects v=E/Bv=E/B regardless of mass and charge, which is exactly what you want in front of a mass analyser that will then separate by m/qm/q. The same E×B\vv E\times\vv B drift governs charged particles in the magnetosphere, the confinement of tokamak plasmas, and the Hall effect.

12 · Your turn

Problem 1 · the field of a charge in uniform motion

A charge qq moves with constant velocity vx^v\hat{\vv x}, passing the origin at t=0t=0. By boosting the Coulomb field from its rest frame, show that at t=0t=0 the lab field at position r=(x,y,z)\vv r=(x,y,z) is

E  =  q4πϵ0rr31β2(1β2sin2θ)3/2,B  =  1c2v×E, \vv E \;=\; \frac{q}{4\pi\epsilon_{0}}\,\frac{\vv r}{r^{3}}\,\frac{1-\beta^{2}}{\big(1-\beta^{2}\sin^{2}\theta\big)^{3/2}}, \qquad \vv B \;=\; \frac{1}{c^{2}}\,\vv v\times\vv E,

where θ\theta is the angle between r\vv r and v\vv v. Show the field is still radial from the present position, that it is weakened by γ2\gamma^{2} directly ahead and strengthened by γ\gamma broadside, and that for β1\beta\ll1 the magnetic field reduces to the Biot–Savart expression Chapter 1.1 had to quote. Comment on what the field looks like as γ\gamma\to\infty.

Solution

Set-up. Let SS' be the charge's rest frame, moving at +vx^+v\hat{\vv x}. There B=0\vv B'=\vv 0 and E=q4πϵ0rr3\vv E'=\dfrac{q}{4\pi\epsilon_{0}}\dfrac{\vv r'}{r'^{3}}.

Coordinates. The lab event is (ct,x,y,z)=(0,x,y,z)(ct,x,y,z)=(0,x,y,z). By the Lorentz transformation, x=γ(xvt)=γxx'=\gamma(x-vt)=\gamma x, y=yy'=y, z=zz'=z, so

r=γ2x2+y2+z2. r' = \sqrt{\gamma^{2}x^{2}+y^{2}+z^{2}}.

Fields. Transform back from SS' to the lab. Parallel components are unchanged and perpendicular ones acquire a γ\gamma (with B=0\vv B'=\vv 0 the v×B\vv v\times\vv B terms drop):

Ex=Ex=q4πϵ0γxr3,Ey=γEy=q4πϵ0γyr3,Ez=q4πϵ0γzr3. E_{x}=E'_{x}=\frac{q}{4\pi\epsilon_{0}}\frac{\gamma x}{r'^{3}}, \qquad E_{y}=\gamma E'_{y}=\frac{q}{4\pi\epsilon_{0}}\frac{\gamma y}{r'^{3}}, \qquad E_{z}=\frac{q}{4\pi\epsilon_{0}}\frac{\gamma z}{r'^{3}}.

All three carry the same factor γ/r3\gamma/r'^{3}, so

E=qγ4πϵ0rr3. \vv E = \frac{q\gamma}{4\pi\epsilon_{0}}\frac{\vv r}{r'^{3}}.

The field points radially away from the charge's present position. That is a small miracle, given that information travels at cc and the charge has moved since the field "left". It works only for uniform motion. Accelerate the charge and the radial structure breaks, which is what radiation is.

The angular factor. Write x=rcosθx=r\cos\theta and y2+z2=r2sin2θy^{2}+z^{2}=r^{2}\sin^{2}\theta:

r2=γ2r2cos2θ+r2sin2θ=γ2r2[cos2θ+sin2θγ2]=γ2r2[1β2sin2θ], r'^{2} = \gamma^{2}r^{2}\cos^{2}\theta + r^{2}\sin^{2}\theta = \gamma^{2}r^{2}\Big[\cos^{2}\theta + \frac{\sin^{2}\theta}{\gamma^{2}}\Big] = \gamma^{2}r^{2}\big[1-\beta^{2}\sin^{2}\theta\big],

using 1/γ2=1β21/\gamma^{2}=1-\beta^{2} and cos2+sin2=1\cos^{2}+\sin^{2}=1. Then r3=γ3r3(1β2sin2θ)3/2r'^{3}=\gamma^{3}r^{3}(1-\beta^{2}\sin^{2}\theta)^{3/2} and

E=q4πϵ0rr3γγ3(1β2sin2θ)3/2=q4πϵ0rr31β2(1β2sin2θ)3/2. \vv E = \frac{q}{4\pi\epsilon_{0}}\frac{\vv r}{r^{3}}\cdot\frac{\gamma}{\gamma^{3}\big(1-\beta^{2}\sin^{2}\theta\big)^{3/2}} = \frac{q}{4\pi\epsilon_{0}}\frac{\vv r}{r^{3}}\cdot\frac{1-\beta^{2}}{\big(1-\beta^{2}\sin^{2}\theta\big)^{3/2}}. \quad\checkmark

The two special directions. Ahead or behind (θ=0,π\theta=0,\pi): the factor is 1β2=1/γ21-\beta^{2}=1/\gamma^{2}, so the field is weaker than Coulomb by γ2\gamma^{2}. Broadside (θ=π/2\theta=\pi/2): the factor is (1β2)/(1β2)3/2=(1β2)1/2=γ(1-\beta^{2})/(1-\beta^{2})^{3/2}=(1-\beta^{2})^{-1/2}=\gamma, so the field is stronger by γ\gamma. These are exactly (2.6.92) and (2.6.93), which §10 needed and derived by the same argument in two special cases.

The magnetic field. Transforming B\vv B from a frame where it vanishes gives By=γvEz/c2B_{y}=-\gamma vE'_{z}/c^{2} and Bz=+γvEy/c2B_{z}=+\gamma vE'_{y}/c^{2}, which in view of the expressions above is exactly B=v×E/c2\vv B=\vv v\times\vv E/c^{2}. That is exact rather than a leading order statement. For β1\beta\ll1 we may replace E\vv E by its Coulomb value:

B1c2v×qr^4πϵ0r2=μ04πqv×r^r2, \vv B \approx \frac{1}{c^{2}}\vv v\times\frac{q\hat{\vv r}}{4\pi\epsilon_{0}r^{2}} = \frac{\mu_{0}}{4\pi}\frac{q\,\vv v\times\hat{\vv r}}{r^{2}},

using 1/(ϵ0c2)=μ01/(\epsilon_{0}c^{2})=\mu_{0}. That is the Biot–Savart law for a point charge, which was the second of the two results Chapter 1.1 §3.3 had to quote. It is now derived, so both of that chapter's ⚑ items are discharged.

The pancake. As γ\gamma\to\infty the forward field dies like 1/γ21/\gamma^{2} while the broadside field grows like γ\gamma. The angular width over which the field is appreciable is set by β2sin2θ1\beta^{2}\sin^{2}\theta\approx1, that is, by sinθ1/γ\sin\theta\approx1/\gamma. So the field collapses into a disc perpendicular to the motion, of angular thickness 1/γ\sim1/\gamma, travelling with the charge. At the LHC, γ7000\gamma\approx7000 for protons, so each beam's field is squashed into a pancake about 10410^{-4} radians thick. That is why an ultrarelativistic charged particle acts on a target like a pulse rather than a slowly growing force, and it is the starting point for the equivalent-photon approximation used to describe ultraperipheral collisions in Part V.

Problem 2 · a plane wave has nothing to give up

For a plane electromagnetic wave in vacuum, EB\vv E\perp\vv B and E=cBE=cB. Show that both invariants of §7 vanish. Then use that to prove three things. No boost can eliminate either field. No boost can make the fields non-perpendicular. And no boost can bring the wave to rest. What can a boost do to a light wave?

Solution

The invariants. FμνFμν=2(B2E2/c2)=2(B2B2)=0F_{\mu\nu}F^{\mu\nu}=2(B^{2}-E^{2}/c^{2})=2(B^{2}-B^{2})=0, and ϵμνρσFμνFρσ=(8/c)EB=0\epsilon_{\mu\nu\rho\sigma}F^{\mu\nu}F^{\rho\sigma}=(8/c)\vv E\cdot\vv B=0 since the fields are perpendicular. Both are zero. By the invariance theorem they are zero in every frame.

(i) Neither field can be eliminated. Suppose some frame had B=0\vv B'=\vv 0. Then in that frame FμνFμν=2E2/c2F_{\mu\nu}F^{\mu\nu}=-2E'^{2}/c^{2}, which must equal zero, forcing E=0E'=0 as well. So the only way to lose the magnetic field is to lose the whole wave. And you cannot do that, since a tensor that vanishes in one frame vanishes in all of them (Chapter 2.4 §6), so a wave that exists at all exists for everyone. The same argument runs identically for E=0\vv E'=\vv 0.

(ii) The fields stay perpendicular and stay in ratio. EB=0\vv E'\cdot\vv B'=0 in every frame because the pseudoscalar is invariant, and E=cBE'=cB' in every frame because the scalar is. A light wave looks like a light wave to everybody.

(iii) It cannot be brought to rest. "At rest" would mean a static field configuration, for which E\vv E and B\vv B are independent and generically EcBE\neq cB. Here is the sharper version. The wave's four-wavevector kμ=(ω/c,k)k^{\mu}=(\omega/c,\vv k) satisfies kk=0k\cdot k=0 (Chapter 2.5 §7.1), and a null four-vector cannot be boosted to a purely timelike one. That is the same statement as "Δs2=0\Delta s^{2}=0 is preserved" in Chapter 2.3. The vanishing of both field invariants is the field-theoretic face of the same fact, which is that null is a Lorentz-invariant category.

What a boost can do. Change the amplitude and the frequency, together and in the same ratio. From §5, boosting along the propagation direction multiplies both E\vv E_{\perp} and B\vv B_{\perp} by the Doppler factor (1β)/(1+β)\sqrt{(1-\beta)/(1+\beta)}, exactly the factor Chapter 2.5 §7.2 derived for ω\omega. So a light wave can be made as weak and as red as you like, approaching zero without ever reaching it. It can also be aimed differently, since a boost transverse to k\vv k changes the propagation direction, which is aberration. What survives every boost is the wave's identity as a wave.

The physical moral. The vanishing of both invariants is what makes light structurally different from every static field. It is also why the photon is massless. The null condition on kμk^{\mu} is E=pcE=pc (Chapter 2.5 §4.3), and that means m=0m=0.

Problem 3 · charge conservation as an output, not an input

Show that νμFμν=0\partial_{\nu}\partial_{\mu}F^{\mu\nu}=0 identically, for any antisymmetric FμνF^{\mu\nu}. Deduce that νjν=0\partial_{\nu}j^{\nu}=0 follows from the field equation rather than being an extra assumption. Then explain what this says about a hypothetical universe in which charge is not conserved, and connect it to Chapter 0.7's Problem 4 on the displacement current.

Solution

The identity. The object νμ\partial_{\nu}\partial_{\mu} is symmetric under μν\mu\leftrightarrow\nu, because mixed partials commute, and FμνF^{\mu\nu} is antisymmetric. Contracting a symmetric object with an antisymmetric one over both indices gives zero, which is the argument of (2.6.71) repeated. Explicitly, relabel the two dummy indices:

XνμFμν=μνFνμ=νμFμν=X    X=0. X \equiv \partial_{\nu}\partial_{\mu}F^{\mu\nu} = \partial_{\mu}\partial_{\nu}F^{\nu\mu} = -\partial_{\nu}\partial_{\mu}F^{\mu\nu} = -X \;\Longrightarrow\; X=0.

The consequence. Take ν\partial_{\nu} of the field equation (2.6.19):

0=νμFμν=μ0νjννjν=0. 0 = \partial_{\nu}\partial_{\mu}F^{\mu\nu} = \mu_{0}\,\partial_{\nu}j^{\nu} \qquad\Longrightarrow\qquad \partial_{\nu}j^{\nu} = 0.

Charge conservation is a theorem rather than a postulate. Section 1 assumed it in order to build jμj^{\mu}, and the theory then hands it straight back.

What it forbids. This is stronger than it looks. It says you cannot write down a consistent electrodynamics with a non-conserved source. If someone hands you a jμj^{\mu} with νjν0\partial_{\nu}j^{\nu}\neq0, the equation μFμν=μ0jν\partial_{\mu}F^{\mu\nu}=\mu_{0}j^{\nu} has no solutions at all. Not "solutions with strange behaviour". None. A universe in which charge could appear from nowhere could not have Maxwell's equations, in any frame, even approximately. The rigidity comes entirely from the antisymmetry of FF, which came from gauge invariance (the callout at the end of §2), which is why Chapter 6.3 can say that gauge symmetry and charge conservation are two descriptions of one fact.

The displacement current. Chapter 0.7's Problem 4 asked you to show that ×B=μ0J\nabla\times\vv B=\mu_{0}\vv J is inconsistent with charge conservation, because taking the divergence gives 0=μ0J0=\mu_{0}\nabla\cdot\vv J, which contradicts (2.6.1) whenever tρ0\partial_{t}\rho\neq0. It also showed that the unique minimal repair is Maxwell's μ0ϵ0tE\mu_{0}\epsilon_{0}\partial_{t}\vv E.

This problem is the four-dimensional version of that argument, and it is now visibly the same argument. The three-dimensional identity (×B)=0\nabla\cdot(\nabla\times\vv B)=0 that forced the repair is the spatial part of νμFμν=0\partial_{\nu}\partial_{\mu}F^{\mu\nu}=0. In the covariant formulation the displacement current is not a repair. It was never absent.

Problem 4 · two parallel wires, both ways

Two long parallel wires a distance dd apart carry currents I1I_{1} and I2I_{2} in the same direction. Model each of them as in §6, with lattice +λ0+\lambda_{0} at rest and electrons λ0-\lambda_{0} drifting at vax^-v_{a}\hat{\vv x}, so that Ia=λ0vaI_{a}=\lambda_{0}v_{a}. (a) Compute the force per unit length magnetically, in the lab. (b) Now boost to the rest frame of wire 2's conduction electrons and recompute it electrostatically, being careful about how "per unit length" transforms. (c) Check that wire 2's lattice feels no net force in that frame either, as it must not. (d) Put in I1=I2=1 AI_{1}=I_{2}=1\ \mathrm{A} and d=1 cmd=1\ \mathrm{cm}.

Solution

(a) The lab. Wire 1 produces B1=μ0I12πdz^\vv B_{1}=\dfrac{\mu_{0}I_{1}}{2\pi d}\hat{\vv z} at wire 2 (taking wire 2 to lie in the +y^+\hat{\vv y} direction from wire 1). Only wire 2's electrons move, with charge per unit length λ0-\lambda_{0} and velocity v2x^-v_{2}\hat{\vv x}:

fL=(λ0)(v2x^)×μ0I12πdz^=λ0v2μ0I12πd(x^×z^)=μ0I1I22πdy^. \frac{\vv f}{L} = (-\lambda_{0})(-v_{2}\hat{\vv x})\times\frac{\mu_{0}I_{1}}{2\pi d}\hat{\vv z} = \lambda_{0}v_{2}\frac{\mu_{0}I_{1}}{2\pi d}\big(\hat{\vv x}\times\hat{\vv z}\big) = -\frac{\mu_{0}I_{1}I_{2}}{2\pi d}\,\hat{\vv y}.

Attractive, magnitude μ0I1I2/2πd\mu_{0}I_{1}I_{2}/2\pi d. The lattice of wire 2 is at rest and wire 1 is neutral, so nothing else contributes.

(b) The co-moving frame. Boost at w=v2w=-v_{2} (so wire 2's electrons are at rest). By (2.6.45) with uwu\to w, wire 1 acquires

λ1=γ2λ0wv1c2=+γ2λ0v1v2c2=+γ2I1v2c2,γ2=(1v22c2)1/2, \lambda'_{1} = -\gamma_{2}\lambda_{0}\frac{wv_{1}}{c^{2}} = +\gamma_{2}\frac{\lambda_{0}v_{1}v_{2}}{c^{2}} = +\gamma_{2}\frac{I_{1}v_{2}}{c^{2}}, \qquad \gamma_{2}=\Big(1-\frac{v_{2}^{2}}{c^{2}}\Big)^{-1/2},

That is a positive charge density, and wire 2's electrons are negative, so the force is attractive, agreeing with (a) in direction.

Wire 2's electrons are now at rest, so their density is their proper density, λ0/γ2-\lambda_{0}/\gamma_{2} (they were contracted by γ2\gamma_{2} in the lab). The electric field from wire 1 at distance dd is Ey=λ1/2πϵ0dE'_{y}=\lambda'_{1}/2\pi\epsilon_{0}d, so the force per unit length measured in this frame is

fyL=(λ0γ2)γ2λ0v1v2/c22πϵ0d=λ02v1v22πϵ0c2d=μ0I1I22πd. \frac{f'_{y}}{L'} = \left(-\frac{\lambda_{0}}{\gamma_{2}}\right)\frac{\gamma_{2}\lambda_{0}v_{1}v_{2}/c^{2}}{2\pi\epsilon_{0}d} = -\frac{\lambda_{0}^{2}v_{1}v_{2}}{2\pi\epsilon_{0}c^{2}d} = -\frac{\mu_{0}I_{1}I_{2}}{2\pi d}.

The "per unit length" bookkeeping. That looks like the same number as (a), which would be wrong, because transverse forces are not equal in the two frames. They differ by γ\gamma, as (2.6.48) says. The resolution is that lengths differ too. Take a lab segment of length LL containing NN electrons. Those same electrons are at rest in the primed frame, so they occupy L=γ2LL'=\gamma_{2}L there. Hence the total force on them is

Fy=fyLL=μ0I1I22πdγ2LFy=Fyγ2=μ0I1I22πdL. F'_{y} = \frac{f'_{y}}{L'}\cdot L' = -\frac{\mu_{0}I_{1}I_{2}}{2\pi d}\,\gamma_{2}L \qquad\Longrightarrow\qquad F_{y} = \frac{F'_{y}}{\gamma_{2}} = -\frac{\mu_{0}I_{1}I_{2}}{2\pi d}\,L. \quad\checkmark

Two factors of γ2\gamma_{2} appeared there, one from length contraction and one from force transformation, and they cancel. That is exactly why the force per unit length happens to come out numerically the same in both frames. The coincidence is worth noticing precisely so that you do not mistake it for a proof. The frame-covariant statement is the total force on a given set of charges.

(c) The lattice of wire 2. In the primed frame it moves at +v2x^+v_{2}\hat{\vv x}, has density γ2λ0\gamma_{2}\lambda_{0}, and sits in both an electric and a magnetic field. From §5, Bz=γ2(BzwEy/c2)=γ2μ0I1/2πdB'_{z}=\gamma_{2}(B_{z}-wE_{y}/c^{2})=\gamma_{2}\mu_{0}I_{1}/2\pi d since Ey=0E_{y}=0 in the lab. So

electric:(γ2λ0)γ2λ0v1v2/c22πϵ0dy^=+γ22μ0I1I22πdy^,magnetic:(γ2λ0)(v2x^)×(γ2μ0I12πdz^)=γ22μ0I1I22πdy^. \begin{aligned} \text{electric:}\quad &(\gamma_{2}\lambda_{0})\frac{\gamma_{2}\lambda_{0}v_{1}v_{2}/c^{2}}{2\pi\epsilon_{0}d}\,\hat{\vv y} = +\gamma_{2}^{2}\frac{\mu_{0}I_{1}I_{2}}{2\pi d}\hat{\vv y},\\[3pt] \text{magnetic:}\quad &(\gamma_{2}\lambda_{0})(v_{2}\hat{\vv x})\times\Big(\gamma_{2}\frac{\mu_{0}I_{1}}{2\pi d}\hat{\vv z}\Big) = -\gamma_{2}^{2}\frac{\mu_{0}I_{1}I_{2}}{2\pi d}\hat{\vv y}. \end{aligned}

They cancel exactly. The lattice feels nothing, as it must, since in the lab it feels nothing and "zero force" is a frame-independent statement.

Notice how the cancellation works. In this frame the lattice is a current, and the repulsion from the now-charged wire 1 is precisely balanced by the magnetic attraction between the two currents. Neither description is more true than the other.

(d) Numbers.

FL=μ0I1I22πd=(4π×107)(1)(1)2π(0.01)=2×105 Nm1. \frac{F}{L} = \frac{\mu_{0}I_{1}I_{2}}{2\pi d} = \frac{(4\pi\times10^{-7})(1)(1)}{2\pi(0.01)} = 2\times10^{-5}\ \mathrm{N\,m^{-1}}.

Twenty micronewtons per metre. That is small, and it is measurable with a torsion balance. Until 2019 this configuration defined the ampere, as the current which, in two infinite parallel wires one metre apart, produces a force of 2×107 Nm12\times10^{-7}\ \mathrm{N\,m^{-1}}. An entire base unit of the SI was defined by a relativistic correction to Coulomb's law of relative order 102610^{-26}.

The brick you just laid

Six objects, and Part II is finished.

The four-current jμ=(cρ,J)j^{\mu}=(c\rho,\vv J), built by requiring μjμ=0\partial_{\mu}j^{\mu}=0 to reproduce Chapter 0.7's continuity equation. After that, charge conservation is one contracted index, and therefore true in every frame.

The field tensor Fμν=μAννAμF^{\mu\nu}=\partial^{\mu}A^{\nu}-\partial^{\nu}A^{\mu}, whose six independent components are E/c-\vv E/c in the first row and ϵijkBk-\epsilon^{ijk}B^{k} in the spatial block. Those are the six that Chapter 2.4 §7.2 counted and promised.

Maxwell's four equations as two. Here μFμν=μ0jν\partial_{\mu}F^{\mu\nu}=\mu_{0}j^{\nu} expands to Gauss and Ampère–Maxwell, and [λFμν]=0\partial_{[\lambda}F_{\mu\nu]}=0 expands to B=0\nabla\cdot\vv B=0 and Faraday. The second of the two is not physics at all. It is six terms cancelling by Clairaut, which is Chapter 0.7's (×A)=0\nabla\cdot(\nabla\times\vv A)=0 and ×ϕ=0\nabla\times\nabla\phi=\vv 0 fused into one identity.

Gauge freedom AμAμ+μχA^{\mu}\to A^{\mu}+\partial^{\mu}\chi, which is Chapter 1.2's total-derivative freedom in field form, spent on the Lorenz gauge to give Aμ=μ0jμ\Box A^{\mu}=\mu_{0}j^{\mu}.

Two invariants, 2(B2E2/c2)2(B^{2}-E^{2}/c^{2}) and (8/c)EB(8/c)\vv E\cdot\vv B, which classify field configurations the way the sign of Δs2\Delta s^{2} classifies intervals.

The Lagrangian 14μ0FμνFμνjμAμ-\tfrac{1}{4\mu_{0}}F_{\mu\nu}F^{\mu\nu}-j_{\mu}A^{\mu}, entry three in Chapter 1.2's table, which on variation returns the field equations and which Chapter 6.4 will copy verbatim with a non-abelian FF.

And the energy–momentum tensor, whose T00T^{00} is 12(ϵ0E2+B2/μ0)\tfrac12(\epsilon_{0}E^{2}+B^{2}/\mu_{0}) and whose T0iT^{0i} is S/c\vv S/c, giving field momentum density g=ϵ0E×B\vv g=\epsilon_{0}\vv E\times\vv B.

And two long-standing debts, both settled by name.

Chapter 2.1 posed a contradiction between Maxwell and Galileo and offered a fork. The resolution is that Maxwell's equations single out a speed, which lives in the metric, and never singled out a frame. So branch (A) was not merely unsupported by Michelson and Morley. It was structurally impossible.

Chapter 1.1 showed Newton's third law failing for two moving charges by μ0q1q2v2y^/4πd2-\mu_{0}q_{1}q_{2}v^{2}\hat{\vv y}/4\pi d^{2} and wrote an IOU. Section 10 found the missing momentum sitting in the field at μ0q1q2v(2x^+y^)/8πd\mu_{0}q_{1}q_{2}v(2\hat{\vv x}+\hat{\vv y})/8\pi d, and showed that its rate of change cancels the mechanical one exactly, in both components, including a 32\tfrac32 that 1.1 never saw. That forced you to accept that a field is not a bookkeeping device but a physical system.

And the thesis, which is the reason the chapter exists. Electromagnetism did not need to be made compatible with relativity. It was relativistic before the word existed, which is exactly why it produced a frame-independent cc and broke the physics of 1900. What needed fixing was mechanics, and Chapters 2.2 to 2.5 fixed it.

The magnetic field is not a second force alongside the electric one. It is what an electric field looks like from a moving frame, an effect of relative order uvd/c2u\,v_{d}/c^{2}, which for two ordinary currents in copper is about 102610^{-26}. It is visible at all only because the positive and negative charge in matter cancel to at least that precision. A discrepancy in the twenty-sixth significant figure, and it defines a base unit of the SI.

Two fields, one tensor. Four equations, two lines. And a cc that belongs to spacetime rather than to any medium. That is the whole of Part II in one chapter, and it is the last thing you need before geometry starts to move.

Where this gets spent.

  • TμνT^{\mu\nu}Chapter 3.6, where it is what sits on the right-hand side of the Einstein field equations. Energy gravitates, and now you know what "energy" means as a tensor and why it had to be symmetric.
  • F=dAF=\dd A and [λFμν]=0\partial_{[\lambda}F_{\mu\nu]}=0 → Chapter 3.5, where they become F=dAF=\dd A and dF=0\dd F=0 literally, and d2=0\dd^{2}=0 stops being an analogy.
  • L=14μ0F2j ⁣ ⁣A\mathcal{L}=-\tfrac{1}{4\mu_{0}}F^{2}-j\!\cdot\!A → Chapter 5.2, quantised, where the field's normal modes become photons. Then Chapter 5.8, where jμAμ-j_{\mu}A^{\mu} becomes the single vertex of QED, and every Feynman diagram in the theory is built from copies of it.
  • Gauge freedom → Chapter 6.3, promoted from a convenience to the generating principle. Demand it locally and this entire chapter is the output rather than the input. Then Chapter 6.4 repeats the construction with a non-commuting group and gets Yang–Mills, using §9's Lagrangian unchanged except for a commutator.
  • Charge invariance and the E\vv EB\vv B mixing → Chapter 5.5, where the Dirac field's coupling to AμA^{\mu} is fixed by exactly these transformation properties.
  • The whole structure → Chapter 7.5, where the gauge fields of the Standard Model appear as excitations of an open string, and Chapter 7.7, where FμνFμνF_{\mu\nu}F^{\mu\nu} turns out to be the leading term of an expansion on a D-brane whose higher terms are the Born–Infeld action.