Part III · General Relativity — Chapter 3.5

Forms, Lie Derivatives, Killing Vectors

A chapter that mostly collects. Three promises made earlier in the book fall due at once, and the tool that makes the next chapter but one solvable is built at the end.

Where we are

Chapter 3.4 finished the hard part of Part III. Curvature exists, it is measurable from inside, and it is the tide. One contracted identity has already decided the shape of the field equations. What is still missing before Chapter 3.6 can be written is not more geometry. It is more language, and this chapter supplies exactly two pieces of it.

The first piece is the machinery of differential forms. Its job is integration. Everything integrated so far in this book has been integrated over a flat region with an obvious notion of volume. On a manifold neither the region nor the volume comes for free. Chapter 3.6 has to integrate a scalar over all of spacetime and then throw away a boundary term, and that cannot be done honestly with the tools now in hand.

The second piece is the Lie derivative, and with it the notion of a Killing vector. A Killing vector is a direction in which the geometry does not change. Chapter 3.7 solves the field equations outside a spherical star. The only reason that computation fits in a chapter rather than a filing cabinet is that two Killing vectors reduce four coupled second-order equations to one first-order equation with two constants of integration. Those constants are Noether's, wearing geometry's clothes, and §9 derives them.

Three debts are collected on the way, and they are not incidental. Chapter 0.7 §5.4 promised that its four integral theorems were one theorem awaiting a language general enough for curved spaces. Section 5 below delivers it. Chapter 0.7 §§7.1 and 7.2 proved two vector identities by grinding through components and said they were secretly the same statement. Chapter 2.6 §3.4 showed that two of Maxwell's four equations are not physics but bookkeeping. Section 3 below shows that those are one fact, and says what fact it is. Chapter 0.7 §7.3 quoted the Poincaré lemma and said that this chapter would prove it, and §4 does.

Here is the route. Section 1 asks what an integrand has to look like, and finds antisymmetry forced on it. Section 2 builds a derivative for such objects that needs no connection at all. Section 3 is the collection. Section 4 asks when a thing that could have come from a derivative actually did, and finds that the answer depends on the shape of the region rather than on the thing. Section 5 is the general Stokes theorem. Section 6 builds the volume element and, with it, a divergence theorem that works on a manifold. Chapter 3.6 needs both. Sections 7 and 8 build the Lie derivative and Killing's equation, and §9 turns a Killing vector into a conserved quantity. Section 10 rewrites all of electromagnetism in three symbols, as a test of the language.

Conventions. Signature (+,,,)(+,-,-,-) and the Riemann and Ricci sign conventions stated in Chapter 3.4 are unchanged. GG and cc stay explicit. Indices in square brackets denote antisymmetrisation with a 1/p!1/p!, defined in §1.

Tools you'll need.  Chapter 0.7 heavily: §5 (the three integral theorems and the table in §5.4), §7.1 and §7.2 (the two second-derivative identities), §7.3 (the vector potential and the quoted Poincaré lemma), and §2.4's closed-but-not-exact counterexample on the punctured plane. Chapter 2.6 §2.3 (the field tensor), §3.2 (the homogeneous Maxwell pair) and §3.4 (why two of Maxwell's equations are bookkeeping). Chapter 0.6 §8 (change of variables and the Jacobian determinant) and §8.3, which already told you the answer to §6 below. Chapter 0.4 §5, the determinant as the unique alternating multilinear function of the columns. Section 1 is that theorem doing physics. Chapter 3.2 §6 (one-forms and the basis dxμ\dd x^{\mu}) and §8 (the commutator of vector fields). Chapter 3.3 §5 (the covariant derivative), §7 (metric compatibility and the symmetry of the connection) and §8 (geodesics). Chapter 1.4 §2 in full. Section 9 below is that theorem, restated.

1 · What integration wants, and why it wants antisymmetry

Here is the destination, stated before we set out. We want objects that can be integrated over a pp-dimensional surface sitting inside a manifold, with the answer depending on the surface and not on how the surface was labelled. We will find that this single demand forces the integrand to be totally antisymmetric, and that nothing else is forced.

1.1 · The case that already works: a line integral

Chapter 3.2 §6 built one-forms, so take one: ω=ωμdxμ\omega=\omega_{\mu}\,\dd x^{\mu}. Take also a curve CC, given in some chart by xμ(λ)x^{\mu}(\lambda) for λ[a,b]\lambda\in[a,b]. Define

Cω    abωμ(x(λ))dxμdλ  dλ. \int_{C}\omega \;\equiv\; \int_{a}^{b}\omega_{\mu}\big(x(\lambda)\big)\,\dv{x^{\mu}}{\lambda}\;\dd\lambda. (3.5.1)

Two things about (3.5.1) deserve to be said out loud, because the whole chapter is built on them.

No metric was used. There is no length, no angle and no volume anywhere in (3.5.1). A one-form eats a vector and returns a number, and the tangent dxμ/dλ\dd x^{\mu}/\dd\lambda is a vector. The pairing between them is the duality of Chapter 3.2 §6, which predates any metric.

The answer does not depend on λ\lambda. Let us check that. Reparametrise the curve, λ=λ(σ)\lambda=\lambda(\sigma), with dλ/dσ>0\dd\lambda/\dd\sigma>0. Two factors in the integrand respond. The chain rule turns the tangent into dxμ/dσ=(dxμ/dλ)(dλ/dσ)\dd x^{\mu}/\dd\sigma = (\dd x^{\mu}/\dd\lambda)(\dd\lambda/\dd\sigma). The ordinary change-of-variable rule for a single integral turns the measure into dλ=(dλ/dσ)dσ\dd\lambda=(\dd\lambda/\dd\sigma)\,\dd\sigma. Put the two together:

ωμdxμdσdσ  =  ωμdxμdλdλdσdσ  =  ωμdxμdλdλ. \omega_{\mu}\dv{x^{\mu}}{\sigma}\,\dd\sigma \;=\; \omega_{\mu}\dv{x^{\mu}}{\lambda}\,\dv{\lambda}{\sigma}\,\dd\sigma \;=\; \omega_{\mu}\dv{x^{\mu}}{\lambda}\,\dd\lambda. (3.5.2)

Look at what did the work there. The Jacobian of the reparametrisation appeared twice: once from the tangent vector and once, inverted, from the measure. The two cancelled. That cancellation is the entire mechanism, and everything in §1.2 is the same cancellation in more dimensions.

1.2 · Two dimensions, where something has to give

Now try a surface. Parametrise a two-dimensional patch by (u1,u2)xμ(u)(u^{1},u^{2})\mapsto x^{\mu}(u), and suppose the integrand is some rank-2 tensor TμνT_{\mu\nu}. The only expression available with the right index structure is

Tμνxμu1xνu2  du1du2. \int T_{\mu\nu}\,\pdv{x^{\mu}}{u^{1}}\,\pdv{x^{\nu}}{u^{2}}\;\dd u^{1}\dd u^{2}. (3.5.3)

Ask what happens under a change of parameters uuu\mapsto u'. Write the Jacobian matrix Jab=ua/ubJ^{a}{}_{b}=\partial u^{a}/\partial u'^{b}, with a,b{1,2}a,b\in\{1,2\}. Two things change, and we take them one at a time.

Change one: the measure. By Chapter 0.6 §8, du1du2=detJdu1du2\dd u^{1}\dd u^{2}=\abs{\det J}\,\dd u'^{1}\dd u'^{2}. For an oriented integral we drop the absolute value and keep detJ\det J, agreeing to count a patch negatively if the new parameters reverse its handedness. That is a definition, and it is the same choice made in Chapter 0.7 when Stokes' theorem acquired a right-hand rule.

Change two: the two tangent vectors. By the chain rule, xμ/ub=(xμ/ua)Jab\partial x^{\mu}/\partial u'^{b}=(\partial x^{\mu}/\partial u^{a})J^{a}{}_{b}. So the pair of tangents feeding TT becomes

Tμνxμu1xνu2  =  Ja1Jb2  Tμνxμuaxνub. T_{\mu\nu}\,\pdv{x^{\mu}}{u'^{1}}\,\pdv{x^{\nu}}{u'^{2}} \;=\; J^{a}{}_{1}\,J^{b}{}_{2}\;T_{\mu\nu}\,\pdv{x^{\mu}}{u^{a}}\,\pdv{x^{\nu}}{u^{b}}. (3.5.4)

Abbreviate TabTμν(xμ/ua)(xν/ub)T_{ab}\equiv T_{\mu\nu}(\partial x^{\mu}/\partial u^{a})(\partial x^{\nu}/\partial u^{b}). That is just TT evaluated on the two coordinate tangents, a 2×22\times2 array of numbers. Our demand is that (3.5.3) come out unchanged. Put beside change one, that demand says we need

Ja1Jb2Tab  =  (detJ)T12for every invertible J. J^{a}{}_{1}J^{b}{}_{2}\,T_{ab} \;=\; \big(\det J\big)\,T_{12} \qquad\text{for every invertible }J. (3.5.5)

Expand both sides. The left has four terms. The right has two, once we write out detJ=J11J22J21J12\det J=J^{1}{}_{1}J^{2}{}_{2}-J^{2}{}_{1}J^{1}{}_{2}:

J11J12T11+J11J22T12+J21J12T21+J21J22T22=  J11J22T12    J21J12T12. \begin{aligned} J^{1}{}_{1}J^{1}{}_{2}T_{11} + J^{1}{}_{1}J^{2}{}_{2}T_{12} + J^{2}{}_{1}J^{1}{}_{2}T_{21} + J^{2}{}_{1}J^{2}{}_{2}T_{22}\\ =\; J^{1}{}_{1}J^{2}{}_{2}T_{12} \;-\; J^{2}{}_{1}J^{1}{}_{2}T_{12}. \end{aligned} (3.5.6)

The four products J11J12J^{1}{}_{1}J^{1}{}_{2}, J11J22J^{1}{}_{1}J^{2}{}_{2}, J21J12J^{2}{}_{1}J^{1}{}_{2} and J21J22J^{2}{}_{1}J^{2}{}_{2} can be varied independently, so their coefficients must match one at a time. Reading them off in order:

T11=0,T12=T12,T21=T12,T22=0. T_{11}=0, \qquad T_{12}=T_{12}, \qquad T_{21}=-\,T_{12}, \qquad T_{22}=0. (3.5.7)

Let's read what those four conditions say. The first and last say the diagonal vanishes. The third says the off-diagonal entries are opposite. Both statements are contained in the single condition Tab=TbaT_{ab}=-T_{ba}, since setting a=ba=b in it gives Taa=TaaT_{aa}=-T_{aa} and hence Taa=0T_{aa}=0.

One step remains, from the array TabT_{ab} back to the tensor. By choosing the patch we can make xμ/ua\partial x^{\mu}/\partial u^{a} point in any two independent directions we like. So antisymmetry of TabT_{ab} for every patch forces antisymmetry of TμνT_{\mu\nu} itself.

What was actually proved, and where you have met it before

The requirement was not imposed for elegance. We asked that a surface integral depend on the surface and not on the labelling, discovered that this forces the integrand to change by detJ\det J, and then noticed that the only multilinear functions of two vectors that reproduce detJ\det J are the antisymmetric ones.

That last sentence is Chapter 0.4 §5, where the determinant was characterised as the unique alternating multilinear function of the columns normalised to 11 on the identity. Here that characterisation is no longer a fact about matrices. It is the reason integrands in geometry are antisymmetric. Chapter 0.4 built the tool eleven chapters before there was anything to use it on.

1.3 · Differential forms, and the wedge

Time for the definition. A pp-form is a totally antisymmetric tensor of type (0,p)(0,p): an object ωμ1μp\omega_{\mu_{1}\cdots\mu_{p}} that changes sign whenever any two of its indices are exchanged. The first few degrees are things we already have names for.

  • A 00-form is a function.
  • A one-form is a covector field.
  • A 22-form is what §1.2 just showed a surface integral wants.

Two pieces of notation, both used constantly from here on.

Antisymmetrisation brackets. For any array,

T[μ1μp]    1p!permutations πsgn(π)  Tμπ(1)μπ(p), T_{[\mu_{1}\cdots\mu_{p}]} \;\equiv\; \frac{1}{p!}\sum_{\text{permutations }\pi}\operatorname{sgn}(\pi)\;T_{\mu_{\pi(1)}\cdots\mu_{\pi(p)}}, (3.5.8)

with sgn(π)=±1\operatorname{sgn}(\pi)=\pm1 the sign of the permutation (Chapter 0.4 §5). The 1/p!1/p! makes the operation idempotent: antisymmetrising something already antisymmetric returns it unchanged, since all p!p! terms are then equal to the original. For two indices, T[μν]=12(TμνTνμ)T_{[\mu\nu]}=\half(T_{\mu\nu}-T_{\nu\mu}), which is Chapter 2.4 §7.1's antisymmetric part.

The wedge product. Given a pp-form α\alpha and a qq-form β\beta, their tensor product αμ1μpβν1νq\alpha_{\mu_{1}\cdots\mu_{p}}\beta_{\nu_{1}\cdots\nu_{q}} is not antisymmetric across the two groups. Antisymmetrise it and normalise:

(αβ)μ1μp+q    (p+q)!p!q!  α[μ1μpβμp+1μp+q]. \big(\alpha\wedge\beta\big)_{\mu_{1}\cdots\mu_{p+q}} \;\equiv\; \frac{(p+q)!}{p!\,q!}\;\alpha_{[\mu_{1}\cdots\mu_{p}}\,\beta_{\mu_{p+1}\cdots\mu_{p+q}]}. (3.5.9)

The combinatorial factor is chosen so that basis elements multiply without stray numbers left over. The smallest case shows what it does. For two one-forms, (p,q)=(1,1)(p,q)=(1,1) and the factor is 2!/1!1!=22!/1!1!=2, so

(αβ)μν  =  2α[μβν]  =  αμβνανβμ. \big(\alpha\wedge\beta\big)_{\mu\nu} \;=\; 2\,\alpha_{[\mu}\beta_{\nu]} \;=\; \alpha_{\mu}\beta_{\nu} - \alpha_{\nu}\beta_{\mu}. (3.5.10)

The sign rule, derived. Swap the roles of α\alpha and β\beta in (3.5.9). To turn β[μ1μqαμq+1μp+q]\beta_{[\mu_{1}\cdots\mu_{q}}\alpha_{\mu_{q+1}\cdots\mu_{p+q}]} back into the original ordering, every one of the pp indices of α\alpha must be moved past every one of the qq indices of β\beta. That is pqpq transpositions, each costing a minus sign inside a totally antisymmetric bracket. Hence

αβ  =  (1)pqβα. \alpha\wedge\beta \;=\; (-1)^{pq}\,\beta\wedge\alpha. (3.5.11)

Read off the two cases we will use. Two one-forms anticommute, αβ=βα\alpha\wedge\beta=-\beta\wedge\alpha, and in particular αα=0\alpha\wedge\alpha=0 for any one-form. A one-form and a 22-form commute. Nothing here is a convention we adopted: (3.5.11) was counted.

The basis. Every pp-form can be written in the basis built by wedging coordinate differentials,

ω  =  1p!ωμ1μp  dxμ1dxμp, \omega \;=\; \frac{1}{p!}\,\omega_{\mu_{1}\cdots\mu_{p}}\;\dd x^{\mu_{1}}\wedge\cdots\wedge\dd x^{\mu_{p}}, (3.5.12)

the 1/p!1/p! compensating for the fact that the sum runs over all orderings of each index set and each ordering contributes the same thing. Equivalently, sum over strictly increasing index sets with no factorial. For example a 22-form in four dimensions is

ω  =  μ<νωμνdxμdxν  =  ω01dx0dx1+ω02dx0dx2+, \omega \;=\; \sum_{\mu<\nu}\omega_{\mu\nu}\,\dd x^{\mu}\wedge\dd x^{\nu} \;=\; \omega_{01}\,\dd x^{0}\wedge\dd x^{1} + \omega_{02}\,\dd x^{0}\wedge\dd x^{2} + \cdots, (3.5.13)

six terms in all.

The count. A pp-form in nn dimensions has one independent component for each way of choosing pp distinct indices from nn. They must be distinct, because a repeated index makes the component vanish by antisymmetry. And only the choice matters, not the order, because reordering changes only the sign. So the number of components is (np)\binom{n}{p}, which is Chapter 2.4 §7.2's count. In four dimensions:

pp001122334455
components (4p)\binom{4}{p}114466441100

Two entries in that row are worth pausing on. The 66 is the number of independent components of the electromagnetic field tensor (Chapter 2.6 §2.3). The field tensor is therefore a 22-form, and not merely an antisymmetric array, which is what §10 collects.

The 00 is not an accident of arithmetic either. A 55-form in four dimensions must have a repeated index somewhere, by the pigeonhole principle, so every one of its components vanishes. In nn dimensions there are no forms of degree above nn, which is why the tower above stops.

In plain terms 3.5.1

An integral over a patch of surface ought to depend on the patch and on nothing else, and in particular not on the grid of labels somebody painted on it. Relabelling stretches the cells of the grid by a factor the earlier parts identified as a determinant, and it stretches the two edge directions feeding the integrand as well; for the two effects to cancel, the integrand must respond to its slots exactly as a determinant responds to two columns, changing sign whenever they are swapped.

So antisymmetry is not a stylistic preference or a tidy formalism. It is what survives the demand that an answer be about a region rather than about a description of one, which is the demand governing this book since its first chapter on coordinates. The objects obeying it are called forms, they carry one slot per dimension of the thing they are integrated over, and the product building larger ones out of smaller inherits its sign rule from the same counting.

One arithmetic fact will be wanted shortly. The number of independent entries such an object can have in four dimensions runs one, four, six, four, one, and then stops dead. The six is the count of electric and magnetic components together, which is no coincidence, and the stopping is why there is a largest possible integral rather than an endless tower.

2 · The exterior derivative, which needs no connection

Again the destination first. We want a derivative that turns a pp-form into a (p+1)(p+1)-form and is a tensor without any connection being chosen. We will find that antisymmetrising the ordinary partial derivative does exactly this, and that the reason it works is the same symmetry of the connection that Chapter 3.3 §7.2 called the torsion-free condition.

2.1 · The definition, and the three familiar cases

Define the exterior derivative of a pp-form ω\omega by

(dω)μ0μ1μp    (p+1)  [μ0ωμ1μp], \big(\dd\omega\big)_{\mu_{0}\mu_{1}\cdots\mu_{p}} \;\equiv\; (p+1)\;\partial_{[\mu_{0}}\,\omega_{\mu_{1}\cdots\mu_{p}]}, (3.5.14)

That is the component form. We will want the same definition written in the wedge-product basis of (3.5.12) as well, because that is the version §5's grind box computes with, so here it is:

dω  =  1p!(νωμ1μp)  dxνdxμ1dxμp. \dd\omega \;=\; \frac{1}{p!}\,\big(\partial_{\nu}\omega_{\mu_{1}\cdots\mu_{p}}\big)\;\dd x^{\nu}\wedge\dd x^{\mu_{1}}\wedge\cdots\wedge\dd x^{\mu_{p}}. (3.5.15)

The factor (p+1)(p+1) in (3.5.14) cancels the 1/(p+1)!1/(p+1)! hidden in the bracket, leaving a plain alternating sum. Take the three smallest cases one at a time.

p=0p=0. A 00-form is a function ff, the bracket does nothing, and

(df)μ  =  μf. \big(\dd f\big)_{\mu} \;=\; \partial_{\mu}f. (3.5.16)

That is the gradient one-form of Chapter 0.6 §4. It is the object that chapter insisted was a row and not a column, and which needed a metric before it could be turned into the arrow called f\nabla f. Here it appears with no metric at all. That is the first sign that d\dd is cheaper than \nabla.

p=1p=1. With two indices the bracket is 12(        )\half(\;\cdot\;-\;\cdot\;) and the prefactor is 22:

(dω)μν  =  2[μων]  =  μων    νωμ. \big(\dd\omega\big)_{\mu\nu} \;=\; 2\,\partial_{[\mu}\omega_{\nu]} \;=\; \partial_{\mu}\omega_{\nu} \;-\; \partial_{\nu}\omega_{\mu}. (3.5.17)

In flat three-dimensional space with Cartesian coordinates, identify a vector field F\vv F with the one-form ωi=Fi\omega_{i}=F^{i}. Then (dω)12=xFyyFx(\dd\omega)_{12}=\partial_{x}F^{y}-\partial_{y}F^{x}, which is (×F)z(\nabla\times\vv F)_{z}, and the other two independent components are the other two components of the curl. The exterior derivative of a one-form is the curl.

p=2p=2 in three dimensions. Identify a vector field with a 22-form by ωjk=ϵijkFi\omega_{jk}=\epsilon^{ijk}F^{i}. This is the same identification Chapter 0.7 §3.1 made implicitly when it wrote a flux as FdA\vv F\cdot\dd\vv A, since a flux integrand eats two edge vectors. Then

(dω)123  =  1ω23+2ω31+3ω12  =  xFx+yFy+zFz  =  F, \big(\dd\omega\big)_{123} \;=\; \partial_{1}\omega_{23} + \partial_{2}\omega_{31} + \partial_{3}\omega_{12} \;=\; \partial_{x}F^{x}+\partial_{y}F^{y}+\partial_{z}F^{z} \;=\; \nabla\cdot\vv F, (3.5.18)

where the three-term alternating sum has collapsed to three terms rather than six because the bracket's other three terms are equal to these by antisymmetry of ω\omega. The exterior derivative of a 22-form is the divergence.

Familiar ground — this is Chapter 0.7's diagram

Chapter 0.7 §7.4 drew the chain

functions      vector fields   ×   vector fields      functions \text{functions} \;\xrightarrow{\ \nabla\ }\; \text{vector fields} \;\xrightarrow{\ \nabla\times\ }\; \text{vector fields} \;\xrightarrow{\ \nabla\cdot\ }\; \text{functions}

and said that in this chapter's language the three objects would be forms of degree 0,1,20,1,2 and all three arrows would be one operator. That is now shown rather than promised. (3.5.16), (3.5.17) and (3.5.18) are the three arrows, and each of them is (3.5.14) at a different value of pp. The chain also has a fourth arrow that Chapter 0.7 could not draw, taking 33-forms to 44-forms. It vanishes in three dimensions, and it does not vanish in four.

2.2 · Why dω\dd\omega is a tensor, with no connection anywhere

Chapter 3.3 §4 proved that μVν\partial_{\mu}V^{\nu} is not a tensor, and the whole of that chapter's §5 was the repair. Antisymmetrising escapes the problem entirely, and the reason is one line.

Write the covariant derivative of a one-form, from Chapter 3.3 (3.3.30):

μων  =  μων    Γλμνωλ. \nabla_{\mu}\omega_{\nu} \;=\; \partial_{\mu}\omega_{\nu} \;-\; \Gamma^{\lambda}{}_{\mu\nu}\,\omega_{\lambda}. (3.5.19)

Our move is to antisymmetrise both sides in μν\mu\nu and see what survives. On the left the result is [μων]\nabla_{[\mu}\omega_{\nu]}. On the right, the first term gives [μων]\partial_{[\mu}\omega_{\nu]} and the second gives Γλ[μν]ωλ-\Gamma^{\lambda}{}_{[\mu\nu]}\omega_{\lambda}.

Now look at that second term. Γλμν\Gamma^{\lambda}{}_{\mu\nu} is symmetric in its two lower indices. That is the torsion-free condition, imposed in Chapter 3.3 §7.2 and derived there to be equivalent to the vanishing of Γλ[μν]\Gamma^{\lambda}{}_{[\mu\nu]}. So the second term is antisymmetrising something symmetric, which gives zero. What is left is

(dω)μν  =  2[μων]  =  2[μων]. \big(\dd\omega\big)_{\mu\nu} \;=\; 2\,\partial_{[\mu}\omega_{\nu]} \;=\; 2\,\nabla_{[\mu}\omega_{\nu]}. (3.5.20)

The same argument runs at every pp. Each of the pp correction terms in the covariant derivative of a pp-form carries a Γ\Gamma with the differentiating index and one form index sitting on its two lower slots, and antisymmetrising over those two kills it.

Let's look at what (3.5.20) is actually saying. It is worth stating twice, once in each direction. Read left to right, it says that the plain partial derivative, antisymmetrised, is secretly covariant and is therefore a tensor. Read right to left, it says that d\dd can be computed with partial derivatives and no Christoffel symbols at all, even on a curved manifold, and that the answer does not depend on which connection you would have used.

The exterior derivative exists before the metric does. Chapter 3.6 will lean on this without saying so. Chapter 6.3, where a connection on an internal space is built from scratch, leans on it harder.

⚠ The belief this section exists to prevent

That d\dd needs no connection does not mean it needs no geometry. The two claims get run together constantly, and the difference matters as soon as you try to integrate anything. What §2.2 proved is narrow and should be held at exactly its width: the Γ\Gamma terms cancel by antisymmetry, so d\dd is defined on a bare manifold, before a metric is chosen and independently of which connection anyone would have chosen. Nothing more.

Everything else in this chapter is paid for. Section 6 builds the object that lets an integral know how much region it is covering, and that object is g\sqrt{\abs{g}}, which is pure metric. Section 8 asks which flows leave distances alone, a question that cannot even be posed without a metric. And §10 writes electromagnetism in three symbols, of which exactly one requires a geometry. The homogeneous half of Maxwell's equations is dF=0\dd F=0 and holds on any manifold whatsoever. The sourced half needs the star, and the star is built from gμνg_{\mu\nu}.

The slogan to carry is not "forms need no geometry". It is that differentiating is free and measuring is not. That is the same division Chapter 3.2 §6 drew between the object a function hands you for nothing and the arrow of steepest ascent it cannot.

In plain terms 3.5.2

Two chapters ago, differentiating a field turned out to be impossible without extra equipment, because the values being compared live at different points and nothing brings one to the other for free. Building that equipment took a whole chapter. Here a cheaper route opens, for the special class of objects the previous section forced on us, and the reason it opens is worth keeping.

Take the naive derivative, which fails because it drags along a term built from the comparison coefficients, and antisymmetrise the result over the differentiating slot and one other. The offending term carries those two slots in a pair of positions where the comparison coefficients are known to be symmetric, and antisymmetrising anything symmetric annihilates it. The failure cancels rather than being repaired, so this derivative can be computed with no comparison rule at all, and it returns the same answer whichever rule you would have chosen.

The payoff is that the three operations the toolkit taught separately, the one measuring steepness, the one measuring swirl and the one measuring outflow, stop being three operations. They are one operation applied to objects carrying one, two and three slots, and the reason they looked different is that vectors in three dimensions can stand in for all three types. That coincidence is over, and what replaces it survives in any number of dimensions and on any curved space.

3 · d2=0\dd^{2}=0, and three debts collected at once

This is the section the chapter exists for. Here is the destination and the collection together. We will prove one line, that applying the exterior derivative twice always gives zero. Then we show that this single line is Chapter 0.7's identity that the curl of a gradient vanishes, Chapter 0.7's identity that the divergence of a curl vanishes, and Chapter 2.6's discovery that two of Maxwell's four equations carry no physical content. Those are not three analogous facts. They are one fact seen three times.

3.1 · The proof

Take any pp-form ω\omega and apply (3.5.14) twice. The result has p+2p+2 indices, and every term in it carries two partial derivatives:

(ddω)μ0μ1μp+1  =  (p+2)(p+1)  [μ0μ1ωμ2μp+1]. \big(\dd\dd\omega\big)_{\mu_{0}\mu_{1}\cdots\mu_{p+1}} \;=\; (p+2)(p+1)\;\partial_{[\mu_{0}}\partial_{\mu_{1}}\omega_{\mu_{2}\cdots\mu_{p+1}]}. (3.5.21)

Look only at the two derivative indices. They sit inside a total antisymmetrisation, so the expression changes sign when they are exchanged. But μ0μ1=μ1μ0\partial_{\mu_{0}}\partial_{\mu_{1}}=\partial_{\mu_{1}}\partial_{\mu_{0}} by Clairaut's theorem (Chapter 0.6 §6.1), so the expression is unchanged when they are exchanged. A quantity equal to minus itself is zero. Hence

  dd  =  0  on every form, of every degree, on every manifold. \boxed{\;\dd\circ\dd \;=\; 0\;}\qquad\text{on every form, of every degree, on every manifold.} (3.5.22)

That is the whole proof, and its brevity is the point rather than a shortcut. Nothing was assumed except that ω\omega is twice continuously differentiable, which is Clairaut's hypothesis and was stated out loud in Chapter 0.6 §6.1.

3.2 · Debt one — the two vector identities of Chapter 0.7

The way to collect a debt is to look at dd=0\dd\dd=0 at one degree at a time. Start at the bottom. Set p=0p=0, so that ω=f\omega=f is a function and df\dd f is the gradient one-form (3.5.16). Then (3.5.22) reads, using the p=1p=1 expansion (3.5.17),

(ddf)μν  =  μνf    νμf  =  0. \big(\dd\dd f\big)_{\mu\nu} \;=\; \partial_{\mu}\partial_{\nu}f \;-\; \partial_{\nu}\partial_{\mu}f \;=\; 0. (3.5.23)

In flat three-dimensional space the three independent components of a 22-form are the three components of a vector, and §2.1 identified d\dd on a one-form with the curl. So (3.5.23) is ×f=0\nabla\times\nabla f=\vv 0, which is (0.7.54). That was proved there in exactly this way, by writing out one component and cancelling two terms.

Now set p=1p=1. Then dω\dd\omega is the curl, ddω\dd\dd\omega is the divergence of it, and (3.5.22) is (×A)=0\nabla\cdot(\nabla\times\vv A)=0, which is (0.7.56). Chapter 0.7 proved that one by writing out six terms and cancelling them in pairs. Here the six terms are the six permutations inside the bracket of (3.5.21), and they cancel for the same reason.

So the two identities of Chapter 0.7 §7 are (3.5.22) at p=0p=0 and p=1p=1. Chapter 0.7 §7.4 asserted this and flagged it as a promise. It is now a theorem.

3.3 · Debt two — half of Maxwell's equations

Now four dimensions. Assemble the electromagnetic potential of Chapter 2.6 §2.2 into a one-form,

A  =  Aμdxμ, A \;=\; A_{\mu}\,\dd x^{\mu}, (3.5.24)

Our aim is to recognise something familiar, so the next step is to take the exterior derivative of that one-form and read off what the components are. The p=1p=1 case of the definition, (3.5.17), gives them straight away:

(dA)μν  =  μAννAμ  =  Fμν, \big(\dd A\big)_{\mu\nu} \;=\; \partial_{\mu}A_{\nu} - \partial_{\nu}A_{\mu} \;=\; F_{\mu\nu}, (3.5.25)

That is Chapter 2.6's definition (2.6.9) of the field tensor, character for character. So F=dAF=\dd A. The electromagnetic field is the exterior derivative of the potential, and its six components are exactly the six independent components a 22-form in four dimensions is allowed by §1.3's table.

Now apply (3.5.22) to it:

dF  =  ddA  =  0. \dd F \;=\; \dd\dd A \;=\; 0. (3.5.26)

Write out what dF=0\dd F=0 says in components. FF is a 22-form, so dF\dd F is a 33-form with (43)=4\binom{4}{3}=4 independent components, and by the same collapse as in (3.5.18) each is a three-term cyclic sum:

(dF)λμν  =  λFμν+μFνλ+νFλμ  =  0. \big(\dd F\big)_{\lambda\mu\nu} \;=\; \partial_{\lambda}F_{\mu\nu} + \partial_{\mu}F_{\nu\lambda} + \partial_{\nu}F_{\lambda\mu} \;=\; 0. (3.5.27)

That is (2.6.25) exactly. Chapter 2.6 §3.2 expanded its four components and found B=0\nabla\cdot\vv B=0 together with the three components of Faraday's law. Chapter 2.6 §3.4 then substituted F=dAF=\dd A, watched six terms cancel in pairs by Clairaut, and concluded that given potentials, those two Maxwell equations cannot fail and carry no information about how electromagnetism works.

Here is the same conclusion in one symbol. The six terms Chapter 2.6 cancelled are the six permutations in (3.5.21) at p=1p=1, and the cancellation is (3.5.22). Chapter 2.6's warning box promised precisely this and called it a promise rather than a derivation. The promise is now kept.

3.4 · What the fact actually is

Three results have just been shown to be one result. It is worth saying what the one result is, because it is not a fact about derivatives at all.

Chapter 0.7 §5.4 made the observation, filed it, and moved on: a boundary has no boundary of its own. The boundary of a disc is a circle, and a circle has no endpoints. The boundary of a ball is a sphere, and a sphere has no edge. In symbols, (M)=\partial(\partial M)=\emptyset.

Section 5 below proves the general Stokes theorem, Mdω=Mω\int_{M}\dd\omega=\oint_{\partial M}\omega, which pairs the operator d\dd with the operation of taking a boundary. Grant it for a moment and run the pairing twice. For any (p2)(p-2)-form ω\omega and any pp-dimensional region MM,

Mddω  =  Mdω  =  (M)ω  =  0, \int_{M}\dd\dd\omega \;=\; \oint_{\partial M}\dd\omega \;=\; \oint_{\partial(\partial M)}\omega \;=\; 0, (3.5.28)

The last step holds because there is nothing left to integrate over. And since the chain holds for every region MM, however small, the integrand itself must vanish: ddω=0\dd\dd\omega=0.

Read that chain in reverse and it says the same thing. The statement that applying the derivative twice gives nothing is the algebraic shadow of the statement that a boundary has no boundary. A fact about regions is showing up as a fact about derivatives.

That is why the two vanishings of Chapter 0.7 felt like a coincidence, and why it felt like a second coincidence that two of Maxwell's equations turned out to be empty. The real situation is this. A single geometric triviality about shapes has three separate algebraic disguises, and in each case we met the disguise before we met the shape.

Recap — what went in, what came out

In: the definition of d\dd as an antisymmetrised partial derivative, and Clairaut's theorem.

Out: dd=0\dd\dd=0, in three lines. Specialising to three dimensions gives both of Chapter 0.7 §7's identities. Specialising to the electromagnetic potential in four dimensions gives both of Maxwell's homogeneous equations and, with them, Chapter 2.6's verdict that half of Maxwell is bookkeeping.

What it cost: nothing beyond the machinery of §1 and §2. Note in particular that no metric, no connection and no curvature entered at any point. The result holds on any manifold whatever, flat or curved, before any geometry is put on it.

In plain terms 3.5.3

Apply the new derivative twice and the answer is always nothing. The proof takes three lines: the two derivative slots sit inside a construction that changes sign when any two slots are swapped, while the derivatives themselves do not care about their order, so the expression equals minus itself and therefore equals zero.

What that line collects is the reason the section exists. The toolkit proved, by grinding through components, that a field of steepest ascent has no swirl and that a field of swirl has no outflow, and remarked that the two proofs looked alike. The chapter on electromagnetism found that two of the four famous equations follow from the existence of potentials and say nothing about how electricity behaves. All three are one statement, and the six terms cancelling in each case are literally the same six terms.

And the statement is not really about differentiation. The edge of a filled disc is a closed loop, and a closed loop has no ends; the skin of a ball is a sphere, and a sphere has no rim. Taking a boundary twice leaves nothing. Because a later theorem pairs the derivative with the boundary, that triviality about shapes must show up as a triviality about derivatives, and it does. What looked like a run of coincidences is one plain fact met three times in the wrong order.

4 · Closed, exact, and the shape of the region

Section 3 answered one direction of a question. Here is the other, and it is the direction where the topology of the region stops being a technicality.

Two definitions. A form ω\omega is closed if dω=0\dd\omega=0. It is exact if ω=dα\omega=\dd\alpha for some form α\alpha of one degree lower. Equation (3.5.22) says exact implies closed, always, with no hypotheses. The question is the converse.

You have met this question three times already, without its name attached.

  • Chapter 0.7 §2.3 asked when a curl-free field is a gradient. In the new language, that is asking when a closed one-form is exact.
  • Chapter 0.7 §7.3 asked when a divergence-free field is a curl, which is a closed 22-form being exact. It quoted the answer as the Poincaré lemma and said this chapter would prove it.
  • Chapter 3.4 §8 met the same distinction in geometric dress, as the gap between having vanishing curvature and being a piece of flat space.

All three are the same question, and in each of them the obstruction is the shape of the region.

4.1 · The counterexample, so that the theorem has something to exclude

Chapter 0.7 §2.4 built the counterexample in full and it is worth having in the new notation. On the plane with the origin removed, define

ω  =  xdyydxx2+y2. \omega \;=\; \frac{x\,\dd y - y\,\dd x}{x^{2}+y^{2}}. (3.5.29)

It is closed. Compute (dω)xy=xωyyωx(\dd\omega)_{xy}=\partial_{x}\omega_{y}-\partial_{y}\omega_{x} with ωx=y/(x2+y2)\omega_{x}=-y/(x^{2}+y^{2}) and ωy=x/(x2+y2)\omega_{y}=x/(x^{2}+y^{2}). Differentiating the first by the quotient rule,

xωy  =  (x2+y2)x(2x)(x2+y2)2  =  y2x2(x2+y2)2, \partial_{x}\omega_{y} \;=\; \frac{(x^{2}+y^{2}) - x(2x)}{(x^{2}+y^{2})^{2}} \;=\; \frac{y^{2}-x^{2}}{(x^{2}+y^{2})^{2}}, (3.5.30)

and the identical computation on ωx\omega_{x} gives the same thing, so the difference is zero everywhere the form is defined.

It is not exact. If ω\omega were df\dd f then ω\oint\omega around any closed loop would be f(end)f(start)=0f(\text{end})-f(\text{start})=0. But on the unit circle, parametrised by (x,y)=(cosθ,sinθ)(x,y)=(\cos\theta,\sin\theta), the definition (3.5.1) gives ωμdxμ/dθ=(cosθ)(cosθ)(sinθ)(sinθ)=1\omega_{\mu}\dd x^{\mu}/\dd\theta = (\cos\theta)(\cos\theta)-(\sin\theta)(-\sin\theta)=1, so ω=2π0\oint\omega=2\pi\neq0.

Nothing is wrong with ω\omega. Something is wrong with the region. The punctured plane has a loop that cannot be shrunk to a point without crossing the missing origin, and that loop is exactly what ω=2π\oint\omega=2\pi detects. Closed but not exact is a measurement of the hole.

4.2 · The Poincaré lemma, for one-forms, proved

Call a region star-shaped about a point if the straight segment from that point to any point of the region stays inside. A ball is star-shaped. The punctured plane is not.

Claim. On a star-shaped region about the origin, every closed one-form is exact.

The construction. Given a closed ω\omega, define a function by integrating ω\omega out along the straight ray from the origin to xx:

f(x)    01ωμ(tx)  xμ  dt. f(x) \;\equiv\; \int_{0}^{1}\omega_{\mu}(t x)\;x^{\mu}\;\dd t. (3.5.31)

This is legitimate precisely because the region is star-shaped: the ray txtx for t[0,1]t\in[0,1] stays inside, so ω\omega is defined all along it.

Line 1: differentiate under the integral sign. Chapter 0.2 §4.4 licensed this move for continuous integrands. We want νf\partial_{\nu}f, and there are two places for the derivative to land. It hits the argument txtx of ω\omega by the chain rule, and it hits the explicit factor xμx^{\mu} by the product rule:

νf(x)  =  01[  t(νωμ)(tx)  xμ  +  ων(tx)  ]dt. \partial_{\nu}f(x) \;=\; \int_{0}^{1}\Big[\;t\,\big(\partial_{\nu}\omega_{\mu}\big)(tx)\;x^{\mu} \;+\; \omega_{\nu}(tx)\;\Big]\dd t. (3.5.32)

Two small things to account for there. The factor tt in the first term is the derivative of the argument txνtx^{\nu} with respect to xνx^{\nu}. The second term came from νxμ=δμν\partial_{\nu}x^{\mu}=\delta^{\mu}{}_{\nu}.

Line 2: use closedness, and only here. Closed means νωμ=μων\partial_{\nu}\omega_{\mu}=\partial_{\mu}\omega_{\nu}. That lets us swap the two labels inside the first term, which is the one move the hypothesis buys us:

νf(x)  =  01[  t(μων)(tx)  xμ  +  ων(tx)  ]dt. \partial_{\nu}f(x) \;=\; \int_{0}^{1}\Big[\;t\,\big(\partial_{\mu}\omega_{\nu}\big)(tx)\;x^{\mu} \;+\; \omega_{\nu}(tx)\;\Big]\dd t. (3.5.33)

Line 3: recognise a product rule in tt. The bracket in (3.5.33) looks like the output of a derivative, so let's go looking for the thing it is the derivative of. Consider the function ttων(tx)t\mapsto t\,\omega_{\nu}(tx) and differentiate it with respect to tt. The product rule gives one term from the explicit tt and one from the argument, and the chain rule supplies xμx^{\mu} in the second:

ddt[tων(tx)]  =  ων(tx)  +  t(μων)(tx)xμ. \dv{}{t}\Big[\,t\,\omega_{\nu}(tx)\,\Big] \;=\; \omega_{\nu}(tx) \;+\; t\,\big(\partial_{\mu}\omega_{\nu}\big)(tx)\,x^{\mu}. (3.5.34)

The right-hand side of (3.5.34) is exactly the bracket in (3.5.33).

Line 4: integrate. The integrand is now a total derivative in tt, so the Fundamental Theorem of Calculus (Chapter 0.2 §2) evaluates it at the two endpoints and we are done:

νf(x)  =  [tων(tx)]t=0t=1  =  ων(x). \partial_{\nu}f(x) \;=\; \Big[\,t\,\omega_{\nu}(tx)\,\Big]_{t=0}^{t=1} \;=\; \omega_{\nu}(x). (3.5.35)

So ω=df\omega=\dd f. \blacksquare

Read line by line, the proof is short enough to hold in the head. Integrate the form radially, differentiate, use closedness once to swap two derivative labels, notice that what remains is a total derivative in the radial parameter, and apply the Fundamental Theorem. The star-shaped hypothesis is spent in exactly one place, and that place is writing down (3.5.31) at all.

Grind box — the same proof for a pp-form, and hence the vector potential

The construction generalises with one extra factor of tt per index. For a pp-form ω\omega on a region star-shaped about the origin, define the (p1)(p-1)-form

(Kω)μ2μp(x)    01tp1  xμ1ωμ1μ2μp(tx)  dt. \big(K\omega\big)_{\mu_{2}\cdots\mu_{p}}(x) \;\equiv\; \int_{0}^{1}t^{\,p-1}\;x^{\mu_{1}}\,\omega_{\mu_{1}\mu_{2}\cdots\mu_{p}}(tx)\;\dd t.

The claim is the homotopy formula

d(Kω)  +  K(dω)  =  ω, \dd\big(K\omega\big) \;+\; K\big(\dd\omega\big) \;=\; \omega,

from which the lemma is immediate: if dω=0\dd\omega=0 the second term dies and ω=d(Kω)\omega=\dd(K\omega) is exact.

The verification is the four lines of the main text with more indices. Rather than sketch that, we do one case completely and explicitly, namely the case the promise was about: p=2p=2 in three dimensions, where the claim is the existence of the vector potential.

Setting up. Under §2.1's identification a 22-form corresponds to a vector field B\vv B, and KK applied to it reads, in components,

Ak(x)  =  01t  ϵkijBi(tx)xj  dt,that isA(x)  =  01t  B(tx)×x  dt. A_{k}(\vv x) \;=\; \int_{0}^{1} t\;\epsilon_{kij}\,B_{i}(t\vv x)\,x_{j}\;\dd t, \qquad\text{that is}\qquad \vv A(\vv x) \;=\; \int_{0}^{1} t\;\vv B(t\vv x)\times\vv x\;\dd t.

Take the curl, differentiating under the integral sign. The derivative hits the argument txt\vv x by the chain rule, producing a factor tt, and separately hits the explicit xjx_{j}:

(×A)m  =  01ϵmnkϵkij[t2(nBi)(tx)xj  +  tBi(tx)δjn]dt. \big(\nabla\times\vv A\big)_{m} \;=\; \int_{0}^{1}\epsilon_{mnk}\epsilon_{kij}\Big[\,t^{2}\,\big(\partial_{n}B_{i}\big)(t\vv x)\,x_{j} \;+\; t\,B_{i}(t\vv x)\,\delta_{jn}\Big]\dd t.

Use the epsilon identity ϵmnkϵkij=δmiδnjδmjδni\epsilon_{mnk}\epsilon_{kij}=\delta_{mi}\delta_{nj}-\delta_{mj}\delta_{ni}, which Chapter 0.7 §4.4's grind box established. The first bracket becomes t2[(jBm)xj(iBi)xm]t^{2}\big[(\partial_{j}B_{m})x_{j}-(\partial_{i}B_{i})x_{m}\big], and here is the one place the hypothesis is spent: iBi=B=0\partial_{i}B_{i}=\nabla\cdot\vv B=0, so the second piece dies and the first is t2(x)Bmt^{2}(\vv x\cdot\nabla)B_{m}. The second bracket becomes t[3BmBm]=2tBmt\big[3B_{m}-B_{m}\big]=2tB_{m}, the 33 coming from δnn=3\delta_{nn}=3.

Recognise a product rule in tt. Differentiating t2Bm(tx)t^{2}B_{m}(t\vv x) with respect to tt gives 2tBm(tx)+t2xj(jBm)(tx)2tB_{m}(t\vv x)+t^{2}x_{j}(\partial_{j}B_{m})(t\vv x), which is exactly the integrand. So by the Fundamental Theorem of Calculus

(×A)m  =  [t2Bm(tx)]t=0t=1  =  Bm(x). \big(\nabla\times\vv A\big)_{m} \;=\; \Big[t^{2}B_{m}(t\vv x)\Big]_{t=0}^{t=1} \;=\; B_{m}(\vv x). \qquad\blacksquare

What this settles. A divergence-free field on a star-shaped region is the curl of something, and the something is written down above. That is the existence of the vector potential, quoted at Chapter 0.7 §7.3 with the note that Chapter 3.5 would prove it. ⚑ The general pp runs the same way. The epsilon identity is replaced by the expansion of dω\dd\omega, the hypothesis dω=0\dd\omega=0 is spent at the same single step, and what survives is again a total derivative in tt whose endpoints give ω\omega. But that is a sketch of a proof and not a proof. What is derived here is the case p=2p=2 in three dimensions. The general statement is quoted, and this is the one place in this chapter where that happens.

4.3 · What is left over is topology

Put the two halves together. Exact always implies closed, by (3.5.22). Closed implies exact on a star-shaped region, by §4.2. On a region that is not star-shaped the implication can fail, by §4.1. So the collection of closed forms that are not exact, counted degree by degree, measures how far the region is from being simple. It is a measurement made with calculus that returns an answer about shape.

That measurement has a name, de Rham cohomology, promised by that name at Chapter 0.7 §7.3. We will not develop it. What matters here is that the same distinction has now appeared four times in this book, and in every case the local information is impeccable while the global conclusion fails.

  • The curl-free field with a non-zero circulation.
  • The divergence-free field with no potential.
  • Chapter 3.4 §8's cone, flat everywhere and yet not a plane.
  • The angle form (3.5.29) above.

Chapter 3.9 meets the same distinction once more, on a universe that can be flat everywhere and still closed. Part VI meets it as the reason certain field configurations cannot be smoothly undone.

In plain terms 3.5.4

Having shown that anything which is a derivative gives nothing when differentiated again, the natural next question is the converse: if differentiating something gives nothing, did it come from a derivative. The answer is yes on any region that can be shrunk to a point without leaving it, and the proof is constructive rather than abstract. Integrate the object outward along straight rays from the centre, and the resulting function has exactly the original as its derivative, with the hypothesis spent at precisely one step.

On a region with a hole the answer is no, and the failure is measurable. The toolkit's old example, the field circulating around a missing origin, differentiates to nothing everywhere it is defined and yet accumulates a full turn around any loop encircling the gap. Nothing is wrong with it. What the loop integral measures is the hole, using nothing but calculus, which is a strange and useful thing for calculus to be able to do.

Two debts are settled here. The toolkit quoted the theorem guaranteeing that a field with no outflow is the swirl of something else, promised a proof in this chapter, and now has one, complete with a formula for the potential. And the same local-versus-global gap that left a rolled paper cone flat everywhere and still not a plane turns out to be this gap rather than a relative of it.

5 · One theorem instead of four

Chapter 0.7 §5.4 lined up the Fundamental Theorem of Calculus, the gradient theorem, Stokes' theorem and the divergence theorem in a table, observed that every row says the same thing, wrote down Mdω=Mω\int_{M}\dd\omega=\oint_{\partial M}\omega as (0.7.46), and said plainly: we are quoting the general form, not deriving it. The machinery is Chapter 3.5's job. The machinery now exists. This section pays.

Here is the destination. We define what it means to integrate a pp-form over a pp-dimensional oriented region, check that the definition does not depend on the labelling, prove the theorem on a cube by the Fundamental Theorem of Calculus applied one face-pair at a time, and extend it to general regions by the chopping-and-cancelling argument of Chapter 0.7. This time that argument cancels exactly rather than approximately, because the antisymmetry of §1 does the work that hand-waving did before.

5.1 · The integral of a pp-form

Let MM be a pp-dimensional region inside the manifold, parametrised by (u1,,up)xμ(u)(u^{1},\ldots,u^{p})\mapsto x^{\mu}(u) over some region UU of parameter space. Define

Mω    Uωμ1μp(x(u))  xμ1u1xμpup  du1dup. \int_{M}\omega \;\equiv\; \int_{U}\omega_{\mu_{1}\cdots\mu_{p}}\big(x(u)\big)\;\pdv{x^{\mu_{1}}}{u^{1}}\cdots\pdv{x^{\mu_{p}}}{u^{p}}\;\dd u^{1}\cdots\dd u^{p}. (3.5.36)

We still owe the check that this is well defined. Section 1.2 did the two-dimensional case, and the general case is the same three sentences. Changing parameters multiplies the product of tangent vectors by pp copies of the Jacobian matrix, contracted into a totally antisymmetric object. By Chapter 0.4 §5 that is precisely detJ\det J times the original. The measure supplies detJ\det J the other way. They cancel. The integral belongs to the region.

Two remarks that will be needed later.

First, no metric appears in (3.5.36), so this notion of integration exists on a bare manifold. Section 6 is where a metric is finally wanted, and it is wanted there for a different job.

Second, reversing the orientation of the region reverses the sign of the integral, since it swaps two of the uu's and the integrand is antisymmetric. That second remark is what makes the whole of §5.3 work.

5.2 · The theorem on a square, in full

Take MM to be the unit square 0u1,u210\le u^{1},u^{2}\le1 and ω\omega a one-form on it, written ω=Pdu1+Qdu2\omega=P\,\dd u^{1}+Q\,\dd u^{2}. Then by (3.5.17) the only independent component of dω\dd\omega is 1Q2P\partial_{1}Q-\partial_{2}P, so

Mdω  =  01 ⁣ ⁣01(Qu1Pu2)du1du2. \int_{M}\dd\omega \;=\; \int_{0}^{1}\!\!\int_{0}^{1}\Big(\pdv{Q}{u^{1}} - \pdv{P}{u^{2}}\Big)\,\dd u^{1}\,\dd u^{2}. (3.5.37)

The QQ term. Do the u1u^{1} integral first and apply the Fundamental Theorem of Calculus in that variable alone, holding u2u^{2} fixed:

01 ⁣ ⁣01Qu1du1du2  =  01[Q(1,u2)Q(0,u2)]du2. \int_{0}^{1}\!\!\int_{0}^{1}\pdv{Q}{u^{1}}\,\dd u^{1}\dd u^{2} \;=\; \int_{0}^{1}\Big[\,Q(1,u^{2}) - Q(0,u^{2})\,\Big]\dd u^{2}. (3.5.38)

The PP term. Do the u2u^{2} integral first, same move, and keep the minus sign:

01 ⁣ ⁣01Pu2du2du1  =  01[P(u1,1)P(u1,0)]du1. -\int_{0}^{1}\!\!\int_{0}^{1}\pdv{P}{u^{2}}\,\dd u^{2}\dd u^{1} \;=\; -\int_{0}^{1}\Big[\,P(u^{1},1) - P(u^{1},0)\,\Big]\dd u^{1}. (3.5.39)

The boundary. The square's boundary is four edges. Traverse them counterclockwise, which is the orientation induced by the square's own, and evaluate (3.5.1) on each. On the bottom edge u2=0u^{2}=0 with u1u^{1} running from 00 to 11, the tangent has only a u1u^{1} component, so only PP contributes and the edge gives +01P(u1,0)du1+\int_{0}^{1}P(u^{1},0)\,\dd u^{1}. On the right edge only QQ contributes, upward, giving +01Q(1,u2)du2+\int_{0}^{1}Q(1,u^{2})\,\dd u^{2}. The top edge runs backwards and gives 01P(u1,1)du1-\int_{0}^{1}P(u^{1},1)\,\dd u^{1}. The left edge runs downward and gives 01Q(0,u2)du2-\int_{0}^{1}Q(0,u^{2})\,\dd u^{2}. Adding the four:

Mω  =  01 ⁣[P(u1,0)P(u1,1)]du1  +  01 ⁣[Q(1,u2)Q(0,u2)]du2, \oint_{\partial M}\omega \;=\; \int_{0}^{1}\!\Big[P(u^{1},0)-P(u^{1},1)\Big]\dd u^{1} \;+\; \int_{0}^{1}\!\Big[Q(1,u^{2})-Q(0,u^{2})\Big]\dd u^{2}, (3.5.40)

which is (3.5.38) plus (3.5.39) term for term. Hence Mdω=Mω\int_{M}\dd\omega=\oint_{\partial M}\omega on the square. \blacksquare

Notice what proved it: the Fundamental Theorem of Calculus, used twice, in one variable at a time. Everything else was bookkeeping about which edge runs which way, and that bookkeeping is what the antisymmetry of §1 automates.

Grind box — the same argument on a pp-dimensional cube, with the signs counted

Let MM be the unit cube 0ua10\le u^{a}\le1 in pp dimensions and let ω\omega be a (p1)(p-1)-form. By linearity it is enough to treat one basis term at a time, so take

ω  =  f  du1duk^dup, \omega \;=\; f\;\dd u^{1}\wedge\cdots\wedge\widehat{\dd u^{k}}\wedge\cdots\wedge\dd u^{p},

the hat meaning that the kk-th differential is omitted.

The left side. Applying (3.5.15), every term of dω\dd\omega except the one carrying kf\partial_{k}f contains a repeated differential and vanishes. Moving duk\dd u^{k} from the front into its home slot costs k1k-1 transpositions, so

dω  =  fuk  dukdu1duk^  =  (1)k1fuk  du1dup. \dd\omega \;=\; \pdv{f}{u^{k}}\;\dd u^{k}\wedge\dd u^{1}\wedge\cdots\widehat{\dd u^{k}}\cdots \;=\; (-1)^{k-1}\,\pdv{f}{u^{k}}\;\dd u^{1}\wedge\cdots\wedge\dd u^{p}.

Integrating over the cube and doing the uku^{k} integral by the Fundamental Theorem,

Mdω  =  (1)k1[fuk=1fuk=0]  du1duk^dup. \int_{M}\dd\omega \;=\; (-1)^{k-1}\int\Big[f\big|_{u^{k}=1} - f\big|_{u^{k}=0}\Big]\;\dd u^{1}\cdots\widehat{\dd u^{k}}\cdots\dd u^{p}.

The right side. The cube has 2p2p faces. On a face uj=constu^{j}=\text{const} with jkj\neq k, the parametrisation of the face has duj=0\dd u^{j}=0, and ω\omega contains duj\dd u^{j}, so ω\omega restricted to that face is zero and it contributes nothing. Only the two faces uk=0u^{k}=0 and uk=1u^{k}=1 survive, and on them ω\omega restricts to ff times the volume element of the face. The orientation induced on the face uk=1u^{k}=1 by the cube is the one for which the outward direction comes first, which is (1)k1(-1)^{k-1} times the natural ordering u1,,uk^,,upu^{1},\ldots,\widehat{u^{k}},\ldots,u^{p}. The face uk=0u^{k}=0 has outward direction reversed and so carries the opposite sign. Adding the two gives exactly the previous display. \blacksquare

At p=2p=2, k=1k=1 gives the QQ term of the main text and k=2k=2 gives the PP term, minus sign and all.

5.3 · General regions, and the cancellation made exact

Assume MM admits such a decomposition. By that we mean it can be chopped into finitely many cells, each smoothly mapped to a cube, with piecewise-smooth boundary. That is a real hypothesis, and the callout below says what it costs.

Then here is the argument. Chop MM into small cells, each of which can be mapped to a cube. Write the theorem for each cell and add. On the right, every interior face is shared by exactly two neighbouring cells. The orientation each cell induces on that shared face is the outward one for that cell, so the two neighbours induce opposite orientations on it. By the second remark of §5.1 the two contributions are exact negatives, and they cancel. What survives is the sum over faces lying on the outer boundary. Hence

  Mdω  =  Mω.   \boxed{\;\int_{M}\dd\omega \;=\; \oint_{\partial M}\omega.\;} (3.5.41)
Familiar ground — and the gap Chapter 0.7 admitted, now closed

The chopping argument is Chapter 0.7 §5.2's, word for word: add up local changes, watch the interior cancel in pairs, keep what is left on the boundary. Chapter 0.7's grind box then admitted that the argument as given there was the right picture and not a proof, because the local statement in each cell was the definition of divergence read backwards and carried an error term whose sum over many cells was not controlled.

That gap is now closed. Here the statement in each cell is not an approximation with a remainder. It is an exact identity, proved in §5.2 by the Fundamental Theorem of Calculus, and the sum of exact identities is an exact identity. Chapter 0.7 even predicted the mechanism precisely, saying that "the cancellation is enforced algebraically by the antisymmetry of forms instead". That is the second remark of §5.1 being used above.

What remains genuinely technical is the same in both chapters. The cells must fit the region, the region's boundary must be piecewise smooth, and the forms must be continuously differentiable. That is not small print we are hiding. It is the honest hypothesis, and it covers everything in this book.

5.4 · The four rows, recovered

Read (3.5.41) at each degree and each dimension.

p=1p=1, ω\omega a function. The region is a curve, its boundary is two points, and the right-hand side is the difference of ff between them. That is the gradient theorem of Chapter 0.7 §2.2. On a straight interval in one dimension it is the Fundamental Theorem of Calculus itself. So this one degree covers rows one and two of Chapter 0.7's table.

p=2p=2, ω\omega a one-form, in three dimensions. By §2.1, dω\dd\omega is the curl. The region is a surface, its boundary a closed curve, and (3.5.41) reads S(×F)dA=SFdr\int_{S}(\nabla\times\vv F)\cdot\dd\vv A=\oint_{\partial S}\vv F\cdot\dd\vv r. That is (0.7.41), Stokes' theorem, which is row three. Restricting to a flat region in the plane gives Green's theorem (0.7.42), which Chapter 0.7 had already shown was not independent.

p=3p=3, ω\omega a 22-form, in three dimensions. By (3.5.18), dω\dd\omega is the divergence, and (3.5.36) with the identification ωjk=ϵijkFi\omega_{jk}=\epsilon^{ijk}F^{i} turns Vω\oint_{\partial V}\omega into the flux FdA\oint\vv F\cdot\dd\vv A. So (3.5.41) is V(F)dV=VFdA\int_{V}(\nabla\cdot\vv F)\dd V=\oint_{\partial V}\vv F\cdot\dd\vv A, the divergence theorem (0.7.44), which is row four.

Count what has just happened. Four theorems, four names, four sets of hypotheses and four right-hand rules have become one line, proved once, in any number of dimensions, on any manifold, with no metric anywhere. That was Chapter 0.7's promise, and this is it.

In plain terms 3.5.5

Very early in the book, three results with three names and three right-hand rules were lined up in a table beside the fundamental theorem itself and shown to have identical shape: whatever a derivative accumulates throughout a region is bookkept entirely on that region's skin, whether the region is an interval with two ends, a curve, a patch of surface with a rim, or a solid with a hull. The table came with a promise that the four rows would become one line once a language existed for curved spaces of any dimension. That language now exists.

The proof is smaller than its reputation. On a single cube, integrating one variable at a time lets the ordinary Fundamental Theorem of Calculus, the first result of the toolkit, convert each integral into a difference of endpoint values; keeping track of which face runs which way is the remaining work, and the sign conventions built earlier do it automatically. Then chop any region into cubes and add. Every internal wall is counted twice with opposite orientation and cancels exactly.

One honest improvement over the earlier treatment deserves noting. There the cancellation argument came with an admission that it was a picture rather than a proof, because the statement inside each cell was approximate. Here it is exact, so adding many introduces no error, and the earlier chapter's own prediction of the repair is what happens.

6 · The volume element, and a divergence theorem that works on a manifold

Here is the destination. Chapter 3.6 has to integrate a scalar over all of spacetime and then discard a total derivative. Neither operation is available yet: "all of spacetime" needs a chart-independent notion of volume, and "discard a total derivative" needs a divergence theorem. This section builds both, and the tool it uses is Chapter 0.6 §8's Jacobian determinant, which will then be used a second time in Chapter 3.6 for a different purpose.

6.1 · There is essentially one top-degree form

In nn dimensions an nn-form has (nn)=1\binom{n}{n}=1 independent component, by §1.3's count. So any two top-degree forms are proportional, and the whole question is what the single component does under a change of chart.

Take Ω=Ω01(n1)dx0dxn1\Omega=\Omega_{01\cdots(n-1)}\,\dd x^{0}\wedge\cdots\wedge\dd x^{n-1}. Under a chart change xxx\to x', the components of a (0,n)(0,n) tensor pick up nn Jacobian factors:

Ωμ1μn  =  xν1xμ1xνnxμn  Ων1νn. \Omega'_{\mu_{1}\cdots\mu_{n}} \;=\; \pdv{x^{\nu_{1}}}{x'^{\mu_{1}}}\cdots\pdv{x^{\nu_{n}}}{x'^{\mu_{n}}}\;\Omega_{\nu_{1}\cdots\nu_{n}}. (3.5.42)

Take the one independent component, μ1μn=01(n1)\mu_{1}\cdots\mu_{n}=01\cdots(n-1). The right-hand side is then a totally antisymmetric contraction of nn copies of the Jacobian matrix, which by Chapter 0.4 §5's characterisation of the determinant is det(x/x)\det(\partial x/\partial x') times Ω01(n1)\Omega_{01\cdots(n-1)}. So the single component of a top-degree form transforms by one factor of the Jacobian determinant, and by nothing else.

6.2 · The metric determinant supplies exactly that factor

So what we need is a quantity that carries exactly one factor of the Jacobian determinant, to be cancelled against the one the coordinate volume carries. The metric supplies it, and the rest of this section is the check.

Write gdetgμνg\equiv\det g_{\mu\nu}, a number at each point, negative in Lorentzian signature because the metric has one positive and three negative eigenvalues. Chapter 3.3 §1 derived the transformation law of the metric components, (3.3.3):

gμν  =  xαxμxβxν  gαβ. g'_{\mu\nu} \;=\; \pdv{x^{\alpha}}{x'^{\mu}}\,\pdv{x^{\beta}}{x'^{\nu}}\;g_{\alpha\beta}. (3.5.43)

Read (3.5.43) as a matrix equation: g=J ⁣gJg'=J^{\!\top}gJ with Jαμ=xα/xμJ^{\alpha}{}_{\mu}=\partial x^{\alpha}/\partial x'^{\mu}. Take determinants, using det(AB)=detAdetB\det(AB)=\det A\det B and detJ ⁣=detJ\det J^{\!\top}=\det J from Chapter 0.4 §5:

g  =  (detJ)2gg  =  detJ  g. g' \;=\; \big(\det J\big)^{2}\,g \qquad\Longrightarrow\qquad \sqrt{-g'} \;=\; \abs{\det J}\;\sqrt{-g}. (3.5.44)

Two details there are worth naming. The square is why the sign of detJ\det J drops out, and the modulus is why taking the square root is harmless.

Now we have half of what we want. The other half is the measure. Chapter 0.6 §8 established dnx=detJ1dnx\dd^{n}x=\abs{\det J}^{-1}\dd^{n}x', so the coordinate volume shrinks by exactly the factor the metric determinant grows by. That is the pairing we were hoping for, so multiply the two together:

  g  dnx  =  g  dnx.   \boxed{\;\sqrt{-g}\;\dd^{n}x \;=\; \sqrt{-g'}\;\dd^{n}x' .\;} (3.5.45)

This is the invariant volume element. For any scalar ff, the quantity fgdnx\int f\sqrt{-g}\,\dd^{n}x is a number belonging to the region, not to the chart. Chapter 0.6 §8.3 told you this was coming and even checked the answer in two flat cases: plane polar coordinates give detgij=1r2=r\sqrt{\det g_{ij}}=\sqrt{1\cdot r^{2}}=r, which is (0.6.52), and spherical coordinates give r2sinθr^{2}\sin\theta, which is (0.6.53). Both were obtained there by cutting up regions by hand. (3.5.45) is the reason they came out that way. The sphere of radius aa gives a2a2sin2θ=a2sinθ\sqrt{a^{2}\cdot a^{2}\sin^{2}\theta}=a^{2}\sin\theta, which is the area element Chapter 3.4 §1 used to compute the region enclosed by its loop.

It is convenient to have the corresponding tensor. Define

εμνρσ    g  [μνρσ], \varepsilon_{\mu\nu\rho\sigma} \;\equiv\; \sqrt{-g}\;[\mu\nu\rho\sigma], (3.5.46)

Here [μνρσ][\mu\nu\rho\sigma] is the permutation symbol, +1+1 for an even permutation of 01230123, 1-1 for an odd one, 00 if any index repeats. The symbol on its own is not a tensor. It has the same entries in every chart, which no tensor with four indices does. But multiplying by g\sqrt{-g} supplies exactly the one Jacobian factor §6.1 showed a top-degree form needs, so εμνρσ\varepsilon_{\mu\nu\rho\sigma} is a tensor. The volume element is then the 44-form ε=g  dx0dx1dx2dx3\varepsilon=\sqrt{-g}\;\dd x^{0}\wedge\dd x^{1}\wedge\dd x^{2}\wedge\dd x^{3}.

6.3 · A formula for the derivative of a determinant

One more tool is needed, here and again in Chapter 3.6 §4. Here is what we are after: we want the change in detM\det M produced by a small change in the entries of MM.

Line 1: the cofactor expansion. Expanding a determinant along its aa-th row gives detM=bMabCab\det M=\sum_{b}M_{ab}C_{ab}, where the cofactor CabC_{ab} contains no entry from row aa. Because the cofactor does not contain MabM_{ab}, differentiating with respect to a single entry is immediate: (detM)/Mab=Cab\partial(\det M)/\partial M_{ab}=C_{ab}.

Line 2: the cofactors are the inverse. We would rather not carry cofactors around, and we do not have to. The adjugate formula of Chapter 0.4 says (M1)ba=Cab/detM(M^{-1})_{ba}=C_{ab}/\det M, so Cab=detM(M1)baC_{ab}=\det M\,(M^{-1})_{ba}.

Line 3: chain rule. Now put the two lines together for a general variation. For any δMab\delta M_{ab} of the entries,

δ(detM)  =  a,bCabδMab  =  detM  a,b(M1)baδMab  =  detM  tr ⁣(M1δM). \delta\big(\det M\big) \;=\; \sum_{a,b}C_{ab}\,\delta M_{ab} \;=\; \det M\;\sum_{a,b}\big(M^{-1}\big)_{ba}\,\delta M_{ab} \;=\; \det M\;\operatorname{tr}\!\big(M^{-1}\delta M\big). (3.5.47)

This is Jacobi's formula. The case we will use is the metric, so apply it with M=gμνM=g_{\mu\nu} and with δ\delta standing for any derivation whatever. That gives the version §6.4 reaches for:

δg  =  g  gμνδgμν. \delta g \;=\; g\;g^{\mu\nu}\,\delta g_{\mu\nu}. (3.5.48)

6.4 · The divergence, and the theorem Chapter 3.6 needs

The contracted connection. Take the Christoffel formula (3.3.50) and contract the upper index with the first lower one:

Γμμλ  =  12gμσ(μgσλ+λgσμσgμλ). \Gamma^{\mu}{}_{\mu\lambda} \;=\; \half\,g^{\mu\sigma}\Big(\partial_{\mu}g_{\sigma\lambda} + \partial_{\lambda}g_{\sigma\mu} - \partial_{\sigma}g_{\mu\lambda}\Big). (3.5.49)

The first and third terms cancel, and the cancellation is worth seeing rather than taking on trust. Relabel the two summed indices in the third term, swapping the names μ\mu and σ\sigma, which is legitimate because both are dummies. It becomes gσμμgσλ-g^{\sigma\mu}\partial_{\mu}g_{\sigma\lambda}, and gσμ=gμσg^{\sigma\mu}=g^{\mu\sigma} by symmetry of the metric, so it is minus the first term exactly. What is left is

Γμμλ  =  12gμσλgσμ  =  12gλg  =  λlng, \Gamma^{\mu}{}_{\mu\lambda} \;=\; \half\,g^{\mu\sigma}\,\partial_{\lambda}g_{\sigma\mu} \;=\; \frac{1}{2g}\,\partial_{\lambda}g \;=\; \partial_{\lambda}\ln\sqrt{-g}, (3.5.50)

the middle step being (3.5.48) with δ=λ\delta=\partial_{\lambda}, and the last being λlng=12λln(g)=λg/(2g)\partial_{\lambda}\ln\sqrt{-g}=\half\partial_{\lambda}\ln(-g)=\partial_{\lambda}g/(2g), where the minus signs inside the logarithm cancel between numerator and denominator.

The divergence. Now write out μVμ\nabla_{\mu}V^{\mu} from Chapter 3.3 §5 and substitute:

μVμ  =  μVμ+ΓμμλVλ  =  μVμ+Vμμlng  =  1gμ ⁣(gVμ), \nabla_{\mu}V^{\mu} \;=\; \partial_{\mu}V^{\mu} + \Gamma^{\mu}{}_{\mu\lambda}V^{\lambda} \;=\; \partial_{\mu}V^{\mu} + V^{\mu}\,\partial_{\mu}\ln\sqrt{-g} \;=\; \frac{1}{\sqrt{-g}}\,\partial_{\mu}\!\Big(\sqrt{-g}\,V^{\mu}\Big), (3.5.51)

the last step being the product rule read backwards. The covariant divergence of a vector field is an ordinary divergence in disguise, with the volume factor tucked inside. No Christoffel symbol survives, which is why this identity is worth having.

The theorem. Multiply (3.5.51) by g\sqrt{-g} and integrate. The result is an ordinary partial derivative under an ordinary integral sign, and §5 knows how to convert one of those into a boundary term. To do that we need something for §5 to act on, so build the 33-form ΘνρσεμνρσVμ\Theta_{\nu\rho\sigma}\equiv\varepsilon_{\mu\nu\rho\sigma}V^{\mu} and take its exterior derivative. Using (3.5.15) and the permutation-symbol identity 13![μνρσ][λνρσ]=δμλ\tfrac{1}{3!}[\mu\nu\rho\sigma][\lambda\nu\rho\sigma]=\delta^{\mu}{}_{\lambda}, every index sum collapses and

dΘ  =  μ ⁣(gVμ)  dx0dx1dx2dx3  =  (μVμ)ε. \dd\Theta \;=\; \partial_{\mu}\!\Big(\sqrt{-g}\,V^{\mu}\Big)\;\dd x^{0}\wedge\dd x^{1}\wedge\dd x^{2}\wedge\dd x^{3} \;=\; \big(\nabla_{\mu}V^{\mu}\big)\,\varepsilon. (3.5.52)

Feeding that into (3.5.41) gives the divergence theorem for a curved manifold:

  Ω(μVμ)g  d4x  =  ΩΘ.   \boxed{\;\int_{\Omega}\big(\nabla_{\mu}V^{\mu}\big)\,\sqrt{-g}\;\dd^{4}x \;=\; \oint_{\partial\Omega}\Theta.\;} (3.5.53)

The only use Chapter 3.6 makes of (3.5.53) is this: if a variation of the metric is required to vanish near the boundary, a term of the form μVμ\nabla_{\mu}V^{\mu} in the integrand contributes nothing at all. That single sentence is what turns a four-line variation into a field equation, and §4 of the next chapter also asks what it costs when the variation does not vanish there.

In plain terms 3.5.6

An integral needs to know how much region it is adding over, and on a curved space with arbitrary labels the naive answer is wrong by a factor that changes from place to place. The correction is a single number: the square root of the determinant of the rule for measuring distances. Relabelling squeezes the boxes of the coordinate grid by a determinant and stretches the distance rule by the square of the same determinant, so the square root of the second exactly undoes the first.

Two old results are collected in passing. The extra radial weight in polar coordinates, and the more elaborate weight in spherical coordinates, were obtained much earlier by dissecting regions by hand. Both are this one determinant, and so is the area element used to measure the patch enclosed by the circuit in the previous chapter.

The section ends by building the equipment the next chapter cannot do without. When the repaired derivative is applied to a field and the result summed over its slot, all the comparison coefficients cancel and what remains is an ordinary derivative times the volume factor. Combined with the previous section's theorem, any term of that shape inside an integral is bookkept on the boundary, and vanishes whenever the thing being varied is held fixed there. The next chapter's central computation lives or dies on that sentence.

7 · Dragging a tensor along a flow

The rest of the chapter changes subject, so here is the new destination. Chapter 3.3 §4 showed that the ordinary derivative of a tensor is not a tensor and needed a whole chapter of repair. Section 2 found one escape from that, available only to antisymmetric objects. Here is a second escape, available to everything, at the price of choosing a vector field first. It costs no connection, and in §8 it becomes the definition of a symmetry.

7.1 · The flow of a vector field

A vector field XX assigns a direction to every point, so it defines a system of ordinary differential equations,

dxμdε  =  Xμ(x(ε)), \dv{x^{\mu}}{\varepsilon} \;=\; X^{\mu}\big(x(\varepsilon)\big), (3.5.54)

The solutions of that system are the integral curves of XX, the paths that follow the arrows. Chapter 0.8 §1 guarantees a unique solution through each point for smooth XX. Following every point along its own curve for a parameter distance ε\varepsilon defines a map of the manifold to itself, ϕε\phi_{\varepsilon}, called the flow. Solving (3.5.54) to first order by Taylor's theorem (Chapter 0.3 §1),

ϕε(x)μ  =  xμ  +  εXμ(x)  +  O(ε2). \phi_{\varepsilon}(x)^{\mu} \;=\; x^{\mu} \;+\; \varepsilon\,X^{\mu}(x) \;+\; O(\varepsilon^{2}). (3.5.55)

The idea. A tensor at the point ϕε(x)\phi_{\varepsilon}(x) still cannot be compared with a tensor at xx. That was Chapter 3.2 §5's central prohibition and nothing has repealed it. But the flow supplies a map between the two points, and any smooth map between points carries tensors along with it, by contracting each index with the map's Jacobian. So here is the recipe: carry the tensor at ϕε(x)\phi_{\varepsilon}(x) back to xx along the flow, subtract the tensor already at xx, divide by ε\varepsilon, and let ε0\varepsilon\to0. That is the Lie derivative LX\mathcal{L}_{X}.

The price is visible in the recipe. The answer depends on the whole field XX near xx, not just on XX at xx, because a nearby direction is needed to define the map. So LXT\mathcal{L}_{X}T is not a derivative "in the direction XX" in the way XT\nabla_{X}T is. It is a comparison of TT with the version of itself that the flow of XX would produce.

7.2 · Scalars, then vectors, with every index step shown

A scalar. Carrying a number requires no Jacobian, so the recipe reads

LXf  =  limε0f(x+εX)f(x)ε  =  Xμμf  =  X(f), \mathcal{L}_{X}f \;=\; \lim_{\varepsilon\to0}\frac{f\big(x+\varepsilon X\big)-f(x)}{\varepsilon} \;=\; X^{\mu}\,\partial_{\mu}f \;=\; X(f), (3.5.56)

which is the directional derivative, and also the way Chapter 3.2 §4 defined a tangent vector in the first place.

A vector. Now the Jacobian is needed. From (3.5.55), writing x=ϕε(x)x'=\phi_{\varepsilon}(x),

Λμν    xμxν  =  δμν  +  ενXμ  +  O(ε2). \Lambda^{\mu}{}_{\nu} \;\equiv\; \pdv{x'^{\mu}}{x^{\nu}} \;=\; \delta^{\mu}{}_{\nu} \;+\; \varepsilon\,\partial_{\nu}X^{\mu} \;+\; O(\varepsilon^{2}). (3.5.57)

Carrying a vector forward along the flow multiplies its components by (3.5.57). We want to carry backward, so we need the inverse matrix, which to first order is obtained by flipping the sign of the small term:

(Λ1)μν  =  δμν    ενXμ  +  O(ε2), \big(\Lambda^{-1}\big)^{\mu}{}_{\nu} \;=\; \delta^{\mu}{}_{\nu} \;-\; \varepsilon\,\partial_{\nu}X^{\mu} \;+\; O(\varepsilon^{2}), (3.5.58)

as one checks by multiplying the two and seeing the ε\varepsilon terms cancel.

Now assemble. The vector to be carried back is YY evaluated at the shifted point, and Taylor's theorem gives Yν(x+εX)=Yν(x)+εXλλYν+O(ε2)Y^{\nu}(x+\varepsilon X)=Y^{\nu}(x)+\varepsilon X^{\lambda}\partial_{\lambda}Y^{\nu}+O(\varepsilon^{2}). Contract with (3.5.58), keeping terms to first order only:

(dragged back Y)μ  =  (δμνενXμ)(Yν+εXλλYν)=  Yμ  +  εXλλYμ    εYννXμ  +  O(ε2). \begin{aligned} \big(\text{dragged back }Y\big)^{\mu} \;&=\; \Big(\delta^{\mu}{}_{\nu}-\varepsilon\,\partial_{\nu}X^{\mu}\Big)\Big(Y^{\nu}+\varepsilon\,X^{\lambda}\partial_{\lambda}Y^{\nu}\Big)\\[3pt] &=\; Y^{\mu} \;+\; \varepsilon\,X^{\lambda}\partial_{\lambda}Y^{\mu} \;-\; \varepsilon\,Y^{\nu}\partial_{\nu}X^{\mu} \;+\; O(\varepsilon^{2}). \end{aligned} (3.5.59)

Two things happened in the second line and both are worth naming. The δμν\delta^{\mu}{}_{\nu} hit both factors, producing the first and second terms. The ενXμ-\varepsilon\partial_{\nu}X^{\mu} hit only YνY^{\nu}, because its other product is second order.

That is the dragged-back vector. The recipe now says to finish the job: subtract YμY^{\mu}, divide by ε\varepsilon and take the limit.

  (LXY)μ  =  XλλYμ    YλλXμ  =  [X,Y]μ.   \boxed{\;\big(\mathcal{L}_{X}Y\big)^{\mu} \;=\; X^{\lambda}\,\partial_{\lambda}Y^{\mu} \;-\; Y^{\lambda}\,\partial_{\lambda}X^{\mu} \;=\; [X,Y]^{\mu}.\;} (3.5.60)

Let's look at that last equality, because it is not something we arranged. It is (3.2.32), the components of the commutator of two vector fields, which Chapter 3.2 §8 built by an entirely different route: it demanded that the composition of two derivations be a derivation. Two independent constructions have produced the same object, and the bracket has now acquired a second meaning worth carrying. [X,Y][X,Y] measures the failure of YY to be carried into itself by the flow of XX.

7.3 · Covectors, and then everything

A covector is carried the other way, so its components are contracted with (3.5.57) rather than its inverse. Repeating the computation with the sign of the Jacobian term flipped and the index in the other position,

(LXω)μ  =  Xλλωμ  +  ωλμXλ. \big(\mathcal{L}_{X}\omega\big)_{\mu} \;=\; X^{\lambda}\,\partial_{\lambda}\omega_{\mu} \;+\; \omega_{\lambda}\,\partial_{\mu}X^{\lambda}. (3.5.61)

The pattern is now visible and it is the same one Chapter 3.3 §5.4 found for the covariant derivative: one transport term, plus one correction per index, with the sign set by whether the index is up or down. For the case this chapter needs, a (0,2)(0,2) tensor,

(LXT)μν  =  XλλTμν  +  TλνμXλ  +  TμλνXλ. \big(\mathcal{L}_{X}T\big)_{\mu\nu} \;=\; X^{\lambda}\,\partial_{\lambda}T_{\mu\nu} \;+\; T_{\lambda\nu}\,\partial_{\mu}X^{\lambda} \;+\; T_{\mu\lambda}\,\partial_{\nu}X^{\lambda}. (3.5.62)

7.4 · The partial derivatives may be replaced by covariant ones

Something in the last three equations should be bothering you. Equations (3.5.60) to (3.5.62) contain plain partial derivatives, and by Chapter 3.3 §4 a plain partial derivative of a tensor is not a tensor. The combinations above are tensors anyway, because the construction that produced them was manifestly tensorial. But there is a cleaner way to see it, and §8 needs the result.

Claim. In (3.5.60) and (3.5.62), every \partial may be replaced by \nabla without changing the answer, for any torsion-free connection.

Proof for the vector case. Write the two covariant derivatives out using Chapter 3.3 §5:

XλλYμYλλXμ  =  XλλYμ+XλΓμλνYν    YλλXμYλΓμλνXν. X^{\lambda}\nabla_{\lambda}Y^{\mu} - Y^{\lambda}\nabla_{\lambda}X^{\mu} \;=\; X^{\lambda}\partial_{\lambda}Y^{\mu} + X^{\lambda}\Gamma^{\mu}{}_{\lambda\nu}Y^{\nu} \;-\; Y^{\lambda}\partial_{\lambda}X^{\mu} - Y^{\lambda}\Gamma^{\mu}{}_{\lambda\nu}X^{\nu}. (3.5.63)

Look only at the two Γ\Gamma terms. In the second of them, relabel the two dummy indices λ\lambda and ν\nu, swapping their names. That is legitimate, since both are summed. It becomes XλΓμνλYν-X^{\lambda}\Gamma^{\mu}{}_{\nu\lambda}Y^{\nu}. Now use Γμνλ=Γμλν\Gamma^{\mu}{}_{\nu\lambda}=\Gamma^{\mu}{}_{\lambda\nu}, the torsion-free symmetry of Chapter 3.3 §7.2, and it is exactly minus the first. They cancel, and what remains is (3.5.60). \blacksquare

The (0,2)(0,2) case is the identical two moves applied to three Γ\Gamma terms, and it is done in the next section where it is used.

In plain terms 3.5.7

Comparing a field's values at two different points has been forbidden since the chapter on manifolds, and the chapters since bought their way round the ban by installing a rule for carrying things from place to place. Here is a second way to buy it, and the currency is different. Choose a direction field, let every point drift along its own arrow for a moment, and use that drift as transport. Carry the field back from where it drifted, and compare with what was already there.

Two features distinguish this from the earlier construction. It needs no rule for comparing directions and no notion of distance, so it is available on a bare space before any geometry is laid down. And it depends on the chosen field everywhere nearby rather than merely at the point, which is why the answer compares the field with a dragged copy of itself rather than giving a rate of change along a line.

For two direction fields the result is an object built several chapters ago by a different argument, namely the failure of two successive drifts to commute. That agreement is worth more than its algebra, since the bracket measuring whether two operations commute also measures whether one field is carried into itself by the other's flow. The next section drags the distance rule itself, and asks when dragging leaves it alone.

8 · Killing vectors: the directions in which nothing changes

Here is the destination. We ask which vector fields have flows that leave the metric exactly as they found it, translate that demand into an equation on the field, and check it on two spaces. The equation is the geometric form of the word "symmetry", and §9 turns each solution of it into a conserved quantity.

8.1 · The definition, and Killing's equation

A vector field ξ\xi is a Killing vector field if dragging the metric along its flow changes nothing:

Lξg  =  0. \mathcal{L}_{\xi}\,g \;=\; 0. (3.5.64)

That is compact, and compact is no use for computing with. Our next goal is a condition on the components of ξ\xi, so write the Lie derivative out with the (0,2)(0,2) formula (3.5.62), taking TT to be the metric itself:

ξλλgμν  +  gλνμξλ  +  gμλνξλ  =  0. \xi^{\lambda}\,\partial_{\lambda}g_{\mu\nu} \;+\; g_{\lambda\nu}\,\partial_{\mu}\xi^{\lambda} \;+\; g_{\mu\lambda}\,\partial_{\nu}\xi^{\lambda} \;=\; 0. (3.5.65)

Now simplify, one move at a time.

Move 1: replace the partial derivatives by covariant ones. Section 7.4 licensed this. To see it explicitly here, expand the three covariant derivatives: ξλλgμν\xi^{\lambda}\nabla_{\lambda}g_{\mu\nu} contributes ξλΓσλμgσνξλΓσλνgμσ-\xi^{\lambda}\Gamma^{\sigma}{}_{\lambda\mu}g_{\sigma\nu}-\xi^{\lambda}\Gamma^{\sigma}{}_{\lambda\nu}g_{\mu\sigma} beyond the partial derivative, while gλνμξλg_{\lambda\nu}\nabla_{\mu}\xi^{\lambda} contributes +gλνΓλμσξσ+g_{\lambda\nu}\Gamma^{\lambda}{}_{\mu\sigma}\xi^{\sigma} and gμλνξλg_{\mu\lambda}\nabla_{\nu}\xi^{\lambda} contributes +gμλΓλνσξσ+g_{\mu\lambda}\Gamma^{\lambda}{}_{\nu\sigma}\xi^{\sigma}. Relabelling the dummies λσ\lambda\leftrightarrow\sigma in the last two and using the symmetry of Γ\Gamma in its lower indices makes them cancel the first two exactly.

Move 2: kill the first term. Metric compatibility says λgμν=0\nabla_{\lambda}g_{\mu\nu}=0. That was imposed in Chapter 3.3 §7.1 as Demand 1, and written out in components there as (3.3.41). So the first term is gone altogether.

Move 3: pull the metric inside the derivative. The same fact, g=0\nabla g=0, lets the metric pass through a covariant derivative, so gλνμξλ=μ(gλνξλ)=μξνg_{\lambda\nu}\nabla_{\mu}\xi^{\lambda}=\nabla_{\mu}\big(g_{\lambda\nu}\xi^{\lambda}\big)=\nabla_{\mu}\xi_{\nu}, which is the index-lowering of Chapter 3.3 §1.1 performed under the derivative. The third term becomes νξμ\nabla_{\nu}\xi_{\mu} by the same move. What is left is

  μξν  +  νξμ  =  0.   \boxed{\;\nabla_{\mu}\xi_{\nu} \;+\; \nabla_{\nu}\xi_{\mu} \;=\; 0.\;} (3.5.66)

This is Killing's equation. It says the covariant derivative of ξ\xi, with both indices down, is antisymmetric.

Count the equations against the unknowns, because the count is the interesting part. In four dimensions Killing's equation is ten equations, one for each symmetric index pair, imposed on four unknown functions. That is heavily overdetermined. It is why most metrics have no Killing vectors at all, and why the ones that do are special.

The practical test, which is the one Chapter 3.7 uses

Suppose that in some chart the metric components do not depend on one of the coordinates, say xax^{a}. The index aa here names one particular coordinate and is not being summed. Take ξ=a\xi=\partial_{a}, whose components are ξμ=δμa\xi^{\mu}=\delta^{\mu}{}_{a}, constants. Then in (3.5.65) the second and third terms vanish because the derivative of a constant is zero, and the first is agμν\partial_{a}g_{\mu\nu}, which vanishes by hypothesis.

So every coordinate the metric does not mention supplies a Killing vector, and finding Killing vectors is often no harder than looking at the metric. The converse is not true, and the failure of the converse matters. A Killing vector need not be a coordinate direction in the chart you happen to be using, and two of the sphere's three below are invisible in the usual chart.

8.2 · Flat spacetime has ten

Take gμν=ημνg_{\mu\nu}=\eta_{\mu\nu} in Cartesian coordinates, where every component is constant and every Γ\Gamma vanishes, so =\nabla=\partial and (3.5.66) reads μξν+νξμ=0\partial_{\mu}\xi_{\nu}+\partial_{\nu}\xi_{\mu}=0.

The four translations. ξ=a\xi=\partial_{a} for each a=0,1,2,3a=0,1,2,3 has constant components, so both terms vanish. Four solutions.

The six rotations and boosts. Try ξμ=ωμνxν\xi_{\mu}=\omega_{\mu\nu}x^{\nu} with ω\omega a constant array. Then μξν=ωνμ\partial_{\mu}\xi_{\nu}=\omega_{\nu\mu}, and (3.5.66) becomes ωνμ+ωμν=0\omega_{\nu\mu}+\omega_{\mu\nu}=0: the array must be antisymmetric. An antisymmetric 4×44\times4 array has (42)=6\binom{4}{2}=6 independent entries, by §1.3's count. Six more solutions.

Ten in all, and they are exactly the ten generators of the Poincaré group: four translations, three rotations and three boosts. Six of those ten are the Lorentz group, which Chapter 2.3 obtained by demanding that the interval be preserved and counted by hand in its Problem 4. The four translations are new here, and they arrive on the same footing as the rest.

Take a moment over what that means. Killing vectors are the infinitesimal version of "transformations that leave the geometry alone". In flat spacetime that turns out to be the whole of special relativity's symmetry group, recovered from a differential equation.

8.3 · The sphere has three, and one obvious candidate is not among them

Take the sphere of radius aa, with gθθ=a2g_{\theta\theta}=a^{2} and gϕϕ=a2sin2θg_{\phi\phi}=a^{2}\sin^{2}\theta.

The easy one. No metric component mentions ϕ\phi, so by the practical test ξ=ϕ\xi=\partial_{\phi} is Killing. It is the rotation about the polar axis.

The two hidden ones. The sphere is no more special about its poles than about anywhere else, so rotations about the other two axes must also be symmetries. In this chart they are not coordinate directions, so the practical test cannot find them and we have to write them down ourselves. They are

ξ(2)  =  cosϕ  θ    cotθsinϕ  ϕ,ξ(3)  =  sinϕ  θ    cotθcosϕ  ϕ, \xi_{(2)} \;=\; \cos\phi\;\partial_{\theta} \;-\; \cot\theta\,\sin\phi\;\partial_{\phi}, \qquad \xi_{(3)} \;=\; -\sin\phi\;\partial_{\theta} \;-\; \cot\theta\,\cos\phi\;\partial_{\phi}, (3.5.67)

and substituting either into (3.5.65) gives zero in all three components. (Checked symbolically: for both fields all three components of Lξg\mathcal{L}_{\xi}g vanish identically, and the covariant form (3.5.66) gives the same array as the partial-derivative form (3.5.65), which is Move 1 verified rather than trusted.)

The candidate that fails. Try ξ=θ\xi=\partial_{\theta}, "push everything south". Its components are constant, so (3.5.65) reduces to its first term, θgμν\partial_{\theta}g_{\mu\nu}, and only one metric component depends on θ\theta:

(Lθg)ϕϕ  =  θ(a2sin2θ)  =  a2sin2θ,all other components zero. \big(\mathcal{L}_{\partial_{\theta}}\,g\big)_{\phi\phi} \;=\; \partial_{\theta}\big(a^{2}\sin^{2}\theta\big) \;=\; a^{2}\sin2\theta, \qquad\text{all other components zero.} (3.5.68)

So θ\partial_{\theta} is not a Killing vector. But (3.5.68) does more than say so. It says precisely how the field fails. The flow leaves north–south distances untouched and stretches east–west distances, at a rate that vanishes at the equator, where sin2θ=0\sin2\theta=0, and is largest at colatitude 4545^{\circ}. That is a sharper statement than "the sphere is not symmetric under pushing south", and the figure measures it.

0.0°
35°
left ξ=∂/∂φ : N–S edge 0.383972435 → 0.383972435 | E–W edge 0.280302322 → 0.280302322
right ξ=∂/∂θ : N–S edge 0.383972435 → 0.383972435 | E–W edge 0.280302322 → 0.280302322
at colatitude θ₀+s = 35.0° : measured d(ℓ²_EW)/ds ÷ (Δφ)² = 0.939691395 predicted a² sin 2(θ₀+s) = 0.939692621
A flow the metric cannot feel, and one it can. Both globes carry the same little four-cornered patch, drawn in blue at its starting place and in purple after the flow has run for parameter distance ss; the thin green curves are the flow lines. On the left the field is ϕ\partial_{\phi}, which turns the globe about its axis; on the right it is θ\partial_{\theta}, which slides every point toward the south pole. The two readouts give the measured lengths of the patch's north–south and east–west edges, computed as great-circle distances between the moved corners — an intrinsic measurement, using no picture from outside. On the left both are unchanged in every digit displayed, for every ss and every starting latitude, which is (3.5.64) being true. On the right the north–south edge is likewise unchanged while the east–west edge changes, and the third readout compares its measured rate of change against a2sin2θa^{2}\sin2\theta from (3.5.68). They agree. A Killing vector is not a direction along which the picture looks the same; it is a direction along which every length anyone could measure comes out the same. The right-hand flow is perfectly respectable — it is just not a symmetry, and the single non-zero component of its Lie derivative says exactly which measurement notices.
In plain terms 3.5.8

Drag the rule for measuring distances along a chosen field, demand that the result be nothing, and you have what this book means by a symmetry of a space. Written out and cleaned up using the fact that the repaired derivative ignores the distance rule, the demand becomes one compact equation on the chosen field. It is badly overdetermined, ten conditions on four functions, which is why a space picked at random has no symmetries at all, and why the ones that do are the ones anybody can solve.

Flat spacetime turns out to have exactly ten independent solutions, and they are the four shifts, the three turns and the three boosts of the previous part. That closes a loop: the transformations obtained there by insisting the interval be preserved reappear as solutions of a differential equation, with nothing about relativity fed in.

The globe supplies the cautionary case. Turning it about its axis leaves every measurable distance as it was, and so does turning it about either of the other two axes, though those are invisible in the usual grid of latitude and longitude. Sliding everything southward is not a symmetry, and the calculation says precisely how it fails: north-south spacings are untouched while east-west spacings stretch, most sharply halfway to the pole and not at all at the equator. The picture measures that rate and returns the predicted number.

9 · A conserved quantity along every geodesic

This is the section Chapter 3.7 spends. The destination has to be stated twice, because the result is obtained twice. First: for any Killing vector ξ\xi and any geodesic with tangent uu, the single number ξμuμ\xi_{\mu}u^{\mu} is the same at every point of the geodesic. Four lines. Second: that number is the Noether charge of Chapter 1.4 for the symmetry generated by ξ\xi, so the theorem is not an analogue of Noether's but an instance of it.

9.1 · The four-line derivation

Let xμ(τ)x^{\mu}(\tau) be a geodesic with tangent uμ=dxμ/dτu^{\mu}=\dd x^{\mu}/\dd\tau. The geodesic equation of Chapter 3.3 §8 then reads uννuμ=0u^{\nu}\nabla_{\nu}u^{\mu}=0, which says the curve parallel-transports its own tangent. Let ξ\xi satisfy Killing's equation (3.5.66). We want to show that a particular number does not change along the curve, so define QξμuμQ\equiv\xi_{\mu}u^{\mu} and differentiate it.

Line 1. QQ is a scalar, so its rate of change along the curve is the directional derivative of it, and for a scalar the covariant and ordinary derivatives agree:

dQdτ  =  uνν(ξμuμ). \dv{Q}{\tau} \;=\; u^{\nu}\,\nabla_{\nu}\big(\xi_{\mu}u^{\mu}\big). (3.5.69)

Line 2: Leibniz. There are two factors inside that derivative, so split it. The covariant derivative obeys the product rule (Chapter 3.3 §5.4), which gives

dQdτ  =  uνuμνξμ  +  ξμuννuμ. \dv{Q}{\tau} \;=\; u^{\nu}u^{\mu}\,\nabla_{\nu}\xi_{\mu} \;+\; \xi_{\mu}\,u^{\nu}\nabla_{\nu}u^{\mu}. (3.5.70)

Line 3: the second term is the geodesic equation. Look at it: it is exactly uννuμu^{\nu}\nabla_{\nu}u^{\mu}, which is zero. So it vanishes and only the first survives:

dQdτ  =  uμuννξμ. \dv{Q}{\tau} \;=\; u^{\mu}u^{\nu}\,\nabla_{\nu}\xi_{\mu}. (3.5.71)

Line 4: only the symmetric part survives, and it is zero. This is the index step to spell out. The factor uμuνu^{\mu}u^{\nu} is symmetric under exchanging μ\mu and ν\nu, because it is a product of two copies of the same object. Contracting a symmetric array with any array picks out only the symmetric part of the second, since the antisymmetric part contributes equal and opposite amounts under the exchange. Concretely, relabel the dummies μν\mu\leftrightarrow\nu in (3.5.71), add the result to the original and divide by two:

dQdτ  =  12uμuν(νξμ+μξν)  =  0, \dv{Q}{\tau} \;=\; \half\,u^{\mu}u^{\nu}\Big(\nabla_{\nu}\xi_{\mu} + \nabla_{\mu}\xi_{\nu}\Big) \;=\; 0, (3.5.72)

And the bracket in that last line vanishes, because it is the left-hand side of Killing's equation (3.5.66). So the rate of change of QQ is zero at every point of the curve, which is what we set out to prove:

  Q  =  ξμuμ  =  constant along every geodesic.   \boxed{\;Q \;=\; \xi_{\mu}\,u^{\mu} \;=\; \text{constant along every geodesic.}\;} (3.5.73)

\blacksquare Four lines, and the only inputs were the geodesic equation and Killing's equation. Note where each was spent: the geodesic equation killed the second term of (3.5.70), and Killing's equation killed the first. Neither could have done the other's job.

Checked numerically

The three Killing fields of §8.3 were used to build three charges QQ, and a geodesic of the unit sphere was integrated from a generic starting point and direction over a long parameter interval. All three charges held constant to about one part in 101410^{14}, which is the integrator's own accuracy. The same computation was then run with ξ=θ\xi=\partial_{\theta}, which is not a Killing vector, and that charge changed by 82%82\% over the same interval. The theorem is doing work, not bookkeeping.

9.2 · The same result as Noether's theorem

Now the identification the plan of this book has been pointing at. Chapter 1.4 proved that every continuous symmetry of an action yields a conserved quantity, and gave it as (1.4.13): Q=ipiKiFQ=\sum_{i}p_{i}K_{i}-F, with KiK_{i} the infinitesimal change in the coordinates and pip_{i} the canonical momentum. We now run that theorem on a particle in a curved spacetime and show that it returns (3.5.73).

The action. Chapter 3.3 §3 gave S=mc2dτS=-mc^{2}\int\dd\tau as (3.3.15), and Chapter 3.3 §8.2 showed that extremising it produces the geodesic equation, working with the integrand F=gαβx˙αx˙βF=\sqrt{g_{\alpha\beta}\dot x^{\alpha}\dot x^{\beta}}. For the present purpose the square root is a nuisance, and it can be removed. Compare the Euler–Lagrange equations of FF with those of 12F2\half F^{2}: by the product and chain rules,

ddτ(12F2)x˙μ(12F2)xμ  =  F[ddτFx˙μFxμ]  +  dFdτFx˙μ. \dv{}{\tau}\pdv{\big(\half F^{2}\big)}{\dot x^{\mu}} - \pdv{\big(\half F^{2}\big)}{x^{\mu}} \;=\; F\left[\dv{}{\tau}\pdv{F}{\dot x^{\mu}} - \pdv{F}{x^{\mu}}\right] \;+\; \dv{F}{\tau}\,\pdv{F}{\dot x^{\mu}}. (3.5.74)

Along a timelike geodesic parametrised by proper time, F2=gμνuμuν=c2F^{2}=g_{\mu\nu}u^{\mu}u^{\nu}=c^{2} is constant by (3.3.58), so dF/dτ=0\dd F/\dd\tau=0 and the last term drops. The two brackets then vanish together. So on proper-time-parametrised curves the two Lagrangians have the same extremals, and we may use

L  =  12mgμν(x)x˙μx˙ν,x˙μdxμdτ. L \;=\; \half\,m\,g_{\mu\nu}(x)\,\dot x^{\mu}\dot x^{\nu}, \qquad \dot x^{\mu}\equiv\dv{x^{\mu}}{\tau}. (3.5.75)

The canonical momentum. Differentiate (3.5.75) with respect to x˙λ\dot x^{\lambda}. The velocity appears twice, so the product rule gives two terms, and the symmetry of gg makes them equal:

pλ  =  Lx˙λ  =  12m(gλνx˙ν+gμλx˙μ)  =  mgλμuμ  =  muλ. p_{\lambda} \;=\; \pdv{L}{\dot x^{\lambda}} \;=\; \half m\Big(g_{\lambda\nu}\dot x^{\nu} + g_{\mu\lambda}\dot x^{\mu}\Big) \;=\; m\,g_{\lambda\mu}u^{\mu} \;=\; m\,u_{\lambda}. (3.5.76)

So the momentum conjugate to a coordinate is the four-velocity with its index lowered, times the mass. That is Chapter 2.5's four-momentum, now with a position-dependent metric doing the lowering.

The symmetry. Take the transformation xμxμ+ϵξμ(x)x^{\mu}\to x^{\mu}+\epsilon\,\xi^{\mu}(x), so that Chapter 1.4's KK is ξμ\xi^{\mu}. Compute the change in LL. Two things vary: the metric, because it is evaluated at a shifted point, and the velocities, because the shift depends on position:

δL  =  12mϵ[  ξλ(λgμν)x˙μx˙ν  +  2gμνx˙μd(ξν)dτ  ]. \delta L \;=\; \half m\,\epsilon\,\Big[\;\xi^{\lambda}\big(\partial_{\lambda}g_{\mu\nu}\big)\dot x^{\mu}\dot x^{\nu} \;+\; 2\,g_{\mu\nu}\,\dot x^{\mu}\,\dv{\big(\xi^{\nu}\big)}{\tau}\;\Big]. (3.5.77)

Handle the second term. By the chain rule dξν/dτ=x˙λλξν\dd\xi^{\nu}/\dd\tau=\dot x^{\lambda}\partial_{\lambda}\xi^{\nu}, so that term is 2gμνx˙μx˙λλξν2g_{\mu\nu}\dot x^{\mu}\dot x^{\lambda}\partial_{\lambda}\xi^{\nu}. Relabel its dummy indices so that the two velocities carry the labels μ\mu and ν\nu: swapping the names ν\nu and λ\lambda turns it into 2gμλx˙μx˙ννξλ2g_{\mu\lambda}\dot x^{\mu}\dot x^{\nu}\partial_{\nu}\xi^{\lambda}. Using the symmetry of gx˙x˙g\dot x\dot x in μν\mu\nu, that equals (gλνμξλ+gμλνξλ)x˙μx˙ν\big(g_{\lambda\nu}\partial_{\mu}\xi^{\lambda}+g_{\mu\lambda}\partial_{\nu}\xi^{\lambda}\big)\dot x^{\mu}\dot x^{\nu}. Substituting back,

δL  =  12mϵ  (ξλλgμν+gλνμξλ+gμλνξλ=  (Lξg)μν)x˙μx˙ν. \delta L \;=\; \half m\,\epsilon\;\Big(\underbrace{\xi^{\lambda}\partial_{\lambda}g_{\mu\nu} + g_{\lambda\nu}\partial_{\mu}\xi^{\lambda} + g_{\mu\lambda}\partial_{\nu}\xi^{\lambda}}_{\textstyle =\;\big(\mathcal{L}_{\xi}g\big)_{\mu\nu}}\Big)\,\dot x^{\mu}\dot x^{\nu}. (3.5.78)

Look at the underbraced bracket. It is (3.5.65) exactly. The change in the Lagrangian under a coordinate shift is the Lie derivative of the metric, contracted with two velocities, which is the cleanest possible statement of what a Killing vector is for. If ξ\xi is Killing then δL=0\delta L=0, so the symmetry hypothesis of Chapter 1.4 §1.2 is satisfied with the allowed total derivative FF equal to zero.

The charge. Feed Kμ=ξμK^{\mu}=\xi^{\mu}, pμ=muμp_{\mu}=mu_{\mu} and F=0F=0 into (1.4.13):

QNoether  =  pμξμ    0  =  mξμuμ, Q_{\text{Noether}} \;=\; p_{\mu}\,\xi^{\mu} \;-\; 0 \;=\; m\,\xi_{\mu}u^{\mu}, (3.5.79)

which is mm times (3.5.73). The two derivations agree, and they agree for a reason rather than by luck: §9.1's Line 4 discarded the antisymmetric part of ξ\nabla\xi, and §9.2's (3.5.78) discarded it too, both times because the velocity appears twice.

Why this is the chapter's most expensive result

Chapter 3.7 has to solve the geodesic equation outside a spherical mass. Written out, that is four coupled second-order differential equations for t(τ),r(τ),θ(τ),ϕ(τ)t(\tau),r(\tau),\theta(\tau),\phi(\tau) with position-dependent coefficients, which is not a tractable problem by hand.

The Schwarzschild metric does not mention tt and does not mention ϕ\phi. By §8.1's practical test that is two Killing vectors, and by (3.5.73) two constants of the motion. Chapter 3.7 will call them EE and LL, and by (3.5.79) they are precisely the energy and angular momentum of Chapter 1.4 §3. Two of the four equations are thereby integrated once and for all. Spherical symmetry disposes of θ\theta. What is left is a single first-order equation for r(τ)r(\tau), which is Chapter 0.8's business and which Chapter 3.7 reads as motion in an effective potential.

Without this section, Chapter 3.7 would be a wall of algebra with no visible structure. With it, that chapter is mostly Chapter 1.3's phase-plane reasoning applied to one variable. It is worth noticing that the saving is not computational cleverness. It is symmetry, converted into conserved quantities by a theorem proved four parts ago for pendulums.

In plain terms 3.5.9

Take any direction along which the geometry does not change, and any freely falling body. The component of the body's velocity along that direction, measured with the geometry's own rule for taking components, never changes. The proof runs to four lines and each input is spent once: the free-fall condition removes one term and the no-change condition the other, and neither could have removed the term the other did.

Then comes the identification making this more than a trick. Feed the same situation to the theorem proved long ago for mechanical systems, which says every continuous symmetry of the quantity attached to a history supplies something conserved. The change in that quantity under sliding the coordinates along the chosen direction turns out to be the dragging operation of two sections ago, applied to the distance rule and contracted twice with the velocity. So the geometric and mechanical conditions are one condition, and the conserved quantity produced is the same quantity.

This is where the chapter earns its place. Two chapters from now the geometry outside a star will mention neither time nor the angle around the axis, which by the practical test hands over two constants at once. Those constants are the energy and the angular momentum, and having them reduces a tangle of four coupled equations to one equation in one variable, of exactly the kind the toolkit taught how to read.

10 · Electromagnetism in three symbols

A language is worth having if it says something short that was long. Here is the test, and here is the destination. All of classical electromagnetism, including charge conservation and gauge freedom and the version valid in curved spacetime, written as F=dAF=\dd A, dF=0\dd F=0, d ⁣ ⁣F=μ0 ⁣J\dd\!\star\!F=\mu_{0}\star\!J. And, more useful than the compression itself, a clean statement of which half of the subject needs a metric and which half does not.

10.1 · The one new operation

Section 1.3's table is symmetric: in four dimensions the counts run 1,4,6,4,11,4,6,4,1, so pp-forms and (4p)(4-p)-forms have the same number of components. An operation exchanging them therefore has a chance of existing, and the volume tensor (3.5.46) supplies it. Define the Hodge star of a pp-form by contracting it into the last pp slots of ε\varepsilon:

(ω)μ1μ4p    1p!  εμ1μ4pν1νp  ων1νp. \big(\star\omega\big)_{\mu_{1}\cdots\mu_{4-p}} \;\equiv\; \frac{1}{p!}\;\varepsilon_{\mu_{1}\cdots\mu_{4-p}\,\nu_{1}\cdots\nu_{p}}\;\omega^{\nu_{1}\cdots\nu_{p}}. (3.5.80)

Note what just happened. The indices of ω\omega were raised to be contracted, and raising needs the inverse metric. So \star is the first operation in this chapter that requires a metric. Everything before it needed none: forms, the wedge, d\dd, d2=0\dd^{2}=0, Stokes. Keep that distinction in view for §10.3, where it separates the two halves of Maxwell's equations.

The two cases needed are p=2p=2 and p=1p=1:

(F)ρσ  =  12ερσμνFμν,(J)λρσ  =  ελρσμjμ. \big(\star F\big)_{\rho\sigma} \;=\; \half\,\varepsilon_{\rho\sigma\mu\nu}\,F^{\mu\nu}, \qquad \big(\star J\big)_{\lambda\rho\sigma} \;=\; \varepsilon_{\lambda\rho\sigma\mu}\,j^{\mu}. (3.5.81)

10.2 · The three statements

First. Section 3.3 already established F=dAF=\dd A, which reproduced Chapter 2.6's definition (2.6.9) character for character. Six components from four.

Second. dF=ddA=0\dd F=\dd\dd A=0 by (3.5.22), and expanding it gave (2.6.25), which is B=0\nabla\cdot\vv B=0 and Faraday's law. Four components.

Third. The inhomogeneous half. Chapter 2.6's (2.6.19) is μFμν=μ0jν\partial_{\mu}F^{\mu\nu}=\mu_{0}j^{\nu}. The claim is that this is

  d ⁣ ⁣F  =  μ0 ⁣J.   \boxed{\;\dd\!\star\!F \;=\; \mu_{0}\,\star\!J.\;} (3.5.82)

Start with the shape of the claim. Both sides are 33-forms, with (43)=4\binom{4}{3}=4 components, which matches the four equations we are trying to reproduce. Verifying it is then a matter of expanding both sides in flat coordinates and comparing. That computation is short but sign-infested, so it sits in the grind box. The result is that (3.5.82) holds exactly with the conventions fixed in (3.5.46) and (3.5.80), namely contracted indices last on ε\varepsilon, and ε0123=+g\varepsilon_{0123}=+\sqrt{-g}.

Grind box — checking d ⁣ ⁣F=μ0 ⁣J\dd\!\star\!F=\mu_{0}\star\!J, and why the sign is a convention worth fixing loudly

Work in flat coordinates where g=1\sqrt{-g}=1 and εμνρσ\varepsilon_{\mu\nu\rho\sigma} is the plain permutation symbol with ε0123=+1\varepsilon_{0123}=+1. Take one component of each side, say (λρσ)=(0,1,2)(\lambda\rho\sigma)=(0,1,2).

Left side. By (3.5.14) at p=2p=2, the component is the three-term cyclic sum 0(F)12+1(F)20+2(F)01\partial_{0}(\star F)_{12}+\partial_{1}(\star F)_{20}+\partial_{2}(\star F)_{01}. Each entry of F\star F is, by (3.5.81), one term of FF with both indices up and the complementary index pair: (F)12=12ε12μνFμν=ε1203F03=F03(\star F)_{12}=\half\varepsilon_{12\mu\nu}F^{\mu\nu}=\varepsilon_{1203}F^{03}=F^{03}, using ε1203=+1\varepsilon_{1203}=+1 and the fact that the two surviving orderings (0,3)(0,3) and (3,0)(3,0) contribute equally and cancel the 12\half. Likewise (F)20=ε2013F13=F13(\star F)_{20}=\varepsilon_{2013}F^{13}=F^{13} and (F)01=ε0123F23=F23(\star F)_{01}=\varepsilon_{0123}F^{23}=F^{23}. So the left side is 0F03+1F13+2F23\partial_{0}F^{03}+\partial_{1}F^{13}+\partial_{2}F^{23}, which is μFμ3\partial_{\mu}F^{\mu3} with the μ=3\mu=3 term absent. That missing term is 3F33=0\partial_{3}F^{33}=0 by antisymmetry, so the left side is exactly μFμ3\partial_{\mu}F^{\mu3}.

Right side. μ0(J)012=μ0ε012μjμ=μ0ε0123j3=μ0j3\mu_{0}(\star J)_{012}=\mu_{0}\varepsilon_{012\mu}j^{\mu}=\mu_{0}\varepsilon_{0123}j^{3}=\mu_{0}j^{3}.

Equating gives μFμ3=μ0j3\partial_{\mu}F^{\mu3}=\mu_{0}j^{3}, which is the ν=3\nu=3 component of (2.6.19). The other three components come out the same way. The whole system was also verified symbolically, for a general Aμ(x)A_{\mu}(x), with all four triples checked and all four agreeing in sign.

Why the loud convention. Had we instead defined \star on a one-form by putting the contracted index first, writing (J)λρσ=εμλρσjμ(\star J)_{\lambda\rho\sigma}=\varepsilon_{\mu\lambda\rho\sigma}j^{\mu}, every component would flip sign, since moving an index past three others costs (1)3(-1)^{3}, and (3.5.82) would carry a minus. Books differ here exactly as they differ over the sign of the Riemann tensor. The rule adopted above puts contracted indices last, at every degree. That rule is self-consistent and is what makes the plus sign correct. This is worth knowing when you reach for a formula from another book: check its convention first.

10.3 · What the compression reveals

(i) Charge conservation is free. Apply d\dd to (3.5.82). The left side is dd(F)=0\dd\dd(\star F)=0 by (3.5.22). Hence

d ⁣ ⁣J  =  0, \dd\!\star\!J \;=\; 0, (3.5.83)

That 33-form has a single independent component, and expanding it the same way the grind box did gives μjμ=0\partial_{\mu}j^{\mu}=0. That is the continuity equation of Chapter 0.7 §6 and Chapter 2.6 §1. Charge conservation is not an extra postulate. It is d2=0\dd^{2}=0 applied to the field equation. Chapter 0.7's Problem 4 found that Maxwell's displacement current was forced by charge conservation. This is that argument run backwards and compressed into one symbol.

(ii) Gauge freedom is free too. Replacing AA+dχA\to A+\dd\chi for any function χ\chi leaves F=dAF=\dd A unchanged, again by (3.5.22). That is Chapter 2.6 §4.1's gauge transformation, and the reason it exists is the same one line.

Section 4 has something to add here. The potentials giving a particular FF differ by a closed one-form, and on a region without holes a closed one-form is exact and hence of the form dχ\dd\chi. So §4's topology is the fine print on the phrase "the potential is determined up to a gauge transformation".

(iii) The asymmetry between the two halves, named. This is the observation the whole section was for. Look at what each statement needs:

StatementNeeds a metric?Status
F=dAF=\dd Anodefinition
dF=0\dd F=0noidentity — cannot fail
d ⁣ ⁣F=μ0 ⁣J\dd\!\star\!F=\mu_{0}\star\!Jyes, inside \starfield equation — physics

Chapter 2.6 §3.4 noticed that two of Maxwell's equations were bookkeeping and two were physics, and Chapter 2.6's forward-looking warning box predicted that in this language "the visible asymmetry between dF\dd F and d ⁣ ⁣F\dd\!\star\!F" would be the reason. It is: the star is where the geometry lives, and only the sourced half of Maxwell touches it.

(iv) Curved spacetime costs nothing. Because d\dd needs no connection (§2.2) and ε\varepsilon already carries g\sqrt{-g} (§6.2), the three statements above are the equations of electromagnetism on any curved manifold, unchanged. In components the sourced half becomes, using §6.4's trick,

μFμν  =  1gμ ⁣(gFμν)  =  μ0jν. \nabla_{\mu}F^{\mu\nu} \;=\; \frac{1}{\sqrt{-g}}\,\partial_{\mu}\!\Big(\sqrt{-g}\,F^{\mu\nu}\Big) \;=\; \mu_{0}\,j^{\nu}. (3.5.84)

The middle equality is worth its two lines. Expanding μFμν\nabla_{\mu}F^{\mu\nu} gives μFμν+ΓμμλFλν+ΓνμλFμλ\partial_{\mu}F^{\mu\nu}+\Gamma^{\mu}{}_{\mu\lambda}F^{\lambda\nu}+\Gamma^{\nu}{}_{\mu\lambda}F^{\mu\lambda}. The last term vanishes: Γνμλ\Gamma^{\nu}{}_{\mu\lambda} is symmetric in μλ\mu\lambda and FμλF^{\mu\lambda} is antisymmetric, so the contraction is zero. The first two are exactly (3.5.51) applied to each fixed value of ν\nu. So on a curved manifold Maxwell's equations are Maxwell's equations with \partial\to\nabla, and the only visible change is a factor of g\sqrt{-g}. That is the whole of "electromagnetism in a gravitational field", in one line, and we have it before the gravitational field equations have even been written down.

In plain terms 3.5.10

As a test of the new language, classical electromagnetism is rewritten in it and comes to three short statements: the field is the derivative of the potential, the derivative of the field is nothing, and the derivative of the field's complement is the current. The compression is pleasant but not the point. What it exposes is which parts of the subject depend on geometry and which do not.

Only one operation in the three requires knowing distances and angles, namely the one swapping a description in terms of two slots for one in terms of the complementary two. That operation appears in exactly one of the three, the one with a source in it. So the half of Maxwell's equations the previous part identified as empty bookkeeping is also the half never touching the geometry, and the half carrying the physics is the half that does. The earlier chapter noticed the first fact and predicted the second; here they are the same fact.

Two things then come for free. Applying the derivative twice to the sourced statement gives nothing on the left, forcing the current to be conserved, so charge conservation follows from the field equation rather than being an extra law. And adding the derivative of any function to the potential changes no field, which is the freedom called gauge, arriving from the same line. Both survive intact on a curved space.

11 · Worked examples

Worked example 1 — the area of a region, from its boundary alone

Use (3.5.41) to show that the area enclosed by a closed plane curve is 12(xdyydx)\half\oint(x\,\dd y - y\,\dd x), and evaluate it for a circle and for a triangle.

The form. On the plane take the one-form ω=12(xdyydx)\omega=\half\big(x\,\dd y - y\,\dd x\big). Its exterior derivative, by (3.5.17) with ωx=y/2\omega_{x}=-y/2 and ωy=x/2\omega_{y}=x/2, is

(dω)xy  =  x ⁣(x2)y ⁣(y2)  =  12+12  =  1, \big(\dd\omega\big)_{xy} \;=\; \partial_{x}\!\left(\frac x2\right) - \partial_{y}\!\left(\frac{-y}{2}\right) \;=\; \half+\half \;=\; 1,

so dω=dxdy\dd\omega=\dd x\wedge\dd y, the area element. Feeding that into (3.5.41) with MM the enclosed region:

Area(M)  =  Mdxdy  =  Mdω  =  Mω  =  12M(xdyydx). \text{Area}(M) \;=\; \int_{M}\dd x\wedge\dd y \;=\; \int_{M}\dd\omega \;=\; \oint_{\partial M}\omega \;=\; \half\oint_{\partial M}\big(x\,\dd y - y\,\dd x\big).

A circle of radius RR. Parametrise the boundary by (Rcost,Rsint)(R\cos t,R\sin t). Then xdy=RcostRcostdtx\,\dd y=R\cos t\cdot R\cos t\,\dd t and ydx=Rsint(Rsint)dt-y\,\dd x=-R\sin t\cdot(-R\sin t)\,\dd t, so the integrand is 12R2(cos2t+sin2t)dt=12R2dt\half R^{2}(\cos^{2}t+\sin^{2}t)\,\dd t=\half R^{2}\dd t and the integral over t[0,2π]t\in[0,2\pi] is πR2\pi R^{2}. ✓

A triangle with vertices a,b,c\vv a,\vv b,\vv c. On the straight edge from a\vv a to b\vv b, parametrise x=a+t(ba)\vv x=\vv a+t(\vv b-\vv a). Then the integrand 12(xdyydx)\half(x\,\dd y-y\,\dd x) works out to the constant 12(axbyaybx)dt\half(a_{x}b_{y}-a_{y}b_{x})\,\dd t, since the terms linear in tt cancel between the two products. Summing the three edges gives 12[(axbyaybx)+(bxcybycx)+(cxaycyax)]\half\big[(a_{x}b_{y}-a_{y}b_{x})+(b_{x}c_{y}-b_{y}c_{x})+(c_{x}a_{y}-c_{y}a_{x})\big], which is the surveyor's formula for the area of a triangle. ✓

The remark worth keeping. Compare ω\omega here with the angle form (3.5.29) of §4.1. They have the same numerator. The only difference is that §4.1 divides by x2+y2x^{2}+y^{2}, which is what makes it closed and what makes it singular at the origin. Here dω=10\dd\omega=1\neq0, so ω\omega is not closed, there is no contradiction with §4, and the integral measures area rather than winding. A planimeter is this formula made of brass: it is the mechanical instrument that measures the area of a shape by being traced around its outline.

Worked example 2 — the flat plane in polar coordinates: three Killing vectors, three conserved quantities

Find all the Killing vectors of ds2=dr2+r2dϕ2\dd s^{2}=\dd r^{2}+r^{2}\dd\phi^{2}, and compute the conserved quantities they supply along a geodesic. This is a case where the answer is known in advance, which is the point of doing it.

The easy one. No metric component mentions ϕ\phi, so by §8.1's practical test ξ(1)=ϕ\xi_{(1)}=\partial_{\phi} is Killing.

The two the chart hides. The plane is flat, so by §8.2 it has translations too. But r\partial_{r} is not a translation, and indeed (Lrg)ϕϕ=r(r2)=2r0\big(\mathcal{L}_{\partial_{r}}g\big)_{\phi\phi}=\partial_{r}(r^{2})=2r\neq0, so it is not Killing. Convert the Cartesian translations instead. With x=rcosϕx=r\cos\phi and y=rsinϕy=r\sin\phi, the chain rule gives

ξ(2)=x  =  cosϕ  r    sinϕr  ϕ,ξ(3)=y  =  sinϕ  r  +  cosϕr  ϕ. \xi_{(2)}=\partial_{x} \;=\; \cos\phi\;\partial_{r} \;-\; \frac{\sin\phi}{r}\;\partial_{\phi}, \qquad \xi_{(3)}=\partial_{y} \;=\; \sin\phi\;\partial_{r} \;+\; \frac{\cos\phi}{r}\;\partial_{\phi}.

Substituting either into (3.5.65) gives zero in all three components, and this was verified symbolically. So there are three Killing vectors, which is n(n+1)/2n(n+1)/2 at n=2n=2: two translations and one rotation, exactly the symmetries of the Euclidean plane.

The conserved quantities. A geodesic of the plane is a straight line; parametrise it by arc length ss, so u=(r˙,ϕ˙)u=(\dot r,\dot\phi). Lower the index on each Killing vector with g=diag(1,r2)g=\mathrm{diag}(1,r^{2}) and contract:

Q(1)  =  gϕϕξ(1)ϕϕ˙  =  r2ϕ˙, Q_{(1)} \;=\; g_{\phi\phi}\,\xi^{\phi}_{(1)}\,\dot\phi \;=\; r^{2}\dot\phi, Q(2)  =  cosϕ  r˙    sinϕrr2ϕ˙  =  cosϕr˙rsinϕϕ˙  =  dds(rcosϕ)  =  x˙. Q_{(2)} \;=\; \cos\phi\;\dot r \;-\; \frac{\sin\phi}{r}\cdot r^{2}\dot\phi \;=\; \cos\phi\,\dot r - r\sin\phi\,\dot\phi \;=\; \dv{}{s}\big(r\cos\phi\big) \;=\; \dot x.

The last step is the product rule read backwards. Similarly Q(3)=y˙Q_{(3)}=\dot y. So the three charges are x˙\dot x, y˙\dot y and r2ϕ˙r^{2}\dot\phi. Those are the two components of the velocity and the angular momentum per unit mass about the origin, all constant along a straight line, which they visibly are. The machinery returned the answer everyone already knew, which is what one wants from a first test of it. And note that Q(1)=r2ϕ˙Q_{(1)}=r^{2}\dot\phi is Kepler's second law, which Chapter 1.4 §3.3 obtained from rotational symmetry and which arrives here from the geometry of the plane.

12 · Your turn

Problem 1 — why vector calculus is a three-dimensional accident

(a) Show that for two one-forms in three dimensions, the components of αβ\alpha\wedge\beta are the components of the cross product α×β\vv\alpha\times\vv\beta, up to the identification ωjk=ϵijkFi\omega_{jk}=\epsilon^{ijk}F^{i} used in §2.1. (b) Using §1.3's count, explain why that identification is available in three dimensions and in no other. (c) In four dimensions, how many independent components does the wedge of two one-forms have, and why is there no "cross product" of two four-vectors? (d) Chapter 0.7 §4.4 said the curl is a vector only in three dimensions and promised the reason later. State it.

Solution

(a) By (3.5.10), (αβ)jk=αjβkαkβj(\alpha\wedge\beta)_{jk}=\alpha_{j}\beta_{k}-\alpha_{k}\beta_{j}. Its (1,2)(1,2) component is α1β2α2β1\alpha_{1}\beta_{2}-\alpha_{2}\beta_{1}, which is (α×β)3(\vv\alpha\times\vv\beta)^{3}; the other two independent components are the other two components of the cross product, in cyclic order.

(b) The identification pairs a 22-form with a vector, so it needs (n2)=(n1)\binom{n}{2}=\binom{n}{1}, that is n(n1)/2=nn(n-1)/2=n, whose only solution with n>1n>1 is n=3n=3. Three dimensions is the unique case where a pair of directions and a single direction are the same amount of information.

(c) Six, by §1.3's table. A six-component object cannot be a four-vector, so there is no cross product. The six-component antisymmetric object in four dimensions is exactly the type of the electromagnetic field tensor, which is why E\vv E and B\vv B mix under boosts (Chapter 2.6 §5): they are two halves of one 22-form, not two vectors.

(d) The curl of a vector field is really the exterior derivative of a one-form, which is a 22-form. Calling the result a vector uses (b)'s coincidence, and outside three dimensions there is nothing to call it. Chapter 0.7 §4.3's splitting of the Jacobian into a symmetric and an antisymmetric part is the same statement: the antisymmetric part has n(n1)/2n(n-1)/2 components and only accidentally has nn.

Problem 2 — a hole in space, measured by a 22-form

On R3\R^{3} with the origin removed, define the 22-form corresponding under §2.1's identification to the radial field B=r^/r2\vv B=\hat{\vv r}/r^{2}. (a) Show dω=0\dd\omega=0. (b) Show Sω=4π\oint_{S}\omega=4\pi over any sphere centred on the origin. (c) Conclude that ω\omega is closed but not exact, and say what fails in §4.2's proof. (d) Chapter 0.7 §7.3 said a divergence-free field is a curl "on regions without cavities". Restate that caveat in the language of this chapter, and say what physical object this counterexample would be.

Solution

(a) Under the identification, dω\dd\omega is B\nabla\cdot\vv B times the volume form ((3.5.18)). For B=r/r3\vv B=\vv r/r^{3}, i(xi/r3)=3/r33xixi/r5=3/r33/r3=0\partial_{i}(x^{i}/r^{3})=3/r^{3}-3x^{i}x^{i}/r^{5}=3/r^{3}-3/r^{3}=0 away from the origin.

(b) On a sphere of radius RR the field is r^/R2\hat{\vv r}/R^{2}, normal to the surface and of constant magnitude, so the flux is (1/R2)×4πR2=4π(1/R^{2})\times 4\pi R^{2}=4\pi, independent of RR.

(c) If ω=dα\omega=\dd\alpha then by (3.5.41) Sω=Sdα=Sα=0\oint_{S}\omega=\oint_{S}\dd\alpha=\oint_{\partial S}\alpha=0, since a sphere has no boundary. That contradicts (b), so ω\omega is not exact. What fails in §4.2 is the hypothesis: the region is not star-shaped, because every ray from a would-be centre to the far side passes through the missing origin, so the integral defining KωK\omega is not available.

(d) "Divergence-free implies a curl" is "closed 22-forms are exact", which holds when the region has no enclosed cavity. That is precisely the condition making the 22-form's integral over every closed surface vanish. The counterexample is a magnetic monopole: a field with B=0\nabla\cdot\vv B=0 everywhere it is defined and yet no globally defined vector potential. That is why monopoles, if they exist, force a rethink of the potential rather than of Maxwell's equations, and it is the seed of a topological argument Part VI returns to.

Problem 3 — a symmetry that stretches

Take the dilation field X=xμμX=x^{\mu}\partial_{\mu} on flat spacetime. (a) Compute LXημν\mathcal{L}_{X}\eta_{\mu\nu} and show it is 2ημν2\eta_{\mu\nu}, so XX is not Killing. (b) Show nonetheless that the flow of XX preserves all angles, in the sense that it rescales every length by the same factor at a given point. Such a field is called a conformal Killing vector. (c) Does ξμuμ\xi_{\mu}u^{\mu} remain constant along a geodesic for such an XX? Redo §9.1 and identify exactly where it fails and what it fails by. (d) For which kind of particle would the failure vanish?

Solution

(a) Xμ=xμX^{\mu}=x^{\mu}, so νXμ=δμν\partial_{\nu}X^{\mu}=\delta^{\mu}{}_{\nu}. In (3.5.62) the first term vanishes because η\eta is constant, and the other two give ημν+ημν=2ημν\eta_{\mu\nu}+\eta_{\mu\nu}=2\eta_{\mu\nu}.

(b) LXg=2g\mathcal{L}_{X}g=2g says the dragged metric is a multiple of the original, and multiplying a metric by a positive number multiplies every squared length by it while leaving every ratio, and hence every angle, alone.

(c) Lines 1 to 3 of §9.1 are unchanged, so dQ/dτ=12uμuν(μξν+νξμ)=12uμuν(LXg)μν\dd Q/\dd\tau=\half u^{\mu}u^{\nu}(\nabla_{\mu}\xi_{\nu}+\nabla_{\nu}\xi_{\mu})=\half u^{\mu}u^{\nu}(\mathcal{L}_{X}g)_{\mu\nu}. For a conformal Killing vector that bracket is 2Ωgμν2\Omega g_{\mu\nu} for some function Ω\Omega, here the constant 11, so dQ/dτ=Ωgμνuμuν\dd Q/\dd\tau=\Omega\,g_{\mu\nu}u^{\mu}u^{\nu}. The failure is proportional to the norm of the tangent.

(d) A massless one. For a null geodesic gμνuμuν=0g_{\mu\nu}u^{\mu}u^{\nu}=0, so the failure vanishes and QQ is conserved after all. Conformal symmetries give conserved quantities for light but not for matter, which is why conformal invariance is a symmetry of the massless theory only. Chapter 7.3 builds an entire chapter on that fact.

Problem 4 — Clairaut's relation, and a rehearsal for Chapter 3.7

On the sphere of radius aa, take ξ=ϕ\xi=\partial_{\phi}. (a) Write down the conserved quantity (3.5.73) supplies for a geodesic parametrised by arc length. (b) Let ψ\psi be the angle between the geodesic and the local circle of latitude. Show that the conserved quantity is asinθcosψa\sin\theta\cos\psi, so that sinθcosψ\sin\theta\cos\psi is constant along any great circle. This is Clairaut's relation. (c) Deduce the highest latitude a given great circle reaches, and check against Chapter 3.3 §9. (d) Say which two features of the Schwarzschild metric will play the roles of θ\theta-independence and ϕ\phi-independence in Chapter 3.7, and what the two charges will be called.

Solution

(a) ξμ=(0,1)\xi^{\mu}=(0,1), so ξμ=gμϕ\xi_{\mu}=g_{\mu\phi} and Q=ξμuμ=gϕϕϕ˙=a2sin2θ  ϕ˙Q=\xi_{\mu}u^{\mu}=g_{\phi\phi}\dot\phi=a^{2}\sin^{2}\theta\;\dot\phi.

(b) With arc length as parameter, a2θ˙2+a2sin2θϕ˙2=1a^{2}\dot\theta^{2}+a^{2}\sin^{2}\theta\,\dot\phi^{2}=1. The unit tangent has an eastward component asinθϕ˙a\sin\theta\,\dot\phi and a southward component aθ˙a\dot\theta, so by definition of ψ\psi the eastward component is cosψ\cos\psi. Hence Q=asinθ(asinθϕ˙)=asinθcosψQ=a\sin\theta\cdot\big(a\sin\theta\,\dot\phi\big)=a\sin\theta\cos\psi.

(c) The geodesic is highest, meaning θ\theta is smallest, where it runs due east at ψ=0\psi=0, since sinθ\sin\theta must then be as small as possible for fixed QQ. So sinθmin=Q/a=sinθ0cosψ0\sin\theta_{\min}=Q/a=\sin\theta_{0}\cos\psi_{0} for any starting point. Chapter 3.3 §9 showed the geodesics are great circles, and a great circle inclined at angle ι\iota to the equator reaches colatitude θmin=π/2ι\theta_{\min}=\pi/2-\iota; both statements say the same thing, since a great circle crosses the equator at exactly the angle ι\iota. ✓

(d) The Schwarzschild metric will depend on neither tt nor ϕ\phi. By §8.1's test those give Killing vectors t\partial_{t} and ϕ\partial_{\phi}, and by (3.5.73) two conserved charges, which Chapter 3.7 calls the energy per unit mass EE and the angular momentum per unit mass LL. Part (a) here is the calculation of LL, done on a sphere to keep the algebra visible.

The brick you just laid

You have a language for integration on a curved space. Demanding that an integral belong to a region and not to a labelling forced its integrand to be totally antisymmetric. That was the determinant of Chapter 0.4, doing physics. The objects obeying the demand are differential forms, and antisymmetrising the ordinary derivative gives them a derivative d\dd that is a tensor with no connection anywhere, because the connection's chart-dependent piece is symmetric in exactly the two slots being antisymmetrised.

Three debts, collected. Applying d\dd twice gives zero, in three lines, and that single statement is Chapter 0.7's ×=0\nabla\times\nabla=0, Chapter 0.7's ×=0\nabla\cdot\nabla\times=0 and Chapter 2.6's discovery that half of Maxwell's equations are bookkeeping. It is also the algebraic shadow of the fact that a boundary has no boundary. The converse, that closed implies exact, holds on a region you can shrink to a point, and it is proved in four lines here. That converse is Chapter 0.7 §7.3's quoted Poincaré lemma, and it hands you a formula for the vector potential. And Mdω=Mω\int_{M}\dd\omega=\oint_{\partial M}\omega is Chapter 0.7's four theorems, one line, proved on a cube by the Fundamental Theorem of Calculus and extended by the cancellation argument that chapter described and could not complete.

Two tools for what comes next. The invariant volume element gdnx\sqrt{-g}\,\dd^{n}x makes "integrate over spacetime" mean something. Jacobi's formula for the derivative of a determinant gives μVμ=μ(gVμ)/g\nabla_{\mu}V^{\mu}=\partial_{\mu}(\sqrt{-g}V^{\mu})/\sqrt{-g}, and that is the sentence that lets Chapter 3.6 throw a term away.

And the reason Chapter 3.7 is possible. The Lie derivative compares a tensor with a copy of itself dragged along a flow, and it needs no connection. Setting Lξg=0\mathcal{L}_{\xi}g=0 defines a Killing vector and yields (3.5.66). Flat spacetime has ten of them and they are the Poincaré generators of Part II. Every Killing vector makes ξμuμ\xi_{\mu}u^{\mu} constant along every geodesic, in four lines. Running Chapter 1.4's theorem on the same system returns the same charge, with δL\delta L turning out to be Lξg\mathcal{L}_{\xi}g contracted with two velocities. Noether's theorem and Killing's equation are one thing.

Where this gets spent. Chapter 3.6 uses §5's Stokes theorem, §6's volume element and determinant formula, and §6.4's divergence identity, to vary the Einstein–Hilbert action and derive the field equations a second time. Chapter 3.7 uses §9 twice, and would be unreadable without it. Chapter 3.9 uses the absence of a timelike Killing vector to explain why energy is not conserved in an expanding universe, collecting Chapter 1.4 §4.3's honest note. Chapter 6.3 rebuilds §10's F=dAF=\dd A with an internal space in place of spacetime, at which point the gauge freedom that fell out of d2=0\dd^{2}=0 here becomes the organising principle of every force in nature.