Part III · General Relativity — Chapter 3.4

Curvature

A test for curvature that can be run from inside, with no outside to look in from. It measures the tide.

Where we are

Chapter 3.3 built two objects and issued one warning. The metric measures. The connection compares. The warning was that neither position-dependent metric components nor non-zero connection coefficients count as evidence of curvature. The flat plane in polar coordinates has both, and it is the flat plane. So no test for curvature yet exists, and this chapter builds one.

The test has to satisfy a hard constraint that Chapter 3.2 §1.2 imposed and never relaxed. It must be conducted entirely from within the space. Nothing may refer to an ambient room, a normal vector sticking out, or a picture drawn from outside. That rules out every intuitive definition of curvature anyone arrives with, and it leaves exactly one available instrument, which Chapter 3.3 §6 has just finished building. Carry a vector around a closed loop and see whether it comes back the way it left. Every step of that procedure stays inside the space. Transport is intrinsic, a closed loop is intrinsic, and the comparison at the end is between two vectors at the same point, which is what makes it legitimate. If the answer is ever "no", the space is curved, and how badly is a number.

Here is the route, announced in advance. Section 1 sets up the loop test and runs it on a sphere. The answer is derived analytically from Chapter 3.3's Christoffel symbols and measured numerically in the figure. The two agree, and what they agree on is the enclosed area divided by the square of the radius. Section 2 is the longest computation in this book so far: the commutator of two covariant derivatives, which is the loop test made infinitesimal. Section 3 collects the payoff. The pieces of the connection that are not tensorial cancel, and what is left is a genuine tensor. Section 4 is the centre of Part III. Two nearby geodesics drift apart at a rate set by that tensor, and setting the result beside Chapter 3.1's Newtonian tidal equation shows they are the same equation. Sections 5 and 6 work out the symmetries, derive the count of twenty independent components in four dimensions, and take the traces. Section 7 derives an identity whose contracted form produces a tensor with zero divergence, and that single fact dictates the shape of the field equations in Chapter 3.6. Section 8 closes the loop opened in Chapter 3.3 §2.

Sign convention, stated loudly and once. This book defines the Riemann tensor by

[μ,ν]Vρ  =  RρσμνVσ,Rμν  =  Rλμλν, \big[\nabla_{\mu},\nabla_{\nu}\big]V^{\rho} \;=\; R^{\rho}{}_{\sigma\mu\nu}\,V^{\sigma}, \qquad R_{\mu\nu} \;=\; R^{\lambda}{}_{\mu\lambda\nu},

with signature (+,,,)(+,-,-,-) throughout. Books differ, in both places independently. Some define Riemann with the opposite overall sign. Some contract the first and third indices in the other order. With the conventions above, a sphere has positive Ricci scalar and matter has positive energy density on the right-hand side of Einstein's equations, which is the combination we want. Before importing a formula from anywhere else, check both signs. GG and cc remain explicit.

Tools you'll need.  Chapter 3.3 throughout: the covariant derivative of its §5, its extension to tensors (3.3.31) and (3.3.32), the inhomogeneous transformation law (3.3.26) of the connection, parallel transport (3.3.35), the Christoffel formula (3.3.50), the geodesic equation (3.3.57) and the sphere's connection coefficients from §9. Chapter 3.1 §4 in detail, meaning the tidal equation, the point-mass tidal matrix, its traceless character in vacuum, and the ring that deforms into an ellipse. Section 4 below is where all of that is collected. Chapter 3.2 §8, the commutator of vector fields, whose second-derivative cancellation is the template for §2, and §7 for the tensor transformation law. Chapter 2.4 §5.1 (contraction), §6 (a tensor equation true in one chart is true in all) and §7.2 (counting independent components of a symmetric or antisymmetric array). Chapter 0.6 §6.1, Clairaut's theorem, used four separate times below. Chapter 0.5 §6, the spectral theorem, used once in §5.

1 · The loop test, and the sphere

Here is the test in full. Take a closed path, plant a vector at a point on it, carry the vector all the way round by the parallel-transport equation (3.3.35), and compare the arrival with the departure. Both vectors live in the same tangent space, so the comparison is legitimate in the strict sense of Chapter 3.2 §5.1. Transport preserves length by metric compatibility, so the only thing that can have happened along the way is a rotation. The angle of that rotation is the measurement.

Check first that the test gives nothing where it should. In flat space with Cartesian coordinates the connection vanishes, so (3.3.35) reads dVμ/dλ=0\dd V^{\mu}/\dd\lambda=0 and the vector comes back identical. The angle is zero for every loop. Any non-zero answer is therefore a fact about the space and not about the chart, provided we also check that flat space still gives zero when the chart is curvilinear. Section 8 runs that check.

1.1 · A loop on a sphere, computed by hand

Here is where we are going, announced before we set off. We choose a closed path on a sphere of radius aa, integrate the transport equation along it in three pieces, and find that the vector returns rotated through an angle exactly equal to the area enclosed divided by a2a^{2}. Everything we need comes from the Christoffel symbols computed in Chapter 3.3 §9, which were

Γθϕϕ  =  sinθcosθ,Γϕθϕ  =  Γϕϕθ  =  cotθ,all others zero, \Gamma^{\theta}{}_{\phi\phi} \;=\; -\sin\theta\cos\theta, \qquad \Gamma^{\phi}{}_{\theta\phi} \;=\; \Gamma^{\phi}{}_{\phi\theta} \;=\; \cot\theta, \qquad\text{all others zero,} (3.4.1)

with θ\theta the angle from the north pole and ϕ\phi the azimuth.

A convenient pair of numbers. Before integrating anything we should choose components that are easy to read. Components carry the length of their basis vector, and here the two basis vectors have different lengths: θ=a\abs{\partial_{\theta}}=a and ϕ=asinθ\abs{\partial_{\phi}}=a\sin\theta. So define the orthonormal components

P    aVθ,U    asinθ  Vϕ,so thatV2  =  P2+U2. P \;\equiv\; a\,V^{\theta}, \qquad U \;\equiv\; a\sin\theta\;V^{\phi}, \qquad\text{so that}\qquad \abs{V}^{2} \;=\; P^{2}+U^{2}. (3.4.2)

That last equality is just gμνVμVν=a2(Vθ)2+a2sin2θ(Vϕ)2g_{\mu\nu}V^{\mu}V^{\nu}=a^{2}(V^{\theta})^{2}+a^{2}\sin^{2}\theta\,(V^{\phi})^{2} regrouped. In this pair, PP is how much of the vector points south and UU how much points east, each measured in ordinary length.

Leg one, along a circle of latitude θ=θ0\theta=\theta_{0}, from ϕ=0\phi=0 to ϕ=Δϕ\phi=\Delta\phi. The tangent to this leg has only a ϕ\phi component, so in (3.3.35) the sum over ν\nu keeps only the term ν=ϕ\nu=\phi:

dVθdϕ  =  ΓθϕϕVϕ  =  sinθ0cosθ0  Vϕ,dVϕdϕ  =  ΓϕϕθVθ  =  cotθ0  Vθ. \dv{V^{\theta}}{\phi} \;=\; -\,\Gamma^{\theta}{}_{\phi\phi}V^{\phi} \;=\; \sin\theta_{0}\cos\theta_{0}\;V^{\phi}, \qquad \dv{V^{\phi}}{\phi} \;=\; -\,\Gamma^{\phi}{}_{\phi\theta}V^{\theta} \;=\; -\cot\theta_{0}\;V^{\theta}. (3.4.3)

Those are equations for VθV^{\theta} and VϕV^{\phi}, and what we want are equations for PP and UU, so rewrite them. Since θ0\theta_{0} is constant along this leg, the factors aa and asinθ0a\sin\theta_{0} are constants and pass straight through the derivative:

dPdϕ  =  asinθ0cosθ0Vϕ  =  cosθ0U,dUdϕ  =  asinθ0cotθ0Vθ  =  cosθ0P, \begin{aligned} \dv{P}{\phi} \;&=\; a\sin\theta_{0}\cos\theta_{0}\,V^{\phi} \;=\; \cos\theta_{0}\,U,\\[4pt] \dv{U}{\phi} \;&=\; -\,a\sin\theta_{0}\cot\theta_{0}\,V^{\theta} \;=\; -\cos\theta_{0}\,P, \end{aligned} (3.4.4)

In each line the constant was regrouped to build the other orthonormal component. In the first, asinθ0Vϕ=Ua\sin\theta_{0}\cdot V^{\phi}=U. In the second, acosθ0Vθa\cos\theta_{0}\cdot V^{\theta} splits as cosθ0(aVθ)=cosθ0P\cos\theta_{0}\cdot(aV^{\theta})=\cos\theta_{0}P, using sinθ0cotθ0=cosθ0\sin\theta_{0}\cot\theta_{0}=\cos\theta_{0}.

Let's take two readings of (3.4.4) before integrating it. The first reading is that P2+U2P^{2}+U^{2} is constant, since d(P2+U2)/dϕ=2Pcosθ0U2Ucosθ0P=0\dd(P^{2}+U^{2})/\dd\phi = 2P\cos\theta_{0}U - 2U\cos\theta_{0}P = 0. That is metric compatibility showing up as a conserved quantity, exactly as Chapter 3.3 §7.1 said it would.

The second reading is that the pair (P,U)(P,U) obeys precisely the equations of a rotation. To see it, write a rotation through angle α\alpha as PPcosαUsinαP\to P\cos\alpha-U\sin\alpha, UPsinα+UcosαU\to P\sin\alpha+U\cos\alpha and differentiate, which gives dP/dα=U\dd P/\dd\alpha=-U and dU/dα=P\dd U/\dd\alpha=P. Matching that against (3.4.4) gives dα/dϕ=cosθ0\dd\alpha/\dd\phi=-\cos\theta_{0}, so over the whole leg the vector turns through

αleg 1  =  cosθ0  Δϕ, \alpha_{\text{leg 1}} \;=\; -\,\cos\theta_{0}\;\Delta\phi, (3.4.5)

a rotation measured positively from the southward direction toward the eastward one.

Legs two and three, along meridians. Now the tangent has only a θ\theta component, so (3.3.35) keeps only the term ν=θ\nu=\theta:

dVθdθ  =  ΓθθρVρ  =  0,dVϕdθ  =  ΓϕθϕVϕ  =  cotθ  Vϕ, \dv{V^{\theta}}{\theta} \;=\; -\,\Gamma^{\theta}{}_{\theta\rho}V^{\rho} \;=\; 0, \qquad \dv{V^{\phi}}{\theta} \;=\; -\,\Gamma^{\phi}{}_{\theta\phi}V^{\phi} \;=\; -\cot\theta\;V^{\phi}, (3.4.6)

The first of those vanishes because both Γθθθ\Gamma^{\theta}{}_{\theta\theta} and Γθθϕ\Gamma^{\theta}{}_{\theta\phi} are zero in (3.4.1). So P=aVθP=aV^{\theta} is constant outright. The second equation separates into dVϕ/Vϕ=cotθdθ\dd V^{\phi}/V^{\phi}=-\cot\theta\,\dd\theta, and integrating both sides gives lnVϕ=lnsinθ+const\ln V^{\phi}=-\ln\sin\theta+\text{const}, that is Vϕ1/sinθV^{\phi}\propto1/\sin\theta. Hence U=asinθVϕU=a\sin\theta\,V^{\phi} is constant too. Along a meridian, nothing rotates, and that is what we should expect. A meridian is a great circle and therefore a geodesic, and Chapter 3.3 §8 built geodesics to be the curves along which the angle to the direction of travel is held fixed.

The one subtlety, which is the frames at the pole. If the loop passes through the north pole we must be careful, because (e^θ,e^ϕ)(\hat e_{\theta},\hat e_{\phi}) is not a single frame there. Watch what happens to it. As θ0\theta\to0 along the meridian of azimuth ϕ\phi, the southward direction e^θ\hat e_{\theta} becomes the horizontal direction pointing away from the pole at azimuth ϕ\phi. In the ambient description, used here only to name directions, e^θ(cosϕ,sinϕ,0)\hat e_{\theta}\to(\cos\phi,\sin\phi,0) and e^ϕ(sinϕ,cosϕ,0)\hat e_{\phi}\to(-\sin\phi,\cos\phi,0). So the frame belonging to azimuth ϕ\phi is the frame belonging to azimuth 00 turned through +ϕ+\phi. The consequence is the piece we need. A single fixed vector at the pole, described first in the frame of azimuth Δϕ\Delta\phi and then in the frame of azimuth 00, has its component pair rotated by +Δϕ+\Delta\phi between the two descriptions.

Assemble the loop. Start at A=(θ0,ϕ=0)A=(\theta_{0},\,\phi=0). Go east along the latitude circle to (θ0,Δϕ)(\theta_{0},\,\Delta\phi), then north up that meridian to the pole, then south down the meridian of azimuth 00 back to AA. Now add the rotations in the order they were picked up:

α  =  cosθ0Δϕleg 1  +  0leg 2  +  +Δϕframe change at the pole  +  0leg 3  =  Δϕ(1cosθ0). \alpha \;=\; \underbrace{-\cos\theta_{0}\,\Delta\phi}_{\text{leg 1}} \;+\; \underbrace{0}_{\text{leg 2}} \;+\; \underbrace{+\,\Delta\phi}_{\text{frame change at the pole}} \;+\; \underbrace{0}_{\text{leg 3}} \;=\; \Delta\phi\,\big(1-\cos\theta_{0}\big). (3.4.7)

The area enclosed. To make anything of that angle we need the area of the region the loop bounds, so compute it next. The region is the wedge of the polar cap between azimuths 00 and Δϕ\Delta\phi, running from the pole down to colatitude θ0\theta_{0}. The area element on the sphere is gdθdϕ=a2sinθdθdϕ\sqrt{g}\,\dd\theta\,\dd\phi = a^{2}\sin\theta\,\dd\theta\,\dd\phi, since the metric is diagonal with gθθgϕϕ=a4sin2θg_{\theta\theta}g_{\phi\phi}=a^{4}\sin^{2}\theta. Integrating over the wedge,

A  =  0Δϕ ⁣ ⁣0θ0a2sinθ  dθdϕ  =  a2Δϕ(1cosθ0). \mathcal A \;=\; \int_{0}^{\Delta\phi}\!\!\int_{0}^{\theta_{0}} a^{2}\sin\theta\;\dd\theta\,\dd\phi \;=\; a^{2}\,\Delta\phi\,\big(1-\cos\theta_{0}\big). (3.4.8)

Both numbers are now in hand, and each of them is the same product Δϕ(1cosθ0)\Delta\phi\,(1-\cos\theta_{0}), once with a factor a2a^{2} in front and once without. So the ratio of (3.4.7) to (3.4.8) carries no trace of the loop at all, and rearranging that ratio gives the result in its usual form:

  α  =  Aa2.   \boxed{\;\alpha \;=\; \frac{\mathcal A}{a^{2}}.\;} (3.4.9)
Recap — what went in, what came out

In: the sphere's Christoffel symbols, the parallel-transport equation, one separable first-order equation, and care about which frame is being used at the pole.

Out: a vector carried round that loop returns rotated, in the same sense as the loop was circulated, through an angle equal to the enclosed area divided by a2a^{2}.

What is remarkable about it. The rotation depends on the loop only through the area it encloses. Not its shape, not its position on the sphere, not how the transport was parametrised. So the quantity that belongs to the space rather than to the loop is the ratio α/A=1/a2\alpha/\mathcal A = 1/a^{2}, which survives as the loop shrinks to nothing. That number is what we are going to call curvature, and §9 recovers exactly 1/a21/a^{2} from the tensor built in §2.

60.00°
90.0°
measured rotation of the returned vector: 44.999998°
enclosed area / a² (theory, §1.1): 45.000000° [7.8540e-1 sr]
difference: -1.73e-6° ratio angle/area = 1.000000 = curvature of the unit sphere
checks: |V| drift 3.3e-14 tangency at the base point 1.5e-15
Curvature, measured. The green loop runs east along a circle of latitude, north to the pole, and south again — the path of §1.1 — and the region it encloses is shaded. A vector is planted at the base point (black) and carried round; the blue arrows are the transported vector at stations along the way, and the purple arrow is what comes back. Every arrow is produced by chopping the path into six thousand short great-circle steps and applying the rotation that carries each step's start to its end, which is the transport equation integrated numerically; no formula from the text enters the computation. The readouts give the measured angle between departure and arrival, the enclosed area from (3.4.8), and their ratio. The angle and the area agree to five or six decimal places at every setting, and the ratio sits at 1.0000001.000000 — the curvature of the unit sphere — whatever the loop. Press halve the loop repeatedly: the area falls by about a factor of eight each time and the angle falls with it, while the ratio does not move. That invariance under shrinking is the whole point, because it says the quantity being measured belongs to the point rather than to the loop, and it is what makes an infinitesimal version of this test possible at all. Section 2 builds that infinitesimal version; §9 computes the sphere's curvature tensor from the connection and gets 1/a21/a^{2} back.
In plain terms 3.4.1

Every intuitive account of curvature smuggles in a viewpoint from outside — a rubber surface seen to sag, a ball seen to bulge, a direction sticking into a surrounding room. None is available here, since there is no room and no outside, and the difficulty of this chapter is finding a test a creature confined to the surface could run. One instrument survives the restriction, and the previous chapter finished building it: carry a direction around a closed circuit and see whether it comes home pointing the way it left.

Everything about that procedure stays inside. The circuit lies on the surface, the carrying rule refers only to the surface, and the final comparison is between two directions at one and the same place, which was always the one permitted comparison. Since carrying preserves lengths, the only thing that can have happened is a turn, whose size anyone on the surface can measure.

Running the test on a globe with pencil and paper gives an answer of startling neatness. The turn depends on the circuit only through the area it surrounds, not on its shape or position. So the quantity belonging to the surface rather than to any circuit is the turn divided by the area, and that ratio survives as the circuit shrinks to a point. It is the reciprocal of the radius squared, and this chapter turns it into a proper object.

2 · The commutator, and the tensor it produces

This is the longest computation in the book so far. The plan comes first, so that you can follow the whole argument without opening a single grind box.

Where we are going, and why the commutator

Why this object. Section 1's loop had two directions in it. You go one way, then the other, then back the first way, then back the second. Shrink that loop and what is left is precisely the difference between differentiating along μ\mu then ν\nu and differentiating along ν\nu then μ\mu. So the infinitesimal loop test is the commutator [μ,ν][\nabla_{\mu},\nabla_{\nu}] applied to a vector, and that is what we compute.

Three steps. First, νVρ\nabla_{\nu}V^{\rho} is a tensor with one upper and one lower index, so name it and put it aside. Second, apply μ\nabla_{\mu} to that object, which by Chapter 3.3 (3.3.32) needs a derivative and one connection term per index, and then expand the derivative of the product inside. That gives six terms in all. Third, subtract the same expression with μ\mu and ν\nu exchanged, and identify which terms are symmetric in μν\mu\nu and therefore die.

What to watch for. Every term containing a derivative of VV cancels. That is the surprise. What is left is proportional to VV itself with no derivatives on it. That is what makes the answer a tensor rather than a differential operator, and it is what §3 cashes in.

2.1 · The computation

Step one. Write WρννVρ=νVρ+ΓρνσVσW^{\rho}{}_{\nu}\equiv\nabla_{\nu}V^{\rho}=\partial_{\nu}V^{\rho}+\Gamma^{\rho}{}_{\nu\sigma}V^{\sigma}. This is a type (1,1)(1,1) tensor by the construction of Chapter 3.3 §5.

Step two. Now differentiate that object covariantly. Equation (3.3.32) gives one +Γ+\Gamma for the upper index and one Γ-\Gamma for the lower one:

μνVρ  =  μWρν  +  ΓρμλWλν    ΓλμνWρλ. \nabla_{\mu}\nabla_{\nu}V^{\rho} \;=\; \partial_{\mu}W^{\rho}{}_{\nu} \;+\; \Gamma^{\rho}{}_{\mu\lambda}\,W^{\lambda}{}_{\nu} \;-\; \Gamma^{\lambda}{}_{\mu\nu}\,W^{\rho}{}_{\lambda}. (3.4.10)

Step three. Put WW back in. In the first term the product rule of Chapter 0.1 §4 acts on ΓρνσVσ\Gamma^{\rho}{}_{\nu\sigma}V^{\sigma} and produces two pieces. In the third term the bracket WρλW^{\rho}{}_{\lambda} is exactly λVρ\nabla_{\lambda}V^{\rho}, and we leave it that way, which keeps the expression short and makes the cancellation visible. Labelling the six terms for reference:

μνVρ  =  μνVρT1  +  (μΓρνσ)VσT2  +  ΓρνσμVσT3+  ΓρμλνVλT4  +  ΓρμλΓλνσVσT5    ΓλμνλVρT6. \begin{aligned} \nabla_{\mu}\nabla_{\nu}V^{\rho} \;=\; & \underbrace{\partial_{\mu}\partial_{\nu}V^{\rho}}_{T_{1}} \;+\; \underbrace{\big(\partial_{\mu}\Gamma^{\rho}{}_{\nu\sigma}\big)V^{\sigma}}_{T_{2}} \;+\; \underbrace{\Gamma^{\rho}{}_{\nu\sigma}\,\partial_{\mu}V^{\sigma}}_{T_{3}}\\[4pt] & +\; \underbrace{\Gamma^{\rho}{}_{\mu\lambda}\,\partial_{\nu}V^{\lambda}}_{T_{4}} \;+\; \underbrace{\Gamma^{\rho}{}_{\mu\lambda}\Gamma^{\lambda}{}_{\nu\sigma}V^{\sigma}}_{T_{5}} \;-\; \underbrace{\Gamma^{\lambda}{}_{\mu\nu}\,\nabla_{\lambda}V^{\rho}}_{T_{6}}. \end{aligned} (3.4.11)

The grind box below writes that out with nothing folded away, for anyone who wants to see every index in place. The argument does not need it.

Grind box — the same six terms, with every substitution written out

Start from (3.4.10) and substitute Wρν=νVρ+ΓρνσVσW^{\rho}{}_{\nu}=\partial_{\nu}V^{\rho}+\Gamma^{\rho}{}_{\nu\sigma}V^{\sigma} in all three places, being careful to rename the summed index in each so that no letter does two jobs.

The first term.

μWρν  =  μ(νVρ+ΓρνσVσ)  =  μνVρ  +  (μΓρνσ)Vσ  +  ΓρνσμVσ, \partial_{\mu}W^{\rho}{}_{\nu} \;=\; \partial_{\mu}\Big(\partial_{\nu}V^{\rho}+\Gamma^{\rho}{}_{\nu\sigma}V^{\sigma}\Big) \;=\; \partial_{\mu}\partial_{\nu}V^{\rho} \;+\; \big(\partial_{\mu}\Gamma^{\rho}{}_{\nu\sigma}\big)V^{\sigma} \;+\; \Gamma^{\rho}{}_{\nu\sigma}\,\partial_{\mu}V^{\sigma},

where the product rule was applied to ΓρνσVσ\Gamma^{\rho}{}_{\nu\sigma}V^{\sigma}, both factors being functions of position. These are T1T_{1}, T2T_{2} and T3T_{3}.

The second term. Here Wλν=νVλ+ΓλνσVσW^{\lambda}{}_{\nu}=\partial_{\nu}V^{\lambda}+\Gamma^{\lambda}{}_{\nu\sigma}V^{\sigma}, so

ΓρμλWλν  =  ΓρμλνVλ  +  ΓρμλΓλνσVσ, \Gamma^{\rho}{}_{\mu\lambda}W^{\lambda}{}_{\nu} \;=\; \Gamma^{\rho}{}_{\mu\lambda}\,\partial_{\nu}V^{\lambda} \;+\; \Gamma^{\rho}{}_{\mu\lambda}\Gamma^{\lambda}{}_{\nu\sigma}V^{\sigma},

which are T4T_{4} and T5T_{5}. Note that λ\lambda is summed inside the double product. It is the index shared by the two connection factors, and it is what makes T5T_{5} a matrix product rather than an entry-by-entry one.

The third term. WρλW^{\rho}{}_{\lambda} is by definition λVρ\nabla_{\lambda}V^{\rho}, and we deliberately do not expand it, because the whole term is about to vanish for a reason that has nothing to do with its contents. Written out it would be ΓλμνλVρΓλμνΓρλσVσ-\Gamma^{\lambda}{}_{\mu\nu}\partial_{\lambda}V^{\rho}-\Gamma^{\lambda}{}_{\mu\nu}\Gamma^{\rho}{}_{\lambda\sigma}V^{\sigma}, which is seven terms in total rather than six. Folding that pair back into λVρ\nabla_{\lambda}V^{\rho} is the only compression made anywhere in this computation.

Collecting gives (3.4.11) exactly.

2.2 · Antisymmetrising, one term at a time

Now write the same expression with μ\mu and ν\nu exchanged and subtract. We go through the six terms in order. Each one is a separate small argument, and each is named.

T1T_{1} dies by Clairaut. The term is μνVρ\partial_{\mu}\partial_{\nu}V^{\rho}, an ordinary second partial derivative of a set of smooth functions. Clairaut's theorem (Chapter 0.6 §6.1) says mixed partials commute, so exchanging μ\mu and ν\nu leaves it unchanged and the subtraction kills it. This is the same cancellation that Chapter 3.2 §8.3 used to build the Lie bracket.

T2T_{2} survives. Exchanging μ\mu and ν\nu turns (μΓρνσ)Vσ(\partial_{\mu}\Gamma^{\rho}{}_{\nu\sigma})V^{\sigma} into (νΓρμσ)Vσ(\partial_{\nu}\Gamma^{\rho}{}_{\mu\sigma})V^{\sigma}, which is a genuinely different object. The difference is (μΓρνσνΓρμσ)Vσ\big(\partial_{\mu}\Gamma^{\rho}{}_{\nu\sigma}-\partial_{\nu}\Gamma^{\rho}{}_{\mu\sigma}\big)V^{\sigma}.

T3T_{3} and T4T_{4} die together. This is the one that needs a moment. Individually neither is symmetric, but their sum is. Exchange μ\mu and ν\nu in T3=ΓρνσμVσT_{3}=\Gamma^{\rho}{}_{\nu\sigma}\partial_{\mu}V^{\sigma} to get ΓρμσνVσ\Gamma^{\rho}{}_{\mu\sigma}\partial_{\nu}V^{\sigma}. Now relabel the summed index σλ\sigma\to\lambda, which changes nothing since it is a dummy, and the result is ΓρμλνVλ\Gamma^{\rho}{}_{\mu\lambda}\partial_{\nu}V^{\lambda}, which is T4T_{4} exactly. By the same two moves, exchanging μν\mu\leftrightarrow\nu in T4T_{4} gives T3T_{3}. So the pair T3+T4T_{3}+T_{4} maps to T4+T3T_{4}+T_{3} and is symmetric, and the subtraction removes it.

T5T_{5} survives. Exchanging gives ΓρνλΓλμσVσ\Gamma^{\rho}{}_{\nu\lambda}\Gamma^{\lambda}{}_{\mu\sigma}V^{\sigma}, and the two connection factors do not commute as index objects, so the difference is (ΓρμλΓλνσΓρνλΓλμσ)Vσ\big(\Gamma^{\rho}{}_{\mu\lambda}\Gamma^{\lambda}{}_{\nu\sigma}-\Gamma^{\rho}{}_{\nu\lambda}\Gamma^{\lambda}{}_{\mu\sigma}\big)V^{\sigma}.

T6T_{6} dies by vanishing torsion. The term is ΓλμνλVρ-\Gamma^{\lambda}{}_{\mu\nu}\nabla_{\lambda}V^{\rho}. Exchanging μ\mu and ν\nu affects only the two lower indices of the connection, and Chapter 3.3's Demand 2 made those symmetric: Γλμν=Γλνμ\Gamma^{\lambda}{}_{\mu\nu}=\Gamma^{\lambda}{}_{\nu\mu}. So the term is unchanged and the subtraction kills it. Had we allowed torsion, this term would have survived and contributed TλμνλVρ-T^{\lambda}{}_{\mu\nu}\nabla_{\lambda}V^{\rho} to the answer, which is exactly how the general formula reads in books that keep torsion.

Two terms have survived, T2T_{2} and T5T_{5}. Collect them, and give the array they multiply a name:

  [μ,ν]Vρ  =  RρσμνVσ,   \boxed{\;\big[\nabla_{\mu},\nabla_{\nu}\big]V^{\rho} \;=\; R^{\rho}{}_{\sigma\mu\nu}\,V^{\sigma},\;} (3.4.12)

That equation defines the symbol RρσμνR^{\rho}{}_{\sigma\mu\nu} by what it does. The two terms that survived the subtraction say what its components are:

  Rρσμν    μΓρνσ    νΓρμσ  +  ΓρμλΓλνσ    ΓρνλΓλμσ.   \boxed{\;R^{\rho}{}_{\sigma\mu\nu} \;\equiv\; \partial_{\mu}\Gamma^{\rho}{}_{\nu\sigma} \;-\; \partial_{\nu}\Gamma^{\rho}{}_{\mu\sigma} \;+\; \Gamma^{\rho}{}_{\mu\lambda}\Gamma^{\lambda}{}_{\nu\sigma} \;-\; \Gamma^{\rho}{}_{\nu\lambda}\Gamma^{\lambda}{}_{\mu\sigma}.\;} (3.4.13)

This is the Riemann curvature tensor. Before doing anything else with (3.4.13), notice its shape. It is antisymmetric in μν\mu\nu by construction, since swapping those two letters exchanges the first two terms with each other and the last two with each other, reversing the sign of the whole. And it is built from the connection and its first derivatives, hence from the metric and its first and second derivatives.

Recap — what went in, what came out

In: the covariant derivative of a (1,1)(1,1) tensor, the product rule, Clairaut's theorem, one relabelling of a dummy index, and the symmetry of the connection in its lower indices.

Out: the difference between differentiating in two orders is not a differential operator at all. It is multiplication by an array. Four of the six terms cancelled, and they were precisely the four containing derivatives of VV.

Cost: nothing was assumed beyond Chapter 3.3's two demands, and torsion is the only place a different choice would change the answer.

In plain terms 3.4.2

Shrinking the circuit test to a point converts it into a question about the order of two operations. Going a short way in one direction and then a short way in another, against the same two moves in the opposite order, traces out a small closed circuit, so the failure of the two orders to agree is exactly what the circuit measures. In the language now available, that failure is the difference between differentiating one way then the other and the other way then the first.

Computing it is the longest piece of algebra in the book so far, and it is worth knowing what to look for. Six terms appear. One is an ordinary second derivative and disappears because the order of those never matters. Two more form a pair that swaps into itself when the directions are exchanged, so the pair cancels against its mirror image. A fourth goes because those coefficients were required to be symmetric in their lower slots.

What matters is not the algebra but what survives it. Every term carrying a derivative of the transported field has vanished, and what remains is proportional to the field itself. An operation that looked certain to produce a differential operator has produced plain multiplication by an array of numbers, and that array depends on the geometry alone. It is the object this chapter is about.

3 · Why the result is a tensor

Now the payoff for §2's length. Look at what (3.4.13) is made of. Every ingredient on the right-hand side is a badly behaved object.

  • The connection is not a tensor, by Chapter 3.3 §5.3.
  • A derivative of a non-tensor is not a tensor either.
  • A product of two non-tensors is not a tensor.

Yet the combination of them is perfectly well behaved. Here is the argument, which takes three lines because §2 did the work.

The left side is a tensor. Chapter 3.3 §5 established that \nabla maps tensors to tensors, one rank higher. So νVρ\nabla_{\nu}V^{\rho} is a (1,1)(1,1) tensor, and μνVρ\nabla_{\mu}\nabla_{\nu}V^{\rho} is a (1,2)(1,2) tensor. A difference of two tensors of the same type is a tensor of that type (Chapter 2.4 §5.1). Hence [μ,ν]Vρ[\nabla_{\mu},\nabla_{\nu}]V^{\rho} is a (1,2)(1,2) tensor, whatever VV is.

The right side has no derivatives of VV. That is what §2's four cancellations bought. If even one of T1T_{1}, T3T_{3}, T4T_{4}, T6T_{6} had survived, the right side would depend on how VV varies near the point and the next step would be unavailable.

Therefore the array is a tensor. We have a quantity RρσμνR^{\rho}{}_{\sigma\mu\nu} such that RρσμνVσR^{\rho}{}_{\sigma\mu\nu}V^{\sigma} is a (1,2)(1,2) tensor for every vector field VV. That is the hypothesis of the quotient theorem of Chapter 2.4 §5.1, and its conclusion is that RρσμνR^{\rho}{}_{\sigma\mu\nu} is itself a tensor, of type (1,3)(1,3). The same thing can be said in the language of Chapter 3.2's Problem 3. The map V[μ,ν]VV\mapsto[\nabla_{\mu},\nabla_{\nu}]V is function-linear, meaning that if you multiply VV by any function hh the answer is multiplied by hh, with no leftover term in h\partial h. Function-linearity is exactly what distinguishes a tensor from a differential operator.

What just happened, said plainly

Four separately non-tensorial objects have been combined so that all their bad behaviour cancels. The four are Γ\partial\Gamma in two orderings and ΓΓ\Gamma\Gamma in two orderings. This is the second time the book has done this, and the first time was a rehearsal. Chapter 3.2 §8 found that μYν\partial_{\mu}Y^{\nu} is not a tensor while the antisymmetric combination XμμYνYμμXνX^{\mu}\partial_{\mu}Y^{\nu}-Y^{\mu}\partial_{\mu}X^{\nu} is, and its recap said explicitly that the trick would recur "much more consequentially, in Chapter 3.4". This is that recurrence.

The mechanism is the same both times. The offending piece of the transformation law is symmetric in the two indices being antisymmetrised, so antisymmetrising annihilates it. For the connection that offending piece carries 2xν/xαxγ\partial^{2}x'^{\nu}/\partial x^{\alpha}\partial x^{\gamma}, which is symmetric by Clairaut. Curvature is what is left of the connection after the part that depends on the chart has been subtracted away.

3.1 · A concrete demonstration: the flat plane in polar coordinates

Argument by quotient theorem is airtight and slightly bloodless. Here is the cancellation happening in a case where every number is known. Chapter 3.3 §7.8 found that the flat plane in polar coordinates has

Γrθθ  =  r,Γθrθ  =  Γθθr  =  1r,all others zero, \Gamma^{r}{}_{\theta\theta} \;=\; -\,r, \qquad \Gamma^{\theta}{}_{r\theta} \;=\; \Gamma^{\theta}{}_{\theta r} \;=\; \frac{1}{r}, \qquad\text{all others zero,} (3.4.14)

and Chapter 3.3 §9's second worked example showed that every one of those numbers is the inhomogeneous term of the transformation law, since the connection vanishes in the Cartesian chart. So this connection is entirely made of the part that is not tensorial. If §3's claim is right, the curvature must be exactly zero. Compute it.

Take ρ=r\rho=r, σ=θ\sigma=\theta, μ=r\mu=r, ν=θ\nu=\theta in (3.4.13) and write the four terms out:

Rrθrθ  =  rΓrθθ    θΓrrθ  +  ΓrrλΓλθθ    ΓrθλΓλrθ. R^{r}{}_{\theta r\theta} \;=\; \partial_{r}\Gamma^{r}{}_{\theta\theta} \;-\; \partial_{\theta}\Gamma^{r}{}_{r\theta} \;+\; \Gamma^{r}{}_{r\lambda}\Gamma^{\lambda}{}_{\theta\theta} \;-\; \Gamma^{r}{}_{\theta\lambda}\Gamma^{\lambda}{}_{r\theta}. (3.4.15)

Evaluate the four in turn, using (3.4.14). The first is r(r)=1\partial_{r}(-r)=-1. The second is zero, because Γrrθ=0\Gamma^{r}{}_{r\theta}=0. The third is zero, because both Γrrr\Gamma^{r}{}_{rr} and Γrrθ\Gamma^{r}{}_{r\theta} vanish. The fourth has two values of λ\lambda: λ=r\lambda=r gives ΓrθrΓrrθ=0\Gamma^{r}{}_{\theta r}\Gamma^{r}{}_{r\theta}=0, and λ=θ\lambda=\theta gives ΓrθθΓθrθ=(r)(1/r)=1\Gamma^{r}{}_{\theta\theta}\Gamma^{\theta}{}_{r\theta}=(-r)(1/r)=-1, entering with the minus sign in front. So

Rrθrθ  =  1Γ    0  +  0  (1)ΓΓ  =  0. R^{r}{}_{\theta r\theta} \;=\; \underbrace{-1}_{\partial\Gamma} \;-\;0\;+\;0\; \underbrace{-\,(-1)}_{-\Gamma\Gamma} \;=\; 0. (3.4.16)

The derivative term and the quadratic term are equal and opposite. Do it once more with the other index assignment, ρ=θ\rho=\theta, σ=r\sigma=r, μ=θ\mu=\theta, ν=r\nu=r:

Rθrθr  =  θΓθrr=0    rΓθθr=r(1/r)=1/r2  +  ΓθθλΓλrr=0    ΓθrλΓλθr=(1/r)(1/r)  =  1r21r2  =  0. R^{\theta}{}_{r\theta r} \;=\; \underbrace{\partial_{\theta}\Gamma^{\theta}{}_{rr}}_{=\,0} \;-\; \underbrace{\partial_{r}\Gamma^{\theta}{}_{\theta r}}_{=\,\partial_{r}(1/r)\,=\,-1/r^{2}} \;+\; \underbrace{\Gamma^{\theta}{}_{\theta\lambda}\Gamma^{\lambda}{}_{rr}}_{=\,0} \;-\; \underbrace{\Gamma^{\theta}{}_{r\lambda}\Gamma^{\lambda}{}_{\theta r}}_{=\,(1/r)(1/r)} \;=\; \frac{1}{r^{2}} - \frac{1}{r^{2}} \;=\; 0. (3.4.17)

Again exactly. Every remaining component vanishes term by term for the same kind of reason, and the whole tensor is identically zero. (Confirmed by computer algebra: all sixteen components of RρσμνR^{\rho}{}_{\sigma\mu\nu} for this metric are identically zero as functions of rr and θ\theta.) A connection made entirely of chart-dependent junk produces exactly no curvature, which is what a tensor built to discard chart-dependent junk had better do.

In plain terms 3.4.3

Length in a derivation should buy something, and the long computation just finished bought this. The comparison coefficients of the last chapter were notoriously not honest measuring devices: a good choice of description makes them vanish everywhere and a bad one brings them back. Anything built naively from them inherits that disease. The array produced by the two-order calculation does not, because the two orders differ by an antisymmetry while the diseased part of the coefficients is symmetric in exactly the slots being antisymmetrised.

There is a short and airtight way to see it that avoids the algebra altogether. The thing computed was a difference of two honest objects, so it is honest. It turned out proportional to the carried field itself, with no derivatives of that field anywhere. Something that turns any field into an honest object by plain multiplication must itself be honest, and that is a theorem from the tensor chapter rather than a hopeful remark.

The flat page in circular labels makes the cancellation visible with actual numbers. There the coefficients are non-zero and consist of nothing but chart-dependent junk, since they vanish in square labels. Compute the new array from them and the derivative piece and the quadratic piece come out equal and opposite in every component, leaving zero. That is the third and final form of the previous chapter's warning, now a calculation rather than a promise.

4 · Geodesic deviation, and where Part III's thesis lands

Everything so far has been mathematics. This section is where it becomes physics, and it is the section the whole of Part III has been walking toward, so it goes slowly.

Where we are going

Chapter 3.1 §4 derived, from Newton alone, that two nearby freely falling particles accelerate relative to one another at a rate proportional to their separation and to the second derivatives of the gravitational potential. It then said, in §4.6: "Chapter 3.4 constructs exactly such an object … and Chapter 3.4 §4 will derive an equation of the form ξ¨Rξ\ddot\xi\sim-R\xi that is the tidal equation written covariantly. When the two are set side by side, the identification is unavoidable."

Here is that section. We take two neighbouring geodesics, derive how their separation evolves, and find the Riemann tensor sitting where Newton had the second derivatives of the potential.

4.1 · The setup, and one identity

Consider a one-parameter family of geodesics xμ(τ,s)x^{\mu}(\tau,s). For each fixed ss, the curve traced out as τ\tau varies is a geodesic parametrised by its own proper time, and ss is the label saying which geodesic we are on. Two vector fields live naturally on the two-dimensional surface those curves sweep out, so define them both now:

uμ    xμτ(the four-velocity),ξμ    xμs(the separation). u^{\mu} \;\equiv\; \pdv{x^{\mu}}{\tau} \qquad\text{(the four-velocity)}, \qquad \xi^{\mu} \;\equiv\; \pdv{x^{\mu}}{s} \qquad\text{(the separation)}. (3.4.18)

Because τ\tau and ss are coordinates on that surface, uu and ξ\xi are coordinate vector fields, and Chapter 3.2 §8.4 proved that coordinate vector fields commute:

[u,ξ]  =  0. \big[u,\xi\big] \;=\; 0. (3.4.19)

One identity is needed before starting, and it is short. Compute the components of uξξu\nabla_{u}\xi-\nabla_{\xi}u directly:

uννξμξννuμ  =  uννξμξννuμ  +  Γμνρuνξρ    Γμνρξνuρ. u^{\nu}\nabla_{\nu}\xi^{\mu} - \xi^{\nu}\nabla_{\nu}u^{\mu} \;=\; u^{\nu}\partial_{\nu}\xi^{\mu} - \xi^{\nu}\partial_{\nu}u^{\mu} \;+\; \Gamma^{\mu}{}_{\nu\rho}\,u^{\nu}\xi^{\rho} \;-\; \Gamma^{\mu}{}_{\nu\rho}\,\xi^{\nu}u^{\rho}. (3.4.20)

In the last term, relabel the two summed indices, exchanging the names ν\nu and ρ\rho. That is free, since both are dummies. It becomes Γμρνξρuν\Gamma^{\mu}{}_{\rho\nu}\xi^{\rho}u^{\nu}, and now use the symmetry of the connection in its lower indices to write Γμρν=Γμνρ\Gamma^{\mu}{}_{\rho\nu}=\Gamma^{\mu}{}_{\nu\rho}. The last two terms of (3.4.20) are then identical and cancel. What remains is the Lie bracket of Chapter 3.2 (3.2.32), so

uξ    ξu  =  [u,ξ]  =  0,that isuξ  =  ξu. \nabla_{u}\xi \;-\; \nabla_{\xi}u \;=\; \big[u,\xi\big] \;=\; 0, \qquad\text{that is}\qquad \nabla_{u}\xi \;=\; \nabla_{\xi}u. (3.4.21)

4.2 · The derivation, in four lines

We want the second derivative of the separation along the geodesic, which is what "relative acceleration" means. Write D/dτuD/\dd\tau\equiv\nabla_{u} for the derivative along the curve.

Line 1. Apply u\nabla_{u} twice and use (3.4.21) to swap the inner pair:

D2ξμdτ2  =  uuξμ  =  uξuμ. \frac{D^{2}\xi^{\mu}}{\dd\tau^{2}} \;=\; \nabla_{u}\nabla_{u}\xi^{\mu} \;=\; \nabla_{u}\nabla_{\xi}u^{\mu}. (3.4.22)

Line 2. What we want on the right is §2's commutator, so rewrite that expression as the opposite ordering plus the difference between the two orderings:

uξuμ  =  ξuuμ  +  (uξξu)uμ. \nabla_{u}\nabla_{\xi}u^{\mu} \;=\; \nabla_{\xi}\nabla_{u}u^{\mu} \;+\; \Big(\nabla_{u}\nabla_{\xi}-\nabla_{\xi}\nabla_{u}\Big)u^{\mu}. (3.4.23)

Line 3. The first term on the right is zero. That is the only place the geodesic property is used, and it is used exactly once: uuμ=0\nabla_{u}u^{\mu}=0 is the geodesic equation (3.3.57), so differentiating it in any direction still gives zero.

Line 4. Convert the remaining commutator of directional derivatives into the commutator of coordinate derivatives. Expanding uξV=uμμ(ξννV)\nabla_{u}\nabla_{\xi}V=u^{\mu}\nabla_{\mu}(\xi^{\nu}\nabla_{\nu}V) with the Leibniz rule and subtracting the other ordering gives

(uξξu)Vμ  =  (uξνξuν)=  0νVμ  +  uαξβ[α,β]Vμ, \Big(\nabla_{u}\nabla_{\xi}-\nabla_{\xi}\nabla_{u}\Big)V^{\mu} \;=\; \underbrace{\big(\nabla_{u}\xi^{\nu}-\nabla_{\xi}u^{\nu}\big)}_{=\;0}\,\nabla_{\nu}V^{\mu} \;+\; u^{\alpha}\xi^{\beta}\big[\nabla_{\alpha},\nabla_{\beta}\big]V^{\mu}, (3.4.24)

and the first bracket vanishes by (3.4.21). Setting V=uV=u and using the definition (3.4.12) of the Riemann tensor:

D2ξμdτ2  =  uαξβRμγαβuγ. \frac{D^{2}\xi^{\mu}}{\dd\tau^{2}} \;=\; u^{\alpha}\xi^{\beta}\,R^{\mu}{}_{\gamma\alpha\beta}\,u^{\gamma}. (3.4.25)

Tidy the indices. Rearrange the factors, which are just numbers, and then use the antisymmetry of RR in its last two indices, Rμγαβ=RμγβαR^{\mu}{}_{\gamma\alpha\beta}=-R^{\mu}{}_{\gamma\beta\alpha}, noted under (3.4.13). Renaming γν\gamma\to\nu, βρ\beta\to\rho, ασ\alpha\to\sigma gives the standard form:

  D2ξμdτ2  =  Rμνρσ  uνξρuσ.   \boxed{\;\frac{D^{2}\xi^{\mu}}{\dd\tau^{2}} \;=\; -\,R^{\mu}{}_{\nu\rho\sigma}\;u^{\nu}\,\xi^{\rho}\,u^{\sigma}.\;} (3.4.26)

This is the equation of geodesic deviation. Read it before comparing it with anything. Take two particles, each moving as freely as it is possible to move, with no forces acting and no potentials anywhere, each following the straightest path available to it. They nevertheless accelerate relative to one another, and the rate is fixed by the curvature of the space they are in, linearly in their separation. In flat space R=0R=0, and neighbouring straight lines drift apart at a constant rate and no faster, which is what straight lines do.

Recap — what went in, what came out

In: a family of geodesics, the fact that coordinate vector fields commute, vanishing torsion (used once, in (3.4.21)), the geodesic equation (used once, in Line 3), and §2's commutator.

Out: the relative acceleration of two neighbouring free particles is minus the Riemann tensor contracted with two velocities and one separation.

Worth noticing: not one word of this derivation mentioned gravity, mass, or force. It is a statement about geodesics on any manifold with a metric.

4.3 · Setting it beside Newton

Now take (3.4.26) to the Newtonian limit, under exactly the three assumptions of Chapter 3.3 §8.3: slow motion, a static field, and a weak one.

Slow motion makes uμ(c,0,0,0)u^{\mu}\approx(c,0,0,0), since u0=cdt/dτcu^{0}=c\,\dd t/\dd\tau\approx c and the spatial components are smaller by v/cv/c. It also makes τt\tau\approx t. The separation between two particles at the same coordinate time is purely spatial, ξμ=(0,ξj)\xi^{\mu}=(0,\xi^{j}).

Substituting into (3.4.26), the sums over ν\nu and σ\sigma collapse to the single value 00 and the sum over ρ\rho to spatial values:

d2ξidt2  =  c2Ri0j0  ξj. \frac{\dd^{2}\xi^{i}}{\dd t^{2}} \;=\; -\,c^{2}\,R^{i}{}_{0j0}\;\xi^{j}. (3.4.27)

And here is Chapter 3.1's equation (3.1.14), derived there from Newton's law of gravitation and one linearisation, with no relativity anywhere in it:

d2didt2  =  Φ,ij  dj,Φ,ij  =  ijΦ. \frac{\dd^{2}d^{\,i}}{\dd t^{2}} \;=\; -\,\Phi_{,ij}\;d^{\,j}, \qquad \Phi_{,ij} \;=\; \partial_{i}\partial_{j}\Phi. (3.4.28)

Put (3.4.27) and (3.4.28) side by side. They are the same equation. The separation vector plays the same role, the second time derivative plays the same role, and in the place where Newton has the matrix of second derivatives of the potential, general relativity has

  c2Ri0j0  =  Φ,ij.   \boxed{\;c^{2}\,R^{i}{}_{0j0} \;=\; \Phi_{,ij}.\;} (3.4.29)

4.4 · The identification, checked rather than assumed

Two equations having the same shape is suggestive, not conclusive. So compute Ri0j0R^{i}{}_{0j0} directly from the weak-field metric and see whether (3.4.29) comes out. Everything needed is already derived: Chapter 3.1 §6.5 obtained g00=1+2Φ/c2g_{00}=1+2\Phi/c^{2} from the accelerating cabin and the Doppler effect, using no general relativity, and Chapter 3.3 §8.3 turned that into

Γi00  =  iΦc2  +  O ⁣(Φ2c4). \Gamma^{i}{}_{00} \;=\; \frac{\partial_{i}\Phi}{c^{2}} \;+\; O\!\left(\frac{\Phi^{2}}{c^{4}}\right). (3.4.30)

With that connection coefficient in hand we can compute the curvature component directly. Take (3.4.13), the definition from §2, and set ρ=i\rho=i, σ=0\sigma=0, μ=j\mu=j and ν=0\nu=0:

Ri0j0  =  jΓi00    0Γij0=0 (static)  +  ΓijλΓλ00    Γi0λΓλj0each a product of two first-order quantities. R^{i}{}_{0j0} \;=\; \partial_{j}\Gamma^{i}{}_{00} \;-\; \underbrace{\partial_{0}\Gamma^{i}{}_{j0}}_{=\,0\ \text{(static)}} \;+\; \underbrace{\Gamma^{i}{}_{j\lambda}\Gamma^{\lambda}{}_{00} \;-\; \Gamma^{i}{}_{0\lambda}\Gamma^{\lambda}{}_{j0}}_{\text{each a product of two first-order quantities}}. (3.4.31)

The second term vanishes because nothing depends on time. The two quadratic terms are products of two connection coefficients, each of which is first order in Φ/c2\Phi/c^{2}, so they are second order and are dropped at this accuracy. What survives is the first term, and by (3.4.30),

Ri0j0  =  j ⁣(iΦc2)  =  Φ,ijc2, R^{i}{}_{0j0} \;=\; \partial_{j}\!\left(\frac{\partial_{i}\Phi}{c^{2}}\right) \;=\; \frac{\Phi_{,ij}}{c^{2}}, (3.4.32)

which is (3.4.29) exactly. (Confirmed symbolically for a general potential Φ(x,y,z)\Phi(x,y,z), all nine components at once.) The identification was not fitted to make the two equations match. It was computed, from a metric component that Chapter 3.1 obtained without using any of this machinery.

The thesis of Part III

Chapter 3.1 established that the gravitational field can be deleted at any point by a change of coordinates, and that what cannot be deleted is the relative acceleration of two nearby freely falling bodies. It called that residue the entire physical content of gravity.

Equation (3.4.29) says what the residue is.

Tidal force is curvature, and curvature is what gravity is.

Not "is described by", not "is analogous to". The array of numbers you would measure by watching two neighbouring falling specks drift, and the array of numbers you compute by carrying a vector around an infinitesimal loop and asking how far it has turned, are the same array. The ocean going out twice a day and the failure of parallel transport to be path-independent are one phenomenon measured two ways.

4.5 · Three things that follow immediately

(i) The trace, and Laplace's equation. Contract (3.4.29) on ii and jj. On the left, the sum Ri0i0R^{i}{}_{0i0} is nearly the Ricci tensor: by the definition Rμν=RλμλνR_{\mu\nu}=R^{\lambda}{}_{\mu\lambda\nu} stated in this chapter's conventions,

R00  =  Rλ0λ0  =  R0000=0  +  Ri0i0  =  Φ,iic2  =  2Φc2, R_{00} \;=\; R^{\lambda}{}_{0\lambda0} \;=\; \underbrace{R^{0}{}_{000}}_{=\,0} \;+\; R^{i}{}_{0i0} \;=\; \frac{\Phi_{,ii}}{c^{2}} \;=\; \frac{\nabla^{2}\Phi}{c^{2}}, (3.4.33)

where R0000R^{0}{}_{000} vanishes because RR is antisymmetric in its last two indices and they are equal. Poisson's equation 2Φ=4πGρ\nabla^{2}\Phi=4\pi G\rho (Chapter 0.7 §7) then gives

R00  =  4πGρc2,and in vacuumR00  =  0. R_{00} \;=\; \frac{4\pi G\rho}{c^{2}}, \qquad\text{and in vacuum}\qquad R_{00} \;=\; 0. (3.4.34)

Chapter 3.1 §4.4 promised that its tracelessness result "is the covariant descendant of 2Φ=0\nabla^{2}\Phi=0". Equation (3.4.34) is that descendant, and Chapter 3.6 will promote it to the full field equations.

(ii) The ellipse. Chapter 3.1 §4.5 said: "that shape is what a tidal field is, and Chapter 3.4 will recover this very ellipse from the curvature tensor." Worked example 2 below does exactly that, computing Ri0j0R^{i}{}_{0j0} for a point mass and finding the eigenvalues (2GM/r3,+GM/r3,+GM/r3)/c2(-2GM/r^{3},\,+GM/r^{3},\,+GM/r^{3})/c^{2} that produce the stretch-and-squeeze.

(iii) The units are right. Φ,ij\Phi_{,ij} has dimensions s2\mathrm{s^{-2}}, so Ri0j0R^{i}{}_{0j0} has dimensions m2\mathrm{m^{-2}}. Curvature is one over a length squared, which is what §1's 1/a21/a^{2} also was. For the Earth's surface field, GM/r3=1.54×106s2GM/r^{3}=1.54\times10^{-6}\,\mathrm{s^{-2}} and dividing by c2c^{2} gives a curvature of about 1.7×1023m21.7\times10^{-23}\,\mathrm{m^{-2}}: a radius of curvature of roughly 2×10112\times10^{11} metres, which is why nobody noticed.

In plain terms 3.4.4

Take two specks of dust released near each other, each moving as freely as anything can move, each following the straightest path its surroundings allow, with nothing pushing either. Ask how their separation changes. The answer, derived here from geometry alone with no mention of gravity or mass or force, is that the separation accelerates at a rate proportional to itself, the constant supplied by the array the previous sections built out of the circuit test.

Now set that beside a result obtained at the start of this part from Newton alone. Two specks in free fall were shown there to drift apart or together at a rate proportional to their separation, the constant supplied by how the pull varies with place. The two statements are one statement. Where the older physics wrote the way the pull varies, the new geometry writes the amount by which a direction fails to come home unchanged, and computing the second from the first confirms it rather than suggesting it.

So the thing you feel as gravity when the tide comes in, and the thing curvature measures, are one object. The ocean rising, two falling specks separating, and a direction coming back turned after a trip around a small loop are three descriptions of one feature of the world. That sentence is what this whole part has been walking toward, and everything after it is bookkeeping by comparison.

5 · The symmetries, and how many numbers there really are

The array RρσμνR^{\rho}{}_{\sigma\mu\nu} has 44=2564^{4}=256 components in four dimensions. Almost all of them are redundant. This section derives the symmetries, then counts.

5.1 · The four symmetries, derived

Work with the fully lowered version RρσμνgρλRλσμνR_{\rho\sigma\mu\nu}\equiv g_{\rho\lambda}R^{\lambda}{}_{\sigma\mu\nu}, which is legitimate and unambiguous by Chapter 3.3's Problem 2.

(S1) Antisymmetry in the last pair: Rρσμν=RρσνμR_{\rho\sigma\mu\nu}=-R_{\rho\sigma\nu\mu}. Immediate from (3.4.13), as noted there: exchanging μ\mu and ν\nu exchanges the first two terms and the last two, flipping the overall sign. Lowering an index does not touch μν\mu\nu.

(S2) Antisymmetry in the first pair: Rρσμν=RσρμνR_{\rho\sigma\mu\nu}=-R_{\sigma\rho\mu\nu}. This one needs an argument, and the argument is pretty. We get at it by first extending the commutator to covectors. For any covector ω\omega and vector VV, the contraction ωρVρ\omega_{\rho}V^{\rho} is a scalar, and [μ,ν][\nabla_{\mu},\nabla_{\nu}] annihilates scalars, which is Chapter 3.3 (3.3.40) with vanishing torsion. Apply Leibniz to that vanishing expression:

0  =  [μ,ν](ωρVρ)  =  ([μ,ν]ωρ)Vρ  +  ωρRρσμνVσ. 0 \;=\; \big[\nabla_{\mu},\nabla_{\nu}\big]\big(\omega_{\rho}V^{\rho}\big) \;=\; \Big(\big[\nabla_{\mu},\nabla_{\nu}\big]\omega_{\rho}\Big)V^{\rho} \;+\; \omega_{\rho}\,R^{\rho}{}_{\sigma\mu\nu}V^{\sigma}. (3.4.35)

We want to strip VV off both terms, so relabel the summed indices in the second one, exchanging ρ\rho and σ\sigma, and both terms then carry VρV^{\rho}. Since VV is arbitrary, the coefficients must match:

[μ,ν]ωρ  =  Rσρμνωσ. \big[\nabla_{\mu},\nabla_{\nu}\big]\omega_{\rho} \;=\; -\,R^{\sigma}{}_{\rho\mu\nu}\,\omega_{\sigma}. (3.4.36)

Iterating Leibniz once more gives the rule for a (0,2)(0,2) tensor: one such term per lower index,

[μ,ν]Tρσ  =  RλρμνTλσ    RλσμνTρλ. \big[\nabla_{\mu},\nabla_{\nu}\big]T_{\rho\sigma} \;=\; -\,R^{\lambda}{}_{\rho\mu\nu}T_{\lambda\sigma} \;-\; R^{\lambda}{}_{\sigma\mu\nu}T_{\rho\lambda}. (3.4.37)

Now put T=gT=g. The left side is zero, because ρgμν=0\nabla_{\rho}g_{\mu\nu}=0 identically and so is any further derivative of it. The right side becomes, on lowering,

0  =  Rσρμν    Rρσμν, 0 \;=\; -\,R_{\sigma\rho\mu\nu} \;-\; R_{\rho\sigma\mu\nu}, (3.4.38)

which is (S2). So metric compatibility, the very demand that fixed the connection in Chapter 3.3, is what makes the curvature antisymmetric in its first pair.

(S3) The first Bianchi identity: Rρσμν+Rρμνσ+Rρνσμ=0R_{\rho\sigma\mu\nu}+R_{\rho\mu\nu\sigma}+R_{\rho\nu\sigma\mu}=0. Cyclic in the last three indices. Derive it directly from (3.4.13). Write the three terms out with the upper index kept:

Rρσμν  =  μΓρνσνΓρμσ  +  ΓρμλΓλνσΓρνλΓλμσ,Rρμνσ  =  νΓρσμσΓρνμ  +  ΓρνλΓλσμΓρσλΓλνμ,Rρνσμ  =  σΓρμνμΓρσν  +  ΓρσλΓλμνΓρμλΓλσν. \begin{aligned} R^{\rho}{}_{\sigma\mu\nu} \;=\;& \partial_{\mu}\Gamma^{\rho}{}_{\nu\sigma}-\partial_{\nu}\Gamma^{\rho}{}_{\mu\sigma} \;+\; \Gamma^{\rho}{}_{\mu\lambda}\Gamma^{\lambda}{}_{\nu\sigma}-\Gamma^{\rho}{}_{\nu\lambda}\Gamma^{\lambda}{}_{\mu\sigma},\\[3pt] R^{\rho}{}_{\mu\nu\sigma} \;=\;& \partial_{\nu}\Gamma^{\rho}{}_{\sigma\mu}-\partial_{\sigma}\Gamma^{\rho}{}_{\nu\mu} \;+\; \Gamma^{\rho}{}_{\nu\lambda}\Gamma^{\lambda}{}_{\sigma\mu}-\Gamma^{\rho}{}_{\sigma\lambda}\Gamma^{\lambda}{}_{\nu\mu},\\[3pt] R^{\rho}{}_{\nu\sigma\mu} \;=\;& \partial_{\sigma}\Gamma^{\rho}{}_{\mu\nu}-\partial_{\mu}\Gamma^{\rho}{}_{\sigma\nu} \;+\; \Gamma^{\rho}{}_{\sigma\lambda}\Gamma^{\lambda}{}_{\mu\nu}-\Gamma^{\rho}{}_{\mu\lambda}\Gamma^{\lambda}{}_{\sigma\nu}. \end{aligned} (3.4.39)

Add the three lines and go through the six derivative terms first. Take μΓρνσ\partial_{\mu}\Gamma^{\rho}{}_{\nu\sigma} from the first line and μΓρσν-\partial_{\mu}\Gamma^{\rho}{}_{\sigma\nu} from the third: the two connections have their lower indices in opposite orders, and Chapter 3.3's Demand 2 says the order does not matter, so the pair cancels. The same argument pairs νΓρμσ-\partial_{\nu}\Gamma^{\rho}{}_{\mu\sigma} from the first line with +νΓρσμ+\partial_{\nu}\Gamma^{\rho}{}_{\sigma\mu} from the second, and σΓρνμ-\partial_{\sigma}\Gamma^{\rho}{}_{\nu\mu} from the second with +σΓρμν+\partial_{\sigma}\Gamma^{\rho}{}_{\mu\nu} from the third. All six derivative terms are gone, in three pairs.

Now the six quadratic terms, by exactly the same mechanism. Pair +ΓρμλΓλνσ+\Gamma^{\rho}{}_{\mu\lambda}\Gamma^{\lambda}{}_{\nu\sigma} from the first line with ΓρμλΓλσν-\Gamma^{\rho}{}_{\mu\lambda}\Gamma^{\lambda}{}_{\sigma\nu} from the third: identical after using the symmetry of Γλνσ\Gamma^{\lambda}{}_{\nu\sigma}, so they cancel. Pair ΓρνλΓλμσ-\Gamma^{\rho}{}_{\nu\lambda}\Gamma^{\lambda}{}_{\mu\sigma} with +ΓρνλΓλσμ+\Gamma^{\rho}{}_{\nu\lambda}\Gamma^{\lambda}{}_{\sigma\mu}, and ΓρσλΓλνμ-\Gamma^{\rho}{}_{\sigma\lambda}\Gamma^{\lambda}{}_{\nu\mu} with +ΓρσλΓλμν+\Gamma^{\rho}{}_{\sigma\lambda}\Gamma^{\lambda}{}_{\mu\nu}. All six vanish. The cyclic sum is zero, and every single cancellation used the same one fact, that the connection is symmetric in its lower indices.

(S4) Pair exchange: Rρσμν=RμνρσR_{\rho\sigma\mu\nu}=R_{\mu\nu\rho\sigma}. This is a consequence of the first three rather than an independent fact, and the proof is a four-line combinatorial manoeuvre with no new ingredient in it. It sits in the grind box, and the conclusion is what we use.

Grind box — pair exchange from (S1), (S2) and (S3)

Write the first Bianchi identity four times, each time cycling the last three indices of a different starting arrangement. Abbreviate the four indices as a,b,c,da,b,c,d.

(i)Rabcd+Racdb+Radbc=0,(ii)Rbcda+Rbdac+Rbacd=0,(iii)Rcdab+Rcabd+Rcbda=0,(iv)Rdabc+Rdbca+Rdcab=0. \begin{aligned} \text{(i)}\quad & R_{abcd}+R_{acdb}+R_{adbc}=0,\\ \text{(ii)}\quad & R_{bcda}+R_{bdac}+R_{bacd}=0,\\ \text{(iii)}\quad & R_{cdab}+R_{cabd}+R_{cbda}=0,\\ \text{(iv)}\quad & R_{dabc}+R_{dbca}+R_{dcab}=0. \end{aligned}

Form the combination (i) − (ii) − (iii) + (iv). Each of the four is separately zero, so the combination is zero. Now sort its twelve terms into two groups, using only (S1) and (S2) to move indices about. The first Bianchi identity is not used again from here.

Group one: four terms that cancel in pairs.

  • +Racdb+R_{acdb} from (i) against Rcabd-R_{cabd} from (iii). Apply (S2) to the first, Racdb=RcadbR_{acdb}=-R_{cadb}, and then (S1) to the last pair, Rcadb=+Rcabd-R_{cadb}=+R_{cabd}. The two terms are equal and oppositely signed, so they cancel.
  • +Radbc+R_{adbc} from (i) against +Rdabc+R_{dabc} from (iv). Here (S2) alone gives Radbc=RdabcR_{adbc}=-R_{dabc}, so the sum is zero.
  • Rbcda-R_{bcda} from (ii) against Rcbda-R_{cbda} from (iii). Again (S2): Rbcda=RcbdaR_{bcda}=-R_{cbda}, so the two add to zero.
  • Rbdac-R_{bdac} from (ii) against +Rdbca+R_{dbca} from (iv). Apply (S2), Rbdac=RdbacR_{bdac}=-R_{dbac}, then (S1) on the last pair, Rdbac=+Rdbca-R_{dbac}=+R_{dbca}. Equal and oppositely signed, so they cancel.

Group two: four terms that double up.

  • +Rabcd+R_{abcd} from (i) and Rbacd-R_{bacd} from (ii). By (S2), Rbacd=+Rabcd-R_{bacd}=+R_{abcd}, so together they give 2Rabcd2R_{abcd}.
  • Rcdab-R_{cdab} from (iii) and +Rdcab+R_{dcab} from (iv). By (S2), Rdcab=RcdabR_{dcab}=-R_{cdab}, so together they give 2Rcdab-2R_{cdab}.

Everything has now been accounted for, and what the vanishing combination says is

2Rabcd    2Rcdab  =  0,that isRabcd  =  Rcdab. 2R_{abcd} \;-\; 2R_{cdab} \;=\; 0, \qquad\text{that is}\qquad R_{abcd} \;=\; R_{cdab}. \qquad\blacksquare

Notice that the three inputs were used in different places: the first Bianchi identity supplied the four vanishing sums, and the two antisymmetries did all the rearranging. That is why (S4) is a consequence rather than an independent assumption.

Numerical confirmation of the algebra. Taking a random four-index array with (S1) and (S2) imposed and nothing else, the combination (i) − (ii) − (iii) + (iv) was evaluated for all 256256 index assignments and compared with 2(RabcdRcdab)2\big(R_{abcd}-R_{cdab}\big). The two agreed to 9×10169\times10^{-16}. So the rearrangement above is an identity in (S1) and (S2) alone, exactly as claimed.

Numerical confirmation. A deliberately generic four-dimensional metric was used, one with no symmetry at all, with off-diagonal entries and all four coordinates appearing. The Riemann tensor was computed from (3.4.13) and all four identities checked at a generic point. The largest violation of any of them was 1.2×10131.2\times10^{-13} against components of size 0.190.19, which is round-off.

5.2 · The count, derived

Now count, in nn dimensions, and then set n=4n=4. The argument has three stages and each is a sentence.

Stage 1. Treat the index pairs as single labels. By (S1) and (S2), RabcdR_{abcd} is antisymmetric within abab and within cdcd. An antisymmetric pair of indices in nn dimensions takes

N  =  (n2)  =  n(n1)2 N \;=\; \binom{n}{2} \;=\; \frac{n(n-1)}{2} (3.4.40)

independent values, which is Chapter 2.4 §7.2's count of an antisymmetric array. In four dimensions N=6N=6. So RR may be regarded as an N×NN\times N array M[ab],[cd]M_{[ab],[cd]}, which is 3636 numbers in four dimensions rather than 256256.

Stage 2. Pair exchange makes that array symmetric. By (S4), MAB=MBAM_{AB}=M_{BA} where AA and BB are the pair labels. A symmetric N×NN\times N array has N(N+1)/2N(N+1)/2 independent entries, which for N=6N=6 is 2121.

Stage 3. The first Bianchi identity removes exactly one more. Here is the step that has to be done carefully, because most of the cyclic identity is already implied by the other symmetries. The claim is that, given (S1), (S2) and (S4), the cyclic identity (S3) is equivalent to the single statement

R[abcd]  =  0, R_{[abcd]} \;=\; 0, (3.4.41)

the vanishing of the completely antisymmetric part. The reason is that the cyclic sum Rabcd+Racdb+RadbcR_{abcd}+R_{acdb}+R_{adbc}, once the pair symmetries are imposed, is itself totally antisymmetric in all four indices, and any totally antisymmetric four-index array in nn dimensions has

(n4)  =  n(n1)(n2)(n3)24 \binom{n}{4} \;=\; \frac{n(n-1)(n-2)(n-3)}{24} (3.4.42)

independent components, again by Chapter 2.4 §7.2. In four dimensions that count is (44)=1\binom44=1. There is exactly one totally antisymmetric four-index array up to scale, namely the Levi-Civita symbol of Chapter 2.4 §8.1, so (S3) imposes precisely one further condition.

The answer.

#  =  N(N+1)2    (n4)  =  n(n1)4(n(n1)2+1)    (n4)=    n2(n21)12   \begin{aligned} \#\;&=\; \frac{N(N+1)}{2} \;-\; \binom{n}{4} \;=\; \frac{n(n-1)}{4}\left(\frac{n(n-1)}{2}+1\right) \;-\; \binom{n}{4}\\[6pt] &=\; \boxed{\;\frac{n^{2}\big(n^{2}-1\big)}{12}\;} \end{aligned} (3.4.43)

The boxed expression is what that difference collapses to once the binomials are written out and simplified. Now put in the number of dimensions we actually live in:

42(421)12  =  161512  =    20   \frac{4^{2}\big(4^{2}-1\big)}{12} \;=\; \frac{16\cdot15}{12} \;=\; \boxed{\;20\;} (3.4.44)

independent components, or directly, 211=2021-1=20.

The count in low dimensions, and one check
nnN=(n2)N=\binom n2N(N+1)/2N(N+1)/2(n4)\binom n4independentn2(n21)/12n^{2}(n^{2}-1)/12
211011
336066
462112020
5105555050
61512015105105

In two dimensions there is a single curvature number at each point, which is why §1's sphere had one answer and why surfaces are so much easier than spacetimes. The whole table was confirmed independently. The four symmetry conditions were written down as linear constraints on a general n4n^{4}-component array and the rank of the resulting system computed, and the number of free components came out 1,6,20,50,1051,6,20,50,105 exactly.

5.3 · The same twenty, counted a completely different way

A number derived once is a number. Derived twice by independent routes, it is a fact. Here is the second route, which also settles two words that will be used together for the rest of Part III. A local inertial frame is a region small enough that gravity has no detectable effect across it. The locally inertial coordinates are the particular coordinates, constructed below, that are adapted to such a region. The frame is the physical situation, and the coordinates are the bookkeeping. Both will be needed again in §7.

Locally inertial coordinates

At any point pp one can choose coordinates in which

gμν(p)=ημνandρgμν(p)=0, g_{\mu\nu}(p)=\eta_{\mu\nu} \qquad\text{and}\qquad \partial_{\rho}g_{\mu\nu}(p)=0,

and therefore Γλμν(p)=0\Gamma^{\lambda}{}_{\mu\nu}(p)=0. One cannot in general also arrange ρσgμν(p)=0\partial_{\rho}\partial_{\sigma}g_{\mu\nu}(p)=0.

Both halves are constructions rather than assertions.

Getting g=ηg=\eta at pp. gμν(p)g_{\mu\nu}(p) is a real symmetric non-degenerate matrix, so by the spectral theorem (Chapter 0.5 §6) there is an orthogonal change of basis diagonalising it, and rescaling each basis vector by the square root of the modulus of its eigenvalue turns the diagonal entries into ±1\pm1. The signature fixes the pattern to (+,,,)(+,-,-,-). That is a linear change of coordinates, so it is a legitimate chart change.

Getting Γ=0\Gamma=0 at pp. Put pp at the origin and define new coordinates

yμ  =  xμ  +  12Γμαβ(p)  xαxβ. y^{\mu} \;=\; x^{\mu} \;+\; \half\,\Gamma^{\mu}{}_{\alpha\beta}(p)\;x^{\alpha}x^{\beta}. (3.4.45)

Two facts about that substitution matter. At the origin yμ/xα=δμα\partial y^{\mu}/\partial x^{\alpha}=\delta^{\mu}{}_{\alpha}, so the linear part is untouched and g(p)g(p) is still η\eta. And 2yμ/xαxβ=Γμαβ(p)\partial^{2}y^{\mu}/\partial x^{\alpha}\partial x^{\beta}=\Gamma^{\mu}{}_{\alpha\beta}(p) everywhere. Feed those two into Chapter 3.3's transformation law (3.3.26). At pp every Jacobian is a Kronecker delta, so the law collapses to

Γμαβ(p)  =  Γμαβ(p)    Γμαβ(p)  =  0. \Gamma'^{\mu}{}_{\alpha\beta}(p) \;=\; \Gamma^{\mu}{}_{\alpha\beta}(p) \;-\; \Gamma^{\mu}{}_{\alpha\beta}(p) \;=\; 0. (3.4.46)

This is precisely why Γ\Gamma had to be a non-tensor: a tensor vanishing in one chart vanishes in all, and no such construction would be possible. And ρgμν(p)=0\partial_{\rho}g_{\mu\nu}(p)=0 follows at once, since metric compatibility reads ρgμν=Γλρμgλν+Γλρνgμλ\partial_{\rho}g_{\mu\nu}=\Gamma^{\lambda}{}_{\rho\mu}g_{\lambda\nu}+\Gamma^{\lambda}{}_{\rho\nu}g_{\mu\lambda} and the right side is zero at pp.

Why the second derivatives cannot all be removed, and where the count comes from. Push the construction to the next order, with a cubic term 16Cμαβγxαxβxγ\tfrac16 C^{\mu}{}_{\alpha\beta\gamma}x^{\alpha}x^{\beta}x^{\gamma} whose coefficients are symmetric in the three lower indices. Then count what is available against what is demanded, in four dimensions:

  • Available. A symmetric triple of indices from four values takes (4+23)=20\binom{4+2}{3}=20 values, and μ\mu takes 44, so CC has 4×20=804\times20=80 free numbers.
  • Demanded. Setting αβgμν(p)=0\partial_{\alpha}\partial_{\beta}g_{\mu\nu}(p)=0 is one condition for each symmetric pair αβ\alpha\beta (10 values) and each symmetric pair μν\mu\nu (10 values), so 10×10=10010\times10=100 conditions.

Eighty knobs, one hundred demands. Twenty combinations of second derivatives of the metric survive every possible change of coordinates. That is the same twenty as before, arrived at with no mention of Bianchi identities or antisymmetric pairs. It also says what the twenty are. They are the irreducible second-derivative content of the metric, the part no observer can transform away.

The same counting at the two lower orders is worth doing as well, because both answers are ones you have seen. The linear part AμαA^{\mu}{}_{\alpha} has 1616 free numbers and g(p)=ηg(p)=\eta is 1010 conditions, leaving 1610=616-10=6. That is the dimension of the Lorentz group, three boosts and three rotations, exactly as Part II found. The quadratic part has 4×10=404\times10=40 numbers against ρgμν(p)=0\partial_{\rho}g_{\mu\nu}(p)=0's 4×10=404\times10=40 conditions, matching exactly, which is why Γ(p)\Gamma(p) can always be removed and never more than that.

Familiar ground — this is Chapter 3.1's falling laboratory

Chapter 3.1 §5 asked how big a freely falling laboratory may be before tidal effects become detectable, and answered with an explicit bound. Section 5.3 is the same statement made exact. Coordinates exist in which, at one event, the metric is Minkowski's and its first derivatives vanish. Inside such a frame, to first order, physics is the physics of Part II and there is no gravity at all. What cannot be removed is the second derivatives, twenty of them, and §4 identified those with the tidal field. The size of the laboratory is set by how far you can go before the second-order terms matter, which is Chapter 3.1's bound, now with a name for the obstruction.

In plain terms 3.4.5

An array with four slots in four dimensions carries two hundred and fifty-six numbers, almost all duplicates. Three properties cut the list down, and each is derived rather than declared. Swapping the last two slots flips the sign, which was built into the construction. Swapping the first two also flips the sign, and that traces back to the demand that carrying a direction preserves its length. A third relation ties together the three ways of cycling the last three slots.

Counting then goes in three stages. A pair of slots that flips sign under exchange behaves like a single label taking six values, so the array is a six-by-six table. A further symmetry makes that table symmetric. The cyclic relation removes one entry more, since what it says is that the completely antisymmetric part vanishes, and in four dimensions there is only one such part. Twenty survive.

The same twenty appear again by a route with nothing in common with the first. Ask how much of the geometry a well-chosen observer can make disappear near a chosen event. The rule for measuring distances can be made to look flat there, and its rate of change made to vanish, both by explicit construction. The rates of change of those rates cannot: the knobs fall short by exactly twenty. Those twenty are what no viewpoint removes, which is what the first chapter said about the tide.

6 · Ricci, the scalar, and what contraction throws away

Twenty numbers per point is a lot to carry into a field equation. Contracting reduces them, and the question is which contraction and at what cost.

6.1 · There is essentially only one contraction

A (1,3)(1,3) tensor can be contracted by summing its upper index against any one of its three lower ones. Take them in turn.

Against the first lower index. Rλλμν=gλρRρλμνR^{\lambda}{}_{\lambda\mu\nu}=g^{\lambda\rho}R_{\rho\lambda\mu\nu}. The array RρλμνR_{\rho\lambda\mu\nu} is antisymmetric in ρλ\rho\lambda by (S2) and gλρg^{\lambda\rho} is symmetric, and contracting a symmetric array with an antisymmetric one over both indices gives zero. That is Chapter 2.4 §7's standard argument, since relabelling the two summed indices reproduces the expression with the opposite sign. So this contraction vanishes identically.

Against the third lower index. Rλμνλ=RλμλνR^{\lambda}{}_{\mu\nu\lambda}=-R^{\lambda}{}_{\mu\lambda\nu} by (S1), which is minus the remaining case.

Against the second lower index. This is the only one that is both non-zero and not a duplicate:

  Rμν    Rλμλν   \boxed{\;R_{\mu\nu} \;\equiv\; R^{\lambda}{}_{\mu\lambda\nu}\;} (3.4.47)

That object is the Ricci tensor. Its sign is a convention, and this book's is stated in the opening callout.

It is symmetric. Lower the contracted index and use pair exchange:

Rμν  =  gλσRσμλν  =(S4)  gλσRλνσμ  =  gσλRσνλμ  =  Rνμ, R_{\mu\nu} \;=\; g^{\lambda\sigma}R_{\sigma\mu\lambda\nu} \;\overset{\text{(S4)}}{=}\; g^{\lambda\sigma}R_{\lambda\nu\sigma\mu} \;=\; g^{\sigma\lambda}R_{\sigma\nu\lambda\mu} \;=\; R_{\nu\mu}, (3.4.48)

where the third step relabelled the two summed indices λσ\lambda\leftrightarrow\sigma and used the symmetry of gλσg^{\lambda\sigma}. So Ricci has 1010 independent components in four dimensions, not 1616.

Contracting once more, with the inverse metric, gives a single number at each point:

R    gμνRμν, R \;\equiv\; g^{\mu\nu}R_{\mu\nu}, (3.4.49)

the Ricci scalar or scalar curvature. Being a scalar, it is the same number for every observer, which makes it the simplest honest statement about curvature available anywhere.

6.2 · What was discarded

Riemann has 2020 components and Ricci has 1010. So contraction throws away exactly 1010 components of information in four dimensions. That leftover has a name.

⚑ The Weyl tensor, named but not developed

The part of the Riemann tensor not determined by the Ricci tensor is the Weyl tensor CρσμνC_{\rho\sigma\mu\nu}: the piece with all its traces removed. In four dimensions it carries the 2010=1020-10=10 components contraction discards. ⚑ We quote three of its properties and use none of them: it has the same symmetries as Riemann, every contraction of it vanishes, and in the mixed form CρσμνC^{\rho}{}_{\sigma\mu\nu} it is unchanged when the metric is multiplied by an arbitrary positive function of position, so it records the shape of the light cones and not the scale of anything. Chapter 7.3 meets that last property again under the name conformal invariance.

Two consequences of the counting are worth having now. In three dimensions Riemann has 66 components and so does Ricci, so the Weyl tensor vanishes identically and Ricci determines the whole curvature. That is why gravity in three dimensions has no propagating waves. In two dimensions there is one component altogether, and it must be proportional to RR. Problem 1 works out the constant.

6.3 · Why this matters immediately: vacuum is not flat

Here is the reason the discarded part cannot be dismissed. Section 4.5 derived R00=4πGρ/c2R_{00}=4\pi G\rho/c^{2}, so outside a mass, where ρ=0\rho=0, the Ricci tensor has a vanishing component, and Chapter 3.6 will show the whole of it vanishes there. Yet the tidal field outside a mass is emphatically not zero. Chapter 3.1 computed it, and §4 identified it with Ri0j0R^{i}{}_{0j0}, which by (3.4.29) is Φ,ij/c2\Phi_{,ij}/c^{2} and is non-zero at every point outside the body.

Both statements hold because Chapter 3.1 §4.4's tidal matrix is traceless in vacuum and non-zero. The trace is the Ricci part, and the traceless remainder is the Weyl part. So:

vacuum    Rμν=0,butRρσμν    0. \text{vacuum} \;\Longrightarrow\; R_{\mu\nu}=0, \qquad\text{but}\qquad R_{\rho\sigma\mu\nu} \;\neq\; 0. (3.4.50)

Ricci flat is not flat. If it were, there would be no gravity outside any body, no orbits, no tides and no gravitational waves. The whole of Chapter 3.7 lives in a region where Rμν=0R_{\mu\nu}=0 and the curvature is what makes Mercury precess.

In plain terms 3.4.6

Twenty numbers at every point is more than a field equation can carry, and there is a standard way of boiling an array down: sum one slot against another. Doing it here is more constrained than it looks, since two of the three summations give either nothing or a copy of the third, for reasons traceable to the antisymmetries. So there is one way, and it gives a symmetric array of ten numbers, which boils once more into a single number everybody agrees on.

The array has fallen apart into independent pieces again, as the toolkit kept promising it would, and the pieces are not equally interesting. Ten numbers went into the pot and ten did not, and the ones left out have every trace removed. They carry the effects surviving where there is no matter, so the omission matters. Outside any body the summed-down array vanishes while the full array does not, and that gap is where orbits, tides and gravitational waves live.

That distinction was visible long before the machinery existed. The first chapter of this part found that a small ball of falling dust changes shape while holding its volume, and identified the volume statement with the law of gravity in empty space. The volume part is what gets summed down; the shape-changing part survives the summation. Empty space is gravitationally empty in the first sense and thoroughly occupied in the second.

7 · The second Bianchi identity, and a tensor with no divergence

This section is short and its result is the most consequential single line in the chapter, so its importance is flagged in advance.

Why this section matters more than it looks

Chapter 2.6 built the energy–momentum tensor TμνT^{\mu\nu} and proved it satisfies μTμν=0\partial_{\mu}T^{\mu\nu}=0, which is the local conservation of energy and momentum. On a curved manifold that becomes μTμν=0\nabla_{\mu}T^{\mu\nu}=0.

Any field equation of the form (geometry) == (constant) ×  Tμν\times\;T^{\mu\nu} therefore requires the geometrical side to have vanishing divergence as an identity, true for every metric whatever, not merely on shell. Otherwise the equation would impose an extra condition on matter that nothing justifies.

This section produces the unique combination of RμνR_{\mu\nu} and RgμνR\,g_{\mu\nu} with that property. Chapter 3.6 does not get to choose the left-hand side of Einstein's equations. This section chooses it.

7.1 · The identity

The identity we are after says something about how the curvature varies from point to point. Here is the claim:

λRρσμν  +  μRρσνλ  +  νRρσλμ  =  0, \nabla_{\lambda}R^{\rho}{}_{\sigma\mu\nu} \;+\; \nabla_{\mu}R^{\rho}{}_{\sigma\nu\lambda} \;+\; \nabla_{\nu}R^{\rho}{}_{\sigma\lambda\mu} \;=\; 0, (3.4.51)

cyclic in the three lower indices λμν\lambda\mu\nu, which are the derivative index and the antisymmetric pair together. The proof uses §5.3's locally inertial coordinates and is four lines.

Line 1. Fix a point pp and adopt coordinates with Γλμν(p)=0\Gamma^{\lambda}{}_{\mu\nu}(p)=0, which §5.3 constructed. At pp, and only at pp, the definition (3.4.13) loses its quadratic terms:

Rρσμνp  =  μΓρνσ    νΓρμσ. R^{\rho}{}_{\sigma\mu\nu}\big|_{p} \;=\; \partial_{\mu}\Gamma^{\rho}{}_{\nu\sigma} \;-\; \partial_{\nu}\Gamma^{\rho}{}_{\mu\sigma}. (3.4.52)

Line 2. The covariant derivative of anything reduces to the ordinary derivative at pp, since every correction term carries a factor of Γ(p)=0\Gamma(p)=0. Note carefully what is not being claimed: Γ\partial\Gamma is not zero at pp, only Γ\Gamma itself is, which is why (3.4.52) has content. Differentiating,

λRρσμνp  =  λμΓρνσ    λνΓρμσ, \nabla_{\lambda}R^{\rho}{}_{\sigma\mu\nu}\big|_{p} \;=\; \partial_{\lambda}\partial_{\mu}\Gamma^{\rho}{}_{\nu\sigma} \;-\; \partial_{\lambda}\partial_{\nu}\Gamma^{\rho}{}_{\mu\sigma}, (3.4.53)

the derivatives of the quadratic terms also vanishing at pp because each is (Γ)Γ+Γ(Γ)(\partial\Gamma)\Gamma+\Gamma(\partial\Gamma) and every term carries an undifferentiated Γ\Gamma.

Line 3. Write the cyclic sum. Each of the three terms contributes two pieces, so there are six, and they cancel in pairs by Clairaut's theorem:

(λμΓρνσλνΓρμσ)+  (μνΓρλσμλΓρνσ)+  (νλΓρμσνμΓρλσ)  =  0, \begin{aligned} &\big(\partial_{\lambda}\partial_{\mu}\Gamma^{\rho}{}_{\nu\sigma} - \partial_{\lambda}\partial_{\nu}\Gamma^{\rho}{}_{\mu\sigma}\big)\\ +\;&\big(\partial_{\mu}\partial_{\nu}\Gamma^{\rho}{}_{\lambda\sigma} - \partial_{\mu}\partial_{\lambda}\Gamma^{\rho}{}_{\nu\sigma}\big)\\ +\;&\big(\partial_{\nu}\partial_{\lambda}\Gamma^{\rho}{}_{\mu\sigma} - \partial_{\nu}\partial_{\mu}\Gamma^{\rho}{}_{\lambda\sigma}\big) \;=\; 0, \end{aligned} (3.4.54)

The first term of each line cancels the second term of the line below it. For instance, λμΓρνσ\partial_{\lambda}\partial_{\mu}\Gamma^{\rho}{}_{\nu\sigma} goes against μλΓρνσ-\partial_{\mu}\partial_{\lambda}\Gamma^{\rho}{}_{\nu\sigma}, and those two are equal because mixed partials commute. The third line's second term goes against the first line's first term after the same move.

Line 4. Equation (3.4.51) is an equation between tensors, since every term in it is a covariant derivative of a tensor. It has just been shown true at pp in one particular chart. By Chapter 2.4 §6, a tensor equation true in one chart is true in every chart. And pp was arbitrary. Hence (3.4.51) holds everywhere, in every chart. \blacksquare

(Confirmed numerically: for the generic four-dimensional metric of §5.1's grind box, the largest value of the cyclic sum at a generic point was 1.3×10111.3\times10^{-11}, against individual terms of size 0.070.07.)

7.2 · Contract it twice

Lower the first index of (3.4.51), which is legitimate because g=0\nabla g=0 lets the metric pass through the derivative:

λRρσμν  +  μRρσνλ  +  νRρσλμ  =  0. \nabla_{\lambda}R_{\rho\sigma\mu\nu} \;+\; \nabla_{\mu}R_{\rho\sigma\nu\lambda} \;+\; \nabla_{\nu}R_{\rho\sigma\lambda\mu} \;=\; 0. (3.4.55)

First contraction. Multiply by gρμg^{\rho\mu} and sum. Take the three terms one at a time, saying what each contraction produces.

Term one: gρμRρσμν=Rμσμν=Rσνg^{\rho\mu}R_{\rho\sigma\mu\nu}=R^{\mu}{}_{\sigma\mu\nu}=R_{\sigma\nu}, by the definition (3.4.47). So it becomes λRσν\nabla_{\lambda}R_{\sigma\nu}.

Term two: the gρμg^{\rho\mu} meets the derivative index, giving gρμμRρσνλ=ρRρσνλg^{\rho\mu}\nabla_{\mu}R_{\rho\sigma\nu\lambda}=\nabla^{\rho}R_{\rho\sigma\nu\lambda}.

Term three: gρμRρσλμ=Rμσλμg^{\rho\mu}R_{\rho\sigma\lambda\mu}=R^{\mu}{}_{\sigma\lambda\mu}, which by (S1) is Rμσμλ=Rσλ-R^{\mu}{}_{\sigma\mu\lambda}=-R_{\sigma\lambda}. So it becomes νRσλ-\nabla_{\nu}R_{\sigma\lambda}.

λRσν  +  ρRρσνλ    νRσλ  =  0. \nabla_{\lambda}R_{\sigma\nu} \;+\; \nabla^{\rho}R_{\rho\sigma\nu\lambda} \;-\; \nabla_{\nu}R_{\sigma\lambda} \;=\; 0. (3.4.56)

Second contraction. Multiply by gσνg^{\sigma\nu} and sum, again one term at a time.

Term one: gσνλRσν=λRg^{\sigma\nu}\nabla_{\lambda}R_{\sigma\nu}=\nabla_{\lambda}R, the Ricci scalar.

Term two: gσνRρσνλg^{\sigma\nu}R_{\rho\sigma\nu\lambda}. Use (S2) to write Rρσνλ=RσρνλR_{\rho\sigma\nu\lambda}=-R_{\sigma\rho\nu\lambda}, and then the contraction of the first and third indices is the Ricci tensor by (3.4.47), giving Rρλ-R_{\rho\lambda}. So the term is ρRρλ-\nabla^{\rho}R_{\rho\lambda}.

Term three: gσννRσλ=σRσλ-g^{\sigma\nu}\nabla_{\nu}R_{\sigma\lambda}=-\nabla^{\sigma}R_{\sigma\lambda}.

λR    ρRρλ    σRσλ  =  0ρRρλ  =  12λR. \nabla_{\lambda}R \;-\; \nabla^{\rho}R_{\rho\lambda} \;-\; \nabla^{\sigma}R_{\sigma\lambda} \;=\; 0 \qquad\Longrightarrow\qquad \nabla^{\rho}R_{\rho\lambda} \;=\; \half\,\nabla_{\lambda}R. (3.4.57)

The last step added the two identical terms. They are identical because ρ\rho and σ\sigma are both summed, and the name of a dummy index is private.

7.3 · The Einstein tensor

Our goal now is to write (3.4.57) as a single divergence, so that the whole content sits inside one \nabla. The metric passes through the derivative once more, so λR=ρ(gρλR)\nabla_{\lambda}R=\nabla^{\rho}\big(g_{\rho\lambda}R\big), and the equation reads ρ(Rρλ12gρλR)=0\nabla^{\rho}\big(R_{\rho\lambda}-\half g_{\rho\lambda}R\big)=0. Define

  Gμν    Rμν    12Rgμν,μGμν  =  0.   \boxed{\;G_{\mu\nu} \;\equiv\; R_{\mu\nu} \;-\; \half\,R\,g_{\mu\nu}, \qquad \nabla^{\mu}G_{\mu\nu} \;=\; 0.\;} (3.4.58)

This is the Einstein tensor, and three things about it are worth stating separately.

  • It is symmetric, because both RμνR_{\mu\nu} and gμνg_{\mu\nu} are.
  • It is built from the metric and its first two derivatives and nothing else.
  • Its divergence vanishes identically, for every metric, as a consequence of (3.4.51) and not of any equation of motion.

(Confirmed numerically for the generic metric: all four components of μGμν\nabla^{\mu}G_{\mu\nu} came out below 4×10114\times10^{-11} against derivatives of Ricci of size 0.110.11.)

This is the shape of the field equations, three chapters early

Put the pieces together. Matter carries a symmetric tensor TμνT^{\mu\nu} with μTμν=0\nabla_{\mu}T^{\mu\nu}=0. Geometry supplies exactly one symmetric tensor built from gg and its first two derivatives whose divergence vanishes identically, namely GμνG^{\mu\nu}. As Chapter 3.6 will make precise, it also supplies gμνg^{\mu\nu} itself, whose divergence vanishes because g=0\nabla g=0. So the most general equation of the permitted form is

Gμν  +  Λgμν  =  κTμν G^{\mu\nu} \;+\; \Lambda\,g^{\mu\nu} \;=\; \kappa\,T^{\mu\nu}

for constants Λ\Lambda and κ\kappa. Chapter 3.6 fixes κ=8πG/c4\kappa=8\pi G/c^{4} by demanding Newton's limit, quotes ⚑ Lovelock's theorem for the uniqueness, and discusses Λ\Lambda. The shape of the answer, however, was decided here, by one contracted identity.

In plain terms 3.4.7

Some results earn their place by what they forbid rather than by what they produce, and this is one of them. Differentiating the curvature array and adding up three cyclic arrangements gives exactly zero, always, for every geometry, as an identity rather than as a condition. Proving it is a matter of standing at a point in the frame where the comparison coefficients vanish, at which the curvature is a plain difference of derivatives, and watching six second derivatives cancel in pairs because the order of ordinary differentiation never matters.

Summing that identity down twice turns it into a statement about the boiled-down array: a particular combination of it and the single curvature number has no divergence whatever. Now recall what the previous part established about matter. Energy and momentum are packaged in one symmetric object whose divergence vanishes, and that vanishing is the local statement that nothing is created or destroyed.

Setting those two facts side by side almost writes the law of gravity by itself. If geometry is to be equated with matter, then whatever stands on the geometry side must have vanishing divergence automatically, for every geometry, or else the equation would quietly impose an extra demand on matter that nothing justifies. Exactly one combination qualifies, and three chapters from now it will be sitting on the left of the field equations, put there not by taste but by this identity.

8 · Flat if and only if the curvature vanishes

Chapter 3.3 §2 promised a genuine test and warned that neither varying metric components nor non-zero connection coefficients was one. Here is the test, in the form of a theorem with one easy half and one hard half.

The easy half, proved. Suppose there exist coordinates in which gμνg_{\mu\nu} is constant, say equal to ημν\eta_{\mu\nu}. Then σgμν=0\partial_{\sigma}g_{\mu\nu}=0 everywhere in that chart. By the Christoffel formula (3.3.50), every Γλμν\Gamma^{\lambda}{}_{\mu\nu} therefore vanishes everywhere in that chart. By (3.4.13), every component of RρσμνR^{\rho}{}_{\sigma\mu\nu} then vanishes in that chart too, each of the four terms being either a derivative of zero or a product containing zero. And RR is a tensor by §3, so a tensor vanishing in one chart vanishes in every chart (Chapter 2.4 §6). Hence Rρσμν=0R^{\rho}{}_{\sigma\mu\nu}=0 in all charts. \blacksquare

Note what that argument settles. The flat plane in polar coordinates has non-constant metric components and non-zero connection coefficients, yet it must have zero curvature, because Cartesian coordinates exist for it. Section 3.1's explicit computation was the verification. This is the reason.

⚑ The converse, quoted

If Rρσμν=0R^{\rho}{}_{\sigma\mu\nu}=0 throughout a region that is simply connected, then coordinates exist in which gμν=ημνg_{\mu\nu}=\eta_{\mu\nu} throughout. Simply connected means that every closed loop in the region can be shrunk to a point without leaving it. We quote this result and do not prove it.

The idea is visible from §1 even without the proof. Vanishing curvature makes parallel transport path-independent for loops that can be shrunk away, since the loop test is the integral of the curvature over any surface the loop bounds. Path-independence then lets one transport a chosen basis at one point to every other point unambiguously, producing a frame that is parallel everywhere. The work is in showing that this frame is the coordinate basis of an actual chart.

The topological caveat is not a technicality. Chapter 3.3's Problem 1 built a cone, whose curvature vanishes at every point away from the apex and which is nonetheless not a piece of the plane: a circle around the apex has circumference 2πkR2\pi kR rather than 2πR2\pi R. The region excluding the apex is not simply connected, loops encircling it cannot be shrunk, and the theorem does not apply. Chapter 3.9 meets the same distinction on cosmological scales, where a universe can be flat everywhere and still closed.

In plain terms 3.4.8

A question raised earlier and left hanging can now be settled. A space is genuinely flat, in the sense that one unbending grid covers all of it, exactly when the array built out of the circuit test vanishes everywhere. One direction is easy: lay down such a grid and the rule for distances has constant entries, so the comparison coefficients vanish and the array vanishes, and since the array is an honest measuring device its vanishing in one description means its vanishing in all.

The other direction is harder and is quoted rather than derived, with the reason for believing it sketched. If the array vanishes then carrying a direction around any circuit that can be shrunk away brings it back unchanged, so a direction chosen at one place can be carried unambiguously to every other place, which is the raw material a flat grid is made of.

One qualification in that sentence does real work rather than decorating the statement. Circuits that cannot be shrunk are exempt, and spaces containing them can be flat everywhere while refusing to be a plain grid. A cone rolled from paper is flat away from its tip, having been made without stretching anything, and yet a circle around the tip has the wrong circumference. Local information does not always add up to global information, and the last chapter of this part turns that on the universe.

9 · Worked examples

Worked example 1 — the sphere's curvature, and the figure checked

Compute the Riemann tensor, the Ricci tensor and the Ricci scalar of the sphere of radius aa from the Christoffel symbols of Chapter 3.3 §9, and check the result against §1's measured holonomy.

The one independent component. By §5.2 a two-dimensional space has exactly one, and by the symmetries it may be taken to be RθϕθϕR^{\theta}{}_{\phi\theta\phi}. From (3.4.13) with ρ=θ\rho=\theta, σ=ϕ\sigma=\phi, μ=θ\mu=\theta, ν=ϕ\nu=\phi:

Rθϕθϕ  =  θΓθϕϕ    ϕΓθθϕ  +  ΓθθλΓλϕϕ    ΓθϕλΓλθϕ. R^{\theta}{}_{\phi\theta\phi} \;=\; \partial_{\theta}\Gamma^{\theta}{}_{\phi\phi} \;-\; \partial_{\phi}\Gamma^{\theta}{}_{\theta\phi} \;+\; \Gamma^{\theta}{}_{\theta\lambda}\Gamma^{\lambda}{}_{\phi\phi} \;-\; \Gamma^{\theta}{}_{\phi\lambda}\Gamma^{\lambda}{}_{\theta\phi}.

Take the four terms in order, using (3.4.1).

The first: θ(sinθcosθ)\partial_{\theta}(-\sin\theta\cos\theta). Write sinθcosθ=12sin2θ\sin\theta\cos\theta=\half\sin2\theta, whose derivative is cos2θ\cos2\theta, so the term is cos2θ=sin2θcos2θ-\cos2\theta = \sin^{2}\theta-\cos^{2}\theta.

The second: Γθθϕ=0\Gamma^{\theta}{}_{\theta\phi}=0, so the term is zero.

The third: both Γθθθ\Gamma^{\theta}{}_{\theta\theta} and Γθθϕ\Gamma^{\theta}{}_{\theta\phi} vanish, so the term is zero.

The fourth: the sum over λ\lambda keeps only λ=ϕ\lambda=\phi, since Γθϕθ=0\Gamma^{\theta}{}_{\phi\theta}=0. It gives ΓθϕϕΓϕθϕ=(sinθcosθ)(cotθ)=+cos2θ-\Gamma^{\theta}{}_{\phi\phi}\Gamma^{\phi}{}_{\theta\phi} = -(-\sin\theta\cos\theta)(\cot\theta) = +\cos^{2}\theta.

Adding: sin2θcos2θ+cos2θ=sin2θ\sin^{2}\theta-\cos^{2}\theta+\cos^{2}\theta=\sin^{2}\theta. So

Rθϕθϕ  =  sin2θ. R^{\theta}{}_{\phi\theta\phi} \;=\; \sin^{2}\theta.

Lower the first index with gθθ=a2g_{\theta\theta}=a^{2}:

Rθϕθϕ  =  a2sin2θ. R_{\theta\phi\theta\phi} \;=\; a^{2}\sin^{2}\theta.

Ricci and the scalar. By (3.4.47), Rϕϕ=Rλϕλϕ=Rθϕθϕ=sin2θR_{\phi\phi}=R^{\lambda}{}_{\phi\lambda\phi}=R^{\theta}{}_{\phi\theta\phi}=\sin^{2}\theta, the λ=ϕ\lambda=\phi term vanishing by antisymmetry in the last two indices. Similarly Rθθ=RϕθϕθR_{\theta\theta}=R^{\phi}{}_{\theta\phi\theta}, which the same computation with the roles exchanged gives as 11. So

Rμν  =  (100sin2θ)  =  1a2gμν,R  =  gμνRμν  =  1a2δμμ  =  2a2. R_{\mu\nu} \;=\; \begin{pmatrix}1&0\\0&\sin^{2}\theta\end{pmatrix} \;=\; \frac{1}{a^{2}}\,g_{\mu\nu}, \qquad R \;=\; g^{\mu\nu}R_{\mu\nu} \;=\; \frac{1}{a^{2}}\,\delta^{\mu}{}_{\mu} \;=\; \frac{2}{a^{2}}.

Positive, as promised in this chapter's sign convention. (All of the above was reproduced independently by computer algebra, including the full Riemann tensor.)

The check against §1. Problem 1 below shows that in two dimensions Rμνρσ=K(gμρgνσgμσgνρ)R_{\mu\nu\rho\sigma}=K\big(g_{\mu\rho}g_{\nu\sigma}-g_{\mu\sigma}g_{\nu\rho}\big) with K=R/2K=R/2, and that the holonomy of a small loop is KK times the area enclosed. Here K=122/a2=1/a2K=\tfrac12\cdot 2/a^{2}=1/a^{2}, so the predicted holonomy is A/a2\mathcal A/a^{2}. That is (3.4.9), derived in §1 by integrating the transport equation and measured in the figure to six decimal places. Three routes, one number.

Worked example 2 — Chapter 3.1's ellipse, recovered from the curvature tensor

Chapter 3.1 §4.5 promised that "Chapter 3.4 will recover this very ellipse from the curvature tensor". Do so.

The components. By (3.4.32), in the weak static field, Ri0j0=Φ,ij/c2R^{i}{}_{0j0}=\Phi_{,ij}/c^{2}. Chapter 3.1 §4.3 computed that Hessian for a point mass, Φ=GM/r\Phi=-GM/r, in two lines of differentiation, obtaining

Φ,ij  =  GMr3(δij3ninj),ni=xir. \Phi_{,ij} \;=\; \frac{GM}{r^{3}}\Big(\delta_{ij}-3n^{i}n^{j}\Big), \qquad n^{i}=\frac{x^{i}}{r}.

So the curvature components a falling observer measures are

Ri0j0  =  GMc2r3(δij3ninj). R^{i}{}_{0j0} \;=\; \frac{GM}{c^{2}r^{3}}\Big(\delta_{ij}-3n^{i}n^{j}\Big).

The eigenvalues. Chapter 3.1's grind box diagonalised exactly this matrix: applying it to n^\hat{\vv n} gives 2n^-2\hat{\vv n}, and applying it to any vector perpendicular to n^\hat{\vv n} leaves it alone. So the spectrum is (2,+1,+1)(-2,+1,+1) times GM/(c2r3)GM/(c^{2}r^{3}), whatever direction the mass lies in.

The motion. Feed that into the geodesic deviation equation in the form (3.4.27), ξ¨i=c2Ri0j0ξj\ddot\xi^{i}=-c^{2}R^{i}{}_{0j0}\xi^{j}. Along n^\hat{\vv n} the eigenvalue is 2GM/(c2r3)-2GM/(c^{2}r^{3}), so ξ¨=+2GMξ/r3\ddot\xi = +2GM\xi/r^{3} and the separation grows. Across it the eigenvalue is +GM/(c2r3)+GM/(c^{2}r^{3}), so ξ¨=GMξ/r3\ddot\xi=-GM\xi/r^{3} and the separation shrinks. Integrating twice from rest gives

ξ(t)=ρ(1+Kt2),ξ(t)=ρ(112Kt2),K=GMr3, \xi_{\parallel}(t) = \rho\big(1+Kt^{2}\big), \qquad \xi_{\perp}(t) = \rho\big(1-\tfrac12Kt^{2}\big), \qquad K=\frac{GM}{r^{3}},

which is Chapter 3.1 (3.1.24) character for character: the ring stretched along the field and squeezed across it, the ellipse in that chapter's figure.

What is new, and what is not. The numbers are the same numbers. What has changed is what they are components of. In Chapter 3.1 they were second derivatives of a potential, a quantity with no coordinate-independent meaning and no place in a relativistic theory. Here they are components of a tensor, and the tensor exists whether or not the field is weak, whether or not motion is slow, and whether or not any potential can be defined. The ellipse was always a picture of the curvature, and Chapter 3.1 drew it four chapters before there was a name for what it was drawing.

One number. At the Earth's surface, K=GM/R3=1.54×106s2K=GM_{\oplus}/R_{\oplus}^{3}=1.54\times10^{-6}\,\mathrm{s^{-2}}, so the largest curvature component is K/c2=1.7×1023m2K/c^{2}=1.7\times10^{-23}\,\mathrm{m^{-2}}. Two objects released one metre apart vertically separate by an extra millimetre after about 2525 seconds. The effect is small. The identification is exact.

10 · Your turn

Problem 1 — curvature in two dimensions

(a) Using only (S1), (S2), (S4) and the fact that a two-dimensional space has one independent component, show that Rμνρσ=K(gμρgνσgμσgνρ)R_{\mu\nu\rho\sigma}=K\big(g_{\mu\rho}g_{\nu\sigma}-g_{\mu\sigma}g_{\nu\rho}\big) for some function KK. (b) Contract twice to show K=R/2K=R/2. (c) Verify on the sphere of radius aa and on the flat plane in polar coordinates. (d) Assuming the small-loop result of Problem 3, show that the holonomy of a small loop in two dimensions is KK times the enclosed area, and hence recover §1's answer.

Solution

(a) The right-hand side is the most general array with the required symmetries built from the metric alone: it is antisymmetric in μν\mu\nu, antisymmetric in ρσ\rho\sigma, and symmetric under exchanging the pairs, as one checks by inspection. In two dimensions the space of arrays with those symmetries is one-dimensional by §5.2, and the displayed combination is a non-zero member of it, so every member is a multiple of it. The multiple may vary from point to point, hence a function KK.

(b) Contract μ\mu with ρ\rho using gμρg^{\mu\rho}: the first term gives gμρgμρgνσ=ngνσg^{\mu\rho}g_{\mu\rho}g_{\nu\sigma}=n\,g_{\nu\sigma} and the second gives gμρgμσgνρ=gνσg^{\mu\rho}g_{\mu\sigma}g_{\nu\rho}=g_{\nu\sigma}, so Rνσ=K(n1)gνσR_{\nu\sigma}=K(n-1)g_{\nu\sigma}. Contract again with gνσg^{\nu\sigma}: R=Kn(n1)R=K\,n(n-1). In n=2n=2 that is R=2KR=2K, so K=R/2K=R/2.

(c) Sphere: R=2/a2R=2/a^{2} from Worked example 1, so K=1/a2K=1/a^{2}, constant, as a sphere should be. Checking the formula directly, K(gθθgϕϕ0)=(1/a2)(a2)(a2sin2θ)=a2sin2θ=RθϕθϕK(g_{\theta\theta}g_{\phi\phi}-0)= (1/a^{2})(a^{2})(a^{2}\sin^{2}\theta)=a^{2}\sin^{2}\theta=R_{\theta\phi\theta\phi} ✓. Flat polar: §3.1 found every component zero, so K=0K=0 and R=0R=0.

(d) By Problem 3 the change in a vector after a small loop of area dA\dd A in the plane spanned by two directions is RR contracted with the loop's area element. In two dimensions there is only one plane, and substituting the form from (a) gives a rotation by KdAK\,\dd A. Adding up over a large region gives KdA\int K\,\dd A, which on a sphere of radius aa is A/a2\mathcal A/a^{2}. That is equation (3.4.9), obtained from the tensor rather than from integrating along the path.

Problem 2 — a space that looks the same everywhere, and what §7 forces

Suppose Rμνρσ=K(gμρgνσgμσgνρ)R_{\mu\nu\rho\sigma}=K\big(g_{\mu\rho}g_{\nu\sigma}-g_{\mu\sigma}g_{\nu\rho}\big) in nn dimensions, with KK a function of position. (a) Compute RμνR_{\mu\nu}, RR and GμνG_{\mu\nu}. (b) Impose μGμν=0\nabla^{\mu}G_{\mu\nu}=0 from §7 and deduce that for n3n\geq3 the function KK must be a constant. (c) Say why n=2n=2 is exempt, and what that means for the sphere.

Solution

(a) From Problem 1(b), Rμν=K(n1)gμνR_{\mu\nu}=K(n-1)g_{\mu\nu} and R=Kn(n1)R=Kn(n-1). Hence

Gμν=K(n1)gμν12Kn(n1)gμν=K(n1)(1n2)gμν. G_{\mu\nu} = K(n-1)g_{\mu\nu} - \half Kn(n-1)g_{\mu\nu} = K(n-1)\left(1-\frac n2\right)g_{\mu\nu}.

(b) Take the divergence. The metric passes through \nabla untouched, so

0=μGμν=(n1)(1n2)νK. 0 = \nabla^{\mu}G_{\mu\nu} = (n-1)\left(1-\frac n2\right)\nabla_{\nu}K.

For n3n\geq3 the numerical factor (n1)(1n/2)(n-1)(1-n/2) is non-zero, so νK=0\nabla_{\nu}K=0 and KK is the same at every point. This is Schur's lemma, and it is a genuine consequence of §7 rather than an assumption: a space that looks the same in every direction at every point must have the same curvature at every point, and the reason is the contracted Bianchi identity.

(c) In n=2n=2 the factor (1n/2)(1-n/2) vanishes, so Gμν0G_{\mu\nu}\equiv0 identically and no condition on KK follows. Two-dimensional surfaces may therefore have curvature that varies from place to place while still having the form of (a). An egg is one. The sphere's constancy is a fact about the sphere, not a theorem. This vanishing of GμνG_{\mu\nu} in two dimensions is also why gravity has nothing to say in two dimensions, a point Chapter 7.2 will need.

Problem 3 — the loop test, from the commutator

Transport a vector VV around a small closed parallelogram whose sides are εaμ\varepsilon a^{\mu} and δbν\delta b^{\nu}. (a) Argue that the change in VV after the circuit is, to leading order, ΔVρ=εδRρσμνVσaμbν\Delta V^{\rho}=-\varepsilon\delta\,R^{\rho}{}_{\sigma\mu\nu}V^{\sigma}a^{\mu}b^{\nu}, using the fact that transporting along aa then bb differs from bb then aa by the commutator. (b) Check the scaling: the change is second order in the size of the loop, which is first order in its area. Why does that make "curvature per unit area" the right thing to measure? (c) Apply this to a small loop on the sphere near the equator, with sides along θ\theta and ϕ\phi, and confirm the answer agrees with §1.

Solution

(a) Parallel transport along a short displacement εa\varepsilon a is, to first order, the operation 1εaμΓμ1-\varepsilon a^{\mu}\Gamma_{\mu} acting on the components, so transporting along the four sides in order composes four such operations. Everything at first order cancels, because each side is traversed once in each direction. What survives at order εδ\varepsilon\delta is the commutator of the two transport operations, and §2 computed exactly that commutator: it is RρσμνVσaμbν-R^{\rho}{}_{\sigma\mu\nu}V^{\sigma}a^{\mu}b^{\nu}, the minus sign arising because transport is generated by Γ-\Gamma while the derivative carries +Γ+\Gamma. The overall sign also depends on which way round the circuit is taken, and the antisymmetry of RR in μν\mu\nu is exactly what makes the answer reverse when the circuit is reversed, as it must.

(b) The change is proportional to εδ\varepsilon\delta, and the loop's area is also proportional to εδ\varepsilon\delta, while its perimeter is proportional to ε+δ\varepsilon+\delta. So the ratio (change)/(area) approaches a finite non-zero limit as the loop shrinks, whereas (change)/(perimeter) goes to zero. Area is the right denominator, which is why §1's ratio was independent of the loop and why the figure's halve the loop button leaves it unmoved.

(c) Near the equator take a=θa=\partial_{\theta} and b=ϕb=\partial_{\phi}, with ε=Δθ\varepsilon=\Delta\theta and δ=Δϕ\delta=\Delta\phi. Using Rθϕθϕ=sin2θR^{\theta}{}_{\phi\theta\phi}=\sin^{2}\theta and its partners, the rotation angle works out to sinθΔθΔϕ\sin\theta\,\Delta\theta\,\Delta\phi, and the area of the same little patch is a2sinθΔθΔϕa^{2}\sin\theta\,\Delta\theta\,\Delta\phi. The ratio is 1/a21/a^{2}, matching (3.4.9) and Worked example 1's KK.

Problem 4 — empty is not flat

(a) In four dimensions, how many components of the Riemann tensor are left undetermined by the statement Rμν=0R_{\mu\nu}=0? (b) Using §4's weak-field results, show explicitly that outside a point mass R00=0R_{00}=0 while Ri0j00R^{i}{}_{0j0}\neq0, and identify which of Chapter 3.1's results this is. (c) In how many dimensions would "empty implies flat" be true, and why?

Solution

(a) Riemann has 2020 and Ricci has 1010. Setting Ricci to zero imposes 1010 conditions, so 1010 components remain free. They are the Weyl components of §6.2.

(b) By (3.4.33), R00=2Φ/c2R_{00}=\nabla^{2}\Phi/c^{2}, and for Φ=GM/r\Phi=-GM/r Chapter 3.1 §4.4 computed 2Φ=0\nabla^{2}\Phi=0 away from the origin. That was its equation (3.1.22), the tracelessness of the tidal matrix. Meanwhile Worked example 2 gives Ri0j0=GM(δij3ninj)/(c2r3)R^{i}{}_{0j0}=GM(\delta_{ij}-3n^{i}n^{j})/(c^{2}r^{3}), whose eigenvalues are (2,1,1)GM/(c2r3)(-2,1,1)GM/(c^{2}r^{3}) and are plainly non-zero. So the trace vanishes and the matrix does not. This is exactly Chapter 3.1 §4.4's observation that the tidal matrix of any vacuum region is traceless, seen from the other side.

(c) In three dimensions. There §6.2's counting gives Riemann 66 components and Ricci 66, so Ricci determines the whole of Riemann and Rμν=0R_{\mu\nu}=0 forces Rρσμν=0R_{\rho\sigma\mu\nu}=0. Also trivially in two, where there is one component and it is RR. In four and above the Weyl part exists and the implication fails, which is the reason gravity in our world has vacuum solutions at all, and hence orbits, tides and waves.

The brick you just laid

You have a test for curvature that runs entirely from inside. Carry a vector around a closed loop, and the angle it comes back turned through, divided by the area enclosed, is a property of the place. On a sphere of radius aa that ratio is 1/a21/a^{2}, derived in §1 by integrating the transport equation along three legs and measured in the figure to six decimals.

Shrinking the loop turns the test into the commutator of two covariant derivatives, and §2 computed it in full. Six terms appear and four of them cancel, and the four that cancel are exactly the ones containing derivatives of the transported vector. What is left is multiplication by the Riemann tensor (3.4.13), which is a genuine tensor even though every ingredient of it is not, because the chart-dependent part of the connection is symmetric in the two indices being antisymmetrised. The flat plane in polar coordinates demonstrates the cancellation with numbers.

The thesis of Part III, landed. Two neighbouring geodesics obey (3.4.26), and in the Newtonian limit that equation is Chapter 3.1's tidal equation with c2Ri0j0c^{2}R^{i}{}_{0j0} standing where Φ,ij\Phi_{,ij} stood. That identification was not fitted but computed, from a metric component Chapter 3.1 obtained with no general relativity at all. Tidal force is curvature, and curvature is what gravity is. Chapter 3.1's ellipse is a picture of Ri0j0R^{i}{}_{0j0}, and Worked example 2 draws it again from the tensor.

The tensor has 2020 independent components in four dimensions, and we derived that twice, once from the symmetries and once by counting how much of the second derivatives of the metric no choice of coordinates can remove. Contraction gives the Ricci tensor and the Ricci scalar and discards ten components, the Weyl part, which is the part that survives in vacuum and carries the tides outside a mass. And the second Bianchi identity, contracted twice, produces μGμν=0\nabla^{\mu}G_{\mu\nu}=0 identically.

Where this gets spent. That last identity is the reason Chapter 3.6's field equations look the way they do. The geometry side must be divergence-free for every metric, and Gμν+ΛgμνG_{\mu\nu}+\Lambda g_{\mu\nu} is what qualifies. Chapter 3.5 builds the Lie derivative and Killing vectors, whose conserved quantities make Chapter 3.7 tractable. Chapter 3.7 solves Rμν=0R_{\mu\nu}=0 outside a spherical mass and finds Mercury's precession and the bending of light, paying Chapter 3.1's factor of two. Chapter 3.9 takes vanishing curvature and non-trivial topology seriously. And in Part VI the commutator of two covariant derivatives is computed again with an internal space in place of the tangent space, at which point (3.4.13) becomes the Yang–Mills field strength and this chapter is read a second time in different clothes.