Part III · General Relativity — Chapter 3.6

The Einstein Field Equations

The law of gravity, derived three times: cornered by a conservation requirement, extremised from an action, and its one free constant fixed by demanding that apples fall.

Where we are

Everything Part III has built so far describes how matter moves in a geometry that is handed to us. Chapter 3.3 §8.3 obtained Newton's law of falling from the geodesic equation. Chapter 3.4 §4 identified the tide with curvature. Both chapters closed by naming the same missing piece: nothing yet says which geometry a given lump of matter produces. This chapter says it.

The field equations are not guessed here, and they were not guessed historically either, though the route taken below is cleaner than the route taken in 1915. The argument is a cornering.

Chapter 2.6 built the object that carries energy and momentum, and it proved that the object has vanishing divergence. Anything set equal to that object must therefore have vanishing divergence too, identically, for every geometry whatever. Chapter 3.4 §7 constructed the only combination of the curvature that does. There is essentially nothing else available. What is left is Einstein's equation with one undetermined constant in it.

Here is the route. Section 1 establishes what is on the right-hand side, and it is not mass. Section 2 lists the constraints the equation must satisfy, and says where each one comes from. Section 3 is the cornering itself.

Section 4 derives the same equation a second way, by writing down the simplest scalar that can be built from a metric and extremising it. That scalar is the entry Chapter 1.2 §8.1 put in its table of actions six chapters early. Section 5 fixes the constant by demanding Newton's limit, and it is the section where a reader decides whether to believe the whole edifice. Section 6 examines the one term the argument of §3 cannot exclude. Section 7 counts equations against unknowns, and finds that the count works out for a reason.

Conventions. The signature is (+,,,)(+,-,-,-). The Riemann and Ricci signs are as stated loudly in Chapter 3.4. Both GG and cc are written out. Three sign traps are met below and each is flagged where it occurs: the sign of g00g_{00} in the weak-field metric, and the sign of the cosmological term.

Tools you'll need  — Chapter 3.4 §7 above all: the second Bianchi identity, its double contraction (3.4.57), and the Einstein tensor (3.4.58) with μGμν=0\nabla^{\mu}G_{\mu\nu}=0. Also 3.4 §6 (Ricci and the scalar), §5.3 (locally inertial coordinates) and §4.5's R00=2Φ/c2R_{00}=\nabla^{2}\Phi/c^{2}, equation (3.4.33). Chapter 2.6 §10 for the energy–momentum tensor, its components, and μTμν=0\partial_{\mu}T^{\mu\nu}=0. Chapter 3.5 §6 for the volume element, Jacobi's formula (3.5.47) and the divergence identity (3.5.51), and §5 for Stokes' theorem. Chapter 3.3 §7 (the Christoffel formula and metric compatibility) and §8.3 (the Newtonian limit, and its Γi00=iΦ/c2\Gamma^{i}{}_{00}=\partial_{i}\Phi/c^{2}, equation (3.3.71)). Chapter 3.1 §6.5, the derived weak-field component g00=1+2Φ/c2g_{00}=1+2\Phi/c^{2}, equation (3.1.39). Chapter 1.2 §3 (varying an action) and §8.1 (the table this chapter's fourth line comes from). Chapter 0.7 §7.5 for Poisson's equation.

1 · What sits on the right-hand side, and why it is not mass

Here is where this section is going. We identify the source of gravity as the energy–momentum tensor of Chapter 2.6. We promote its conservation law to a curved manifold, and we write down its form for the two kinds of matter this book needs. The important consequence is that pressure is a source of gravity, and no Newtonian intuition supplies that.

1.1 · The candidate, and why nothing simpler will do

In Newtonian gravity the source is the mass density ρ\rho, a single function. That cannot survive into a relativistic theory, and Chapter 2.5 §9 already established the reason. Mass is not additive, and it is not separately conserved.

Energy is both of those things, so energy is the natural replacement. But Chapter 2.5 §3 showed that energy is one component of a four-vector. A theory that sources gravity with energy alone would therefore source it with a frame-dependent quantity, and two observers in relative motion would compute different geometries for the same physical situation. That is not a difficulty to be worked around. It is a contradiction.

Chapter 2.6 §10 built the object that fixes this. The energy–momentum tensor TμνT^{\mu\nu} arose there as the Noether current of spacetime translations. Its components were computed there too, and they are worth having in front of us again:

ComponentWhat it is
T00T^{00}energy density
T0iT^{0i}c×c\times momentum density, equivalently energy flux /c/c
TijT^{ij}the stress: flux of ii-momentum across a surface of constant xjx^{j}

The tensor is symmetric, Tμν=TνμT^{\mu\nu}=T^{\nu\mu}. Chapter 2.6 §10.1 checked that explicitly for the electromagnetic field, after fixing the Noether current with an improvement term. The same chapter, in §10.3, went on to prove

μTμν  =  0 \partial_{\mu}T^{\mu\nu} \;=\; 0 (3.6.1)

in flat spacetime for a closed system. That is four equations, one for each value of ν\nu. Together they express local conservation of energy and of each component of momentum.

1.2 · The same statement on a curved manifold

Equation (3.6.1) is not yet a statement about a manifold. The trouble is that μ\partial_{\mu} of a tensor is not itself a tensor, as Chapter 3.3 §4 showed. We need a replacement that is, and it is this:

μTμν  =  0, \nabla_{\mu}T^{\mu\nu} \;=\; 0, (3.6.2)

and the reason for the replacement is not an appeal to elegance. It is Chapter 3.4 §5.3's locally inertial coordinates.

At any chosen point there is a chart in which gμν=ημνg_{\mu\nu}=\eta_{\mu\nu} and Γλμν=0\Gamma^{\lambda}{}_{\mu\nu}=0. In that chart \nabla and \partial coincide, so (3.6.2) reduces to (3.6.1), and the equivalence principle says the flat law must hold there. Now use the second half of the argument. Equation (3.6.2) is an equation between tensors, so by Chapter 2.4 §6, holding in one chart at a point means holding in every chart at that point. And the point was arbitrary.

⚠ The one place this rule is ambiguous, said out loud

The recipe just used was to take the flat-space law and replace \partial by \nabla. It is not always well defined. Suppose the flat law contains two derivatives. The two orderings μν\nabla_{\mu}\nabla_{\nu} and νμ\nabla_{\nu}\nabla_{\mu} agree in flat space, and on a curved manifold they differ by a term containing the Riemann tensor, exactly as Chapter 3.4 §2 computed. So the recipe leaves the coefficient of a possible curvature term undetermined, and two laws that look different can reduce to the same flat one.

Here there is no ambiguity, because (3.6.1) has exactly one derivative and there is nothing to reorder. The ambiguity does bite elsewhere, and the standard example is the wave equation for a field in curved spacetime. There the missing coefficient has to be fixed by some other principle. This book meets the question again when quantum fields are put on a curved background.

1.3 · Dust, and then a fluid with pressure

Dust. Take non-interacting particles all moving together, with rest-mass density ρ\rho measured in their own rest frame and four-velocity field uμu^{\mu}. In the rest frame the only thing present is energy density ρc2\rho c^{2}, so T00=ρc2T^{00}=\rho c^{2} and every other component vanishes. There is exactly one symmetric tensor built from ρ\rho and uu with that property. At rest uμ=(c,0,0,0)u^{\mu}=(c,0,0,0), so the tensor we want is this one:

Tμν  =  ρuμuν T^{\mu\nu} \;=\; \rho\,u^{\mu}u^{\nu} (3.6.3)

It gives T00=ρc2T^{00}=\rho c^{2} and nothing else, which is what was required, and it is built from tensors, so it is one. Its trace is worth recording while we are here. Using uu=c2u\cdot u=c^{2} from Chapter 2.5, TgμνTμν=ρc2T\equiv g_{\mu\nu}T^{\mu\nu}=\rho c^{2}.

Perfect fluid. Now allow isotropic pressure pp. In the rest frame the stress block is Tij=pδijT^{ij}=p\,\delta^{ij}. That is what the word "pressure" means: a flux of ii-momentum in the ii direction, the same in every direction. The energy density is still ρc2\rho c^{2}, so in the rest frame Tμν=diag(ρc2,p,p,p)T^{\mu\nu}=\mathrm{diag}(\rho c^{2},p,p,p). We want a tensor that reproduces that, and only two objects are available to build it out of, uμu^{\mu} and gμνg^{\mu\nu}:

Tμν  =  (ρ+pc2)uμuν    pgμν. T^{\mu\nu} \;=\; \Big(\rho + \frac{p}{c^{2}}\Big)u^{\mu}u^{\nu} \;-\; p\,g^{\mu\nu}. (3.6.4)

Let's check it component by component, in the rest frame, where gμν=ημνg^{\mu\nu}=\eta^{\mu\nu} and uμuνu^{\mu}u^{\nu} has only a 0000 entry, equal to c2c^{2}. The 0000 component is (ρ+p/c2)c2p1=ρc2+pp=ρc2(\rho+p/c^{2})c^{2}-p\cdot1=\rho c^{2}+p-p=\rho c^{2} ✓. The 1111 component is 0p(1)=+p0-p\cdot(-1)=+p ✓, the minus sign coming from η11=1\eta^{11}=-1, and likewise for 2222 and 3333. Off-diagonal entries vanish on both sides. ✓ Setting p=0p=0 returns (3.6.3).

Its trace is worth computing now, because §5 will need it. Use gμνuμuν=c2g_{\mu\nu}u^{\mu}u^{\nu}=c^{2} together with gμνgμν=δμμ=4g_{\mu\nu}g^{\mu\nu}=\delta^{\mu}{}_{\mu}=4 in four dimensions, and the trace comes out as

T  =  (ρ+pc2)c2    4p  =  ρc2    3p. T \;=\; \Big(\rho+\frac{p}{c^{2}}\Big)c^{2} \;-\; 4p \;=\; \rho c^{2} \;-\; 3p. (3.6.5)

1.4 · Three consequences worth naming before we start

(i) Pressure gravitates. TμνT^{\mu\nu} contains pp, so pressure appears in the source. This has no Newtonian counterpart at all. In 2Φ=4πGρ\nabla^{2}\Phi=4\pi G\rho the pressure of a gas is invisible. Section 5 will find exactly how it enters, and the answer is that the combination doing the gravitating is ρ+3p/c2\rho+3p/c^{2}.

(ii) A hot body weighs more than a cold one. Not because its atoms have gained mass, but because T00T^{00} counts every form of energy, including the kinetic energy of thermal motion and the energy stored in the fields binding it. Chapter 2.5 §9 showed the same arithmetic from the other side, in the mass deficit of a bound system.

(iii) A gas of radiation has a fixed pressure. Chapter 2.6 §10 built the electromagnetic energy–momentum tensor, and its closing summary recorded that the tensor is traceless, ημνTμν=0\eta_{\mu\nu}T^{\mu\nu}=0. So set (3.6.5) to zero and solve for the pressure, which gives

p  =  ρc23for radiation, p \;=\; \frac{\rho c^{2}}{3} \qquad\text{for radiation,} (3.6.6)

and that result is derived rather than quoted. The equation of state of light follows from the tracelessness of its stress tensor and from nothing else. Section 5 and Chapter 3.9 both use it, and Chapter 7.3 builds a symmetry principle on the tracelessness itself.

In plain terms 3.6.1

Mass cannot be what gravity responds to, and the reason is not subtle once the previous part is in hand. Mass is not additive, is not separately conserved, and picking out energy alone would mean picking out one entry of a four-part object, so two observers in relative motion would compute different geometries for the same physical arrangement. What is needed is an object packaging energy, momentum and the flow of both, and the earlier part built exactly that when it asked what is conserved because the laws are the same here as there.

That object has ten independent entries and every one of them sources gravity. Energy density is only the corner entry. The diagonal spatial entries are pressure, the off-diagonal ones are shear and momentum flow, and all are on the same footing. The statement that the whole thing has no divergence is the local claim that energy and momentum are never created or destroyed, only moved, and it carries over to a curved arena unchanged because it contains only one derivative and so leaves nothing to be ambiguous about.

One consequence deserves flagging before it arrives. Pressure gravitates. Squeeze a gas without adding any material to it and the source term grows, which nothing in Newtonian gravity would lead anyone to expect and which matters enormously for dying stars and for the expansion history of the universe.

2 · What the equation is allowed to look like

Before writing anything down, let's list what the answer has to satisfy. Each item below is a constraint with a stated origin rather than a matter of taste. Section 3 will find that together they leave almost nothing to choose.

(C1) It must be an equation between tensors. The authority is Chapter 3.2 §7 together with Chapter 2.4 §6. A tensor equation true in one chart is true in all of them, and an equation that is not tensorial holds only in some charts. On a manifold there are no preferred charts for it to hold in. So a non-tensorial gravitational law would be a law about coordinates rather than about the world.

(C2) The source is a symmetric rank-2 tensor, so the geometry side must be one too. That is the content of §1. Since TμνT^{\mu\nu} is symmetric with ten independent components, whatever equals it has the same type. This immediately rules out the Riemann tensor itself, which has four indices, and the Ricci scalar alone, which has none.

(C3) It must contain no derivatives of the metric beyond the second. There are two independent reasons for this, and they are worth keeping separate.

The first is the Newtonian limit, and it is the honest one. Newtonian gravity is 2Φ=4πGρ\nabla^{2}\Phi=4\pi G\rho, which is second order in the potential. Chapter 3.1 §6.5 derived g00=1+2Φ/c2g_{00}=1+2\Phi/c^{2}, so the potential lives inside the metric. An equation that reproduces Poisson must therefore be second order in gg. Third or fourth derivatives would produce extra terms that Newtonian gravity does not have and that experiment does not see.

The second is the structure of every field equation so far in this book. Chapter 1.2 §8 showed that an action built from a field and its first derivatives yields an equation of motion with second derivatives, and every entry of that chapter's table of actions has that form. Section 4 below finds that the gravitational action is an exception in one specific and interesting way. The exception costs exactly one boundary term.

⚑ A third reason, quoted and deferred

There is a deeper argument against higher derivatives that this book cannot yet make. A theory whose equations of motion contain more than two time derivatives generically carries extra degrees of freedom whose energy is unbounded below. The system can then lower its energy without limit by exciting them. This is Ostrogradsky's theorem, and we quote it rather than proving it, because the proof is Chapter 1.3's Hamiltonian machinery applied to a case that chapter did not cover.

It is flagged here because Chapter 7.1 leans on it. The reason one cannot repair quantum gravity by adding higher-derivative terms to the action is that the repair buys a ghost. For now, take (C3) on the Newtonian-limit argument alone, which is self-contained.

(C4) Its divergence must vanish identically. This is the constraint that does the work. Equation (3.6.2) says μTμν=0\nabla_{\mu}T^{\mu\nu}=0 for every physical matter configuration. Suppose the geometry side of the field equation had a divergence that vanished only sometimes. Then setting the two sides equal would impose a further condition on the matter, and that is a condition no experiment supports and no principle supplies.

So the geometrical side must have vanishing divergence as an identity. That means automatically, for every metric anybody could write down, whether or not it solves anything.

(C5) It must reduce to Newton where Newton works. Chapter 3.3 §8.3 established the correspondence limit for the motion of a body. The same limit for the field equation is §5's business, and it is what fixes the one remaining constant.

Notice the shape of the argument

Not one of (C1) to (C5) is an aesthetic preference. (C1) comes from the absence of preferred charts, (C2) from what matter is, (C3) from the requirement of reproducing an experimentally established law, (C4) from conservation of energy, (C5) from the same source as (C3).

The book's recurring claim is that most of the objects in it were forced rather than invented, and this chapter is the strongest case of it. Section 3 shows that (C1) to (C4) leave a two-parameter family of possible equations. Constraint (C5) then fixes one of the two parameters, and the remaining one is the subject of §6.

In plain terms 3.6.2

Before anything is designed there is a specification, and this one is short enough to hold in the head. The law must relate objects of the same kind on both sides, since otherwise it would hold only in some descriptions and be a statement about the description rather than the world. The matter side is a symmetric array with two slots, so the geometry side must be one too, which alone eliminates most of what the previous chapters built. The law must involve the shape of space and its first two rates of change and nothing beyond, because the old law it reproduces has two derivatives of the potential and the potential sits inside the shape.

Then the demanding one. Matter's bookkeeping object has no divergence, the statement that nothing is created or destroyed. If it is set equal to something built out of geometry, that something must have no divergence either, and not merely for the geometries that solve the equation. It must have none for every conceivable geometry, as an identity, or the law would quietly demand of matter something no experiment has seen matter obey.

Nothing on that list is a preference. Each traces back to the absence of preferred coordinates, or to what matter demonstrably is, or to the need to reproduce a law three centuries of astronomy confirmed. What remains after the list is applied is very nearly unique.

3 · The cornering — why GμνG_{\mu\nu} and nothing else

Here is where this section is going. We list every symmetric rank-2 tensor that can be built from the metric and its first two derivatives. We impose (C4), and we find that a one-parameter family survives, up to overall scale. That family is the Einstein tensor plus a multiple of the metric, and there is nothing else.

3.1 · The list of candidates

Start with what is available. By (C3) the ingredients are gμνg_{\mu\nu}, its inverse, and its derivatives up to second order. Chapter 3.4 §2 showed that the only tensor that can be built from the metric and its first two derivatives is the Riemann tensor. The first derivatives alone give the connection, which is not a tensor, and the second derivatives enter tensorially only in the combination (3.4.13). So the ingredients are gμνg_{\mu\nu} and RρσμνR^{\rho}{}_{\sigma\mu\nu}.

Now impose (C2), which asks for something symmetric with two indices. From RρσμνR^{\rho}{}_{\sigma\mu\nu} the only contraction producing a rank-2 tensor is the Ricci tensor. Chapter 3.4 §6.1 showed that it is essentially unique, and §6 of that chapter showed that it is symmetric. Contracting once more gives the Ricci scalar RR, which can be multiplied by gμνg_{\mu\nu} to make a symmetric rank-2 object. And gμνg_{\mu\nu} itself qualifies, since it needs no derivatives at all. The complete list is therefore three terms, and the most general candidate we can write down is

ARμν  +  BRgμν  +  Cgμν  =  κTμν, A\,R_{\mu\nu} \;+\; B\,R\,g_{\mu\nu} \;+\; C\,g_{\mu\nu} \;=\; \kappa\,T_{\mu\nu}, (3.6.7)

with AA, BB, CC and κ\kappa constants. One caveat about that list should be said out loud now. If terms quadratic in the curvature were allowed, such as R2gμνR^{2}g_{\mu\nu} or RμαRανR_{\mu\alpha}R^{\alpha}{}_{\nu}, the list would be longer. Those terms contain squares of second derivatives, and (C3) as it was stated asked for the geometry side to be linear in the second derivatives, so they are out. Section 3.3 says what happens if that restriction is relaxed.

3.2 · Imposing (C4), one term at a time

We want to know what (C4) costs, so take the divergence of the left-hand side of (3.6.7), one term at a time. The index is raised with the metric, and the metric passes through \nabla untouched by metric compatibility (Chapter 3.3 §7.1).

Third term. μ(Cgμν)=Cμgμν=0\nabla^{\mu}\big(C g_{\mu\nu}\big)=C\,\nabla^{\mu}g_{\mu\nu}=0, again by metric compatibility. So this term is divergence-free whatever CC is, and (C4) says nothing about it at all. Remember that. It is §6.

First term. Chapter 3.4 §7.2 contracted the second Bianchi identity twice and obtained (3.4.57):

μRμν  =  12νR. \nabla^{\mu}R_{\mu\nu} \;=\; \half\,\nabla_{\nu}R. (3.6.8)

So the first term contributes A12νRA\cdot\half\nabla_{\nu}R to the divergence.

Second term. μ(BRgμν)=BgμνμR=BνR\nabla^{\mu}\big(B R g_{\mu\nu}\big)=B\,g_{\mu\nu}\nabla^{\mu}R=B\,\nabla_{\nu}R, the metric having again passed through the derivative and then lowered the free index on μR\nabla^{\mu}R.

Add them. Putting the three contributions together, the total divergence of the left-hand side is

μ(ARμν+BRgμν+Cgμν)  =  (A2+B)νR. \nabla^{\mu}\Big(A R_{\mu\nu} + B R g_{\mu\nu} + C g_{\mu\nu}\Big) \;=\; \left(\frac{A}{2} + B\right)\nabla_{\nu}R. (3.6.9)

Now the crucial step. It is a logical one rather than an algebraic one. Constraint (C4) demands that (3.6.9) vanish identically, meaning for every metric and not merely for solutions.

But νR\nabla_{\nu}R is not identically zero. The sphere of Chapter 3.4 has constant RR, and a generic metric does not, and one has only to write down a metric with a position-dependent Ricci scalar to see it. So the only way to make the right-hand side vanish for every metric is for the coefficient in front of it to vanish:

B  =  A2. B \;=\; -\,\frac{A}{2}. (3.6.10)

That single condition ties BB to AA, so the next move is to put it back where it came from. Substituting the relation into (3.6.7) and factoring out the overall AA,

A(Rμν12Rgμν)  +  Cgμν  =  κTμν. A\left(R_{\mu\nu} - \half R\,g_{\mu\nu}\right) \;+\; C\,g_{\mu\nu} \;=\; \kappa\,T_{\mu\nu}. (3.6.11)

The bracket is exactly the Einstein tensor GμνG_{\mu\nu} of Chapter 3.4 (3.4.58). All that is left is to name the constants, so divide through by AA, absorb it into κ\kappa, and write λC/A\lambda\equiv C/A:

  Gμν  +  λgμν  =  κTμν.   \boxed{\;G_{\mu\nu} \;+\; \lambda\,g_{\mu\nu} \;=\; \kappa\,T_{\mu\nu}.\;} (3.6.12)

Two constants remain. Section 5 fixes κ\kappa by (C5). Section 6 examines λ\lambda, which (C5) constrains only weakly and which the argument of this section cannot exclude at all.

Recap — what went in, what came out

In: the requirement that the equation be tensorial and symmetric, that it use the metric and its first two derivatives linearly, and that its geometry side be divergence-free identically. Plus one input from three chapters ago: the twice-contracted Bianchi identity.

Out: the form of Einstein's equation, with two undetermined constants. Notice that the factor of 12\half in the Einstein tensor was not chosen and was not fitted to data. It is the number that makes (3.6.9) vanish. It comes from the 12\half in the twice-contracted Bianchi identity, which in turn came from adding two identical dummy terms in Chapter 3.4 §7.2.

What it cost: the assumption that no terms quadratic in the curvature appear. That is the one hypothesis of this section which the earlier chapters do not force, and §3.3 says what is known about relaxing it.

3.3 · What happens if the hypotheses are weakened

⚑ Lovelock's theorem, quoted

Section 3.1 assumed that the geometry side is linear in the second derivatives of the metric, which excluded terms like R2gμνR^{2}g_{\mu\nu}. That assumption can be dropped.

Lovelock's theorem (1971). In four spacetime dimensions, the only symmetric, rank-2, identically divergence-free tensor that can be constructed from the metric and its first and second derivatives is aGμν+bgμνa\,G_{\mu\nu}+b\,g_{\mu\nu} for constants aa and bb. We quote this and do not prove it. The proof is a classification argument of a kind this book does not develop.

Two of the hypotheses deserve emphasis, because they are where the theorem's content lies.

The dimension matters. In five or more dimensions there are additional terms that are also identically divergence-free and also give second-order equations. They are built from the square of the curvature in a particular combination, and the first of them is called the Gauss–Bonnet term. In four dimensions that combination contributes nothing to the equations of motion at all, which is a fact about four dimensions rather than a fact about gravity. Chapter 7.8 meets these terms where they are live.

And the derivative order matters. If third derivatives are permitted, the list grows again, and Ostrogradsky's theorem is the reason not to permit them.

Subject to the flagged theorem, then, (3.6.12) is not a possible law of gravity in four dimensions. It is the possible law, up to the two constants.

In plain terms 3.6.3

This is the section where the law of gravity gets cornered rather than proposed. Start by listing everything with the right shape that can be assembled out of the geometry: there are three items, the boiled-down curvature array, the single curvature number multiplied by the distance rule, and the distance rule by itself. Write the most general mixture of the three, with unknown coefficients, and then impose the requirement that its divergence vanish for every conceivable geometry.

Two of the three items have divergences that are automatically nothing, so they impose no condition. The other two have divergences proportional to the same quantity, which is not generally zero, so the only escape is for their coefficients to cancel. That single demand fixes the ratio of the first two coefficients, and the combination it produces is exactly the one built three chapters earlier from a differentiated identity. The famous factor of one half was never chosen; it is what makes the cancellation work.

What survives is one equation with two unfixed numbers in it, and the rest of the chapter is about those two numbers. A quoted theorem strengthens this considerably: even if the restriction to the mildest possible dependence on second rates of change is dropped, nothing new appears in four dimensions, though it does in five. The law is not one option among many. It is what is left.

4 · The same answer from an action

Chapter 1.2 §8.1 printed a table of eight actions and promised that each would be constructed in its own chapter. Line four read Gravity (Einstein–Hilbert), Chapter 3.6. This section constructs that action and varies it. The equation that comes out is (3.6.12), obtained a second time by a route that shares no step with §3.

Let's lay out the structure before starting, because this is the longest computation in Part III after Chapter 3.4 §2. The integrand is RgR\sqrt{-g}, and the thing being varied is the metric. Write R=gμνRμνR=g^{\mu\nu}R_{\mu\nu}, and the product rule splits the variation into exactly three pieces:

δ(g  gμνRμν)  =  (δg)Rpiece 1  +  g  Rμνδgμνpiece 2  +  g  gμνδRμνpiece 3. \delta\Big(\sqrt{-g}\;g^{\mu\nu}R_{\mu\nu}\Big) \;=\; \underbrace{\big(\delta\sqrt{-g}\big)\,R}_{\text{piece 1}} \;+\; \underbrace{\sqrt{-g}\;R_{\mu\nu}\,\delta g^{\mu\nu}}_{\text{piece 2}} \;+\; \underbrace{\sqrt{-g}\;g^{\mu\nu}\,\delta R_{\mu\nu}}_{\text{piece 3}}. (3.6.13)

Piece 2 needs no work at all. It is already in the form (something) times δgμν\delta g^{\mu\nu}, which is exactly what a variational calculation wants. Piece 1 is a determinant identity, and grind box A does it.

Piece 3 looks the worst and is the most interesting. It is a total derivative, so it contributes nothing to the equations of motion and everything to a discussion of boundaries, and it takes two grind boxes, B and C. The reasoning stays here. Only the algebra is folded away.

4.1 · Why RR, and the action

An action is a number, so what we need this time is a scalar rather than a tensor. Constraint (C3) still says at most two derivatives of the metric, so the question is what scalars can be built under that restriction.

A scalar has no free indices, so every index must be contracted. With one factor of the Riemann tensor there is essentially one way to do that. Contract the first and third indices to get the Ricci tensor, and Chapter 3.4 §6.1 showed that the other choices give zero or the same thing with a sign. Then contract the remaining two indices with the inverse metric to get RR.

With no factor of Riemann there is only a constant. With two factors of Riemann, meaning R2R^{2} or RμνRμνR_{\mu\nu}R^{\mu\nu} or RμνρσRμνρσR_{\mu\nu\rho\sigma}R^{\mu\nu\rho\sigma}, one has products of second derivatives, and by the same counting as §3.3 those are excluded by (C3) as stated. So the complete list of admissible scalars is a constant and RR, and the action is

S  =  SEH+Sm,SEH  =  α ⁣ ⁣(R2λ)g  d4x, S \;=\; S_{\text{EH}} + S_{\text{m}}, \qquad S_{\text{EH}} \;=\; \alpha\!\int\!\big(R - 2\lambda\big)\,\sqrt{-g}\;\dd^{4}x, (3.6.14)

with α\alpha and λ\lambda constants, and with SmS_{\text{m}} whatever action describes the matter. The factor 2-2 in front of λ\lambda is pure convention. It is chosen so that the symbol matches the λ\lambda of (3.6.12), and we shall see it do so.

The volume element gd4x\sqrt{-g}\,\dd^{4}x is Chapter 3.5 §6.2's. It is there because without it the integral would depend on the chart, which by (C1) is not allowed. Chapter 0.6 §8.3 said this six chapters ago and named this action as the reason.

Set λ=0\lambda=0 until §6, to keep the derivation clean. The constant is restored in three lines at the end of §4.4.

4.2 · Piece 1: the variation of the volume element

Piece 1 asks how the volume element responds to a change in the metric. The claim is

δg  =  12g  gμνδgμν. \delta\sqrt{-g} \;=\; -\,\half\,\sqrt{-g}\;g_{\mu\nu}\,\delta g^{\mu\nu}. (3.6.15)

Everything needed for it is already in hand. Chapter 3.5 §6.3 derived Jacobi's formula (3.5.47), which gives the change in a determinant produced by a change in its entries, and it derived the specialisation to the metric, (3.5.48). The only additional ingredient is the relation between varying gμνg_{\mu\nu} and varying gμνg^{\mu\nu}, which comes from differentiating the statement that the two are inverse. Grind box A does it in five lines.

Grind box A — δg\delta\sqrt{-g}, from Jacobi's formula

Line 1. The two variations are related. The metric and its inverse satisfy gμαgαν=δμνg^{\mu\alpha}g_{\alpha\nu}=\delta^{\mu}{}_{\nu}. The right-hand side is a constant array, so varying both sides gives

(δgμα)gαν  +  gμαδgαν  =  0. \big(\delta g^{\mu\alpha}\big)g_{\alpha\nu} \;+\; g^{\mu\alpha}\,\delta g_{\alpha\nu} \;=\; 0.

Multiply by gνβg^{\nu\beta} and sum over ν\nu, which turns gανgνβg_{\alpha\nu}g^{\nu\beta} into δβα\delta^{\beta}{}_{\alpha} in the first term:

δgμβ  =  gμαgνβδgαν. \delta g^{\mu\beta} \;=\; -\,g^{\mu\alpha}g^{\nu\beta}\,\delta g_{\alpha\nu}.

Line 2. Contract it. Multiply the previous display by gμβg_{\mu\beta} and sum. On the right, gμβgμα=δαβg_{\mu\beta}g^{\mu\alpha}=\delta^{\alpha}{}_{\beta} and then δαβgνβ=gνα\delta^{\alpha}{}_{\beta}g^{\nu\beta}=g^{\nu\alpha}, so

gμνδgμν  =  gμνδgμν. g_{\mu\nu}\,\delta g^{\mu\nu} \;=\; -\,g^{\mu\nu}\,\delta g_{\mu\nu}.

The two contracted variations are negatives of each other. This is the step that is easiest to get backwards, and getting it backwards costs a sign in the field equations.

Line 3. Jacobi's formula. Chapter 3.5's (3.5.48) applied to the metric reads δg=ggμνδgμν\delta g = g\,g^{\mu\nu}\delta g_{\mu\nu}, where g=detgμνg=\det g_{\mu\nu}. Using line 2 to swap which variation appears,

δg  =  g  gμνδgμν. \delta g \;=\; -\,g\;g_{\mu\nu}\,\delta g^{\mu\nu}.

Line 4. The square root. δg=δ(g)2g=δg2g\delta\sqrt{-g}=\dfrac{\delta(-g)}{2\sqrt{-g}}=\dfrac{-\delta g}{2\sqrt{-g}}. Substituting line 3,

δg  =  g  gμνδgμν2g. \delta\sqrt{-g} \;=\; \frac{g\;g_{\mu\nu}\,\delta g^{\mu\nu}}{2\sqrt{-g}}.

Line 5. Tidy the determinant. Since gg is negative, g=(g)2-g=(\sqrt{-g})^{2}, so g=(g)2g=-(\sqrt{-g})^{2} and therefore g/g=gg/\sqrt{-g}=-\sqrt{-g}. Hence

δg  =  12g  gμνδgμν, \delta\sqrt{-g} \;=\; -\,\half\,\sqrt{-g}\;g_{\mu\nu}\,\delta g^{\mu\nu},

which is (3.6.15). (This was checked symbolically against a direct differentiation of detg\sqrt{-\det g} for a four-dimensional metric family with off-diagonal entries. The two agreed exactly, as did the equivalent form δg=+12ggμνδgμν\delta\sqrt{-g}=+\half\sqrt{-g}\,g^{\mu\nu}\delta g_{\mu\nu}.)

Now combine that result with piece 2, which was already in the shape a variational calculation wants. Together the two of them produce this:

(δg)R  +  gRμνδgμν  =  g(Rμν12Rgμν)δgμν  =  g  Gμνδgμν. \big(\delta\sqrt{-g}\big)R \;+\; \sqrt{-g}\,R_{\mu\nu}\delta g^{\mu\nu} \;=\; \sqrt{-g}\left(R_{\mu\nu} - \half R\,g_{\mu\nu}\right)\delta g^{\mu\nu} \;=\; \sqrt{-g}\;G_{\mu\nu}\,\delta g^{\mu\nu}. (3.6.16)

The Einstein tensor has appeared without being sought. In §3 the factor of 12\half came from the contracted Bianchi identity. Here it comes from the derivative of a determinant. The two routes have nothing in common, and they produce the same coefficient, which is the strongest kind of check a derivation can have.

4.3 · Piece 3: the Palatini identity, derived

What remains is gμνδRμνg^{\mu\nu}\delta R_{\mu\nu}, which is piece 3. The strategy is to show that it is a total derivative, and it takes two steps.

Step one is the Palatini identity,

δRμν  =  λδΓλνμ    νδΓλλμ, \delta R_{\mu\nu} \;=\; \nabla_{\lambda}\,\delta\Gamma^{\lambda}{}_{\nu\mu} \;-\; \nabla_{\nu}\,\delta\Gamma^{\lambda}{}_{\lambda\mu}, (3.6.17)

and the reason this identity can be derived rather than quoted is a fact worth stating on its own. Although the connection is not a tensor, the difference of two connections is.

Chapter 3.3 §5.2 derived the transformation law (3.3.26) and pointed out that its offending inhomogeneous term depends only on the change of chart, not on the metric. So take the difference of two connections belonging to two nearby metrics. That term cancels, and what is left transforms as a (1,2)(1,2) tensor. A variation δΓ\delta\Gamma is exactly such a difference. Grind box B uses that fact together with Chapter 3.4 §5.3's locally inertial coordinates.

Grind box B — the Palatini identity, in four lines

Line 1. The definition. Contracting Chapter 3.4's (3.4.13) on the first and third indices gives the Ricci tensor written out in terms of the connection:

Rμν  =  λΓλνμ    νΓλλμ  +  ΓλλρΓρνμ    ΓλνρΓρλμ. R_{\mu\nu} \;=\; \partial_{\lambda}\Gamma^{\lambda}{}_{\nu\mu} \;-\; \partial_{\nu}\Gamma^{\lambda}{}_{\lambda\mu} \;+\; \Gamma^{\lambda}{}_{\lambda\rho}\Gamma^{\rho}{}_{\nu\mu} \;-\; \Gamma^{\lambda}{}_{\nu\rho}\Gamma^{\rho}{}_{\lambda\mu}.

Line 2. Vary it. The first two terms give λδΓλνμνδΓλλμ\partial_{\lambda}\delta\Gamma^{\lambda}{}_{\nu\mu}-\partial_{\nu}\delta\Gamma^{\lambda}{}_{\lambda\mu}. The quadratic terms give four contributions by the product rule, and every one of them carries an undifferentiated Γ\Gamma as a factor.

Line 3. Evaluate at a point in locally inertial coordinates. Chapter 3.4 §5.3 constructed, around any chosen point pp, a chart in which Γλμν(p)=0\Gamma^{\lambda}{}_{\mu\nu}(p)=0 while Γ(p)0\partial\Gamma(p)\neq0. In that chart, at that point, all four quadratic contributions vanish, and

δRμνp  =  λδΓλνμ    νδΓλλμ. \delta R_{\mu\nu}\big|_{p} \;=\; \partial_{\lambda}\,\delta\Gamma^{\lambda}{}_{\nu\mu} \;-\; \partial_{\nu}\,\delta\Gamma^{\lambda}{}_{\lambda\mu}.

Note what is not being claimed: Γ\partial\Gamma is not zero at pp, only Γ\Gamma is, which is exactly the point that made Chapter 3.4 §7.1's proof of the Bianchi identity work.

Line 4. Promote to covariant derivatives, and then to every chart. Since δΓ\delta\Gamma is a tensor, its covariant derivative is defined, and at pp in this chart every correction term in that covariant derivative carries a factor of Γ(p)=0\Gamma(p)=0. So the partial derivatives above may be replaced by covariant ones without changing anything, giving (3.6.17) at pp in this chart. But (3.6.17) is an equation between tensors, so by Chapter 2.4 §6 it holds in every chart. And the point pp was arbitrary. \blacksquare

(This was checked symbolically. For a four-dimensional metric family with off-diagonal entries, all sixteen components of the two sides of (3.6.17) agreed exactly at randomly chosen points.)

Step two contracts the Palatini identity with the inverse metric and recognises a divergence. The move that makes it work is metric compatibility. Since gμνg^{\mu\nu} passes through λ\nabla_{\lambda} untouched, it can be taken inside the derivative, and then the whole expression is the divergence of something. Grind box C does the index bookkeeping.

Grind box C — the contraction, and the vector whose divergence it is

Contract (3.6.17) with gμνg^{\mu\nu}:

gμνδRμν  =  gμνλδΓλνμ    gμννδΓλλμ. g^{\mu\nu}\,\delta R_{\mu\nu} \;=\; g^{\mu\nu}\,\nabla_{\lambda}\delta\Gamma^{\lambda}{}_{\nu\mu} \;-\; g^{\mu\nu}\,\nabla_{\nu}\delta\Gamma^{\lambda}{}_{\lambda\mu}.

Take the metric inside. By metric compatibility (Chapter 3.3 §7.1), λgμν=0\nabla_{\lambda}g^{\mu\nu}=0, so gμνλX=λ(gμνX)g^{\mu\nu}\nabla_{\lambda}X=\nabla_{\lambda}(g^{\mu\nu}X) for anything XX. Applying that to both terms:

gμνδRμν  =  λ(gμνδΓλνμ)    ν(gμνδΓλλμ). g^{\mu\nu}\,\delta R_{\mu\nu} \;=\; \nabla_{\lambda}\Big(g^{\mu\nu}\,\delta\Gamma^{\lambda}{}_{\nu\mu}\Big) \;-\; \nabla_{\nu}\Big(g^{\mu\nu}\,\delta\Gamma^{\lambda}{}_{\lambda\mu}\Big).

Give the two terms the same derivative index. The first differentiates with respect to λ\lambda and the second with respect to ν\nu. Both are summed, so both names are private. Rename νλ\nu\to\lambda in the second term, which forces its other ν\nu to become λ\lambda as well, and then rename its internal dummy λμ\lambda\to\mu and its μν\mu\to\nu to avoid a clash. The result is that both terms are λ\nabla_{\lambda} of something, and

gμνδRμν  =  λvλ,vλ    gμνδΓλμν    gλνδΓμμν, g^{\mu\nu}\,\delta R_{\mu\nu} \;=\; \nabla_{\lambda}v^{\lambda}, \qquad v^{\lambda} \;\equiv\; g^{\mu\nu}\,\delta\Gamma^{\lambda}{}_{\mu\nu} \;-\; g^{\lambda\nu}\,\delta\Gamma^{\mu}{}_{\mu\nu},

where we have used the symmetry of δΓ\delta\Gamma in its two lower indices to write δΓλνμ\delta\Gamma^{\lambda}{}_{\nu\mu} as δΓλμν\delta\Gamma^{\lambda}{}_{\mu\nu}. That symmetry is inherited from the symmetry of Γ\Gamma itself, Chapter 3.3 §7.2.

Then the volume factor. Chapter 3.5 §6.4's identity (3.5.51) says gλvλ=λ(gvλ)\sqrt{-g}\,\nabla_{\lambda}v^{\lambda}=\partial_{\lambda}\big(\sqrt{-g}\,v^{\lambda}\big), an ordinary partial derivative of an ordinary product. So piece 3 of (3.6.13) is exactly

g  gμνδRμν  =  λ(g  vλ). \sqrt{-g}\;g^{\mu\nu}\,\delta R_{\mu\nu} \;=\; \partial_{\lambda}\Big(\sqrt{-g}\;v^{\lambda}\Big).

(The whole identity was checked symbolically. For four-dimensional and three-dimensional metric families, δ(gR)\delta(\sqrt{-g}R) computed by direct differentiation agreed exactly, at random points, with gGμνδgμν+λ(gvλ)\sqrt{-g}\,G_{\mu\nu}\delta g^{\mu\nu}+\partial_{\lambda}(\sqrt{-g}\,v^{\lambda}) using the vλv^{\lambda} displayed above.)

4.4 · Assembling, and the field equations

All three pieces are now in hand, so put them together. Using (3.6.16) for pieces 1 and 2, and grind box C for piece 3,

δSEH  =  α ⁣ ⁣[  g  Gμνδgμν  +  λ(gvλ)]d4x. \delta S_{\text{EH}} \;=\; \alpha\!\int\!\left[\;\sqrt{-g}\;G_{\mu\nu}\,\delta g^{\mu\nu} \;+\; \partial_{\lambda}\Big(\sqrt{-g}\,v^{\lambda}\Big)\right]\dd^{4}x. (3.6.18)

The second term is an ordinary divergence, so by Chapter 3.5 §5's theorem it equals an integral over the boundary of the region. Adopt the same boundary condition Chapter 1.2 §3.3 adopted for the pendulum: the variation is required to vanish, along with its first derivatives, outside some bounded region. Then vλv^{\lambda} vanishes on the boundary and the term contributes nothing. Section 4.5 asks what that condition costs, because it is not free.

The matter side. Define the energy–momentum tensor by how the matter action responds to a change in the geometry:

δSm    12 ⁣ ⁣g  Tμνδgμν  d4x,that isTμν  =  2gδSmδgμν. \delta S_{\text{m}} \;\equiv\; \half\!\int\!\sqrt{-g}\;T_{\mu\nu}\,\delta g^{\mu\nu}\;\dd^{4}x, \qquad\text{that is}\qquad T_{\mu\nu} \;=\; \frac{2}{\sqrt{-g}}\,\frac{\delta S_{\text{m}}}{\delta g^{\mu\nu}}. (3.6.19)

This is a definition, so by itself it proves nothing. What makes it legitimate is that it reproduces the tensor Chapter 2.6 built by an entirely different argument. Worked example 1 checks exactly that, by feeding Chapter 2.6's electromagnetic Lagrangian (2.6.72) into (3.6.19) and recovering (2.6.80), sign included. The definition also delivers a symmetric tensor automatically, because δgμν\delta g^{\mu\nu} is symmetric. Chapter 2.6 had to arrange that by hand with an improvement term.

Stationarity. Demanding δS=δSEH+δSm=0\delta S=\delta S_{\text{EH}}+\delta S_{\text{m}}=0 for every allowed δgμν\delta g^{\mu\nu}, and applying the fundamental lemma of the calculus of variations (Chapter 1.2 §3.4) to strip off the arbitrary variation,

αGμν  +  12Tμν  =  0Gμν  =  12α  Tμν    κTμν. \alpha\,G_{\mu\nu} \;+\; \half\,T_{\mu\nu} \;=\; 0 \qquad\Longrightarrow\qquad G_{\mu\nu} \;=\; -\,\frac{1}{2\alpha}\;T_{\mu\nu} \;\equiv\; \kappa\,T_{\mu\nu}. (3.6.20)

That is (3.6.12) with λ=0\lambda=0, obtained a second time. Note that α\alpha is still undetermined. The action principle fixes the form of the law and cannot fix its coupling, exactly as Chapter 2.6 §9 could not fix μ0\mu_{0} from the shape of the electromagnetic Lagrangian alone.

Restoring λ\lambda, in three lines. Put the constant back in (3.6.14). Its variation involves only δg\delta\sqrt{-g}, so by (3.6.15),

δ ⁣ ⁣(2λ)g  d4x  =  2λ ⁣ ⁣(12ggμνδgμν)d4x  =  λ ⁣ ⁣g  gμνδgμνd4x, \delta\!\int\!\big(-2\lambda\big)\sqrt{-g}\;\dd^{4}x \;=\; -2\lambda\!\int\!\left(-\half\sqrt{-g}\,g_{\mu\nu}\delta g^{\mu\nu}\right)\dd^{4}x \;=\; \lambda\!\int\!\sqrt{-g}\;g_{\mu\nu}\,\delta g^{\mu\nu}\,\dd^{4}x, (3.6.21)

which adds αλgμν\alpha\lambda\,g_{\mu\nu} to the left of (3.6.20), and hence λgμν\lambda g_{\mu\nu} to the left of the field equation. That reproduces (3.6.12) in full, and it justifies the 2-2 in the action. The cosmological term is the constant that can be added to any Lagrangian. Section 6 takes that sentence seriously.

4.5 · The boundary term is not free

Two things were quietly discarded above and both deserve to be named.

Why the equations are second order at all. Constraint (C3) said the field equations must not contain derivatives of gg beyond the second, and Chapter 1.2 §8 established that an action built from a field and its first derivatives delivers exactly that. But RR contains second derivatives of the metric, so (3.6.14) is not of that form, and one would naively expect fourth-order equations.

It does not happen, and grind box C says why. Every second derivative of the metric in δ(gR)\delta(\sqrt{-g}R) sits inside λ(gvλ)\partial_{\lambda}(\sqrt{-g}v^{\lambda}), which is a total derivative and drops out of the equations of motion entirely. The Einstein–Hilbert action is second order in disguise, and that is a special property of RR rather than a general feature of curvature scalars. It is precisely why adding R2R^{2} to the action does give fourth-order equations.

What the discarded term costs. Because the second derivatives live in the boundary term, making δS\delta S vanish required more than fixing the metric on the boundary. It required fixing the metric's normal derivative there too, and that is one condition too many. The standard fix is to add to (3.6.14) a boundary integral whose variation cancels the unwanted piece. What is left is a variational principle in which only gμνg_{\mu\nu} itself is held fixed on the boundary.

⚑ Named and deferred: the Gibbons–Hawking–York term

The required addition is SGHYΩKhd3xS_{\text{GHY}}\propto\oint_{\partial\Omega}K\sqrt{\abs{h}}\,\dd^{3}x, where hh is the metric induced on the boundary and KK is the trace of its extrinsic curvature. That trace measures the rate at which the boundary's normal direction turns as one moves along it. We are quoting the form and not deriving it, because extrinsic curvature is machinery this book has deliberately avoided. Chapter 3.2 §1 forbade any reference to an ambient space, and extrinsic curvature is precisely what a surface looks like from outside.

Why it is worth flagging rather than ignoring. For the problems of Chapters 3.7 and 3.8 the boundary term contributes nothing and can be forgotten. It stops being ignorable exactly when the value of the action itself is the thing one wants, rather than the equations it produces. The outstanding case is the thermodynamics of black holes, where the numerical value of the gravitational action supplies the entropy. Chapter 3.9 §7 quotes that entropy, S=A/4S=A/4 in appropriate units, and says the derivation is beyond this book. This boundary term is one of the places where the derivation happens, and Chapter 7.9 returns to it.

In plain terms 3.6.4

The same law arrives again by a road sharing no step with the first. Attach a single number to each possible shape of spacetime, namely the total curvature added up with the correct volume weighting, and ask which shape makes that number stationary. There is essentially one number available to attach, because insisting on no more than two rates of change leaves exactly one scalar and a constant.

Varying it splits into three parts. The first asks how the volume weighting responds, and the rule for differentiating a determinant answers it. The second needs no work. The third looks worst and is a total derivative, contributing nothing inside the region and everything on its edge. Combining the first two, the object cornered in the previous section appears unbidden, and the notorious factor of one half arrives from the derivative of a determinant rather than from a differentiated identity. Two unrelated routes, one coefficient.

The discarded edge term is not tidy-up. It is where the second rates of change hide, which is why the equations come out second order though the number attached is not. Setting it aside also demands more of the boundary than one is entitled to, and repairing that costs an extra term whose value matters in one place above all: the entropy of a black hole is what the gravitational number evaluates to, and that story waits for the last part.

5 · The constant, fixed by demanding that apples fall

This is the section where the whole construction is put on trial. Everything so far has been structural. The equation has the form it has because of conservation, because of symmetry, and because of the number of derivatives allowed. None of that touches the world. Here it does.

The route has five steps and each one is small, so here they are before we start. We take a weak, static gravitational field and slow-moving matter. We rearrange the field equation into a form where the Ricci tensor stands alone. We evaluate its 0000 component on the left from the metric, and then the same component on the right from the matter. Finally we compare the result with Poisson's equation. Out comes κ=8πG/c4\kappa=8\pi G/c^{4}.

⚠ The sign of g00g_{00}, stated before it is used

Everything below rests on the weak-field metric component. In this book's signature (+,,,)(+,-,-,-) it is

g00    1  +  2Φc2, g_{00} \;\approx\; 1 \;+\; \frac{2\Phi}{c^{2}},

with a plus sign, and with Φ\Phi the Newtonian potential, which is negative near a mass. This is not a convention adopted here for convenience. Chapter 3.1 §6.5 derived it, as equation (3.1.39), from the gravitational redshift of a signal climbing inside an accelerating cabin, using special relativity and the equivalence principle and no general relativity whatever.

The sense of it can be checked directly. A clock lower in a well runs slow, so the coefficient relating proper time to coordinate time must be smaller there, and Φ\Phi is more negative there. The signs agree.

Books using (,+,+,+)(-,+,+,+) write g00=(1+2Φ/c2)g_{00}=-\big(1+2\Phi/c^{2}\big). That is the same physics in the other signature, and copying it into this chapter would flip the sign of the coupling constant derived below. This is the second of the two sign traps announced at the start of the chapter. The first was the Riemann convention of Chapter 3.4, and the third is in §6.

5.1 · Step 1 — trace-reverse the field equation

The field equation (3.6.20) has the Ricci scalar buried inside GμνG_{\mu\nu}, and that is inconvenient. We want the scalar out where we can see it, so contract the equation with gμνg^{\mu\nu}. Using gμνRμν=Rg^{\mu\nu}R_{\mu\nu}=R and gμνgμν=4g^{\mu\nu}g_{\mu\nu}=4,

gμνGμν  =  R    12R4  =  R,soR  =  κT, g^{\mu\nu}G_{\mu\nu} \;=\; R \;-\; \half R\cdot 4 \;=\; -\,R, \qquad\text{so}\qquad -R \;=\; \kappa\,T, (3.6.22)

with TgμνTμνT\equiv g^{\mu\nu}T_{\mu\nu}. That is an expression for RR in terms of the matter, so we can now eliminate the scalar entirely. Substitute R=κTR=-\kappa T back into Rμν=Gμν+12RgμνR_{\mu\nu}=G_{\mu\nu}+\half R\,g_{\mu\nu}, and the result is

  Rμν  =  κ(Tμν12Tgμν).   \boxed{\;R_{\mu\nu} \;=\; \kappa\left(T_{\mu\nu} - \half\,T\,g_{\mu\nu}\right).\;} (3.6.23)

This is the trace-reversed form, and it is exactly equivalent to (3.6.20). Contract it and the previous step runs backwards. It is the form to use whenever the Ricci tensor is easier to compute than the Einstein tensor, which is most of the time, and Chapter 3.7 uses it from its first line.

One consequence comes free and matters a great deal. In vacuum, where Tμν=0T_{\mu\nu}=0, the field equations reduce to Rμν=0R_{\mu\nu}=0. Chapter 3.4 §6.3 already showed that this does not mean flat, because the Weyl part of the curvature survives. Chapter 3.7 solves Rμν=0R_{\mu\nu}=0.

5.2 · Step 2 — the assumptions, each named as it is made

Assumption 1, weak field. We write gμν=ημν+hμνg_{\mu\nu}=\eta_{\mu\nu}+h_{\mu\nu} with every hμν1\abs{h_{\mu\nu}}\ll1, and we keep only terms linear in hh. In particular h00=2Φ/c2h_{00}=2\Phi/c^{2}.

Assumption 2, static. Nothing depends on time, so 0gμν=0\partial_{0}g_{\mu\nu}=0.

Assumption 3, slow matter. The matter's four-velocity is uμ(c,0,0,0)u^{\mu}\approx(c,0,0,0), so in (3.6.4) we keep the leading term only. Pressure is taken to be small compared with rest-energy density, pρc2p\ll\rho c^{2}. Section 5.5 removes this one.

5.3 · Step 3 — the left-hand side, R00R_{00}

We want the 0000 component of the left-hand side, so write out the Ricci tensor in terms of the connection, exactly as grind box B did, and then set μ=ν=0\mu=\nu=0:

R00  =  λΓλ00    0Γλλ0  +  ΓλλρΓρ00    Γλ0ρΓρλ0. R_{00} \;=\; \partial_{\lambda}\Gamma^{\lambda}{}_{00} \;-\; \partial_{0}\Gamma^{\lambda}{}_{\lambda0} \;+\; \Gamma^{\lambda}{}_{\lambda\rho}\Gamma^{\rho}{}_{00} \;-\; \Gamma^{\lambda}{}_{0\rho}\Gamma^{\rho}{}_{\lambda0}. (3.6.24)

Now take those four terms one at a time. Three of them turn out to be nothing.

The two quadratic terms are second order. Each Γ\Gamma is built from one derivative of the metric, so by Assumption 1 each one is first order in hh. A product of two of them is second order, and it is dropped.

The second term vanishes. It carries 0\partial_{0}, and by Assumption 2 nothing depends on x0x^{0}.

The first term is the whole answer. Split the sum over λ\lambda into its time and space parts. The time part is 0Γ000\partial_{0}\Gamma^{0}{}_{00}, which vanishes by Assumption 2 again. What is left is the spatial sum, and Chapter 3.3 §8.3 has already computed the connection coefficient that sum needs. It is equation (3.3.71), obtained there from exactly these three assumptions:

Γi00  =  iΦc2  +  O ⁣(Φ2c4). \Gamma^{i}{}_{00} \;=\; \frac{\partial_{i}\Phi}{c^{2}} \;+\; O\!\left(\frac{\Phi^{2}}{c^{4}}\right). (3.6.25)

Now put that coefficient into the surviving term and carry out the remaining differentiation, which finishes this side of the equation:

R00  =  iΓi00  =  iiΦc2  =  2Φc2, R_{00} \;=\; \partial_{i}\Gamma^{i}{}_{00} \;=\; \frac{\partial_{i}\partial_{i}\Phi}{c^{2}} \;=\; \frac{\nabla^{2}\Phi}{c^{2}}, (3.6.26)

where the last step recognises the Laplacian of Chapter 0.7 §7.5. At this order indices are raised and lowered with δij\delta_{ij}, so the distinction between i\partial_{i} and i\partial^{i} costs only a sign that appears twice and cancels.

Familiar ground — this number has been obtained once already, by another route

Equation (3.6.26) is Chapter 3.4's (3.4.33). That chapter got it by a completely different argument. It derived the geodesic deviation equation, set that beside Chapter 3.1's Newtonian tidal equation, read off c2Ri0j0=ijΦc^{2}R^{i}{}_{0j0}=\partial_{i}\partial_{j}\Phi, and then took the trace. Here the same result comes straight from the definition of the Ricci tensor and the Christoffel symbols of Chapter 3.3.

The agreement is not a coincidence, and it is not circular either. The two routes share only the weak-field metric component, which Chapter 3.1 obtained from the redshift with no general relativity in it at all.

(This was confirmed symbolically as well. Computing R00R_{00} from the full metric g00=1+2Φ/c2g_{00}=1+2\Phi/c^{2}, gij=(1+2sΦ/c2)δijg_{ij}=-(1+2s\Phi/c^{2})\delta_{ij} and expanding to first order gives 2Φ/c2\nabla^{2}\Phi/c^{2} for every value of ss. So the spatial part of the metric, which none of these arguments has pinned down, does not affect this component at this order.)

5.4 · Step 4 — the right-hand side, and the comparison

Now for the right-hand side. Use the trace-reversed form (3.6.23) with μ=ν=0\mu=\nu=0. By Assumption 3 the matter is dust at rest, so from (3.6.3) and uμ=gμνuν(c,0,0,0)u_{\mu}=g_{\mu\nu}u^{\nu}\approx(c,0,0,0),

T00  =  ρu0u0  =  ρc2,T  =  ρc2. T_{00} \;=\; \rho\,u_{0}u_{0} \;=\; \rho c^{2}, \qquad T \;=\; \rho c^{2}. (3.6.27)

Those are the two ingredients the right-hand side asks for, so assemble the bracket that appears in (3.6.23):

T0012Tg00  =  ρc2    12ρc2(1+2Φc2)  =  12ρc2    ρΦ    12ρc2, T_{00} - \half T g_{00} \;=\; \rho c^{2} \;-\; \half\rho c^{2}\left(1+\frac{2\Phi}{c^{2}}\right) \;=\; \half\,\rho c^{2} \;-\; \rho\,\Phi \;\approx\; \half\,\rho c^{2}, (3.6.28)

where the last step drops ρΦ\rho\Phi, because it is smaller than 12ρc2\half\rho c^{2} by the factor 2Φ/c22\Phi/c^{2}, and Assumption 1 declared that negligible.

Let's look at what that line is actually saying. Note the arithmetic: 112=121-\half=\half. The full energy density enters, the trace-reversal removes half of it, and that surviving factor of one half is exactly what turns the 4π4\pi of Poisson's equation into the 8π8\pi of the final answer.

Both sides of the 0000 equation are now in hand, so set them equal. That means (3.6.26) on the left and κ\kappa times (3.6.28) on the right:

2Φc2  =  κ12ρc22Φ  =  12κρc4. \frac{\nabla^{2}\Phi}{c^{2}} \;=\; \kappa\cdot\half\,\rho\,c^{2} \qquad\Longrightarrow\qquad \nabla^{2}\Phi \;=\; \half\,\kappa\,\rho\,c^{4}. (3.6.29)

That is a Poisson equation with an undetermined coefficient in it, so compare it against the real Poisson equation, which Chapter 0.7 §7.5 derived from Newton's inverse-square law and the divergence theorem:

2Φ  =  4πGρ. \nabla^{2}\Phi \;=\; 4\pi G\rho. (3.6.30)

The two agree for every density only if the coefficients of ρ\rho agree, which requires 12κc4=4πG\half\kappa c^{4}=4\pi G. Solving that for κ\kappa,

  κ  =  8πGc4   \boxed{\;\kappa \;=\; \frac{8\pi G}{c^{4}}\;} (3.6.31)

Put that value of κ\kappa back into the boxed result of §3, and there, at last, are the field equations of general relativity:

  Gμν  +  λgμν  =  8πGc4  Tμν.   \boxed{\;G_{\mu\nu} \;+\; \lambda\,g_{\mu\nu} \;=\; \frac{8\pi G}{c^{4}}\;T_{\mu\nu}.\;} (3.6.32)

And by (3.6.20), α=1/(2κ)=c4/(16πG)\alpha=-1/(2\kappa)=-c^{4}/(16\pi G), so the action is fixed too:

SEH  =  c416πG ⁣(R2λ)g  d4x. S_{\text{EH}} \;=\; -\,\frac{c^{4}}{16\pi G}\int\!\big(R-2\lambda\big)\sqrt{-g}\;\dd^{4}x. (3.6.33)
⚑ The sign in front of the action, and Chapter 1.2's table

Chapter 1.2 §8.1 previewed this action as S=12κRgd4xS=\frac{1}{2\kappa}\int R\sqrt{-g}\,\dd^{4}x, quoted forward and flagged there as quoted. Equation (3.6.33) has the opposite overall sign, and the discrepancy is a convention rather than a disagreement.

Here is why. Replace gμνg_{\mu\nu} everywhere by gμν-g_{\mu\nu}, which is what changing signature does. The Christoffel symbols are unchanged, because the formula (3.3.50) contains one inverse metric and one derivative of the metric, and the two sign changes cancel. Hence RρσμνR^{\rho}{}_{\sigma\mu\nu} and RμνR_{\mu\nu} are unchanged as well.

But R=gμνRμνR=g^{\mu\nu}R_{\mu\nu} flips sign, since the inverse metric does. In Gμν=Rμν12RgμνG_{\mu\nu}=R_{\mu\nu}-\half R g_{\mu\nu} the second term picks up two sign flips, which cancel, so GμνG_{\mu\nu} is unchanged. In gR\sqrt{-g}R there is only one flip, so that quantity changes sign. The field equations are signature-independent. The Lagrangian producing them is not.

Landau and Lifshitz, who use (+,,,)(+,-,-,-) as this book does, write the gravitational action with the minus sign, in agreement with (3.6.33). Books using (,+,+,+)(-,+,+,+) write it with the plus, and Chapter 1.2's table quoted that form. So check the signature before importing a gravitational Lagrangian from anywhere, exactly as Chapter 3.4 said to check both Riemann conventions before importing a curvature formula.

5.5 · What the number means, and one generalisation

The size of κ\kappa. Numerically κ=8πG/c4=2.08×1043 m1kg1s2\kappa=8\pi G/c^{4}=2.08\times10^{-43}\ \mathrm{m^{-1}\,kg^{-1}\,s^{2}}. The reciprocal is more telling: 1/κ=c4/8πG=4.8×1042 N1/\kappa=c^{4}/8\pi G=4.8\times10^{42}\ \mathrm{N}.

Read (3.6.32) as (curvature) =κ×=\kappa\times(stress), and 1/κ1/\kappa becomes the stiffness of spacetime, meaning the stress needed to produce unit curvature. It is about 104210^{42} newtons. That is why an object as massive as the Earth bends spacetime by only the 1.7×1023 m21.7\times10^{-23}\ \mathrm{m^{-2}} that Chapter 3.4 §4.5 computed, and why gravity looked for three centuries like the weakest thing in physics rather than like geometry.

Restoring the pressure. Assumption 3 dropped pp, and we can now afford to put it back. Use the perfect fluid (3.6.4) instead of dust, and take its trace from (3.6.5), which is T=ρc23pT=\rho c^{2}-3p. Then, still at rest and to leading order, the bracket becomes

T0012Tg00  =  ρc2    12(ρc23p)  =  12(ρc2+3p), T_{00} - \half T g_{00} \;=\; \rho c^{2} \;-\; \half\big(\rho c^{2}-3p\big) \;=\; \half\Big(\rho c^{2} + 3p\Big), (3.6.34)

Nothing else in the calculation changes, so run steps 3 and 4 again with this bracket in place of 12ρc2\half\rho c^{2}. What comes out is

  2Φ  =  4πG(ρ+3pc2).   \boxed{\;\nabla^{2}\Phi \;=\; 4\pi G\left(\rho + \frac{3p}{c^{2}}\right).\;} (3.6.35)

This is the promise of §1.4 kept. Pressure gravitates, and it does so three times over. Squeeze a gas without adding anything to it and the source of its gravitational field grows.

The effect is invisible in ordinary matter. There p/c2p/c^{2} is smaller than ρ\rho by the square of the ratio of thermal speeds to cc, which for the Sun's centre is about 10610^{-6}. But it is not invisible everywhere. In a neutron star the pressure term is a sizeable fraction of the source, and it works against the star. Adding pressure to resist collapse also adds to the gravity doing the collapsing. That is part of the reason a sufficiently massive star has nothing left that can hold it up.

And for radiation, where (3.6.6) gives p=ρc2/3p=\rho c^{2}/3, the bracket becomes 2ρ2\rho. A gas of light gravitates twice as strongly as the same energy density in cold dust. Chapter 3.9 needs that.

The same expression also contains the seed of §6, and it takes one line to see. Suppose a substance could have p<ρc2/3p<-\rho c^{2}/3. Then the bracket would be negative, 2Φ\nabla^{2}\Phi would have the wrong sign, and the gravity of that substance would push rather than pull. Nothing in (3.6.35) forbids it.

0.00
w = 0.00 → source = ρ(1 + 3w) = 1.000 ρ (cold dust — pressureless matter)
attractive, 1.00× the Newtonian value for the same energy density
How much of the field equations survives at everyday field strengths, and what survives that Newton did not have. The curve is the source of (3.6.35) in units of ρ\rho, namely 1+3w1+3w with w=p/ρc2w=p/\rho c^{2}; Newtonian gravity has only the flat line at height 11, marked grey. Everything else on the plot is a relativistic correction that does not vanish as cc\to\infty, because p/c2p/c^{2} is what enters and pp itself can be large. The ball on the right is a small uniform lump of the fluid, with arrows showing the acceleration of test particles released around it; their length is 1+3w\abs{1+3w} and their direction is set by its sign. Cold dust and radiation both pull, radiation twice as hard for the same energy density. At w=13w=-\tfrac13 the arrows vanish altogether: such a fluid has energy, and does not gravitate. Below that, in the shaded region, the arrows point outward and the substance repels. Nothing derived in this chapter excludes that region, which is the entire content of §6 — and the observed universe sits at w1w\approx-1, the far left of the plot.
In plain terms 3.6.5

Everything so far has been shape without scale, since the constant tying geometry to matter was carried along unfilled, and filling it in is where the construction becomes a theory of the world rather than bookkeeping. Take a weak, unchanging field and slow matter, and rearrange the law so the boiled-down curvature stands alone on the left. Its timekeeping entry works out to be the ordinary second-derivative operator applied to the potential, divided by the square of the speed of light, which the curvature chapter had already reached by a different argument about drifting dust.

The matter side gives the energy density less half its own trace, which for slow cold matter is half the energy density. Comparing with the three-centuries-old equation relating potential to density fixes it at eight pi times the gravitational constant over the fourth power of the speed of light. Its reciprocal is more eloquent: about ten to the forty-two newtons, the stress needed to bend spacetime usefully, and the reason gravity looks like the feeblest thing in physics.

One term survives that Newton had no way to see. Because the source is the whole bookkeeping object and not merely its topmost entry, pressure appears in it, tripled. Compressing a gas increases the gravity it makes, light pulls twice as hard as cold matter of the same energy, and a substance with sufficiently negative pressure would push rather than pull.

6 · The one term the argument cannot exclude

Now return to λ\lambda. Section 3 found it and could say nothing about it, because constraint (C4) was silent: μgμν=0\nabla^{\mu}g_{\mu\nu}=0 automatically. Section 4 found it again from the other direction, as the constant that may be added to any Lagrangian. This section asks what it is.

6.1 · It is a substance with negative pressure

The way to find out what a term means is usually to move it to the other side and read it as matter. So move it from the geometry side of (3.6.32), which costs a sign:

Gμν  =  κTμν    λgμν  =  κ(Tμν    λκgμν). G_{\mu\nu} \;=\; \kappa\,T_{\mu\nu} \;-\; \lambda\,g_{\mu\nu} \;=\; \kappa\left(T_{\mu\nu} \;-\; \frac{\lambda}{\kappa}\,g_{\mu\nu}\right). (3.6.36)

Now read that bracket as a total energy–momentum tensor. Whatever λ\lambda turns out to be, its contribution to the source is

Tμνvac  =  λκ  gμν. T^{\text{vac}}_{\mu\nu} \;=\; -\,\frac{\lambda}{\kappa}\;g_{\mu\nu}. (3.6.37)

We want to know what substance would produce that, so compare it with the perfect fluid (3.6.4), with both indices lowered: Tμν=(ρ+p/c2)uμuνpgμνT_{\mu\nu}=(\rho+p/c^{2})u_{\mu}u_{\nu}-p\,g_{\mu\nu}. For (3.6.37) to have this form, two conditions must hold, and we take them one at a time. The uμuνu_{\mu}u_{\nu} term is absent from (3.6.37), so its coefficient must vanish:

ρ+pc2  =  0p  =  ρc2. \rho + \frac{p}{c^{2}} \;=\; 0 \qquad\Longrightarrow\qquad p \;=\; -\,\rho\,c^{2}. (3.6.38)

That is the first condition. Matching the remaining term, pgμν=(λ/κ)gμν-p\,g_{\mu\nu}=-(\lambda/\kappa)g_{\mu\nu}, gives p=λ/κp=\lambda/\kappa, and then (3.6.38) converts that into a density:

ρvac  =  λκc2  =  λc28πG. \rho_{\text{vac}} \;=\; -\,\frac{\lambda}{\kappa\,c^{2}} \;=\; -\,\frac{\lambda\,c^{2}}{8\pi G}. (3.6.39)

Three things follow immediately, and not one of them was put in by hand.

(i) The equation of state is fixed, not chosen. A cosmological term is a fluid with p=ρc2p=-\rho c^{2}, which is w=1w=-1 in the figure's notation. It is the far-left point of that plot. By (3.6.35) its source term is ρ+3p/c2=2ρ\rho+3p/c^{2}=-2\rho. So with a positive energy density it repels, and it repels twice as hard as the same energy density in dust attracts.

(ii) It is the only fluid that looks the same to everybody. Equation (3.6.37) is proportional to gμνg_{\mu\nu}, and the metric is the one tensor whose components are the same in every local inertial frame. Every other fluid singles out a rest frame, through uμu^{\mu}. This one does not. That is exactly what one would demand of the energy of empty space, since there is no such thing as moving relative to the vacuum.

(iii) Its density is constant. That is not an assumption. It follows because λ\lambda is a constant and κ\kappa is a constant. Expand the universe and ordinary matter thins out while this does not, which is why a term negligible in the early universe can come to dominate the late one. Chapter 3.9 works that out.

⚠ The third sign trap, and the symbol Λ\Lambda

Equation (3.6.39) carries a minus sign. With the term written as +λgμν+\lambda g_{\mu\nu} on the left, a positive vacuum energy density corresponds to a negative λ\lambda. That is a consequence of our signature, and it would be a permanent nuisance, so we do what everyone does and absorb it into the definition. Write Λλ\Lambda\equiv-\lambda, so that

Gμν    Λgμν  =  8πGc4Tμν,Λ  =  8πGc2ρvac. G_{\mu\nu} \;-\; \Lambda\,g_{\mu\nu} \;=\; \frac{8\pi G}{c^{4}}\,T_{\mu\nu}, \qquad \Lambda \;=\; \frac{8\pi G}{c^{2}}\,\rho_{\text{vac}}.

With this definition Λ>0\Lambda>0 means positive vacuum energy and gravitational repulsion, and Λ\Lambda has dimensions of one over length squared, as a curvature should.

Books using (,+,+,+)(-,+,+,+) write Gμν+Λgμν=κTμνG_{\mu\nu}+\Lambda g_{\mu\nu}=\kappa T_{\mu\nu} and mean the same Λ\Lambda. The flip is the one traced in §5.4. Under gμνgμνg_{\mu\nu}\to-g_{\mu\nu} the tensor GμνG_{\mu\nu} is unchanged, TμνT_{\mu\nu} is unchanged, and gμνg_{\mu\nu} is not. Since Λ\Lambda is a measured number rather than a convention, it is the placement of the sign that has to move. Chapter 3.4's preview wrote the term as +Λgμν+\Lambda g^{\mu\nu} "for constants Λ\Lambda and κ\kappa". That is this same term with the naming still open, and it is settled here.

6.2 · Why nothing forbids it

The quickest way to see that nothing rules the term out is to go back through the chapter and check every constraint in §2 against Λgμν\Lambda g_{\mu\nu} in turn.

ConstraintDoes Λgμν\Lambda g_{\mu\nu} satisfy it?
(C1) tensorialyes, since gμνg_{\mu\nu} is a tensor
(C2) symmetric, rank 2yes
(C3) at most two derivatives of ggyes, it has none
(C4) identically divergence-freeyes, by metric compatibility, μgμν=0\nabla^{\mu}g_{\mu\nu}=0
(C5) Newtonian limityes, provided Λ\Lambda is small enough, as below

Every row passes. And §4 makes the point from the other side. In the action, the cosmological term is a constant added to the Lagrangian, and nothing forbids adding a constant to a Lagrangian.

In ordinary mechanics, adding a constant to the Lagrangian changes nothing, because the extra contribution to the action is itself a fixed number, and a fixed number has no variation. Here it is not fixed. The constant is multiplied by g\sqrt{-g}, which depends on the very thing being varied. So gravity, alone among the theories in this book, notices a constant in its Lagrangian. What it notices is this term.

What (C5) says is a bound rather than a prohibition. Redo §5 keeping Λ\Lambda. The extra term contributes Λg00Λ-\Lambda g_{00}\approx-\Lambda to the left of the field equation, so (3.6.29) becomes 2Φ=4πGρΛc2\nabla^{2}\Phi=4\pi G\rho-\Lambda c^{2}. For that to be indistinguishable from Poisson's equation in the Solar System, Λc2\Lambda c^{2} must be tiny compared with 4πGρ4\pi G\rho there, which it is, by an enormous margin. The constant is not excluded. It is bounded.

6.3 · What it is measured to be, and the problem that leaves

⚑ Two numbers, quoted as observations

Observations of distant supernovae and of the cosmic microwave background give Λ1.1×1052 m2\Lambda\approx1.1\times10^{-52}\ \mathrm{m^{-2}}, corresponding by (3.6.39) to ρvac6×1027 kgm3\rho_{\text{vac}}\approx6\times10^{-27}\ \mathrm{kg\,m^{-3}}. That is about four hydrogen atoms per cubic metre, and roughly 70%70\% of the total energy density of the present universe. These are quoted as measurements. Chapter 3.9 derives what a universe with this term does. It does not derive the number.

Two things about that value are worth stating now. It is fantastically small in any natural unit. Expressed as a length, Λ1/2\Lambda^{-1/2} is about 102610^{26} metres, which is comparable to the size of the observable universe and about 106110^{61} times the Planck length.

And it is not zero, which is worse. A vanishing constant might be explained by a symmetry forbidding it. A tiny non-zero one has to be explained by something that gets the number right. Nothing in this book explains it, and nothing anywhere else does either. Chapter 7.9 is where the failure is accounted for honestly.

One historical note, because it is usually told wrongly. Einstein introduced this term in 1917 to permit a static universe, since with Λ=0\Lambda=0 the equations have no static solution containing matter. The balance he found is unstable, as Chapter 3.9 §3 shows by integrating the equations. Once the expansion of the universe was established, the motivation evaporated.

What did not evaporate is the term. It was never inserted. It was always allowed. Removing it requires an extra assumption, and that assumption turned out to be false.

In plain terms 3.6.6

One extra piece slipped through the cornering untouched, namely the distance rule itself multiplied by a constant. It slipped through because the requirement doing all the work, that the geometry side have no divergence, is automatically satisfied by the distance rule. Seen from the action, the same piece is a constant added to the quantity being extremised, and no principle in this book forbids that.

Moved across to the matter side, the term describes a substance, and its properties are forced rather than assumed. Its pressure must be exactly the negative of its energy density, which is the unique choice that looks identical to every observer, as the energy of empty space ought to. Because the source of gravity contains three times the pressure, such a substance has a negative source and therefore pushes rather than pulls. And its density cannot dilute as the universe grows, since it is built from constants, so a contribution negligible early on can come to dominate later.

The observed value is small beyond ordinary description, and small is worse than zero. A quantity forced to vanish can be explained by a principle forbidding it; a quantity very small and not zero needs an explanation producing the actual number, and none exists. This is the single most embarrassing number in physics and the last part of the book returns to it without pretending to fix it.

7 · Ten equations, four identities, and why that is exactly right

The field equations are written. Before spending them, let's count them, because the count reveals something the equations do not say out loud.

7.1 · The count

Unknowns. The metric gμνg_{\mu\nu} is symmetric with two indices in four dimensions, so it has 45/2=104\cdot5/2=10 independent components. That is ten functions of four variables.

Equations. Both sides of (3.6.32) are symmetric rank-2 tensors, so it is also ten equations. Ten equations for ten unknowns looks like a well-posed problem, and looks determined.

But four of the ten are not independent. Chapter 3.4 §7.3 proved μGμν=0\nabla^{\mu}G_{\mu\nu}=0 identically, and μ(Λgμν)=0\nabla^{\mu}(\Lambda g_{\mu\nu})=0 as well, so the divergence of the left-hand side of (3.6.32) is zero whatever the metric. That is four differential relations among the ten equations, one for each value of ν\nu. Only six of the ten carry independent information about how the metric evolves.

And that shortfall is exactly right. A solution gμν(x)g_{\mu\nu}(x) can be relabelled by any smooth change of chart xx(x)x\to x'(x), which is four arbitrary functions. The relabelled metric describes the same geometry, so it must also be a solution.

So the equations cannot determine all ten components. If they did, they would forbid the relabelling, and by (C1) the relabelling is not physical. Four functions' worth of freedom has to remain undetermined, and four is exactly what the Bianchi identities leave undetermined.

The same structure, met before

This is not a peculiarity of gravity. Chapter 2.6 §3.1 wrote Maxwell's sourced equations as μFμν=μ0jν\partial_{\mu}F^{\mu\nu}=\mu_{0}j^{\nu}, which is four equations for the four components of AμA_{\mu}. Taking ν\partial_{\nu} of both sides gives νμFμν=0\partial_{\nu}\partial_{\mu}F^{\mu\nu}=0 automatically, since FF is antisymmetric and the two derivatives commute. So one of the four equations is not independent. Three carry information, and the missing one matches the one function of gauge freedom, AA+dχA\to A+\dd\chi, that Chapter 3.5 §10.3 showed is always available.

The dictionary is exact: one identity and one gauge function in electromagnetism, four identities and four coordinate functions in gravity. In both cases the identity is forced by the structure of the left-hand side, and in both cases it exists to make room for a freedom that no measurement can see. Chapter 6.3 makes this correspondence into a principle.

7.2 · Which four are the constraints

Let's look at where the four identities bite. Write out μGμ0=0\nabla_{\mu}G^{\mu 0}=0. It contains 0G00\partial_{0}G^{00} and iGi0\partial_{i}G^{i0}, so it expresses the time derivative of G00G^{00} in terms of quantities involving one fewer time derivative. Since GμνG_{\mu\nu} contains second derivatives of the metric, the consequence is that the four equations G0ν=κT0νG^{0\nu}=\kappa T^{0\nu} contain no second time derivatives at all.

So they are not evolution equations. They are conditions that the initial data must satisfy, and the identities then guarantee that if the conditions hold at one time they hold at every later time.

Once again electromagnetism did this first. In Chapter 2.6, Gauss's law E=ρ/ϵ0\nabla\cdot\vv E=\rho/\epsilon_{0} contains no time derivative of E\vv E. It is a constraint on initial data, preserved by the other equations because charge is conserved. Gravity has six evolution equations and four constraints. Electromagnetism has three and one.

7.3 · The equations are nonlinear, and that is physics

One last structural fact, and it is the one that makes Part III hard and Chapter 7.1 harder. GμνG_{\mu\nu} is built from Γ\Gamma and its derivatives, and Γ\Gamma contains g1gg^{-1}\partial g. So the Einstein tensor contains the inverse metric multiplied by derivatives of the metric, over and over. It is nonlinear in gμνg_{\mu\nu}, badly so.

What the nonlinearity means physically. Superposition fails, so the field of two masses is not the sum of their separate fields. The reason is not technical. Gravitational fields carry energy, and by §1 everything carrying energy sources gravity, so gravity gravitates. There is no way to write a linear theory with that property, because linearity means the source is independent of the field.

Contrast electromagnetism, which is linear. Chapter 2.6's μFμν=μ0jν\partial_{\mu}F^{\mu\nu}=\mu_{0}j^{\nu} is linear in AA, and two solutions superpose. The reason is that the electromagnetic field is not itself charged. Chapter 6.4 builds a theory where the field does carry the charge it responds to, and finds equations that look startlingly like these.

Two consequences are worth filing. Exact solutions are rare and precious, which is why Chapter 3.7's Schwarzschild solution is a landmark rather than an exercise. And the standard technique of physics, which is to expand about a simple solution and keep the first correction, becomes the only technique available. That is why Chapter 3.1's warning that nearly everything is an approximation applies here with particular force.

In plain terms 3.6.7

Counting is worth doing before solving. The unknown is a symmetric array with ten entries and the law supplies ten equations, a matched set until one notices that the identity driving the chapter makes four of them redundant. Only six carry information about how the geometry develops.

That shortfall is not a defect and could not have been otherwise. Any solution can be repainted with different coordinate labels, which takes four arbitrary functions and changes nothing measurable, so the law is obliged to leave four functions' worth undetermined. The identity exists to make room for the freedom. The four leftover equations are not useless; they are conditions the starting data must satisfy, automatically preserved thereafter, exactly as the law relating electric field to charge constrains starting data rather than governing its development.

The last observation is the expensive one. The geometry side is nonlinear in the geometry, so the field of two masses is not the field of one added to the field of the other. There is no way to avoid this, because everything carrying energy is a source, and the gravitational field carries energy, so gravity is a source of itself. Electromagnetism escapes because its field carries no charge. That difference is why exact solutions are rare, why approximation is the normal state of affairs, and why the last part of the book finds gravity so much harder than everything else.

8 · Worked examples

Worked example 1 — the variational definition of TμνT_{\mu\nu} reproduces Chapter 2.6's tensor

Feed Chapter 2.6's electromagnetic Lagrangian into (3.6.19) and check that the result is (2.6.80), sign included. Without this check, the definition (3.6.19) would be no more than an assertion.

The action. Chapter 2.6 §9 gave the electromagnetic Lagrangian density as (2.6.72), and its source-free part is L=FμνFμν/4μ0\mathcal{L}=-F_{\mu\nu}F^{\mu\nu}/4\mu_{0}. Written out on a manifold, the action is therefore

Sm  =  14μ0 ⁣gμαgνβFμνFαβ  g  d4x. S_{\text{m}} \;=\; -\frac{1}{4\mu_{0}}\int\! g^{\mu\alpha}g^{\nu\beta}F_{\mu\nu}F_{\alpha\beta}\;\sqrt{-g}\;\dd^{4}x.

The indices have been written out explicitly because the metric dependence is the whole point. Note that Fμν=μAννAμF_{\mu\nu}=\partial_{\mu}A_{\nu}-\partial_{\nu}A_{\mu} carries no metric dependence at all, by Chapter 3.5 §2.2. It is dA\dd A, and the exterior derivative needs no connection. So only the two inverse metrics and the volume element vary.

Vary the contraction. Two inverse metrics, so the product rule gives two terms:

δ(FαβFαβ)  =  δgμαgνβFμνFαβ  +  gμαδgνβFμνFαβ  =  2FμβFνβ  δgμν, \delta\big(F_{\alpha\beta}F^{\alpha\beta}\big) \;=\; \delta g^{\mu\alpha}\,g^{\nu\beta}F_{\mu\nu}F_{\alpha\beta} \;+\; g^{\mu\alpha}\,\delta g^{\nu\beta}\,F_{\mu\nu}F_{\alpha\beta} \;=\; 2\,F_{\mu\beta}F_{\nu}{}^{\beta}\;\delta g^{\mu\nu},

The two terms are equal after relabelling dummies and using the antisymmetry of FF twice, which supplies two minus signs that cancel.

Vary the volume element with (3.6.15). Collecting:

δSm  =  14μ0 ⁣g[2FμβFνβ    12gμνFαβFαβ]δgμν  d4x. \delta S_{\text{m}} \;=\; -\frac{1}{4\mu_{0}}\int\!\sqrt{-g}\left[\,2F_{\mu\beta}F_{\nu}{}^{\beta} \;-\; \half g_{\mu\nu}F_{\alpha\beta}F^{\alpha\beta}\right]\delta g^{\mu\nu}\;\dd^{4}x.

Read off TμνT_{\mu\nu} by comparing with (3.6.19), which says the bracket times g\sqrt{-g} equals 12gTμν\half\sqrt{-g}\,T_{\mu\nu} divided by the prefactor:

Tμν  =  1μ0(FμβFνβ    14gμνFαβFαβ). T_{\mu\nu} \;=\; -\frac{1}{\mu_{0}}\left(F_{\mu\beta}F_{\nu}{}^{\beta} \;-\; \frac14\,g_{\mu\nu}F_{\alpha\beta}F^{\alpha\beta}\right).

Compare with Chapter 2.6. Equation (2.6.80) reads Tμν=1μ0(FμλFλν+14ημνFαβFαβ)T^{\mu\nu}=\frac{1}{\mu_{0}}\big(F^{\mu\lambda}F_{\lambda}{}^{\nu}+\frac14\eta^{\mu\nu}F_{\alpha\beta}F^{\alpha\beta}\big). Lower both free indices, and then handle the first term: FμλFλν=FμλFνλ=FμβFνβF_{\mu}{}^{\lambda}F_{\lambda\nu} = -F_{\mu}{}^{\lambda}F_{\nu\lambda} = -F_{\mu\beta}F_{\nu}{}^{\beta}. The first step uses the antisymmetry of FF, and the second raises and lowers the summed index, which is free of charge. So Chapter 2.6's tensor is 1μ0(FμβFνβ14gμνF2)-\frac{1}{\mu_{0}}\big(F_{\mu\beta}F_{\nu}{}^{\beta}-\frac14 g_{\mu\nu}F^{2}\big). Identical.

The sign, checked against a physical number. For a pure electric field along xx, F01=Ex/cF_{01}=E_{x}/c, so F0βF0β=g11F01F01=E2/c2F_{0\beta}F_{0}{}^{\beta}=g^{11}F_{01}F_{01}=-E^{2}/c^{2}, and by Chapter 2.6 §7.1 FαβFαβ=2(B2E2/c2)F_{\alpha\beta}F^{\alpha\beta}=2(B^{2}-E^{2}/c^{2}). Substituting, T00=1μ0(E2/c212(B2E2/c2))=ϵ0E2/2+B2/2μ0T_{00}=-\frac{1}{\mu_{0}}\big(-E^{2}/c^{2}-\half(B^{2}-E^{2}/c^{2})\big)=\epsilon_{0}E^{2}/2+B^{2}/2\mu_{0}, which is positive, and which is Chapter 2.6's (2.6.82). So the definition (3.6.19) has the sign that makes energy density positive in this book's signature. (Books using (,+,+,+)(-,+,+,+) define TμνT_{\mu\nu} with the opposite sign, for the same reason the action carries the opposite sign there.)

Worked example 2 — how curved is the inside of an ordinary object?

Use the trace of the field equations to compute the Ricci scalar inside a body of uniform density, and evaluate it for air, water, the Earth and a neutron star.

The formula. From (3.6.22) we have R=κTR=-\kappa T, and for slow cold matter (3.6.3) gives T=ρc2T=\rho c^{2}. Putting the two together,

R  =  8πGc4  ρc2  =  8πGρc2. R \;=\; -\,\frac{8\pi G}{c^{4}}\;\rho c^{2} \;=\; -\,\frac{8\pi G\rho}{c^{2}}.

Note that this is a local statement. The Ricci scalar at a point depends only on the density at that point, not on how much matter is elsewhere. The mass of the Earth as a whole does not appear. Everything non-local about gravity lives in the Weyl part of the curvature that Chapter 3.4 §6.2 named and set aside, which is exactly why vacuum is not flat.

Numbers. Curvature has dimensions of one over length squared, so the quantity worth putting beside it is the radius of curvature R1/2\abs{R}^{-1/2}.

Materialρ (kgm3)\rho\ (\mathrm{kg\,m^{-3}})R (m2)R\ (\mathrm{m^{-2}})R1/2\abs{R}^{-1/2}
Air at sea level1.21.22.3×1026-2.3\times10^{-26}6.6×1012 m6.6\times10^{12}\ \mathrm{m}
Water1.0×1031.0\times10^{3}1.9×1023-1.9\times10^{-23}2.3×1011 m2.3\times10^{11}\ \mathrm{m}
Earth, mean5.5×1035.5\times10^{3}1.0×1022-1.0\times10^{-22}9.9×1010 m9.9\times10^{10}\ \mathrm{m}
Neutron star, core5×10175\times10^{17}9.3×109-9.3\times10^{-9}1.0×104 m1.0\times10^{4}\ \mathrm{m}

What to take from the table. Inside the Earth, the radius of curvature of spacetime is about two-thirds of an astronomical unit. The geometry departs from flatness on a scale comparable to the Earth's distance from the Sun. That is why nobody noticed for three hundred years, and it is the same conclusion Chapter 3.4 §4.5 reached from the tidal side with the closely similar number 1.7×1023 m21.7\times10^{-23}\ \mathrm{m^{-2}}.

The last row is the interesting one. For a neutron star the radius of curvature drops to about ten kilometres, which is the size of the object itself. When the radius of curvature becomes comparable to the body producing it, no expansion in Φ/c2\Phi/c^{2} is available, and the full nonlinear equations are needed. That is the boundary of the regime §5 assumed, stated as a length rather than as a small parameter.

9 · Your turn

Problem 1 — three forms of the same equation

(a) Derive the trace-reversed form (3.6.23) in the presence of Λ\Lambda, showing that Rμν=κ(Tμν12Tgμν)ΛgμνR_{\mu\nu}=\kappa\big(T_{\mu\nu}-\half Tg_{\mu\nu}\big)-\Lambda g_{\mu\nu}. (b) Deduce the vacuum equations with and without Λ\Lambda, and give RR in each case. (c) Chapter 3.7 solves the case Λ=0\Lambda=0, Tμν=0T_{\mu\nu}=0. Say in one sentence why that is a sensible thing to do for the space outside the Sun even though Λ0\Lambda\neq0 in the universe. (d) In two dimensions Gμν0G_{\mu\nu}\equiv0 identically (Chapter 3.4 Problem 2). What does that say about the field equations there?

Solution

(a) Contract GμνΛgμν=κTμνG_{\mu\nu}-\Lambda g_{\mu\nu}=\kappa T_{\mu\nu} with gμνg^{\mu\nu}: R4Λ=κT-R-4\Lambda=\kappa T, so R=κT4ΛR=-\kappa T-4\Lambda. Substitute into Rμν=Gμν+12Rgμν=κTμν+Λgμν+12(κT4Λ)gμνR_{\mu\nu}=G_{\mu\nu}+\half Rg_{\mu\nu}=\kappa T_{\mu\nu}+\Lambda g_{\mu\nu}+\half(-\kappa T-4\Lambda)g_{\mu\nu}, and collect: the Λ\Lambda terms give Λ2Λ=Λ\Lambda-2\Lambda=-\Lambda, leaving the stated result.

(b) With Λ=0\Lambda=0 and Tμν=0T_{\mu\nu}=0 we get Rμν=0R_{\mu\nu}=0 and hence R=0R=0. With Λ0\Lambda\neq0 and Tμν=0T_{\mu\nu}=0 we get Rμν=ΛgμνR_{\mu\nu}=-\Lambda g_{\mu\nu}, and contracting gives R=4ΛR=-4\Lambda. Note that the second is not flat and not even Ricci-flat. Empty space with a cosmological constant is curved, which is the whole of Chapter 3.9's late-time behaviour.

(c) Because Λ1052 m2\Lambda\approx10^{-52}\ \mathrm{m^{-2}}, while the curvature scales relevant to planetary orbits are set by GM/c2r3GM_{\odot}/c^{2}r^{3}, which at Mercury's orbit is about 1030 m210^{-30}\ \mathrm{m^{-2}}. That is larger by twenty-two orders of magnitude. Neglecting Λ\Lambda inside the Solar System is not an approximation anyone can detect.

(d) The left-hand side vanishes identically, so the equations read 0=κTμν0=\kappa T_{\mu\nu}. That is not a field equation but a prohibition, since it says there can be no matter. Gravity has no dynamics in two dimensions. Chapter 7.2, which works entirely in two dimensions, is therefore not doing gravity, and that turns out to be a feature.

Problem 2 — geometry tells matter how to move, and it is not a separate law

The field equations imply μTμν=0\nabla_{\mu}T^{\mu\nu}=0, since the left-hand side is identically divergence-free. Take dust, Tμν=ρuμuνT^{\mu\nu}=\rho u^{\mu}u^{\nu}. (a) Expand μ(ρuμuν)=0\nabla_{\mu}(\rho u^{\mu}u^{\nu})=0 by the product rule. (b) Contract the result with uνu_{\nu} and use uu=c2u\cdot u=c^{2} to show μ(ρuμ)=0\nabla_{\mu}(\rho u^{\mu})=0. (c) Substitute back and deduce the geodesic equation. (d) Say what has just been proved about the logical structure of general relativity.

Solution

(a) μ(ρuμuν)=(μ(ρuμ))uν+ρuμμuν=0\nabla_{\mu}(\rho u^{\mu}u^{\nu})=\big(\nabla_{\mu}(\rho u^{\mu})\big)u^{\nu}+\rho\,u^{\mu}\nabla_{\mu}u^{\nu}=0.

(b) Contract with uνu_{\nu}. The first term gives (μ(ρuμ))uνuν=c2μ(ρuμ)\big(\nabla_{\mu}(\rho u^{\mu})\big)u_{\nu}u^{\nu}=c^{2}\nabla_{\mu}(\rho u^{\mu}). The second gives ρuμuνμuν=12ρuμμ(uνuν)=12ρuμμ(c2)=0\rho\,u^{\mu}u_{\nu}\nabla_{\mu}u^{\nu}=\half\rho\,u^{\mu}\nabla_{\mu}\big(u_{\nu}u^{\nu}\big)=\half\rho\,u^{\mu}\nabla_{\mu}(c^{2})=0, where the metric was taken through the derivative to combine the two factors, and the middle step is the product rule read backwards. So μ(ρuμ)=0\nabla_{\mu}(\rho u^{\mu})=0, which is the continuity equation for the dust, saying that no particles are created or destroyed.

(c) With the first term of (a) gone, ρuμμuν=0\rho\,u^{\mu}\nabla_{\mu}u^{\nu}=0, and dividing by ρ\rho where it is non-zero gives uμμuν=0u^{\mu}\nabla_{\mu}u^{\nu}=0. That is the geodesic equation in the form Chapter 3.3 §8.1 derived it: the tangent parallel-transports itself.

(d) The equation of motion for matter is a consequence of the field equations, not an additional postulate. In Newtonian gravity, "the field satisfies Poisson's equation" and "a particle accelerates according to Φ-\nabla\Phi" are two independent laws. Here the second follows from the first, through the Bianchi identity. Geometry does not merely tell matter how to move. It is not permitted to say anything else.

Problem 3 — the spatial metric, and a grievance from Chapter 3.1 settled

Chapter 3.1 §6.5 obtained g00=1+2Φ/c2g_{00}=1+2\Phi/c^{2} from the redshift and then complained: "this argument has said absolutely nothing about the spatial components of the geometry." Settle it. Take the static weak-field ansatz g00=1+2Φ/c2g_{00}=1+2\Phi/c^{2}, gij=(1+2Ψ/c2)δijg_{ij}=-\big(1+2\Psi/c^{2}\big)\delta_{ij} with Φ,Ψ0\Phi,\Psi\to0 far away, and a dust source at rest. (a) Write down the ijij components of (3.6.23) for iji\neq j and for i=ji=j. (b) Given that to first order Rij=[ij(Φ+Ψ)+δij2Ψ]/c2R_{ij}=-\big[\partial_{i}\partial_{j}(\Phi+\Psi)+\delta_{ij}\nabla^{2}\Psi\big]/c^{2}, deduce Ψ=Φ\Psi=-\Phi. (c) Write the metric. (d) Say which later result this is needed for.

Solution

(a) For dust at rest, Tij=0T_{ij}=0 and T=ρc2T=\rho c^{2}, so Rij=κ(012ρc2gij)=+12κρc2δijR_{ij}=\kappa\big(0-\half\rho c^{2}g_{ij}\big)=+\half\kappa\rho c^{2}\,\delta_{ij} to leading order, using gij=δijg_{ij}=-\delta_{ij}. So the off-diagonal components must vanish and the three diagonal ones must be equal.

(b) Write SΦ+ΨS\equiv\Phi+\Psi. Off-diagonal, iji\neq j: the stated expression gives ijS=0\partial_{i}\partial_{j}S=0. Diagonal: the three components are equal to each other, so x2S=y2S=z2S\partial_{x}^{2}S=\partial_{y}^{2}S=\partial_{z}^{2}S. Call the common value kk, so that 2S=3k\nabla^{2}S=3k. The diagonal equation itself reads (k+2Ψ)/c2=12κρc2-\big(k+\nabla^{2}\Psi\big)/c^{2}=\half\kappa\rho c^{2}, while R00=2Φ/c2R_{00}=\nabla^{2}\Phi/c^{2} gives 2Φ/c2=12κρc2\nabla^{2}\Phi/c^{2}=\half\kappa\rho c^{2} for the same right-hand side. Subtracting, k+2Ψ=2Φk+\nabla^{2}\Psi=-\nabla^{2}\Phi, that is k=2S=3kk=-\nabla^{2}S=-3k, so k=0k=0. Every second derivative of SS therefore vanishes, which makes SS a linear function of position. Requiring both potentials to vanish at infinity then leaves only S=0S=0. Hence Ψ=Φ\Psi=-\Phi.

(c) ds2=(1+2Φc2)c2dt2(12Φc2)(dx2+dy2+dz2).\dd s^{2}=\Big(1+\dfrac{2\Phi}{c^{2}}\Big)c^{2}\dd t^{2}-\Big(1-\dfrac{2\Phi}{c^{2}}\Big)\big(\dd x^{2}+\dd y^{2}+\dd z^{2}\big). Note the signs: time is stretched by 1+2Φ/c21+2\Phi/c^{2} and space by 12Φ/c21-2\Phi/c^{2}, so with Φ<0\Phi<0 near a mass, clocks run slow and rulers are shortened.

(d) Chapter 3.1 §7.3 computed the deflection of light past the Sun using only the equivalence principle and got exactly half the observed value, promising that the missing half would be identified. It is the spatial part of (c). A light ray is affected by the geometry of space as much as by the geometry of time, and the cabin argument saw only the latter. Chapter 3.8 §§3–4 do the integral and collect the factor of two.

Problem 4 — the vacuum as a fluid, and Einstein's static universe

(a) Using (3.6.39) and Λ=1.1×1052 m2\Lambda=1.1\times10^{-52}\ \mathrm{m^{-2}}, compute ρvac\rho_{\text{vac}} in kilograms per cubic metre and in hydrogen atoms per cubic metre. (b) At what distance from the Sun does the repulsive source 2ρvac-2\rho_{\text{vac}} match the Sun's own ρ\rho averaged over a sphere of that radius? Put another way, where does Λ\Lambda start to matter? (c) Show, using (3.6.35) with Λ\Lambda restored, that a static uniform universe of density ρ\rho requires Λ=4πGρ/c2\Lambda=4\pi G\rho/c^{2}. (d) Argue that this balance is unstable.

Solution

(a) ρvac=Λc2/8πG=(1.1×1052)(8.99×1016)/(1.68×109)=5.9×1027 kgm3\rho_{\text{vac}}=\Lambda c^{2}/8\pi G=(1.1\times10^{-52})(8.99\times10^{16})/(1.68\times10^{-9}) =5.9\times10^{-27}\ \mathrm{kg\,m^{-3}}. Dividing by the hydrogen mass 1.67×1027 kg1.67\times10^{-27}\ \mathrm{kg} gives about 3.53.5 atoms per cubic metre.

(b) The Sun's mass spread over a sphere of radius rr has mean density 3M/4πr33M_{\odot}/4\pi r^{3}. Setting that equal to 2ρvac2\rho_{\text{vac}} and solving, r=(3M/8πρvac)1/3r=\big(3M_{\odot}/8\pi\rho_{\text{vac}}\big)^{1/3}. With M=2.0×1030 kgM_{\odot}=2.0\times10^{30}\ \mathrm{kg} this is about 3×1018 m3\times10^{18}\ \mathrm{m}, or roughly 100100 parsecs. So Λ\Lambda is irrelevant on any scale smaller than a substantial piece of a galaxy, and dominant on scales much larger. That is what §6.2's bound said in other words.

(c) With Λ\Lambda restored, §6.2 gave 2Φ=4πGρΛc2\nabla^{2}\Phi=4\pi G\rho-\Lambda c^{2}. A static uniform universe has no preferred point and hence a potential with no second derivative anywhere, so 2Φ=0\nabla^{2}\Phi=0, giving Λ=4πGρ/c2\Lambda=4\pi G\rho/c^{2}.

(d) Compress the universe slightly. Ordinary matter's density ρ\rho rises, since the same matter now occupies less volume, while Λ\Lambda is a constant and does not. So the source 4πGρΛc24\pi G\rho-\Lambda c^{2} becomes positive, gravity wins, and the compression accelerates. Expand slightly and the reverse happens. The equilibrium is a pencil balanced on its point. Chapter 3.9 §3 shows the same thing by integrating the equations rather than arguing from them.

The brick you just laid

You have the law of gravity, and it was not guessed. Matter's energy, momentum, pressure and stress are packaged in one symmetric tensor whose divergence vanishes. Anything set equal to that tensor must be divergence-free identically, for every metric. The only symmetric rank-2 objects available from a metric and its first two derivatives are RμνR_{\mu\nu}, RgμνRg_{\mu\nu} and gμνg_{\mu\nu}. Imposing the divergence condition on the most general combination of those three fixes the ratio of the first two and leaves the third free. That is (3.6.12), and ⚑ Lovelock's theorem says there is nothing else in four dimensions.

The same equation, from an action. There is essentially one scalar that can be built from a metric with at most two derivatives, so the action writes itself. Varying it splits into three pieces, and the Einstein tensor appears with its factor of 12\half coming this time from the derivative of a determinant instead of from a contracted Bianchi identity. The third piece is a total derivative, which is why the equations are second order even though the action is not. The term thrown away is the one ⚑ Gibbons–Hawking–York repairs, which Chapter 7.9 needs when the value of the action becomes an entropy.

The constant, fixed. Weak field, static, slow matter. R00R_{00} works out to 2Φ/c2\nabla^{2}\Phi/c^{2} using g00=1+2Φ/c2g_{00}=1+2\Phi/c^{2}, with the plus sign belonging to this book's signature and derived in Chapter 3.1 from the redshift with no general relativity in it. The trace-reversed source is 12ρc2\half\rho c^{2}, and matching against Poisson's equation gives κ=8πG/c4\kappa=8\pi G/c^{4}. Restoring the pressure gives 2Φ=4πG(ρ+3p/c2)\nabla^{2}\Phi=4\pi G(\rho+3p/c^{2}), so pressure gravitates, radiation gravitates twice as hard as dust of the same energy density, and a substance with p<ρc2/3p<-\rho c^{2}/3 would push.

And the term nothing excludes. Λgμν\Lambda g_{\mu\nu} passes every constraint in the chapter, it is the constant one may add to any Lagrangian, and it behaves as a fluid with p=ρc2p=-\rho c^{2}. That is the unique equation of state which looks the same to every observer, and the one that does not dilute as space grows. It is measured to be about 1052 m210^{-52}\ \mathrm{m^{-2}}, tiny and not zero, and that is a problem nobody has solved.

Where this gets spent. Chapter 3.7 sets Tμν=0T_{\mu\nu}=0 and solves Rμν=0R_{\mu\nu}=0 outside a spherical mass, using Chapter 3.5's Killing vectors to make the geodesics tractable, and pays the factor-of-two debt from Chapter 3.1 using the spatial metric of Problem 3. Chapter 3.9 puts a perfect fluid on the right and a homogeneous, isotropic metric on the left, and there the pressure term and the cosmological term both do real work. And Chapter 7.1 asks what happens when the same equations are quantised. At that point the nonlinearity of §7.3, which is gravity gravitating, stops being an inconvenience and becomes the obstruction.