Part III · General Relativity — Chapter 3.7

Schwarzschild: The Solution and Its Orbits

Solve Einstein's equations exactly, by hand, in one sitting. Then extract from the answer an orbit Newton has no name for, and a planet's slow turning that nobody could explain for sixty years.

Where we are

Chapter 3.6 finished the law and spent nothing. You have Gμν=κTμνG_{\mu\nu}=\kappa T_{\mu\nu} with κ=8πG/c4\kappa=8\pi G/c^{4}, derived twice. You have its trace-reversed form, which in empty space reads Rμν=0R_{\mu\nu}=0. And you have a warning from Chapter 3.4 §6.3 that empty is not the same as flat. What you have not done is solve anything. This chapter solves something.

The target is the geometry outside a spherical star: our Sun, a neutron star, a black hole. It is the first exact solution of Einstein's equations, found by Karl Schwarzschild within weeks of the field equations being published and while he was serving on the Russian front, and it is short enough to derive completely. Nothing below is quoted. The metric is constructed, the connection is computed, the Ricci tensor is computed, the equations are solved, and the two constants of integration are fixed by two physical demands that are named as they are used.

Then the solution gets spent. Chapter 3.5 §9 promised, twice and in writing, that its Killing vectors would make the orbits of this chapter tractable. Section 5 collects that promise by name and turns four coupled second-order equations into one first-order equation for a single variable. That equation reads as motion in a landscape, and the landscape contains exactly one term that Newton's does not. Everything else in this chapter is that term's consequences: a smallest possible circular orbit, below which no circular orbit exists at any radius whatever, and an ellipse that fails to close by forty-three seconds of arc per century.

The route. Section 1 says what the two symmetry words mean as statements about geometry rather than about pictures. Section 2 turns them into an ansatz with two unknown functions in it, and names the one step in that reduction which is a choice. Section 3 solves Rμν=0R_{\mu\nu}=0. Section 4 shows that the staticity assumed in §1 was not needed and comes back on its own, which has a startling consequence for collapsing stars. Section 5 collects the conserved quantities, §6 builds the effective potential, §7 finds the innermost stable circular orbit, and §8 derives the precession and computes Mercury's number with every conversion on the page.

Conventions. The signature is (+,,,)(+,-,-,-), and the Riemann and Ricci sign conventions are the ones stated loudly in Chapter 3.4. We write x0=ctx^{0}=ct, so that g00g_{00} multiplies (dx0)2=c2dt2(\dd x^{0})^{2}=c^{2}\dd t^{2}. Both GG and cc stay explicit everywhere. The Schwarzschild radius is written 2GM/c22GM/c^{2} and never 2M2M. A prime means d/dr\dd/\dd r throughout.

Tools you'll need  — Chapter 3.6 §5.1, which supplies the trace-reversed field equations (3.6.23) and the remark that in vacuum they read Rμν=0R_{\mu\nu}=0. Chapter 3.5 §8.1 for Killing's equation and its practical test, which says that every coordinate the metric does not mention supplies a Killing vector, and §9 in full, which is the section this chapter spends. Chapter 3.4 §6.1 for the Ricci tensor and its definition (3.4.47), and §6.3 for Ricci-flat against flat. Chapter 3.3 §7 (the Christoffel formula, (3.3.50)), §8 (geodesics, both derivations, and the constancy of gμνuμuνg_{\mu\nu}u^{\mu}u^{\nu} along one) and §3 (proper time). Chapter 2.5 §1 for the normalisation uu=c2u\cdot u=c^{2}, equation (2.5.9), in flat spacetime. Chapter 3.2 §2 for what a chart is and what it is not. Chapter 3.1 §6.5, the derived weak-field component g00=1+2Φ/c2g_{00}=1+2\Phi/c^{2} (3.1.39), which fixes this chapter's mass. Chapter 0.8 §2.2, the integrating factor, which does the one integration in §3, and §4.3 for reading an energy equation as motion in a landscape, a reading Chapter 1.3 §4 turned into phase portraits. Chapter 0.3 §2.4, the binomial series, for the one expansion in §8.4. Chapter 1.4 §3 for energy and angular momentum as Noether charges, and Worked example 2 there, on the Laplace–Runge–Lenz vector, which proved that Kepler ellipses close and predicted that any departure from an exact inverse square would make them precess. This chapter supplies the departure.

1 · What "spherically symmetric and static" buys

Here is where this section is going. We are going to translate two English adjectives into conditions on the metric, using the machinery of Chapter 3.5 rather than a picture. The translation matters because "looks the same from every direction" is a statement about a drawing, and a drawing cannot be substituted into a field equation. "There are three vector fields whose flows leave the metric unchanged" can be.

1.1 · Spherical symmetry, as a statement about the geometry

Chapter 3.5 §8.1 defined a Killing vector field ξ\xi as one whose flow leaves the metric exactly as it found it, Lξg=0\mathcal{L}_{\xi}g=0, equivalently μξν+νξμ=0\nabla_{\mu}\xi_{\nu}+\nabla_{\nu}\xi_{\mu}=0. That chapter also emphasised the distinction this section needs: a Killing vector is not a direction along which the picture looks the same, it is a direction along which every length anybody could measure comes out the same.

A spacetime is spherically symmetric when it admits three Killing vector fields ξ(1),ξ(2),ξ(3)\xi_{(1)},\xi_{(2)},\xi_{(3)}, each spacelike, whose commutators reproduce the algebra of rotations:

[ξ(i),ξ(j)]  =  kϵijkξ(k). \big[\xi_{(i)},\,\xi_{(j)}\big] \;=\; \sum_{k}\epsilon_{ijk}\,\xi_{(k)}. (3.7.1)

Two things about this definition are worth pausing on, because it looks more elaborate than the English phrase it replaces.

The bracket is Chapter 3.2's commutator of vector fields, and by Chapter 3.5 §7.2 it is also the Lie derivative of one field along the other. Requiring (3.7.1) rather than merely requiring three Killing fields is what distinguishes rotations from, say, three translations: it is the statement that these three symmetries fail to commute in the particular pattern that turns about three perpendicular axes fail to commute.

You have already met these three fields. Chapter 3.5 §8.3 computed them on a sphere of radius aa: the obvious one φ\partial_{\varphi}, and two more, cosφθcotθsinφφ\cos\varphi\,\partial_{\theta}-\cot\theta\sin\varphi\,\partial_{\varphi} and sinφθcotθcosφφ-\sin\varphi\,\partial_{\theta}-\cot\theta\cos\varphi\,\partial_{\varphi}, which are rotations about the other two axes and are invisible in the usual chart. Those are the ξ(i)\xi_{(i)} of (3.7.1), extended off the sphere.

1.2 · Static, which is stronger than stationary

A spacetime is stationary when it admits a further Killing field ξ(0)\xi_{(0)} that is timelike, meaning gμνξ(0)μξ(0)ν>0g_{\mu\nu}\xi_{(0)}^{\mu}\xi_{(0)}^{\nu}>0 in this book's signature. In a chart adapted to it, that says the metric components do not depend on x0x^{0}. Nothing about the geometry changes with time.

A spacetime is static when it is stationary and additionally invariant under reversing the time direction, x0x0x^{0}\to-x^{0}.

The second condition is not decoration and it is not implied by the first. Here is a case that separates them, and it is the case that matters physically. Take the geometry outside a rotating star. Nothing about it changes with time, so it is stationary. But reversing time reverses the sense of the rotation, which is a different physical situation, so it is not static. The exact solution for a rotating mass is Kerr's, found in 1963, and it is a different and much longer calculation which this book does not perform. What follows is the non-rotating case.

What the time reversal buys, in one sentence

Under x0x0x^{0}\to-x^{0} the components g0ig_{0i} pick up a minus sign, since exactly one of their two indices refers to the reversed coordinate, while g00g_{00} and gijg_{ij} do not. Demanding that the metric be unchanged therefore forces

g0i  =  g0ig0i=0. g_{0i} \;=\; -\,g_{0i} \qquad\Longrightarrow\qquad g_{0i}=0 .

Staticity is precisely the statement that there are no cross terms pairing time with a direction in space. Without it, §2's ansatz would carry an extra unknown function, and that function is where the dragging of inertial frames by a rotating body lives. Chapter 3.8 uses the distinction again when it asks what kind of surface a horizon is.

1.3 · These are hypotheses, and they are being chosen

Nothing derived so far says that any actual object is spherically symmetric or static. The Sun rotates. The Earth is oblate. A collapsing star is neither. What §§2–3 solve is a special case chosen because it is exactly solvable, and the honest description of the result is this is the geometry outside a non-rotating spherical body. The accuracy of everything downstream is limited by how well a given body satisfies that.

Section 4 removes one of the two hypotheses at no cost, which is a considerably better outcome than it sounds.

In plain terms 3.7.1

Two words are doing all the work in this chapter and both have to be turned into statements about the geometry rather than about a drawing. Spherically symmetric means there exist three directions of dragging along which every measurable distance comes out unchanged, and that those three combine among themselves exactly the way turns about three perpendicular axes do. Static means two further things at once: there is another such direction, one along which a clock could actually tick, and running that direction backwards changes nothing either.

The second half of that is not decoration. A geometry with an unchanging time direction but no reversal symmetry is called stationary instead, and the difference is real rather than pedantic, because the space outside a spinning star is stationary and not static: reverse time and the spin reverses with it. What the reversal buys is the disappearance of every entry in the distance rule that pairs time with a direction in space, and without it the calculation below would carry another unknown function.

Naming a symmetry as a direction of dragging is what makes it something to calculate with rather than to look at. The previous chapter turned each one into a quantity conserved along every free path, and the same three-line test used there, that a label the distance rule never mentions supplies a symmetry, will identify the ones needed here by inspection.

2 · The ansatz, and why it has only two unknown functions

Here is where this section is going. Einstein's equations are ten coupled nonlinear equations for ten unknown functions of four variables. We are going to use §1's symmetries to reduce that to two unknown functions of one variable. Then we will say exactly which step in the reduction was a choice rather than a consequence, because that step is the one Chapter 3.8 has to un-learn.

2.1 · What the symmetries remove

Work in a chart adapted to the symmetries, which exists by Chapter 3.2 §2 and which we build one piece at a time.

Stationarity removes the x0x^{0} dependence. Chapter 3.5 §8.1's practical test runs in reverse here: if 0\partial_{0} is a Killing field, there is a chart in which no metric component depends on x0x^{0}. Ten functions of four variables become ten functions of three.

Time reversal removes the cross terms. By §1.2, g0i=0g_{0i}=0. Six functions remain in the spatial block, plus g00g_{00}: seven.

Spherical symmetry removes the angular dependence and most of the spatial block. The three rotational Killing fields act on the spatial slices. Their orbits are the sets of points reachable from a given point by rotating, and those sets are two-dimensional spheres. The metric restricted to any one of them must be invariant under all three rotations. Chapter 3.5 §8.3's calculation on the sphere is exactly the statement that up to overall scale there is one such metric, the round one:

dΩ2    dθ2  +  sin2θ  dφ2. \dd\Omega^{2} \;\equiv\; \dd\theta^{2} \;+\; \sin^{2}\theta\;\dd\varphi^{2}. (3.7.2)

So each sphere contributes SdΩ2-S\,\dd\Omega^{2} for a single positive function SS, one number per sphere rather than three functions of angle. Rotational invariance also forbids any cross term pairing dθ\dd\theta or dφ\dd\varphi with the remaining radial direction, since such a term would pick out a direction on the sphere and there is none to pick. What is left is

ds2  =  A(dx0)2    B~dρ2    SdΩ2, \dd s^{2} \;=\; A\,\big(\dd x^{0}\big)^{2} \;-\; \tilde B\,\dd\rho^{2} \;-\; S\,\dd\Omega^{2}, (3.7.3)

with AA, B~\tilde B and SS three functions of the radial label ρ\rho alone. Three unknown functions of one variable. One more move takes it to two.

2.2 · The areal radius — the load-bearing choice

The label ρ\rho in (3.7.3) is arbitrary: relabel ρρ~(ρ)\rho\to\tilde\rho(\rho) and the same geometry is described by three different functions. That freedom is one function's worth, and we spend it.

Here is what we spend it on. The sphere labelled ρ\rho has, by (3.7.3), metric SdΩ2S\,\dd\Omega^{2}, so its area is Ssinθdθdφ=4πSS\int\sin\theta\,\dd\theta\,\dd\varphi=4\pi S. Define a new radial label rr by

r    S,so that the sphere at r has area exactly 4πr2. r \;\equiv\; \sqrt{S}, \qquad\text{so that the sphere at $r$ has area exactly } 4\pi r^{2}. (3.7.4)

This is the areal radius. One condition is being assumed and it should be said out loud: SS must be strictly monotonic in the old radial label, so that r=Sr=\sqrt{S} is invertible and can serve as a coordinate at all. Where it is not, this relabelling is unavailable, and the solution derived below does not describe that region. The place where it fails has a name, a throat, meaning a place where the spheres stop growing as you move outward. Nothing in this chapter meets such a place, but the assumption is real and it is being spent here rather than later.

With the areal radius, (3.7.3) becomes

  ds2  =  A(r)c2dt2    B(r)dr2    r2dΩ2.   \boxed{\;\dd s^{2} \;=\; A(r)\,c^{2}\dd t^{2} \;-\; B(r)\,\dd r^{2} \;-\; r^{2}\,\dd\Omega^{2}.\;} (3.7.5)

Two unknown functions of one variable, AA and BB, both dimensionless. That is the ansatz §3 solves.

⚠ What rr is, and the belief this box exists to prevent

The definition (3.7.4) is a choice, made because it is convenient, and it is the single most misread line in this subject. Read it again: rr is defined so that the sphere at rr has area 4πr24\pi r^{2}. Nothing else.

rr is not the distance to the centre. The distance measured radially, at fixed time, from the sphere at r1r_{1} to the sphere at r2r_{2} is, by (3.7.5), the integral r1r2B(r)dr\int_{r_{1}}^{r_{2}}\sqrt{B(r)}\,\dd r, and §3 finds B>1B>1 everywhere outside the mass, so that distance is larger than r2r1r_{2}-r_{1}. The label and the length disagree, and BB is the exchange rate between them.

There need not even be a centre in the region we are solving. Equation (3.7.5) describes the vacuum outside a body. Inside the body the solution is different and this chapter does not compute it. So "the distance to the centre" is a quantity about a region we are not describing.

Chapter 3.8 §6 shows that treating rr as a distance is exactly what makes the surface at r=2GM/c2r=2GM/c^{2} look catastrophic. It is not catastrophic, and the confusion starts here, with a choice made for convenience and then forgotten.

2.3 · Ten equations, two unknowns, and why that is not yet a contradiction

Substituting (3.7.5) into Rμν=0R_{\mu\nu}=0 gives ten equations for two functions. That is a wildly overdetermined system and it would be reasonable to expect it to have no solution at all. Section 3 finds that six of the ten equations are 0=00=0 identically, by the symmetry, that a seventh is a multiple of another, and that the surviving three are consistent. The consistency is not automatic and Chapter 3.6 §7.1 says where it comes from: the contracted Bianchi identity supplies four relations among the ten field equations for every metric, so the count of genuinely independent equations was never ten.

In plain terms 3.7.2

Exact solutions are rare, as the last chapter said, and symmetry is how the rare ones happen. Ten unknown functions of four variables become two functions of one variable. Following how is worth more than the answer. Independence of time, and of time's direction, removes every entry pairing time with space. Spherical symmetry makes every sphere a perfectly round one and forces the angular part into the single combination a globe uses, with one number in front of it. What survives is one factor multiplying the time part of the distance rule and one multiplying the radial part, each depending on the radial label alone.

Then comes a move that is a decision rather than a deduction, and it has to be said out loud. The radial label is defined by declaring that the sphere carrying it has area four pi times that label squared. Nothing says the label equals the distance anybody would measure walking inward from one sphere to a smaller one, and here it emphatically does not, since the radial factor in the distance rule is what converts one into the other.

Read the label as a distance to the centre and the surface met in the next chapter looks like a catastrophe. It is not one, and the confusion begins right here, in a definition adopted because it makes the algebra short and then quietly reinterpreted as a measurement.

3 · Solving Rμν=0R_{\mu\nu}=0

Here is the route, because there are five steps and each one is small.

  • Compute the connection for (3.7.5).
  • Compute the Ricci tensor from it.
  • Find that one particular combination of two of the equations collapses to the statement that (AB)=0(AB)'=0.
  • Fix the resulting constant by demanding that the geometry become flat far away.
  • Substitute B=1/AB=1/A into the angular equation, which becomes (rA)=1(rA)'=1, and integrate that once.

Two constants of integration appear along the way, and each one is fixed by naming a physical requirement. Out comes an exact solution.

Chapter 3.6 §5.1 supplies the equation we are solving. The trace-reversed field equations (3.6.23) with Tμν=0T_{\mu\nu}=0 and λ=0\lambda=0 read

Rμν  =  0. R_{\mu\nu} \;=\; 0 . (3.7.6)

These are the equations for empty space. They apply outside the star, not inside it, and Chapter 3.4 §6.3 already warned that they do not say the geometry is flat.

3.1 · The connection

Chapter 3.3 §7 derived the Christoffel formula (3.3.50),

Γλμν  =  12gλσ(μgσν+νgσμσgμν), \Gamma^{\lambda}{}_{\mu\nu} \;=\; \half\,g^{\lambda\sigma}\Big(\partial_{\mu}g_{\sigma\nu} + \partial_{\nu}g_{\sigma\mu} - \partial_{\sigma}g_{\mu\nu}\Big), (3.7.7)

and the metric (3.7.5) is diagonal, so the inverse metric is diagonal too, with entries g00=1/Ag^{00}=1/A, g11=1/Bg^{11}=-1/B, g22=1/r2g^{22}=-1/r^{2} and g33=1/(r2sin2θ)g^{33}=-1/(r^{2}\sin^{2}\theta). That collapses the sum over σ\sigma in (3.7.7) to a single term every time. Grind box A does the arithmetic. The answer is nine non-zero components, counting each symmetric pair once:

Γ001=A2A,Γ100=A2B,Γ111=B2B,Γ122=rB,Γ133=rsin2θB,Γ212=1r,Γ233=sinθcosθ,Γ313=1r,Γ323=cotθ. \begin{aligned} &\Gamma^{0}{}_{01}=\frac{A'}{2A}, &\quad &\Gamma^{1}{}_{00}=\frac{A'}{2B}, &\quad &\Gamma^{1}{}_{11}=\frac{B'}{2B}, \\[3pt] &\Gamma^{1}{}_{22}=-\frac{r}{B}, &\quad &\Gamma^{1}{}_{33}=-\frac{r\sin^{2}\theta}{B}, &\quad &\Gamma^{2}{}_{12}=\frac{1}{r}, \\[3pt] &\Gamma^{2}{}_{33}=-\sin\theta\cos\theta, &\quad &\Gamma^{3}{}_{13}=\frac{1}{r}, &\quad &\Gamma^{3}{}_{23}=\cot\theta . \end{aligned} (3.7.8)

Four of the nine are exactly the connection coefficients of flat space in spherical coordinates: the two carrying 1/r1/r, the one carrying cotθ\cot\theta and the one carrying sinθcosθ\sin\theta\cos\theta. Chapter 3.3 §2 computed those four as the standing warning that varying components do not mean curvature. They are here because the chart is spherical, not because the spacetime is curved. The other five carry AA and BB, and they are where the physics lives.

Grind box A — the nine Christoffel symbols, from (3.3.50)

Write xμ=(x0,r,θ,φ)x^{\mu}=(x^{0},r,\theta,\varphi) and recall that no metric component depends on x0x^{0} or on φ\varphi, and that only g33=r2sin2θg_{33}=-r^{2}\sin^{2}\theta depends on θ\theta.

Γ001\Gamma^{0}{}_{01}. Take λ=0\lambda=0, so σ=0\sigma=0 and g00=1/Ag^{00}=1/A. With μ=0\mu=0, ν=1\nu=1 the bracket in (3.7.7) is 0g01+1g000g01=rA=A\partial_{0}g_{01}+\partial_{1}g_{00}-\partial_{0}g_{01}=\partial_{r}A=A'. Hence Γ001=12(1/A)A=A/2A\Gamma^{0}{}_{01}=\half(1/A)A'=A'/2A.

Γ100\Gamma^{1}{}_{00}. Take λ=1\lambda=1, so g11=1/Bg^{11}=-1/B. With μ=ν=0\mu=\nu=0 the bracket is 20g101g00=A2\partial_{0}g_{10}-\partial_{1}g_{00}=-A'. Hence Γ100=12(1/B)(A)=A/2B\Gamma^{1}{}_{00}=\half(-1/B)(-A')=A'/2B.

Γ111\Gamma^{1}{}_{11}. Same λ\lambda, and the bracket is 1g11=B\partial_{1}g_{11}=-B', giving 12(1/B)(B)=B/2B\half(-1/B)(-B')=B'/2B.

Γ122\Gamma^{1}{}_{22}. Same λ\lambda again. With μ=ν=2\mu=\nu=2 the bracket is 22g121g22=+2r2\partial_{2}g_{12}-\partial_{1}g_{22}=+2r, giving 12(1/B)(2r)=r/B\half(-1/B)(2r)=-r/B. The same computation with 3333 in place of 2222 carries the extra sin2θ\sin^{2}\theta through and gives Γ133=rsin2θ/B\Gamma^{1}{}_{33}=-r\sin^{2}\theta/B.

Γ212\Gamma^{2}{}_{12}. Take λ=2\lambda=2, g22=1/r2g^{22}=-1/r^{2}. With μ=1\mu=1, ν=2\nu=2 the bracket is 1g22=2r\partial_{1}g_{22}=-2r, giving 12(1/r2)(2r)=1/r\half(-1/r^{2})(-2r)=1/r. Identically for Γ313\Gamma^{3}{}_{13}, the extra sin2θ\sin^{2}\theta cancelling between g33g^{33} and 1g33\partial_{1}g_{33}.

Γ233\Gamma^{2}{}_{33}. λ=2\lambda=2, μ=ν=3\mu=\nu=3: the bracket is 2g33=+2r2sinθcosθ-\partial_{2}g_{33}=+2r^{2}\sin\theta\cos\theta, giving 12(1/r2)(2r2sinθcosθ)=sinθcosθ\half(-1/r^{2})(2r^{2}\sin\theta\cos\theta)=-\sin\theta\cos\theta.

Γ323\Gamma^{3}{}_{23}. λ=3\lambda=3, μ=2\mu=2, ν=3\nu=3: the bracket is 2g33=2r2sinθcosθ\partial_{2}g_{33}=-2r^{2}\sin\theta\cos\theta, and 12(1/(r2sin2θ))(2r2sinθcosθ)=cotθ\half\big(-1/(r^{2}\sin^{2}\theta)\big)(-2r^{2}\sin\theta\cos\theta)=\cot\theta.

Everything else vanishes. A component Γλμν\Gamma^{\lambda}{}_{\mu\nu} is non-zero only if the bracket contains a surviving derivative, and the only non-vanishing derivatives of metric components are r\partial_{r} of g00,g11,g22,g33g_{00},g_{11},g_{22},g_{33} and θ\partial_{\theta} of g33g_{33}. Running through the possibilities leaves the nine above.

One contraction, for later. Chapter 3.5 §6.4 derived Γλλμ=μlng\Gamma^{\lambda}{}_{\lambda\mu}=\partial_{\mu}\ln\sqrt{-g} from Jacobi's formula, and here g=ABr4sin2θ-g=AB\,r^{4}\sin^{2}\theta, so g=ABr2sinθ\sqrt{-g}=\sqrt{AB}\,r^{2}\sin\theta and

Γλλ1  =  A2A+B2B+2r,Γλλ2  =  cotθ, \Gamma^{\lambda}{}_{\lambda 1} \;=\; \frac{A'}{2A} + \frac{B'}{2B} + \frac{2}{r}, \qquad \Gamma^{\lambda}{}_{\lambda 2} \;=\; \cot\theta,

and the other two vanish. Adding the list above directly gives the same thing, which is the Jacobi formula checked rather than trusted.

(All nine were confirmed symbolically from (3.7.7) for general A(r)A(r) and B(r)B(r), and no other component is non-zero.)

3.2 · The Ricci tensor, and the count

Chapter 3.4 §6.1 defined the Ricci tensor and wrote it out in terms of the connection, equation (3.4.47):

Rμν  =  λΓλνμ    νΓλλμ  +  ΓλλρΓρνμ    ΓλνρΓρλμ. R_{\mu\nu} \;=\; \partial_{\lambda}\Gamma^{\lambda}{}_{\nu\mu} \;-\; \partial_{\nu}\Gamma^{\lambda}{}_{\lambda\mu} \;+\; \Gamma^{\lambda}{}_{\lambda\rho}\Gamma^{\rho}{}_{\nu\mu} \;-\; \Gamma^{\lambda}{}_{\nu\rho}\Gamma^{\rho}{}_{\lambda\mu}. (3.7.9)

Feeding (3.7.8) into it is a page of bookkeeping, folded into grind box B. The results are

R00=A2B    AB4B2    (A)24AB  +  ArB,R11=A2A  +  AB4AB  +  (A)24A2  +  BrB,R22=1    1B  +  rB2B2    rA2AB,R33  =  sin2θ  R22, \begin{aligned} R_{00} &= \frac{A''}{2B} \;-\; \frac{A'B'}{4B^{2}} \;-\; \frac{(A')^{2}}{4AB} \;+\; \frac{A'}{rB}, \\[6pt] R_{11} &= -\frac{A''}{2A} \;+\; \frac{A'B'}{4AB} \;+\; \frac{(A')^{2}}{4A^{2}} \;+\; \frac{B'}{rB}, \\[6pt] R_{22} &= 1 \;-\; \frac{1}{B} \;+\; \frac{rB'}{2B^{2}} \;-\; \frac{rA'}{2AB}, \qquad R_{33} \;=\; \sin^{2}\theta\;R_{22}, \end{aligned} (3.7.10)

with every off-diagonal component zero.

Now the count, which stays out here in the main text. Ten equations went in. Six of them, the off-diagonal ones, are 0=00=0 and say nothing. Of the remaining four, the last two are proportional, so R33=0R_{33}=0 carries no information beyond R22=0R_{22}=0. That proportionality is spherical symmetry showing up in the answer, since θ\theta and φ\varphi are on the same footing. What is left is three equations for two unknown functions.

Three equations for two unknowns is one equation too many, and a system in that position has no right to possess a solution. It does, and the reason was supplied a chapter ago: by Chapter 3.6 §7.1 the contracted Bianchi identity imposes four differential relations among the ten field equations for every metric whatever, so the three surviving equations are not independent of one another. That is worth a sentence rather than a shrug, because it is the same identity that fixed the factor of 12\half in the Einstein tensor, doing a second job here.

Grind box B — the Ricci components, term by term

R00R_{00}, all four terms of (3.7.9).

Term 1. λΓλ00\partial_{\lambda}\Gamma^{\lambda}{}_{00}. The only non-zero Γλ00\Gamma^{\lambda}{}_{00} is Γ100=A/2B\Gamma^{1}{}_{00}=A'/2B, so this is

r ⁣(A2B)  =  A2B    AB2B2. \partial_{r}\!\left(\frac{A'}{2B}\right) \;=\; \frac{A''}{2B} \;-\; \frac{A'B'}{2B^{2}}.

Term 2. 0Γλλ0=0-\partial_{0}\Gamma^{\lambda}{}_{\lambda 0}=0: nothing depends on x0x^{0}.

Term 3. ΓλλρΓρ00\Gamma^{\lambda}{}_{\lambda\rho}\Gamma^{\rho}{}_{00}. Only ρ=1\rho=1 contributes, and grind box A's contraction gives

(A2A+B2B+2r)A2B  =  (A)24AB+AB4B2+ArB. \left(\frac{A'}{2A}+\frac{B'}{2B}+\frac{2}{r}\right)\frac{A'}{2B} \;=\; \frac{(A')^{2}}{4AB}+\frac{A'B'}{4B^{2}}+\frac{A'}{rB}.

Term 4. Γλ0ρΓρλ0-\Gamma^{\lambda}{}_{0\rho}\Gamma^{\rho}{}_{\lambda 0}. The only connection coefficients with a 00 among their indices are Γ001\Gamma^{0}{}_{01} and Γ100\Gamma^{1}{}_{00}, so the double sum has exactly two surviving terms, (λ,ρ)=(0,1)(\lambda,\rho)=(0,1) and (1,0)(1,0), and they are equal:

2A2AA2B  =  (A)22AB. -2\cdot\frac{A'}{2A}\cdot\frac{A'}{2B} \;=\; -\frac{(A')^{2}}{2AB}.

Add. The two ABA'B' terms combine as 12+14=14-\tfrac12+\tfrac14=-\tfrac14, and the two (A)2(A')^{2} terms as +1412=14+\tfrac14-\tfrac12=-\tfrac14, giving the first line of (3.7.10).

R11R_{11}. Same four terms. Term 1 is r(B/2B)=B/2B(B)2/2B2\partial_{r}(B'/2B)=B''/2B-(B')^{2}/2B^{2}. Term 2 is r[A/2A+B/2B+2/r]-\partial_{r}\big[A'/2A+B'/2B+2/r\big], which cancels the BB'' of term 1 exactly and leaves A/2A+(A)2/2A2+(B)2/2B2+2/r2-A''/2A+(A')^{2}/2A^{2}+(B')^{2}/2B^{2}+2/r^{2}. Term 3 is [A/2A+B/2B+2/r](B/2B)\big[A'/2A+B'/2B+2/r\big](B'/2B). Term 4 sums the squares (A)2/4A2(B)2/4B22/r2-(A')^{2}/4A^{2}-(B')^{2}/4B^{2}-2/r^{2}, the last from the two coefficients equal to 1/r1/r. Collecting gives the second line of (3.7.10). Note that the 2/r22/r^{2} cancels between terms 2 and 4, which is why no such term appears in the answer.

R22R_{22}. Term 1 is r(r/B)=1/B+rB/B2\partial_{r}(-r/B)=-1/B+rB'/B^{2}. Term 2 is θcotθ=+1+cot2θ-\partial_{\theta}\cot\theta=+1+\cot^{2}\theta. Term 3 is [A/2A+B/2B+2/r](r/B)\big[A'/2A+B'/2B+2/r\big](-r/B). Term 4 is [2(r/B)(1/r)+cot2θ]=2/Bcot2θ-\big[2(-r/B)(1/r)+\cot^{2}\theta\big]=2/B-\cot^{2}\theta. The cot2θ\cot^{2}\theta cancels between terms 2 and 4, which it must, since R22R_{22} cannot depend on θ\theta. What survives is the third line of (3.7.10).

R33R_{33} and the off-diagonal components. R33R_{33} repeats the R22R_{22} computation with the extra sin2θ\sin^{2}\theta carried along, giving sin2θR22\sin^{2}\theta\,R_{22}. Every off-diagonal component vanishes: any RμνR_{\mu\nu} with μν\mu\neq\nu requires a connection coefficient carrying both a μ\mu and a ν\nu, and inspection of (3.7.8) shows the surviving sums are empty or cancel in pairs.

(All of (3.7.10) was reproduced symbolically from (3.7.9) for general A(r)A(r), B(r)B(r), together with the vanishing of all six off-diagonal components and the identity R33=sin2θR22R_{33}=\sin^{2}\theta\,R_{22}.)

3.3 · The line the whole solution turns on

Look at R00R_{00} and R11R_{11} in (3.7.10). Each contains AA'', and each contains the same three other structures. Weight the first by BB and the second by AA and add:

BR00  +  AR11  =    A2A20      AB4B+AB4B0      (A)24A+(A)24A0+  Ar  +  ABrB. \begin{aligned} B\,R_{00} \;+\; A\,R_{11} \;=\;\; & \underbrace{\frac{A''}{2}-\frac{A''}{2}}_{0} \;\; \underbrace{-\;\frac{A'B'}{4B}+\frac{A'B'}{4B}}_{0} \;\; \underbrace{-\;\frac{(A')^{2}}{4A}+\frac{(A')^{2}}{4A}}_{0} \\[8pt] &+\; \frac{A'}{r} \;+\; \frac{AB'}{rB}. \end{aligned} (3.7.11)

Every second derivative has gone, and so has every quadratic term. The two survivors are the two halves of a product rule:

BR00  +  AR11  =  AB+ABrB  =  (AB)rB. B\,R_{00} \;+\; A\,R_{11} \;=\; \frac{A'B + AB'}{rB} \;=\; \frac{(AB)'}{rB}. (3.7.12)

Setting the vacuum equations (3.7.6) to zero makes the left-hand side vanish, and rr and BB are finite and non-zero in the region where we are solving, so

(AB)  =  0A(r)B(r)  =  constant. (AB)' \;=\; 0 \qquad\Longrightarrow\qquad A(r)\,B(r) \;=\; \text{constant}. (3.7.13)

That single line is why the Schwarzschild solution can be found by hand. A pair of coupled second-order equations has been replaced, exactly and with no approximation, by an algebraic relation between the two unknowns. It is worth noticing why the combination works: R00R_{00} and R11R_{11} differ from one another only in which of AA and BB sits in each denominator, so weighting by the opposite function makes the three complicated structures identical up to sign. Nothing was clever about the choice. It was forced by the shape of (3.7.10).

3.4 · The first constant, fixed by a physical demand

Equation (3.7.13) gives a constant, and algebra cannot say which one. Name the physical input that does: asymptotic flatness. Far from an isolated body the geometry should approach that of empty flat spacetime, so that a distant observer's clocks and rulers are the ordinary ones of Part II. In the chart (3.7.5) that means

A(r)1andB(r)1as r, A(r)\to1 \quad\text{and}\quad B(r)\to1 \qquad\text{as } r\to\infty, (3.7.14)

since with A=B=1A=B=1 the line element becomes c2dt2dr2r2dΩ2c^{2}\dd t^{2}-\dd r^{2}-r^{2}\dd\Omega^{2}, which is Minkowski spacetime in spherical coordinates: non-constant components and no curvature whatever, which is the lesson Chapter 3.3 §2 delivered early using the polar plane. The product ABAB therefore tends to 11, and being constant it equals 11 everywhere:

  B  =  1A.   \boxed{\;B \;=\; \frac{1}{A}.\;} (3.7.15)

One unknown function remains. Note carefully what was spent: not a calculation, but a decision about the physical situation being described. A different boundary condition would describe a different situation, and Problem 1 asks what a different constant would have meant.

3.5 · One integration

Put B=1/AB=1/A into the third line of (3.7.10). Each of the four terms simplifies, and it is worth doing them one at a time. With B=1/AB=1/A we have B=A/A2B'=-A'/A^{2} and B2=1/A2B^{2}=1/A^{2}, so

rB2B2  =  r2(AA2)A2  =  rA2,rA2AB  =  rA2AA  =  rA2,11B  =  1A. \begin{aligned} \frac{rB'}{2B^{2}} &\;=\; \frac{r}{2}\left(-\frac{A'}{A^{2}}\right)A^{2} \;=\; -\frac{rA'}{2}, \\[3pt] -\frac{rA'}{2AB} &\;=\; -\frac{rA'}{2A}\cdot A \;=\; -\frac{rA'}{2}, \\[3pt] 1-\frac{1}{B} &\;=\; 1-A. \end{aligned} (3.7.16)

The two derivative terms are equal, so they add rather than cancel, and

R22  =  rA    A  +  1  =  0rA+A  =  1. R_{22} \;=\; -\,rA' \;-\; A \;+\; 1 \;=\; 0 \qquad\Longleftrightarrow\qquad rA' + A \;=\; 1. (3.7.17)

Name the technique. The left-hand side is already a product rule: rA+A=(rA)rA'+A=(rA)'. So

  (rA)  =  1.   \boxed{\;\big(rA\big)' \;=\; 1.\;} (3.7.18)

This is a first-order linear equation of the kind Chapter 0.8 §2.2 handled with an integrating factor, and here the integrating factor has already done its work: rr is exactly the factor that turns A+A/r=1/rA'+A/r=1/r into an exact derivative. One integration, and one constant of integration C1C_{1}:

rA  =  r+C1A(r)  =  1  +  C1r. rA \;=\; r + C_{1} \qquad\Longrightarrow\qquad A(r) \;=\; 1 \;+\; \frac{C_{1}}{r}. (3.7.19)

Notice that (3.7.14) is satisfied automatically for any C1C_{1}: the solution is asymptotically flat whatever the constant is, so asymptotic flatness has nothing left to say and something else must fix C1C_{1}.

3.6 · The second constant, and where the mass comes in

Nothing in (3.7.6) mentions a mass. The equations are the equations of empty space, and they are the same equations whether the object at the centre is the Sun or a grain of dust. The mass can therefore enter in only one place: as the constant of integration, fixed by matching to something known.

The known thing is the Newtonian limit. Chapter 3.1 §6.5 derived, from the gravitational redshift in an accelerating cabin and with no general relativity used at all, that a weak static field has

g00  =  1  +  2Φc2, g_{00} \;=\; 1 \;+\; \frac{2\Phi}{c^{2}}, (3.7.20)

equation (3.1.39), with the plus sign belonging to this book's signature, and Chapter 3.6 §5 used the same component to fix κ\kappa. Far from a mass MM the Newtonian potential is Φ=GM/r\Phi=-GM/r, so

1+C1r  =  A  =  g00    12GMc2rC1  =  2GMc2. 1 + \frac{C_{1}}{r} \;=\; A \;=\; g_{00} \;\simeq\; 1 - \frac{2GM}{c^{2}r} \qquad\Longrightarrow\qquad C_{1} \;=\; -\,\frac{2GM}{c^{2}}. (3.7.21)

Both sides carry a term falling off as 1/r1/r and matching their coefficients fixes the constant. And MM is defined by this equation: it is the mass a distant observer would infer from the orbits of test bodies, and nothing in the derivation refers to what is inside. Substituting into (3.7.5) with (3.7.15),

  ds2  =  (12GMc2r)c2dt2    dr2  12GMc2r      r2(dθ2+sin2θdφ2).   \boxed{\;\dd s^{2} \;=\; \left(1-\frac{2GM}{c^{2}r}\right)c^{2}\dd t^{2} \;-\; \frac{\dd r^{2}}{\;1-\dfrac{2GM}{c^{2}r}\;} \;-\; r^{2}\Big(\dd\theta^{2}+\sin^{2}\theta\,\dd\varphi^{2}\Big).\;} (3.7.22)

That is the Schwarzschild metric, and it is exact. No expansion was made, no small parameter was introduced, and the only approximation anywhere in the derivation was in the matching of (3.7.21), which was performed at large rr where it is valid and which fixed a constant that is then exact everywhere.

Recap — what went in, what came out, what it cost

In: the two symmetry assumptions of §1, the vacuum field equations of Chapter 3.6, the Christoffel formula of Chapter 3.3, the Ricci definition of Chapter 3.4, asymptotic flatness, and the weak-field component of Chapter 3.1.

Out: an exact solution, with the entire influence of the central body carried by one number. Two very different solar systems with the same MM have identical geometries outside them, whatever their internal structure. Section 4 will strengthen that statement considerably.

What it cost: the ansatz. Everything downstream describes a non-rotating, spherical source. It cost no approximation, which is rarer in this book than it may look: Chapter 3.6 §7.3 explained that the field equations are nonlinear, that exact solutions are correspondingly rare, and that expanding about a simple solution is normally the only technique available. This is one of the exceptions, and it is the reason it is a landmark rather than an exercise.

3.7 · The Schwarzschild radius, as a length

The metric (3.7.22) contains one length, the value of rr at which the coefficient of c2dt2c^{2}\dd t^{2} vanishes:

  rs    2GMc2.   \boxed{\;r_{s} \;\equiv\; \frac{2GM}{c^{2}}.\;} (3.7.23)

Put numbers into it, because as a symbol it is easy to mistake for something enormous. With c2=8.988×1016 m2s2c^{2}=8.988\times10^{16}\ \mathrm{m^{2}\,s^{-2}} throughout:

The Sun. GM=1.3271×1020 m3s2GM_{\odot}=1.3271\times10^{20}\ \mathrm{m^{3}\,s^{-2}}, so

rs,  =  2(1.3271×1020)8.988×1016  =  2953 m  =  2.95 km. r_{s,\odot} \;=\; \frac{2\,(1.3271\times10^{20})}{8.988\times10^{16}} \;=\; 2953\ \mathrm{m} \;=\; 2.95\ \mathrm{km}. (3.7.24)

The Earth. Now run the same arithmetic with the Earth's measured value in place of the Sun's, GM=3.986×1014 m3s2GM_{\oplus}=3.986\times10^{14}\ \mathrm{m^{3}\,s^{-2}}, so

rs,  =  2(3.986×1014)8.988×1016  =  8.87×103 m  =  8.9 mm. r_{s,\oplus} \;=\; \frac{2\,(3.986\times10^{14})}{8.988\times10^{16}} \;=\; 8.87\times10^{-3}\ \mathrm{m} \;=\; 8.9\ \mathrm{mm}. (3.7.25)

Nine millimetres. That is the number worth carrying, because it converts a symbol into a length you could hold: compress the entire Earth inside a sphere the size of a small grape and the coefficient of c2dt2c^{2}\dd t^{2} at its surface reaches zero. And because rsr_{s} is linear in MM, the number scales straight up with the mass. A hole of 10910^{9} solar masses, the class imaged at the centre of M87, has rs=2.95×109 kmr_{s}=2.95\times10^{9}\ \mathrm{km}, which is 19.719.7 astronomical units, comparable to the orbit of Uranus.

The ratio that governs how big the relativistic corrections are is rs/rr_{s}/r. At the Sun's surface, rs,/R=2953/(6.957×108)=4.2×106r_{s,\odot}/R_{\odot}=2953/(6.957\times10^{8})=4.2\times10^{-6}: four parts in a million. That number is why Newtonian gravity survived three centuries of astronomy, and §8 shows what it takes to detect the difference anyway.

In plain terms 3.7.3

Three independent equations for two unknown functions is one equation more than the count allows, and a system in that position has no business being consistent. It is, and not by luck: the differentiated identity that shaped the field equations one chapter ago guarantees relations among them.

The solution turns on a single line. Take the equation belonging to the time direction and the one belonging to the radial direction, weight each by the other's factor, and add. Every second derivative cancels, every quadratic term cancels, and what survives is the rate of change of the product of the two unknowns, set equal to nothing. So that product is the same everywhere, and demanding that the geometry become the ordinary flat one far from the body makes it one. Two unknowns have collapsed into one, by a physical requirement rather than by algebra. The angular equation then says that the radial label times the surviving function has derivative one, the first kind of equation the toolkit taught anybody to integrate.

Nothing in the empty-space equations knows what a mass is; they are the same equations around a star and around a speck of dust. The constant left over from that integration acquires meaning only when the answer is set beside the old inverse-square law far away, and it is at that moment, and not earlier, that a mass enters at all.

4 · Birkhoff's theorem, and what it says about a collapsing star

Here is where this section is going. Section 1 assumed staticity and §3 used it in two places: to drop the x0x^{0} dependence and to drop the cross terms. We now throw the assumption away, keep only spherical symmetry, and redo the calculation. The result is that staticity comes back on its own, because the field equations force it. That is Birkhoff's theorem, and it is derived here rather than quoted.

4.1 · Redo it with time allowed in

Keep spherical symmetry, which by §2.1 still forces every sphere to be round, still forbids angular cross terms, and still permits the areal radius (3.7.4). Drop staticity entirely and allow both remaining functions to depend on time:

ds2  =  A(x0,r)(dx0)2    B(x0,r)dr2    r2dΩ2. \dd s^{2} \;=\; A(x^{0},r)\,\big(\dd x^{0}\big)^{2} \;-\; B(x^{0},r)\,\dd r^{2} \;-\; r^{2}\,\dd\Omega^{2}. (3.7.26)

Two remarks on what this ansatz does and does not assume, since the whole force of the theorem rests on them. It does not assume the matter is static: the source can be pulsing, collapsing or exploding, provided it stays spherical. It does still exclude a cross term g01g_{01}, and that is not a further assumption about the geometry. With g010g_{01}\neq0 the time part of the line element reads A(dx0)2+2g01dx0drA(\dd x^{0})^{2}+2g_{01}\dd x^{0}\dd r. Complete the square and supply an integrating factor, which exists because only one other variable is present. A function T(x0,r)T(x^{0},r) then appears for which the whole of it is A~c2dT2\tilde A\,c^{2}\dd T^{2}. That is a change of chart, precisely the operation Chapter 3.2 §2 defined, and it costs nothing.

Recompute. Every component of (3.7.8) that survives acquires possible time derivatives, three new connection coefficients appear carrying 0A\partial_{0}A and 0B\partial_{0}B, and one Ricci component that was identically zero before is no longer so. Grind box C computes it. The result is remarkably clean: writing B˙0B\dot B\equiv\partial_{0}B,

R01  =  B˙rB. R_{01} \;=\; \frac{\dot B}{r\,B}. (3.7.27)

4.2 · The theorem

Setting (3.7.27) to zero, and noting again that rr and BB are finite and non-zero where we are solving, gives

B˙  =  0B  =  B(r) alone. \dot B \;=\; 0 \qquad\Longrightarrow\qquad B \;=\; B(r) \ \text{alone}. (3.7.28)

One component of the vacuum equations has forbidden the radial function to depend on time, with no assumption made. The rest follows in three moves.

Move 1. With BB a function of rr alone, the combination of §3.3 goes through unchanged. Grind box C confirms that BR00+AR11B R_{00}+A R_{11} is still r(AB)/(rB)\partial_{r}(AB)/(rB) even when AA depends on time. So r(AB)=0\partial_{r}(AB)=0 and

A(x0,r)B(r)  =  f(x0) A(x^{0},r)\,B(r) \;=\; f(x^{0}) (3.7.29)

for some function of time alone. Note what has changed: the "constant" of §3.3 was constant in rr, and nothing said it could not vary with x0x^{0}.

Move 2. The time dependence is not in the geometry, it is in the clock. From (3.7.29), A=f(x0)/B(r)A=f(x^{0})/B(r), so the time part of the line element is f(x0)(dx0)2/B(r)f(x^{0})(\dd x^{0})^{2}/B(r). Define a new time coordinate TT by

cdT    f(x0)  dx0, c\,\dd T \;\equiv\; \sqrt{f(x^{0})}\;\dd x^{0}, (3.7.30)

which is legitimate because ff depends on x0x^{0} alone, so the right-hand side is an exact differential and TT is a genuine coordinate. Then f(x0)(dx0)2=c2dT2f(x^{0})(\dd x^{0})^{2}=c^{2}\dd T^{2} and

ds2  =  c2dT2B(r)    B(r)dr2    r2dΩ2, \dd s^{2} \;=\; \frac{c^{2}\dd T^{2}}{B(r)} \;-\; B(r)\,\dd r^{2} \;-\; r^{2}\dd\Omega^{2}, (3.7.31)

which has no time dependence at all and no cross terms: static. The whole of the apparent time dependence was a statement about how a distant observer had chosen to run their clock.

Move 3. Equation (3.7.31) is the ansatz (3.7.5) with A=1/BA=1/B already imposed, so §§3.5–3.6 apply verbatim: (rA)=1(rA)'=1, one integration, and the constant fixed by the Newtonian limit. The solution is (3.7.22) again.

  Birkhoff’s theorem.Every spherically symmetric solution of Rμν=0is the Schwarzschild solution; there are no others.   \boxed{\;\begin{gathered}\textbf{Birkhoff's theorem.}\\ \text{Every spherically symmetric solution of }R_{\mu\nu}=0\\ \text{is the Schwarzschild solution; there are no others.}\end{gathered}\;} (3.7.32)
Grind box C — the time-dependent components, and R01R_{01}

The new connection coefficients. With AA and BB both functions of (x0,r)(x^{0},r), three components join the list (3.7.8) and two acquire new pieces. From (3.7.7), with x˙0\dot{\phantom{x}}\equiv\partial_{0}:

Γ000=A˙2A,Γ011=B˙2A,Γ101=B˙2B, \Gamma^{0}{}_{00}=\frac{\dot A}{2A}, \qquad \Gamma^{0}{}_{11}=\frac{\dot B}{2A}, \qquad \Gamma^{1}{}_{01}=\frac{\dot B}{2B},

each obtained the same way as in grind box A, the only difference being which derivative survives in the bracket. For instance Γ101\Gamma^{1}{}_{01} takes λ=1\lambda=1, g11=1/Bg^{11}=-1/B, and the bracket 0g11+1g101g01=B˙\partial_{0}g_{11}+\partial_{1}g_{10}-\partial_{1}g_{01}=-\dot B, giving 12(1/B)(B˙)=B˙/2B\half(-1/B)(-\dot B)=\dot B/2B. The nine of (3.7.8) keep their forms with AA and BB now functions of two variables.

The 0101 component of the Ricci tensor. Put μ=0\mu=0, ν=1\nu=1 into (3.7.9):

R01  =  λΓλ10    1Γλλ0  +  ΓλλρΓρ10    Γλ1ρΓρλ0. R_{01} \;=\; \partial_{\lambda}\Gamma^{\lambda}{}_{10} \;-\; \partial_{1}\Gamma^{\lambda}{}_{\lambda 0} \;+\; \Gamma^{\lambda}{}_{\lambda\rho}\Gamma^{\rho}{}_{10} \;-\; \Gamma^{\lambda}{}_{1\rho}\Gamma^{\rho}{}_{\lambda 0}.

Take them in turn. The only non-zero Γλ10\Gamma^{\lambda}{}_{10} are Γ001=A/2A\Gamma^{0}{}_{01}=A'/2A and Γ101=B˙/2B\Gamma^{1}{}_{01}=\dot B/2B, so term 1 is 0(A/2A)+1(B˙/2B)\partial_{0}(A'/2A)+\partial_{1}(\dot B/2B). Using the contraction of grind box A, now reading Γλλ0=A˙/2A+B˙/2B\Gamma^{\lambda}{}_{\lambda 0}=\dot A/2A+\dot B/2B, term 2 is r(A˙/2A+B˙/2B)-\partial_{r}\big(\dot A/2A+\dot B/2B\big). The two B˙/2B\dot B/2B pieces cancel outright between terms 1 and 2, and the two remaining pieces, 0(A/2A)\partial_{0}(A'/2A) and r(A˙/2A)-\partial_{r}(\dot A/2A), cancel as well. The quotient rule turns the first into 0rA/2AAA˙/2A2\partial_{0}\partial_{r}A/2A-A'\dot A/2A^{2} and the second into r0A/2AA˙A/2A2\partial_{r}\partial_{0}A/2A-\dot A A'/2A^{2}, and these are the same expression, because mixed partial derivatives commute. Terms 1 and 2 together give zero.

Term 3 is ΓλλρΓρ10\Gamma^{\lambda}{}_{\lambda\rho}\Gamma^{\rho}{}_{10} summed over ρ=0,1\rho=0,1:

(A˙2A+B˙2B)A2A  +  (A2A+B2B+2r)B˙2B. \left(\frac{\dot A}{2A}+\frac{\dot B}{2B}\right)\frac{A'}{2A} \;+\; \left(\frac{A'}{2A}+\frac{B'}{2B}+\frac{2}{r}\right)\frac{\dot B}{2B}.

Term 4 is Γλ1ρΓρλ0-\Gamma^{\lambda}{}_{1\rho}\Gamma^{\rho}{}_{\lambda 0}, whose surviving contributions come from (λ,ρ)=(0,0),(0,1),(1,0),(1,1)(\lambda,\rho)=(0,0),(0,1),(1,0),(1,1) and the two angular pairs, the latter giving nothing because Γ220=Γ330=0\Gamma^{2}{}_{20}=\Gamma^{3}{}_{30}=0:

[A2AA˙2A(0,0)+B˙2AA2B(0,1)+B˙2BA2A(1,0)+B2BB˙2B(1,1)], -\left[\underbrace{\frac{A'}{2A}\frac{\dot A}{2A}}_{(0,0)} + \underbrace{\frac{\dot B}{2A}\frac{A'}{2B}}_{(0,1)} + \underbrace{\frac{\dot B}{2B}\frac{A'}{2A}}_{(1,0)} + \underbrace{\frac{B'}{2B}\frac{\dot B}{2B}}_{(1,1)}\right] ,

the two angular pairs contributing nothing because Γ220=Γ330=0\Gamma^{2}{}_{20}=\Gamma^{3}{}_{30}=0. Expand term 3 and set it beside this: the four products match one for one and cancel, and the only survivor is the piece term 3 carries with the 2/r2/r in it. Hence

R01  =  2rB˙2B  =  B˙rB, R_{01} \;=\; \frac{2}{r}\cdot\frac{\dot B}{2B} \;=\; \frac{\dot B}{rB},

which is (3.7.27).

(Confirmed symbolically for general A(x0,r)A(x^{0},r), B(x0,r)B(x^{0},r): R01R_{01} came out exactly B˙/(rB)\dot B/(rB), and with BB restricted to depend on rr alone the combination BR00+AR11BR_{00}+AR_{11} came out exactly r(AB)/(rB)\partial_{r}(AB)/(rB), as Move 1 claims.)

4.3 · What the theorem says about a star

The consequences are large and they are not obvious from the algebra.

A pulsating star has a static exterior. Take a spherical star that expands and contracts. A Cepheid variable does exactly this, and so does a star ringing after a disturbance. Outside it, the geometry is (3.7.22) with a fixed MM, unchanging, for all time. A planet orbiting it feels nothing at all.

There is no spherically symmetric gravitational radiation. A radiating system must have a time-dependent field outside it, and Birkhoff says a spherical one cannot. Electromagnetism forbids the corresponding thing, since a spherically pulsing charge does not radiate. But it forbids it by a weaker route, conservation of charge, which says nothing about the field elsewhere. Here the statement is stronger: the entire exterior geometry is unchanged, because it is the only solution the equations possess. That is why the gravitational waves detected in this century come from pairs of compact objects spiralling about each other, which are as far from spherical as a system can be, rather than from single objects ringing.

The field of a collapsing star does not change as it collapses. This one is worth stating slowly. A spherical star collapses through its own Schwarzschild radius. During the collapse the matter distribution changes violently. Outside the star's surface, the geometry is (3.7.22) with the same MM throughout, before, during and after. Nothing propagates outward, nothing announces the collapse, and a distant orbit is undisturbed. Whatever a horizon is, its formation is not an event that anything outside can feel. That makes the question of what it is considerably more interesting, and Chapter 3.8 asks it.

And a debt of Chapter 3.4 is collected. That chapter's §6.3 insisted that Ricci-flat does not mean flat, and named the Weyl part of the curvature as what survives contraction. Here is the example. Equation (3.7.22) has Rμν=0R_{\mu\nu}=0 everywhere it applies, since it was constructed to. And it is not flat: it bends light, holds planets in orbit, and §7 shows it does something no flat geometry could. Every bit of that is Weyl curvature. The tidal deviation of Chapter 3.1, which Chapter 3.4 §4 identified with curvature, is entirely carried by the part the Ricci tensor throws away.

In plain terms 3.7.4

Redo the calculation without assuming anything about time and the assumption comes back unbidden, cornered by the equations rather than imposed. Let both unknown functions depend on time as well as on the radial label, and one component of the empty-space equations, the one mixing time with radius, forbids the radial function any time dependence. The rest of the argument then runs unaltered, and the only surviving trace of time is absorbed by rescaling the distant observer's clock.

So the geometry outside any spherical arrangement of matter is the one already found, whatever that matter does. A star may pulse, collapse, rebound or detonate; if it stays spherical, nothing outside changes. There is therefore no spherically symmetric gravitational radiation, which is why the waves picked up in this century come from pairs of bodies swinging about each other rather than single ones ringing. And a star falling in on itself sends no warning outward: the field it leaves is the field it always had.

A second reading matters more for what follows: empty is not the same as flat. The solution has vanishing boiled-down curvature everywhere outside the mass, since that is what it was built to have, and it is nonetheless not flat: it bends light, holds planets, and further in acquires a surface crossable in one direction only. All of that lives in the part of the curvature the boiling-down discards.

5 · Two Killing vectors, two conserved quantities

The solution exists. Now spend it, which means asking how a body moves in it. Here is where this section is going. Free motion in a metric is four coupled second-order equations for t(τ),r(τ),θ(τ),φ(τ)t(\tau),r(\tau),\theta(\tau),\varphi(\tau), with position-dependent coefficients, and that is not a problem anybody solves by hand. We reduce it to one first-order equation in one variable, using symmetry three times: once to fix the plane, and twice to integrate an equation once for free. Chapter 3.5 §9 promised exactly this, twice and by name, and this section is where the promise falls due.

5.1 · The promise, and the two charges

Chapter 3.5 §8.1 gave the practical test: every coordinate the metric does not mention supplies a Killing vector. Look at (3.7.22). It does not mention tt, and it does not mention φ\varphi. So

ξ(t)  =  t,ξ(φ)  =  φ \xi_{(t)} \;=\; \partial_{t}, \qquad \xi_{(\varphi)} \;=\; \partial_{\varphi} (3.7.33)

are Killing vector fields. And Chapter 3.5 §9.1 proved, in four lines, that for any Killing field ξ\xi and any geodesic with tangent uμ=dxμ/dτu^{\mu}=\dd x^{\mu}/\dd\tau, the number ξμuμ\xi_{\mu}u^{\mu} is the same at every point of the geodesic. That is equation (3.5.73). Two Killing fields, two constants.

Compute them. Lowering an index means contracting with the metric, and the metric is diagonal, so ξμ=gμνξν\xi_{\mu}=g_{\mu\nu}\xi^{\nu} picks out one component each time.

The time charge. In the coordinates xμ=(x0,r,θ,φ)x^{\mu}=(x^{0},r,\theta,\varphi) with x0=ctx^{0}=ct, the field t\partial_{t} has components ξμ=(c,0,0,0)\xi^{\mu}=(c,0,0,0), since t=c0\partial_{t}=c\,\partial_{0}. Then ξ0=g00c=cA\xi_{0}=g_{00}\,c=cA and

ξμuμ  =  ξ0u0  =  cAcdtdτ  =  c2(12GMc2r)dtdτ    E. \xi_{\mu}u^{\mu} \;=\; \xi_{0}u^{0} \;=\; cA\cdot c\,\dv{t}{\tau} \;=\; c^{2}\left(1-\frac{2GM}{c^{2}r}\right)\dv{t}{\tau} \;\equiv\; E. (3.7.34)

The angular charge. Now do the same for the rotational Killing field. The field φ\partial_{\varphi} has components ξμ=(0,0,0,1)\xi^{\mu}=(0,0,0,1), so lowering the index gives ξ3=g33=r2sin2θ\xi_{3}=g_{33}=-r^{2}\sin^{2}\theta, and

ξμuμ  =  r2sin2θdφdτ    L. -\,\xi_{\mu}u^{\mu} \;=\; r^{2}\sin^{2}\theta\,\dv{\varphi}{\tau} \;\equiv\; L. (3.7.35)

The minus sign is a convention chosen so that LL is positive for motion in the direction of increasing φ\varphi. The conserved quantity is ξμuμ\xi_{\mu}u^{\mu} either way.

What they are. Chapter 3.5 §9.2 ran Chapter 1.4's Noether theorem on the same system and found that the charge belonging to ξ\xi is mξμuμm\,\xi_{\mu}u^{\mu}, which is equation (3.5.79). So mEmE is the Noether charge of time-translation invariance and mLmL is the Noether charge of rotational invariance: EE is the energy per unit mass and LL the angular momentum per unit mass, and they are the same two quantities Chapter 1.4 §3 obtained for planets long before any of this geometry existed. Their names were not assigned by analogy. They were derived and then recognised.

Two sanity checks on (3.7.34) and (3.7.35)

Far away. As rr\to\infty the factor in (3.7.34) tends to 11 and dt/dτγ\dd t/\dd\tau\to\gamma, so Eγc2E\to\gamma c^{2}, which is Chapter 2.5's energy per unit mass including the rest energy. A body at rest at infinity has E=c2E=c^{2} exactly. A bound body has E<c2E\lt c^{2}.

In the equatorial plane. With θ=π/2\theta=\pi/2, equation (3.7.35) reads L=r2dφ/dτL=r^{2}\dd\varphi/\dd\tau. That is Kepler's second law, equal areas in equal times, with proper time in place of coordinate time. Chapter 3.5 Worked example 2 obtained the same expression on the flat plane from the same Killing field, and Chapter 1.4 §3.3 obtained it from rotational symmetry. Three routes, one quantity.

5.2 · The equatorial plane, and why taking it costs nothing

Both charges still contain θ\theta. Dispose of it, and say why the disposal is legitimate rather than convenient.

The argument is a symmetry argument and it has two steps. First, take any geodesic and any point pp on it. The tangent uu at pp and the radial direction at pp span a plane through the centre. By spherical symmetry there is a rotation carrying that plane to the equatorial plane θ=π/2\theta=\pi/2, and a rotation is an isometry, so it carries geodesics to geodesics. Therefore every geodesic can be rotated so that at pp it lies in the equatorial plane with no motion out of it, dθ/dτ=0\dd\theta/\dd\tau=0.

Second, check that it stays there. The θ\theta component of the geodesic equation, using Γ212=1/r\Gamma^{2}{}_{12}=1/r and Γ233=sinθcosθ\Gamma^{2}{}_{33}=-\sin\theta\cos\theta from (3.7.8), is

d2θdτ2  +  2rdrdτdθdτ    sinθcosθ(dφdτ)2  =  0. \dvn{2}{\theta}{\tau} \;+\; \frac{2}{r}\,\dv{r}{\tau}\dv{\theta}{\tau} \;-\; \sin\theta\cos\theta\left(\dv{\varphi}{\tau}\right)^{2} \;=\; 0. (3.7.36)

At θ=π/2\theta=\pi/2 the last term vanishes because cos(π/2)=0\cos(\pi/2)=0, and if dθ/dτ=0\dd\theta/\dd\tau=0 there as well then the middle term vanishes too, leaving d2θ/dτ2=0\dd^{2}\theta/\dd\tau^{2}=0. Zero value, zero first derivative and zero second derivative: by the uniqueness theorem for the initial-value problem, Chapter 0.8 §1.1, the solution is θπ/2\theta\equiv\pi/2 for all τ\tau.

This is exactly the statement that a planet's orbit lies in a plane, and the reason is exactly the Newtonian one: Chapter 1.4 §3.3's conservation of the angular momentum vector, whose fixed direction is the normal to the orbital plane. From here on, θ=π/2\theta=\pi/2 and (3.7.35) reads L=r2dφ/dτL=r^{2}\dd\varphi/\dd\tau.

5.3 · The count, before and after

UnknownsEquations
Geodesic equation as writtent,r,θ,φt,r,\theta,\varphifour, coupled, second order
after spherical symmetry (§5.2)t,r,φt,r,\varphithree, second order
after the two Killing chargesrrone, first order

That is the whole of the reduction, and it was bought entirely with symmetry. Chapter 3.5 §9 said it in advance: "Without this section, Chapter 3.7 would be a wall of algebra with no visible structure. With it, that chapter is mostly Chapter 1.3's phase-plane reasoning applied to one variable." The next section supplies the one remaining equation.

In plain terms 3.7.5

A promise made two chapters ago falls due here, and it is collected in full. The rule for measuring distances outside a spherical mass mentions neither the time label nor the angle around the axis, and by the practical test established there, every label the rule fails to mention supplies a direction along which the geometry does not change. Each hands over a number that is the same at every point of every freely falling path.

The two numbers are the energy per unit mass and the angular momentum per unit mass. They are the same two constants the theorem relating symmetries to conservation produced for planets long before this geometry existed, and nothing about them was assumed; they fall out of the metric and are then recognised. The second, written out, is the equal-areas rule of the seventeenth century, with the traveller's own clock in place of the astronomer's.

The saving is best stated as arithmetic. Free motion is four coupled second-order equations for four unknown functions of the traveller's own time. Spherical symmetry turns any orbit into a single plane, disposing of one. The two conserved quantities integrate two more, once and for all. What is left is one first-order equation in one variable, of exactly the form a marble rolling in a valley obeys. That is not computational cleverness; it is symmetry, converted into constants by a theorem proved for pendulums.

6 · The effective potential

Here is where this section is going. One line of algebra turns the normalisation of the four-velocity into an energy equation for r(τ)r(\tau). We then read that equation as a particle moving in a landscape, compare the landscape with Newton's term by term, and find that exactly one term is new. Everything in §§7 and 8 is that term's doing.

6.1 · One line of algebra

The four-velocity of a massive body satisfies

gμνuμuν  =  c2, g_{\mu\nu}u^{\mu}u^{\nu} \;=\; c^{2}, (3.7.37)

which is Chapter 2.5 §1's uu=c2u\cdot u=c^{2}, equation (2.5.9), carried onto a manifold in Chapter 3.3 §8. There it was shown to be constant along any geodesic, so that imposing it once at one point imposes it everywhere. Write it out for (3.7.22) in the equatorial plane, with f(r)12GM/c2rf(r)\equiv1-2GM/c^{2}r as shorthand:

f(u0)2    1f(drdτ)2    r2(dφdτ)2  =  c2. f\,\big(u^{0}\big)^{2} \;-\; \frac{1}{f}\left(\dv{r}{\tau}\right)^{2} \;-\; r^{2}\left(\dv{\varphi}{\tau}\right)^{2} \;=\; c^{2}. (3.7.38)

Now substitute the two charges. From (3.7.34), u0=cdt/dτ=E/(cf)u^{0}=c\,\dd t/\dd\tau=E/(cf), so the first term is fE2/(c2f2)=E2/(c2f)f\cdot E^{2}/(c^{2}f^{2})=E^{2}/(c^{2}f). From (3.7.35), dφ/dτ=L/r2\dd\varphi/\dd\tau=L/r^{2}, so the third term is L2/r2L^{2}/r^{2}. Hence

E2c2f    1f(drdτ)2    L2r2  =  c2. \frac{E^{2}}{c^{2}f} \;-\; \frac{1}{f}\left(\dv{r}{\tau}\right)^{2} \;-\; \frac{L^{2}}{r^{2}} \;=\; c^{2}. (3.7.39)

We want the radial derivative on its own, since r(τ)r(\tau) is the one unknown we are chasing. So multiply through by ff and rearrange for the radial term:

  (drdτ)2  =  E2c2    (12GMc2r) ⁣(c2+L2r2).   \boxed{\;\left(\dv{r}{\tau}\right)^{2} \;=\; \frac{E^{2}}{c^{2}} \;-\; \left(1-\frac{2GM}{c^{2}r}\right)\!\left(c^{2}+\frac{L^{2}}{r^{2}}\right).\;} (3.7.40)

That is the whole of orbital dynamics in general relativity. One equation, first order, one unknown function, two constants. Everything from here to the end of the chapter is extracted from it.

6.2 · Reading it as a landscape

Expand the product in (3.7.40), one term at a time. The four products are c2c^{2}, L2/r2L^{2}/r^{2}, 2GM/r-2GM/r and 2GML2/(c2r3)-2GML^{2}/(c^{2}r^{3}), so

(drdτ)2  =  E2c2c2  +  2GMr    L2r2  +  2GML2c2r3. \left(\dv{r}{\tau}\right)^{2} \;=\; \frac{E^{2}}{c^{2}} - c^{2} \;+\; \frac{2GM}{r} \;-\; \frac{L^{2}}{r^{2}} \;+\; \frac{2GML^{2}}{c^{2}r^{3}}. (3.7.41)

Divide by two and move everything except the derivative to the right:

12(drdτ)2  +  Veff(r)  =  E,E    E2c42c2, \half\left(\dv{r}{\tau}\right)^{2} \;+\; V_{\text{eff}}(r) \;=\; \mathcal{E}, \qquad \mathcal{E} \;\equiv\; \frac{E^{2}-c^{4}}{2c^{2}}, (3.7.42)

Everything depending on position alone has been gathered into the single function VeffV_{\text{eff}}. Here is what that function is:

  Veff(r)  =  GMr  +  L22r2    GML2c2r3.   \boxed{\;V_{\text{eff}}(r) \;=\; -\,\frac{GM}{r} \;+\; \frac{L^{2}}{2r^{2}} \;-\; \frac{GML^{2}}{c^{2}r^{3}}.\;} (3.7.43)

Equation (3.7.42) is the statement that a quantity quadratic in a velocity, plus a function of position alone, is constant. That is the form Chapter 1.3 §4 used to draw phase portraits, and the form Chapter 0.8 §4.3 used for the oscillator. The technique is read off the turning points: motion is possible only where EVeff(r)\mathcal{E}\ge V_{\text{eff}}(r), the body turns around where the two are equal, and it orbits in a circle where VeffV_{\text{eff}} is stationary.

Two cautions about what E\mathcal{E} is, since it looks like an energy and is not quite one. It is built from EE, which already includes the rest energy per unit mass c2c^{2}. The combination (E2c4)/2c2(E^{2}-c^{4})/2c^{2} removes it, and for a slow body Ec2+εE\simeq c^{2}+\varepsilon gives Eε\mathcal{E}\simeq\varepsilon, the Newtonian energy per unit mass. And the "time" in dr/dτ\dd r/\dd\tau is the traveller's proper time, not the distant observer's tt. Both facts matter and neither affects the shape of VeffV_{\text{eff}}.

6.3 · Term by term against Newton

Build the Newtonian comparison from scratch, in two lines, so that nothing is being taken on trust. Newtonian energy per unit mass for a body in a plane is ε=12(r˙2+r2φ˙2)GM/r\varepsilon=\half\big(\dot r^{2}+r^{2}\dot\varphi^{2}\big)-GM/r, and Chapter 1.4 §3.3's angular momentum per unit mass is L=r2φ˙L=r^{2}\dot\varphi, so r2φ˙2=L2/r2r^{2}\dot\varphi^{2}=L^{2}/r^{2}. Eliminating φ˙\dot\varphi,

ε  =  12r˙2  +  VN(r),VN(r)  =  GMr  +  L22r2. \varepsilon \;=\; \half\dot r^{2} \;+\; V_{\text{N}}(r), \qquad V_{\text{N}}(r) \;=\; -\,\frac{GM}{r} \;+\; \frac{L^{2}}{2r^{2}}. (3.7.44)

The second term is the centrifugal barrier, and it is a piece of kinetic energy wearing a potential's clothes. Chapter 1.2's Problem 3 on the bead on a rotating hoop already made that point, and it is the reason no fictitious force has to be invoked to explain the barrier.

Set the two side by side.

TermNewtonSchwarzschildWhat it does
attractionGM/r-GM/rGM/r-GM/ridentical, and it pulls inward
centrifugal barrier+L2/2r2+L^{2}/2r^{2}+L2/2r2+L^{2}/2r^{2}identical, and it keeps a rotating body out
the new termGML2/c2r3-GML^{2}/c^{2}r^{3}attractive, and it beats the barrier at small rr

The first two terms are not approximately the same as Newton's. They are the same, with no correction of any order. The entire difference between the two theories, for a massive body in orbit, is the third term.

Three things about it deserve naming before §7 exploits them.

It is attractive. The minus sign means it deepens the well, and it deepens it in the region where the barrier was supposed to be doing its job.

It carries L2L^{2}. A body falling straight in, L=0L=0, does not feel it at all. Problem 2 shows that radial infall in this metric obeys the Newtonian equation exactly. The new physics is entirely about bodies with sideways motion.

It falls off as 1/r31/r^{3}, so it wins at small rr. Compare it with the barrier: the ratio is

GML2/c2r3L2/2r2  =  2GMc2r  =  rsr, \frac{\big|{-GML^{2}}/{c^{2}r^{3}}\big|}{L^{2}/2r^{2}} \;=\; \frac{2GM}{c^{2}r} \;=\; \frac{r_{s}}{r}, (3.7.45)

independent of LL altogether. So the new term is a fraction rs/rr_{s}/r of the barrier: negligible far out, comparable at rrsr\sim r_{s}, dominant inside. Take Mercury's perihelion, a distance a(1e)=4.60×1010 ma(1-e)=4.60\times10^{10}\ \mathrm{m} from the orbital elements §8.5 lists. There rs,/r=2953/(4.60×1010)=6.4×108r_{s,\odot}/r=2953/(4.60\times10^{10})=6.4\times10^{-8}, which is the size of the effect §8 has to dig out of the data.

In plain terms 3.7.6

A freely falling body's velocity has a fixed length, and squaring that statement with the two constants in hand gives one equation relating the radial position to the traveller's own time. Rearranged, it says precisely what an elementary mechanics problem says about a marble in a valley: a term quadratic in the rate of change, plus a landscape depending on position alone, adds up to a fixed number. The apparatus of turning points applies unaltered.

Set that landscape beside the Newtonian one term by term. The attraction is there, unchanged. The barrier that keeps a body with sideways motion from reaching the centre is there, unchanged, and it is what made orbits stable for three centuries. Then there is a third term with no Newtonian counterpart whatever: attractive like the first, carrying the square of the angular momentum like the second, and falling off as the cube of the distance: undetectable far out and decisive close in.

That one extra term is the entire difference between the two theories for a body in orbit, and everything the chapter extracts afterwards comes from it. Notice what it does to the barrier. The barrier still rises as a body comes inward, and then, at short enough range, the new term wins and the barrier collapses. A wall that fails at close quarters is a different sort of object from a wall.

7 · The innermost stable circular orbit

Here is where this section is going, and it is the chapter's thesis. Circular orbits sit where VeffV_{\text{eff}} is stationary. We solve dVeff/dr=0\dd V_{\text{eff}}/\dd r=0 and find a quadratic, which therefore has two roots, one root or none. In Newtonian gravity the same question has exactly one root for every radius. The discriminant tells us when the roots exist, and the answer is that below a certain angular momentum there is no circular orbit at any radius whatever. The two roots merge at r=6GM/c2r=6GM/c^{2}, and that radius has a name and an observational consequence.

7.1 · Where the landscape is level

Differentiate (3.7.43) term by term:

dVeffdr  =  GMr2    L2r3  +  3GML2c2r4. \dv{V_{\text{eff}}}{r} \;=\; \frac{GM}{r^{2}} \;-\; \frac{L^{2}}{r^{3}} \;+\; \frac{3GML^{2}}{c^{2}r^{4}}. (3.7.46)

Set it to zero and multiply by r4r^{4}, which is legitimate for r>0r>0 and clears every denominator at once:

  GMr2    L2r  +  3GML2c2  =  0.   \boxed{\;GM\,r^{2} \;-\; L^{2}\,r \;+\; \frac{3GML^{2}}{c^{2}} \;=\; 0.\;} (3.7.47)

A quadratic in rr. Stop and compare with Newton before solving it. Differentiating (3.7.44) gives GM/r2L2/r3=0GM/r^{2}-L^{2}/r^{3}=0, which after multiplying by r3r^{3} is linear:

GMr    L2  =  0r  =  L2GM. GM\,r \;-\; L^{2} \;=\; 0 \qquad\Longrightarrow\qquad r \;=\; \frac{L^{2}}{GM}. (3.7.48)

One root, always, for every L>0L>0. Turn it round: name any radius you like and L=GMrL=\sqrt{GMr} puts a circular orbit there. Newtonian gravity has a circular orbit at every radius, and each is a minimum of VNV_{\text{N}}, hence stable. Equation (3.7.47) is a different kind of object, and the difference is the term carrying 1/c21/c^{2}.

7.2 · The discriminant, and a floor on the angular momentum

Solve (3.7.47) by the quadratic formula, with a=GMa=GM, b=L2b=-L^{2} and   c0=3GML2/c2\;c_{0}=3GML^{2}/c^{2}:

r±  =  L2±L412G2M2L2/c22GM  =  L(Lc±L2c212G2M2)2GMc, r_{\pm} \;=\; \frac{L^{2}\pm\sqrt{L^{4}-12G^{2}M^{2}L^{2}/c^{2}}}{2GM} \;=\; \frac{L\Big(Lc\pm\sqrt{L^{2}c^{2}-12G^{2}M^{2}}\Big)}{2GMc}, (3.7.49)

the second form obtained by taking one factor of LL out of the square root, which is legal since L>0L>0, and then a factor of 1/c1/c out of what remains.

Real roots require the quantity under the square root to be non-negative:

L2c2    12G2M2    0L    Lmin    23GMc. L^{2}c^{2} \;-\; 12G^{2}M^{2} \;\ge\; 0 \qquad\Longleftrightarrow\qquad L \;\ge\; L_{\min} \;\equiv\; 2\sqrt3\,\frac{GM}{c}. (3.7.50)

Read that as a physical statement, because it is a startling one.

⚠ What "no circular orbit" does and does not mean

If L<23GM/cL\lt2\sqrt3\,GM/c, equation (3.7.47) has no real solution. There is then no radius whatever at which a body with that angular momentum can move in a circle. Not an unstable circle, not a marginal one, none. The effective potential has no stationary point at all, and Problem 3 asks you to confirm that in that case it increases monotonically with rr, so that a body released anywhere spirals inward and is captured.

What it does not mean. It does not mean a body with small angular momentum cannot exist far from the mass. It can, and it will be on a plunging trajectory rather than an orbit. It does not mean the body is moving faster than light or doing anything else forbidden. And it is not a statement about small rr alone. The failure is global, which is exactly what makes it different from the Newtonian case.

Newtonian gravity has no analogue of this at all. There, (3.7.48) supplies a circular orbit for every L>0L>0 and every radius, without exception, and the barrier L2/2r2L^{2}/2r^{2} always eventually wins over GM/r-GM/r as r0r\to0. Here the new term wins instead.

7.3 · The two roots, and where they merge

When both roots exist they are different orbits with different characters. Factorise (3.7.46) using them: since the numerator of dVeff/dr\dd V_{\text{eff}}/\dd r is GM(rr)(rr+)GM(r-r_{-})(r-r_{+}) divided by r4r^{4}, with r<r+r_{-}\lt r_{+}, the derivative is positive for r<rr\lt r_{-}, negative between the roots and positive again for r>r+r>r_{+}. So

  • rr_{-} is a maximum of VeffV_{\text{eff}}, so it is an unstable circular orbit, destroyed by any disturbance.
  • r+r_{+} is a minimum, so it is a stable circular orbit, which a nudged body oscillates about.

As LL decreases toward LminL_{\min} the square root in (3.7.49) shrinks and the two roots approach each other. At L=LminL=L_{\min} it vanishes and they coincide:

rISCO  =  Lmin22GM  =  12GM12G2M2c2  =    6GMc2  =  3rs.   r_{\text{ISCO}} \;=\; \frac{L_{\min}^{2}}{2GM} \;=\; \frac{1}{2GM}\cdot\frac{12G^{2}M^{2}}{c^{2}} \;=\; \boxed{\;\frac{6GM}{c^{2}} \;=\; 3\,r_{s}.\;} (3.7.51)

This is the innermost stable circular orbit, the ISCO. At exactly rISCOr_{\text{ISCO}} both dVeff/dr\dd V_{\text{eff}}/\dd r and d2Veff/dr2\dd^{2}V_{\text{eff}}/\dd r^{2} vanish, because the maximum and the minimum have annihilated into an inflection. So the orbit there is marginally stable: neither restored nor expelled at leading order. Inside it there are no stable circular orbits at all, and Problem 3 shows that the unstable branch rr_{-} runs from rISCOr_{\text{ISCO}} inward and never reaches below 3GM/c23GM/c^{2}.

Recap — two lines of algebra, one new fact about the world

In: the effective potential (3.7.43), whose only departure from Newton is one term, together with the definition of a circular orbit as a stationary point.

Out: a chain of consequences, each one a single step from the last.

  • A quadratic instead of a linear equation.
  • Hence two circular orbits instead of one at each angular momentum, an unstable inner one and a stable outer one.
  • Hence a discriminant, and hence a minimum angular momentum 23GM/c2\sqrt3\,GM/c below which circular motion is impossible everywhere.
  • Hence a smallest stable circular orbit at 6GM/c26GM/c^{2}, whose position depends on the mass and on nothing else.

What it cost: nothing beyond §6. No approximation was made anywhere in §7. The ISCO is an exact statement about an exact solution, unlike the precession of §8, which is perturbative.

7.4 · The number, and what it is the inner edge of

Put a mass in. Take a black hole of ten solar masses, the class produced by the collapse of a massive star. Its ISCO sits at

rISCO  =  6(10)(1.3271×1020)8.988×1016  =  8.86×104 m  =  88.6 km, r_{\text{ISCO}} \;=\; \frac{6\,(10)\,(1.3271\times10^{20})}{8.988\times10^{16}} \;=\; 8.86\times10^{4}\ \mathrm{m} \;=\; 88.6\ \mathrm{km}, (3.7.52)

with rs=29.5 kmr_{s}=29.5\ \mathrm{km} for comparison. And Lmin=23GM/c=3.464GM/cL_{\min}=2\sqrt3\,GM/c=3.464\,GM/c, which for the same hole is 3.464×(1.3271×1021)/(2.998×108)=1.53×1013 m2s13.464\times(1.3271\times10^{21})/(2.998\times10^{8})=1.53\times10^{13}\ \mathrm{m^{2}\,s^{-1}}.

The observational hook. Gas falling toward a compact object does not fall straight in. It has angular momentum, so it settles into a disc and spirals slowly inward as friction removes that angular momentum. At each stage it occupies the stable circular orbit appropriate to its current LL. When it reaches rISCOr_{\text{ISCO}} there is nowhere further to go: the next orbit inward does not exist, and the gas leaves the disc and plunges. An accretion disc therefore has an inner edge, at 6GM/c26GM/c^{2}, and since the temperature and the emitted spectrum depend on how deep the gas gets, the inner edge is a measurable feature and it is one of the ways the mass of a black hole is estimated. Worked example 2 computes how much energy is released on the way in, and the answer is startling.

Familiar ground — two equilibria annihilating, which you have seen in a different subject

The mathematics of §7.2 is not special to gravity, and you have met it in a form closer to home. Take a tumour growing logistically and treated at a fixed absolute rate, meaning a therapy that removes DD cells per unit time independent of burden. The population obeys dN/dt=ρN(1N/K)D\dd N/\dd t=\rho N\big(1-N/K\big)-D, and its equilibria satisfy ρN2/KρN+D=0\rho N^{2}/K-\rho N+D=0: a quadratic, with roots

N±  =  K2(1±14DρK). N_{\pm} \;=\; \frac{K}{2}\left(1\pm\sqrt{1-\frac{4D}{\rho K}}\,\right).

Everything §7 says transfers line for line. Two equilibria exist when the discriminant is positive: a lower unstable one, below which the tumour is driven to extinction and above which it grows back, and an upper stable one, the burden the treatment holds it at. As DD rises the two approach each other. At D=ρK/4D=\rho K/4 they merge at N=K/2N=K/2, and above that there is no equilibrium at any burden whatever and the population is driven to zero from any starting point. The merged root is where both the first and the second derivative of the right-hand side vanish, exactly as at the ISCO. That merging is the reason a dose–response curve for such a therapy has a threshold rather than a gradual improvement, and it is the same algebra as (3.7.50).

What breaks. Two things, and both matter. In the tumour model DD is a control that an oncologist turns. Here LL is a constant of the motion, a label on a whole family of landscapes, and nothing turns it. A given body keeps its LL forever unless something exerts a torque, which is why the slow inward drift of a disc requires friction and does not happen on its own. And the tumour model is a first-order equation whose equilibria are where the rate vanishes, while (3.7.42) is second order and its equilibria are where a potential is stationary. The shared object is the quadratic and its discriminant, and it is shared exactly. The dynamics either side of it are different.

4.100
L = 4.1000 GM/c → discriminant L² − 12 = 4.81000 ≥ 0 : circular orbits exist
unstable circular orbit at r = 3.90900 · stable circular orbit at r = 12.90100 · barrier height V(r₋) − V(r₊) = 0.047647
closed form of §7.2 : r₋ = 3.90900 , r₊ = 12.90100 |numerical − formula| = 5.7e-10 , 1.0e-9
The barrier that fails. Inside this figure only, lengths are in units of GM/c2GM/c^{2}, angular momentum in units of GM/cGM/c and the potential in units of c2c^{2}, so that the Schwarzschild radius is at 22 and the ISCO at 66; nothing else changes. The dashed grey curve is the Newtonian effective potential (3.7.44) and the solid blue curve is (3.7.43); they differ by the single term L2/r3-L^{2}/r^{3}. Both are computed point by point from those formulas, and the marked extrema are found by bisecting the numerically evaluated derivative rather than by plotting the closed-form roots — the third readout compares the two, and their agreement is (3.7.49) checked rather than assumed. Slide LL downward and watch the two blue extrema, the unstable maximum and the stable minimum, travel toward each other. At L=233.464L=2\sqrt3\approx3.464 they meet at r=6r=6 and vanish; below that the blue curve rises monotonically from -\infty, there is no barrier anywhere, and every trajectory reaches the centre. The grey curve does none of this — its single extremum sits at r=L2r=L^{2} for every LL, however small, and the Newtonian barrier always wins in the end.
In plain terms 3.7.7

Circular orbits sit where the landscape is level. Setting its slope to nothing and clearing denominators leaves a quadratic, a different sort of object from what the old theory offers, where the same question has exactly one answer for every radius anyone cares to name. A quadratic has two solutions, or one, or none, and which occurs is decided by how much angular momentum the body carries.

Two is the ordinary case: an outer minimum, where a body sits stably and returns after a nudge, and an inner maximum, a circular orbit any disturbance destroys. As the angular momentum is reduced the two travel toward one another. At one particular value they meet, and below it there is no solution at all: no circular orbit anywhere, at any radius, however the body is arranged. The landscape has turned into a slide.

The meeting point sits at three times the length fixed by the mass alone, and nothing in Newtonian gravity resembles it. There, every radius admits a circular orbit and every one of them is stable. Here there is a floor, its position set by the mass and by nothing else, and gas spiralling inward must stop orbiting and fall when it arrives. That is a prediction about where the inner edge of a disc around a black hole lies, and why such discs have an inner edge at all.

8 · The orbit equation, and perihelion precession

Here is where this section is going. Section 6's equation describes rr against proper time, which is not what an astronomer measures. We change variables so that it describes the shape of the orbit, meaning rr against φ\varphi. We derive the Newtonian version first, where the answer is a closed ellipse. The relativistic version differs by one term. Linearising about a circular orbit turns that term into a change of frequency rather than of shape, so the ellipse still repeats but slightly slower than once per turn, and the whole figure rotates. Then we compute the rate for Mercury, with every conversion on the page.

8.1 · From motion to shape

Two changes of variable, and both are standard for central-force problems because both remove a nuisance.

Change one: use φ\varphi as the independent variable instead of τ\tau. The chain rule and (3.7.35) give the conversion

ddτ  =  dφdτddφ  =  Lr2ddφ, \dv{}{\tau} \;=\; \dv{\varphi}{\tau}\,\dv{}{\varphi} \;=\; \frac{L}{r^{2}}\,\dv{}{\varphi}, (3.7.53)

which is available precisely because LL is conserved. That is a Killing vector paying for itself a second time.

Change two: use u1/ru\equiv1/r instead of rr. The motivation is that the attraction goes as 1/r21/r^{2} and the barrier as 1/r21/r^{2}, so in terms of uu the equation will have polynomial coefficients. Then r=1/ur=1/u, and

drdτ  =  Lr2drdφ  =  Lu2(1u2dudφ)  =  Ldudφ, \dv{r}{\tau} \;=\; \frac{L}{r^{2}}\dv{r}{\varphi} \;=\; Lu^{2}\cdot\left(-\frac{1}{u^{2}}\dv{u}{\varphi}\right) \;=\; -\,L\,\dv{u}{\varphi}, (3.7.54)

the middle step using dr/dφ=d(1/u)/dφ=u2du/dφ\dd r/\dd\varphi=\dd(1/u)/\dd\varphi=-u^{-2}\dd u/\dd\varphi. The factors of u2u^{2} cancel exactly, which is why the substitution is worth making.

Put (3.7.54) into (3.7.41) and divide by L2L^{2}:

(dudφ)2  =  E2/c2c2L2  +  2GML2u    u2  +  2GMc2u3. \left(\dv{u}{\varphi}\right)^{2} \;=\; \frac{E^{2}/c^{2}-c^{2}}{L^{2}} \;+\; \frac{2GM}{L^{2}}\,u \;-\; u^{2} \;+\; \frac{2GM}{c^{2}}\,u^{3}. (3.7.55)

We want a second-order equation for the shape rather than a first-order one carrying a square, because the second-order kind is the one we can recognise on sight. So differentiate with respect to φ\varphi and divide through by 2du/dφ2\,\dd u/\dd\varphi:

  d2udφ2  +  u  =  GML2  +  3GMc2u2.   \boxed{\;\dvn{2}{u}{\varphi} \;+\; u \;=\; \frac{GM}{L^{2}} \;+\; \frac{3GM}{c^{2}}\,u^{2}.\;} (3.7.56)

The division by du/dφ\dd u/\dd\varphi is legitimate wherever that derivative is non-zero, which is everywhere except at the turning points and on exactly circular orbits. And a circular orbit, having uu constant, satisfies (3.7.56) directly, so nothing has been lost.

8.2 · Newton first, and why an ellipse closes

Delete the term carrying 1/c21/c^{2}. That is exactly the Newtonian limit, since the deleted term is the only place cc appears. Then (3.7.56) becomes

d2udφ2  +  u  =  GML2. \dvn{2}{u}{\varphi} \;+\; u \;=\; \frac{GM}{L^{2}}. (3.7.57)

This is the simplest differential equation in physics. It is the harmonic oscillator of Chapter 0.8 §4 with a constant driving term, and its general solution is a particular constant plus the homogeneous solution:

u(φ)  =  GML2(1+ecosφ),that isr(φ)  =  p1+ecosφ,p    L2GM, \begin{aligned} u(\varphi) \;&=\; \frac{GM}{L^{2}}\Big(1 + e\cos\varphi\Big), \\[3pt] \text{that is}\qquad r(\varphi) \;&=\; \frac{p}{1+e\cos\varphi}, \qquad p \;\equiv\; \frac{L^{2}}{GM}, \end{aligned} (3.7.58)

with the two integration constants written as an amplitude ee and a choice of where φ=0\varphi=0 sits, taken at the point of closest approach. That is the polar equation of a conic, an ellipse for e<1e\lt1. It is the same expression Chapter 1.4's Worked example 2 obtained by an entirely different route, from the Laplace–Runge–Lenz vector, with p=L2/mkp=L^{2}/mk in that chapter's variables.

Why it closes. The right-hand side of (3.7.58) has period exactly 2π2\pi in φ\varphi, because the coefficient of uu in (3.7.57) is exactly 11. So after one full turn the body is at the same radius moving the same way, and the orbit is a closed curve traced over and over. Chapter 1.4 said the same thing in the language of symmetry: the Laplace–Runge–Lenz vector points from the focus to the perihelion, it is conserved for an exact inverse-square law and for nothing else, and its conservation is the statement that the axis of the ellipse never moves. That chapter then wrote, in advance: "Adding any deviation from exact 1/r1/r destroys A\vv A and the orbit precesses. That is exactly what the λ\lambda dial did to the closed orbit in §4.4's figure, and exactly what the Sun's other planets, and then general relativity, do to Mercury (Chapter 3.7)." Here is the deviation.

8.3 · The new term, linearised

Restore it. The full equation (3.7.56) is nonlinear, and we are not going to solve it exactly. Instead we do what Chapter 3.6 §7.3 warned would become the normal state of affairs once the field equations turned out to be nonlinear: expand about a simple solution and keep the first correction, naming the small parameter as we go.

Step 1. The circular solution. A constant u=ucu=u_{c} has d2u/dφ2=0\dd^{2}u/\dd\varphi^{2}=0, so (3.7.56) requires

uc  =  GML2  +  3GMc2uc2, u_{c} \;=\; \frac{GM}{L^{2}} \;+\; \frac{3GM}{c^{2}}\,u_{c}^{2}, (3.7.59)

which is §7's circular-orbit condition wearing different symbols. Multiply by rc2=1/uc2r_{c}^{2}=1/u_{c}^{2} and it becomes (3.7.47) exactly.

Step 2. Perturb. Write u=uc+wu=u_{c}+w with wuc\abs{w}\ll u_{c}. Substituting into (3.7.56), the square on the right expands as (uc+w)2=uc2+2ucw+w2(u_{c}+w)^{2}=u_{c}^{2}+2u_{c}w+w^{2}:

d2wdφ2  +  uc  +  w  =  GML2  +  3GMc2(uc2+2ucw+w2). \dvn{2}{w}{\varphi} \;+\; u_{c} \;+\; w \;=\; \frac{GM}{L^{2}} \;+\; \frac{3GM}{c^{2}}\Big(u_{c}^{2} + 2u_{c}w + w^{2}\Big). (3.7.60)

The terms with no ww in them cancel by (3.7.59). The term in w2w^{2} is dropped, and this is the only approximation in the section. Be careful about why it is safe, because the obvious reason is the wrong one.

The obvious reason would be that w/ucw/u_{c}, the fractional variation of 1/r1/r around the orbit, is small. For Mercury that quantity is the eccentricity e=0.206e=0.206, which is not small at all. A 10%10\% error in Δφ\Delta\varphi would be four arcseconds per century, a hundred times the precision §8.5 claims. So that reason will not do.

What actually makes the neglect harmless is that the dropped term is quadratic. Let's take that slowly, because the chapter's headline number rests on it. Square a sinusoid and you get a constant plus a term at twice the frequency, and neither of those is at the frequency of the oscillation itself. A term that carries nothing at the fundamental frequency cannot shift that frequency at first order in the amplitude: it only displaces the orbit slightly and adds a small wobble at double the rate. The shift in ω2\omega^{2} therefore starts at second order, at O((3GMuc/c2)2)O\big((3GMu_{c}/c^{2})^{2}\big), and that is second order in exactly the small parameter §8.4 names.

Now put a number on it. For Mercury that is a relative change in Δφ\Delta\varphi of order 10810^{-8}, a millionth of an arcsecond per century. The comparison in §8.5 is against a measurement good to one part in a thousand, so the neglect is legitimate there for any eccentricity a planet actually has. What is left is

d2wdφ2  +  ω2w  =  0,ω2    1    6GMc2uc. \dvn{2}{w}{\varphi} \;+\; \omega^{2}\,w \;=\; 0, \qquad \omega^{2} \;\equiv\; 1 \;-\; \frac{6GM}{c^{2}}\,u_{c}. (3.7.61)

Step 3. Read it. Equation (3.7.61) is the harmonic oscillator again, but the coefficient of ww is no longer 11. Its solution w=w0cos(ωφ)w=w_{0}\cos(\omega\varphi) has period 2π/ω2\pi/\omega in φ\varphi rather than 2π2\pi, and ω<1\omega\lt1 because the correction is subtracted. So the radial distance returns to its minimum later than one full turn.

Finally, replace ucu_{c} by its leading value. From (3.7.59), uc=GM/L2u_{c}=GM/L^{2} up to a relative correction of order 3GMuc/c23GMu_{c}/c^{2}, which we have already agreed to neglect at the same order:

  ω2  =  1    6G2M2c2L2.   \boxed{\;\omega^{2} \;=\; 1 \;-\; \frac{6G^{2}M^{2}}{c^{2}L^{2}}.\;} (3.7.62)

8.4 · The advance per orbit

The perihelion recurs when ωφ\omega\varphi has advanced by 2π2\pi, that is after Δφtotal=2π/ω\Delta\varphi_{\text{total}}=2\pi/\omega. Subtract the full turn to get the advance:

Δφ  =  2πω2π  =  2π[(16G2M2c2L2)1/21]  2π3G2M2c2L2  =  6πG2M2c2L2, \begin{aligned} \Delta\varphi \;&=\; \frac{2\pi}{\omega} - 2\pi \;=\; 2\pi\left[\left(1-\frac{6G^{2}M^{2}}{c^{2}L^{2}}\right)^{-1/2}-1\right] \\[4pt] &\simeq\; 2\pi\cdot\frac{3G^{2}M^{2}}{c^{2}L^{2}} \;=\; \frac{6\pi G^{2}M^{2}}{c^{2}L^{2}}, \end{aligned} (3.7.63)

using the binomial expansion (1x)1/21+x/2(1-x)^{-1/2}\simeq1+x/2 of Chapter 0.3 §2.4, valid because the small parameter is

x  =  6G2M2c2L2  =  6GMc2p(using L2=GMp), x \;=\; \frac{6G^{2}M^{2}}{c^{2}L^{2}} \;=\; \frac{6GM}{c^{2}p} \qquad\text{(using $L^{2}=GMp$)}, (3.7.64)

which is 1.6×1071.6\times10^{-7} for Mercury. Now convert LL into something an astronomer reports.

⚑ The one imported relation, and why using it here is legitimate

Equation (3.7.64) substituted L2=GMpL^{2}=GMp, which is the Newtonian relation (3.7.58) between the angular momentum and the semi-latus rectum. In the relativistic problem that relation is not exact. Section 7's (3.7.47) already shows the correspondence between LL and orbital size is modified.

Why it may be used anyway. The quantity being computed, Δφ\Delta\varphi, is already first order in the small parameter GM/c2pGM/c^{2}p, since it vanishes identically when that parameter is zero. Replacing L2=GMpL^{2}=GMp by L2=GMp(1+O(GM/c2p))L^{2}=GMp\big(1+O(GM/c^{2}p)\big) therefore changes Δφ\Delta\varphi at second order, which for Mercury is a relative change of about one part in 10710^{7}, or roughly a millionth of an arcsecond per century. That is far below the precision of the comparison in §8.5. Using a lowest-order relation inside a first-order result is consistent. Using it inside a zeroth-order result would not be.

This is the same discipline Chapter 3.1 §7.2 used when it integrated the light deflection along the undeflected straight line: evaluate a small correction along the uncorrected solution, so that what is neglected is the correction to the correction.

With L2=GMpL^{2}=GMp, equation (3.7.63) becomes

  Δφ  =  6πGMc2p  radians per orbit,p  =  a(1e2).   \boxed{\;\Delta\varphi \;=\; \frac{6\pi GM}{c^{2}\,p}\ \ \text{radians per orbit}, \qquad p \;=\; a\big(1-e^{2}\big).\;} (3.7.65)

The second relation is pure conic geometry and takes two lines. From (3.7.58), the closest and furthest distances are rmin=p/(1+e)r_{\min}=p/(1+e) at φ=0\varphi=0 and rmax=p/(1e)r_{\max}=p/(1-e) at φ=π\varphi=\pi. Their sum is the major axis 2a2a, so

2a  =  p(11+e+11e)  =  2p1e2p  =  a(1e2). 2a \;=\; p\left(\frac{1}{1+e}+\frac{1}{1-e}\right) \;=\; \frac{2p}{1-e^{2}} \qquad\Longrightarrow\qquad p \;=\; a\big(1-e^{2}\big). (3.7.66)

8.5 · Mercury, with every conversion shown

Mercury's orbit, as measured: semi-major axis a=5.7909×1010 ma=5.7909\times10^{10}\ \mathrm{m}, eccentricity e=0.20563e=0.20563, orbital period P=87.969P=87.969 days. And GM=1.3271×1020 m3s2GM_{\odot}=1.3271\times10^{20}\ \mathrm{m^{3}\,s^{-2}}, c2=8.9876×1016 m2s2c^{2}=8.9876\times10^{16}\ \mathrm{m^{2}\,s^{-2}}.

Step 1. The semi-latus rectum. e2=0.042284e^{2}=0.042284, so

p  =  a(1e2)  =  (5.7909×1010)(0.957716)  =  5.5460×1010 m. p \;=\; a\big(1-e^{2}\big) \;=\; \big(5.7909\times10^{10}\big)\big(0.957716\big) \;=\; 5.5460\times10^{10}\ \mathrm{m}. (3.7.67)

Step 2. Radians per orbit. Substitute into (3.7.65). The numerator is 6πGM=18.8496×1.3271×1020=2.5016×1021 m3s26\pi GM_{\odot}=18.8496\times1.3271\times10^{20}=2.5016\times10^{21}\ \mathrm{m^{3}\,s^{-2}}. The denominator is c2p=(8.9876×1016)(5.5460×1010)=4.9845×1027 m3s2c^{2}p=\big(8.9876\times10^{16}\big)\big(5.5460\times10^{10}\big)=4.9845\times10^{27}\ \mathrm{m^{3}\,s^{-2}}. So

Δφ  =  2.5016×10214.9845×1027  =  5.0187×107 rad per orbit. \Delta\varphi \;=\; \frac{2.5016\times10^{21}}{4.9845\times10^{27}} \;=\; 5.0187\times10^{-7}\ \mathrm{rad\ per\ orbit}. (3.7.68)

Half a microradian. Per orbit, this is undetectable. The effect is extracted by waiting.

Step 3. Orbits per century. A Julian century is 100×365.25=36525100\times365.25=36525 days, so

N  =  36525 d87.969 d  =  415.20 orbits per century. N \;=\; \frac{36525\ \mathrm{d}}{87.969\ \mathrm{d}} \;=\; 415.20\ \text{orbits per century}. (3.7.69)

Step 4. Multiply, and convert. The advances add, since each orbit turns the ellipse by the same angle:

Δφcentury  =  (5.0187×107)(415.20)  =  2.0838×104 rad. \Delta\varphi_{\text{century}} \;=\; \big(5.0187\times10^{-7}\big)\big(415.20\big) \;=\; 2.0838\times10^{-4}\ \mathrm{rad}. (3.7.70)

One radian is 180/π180/\pi degrees and each degree is 36003600 arcseconds, so 1 rad=(180/π)(3600)=2.06265×1051\ \mathrm{rad}=(180/\pi)(3600)=2.06265\times10^{5} arcseconds. Hence

  Δφcentury  =  (2.0838×104)(2.06265×105)=  42.98 per century.   \boxed{\;\begin{aligned}\Delta\varphi_{\text{century}} \;&=\; \big(2.0838\times10^{-4}\big)\big(2.06265\times10^{5}\big) \\[3pt] &=\; 42.98''\ \text{per century}.\end{aligned}\;} (3.7.71)
⚑ The measured number, quoted as a measurement

Mercury's perihelion is observed to move by a large amount, most of which is not this effect at all. The bulk of the apparent motion is the precession of the equinoxes, a rotation of the coordinate frame in which positions are reported, and therefore not a fact about Mercury. Most of the rest is the Newtonian gravitational pull of the other planets, chiefly Venus and Jupiter, which is computed from Newton's law without any relativity. Subtracting both leaves a residual that resisted explanation from Le Verrier's announcement of it in 1859 until 1915.

That residual is measured to be 42.98±0.0442.98\pm0.04 arcseconds per century. The number is quoted here, not derived: extracting it requires the observational record and the Newtonian planetary perturbation theory, neither of which this book builds.

Equation (3.7.71) gives 42.9842.98, which sits inside the measurement's own uncertainty of about one part in a thousand. It was obtained from the mass of the Sun, the shape of Mercury's orbit, and a metric derived in §3, with no adjustable parameter anywhere. Nothing was fitted. This is the calculation Einstein performed in November 1915, before the field equations were published, and by his own account it gave him palpitations.

8.6 · Why Mercury, and not the Earth

Run the same arithmetic for the Earth: a=1.4960×1011 ma=1.4960\times10^{11}\ \mathrm{m}, e=0.01671e=0.01671, P=365.26P=365.26 days. Then p=a(1e2)=1.4956×1011 mp=a(1-e^{2})=1.4956\times10^{11}\ \mathrm{m}, and

Δφ  =  2.5016×1021(8.9876×1016)(1.4956×1011)  =  1.861×107 rad, \Delta\varphi \;=\; \frac{2.5016\times10^{21}}{\big(8.9876\times10^{16}\big)\big(1.4956\times10^{11}\big)} \;=\; 1.861\times10^{-7}\ \mathrm{rad}, (3.7.72)

and with N=36525/365.26=100.0N=36525/365.26=100.0 orbits per century the total is 1.861×105 rad=3.841.861\times10^{-5}\ \mathrm{rad}=3.84 arcseconds per century, smaller than Mercury's by a factor of 11.211.2.

The reason is visible in the formula and it explains the whole history of the problem. Two factors work in the same direction:

Δφcentury    1pper orbit×1Porbits    1a(1e2)1a3/2  =  1a5/2(1e2), \Delta\varphi_{\text{century}} \;\propto\; \underbrace{\frac{1}{p}}_{\text{per orbit}}\times\underbrace{\frac{1}{P}}_{\text{orbits}} \;\propto\; \frac{1}{a\big(1-e^{2}\big)}\cdot\frac{1}{a^{3/2}} \;=\; \frac{1}{a^{5/2}\big(1-e^{2}\big)}, (3.7.73)

where Kepler's third law Pa3/2P\propto a^{3/2}, from Chapter 1.4 §3, supplied the second step. The inner planet wins twice. Its orbit is smaller, so each turn shifts the ellipse further, and its year is shorter, so it takes more turns per century. Eccentricity helps a third time, weakly, and it helps in another way the formula does not show: a nearly circular orbit has no well-defined perihelion to track. Venus, whose eccentricity is 0.00680.0068, is a poor target for that reason despite being closer to the Sun than the Earth. Mercury is the innermost planet and much the most eccentric. It is the only place in the Solar System where the effect was ever going to be seen first.

In plain terms 3.7.8

Changing variables to the reciprocal of the distance converts the orbit from a statement about motion into one about shape, and in the old theory the resulting equation is the simplest in physics: a thing whose second derivative plus itself equals a constant. Its solutions repeat exactly once per turn, which is why an ellipse closes rather than drifting. An earlier chapter said the same from the side of symmetry, with a vector from the focus to the closest approach that never moves.

The relativistic version adds a single term, quadratic in the reciprocal distance and very small. Treated as a small correction about a circular solution, its effect is not a change of shape but of rate: the pattern still repeats, slightly slower than once per turn. The body therefore reaches its closest approach a little beyond where it did last time, and the whole ellipse turns, forever, at a fixed rate.

Two features deserve emphasis. This is a perturbative result, legitimate because the correction is one part in ten million even for the innermost planet, and worthless otherwise. And the turning per orbit depends on the orbit through one quantity only, its width measured across the focus, so a tight eccentric orbit gives the largest shift per turn and completes the most turns per century. That is why the discrepancy showed up at Mercury and nowhere else.

9 · Worked examples

Worked example 1 — Kepler's third law, exactly

Find the angular velocity of a circular orbit in the Schwarzschild geometry, measured in the distant observer's time tt. The answer is dφ/dt=GM/r3\dd\varphi/\dd t=\sqrt{GM/r^{3}}, which is Kepler's third law with no correction of any order. That is surprising enough to be worth three qualifications.

Use the geodesic equation directly. Chapter 3.3 §8 gives it as

d2xλdτ2+Γλμνdxμdτdxνdτ  =  0. \dvn{2}{x^{\lambda}}{\tau} + \Gamma^{\lambda}{}_{\mu\nu}\,\dv{x^{\mu}}{\tau}\dv{x^{\nu}}{\tau} \;=\; 0.

Take λ=1\lambda=1, the radial component, on a circular orbit in the equatorial plane, where rr is constant so dr/dτ=0\dd r/\dd\tau=0 and d2r/dτ2=0\dd^{2}r/\dd\tau^{2}=0. The only surviving connection coefficients from (3.7.8) with both lower indices among {0,3}\{0,3\} are Γ100\Gamma^{1}{}_{00} and Γ133\Gamma^{1}{}_{33}, so

Γ100(dx0dτ)2+Γ133(dφdτ)2  =  0. \Gamma^{1}{}_{00}\left(\dv{x^{0}}{\tau}\right)^{2} + \Gamma^{1}{}_{33}\left(\dv{\varphi}{\tau}\right)^{2} \;=\; 0.

Substitute the coefficients. With B=1/AB=1/A and θ=π/2\theta=\pi/2, Γ100=A/2B=AA/2\Gamma^{1}{}_{00}=A'/2B=AA'/2 and Γ133=r/B=rA\Gamma^{1}{}_{33}=-r/B=-rA. Dividing the whole equation by AA, which is non-zero outside rsr_{s}:

A2(dx0dτ)2  =  r(dφdτ)2(dφdx0)2  =  A2r. \frac{A'}{2}\left(\dv{x^{0}}{\tau}\right)^{2} \;=\; r\left(\dv{\varphi}{\tau}\right)^{2} \qquad\Longrightarrow\qquad \left(\dv{\varphi}{x^{0}}\right)^{2} \;=\; \frac{A'}{2r}.

Evaluate AA'. With A=12GM/c2rA=1-2GM/c^{2}r we get A=2GM/c2r2A'=2GM/c^{2}r^{2}, so

(dφdx0)2  =  GMc2r3  (dφdt)2  =  GMr3,   \left(\dv{\varphi}{x^{0}}\right)^{2} \;=\; \frac{GM}{c^{2}r^{3}} \qquad\Longrightarrow\qquad \boxed{\;\left(\dv{\varphi}{t}\right)^{2} \;=\; \frac{GM}{r^{3}},\;}

using x0=ctx^{0}=ct to convert. Writing Ω=dφ/dt\Omega=\dd\varphi/\dd t and T=2π/ΩT=2\pi/\Omega for the orbital period, T2=4π2r3/GMT^{2}=4\pi^{2}r^{3}/GM: Kepler's third law, unmodified, exact.

Three qualifications, and each of them matters.

(i) tt is the distant observer's time, not the orbiting body's. Per unit of the orbiting body's proper time the answer is different, and the difference is not small close in. What is exact is the relation between the areal radius and the period as booked by someone far away.

(ii) rr is the areal radius, which by §2.2 is not the distance to the centre. Kepler's law is exact in terms of the label defined by the area of the sphere, and would not be in terms of a radial distance.

(iii) It says nothing about stability. The relation holds for every r>3GM/c2r>3GM/c^{2} where a circular orbit exists at all, including the unstable branch. An exact and elegant formula for the period of an orbit that any disturbance destroys is a useful reminder that "exact" and "physically realisable" are different properties.

The result is a genuine coincidence of the areal-radius chart rather than a deep fact, and it is worth knowing because it makes the deviations elsewhere easier to trust: if Kepler's third law survives untouched, the 4343 arcseconds of §8 cannot be blamed on a sloppy definition of orbital period.

Worked example 2 — how much energy an accretion disc releases, and why it beats fusion

Compute the conserved energy EE of a circular orbit as a function of rr, evaluate it at the ISCO, and interpret the difference from the rest energy.

The energy of a circular orbit. On a circular orbit dr/dτ=0\dd r/\dd\tau=0, so (3.7.42) reads E=Veff(r)\mathcal{E}=V_{\text{eff}}(r), that is

E2c42c2  =  Veff(r)E2  =  c4+2c2Veff(r). \frac{E^{2}-c^{4}}{2c^{2}} \;=\; V_{\text{eff}}(r) \qquad\Longrightarrow\qquad E^{2} \;=\; c^{4} + 2c^{2}V_{\text{eff}}(r).

We also need LL for that orbit. Solve (3.7.47) for L2L^{2} rather than for rr, which is easier because the equation is linear in L2L^{2}:

L2(r3GMc2)  =  GMr2L2  =  GMr  13GMc2r  . L^{2}\left(r - \frac{3GM}{c^{2}}\right) \;=\; GM\,r^{2} \qquad\Longrightarrow\qquad L^{2} \;=\; \frac{GM\,r}{\;1-\dfrac{3GM}{c^{2}r}\;}.

Two things are already visible. L2L^{2} is positive only for r>3GM/c2r>3GM/c^{2}, so there is no circular orbit of any kind inside 3GM/c23GM/c^{2}. And LL\to\infty as rr approaches that radius from outside. Chapter 3.8 §2 identifies what is happening there.

Substitute and simplify. Putting this L2L^{2} into (3.7.43) and then into the expression for E2E^{2} gives, after collecting terms over the common denominator r3GM/c2r-3GM/c^{2},

  Ec2  =  12GMc2r  13GMc2r  .   \boxed{\;\frac{E}{c^{2}} \;=\; \frac{1-\dfrac{2GM}{c^{2}r}}{\sqrt{\;1-\dfrac{3GM}{c^{2}r}\;}}.\;}

Check it far away: as rr\to\infty both the numerator and the denominator tend to 11, so Ec2E\to c^{2}, the rest energy per unit mass of a body at rest at infinity, as (3.7.34)'s sanity check demanded. ✓

At the ISCO. Put r=6GM/c2r=6GM/c^{2}, so that 2GM/c2r=1/32GM/c^{2}r=1/3 and 3GM/c2r=1/23GM/c^{2}r=1/2:

EISCOc2  =  113112  =  2/31/2  =  223  =  0.942809. \frac{E_{\text{ISCO}}}{c^{2}} \;=\; \frac{1-\tfrac13}{\sqrt{1-\tfrac12}} \;=\; \frac{2/3}{1/\sqrt2} \;=\; \frac{2\sqrt2}{3} \;=\; 0.942809.

What that number is. A parcel of gas starting at rest far away has E=c2E=c^{2}. By the time it reaches the ISCO it has E=0.942809c2E=0.942809\,c^{2}. Since EE is conserved along a geodesic, the difference cannot have gone anywhere by itself. It is carried away by whatever friction moved the gas from one circular orbit to the next. So the fraction of the rest energy released on the way in is

1223  =  0.0572  =  5.72%. 1 - \frac{2\sqrt2}{3} \;=\; 0.0572 \;=\; 5.72\%.

Set that beside the alternative. Hydrogen fusing to helium releases the mass deficit of Chapter 2.5 §9: four protons weigh 4×1.00728=4.029124\times1.00728=4.02912 atomic mass units and a helium-4 nucleus weighs 4.001514.00151, so the fraction converted is (4.029124.00151)/4.02912=0.69%(4.02912-4.00151)/4.02912=0.69\%. Accretion onto a non-rotating black hole, stopping at the ISCO, releases over eight times as much from the same mass, and the fuel does not need to be hydrogen. This is the reason a quasar, powered by gas falling onto a hole, can outshine the entire galaxy of stars around it. The gravitational field is not a store of energy being consumed. It is a fixed geometry, and what is being spent is the rest mass of the gas.

And a limit worth naming. Nothing in §7 fixed the ISCO at 6GM/c26GM/c^{2} for a rotating hole. Section 1.2 excluded rotation from the start, the exterior geometry of a spinning body is not (3.7.22), and both the innermost stable orbit and the fraction above would have to be recomputed from a solution this book does not build. What is derived here is the non-rotating figure, and nothing in this chapter licenses quoting it for a spinning hole.

Worked example 3 — the same formula at the centre of the Galaxy

Apply (3.7.65) to the star S2, which orbits the compact object at the centre of the Milky Way. The point is that the formula was derived with no assumption of weakness, so it should be tried where the effect is large.

The inputs, all measured quantities of the same kind as Mercury's: the central mass is M=4.3×106MM=4.3\times10^{6}M_{\odot}, so GM=4.3×106×1.3271×1020=5.707×1026 m3s2GM=4.3\times10^{6}\times1.3271\times10^{20}=5.707\times10^{26}\ \mathrm{m^{3}\,s^{-2}}. The orbit has a=970 AUa=970\ \mathrm{AU} and e=0.88e=0.88.

The semi-latus rectum. 1e2=10.7744=0.22561-e^{2}=1-0.7744=0.2256, so p=970×0.2256=218.8 AUp=970\times0.2256=218.8\ \mathrm{AU}, and with 1 AU=1.496×1011 m1\ \mathrm{AU}=1.496\times10^{11}\ \mathrm{m} that is p=3.273×1013 mp=3.273\times10^{13}\ \mathrm{m}.

The advance per orbit.

Δφ  =  6πGMc2p  =  18.850×5.707×1026(8.9876×1016)(3.273×1013)  =  1.0758×10282.9418×1030  =  3.66×103 rad. \Delta\varphi \;=\; \frac{6\pi GM}{c^{2}p} \;=\; \frac{18.850\times5.707\times10^{26}}{\big(8.9876\times10^{16}\big)\big(3.273\times10^{13}\big)} \;=\; \frac{1.0758\times10^{28}}{2.9418\times10^{30}} \;=\; 3.66\times10^{-3}\ \mathrm{rad}.

Convert: 3.66×103×(180/π)=0.20953.66\times10^{-3}\times(180/\pi)=0.2095 degrees, which is 12.6 arcminutes per orbit\boxed{12.6\ \text{arcminutes per orbit}}. That is per orbit, not per century, and arcminutes, not arcseconds. The same formula that produced 5×1075\times10^{-7} radians for Mercury produces 73007300 times as much here.

Where the factor comes from. Entirely from GM/pGM/p. The mass is larger by 4.3×1064.3\times10^{6} and the semi-latus rectum is larger by 218.8/0.371=590218.8/0.371=590, and 4.3×106/590=73004.3\times10^{6}/590=7300. ✓ It is worth checking that the perturbative treatment of §8.3 still applies. The compactness at closest approach is rs/rmin=2GM/(c2a(1e))r_{s}/r_{\min}=2GM/\big(c^{2}a(1-e)\big), and with a(1e)=116 AU=1.74×1013 ma(1-e)=116\ \mathrm{AU}=1.74\times10^{13}\ \mathrm{m} and rs=2GM/c2=1.27×1010 mr_{s}=2GM/c^{2}=1.27\times10^{10}\ \mathrm{m} that is 7.3×1047.3\times10^{-4}. That is some eleven thousand times Mercury's 6.4×1086.4\times10^{-8}, and still four orders of magnitude below unity. The expansion is safe, though a good deal less safe than in the Solar System.

What this example does and does not establish. It computes a prediction. Whether the prediction is confirmed is a question about instruments, since resolving a star's orbit around an object eight kiloparsecs away requires interferometry at the milliarcsecond level, and this chapter derives no observational result. What it does show is that (3.7.65) was never a formula about Mercury. It is a formula about GM/c2pGM/c^{2}p, and there are places in the sky where that number is four orders of magnitude larger than in the Solar System.

10 · Your turn

Problem 1 — the two constants of integration, and what each one bought

The derivation of §3 produced two constants and fixed them by two different physical demands. (a) Suppose (AB)=0(AB)'=0 had been solved as AB=kAB=k for some constant k1k\neq1. Carry kk through §§3.5–3.6 and write the resulting line element. (b) Show that a rescaling of the time coordinate removes kk entirely, and say precisely what a distant observer would have to believe for k1k\neq1 to be the natural choice. (c) The constant C1C_{1} of (3.7.19) was fixed by matching to the Newtonian potential. What would C1>0C_{1}>0 describe, and is it excluded by anything derived in this chapter? (d) Chapter 3.6 §6 permitted a cosmological term. Say in one sentence why ignoring it in (3.7.6) is legitimate for the Solar System.

Solution

(a) With AB=kAB=k we have B=k/AB=k/A, and the third line of (3.7.10) becomes, by the same three substitutions as (3.7.16), R22=rA/kA/k+1R_{22}=-rA'/k-A/k+1. Setting it to zero gives (rA)=k(rA)'=k, so A=k+C1/rA=k+C_{1}/r, and

ds2  =  (k+C1r)c2dt2    kdr2k+C1/r    r2dΩ2. \dd s^{2} \;=\; \left(k+\frac{C_{1}}{r}\right)c^{2}\dd t^{2} \;-\; \frac{k\,\dd r^{2}}{k+C_{1}/r} \;-\; r^{2}\dd\Omega^{2}.

(b) Define t~=kt\tilde t=\sqrt{k}\,t. Then kc2dt2=c2dt~2k\,c^{2}\dd t^{2}=c^{2}\dd\tilde t^{2}, and dividing numerator and denominator of the radial term by kk shows the metric is (3.7.22) with C1/kC_{1}/k in place of C1C_{1}. So kk is not physics. It records that the coordinate tt was not chosen to be the proper time of an observer at infinity. For k1k\neq1 to be natural, a distant observer would have to be using a clock running at a rate k\sqrt{k} times their own proper time, which nobody does.

(c) By (3.7.21), C1=2GM/c2C_{1}=-2GM/c^{2}, so C1>0C_{1}>0 means M<0M\lt0: a negative mass, whose exterior geometry would repel test bodies. Nothing in this chapter excludes it, because Rμν=0R_{\mu\nu}=0 knows nothing about mass and the sign entered only through the matching. What excludes it is the observed sign of gravity, plus the energy conditions that Chapter 3.9 flags, neither of which is derived here.

(d) Because Λ1052 m2\Lambda\approx10^{-52}\ \mathrm{m^{-2}} while the curvature scale set by the Sun at Mercury's orbit is GM/c2r31030 m2GM_{\odot}/c^{2}r^{3}\approx10^{-30}\ \mathrm{m^{-2}}, larger by twenty-two orders of magnitude. That is the same estimate Chapter 3.6's Problem 1(c) made.

Problem 2 — radial infall, and a coincidence that is not one

Take L=0L=0 in (3.7.40), and a body released from rest at infinity, so that E=c2E=c^{2} by §5.1's sanity check. (a) Show that dr/dτ=2GM/r\dd r/\dd\tau=-\sqrt{2GM/r} exactly, which is the Newtonian escape-speed expression with no correction of any order. (b) Integrate to find the proper time to fall from rest at r0r_{0} to r=0r=0, and check the Newtonian formula is reproduced exactly. (c) Evaluate for a body falling from the Schwarzschild radius of a ten-solar-mass hole. (d) State carefully what has and has not been shown about the surface at r=rsr=r_{s}.

Solution

(a) With L=0L=0, (3.7.40) is (dr/dτ)2=E2/c2(12GM/c2r)c2(\dd r/\dd\tau)^{2}=E^{2}/c^{2}-(1-2GM/c^{2}r)c^{2}. Put E=c2E=c^{2}: the first term is c2c^{2}, and expanding the second gives c2+2GM/r-c^{2}+2GM/r. The c2c^{2} cancels exactly and (dr/dτ)2=2GM/r(\dd r/\dd\tau)^{2}=2GM/r. Taking the inward root, dr/dτ=2GM/r\dd r/\dd\tau=-\sqrt{2GM/r}. Note that no expansion was made: this is exact in the full metric.

(b) Separate and integrate, which is Chapter 0.8 §2.1's technique:

τ  =  0r0dr2GM/r  =  12GM0r0 ⁣r  dr  =  23r03/22GM. \tau \;=\; \int_{0}^{r_{0}}\frac{\dd r}{\sqrt{2GM/r}} \;=\; \frac{1}{\sqrt{2GM}}\int_{0}^{r_{0}}\!\sqrt{r}\;\dd r \;=\; \frac{2}{3}\,\frac{r_{0}^{3/2}}{\sqrt{2GM}}.

That is the Newtonian free-fall time from rest at infinity, unchanged.

(c) For M=10MM=10M_{\odot}, GM=1.3271×1021GM=1.3271\times10^{21} and rs=2GM/c2=2.953×104r_{s}=2GM/c^{2}=2.953\times10^{4} m. Then rs3/2=5.075×106r_{s}^{3/2}=5.075\times10^{6}, 2GM=5.152×1010\sqrt{2GM}=5.152\times10^{10}, and τ=23(5.075×106)/(5.152×1010)=6.6×105 s\tau=\tfrac23(5.075\times10^{6})/(5.152\times10^{10})=6.6\times10^{-5}\ \mathrm{s}. Sixty-six microseconds from the horizon to the centre.

(d) What has been shown: the proper time to cross r=rsr=r_{s} and reach r=0r=0 is finite, and nothing in dr/dτ=2GM/r\dd r/\dd\tau=-\sqrt{2GM/r} misbehaves at r=rsr=r_{s}, where that expression is perfectly smooth. What has not been shown: anything at all about the coordinate tt, which does not appear in the calculation, or about whether the traveller can send signals out. Both of those are Chapter 3.8's business, and the contrast between them is the point of that chapter.

Problem 3 — the inner branch, and the barrier that fails

(a) For L<23GM/cL\lt2\sqrt3\,GM/c, show that dVeff/dr>0\dd V_{\text{eff}}/\dd r>0 for every r>0r>0, and describe the motion of a body released from rest far away. (b) Show that the unstable root rr_{-} of (3.7.49) satisfies 3GM/c2<r6GM/c23GM/c^{2}\lt r_{-}\le6GM/c^{2} for every admissible LL, and find its limit as LL\to\infty. (c) Show that the stable root satisfies r+6GM/c2r_{+}\ge6GM/c^{2} and that r+L2/GMr_{+}\to L^{2}/GM for large LL, recovering (3.7.48). (d) Say in one sentence what happens to (3.7.43) as r0r\to0 for any L0L\neq0, and contrast it with (3.7.44).

Solution

(a) By (3.7.46), dVeff/dr\dd V_{\text{eff}}/\dd r has the sign of the quadratic (3.7.47), whose leading coefficient GMGM is positive. A positive-leading quadratic with negative discriminant is positive everywhere, so dVeff/dr>0\dd V_{\text{eff}}/\dd r>0 for all r>0r>0: the potential rises monotonically from -\infty at the origin to 00 at infinity. A body released from rest far away has E0\mathcal{E}\approx0, and since Veff<0V_{\text{eff}}\lt0 everywhere the turning-point condition E=Veff\mathcal{E}=V_{\text{eff}} is never met. It falls all the way in, whatever its angular momentum, provided only that L<LminL\lt L_{\min}.

(b) Write Lc/GM\ell\equiv Lc/GM, so (3.7.49) reads r=(2212)GM/2c2r_{-}=\big(\ell^{2}-\ell\sqrt{\ell^{2}-12}\big)GM/2c^{2}, defined for 23\ell\ge2\sqrt3. At =23\ell=2\sqrt3 the root is 12GM/2c2=6GM/c212GM/2c^{2}=6GM/c^{2}. For large \ell, expand the square root: 212=2112/226\ell\sqrt{\ell^{2}-12}=\ell^{2}\sqrt{1-12/\ell^{2}}\simeq\ell^{2}-6, so r3GM/c2r_{-}\to3GM/c^{2} from above. Differentiating shows rr_{-} decreases monotonically in \ell, so it is confined to the stated interval. The limiting radius 3GM/c23GM/c^{2} is where Worked example 2's L2L^{2} diverges. It is the same radius, reached from the other side.

(c) r+=(2+212)GM/2c2r_{+}=\big(\ell^{2}+\ell\sqrt{\ell^{2}-12}\big)GM/2c^{2} equals 6GM/c26GM/c^{2} at =23\ell=2\sqrt3 and increases with \ell. For large \ell the same expansion gives r+(226)GM/2c22GM/c2=L2/GMr_{+}\simeq(2\ell^{2}-6)GM/2c^{2}\to\ell^{2}GM/c^{2}=L^{2}/GM, which is (3.7.48). The Newtonian answer is the large-angular-momentum limit, as it must be, since large LL means a wide orbit.

(d) As r0r\to0 the term GML2/c2r3-GML^{2}/c^{2}r^{3} dominates both others, so VeffV_{\text{eff}}\to-\infty. By contrast VN+V_{\text{N}}\to+\infty, dominated by the barrier L2/2r2L^{2}/2r^{2}. The Newtonian barrier is infinitely high and the relativistic one is finite and then falls away, which in one line is the whole difference between the two theories close to a mass.

Problem 4 — where the ISCO is real, and where it is fiction

The solution (3.7.22) holds only in vacuum, so a feature at radius rr is physically present only if rr lies outside the body. (a) Compute rISCOr_{\text{ISCO}} for the Sun (M=MM=M_{\odot}, R=6.96×108R=6.96\times10^{8} m), for a white dwarf (0.6M0.6M_{\odot}, R7×106R\approx7\times10^{6} m) and for a neutron star (1.4M1.4M_{\odot}, R1.1×104R\approx1.1\times10^{4} m). In which cases does the ISCO lie outside the body? (b) For the Sun, what is at r=6GM/c2r=6GM_{\odot}/c^{2} in reality? (c) A spherical star collapses to a black hole. Explain, using §4, why the exterior geometry, and therefore the ISCO, is unchanged throughout the collapse. (d) What, then, actually changes for an orbiting body as the collapse proceeds?

Solution

(a) rISCO=6GM/c2=6×1476.6 mr_{\text{ISCO}}=6GM/c^{2}=6\times1476.6\ \mathrm{m} per solar mass, that is 8.86 km8.86\ \mathrm{km} per solar mass. Take the three cases in turn.

  • Sun: 8.86 km8.86\ \mathrm{km} against R=6.96×105 kmR=6.96\times10^{5}\ \mathrm{km}, deep inside, by five orders of magnitude.
  • White dwarf: 5.3 km5.3\ \mathrm{km} against 7×103 km7\times10^{3}\ \mathrm{km}, inside.
  • Neutron star: 12.4 km12.4\ \mathrm{km} against about 11 km11\ \mathrm{km}, outside, marginally.

So of the three, only the neutron star has a real ISCO, and it sits barely above its surface. That is why the inner edge of an accretion disc around a neutron star is a sensitive probe of the star's radius.

(b) Solar material, about 9 km9\ \mathrm{km} from the centre of the Sun, at a density and temperature the vacuum equations know nothing about. The metric there is a solution of Gμν=κTμνG_{\mu\nu}=\kappa T_{\mu\nu} with Tμν0T_{\mu\nu}\neq0, which this chapter does not compute. There is no ISCO, no horizon and no Schwarzschild radius inside the Sun. Those are features of a solution that does not apply there.

(c) Birkhoff, (3.7.32). The exterior is spherically symmetric and in vacuum throughout the collapse, so it is the Schwarzschild solution with some MM. And MM is fixed by the far-field matching (3.7.21), which cannot change while nothing crosses the region. So the geometry outside the collapsing surface is the same at every stage.

(d) Nothing about the geometry where the orbiting body is. What changes is how much of the Schwarzschild geometry is real: as the surface falls past 6GM/c26GM/c^{2}, the ISCO stops being a point inside the star and becomes a genuine feature of the vacuum. A distant orbit feels no disturbance at any moment. That is the content of §4.3, and the reason the formation of a horizon is not an event that announces itself.

The brick you just laid

You have solved Einstein's equations. Two symmetry conditions, written as statements about Killing vectors rather than about pictures, cut ten unknown functions of four variables down to two functions of one. The areal radius, (3.7.4), did the last step as a choice that Chapter 3.8 will have to un-make. Nine Christoffel symbols, three independent Ricci equations, and then one line: BR00+AR11BR_{00}+AR_{11} collapses to (AB)/(rB)(AB)'/(rB), so ABAB is constant, and asymptotic flatness makes it one. The angular equation becomes (rA)=1(rA)'=1, and one integration gives A=1+C1/rA=1+C_{1}/r. Then the Newtonian limit of Chapter 3.1, namely g00=1+2Φ/c2g_{00}=1+2\Phi/c^{2}, derived there from redshift alone, fixes C1=2GM/c2C_{1}=-2GM/c^{2}. The mass was never in the field equations. It arrived as a constant of integration. And rs=2GM/c2r_{s}=2GM/c^{2} is 2.952.95 km for the Sun and 8.98.9 mm for the Earth.

Birkhoff, derived rather than quoted. Drop staticity, allow both functions to depend on time, and the 0101 component of Rμν=0R_{\mu\nu}=0 reads B˙/rB\dot B/rB. So BB cannot depend on time, and the residual time dependence in AA is removable by rescaling a clock. The Schwarzschild solution is the only spherically symmetric vacuum solution. A star may pulse, collapse or explode and its exterior does not change. There is no spherically symmetric gravitational radiation. And Chapter 3.4's insistence that Ricci-flat is not flat has its example.

Chapter 3.5's promise, collected by name. The metric mentions neither tt nor φ\varphi, so t\partial_{t} and φ\partial_{\varphi} are Killing vectors, and Chapter 3.5 §9's theorem converts each into a constant along every geodesic: E=c2(1rs/r)dt/dτE=c^{2}(1-r_{s}/r)\dd t/\dd\tau and L=r2dφ/dτL=r^{2}\dd\varphi/\dd\tau, which Chapter 1.4's Noether theorem identifies as energy and angular momentum per unit mass. Four coupled second-order equations became one first-order equation, and the normalisation uu=c2u\cdot u=c^{2} turned it into motion in a landscape.

One new term, and everything that follows from it. The effective potential is Newton's plus GML2/c2r3-GML^{2}/c^{2}r^{3}, a fraction rs/rr_{s}/r of the centrifugal barrier. Circular orbits satisfy a quadratic where Newton had a linear equation, so there are two of them, an unstable inner one and a stable outer one. The discriminant vanishes at L=23GM/cL=2\sqrt3\,GM/c, below which no circular orbit exists at any radius at all. The two merge at rISCO=6GM/c2=3rsr_{\text{ISCO}}=6GM/c^{2}=3r_{s}, which is 88.688.6 km for a ten-solar-mass hole and is the inner edge of an accretion disc. Getting there releases 122/3=5.72%1-2\sqrt2/3=5.72\% of the rest mass, eight times what fusion yields.

And an ellipse that does not close. In terms of u=1/ru=1/r the orbit obeys u+u=GM/L2+3GMu2/c2u''+u=GM/L^{2}+3GMu^{2}/c^{2}, whose Newtonian truncation gives the closed conic of Chapter 1.4's Laplace–Runge–Lenz vector. Linearising about the circular solution turns the new term into a frequency shift, ω2=16G2M2/c2L2\omega^{2}=1-6G^{2}M^{2}/c^{2}L^{2}, so the radius repeats every 2π/ω2\pi/\omega rather than every 2π2\pi and the perihelion advances by 6πGM/c2p6\pi GM/c^{2}p per orbit. For Mercury that is 5.02×1075.02\times10^{-7} radians per orbit, 415.2415.2 orbits per century, 42.9842.98 arcseconds per century, against a measured residual of ⚑ 42.98±0.0442.98\pm0.04. Nothing was fitted.

Where this gets spent. Chapter 3.8 takes the same metric and sends light through it: the effective potential of §6 with c20c^{2}\to0 in one place, an unstable circular orbit for photons where Worked example 2's L2L^{2} diverged, a deflection of 4GM/c2b4GM/c^{2}b, and the collection of Chapter 3.1 §7.3's confessed factor of two. It then asks what the surface at r=rsr=r_{s} is, and the answer turns entirely on §2.2's warning that rr is a label chosen for convenience. The coordinates fail there and the geometry does not, and an invariant built from the curvature settles it. Chapter 3.9 puts matter back on the right-hand side and solves the equations for the universe rather than for a star, using the same three moves: symmetry to write an ansatz, the field equations to get an ordinary differential equation, and a boundary condition to fix a constant.