Part III · General Relativity — Chapter 3.7
Schwarzschild: The Solution and Its Orbits
Solve Einstein's equations exactly, by hand, in one sitting. Then extract from the answer an orbit Newton has no name for, and a planet's slow turning that nobody could explain for sixty years.
Chapter 3.6 finished the law and spent nothing. You have with , derived twice. You have its trace-reversed form, which in empty space reads . And you have a warning from Chapter 3.4 §6.3 that empty is not the same as flat. What you have not done is solve anything. This chapter solves something.
The target is the geometry outside a spherical star: our Sun, a neutron star, a black hole. It is the first exact solution of Einstein's equations, found by Karl Schwarzschild within weeks of the field equations being published and while he was serving on the Russian front, and it is short enough to derive completely. Nothing below is quoted. The metric is constructed, the connection is computed, the Ricci tensor is computed, the equations are solved, and the two constants of integration are fixed by two physical demands that are named as they are used.
Then the solution gets spent. Chapter 3.5 §9 promised, twice and in writing, that its Killing vectors would make the orbits of this chapter tractable. Section 5 collects that promise by name and turns four coupled second-order equations into one first-order equation for a single variable. That equation reads as motion in a landscape, and the landscape contains exactly one term that Newton's does not. Everything else in this chapter is that term's consequences: a smallest possible circular orbit, below which no circular orbit exists at any radius whatever, and an ellipse that fails to close by forty-three seconds of arc per century.
The route. Section 1 says what the two symmetry words mean as statements about geometry rather than about pictures. Section 2 turns them into an ansatz with two unknown functions in it, and names the one step in that reduction which is a choice. Section 3 solves . Section 4 shows that the staticity assumed in §1 was not needed and comes back on its own, which has a startling consequence for collapsing stars. Section 5 collects the conserved quantities, §6 builds the effective potential, §7 finds the innermost stable circular orbit, and §8 derives the precession and computes Mercury's number with every conversion on the page.
Conventions. The signature is , and the Riemann and Ricci sign conventions are the ones stated loudly in Chapter 3.4. We write , so that multiplies . Both and stay explicit everywhere. The Schwarzschild radius is written and never . A prime means throughout.
Tools you'll need — Chapter 3.6 §5.1, which supplies the trace-reversed field equations (3.6.23) and the remark that in vacuum they read . Chapter 3.5 §8.1 for Killing's equation and its practical test, which says that every coordinate the metric does not mention supplies a Killing vector, and §9 in full, which is the section this chapter spends. Chapter 3.4 §6.1 for the Ricci tensor and its definition (3.4.47), and §6.3 for Ricci-flat against flat. Chapter 3.3 §7 (the Christoffel formula, (3.3.50)), §8 (geodesics, both derivations, and the constancy of along one) and §3 (proper time). Chapter 2.5 §1 for the normalisation , equation (2.5.9), in flat spacetime. Chapter 3.2 §2 for what a chart is and what it is not. Chapter 3.1 §6.5, the derived weak-field component (3.1.39), which fixes this chapter's mass. Chapter 0.8 §2.2, the integrating factor, which does the one integration in §3, and §4.3 for reading an energy equation as motion in a landscape, a reading Chapter 1.3 §4 turned into phase portraits. Chapter 0.3 §2.4, the binomial series, for the one expansion in §8.4. Chapter 1.4 §3 for energy and angular momentum as Noether charges, and Worked example 2 there, on the Laplace–Runge–Lenz vector, which proved that Kepler ellipses close and predicted that any departure from an exact inverse square would make them precess. This chapter supplies the departure.
1 · What "spherically symmetric and static" buys
Here is where this section is going. We are going to translate two English adjectives into conditions on the metric, using the machinery of Chapter 3.5 rather than a picture. The translation matters because "looks the same from every direction" is a statement about a drawing, and a drawing cannot be substituted into a field equation. "There are three vector fields whose flows leave the metric unchanged" can be.
1.1 · Spherical symmetry, as a statement about the geometry
Chapter 3.5 §8.1 defined a Killing vector field as one whose flow leaves the metric exactly as it found it, , equivalently . That chapter also emphasised the distinction this section needs: a Killing vector is not a direction along which the picture looks the same, it is a direction along which every length anybody could measure comes out the same.
A spacetime is spherically symmetric when it admits three Killing vector fields , each spacelike, whose commutators reproduce the algebra of rotations:
Two things about this definition are worth pausing on, because it looks more elaborate than the English phrase it replaces.
The bracket is Chapter 3.2's commutator of vector fields, and by Chapter 3.5 §7.2 it is also the Lie derivative of one field along the other. Requiring (3.7.1) rather than merely requiring three Killing fields is what distinguishes rotations from, say, three translations: it is the statement that these three symmetries fail to commute in the particular pattern that turns about three perpendicular axes fail to commute.
You have already met these three fields. Chapter 3.5 §8.3 computed them on a sphere of radius : the obvious one , and two more, and , which are rotations about the other two axes and are invisible in the usual chart. Those are the of (3.7.1), extended off the sphere.
1.2 · Static, which is stronger than stationary
A spacetime is stationary when it admits a further Killing field that is timelike, meaning in this book's signature. In a chart adapted to it, that says the metric components do not depend on . Nothing about the geometry changes with time.
A spacetime is static when it is stationary and additionally invariant under reversing the time direction, .
The second condition is not decoration and it is not implied by the first. Here is a case that separates them, and it is the case that matters physically. Take the geometry outside a rotating star. Nothing about it changes with time, so it is stationary. But reversing time reverses the sense of the rotation, which is a different physical situation, so it is not static. The exact solution for a rotating mass is Kerr's, found in 1963, and it is a different and much longer calculation which this book does not perform. What follows is the non-rotating case.
Under the components pick up a minus sign, since exactly one of their two indices refers to the reversed coordinate, while and do not. Demanding that the metric be unchanged therefore forces
Staticity is precisely the statement that there are no cross terms pairing time with a direction in space. Without it, §2's ansatz would carry an extra unknown function, and that function is where the dragging of inertial frames by a rotating body lives. Chapter 3.8 uses the distinction again when it asks what kind of surface a horizon is.
1.3 · These are hypotheses, and they are being chosen
Nothing derived so far says that any actual object is spherically symmetric or static. The Sun rotates. The Earth is oblate. A collapsing star is neither. What §§2–3 solve is a special case chosen because it is exactly solvable, and the honest description of the result is this is the geometry outside a non-rotating spherical body. The accuracy of everything downstream is limited by how well a given body satisfies that.
Section 4 removes one of the two hypotheses at no cost, which is a considerably better outcome than it sounds.
Two words are doing all the work in this chapter and both have to be turned into statements about the geometry rather than about a drawing. Spherically symmetric means there exist three directions of dragging along which every measurable distance comes out unchanged, and that those three combine among themselves exactly the way turns about three perpendicular axes do. Static means two further things at once: there is another such direction, one along which a clock could actually tick, and running that direction backwards changes nothing either.
The second half of that is not decoration. A geometry with an unchanging time direction but no reversal symmetry is called stationary instead, and the difference is real rather than pedantic, because the space outside a spinning star is stationary and not static: reverse time and the spin reverses with it. What the reversal buys is the disappearance of every entry in the distance rule that pairs time with a direction in space, and without it the calculation below would carry another unknown function.
Naming a symmetry as a direction of dragging is what makes it something to calculate with rather than to look at. The previous chapter turned each one into a quantity conserved along every free path, and the same three-line test used there, that a label the distance rule never mentions supplies a symmetry, will identify the ones needed here by inspection.
2 · The ansatz, and why it has only two unknown functions
Here is where this section is going. Einstein's equations are ten coupled nonlinear equations for ten unknown functions of four variables. We are going to use §1's symmetries to reduce that to two unknown functions of one variable. Then we will say exactly which step in the reduction was a choice rather than a consequence, because that step is the one Chapter 3.8 has to un-learn.
2.1 · What the symmetries remove
Work in a chart adapted to the symmetries, which exists by Chapter 3.2 §2 and which we build one piece at a time.
Stationarity removes the dependence. Chapter 3.5 §8.1's practical test runs in reverse here: if is a Killing field, there is a chart in which no metric component depends on . Ten functions of four variables become ten functions of three.
Time reversal removes the cross terms. By §1.2, . Six functions remain in the spatial block, plus : seven.
Spherical symmetry removes the angular dependence and most of the spatial block. The three rotational Killing fields act on the spatial slices. Their orbits are the sets of points reachable from a given point by rotating, and those sets are two-dimensional spheres. The metric restricted to any one of them must be invariant under all three rotations. Chapter 3.5 §8.3's calculation on the sphere is exactly the statement that up to overall scale there is one such metric, the round one:
So each sphere contributes for a single positive function , one number per sphere rather than three functions of angle. Rotational invariance also forbids any cross term pairing or with the remaining radial direction, since such a term would pick out a direction on the sphere and there is none to pick. What is left is
with , and three functions of the radial label alone. Three unknown functions of one variable. One more move takes it to two.
2.2 · The areal radius — the load-bearing choice
The label in (3.7.3) is arbitrary: relabel and the same geometry is described by three different functions. That freedom is one function's worth, and we spend it.
Here is what we spend it on. The sphere labelled has, by (3.7.3), metric , so its area is . Define a new radial label by
This is the areal radius. One condition is being assumed and it should be said out loud: must be strictly monotonic in the old radial label, so that is invertible and can serve as a coordinate at all. Where it is not, this relabelling is unavailable, and the solution derived below does not describe that region. The place where it fails has a name, a throat, meaning a place where the spheres stop growing as you move outward. Nothing in this chapter meets such a place, but the assumption is real and it is being spent here rather than later.
With the areal radius, (3.7.3) becomes
Two unknown functions of one variable, and , both dimensionless. That is the ansatz §3 solves.
The definition (3.7.4) is a choice, made because it is convenient, and it is the single most misread line in this subject. Read it again: is defined so that the sphere at has area . Nothing else.
is not the distance to the centre. The distance measured radially, at fixed time, from the sphere at to the sphere at is, by (3.7.5), the integral , and §3 finds everywhere outside the mass, so that distance is larger than . The label and the length disagree, and is the exchange rate between them.
There need not even be a centre in the region we are solving. Equation (3.7.5) describes the vacuum outside a body. Inside the body the solution is different and this chapter does not compute it. So "the distance to the centre" is a quantity about a region we are not describing.
Chapter 3.8 §6 shows that treating as a distance is exactly what makes the surface at look catastrophic. It is not catastrophic, and the confusion starts here, with a choice made for convenience and then forgotten.
2.3 · Ten equations, two unknowns, and why that is not yet a contradiction
Substituting (3.7.5) into gives ten equations for two functions. That is a wildly overdetermined system and it would be reasonable to expect it to have no solution at all. Section 3 finds that six of the ten equations are identically, by the symmetry, that a seventh is a multiple of another, and that the surviving three are consistent. The consistency is not automatic and Chapter 3.6 §7.1 says where it comes from: the contracted Bianchi identity supplies four relations among the ten field equations for every metric, so the count of genuinely independent equations was never ten.
Exact solutions are rare, as the last chapter said, and symmetry is how the rare ones happen. Ten unknown functions of four variables become two functions of one variable. Following how is worth more than the answer. Independence of time, and of time's direction, removes every entry pairing time with space. Spherical symmetry makes every sphere a perfectly round one and forces the angular part into the single combination a globe uses, with one number in front of it. What survives is one factor multiplying the time part of the distance rule and one multiplying the radial part, each depending on the radial label alone.
Then comes a move that is a decision rather than a deduction, and it has to be said out loud. The radial label is defined by declaring that the sphere carrying it has area four pi times that label squared. Nothing says the label equals the distance anybody would measure walking inward from one sphere to a smaller one, and here it emphatically does not, since the radial factor in the distance rule is what converts one into the other.
Read the label as a distance to the centre and the surface met in the next chapter looks like a catastrophe. It is not one, and the confusion begins right here, in a definition adopted because it makes the algebra short and then quietly reinterpreted as a measurement.
3 · Solving
Here is the route, because there are five steps and each one is small.
- Compute the connection for (3.7.5).
- Compute the Ricci tensor from it.
- Find that one particular combination of two of the equations collapses to the statement that .
- Fix the resulting constant by demanding that the geometry become flat far away.
- Substitute into the angular equation, which becomes , and integrate that once.
Two constants of integration appear along the way, and each one is fixed by naming a physical requirement. Out comes an exact solution.
Chapter 3.6 §5.1 supplies the equation we are solving. The trace-reversed field equations (3.6.23) with and read
These are the equations for empty space. They apply outside the star, not inside it, and Chapter 3.4 §6.3 already warned that they do not say the geometry is flat.
3.1 · The connection
Chapter 3.3 §7 derived the Christoffel formula (3.3.50),
and the metric (3.7.5) is diagonal, so the inverse metric is diagonal too, with entries , , and . That collapses the sum over in (3.7.7) to a single term every time. Grind box A does the arithmetic. The answer is nine non-zero components, counting each symmetric pair once:
Four of the nine are exactly the connection coefficients of flat space in spherical coordinates: the two carrying , the one carrying and the one carrying . Chapter 3.3 §2 computed those four as the standing warning that varying components do not mean curvature. They are here because the chart is spherical, not because the spacetime is curved. The other five carry and , and they are where the physics lives.
Grind box A — the nine Christoffel symbols, from (3.3.50)
Write and recall that no metric component depends on or on , and that only depends on .
. Take , so and . With , the bracket in (3.7.7) is . Hence .
. Take , so . With the bracket is . Hence .
. Same , and the bracket is , giving .
. Same again. With the bracket is , giving . The same computation with in place of carries the extra through and gives .
. Take , . With , the bracket is , giving . Identically for , the extra cancelling between and .
. , : the bracket is , giving .
. , , : the bracket is , and .
Everything else vanishes. A component is non-zero only if the bracket contains a surviving derivative, and the only non-vanishing derivatives of metric components are of and of . Running through the possibilities leaves the nine above.
One contraction, for later. Chapter 3.5 §6.4 derived from Jacobi's formula, and here , so and
and the other two vanish. Adding the list above directly gives the same thing, which is the Jacobi formula checked rather than trusted.
(All nine were confirmed symbolically from (3.7.7) for general and , and no other component is non-zero.)
3.2 · The Ricci tensor, and the count
Chapter 3.4 §6.1 defined the Ricci tensor and wrote it out in terms of the connection, equation (3.4.47):
Feeding (3.7.8) into it is a page of bookkeeping, folded into grind box B. The results are
with every off-diagonal component zero.
Now the count, which stays out here in the main text. Ten equations went in. Six of them, the off-diagonal ones, are and say nothing. Of the remaining four, the last two are proportional, so carries no information beyond . That proportionality is spherical symmetry showing up in the answer, since and are on the same footing. What is left is three equations for two unknown functions.
Three equations for two unknowns is one equation too many, and a system in that position has no right to possess a solution. It does, and the reason was supplied a chapter ago: by Chapter 3.6 §7.1 the contracted Bianchi identity imposes four differential relations among the ten field equations for every metric whatever, so the three surviving equations are not independent of one another. That is worth a sentence rather than a shrug, because it is the same identity that fixed the factor of in the Einstein tensor, doing a second job here.
Grind box B — the Ricci components, term by term
, all four terms of (3.7.9).
Term 1. . The only non-zero is , so this is
Term 2. : nothing depends on .
Term 3. . Only contributes, and grind box A's contraction gives
Term 4. . The only connection coefficients with a among their indices are and , so the double sum has exactly two surviving terms, and , and they are equal:
Add. The two terms combine as , and the two terms as , giving the first line of (3.7.10).
. Same four terms. Term 1 is . Term 2 is , which cancels the of term 1 exactly and leaves . Term 3 is . Term 4 sums the squares , the last from the two coefficients equal to . Collecting gives the second line of (3.7.10). Note that the cancels between terms 2 and 4, which is why no such term appears in the answer.
. Term 1 is . Term 2 is . Term 3 is . Term 4 is . The cancels between terms 2 and 4, which it must, since cannot depend on . What survives is the third line of (3.7.10).
and the off-diagonal components. repeats the computation with the extra carried along, giving . Every off-diagonal component vanishes: any with requires a connection coefficient carrying both a and a , and inspection of (3.7.8) shows the surviving sums are empty or cancel in pairs.
(All of (3.7.10) was reproduced symbolically from (3.7.9) for general , , together with the vanishing of all six off-diagonal components and the identity .)
3.3 · The line the whole solution turns on
Look at and in (3.7.10). Each contains , and each contains the same three other structures. Weight the first by and the second by and add:
Every second derivative has gone, and so has every quadratic term. The two survivors are the two halves of a product rule:
Setting the vacuum equations (3.7.6) to zero makes the left-hand side vanish, and and are finite and non-zero in the region where we are solving, so
That single line is why the Schwarzschild solution can be found by hand. A pair of coupled second-order equations has been replaced, exactly and with no approximation, by an algebraic relation between the two unknowns. It is worth noticing why the combination works: and differ from one another only in which of and sits in each denominator, so weighting by the opposite function makes the three complicated structures identical up to sign. Nothing was clever about the choice. It was forced by the shape of (3.7.10).
3.4 · The first constant, fixed by a physical demand
Equation (3.7.13) gives a constant, and algebra cannot say which one. Name the physical input that does: asymptotic flatness. Far from an isolated body the geometry should approach that of empty flat spacetime, so that a distant observer's clocks and rulers are the ordinary ones of Part II. In the chart (3.7.5) that means
since with the line element becomes , which is Minkowski spacetime in spherical coordinates: non-constant components and no curvature whatever, which is the lesson Chapter 3.3 §2 delivered early using the polar plane. The product therefore tends to , and being constant it equals everywhere:
One unknown function remains. Note carefully what was spent: not a calculation, but a decision about the physical situation being described. A different boundary condition would describe a different situation, and Problem 1 asks what a different constant would have meant.
3.5 · One integration
Put into the third line of (3.7.10). Each of the four terms simplifies, and it is worth doing them one at a time. With we have and , so
The two derivative terms are equal, so they add rather than cancel, and
Name the technique. The left-hand side is already a product rule: . So
This is a first-order linear equation of the kind Chapter 0.8 §2.2 handled with an integrating factor, and here the integrating factor has already done its work: is exactly the factor that turns into an exact derivative. One integration, and one constant of integration :
Notice that (3.7.14) is satisfied automatically for any : the solution is asymptotically flat whatever the constant is, so asymptotic flatness has nothing left to say and something else must fix .
3.6 · The second constant, and where the mass comes in
Nothing in (3.7.6) mentions a mass. The equations are the equations of empty space, and they are the same equations whether the object at the centre is the Sun or a grain of dust. The mass can therefore enter in only one place: as the constant of integration, fixed by matching to something known.
The known thing is the Newtonian limit. Chapter 3.1 §6.5 derived, from the gravitational redshift in an accelerating cabin and with no general relativity used at all, that a weak static field has
equation (3.1.39), with the plus sign belonging to this book's signature, and Chapter 3.6 §5 used the same component to fix . Far from a mass the Newtonian potential is , so
Both sides carry a term falling off as and matching their coefficients fixes the constant. And is defined by this equation: it is the mass a distant observer would infer from the orbits of test bodies, and nothing in the derivation refers to what is inside. Substituting into (3.7.5) with (3.7.15),
That is the Schwarzschild metric, and it is exact. No expansion was made, no small parameter was introduced, and the only approximation anywhere in the derivation was in the matching of (3.7.21), which was performed at large where it is valid and which fixed a constant that is then exact everywhere.
In: the two symmetry assumptions of §1, the vacuum field equations of Chapter 3.6, the Christoffel formula of Chapter 3.3, the Ricci definition of Chapter 3.4, asymptotic flatness, and the weak-field component of Chapter 3.1.
Out: an exact solution, with the entire influence of the central body carried by one number. Two very different solar systems with the same have identical geometries outside them, whatever their internal structure. Section 4 will strengthen that statement considerably.
What it cost: the ansatz. Everything downstream describes a non-rotating, spherical source. It cost no approximation, which is rarer in this book than it may look: Chapter 3.6 §7.3 explained that the field equations are nonlinear, that exact solutions are correspondingly rare, and that expanding about a simple solution is normally the only technique available. This is one of the exceptions, and it is the reason it is a landmark rather than an exercise.
3.7 · The Schwarzschild radius, as a length
The metric (3.7.22) contains one length, the value of at which the coefficient of vanishes:
Put numbers into it, because as a symbol it is easy to mistake for something enormous. With throughout:
The Sun. , so
The Earth. Now run the same arithmetic with the Earth's measured value in place of the Sun's, , so
Nine millimetres. That is the number worth carrying, because it converts a symbol into a length you could hold: compress the entire Earth inside a sphere the size of a small grape and the coefficient of at its surface reaches zero. And because is linear in , the number scales straight up with the mass. A hole of solar masses, the class imaged at the centre of M87, has , which is astronomical units, comparable to the orbit of Uranus.
The ratio that governs how big the relativistic corrections are is . At the Sun's surface, : four parts in a million. That number is why Newtonian gravity survived three centuries of astronomy, and §8 shows what it takes to detect the difference anyway.
Three independent equations for two unknown functions is one equation more than the count allows, and a system in that position has no business being consistent. It is, and not by luck: the differentiated identity that shaped the field equations one chapter ago guarantees relations among them.
The solution turns on a single line. Take the equation belonging to the time direction and the one belonging to the radial direction, weight each by the other's factor, and add. Every second derivative cancels, every quadratic term cancels, and what survives is the rate of change of the product of the two unknowns, set equal to nothing. So that product is the same everywhere, and demanding that the geometry become the ordinary flat one far from the body makes it one. Two unknowns have collapsed into one, by a physical requirement rather than by algebra. The angular equation then says that the radial label times the surviving function has derivative one, the first kind of equation the toolkit taught anybody to integrate.
Nothing in the empty-space equations knows what a mass is; they are the same equations around a star and around a speck of dust. The constant left over from that integration acquires meaning only when the answer is set beside the old inverse-square law far away, and it is at that moment, and not earlier, that a mass enters at all.
4 · Birkhoff's theorem, and what it says about a collapsing star
Here is where this section is going. Section 1 assumed staticity and §3 used it in two places: to drop the dependence and to drop the cross terms. We now throw the assumption away, keep only spherical symmetry, and redo the calculation. The result is that staticity comes back on its own, because the field equations force it. That is Birkhoff's theorem, and it is derived here rather than quoted.
4.1 · Redo it with time allowed in
Keep spherical symmetry, which by §2.1 still forces every sphere to be round, still forbids angular cross terms, and still permits the areal radius (3.7.4). Drop staticity entirely and allow both remaining functions to depend on time:
Two remarks on what this ansatz does and does not assume, since the whole force of the theorem rests on them. It does not assume the matter is static: the source can be pulsing, collapsing or exploding, provided it stays spherical. It does still exclude a cross term , and that is not a further assumption about the geometry. With the time part of the line element reads . Complete the square and supply an integrating factor, which exists because only one other variable is present. A function then appears for which the whole of it is . That is a change of chart, precisely the operation Chapter 3.2 §2 defined, and it costs nothing.
Recompute. Every component of (3.7.8) that survives acquires possible time derivatives, three new connection coefficients appear carrying and , and one Ricci component that was identically zero before is no longer so. Grind box C computes it. The result is remarkably clean: writing ,
4.2 · The theorem
Setting (3.7.27) to zero, and noting again that and are finite and non-zero where we are solving, gives
One component of the vacuum equations has forbidden the radial function to depend on time, with no assumption made. The rest follows in three moves.
Move 1. With a function of alone, the combination of §3.3 goes through unchanged. Grind box C confirms that is still even when depends on time. So and
for some function of time alone. Note what has changed: the "constant" of §3.3 was constant in , and nothing said it could not vary with .
Move 2. The time dependence is not in the geometry, it is in the clock. From (3.7.29), , so the time part of the line element is . Define a new time coordinate by
which is legitimate because depends on alone, so the right-hand side is an exact differential and is a genuine coordinate. Then and
which has no time dependence at all and no cross terms: static. The whole of the apparent time dependence was a statement about how a distant observer had chosen to run their clock.
Move 3. Equation (3.7.31) is the ansatz (3.7.5) with already imposed, so §§3.5–3.6 apply verbatim: , one integration, and the constant fixed by the Newtonian limit. The solution is (3.7.22) again.
Grind box C — the time-dependent components, and
The new connection coefficients. With and both functions of , three components join the list (3.7.8) and two acquire new pieces. From (3.7.7), with :
each obtained the same way as in grind box A, the only difference being which derivative survives in the bracket. For instance takes , , and the bracket , giving . The nine of (3.7.8) keep their forms with and now functions of two variables.
The component of the Ricci tensor. Put , into (3.7.9):
Take them in turn. The only non-zero are and , so term 1 is . Using the contraction of grind box A, now reading , term 2 is . The two pieces cancel outright between terms 1 and 2, and the two remaining pieces, and , cancel as well. The quotient rule turns the first into and the second into , and these are the same expression, because mixed partial derivatives commute. Terms 1 and 2 together give zero.
Term 3 is summed over :
Term 4 is , whose surviving contributions come from and the two angular pairs, the latter giving nothing because :
the two angular pairs contributing nothing because . Expand term 3 and set it beside this: the four products match one for one and cancel, and the only survivor is the piece term 3 carries with the in it. Hence
which is (3.7.27).
(Confirmed symbolically for general , : came out exactly , and with restricted to depend on alone the combination came out exactly , as Move 1 claims.)
4.3 · What the theorem says about a star
The consequences are large and they are not obvious from the algebra.
A pulsating star has a static exterior. Take a spherical star that expands and contracts. A Cepheid variable does exactly this, and so does a star ringing after a disturbance. Outside it, the geometry is (3.7.22) with a fixed , unchanging, for all time. A planet orbiting it feels nothing at all.
There is no spherically symmetric gravitational radiation. A radiating system must have a time-dependent field outside it, and Birkhoff says a spherical one cannot. Electromagnetism forbids the corresponding thing, since a spherically pulsing charge does not radiate. But it forbids it by a weaker route, conservation of charge, which says nothing about the field elsewhere. Here the statement is stronger: the entire exterior geometry is unchanged, because it is the only solution the equations possess. That is why the gravitational waves detected in this century come from pairs of compact objects spiralling about each other, which are as far from spherical as a system can be, rather than from single objects ringing.
The field of a collapsing star does not change as it collapses. This one is worth stating slowly. A spherical star collapses through its own Schwarzschild radius. During the collapse the matter distribution changes violently. Outside the star's surface, the geometry is (3.7.22) with the same throughout, before, during and after. Nothing propagates outward, nothing announces the collapse, and a distant orbit is undisturbed. Whatever a horizon is, its formation is not an event that anything outside can feel. That makes the question of what it is considerably more interesting, and Chapter 3.8 asks it.
And a debt of Chapter 3.4 is collected. That chapter's §6.3 insisted that Ricci-flat does not mean flat, and named the Weyl part of the curvature as what survives contraction. Here is the example. Equation (3.7.22) has everywhere it applies, since it was constructed to. And it is not flat: it bends light, holds planets in orbit, and §7 shows it does something no flat geometry could. Every bit of that is Weyl curvature. The tidal deviation of Chapter 3.1, which Chapter 3.4 §4 identified with curvature, is entirely carried by the part the Ricci tensor throws away.
Redo the calculation without assuming anything about time and the assumption comes back unbidden, cornered by the equations rather than imposed. Let both unknown functions depend on time as well as on the radial label, and one component of the empty-space equations, the one mixing time with radius, forbids the radial function any time dependence. The rest of the argument then runs unaltered, and the only surviving trace of time is absorbed by rescaling the distant observer's clock.
So the geometry outside any spherical arrangement of matter is the one already found, whatever that matter does. A star may pulse, collapse, rebound or detonate; if it stays spherical, nothing outside changes. There is therefore no spherically symmetric gravitational radiation, which is why the waves picked up in this century come from pairs of bodies swinging about each other rather than single ones ringing. And a star falling in on itself sends no warning outward: the field it leaves is the field it always had.
A second reading matters more for what follows: empty is not the same as flat. The solution has vanishing boiled-down curvature everywhere outside the mass, since that is what it was built to have, and it is nonetheless not flat: it bends light, holds planets, and further in acquires a surface crossable in one direction only. All of that lives in the part of the curvature the boiling-down discards.
5 · Two Killing vectors, two conserved quantities
The solution exists. Now spend it, which means asking how a body moves in it. Here is where this section is going. Free motion in a metric is four coupled second-order equations for , with position-dependent coefficients, and that is not a problem anybody solves by hand. We reduce it to one first-order equation in one variable, using symmetry three times: once to fix the plane, and twice to integrate an equation once for free. Chapter 3.5 §9 promised exactly this, twice and by name, and this section is where the promise falls due.
5.1 · The promise, and the two charges
Chapter 3.5 §8.1 gave the practical test: every coordinate the metric does not mention supplies a Killing vector. Look at (3.7.22). It does not mention , and it does not mention . So
are Killing vector fields. And Chapter 3.5 §9.1 proved, in four lines, that for any Killing field and any geodesic with tangent , the number is the same at every point of the geodesic. That is equation (3.5.73). Two Killing fields, two constants.
Compute them. Lowering an index means contracting with the metric, and the metric is diagonal, so picks out one component each time.
The time charge. In the coordinates with , the field has components , since . Then and
The angular charge. Now do the same for the rotational Killing field. The field has components , so lowering the index gives , and
The minus sign is a convention chosen so that is positive for motion in the direction of increasing . The conserved quantity is either way.
What they are. Chapter 3.5 §9.2 ran Chapter 1.4's Noether theorem on the same system and found that the charge belonging to is , which is equation (3.5.79). So is the Noether charge of time-translation invariance and is the Noether charge of rotational invariance: is the energy per unit mass and the angular momentum per unit mass, and they are the same two quantities Chapter 1.4 §3 obtained for planets long before any of this geometry existed. Their names were not assigned by analogy. They were derived and then recognised.
Far away. As the factor in (3.7.34) tends to and , so , which is Chapter 2.5's energy per unit mass including the rest energy. A body at rest at infinity has exactly. A bound body has .
In the equatorial plane. With , equation (3.7.35) reads . That is Kepler's second law, equal areas in equal times, with proper time in place of coordinate time. Chapter 3.5 Worked example 2 obtained the same expression on the flat plane from the same Killing field, and Chapter 1.4 §3.3 obtained it from rotational symmetry. Three routes, one quantity.
5.2 · The equatorial plane, and why taking it costs nothing
Both charges still contain . Dispose of it, and say why the disposal is legitimate rather than convenient.
The argument is a symmetry argument and it has two steps. First, take any geodesic and any point on it. The tangent at and the radial direction at span a plane through the centre. By spherical symmetry there is a rotation carrying that plane to the equatorial plane , and a rotation is an isometry, so it carries geodesics to geodesics. Therefore every geodesic can be rotated so that at it lies in the equatorial plane with no motion out of it, .
Second, check that it stays there. The component of the geodesic equation, using and from (3.7.8), is
At the last term vanishes because , and if there as well then the middle term vanishes too, leaving . Zero value, zero first derivative and zero second derivative: by the uniqueness theorem for the initial-value problem, Chapter 0.8 §1.1, the solution is for all .
This is exactly the statement that a planet's orbit lies in a plane, and the reason is exactly the Newtonian one: Chapter 1.4 §3.3's conservation of the angular momentum vector, whose fixed direction is the normal to the orbital plane. From here on, and (3.7.35) reads .
5.3 · The count, before and after
| Unknowns | Equations | |
|---|---|---|
| Geodesic equation as written | four, coupled, second order | |
| after spherical symmetry (§5.2) | three, second order | |
| after the two Killing charges | one, first order |
That is the whole of the reduction, and it was bought entirely with symmetry. Chapter 3.5 §9 said it in advance: "Without this section, Chapter 3.7 would be a wall of algebra with no visible structure. With it, that chapter is mostly Chapter 1.3's phase-plane reasoning applied to one variable." The next section supplies the one remaining equation.
A promise made two chapters ago falls due here, and it is collected in full. The rule for measuring distances outside a spherical mass mentions neither the time label nor the angle around the axis, and by the practical test established there, every label the rule fails to mention supplies a direction along which the geometry does not change. Each hands over a number that is the same at every point of every freely falling path.
The two numbers are the energy per unit mass and the angular momentum per unit mass. They are the same two constants the theorem relating symmetries to conservation produced for planets long before this geometry existed, and nothing about them was assumed; they fall out of the metric and are then recognised. The second, written out, is the equal-areas rule of the seventeenth century, with the traveller's own clock in place of the astronomer's.
The saving is best stated as arithmetic. Free motion is four coupled second-order equations for four unknown functions of the traveller's own time. Spherical symmetry turns any orbit into a single plane, disposing of one. The two conserved quantities integrate two more, once and for all. What is left is one first-order equation in one variable, of exactly the form a marble rolling in a valley obeys. That is not computational cleverness; it is symmetry, converted into constants by a theorem proved for pendulums.
6 · The effective potential
Here is where this section is going. One line of algebra turns the normalisation of the four-velocity into an energy equation for . We then read that equation as a particle moving in a landscape, compare the landscape with Newton's term by term, and find that exactly one term is new. Everything in §§7 and 8 is that term's doing.
6.1 · One line of algebra
The four-velocity of a massive body satisfies
which is Chapter 2.5 §1's , equation (2.5.9), carried onto a manifold in Chapter 3.3 §8. There it was shown to be constant along any geodesic, so that imposing it once at one point imposes it everywhere. Write it out for (3.7.22) in the equatorial plane, with as shorthand:
Now substitute the two charges. From (3.7.34), , so the first term is . From (3.7.35), , so the third term is . Hence
We want the radial derivative on its own, since is the one unknown we are chasing. So multiply through by and rearrange for the radial term:
That is the whole of orbital dynamics in general relativity. One equation, first order, one unknown function, two constants. Everything from here to the end of the chapter is extracted from it.
6.2 · Reading it as a landscape
Expand the product in (3.7.40), one term at a time. The four products are , , and , so
Divide by two and move everything except the derivative to the right:
Everything depending on position alone has been gathered into the single function . Here is what that function is:
Equation (3.7.42) is the statement that a quantity quadratic in a velocity, plus a function of position alone, is constant. That is the form Chapter 1.3 §4 used to draw phase portraits, and the form Chapter 0.8 §4.3 used for the oscillator. The technique is read off the turning points: motion is possible only where , the body turns around where the two are equal, and it orbits in a circle where is stationary.
Two cautions about what is, since it looks like an energy and is not quite one. It is built from , which already includes the rest energy per unit mass . The combination removes it, and for a slow body gives , the Newtonian energy per unit mass. And the "time" in is the traveller's proper time, not the distant observer's . Both facts matter and neither affects the shape of .
6.3 · Term by term against Newton
Build the Newtonian comparison from scratch, in two lines, so that nothing is being taken on trust. Newtonian energy per unit mass for a body in a plane is , and Chapter 1.4 §3.3's angular momentum per unit mass is , so . Eliminating ,
The second term is the centrifugal barrier, and it is a piece of kinetic energy wearing a potential's clothes. Chapter 1.2's Problem 3 on the bead on a rotating hoop already made that point, and it is the reason no fictitious force has to be invoked to explain the barrier.
Set the two side by side.
| Term | Newton | Schwarzschild | What it does |
|---|---|---|---|
| attraction | identical, and it pulls inward | ||
| centrifugal barrier | identical, and it keeps a rotating body out | ||
| the new term | — | attractive, and it beats the barrier at small |
The first two terms are not approximately the same as Newton's. They are the same, with no correction of any order. The entire difference between the two theories, for a massive body in orbit, is the third term.
Three things about it deserve naming before §7 exploits them.
It is attractive. The minus sign means it deepens the well, and it deepens it in the region where the barrier was supposed to be doing its job.
It carries . A body falling straight in, , does not feel it at all. Problem 2 shows that radial infall in this metric obeys the Newtonian equation exactly. The new physics is entirely about bodies with sideways motion.
It falls off as , so it wins at small . Compare it with the barrier: the ratio is
independent of altogether. So the new term is a fraction of the barrier: negligible far out, comparable at , dominant inside. Take Mercury's perihelion, a distance from the orbital elements §8.5 lists. There , which is the size of the effect §8 has to dig out of the data.
A freely falling body's velocity has a fixed length, and squaring that statement with the two constants in hand gives one equation relating the radial position to the traveller's own time. Rearranged, it says precisely what an elementary mechanics problem says about a marble in a valley: a term quadratic in the rate of change, plus a landscape depending on position alone, adds up to a fixed number. The apparatus of turning points applies unaltered.
Set that landscape beside the Newtonian one term by term. The attraction is there, unchanged. The barrier that keeps a body with sideways motion from reaching the centre is there, unchanged, and it is what made orbits stable for three centuries. Then there is a third term with no Newtonian counterpart whatever: attractive like the first, carrying the square of the angular momentum like the second, and falling off as the cube of the distance: undetectable far out and decisive close in.
That one extra term is the entire difference between the two theories for a body in orbit, and everything the chapter extracts afterwards comes from it. Notice what it does to the barrier. The barrier still rises as a body comes inward, and then, at short enough range, the new term wins and the barrier collapses. A wall that fails at close quarters is a different sort of object from a wall.
7 · The innermost stable circular orbit
Here is where this section is going, and it is the chapter's thesis. Circular orbits sit where is stationary. We solve and find a quadratic, which therefore has two roots, one root or none. In Newtonian gravity the same question has exactly one root for every radius. The discriminant tells us when the roots exist, and the answer is that below a certain angular momentum there is no circular orbit at any radius whatever. The two roots merge at , and that radius has a name and an observational consequence.
7.1 · Where the landscape is level
Differentiate (3.7.43) term by term:
Set it to zero and multiply by , which is legitimate for and clears every denominator at once:
A quadratic in . Stop and compare with Newton before solving it. Differentiating (3.7.44) gives , which after multiplying by is linear:
One root, always, for every . Turn it round: name any radius you like and puts a circular orbit there. Newtonian gravity has a circular orbit at every radius, and each is a minimum of , hence stable. Equation (3.7.47) is a different kind of object, and the difference is the term carrying .
7.2 · The discriminant, and a floor on the angular momentum
Solve (3.7.47) by the quadratic formula, with , and :
the second form obtained by taking one factor of out of the square root, which is legal since , and then a factor of out of what remains.
Real roots require the quantity under the square root to be non-negative:
Read that as a physical statement, because it is a startling one.
If , equation (3.7.47) has no real solution. There is then no radius whatever at which a body with that angular momentum can move in a circle. Not an unstable circle, not a marginal one, none. The effective potential has no stationary point at all, and Problem 3 asks you to confirm that in that case it increases monotonically with , so that a body released anywhere spirals inward and is captured.
What it does not mean. It does not mean a body with small angular momentum cannot exist far from the mass. It can, and it will be on a plunging trajectory rather than an orbit. It does not mean the body is moving faster than light or doing anything else forbidden. And it is not a statement about small alone. The failure is global, which is exactly what makes it different from the Newtonian case.
Newtonian gravity has no analogue of this at all. There, (3.7.48) supplies a circular orbit for every and every radius, without exception, and the barrier always eventually wins over as . Here the new term wins instead.
7.3 · The two roots, and where they merge
When both roots exist they are different orbits with different characters. Factorise (3.7.46) using them: since the numerator of is divided by , with , the derivative is positive for , negative between the roots and positive again for . So
- is a maximum of , so it is an unstable circular orbit, destroyed by any disturbance.
- is a minimum, so it is a stable circular orbit, which a nudged body oscillates about.
As decreases toward the square root in (3.7.49) shrinks and the two roots approach each other. At it vanishes and they coincide:
This is the innermost stable circular orbit, the ISCO. At exactly both and vanish, because the maximum and the minimum have annihilated into an inflection. So the orbit there is marginally stable: neither restored nor expelled at leading order. Inside it there are no stable circular orbits at all, and Problem 3 shows that the unstable branch runs from inward and never reaches below .
In: the effective potential (3.7.43), whose only departure from Newton is one term, together with the definition of a circular orbit as a stationary point.
Out: a chain of consequences, each one a single step from the last.
- A quadratic instead of a linear equation.
- Hence two circular orbits instead of one at each angular momentum, an unstable inner one and a stable outer one.
- Hence a discriminant, and hence a minimum angular momentum below which circular motion is impossible everywhere.
- Hence a smallest stable circular orbit at , whose position depends on the mass and on nothing else.
What it cost: nothing beyond §6. No approximation was made anywhere in §7. The ISCO is an exact statement about an exact solution, unlike the precession of §8, which is perturbative.
7.4 · The number, and what it is the inner edge of
Put a mass in. Take a black hole of ten solar masses, the class produced by the collapse of a massive star. Its ISCO sits at
with for comparison. And , which for the same hole is .
The observational hook. Gas falling toward a compact object does not fall straight in. It has angular momentum, so it settles into a disc and spirals slowly inward as friction removes that angular momentum. At each stage it occupies the stable circular orbit appropriate to its current . When it reaches there is nowhere further to go: the next orbit inward does not exist, and the gas leaves the disc and plunges. An accretion disc therefore has an inner edge, at , and since the temperature and the emitted spectrum depend on how deep the gas gets, the inner edge is a measurable feature and it is one of the ways the mass of a black hole is estimated. Worked example 2 computes how much energy is released on the way in, and the answer is startling.
The mathematics of §7.2 is not special to gravity, and you have met it in a form closer to home. Take a tumour growing logistically and treated at a fixed absolute rate, meaning a therapy that removes cells per unit time independent of burden. The population obeys , and its equilibria satisfy : a quadratic, with roots
Everything §7 says transfers line for line. Two equilibria exist when the discriminant is positive: a lower unstable one, below which the tumour is driven to extinction and above which it grows back, and an upper stable one, the burden the treatment holds it at. As rises the two approach each other. At they merge at , and above that there is no equilibrium at any burden whatever and the population is driven to zero from any starting point. The merged root is where both the first and the second derivative of the right-hand side vanish, exactly as at the ISCO. That merging is the reason a dose–response curve for such a therapy has a threshold rather than a gradual improvement, and it is the same algebra as (3.7.50).
What breaks. Two things, and both matter. In the tumour model is a control that an oncologist turns. Here is a constant of the motion, a label on a whole family of landscapes, and nothing turns it. A given body keeps its forever unless something exerts a torque, which is why the slow inward drift of a disc requires friction and does not happen on its own. And the tumour model is a first-order equation whose equilibria are where the rate vanishes, while (3.7.42) is second order and its equilibria are where a potential is stationary. The shared object is the quadratic and its discriminant, and it is shared exactly. The dynamics either side of it are different.
Circular orbits sit where the landscape is level. Setting its slope to nothing and clearing denominators leaves a quadratic, a different sort of object from what the old theory offers, where the same question has exactly one answer for every radius anyone cares to name. A quadratic has two solutions, or one, or none, and which occurs is decided by how much angular momentum the body carries.
Two is the ordinary case: an outer minimum, where a body sits stably and returns after a nudge, and an inner maximum, a circular orbit any disturbance destroys. As the angular momentum is reduced the two travel toward one another. At one particular value they meet, and below it there is no solution at all: no circular orbit anywhere, at any radius, however the body is arranged. The landscape has turned into a slide.
The meeting point sits at three times the length fixed by the mass alone, and nothing in Newtonian gravity resembles it. There, every radius admits a circular orbit and every one of them is stable. Here there is a floor, its position set by the mass and by nothing else, and gas spiralling inward must stop orbiting and fall when it arrives. That is a prediction about where the inner edge of a disc around a black hole lies, and why such discs have an inner edge at all.
8 · The orbit equation, and perihelion precession
Here is where this section is going. Section 6's equation describes against proper time, which is not what an astronomer measures. We change variables so that it describes the shape of the orbit, meaning against . We derive the Newtonian version first, where the answer is a closed ellipse. The relativistic version differs by one term. Linearising about a circular orbit turns that term into a change of frequency rather than of shape, so the ellipse still repeats but slightly slower than once per turn, and the whole figure rotates. Then we compute the rate for Mercury, with every conversion on the page.
8.1 · From motion to shape
Two changes of variable, and both are standard for central-force problems because both remove a nuisance.
Change one: use as the independent variable instead of . The chain rule and (3.7.35) give the conversion
which is available precisely because is conserved. That is a Killing vector paying for itself a second time.
Change two: use instead of . The motivation is that the attraction goes as and the barrier as , so in terms of the equation will have polynomial coefficients. Then , and
the middle step using . The factors of cancel exactly, which is why the substitution is worth making.
Put (3.7.54) into (3.7.41) and divide by :
We want a second-order equation for the shape rather than a first-order one carrying a square, because the second-order kind is the one we can recognise on sight. So differentiate with respect to and divide through by :
The division by is legitimate wherever that derivative is non-zero, which is everywhere except at the turning points and on exactly circular orbits. And a circular orbit, having constant, satisfies (3.7.56) directly, so nothing has been lost.
8.2 · Newton first, and why an ellipse closes
Delete the term carrying . That is exactly the Newtonian limit, since the deleted term is the only place appears. Then (3.7.56) becomes
This is the simplest differential equation in physics. It is the harmonic oscillator of Chapter 0.8 §4 with a constant driving term, and its general solution is a particular constant plus the homogeneous solution:
with the two integration constants written as an amplitude and a choice of where sits, taken at the point of closest approach. That is the polar equation of a conic, an ellipse for . It is the same expression Chapter 1.4's Worked example 2 obtained by an entirely different route, from the Laplace–Runge–Lenz vector, with in that chapter's variables.
Why it closes. The right-hand side of (3.7.58) has period exactly in , because the coefficient of in (3.7.57) is exactly . So after one full turn the body is at the same radius moving the same way, and the orbit is a closed curve traced over and over. Chapter 1.4 said the same thing in the language of symmetry: the Laplace–Runge–Lenz vector points from the focus to the perihelion, it is conserved for an exact inverse-square law and for nothing else, and its conservation is the statement that the axis of the ellipse never moves. That chapter then wrote, in advance: "Adding any deviation from exact destroys and the orbit precesses. That is exactly what the dial did to the closed orbit in §4.4's figure, and exactly what the Sun's other planets, and then general relativity, do to Mercury (Chapter 3.7)." Here is the deviation.
8.3 · The new term, linearised
Restore it. The full equation (3.7.56) is nonlinear, and we are not going to solve it exactly. Instead we do what Chapter 3.6 §7.3 warned would become the normal state of affairs once the field equations turned out to be nonlinear: expand about a simple solution and keep the first correction, naming the small parameter as we go.
Step 1. The circular solution. A constant has , so (3.7.56) requires
which is §7's circular-orbit condition wearing different symbols. Multiply by and it becomes (3.7.47) exactly.
Step 2. Perturb. Write with . Substituting into (3.7.56), the square on the right expands as :
The terms with no in them cancel by (3.7.59). The term in is dropped, and this is the only approximation in the section. Be careful about why it is safe, because the obvious reason is the wrong one.
The obvious reason would be that , the fractional variation of around the orbit, is small. For Mercury that quantity is the eccentricity , which is not small at all. A error in would be four arcseconds per century, a hundred times the precision §8.5 claims. So that reason will not do.
What actually makes the neglect harmless is that the dropped term is quadratic. Let's take that slowly, because the chapter's headline number rests on it. Square a sinusoid and you get a constant plus a term at twice the frequency, and neither of those is at the frequency of the oscillation itself. A term that carries nothing at the fundamental frequency cannot shift that frequency at first order in the amplitude: it only displaces the orbit slightly and adds a small wobble at double the rate. The shift in therefore starts at second order, at , and that is second order in exactly the small parameter §8.4 names.
Now put a number on it. For Mercury that is a relative change in of order , a millionth of an arcsecond per century. The comparison in §8.5 is against a measurement good to one part in a thousand, so the neglect is legitimate there for any eccentricity a planet actually has. What is left is
Step 3. Read it. Equation (3.7.61) is the harmonic oscillator again, but the coefficient of is no longer . Its solution has period in rather than , and because the correction is subtracted. So the radial distance returns to its minimum later than one full turn.
Finally, replace by its leading value. From (3.7.59), up to a relative correction of order , which we have already agreed to neglect at the same order:
8.4 · The advance per orbit
The perihelion recurs when has advanced by , that is after . Subtract the full turn to get the advance:
using the binomial expansion of Chapter 0.3 §2.4, valid because the small parameter is
which is for Mercury. Now convert into something an astronomer reports.
Equation (3.7.64) substituted , which is the Newtonian relation (3.7.58) between the angular momentum and the semi-latus rectum. In the relativistic problem that relation is not exact. Section 7's (3.7.47) already shows the correspondence between and orbital size is modified.
Why it may be used anyway. The quantity being computed, , is already first order in the small parameter , since it vanishes identically when that parameter is zero. Replacing by therefore changes at second order, which for Mercury is a relative change of about one part in , or roughly a millionth of an arcsecond per century. That is far below the precision of the comparison in §8.5. Using a lowest-order relation inside a first-order result is consistent. Using it inside a zeroth-order result would not be.
This is the same discipline Chapter 3.1 §7.2 used when it integrated the light deflection along the undeflected straight line: evaluate a small correction along the uncorrected solution, so that what is neglected is the correction to the correction.
With , equation (3.7.63) becomes
The second relation is pure conic geometry and takes two lines. From (3.7.58), the closest and furthest distances are at and at . Their sum is the major axis , so
8.5 · Mercury, with every conversion shown
Mercury's orbit, as measured: semi-major axis , eccentricity , orbital period days. And , .
Step 1. The semi-latus rectum. , so
Step 2. Radians per orbit. Substitute into (3.7.65). The numerator is . The denominator is . So
Half a microradian. Per orbit, this is undetectable. The effect is extracted by waiting.
Step 3. Orbits per century. A Julian century is days, so
Step 4. Multiply, and convert. The advances add, since each orbit turns the ellipse by the same angle:
One radian is degrees and each degree is arcseconds, so arcseconds. Hence
Mercury's perihelion is observed to move by a large amount, most of which is not this effect at all. The bulk of the apparent motion is the precession of the equinoxes, a rotation of the coordinate frame in which positions are reported, and therefore not a fact about Mercury. Most of the rest is the Newtonian gravitational pull of the other planets, chiefly Venus and Jupiter, which is computed from Newton's law without any relativity. Subtracting both leaves a residual that resisted explanation from Le Verrier's announcement of it in 1859 until 1915.
⚑ That residual is measured to be arcseconds per century. The number is quoted here, not derived: extracting it requires the observational record and the Newtonian planetary perturbation theory, neither of which this book builds.
Equation (3.7.71) gives , which sits inside the measurement's own uncertainty of about one part in a thousand. It was obtained from the mass of the Sun, the shape of Mercury's orbit, and a metric derived in §3, with no adjustable parameter anywhere. Nothing was fitted. This is the calculation Einstein performed in November 1915, before the field equations were published, and by his own account it gave him palpitations.
8.6 · Why Mercury, and not the Earth
Run the same arithmetic for the Earth: , , days. Then , and
and with orbits per century the total is arcseconds per century, smaller than Mercury's by a factor of .
The reason is visible in the formula and it explains the whole history of the problem. Two factors work in the same direction:
where Kepler's third law , from Chapter 1.4 §3, supplied the second step. The inner planet wins twice. Its orbit is smaller, so each turn shifts the ellipse further, and its year is shorter, so it takes more turns per century. Eccentricity helps a third time, weakly, and it helps in another way the formula does not show: a nearly circular orbit has no well-defined perihelion to track. Venus, whose eccentricity is , is a poor target for that reason despite being closer to the Sun than the Earth. Mercury is the innermost planet and much the most eccentric. It is the only place in the Solar System where the effect was ever going to be seen first.
Changing variables to the reciprocal of the distance converts the orbit from a statement about motion into one about shape, and in the old theory the resulting equation is the simplest in physics: a thing whose second derivative plus itself equals a constant. Its solutions repeat exactly once per turn, which is why an ellipse closes rather than drifting. An earlier chapter said the same from the side of symmetry, with a vector from the focus to the closest approach that never moves.
The relativistic version adds a single term, quadratic in the reciprocal distance and very small. Treated as a small correction about a circular solution, its effect is not a change of shape but of rate: the pattern still repeats, slightly slower than once per turn. The body therefore reaches its closest approach a little beyond where it did last time, and the whole ellipse turns, forever, at a fixed rate.
Two features deserve emphasis. This is a perturbative result, legitimate because the correction is one part in ten million even for the innermost planet, and worthless otherwise. And the turning per orbit depends on the orbit through one quantity only, its width measured across the focus, so a tight eccentric orbit gives the largest shift per turn and completes the most turns per century. That is why the discrepancy showed up at Mercury and nowhere else.
9 · Worked examples
Find the angular velocity of a circular orbit in the Schwarzschild geometry, measured in the distant observer's time . The answer is , which is Kepler's third law with no correction of any order. That is surprising enough to be worth three qualifications.
Use the geodesic equation directly. Chapter 3.3 §8 gives it as
Take , the radial component, on a circular orbit in the equatorial plane, where is constant so and . The only surviving connection coefficients from (3.7.8) with both lower indices among are and , so
Substitute the coefficients. With and , and . Dividing the whole equation by , which is non-zero outside :
Evaluate . With we get , so
using to convert. Writing and for the orbital period, : Kepler's third law, unmodified, exact.
Three qualifications, and each of them matters.
(i) is the distant observer's time, not the orbiting body's. Per unit of the orbiting body's proper time the answer is different, and the difference is not small close in. What is exact is the relation between the areal radius and the period as booked by someone far away.
(ii) is the areal radius, which by §2.2 is not the distance to the centre. Kepler's law is exact in terms of the label defined by the area of the sphere, and would not be in terms of a radial distance.
(iii) It says nothing about stability. The relation holds for every where a circular orbit exists at all, including the unstable branch. An exact and elegant formula for the period of an orbit that any disturbance destroys is a useful reminder that "exact" and "physically realisable" are different properties.
The result is a genuine coincidence of the areal-radius chart rather than a deep fact, and it is worth knowing because it makes the deviations elsewhere easier to trust: if Kepler's third law survives untouched, the arcseconds of §8 cannot be blamed on a sloppy definition of orbital period.
Compute the conserved energy of a circular orbit as a function of , evaluate it at the ISCO, and interpret the difference from the rest energy.
The energy of a circular orbit. On a circular orbit , so (3.7.42) reads , that is
We also need for that orbit. Solve (3.7.47) for rather than for , which is easier because the equation is linear in :
Two things are already visible. is positive only for , so there is no circular orbit of any kind inside . And as approaches that radius from outside. Chapter 3.8 §2 identifies what is happening there.
Substitute and simplify. Putting this into (3.7.43) and then into the expression for gives, after collecting terms over the common denominator ,
Check it far away: as both the numerator and the denominator tend to , so , the rest energy per unit mass of a body at rest at infinity, as (3.7.34)'s sanity check demanded. ✓
At the ISCO. Put , so that and :
What that number is. A parcel of gas starting at rest far away has . By the time it reaches the ISCO it has . Since is conserved along a geodesic, the difference cannot have gone anywhere by itself. It is carried away by whatever friction moved the gas from one circular orbit to the next. So the fraction of the rest energy released on the way in is
Set that beside the alternative. Hydrogen fusing to helium releases the mass deficit of Chapter 2.5 §9: four protons weigh atomic mass units and a helium-4 nucleus weighs , so the fraction converted is . Accretion onto a non-rotating black hole, stopping at the ISCO, releases over eight times as much from the same mass, and the fuel does not need to be hydrogen. This is the reason a quasar, powered by gas falling onto a hole, can outshine the entire galaxy of stars around it. The gravitational field is not a store of energy being consumed. It is a fixed geometry, and what is being spent is the rest mass of the gas.
And a limit worth naming. Nothing in §7 fixed the ISCO at for a rotating hole. Section 1.2 excluded rotation from the start, the exterior geometry of a spinning body is not (3.7.22), and both the innermost stable orbit and the fraction above would have to be recomputed from a solution this book does not build. What is derived here is the non-rotating figure, and nothing in this chapter licenses quoting it for a spinning hole.
Apply (3.7.65) to the star S2, which orbits the compact object at the centre of the Milky Way. The point is that the formula was derived with no assumption of weakness, so it should be tried where the effect is large.
The inputs, all measured quantities of the same kind as Mercury's: the central mass is , so . The orbit has and .
The semi-latus rectum. , so , and with that is .
The advance per orbit.
Convert: degrees, which is . That is per orbit, not per century, and arcminutes, not arcseconds. The same formula that produced radians for Mercury produces times as much here.
Where the factor comes from. Entirely from . The mass is larger by and the semi-latus rectum is larger by , and . ✓ It is worth checking that the perturbative treatment of §8.3 still applies. The compactness at closest approach is , and with and that is . That is some eleven thousand times Mercury's , and still four orders of magnitude below unity. The expansion is safe, though a good deal less safe than in the Solar System.
What this example does and does not establish. It computes a prediction. Whether the prediction is confirmed is a question about instruments, since resolving a star's orbit around an object eight kiloparsecs away requires interferometry at the milliarcsecond level, and this chapter derives no observational result. What it does show is that (3.7.65) was never a formula about Mercury. It is a formula about , and there are places in the sky where that number is four orders of magnitude larger than in the Solar System.
10 · Your turn
Problem 1 — the two constants of integration, and what each one bought
The derivation of §3 produced two constants and fixed them by two different physical demands. (a) Suppose had been solved as for some constant . Carry through §§3.5–3.6 and write the resulting line element. (b) Show that a rescaling of the time coordinate removes entirely, and say precisely what a distant observer would have to believe for to be the natural choice. (c) The constant of (3.7.19) was fixed by matching to the Newtonian potential. What would describe, and is it excluded by anything derived in this chapter? (d) Chapter 3.6 §6 permitted a cosmological term. Say in one sentence why ignoring it in (3.7.6) is legitimate for the Solar System.
Solution
(a) With we have , and the third line of (3.7.10) becomes, by the same three substitutions as (3.7.16), . Setting it to zero gives , so , and
(b) Define . Then , and dividing numerator and denominator of the radial term by shows the metric is (3.7.22) with in place of . So is not physics. It records that the coordinate was not chosen to be the proper time of an observer at infinity. For to be natural, a distant observer would have to be using a clock running at a rate times their own proper time, which nobody does.
(c) By (3.7.21), , so means : a negative mass, whose exterior geometry would repel test bodies. Nothing in this chapter excludes it, because knows nothing about mass and the sign entered only through the matching. What excludes it is the observed sign of gravity, plus the energy conditions that Chapter 3.9 flags, neither of which is derived here.
(d) Because while the curvature scale set by the Sun at Mercury's orbit is , larger by twenty-two orders of magnitude. That is the same estimate Chapter 3.6's Problem 1(c) made.
Problem 2 — radial infall, and a coincidence that is not one
Take in (3.7.40), and a body released from rest at infinity, so that by §5.1's sanity check. (a) Show that exactly, which is the Newtonian escape-speed expression with no correction of any order. (b) Integrate to find the proper time to fall from rest at to , and check the Newtonian formula is reproduced exactly. (c) Evaluate for a body falling from the Schwarzschild radius of a ten-solar-mass hole. (d) State carefully what has and has not been shown about the surface at .
Solution
(a) With , (3.7.40) is . Put : the first term is , and expanding the second gives . The cancels exactly and . Taking the inward root, . Note that no expansion was made: this is exact in the full metric.
(b) Separate and integrate, which is Chapter 0.8 §2.1's technique:
That is the Newtonian free-fall time from rest at infinity, unchanged.
(c) For , and m. Then , , and . Sixty-six microseconds from the horizon to the centre.
(d) What has been shown: the proper time to cross and reach is finite, and nothing in misbehaves at , where that expression is perfectly smooth. What has not been shown: anything at all about the coordinate , which does not appear in the calculation, or about whether the traveller can send signals out. Both of those are Chapter 3.8's business, and the contrast between them is the point of that chapter.
Problem 3 — the inner branch, and the barrier that fails
(a) For , show that for every , and describe the motion of a body released from rest far away. (b) Show that the unstable root of (3.7.49) satisfies for every admissible , and find its limit as . (c) Show that the stable root satisfies and that for large , recovering (3.7.48). (d) Say in one sentence what happens to (3.7.43) as for any , and contrast it with (3.7.44).
Solution
(a) By (3.7.46), has the sign of the quadratic (3.7.47), whose leading coefficient is positive. A positive-leading quadratic with negative discriminant is positive everywhere, so for all : the potential rises monotonically from at the origin to at infinity. A body released from rest far away has , and since everywhere the turning-point condition is never met. It falls all the way in, whatever its angular momentum, provided only that .
(b) Write , so (3.7.49) reads , defined for . At the root is . For large , expand the square root: , so from above. Differentiating shows decreases monotonically in , so it is confined to the stated interval. The limiting radius is where Worked example 2's diverges. It is the same radius, reached from the other side.
(c) equals at and increases with . For large the same expansion gives , which is (3.7.48). The Newtonian answer is the large-angular-momentum limit, as it must be, since large means a wide orbit.
(d) As the term dominates both others, so . By contrast , dominated by the barrier . The Newtonian barrier is infinitely high and the relativistic one is finite and then falls away, which in one line is the whole difference between the two theories close to a mass.
Problem 4 — where the ISCO is real, and where it is fiction
The solution (3.7.22) holds only in vacuum, so a feature at radius is physically present only if lies outside the body. (a) Compute for the Sun (, m), for a white dwarf (, m) and for a neutron star (, m). In which cases does the ISCO lie outside the body? (b) For the Sun, what is at in reality? (c) A spherical star collapses to a black hole. Explain, using §4, why the exterior geometry, and therefore the ISCO, is unchanged throughout the collapse. (d) What, then, actually changes for an orbiting body as the collapse proceeds?
Solution
(a) per solar mass, that is per solar mass. Take the three cases in turn.
- Sun: against , deep inside, by five orders of magnitude.
- White dwarf: against , inside.
- Neutron star: against about , outside, marginally.
So of the three, only the neutron star has a real ISCO, and it sits barely above its surface. That is why the inner edge of an accretion disc around a neutron star is a sensitive probe of the star's radius.
(b) Solar material, about from the centre of the Sun, at a density and temperature the vacuum equations know nothing about. The metric there is a solution of with , which this chapter does not compute. There is no ISCO, no horizon and no Schwarzschild radius inside the Sun. Those are features of a solution that does not apply there.
(c) Birkhoff, (3.7.32). The exterior is spherically symmetric and in vacuum throughout the collapse, so it is the Schwarzschild solution with some . And is fixed by the far-field matching (3.7.21), which cannot change while nothing crosses the region. So the geometry outside the collapsing surface is the same at every stage.
(d) Nothing about the geometry where the orbiting body is. What changes is how much of the Schwarzschild geometry is real: as the surface falls past , the ISCO stops being a point inside the star and becomes a genuine feature of the vacuum. A distant orbit feels no disturbance at any moment. That is the content of §4.3, and the reason the formation of a horizon is not an event that announces itself.
You have solved Einstein's equations. Two symmetry conditions, written as statements about Killing vectors rather than about pictures, cut ten unknown functions of four variables down to two functions of one. The areal radius, (3.7.4), did the last step as a choice that Chapter 3.8 will have to un-make. Nine Christoffel symbols, three independent Ricci equations, and then one line: collapses to , so is constant, and asymptotic flatness makes it one. The angular equation becomes , and one integration gives . Then the Newtonian limit of Chapter 3.1, namely , derived there from redshift alone, fixes . The mass was never in the field equations. It arrived as a constant of integration. And is km for the Sun and mm for the Earth.
Birkhoff, derived rather than quoted. Drop staticity, allow both functions to depend on time, and the component of reads . So cannot depend on time, and the residual time dependence in is removable by rescaling a clock. The Schwarzschild solution is the only spherically symmetric vacuum solution. A star may pulse, collapse or explode and its exterior does not change. There is no spherically symmetric gravitational radiation. And Chapter 3.4's insistence that Ricci-flat is not flat has its example.
Chapter 3.5's promise, collected by name. The metric mentions neither nor , so and are Killing vectors, and Chapter 3.5 §9's theorem converts each into a constant along every geodesic: and , which Chapter 1.4's Noether theorem identifies as energy and angular momentum per unit mass. Four coupled second-order equations became one first-order equation, and the normalisation turned it into motion in a landscape.
One new term, and everything that follows from it. The effective potential is Newton's plus , a fraction of the centrifugal barrier. Circular orbits satisfy a quadratic where Newton had a linear equation, so there are two of them, an unstable inner one and a stable outer one. The discriminant vanishes at , below which no circular orbit exists at any radius at all. The two merge at , which is km for a ten-solar-mass hole and is the inner edge of an accretion disc. Getting there releases of the rest mass, eight times what fusion yields.
And an ellipse that does not close. In terms of the orbit obeys , whose Newtonian truncation gives the closed conic of Chapter 1.4's Laplace–Runge–Lenz vector. Linearising about the circular solution turns the new term into a frequency shift, , so the radius repeats every rather than every and the perihelion advances by per orbit. For Mercury that is radians per orbit, orbits per century, arcseconds per century, against a measured residual of ⚑ . Nothing was fitted.
Where this gets spent. Chapter 3.8 takes the same metric and sends light through it: the effective potential of §6 with in one place, an unstable circular orbit for photons where Worked example 2's diverged, a deflection of , and the collection of Chapter 3.1 §7.3's confessed factor of two. It then asks what the surface at is, and the answer turns entirely on §2.2's warning that is a label chosen for convenience. The coordinates fail there and the geometry does not, and an invariant built from the curvature settles it. Chapter 3.9 puts matter back on the right-hand side and solves the equations for the universe rather than for a star, using the same three moves: symmetry to write an ansatz, the field equations to get an ordinary differential equation, and a boundary condition to fix a constant.