Reference
The Through-Line
The whole book in plain language, assembled from every “In plain terms” passage in order. No mathematics.
278 passages · about 65,759 words · Parts 0 to IV of eight, complete through cosmology · reads start to finish as one essay · Math Ledger
What follows is the whole book, with the mathematics taken out.
Each passage below was written to close a section of a chapter, in the reader's own language rather than the chapter's. Read in place they are pauses for breath. Read here, one after another, they are a continuous account of how physics arrived at its present picture of the world, with no equation anywhere in it. The finished arc runs from the definition of a derivative to what string theory is actually claiming; what exists today runs as far as the field equations of gravity, and the closing passage at the foot of the page says what that leaves owed.
One idea runs underneath all of it. Ask what stays the same when you change your point of view. Whatever survives is real; whatever does not was a fact about where you were standing. Everything that follows is that sentence, applied to progressively larger questions.
Four moves recur, and their names are worth having before meeting them. The first is a choice of perspective, which is what coordinates, bases, frames and gauges all are, and separating what depends on it from what does not is this book's central habit. The second is a problem falling apart into independent pieces, which happens under a different name in nearly every part and is one theorem spent over and over. The third is that nearly everything is an approximation, since physics does not usually solve its problems but expands them, and the honest version of the method names what was thrown away. The fourth is the thing that had to exist: most of the objects here were not invented but cornered, forced into being by a requirement rather than chosen for convenience.
Part 0 builds tools rather than physics. But the tools are not neutral, and the choices made in building them decide what can be said later.
Part 0 · The Toolkit
Rebuild the floor. No physics yet — but every example comes from physics.
0.1 What a Derivative Really Is
Almost everything physics has to say concerns things that change, so the first thing the subject needs is a way of saying how fast something is changing right now, at a single moment, rather than on average across a stretch of time. That is harder to state than it sounds. Over a stretch you can divide the distance covered by the time taken and get an honest answer, but a moment has no duration, and no distance divided by no time yields nothing at all.
The way out is to stop forcing the arithmetic. Rather than asking what the ratio equals once the interval has shrunk to nothing, we ask what it is heading towards as the interval shrinks, and that question has an answer because the top of the fraction dwindles at least as fast as the bottom does, so the ratio settles down instead of running away.
This is the first appearance of a habit that runs through the whole book. A quantity you cannot reach directly gets defined by what a sequence of quantities you can reach approaches, and such a definition is worth having only when that approach is guaranteed rather than hoped for. A speedometer is evidence that instantaneous speed is real; this chapter is the account of why it is also well defined.
Written out formally, the definition carries a clause that is easy to read past: the derivative exists whenever the limit exists. That clause is not legal boilerplate, and it is the most informative part of the sentence. It demands a single number that the ratio approaches however the interval shrinks, and in particular the same number whether you close in from the left or from the right. Where two different answers arrive, there is no best straight line through the point and nothing for the derivative to be.
Being able to draw a curve without lifting your pen is not enough to guarantee this. A graph with a sharp corner is perfectly unbroken and still gives two answers at the corner, and worse is possible: functions that are unbroken everywhere and have a derivative nowhere, wrinkled at every magnification, so that zooming in never straightens anything out.
Those are not museum pieces. Physics assumes smoothness constantly and almost always silently, and the places where the assumption fails are reliably the places worth studying, among them shock waves, phase transitions, and the paths a quantum particle takes. It is worth knowing from the first page that smoothness is an assumption being made rather than a property of the world.
What happens here is a deliberate replacement, and it deserves to be flagged as one. The definition given at school, that a derivative is the slope of a graph, is being set aside for one that mentions no graph at all: the derivative is the number that makes a straight-line stand-in for a function as good as a straight line can be. On a flat page in one dimension the two agree, which is why slope was ever taught, but slope is the special case that will not survive curved space, several variables, or the later chapters in which the thing being varied is a whole history.
The precision that makes the new definition work lies in how the error is described. To say the leftover is small would be worth almost nothing, since nearly any approximation has a small error when the displacement is small. The claim is far stronger: the leftover shrinks faster than the displacement it accompanies, so that measured against the displacement it disappears entirely. Exactly one coefficient achieves that, and uniqueness is what turns a good guess into a definition.
The leftover does not vanish without trace, though. Divided by the square of the displacement it settles on a definite number rather than dying, so the error is specifically quadratic, and that one observation is the door into the two chapters that follow.
Every rule you were once asked to memorise falls out of that one sentence, usually in a line or two. Deriving them once is worth more than rehearsing them ten times, because the derivations are what carries over into settings where the memorised forms mean nothing.
The product rule is the clearest case. Picture a product as the area of a rectangle and nudge both sides outwards: the area gains a strip along the top, a strip along the side, and a small square in the corner where the two growths overlap. Each strip is proportional to the nudge and survives; the corner is proportional to the nudge twice over and dies. That is the entire rule, and discarding a second-order corner is a move you will make for the rest of the book. Much later, the pieces that refuse to die when a vector is carried around a closed loop turn out to be the curvature of spacetime.
The chain rule is where the reframe pays for itself immediately. Composing two functions near a point amounts to composing their straight-line stand-ins, and straight-line maps compose by multiplication. The same sentence holds with matrices in place of numbers when there are several variables, and in general relativity it becomes the rule that defines a tensor. There is no route to any of that from slope times slope.
The exponential is not a function anyone was fond of; it is one that was cornered into existence. Differentiate any base raised to a power and something striking appears before a base has even been chosen: the rate of growth comes out proportional to the current size, whichever base you picked. All that distinguishes one base from another is an awkward constant left sitting in front, so the natural move is to ask which base makes that constant equal to one.
That request has exactly one answer, and the answer is the number we call e. Notice the direction of the definition. The value 2.718 and so on is not the starting point but the consequence; the starting point is the demand that a quantity be its own rate of change, and the number is whatever has to be true for that demand to be met. A great many objects in this book arrive the same way, forced into existence by a requirement rather than chosen for convenience.
You already rely on the pattern. Anything cleared at a rate proportional to how much is present decays exponentially, because nothing else can. Hold onto that shape: the fundamental equation of quantum mechanics has the identical form with one extra factor in it, and that factor turns shrinking into turning. Three chapters from here you will see exactly why.
Trigonometry needs one genuinely geometric input before any of it can be differentiated, and the input arrives by trapping a circular sector between two triangles and comparing their areas. Everything else follows from that single squeeze.
Where a choice becomes visible is in the units. The clean result, that the sine of a small angle is very nearly the angle itself, holds only when angles are measured in radians, which is the unit in which an angle is a length along a circle rather than a turn cut into a convenient number of parts. Measure in degrees and a conversion factor lodges itself in the derivative and never leaves, so every formula downstream inherits the litter. This is the first instance of something the book returns to constantly, which is that a description always rests on a choice, and that one choice among many makes the structure visible while the rest bury it.
The consequence worth carrying away is a single line. Differentiate the sine twice and you get the sine back with its sign reversed, which is to say a quantity proportional to the negative of its own second rate of change. That is the equation obeyed by anything that oscillates, and from a few chapters onward it never really leaves the book.
Differentiating a second time gives the rate at which the rate is changing, and the usual gloss, namely how sharply a graph bends, undersells it. The sharper statement is the one you see on magnifying the graph: the second derivative is the size of the leftover from replacing a function by a straight line. Bending is what that looks like when drawn; the leading error is what it is.
This is also why mechanics is hard, in a specific and honest way. Newton's law is a statement about the second derivative, so you are handed an acceleration and asked for a position, which means accumulating twice and picking up two unknown constants. Those constants are where the object started and how fast it was going, so a law of motion does not by itself determine a motion; law and starting conditions together do.
One last thing, because it looks like cheating and is not. Physicists write derivatives as though they were fractions and cancel the pieces freely, and the licence is real: the derivative is a map from a small input displacement to a small output one, so the manipulation is an exact statement about such maps rather than an approximation. You leave with the derivative in the form that generalises, and with a reason to keep chasing the leftovers, which two chapters from here becomes a method.
0.2 Integration and Accumulation
Accumulation is the reverse problem, and the one physics spends most of its time on. You are handed a rate and asked for a total: a velocity and asked for a distance travelled, a concentration and asked for a total exposure. The strategy is forced. Chop the interval into pieces short enough that the rate barely changes across each, treat it as constant there, add up the contributions, and ask what happens as the pieces are made finer.
The demand hidden in that last step is severe, and worth appreciating rather than reading past. It is not enough that some sensible way of chopping settles on a number. Every way of chopping has to settle, on the same number, wherever inside each piece you take your sample. A function taking one value at every fraction and a different value at every other number fails this completely, because the answer then depends entirely on where you looked.
There is a dividend hiding in where you take that sample. Put it at the centre of each piece and the drift overshoots on one side by exactly what it undershoots on the other, so the leading error cancels without anyone arranging it. Sampling where the error is antisymmetric is a manoeuvre worth recognising now, because the same cancellation does real work again before this chapter is out.
Nothing so far has connected the two halves of calculus. One concerns how a function behaves in a vanishingly small neighbourhood of a point, the other a total swept out across a whole interval, and on the face of it they share no vocabulary. What emerges is that they are the same question, which is genuinely astonishing and does not become less so with familiarity.
The mechanism is easier than the statement. Let the right-hand end of an accumulation slide forward a little and ask what the total gains. It gains a thin sliver, and that sliver is a rectangle whose height is the rate at that end, plus an error that is second order and therefore dies. It is the corner-square argument of the last chapter in a different hat. So the accumulated total changes at a rate equal to the thing being accumulated.
Read the other way round, the theorem says something structural that will outlast the calculus. Cut a region into small pieces, add up the local change across each, and every interior contribution appears twice with opposite signs and cancels against its neighbour. Only the two ends survive. Local changes added up over a region equal something evaluated purely on that region's boundary, and that sentence, in higher dimensions and better clothes, is the structural fact this book returns to most often.
Because an integral is evaluated by finding something whose rate of change is the integrand, integrating is not really a procedure. It is recognition, and the whole technique consists of reading the differentiation rules backwards. The chain rule read backwards is substitution; the product rule read backwards is integration by parts.
The second deserves to be read as an instruction rather than a formula. It says you may move a derivative off one factor and onto the other, and the price is a change of sign together with a term evaluated only at the two ends. Put that way it stops being a device for computing textbook integrals and becomes the manipulation theoretical physics leans on more than any other. It is the step that converts a statement that some total is stationary into an equation the system must obey at every point, which is where every field equation in this book comes from.
The term at the ends deserves attention of its own, because physicists discard it habitually. Doing so is not a piece of algebra. It is a claim that whatever you are describing is quiet enough far away for that contribution to vanish, and there are important situations where the claim is false and the discarded term is the interesting object. The entropy of a black hole is one of them.
Here is an integral that cannot be done in the ordinary way, and the obstruction is a theorem rather than a shortage of ingenuity: the bell curve has no antiderivative among the usual functions. With no elementary route in, the section takes one that looks like sleight of hand and is not. Square the integral, so that one copy of the problem becomes two independent copies laid at right angles, and read the result as a single integral over a plane.
Why that helps is the part to keep. On the plane the integrand depends only on distance from the origin, and rewriting it in terms of distance and angle forces a factor of the distance into the patch of area, since a patch far out covers more ground than the same span of angle close in. That factor is exactly what the one-dimensional problem lacked. Geometry supplies what algebra could not, and the pi in the answer comes from the angular direction, which is to say from the existence of a circle. Nothing in the original problem could have warned you it was coming.
Learn this one properly, because an unreasonable quantity of physics rests on it. Attach a source to the exponent, differentiate with respect to the source, and you have in miniature the working method by which quantum field theory extracts every prediction it makes.
The infinity written at the top of these integrals is shorthand for a limit, and the limit is over ordinary integrals across larger and larger finite regions. Whether it exists is decided, in nearly every case a physicist meets, by comparison against a power law: a steep enough power converges at the far end, a mild enough one converges at the near end, and the two conditions pull in opposite directions.
That opposition has a consequence worth carrying. No pure power law behaves at both ends at once, so something must change character in between, which is why a genuine physical answer always has a scale buried in it. When quantum field theory produces a quantity that runs away at short distances, that is this same test failing at the far end rather than an exotic disease, and extracting a finite prediction anyway is what renormalisation means.
One honest admission to close on. Integration, unlike differentiation, is not an algorithm, and most functions have no antiderivative you can write down. The handful that can be integrated therefore carry far more weight than their number suggests, and part of the reason bell curves are everywhere is that they are what we can integrate. It is worth being able to say which parts of a formalism are nature and which are our own limitations.
0.3 Series, Approximation, Orders of Magnitude
Stated without softening, the claim this chapter rests on is that physics does not usually solve its problems; it expands them. The systems with an exact closed-form answer could be listed on a postcard, and everything else you have read about was obtained by finding a small quantity, expanding in it, keeping the terms that mattered and knowing what was thrown away. That last clause is not an apology attached to the method. It is what makes it a method rather than a guess.
Which is why the error term is derived here first and the series second. A series offered without an estimate of what truncating it costs is not a result but a wish, and the derivation keeps everything exact at every stage: the remainder is written down in full, as an explicit accumulation, never approximated and never quietly dropped.
The most useful form of that remainder says something almost too convenient to believe. The error you commit by stopping at a given term is itself the very next term of the series, with the relevant derivative evaluated at some point in the interval that you have no way of identifying. You do not need to identify it. A bound on how large that derivative can get is enough, and this is precisely how a working physicist manages to be casual and extremely accurate at the same time.
Only a handful of expansions do most of the real work in physics, and the thing worth noticing about each is what its first correction is for. The leading term of the sine series is the small-angle approximation everyone uses; the term after it reports the size of the error being made, and therefore says exactly when to stop using the approximation. An expansion that comes with its next term is an approximation carrying its own warranty, and that is the whole difference between a rule of thumb and a controlled statement.
The genuine surprise here comes from noticing that the exponential series never asks what kind of number its argument is. Feed it an imaginary quantity and sort the resulting terms by whether their power is even or odd, and the even ones assemble into the cosine while the odd ones assemble into the sine. Those were never three separate functions. They are one object, sorted by parity.
This collects the promise left standing at the end of the first chapter. Multiplying by the imaginary unit is a quarter turn in the plane, so the equation that made a concentration shrink away becomes, once that single factor is inserted, an equation that rotates something at fixed length instead. Quantum mechanics conserves probability because probability is the squared length of a thing being rotated, and rotations do not change lengths.
An infinite sum is a promise rather than a completed act, so there has to be a test for when the promise is kept, and comparing each term with the one before it settles nearly every case that arises in practice.
Then comes the unsettling part. Take a function that is smooth, bounded, positive and completely uneventful along the whole real line, and its series still gives up at a definite distance from the origin, beyond which each additional term makes the answer worse rather than better. Nothing whatsoever happens to the function at that distance. Nothing on the real line accounts for the boundary, and no amount of looking harder along the real line ever will.
The explanation lies off the line entirely. Allow the variable to be complex and the function has two points where it blows up, both at exactly that distance from the origin, and the radius is set by whichever bad point is nearest, including bad points in directions the original problem never visits. This is the first time the description you have been using turns out to be a restricted view of something larger, and it will not be the last. It becomes physics later on: the structure of a scattering amplitude in a region no experiment can reach is what fixes the particle masses and lifetimes measured in the region experiments can.
Everything up to here has concerned series that converge, and most of the series physics actually runs on do not converge, for any value of their parameter, ever. They remain the most accurate predictive tools humanity has built, and the contradiction is only apparent. Dissolving it takes one distinction.
There are two different things a statement about a series can mean. One fixes the parameter and asks whether adding terms forever eventually arrives at the answer. The other fixes the number of terms and asks whether the approximation improves as the parameter is made small. Physics almost always wants the second, because the small parameter is handed to you by nature while the number of terms is your choice, and a series can satisfy the second demand perfectly while failing the first everywhere.
Such a series behaves in a characteristic way. Its terms shrink for a while, reach a smallest term, and thereafter grow without limit, and while you are still in the shrinking regime the error is no larger than the first term left out. So the series has a best place to stop and a best accuracy it can reach, both calculable in advance. The perturbation series of quantum electrodynamics diverges, and it predicts the electron's magnetic moment to twelve significant figures, because nobody is remotely near the turning point and the first neglected term is minute.
The last tool in the chapter costs almost nothing and returns answers that look as though they should have required work. It rests on one observation: the units you measure in are a choice you made, so a law of physics cannot depend on them, and that demand alone constrains what an answer may look like. Applied to a pendulum it delivers the dependence on length and on gravity and rules out any dependence on the bob's mass, before an equation of motion has been written.
Pushed harder the method produces something startling. Take the constant belonging to relativity, the one belonging to quantum mechanics and the one belonging to gravity: there is exactly one length you can build from them, not one among many. It is forced into existence, it is unimaginably small, and it is why a quantum theory of gravity arrives with a built-in scale no experiment can approach. What it can never give you is a pure number, which is why some constants of nature must be measured rather than derived.
A closing caution, since the chapter has been one long argument for expanding things. Having every derivative is not the same as being determined by them: some effects are so flat near the origin that every term of every expansion reports them as zero. They are not fictions. The mass of the proton lives there.
0.4 Vector Spaces and Linear Maps
There is no arrow anywhere in the definition of a vector, and no length and no angle either, which is exactly what gives the subject its reach. What the definition lists is what you are permitted to do to the object: add two of them together, stretch one by a number, and find that the results still belong to the collection you started with. Anything meeting that short list is a vector, and no theorem proved from the list can tell one example from another.
Which is why the examples are so unalike. A polynomial is a vector, a function on an interval is a vector, an array of numbers is a vector, and so is the state of a quantum system. Most consequential of all, the solutions of a linear differential equation form a vector space, because adding two solutions produces a solution and scaling one produces another. Physicists call that superposition and speak of it as a principle of nature. It is nothing of the sort. It is an observation about the equation.
Carry that forward, because it is what makes the abstraction worth its awkwardness. One body of results, proved once and never again, will later solve a room full of coupled pendulums and a hydrogen atom without needing to be told which of the two it is looking at.
Coordinates arrive as a convenience and become a hazard within a page, so it is worth being exact about what they are. Choose a collection of vectors large enough to build everything in sight and lean enough that none of them is redundant, and every vector in the space acquires one list of numbers with respect to that collection. Being able to build everything gets you at least one such list; having no redundancy gets you at most one; a basis is the name for a collection with both properties at once.
The hazard hides in the phrase with respect to. The list is not the vector. It is a description of the vector relative to a choice somebody made, and a different choice yields a completely different list while nothing whatever happens to the thing being described. This is the first appearance of the idea the whole book is organised around, which is the separation of what depends on your point of view from what does not, and the habit is worth installing now, while the stakes are still small.
One thing does survive every choice. Any two bases of the same space contain the same number of vectors, so that count is a property of the space rather than of anybody's description of it. It is the first invariant you meet.
Knowing what a map does to the members of a basis is knowing everything about it, since every other vector is a combination of those few and the map respects combinations. That single observation is where the array of numbers comes from. Record what happens to each basis vector, in coordinates, stack the answers as columns, and you have the array; it is a transcript of that and of nothing else.
Once the array is defined that way, the strange rule for multiplying two of them stops being a rule. Ask which array describes doing one map and then a second, follow the bookkeeping, and rows against columns emerges at the far end with no choice made anywhere. Every odd feature is thereby explained rather than memorised: why the inner sizes must agree, since that number counts the dimensions of the space in the middle; and why nobody need check that the operation is associative, since doing three things in sequence cannot depend on how you bracket the description of it.
The same reading explains the failure that matters most. Two maps applied in opposite orders are two different operations, so the arrays do not commute, and the size of that discrepancy is not a blemish on the notation. Turn a book and then flip it, then flip a book and turn it, and the book finishes somewhere else.
Set two descriptions of one space side by side and a small opposition appears which is worth more than it looks. The new basis vectors are assembled from the old ones by some recipe, and the components of any particular vector are reassembled by the exact reverse of it. They must oppose each other, because their product is the vector, and the vector is taking no part in the change. Lengthen your ruler and every measurement expressed in rulers shrinks by the compensating factor.
For a map from a space to itself the same argument yields a sandwich: translate the new description into the old, apply the map, translate back again. Two arrays related in that way are one map seen twice. Whatever genuinely belongs to the map must be left untouched by the sandwich, and whatever moves was a fact about your bookkeeping rather than about the physics. An individual entry means nothing on its own.
A promise starts here as well. Reflection in a diagonal line looks like a swap of two numbers in one description and like keeping one direction while flipping the other in a better one, and the better description is the one in which the map has fallen apart into independent pieces that no longer interfere. Finding such a description, in general, is what the following chapter exists to do.
Three demands pin the determinant down completely, and not one of them mentions a formula. Ask for a quantity that responds linearly to each edge of a solid separately, that reverses sign when two edges are exchanged, and that assigns the value one to the standard cube. Those are the properties of signed volume, and because only a single quantity can satisfy all three, the determinant and the signed volume of the box spanned by the columns are one thing under two names.
Read it as a statement about the map rather than the array and the useful sentence appears. The determinant is the factor by which a map multiplies every volume, and its sign records whether handedness survived the trip. The celebrated product rule then requires no proof whatever: do one map and then another, and volume is scaled twice, and scaling factors multiply because multiplying is what the word factor means. The algebra people are usually made to suffer proves something visible by inspection once you know what the quantity is.
The vanishing case is the one to keep. A determinant of zero means the image of the cube has been flattened onto something thinner, so a whole direction has been discarded and no map can recover it. That is not a technicality about invertibility. It is the destruction of information, made visible as a collapse.
Adding up the diagonal entries of an array looks like the least promising operation in the chapter, and it is the one this book spends most often. Two facts do the work. The first is that the sum is unaltered when the array is rewritten in another description, which is startling, since the word diagonal has no meaning until a description has been chosen. The second connects it to volume: take a map that barely differs from doing nothing, and the amount by which it multiplies volume is one plus that sum, scaled by how small the disturbance is.
So the trace is the infinitesimal shadow of the determinant. One is a multiplicative statement about a transformation that has happened, the other an additive statement about a tendency, and they are the same quantity caught at two moments — the relation the exponential bears to its own first correction, which is no coincidence.
A surprising amount follows. The rate at which a flowing fluid expands is the trace of its derivative, which is why a certain sum of partial derivatives measures compressibility. Volume in the space of mechanical states is conserved because a particular trace vanishes identically, which is one of the deepest facts in classical mechanics. And there are eight gluons rather than nine because tracelessness is one condition, removing exactly one direction from the candidates.
A direction a map leaves pointing where it was, changing only its length, is the simplest behaviour available, and hunting for such directions is the whole business of the next chapter. Whether any exist depends on which numbers you allowed yourself at the outset, and that is not a technicality. Rotate the plane by a quarter turn and no direction survives pointing along itself, since turning everything is what rotating means. Over the real numbers it has no special directions whatever.
Admit complex numbers and the obstruction evaporates, for a reason almost embarrassing in its brevity: the special directions are the roots of a polynomial, and over the complex numbers every polynomial has roots. That is the second time a description has turned out to be a restricted view of something larger, as the chapter on series said it would not be the last. The rotation that had none now has two, their multipliers phases of unit length, which is Euler's formula arriving unannounced. Rotation and complex phase were never two phenomena.
This settles in advance a question usually raised much later and treated as mysterious. Measurable quantities are going to be these multipliers, so a theory whose maps might have none is no theory of measurement; and evolution in time is going to be a rotation, so probability survives because rotations preserve lengths. Both were decided here, with no physics anywhere in the argument.
0.5 Inner Products, Eigenvectors, Spectral Theorem
Until now the space has had no geometry in it whatever: no lengths, no angles, no way of saying that two directions differ rather than merely being labelled differently. An inner product restores all of that in one stroke, and the question it answers is how much of one vector lies along another. Over the complex numbers the definition is obliged to conjugate one of its two arguments, and that is forced rather than chosen, because without the conjugation a perfectly respectable non-zero vector can be assigned zero length, and the ability to tell two states apart would be lost on the first page.
Out of three short axioms comes an inequality stating that the overlap between two vectors can never exceed the product of their lengths, with equality only when the two are parallel. It has the air of housekeeping. What it actually does is make the word angle legal in a setting where no angles had been defined, and it will do considerably more than that.
Worth saying plainly now, because it is later dressed up beyond recognition. That inequality, applied to two particular vectors assembled from a state and a pair of measurements, is the uncertainty principle. The most quoted sentence of twentieth-century physics is a statement about overlaps, and nothing is added to it afterwards except an interpretation of the letters.
Not all descriptions cost the same, and the difference is large enough to change what is worth attempting. In a general basis, finding the coordinates of a vector means solving a system of equations. Choose the basis so that its members are mutually perpendicular and each of unit length, and the work disappears: every coordinate is a single overlap, computed on its own, without reference to any of the others. Nothing has been approximated. This is why perpendicular bases dominate physics so thoroughly that people forget the other kind exists.
Two consequences arrive immediately. The squared length of a vector becomes the sum of the squared sizes of its coordinates, which later says that the total energy of a vibrating string is the sum of the energies in its modes, and later still that a set of probabilities adds to one. And the claim that the basis has missed no direction can be written as an operation which, inserted anywhere in an expression, changes nothing, so that a difficult object breaks into a sum of easy ones without being altered.
None of it would matter if such bases were rare. They are not. Take any basis at all, subtract from each new vector everything it shares with the directions already settled, and rescale what survives. The procedure never fails, so no theorem ahead has to assume its raw material exists.
Suppose you are stranded outside some subspace and want the point of it nearest to you. Keep the parts of your vector lying along that subspace's own perpendicular directions and discard everything else, and what you kept is the nearest point, uniquely so. The proof is Pythagoras used once. What is left over is perpendicular to the subspace, the error committed by choosing any other point lies inside the subspace, and two perpendicular pieces add as squares, so wandering away can only add.
The sentence to carry off is that nearest point and perpendicular leftover are not two conditions that happen to agree. They are one condition. Every optimisation problem whose answer turns out to be linear algebra is this fact wearing a costume.
The most familiar costume is ordinary least squares. Fitting a model means choosing the combination of predictors closest to the observed outcomes; the combinations you can reach form a subspace; the fitted values are therefore the projection of the data onto it. The equations defining the fit assert that the residual is perpendicular to every predictor, which is why the word normal in their name has always meant perpendicular. The decomposition of variance printed at the foot of every regression table is Pythagoras with the right angle at the fitted values.
Every map acquires a partner the moment its space has an inner product, and the partner is specified by a piece of behaviour rather than by a recipe. It is the map you substitute when you want to shift an operator from one side of an overlap to the other. No basis is mentioned anywhere in that description, which is exactly what makes it worth having. The familiar instruction to transpose an array and conjugate its entries is not the definition; it is what the definition looks like after you commit to a perpendicular basis, and in any other basis it is false.
Three kinds of map are then singled out by how each sits beside its partner. One kind is its own partner. A second has a partner that undoes it. A third merely commutes with its partner, and is the widest family for which the next section's conclusions survive at all.
The second kind earns a word now. A map whose partner undoes it leaves every overlap exactly as it found it, so it alters no length and no angle, and it carries perpendicular bases to perpendicular bases. It is the complex counterpart of a rigid rotation. That is why evolution in time, in any theory where a total probability must stay equal to one, has no choice about what kind of map it is.
Turning and rescaling are what a map does to a vector at once, and the directions in which the turning stops are where the map is at its most legible. Locating them comes down to asking when the map, with a multiple of the identity subtracted from it, stops being invertible, since only then can something non-zero be sent to nothing. That condition is the vanishing of a determinant, which converts the search into finding the roots of a polynomial.
What the polynomial returns depends on which numbers you permit, exactly as the previous chapter warned. It depends on the map too, in a way better watched than read about. Begin with a symmetric object and its two special directions stand at right angles. Add an antisymmetric piece and they lean towards one another, meeting and merging at a definite amount of skew, beyond which they are gone altogether and nothing is left unturned.
That behaviour is what the remainder of the chapter accounts for. Perpendicular special directions are not a general property of maps and may never be assumed; they belong to symmetric ones specifically, and the symmetry does every bit of the work. Notice which way the dependence runs, because an enormous amount will later be resting on it.
At the centre of the toolkit sit three statements, none of them long to prove. A map that is its own partner has multipliers that are real numbers. Its special directions belonging to different multipliers are exactly perpendicular, not merely independent. And there are enough of those directions to describe every vector in the space, with nothing left over.
Taken together they say that such a map, viewed in the right description, is a list of numbers attached to a list of mutually perpendicular directions, each one acting alone and none of them speaking to any other. This is the move most of the book is built from, and it deserves its name: the problem has fallen apart into independent pieces. Coupled oscillators, the modes of a plucked string, the components of a wave, the energy levels of an atom — every one of those is this, and the repetition is not laziness. It is one theorem being spent over and over.
Nothing in the argument is physics. Yet read the three again with the words changed and they say that measurement outcomes are real numbers, that distinct outcomes are perfectly distinguishable rather than merely different, and that any state can be written as a combination of outcomes whose weights add to one. Those are usually offered as postulates about nature. They are theorems about symmetric arrays.
Once an object has been broken into independent directions you can do arithmetic on it one direction at a time, and that is the whole of what it means to take a function of a map. Multiply each direction's number by whatever the function does to that number, and leave the directions themselves alone. This agrees with the obvious alternative of substituting the map into a power series, because perpendicularity kills every cross term, and it goes on working for functions that have no series at all.
The case that matters is the exponential of the imaginary unit times a self-partnered map. Each of the real numbers becomes a phase of unit length, so the result preserves every overlap and belongs to the rotation-like family met earlier. The converse holds as well, so the two families carry the same information and are related by exponentiation, precisely as the real numbers are related to the points of a circle.
That equivalence is more consequential than its length suggests. It is the reason a quantity being observable forces the flow it generates to conserve probability, so that an argument people offer as physics is a line of algebra. It is also the first sighting of the pairing between a generator and the family of transformations it builds, which is how every symmetry in fundamental physics is eventually described.
Two maps of the self-partnered kind fall apart into the same independent directions when, and only when, the order of applying them makes no difference. Half of that is easy: if both are lists of numbers over one shared set of directions, then numbers commute and so must the maps. The other half carries the content. Wherever the first map cannot distinguish between several directions, having given them all the same number, the second goes inside that ambiguity and chooses, and between them they produce one set of directions suiting both.
This is how a physical state comes to be labelled by a handful of numbers, one from each map, the labels fixing the state completely once there are enough maps to leave no ambiguity. The reason such labels can coexist, which is to say the reason a state may have a definite value of one quantity and simultaneously of another, is the argument above and nothing besides.
Run it backwards for the other half. Where two maps fail to commute there is no shared set of directions, so no state exists in which both quantities are sharp, and the size of the failure is what the overlap inequality from the opening of the chapter measures. The two things people find most mysterious about quantum mechanics are one theorem about matrices with an if and only if in it.
0.6 Multivariable Calculus
Freezing every input but one and differentiating in the survivor is the obvious opening move, and it behaves exactly as it did with a single variable, since the frozen inputs are literally constants while the calculation is going on. No new technique is required. There is a trap, though, and it is not a subtle one.
Such derivatives look along the coordinate axes and nowhere else. A point in a plane has a whole circle of directions leading away from it, and a function is at liberty to behave one way along the two axes you inspected and quite differently along everything between them. The standard example has a value depending only on the direction of approach and not at all on the distance travelled, so that it is zero along both axes, a different constant along the diagonal, and every intermediate value on some ray or other. It has no limit at the origin and is not continuous there, and its two axis derivatives exist, are finite, and report nothing amiss.
The moral concerns what a definition must achieve rather than what your daily calculations look like. Sampling finitely many lines cannot control a function, so the honest definition has to constrain every direction at once and constrain them uniformly. That definition was written down in the first chapter, and it needs exactly one alteration.
Here a decision taken on the first page pays for itself, and the debt is worth naming as it is settled. The derivative was defined as the coefficient of the best straight-line stand-in for a function rather than as the slope of a graph, at some cost in apparent fussiness. There is no such thing as the slope of a function of three inputs: there are infinitely many, one per direction, and no principle for preferring any. There is exactly one best linear stand-in. Let the displacement be a vector and the coefficient a linear map, and the definition transfers unchanged.
Everything downstream is a consequence of that substitution. The map, where it exists, is unique, and its array is the table of partial derivatives arranged with one row for each output and one column for each input direction. So the partial derivatives do rebuild the whole derivative, but only once its existence is known by some other route, which is what the trap was there to establish.
A second dividend is easy to overlook. The symbol for a partial derivative conceals which quantity was held fixed, and different choices give genuinely different numbers, as anyone who has lost marks in thermodynamics can attest. The total derivative accepts every displacement at once, so it has nothing to hold fixed and the ambiguity never arises.
For a function returning a single number the linear stand-in is a row rather than a square, and a row of numbers is very hard not to stand upright and call a vector. Do so, then ask which direction increases the function fastest, and the question becomes the one the overlap inequality of the previous chapter answers exactly. An overlap with a direction of unit length is largest when the two are parallel, and its size is then the length of the other. So the steepest direction is the gradient's own, and the rate of increase along it is the gradient's length. The arrow does not merely point uphill; how long it is is how steep the hill is.
The complementary statement comes out of the same computation. Travel along a curve on which the function never changes, and its rate of change along that curve is zero, which says the gradient has no overlap with any direction tangent to such a curve. The gradient is therefore perpendicular to the contours and crosses them by the shortest route available.
The two facts assemble into one picture a walker would recognise. Contours are the lines along which nothing happens, the arrow cuts squarely across them, and it is longest precisely where they crowd together, since crowded contours are where a short step buys the largest change.
Look again at the step where the row was stood upright, because something was spent unnoticed. What differentiation naturally produces is a measuring device: it accepts a displacement and returns the resulting change, needing no notion of length or angle to do so. This is also the honest account of the symbols people have been cancelling since school. Turning it into an arrow required a rule for converting one into the other, and that rule is extra structure, assumed rather than derived. In square coordinates on flat space it is the identity, so device and arrow carry identical numbers and the conversion cannot be seen.
Step outside them and it becomes visible at once. In polar coordinates the conversion divides the angular part by the square of the radius, which is where the notorious reciprocal in the memorised polar formula comes from, one factor from the conversion and one from the angular basis direction not being of unit length. In relativity the conversion reverses the sign of every spatial part, and no coordinates remove it.
The rule to carry forward is short. Differentiation produces measuring devices, arrows are a different species, and the metric converts between them. On a curved space no such object exists until somebody supplies one, and supplying it is what specifying a gravitational field means. The measuring device is always there; the arrow is not.
Composing two functions near a point amounts to composing their linear stand-ins, and that sentence was already the entire chain rule when one variable was in play. All that changes with several is what composing linear maps means, which the linear algebra chapter settled: their arrays are multiplied. The derivation is the earlier one word for word. As a bonus the bookkeeping can no longer go astray: the output space of the inner function is the input space of the outer, so the shapes are forced into agreement.
Written out with indices for the case that governs everything later, in which the coordinates are themselves functions of new coordinates, the same formula says how the components of a rate of change convert from one coordinate system to another. Promoted to a definition, that single sentence is what a tensor is, and the concept holds nothing else.
Two habits appear here of their own accord rather than by decree. One index is always repeated and summed while the others are not, because that is what composing maps does, which is why the summation sign is eventually dropped as redundant. And the repeated index turns up once as a measuring device and once as an arrow — the only pairing yielding a number that survives a change of description, and later the quickest way to catch an error.
One order further in the approximation and every smooth function near a point becomes the same three things: a constant, a linear part, and a quadratic part built from the array of second derivatives. That array is symmetric whenever those derivatives are continuous, and symmetric is exactly the hypothesis the previous chapter spent its length on. So the quadratic part falls apart into independent pieces: along each of a set of mutually perpendicular directions the function is an ordinary parabola, and the numbers belonging to those directions are its curvatures.
The classification then reads itself off. Curving upward in every direction is a minimum, downward in every direction a maximum, and mixed signs a saddle, which is a maximum and a minimum at once, depending on where you stand. A curvature of zero leaves the matter unsettled and hands it to higher order, no idle case: it is where a mechanism for breaking a symmetry eventually lives.
Read curvature as information and it says something you can feel. A sharp peak means its location is well determined and a flat one means it is not, so the curvature of a likelihood surface at its maximum is the precision of an estimate, and the width of a confidence interval is the shallowness of a hill. A long shallow valley is what two nearly redundant predictors look like from the inside.
The rule for optimising under a constraint is famous as a recipe and is not a recipe at all. Being extremal along a surface means that no motion within the surface changes the function to first order, which says the gradient has no component in any direction you are still permitted to move. What remains of it must point straight out of the surface, and pointing straight out of the surface is exactly what the constraint's own gradient does. So the two are parallel, and the number relating them is the multiplier. There is nowhere left to go, and the equation is that sentence written down.
Put geometrically it stops being clever and becomes obvious. At the answer, the contour of the function touches the constraint rather than crossing it. Anywhere the two cross you can slide along the constraint from one side of the contour to the other and change the function, so you were not yet at the answer.
The multiplier is not a device introduced and then discarded. It measures how much the best achievable value improves when the constraint is relaxed by one unit, which makes it a price. Economists call it a shadow price, mechanics the force the constraint exerts, and thermodynamics, in one celebrated instance, the temperature. A multiplier of zero says the constraint was never binding.
Near any point a differentiable map is a linear map, to within an error dying faster than the step you took; and a linear map multiplies every volume by its determinant. Those two facts settle the last question. Chop a region into cells small enough for the first to hold, apply the second to each cell, and the factor appearing in every change of variables is the size of that determinant, which is a local volume magnification and nothing more mysterious.
Two familiar weights thereby stop being items to memorise. The radial factor in plane polar coordinates and the more elaborate one in spherical coordinates both fall out of a short calculation, and the first repays a debt taken out several chapters ago, where it had to be got by cutting up a circle by hand. The same quantity returns in curved spacetime as the square root of the metric's determinant, which is why the volume element of general relativity looks as it does.
That closes the toolkit's account of change. A rate of change is a linear approximation; the description worth having is the one in which a problem separates into pieces that ignore each other; and changing description costs a determinant. What is missing is any account of accumulating a quantity across a region rather than at a point, and that is where the next chapters go.
0.7 Fields, Flux, and the Big Theorems
The object everything from here on is built around differs in kind from anything the toolkit has handled so far. A number is one number and an arrow is one arrow, whereas a field is a value fastened to every point of space at once, so that naming one means naming infinitely many things together and being able to say how they change as you walk from place to place.
That is the shape every fundamental description in this book eventually takes: the electric and magnetic fields, the geometry of spacetime, and the objects whose ripples we call particles. Getting comfortable with the type of thing now costs nothing and saves a great deal later.
One warning belongs at the outset, because the notation works hard to conceal it. The arrow a field attaches to one point and the arrow it attaches to another belong to separate spaces that merely resemble each other, so comparing them, which is what subtracting them to build a derivative requires, is an assumption rather than an operation. On a flat page with square axes the assumption is invisible because it happens to be true. On a curved surface it fails outright, since nothing entitles you to say that a direction at the pole and one at the equator point the same way, and repairing that failure is where gravity eventually comes from.
There is a genuine surprise waiting here, the first place in the book where the shape of a region rather than the behaviour of a field decides a physical fact. Adding a field's push along a path is how you ask how much work it did, and when the field is a gradient the answer depends only on the two ends, which is the accumulation theorem of the second chapter with a curve standing where an interval used to.
Four conditions compete for the name conservative, three of them global and untestable, since nobody can check infinitely many loops, and one local and settled in seconds. The local one does not imply the others. Delete one point from the plane and there is a field with no swirl anywhere it exists whose push around any loop enclosing the gap is a fixed number that never dwindles, because its potential is the angle, and the angle is not a function: go round once and it has grown.
The repair is a condition on the region, not on the field. What makes this physics rather than a curiosity is that the potential of a magnetic field behaves this way outside a solenoid, so an electron beam split around one shifts its interference pattern although the field vanishes everywhere the electrons go. The potential carries something the field strength does not.
Shrink a closed surface down around a point, keep account of how much more leaves it than enters, divide by the volume enclosed, and the number surviving the shrinking is the divergence. Notice what the recipe never mentions: no axes, no components, no choice of description whatsoever. That matters, because the formula it produces looks like a statement about coordinates, and it cannot be one.
What the formula turns out to be is the sum of the diagonal entries of the array of first derivatives, which the linear algebra chapters singled out as the one quantity in an array that no change of description can disturb. Set that beside the other fact established there, that a map barely different from doing nothing multiplies volumes by one plus that sum, and the meaning arrives in a line. Carry a small blob of dye along with the flow and the divergence is the fractional rate at which the blob swells.
Positive means expanding, negative means squeezing, and zero means the flow rearranges without ever compressing. That reading is worth more than the formula, because it converts one of the deepest statements in classical mechanics, that the volume occupied by a spread of possible states never changes as they evolve, into a two-line consequence of the fact that mixed second derivatives are indifferent to their order.
Where the divergence measured swelling, what is left to measure is turning, and turning is detected by walking a small closed loop and totting up how much the field carried you along as you went. Do it in the plane perpendicular to each of three directions and the three answers assemble into a vector, though the assembling is arithmetical luck rather than law: the count of independent antisymmetric quantities matches the count of directions only in three dimensions. In four it does not, which is why the electric and magnetic fields will later have to stop impersonating two vectors and admit to being six entries of one antisymmetric object, and why the cross product goes the same way.
Underneath both operators lies a single decomposition. Any array splits, in exactly one way, into a part unchanged by swapping its two labels and a part that reverses sign, and applied to the derivatives of a flow this says that near any point a flow carries a small blob along, stretches it along three perpendicular axes, and spins it rigidly.
The divergence is the total stretching and the curl is twice the rate of spin. They were never two inventions but two pieces of one derivative, and they do not exhaust it, since the shearing left over is invisible to both: a flow can shred a blob while reporting zero for each.
Three results with three names, three sets of hypotheses and three right-hand rules turn out to be one result stated at three different sizes. The argument establishing each is the one the second chapter used on an interval: chop the region into cells, add up what happens locally inside each of them, and observe that every internal wall is shared by two neighbours who count it with opposite signs, so all of it cancels in pairs and only the outermost skin survives.
What is left is a sentence worth learning in preference to the formulas. Whatever a derivative accumulates throughout a region is bookkept entirely on that region's boundary, and the dimension of the region is free: the two ends of an interval, the two ends of a curve, the rim of a surface, the skin of a solid. Written once, in a language general enough for curved spaces of any dimension, the four statements become one line, and that line is the most reused structural fact in theoretical physics.
A corollary hides in the geometry and will be spent shortly. A boundary has no boundary of its own, since a disc is bounded by a circle and a circle stops nowhere, and a ball by a sphere, which has no edge. Under the correspondence above, that fact about regions must appear as a fact about derivatives, and it does.
Conservation is normally stated globally, as the claim that some grand total, added over everything, never changes. The version obtained here is considerably stronger and considerably stranger, and the whole of the derivation is one physical sentence: the substance is neither made nor destroyed, so the only way the amount inside a region can change is by crossing the boundary. Convert that crossing into an integral over the region, notice that the region was never specified, and a claim about every region collapses into a claim at each point.
The strength of the local statement is easiest to feel by asking what the global one would tolerate. Global conservation is content for a charge to vanish in London at the instant an identical charge appears in Sydney, because the books balance either way. The local statement forbids it outright, since to leave anywhere the substance must cross the surface enclosing it, and to arrive anywhere it must travel through the space in between. Nothing teleports, and that is a far larger assertion about the world than any accountancy of totals.
This is what conservation comes to mean for the rest of the book. Charge, probability, energy and momentum are each conserved in precisely this sense, and the familiar global version is what you get by adding the local one up and assuming nothing interesting is going on infinitely far away.
What remains is a pair of identities, each a single application of the fact that mixed second derivatives ignore their order. The swirl of a gradient vanishes everywhere and always, and so does the outflow of a curl. Neither computation is worth remembering, and both are worth remembering for what they license.
The first is why a field with no sources anywhere can be written as the curl of something else, which is how the magnetic potential comes to exist. The second is why that potential is not unique, since adding the gradient of any function leaves every observable untouched, and that redundancy is the seed from which every force in the Standard Model grows. The two identities are themselves one identity, the algebraic shadow of the remark that a boundary has no boundary.
Composing the two derivatives in the one order that does not give zero produces the operator governing more equations in physics than any other, and it has a plain reading: it reports the amount by which a field at a point falls short of its average over a small sphere around that point. Diffusion is a substance chasing its neighbourhood average, and heat spreading is the identical statement. The calculus of fields is now complete, and what is missing is any way of solving the equations it produces, which is where the next chapter goes.
0.8 Differential Equations and the Oscillator
Every law stated in this book relates a quantity to its own rate of change, and that is not a stylistic habit of physicists but what a law is. To say that nature runs on local rules is to say that the way things stand at one instant fixes the way they stand an instant later, and writing that down produces a differential equation and nothing else. Solving such equations is therefore most of the work.
Two consequences follow from the shape of the equations rather than from any philosophy. The first is that a law does not by itself determine a history. Law and starting conditions together do, and the number of quantities they must supply is exactly the highest rate of change appearing, which is why mechanics wants a position and a velocity rather than one or three.
The second is determinism, which is not an attitude adopted by physicists but a theorem about equations of this form, holding whenever the right-hand side does not vary too violently as the state is nudged. The hypothesis has teeth: where it fails, uniqueness fails with it, and there are innocent-looking equations whose solutions are not fixed by their starting conditions at all. When a model of yours behaves that way, the usual diagnosis is a missing piece of physics rather than an undecided universe.
Multiplying an equation through by a carefully chosen function, so that its left-hand side collapses into a single derivative, has the air of a manoeuvre somebody once had the wit to think of. It is nothing of the sort. Demand that the collapse happen, write out what the demand requires, and the multiplier is pinned down by an equation already solved a page earlier. Nothing was guessed, and this is another object arriving because a requirement cornered it into existence rather than because anyone liked it.
What comes out deserves to be read rather than filed. The state of a system at any moment is the fading memory of where it started plus the fading memory of everything done to it since, each past disturbance still contributing whatever remains of it after the time elapsed. So the entire response is built from one ingredient, namely what the system does after a single sharp isolated kick, and everything else is that ingredient shifted, scaled and added up.
The ingredient has a long future ahead of it. It reappears as the response of any linear system to any disturbance, it carries a condition saying that effects do not precede their causes, and in the last part of the book, with a point of spacetime in place of an instant, it is the object out of which Feynman diagrams are assembled.
An exponential is the one function that differentiation barely touches, since the operation hands the same function back multiplied by a number, and that single fact is why every method for solving linear equations with constant coefficients works. Put the question the linear algebra chapters put to every map, namely which vectors it merely rescales, to differentiation itself, treating functions as the vectors they were shown to be. The answer is the exponentials, one for each multiplier, and nothing besides.
The old instruction to try an exponential, usually presented as a historical accident, is therefore an instruction to move to the description in which differentiation has stopped being differentiation and become multiplication by a number. This is the move already made twice in other clothes, the one that turned a symmetric array into a list of numbers along perpendicular directions, and it works for the same reason: the problem has fallen apart into independent pieces.
Two dividends arrive at once. The solutions form a space whose dimension is the order of the equation, so two independent solutions of a second-order equation are a complete answer rather than a lucky haul, and school algebra's arbitrary constants are coordinates in a basis. A multiplier with an imaginary part means oscillation while its real part means growth or decay, and between them they classify the behaviour of every linear system in the book.
Expand any potential energy about a stable resting point and watch what is left. The constant is irrelevant, the linear term is absent precisely because the point was a resting point, and what survives is quadratic. So the equation being solved here is not a statement about springs at all. Every stable system in nature, examined closely enough, obeys it, which is why a molecule, an atom in a crystal, a pendulum, a mode of the electromagnetic field and a disturbed black hole are all describable by the same few lines, and why solving one equation completely is worth more than surveying a dozen partially.
Three routes to the answer are worth carrying instead of one, because they generalise in different directions. Trying an exponential is the eigenvalue reading. Packing amplitude and phase into a single complex number turns differentiation into multiplication and turns the addition of two oscillations into the addition of two numbers. Multiplying through by the velocity produces a quantity that never changes, which lowers the order of the problem without solving it.
That last route reaches furthest. Plot position against momentum rather than against time and every solution becomes a closed loop, closed precisely because the energy is conserved. The area it encloses is the energy multiplied by the period, and quantum mechanics will leave the ellipse standing and merely ration which of them are permitted.
Everything a resisted oscillator does is settled by one comparison, between the rate at which it would swing unopposed and the rate at which the opposition drains it, and the three outcomes are the three things a real number can be with respect to zero. Under light resistance the multipliers are complex, so the motion rings and fades, the imaginary part setting the pitch and the real part how long the ringing lasts. Under heavy resistance they are real and the motion crawls back to rest without crossing over.
There is a surprise in the heavy case. More resistance makes the return slower rather than faster, so a door closer filled with treacle takes forever, and the quickest return available sits on the boundary between ringing and crawling. That boundary is where two multipliers have collided and the supply of independent solutions has run short, the defective situation the linear algebra chapter warned about, met here in a shock absorber instead of an array.
One dimensionless number does all the summarising. It counts the radians of swinging over which the stored energy falls by a factor of e, so a system rated at a hundred rings for about sixteen cycles before losing most of what it had. It is shortly going to measure something apparently unrelated, namely how sharp the response is when the system is pushed from outside.
Push on an oscillator at a frequency of your choosing rather than its own, wait for its memory of how it started to fade, and what remains is a steady response whose size depends on how near you came. The shape of that response is universal. Nothing about springs enters it, only that the system has one preferred frequency and that its motion decays at some rate, so any system with those two properties answers with the same curve.
The width of the curve is the quantity to watch, because it is exactly twice the decay rate, not approximately. The number that counted radians of ringing therefore also measures the sharpness of the peak, and the two jobs are one job. Read the identity backwards and it becomes an instrument: whatever lasts a short time is obliged to produce a broad bump, so the width of a bump reports a lifetime. Nobody has ever timed an unstable particle with a clock, and nobody could. Their lifetimes are read off the widths of peaks in scattering data.
One caution, since the popular account has it backwards. With no losses whatever there is no steady response at all, because the amplitude grows without limit and never settles. Resonance is a negotiation between the driving and the dissipation, and the sharper it is the more completely the dissipation is in charge.
Two masses joined by a spring look like a harder problem than one mass on a spring, and the appearance is false in a way that matters far beyond two. The forces come from a potential energy, second derivatives are indifferent to their order, so the array governing the motion is symmetric, which is exactly what the spectral theorem asks for.
There are therefore directions along which the masses move as one, each combination obeying the single-oscillator equation alone and ignoring the other entirely. The complicated motion in which energy sloshes between the masses is two simple motions added together and watched in an inconvenient description. This collects a promise made when vector spaces were introduced: one theorem, proved once, solving a room full of coupled pendulums without being told what it is looking at.
Nothing in the argument cared how many masses there were. Take a chain, let their number grow while their spacing shrinks, and the list of displacements becomes a function of position, the array becomes a differential operator, and what survives is the wave equation. A field is what infinitely many coupled oscillators turn into. That sentence is the spine of the last three parts of this book, because it means quantising a field amounts to quantising infinitely many springs, and the whole number counting one mode's rungs is what we call a particle.
Losing linearity costs more than heavier algebra, and being exact about what goes is worth a moment. Every habit acquired so far, whether breaking a problem into modes, or expanding a disturbance into pure tones and treating them separately, rests on one line of algebra true only because the equation was linear. A pendulum swung wide is already outside it, since the period then depends on amplitude, so two solutions of different sizes keep different time and their sum solves nothing.
The consequences are not academic. Gravity is nonlinear because the gravitational field carries energy and energy gravitates, so the field of two stars is not one star's field added to the other's, which is why a problem that is an exercise in Newton's theory took half a century in Einstein's. The strong force is nonlinear for the same reason, which is why beams of light pass through one another and beams of gluons do not.
What replaces superposition is the strategy named at the start of the toolkit: find a linear problem nearby, solve it exactly, and treat the difference as a correction computed term by term. The oscillator is not merely the first system anybody can solve; it is the system every hard problem gets expanded around, which is why the next thing worth having is a way of writing any disturbance whatever as a sum of oscillations.
0.9 Fourier, Delta Functions, Probability
A space whose vectors are functions was promised when vector spaces were first defined, and it arrives here fully equipped. Give two functions an overlap by multiplying one against the other and integrating, and every result about lengths, angles and perpendicularity becomes available at once, no new theorem required. What is missing is a basis, and the one chosen is the collection of pure waves that fit a whole number of times into the interval.
That the waves are mutually perpendicular takes one short integration, and it is the entire technical content of the subject. Once it is done the earlier result about perpendicular bases takes over: each coefficient is a single overlap, computed on its own without reference to the others, so breaking a function into waves costs only a sequence of separate integrals.
Two cautions belong alongside it. Perpendicularity does not establish that the waves are numerous enough to reach every function, and that they are is quoted here rather than proved. And the sense in which a series of smooth waves reproduces a function with a jump is weaker than one would like, because near the jump every partial sum overshoots, by the same nine per cent of the step, however many terms are taken. A sharp edge and a limited supply of frequencies are incompatible, which stops being an annoyance and becomes a theorem shortly.
Let the interval holding the function grow without limit and watch which parts of the description survive. The permitted wavelengths crowd together as the box lengthens, the discrete list of coefficients becomes a function of a continuous label, and the sum over modes becomes an integral, provided the bookkeeping is arranged so each term carries the spacing between neighbouring modes. It is the same expansion along perpendicular directions as before, the directions now labelled by a real number rather than a whole one, and that innocuous-looking change is the one thing here that later takes real work to make legitimate.
What comes through the limit untouched is the statement that total size is preserved. The squared size of a function, added over positions, equals the squared size of its transform, added over wavelengths, which says that transforming is a rotation of the space of functions rather than a computation performed upon them. It preserves lengths, it preserves angles, and it carries perpendicular bases to perpendicular bases.
Hold that in the form stated, because rotations are precisely the maps that keep a total probability equal to one. When a physicist remarks that the transform conserves probability, the content of the remark is that it is a rotation, and that nothing about the object has been altered by transforming it. Only the axes it is being described against have moved.
Any perpendicular basis would have delivered everything so far, so the case for this one has yet to be made, and it rests on a property no other basis possesses. Differentiating a pure wave returns the same wave multiplied by a number. The chosen basis is the one in which differentiation is diagonal, which is why it was chosen and the reason the subject exists.
In that description differentiation stops being differentiation. It becomes multiplication by a number differing from one wave to the next, so a differential equation, which ties the value of a function to the values at its neighbours, becomes an algebraic equation with one instance per wave, answered by dividing. The difficulty has moved into transforming in and back out, and none of it is left in the middle.
By now the manoeuvre should be recognisable, because this is the fourth performance under the fourth different name. A symmetric array became a list of numbers along perpendicular directions. Two coupled masses became two oscillators that ignore each other. An exponential turned out to be what differentiation merely rescales. And a function on the line has become a list of amplitudes, one per wavelength, none interfering with any other. Find the description in which the problem falls apart into independent pieces: one idea, and it has now paid for itself four times over.
Hit a system once, sharply, and record everything it does afterwards. Provided it responds in proportion to what is done to it and behaves the same way whatever the hour, that record is the whole system, because any input can be regarded as a dense succession of sharp hits and the output is whatever survives of each, added together. The operation assembling the answer, in which one function is slid across another and the overlap recorded at every displacement, is expensive: a full integral at every point.
Written in terms of waves it is a product. Transform both functions, multiply them together, transform the result back, and an integral at every point has become one arithmetical operation at each wavelength. That is why the transform earns its place in daily practice, quite separately from why it exists in principle.
The consequence is larger than it looks. In the last part of the book the response to a disturbance created at one point of spacetime and detected at another is this same object with more indices, and the amplitude for a particle to travel, interact somewhere unspecified, and travel onward is a sliding overlap of such responses. Written in terms of wavelengths those overlaps turn into multiplications, which is exactly why the rules for Feynman diagrams tell you to write a factor for each line and multiply.
The single sharp kick relied on a moment ago was never defined, and defining it honestly means giving up the idea that it is a function. Nothing assigns a value to each point of the line and behaves as required, because a function altered at one point leaves every integral where it was, and infinity is not a value anything takes. The way out is to stop asking what it is and state what it does, which is to eat a function and hand back its value at one point.
That is not a partial answer but the whole definition, and objects specified this way are perfectly legitimate. Two habits go with them. Such an object appears only under an integral sign, paired with something well behaved, and the limit producing it, in which a bump narrows while its area is held at one, must be taken after the integration and not before. Take it before and you have a quantity that is zero everywhere and infinite at a point, carrying no information.
Worth extracting too is the device that makes its most quoted representation respectable. Insert a factor rendering a divergent integral convergent, compute the honest answer, and remove the factor at the end. That is neither a dodge nor a trick peculiar to this object. It is the procedure later called renormalisation, met here in its harmless form.
Whatever else a wave may be, it cannot be both brief and pure, and the exact statement of that impossibility is the strongest thing in the toolkit. Measure the spread of a function across positions in the ordinary statistical way, measure the spread of its transform across wavelengths likewise, and the product can never fall below one half. Squeezing either widens the other by the compensating factor, and the bell curve is the unique shape attaining the minimum, a conclusion the argument produces rather than assumes.
Look at what the proof needed, because the list is short and every item was in hand: the definition of the transform, the fact that transforming preserves size, the fact that differentiating has become multiplying, the overlap inequality proved when inner products were introduced and applied here to two particular functions, and one integration by parts. There is nothing else in it, and no physics anywhere: no particles, no measurement, no observer and no Planck's constant.
That inequality is the uncertainty principle, and all quantum mechanics adds to it later is one substitution: that a particle's momentum is the wavelength of a wave. It says nothing about disturbing a thing by looking at it, because no measurement occurs in the derivation and nothing is done to anything. A brief flash of light contains a wide spread of colours whether or not anybody watches.
Probability arrives at the close not as a change of subject but as the same subject in other words. Adding two independent random quantities slides one density across the other and records the overlap, which is exactly the operation of a moment ago, so in the wave description the densities multiply. Repeat the multiplication many times, rescaling as you go, and every feature of the original distribution is ground away except the quadratic one, which survives because the rescaling was chosen to let it. The bell curve is what remains once everything else has been averaged out of existence, and that is why it is everywhere.
The toolkit is finished. What you have now is a way of describing change and a way of accumulating it, the habit of expanding rather than solving, a language of spaces and the maps between them, the theorem that the right description makes a hard problem fall apart into independent pieces, a calculus for quantities defined at every point, and the one basis in which differentiation is arithmetic.
None of it is physics. All of it was assembled for a single job, and the next part begins that job by discarding the thing most people take to be the foundation of the whole subject. It throws away forces, and puts in their place one number attached to each history the world might have followed.
The trade about to be made is not obviously a good one, and the least flattering version of it should be put first.
Newtonian mechanics has a hole in the middle of it. The second law predicts nothing at all until somebody supplies a force, and every force is imported from outside, fitted to data and justified by working. What is about to replace it has a hole in precisely the same place. One function is still supplied by hand, and no argument in this part derives it. The only change — and the whole of Part I is the case that this is the change that matters — is that the function becomes a single number attached to each history the world might follow, in place of a bundle of arrows attached to each body at each instant.
Three things are bought with that, and the third is the one worth the fare. A number does not care which coordinates it is written in, so the equations keep their shape when the description changes, which arrows conspicuously fail to do. A number can be written down for a thing that has no particles to hang arrows on, which is what a field is. And a number can be asked to stay unchanged under some transformation, whereupon it hands back a conserved quantity without being asked twice. That last is the largest result in classical physics, and it cannot even be stated in the language being given up.
Part I · The Action Principle
The most-skipped prerequisite in physics — and the reason people wall out at GR.
1.1 What's Wrong With Forces
The second law does not tell you what the force is, and that omission, almost never said out loud, is where the promised discarding of forces begins. The law carries an undetermined symbol, so that given a force law it predicts a motion and given none it predicts nothing at all, and no observation can contradict it, because any motion whatever can be absorbed by declaring the force to be mass times the measured acceleration.
So Newtonian mechanics is a framework awaiting a force law rather than a theory. The physics lives entirely in the forces, and each is carried in from outside and justified by fitting data, one import for gravity, another for drag. The first law is not the trivial special case it is usually reduced to either: it asserts that frames exist in which the second law takes its simple form, which is worth asserting because in most frames the law is false and survives only by inventing forces nothing exerts.
What eventually replaces all this keeps the same shape, one function supplied by hand, and the difference is that the function will be a single number rather than a bundle of arrows. That difference is what later allows a demand for symmetry to decide what the function is permitted to be.
A hypothesis sits buried in the conservation of energy, never mentioned at school, and it turns what looks like a law of nature into a theorem with a condition attached. The half that assumes nothing says the work done along the actual path equals the change in kinetic energy, and that holds for friction, for magnetism, for a force invented on the spot. It is not yet a conservation law, because the work accumulates along a route and nothing obliges it to collapse into a difference.
It collapses exactly when the force is the gradient of something, which is where the field theorems of the toolkit get spent. A potential energy exists because that condition holds, the total is unchanging because it holds, and it can be checked at a point rather than over infinitely many loops because it holds. Drop the condition and the total genuinely is not conserved, as a projectile fighting the air shows to eleven digits.
The missing energy is in the air, spread over the enormous number of coordinates you declined to track, all of which obey forces that do satisfy the condition. Friction is a bookkeeping decision rather than a hole in the accounts. A better answer arrives three chapters ahead, where it concerns not forces but the laws being the same today as yesterday, with the unsettling corollary that in an expanding universe they are not.
Neither conservation law here is a theorem. Both rest on an assumption about forces, and one of them is about to be broken by a force nobody would call exotic. Total momentum survives because each internal force cancels against its partner, which is the third law. Total angular momentum survives only if those paired forces also point along the line joining the two bodies, a strictly stronger demand that the weaker one does not imply.
Now take two charges, one moving straight at the other and the second moving crosswise. The first produces no magnetic field directly ahead of itself, so the second feels nothing, while the second produces a good field where the first sits, so the first is pushed. One particle is shoved and the other is not, in a system with nothing else in the universe, and the pair's momentum changes at a computable rate.
Either momentum conservation is false, or the momentum is somewhere that is not a particle, and the only other object in the problem is the field. Something that stores momentum, that can be pushed and pushes back, and that holds the missing amount while one charge waits to learn the other has moved, is not a bookkeeping convenience. It is an object, and this is the first place in the book where a field is cornered into being real.
Of the four places the force picture gives way, the one deciding the shape of all that follows is the least dramatic. Write the second law in polar coordinates and two terms appear that were not there, the centrifugal and the Coriolis. No physics was added and the force is the same vector it always was; the extra terms came from the basis directions turning as you move — that is, from the labelling.
How completely they are artefacts is clearest for a particle with no force on it, travelling in a straight line. Its radial coordinate has a second derivative that is not zero, and the honest law needs the extra term to cancel it. So the shape of the law, meaning which terms appear, is correct in one coordinate system and wrong in every other, repaired by hand from a table.
That is liveable in a plane, where the useful coordinate systems are few and tabulated. It cannot be taken where no straight grid exists, and a curved spacetime provably admits none. A law simple only in coordinates that do not exist is a special case rather than a law. The other three failures point the same way: a constraint force carried through a page of algebra and thrown away, a field with a continuum of degrees of freedom, a quantum theory with no force operator.
Fermat's rule for light is older than Newton's laws and works in a manner with nothing in common with them. It says nothing about what happens to a ray at any place and describes no push. It attaches one number, the total travel time, to each entire route the light might take, and then picks the route out by a condition on that number.
Impose the condition on a slab of air above a slab of glass and the empirical law of refraction falls out, angles and interface and all, from a statement mentioning neither. That is the whole template — a global claim about complete paths reproducing a local law at each point. The two features making it work are exactly the two failures of the force picture: what is attached to a path is a number, so relabelling the plane cannot disturb it, and the only input was one value at each point of space, so nothing had to be resolved into components.
One word in the usual statement is wrong and worth correcting now, because the correction is needed repeatedly. The condition imposed is that the number stops changing to first order, not that it is smallest. For two flat slabs the two coincide; put source and detector before a curved enough mirror and the actual ray takes longer than its neighbours, with the physics unaltered.
1.2 Stationary Action
What is being varied from here on is not a number but an entire history, and the change of type has one consequence worth carrying in advance. A machine that eats a whole function and returns a single number has a domain with infinitely many dimensions, one for each instant, loosely speaking, rather than the three or six that a mechanical problem used to offer.
That is why hunting for a stationary point cannot be done the way it was done on a surface. Finding the flat spot of a function of several variables meant setting a finite list of partial derivatives to zero, and here there is no finite list to set. What comes out instead is one condition for every instant of time, and one condition per instant is a differential equation.
So the answer to a problem of this kind is always a differential equation, for a structural reason rather than as a happy fact about the particular integrals involved. Three such machines are already familiar: the length of a curve, the time light takes along a route, and the quantity that is shortly to be given the whole weight of the subject to carry.
Dividing by a function is not an operation, so the picture of a derivative as one change divided by another has nowhere to stand here. The definition that does stand is the one insisted on at the very start of the toolkit, at some cost in apparent fussiness: nudge the input, and identify the part of the response that is linear in the nudge. Nothing in that instruction cares whether the nudge is a small number or a whole function.
So the pedantry of the opening pages is repaid a second time, and for far more than it was repaid the first. Several variables cost only the replacement of one direction by many; here the displacement is an entire function, and the coefficient is no longer a number but a device that eats such functions and returns numbers. The error still dies faster than the nudge, and everything else is unchanged word for word.
The demand being made is stronger than it looks. Any single choice of nudge probes one direction in the space of paths, and something can easily be flat along one direction while falling away steeply in others. Requiring flatness against every admissible nudge at once is requiring it in infinitely many directions simultaneously, and that is what makes the condition strong enough to pick out one path.
The calculation producing the equation of motion is short, and nearly all of its load is carried by one manoeuvre borrowed from the chapter on integration. Once the nudge is inside the integrand, one term carries it as a plain factor and the other carries its rate of change, and the two cannot be argued about separately because the second is determined by the first. Moving the derivative off one factor and onto the other repairs that, at the cost of a sign and a term evaluated only at the two ends.
This is exactly the use advertised for integration by parts several chapters back, the step converting a statement about one whole total into an equation the system must satisfy at every point, which is where every field equation in this book comes from. The end term dies because the endpoints were held fixed and not because of any property of the integrand, worth remembering for the occasions when the ends are free and the surviving term turns out to be where the momentum lives.
One step remains and it is not free. An integral can vanish through cancellation without its integrand vanishing anywhere, and what defeats the cancellation is that the nudge is yours to choose, so you may aim a small bump wherever the integrand is not zero and leave it nothing to cancel against.
Newton's second law is not a rival to the new machinery, nor a separate law that happens to agree with it. It is the same equation, written in one coordinate system with one choice of the scalar, and that choice is the kinetic energy less the potential energy. Turn the crank and the two objects appearing are the momentum and the force, in that order and in those roles.
That reading licenses the definition carried forward from here. Whatever the scalar yields when differentiated with respect to a velocity is thereafter the momentum conjugate to that coordinate, and the equation always says its rate of change equals the derivative of the scalar with respect to the coordinate itself. The name is not sentimental. It is chosen so that the equation keeps Newton's shape whatever the coordinates happen to be, and the payoff shows immediately: where a coordinate fails to appear in the scalar at all, its conjugate momentum cannot change.
One habit here deserves a warning. Position and velocity are differentiated as though they were unrelated letters, which along an actual trajectory they are not. The resolution is that the scalar is defined on a space where they genuinely are unrelated, and their relation is imposed only afterwards, when a path is substituted in.
Where the combination of kinetic minus potential comes from is the question this chapter refuses to dodge, and the honest answer has a hole. Within classical physics no argument begins at Newton's laws and arrives at that combination by necessity. There is the observation that it works, and the knowledge that it is not the only thing that would: adding anything itself a rate of change, or multiplying the whole by a constant, leaves every prediction untouched. The minus sign is not intuitive and nobody should pretend otherwise.
A second correction concerns the word least, wrong since the principle was named. The condition imposed is that the first-order change vanishes, and that signals a minimum, a maximum or a saddle. In a space with infinitely many directions saddles are not the exception, and an oscillator travelling for longer than half a period supplies a stationary path with neighbours of strictly smaller action.
The real explanation is not available yet, and naming it stops anybody filling the gap with a story about nature being economical. In the last part of this book a quantum particle contributes from every route it might take, weighted by a phase built from that same number. Where the number changes at first order the neighbouring contributions point in all directions and cancel; where it does not, a whole neighbourhood shares one phase and survives.
Constraints were the first of the four failures, and they stop being one the moment you may choose your own variables. A bob on a string has one way to move, so describe it by one number, the angle, and every value of that number already puts the bob at the right distance from the pivot. The constraint has not been solved or eliminated but converted into a coordinate, and nothing is left to enforce.
The tension, which cost a page of algebra in the force picture and was then discarded, never appears at all. The reason fits in a sentence: a force perpendicular to the motion does no work, the scalar is built out of energies, and a force doing no work has no energy to contribute. Adding a second link adds one coordinate and one equation, where the force method adds two more unknowns — so only one of them grows with the thing you hoped to be rid of.
When the constraint force is the answer you want, because you are designing the rod and want to know whether it snaps, it returns by the device the toolkit built for optimising under restriction. The multiplier was described there as a price, the sensitivity of the best available value to relaxing the restriction. Here the price is the tension in newtons, and the identification is exact rather than analogy.
A number attached to a path has no components, and having no components is the whole of why it survives. Relabelling which numbers you use to name the points of a path cannot change the number the path was assigned, so the statement that this number stops changing is a fact about the path rather than about the naming. The equations following from it therefore keep their shape under any smooth invertible change of coordinates, including ones nobody has tabulated.
Set that beside the polar-coordinate embarrassment of the previous chapter. The centrifugal and Coriolis terms, which took half a page of differentiating rotating basis directions to produce, now drop out of differentiating a scalar, with no unit vectors touched and no fictitious force to be remembered. The same calculation hands over a conservation law for nothing, since the angle is absent from the scalar and an absent coordinate has a conjugate momentum that cannot move.
This is a property deciding the remainder of the book rather than a convenience. A curved spacetime provably admits no global straight grid, so a formulation simple only in such a grid is unusable there, while a formulation whose entire input is one number per path does not notice. Everything from here to the last chapter consists of choosing a scalar and making it stationary, and this is the reason it can.
Replace the single independent variable by the four of spacetime, let the quantity being varied be defined at every point rather than at every instant, and the four moves run again word for word. The nudge must now vanish on the boundary of a region instead of at two endpoints, the by-parts step becomes the divergence theorem, and the term it leaves behind dies for the same reason as before. Out comes one equation, and it is every field equation in physics.
The chapter closes with eight lines of scalar, one per theory: a free particle, a relativistic one, the electromagnetic field, gravity, a scalar field, the electron, the strong and weak interactions, and a string. Between them they hold nearly everything the rest of this book is about, each is fed to the same machine, and what comes back is the equations of motion of that subject. Not analogues of them, but them.
What has been bought is worth stating flatly. The input to physics is no longer a list of forces to be discovered one at a time but a single number attached to each history, and the output follows from one condition imposed on it. What has not been bought is any reason to prefer one such number to another, which is the whole of what remains, and two chapters from now a symmetry supplies it.
1.3 Hamilton and Phase Space
A curve can be given by listing its points, or by listing all the straight lines that graze it without crossing, and the two lists carry the same information, because a curve bending only one way is the highest of its own tangent lines. Swapping between the two descriptions is a standard operation, and performing it twice returns you exactly to where you began.
Which description you want is decided by what you are able to set. A laboratory has thermometers and no entropy meters, so a chemist labels a state by its temperature, which is the slope of the internal energy against entropy rather than the entropy itself, and the free energies are that one swap performed on one slot or another. What makes it reversible is that the curve bends one way only. Where it does not, the swap flattens the offending stretch permanently, and that flat stretch is a first-order phase transition rather than a defect in the mathematics.
In mechanics the variable traded away is the velocity and the slope accepted in its place is the momentum, and the reversibility condition becomes the statement that the kinetic energy is an honest positive quantity in the velocities. What the trade produces has been met before, as the combination a problem with no explicit clock in it happens to conserve.
The quantity paired with a coordinate here is not always mass times velocity, and the two ways it can differ both matter. For a pendulum described by an angle it is an angular momentum, with the wrong units for a momentum entirely, because what the pairing requires is that coordinate and partner multiply to something with the units of action. For a charge in a magnetic field it is not the mechanical momentum even in Cartesian coordinates, the field's potential having been added to it.
That second case is not notation. The mechanical momentum is what an instrument reads, while the conjugate one shifts when the potential is rewritten in a way that changes no physics, so the conjugate one is not by itself observable. Yet the conjugate one is what quantum mechanics turns into an operator, which is why every magnetic effect in that subject, including the interference shift around a solenoid whose field the electrons never enter, follows from a single substitution.
The other object built here is usually introduced as the total energy. It is the total energy under two hypotheses, both of which fail in ordinary situations, and a bead on a motor-driven hoop is enough to break them. The conserved quantity and the energy come apart there, and it is the first that the formulation runs on.
Trading one second-order equation for two first-order ones is available to anybody and costs nothing, since calling the velocity a new variable does it. All of the content here is in which new variable was chosen. Taking the conjugate partner rather than the velocity produces two equations that are the same equation with the roles of coordinate and momentum exchanged, apart from a single minus sign sitting in one of them.
That sign cannot be removed. Reversing the sign of the momentum only moves it to the other equation, and it is what makes the two halves of the calculation cancel when you ask whether the energy changes, so that the energy does not change. Written compactly the sign becomes an antisymmetric object standing between the two halves of the space, whose effect is to turn a gradient through a quarter turn, so the motion runs along the level surfaces of the energy rather than across them.
Everything structural in the rest of the chapter is that antisymmetry and nothing else. A symmetric object of the same general kind measures lengths and will run the chapters on gravity; this one measures oriented areas instead, which is a different geometry and the reason the space must have an even number of dimensions, an antisymmetric table in an odd number of them being necessarily degenerate.
A single point of this space fixes the entire past and future of the system, which is the whole reason for building it. A configuration by itself is not a state, since a pendulum passing through the bottom of its swing might be going either way at any speed. A configuration together with its conjugate partner is a state — this is the space of states the existence theorem always asked for.
Two things follow. Exactly one solution passes through each point, so trajectories never cross and the space fills with curves like the streamlines of a steady fluid, the only apparent meetings being where nothing moves at all. And the energy is unchanging along each curve, so a pendulum's whole qualitative behaviour reduces to three kinds of curve: closed loops for swinging, open ones for going over the top, and the single curve dividing them, along which the bob takes forever to arrive upside down.
The number worth attaching to a closed loop is the area it encloses, whose rate of change with energy is the period. A promise made when the oscillator was first drawn is collected here: quantum mechanics does not delete these loops, it rations them, permitting only those enclosing a whole number of units of Planck's constant, with half a unit for the lowest. That is why the constant has the units it does.
The centrepiece of the chapter costs three lines and is hard to credit at first hearing. Take any region of this space, regard it as a collection of possible starting conditions, let every point evolve for as long as you like, and the volume occupied afterwards is exactly what it was. Not nearly, not on average, and not for well-behaved systems only.
The shape is unconstrained while the measure is exact, and that combination is what makes the theorem worth having. A blob released near a pendulum's dividing curve is drawn into a filament wound many times round, thinner than any resolution you can afford, and its area does not move. Damping can therefore never be fundamental, since a system spiralling into rest destroys volume; friction is a larger volume-preserving system with some coordinates thrown away. And information is stirred rather than destroyed, because running the equations backwards recovers the original blob exactly.
That last sentence travels furthest. Its quantum successor says evolution preserves overlaps, so two states that begin distinguishable stay distinguishable forever, and an evaporating black hole is a crisis rather than a curiosity because the usual calculation appears to contradict it. The arrow of time, meanwhile, is not in the dynamics, which runs equally happily either way, but in the fact that a filament with an enormous neighbourhood looks like a filled region at finite resolution.
Out of the equations of motion falls a way of combining two quantities on this space to produce a third, and its importance lies almost entirely in what it later becomes. The immediate use is already substantial. The rate of change of any measurable quantity whatever is its combination with the energy, so asking whether something is conserved stops being a matter of solving the motion and becomes a matter of computing one expression and seeing whether it is zero.
Among the coordinates themselves the answers are as simple as they could be. Two coordinates give nothing, two momenta give nothing, and a coordinate paired with its own conjugate momentum gives exactly one. Those three lines and the four properties the operation possesses are the entire structure, and the fourth property is what makes the whole an algebra of the kind the last parts of this book are built from.
Here is what the chapter was for. Replace this operation by the difference between doing two things in one order and in the other, divided by the imaginary unit and Planck's constant, alter nothing else, and classical mechanics becomes quantum mechanics. The relation between a coordinate and its momentum becomes the founding relation of the subject, and the conservation criterion becomes the familiar remark that a quantity is conserved when it commutes with the energy.
Any function on this space, not only the energy, can be put into the equations of motion in the energy's place, and doing so makes it push everything around in a definite way. The energy pushes the system forward in time, which is what the equations of motion said all along, so it is not merely the energy: it generates a translation of the clock.
Do the same with the momentum along a direction and everything slides a little that way. Do it with the angular momentum about an axis and everything turns about that axis, positions and momenta alike, as vectors should. Nothing was assumed about the system anywhere in this, and it collects a promise made when exponentials of self-partnered maps first appeared: a generator and the family it builds are two aspects of one object.
Read the combination of a quantity with the energy in both directions. Left to right it is the rate at which that quantity changes, so its vanishing means the quantity is conserved. Right to left it is the change in the energy under the flow that quantity generates, so its vanishing means the energy is untouched by it. They are one expression up to a sign — so a conserved quantity is not accompanied by a symmetry and not emitted by one, but is the symmetry, seen as something that moves things.
Not every relabelling is permitted here, and the restriction is real. The equations survived any smooth change of coordinates because the momenta were free to follow; once the momenta are independent variables, only changes preserving the pairing between a coordinate and its partner leave the equations alone. Those changes preserve areas and volumes, and evolving the system forward is itself one of them.
Push the freedom to its limit and ask for a relabelling making the new energy identically zero. Every new variable is then constant and the system solved, the price being one partial differential equation for the function generating the change. That function does not merely resemble the number attached to each history. It is that number, regarded as depending on where the path ends, which closes a loop opened two chapters ago.
One remark makes this part look different in retrospect. Feed a wave whose phase is that number divided by Planck's constant into the equation of quantum mechanics, and the leading term is that same partial differential equation, the discarded terms smaller by one factor of the constant. Nearly everything is an approximation, and this is one: an expansion in a named small quantity, with a known first discarded term. Classical mechanics is the geometry of surfaces of constant phase, as ray optics is, and Hamilton had that structure ninety-two years before anyone wrote a wave equation for matter.
1.4 Noether's Theorem
Getting the definition right is most of the work, and the theorem itself is four lines long because of it. A symmetry in the sense required is a family of deformations of the path controlled by one dial, with the do-nothing deformation at zero on the dial, so the transformation can be turned on gradually and differentiated with respect to how much of it is on. Only that first derivative ever matters, and it is called the generator.
The second half of the definition is a loosening, and refusing it would produce nonsense. The natural demand is that the scalar be unchanged, and under a change to a frame moving at a small steady speed it is not, which would force the conclusion that Newtonian mechanics can detect uniform motion inside a sealed cabin. What the scalar changes by is the rate of change of something, and anything of that shape contributes the same fixed amount to every competing path, so it is invisible to the comparison.
One quantifier in the definition does real work later and is the commonest place to get the theorem wrong. The requirement is that the scalar behave this way along every path you could draw, including the wild ones that solve nothing. A transformation leaving it alone only along the paths the system actually takes describes solutions you already have, and yields nothing.
No result in this book matters more than this one, and the whole of its proof is four lines. Expand the change in the scalar under the deformation by the chain rule, which is true along any path whatever, and then restrict attention to a path that solves the equation of motion. What the equation of motion does, and this is the entire trick, is turn that expression into the rate of change of one particular quantity.
The hypothesis has already said that the change is the rate of change of something else. Two expressions for one thing, so their difference has a rate of change of zero, and a quantity whose rate of change is zero is conserved. The two halves of that argument sit on opposite sides of a distinction worth keeping as a slogan: the hypothesis is required along every path, the conclusion holds only along paths that solve the equations, and swapping the two makes the theorem either vacuous or false.
One further case is usually waved at and repays being done properly, namely a transformation moving the clock as well as the coordinates. The bookkeeping then produces an extra term attached to the size of the time shift, and the quantity sitting in that term is the energy. It was not put in by hand. It arrived because the clock moved.
Energy, momentum and angular momentum arrive here as one calculation performed three times with different entries. Shift every clock by the same amount and, provided the scalar has no explicit dependence on the time, the conserved quantity is the energy. Shift every particle by the same displacement and, provided the potential depends only on separations, it is the total momentum. Turn everything through one angle about an axis and, provided the potential depends only on the lengths of those separations, it is the angular momentum about it.
Read as physics rather than bookkeeping, the three hypotheses say the laws are the same at every moment, the same at every place, and the same in every direction. Three separate empirical laws have become one statement about the uniformity of the arena. None is a fact about matter, and no particle, force or material can be invented that violates them, because so long as it lives in a uniform isotropic space and an unchanging time its scalar carries those symmetries and the quantities are conserved whatever else it does.
Notice what has replaced the third law. Momentum conservation followed there from an assumption that internal forces cancel in pairs; it follows here from the observation that a potential between two bodies can depend only on how far apart they are — the same physical content, no longer an axiom about forces.
A theorem earns trust by being falsifiable, and this one is easy to put on trial. Differentiate the energy along a solution while assuming nothing, and out comes the statement that its rate of change is exactly minus the rate at which the scalar depends on the time. The left side is how fast the conservation law fails, the right side how fast the symmetry fails, and they are the same number.
An oscillator driven from outside has a scalar depending on the time, and its energy is not conserved, because something outside is doing work on it. Explicit time dependence is the mathematical signature of there being an outside, and enlarging the system to include whatever drives it restores the symmetry and the law together. Every ordinary violation of energy conservation is an incomplete system.
One case is not of that kind, and there the honest answer beats the tidy one. On the largest scales every distance grows by a common factor depending on the time, so the geometry itself is time-dependent and there is nowhere larger to escape into. Light from a distant galaxy arrives stretched, each photon carrying less energy than it set out with, and that energy has not been transferred or stored anywhere, because no conserved global energy exists for it to belong to. Asking where it went assumes the law the missing symmetry removed.
Some of nature's symmetries cannot be applied a little at a time, and the theorem does not reach them. Reflecting every spatial direction is the clean case: that operation and the do-nothing operation differ by the sign of a determinant, the sign takes only two values, and no continuous path of transformations gets from one to the other without leaving the set of allowed transformations. There is no dial to turn, so there is no derivative at zero, no generator, and no conserved current.
What such a symmetry offers instead exists only once states can carry labels, and it is a multiplicative label, usually plus or minus one, whose product over everything present is the same before a reaction as after. Additive quantities come from symmetries you can apply by degrees and multiplicative ones from symmetries you cannot, and they are different objects obeying different arithmetic.
Nature turns out to violate two of the three, which was a genuine surprise, since nothing in mechanics or electromagnetism gives any reason for the laws to distinguish left from right. Nothing quietly breaks when they fail, because there was no charge to lose and no continuity equation to spoil; the symmetry is merely absent. The one combination that does appear to survive is not luck but a theorem, forced by requirements the later parts of this book impose.
Repeating the proof with a field in place of a coordinate costs nothing and returns much more. The same steps, with a four-dimensional divergence in place of a time derivative, deliver not a number that fails to change but a density and a current, tied by the statement that the density's rate of increase at a point equals what flows out. That is the equation promised when local conservation was first set out.
The gain deserves stating exactly. A global law says a grand total, added over everything, is the same at one moment as another. A local law says the amount inside any region changes only by crossing its boundary, so nothing may leave one place and arrive at another without traversing the ground between. The difference is not stylistic: a global ledger needs a notion of the same moment everywhere, and the next part of this book dismantles it.
The example to keep is a field of complex values whose scalar is unchanged when the phase is turned by one angle everywhere, and whose conserved quantity is, up to a constant, electric charge. That the angle must be the same everywhere should look suspicious for the reason already given, demanding as it does instantaneous agreement across the universe. Asking what happens when the angle may vary from place to place is the question that produces every force in nature.
The last step reverses the arrow, and shows there was never an arrow to reverse. Take the conserved quantity a symmetry produced, feed it back as a generator of motion, and what it generates is the symmetry it came from. Not an analogy but an identity. A conserved quantity and a symmetry are one object under two names, and which you reach for depends on whether you ask what stays the same or what moves.
These quantities close among themselves: combining two returns a third, with coefficients fixed by how the transformations compose. Rotations about different axes do not commute and their combination is not zero; translations do commute and theirs is. That structure is what a symmetry group concretely amounts to, and it makes the book's recurring instruction precise: name a group, write the most general scalar invariant under it, quantise.
Part I is finished. Forces are gone, replaced by one number attached to each history, nature selecting the one where that number stops changing; a single point now fixes a whole future; and every continuous symmetry hands over a conserved quantity that is the symmetry itself. One thing was left broken, when the momentum of two moving charges refused to add up. Repairing it means surrendering something so ordinary nobody has troubled to state it, that everybody shares one clock; the next part opens on the collision between electromagnetism and that assumption.
Two statements were held on excellent evidence by the 1890s, and they cannot both be true.
The first is old and Galilean: the laws of physics read the same in any laboratory moving steadily, confirmed every time nobody in a ship's cabin falls over. The second is that the equations of electricity and magnetism contain a speed, assembled out of two constants measured with a capacitor and a current balance, with no light anywhere in the measurement, and that the speed comes out as the speed of light. A speed in a law is a speed with respect to something, and three decades of increasingly delicate experiments failed to say what.
Part I supplied the instrument for handling a collision of this kind, though not the answer. Writing a theory as one number per history made it possible to ask of any law whether its shape survives a change of description, and to read that off the shape rather than grind through the algebra. The question is about to be put to Maxwell's equations, and the answer is that they survive a transformation nobody had thought to write down.
One thing should be named in advance so that it does not arrive as a conjuring trick. What has to be surrendered is not the ether, which nobody will miss. It is the shared present: the assumption, so deep that physics before this point has no name for it, that two events either happen at the same moment or do not, and that everybody can be told which.
Part II · Special Relativity
Geometry replaces mechanics.
2.1 The Crisis of 1900
A principle and the transformation that implements it are separate claims, and for three centuries nobody noticed, because the second was never perceived as a claim at all but as what the words of the first meant. The principle is Galileo's, from 1632: shut yourself below decks with some flies and a bowl of water, and everything proceeds as it did in port, so no experiment inside a smoothly moving laboratory can reveal how fast it is going.
The implementation is the familiar one. Subtract the distance the other origin has slid, leave the two directions across the motion alone, and give both observers the same reading for when anything happened. That last instruction carries all the weight, because it asserts a single universal clock, so that two events either are simultaneous or are not, with no reference to who is asking. It is the assumption the previous part promised would have to go.
Mechanics does not mind in the slightest. Differentiate twice and the steady relative velocity vanishes, so both observers agree about acceleration; and every force Newtonian physics uses depends on separations between bodies, which the shift leaves untouched. The law reappears in the moving frame with the identical shape rather than one needing repair by extra terms, which is exactly what the turning basis directions of polar coordinates had inflicted on it.
Nothing resembling a speed was put into the four equations of electromagnetism, and a speed came out. Apply the law tying a changing magnetic field to a circulating electric one to its partner, which ties a changing electric field to a circulating magnetic one, and the two fields uncouple, leaving one equation apiece. That equation is the one a chain of masses and springs produced back in the toolkit, and reading off its coefficient gives a speed built from two laboratory constants.
The provenance of those constants is the point. One is fixed by putting a known charge on two metal plates and measuring the voltage across them; the other by running known currents through parallel wires and measuring how hard they pull. Neither experiment contains any light, any motion, or any waves.
No channel exists by which optics could have leaked into a static charge on metal or a steady current in a wire, so the only way the number comes out right is if light is the disturbance those equations describe. Maxwell drew that conclusion in 1862, and optics, two thousand years old and with its own laws and its own instruments, stopped being a separate science. This is what a successful unification looks like from the inside: two numbers measured for unrelated reasons combining into a third that had already been measured for a third.
Sound in air, ripples on a pond, a pulse running down a rope, tremors through rock: every one of them travels at a speed assembled out of properties of the stuff it travels in, and therefore travels at that speed relative to the stuff. Stand in a wind and sound reaches you faster downwind than up, by exactly the wind's speed. So asking what the new speed was measured against was not a foolish question but an obligatory one.
The mathematics says it more sharply than the analogy does. A wave equation is written in some particular set of coordinates and its solutions run at the stated speed in those, so to an observer drifting past at some rate the same disturbance runs at a different one. The equations of electromagnetism therefore single out one frame, contradicting the principle of relativity before any experiment has been performed.
Naming a medium was the sober response rather than a fudge, because it was the only known way for a wave to work. It cost something. The medium had to resist shearing in order to carry a transverse disturbance, so it had to be an elastic solid filling all of space, with a stiffness-to-density ratio ten billion times steel's, through which the planets passed without measurable drag. Nineteenth-century physicists knew that list and it bothered them. An awkward hypothesis making a sharp prediction is still a respectable object.
Analogy has been carrying the argument, and analogy can always be resisted, so the case is now made by computation. Take the wave equation, strip it to one dimension, and rewrite it in coordinates sliding past at a steady rate, which is the chain rule and nothing else. The spatial derivative comes through untouched; the time derivative picks up an extra piece, because holding your position fixed in one frame means drifting backwards in the other.
Two things then go wrong at once. A cross term appears with no counterpart in the original, and it reverses sign when the motion does, which is precisely how an equation announces that it knows which way is forwards. The leading coefficient has also changed. So the equation holds in one frame and no other, and the demonstration is three lines long.
One entry in the calculation did all the damage: the statement that changing where you are does not change what time it is. The fork is therefore genuine, and there are exactly three ways through. Either the principle of relativity fails for electromagnetism and a preferred frame exists after all, or the equations of electromagnetism are wrong, or the transformation is wrong and the shared clock goes with it. The transformed equation is not nonsense; it correctly describes a wave in a medium seen by somebody moving through it. Everything turns on whether light is like that.
An experiment that finds nothing is worth exactly as much as the effect it was built to find. Timing light directly was hopeless, so Michelson compared two perpendicular paths by interference, which resolves a hundredth of a wavelength, and turned the apparatus through a right angle so the arms exchanged roles. The unknown mismatch between their lengths cancels out of that comparison identically.
Both arms are slowed, and missing that is the standard error. The arm along the motion loses more time crawling against the flow than it regains running with it, because the slow leg lasts longer and so counts for more. The arm across is slowed too, by a smaller power, since part of the light's fixed speed budget goes on keeping up sideways. The whole observable is the gap between those two powers, and it is second order in the ratio of the Earth's speed to light's, which is why every earlier experiment sensitive only to first order was guaranteed a null result before it was built.
The predicted shift was about four tenths of a fringe, in an instrument that could see a hundredth. Nothing appeared, at any hour or season, and a hundred and forty years of increasingly ferocious repetition has not made it appear. That is not a discrepancy in a coefficient; it is the absence of something that should have been unmissable by a factor of nearly forty.
FitzGerald proposed a repair of startling economy, and what makes it instructive is that it works. Suppose anything moving through the medium is shortened along its direction of motion by a particular factor, and the two round-trip times become equal, not approximately but identically, at every speed, because the shortening and the slowing are related by an algebraic identity. The predicted shift is then exactly zero.
It was not absurd either. If the forces holding matter together are electromagnetic, and fields are distorted by motion through the medium, atoms in a rod might settle at a different spacing once it is set moving, so the rod is genuinely shorter and the medium genuinely there. With that and one further substitution, the theory reproduced every optical experiment then known. A null result cannot kill a determined theory; what an experiment refutes is a model rather than a concept.
So the medium was not disproved. It stopped doing any work. Every quantity it had been introduced to explain turned out calculable without it, and every attempt to detect it produced a null result patched by a property whose only content was the null result it patched. A frame no experiment can pick out is not a frame of reference. The verdict is methodological rather than empirical, and deserves stating without condescension, since modifying a productive hypothesis minimally under pressure is what anybody should do.
Notice how little is being asked for at the end of all that. Two sentences: the laws take the same form in every inertial frame, all of them and not merely the mechanical ones; and light travels in vacuum at one definite speed, the same for every observer and independent of the motion of its source.
The first sentence was already common property with one word altered, and that word forbids the branch on which mechanics obeys the principle while optics does not. The second is the radical one, and it is not a disguised convention about setting clocks: a round-trip speed can be measured with one clock, one mirror and one pulse, synchronising nothing, so the claim has empirical content and has been checked to absurdity. Photons from particles moving at nearly the limiting speed arrive at that speed, not at twice it.
What makes the second sentence strange is worth locating exactly. A medium delivers independence from the source for free, since once a disturbance is launched the medium takes over and forgets its origin; asserting that independence with no medium to enforce it is a far stronger claim, and asserting that the speed is also the same for every observer is one no medium delivers. The interferometer had already revealed the shape of the answer, one factor along the motion and none across it. What it could not reveal was why.
2.2 The Lorentz Transformation, Derived
Two postulates cannot determine anything on their own, because a constraint is always a constraint on some class of candidates and nobody has yet said which class. Three further assumptions do that work, and the derivation's reputation for being a conjuring trick comes entirely from leaving them unspoken. Each is a physical claim, each is testable, and each could have come out otherwise.
The first is that no place and no moment is special, so the dictionary between two observers' coordinates is the same wherever and whenever they set up. That forces the dictionary to be linear, by an argument better than the usual one. The familiar route, that free particles travel in straight lines for everybody so straight lines must go to straight lines, leaves a loophole, because the maps preserving straightness include ones carrying a denominator, and a denominator that vanishes somewhere marks out a special place and time in the universe.
The second is that no direction is special, which leaves the two transverse directions alone and keeps the sideways coordinates out of the other two equations. The third is that the relation between the observers is symmetric, each receding from the other at one rate. Those three, with the two postulates, are the complete input; if what comes out is surprising, the surprise is living in one of five statements and has nowhere else to hide.
Everything in the derivation is bookkeeping except one step, and that step is the only place where any physics of light enters. Demanding that the moving origin move at the stated rate fixes the spatial equation up to one unknown factor, and costs almost nothing, since the old transformation is the case where that factor is one.
The light condition then goes in, once for a pulse each way along the axis. Adding the two resulting equations forces the same unknown factor into the time equation as sits in the spatial one; subtracting them forces the time equation to acquire a term proportional to position. That term is the end of simultaneity, and it arrives before the remaining factor is known. Symmetry between the observers pins that factor down, and no second candidate appears.
The bonus is larger than the result. Invariance was demanded only of the events light connects, and what comes back is a combination of the time and space separations that is the same for every pair, whatever its value. Quantities invariant for reasons nobody demanded are worth building a subject on. The old transformation survives as the leading term of an expansion, and its two corrections do not enter at the same order: lengths and rates are wrong at second order, while simultaneity is wrong at first order multiplied by a distance, so the largest low-speed effect concerns distant clocks.
Whether a family of transformations closes on itself decides whether you have a theory or a formula. Compose two boosts along one axis and the result is another boost, with a parameter emphatically not the sum of the two, since half the speed of light followed by half the speed of light delivers eight tenths. Closure is neither automatic nor decorative, because had it failed there would be no coherent notion of the frames of physics, and some chain of changes of viewpoint would have carried you outside the theory.
Two smaller facts fall out along the way. The transformation has determinant one, so it preserves four-dimensional volume, which is the determinant doing the job the toolkit gave it. And the whole apparatus is now a matrix acting on four numbers, so composing changes of frame has become multiplying matrices and nothing more.
One caveat is worth flagging because it has no ancestor in older physics. Boosts along a single axis behave; boosts along different axes do not compose into a boost at all, and the product leaves a spatial rotation behind. That rotation is measurable, contributing to the fine structure of atomic spectra, and nothing in Newtonian mechanics predicts anything of the kind. The composition rule also says plainly that velocity is the wrong quantity to be adding.
Three consequences come out of one substitution, and the order in which they are met decides whether the subject ever becomes intelligible. The root fact is that two events at the same time but different places for one observer are not at the same time for another, and the two famous effects follow from it rather than accompany it. Readers who meet them the other way round acquire a picture in which clocks and rulers are mysteriously defective while the word now stays universal, and they rarely recover.
It does not stay universal. Each state of motion carries its own family of nows, cutting through spacetime at different angles, which is what the local conservation laws of the last part were built to survive. It is also why each observer finding the other's clock slow is no contradiction: one clock is compared against a pair synchronised in the other frame, the two observers use different pairs, and the offset between them is exactly the size that balances the books.
Contraction turns out to be a property of the measurement rather than of the rod. Locating both ends at one moment is a demand about simultaneity, so different observers use different pairs of events and get different answers, and a pair the rod's own frame calls simultaneous flips the answer the other way. Nothing was squeezed, and nothing was done to the rod.
The rule replacing the addition of velocities earns its place by passing a test nobody arranged for it. Invariance of the light speed was imposed for one pulse leaving one origin in one of two directions, and the rule that resulted now certifies that anything travelling at that speed in one frame travels at it in every frame, anywhere, at any time, in any direction.
The barrier that follows is not a wall anybody installed. Rate a speed by the quantity measuring its distance from the limit, and combining two speeds multiplies the two ratings together, up to a positive factor. Small positive numbers can be multiplied together forever without reaching zero, so no finite chain of ordinary boosts arrives at the limit, and the whole of the light barrier is the observation that a product of positive numbers is positive.
None of this forbids coordinate speeds above the limit, and it is worth being clear which is which. Sweep a laser across the face of the Moon and the spot crosses it many times faster than light, because the photons landing at one place and those landing next door travelled independently and nothing went sideways. Close a pair of shears at a shallow enough angle and the crossing point outruns light for the same reason. Ask what is being transported, and where the answer is nothing, any speed is permitted.
Compose two rotations of a plane and their angles add, as anybody expects of a family labelled by one number. Boosts along an axis are such a family, so something of theirs adds, and it is not velocity. Take the composition rule seriously and it is a formula everybody has seen, the addition theorem for the hyperbolic tangent, so what adds is the angle whose hyperbolic tangent the velocity is.
Call that angle the rapidity. Rapidities add exactly, with no correction term, and their matrices multiply the way two turns of a plane do. This collects the pairing between a generator and the family it builds, promised when a self-partnered map was first exponentiated and again when conserved quantities were made to push things around: one fixed matrix, exponentiated by the parameter, generates the family.
Why velocities looked defective is now visible, and it is not a fact about nature. Velocity is the slope of a worldline and rapidity is its angle, and slopes have never added: two turns of a plane combine their angles cleanly while their slopes combine by an ugly quotient that shocks nobody, since nobody expected slopes to be the natural label. Three centuries were spent adding the slopes. The barrier stops being mysterious in the same stroke, since the angle runs over every real number and adds without limit while its hyperbolic tangent creeps towards one and never arrives.
One twin comes back younger than the other, and the asymmetry permitting it is physical rather than verbal. The one who stayed occupied a single inertial frame throughout; the one who travelled was at rest in one frame going out and in a different one coming back, and the switch is detectable from inside the ship, with no window needed. One history is bent and the other is not.
Saying so explains why the two may differ without showing where the missing years went, and they can be tracked. The traveller's claim that the other clock was running slow is correct on each leg taken separately. What he may not do is stitch the two legs' notions of the present together, because at the turn his family of nows swings through an enormous span of the other worldline, and that span is precisely the amount the naive accounting was missing.
Every paradox in the subject is this one in different clothes. Somewhere in the statement, someone has treated the simultaneity of two distant events as a fact rather than a labelling, and the cure never varies. List the events; ask which claims are local coincidences of one thing with another at one place, since everybody agrees on those; transform the rest; and the smuggled assumption will be standing in plain view. What this pile of results is waiting for is the idea that organises it.
2.3 Minkowski Geometry
Out of a demand made about one narrow family of events came an equality holding for every pair of them, and that overshoot is what makes a geometry available. The combination of time separation and space separation that all observers compute alike has a name, the interval, and the way to see what it is doing is to ask first what a rotation is.
The answer taught first is a rigid turn about an axis, which is a picture rather than a definition. The definition is algebraic: a rotation is a linear map leaving the sum of the squared coordinates alone. Nothing about turning is needed, and the logic runs from the quadratic expression to the transformations rather than the other way. The quantity those transformations preserve is what the word length means, not a separate fact about the world but the name for whatever observers with differently oriented axes agree on.
Read the invariance again with that in mind, and the whole chapter is one substitution. Flip the sign of the three spatial terms in the quadratic expression, and the rigid motions change from rotations into the transformations of relativity. That is not an analogy; it is the same definition with a different expression inside it. The converse holds too, since the maps preserving the flipped expression are exactly the boosts, up to a reflection of space or of time.
Notation deserves attention when it makes a structure visible rather than merely shorter, and two decisions do that here. The first is to measure time in the same units as distance, so that two coordinates entering the transformation on an equal footing are seen to do so instead of being held apart by a conversion factor. The second is to collect the four coordinates of an event under one symbol carrying a label.
The signs then need somewhere to live, and they live in a small square array with one plus and three minuses down its diagonal, called the metric. Its entire job at this stage is to supply a sign to each term. In ordinary space the corresponding array is all plus signs, which is why nobody ever writes it and why the metric is invisible in first-year physics. The whole difference between the two geometries is the difference between those two arrays.
What the arrangement buys is a machine for manufacturing agreement. Take any two collections of four numbers that transform between frames the way a displacement between events does, combine them through the metric, and the number resulting is one every observer computes alike. The interval is the special case where both collections are the same displacement, and the machine will be fed repeatedly: a velocity shortly, a momentum in the next chapter, a current after that.
Take one event and apply every transformation in the family to it, then ask what curve the image traces. In ordinary space the answer is a circle, since the sum of squares is preserved and the angle sweeps through everything, and a circle is a locus of constant distance from the centre. Flip the sign and the same question returns a hyperbola, with the light rays through the starting point as its asymptotes.
So hyperbolas are the circles of this geometry, the loci of events at fixed interval from a given one. Boosting slides a point along the hyperbola it started on and never carries it to a neighbouring one, which shows up as large changes in both coordinates alongside no change in the interval. The light rays are the degenerate member, where the curve collapses onto its own asymptotes, and that member has no Euclidean counterpart, because a sum of squares vanishes at one point while a difference vanishes along two lines.
This disposes of the complaint that boosted axes look wrong. They are genuinely at right angles in this geometry and fail to look it because the eye applies the geometry of the paper; the invariant hyperbolas supply the tick marks, since they belong to no observer. Longer on the paper means less elapsed time, so the two measures do not merely differ by a scale factor, they run in opposite directions.
In ordinary space the distance between two distinct points is positive and there is nothing further to say. Reverse one sign and the same quantity comes out positive, negative or exactly zero, and which of the three it is turns out to be the most important thing about the pair. The proof that the classification belongs to the pair rather than the observer is two words long: the quantity is invariant, so its sign is.
What has appeared is not another agreed number but an agreed structure: what survives of before and after once the universal present is gone. Events light or matter could reach keep their order in every frame, while events too far apart for anything to cross between have an order different observers reverse without anybody being wrong, since nothing was ever at stake in it. Time order is absolute where it could matter and negotiable where it could not.
The speed limit follows, and it belongs to the geometry rather than to light. Let some influence outrun it and the two events it joins are of the negotiable kind, so some observer records the effect before the cause. Two such influences are worse: one person signals another, the second replies with an influence instantaneous in his own frame, the reply arrives before the first was sent, and she declines to send it. Light travels at the limit because light is massless.
Straight separations have carried everything so far, and no real history is straight. Bend the path, chop it into pieces short enough to be straight, take the interval along each and add them, and what accumulates is the proper time. It is a functional of the whole path rather than a function of its endpoints, in the sense the action was, and it is the arc length of the worldline in the geometry spacetime actually has.
Set it beside the arc length of a curve in the plane and the two expressions differ in one place only, the sign under the root. That single difference does all the work of the argument to come. The quantity is invariant besides, because each piece is, so two observers who disagree about every intermediate number agree about the total.
It is also what a clock reads, which keeps it from being an abstraction: in a clock's own momentary rest frame the interval between neighbouring events on its path is the time it displays. Dilation stops being a statement about defective clocks and becomes one about the length of a curve. One physical assumption hides in the word momentary and deserves naming: an ideal clock's rate depends on its speed and not on its acceleration. That is no theorem, a pendulum in a launching rocket violates it, and it holds experimentally to a degree hard to credit.
The most familiar sentence in geometry is that a straight line is the shortest route between two points, and here it fails in the most interesting available way. Among all the paths a material object could take between two given events, the unaccelerated one accumulates the greatest elapsed time, and every other loses strictly.
The proof is four lines because the frame may be chosen first, and it may be chosen because the answer does not depend on the choice. Work where the two events happen at the same place. The path that stays put accumulates the full coordinate time, while every other pays a penalty at each instant for moving, a penalty never zero and never negative. So the travelling twin ages less for the same reason a detour through the next town adds mileage, except that in this geometry the detour subtracts, and the triangle inequality has turned around.
Two debts are settled by that. The warning that a stationary path need not be a minimum, issued when the action principle arrived, gets its cleanest illustration, since this stationary path is a maximum. And the calculation needed nothing but a functional of that shape and the machinery of the last part, not flatness and not constancy of the array of signs. Replace that array by one whose entries vary from place to place and the identical calculation returns the paths of free fall.
One construction remains, and it is forced rather than chosen. Ordinary velocity has a defect that surfaces once frames matter: its numerator and denominator belong to different worlds, the displacement being part of a four-dimensional object while the elapsed time is one observer's coordinate. The repair is to divide by something nobody disagrees about, and the worldline's length is the one candidate.
What comes out has a fixed length, the same for every particle at every moment whatever it does, so all its freedom is in its direction. Dilation then reads as a fixed-length object tilted from your time axis having a smaller projection along it. The slogan about everything moving through spacetime at one speed gestures at this and is a mnemonic rather than a derivation, since the two pieces combine with a minus sign and both run to infinity rather than trading against each other.
That is as far as three chapters carry it. The speed limit is built into the geometry rather than into any material; observers slice one fabric at different angles and agree on the interval; the straightest history carries the most time. But the word used throughout for a collection of four numbers has been a definition by resemblance, and the up-and-down placement of its labels has been a spelling rule obeyed on trust. Saying what makes such a collection an object rather than a list comes next.
2.4 Tensors, Honestly
Numbers on their own carry no context. A set of components means something only relative to the framework used to measure them, which is to say relative to a choice of axes, so shifting that choice must shift the numbers describing one unchanged reality. That is not a defect in the description. It is what describing anything physical consists of.
The fatal error is the opposite instinct, which is to freeze one set of numbers in place and require everybody to use it. Doing so does not make the object more real. It produces observers computing conflicting values for what was supposed to be a single fact about the world, and this section stages exactly that collapse with four entirely innocent-looking numbers.
So the goal of everything that follows is not to hunt for quantities that refuse to change, since almost nothing does. It is to find quantities that change in a fully predictable way when the choice of axes changes, so that the physical statements assembled out of them come out the same for everybody, whatever axes each of them chose.
Objects here are defined by their behaviour rather than by their appearance, which feels backwards the first time it is met. Nothing dictates what a vector is made of or what it should look like, and the definition says only how its description must adapt when the frame of reference changes.
The prototype for that behaviour is the physical gap between two nearby events. Since the two events happen whether or not anybody is watching, the rule converting one observer's measurement of that gap into another's is fixed by the change of frame alone and by nothing else. Any object transforming by that same rule is a vector, and there is no further qualification to meet.
This behaviour-based definition is worth its initial awkwardness because it abandons any need for flat backgrounds, straight arrows, or a fixed grid on which to draw them. Nothing in it depends on the geometry being simple. That is precisely why it will survive perfectly intact into Part III, when the coordinates begin to curve and the comfortable picture of an arrow stops making sense.
Spacetime turns out to be inhabited by two distinct but complementary types of object. The first is the ordinary vector, which you can visualise as an arrow representing a physical displacement, pointing from here to there. The second is the covector, which is better understood as a measuring device than as a thing being measured. Rather than an arrow, picture a covector as a series of evenly spaced parallel sheets, much like the contour lines on a map.
Pair the two together and the covector counts how many of its sheets the arrow punches through, and that count is the whole of the pairing. The single image explains why their numbers move in opposite directions whenever the units change. If you stretch your measuring ruler, the arrow's numerical components shrink, and so the covector's sheets must crowd closer together in order that the total count of pierced sheets comes out exactly as it did before. That count is an indisputable physical fact, and no choice of ruler is entitled to alter it.
The only reason these two species are so frequently conflated is that their numbers happen to mirror each other perfectly in a simple, flat, Cartesian grid. That agreement is a mathematical coincidence of the flat ruler rather than a law of nature, and it shatters the moment either the geometry or the coordinate system becomes interesting.
Because arrows and sheets are genuinely different species of object, something has to translate between them, and the metric is that translator. Feed the metric a displacement arrow and it hands back the particular stack of sheets that measures lengths in the same way. This is the physical reality behind the notation of moving an index up or down: it is a literal act of translation, performed by a specific piece of machinery, rather than a rearrangement of symbols on the page.
The distinction is not a typographical flourish. Depending on the rules of the space you are working in, the translation can flip the mathematical signs of your components, which means that an index moved carelessly is a sign error already in flight, and typically one that will not be discovered for several lines yet.
Most importantly of all, the translation depends entirely on which metric you hand it. Here that metric is a rigid, unchanging table of values, identical at every point in spacetime. Later it will evolve into a flexible field that varies from place to place, taking different values here than it does over there — and that single conceptual shift, from a fixed table to a field, is the entire foundation of gravity.
A tensor with several slots is nothing more than those two objects layered, and the layering introduces no new idea. Whether a given index sits upstairs, behaving like an arrow, or downstairs, behaving like a stack of sheets, its transformation rule is the basic rule applied to each index in turn, one at a time and in any order. Nothing conceptually new arrives with the extra slots; the same idea is being stacked.
What genuinely matters is understanding which mathematical operations are safe to perform. You may safely add tensors of identical type, and you may safely multiply them together. You may also safely contract them, which means pairing one upper index against one lower index so that both are consumed and the resulting object is simpler than what you started with.
That is why the rules of this mathematics always demand pairing an up with a down. Summing two upper indices together is not a breach of house style, and calling it bad form understates the damage. It produces a statement meaning different things to different observers, which is another way of saying it means nothing at all, and the notation is built so that this particular failure is visible at a glance.
What all of this machinery is for is a single check, run quickly and run often: whether a proposed law of physics is true for everybody or is a quirk of one laboratory. The check has to be cheap, because it is going to be run on every equation in the rest of the book.
Without such a system you would have to rewrite every equation from the viewpoint of a moving observer, grind through the algebra and see whether the structure survived, and then repeat that exercise for every new law anyone proposed. Tensor notation replaces that calculation with an inspection of the page. If both sides of an equation are the same type of tensor, carrying matching indices at the same heights, the law holds in every frame and nothing further need be checked. The frame-independence has become visible in the shape of the equation itself, before a line of computation is performed.
This turns what looks like a writing convention into a design rule for physical law. It also explains the fate of Newton's second law, the one relating force to mass and acceleration. That law was never wrong, and at everyday speeds it remains superb. But its content changes with the observer's frame, and that alone disqualifies it as a fundamental statement about nature rather than a description of one laboratory.
Any object carrying two indices can be neatly divided into two distinct halves: a symmetric part, which remains identical if you swap the order of its indices, and an antisymmetric part, which flips its mathematical sign under that same swap.
The power of this division is that it survives any change of frame. What is symmetric to one observer remains symmetric to all of them, which proves that this is a genuine physical division of the object rather than an artefact of whichever notation someone happened to choose. It is a real seam in the thing itself, which is to say the object has fallen apart into independent pieces, and the two halves go on to have entirely separate careers.
Count how many independent numbers each part requires in four-dimensional spacetime and the symmetric part holds ten, the antisymmetric part six. Both counts turn out to matter enormously. The ten symmetric numbers will become the metric of curved spacetime, which is to say the ten functions that encode gravity itself. The six antisymmetric numbers will turn out to hold the electric and magnetic fields, three components each — the first substantial hint that electricity and magnetism were never two separate forces to begin with.
The strict formatting rules governing these equations are not a matter of fussiness. They are an error detector built into the notation, and it catches almost everything.
For an equation to be valid, any free indices left over on one side must match, in both number and height, the free indices sitting on the other. Indices that have been summed away are internal bookkeeping and may be renamed freely, but a single index name may never appear three times within the same term.
Each of these rules follows from what the objects are rather than from a taste for tidiness, and each can be re-derived from the transformation law in a line. Taking them seriously catches fatal mistakes at once, with no physics computed at all. Almost every algebraic slip made over the coming chapters will announce itself as a mismatched index long before it can corrupt a final number. That makes the check of index balance a constant reflex, performed in the same way and for the same reasons as checking the units of an answer, and it costs about as much.
2.5 Relativistic Dynamics
The object sketched at the close of the last chapter now gets built, and the diagnosis is sharper than the sketch allowed. Divide a displacement between two events by the time one observer's clock assigns to it: the numerator behaves impeccably, being the prototype everything else is measured against, while the denominator belongs to whoever holds the clock. What results is spoiled by a stray factor depending on the particle's own speed, so two particles watched through one change of frame are spoiled by different numbers and no rescaling rescues either.
Divide instead by the reading of the clock the particle carries, which everybody computes and everybody agrees about, and the defect has nowhere to come from. What comes out has the same length for every particle at every instant, and that constraint is not a discovery about matter but the definition of the carried clock rearranged: a curve measured by its own length has a tangent of fixed size.
Differentiate once more and something arrives for nothing. Since the length never changes, its rate of change stands perpendicular to it in the interval's geometry, so a force can turn a fixed-length arrow and can never stretch it. The speed limit stops being a barrier a particle runs into and becomes the observation that turning is the only motion on offer, and turning keeps you on the surface you began on.
Every conservation law is somebody's claim about a process, and what matters is whether the claim survives translation into another observer's language, since one that does not is a report about a laboratory rather than a fact about the world. Three numbers cannot survive, and the failure is not gentle. Two identical lumps of putty approaching at equal speeds and sticking conserve the old momentum exactly in the symmetric frame, while an observer drifting past finds the outgoing total larger than the incoming one by more than a third.
Notice precisely what that is. It is not that the old momentum is slightly wrong at high speed, since in that frame it was right to every order. What fails is agreement about whether the law holds at all, and a disease of that kind admits no correction term. Conserve the four-part object instead and the disease has nowhere to live, since the transformation between observers is linear and carries zero to zero.
Then comes the part that cannot be declined. Suppose you accept the object but wish to conserve only its three familiar entries. A change of frame mixes the fourth into the first three, so demanding that the three vanish for every observer forces the fourth to vanish as well. Frame-independence has selected the object out of all the candidates, and having taken three of its components you are stuck with the fourth.
A fourth conserved quantity stands about unnamed, and the only honest route to its identity is to expand it for a slow particle and watch what it becomes. Out comes a constant belonging to the particle, then the Newtonian kinetic energy, then corrections falling with the square of the speed ratio. Nothing there is a concession: nearly everything is an approximation, and the controlled kind names its next term. What licenses the name is that when a collision leaves the participants the species it found them, the constant cancels and kinetic energy is conserved as before.
The constant is where the interest lies. Newtonian mechanics let you shift the zero of energy by any amount, which is why nobody asked a stationary object's energy; the question had no content. That freedom is withdrawn, since the constant cancels only while the rest masses hold, and processes change them: a nucleus binds, a particle decays, two lumps of putty stick and come out heavier than their parts.
Rest energy is a reservoir rather than an offset, at a fixed exchange rate against motion. Notice how little was requested and how much arrived. Nobody went looking for it, no experiment produced it, no argument about matter suggested it. It was conscripted: the demand that a law read alike for everybody left a fourth component lying about, and what it turned out to be is the equation everyone recites.
Contract the welded energy and momentum with itself and the mass falls out as a number nobody disputes. Observers assign the same electron different energies and different momenta and all compute the same mass, which makes mass a label the object carries rather than something growing with speed, exactly as the interval is the label a pair of events carries. Plot energy against momentum and the hyperbolae drawn two chapters ago for time and position reappear, since a change of frame slides any such object along its own invariant square.
One consequence deserves extracting first, because it makes the massless case work: the velocity is the ratio of momentum to energy, with no dilation factor and no mass in it anywhere. That recipe does not care whether a mass exists. Set the mass to zero and it returns the speed limit exactly, at every energy, with no approximation.
Reading the massless case as the end of a sequence of ever lighter particles is the expensive mistake to avoid. The construction that built momentum from a mass and a carried clock does not survive that limit, since a light ray has no carried clock to divide by, and the phrase naming the time a light ray experiences names nothing. The repair is to invert the order: take the four-part momentum as primary and let mass be the label it carries.
Guessing which number to attach to each history would be a poor way to proceed, so the calculation runs backwards: demand that the machinery of the first part return the repaired equation of motion, and solve for it. What comes back for a free particle, once a harmless constant is set aside, is startling in its plainness. The number attached to a history is the reading of the clock carried along it, multiplied by the mass and the square of the speed limit, with a minus sign in front.
Two statements this book has made separately are therefore one statement. The first part said nature selects the history at which the number stops changing; the chapter before last proved that among all histories joining two events the unaccelerated one carries the most time. The minus sign reconciles them, since making a quantity largest is making its negative smallest.
Something else is settled on the way past. The number being attached is no longer kinetic energy less potential energy, resembling that combination in nothing but its low-speed limit, so the warning issued when the action principle first arrived, that the familiar difference is a special case and not a definition, collects here. Notice too which object is well behaved: neither the integrand nor the element of time multiplying it is the same for every observer, and only their product is.
The second law costs more to replace than it appears to. Differentiate the four-part momentum by the carried clock and you have a force with the right credentials, carrying a constraint nobody imposed: since the four-velocity has fixed length, the force stands perpendicular to it, so only three of its entries are free. Write out the entry that is not free and it says energy changes at the rate the ordinary force does work, which in Newtonian mechanics was a separate derivation and here is the fixed length restated.
Then the damage. Force and acceleration stop pointing the same way. Push a fast particle along its motion and it responds far more sluggishly than to an equal push across it: at a dilation factor of ten, a hundred times less so. That is why transverse focusing of a beam and forward acceleration are engineered as separate problems with separate hardware.
A habit sixty years of textbooks kept dies with it. Absorbing the dilation factors into the mass makes the formulas look Newtonian again, at the price of the mass being two numbers at once, one for a push along the motion and one across. For any other direction no single number works, the two vectors not being parallel, so a matrix is needed, and calling a matrix the mass abandons what the word was for. The phrase will not appear again.
Waves need one more four-part object, and the argument producing it is counting rather than calculation. A crest is not a thing that travels but a set of events at which the field is momentarily largest, so whether a given event is a crest is a question about one point of spacetime that both observers are asked about. Let a detector click once per crest between two events on its own history: the count is a whole number, nobody gains or loses one by moving, and whole numbers cannot transform. So the phase is agreed, and whatever pairs with position to produce it must be four-part too.
Frequency and wavelength are thereby welded together as energy and momentum are, and the whole Doppler effect becomes one multiplication. Light from a source passing at closest approach, where the separation is momentarily unchanging and the old theory predicts no shift, arrives reddened by exactly the dilation factor.
Two null objects have now appeared for unrelated reasons, one for a massless particle and one for a light wave, sharing a transformation law, a null condition and the same relation between time and space parts. One constant relating them relates both halves at once, since a single scalar between two such objects covers everything or nothing. Planck's relation for energy and de Broglie's for momentum stand or fall together, which was available years before anybody believed either.
Add the four-part momenta of any collection of particles, whatever they are doing, and the total is another object of the same kind with an invariant square of its own. That square defines the collection's mass; everybody agrees on it, and since the total is conserved it is conserved too. In the frame where the collection is collectively at rest it is the total energy over the square of the speed limit, which is the single-particle statement about rest energy carried over to many.
What this does to the word mass repays dwelling on. Two pulses of light of equal energy flying apart possess a mass, built entirely from constituents having none, while the same two aimed the same way possess none. Nothing changed but where they pointed. Mass is a property of the total, measuring how far the constituent momenta point in different directions, so randomly directed motion inside a box shows up as mass of the box.
The practical consequence is a piece of accelerator engineering. Only that invariant mass is available for making anything new, and energy tied up in the forward motion of the whole cannot be spent. Doubling the beam energy of a machine whose beams meet head-on doubles what is available; doubling it in one firing into a stationary target multiplies it by the square root of two. That difference in scaling is the whole argument for building colliders.
A bound system, assembled from constituents that began far apart and settled by radiating the excess away, weighs less than its parts by exactly what left. Chemistry does this at one part in a hundred million, far under what Lavoisier's balance could resolve, which is why he pronounced mass conserved. Nuclear binding does it at one part in a thousand, and that gap of five orders is why nuclear energy differs from chemistry in kind.
The usual telling, in which mass converts into energy, is not what happens. The conserved total never changes; only which column it sits in, rest mass or motion. Run the reaction in a sealed reflecting box and weigh it: the mass is unchanged, everything released still inside. It falls only when the heat is let out.
One step further and the correction becomes the main effect. The three quarks in a proton account for about one hundredth of its mass; the rest is their confined motion and the binding field, appearing as mass because their momenta point every way and cancel while the energies add. That number is the one the chapter on expansion said no series would find, lying so flat near the origin that every term reports it as zero. Mass entered as the amount of stuff in a body and leaves as a label on a four-part object, better than ninety-nine per cent of yours borrowed motion.
2.6 Electromagnetism Is Relativity
Charge conservation arrived long ago as a four-term equation whose behaviour under a change of frame was completely opaque: one derivative in time of one function and three in space of three others. Set beside the machinery of the last two chapters, that is the shape of a four-dimensional divergence, provided the four functions assemble into one object. Rather than guess at the arrangement, demand it, and no freedom is left: the charge density, carrying one factor of the speed limit so the entries are commensurable, sits above the three of current.
Charge density is not itself an agreed number, although the charge is. The same charge occupies a contracted volume when you run past it, so the density picks up the dilation factor the four-velocity already carried. Density of charge and density of current are one thing seen from different states of motion, exactly as elapsed time and displacement were, and as energy and momentum were a chapter ago.
The machine assembled three chapters back for manufacturing agreement between observers was promised a velocity, then a momentum, then a current, and it now has all three and wants nothing further. What it returns on this last feeding is charge conservation as a single contraction, the same statement for everybody at a glance, where before it was four terms whose fate under a change of frame nobody could see.
Two of the four equations are spent before the real work starts, and what they buy is the existence of the potentials: one number and one three-part object, out of which both fields are then built. Look at how they are built and every entry is one derivative of one potential component minus a different derivative of a different one. In four dimensions exactly one object has that shape, and it reverses its sign when its two labels are exchanged.
The antisymmetry is forced rather than noticed. Potentials are not unique, since a whole function's worth of freedom may be added without altering any field, and the only way to build something blind to that freedom from one derivative and one potential is to keep the part reversing sign, the leftover being symmetric and impossible to dodge otherwise. Six independent entries follow, and that count was performed a chapter ago with the answer promised and withheld.
Here is the answer. The six are the three electric components and the three magnetic ones, arriving by construction with no room to choose otherwise. They are not two fields that happen to fit inside one container. They are one object sliced by an observer, and which slices somebody calls electric depends on that observer's motion in exactly the way that which part of a separation between events somebody calls time depends on it.
Written in the new language four equations become two, and the compression is not typographical. One of them carries a source and admits a single sensible arrangement of labels; expanding its four entries returns the law relating field to charge and the one relating circulating magnetic field to current. The term Maxwell inserted by hand, whose absence made the set inconsistent with charge conservation, is not inserted at all: it is one entry of a sum obliged to run over all four values because the label is contracted.
The other equation is not physics. Write the field tensor in terms of the potentials and six terms cancel in pairs, each pair differing only in the order of two derivatives. Given potentials, the absence of magnetic sources and the law of induction cannot fail, and carry no information about how electromagnetism works. They are the toolkit's two identities, the vanishing swirl of a gradient and the vanishing outflow of a curl, which in four dimensions are one.
Both lines relate objects of the same type, so each holds in every frame the moment it holds in one, with no chain rule, no cross terms and none of the carnage that transforming the wave equation by hand produced earlier. The content of the subject is thereby one sentence: charges and currents make fields. The rest is the price of using potentials, and the price is nothing.
A whole function's worth of freedom sits inside the potentials and no experiment can see it. Add the four-dimensional gradient of anything smooth and both fields come out unaltered, for the reason that has now done this work three times, that two derivatives are indifferent to their order. Since the fields are what push charges about, potentials differing that way describe the same world. Freedom is an obstacle when you want to solve for something, so it gets spent, and one condition spends enough.
The condition chosen is itself a contraction, so imposing it in one frame imposes it in all, which the other common choice cannot claim. With it in force the four components stop talking to one another and each satisfies a wave equation with its own source. Remove the sources and what remains is the equation that opened this part, obtained in three lines because the bookkeeping was done first instead of last.
That dissolves the question which broke nineteenth-century physics. The operator in that equation is a contraction, so it reads the same for everybody, and the speed enters it only through the array of signs, which is to say through the geometry rather than through any substance filling space. Maxwell's equations do single out a speed. They never singled out a frame, and only the assumption of one universal clock made those two look like the same claim.
If the two fields are entries of one object, changing frames must shuffle them into each other, and the rule needs nothing beyond the transformation law in hand. The entries along the direction of motion come through untouched, the reverse of what happens to a four-part vector, whose parallel component is the one mixing with time; a boost acting on both labels of one entry undoes itself. Across the motion the fields trade, and the combination appearing in the trade is the familiar force per unit charge.
The immediate consequence is that a purely electric field is not a notion anybody can defend as absolute. Let one observer find no magnetism anywhere and another moving past will find some. Magnetism is what an electric field looks like from a moving frame, a slogan about to be turned into arithmetic checkable against a laboratory bench.
What the slogan does not license is the belief that either field can always be transformed away. Some configurations refuse, and two agreed numbers built by contracting every label decide which. The exact statement is weaker and better: the split between electric and magnetic is observer-dependent in precisely the way the split of a displacement into time and space is. Two observers disagree about how much of the object is electric in the same sense that they disagree about how much of a momentum is energy.
The argument making the claim literal is checkable against a current balance. A current-carrying wire is electrically neutral in the laboratory, which is experiment and not assumption, since a live wire does not attract a pith ball, so a charge moving alongside feels a purely magnetic force. Now board that charge. In its own frame nothing moves, so no magnetic force is available, yet it still accelerates towards the wire, whether it strikes not being a matter of opinion.
Something else is doing the work, and there is one candidate. The wire holds two populations of charge, a lattice at rest and a gas of electrons drifting through it, contracting by different factors when frames change, the electrons having been moving already and the lattice not. The cancellation that made the wire neutral is spoiled, the wire carries a net charge, and the attraction is ordinary electrostatics. Both compute the same trajectory and neither description is deeper.
The numbers make it startling. Electrons drift along a copper wire at under a tenth of a millimetre per second, two and a half parts in ten million million of light speed, so magnetism is an effect of that size. Electromagnets lift cars because the unbalanced quantity is colossal: thirteen thousand coulombs of mobile charge per metre, cancelled to that precision by the lattice. A minute fractional imbalance in an enormous cancellation is still a substantial charge.
Since each field separately depends on who is looking, the quantities worth having are those nobody can dispute, which means whatever survives contracting every label away. From one antisymmetric object there are exactly two such numbers: the difference between the squared magnitudes of the fields, and the amount by which they overlap. Everybody computes the same pair, whatever their motion.
The pair sorts field configurations into types not open to argument, as the sign of the interval sorted pairs of events into three kinds. Where the fields stand perpendicular and the electric one dominates, some observer finds no magnetism at all; where the magnetic dominates, some observer finds no electricity; and where the two are not perpendicular, neither can be removed by anybody ever, the best available simplification being a frame in which they lie parallel. The test is run without leaving home, because the quantities tested are agreed.
Light is the case where both numbers vanish, and since they vanish for everybody, no change of frame turns a light wave into anything else. Its frequency may be shifted to whatever you like and its amplitude scaled by any factor, and it can never be brought to rest, because no frame exists in which the two fields become independent. Light sits at the boundary between the electric and magnetic types exactly as a light ray sits on the boundary of the light cone.
Three requirements and a shortage of materials pin down the force on a charge. It must be a four-part object, proportional to the charge, and assembled out of the particle's four-velocity and the one field object available; essentially one candidate meets all three. Expanding it returns the force law the first part of this book had to quote without justification, with the repaired momentum riding along at no extra cost, while its time entry says the particle's energy changes at the rate the electric field does work.
That the magnetic field does no work is therefore not a separate fact to be remembered. It is a remark about where the magnetic entries sit, since they occupy the purely spatial part of the array and so cannot appear in the equation governing energy. The familiar statement and the arrangement of the object are one thing.
Then a check that had to work, and does. Any four-force must stand perpendicular to the four-velocity, a constraint out of the geometry of the interval with no reference to electricity. The candidate satisfies it because the field tensor reverses sign when its labels are swapped, a property out of the freedom in the potentials with no reference to geometry. Two constraints from unrelated directions, matching exactly. Had the field tensor carried any part that did not reverse sign, charges would gain energy from nothing.
Compress everything so far and one line is left: one number formed by contracting the field tensor with itself, and one term coupling the current to the potential. Hand that line to the machinery of the first part and out comes the equation saying charges and currents make fields, constants and all. The other half needs no varying, holding the instant the field is written in terms of potentials.
How nearly the form is forced is what to carry away. Demand a number every observer computes alike and only contractions are admitted. Demand that the redundancy in the potentials stay invisible and one candidate is killed outright, the one that would have given the force's carrier a mass, which is the whole reason light has none. Demand no more than first derivatives and almost nothing is left. Two terms survive, their constants fixed by matching the static force and by a choice of normalisation.
The method is worth more than the result. List the ingredients, list the symmetries, write down every term they permit and order these by how many derivatives they carry, whereupon the leading one is the theory and the rest are corrections. Let the potential carry an internal label and become a matrix, add the one ingredient non-commuting labels supply, and the same line written again is the strong interaction, whose carriers act on one another because that ingredient says they must.
The oldest debt in the book is settled by computation rather than assertion. Two charges in uniform motion, alone in the universe, were seen in the first part to push on each other unequally, momentum appearing from nowhere. Work out what the empty space between them holds and differentiate: the field loses momentum at exactly the rate the particles gain it, awkward factor of three halves and all.
The diagnosis matters more than the arithmetic. The third law assumed all momentum belongs to particles, handed over the instant either moves. Relativity forbids the second clause: once one charge moves and the other cannot learn of it until light crosses the gap, the momentum must be somewhere meanwhile. What holds it there is no table of where a force would be but a physical system with its own degrees of freedom, which is why a later part must quantise it.
Part II has done its work. You hold a geometry where one speed is the same for everybody with nothing to measure it against, a language making frame-independence visible in an equation's shape, a mechanics of four-entry objects, which is where energy came from, and one field with electricity and magnetism as its slices. One thing has not moved: gravity is still a force reaching across empty space, arriving the moment it is sent, and nothing in this part permits an influence with no delay.
Three chapters ago the interval was a fact about a fabric nobody could move. It is about to become the thing that moves.
What Part II delivered was a geometry with a speed limit written into it rather than into any material, and a way of telling, by looking at the shape of an equation, whether a law was about the world or about a laboratory. What it could not deliver was gravity, and the reason is worth stating precisely rather than as a slogan. Newton's law of gravitation names two masses and the distance between them, and the distance between them at a given moment is exactly the kind of quantity the last part showed to be a fact about who is asking. There is no repair available that keeps a force and fixes the timing, because the trouble is not the speed at which gravity travels but the arithmetic in which the law is written.
What follows takes the one measured coincidence that nobody in three centuries could explain, that a body's reluctance to be pushed and its response to gravity are the same number, and treats it as the whole of the subject. The price is that the arena stops being furniture. Distance becomes a field with equations of its own, the toolkit's warning about comparing an arrow here with an arrow there comes due, and by the end of this part the shape of space and the contents of space are two halves of one equation.
Part III · General Relativity
Geometry becomes dynamical.
3.1 The Equivalence Principle
Rebuilding gravity is not optional, for the reason the previous part left standing, and the rebuilding starts not with a new equation but with an old and very strange piece of arithmetic. It is an equality between two numbers that had no reason to be equal, noticed by the author of the theory it sits in and never explained by it.
Mass enters the older physics twice, in roles unrelated to each other. In the first it is stubbornness, the reluctance of a thing to have its motion changed, measured with springs and collisions where gravity plays no part. In the second it is a kind of charge, the strength with which a thing answers gravitational attraction, measured by hanging it from a balance and reading a deflection. Electric charge shows that such a pairing need not hold, since a proton and an electron carry equal charge while differing in stubbornness by a factor of nearly two thousand. For gravity the two numbers agree to a few parts in a thousand million million.
The consequence is a cancellation, and it is the whole foundation of what follows. Because one number governs both the pull and the resistance, it divides out, and what remains describes the place rather than the object standing in it. Everything built here rests on that measured coincidence, which is why the experiment keeps being repeated with better apparatus.
The same symbol stands on both sides of the equation for a falling body, so it cancels, and it takes the body's identity with it. What is left determines a path from a starting place and a starting velocity alone, which means that a feather and a cannonball released together in the absence of air trace out one curve rather than two. Every other force keeps a residue of the object in its equation of motion, most visibly electricity, where the surviving ratio of charge to stubbornness varies over three orders of magnitude across ordinary particles and reaches zero for a neutron.
That cancellation buys something startling. If the pull is the same everywhere, then relabelling positions with a rule that accelerates along with the fall removes the pull from the equations entirely, and because the rule mentions nothing about the body, one relabelling does this for everything in the room at the same moment. A field that can be abolished by an act of bookkeeping, for all objects simultaneously, is not behaving like a force.
The reverse reading is worth holding onto. Begin with no gravity at all and describe events using coordinates that accelerate, and a uniform field appears out of nothing. Acceleration and uniform gravity are therefore not two phenomena resembling each other; they are one situation described twice, and nothing measurable distinguishes them.
Widening a claim about falling stones into a claim about every experiment anybody could perform inside a sealed box is a genuine leap, and it deserves to be recognised as one. What the extended claim says is that a laboratory in free fall, kept small enough and watched briefly enough, is indistinguishable from a laboratory drifting in empty space far from anything, so that chemistry, optics, radioactive decay and the behaviour of clocks all come out as the previous part said they would.
Three separable assertions hide inside that sentence, and separating them is what makes the principle testable rather than rhetorical. The first is the measured fact already in hand, that what the apparatus is built from makes no difference. The second is that how fast the falling laboratory happens to be moving makes no difference. The third, which carries most of the weight, is that where and when the laboratory sits makes no difference, so that no experiment can reveal the local strength of gravity from inside.
None of this was derived, and saying so is not a weakness in the argument but a statement about where its risk lies. Each clause can fail, each has been looked at hard, and the third is checked by the most direct experiment imaginable, which is holding two identical clocks at different heights and asking whether they keep the same time.
So far the story has been about what falling removes; here is what it leaves behind, and it is the only part of gravity really there. Take two neighbouring specks of dust, released and left alone, and ask how they move relative to each other rather than to a frame. Subtract their equations of motion and the pull cancels, being nearly the same at both places; what survives is the rate at which the pull changes from place to place, times the gap.
Work that out for one attracting body and the picture is worth having. Two specks one above the other drift apart, the lower being nearer and falling harder; two side by side drift together, both falling toward one centre along converging lines. A ring of dust stretches along the pull, narrows across it, becoming an ellipse. This is the tide, and the ocean does it twice a day for the same reason.
The stretching and squeezing balance, and that is no coincidence. A small ball of dust changes shape while holding its volume fixed, and that sentence is the law of gravity in empty space in plain clothes. It is not the phase-space volume theorem in new clothes: that one concerned possible states, this one concerns actual dust, and each is a flow with nothing made or lost. No coordinates remove this relative drift, so whatever gravity ultimately is, this is it.
The word local has been carrying a great deal of weight without being asked to pay for it, and it can now be handed a number. Free fall abolishes gravity at one place and one moment exactly; move away from that place, or wait, and the drift computed a moment ago begins to show. Requiring the drift to stay under whatever the instruments in the room can resolve produces an inequality, and the inequality constrains the size of the room multiplied by the square of the time spent in it.
That the constraint falls on a product rather than on a length is the useful part. A falling laboratory is not a small box; it is a small patch of space and time together, and one may be traded for the other. A metre-wide bench near the Earth's surface behaves as though gravity had been switched off for about a second if you can measure to a thousandth of a millimetre, and for thirteen nanoseconds if you are working at the precision of a gravitational-wave detector.
Underneath the numbers sits a structural point about the shape the eventual theory must take. The physics of the previous part is exact only in the limit of a vanishing region, so what replaces Newtonian gravity has to hold in the small and be stitched together across large regions rather than written down globally at a stroke.
Two clocks that never move relative to each other, one on a shelf and one on the floor beneath it, do not agree, and the argument reaching that conclusion uses no general relativity at all. Put a windowless cabin in empty space and accelerate it; send a flash from the floor to the ceiling; the ceiling has picked up speed away from the flash during the crossing, so the light arrives reddened by the ordinary Doppler effect of the previous part. Then apply the principle that a steadily accelerating cabin and a cabin standing still in gravity are one situation described twice.
Reading the result as light losing energy on the climb is the weaker option. The stronger reading counts wave crests. Nothing about the arrangement changes with time, so crests are neither created nor destroyed between floor and ceiling, and exactly as many arrive each second as leave each second. If the receiver nonetheless counts fewer of them in each of its own seconds, the only possibility left is that its seconds and the emitter's seconds are of different lengths.
That is a serious problem for the geometry inherited from the previous part, in which the relation between clock readings and coordinate time is fixed once and holds everywhere. One further caution will matter almost immediately: this argument constrains the timekeeping part of the geometry and says nothing whatever about distances.
Honesty is cheaper now than it would be later, so here is the result together with what is wrong with it. Turn the accelerating cabin on its side and send a flash across it. In the frame where nobody accelerates the flash goes straight, but the cabin rises while the flash is in transit, so anyone inside sees the beam bend downward. Light falls. Adding that small bending up along a ray skimming past the Sun gives a deflection of about nine tenths of an arcsecond.
The measured value is one and three quarter arcseconds, so the calculation is short by a factor of two, and the factor is exact rather than approximate. Nothing has gone wrong arithmetically. What has gone wrong is that every step of the cabin argument concerned durations: how long the crossing takes, how much speed the cabin picks up meanwhile. It therefore uses only the part of the eventual geometry that governs clocks, and no information whatever about how distances are measured near a heavy body.
Light spends its budget evenly between space and time, since it moves at the limiting speed, so the distance part contributes exactly as much as the time part and doubles the answer. A planet crawling along at a millionth of that speed barely samples the distance part, which is why the old theory works so well for everything except light.
The case for treating gravity as the shape of the arena rather than as a force within it rests on a single experimental fact and collapses without it. Because everything falls the same way, exactly one path leaves each event in each direction, and that path can be described without naming what travels along it. A rule of that kind is what a geometry hands you free of charge, as the earlier chapter on stationary action showed when it produced straight lines on a flat sheet and great circles on a globe from nothing but a recipe for measuring length.
Electricity fails the same test, and loudly rather than marginally. Release an electron, a proton, a helium nucleus and a neutron together in one electric field and they trace four different curves, one of them straight. There is no single family of paths belonging to the region, so there is nothing for a geometry to be, and the electric force needs a quite different device that a later part supplies.
Two closing observations point forward. The paths depend on how fast you were going when you set off, so the geometry must live on space and time together rather than on space alone. And the rule for the length of a path was already shown to vary from place to place by the business with the clocks.
3.2 Manifolds
Up to now every arena in this book has been a flat grid you could lay down once and use everywhere, and two separate things now make that untenable. The first is a matter of shape. Nothing anybody measures is a statement about the universe as a whole, since observations happen in bounded regions over bounded stretches of time, so whether space and time together form an endless flat sheet is to be settled rather than assumed. The remedy is to build the arena out of overlapping patches, each looking like the familiar grid, with careful books kept on how neighbours agree.
The second problem is worse and has nothing to do with shape. Picture a small arrow lying flat against the surface of a globe. That picture is drawn in the room the globe is sitting in, and the room is not part of the globe. Spacetime has no room around it and nothing outside from which it could be viewed, and every measurement anyone will ever make is made from within. Assuming an outside means assuming structure no experiment constrains.
There is a quieter version of the same complaint. An arrow from one place to another is a subtraction of positions, and subtracting positions requires an arena whose points can be added and taken away from each other. A curved surface is not that kind of thing.
Covering everything at once with a single description is a luxury, and a globe is the cheapest demonstration that the luxury is often unavailable. Two overlapping charts will do instead, provided there is an explicit dictionary translating the coordinates of one into those of the other wherever both apply. That dictionary is the only thing in the whole construction that ever gets differentiated, which is why the apparatus is lighter than it first appears.
That one chart genuinely cannot suffice deserves a proof rather than an appeal to intuition, and the proof is four lines long. A globe is closed and bounded, so any continuous quantity spread over it has to reach a largest value somewhere. A flat open region has no largest anything, since around each of its points there is always room to move a little further in every direction. A faithful correspondence between the two would have to match a place where a quantity peaks against a place where nothing peaks, and no such correspondence exists.
The construction is also notable for what it lacks. Nothing so far measures a distance or an angle or picks out a straightest path, and this poverty is deliberate rather than an oversight. Distance is going to become the dynamical variable of the theory, the thing that responds to matter and changes with time, so it cannot be welded into the arena that carries it.
The rule for translating between neighbouring patches is what holds a patchwork together, and it is the only object in the construction with any calculus in it. The underlying space is a bare collection of points with nothing there to differentiate, the coordinates assigned to those points are labels, and the one genuinely mathematical object is the dictionary saying how one labelling turns into another. Requiring that dictionary to be infinitely differentiable is the whole of what the word smooth means here.
The requirement earns its keep immediately. Call a quantity spread over the space smooth if it looks smooth when written in the labels of some particular patch, and an obligation appears at once, since somebody using different labels must reach the same verdict or the word describes the bookkeeping rather than the quantity. Writing the second description as the first composed with the dictionary, and remembering that a composition of well-behaved maps is well behaved, settles the matter in one line.
This is the same move the book keeps making in different clothes. Ask what survives a change of description; whatever survives is a fact about the thing, and whatever does not was a fact about the description. Here it is applied to the mildest property imaginable, and having secured it we may talk about smooth quantities on a curved space without naming a preferred way of labelling it.
Arrows need somewhere to live, and the somewhere has been quietly assumed all along. Take it away and the question becomes what a direction was ever for, and the answer is that a direction gets handed to something which reports a rate of change. Walk that way and the temperature rises; walk this way and it falls. Measurable quantities spread over the arena are functions on it, so a direction is a device taking any such quantity and returning the rate at which it changes here. Nothing in that sentence points outside.
The construction is then almost forced. A path through the point, composed with a quantity defined on the space, gives a function of one variable, which the first chapter taught us to differentiate. Doing that in coordinates splits the answer into a part belonging to the path and a part belonging to the quantity, and the first is exactly the list of numbers the discarded arrow would have carried. What was lost was the picture, and the picture was the one piece that needed an outside.
Two demands survive and become the definition: the device must be additive, and it must handle products the way differentiation does. What makes this worth the discomfort is a theorem rather than a hope. Such devices number exactly as many as the space has dimensions, so the replacement is the same size as the thing replaced.
From here on, each point of the space carries its own private stock of directions, and no two stocks are the same stock. A direction at one place is a device reporting rates of change there, and a direction elsewhere reports rates of change elsewhere, so the two are not members of a common collection and cannot be added or subtracted. That is the most consequential thing said so far, and it deserves to be sat with rather than nodded through.
The reason it feels wrong is instructive. On a flat sheet we slide arrows about freely, and that habit works because a flat sheet is secretly two things at once, a place and also a system in which positions can be added and taken away from one another. The second provides a rule for saying an arrow here matches an arrow there, and square coordinates hide the fact that a rule was ever needed. A curved surface offers no such rule, and carrying an arrow across a globe by different routes delivers different verdicts.
This is why differentiating fails. Finding a rate of change means subtracting a value here from a value a short step away, and if the two values belong to separate systems the subtraction is not hard but meaningless. Repairing it is the whole business of the next chapter, and the leftover from the repair is going to be gravity.
The measuring devices from the chapter on tensors reappear here, and this time they arrive with a definition rather than a picture. Alongside the directions at a point sit the machines that eat a direction and return a number, and those machines form a space of their own with as many dimensions. Any measurable quantity spread over the arena produces one automatically: hand it a direction and it reports how fast the quantity changes that way.
Feeding that construction the coordinate labels themselves produces exactly the objects physicists have been writing under integral signs since school and cancelling without justification. They turn out to be the machines paired with the coordinate directions, and the pairing of one with another gives one when the labels match and zero when they do not, which is the condition defining a dual basis. What was an abuse of notation four parts ago is now an exact statement about honest objects.
One distinction has to be kept sharp, because a later chapter will charge for confusing it. The machine built from a quantity is not the gradient of that quantity. Turning it into a genuine direction requires something that converts between the two species, and no such converter exists yet. That converter is the metric, it is a physical field rather than part of the furniture, and the whole of the next chapter is about it.
Nothing about the definition of a tensor had to be replaced, and that is the reward for having defined it by behaviour in the first place rather than by appearance. An object is a tensor if its description responds to a change of labels in one specific controlled way, and the definition never asked what the object looked like, whether the labels were straight, or whether the arena was flat. Every consequence drawn from it in the earlier chapter was drawn at a single point, so every consequence still holds.
Exactly one thing has changed. The matrix relating one labelling to another used to be the same everywhere, because the changes of description allowed in the previous part were rigid ones. Now it varies from place to place, since any smooth relabelling is permitted. At any single point that makes no difference whatever, because a matrix is a matrix. It makes a difference only where a calculation reaches across from one point to a neighbouring one.
There is precisely one such calculation in the whole earlier chapter, and it is taking a derivative. That is the entire damage, and locating it so narrowly is what keeps this part of the book to four mathematical chapters instead of forty. The next chapter repairs that one operation, and what has to be added in order to repair it turns out to be the gravitational field itself.
Assign a direction to every point rather than to one point and a question becomes askable that was not askable before. Follow one field briefly and then the other, then do it in the opposite order, and ask whether you end up in the same place. For directions belonging to a coordinate grid the answer is yes by construction, since going east then north and north then east is what having a grid means. For two arbitrary fields the answer is generally no, and the gap is itself a field of directions.
Getting there needs care. Applying one field and then another measures the bending of a quantity rather than its slope, which is too much machinery for a direction to carry. Doing it in both orders and subtracting cancels the excess, because the excess does not notice which field came first.
The arena is now complete, and it is worth being clear about what it still cannot do. There is no way to measure a length, no way to say two paths meet at a right angle, and above all no way to compare a direction here with a direction there. That last gap blocks differentiation outright, since a rate of change is a comparison between neighbouring places. Supplying the missing comparison is the whole content of the next chapter, and the thing that has to be supplied turns out to be gravity.
3.3 Metric and Connection
The arena handed over at the end of the last chapter was deliberately unfurnished. Nothing in it could measure a distance, nothing could say whether two directions stood at right angles, and nothing could compare a direction here with one over there. The first two gaps close with one object installed at every point, taking two directions and returning a number. Out of that number come lengths, angles, and the sorting of directions into timelike, spacelike and lightlike that the previous part built its geometry on.
Two features of the installation matter. Because it happens point by point, the numbers describing it are functions of position rather than constants, and that dependence is where gravity will eventually live. Because the thing installed eats directions and returns numbers, it belongs to the species of the measuring machines of the tensor chapter, and responds to relabelling in the one controlled way that species does.
An old debt is settled almost in passing, and it is the charge the previous chapter warned would fall due. Ever since the multivariable chapter there has been a distinction between the object a function supplies for free, reporting how fast it changes in a direction, and the arrow of steepest ascent, undefinable until length exists. On the bare arena only the first existed. A length having been supplied, the second exists too, depending on the geometry as much as on the function.
Here is a habit of thought that has to be broken now rather than later. Write down the rule for measuring distances on a flat page using ordinary square coordinates and it looks utterly plain, with the same two numbers everywhere. Describe the very same page using distance-from-the-centre and angle-around-the-centre instead, and the rule acquires a factor that grows with distance and never stops growing. Nothing was done to the paper. Only the labelling changed.
The factor that appeared reports something real but modest, namely that a step of one unit in the angle coordinate carries you further when you are further out. Coordinate directions in such a description have lengths depending on where you stand, and the rule for measuring distances has to say so. That is bookkeeping about labels rather than a fact about the page.
The reason this matters so much is a coincidence of appearance. The rule for distances on a globe has exactly the same shape as the rule for the flat page in circular labels: one plain term and one term carrying a position-dependent factor. One of those surfaces is curved and one is not, and no amount of staring at the two rules will tell you which is which. A real test has to be built from scratch, and building it is the business of the chapter after this one.
The quantity a wristwatch reads was defined in the previous part using the fixed geometry of empty spacetime, and the definition survives the move to a variable geometry with no change of wording, because it only ever referred to the rule for measuring separations. Feed it the new rule and it computes elapsed time along any history whatever. For somebody sitting still, only one entry in the rule survives, and their watch differs from the coordinate clock by the square root of that single entry.
That last sentence is a promise being collected. The equivalence-principle chapter showed, using only an accelerating cabin and the ordinary Doppler effect, that two stationary clocks at different heights tick at different rates. It then wrote down what the timekeeping entry of the geometry would have to be, and left the notation for later. The notation has arrived, and the conclusion stands unchanged.
The number attached to each possible history is likewise inherited without alteration: elapsed watch time, multiplied by the mass and the square of the speed of light, with a minus sign. The previous part cornered it into that form by demanding that everyone compute the same value. What is startling is how little there is. No force appears and no potential energy appears, and the mass factors out and drops from the equations of motion, which is the universality of falling that started this part.
Differentiating a field means comparing its values at neighbouring points, and the previous chapter established that on a curved arena such a comparison is unavailable. One escape presents itself: work in a set of labels, treat the components as ordinary functions of those labels, and differentiate those. The result is a perfectly good array of numbers. The trouble is that the array describes the labelling as much as the field.
Following the calculation through shows precisely where the damage enters. Relating one labelling to another involves a matrix of rates of change, and taking a derivative makes that matrix get differentiated too, which produces a surplus term that a genuine measuring device would never carry. In the previous part the matrix was the same everywhere, so its derivative vanished and the problem never appeared; the moment arbitrary smooth relabellings are permitted, it does.
The failure is worth seeing with numbers rather than symbols, and the flat plane supplies them. Take a field of identical arrows all pointing the same way, which is as unchanging as a field can be. Describe it in circular labels and the components acquire a plain dependence on position, since the arrow points along the radius here and across it there. Differentiate those components and you get something non-zero, from a field that does not change. The array is reporting the turning of the labels, and nothing else.
Rather than inventing a repair and hoping it works, the move is to write down the most general thing that could repair the damage and let the requirement of good behaviour decide what it is. The damage was an unwanted extra term proportional to the field itself, so the repair is an added term proportional to the field itself, carrying unknown coefficients. Call them the comparison coefficients, since their whole business is comparison at a distance; insisting that the sum behave properly under relabelling determines exactly how they must respond to relabelling, and nothing is left to choose.
What emerges has a property that reliably surprises people. They do not respond to relabelling the way an honest measuring device does. They pick up an extra additive piece, so they can be zero in one description of a situation and non-zero in another. That is not a flaw but the entire point, since their assignment is to cancel a quantity that misbehaves in exactly that way, and only something equally badly behaved can do the cancelling.
Anyone who has followed this part should feel a jolt of recognition. A quantity that can be made to vanish everywhere by choosing your description well, but which no choice removes once you look at a large enough region, is precisely how the first chapter of this part described the gravitational field. The comparison coefficients are that field.
Setting the repaired derivative to zero along a path says the field is not changing as you walk that path, and that single sentence supplies the missing comparison. Plant a direction at one end, insist that it stay as unchanged as the arena permits at each step, and you arrive at the other end with a definite answer. The rule is an ordinary equation of the kind the differential-equations chapter solved, so a starting value fixes the result all along the route.
Notice the three words that had to be included. The answer is well defined given a path. Two routes between the same pair of points involve different equations, since the comparison coefficients are evaluated along different tracks, and their answers have no reason to agree. On a globe they visibly do not, and the previous chapter showed as much before any machinery existed to describe it. The disagreement is not a defect and no better construction removes it.
That disagreement is the whole subject of the next chapter. Make the route a small closed loop, carry a direction round it, and compare what returns with what set out. The mismatch shrinks as the loop shrinks, and the rate at which it shrinks belongs to the point rather than to the loop. That quantity is curvature, and by the end of the next chapter it will be the same thing as the tide.
Up to this point the comparison coefficients were almost entirely free. Any set of them behaving correctly under relabelling would do, and there were infinitely many. Two requirements cut that down to exactly one, and neither is a technicality. The first is that carrying a direction along a path must not change its length or its angle to a companion, since a rule that quietly stretched things would be measuring nothing. The second is that two derivatives of an ordinary quantity taken in either order must agree, a property the arena was built to have.
Counting first tells you what to expect. In four dimensions the unknown coefficients number forty and the first requirement supplies forty equations, so a single answer is what should come out. The derivation obliges, by a trick worth naming: write the requirement three times with its indices rotated, add two copies and subtract the third, and four of the six unwanted terms annihilate in pairs because the comparison coefficients are symmetric and so is the measuring device.
What survives is a formula giving the comparison coefficients entirely in terms of the rule for measuring lengths and how it varies from place to place. That is a strong statement about the world rather than about the algebra. Fix how distances are measured and you have, without any further choice, fixed how directions at neighbouring points are matched up.
Two different sentences describe a straight line, and on a flat page they agree so quietly that nobody notices there were two. One says a straight line never turns: keep going in the direction you are already going. The other says it is the shortest route between its ends. Each survives the move to a curved arena, each becomes an equation, and the two equations are derived here separately with almost no shared assumptions.
The first derivation says that the direction of travel, carried forward by the comparison coefficients, is still the direction of travel. The second forgets them entirely, takes the number attached to each history from the previous part, and asks which history makes it stationary, using machinery built for pendulums four parts ago. Both give the same equation with the same coefficients, which is no accident: the rule preserving lengths is the one the length functional generates.
Then the test that matters. Feed in a weak, unchanging field, allow only slow motion, and use the timekeeping entry the equivalence-principle chapter obtained from an accelerating cabin and nothing else. Out comes the law of falling bodies, with the familiar acceleration appearing not as a force but as a coefficient of the geometry. Gravity has stopped pushing and become a fact about the shape of the arena. Still entirely missing is any rule saying which shape a given lump of matter produces.
3.4 Curvature
Every intuitive account of curvature smuggles in a viewpoint from outside — a rubber surface seen to sag, a ball seen to bulge, a direction sticking into a surrounding room. None is available here, since there is no room and no outside, and the difficulty of this chapter is finding a test a creature confined to the surface could run. One instrument survives the restriction, and the previous chapter finished building it: carry a direction around a closed circuit and see whether it comes home pointing the way it left.
Everything about that procedure stays inside. The circuit lies on the surface, the carrying rule refers only to the surface, and the final comparison is between two directions at one and the same place, which was always the one permitted comparison. Since carrying preserves lengths, the only thing that can have happened is a turn, whose size anyone on the surface can measure.
Running the test on a globe with pencil and paper gives an answer of startling neatness. The turn depends on the circuit only through the area it surrounds, not on its shape or position. So the quantity belonging to the surface rather than to any circuit is the turn divided by the area, and that ratio survives as the circuit shrinks to a point. It is the reciprocal of the radius squared, and this chapter turns it into a proper object.
Shrinking the circuit test to a point converts it into a question about the order of two operations. Going a short way in one direction and then a short way in another, against the same two moves in the opposite order, traces out a small closed circuit, so the failure of the two orders to agree is exactly what the circuit measures. In the language now available, that failure is the difference between differentiating one way then the other and the other way then the first.
Computing it is the longest piece of algebra in the book so far, and it is worth knowing what to look for. Six terms appear. One is an ordinary second derivative and disappears because the order of those never matters. Two more form a pair that swaps into itself when the directions are exchanged, so the pair cancels against its mirror image. A fourth goes because those coefficients were required to be symmetric in their lower slots.
What matters is not the algebra but what survives it. Every term carrying a derivative of the transported field has vanished, and what remains is proportional to the field itself. An operation that looked certain to produce a differential operator has produced plain multiplication by an array of numbers, and that array depends on the geometry alone. It is the object this chapter is about.
Length in a derivation should buy something, and the long computation just finished bought this. The comparison coefficients of the last chapter were notoriously not honest measuring devices: a good choice of description makes them vanish everywhere and a bad one brings them back. Anything built naively from them inherits that disease. The array produced by the two-order calculation does not, because the two orders differ by an antisymmetry while the diseased part of the coefficients is symmetric in exactly the slots being antisymmetrised.
There is a short and airtight way to see it that avoids the algebra altogether. The thing computed was a difference of two honest objects, so it is honest. It turned out proportional to the carried field itself, with no derivatives of that field anywhere. Something that turns any field into an honest object by plain multiplication must itself be honest, and that is a theorem from the tensor chapter rather than a hopeful remark.
The flat page in circular labels makes the cancellation visible with actual numbers. There the coefficients are non-zero and consist of nothing but chart-dependent junk, since they vanish in square labels. Compute the new array from them and the derivative piece and the quadratic piece come out equal and opposite in every component, leaving zero. That is the third and final form of the previous chapter's warning, now a calculation rather than a promise.
Take two specks of dust released near each other, each moving as freely as anything can move, each following the straightest path its surroundings allow, with nothing pushing either. Ask how their separation changes. The answer, derived here from geometry alone with no mention of gravity or mass or force, is that the separation accelerates at a rate proportional to itself, the constant supplied by the array the previous sections built out of the circuit test.
Now set that beside a result obtained at the start of this part from Newton alone. Two specks in free fall were shown there to drift apart or together at a rate proportional to their separation, the constant supplied by how the pull varies with place. The two statements are one statement. Where the older physics wrote the way the pull varies, the new geometry writes the amount by which a direction fails to come home unchanged, and computing the second from the first confirms it rather than suggesting it.
So the thing you feel as gravity when the tide comes in, and the thing curvature measures, are one object. The ocean rising, two falling specks separating, and a direction coming back turned after a trip around a small loop are three descriptions of one feature of the world. That sentence is what this whole part has been walking toward, and everything after it is bookkeeping by comparison.
An array with four slots in four dimensions carries two hundred and fifty-six numbers, almost all duplicates. Three properties cut the list down, and each is derived rather than declared. Swapping the last two slots flips the sign, which was built into the construction. Swapping the first two also flips the sign, and that traces back to the demand that carrying a direction preserves its length. A third relation ties together the three ways of cycling the last three slots.
Counting then goes in three stages. A pair of slots that flips sign under exchange behaves like a single label taking six values, so the array is a six-by-six table. A further symmetry makes that table symmetric. The cyclic relation removes one entry more, since what it says is that the completely antisymmetric part vanishes, and in four dimensions there is only one such part. Twenty survive.
The same twenty appear again by a route with nothing in common with the first. Ask how much of the geometry a well-chosen observer can make disappear near a chosen event. The rule for measuring distances can be made to look flat there, and its rate of change made to vanish, both by explicit construction. The rates of change of those rates cannot: the knobs fall short by exactly twenty. Those twenty are what no viewpoint removes, which is what the first chapter said about the tide.
Twenty numbers at every point is more than a field equation can carry, and there is a standard way of boiling an array down: sum one slot against another. Doing it here is more constrained than it looks, since two of the three summations give either nothing or a copy of the third, for reasons traceable to the antisymmetries. So there is one way, and it gives a symmetric array of ten numbers, which boils once more into a single number everybody agrees on.
The array has fallen apart into independent pieces again, as the toolkit kept promising it would, and the pieces are not equally interesting. Ten numbers went into the pot and ten did not, and the ones left out have every trace removed. They carry the effects surviving where there is no matter, so the omission matters. Outside any body the summed-down array vanishes while the full array does not, and that gap is where orbits, tides and gravitational waves live.
That distinction was visible long before the machinery existed. The first chapter of this part found that a small ball of falling dust changes shape while holding its volume, and identified the volume statement with the law of gravity in empty space. The volume part is what gets summed down; the shape-changing part survives the summation. Empty space is gravitationally empty in the first sense and thoroughly occupied in the second.
Some results earn their place by what they forbid rather than by what they produce, and this is one of them. Differentiating the curvature array and adding up three cyclic arrangements gives exactly zero, always, for every geometry, as an identity rather than as a condition. Proving it is a matter of standing at a point in the frame where the comparison coefficients vanish, at which the curvature is a plain difference of derivatives, and watching six second derivatives cancel in pairs because the order of ordinary differentiation never matters.
Summing that identity down twice turns it into a statement about the boiled-down array: a particular combination of it and the single curvature number has no divergence whatever. Now recall what the previous part established about matter. Energy and momentum are packaged in one symmetric object whose divergence vanishes, and that vanishing is the local statement that nothing is created or destroyed.
Setting those two facts side by side almost writes the law of gravity by itself. If geometry is to be equated with matter, then whatever stands on the geometry side must have vanishing divergence automatically, for every geometry, or else the equation would quietly impose an extra demand on matter that nothing justifies. Exactly one combination qualifies, and three chapters from now it will be sitting on the left of the field equations, put there not by taste but by this identity.
A question raised earlier and left hanging can now be settled. A space is genuinely flat, in the sense that one unbending grid covers all of it, exactly when the array built out of the circuit test vanishes everywhere. One direction is easy: lay down such a grid and the rule for distances has constant entries, so the comparison coefficients vanish and the array vanishes, and since the array is an honest measuring device its vanishing in one description means its vanishing in all.
The other direction is harder and is quoted rather than derived, with the reason for believing it sketched. If the array vanishes then carrying a direction around any circuit that can be shrunk away brings it back unchanged, so a direction chosen at one place can be carried unambiguously to every other place, which is the raw material a flat grid is made of.
One qualification in that sentence does real work rather than decorating the statement. Circuits that cannot be shrunk are exempt, and spaces containing them can be flat everywhere while refusing to be a plain grid. A cone rolled from paper is flat away from its tip, having been made without stretching anything, and yet a circle around the tip has the wrong circumference. Local information does not always add up to global information, and the last chapter of this part turns that on the universe.
3.5 Forms, Lie Derivatives, Killing Vectors
An integral over a patch of surface ought to depend on the patch and on nothing else, and in particular not on the grid of labels somebody painted on it. Relabelling stretches the cells of the grid by a factor the earlier parts identified as a determinant, and it stretches the two edge directions feeding the integrand as well; for the two effects to cancel, the integrand must respond to its slots exactly as a determinant responds to two columns, changing sign whenever they are swapped.
So antisymmetry is not a stylistic preference or a tidy formalism. It is what survives the demand that an answer be about a region rather than about a description of one, which is the demand governing this book since its first chapter on coordinates. The objects obeying it are called forms, they carry one slot per dimension of the thing they are integrated over, and the product building larger ones out of smaller inherits its sign rule from the same counting.
One arithmetic fact will be wanted shortly. The number of independent entries such an object can have in four dimensions runs one, four, six, four, one, and then stops dead. The six is the count of electric and magnetic components together, which is no coincidence, and the stopping is why there is a largest possible integral rather than an endless tower.
Two chapters ago, differentiating a field turned out to be impossible without extra equipment, because the values being compared live at different points and nothing brings one to the other for free. Building that equipment took a whole chapter. Here a cheaper route opens, for the special class of objects the previous section forced on us, and the reason it opens is worth keeping.
Take the naive derivative, which fails because it drags along a term built from the comparison coefficients, and antisymmetrise the result over the differentiating slot and one other. The offending term carries those two slots in a pair of positions where the comparison coefficients are known to be symmetric, and antisymmetrising anything symmetric annihilates it. The failure cancels rather than being repaired, so this derivative can be computed with no comparison rule at all, and it returns the same answer whichever rule you would have chosen.
The payoff is that the three operations the toolkit taught separately, the one measuring steepness, the one measuring swirl and the one measuring outflow, stop being three operations. They are one operation applied to objects carrying one, two and three slots, and the reason they looked different is that vectors in three dimensions can stand in for all three types. That coincidence is over, and what replaces it survives in any number of dimensions and on any curved space.
Apply the new derivative twice and the answer is always nothing. The proof takes three lines: the two derivative slots sit inside a construction that changes sign when any two slots are swapped, while the derivatives themselves do not care about their order, so the expression equals minus itself and therefore equals zero.
What that line collects is the reason the section exists. The toolkit proved, by grinding through components, that a field of steepest ascent has no swirl and that a field of swirl has no outflow, and remarked that the two proofs looked alike. The chapter on electromagnetism found that two of the four famous equations follow from the existence of potentials and say nothing about how electricity behaves. All three are one statement, and the six terms cancelling in each case are literally the same six terms.
And the statement is not really about differentiation. The edge of a filled disc is a closed loop, and a closed loop has no ends; the skin of a ball is a sphere, and a sphere has no rim. Taking a boundary twice leaves nothing. Because a later theorem pairs the derivative with the boundary, that triviality about shapes must show up as a triviality about derivatives, and it does. What looked like a run of coincidences is one plain fact met three times in the wrong order.
Having shown that anything which is a derivative gives nothing when differentiated again, the natural next question is the converse: if differentiating something gives nothing, did it come from a derivative. The answer is yes on any region that can be shrunk to a point without leaving it, and the proof is constructive rather than abstract. Integrate the object outward along straight rays from the centre, and the resulting function has exactly the original as its derivative, with the hypothesis spent at precisely one step.
On a region with a hole the answer is no, and the failure is measurable. The toolkit's old example, the field circulating around a missing origin, differentiates to nothing everywhere it is defined and yet accumulates a full turn around any loop encircling the gap. Nothing is wrong with it. What the loop integral measures is the hole, using nothing but calculus, which is a strange and useful thing for calculus to be able to do.
Two debts are settled here. The toolkit quoted the theorem guaranteeing that a field with no outflow is the swirl of something else, promised a proof in this chapter, and now has one, complete with a formula for the potential. And the same local-versus-global gap that left a rolled paper cone flat everywhere and still not a plane turns out to be this gap rather than a relative of it.
Very early in the book, three results with three names and three right-hand rules were lined up in a table beside the fundamental theorem itself and shown to have identical shape: whatever a derivative accumulates throughout a region is bookkept entirely on that region's skin, whether the region is an interval with two ends, a curve, a patch of surface with a rim, or a solid with a hull. The table came with a promise that the four rows would become one line once a language existed for curved spaces of any dimension. That language now exists.
The proof is smaller than its reputation. On a single cube, integrating one variable at a time lets the ordinary Fundamental Theorem of Calculus, the first result of the toolkit, convert each integral into a difference of endpoint values; keeping track of which face runs which way is the remaining work, and the sign conventions built earlier do it automatically. Then chop any region into cubes and add. Every internal wall is counted twice with opposite orientation and cancels exactly.
One honest improvement over the earlier treatment deserves noting. There the cancellation argument came with an admission that it was a picture rather than a proof, because the statement inside each cell was approximate. Here it is exact, so adding many introduces no error, and the earlier chapter's own prediction of the repair is what happens.
An integral needs to know how much region it is adding over, and on a curved space with arbitrary labels the naive answer is wrong by a factor that changes from place to place. The correction is a single number: the square root of the determinant of the rule for measuring distances. Relabelling squeezes the boxes of the coordinate grid by a determinant and stretches the distance rule by the square of the same determinant, so the square root of the second exactly undoes the first.
Two old results are collected in passing. The extra radial weight in polar coordinates, and the more elaborate weight in spherical coordinates, were obtained much earlier by dissecting regions by hand. Both are this one determinant, and so is the area element used to measure the patch enclosed by the circuit in the previous chapter.
The section ends by building the equipment the next chapter cannot do without. When the repaired derivative is applied to a field and the result summed over its slot, all the comparison coefficients cancel and what remains is an ordinary derivative times the volume factor. Combined with the previous section's theorem, any term of that shape inside an integral is bookkept on the boundary, and vanishes whenever the thing being varied is held fixed there. The next chapter's central computation lives or dies on that sentence.
Comparing a field's values at two different points has been forbidden since the chapter on manifolds, and the chapters since bought their way round the ban by installing a rule for carrying things from place to place. Here is a second way to buy it, and the currency is different. Choose a direction field, let every point drift along its own arrow for a moment, and use that drift as transport. Carry the field back from where it drifted, and compare with what was already there.
Two features distinguish this from the earlier construction. It needs no rule for comparing directions and no notion of distance, so it is available on a bare space before any geometry is laid down. And it depends on the chosen field everywhere nearby rather than merely at the point, which is why the answer compares the field with a dragged copy of itself rather than giving a rate of change along a line.
For two direction fields the result is an object built several chapters ago by a different argument, namely the failure of two successive drifts to commute. That agreement is worth more than its algebra, since the bracket measuring whether two operations commute also measures whether one field is carried into itself by the other's flow. The next section drags the distance rule itself, and asks when dragging leaves it alone.
Drag the rule for measuring distances along a chosen field, demand that the result be nothing, and you have what this book means by a symmetry of a space. Written out and cleaned up using the fact that the repaired derivative ignores the distance rule, the demand becomes one compact equation on the chosen field. It is badly overdetermined, ten conditions on four functions, which is why a space picked at random has no symmetries at all, and why the ones that do are the ones anybody can solve.
Flat spacetime turns out to have exactly ten independent solutions, and they are the four shifts, the three turns and the three boosts of the previous part. That closes a loop: the transformations obtained there by insisting the interval be preserved reappear as solutions of a differential equation, with nothing about relativity fed in.
The globe supplies the cautionary case. Turning it about its axis leaves every measurable distance as it was, and so does turning it about either of the other two axes, though those are invisible in the usual grid of latitude and longitude. Sliding everything southward is not a symmetry, and the calculation says precisely how it fails: north-south spacings are untouched while east-west spacings stretch, most sharply halfway to the pole and not at all at the equator. The picture measures that rate and returns the predicted number.
Take any direction along which the geometry does not change, and any freely falling body. The component of the body's velocity along that direction, measured with the geometry's own rule for taking components, never changes. The proof runs to four lines and each input is spent once: the free-fall condition removes one term and the no-change condition the other, and neither could have removed the term the other did.
Then comes the identification making this more than a trick. Feed the same situation to the theorem proved long ago for mechanical systems, which says every continuous symmetry of the quantity attached to a history supplies something conserved. The change in that quantity under sliding the coordinates along the chosen direction turns out to be the dragging operation of two sections ago, applied to the distance rule and contracted twice with the velocity. So the geometric and mechanical conditions are one condition, and the conserved quantity produced is the same quantity.
This is where the chapter earns its place. Two chapters from now the geometry outside a star will mention neither time nor the angle around the axis, which by the practical test hands over two constants at once. Those constants are the energy and the angular momentum, and having them reduces a tangle of four coupled equations to one equation in one variable, of exactly the kind the toolkit taught how to read.
As a test of the new language, classical electromagnetism is rewritten in it and comes to three short statements: the field is the derivative of the potential, the derivative of the field is nothing, and the derivative of the field's complement is the current. The compression is pleasant but not the point. What it exposes is which parts of the subject depend on geometry and which do not.
Only one operation in the three requires knowing distances and angles, namely the one swapping a description in terms of two slots for one in terms of the complementary two. That operation appears in exactly one of the three, the one with a source in it. So the half of Maxwell's equations the previous part identified as empty bookkeeping is also the half never touching the geometry, and the half carrying the physics is the half that does. The earlier chapter noticed the first fact and predicted the second; here they are the same fact.
Two things then come for free. Applying the derivative twice to the sourced statement gives nothing on the left, forcing the current to be conserved, so charge conservation follows from the field equation rather than being an extra law. And adding the derivative of any function to the potential changes no field, which is the freedom called gauge, arriving from the same line. Both survive intact on a curved space.
3.6 The Einstein Field Equations
Mass cannot be what gravity responds to, and the reason is not subtle once the previous part is in hand. Mass is not additive, is not separately conserved, and picking out energy alone would mean picking out one entry of a four-part object, so two observers in relative motion would compute different geometries for the same physical arrangement. What is needed is an object packaging energy, momentum and the flow of both, and the earlier part built exactly that when it asked what is conserved because the laws are the same here as there.
That object has ten independent entries and every one of them sources gravity. Energy density is only the corner entry. The diagonal spatial entries are pressure, the off-diagonal ones are shear and momentum flow, and all are on the same footing. The statement that the whole thing has no divergence is the local claim that energy and momentum are never created or destroyed, only moved, and it carries over to a curved arena unchanged because it contains only one derivative and so leaves nothing to be ambiguous about.
One consequence deserves flagging before it arrives. Pressure gravitates. Squeeze a gas without adding any material to it and the source term grows, which nothing in Newtonian gravity would lead anyone to expect and which matters enormously for dying stars and for the expansion history of the universe.
Before anything is designed there is a specification, and this one is short enough to hold in the head. The law must relate objects of the same kind on both sides, since otherwise it would hold only in some descriptions and be a statement about the description rather than the world. The matter side is a symmetric array with two slots, so the geometry side must be one too, which alone eliminates most of what the previous chapters built. The law must involve the shape of space and its first two rates of change and nothing beyond, because the old law it reproduces has two derivatives of the potential and the potential sits inside the shape.
Then the demanding one. Matter's bookkeeping object has no divergence, the statement that nothing is created or destroyed. If it is set equal to something built out of geometry, that something must have no divergence either, and not merely for the geometries that solve the equation. It must have none for every conceivable geometry, as an identity, or the law would quietly demand of matter something no experiment has seen matter obey.
Nothing on that list is a preference. Each traces back to the absence of preferred coordinates, or to what matter demonstrably is, or to the need to reproduce a law three centuries of astronomy confirmed. What remains after the list is applied is very nearly unique.
This is the section where the law of gravity gets cornered rather than proposed. Start by listing everything with the right shape that can be assembled out of the geometry: there are three items, the boiled-down curvature array, the single curvature number multiplied by the distance rule, and the distance rule by itself. Write the most general mixture of the three, with unknown coefficients, and then impose the requirement that its divergence vanish for every conceivable geometry.
Two of the three items have divergences that are automatically nothing, so they impose no condition. The other two have divergences proportional to the same quantity, which is not generally zero, so the only escape is for their coefficients to cancel. That single demand fixes the ratio of the first two coefficients, and the combination it produces is exactly the one built three chapters earlier from a differentiated identity. The famous factor of one half was never chosen; it is what makes the cancellation work.
What survives is one equation with two unfixed numbers in it, and the rest of the chapter is about those two numbers. A quoted theorem strengthens this considerably: even if the restriction to the mildest possible dependence on second rates of change is dropped, nothing new appears in four dimensions, though it does in five. The law is not one option among many. It is what is left.
The same law arrives again by a road sharing no step with the first. Attach a single number to each possible shape of spacetime, namely the total curvature added up with the correct volume weighting, and ask which shape makes that number stationary. There is essentially one number available to attach, because insisting on no more than two rates of change leaves exactly one scalar and a constant.
Varying it splits into three parts. The first asks how the volume weighting responds, and the rule for differentiating a determinant answers it. The second needs no work. The third looks worst and is a total derivative, contributing nothing inside the region and everything on its edge. Combining the first two, the object cornered in the previous section appears unbidden, and the notorious factor of one half arrives from the derivative of a determinant rather than from a differentiated identity. Two unrelated routes, one coefficient.
The discarded edge term is not tidy-up. It is where the second rates of change hide, which is why the equations come out second order though the number attached is not. Setting it aside also demands more of the boundary than one is entitled to, and repairing that costs an extra term whose value matters in one place above all: the entropy of a black hole is what the gravitational number evaluates to, and that story waits for the last part.
Everything so far has been shape without scale, since the constant tying geometry to matter was carried along unfilled, and filling it in is where the construction becomes a theory of the world rather than bookkeeping. Take a weak, unchanging field and slow matter, and rearrange the law so the boiled-down curvature stands alone on the left. Its timekeeping entry works out to be the ordinary second-derivative operator applied to the potential, divided by the square of the speed of light, which the curvature chapter had already reached by a different argument about drifting dust.
The matter side gives the energy density less half its own trace, which for slow cold matter is half the energy density. Comparing with the three-centuries-old equation relating potential to density fixes it at eight pi times the gravitational constant over the fourth power of the speed of light. Its reciprocal is more eloquent: about ten to the forty-two newtons, the stress needed to bend spacetime usefully, and the reason gravity looks like the feeblest thing in physics.
One term survives that Newton had no way to see. Because the source is the whole bookkeeping object and not merely its topmost entry, pressure appears in it, tripled. Compressing a gas increases the gravity it makes, light pulls twice as hard as cold matter of the same energy, and a substance with sufficiently negative pressure would push rather than pull.
One extra piece slipped through the cornering untouched, namely the distance rule itself multiplied by a constant. It slipped through because the requirement doing all the work, that the geometry side have no divergence, is automatically satisfied by the distance rule. Seen from the action, the same piece is a constant added to the quantity being extremised, and no principle in this book forbids that.
Moved across to the matter side, the term describes a substance, and its properties are forced rather than assumed. Its pressure must be exactly the negative of its energy density, which is the unique choice that looks identical to every observer, as the energy of empty space ought to. Because the source of gravity contains three times the pressure, such a substance has a negative source and therefore pushes rather than pulls. And its density cannot dilute as the universe grows, since it is built from constants, so a contribution negligible early on can come to dominate later.
The observed value is small beyond ordinary description, and small is worse than zero. A quantity forced to vanish can be explained by a principle forbidding it; a quantity very small and not zero needs an explanation producing the actual number, and none exists. This is the single most embarrassing number in physics and the last part of the book returns to it without pretending to fix it.
Counting is worth doing before solving. The unknown is a symmetric array with ten entries and the law supplies ten equations, a matched set until one notices that the identity driving the chapter makes four of them redundant. Only six carry information about how the geometry develops.
That shortfall is not a defect and could not have been otherwise. Any solution can be repainted with different coordinate labels, which takes four arbitrary functions and changes nothing measurable, so the law is obliged to leave four functions' worth undetermined. The identity exists to make room for the freedom. The four leftover equations are not useless; they are conditions the starting data must satisfy, automatically preserved thereafter, exactly as the law relating electric field to charge constrains starting data rather than governing its development.
The last observation is the expensive one. The geometry side is nonlinear in the geometry, so the field of two masses is not the field of one added to the field of the other. There is no way to avoid this, because everything carrying energy is a source, and the gravitational field carries energy, so gravity is a source of itself. Electromagnetism escapes because its field carries no charge. That difference is why exact solutions are rare, why approximation is the normal state of affairs, and why the last part of the book finds gravity so much harder than everything else.
3.7 Schwarzschild: The Solution and Its Orbits
Two words are doing all the work in this chapter and both have to be turned into statements about the geometry rather than about a drawing. Spherically symmetric means there exist three directions of dragging along which every measurable distance comes out unchanged, and that those three combine among themselves exactly the way turns about three perpendicular axes do. Static means two further things at once: there is another such direction, one along which a clock could actually tick, and running that direction backwards changes nothing either.
The second half of that is not decoration. A geometry with an unchanging time direction but no reversal symmetry is called stationary instead, and the difference is real rather than pedantic, because the space outside a spinning star is stationary and not static: reverse time and the spin reverses with it. What the reversal buys is the disappearance of every entry in the distance rule that pairs time with a direction in space, and without it the calculation below would carry another unknown function.
Naming a symmetry as a direction of dragging is what makes it something to calculate with rather than to look at. The previous chapter turned each one into a quantity conserved along every free path, and the same three-line test used there, that a label the distance rule never mentions supplies a symmetry, will identify the ones needed here by inspection.
Exact solutions are rare, as the last chapter said, and symmetry is how the rare ones happen. Ten unknown functions of four variables become two functions of one variable. Following how is worth more than the answer. Independence of time, and of time's direction, removes every entry pairing time with space. Spherical symmetry makes every sphere a perfectly round one and forces the angular part into the single combination a globe uses, with one number in front of it. What survives is one factor multiplying the time part of the distance rule and one multiplying the radial part, each depending on the radial label alone.
Then comes a move that is a decision rather than a deduction, and it has to be said out loud. The radial label is defined by declaring that the sphere carrying it has area four pi times that label squared. Nothing says the label equals the distance anybody would measure walking inward from one sphere to a smaller one, and here it emphatically does not, since the radial factor in the distance rule is what converts one into the other.
Read the label as a distance to the centre and the surface met in the next chapter looks like a catastrophe. It is not one, and the confusion begins right here, in a definition adopted because it makes the algebra short and then quietly reinterpreted as a measurement.
Three independent equations for two unknown functions is one equation more than the count allows, and a system in that position has no business being consistent. It is, and not by luck: the differentiated identity that shaped the field equations one chapter ago guarantees relations among them.
The solution turns on a single line. Take the equation belonging to the time direction and the one belonging to the radial direction, weight each by the other's factor, and add. Every second derivative cancels, every quadratic term cancels, and what survives is the rate of change of the product of the two unknowns, set equal to nothing. So that product is the same everywhere, and demanding that the geometry become the ordinary flat one far from the body makes it one. Two unknowns have collapsed into one, by a physical requirement rather than by algebra. The angular equation then says that the radial label times the surviving function has derivative one, the first kind of equation the toolkit taught anybody to integrate.
Nothing in the empty-space equations knows what a mass is; they are the same equations around a star and around a speck of dust. The constant left over from that integration acquires meaning only when the answer is set beside the old inverse-square law far away, and it is at that moment, and not earlier, that a mass enters at all.
Redo the calculation without assuming anything about time and the assumption comes back unbidden, cornered by the equations rather than imposed. Let both unknown functions depend on time as well as on the radial label, and one component of the empty-space equations, the one mixing time with radius, forbids the radial function any time dependence. The rest of the argument then runs unaltered, and the only surviving trace of time is absorbed by rescaling the distant observer's clock.
So the geometry outside any spherical arrangement of matter is the one already found, whatever that matter does. A star may pulse, collapse, rebound or detonate; if it stays spherical, nothing outside changes. There is therefore no spherically symmetric gravitational radiation, which is why the waves picked up in this century come from pairs of bodies swinging about each other rather than single ones ringing. And a star falling in on itself sends no warning outward: the field it leaves is the field it always had.
A second reading matters more for what follows: empty is not the same as flat. The solution has vanishing boiled-down curvature everywhere outside the mass, since that is what it was built to have, and it is nonetheless not flat: it bends light, holds planets, and further in acquires a surface crossable in one direction only. All of that lives in the part of the curvature the boiling-down discards.
A promise made two chapters ago falls due here, and it is collected in full. The rule for measuring distances outside a spherical mass mentions neither the time label nor the angle around the axis, and by the practical test established there, every label the rule fails to mention supplies a direction along which the geometry does not change. Each hands over a number that is the same at every point of every freely falling path.
The two numbers are the energy per unit mass and the angular momentum per unit mass. They are the same two constants the theorem relating symmetries to conservation produced for planets long before this geometry existed, and nothing about them was assumed; they fall out of the metric and are then recognised. The second, written out, is the equal-areas rule of the seventeenth century, with the traveller's own clock in place of the astronomer's.
The saving is best stated as arithmetic. Free motion is four coupled second-order equations for four unknown functions of the traveller's own time. Spherical symmetry turns any orbit into a single plane, disposing of one. The two conserved quantities integrate two more, once and for all. What is left is one first-order equation in one variable, of exactly the form a marble rolling in a valley obeys. That is not computational cleverness; it is symmetry, converted into constants by a theorem proved for pendulums.
A freely falling body's velocity has a fixed length, and squaring that statement with the two constants in hand gives one equation relating the radial position to the traveller's own time. Rearranged, it says precisely what an elementary mechanics problem says about a marble in a valley: a term quadratic in the rate of change, plus a landscape depending on position alone, adds up to a fixed number. The apparatus of turning points applies unaltered.
Set that landscape beside the Newtonian one term by term. The attraction is there, unchanged. The barrier that keeps a body with sideways motion from reaching the centre is there, unchanged, and it is what made orbits stable for three centuries. Then there is a third term with no Newtonian counterpart whatever: attractive like the first, carrying the square of the angular momentum like the second, and falling off as the cube of the distance: undetectable far out and decisive close in.
That one extra term is the entire difference between the two theories for a body in orbit, and everything the chapter extracts afterwards comes from it. Notice what it does to the barrier. The barrier still rises as a body comes inward, and then, at short enough range, the new term wins and the barrier collapses. A wall that fails at close quarters is a different sort of object from a wall.
Circular orbits sit where the landscape is level. Setting its slope to nothing and clearing denominators leaves a quadratic, a different sort of object from what the old theory offers, where the same question has exactly one answer for every radius anyone cares to name. A quadratic has two solutions, or one, or none, and which occurs is decided by how much angular momentum the body carries.
Two is the ordinary case: an outer minimum, where a body sits stably and returns after a nudge, and an inner maximum, a circular orbit any disturbance destroys. As the angular momentum is reduced the two travel toward one another. At one particular value they meet, and below it there is no solution at all: no circular orbit anywhere, at any radius, however the body is arranged. The landscape has turned into a slide.
The meeting point sits at three times the length fixed by the mass alone, and nothing in Newtonian gravity resembles it. There, every radius admits a circular orbit and every one of them is stable. Here there is a floor, its position set by the mass and by nothing else, and gas spiralling inward must stop orbiting and fall when it arrives. That is a prediction about where the inner edge of a disc around a black hole lies, and why such discs have an inner edge at all.
Changing variables to the reciprocal of the distance converts the orbit from a statement about motion into one about shape, and in the old theory the resulting equation is the simplest in physics: a thing whose second derivative plus itself equals a constant. Its solutions repeat exactly once per turn, which is why an ellipse closes rather than drifting. An earlier chapter said the same from the side of symmetry, with a vector from the focus to the closest approach that never moves.
The relativistic version adds a single term, quadratic in the reciprocal distance and very small. Treated as a small correction about a circular solution, its effect is not a change of shape but of rate: the pattern still repeats, slightly slower than once per turn. The body therefore reaches its closest approach a little beyond where it did last time, and the whole ellipse turns, forever, at a fixed rate.
Two features deserve emphasis. This is a perturbative result, legitimate because the correction is one part in ten million even for the innermost planet, and worthless otherwise. And the turning per orbit depends on the orbit through one quantity only, its width measured across the focus, so a tight eccentric orbit gives the largest shift per turn and completes the most turns per century. That is why the discrepancy showed up at Mercury and nowhere else.
3.8 Light, Redshift, and What a Horizon Is
Light carries no wristwatch, and that is the first thing to give way. Every path in the last few chapters was labelled by the reading of a clock carried along it, and for a path on which the interval vanishes there is no such reading: between the emission of a flash and its absorption a billion years later, the flash records nothing. What stands in for the missing clock is a parameter chosen so that the equation of motion keeps its shape, and that choice is fixed only up to a change of scale and a shift of origin.
The freedom is not a nuisance. It is why the two conserved quantities the previous chapter extracted from the two directions in which the geometry does not change stop being separately meaningful for light. Rescale the parameter and the one playing the part of energy and the one playing the part of angular momentum rescale together. Only their ratio is untouched, and once the units are arranged that ratio is a length: the perpendicular distance from the centre at which the ray would have passed had nothing deflected it. Anyone who has watched a beam scatter off a target has met the same quantity doing the same job.
So a planet needs two numbers and a ray needs one. Everything light does outside a round mass is a function of that single length.
Ask where light can be made to go round in a circle and the older theory is nearly silent, because nothing in it makes light respond to gravity at all. Treat a light corpuscle as an ordinary projectile moving at the limiting speed and Newtonian arithmetic does answer, and it is wrong by a factor of three: it puts the circular orbit inside the radius this chapter spends its second half worrying about, where the real one sits comfortably outside. The correct surface is where the barrier keeping light out of the central region reaches its highest point.
That the barrier has a highest point is the whole content of the section, and it says two things. A ray aimed to pass just outside the peak climbs, slows, turns and comes back out, having wrapped part of the way round. A ray aimed just inside goes over the top and never returns. The dividing aim is a single number, and it is not the radius of the circular orbit but something larger, because a ray arriving from far away is bent inward on the way in.
The perch itself is unstable, in the way the top of a hill is unstable. Nothing settles there, which is why the surface is a boundary rather than a place, and why the dark patch a distant instrument would record is larger than the object casting it.
Solving the equation for a ray exactly is not on offer, and the honest move is the one this book has made in every part: expand about something already understood and keep the first correction. What is already understood here is startling in its plainness. Take the mass out and the equation says light travels in a straight line, and that line, written in the polar language the problem forces on us, is the tidiest formula in the chapter. The whole deflection is a small disturbance of it.
The correction is found by feeding the straight line back into the piece of the equation that was dropped and asking what must be added to absorb it. This is legitimate because the disturbance is small: for a ray grazing the Sun the trajectory shifts by a few parts in a million, so evaluating a small term along the uncorrected path makes an error that is the correction to the correction. Where the ray runs off to infinity in either direction, the disturbed solution has tilted the two asymptotes toward each other by a definite angle.
The result is the exact double of the estimate made from the accelerating cabin, and the arithmetic producing one and three-quarter seconds of arc is worth doing on paper once. That angle, and the geometry of a source lying behind a mass, is the reason astronomers wait for eclipses.
Here is the debt from the beginning of this part, and the payment is more interesting than the number. The cabin argument used only the part of the geometry governing clocks, and it could not have done otherwise: every step in it was about how long something took. So repeat this chapter's calculation on a mutilated geometry: keep the distortion of timekeeping, throw away the distortion of distances, and see what survives. Exactly half survives, and it is the value published in nineteen eleven and withdrawn four years later. Do it the other way round, keeping only the distortion of distances, and the other half appears.
The two halves add, and they add exactly, because at this accuracy the disturbance is small enough for its effects to superpose. That much is arithmetic. The interesting part is why a planet does not collect both.
Run the same calculation for a body moving at any speed and the bending contributed by the distortion of distances is fixed, while the bending contributed by the distortion of timekeeping grows without limit as the body slows. Their ratio is the square of the speed divided by the square of the speed of light. For Mercury that is a part in twenty-six million, which is why three centuries of astronomy got away with only the clock half. For light the ratio is one, and nothing can be got away with.
Clocks are the subject again, and the fact about them arrives three times from three directions, with the agreement worth more than any one derivation. The first route uses the direction in which the geometry does not change: a quantity built from it is the same at every point of the ray, while the frequency an observer actually measures involves that quantity divided by the rate of their own clock, and the two differ by exactly the factor this section is about. The second counts wave crests, as the equivalence-principle chapter did, and needs only the relation between a stationary clock and the coordinate label. The third takes the answer far from the mass, where it collapses onto the accelerating-cabin formula obtained with no general relativity in it at all. Three routes, one answer, and the third closes a loop opened eight chapters ago.
Then satellite navigation, usually presented as two competing corrections added by hand. It is not two corrections. There is one geometry, one square root and one expansion of it; the term involving height and the term involving speed drop out of the same line. They may be added because each is around a ten-billionth while their product is around a hundred-billion-billionth, which over a day comes to a few femtoseconds. Neglecting the whole thing costs eleven kilometres a day.
A distinction prepared when the flat grid became a patchwork, and flagged again one chapter ago, carries everything here: whether a failure is a fact about the labels rather than a fact about the thing described. At one particular radius two of the numbers describing the geometry misbehave, one dropping to zero and the other running away to infinity. The temptation to read that as a catastrophe is enormous and should be resisted, because the same thing happens at the origin of a flat page in circular labels, and nothing is wrong with the page.
Three calculations settle it, and none is an argument from taste. Somebody falling inward crosses the offending radius after a perfectly ordinary and finite interval on their own watch, while the coordinate label assigned to the crossing runs away without limit, so only one of them is about the traveller. The distance down to the surface, measured with the geometry's own ruler, is finite as well, despite the ruler component being the one that diverges.
And then the number that cannot be argued with. There is a quantity built from the curvature that every observer computes the same, whatever the labels, and at the offending radius it is finite. It becomes infinite in one place only, the centre. Bigger holes are gentler at the edge, so gentle that for the largest nothing detectable happens at crossing.
If the trouble is with the labels, change the labels. The recipe is the one the manifold chapter set out and never used in anger: find a better chart on the overlap and check that the geometry written in it behaves. The better chart is built by following the light. Ask what path an inward radial flash takes, integrate that condition to get a new radial ruler, and use the arrival label of the flash itself as the new time. Written in those coordinates the geometry contains no bad number at that radius at all, and the quantity measuring whether a description has collapsed is as healthy there as anywhere else.
What the new chart makes visible is the thing genuinely there. At each radius the two directions a flash can take are computable, and following them inward one watches them lean over together. Far out they lean opposite ways, which is what being able to leave means. At the special radius the outward one stops leaning outward and stands still. Inside it, both lean inward. None of this is drawn; it is a pair of slopes read off an equation.
An accelerating rocket in empty space has a boundary of the same local kind, worked out in full earlier. The honest difference is that the rocket's boundary belongs to the rocket and dissolves when it stops accelerating, and this one belongs to nobody.
The geometry does fail in exactly one place, and by now the criterion for saying so is not a matter of opinion. The invariant that stayed finite at the surface everyone calls the edge of a black hole runs away to infinity at the centre, and being an invariant it does so in every description at once. No cleverer chart is available, because the failure is not in the labels.
What should be said about it is less than people expect. The theory does not describe what happens there. It stops, in the specific sense that the paths of falling matter end after a finite interval of their own time with no continuation the equations can supply. That is not a paradox and not a scandal; it is a theory reporting the edge of its own domain, which is the most useful thing an incomplete theory can do, and its own equations are what report it.
Nor is it a peculiarity of the perfectly round solution read here. Theorems exist showing such an ending is forced under conditions assuming no symmetry whatever, and those conditions include an assumption about matter this book has leaned on silently everywhere and never named. Naming it is the last thing done, because knowing which assumption carries the weight is the difference between a result and a slogan.
3.9 Cosmology, and a Loose Thread
Every geometry in this part so far has been built around a lump: a star, a hole, a mass with empty space around it. The largest object there is has no outside, so the method must be turned inside out, and what replaces the lump is an assumption about how boring the universe is. Two statements do the work. There is no special place, meaning the geometry can be slid in any direction without changing what anyone would measure. And there is no special direction, meaning it can be turned about any axis with the same result. In the language of the last few chapters these are directions of dragging along which the distance rule is unchanged, six of them.
Two honesty notes belong with the assumption rather than after it. It is an assumption, supported by evidence and not proved: the oldest light in the sky has the same temperature in every direction to one part in a hundred thousand, and surveys stop looking clumpy once the boxes being compared are large enough. And it is plainly false on small scales, since you are not a smooth fluid and neither is the galaxy. What is being described is an average taken over regions vast enough that the lumps stop mattering, which is a decision about what question is being asked rather than a fact about the world.
Assuming so little turns out to determine almost everything. Slice the universe into moments at which conditions are the same everywhere, let each slice be a geometry with no special place and no special direction, and ask which geometries qualify. Insisting that the curvature be the same at every point reduces the whole question to one first-order equation of the kind the toolkit taught anyone to integrate, plus a demand that nothing go wrong at the origin. Its solution comes in exactly three flavours, distinguished by the sign of one constant: the curvature is positive, zero, or negative, and no other possibility exists.
What is left over is a single function of one variable, a scale factor multiplying every spatial distance at once. It is worth being clear about what that function is not. It is not the radius of anything, except in the positively curved case where a radius happens to exist, and its absolute value means nothing, since doubling it while halving every coordinate describes the same geometry. Only ratios of it at two moments are measurable.
One further separation matters, and the curvature chapter promised that this one would turn it on the universe. The sign of the curvature is a local measurement; whether space closes up on itself is not. A geometry can be flat at every point and still finite, as a cylinder rolled from a flat page is.
Two symmetries and one fluid leave the field equations with almost nothing to say, and what they do say is two equations where there appear to be three. The timekeeping entry gives a first-order statement: the square of the fractional growth rate equals the density, up to constants, less a curvature term. Any one of the spatial entries gives a second-order statement about acceleration, but only after the first is substituted into it, which is the first sign they are not independent. Then the check that settles it. Differentiate the first, subtract twice the growth rate times the second, and everything cancels except a multiple of the statement that the fluid thins as its volume grows.
That is not an accident, and you have met it before in other clothes. The identity that forced the shape of the field equations three chapters ago says the geometry side has no divergence, so the matter side cannot either, so the matter's own equation of motion is contained in the field equations rather than added to them. Here that abstraction becomes arithmetic you can check in four lines.
The pencil on its point arrives at the end. There is a solution in which the universe sits still, matter's pull exactly cancelled by the repulsive term, and it is why that term was invented; disturbing it by any amount produces growth doubling on a timescale of billions of years.
Everything a universe does is decided by what is in it, and the deciding quantity is the ratio of pressure to energy density. Cold matter has none worth mentioning and thins out as its volume grows, which is the arithmetic of counting: the same particles, more room. A gas of light thins out one factor faster, and that extra factor is precisely the stretching of each wavelength, arriving here as an energy loss. The energy of empty space, by the argument of three chapters ago, cannot thin out at all, because it is built from constants.
Feed each into the growth equation and the answers are three lines of algebra apiece, because the problem has fallen apart into independent pieces again. Light alone gives growth like the square root of time, matter alone like the two-thirds power, and the vacuum term alone gives growth by a fixed factor per unit time, without end. Since the three dilute at different rates, whichever falls off fastest wins early and whichever falls off slowest wins late, so the history has a fixed running order regardless of the amounts: light, then matter, then the term that does not dilute.
The last of those deserves stating plainly. Growth by a fixed factor per unit time never slows, so a universe containing any positive amount of that term is eventually dominated by it, whatever else is in it.
Light from far away is followed the way the previous chapters followed it, by writing down the path a flash actually takes and asking what a receiver measures along it. Done here the answer is startlingly clean. The measured frequency falls in exact proportion to the scale factor, so the stretch of a spectral line is the ratio of the scale factor now to its value when the light set out. Distant spectra arrive reddened, so that ratio exceeds one and the scale factor is growing. No velocity appears anywhere in the derivation, and none appears in the answer.
Galaxies at rest with respect to the smoothed-out matter are in free fall and keep the same coordinates forever; nothing pushes them apart and nothing is moving through anything. What grows is the conversion between coordinate differences and measured distances. Distant galaxies therefore have separations growing faster than light, which breaks no rule: the speed limit compares two things at the same place, and velocities at different places cannot be compared without carrying one to the other along a path.
The rubber balloon gets the important part right and one thing badly wrong. Dots on it do recede from one another with no dot at the centre, and the surface has no edge. But the rubber sits in a room, and the room is what makes expanding into anything mean anything. There is no room.
This is the oldest promise in the book, made where a stretched photon had nowhere to put the energy it lost, and the answer is not a puzzle but a theorem returning nothing when handed nothing. Conservation of energy was never a law standing on its own. It was the output of an argument whose input is a symmetry: laws that do not change with time yield a quantity that does not change with time. In an expanding universe the geometry itself changes with time, so the input is absent. Nothing has broken. A machine has been offered no raw material and has produced no product.
The absence can be established rather than asserted. A direction of dragging that leaves the geometry alone must leave alone every number computed from the geometry, and one such number is fixed by the field equations to be proportional to the density of matter, which falls as the universe grows. No such direction can therefore point along time.
The photon gas makes it concrete. Take a region expanding along with the matter and count the energy of the light inside: it falls in proportion to the stretching, with no piston, no surroundings and no reservoir to receive it. What does survive is the local statement, exactly and always: nothing is created or destroyed at any point. Adding that up over a large region is the step that fails.
One number in this book is left deliberately unpaid, and it is worth knowing precisely what is owed. The surface built in the previous chapter has an entropy, and that entropy is proportional to its area rather than to the volume it encloses: a quarter of the area, measured in the natural unit built from the gravitational constant, the quantum of action and the speed of light. Everything about that is strange. Entropy counts arrangements, arrangements live in volumes, and the answer came out one dimension short. Nothing in this part can produce it, and the last part of the book is where it is counted.
The other thing to say can only be said now, before it happens. The next three parts are done on flat spacetime, and gravity does not reappear for thirty chapters. That is not the book abandoning what it has built. It is an accurate report of where the subject stands: the curvature near anything a laboratory contains is far below the precision of any experiment ever performed on a particle, so flat spacetime is not a convenience but a measured approximation with a stated error.
Where that stops being legitimate is also known, and can be named: the centre of the object taken apart last chapter, the first instants of the expansion described here, and the value of the one constant nobody can explain.
A run of results in Part 0 was proved, admired and then left lying about with no visible use. They are all collected here, at once.
That a self-partnered map has real multipliers and perpendicular special directions. That a rotation with no real special direction acquires two the instant complex numbers are admitted, and that their multipliers are phases. That a map whose partner undoes it changes no length, so anything evolving by one keeps its total for ever. That a function on an interval is a sum over a basis of waves, and that the total of the whole is the sum of the totals of the parts. None of that was physics when it was proved. All of it is the physics now.
What forces the change is that two assumptions carried unexamined through everything so far both fail, and they fail together. The first is that a thing has a definite state and that measuring it is a matter of sufficient care. The second is that a quantity nobody has measured nevertheless has a value. Neither survives contact with experiment at the scale of an atom, and what replaces them is not a repair to mechanics but a different kind of object altogether: a direction in an abstract space, evolving by a rotation, with measurement a projection onto an axis and the prediction the squared length of the shadow.
Part IV · Quantum Mechanics
Linear algebra, taken absolutely seriously.
4.1 What Classical Physics Cannot Do
A kiln with a small hole in it glows, and the colour of the glow depends on how hot the kiln is and on nothing else at all — not the clay, not the shape, not the polish of the inside. That claim sounds too strong to be free, so the argument is worth having: it converts a question about pottery into one with exactly one correct answer.
Set two hot boxes side by side at the same temperature and join them through a pipe carrying a filter that passes one narrow band of colour and nothing else. If the two boxes held different amounts of that colour, more of it would cross one way than the other, one box would grow hotter than the other with nothing whatever being done to either, and heat would have flowed from cold to hot on its own. Thermodynamics forbids that outright. So the two amounts must match, band by band, and since the boxes were arbitrary, the amount of each colour inside any hot cavity is one universal function of colour and temperature.
That is what makes the problem worth attacking rather than merely interesting. Nothing is left to argue about — no material, no geometry, no surface finish — and any theory of light and heat whatsoever has to deliver that single function. The rest of this chapter is what the classical theory delivers when it is asked.
The way to count what a box of light can do is one you already own from the chapter on oscillators. A field in a box is a collection of standing waves, each with its own frequency, each swinging independently of every other, and that decomposition is the same trick that turned two masses on three springs into two separate motions and then turned a chain of masses into a continuous string. Nothing new is needed for three dimensions: there are three whole numbers to choose instead of one, because the wave has to fit a whole number of half-wavelengths across the box in each direction.
Counting them is then geometry rather than physics. Each choice of three whole numbers is a point on a lattice, one to each little cube, so the number of standing waves with pitch below some limit is the number of lattice points inside a ball, which is the volume of that ball divided by the volume of one cube. Ball volumes grow as the cube of the radius, so the running total grows as the cube of the frequency, and the number of new modes appearing in each thin slice of frequency therefore grows as the square. That one fact is what makes the classical prediction diverge, and it comes from nothing deeper than the surface of a sphere growing as the square of its radius.
Now put one shared quantity of heat into each of those modes. That is what the classical theory says to do, and it is not a guess: the same maximum-entropy calculation that produced the Boltzmann weights hands back the same average energy for every mode whose energy is quadratic in its amplitude, whatever its frequency, because rescaling an integration variable does not care what it is rescaling. Multiply the count of modes by the energy of each, and the prediction is finished in a single line.
What comes out is a spectrum with no peak. It rises as the square of the frequency and never turns over, so the total energy in the box is the area under a curve that climbs forever, and that area is infinite. Not large — infinite, and infinite in a way no care with the constants repairs, because the divergence lives in the tail where the modes are and nothing done at low frequency can reach it. Taken at its word, a cavity at room temperature holds unbounded energy and would empty itself into the ultraviolet the moment anyone opened it.
The failure has a shape worth recognising, since it is the same one that makes a Lorentzian have no mean: a density falling off too slowly to be summed, with everything going wrong out where nobody was looking.
Einstein's route to the answer is the one taken here, and the reason is worth stating. It needs no statistics of light at all, only the bookkeeping of atoms trading energy with radiation, and that matters because the statistics of light cannot honestly be written down until the end of this part. Take atoms with two levels sitting in the radiation and write every process that could change how many sit in each, at a rate proportional to how many are available to undergo it. Two are obvious: an atom absorbs and goes up, an atom drops on its own and emits.
The argument will not close with only those two, and that is the moment worth stopping on. Demand that the populations hold steady with each process exactly balancing its own reverse, feed in the ratio of populations the Boltzmann weights supply, and then drive the temperature up without limit. The radiation must grow without limit too; the two-process bookkeeping says it cannot, no matter what values the two rates are given.
The repair is forced rather than chosen. There has to be a third process in which radiation already present makes an excited atom drop and emit, at a rate proportional to how much radiation is present. Nobody had seen such a thing. The algebra insisted on it, and it is the mechanism of every laser since.
Two unknown ratios survive the balancing, and two demands fix them. Requiring the answer to keep growing as the temperature rises makes absorption and the new third process exactly equally strong. Requiring it to reproduce the classical law at low frequency, where that law is right and measurably so, fixes the other. What is left is a formula with one new constant in it, and the honest description is not that a law was derived from nothing: the shape came from the balancing, and the numbers in front came from a limit supplied by the theory being replaced. A second, independent route arrives near the end of this part, and it is that route which makes the first more than a fit.
The new constant is at this stage nothing but a fitted number carrying the units of energy multiplied by time. Leave it that way. What it does is set, for each temperature, a frequency above which the shared quantity of heat is no longer enough to stir a mode at all, so those modes sit the whole thing out. That is what cures the divergence: the modes still pile up as the square of the frequency, but their occupancy dies faster than any power.
Two measurements the formula was never fitted to — where the peak sits, and how the total grows with temperature — come out right, the second of them to nine figures.
Two experiments now, of quite different kinds, saying different things. The first is about arrival. Shine light on a metal and electrons come off with a maximum energy that depends on the colour and not at all on the brightness, and below a threshold colour nothing comes off at all however long anyone waits. A wave delivering its energy smoothly and evenly cannot do that, and the classical estimate of how long one atom would need to gather enough energy from a dim beam runs to seconds, against the nanoseconds measured. Energy arrives in lumps whose size is set by the frequency, through the same constant that turned up in the cavity.
The second experiment is about collisions. Bounce X-rays off electrons and the wavelength shifts by an amount fixed by the scattering angle and by nothing else — not by the incoming wavelength, not by the target material. A wave shaking an electron makes it re-radiate at the driving frequency and produces no shift whatever, so that is not a small discrepancy but a contradiction. Treat the encounter instead as two particles colliding, with the relativistic bookkeeping of the second part used unchanged and nothing added, and the measured shift falls out exactly.
So the lump carries momentum as well as energy, in the proportion the wave description had already quietly suggested, and a suspicion recorded two parts ago becomes a law.
The last two failures are about atoms, and they are worse than the first two, because they are not quantitative errors but structural ones. Heated gases emit at sharp isolated frequencies, and for hydrogen those frequencies fit one arithmetic formula built out of two whole numbers with a single constant in front, to a part in a hundred thousand across the whole visible series. Nothing in classical physics produces whole numbers. A charge orbiting freely can circle at any radius at all, so it can emit at any frequency at all, and a continuous smear is precisely what is not seen.
Worse still, it cannot orbit. An accelerating charge radiates, which is the same theory that explained the radio transmitter and is not optional, so an orbiting electron bleeds energy continuously, spirals in, and reaches the nucleus. Doing that integral gives sixteen picoseconds, after two hundred thousand turns, with the emitted frequency sliding upward the entire way. Every atom in the universe should have collapsed long before anyone could look at one.
So the object needs replacing rather than adjusting, and what replaces it is a state that is a direction in an abstract space instead of a position and a velocity. The mathematics for that was built nine chapters ago and has been waiting since. The next chapter puts the physics back into it.
4.2 The Linear Algebra of Quantum States
Something changes in this chapter, and it is worth naming before it happens rather than afterwards. Everything so far has been cornered. Geometry was not chosen; it was what survived once you insisted that no observer's description could be privileged. Even the field equations of gravity were forced, in the sense that the alternatives were eliminated one by one until a single form was left standing. What arrives now cannot be got that way, and no honest telling pretends otherwise.
So the method changes. Certain statements are going to be put down as assumptions, each in its own box, each marked, each stated before anything leans on it. That is not a weakness in the presentation. It is the difference between a subject that was derived and a subject that was discovered, and confusing the two is how people come away believing quantum mechanics was deduced from something more familiar. It was not. It was guessed, tested for a century, and never yet found wrong.
What makes the chapter worth reading anyway is how little has to be assumed. The mathematics is finished — it was finished nine chapters ago, in a chapter with no physics in it at all. What gets added here is short enough to list on one page, and the rest of the part is that short list being spent.
Here is the debt being settled. Nine chapters ago a chapter of pure linear algebra was written with the announcement that it was quantum mechanics with the physics taken out, and that when the physics came back the work would already be done. What that amounts to is a table with two columns. On the left sits a theorem about self-partnered maps on a space with an inner product, proved with no physics in the room. On the right sits a sentence about measurement. The rows are not analogies and the right-hand column is not an interpretation of the left; each row is one statement appearing twice, in two vocabularies.
Read the table and notice how little is added. Measured values are real because the multipliers of such a map are real. Different outcomes are never mistaken for one another because the special directions belonging to different multipliers are exactly perpendicular. Any state at all is a combination of outcomes because there are enough of those directions to span everything. Weights adding to one is the statement that the squared length of a vector is the sum of the squared sizes of its coordinates.
What has to be supplied on top is small enough to count: seven assertions in this chapter, one more in the last chapter of this part, and one experimental fact. Everything else in the table was earned long ago.
The first assertion is about what a state is, and it has two clauses that do different work. A state is a direction in the space, and its length is fixed at one. Fixing the length is not a restriction on nature but a choice of bookkeeping, because the weights are going to be squared lengths and they have to add to one; doubling every component would double every weight and describe nothing new.
The second clause is stranger and more consequential. Multiplying the whole state by a phase — a complex number of size one, applied to every component at once — changes nothing that can ever be measured, so states differing only by such a factor are the same state rather than two states that happen to agree. What is emphatically not invisible is a phase applied to one component and not another. That relative phase is the entire difference between two beams that reinforce and two that cancel, and it is what a wave description was carrying all along.
So the correct object is not a vector but a direction with its overall phase forgotten, and the counting bears this out: a two-state system has two real parameters to its name rather than four, which is why every account of one draws a sphere.
The second assertion is the one the whole design was aimed at. A measurable quantity is represented by a map that is its own partner, and that single sentence is the last thing needing to be assumed about measurement outcomes. Everything usually presented alongside it as a separate postulate is a theorem already proved: that the readings are real numbers, that distinct readings are perfectly distinguishable rather than merely different, that any state can be resolved into the possible readings, that the weights add to one.
One honest qualification belongs here rather than later. The word used is Hermitian, and in a space of functions that word is not quite strong enough — the stronger property involves specifying which functions the map is allowed to act on, and the difference is physical rather than pedantic. A particle confined to a half-line turns out to have no momentum observable at all. The repair has its own chapter.
The other dividend is labels. Two quantities can be sharp at once exactly when the order of the two maps does not matter, and when there are enough such quantities to leave no ambiguity, the list of their values names the state uniquely. Those numbers are what atoms are catalogued by.
This is where the book stops cornering and starts assuming, and the moment deserves to be marked rather than slipped past. Everything so far has been forced: the action principle out of Newton, the geometry of spacetime out of two facts about light, the field equations out of free fall and one identity. The rule connecting a state to the odds of each outcome is not like that. It is put in by hand, because nothing available implies it and nothing since has managed to.
The rule itself is short. Project the state onto the directions belonging to a given reading, and the squared length of what survives is the probability of getting that reading. Non-negative because it is a squared length; adding to one because the projections reconstruct the whole vector with nothing left over; unaffected by an overall phase because that phase has size one and squaring removes it.
There is a celebrated theorem which comes close to deriving it, and its hypotheses are worth naming because they are where the content hides: it assumes that probabilities are already assigned to projections, additively, and it needs at least three dimensions. Assume that much and the squared length is forced. Which is to say the theorem explains the formula given the framework, and the framework is most of what one wanted explained.
A second assertion is needed and most accounts fold it into the first, which costs them the ability to say later which of the two is in trouble. The first says what the odds are. The second says what is left afterwards: whatever part of the state pointed along the directions belonging to the reading obtained, rescaled back to length one. The rest is gone.
The mathematics of that is something done every day. Restrict attention to the cases consistent with what was observed, then renormalise so the surviving weights add to one again — this is conditioning, and the algebra is identical, line for line. The break comes from what is being restricted. Here it is amplitudes rather than probabilities, and amplitudes can cancel, so a component that survives conditioning may still be annulled by another that also survived. Send a beam through a filter, and whether a second filter passes anything depends on whether the first one was looked at.
What none of this settles is why anything is left at all — why one reading occurs rather than the others, and what physical process the rescaling describes. That question is not answered in this chapter or in this book, and the last chapter of the part says exactly which part of it decoherence addresses and which part it leaves alone.
Now motion, and here almost everything is derived rather than assumed. Total probability must remain one for as long as the system exists, so whatever moves a state forward in time cannot change its length, and a map that preserves all lengths preserves all overlaps and is a rotation of the space. That much is not a physical hypothesis; it is the bookkeeping of the state postulate followed to its conclusion.
Then the correspondence from the toolkit chapter takes over. Rotations of this kind are exactly the exponentials of self-partnered maps, so there is a generator, and it is an observable. The whole question is which one, and the answer is the only genuinely new thing here: the generator is the energy. Nothing forces that. It is an identification with an experiment behind it, in exactly the sense that the constant relating colour to energy was an identification in the last chapter, and it is worth marking as such because it is the single input.
Differentiate the rotation and an equation of motion appears, which the next chapter takes seriously. The pattern underneath is one already met in the classical setting: a conserved quantity and the motion it generates are one object, and putting operators in place of functions repeats every word.
Position and momentum need a relation to each other, and the classical theory supplies one already: the bracket that measured how a quantity changes under the flow another one generates gave a particular answer for those two. The assertion here is that the same relation holds with the bracket replaced by the failure of two maps to commute, times a constant. That is a real assumption, and whether it can be extended to every quantity at once is a question with a sharp and negative answer several chapters ahead.
What the assertion costs is immediate and enormous, and the argument is three lines long. Adding up the diagonal entries of a product does not care about the order of multiplication, so the diagonal sum of the failure-to-commute is zero for any two square arrays whatever. The relation demands that it equal a fixed non-zero number on every diagonal entry, whose sum is therefore not zero. No finite list of numbers can do this. Not approximately either: the best possible attempt at any size misses by as much as the target itself.
So the space cannot be finite-dimensional, and this is settled before any physics is done. Two chapters of mathematics follow, and they were not chosen for thoroughness. They were forced here.
Put two systems side by side and ask what describes the pair. The answer is the one place where quantum mechanics departs from ordinary intuition by an amount that can be written as a number. Classically, describing two things means describing each and listing both, so the descriptions add. Here a basis for the pair is a list of pairs of basis states, one drawn from each, so the dimensions multiply.
That difference is entirely responsible for the phenomenon everyone finds strange. States describing each system separately and then pairing them do exist, but they form a vanishingly thin subset — counted properly, of dimension roughly the sum where the whole space has dimension the product. Almost every state of the pair is therefore not of that form, which means it assigns no state at all to either half on its own. A concrete two-by-two example can be checked by hand in three lines: four conditions on four numbers, and no solution.
The practical scale of this is worth a number. Three hundred two-state systems require more amplitudes to specify than there are atoms in the observable universe, while a description of each one separately needs six hundred angles. That gap is where the interest in building such machines comes from, and the last chapter of this part is about what fills it.
The point of this section is that there is one calculation, and it is finished. Take two states of nearly the same energy with something connecting them, which is the smallest interesting arrangement there is, and solve it once. What comes out is a single expression: the chance of finding the system in the other state swings back and forth forever, at a rate set by the separation of the two energy levels, with a depth set by how evenly the connection mixes them.
Now read the answer three times. In a molecule of ammonia the two states are the nitrogen atom sitting on either side of its three hydrogens, the connection is its ability to pass through them, and the swing takes forty-two trillionths of a second — which corresponds to a microwave line that was used to build the first device of its kind. In a neutrino the two states are the two identities it can be detected with, the connection is the mismatch between those and the two definite masses, and the swing takes six hundred kilometres of flight. In an engineered two-level circuit both numbers are chosen by the designer, and half a swing at twenty nanoseconds is what turns one state into the other.
Three subjects, three sets of units, one matrix. That is what the chapter has been claiming.
4.3 Function Spaces: Measure, L², and Completeness
The integral you were taught works by chopping the horizontal axis into thin strips and adding up the areas of rectangles. It is a good definition and it handles every function a physics problem is likely to hand you. It has one flaw, and the flaw is not about any particular function. It is about what happens when you take a sequence of functions and ask what they are approaching.
A sequence of numbers can be seen to be settling down without anyone knowing what it is settling down to. Every term past a certain point is within a hairsbreadth of every other term: that is a self-contained test, and it needs no knowledge of the answer. Whether a sequence that passes the test actually arrives somewhere is not a fact about the sequence. It is a fact about the world the sequence lives in. Inside the fractions alone, the decimal expansion of the square root of two passes the test and arrives nowhere, because the place it was heading is not a fraction. Adding the missing destinations is what the real numbers are for.
Exactly the same thing goes wrong one level up, with functions instead of numbers. There is a sequence of the most ordinary functions imaginable, each one flat except on a few short intervals, which passes the settling-down test and has nothing to settle on. The destination it wants is a function so ragged that the strip-and-rectangle definition cannot assign it an area at all. That is not an oddity to be stepped around. Quantum mechanics is going to say that a state is a vector in a space of functions, and that the limit of a sequence of states is a state. In this space that sentence is false.
So the definition of area has to be replaced, and the replacement starts one step further back than you might expect. Before asking what the area under a curve is, ask what the size of a set of points on the line is. Get that right and everything else follows.
Before you can improve on the idea of area you have to improve on the idea of length. The question is how big a set of points on a line is, when the set is not an interval and may be scattered all through one.
Three requirements settle it almost entirely. The size of an interval is its length. Sliding a set along the line leaves its size alone. And if a set is cut into a list of non-overlapping pieces, even an endless list, the sizes of the pieces add up to the size of the whole. That third requirement is the only one with any bite, and the word "endless" is where the bite is. Every difficulty in the previous section was an endless process, so a rule that only adds finitely many pieces is no use.
Meeting the three requirements takes a construction, and this chapter quotes it rather than performing it, with the quotation marked in place. What the quotation includes is worth knowing, because one clause of it is spent later on something that looks unrelated: any set that can be measured can be trapped between an open set and a closed one that are as close to it in size as desired. That clause is what will eventually let a wildly discontinuous function be traded for a smooth one.
Two consequences follow at once. A set of separate points, even endlessly many of them, has size zero, because you can cover the first with a cover of length one-half of your allowance, the second with one quarter, and so on forever, and the total is your allowance however small you made it. So the rationals, dense as they are, have no size at all. And the price of the whole scheme, which is the honest part: there are sets to which no size can be given, and the proof that they exist uses nothing but the three requirements. The restriction is real rather than a technician's fussiness.
The new definition of area is one idea. Instead of chopping the horizontal axis into thin strips and asking how tall the function is over each strip, chop the vertical axis into thin bands and ask how wide the set of places is where the function lands in each band. Then add up height times width as before. Counting a pile of coins by walking along it is the first method; sorting the pile into denominations and multiplying is the second. For a finite pile the two totals agree, and they agree here too, wherever the old definition worked at all.
What changes is which functions are eligible. The old definition needed the function to be reasonably well behaved along the horizontal axis. The new one needs only that each of the bands corresponds to a set whose size is defined, and the collection of such sets was built in the previous section to be closed under endless unions and intersections. That is not a coincidence. It is what makes the class of eligible functions survive limits, which is the exact property the old class lacked and the whole reason for the rebuild.
The function that took the value one at every fraction and zero elsewhere now has an area, and the calculation is a single multiplication: the height is one, the set of fractions has size zero, the product is zero. A page of frustration in the earlier chapter becomes a line.
One thing is given up, and it should be said rather than discovered later. Certain integrals converge only because positive and negative contributions arrive alternately and cancel. Sorting the values before adding them destroys that arrangement, so such integrals are not admitted by the new definition and have to be written as limits of ordinary ones. That is the price of never depending on the order of the terms, and the next section is what the price buys.
Everything so far has been preparation for two statements, and these two are what the rebuilt integral is actually for. Both answer the same question: if a sequence of functions is approaching a limit, does the area under them approach the area under the limit?
The first says yes whenever the functions only ever increase, filling in from below and never retreating. Nothing else is required, and the proof leans on exactly one thing: that the size of a growing family of sets is the limit of their sizes. That was the awkward clause about endless lists in the definition of size, and this is what it was for.
The second is the one used daily. It says yes whenever the whole sequence can be held under a single fixed function of finite area, a ceiling that does not move as the sequence evolves. Without such a ceiling the statement is simply false, and the way it fails is always the same: the area runs away somewhere the limit cannot see it, up a spike that grows taller as it grows narrower, or off toward infinity in a bump that never shrinks but keeps moving. Each of those has a fixed amount of area at every stage and none in the limit. A ceiling of finite area is a fence against both.
The second statement is the one an earlier chapter borrowed and promised to repay. It was used to justify differentiating an integral by differentiating inside it, which is the trick that produced every moment of the Gaussian without solving a single new integral. The justification is now complete, and the bound it needs was written down at the time.
The space of states can now be assembled. Take the functions whose squared size has a finite total, and give any two of them an overlap by multiplying one against the conjugate of the other and integrating. The choice of the square rather than some other power is not arbitrary: it is the only power for which the overlap of two members is guaranteed to be finite, which is why physics lives here.
Three requirements were laid down long ago for anything calling itself an overlap. Two of them hold immediately. The third demands that only the zero vector have zero length, and it fails, on precisely the function this chapter began with. A function supported on the fractions alone has zero length without being zero.
The failure is exactly as wide as a set of size zero, so the fix is to stop distinguishing functions that agree except on such a set. Two functions differing at a scatter of isolated points are declared to be the same vector. This sounds like an evasion and is not: it is the same convention already in force for probability densities, where changing the density at a few points changes no probability, no average and no likelihood, and nobody regards the density as ill-defined on that account.
The price is worth naming, because it is charged later. A vector in this space has no value at any particular point. The question of what the wavefunction equals at one exact position has no answer, and every honest statement about a continuous coordinate is an integral over a region. In return, the overlap becomes a genuine measure of distance, and every theorem about lengths and angles proved for abstract spaces becomes available here with no work at all, because none of those proofs ever asked what the vectors were.
This is the centre of the chapter. The claim is that the enlarged space has no gaps: any sequence of states that is settling down settles down onto a state, and never onto something outside the space. That statement is what an earlier chapter took entirely on credit every time it said that a limit of states is a state, and it is the property whose absence made the old integral unusable.
The theorem itself is quoted rather than proved, and the quotation is marked. The shape of the argument is short enough to describe: thin the sequence so that consecutive terms are twice as close as the last pair, add up the gaps, and the two convergence theorems of the previous section show both that the total settles at almost every point and that what it settles on belongs to the space.
The satisfying part is what happens to the two sequences from the beginning of the chapter. The first, built from the fractions one at a time, now converges, and the object it converges to is the zero vector, with an area of exactly zero. That was promised eight chapters ago and it is now delivered.
The second is the one that mattered. It converged nowhere before, and now it converges to a perfectly respectable member of the space, an object at a positive distance from every function the old theory could handle. The gap that was exhibited was real, and what fills it is genuinely new rather than a relabelling. That is what it means for the enlargement to have been exactly the right size.
In finitely many dimensions a basis is a set of perpendicular unit vectors, and you know it is big enough by counting: three of them in three dimensions, and there is nothing more to check. Counting is not available here, so the question of whether a perpendicular set is big enough has to be answered some other way.
Four different answers are in circulation and they all get called completeness. The finite combinations reach everywhere. Every vector is the sum of its components. The squared length is the sum of the squared components. And nothing but zero is perpendicular to the whole set. They look like four separate claims, and the section proves they are one claim in four costumes, which is useful because the last is usually the only one you can check and the second is always the one you want to use.
The proof leans on the previous section in exactly one place: to know that a list of components with a finite total square actually assembles into a vector rather than pointing at a gap. That is where completeness of the space is spent, and it is the reason the two senses of the word have to be kept apart.
One more thing has to be true before any of it is usable, and it is easy to overlook. A basis here could in principle be too big to write as a list. It is not, and the argument is a pleasant one: any two distinct perpendicular unit vectors are a fixed distance apart, so small balls around them never overlap, so a countable supply of approximating points cannot serve more than countably many of them. Bases can therefore be numbered, expansions are ordinary series, and everything a physicist writes with a summation sign is legitimate.
Nine chapters ago the pure waves were shown to be mutually perpendicular, and it was said plainly that being perpendicular is not the same as being numerous enough, and that the second claim was being borrowed. This section returns it.
The route has two halves. First, any state in the space can be approached by a continuous function, which re-uses the one quoted clause about squeezing sets between open and closed ones. That clause was spent a section earlier, to show the space has a countable basis at all, and this is the second job it does. Second, any continuous repeating function can be approached by a finite combination of pure waves, and the argument for that is constructive rather than abstract: average the partial sums instead of taking them, and the resulting weight becomes a positive bump of unit area narrowing onto a point. That bump is the same device used earlier to make sense of an idealised spike, reused without alteration.
Putting the halves together, the waves reach everything, so they are a basis, and the borrowing is repaid.
What remains is a warning that is also the most interesting fact in the section. The sense in which a series of smooth waves reproduces a function with a jump is not the sense you would guess. Measured by total squared discrepancy the approximation improves without limit, falling in proportion to the inverse square root of the number of terms. Measured by the worst error near the jump it does not improve at all: every partial sum overshoots by about nine per cent of the step, forever. Both are true, because the overshoot keeps its height and loses its width, and an integral cannot see a fixed height on a vanishing width. Which of the two is the physically meaningful one is not a matter of taste. Every measurable prediction is an integral over a region, so the first is the one that counts, and the second is what happens to anyone who asks the theory for the value of a wavefunction at a point.
4.4 Domains, and the Adjoint's Domain
A chapter of Part 0 built the whole theory of matrices and eigenvectors, and it ended with an honest confession: four of its arguments had used the fact that there were only finitely many directions to work with. Quantum mechanics does not have finitely many directions, and the previous chapter proved that in three lines rather than asserting it. So the confession has to be settled.
Of the four, one has already been paid. The freedom to rearrange a double sum was bought in the previous chapter: one of its convergence theorems has a corollary that lets a sum be moved through an integral whenever the terms are all positive, and rearranging a double sum is that same move in disguise. Two more belong to the next chapter, and they are both about the same thing: in a space of functions an operator may have no eigenvectors whatsoever, and the machinery built for matrices starts by finding one.
The item that belongs here is the strangest of the four, and it has no counterpart in the finite world at all. A map can be perfectly one-to-one and still miss part of the space it maps into. Nothing is lost and yet something is unreachable. In a room with finitely many directions that cannot happen, so the situation has no picture attached, and the way to get one is to watch it happen. It happens twice in this chapter, and the second time it is the reason a particle held on one side of a wall has no measurable momentum at all.
Some operations stretch, and the useful question is whether there is a worst case. If there is a single number such that no input is ever stretched by more than that factor, the operation is stable: a small error going in produces a small error coming out, and the two smallnesses are tied together by that one number. If there is no such number, small errors can be amplified without limit, and stability is gone.
Differentiation has no such number, and the reason is worth keeping rather than the proof. Wiggles of any frequency have the same size as functions and get bigger the faster they wiggle, so the faster the wiggle the larger the derivative, with no ceiling. Anyone who has tried to take a derivative of noisy measurements has met this in person. It is also the correct statement about momentum, because there is no largest momentum a particle can have, and an operator representing an unlimited quantity had better be unlimited.
Then comes the move that decides everything else. You might hope to write momentum down some other way and recover stability. A theorem says no. Any operation that acts on every state whatever, and that has the symmetry a measurable quantity must have, is automatically stable. Turn that around: an unstable measurable quantity cannot act on every state. So the list of states that momentum is allowed to act on is not a hedge and not a technicality. It is compulsory, and the rest of the chapter is about what it costs.
An operator is no longer a rule. It is a rule together with a list of the states it is allowed to be applied to, and changing the list changes the operator even though the rule is written with the same symbols. That sounds like fussiness and it is the opposite. The list is where the physics of a boundary lives, and two lists with the same rule will shortly give a box two different sets of energy levels.
The partner of an operator, its adjoint, was defined in Part 0 by a single requirement: moving the operator from one side of an inner product to the other must not change the answer. In a space of finitely many directions that requirement pins down a matrix and there is nothing further to discuss. In a space of functions the same requirement has to be read as a question asked of each state separately: is there a vector that does this job for this state? For some states there is and for others there is not, so the requirement quietly hands the partner its own list, which nobody chose and which has no reason to match the first.
One more thing, and it is the thing that makes the rest predictable. Lengthen the first list and the second list gets shorter, because there are now more conditions for a state to satisfy. The two lists move towards each other. An operator worthy of representing a measurement is one where they coincide, and the interesting question is how many ways there are to make that happen.
Two conditions have been separated that Part 0 could treat as one. The weaker says that moving the operator from one side of an inner product to the other changes nothing, so long as both states involved are on the allowed list. The stronger says that, and adds that the partner operator's allowed list is the same list, not a longer one.
The uncomfortable part, and it is worth sitting with, is that the weaker condition already delivers everything a working physicist would check on a single measurement. Measured values come out real. Different values belong to perpendicular states. Averages are real numbers. All of that is in hand without the stronger condition, so no calculation you would naturally perform can tell you which of the two you have.
What the stronger condition buys is not a reassurance but a structure, and it is a statement about every state at once rather than about any one of them. Only for the stronger one is there a guarantee that the possible outcomes of a measurement account for the whole space, and only for the stronger one does the operator generate a flow in time. Lose it and you do not get wrong answers. You get no answers, because the machinery that turns an operator into predictions has a hypothesis it does not meet. The next section shows this happening to the most familiar operator in the subject, on the most familiar interval.
Everything in this section is one line of calculus done honestly. When you move a derivative from one function to the other inside an integral, a leftover appears at the two ends of the interval. In Part 0 it was always thrown away, with the note that throwing it away is a physical claim. Here it is the only thing being looked at, because it is precisely the obstruction to a momentum operator being a measurable quantity.
On an infinite line the leftover disappears by itself. Anything with finite total probability and a finite spread of momentum has to die away at both ends, so there is nothing to arrange and nothing to choose, and momentum is a perfectly ordinary observable. On a segment the leftover is the difference between the two ends, and it can be cancelled by insisting that whatever leaves one end arrives at the other with a fixed change of phase. The size of that change is a real number, free to be anything, and every value of it gives a different genuine momentum operator with a different ladder of allowed values. The condition at the wall is not bookkeeping. It is a physical input, and the measurable momenta depend on it.
On a half-line the arithmetic gives a different kind of answer. The leftover involves one end only, and the sole way to cancel it is to demand that every state vanish there, which is a demand so strong that the partner operator is left with a larger list of states than the operator itself and no amount of adjustment closes the gap. The conclusion is not a technicality: a particle held on one side of an impenetrable wall has no momentum observable at all. Not one that is hard to compute. One that does not exist.
The previous section answered a question three times by direct calculation, and got three different kinds of answer: one operator, a circle of operators, and no operator. A theorem exists that predicts which of the three you are about to get, and it does so by a count that takes a single line.
The recipe is this. Take the partner operator, which has the longer list of allowed states, and look for states it sends to themselves multiplied by a purely imaginary number. The original operator can never do this, because it always returns real multiples, so anything the partner manages here is a direct measurement of how much larger the partner is. Count these states above the axis and below it separately, and the pair of counts decides which of the three answers you are going to get. It gets all three right here, which is why it can be trusted in the next section, where the hand calculation would be considerably longer.
Energy involves a second derivative rather than a first, so at each wall there are two numbers to worry about instead of one: the value of the wavefunction and its slope. Running the counting recipe gives two independent directions of freedom at each end rather than one, and the family of legitimate energy operators for a particle in a box turns out to have four adjustable real numbers in it. Four, not one. The single knob belonged to momentum, where each wall carries one number, and it does not carry over.
Every boundary condition anyone writes down for a box is one point of that family. The textbook condition that the wavefunction vanishes at both walls is one point. Demanding instead that the slope vanish is another. Gluing the two ends together into a ring is a third. A wall that is neither perfectly hard nor perfectly soft sits somewhere in between, on a curve through the family that can be solved and plotted.
And the levels move. Turning the knob slides every energy continuously, and past a certain setting a level drops below zero altogether and becomes a state stuck to the wall rather than rattling around inside. Same particle, same box, same formula, different numbers on the spectrometer. That is the sense in which the condition at the boundary is not bookkeeping. It is a physical property of the wall, it is one of the four numbers, and it is measurable.
4.5 The Spectral Theorem in Infinite Dimensions
The previous chapter established which mathematical objects are allowed to represent measurable quantities, and it did that by being careful about what each object is permitted to act on. What it did not supply is the list of numbers the instrument can actually read. In the finite world that list was the set of special multipliers belonging to the special directions, and finding it was the whole content of the central theorem of the toolkit.
Here is the trouble, and it is not a subtlety. For a particle on a line there are no special directions. Ask for a state in which the position has one definite value and the honest answer is that no such state exists, because a state concentrated at a single point is indistinguishable from nothing at all. Ask for a state of definite momentum and the answer is a perfectly good wave that spreads over the whole line with undiminished size, which means it is not a state either. So the recipe of finding the special directions and reading off their multipliers has nothing to work on, and yet a position measurement plainly returns a number.
The repair takes three steps, and each one loosens a requirement rather than replacing it. Ask which candidate values cannot be undone instead of which ones have states attached. Ask for a change of description that turns the quantity into ordinary multiplication instead of into a list. And replace the sum over separate values by an integral over a continuum of them. The one step that this book quotes rather than builds is the middle one, and the two sections after it check that quotation by hand on the three quantities everything later is built out of.
The word that had to be widened is the one naming the list of possible readings. In the finite world a candidate reading earned its place by having a state attached to it, a state in which the quantity has that value and no other. The widened test asks something weaker and more robust: subtract the candidate reading from the quantity and ask whether what is left can be reliably undone. If it can, that reading is impossible. If it cannot, the reading is possible, and this version of the question makes sense whether or not any state is attached.
Running the test on position gives every real number, which is the answer anyone would want, and it gives it without a single state of definite position existing. What stands in the missing state's place is a sequence of states in which the position is pinned down more and more tightly, none of them perfect and each of them something a laboratory could prepare. The theory is refusing only the idealisation, not the measurement.
Two further things fall out and both are reassurances the theory owed. The readings of a genuine observable are real numbers rather than complex ones, which had been assumed since the quantity was first written down and is now proved. And the ways the test can fail reduce from three to two, so every possible reading either has a state behind it or has a sequence of states closing in on it. Nothing sits in a third category with no physical account at all.
The theorem that replaces the central result of the toolkit says something that sounds almost disappointing until you see what it buys. Every legitimate observable, however complicated, is the operation of multiplying by an ordinary real-valued function, once you have found the right way of describing the states. That is all. The complexity of the operator has been moved entirely into the change of description, and the operator itself has become the simplest thing a mathematician can write down.
Compare that with the finite version and the family resemblance is exact. There the claim was that a quantity is a list of numbers attached to a set of mutually perpendicular directions, and a list of numbers is nothing but a function defined on a finite set of labels. Widening the labels from a finite set to a continuum is the entire generalisation. What genuinely disappears is the set of special directions, and it disappears because for a particle on a line there are none, so the new statement has been written so as never to mention them.
This is the one result in this part of the book that is used but not proved, and the mark on it is the honest admission that its proof would cost three chapters of pure analysis. The next two sections are what makes borrowing it defensible. Rather than trust the general claim, we produce the change of description and the multiplying function by hand for each of the three quantities that everything later is built out of, so the borrowed statement is tested where it carries the most weight.
Borrowing a theorem is only respectable if you check it, and this book checks the borrowed one by hand on three quantities. The first is position, where the check is a matter of noticing that the quantity was already defined as multiplication by a function, so the general claim is being tested against the example it was modelled on. That is not a weakness in the check. It is the reason the general claim is plausible in the first place.
The second is momentum, and there the change of description is the transform between position and wavelength that was built in Chapter 0.9 and shown there to preserve every length and angle. In the wavelength description momentum is multiplication by an ordinary function, one line of algebra confirms it, and a numerical experiment on a four-thousand-point grid confirms it again to four figures. The discrepancy that remains shrinks by a factor of four each time the grid is refined, which is the signature of the crude derivative being wrong rather than the transform.
The third is the energy of an oscillating particle, and it differs from the other two in a way worth pausing on. This quantity does have states of definite value, an entire family of them, so its labels form a discrete list rather than a continuum. The general statement was written to allow either, which is why it speaks of a measure rather than of an interval. All that is missing is the guarantee that the family is large enough to describe every state, and the next section proves that rather than assuming it.
An oscillating particle is the one quantity in this chapter that behaves the way the finite theory said everything would. It has states of definite energy, they are honest members of the space, and their energies form an evenly spaced ladder starting not at zero but half a rung up. All of that came out of a single compact function whose expansion generates the whole family, with no physical input beyond the shape of the energy.
The claim that costs something is not that these states exist but that there are enough of them. Being mutually perpendicular is easy; being numerous enough to describe every state is the hard half, and the previous chapter was explicit that these are different claims and that the second is what everyone actually uses. The proof runs by contradiction. Assume some state is perpendicular to every member of the family, deduce that it is perpendicular to every power of position damped by a Gaussian, expand a wave in powers, and conclude that the transform of that state vanishes at every wavelength. A state with no content at any wavelength has no size, so it was nothing to begin with.
The step that has to be justified is exchanging an infinite sum with an integral, and the licence is the convergence theorem the previous chapter proved for exactly this purpose. That is worth registering as the shape of the whole enterprise: a result quoted in Part 0 and used freely ever since is now earned, and what earns it is a theorem that was itself built rather than borrowed.
Once a quantity has become multiplication by a function, there is a natural way to ask how much of a given state sits in any range of values. Restrict the function to that range by multiplying by something that equals one inside it and zero outside, and what you have is a device that keeps the part of the state belonging to those readings and discards the rest. That device is a projection in the old sense, and there is one for every range rather than one for every value.
Pairing that family with a particular state produces an ordinary probability distribution over the possible readings, and from there everything the finite theory did carries over with sums replaced by integrals. The average reading is the average of the values against that distribution. The list of values times their projections becomes an integral of the same shape. And the probability rule that was stated for individual outcomes becomes a rule about ranges of outcomes, which is what it always had to be for a continuous quantity, since no continuous quantity has a positive probability of taking any one exact value.
This is where the promise made in the toolkit, Chapter 0.5, is finally settled. That chapter said the projection form of its central theorem was the one that would survive, that the sum would become an integral, and that the claim about having enough directions would have to be re-earned rather than assumed. All three have now happened, and the third happened three separate times, once for the space itself and once for each of the two families of waves the book expands in.
Two symbols appear on nearly every page of every quantum mechanics book, and neither of them names anything in the space of states. One would be a wave of definite wavelength stretching over the whole line with undiminished size, which has infinite total intensity. The other would be a state concentrated at a single point, which in this space is indistinguishable from nothing. Calculations using them nevertheless come out right, every time, and that is the fact needing an explanation.
The explanation is that both symbols abbreviate a procedure rather than denoting an object. Confine the particle to a large ring, or chop the line into small cells, and in either case the troublesome object becomes an ordinary state and the space acquires an ordinary perpendicular family to expand in. Do the calculation there, where nothing is idealised. Then arrange the bookkeeping so that each term carries the spacing between neighbouring labels, and let the spacing shrink. Sums turn into integrals, and the spike that everyone writes appears as the limit of a narrowing bump of fixed area, which is how it was introduced in the first place.
Two things are worth keeping from this. The procedure always works, so the abbreviation is safe, and that is the whole licence for the notation. And the comfortable reading of the notation, in which these families are perpendicular bases like any other, is false rather than imprecise. There are only ever countably many directions in this space, and the points of a line cannot be counted. The general theory in which such objects are built properly rather than reached by a limit belongs to a later part of the book, and nothing here needs it.
What this chapter has really produced is permission. Four moves are made constantly in the rest of the subject, and before now none of them was justified for a quantity with no states of definite value. Slipping a complete set of alternatives into the middle of an expression is allowed when the set is genuinely complete, or when it comes from the family of projections built earlier, or when it is shorthand for the box-and-limit routine. Expanding a state in states of definite value is allowed only when such states exist and are numerous enough, which rules out position and momentum altogether.
Writing down an exponential of a quantity is allowed for any legitimate observable, and the next section shows that the particular exponential describing the passage of time preserves total probability exactly. Throwing away the leftover from an integration by parts is allowed only when both functions belong to the set the operator is permitted to act on, which is the whole subject of the previous chapter and the item that goes wrong most often.
The list of what remains forbidden is short and worth keeping. Two quantities cannot be multiplied without first asking where the product is defined, and, less obviously, they cannot be added either: the sum of two legitimate observables is not automatically a legitimate observable, which is why the energy of a particle in a potential has to be examined rather than assumed. A quantity that merely gives real averages does not qualify for any of this, since every result here demands the stronger condition. And the two most-used symbols in the subject are abbreviations rather than states, so nothing that treats them as states is safe.
Time evolution has to preserve total probability, so whatever moves a state forward by a given interval must be a rotation of the space of states. Chapter 4.2 showed that every such rotation is an exponential of something, and that the something is a legitimate observable with the dimensions of energy. That argument was made by differentiating a product rule, and in a space where the observable is unbounded and defined only on part of the space, differentiating a product rule is not a step one may take without care.
This section supplies the care, and it splits into a half that is proved and a half that is quoted. The proved half says that starting from a legitimate energy observable, the exponential really is a rotation, really does compose correctly, really does move continuously, and really does hand the energy back when differentiated. All four are one-line statements once the observable has been turned into multiplication by a function, because an exponential of a real function has size one everywhere.
The quoted half runs the other way and says that any family of rotations that composes correctly and moves continuously arises from exactly one such observable. Put the two together and the statement is a translation rather than a discovery: a flow that conserves probability and a legitimate energy are the same thing named twice. What is left over as genuine physics is the single claim that this particular operator is the energy a calorimeter measures, and that claim was labelled a postulate when it was made and remains one.
4.6 The Schrödinger Equation
Whatever moves a quantum state forward in time has to do three things, and none of them is optional. It has to respect combinations, because the theory says any combination of two possible states is itself a possible state. It has to leave the total probability at one, because that is what a total probability is. Those two together force it to be a rigid rotation of the space of states, preserving not only lengths but every angle, and a rotation is reversible by construction.
It also has to compose. Waiting an hour and then waiting a second hour must be the same as waiting two hours in one go, and that is a real assumption rather than a formality: it says the system is not being interfered with while it evolves. Switch on a laser halfway through and the assumption is false, which is why a later chapter has to build different machinery for a driven system.
And it has to move continuously, with the state after a very short wait very close to the state before it. The technical form of that requirement is weaker than the obvious one, and the weakness is deliberate. Demanding the strong version would restrict the theory to systems whose energy has a ceiling, and almost nothing in physics does.
The three requirements of the previous section turn out to have exactly one solution, and a theorem hands it over. Any family of rotations that composes correctly and moves continuously is the exponential of a single fixed quantity, that quantity is a legitimate observable, and it is determined by the family rather than chosen alongside it. So a system that conserves probability automatically has something playing the role of a generator, whether or not anyone has identified what it is.
Two things then have to be added by hand and neither is mathematics. The first is the name: the generator is the energy, the same energy a calorimeter reads and the same energy classical mechanics computes. That is a postulate, it was labelled as one when it was made, and it is what makes the whole subject calculable, because the classical energy of a system is something you can write down before you know anything quantum about it.
The second is a bookkeeping decision about signs. There are two square roots of minus one and nothing distinguishes them, so calling one of them the imaginary unit is a free choice. Everything after that is not free. Once the choice is made, the sign in the momentum operator and the sign in the exponent are both determined, and the check that they are right is that a particle sent to the right goes to the right. Literature from other fields often makes the opposite choice consistently, and a formula carried across without conjugating it will be wrong in exactly the quantities that matter.
Differentiate the flow and the equation appears. It says that the rate at which the state turns is the energy operator applied to the state, with a factor in front converting an energy into a frequency. Everything on the way here came from two places: probability has to keep adding to one, and the system has to be left alone while it evolves. The single genuine assumption is the name given to the operator that comes out.
Four misreadings are worth heading off, because each of them costs something later. The equation is first order in time, unlike the equation for a vibrating string, so one snapshot fixes everything forwards and backwards. The function it describes is not a disturbance in any material, and for two particles it lives in six dimensions rather than three, which no substance filling the room could do. The function itself is not measurable, and multiplying all of it by a fixed phase changes nothing at all, though changing the phase of one part of a superposition relative to another changes a great deal. And nothing in it describes a measurement, which stays a separate statement, and stays one until the last chapter of this part.
There is one restriction on who the equation applies to. Taking a derivative requires the state to be one the energy operator is allowed to act on, and the previous two chapters were about how seriously that restriction has to be taken. The flow itself applies to every state without exception, and the differential equation is the version that holds where the derivative exists.
The equation says the state turns at a rate set by the energy operator, and so far nothing says which operator that is. The answer taken here is the classical one carried across: kinetic energy built from momentum, plus potential energy built from position. This is a choice. Nothing in the previous three sections implies it, and it is worth being blunt about that, because most presentations slide from the abstract statement to this formula as though the second followed from the first.
What makes the choice safe in this particular case is that no term in it mixes position with momentum. Those two do not commute, so a classical expression containing both has several possible operator versions and no reason to prefer one. Kinetic plus potential has no such term, so there is nothing to decide, which is exactly why the trouble shows up elsewhere: a charged particle in a magnetic field does mix them, and a later chapter proves that no rule for translating classical quantities into operators can be made to work for everything at once.
There is also a condition to check before the expression is admissible. The operator has to be of the strict kind the last two chapters were about, and whether the sum qualifies depends on how badly behaved the potential is and on what happens at the edges. No general theorem is quoted for it here. Every case this book uses is checked where it arises, which is the same standard the previous chapter set for itself.
To compute anything you have to stop talking about operators in the abstract and say what they do to something. The usual choice is to describe a state by a function of position, in which case position acts by multiplying by the coordinate and momentum acts by differentiating and multiplying by a constant. The check that this is the right operator is one line of the product rule: the two derivative terms cancel and the leftover is exactly the relation the previous chapters postulated.
That formula is not the only possibility, and it is worth knowing why nobody worries about the others. Multiplying every state by a position-dependent phase, and adjusting the momentum operator to compensate, gives a different pair satisfying the same relation. A quoted theorem says these copies are the whole story: any honest realisation of the relation is the standard one seen through a rotation of the space, so the choice made here is a choice of coordinates rather than of physics.
The theorem has three conditions and all three earn their place. It needs the relation in its exponentiated form rather than in its raw form, and the difference is not pedantry, since the momentum of a particle on a half-line satisfies the raw relation and is not equivalent to anything. It needs the description to be irreducible, which fails when the particle carries an extra property such as spin. And it needs finitely many degrees of freedom, which fails for a field, where genuinely different descriptions of the same relations exist and the choice between them is physics.
The description in terms of momentum is available on the same footing, and passing between the two is the Fourier transform. The convention chosen nine chapters ago makes that transform preserve total size exactly, so a state normalised in one description is normalised in the other with nothing to adjust.
Substituting the momentum operator into the energy operator, and the energy operator into the equation of motion, turns an abstract statement into an equation for a function of position and time. The rate of change of the function is set by two things added together: how sharply the function is bending, and how high the potential is where the function is sitting. Both are energies, and the constant in front of the time derivative is what converts a rate of turning into an energy.
The second derivative on the right is not an extra ingredient. In the momentum description kinetic energy is multiplication by momentum squared over twice the mass, which is the familiar formula, and the transform between the two descriptions turns multiplication by momentum into differentiation. Differentiating twice is what multiplying by momentum squared becomes.
One consequence is worth carrying forward. Written as an average, the kinetic energy is the total steepness of the function, added up over space, and steepness is never negative. So making a state fit inside a small region forces it to rise and fall quickly, which forces its energy up. That is the entire reason a confined particle cannot be at rest, and the next chapter turns it into numbers.
Set the quantum equation beside the equation for something spreading through tissue and they are the same equation, with one coefficient real in the first case and imaginary in the second. Solve both for a single wavelength and the whole difference appears in one line. In the classical case the amplitude of each wavelength shrinks, fastest for the finest detail, so structure is erased and the profile forgets where it started. In the quantum case the amplitude of each wavelength keeps its size exactly and only its phase advances, so nothing is erased and nothing is forgotten.
That is what the imaginary unit is for. It converts a factor that shrinks into a factor that turns, and a factor that turns is what conserving total probability requires. It also settles which equation can be run backwards. Reversing the classical one amplifies the finest structure without limit, which is why reconstructing a past profile is unstable. Reversing the quantum one changes nothing in size, which is the same statement as the flow having an inverse.
Two corrections are worth carrying away. The quantum equation is not the equation for a vibrating string with an imaginary unit added, since that equation is of a different order in time. And the quantity that spreads like a diffusing substance is not the wavefunction: a free packet widens in proportion to elapsed time rather than to its square root, because the spread in speed was fixed at the beginning and nothing is knocking the particle about along the way.
Total probability staying at one is already known. What this section adds is much stronger and is the thing anyone actually wants: probability does not disappear in one place and reappear in another. It flows, there is a formula for the flow, and the formula together with the density satisfies the same balance law that governs charge and mass, the one saying that whatever is inside a region changes only by crossing the boundary.
The derivation is four lines and one of them carries all the weight. The potential energy terms cancel, and they cancel because the potential is a real quantity. Give it an imaginary part and the cancellation fails, probability drains away at a fixed proportional rate, and the model is describing something leaving the system rather than a closed system. That is the same requirement as the one the last two chapters were built around, wearing different clothes.
The flow itself has a clean reading. Write the wavefunction as a size times a phase, and the flow is the size squared times the rate at which the phase changes across space, divided by the mass. So the flow lives entirely in the phase, which is the part of the wavefunction the probability rule throws away. A wavefunction that is real carries no flow at all, however sharply it is peaked, and two states with identical probability densities can have completely different currents. Adding the flow up over all space gives the average momentum divided by the mass, which is what a current ought to be.
Look for solutions in which time and position appear in separate factors and the equation splits in two. The time factor is a phase turning at a rate proportional to a constant, and the position factor satisfies an equation saying the energy operator returns that same constant times the function. So the separable solutions are exactly the states of definite energy, and finding them is an eigenvalue problem in space alone with no time in it.
They are called stationary because nothing measurable about them changes. The phase factor has size one, so it cancels out of every probability and every average, and an atom left alone in one of these states stays as it is. That does not mean nothing is flowing: a state can have a perfectly steady current running through it and still be stationary, because a steady current changes no density anywhere.
They matter because everything else is built from them. The equation is linear, so any combination of solutions is a solution, and when the states of definite energy are numerous enough every state is such a combination. Solving a quantum problem then means finding the energies and their states once, after which evolving any initial condition is attaching a phase to each piece and adding. Combine two of them and the interference term beats at the difference of the two energies divided by the constant, which is what a spectral line is.
A free particle has no states of definite momentum, because a pure wave of one wavelength stretches over the whole line and cannot be normalised. What exists instead is a bundle of wavelengths added together, narrow in space and narrow in wavelength, and the narrower you make it in one the wider it is in the other. Starting from a bell-shaped bundle, the evolution can be worked out exactly, and the answer is a bell shape at every later time.
Two things happen to it. The centre travels at the ordinary classical speed, momentum over mass, and it is worth knowing that the ripples inside the bundle travel at half that speed, so no visible crest ever moves at the speed of the particle. And the bundle widens, with the widening driven by the spread of speeds it was given at the start rather than by anything happening to it along the way. The arithmetic is the same as for a cloud of classical particles released with a range of speeds.
What is quantum about it is that the spread of speeds cannot be reduced without widening the starting bundle, since their product is fixed. Put numbers in and an electron pinned to a nanometre carries a speed uncertainty of tens of kilometres per second and doubles its width in thirty femtoseconds. Do the same for a speck of dust and the doubling takes a million years. The mass in the denominator is what separates the two, and that is the shape every argument about the classical limit takes.
4.7 Wells, Barriers, and Tunnelling
Solving a quantum problem means finding the energies at which the equation has an acceptable solution, and the states that go with them. When the potential is flat on each of a few intervals, the solution on each interval is something you have known since Chapter 0.8: a wave if the energy is above the potential there, a rising or falling exponential if it is below. All the work is in stitching the pieces together at the joins.
The stitching rule is that the function and its slope both run continuously across a join. That is usually presented as a rule to memorise. It is not a rule. Break the slope and the second derivative acquires a spike of infinite height and zero width, so the energy operator applied to the function produces something that is not a function in the space at all, and the operator is not allowed to act on it. Break the function and it is worse. So the stitching conditions are a description of which functions the energy operator is permitted to act on, which is the thing the last three chapters have been calling a domain.
One case escapes. The argument needs the potential to be a bounded quantity, and an infinitely high wall is not. There the previous chapter's warning applies in full: the formula for the energy does not by itself say what the operator is, and for a box there is a four-parameter family of legitimate answers with different energy levels. The one everybody uses, where the wavefunction is pinned to zero at each wall, is picked out by taking a finite wall and making it higher and higher. As it rises, the wavefunction outside is squeezed into a thinner and thinner sliver, its value at the wall goes to zero, and its slope does not. That is the choice, and it is a statement about the wall rather than about the particle.
Reflecting a function through the origin is an operation you can perform on any state, and it is a legitimate observable. Doing it twice puts everything back, so its only possible readings are and , and the states that read are the functions symmetric about the origin while those reading are the antisymmetric ones. Every function splits into one of each, so between them the two readings cover everything.
This observable is compatible with the energy exactly when the potential is symmetric about the origin, and then it becomes useful. In one dimension a bound state never shares its energy with another state, which is proved above from a quantity built out of two solutions that turns out to be constant and to vanish far away. Reflecting a bound state gives a state of the same energy, and with nothing else at that energy to be, it has to be the original one back again, multiplied by or by . So every bound state of a symmetric potential is symmetric or antisymmetric.
That saves half the work. Instead of solving for an unknown function on the whole line, you solve on half of it twice, once for each symmetry, and the two searches together miss nothing. The same three steps, a symmetry, an operator that commutes with the energy, and a label, run again on rotations, on which transitions an atom is allowed to make, and on the swapping of two identical particles. This is the version with no algebra in it, which is why it repays a slow reading.
Inside a box with nothing to push the particle around, the equation says the wavefunction curves in proportion to its energy, so the solutions are waves. Pinning the wave to zero at both walls means a whole number of half-wavelengths has to fit exactly across the box, which allows only certain wavelengths, which allows only certain energies. That is where the discreteness comes from. It is not in the equation, which is happy at any energy at all. It is in the walls.
The lowest level is not zero, and the reason has to be exact. Zero energy would need a wavefunction with no curvature that also vanishes at both walls, and the only such function is the one that is zero everywhere. That is not a description of a particle sitting still. It is not a description of anything, because a state has to have total probability one and the zero function has total probability zero. So the lowest genuine state is the single arch that rises and falls once across the box, and its energy is what that arch costs.
The numbers come out at the right scale with no fitting whatever. An electron confined to a nanometre has a lowest level of electronvolts and a gap to the next level corresponding to near-infrared light. Confine it to a couple of Ångströms, the size of an atom, and the gap moves into the ultraviolet, which is where atomic transitions actually are. Because every level goes as one over the square of the size, halving the box quadruples every energy in it. That steepness is the whole reason confinement matters at small sizes and is invisible at ordinary ones.
With walls of finite height, the wavefunction no longer has to vanish at the edges. It wiggles inside the well and then decays outside, and the two pieces have to meet with the same value and the same slope. Removing the arbitrary overall size from those two conditions leaves one equation relating the energy to itself through a tangent, which no rearrangement will solve. So the answer is read off a graph and then computed numerically, and that is the normal situation rather than a failure.
The graph is worth more than the numbers. Everything about the well is carried by one combination of its depth and its width, and moving that one number slides a falling curve across a family of rising ones. Each crossing is a state. As the well deepens the falling curve reaches further right and picks up a new rising branch at regular intervals, so states appear one at a time, each arriving with almost no binding and tightening as the well deepens further. Counting the branches gives the number of states without solving for any of them.
Two consequences deserve keeping. In one dimension the first crossing exists no matter how shallow or how narrow the well is, so a square well on a line always traps something. In three dimensions the same picture applies with the first branch removed, for a reason that comes from the origin of a radial coordinate, and a weak enough attraction traps nothing at all. And squeezing the well to a point while making it proportionally deeper leaves exactly one state, at an energy proportional to the square of the well's strength, which is the one bound state in this book you can write down without solving anything.
Above the top of a step there is nothing to trap the particle, so no solution dies away at infinity and none of them is a state in the strict sense. That is expected rather than alarming: the previous two chapters said in advance that an operator can have a whole continuum of readings with no states of definite value behind them. What can still be computed are ratios, and a ratio is what the experiment measures anyway.
The natural ratio to write down compares the size of the transmitted wave with the size of the incoming one, and that is the wrong answer. Size is not the quantity that flows. What flows is a density multiplied by a speed, and the transmitted wave travels at a different speed because it has a different wavelength. So the right comparison is between the two flows, which puts a ratio of wavenumbers in front of the ratio of squared sizes. Getting this wrong is not a detail: the two numbers then fail to add up to one, and the failure is tens of per cent.
Once transmission and reflection are defined as flows, the fact that they add to one is not a separate assumption. The previous chapter proved that probability flows without being created or destroyed, and a state whose density does not change in time must therefore carry the same flow at every point. Everything that goes in comes out on one side or the other. The one physical surprise is the numbers themselves: a classical particle with more than enough energy always gets over a step, and this one is sometimes thrown back, because it is a wave meeting a change in its own wavelength.
Below the top of the wall the equation does not stop having solutions. The wavefunction stops oscillating and starts decaying, and decay takes an infinite distance to reach zero, so a wall of finite width always has something left at the far side. What comes out is an exponential of the width, of the square root of how far the top of the wall is above the energy, and of the square root of the mass. That last one is why this is an electron's phenomenon.
The number for a case you could build: an electron with half the energy it needs, meeting a wall a nanometre thick, gets through about three times in a thousand. Add half a nanometre and the answer drops by a factor of thirty-seven, and it drops by that same factor for every further half nanometre. That steepness is the whole practical content. It is what makes a microscope that reads a surface by tunnel current sensitive to a single atom's worth of height, and it is why radioactive half-lives of the same kind of decay can differ by twenty-four powers of ten while the energies differ by a factor of two.
Above the top something stranger happens. The decaying exponential turns into an oscillation, the formula acquires a sine, and a sine has zeros. At particular energies the wall becomes perfectly transparent, with nothing at all reflected, because the reflections from its two faces cancel each other exactly. Those energies are the ones at which a whole number of half-waves fits across the wall, which is the same condition that gives the levels of a box. Following the same formula down to negative energy, the places where it blows up are exactly the bound states of a well of the same size. Trapping and scattering are one problem read on two halves of one axis.
4.8 The Oscillator, and the Ladder
The energy of an oscillator is a sum of two squares, one for the motion and one for the stretch. For ordinary numbers a sum of two squares can be written as a product of two factors, and the reason is that the two cross terms in the product cancel each other. That trick is worth wanting here, because a product of simple things is far easier to work with than a sum of squares of complicated things.
Try it with position and momentum and the cancellation almost happens. The two cross terms do not quite kill each other, because the order in which you write position and momentum matters, and the difference between the two orders is the one thing quantum mechanics insisted on at the start. So the factorisation is correct apart from a single leftover term, and the leftover is precisely that difference.
Now look at the size of the leftover in units of energy. It is half of Planck's constant times the frequency, and it is exactly the amount by which the lowest state of an oscillator sits above the bottom of its well. That number has not been calculated yet. It has fallen out of an attempt to write a sum of two squares as a product, which is the closest this book comes to saying that the zero-point energy is a piece of arithmetic rather than a piece of physics.
Two operations have been built. One raises the energy of a state by a fixed step and the other lowers it by the same step, and the fixed step is Planck's constant times the frequency. That much is a matter of pushing one operator past another and collecting the remainder, and it involves no picture of the system at all.
The obvious worry is that lowering can be repeated forever, which would give states of arbitrarily negative energy and no lowest one. What rules that out is a fact about lengths. Apply the lowering operation to a state that sits on a definite rung, and the squared length of the result turns out to be exactly the rung number you started from. A squared length cannot be negative. So no rung is below zero, and the only way the descent can be stopped is by reaching a state the lowering operation sends to nothing at all.
That is the whole reason the energies of an oscillator are a discrete evenly spaced ladder starting half a step above the bottom of the well. The spacing comes from a commutator, the floor comes from a length being non-negative, and the half step comes from the failed factorisation of §2. It is worth noticing how little was used. Nothing here refers to a spring, a wavefunction, or the shape of the potential beyond the two lines it took to set the problem up.
Finding the states of an oscillator would ordinarily mean solving a second-order differential equation once for every energy level, with the requirement that the answer stay finite doing the work of picking out which energies are allowed. The method of this chapter replaces all of that with one first-order equation. Ask which function the lowering operation destroys, and you are asking for a function whose slope is minus its position times itself, which separates and integrates in three lines. The answer is a bell curve.
Every other state is then made by applying the raising operation to that bell curve as many times as you like. Each application multiplies by the coordinate and subtracts a derivative, so each one adds one power to a polynomial sitting in front of the same bell curve. Those polynomials are exactly the ones Chapter 4.5 built from a generating function, and the raising operation turns out to be their recurrence relation, which is a fact worth checking rather than assuming, since the two constructions had no reason to agree.
Two further things come out for nothing. Because a first-order equation has a one-parameter family of solutions, there is exactly one lowest state, and therefore exactly one state on every rung. And because reflecting through the origin reverses both position and momentum, it reverses the raising operation as well, so each rung is the opposite of the one below in its behaviour under reflection. Even, odd, even, odd, all the way up, established without drawing a single one of them.
Chapter 0.8 drew the classical oscillator as a point going round and round an ellipse, with position on one axis and momentum on the other, and worked out the area the ellipse encloses. It came to the energy divided by the frequency, times two pi. That chapter then said, in writing, that a later one would ration those ellipses, and this is the later one.
Feed the energies just derived into that area and the answer is Planck's constant multiplied by a half-integer. So the allowed orbits are not a continuum. They are a set of nested rings, each enclosing exactly one more unit of Planck's constant than the one inside it, and the innermost enclosing half a unit rather than none. That last half unit is the same zero-point energy that appeared when the factorisation failed, now wearing its geometric costume: the state cannot shrink to a point, and the smallest patch of phase space it can occupy has a definite size.
Chapter 1.3 guessed this rule and said so honestly at the time. What has changed is that the rule is now a consequence rather than a guess, in this one case, and the derivation used nothing about orbits. The general version, and an honest account of how badly it does on potentials that are not parabolas, is Chapter 4.10's business.
Every state built so far has a definite energy, and every one of them has a curious property: on average the particle is at the centre and not moving, and it stays that way forever. Nothing so far in the chapter oscillates, which is odd for a chapter about an oscillator.
The state that does oscillate is found by asking for something different. Instead of a state on which the energy takes a definite value, ask for one on which the lowering operation acts as multiplication by a number. That is a legitimate thing to ask for, because the lowering operation is not an observable, so nothing forces the number to be real, and such a state does exist for every complex number you name. Expanding it over the rungs gives coefficients that make its probabilities a Poisson distribution, of the kind that describes counts of independent rare events.
What this state does is the point. Its centre moves exactly along the classical path, with the right amplitude and the right phase, and its width never changes at all, which is the opposite of what a free packet does. It is a bell curve of fixed shape sliding back and forth in the well. The price is that it has no definite energy: its rung is uncertain by the square root of its mean rung, so a large one is sharply defined in relative terms and a small one is not. That is the sense in which a classical oscillation is a quantum state with a great many quanta in it and no idea how many.
4.9 Commutators, Uncertainty, and Symmetry
A statement that has been called the deepest in twentieth-century physics has just been obtained by multiplying an inequality by a constant. That is not a trick, and it is worth being clear about why it is honest.
The toolkit chapter on Fourier analysis proved that a signal cannot be both brief and pure in pitch. Squeeze it in time and its spectrum widens by the reciprocal factor, always, with the product of the two spreads never falling below one half. That result is about waves. It was known to people designing radios before anyone applied it to matter, and nothing in its proof mentions particles or measurement or observation.
Quantum mechanics contributes exactly one sentence to the story: a particle's momentum is the wavenumber of a wave, multiplied by a constant with the units of action. That sentence is a physical claim with an experiment behind it, and the earlier chapter marked it as such. Accept it, multiply, and the famous inequality drops out. The mystery, if you want one, is entirely in the sentence. It is not in the inequality.
One small piece of tidying comes free, and it is worth being exact about how much of it is free. The constant has the units of action, and the classical chapters had already established that a coordinate and the momentum belonging to it always multiply together to give an action, whatever the coordinate happens to be. So the bound written down for an angle and its angular momentum at least makes dimensional sense, in the same form, with the same constant and no conversion factor anywhere. That much is not luck. It is what "conjugate" was defined to mean.
Whether the inequality is then true for such a pair is a separate question, and the dimensions do not settle it. For the angle the answer is no, until the statement is repaired. The next section says which step breaks, and Worked example 2 exhibits the failure and mends it.
The general statement takes three steps and it is worth watching how little each one costs. The spread of a measurement, defined honestly as a standard deviation over repeated identical preparations, was shown two chapters ago to be the length of a particular vector built from the state and the quantity being measured. So take two quantities, build the two vectors, and apply the inequality proved in the toolkit chapter, which says that the overlap of two vectors is never bigger than the product of their lengths. That single application is the uncertainty principle.
What remains is bookkeeping on the overlap. It is a complex number, and its two parts have different meanings: the real part measures how the two quantities co-vary, and the imaginary part measures how badly the two operations fail to commute. Since a complex number is at least as big as either part alone, throwing away the real part costs nothing and leaves the familiar form, with the failure to commute sitting on the right-hand side. Keep the real part instead and you get a slightly stronger statement for the same work.
Two warnings, both earned. The first is that this is a statement about preparation. Every symbol in it refers to a spread across many systems prepared identically and each measured once. It says nothing about an apparatus knocking a particle about, and the theorems that do say something about that are separate results with different hypotheses, named above and not proved here.
The second is that the bound is a floor rather than a forecast. The lowest oscillator state sits exactly on it. The fourth one sits seven times above it, and is no less legitimate for that. And when the right-hand side happens to vanish in some particular state, the theorem has not announced that both quantities are sharp. It has fallen silent.
Two quantities that commute can be sharp together, and the states in which both are sharp are numerous enough to describe everything. That much was proved in the toolkit chapter. The question left open was practical: given a handful of such quantities, how do you know you have enough of them to tell every state apart?
The definition says you have enough when knowing all their values pins the state down. The useful reformulation says you have enough when the list cannot be extended: any further quantity that commutes with all of yours is already some combination of them, so adding it would tell you nothing new. The two conditions are the same condition, and the proof of that is half a page.
In practice nobody checks either one directly. You write down the possible combinations of values, you count them, and you compare that count with the number of independent states available. If the two numbers agree, the labels are enough. That is the whole method, and it is what lets an atomic state be written as three numbers in a bracket, those three labels being the whole of what distinguishes it from every other state.
Everything the theory predicts is an average of the form "state, operator, state". There are two places the time can be kept in such an expression, and both give the same number. Keep it in the state and the state evolves while the measured quantities stand still, which is the picture used so far. Move it onto the quantity instead and the state is frozen while the operators evolve.
This is not two theories. The manoeuvre is exactly the change of basis met in the linear algebra chapter, where one map acquires a different array of numbers in a different basis while remaining one map. Failing to see that is why the two pictures look like rival physics. The check is that the list of possible measured values, which is a property of the map and not of the basis, does not budge.
Differentiating the moved operator gives its equation of motion, and the answer is a bracket with the energy. Set it beside the classical equation of motion for an arbitrary quantity, derived in Part I from Hamilton's equations, and the two are identical except that one uses the classical bracket and the other uses the commutator divided by . Nothing else differs. That single replacement is what the phrase "canonical quantisation" refers to.
The immediate reward is that conservation becomes a computation rather than an investigation. A quantity is conserved when it commutes with the energy, and the statement is stronger than its classical counterpart: not merely the average but the whole distribution of measured values stays put.
Feeding position and momentum into the equation of motion produces two statements that look exactly like the classical laws with averages written over everything. The rate of change of the average position is the average momentum over the mass, and the rate of change of the average momentum is minus the average force. This is the result an earlier chapter stated without proof and attached a warning to, and the warning is the more valuable half.
The warning is this. "The average of the force" is not "the force at the average position". A packet has width, it samples the force over a range, and unless the force varies in a straight line across that range the two quantities differ. The size of the difference is set by the curvature of the force and by the square of the packet's width, and the width cannot be reduced to nothing without the momentum spread exploding.
So the classical-looking reading is exact only when the force is a straight line in position, which means the potential is at most quadratic. The cases are exactly the quadratics: no force at all, a uniform force, a spring, and the inverted spring of a ball balanced on a hilltop. Every other potential in physics fails the test, and that is a large part of why the spring turns up everywhere in this subject as the case that can be trusted.
Two runs of the same integration, differing only in the potential, make the split visible. Both satisfy the theorem to within a millionth. Only the quadratic one has the centre of its packet obeying Newton's second law; the quartic one misses by roughly as much as the force itself. And exactness in the mean is a modest property in any case: a stationary state of the spring sits with its average position at zero forever, satisfying the classical equation with both sides vanishing while behaving nothing like a swinging weight.
A symmetry is anything you can do to a system without changing any prediction. Chase that definition through the machinery and a chain of four links appears, each of them already built. Preserving predictions means preserving lengths, and provided the transformation is linear, which every family the section used visibly is, that forces it to be a rotation of the space of states. Linearity is an assumption at this link rather than a result, and the main text says so where it is used. A continuous family of such rotations has, by the theorem imported earlier in this part, a unique observable behind it that generates it. The effect of the family on any other quantity is a commutator with that generator. And leaving the energy alone, which is what makes the family a symmetry of the dynamics rather than a mere relabelling, is the same equation as the generator being conserved.
So a conserved quantity and a symmetry are one thing, exactly as the Part I chapters concluded for classical mechanics, and the argument is shorter here because the theorem doing the heavy lifting was proved elsewhere.
Three examples, and they are the same three the classical chapter used. Sliding everything along is generated by momentum, and momentum is conserved when the potential is flat. Turning everything about an axis is generated by angular momentum, and the commutator of two such generators is another one, which is the algebraic fingerprint of the fact that turns about different axes do not commute. Waiting is generated by the energy, which is therefore conserved with no computation needed. Discrete symmetries, like reflection, leave no generator behind, because any continuous family you could run through such a symmetry has to be built out of the symmetry itself and hands back nothing that was not already there. What they leave is a two-valued label rather than a continuous charge.
One thing has been missing throughout. Everything here describes how quantum quantities behave among themselves. None of it says how the classical world reappears when the scales get large, and the reason no equation here answers that is that setting the constant to zero in any of them produces nonsense. The next chapter takes the limit properly, and it also proves that the translation between classical and quantum quantities, which has worked every time it has been tried here, cannot be made to work for everything at once.
That is as far as the argument has been carried, and it is worth saying plainly where it stops rather than letting the page run out.
Four things have been built. There is a toolkit, which turned out not to be neutral, since defining a derivative as a linear stand-in rather than as a slope is what allowed one definition to survive several variables, then whole histories, then curved space. There is a mechanics in which forces are not primitive, one number is attached to each history the world might follow, and every continuous symmetry hands back a conserved quantity. There is a geometry in which one speed is the same for everybody with nothing for it to be measured against, and in which electricity and magnetism stop being two fields and become one object, sliced differently by observers in different states of motion. And there is a gravity that is not a force but the shape of the arena, cornered into its final form by the demand that the geometry side of the equation impose no condition on matter that matter has never been observed to obey.
What is owed is larger than what has been delivered. Part III is finished, so the account of gravity has now been run against a star, against light and against the universe, and the last of those returned the news that a growing universe has no conserved energy for anything to belong to. Nothing here has said what happens when a thing has no definite state, though the mathematics that answers it was built in Part 0 for reasons with no physics in them. Nothing has said why there are the particular forces there are, though the shape of the answer has been visible since the redundancy in the electromagnetic potentials was noticed and deliberately left alone. And nothing has said what to do about the one subject that has refused every method that worked for the others, which is the subject this whole arc has been walking towards. Those are the promises made here and not yet kept, and they are kept in the order the book takes them.