Part 0 · The Toolkit — Chapter 0.3
Series, Approximation, Orders of Magnitude
How physicists actually think, and why every theory in this book is the first few terms of another one.
Here is the thesis of this chapter, stated as hard as it can be stated: physics is almost never solved exactly. It is expanded. The list of physically interesting systems with a closed-form solution is embarrassingly short. It runs to two bodies under gravity, the harmonic oscillator, hydrogen, and a handful of others.
Everything else you have ever read about, from planetary orbits to the electron's magnetic moment, was obtained the same way. Someone identified a small parameter, expanded in it, and kept the terms that mattered.
Chapter 0.1 gave you the first term of such an expansion. It noticed, almost in passing, that the leftover error was quadratic. This chapter keeps going.
We turn that observation into a theorem with an error bound, and then into the single most productive habit in theoretical physics: find the small parameter, expand in it, keep what you need, and know what you threw away.
Several sentences in this book that sound like slogans are, underneath, claims about the first surviving term of a series. Here are three of them.
- "Newton is the low-speed limit of Einstein."
- "Classical mechanics is the limit of quantum mechanics."
- "The Standard Model is an effective theory."
Each one is, literally and unromantically, a statement about which term survives first. By the end of this chapter you will be able to read those sentences as mathematics rather than as reassurance.
Tools you'll need — Chapter 0.1: linearisation, higher derivatives, the definition of . Chapter 0.2: the fundamental theorem of calculus and integration by parts, which is the engine of §1.
1 · Taylor's theorem, derived
Almost every book states Taylor's series and then, some pages later, mumbles about a remainder. We will do it the other way round, because the remainder is the entire point. A series without an error estimate is not a result. It is a wish.
Start from the fundamental theorem of calculus, which says nothing more than "the function at the end equals the function at the start plus the accumulated change":
That is exact. No approximation has been made and none will be. Every step below is an identity.
Our next move is to integrate by parts, and there is exactly one trick in the whole derivation. It is worth naming before we use it: the trick is the choice of antiderivative. Integrating gives for any constant , and we get to pick . We pick it so that the upper boundary term vanishes. So take the antiderivative to be , which is zero at . Then
The boundary term at is zero by construction. At it is , and it enters with a minus sign in front of the bracket's lower limit, so it contributes . Now flip the sign inside the remaining integral, and we have
Let's stop and look at what we have. The first two terms are exactly the linear approximation of Chapter 0.1. The third term is the new thing. The vague phrase "smaller than " has been replaced by an explicit integral, and that integral is the error, written down in full, with no limits taken and nothing swept anywhere.
Now do it again. The goal this time is to peel one more polynomial term off that integral, so we integrate by parts a second time. In , let and , whose antiderivative again vanishes at the upper limit:
The underbraced boundary term is the next polynomial term, so put that whole line back into the result of the first round and we have
There is the that the zoom figure in Chapter 0.1 measured numerically. It was never a coincidence. It is the second boundary term.
The pattern is now visible. Each integration by parts peels off one more term of a polynomial and hands back an integral with one more power of and one more derivative on . Doing it times gives Taylor's theorem with the integral form of the remainder:
(As usual and , so the term is just .) Setting and renaming gives the Maclaurin form, which is what you will use nine times out of ten: .
Grind box — the induction, and the Lagrange form of the remainder
The induction. Define, for ,
Equation (0.3.1) says , which starts the induction. Integrating by parts with and antiderivative (again chosen to vanish at ):
Each application converts a remainder into "the next polynomial term plus the next remainder", so applying it times to gives (0.3.6) exactly. The only hypothesis used is that exists and is continuous on the closed interval, which is enough for integration by parts to be legal.
Lagrange form. The integral remainder is exact but not always easy to look at. Note first that, substituting ,
For the weight is non-negative on the whole interval, so the mean value theorem for integrals applies: there is some between and with
(For , reverse the orientation of the interval. The weight is then of one sign again, and the same conclusion follows.) This is the form you will actually use for estimates. You rarely know , but you can almost always bound on the interval.
Read the Lagrange form as a sentence: the error you make by truncating after the term is itself a Taylor term, the very next one, with the derivative evaluated somewhere you don't know. That single fact is why physicists can be so casual and so accurate at the same time. You do not need . You need a bound on the next derivative, and then
Two things control the error: how big the next derivative can get, and how fast falls. The factorial in the denominator is the reason a handful of terms is usually enough. In §4 we will meet the case where the derivatives grow faster than the factorial, which is where physics actually lives.
One warning to carry into §3, since it costs nothing to state now. The Taylor series is what you get by letting in (0.3.6). It equals precisely when . That is not the same as the series converging. Those are two different conditions, and the difference is the subject of a callout later in this chapter that you may find genuinely unsettling.
Stated without softening, the claim this chapter rests on is that physics does not usually solve its problems; it expands them. The systems with an exact closed-form answer could be listed on a postcard, and everything else you have read about was obtained by finding a small quantity, expanding in it, keeping the terms that mattered and knowing what was thrown away. That last clause is not an apology attached to the method. It is what makes it a method rather than a guess.
Which is why the error term is derived here first and the series second. A series offered without an estimate of what truncating it costs is not a result but a wish, and the derivation keeps everything exact at every stage: the remainder is written down in full, as an explicit accumulation, never approximated and never quietly dropped.
The most useful form of that remainder says something almost too convenient to believe. The error you commit by stopping at a given term is itself the very next term of the series, with the relevant derivative evaluated at some point in the interval that you have no way of identifying. You do not need to identify it. A bound on how large that derivative can get is enough, and this is precisely how a working physicist manages to be casual and extremely accurate at the same time.
2 · The essential expansions
Five expansions carry most of the load in physics. We derive each one rather than quoting it. The derivations take one line apiece, and you should never be in the position of half-remembering whether there is a factorial in the binomial series.
2.1 · The exponential
By definition (Chapter 0.1) is its own derivative, so and for every . Feeding that into the Maclaurin form:
Does it converge to ? To find out, use the Lagrange bound. On the interval between and we have , so . Once , every further step multiplies this by . So the bound is eventually smaller than a constant times , and it therefore goes to zero.
The series converges to the function for every real . It also converges for every complex , which is something we are about to exploit.
2.2 · Sine and cosine
Derivatives of cycle with period four: , which at evaluate to . Only odd powers survive, alternating in sign:
the second by the identical argument starting from (values at the origin). Convergence everywhere is immediate: every derivative of or is bounded by in absolute value, so by the same factorial argument.
Notice the leading term of . It is , which is the "small-angle approximation" you have used since school. Now notice the first correction, , which tells you exactly when to stop using it. At rad that correction is , a relative error of . That number is not a rule of thumb. It is a computed quantity.
2.3 · The logarithm
Differentiating repeatedly gives , , , and in general . At that is , and dividing by leaves :
The restriction is not decoration, and §3 explains where it comes from. The leading behaviour is one of the two most-used approximations in quantitative science. It is why small log-hazard-ratios read directly as percentages, a point we return to in the Familiar Ground box.
2.4 · The binomial series, for arbitrary real exponent
This is the workhorse. Let , with any real number at all, not necessarily a positive integer. Then , , and in general
Dividing each of those by is what defines the generalised binomial coefficient, and doing so gives us
When is a non-negative integer the product eventually hits a factor of zero, the series terminates, and you recover the finite binomial theorem you already know. For every other it runs forever. Two cases deserve to be committed to memory, because between them they cover essentially every square root you will ever have to linearise:
The second is with , and it is the relativistic factor. Put and you have, in one line, the expansion that Chapter 0.1 obtained by grinding out two derivatives. Now you also get every higher term for free, which is exactly what Worked example 2 will cash in.
Grind box — where the radius comes from, without complex analysis
For the logarithm you can bypass Taylor's theorem entirely and get an exact remainder from the finite geometric sum. For any and any , multiply out to verify
This is an algebraic identity, not an approximation: the finite sum telescopes when you multiply both sides by . Now integrate from to , using and :
For the integrand is at most , so and the series converges to the function. For the factor is no longer bounded by , so that argument does not run again unchanged. Substitute and bound over the whole interval, which gives as well. The bound degrades as , and that degradation is the analytic shadow of there. For the terms themselves blow up, so the series cannot converge at all, since a convergent series must have terms tending to zero. The radius is exactly , and here the reason is visible on the real line, because runs to at .
The binomial radius. Take the ratio of consecutive coefficients in (0.3.11):
By the ratio test of §3 the series converges for and diverges for , for every non-integer . Again the reason is visible: misbehaves at , one unit from the origin.
2.5 · Reference table
Here are all five in one place, together with the radius each one is good out to and the chapter where this book spends it.
| Function | Series about | Radius | Where it is spent in this book |
|---|---|---|---|
| Everywhere. Time evolution, partition functions, propagators | |||
| Small oscillations (0.8), string modes (7.4) | |||
| Pendulum potential (below), lattice dispersion | |||
| Entropy, log-likelihoods, running couplings (5.11) | |||
| Geometric resummation, and the propagator in 5.6 | |||
| Every square root | |||
| (2.5) | |||
| The Lorentz factor (2.5) |
2.6 · Euler's formula, and why complex numbers are not a convenience
The exponential series (0.3.7) converges for every real , and its convergence proof used only . So it converges just as well for complex arguments.
We therefore define for complex by that series. There is no other sensible choice, since any definition agreeing with on the real line and having a power series must be this one.
Now substitute with real, and use the four-fold cycle , which then repeats:
Splitting the sum into even and odd is legitimate because the series converges absolutely, meaning that the sum of converges on its own. Absolutely convergent series may be rearranged freely without changing their value. So now compare the two brackets with (0.3.8) and read off
which is not an identity between three unrelated functions but a statement that they were the same object all along, sorted by parity. Three consequences, each of which we will use repeatedly.
1. It has unit modulus. . So is a point on the unit circle at angle , and is arc length along that circle. The identity is the special case "walk half way round".
2. Multiplying by is a rotation. Take any complex number , thought of as the point in the plane, and multiply:
Those are precisely the components of the two-dimensional rotation matrix acting on . Rotation, one of the central objects of the rest of this book, is multiplication by a complex exponential.
That fact does not stay small. In Chapter 0.4 it becomes the statement that and are the same group. In Chapter 6.3, the fact that the phase of a quantum field lives on a circle is what generates electromagnetism.
3. It converts decay into oscillation. Chapter 0.1 left you holding a comparison. The clearance equation gives , while quantum time evolution gives . Structurally the two are identical. The only difference is the . Euler's formula now tells you exactly what that does:
A decaying exponential loses amplitude. An imaginary exponential keeps its length and changes only its direction. That is the whole reason quantum mechanics conserves probability: is the squared length of a vector that is being rotated, and rotations preserve length.
This is also why complex numbers are not an accounting convenience in quantum mechanics. Real exponentials can only grow or decay. You need the circle, and the circle is .
One free bonus, to show the machinery paying rent immediately. From , expand both sides with (0.3.14) and equate real and imaginary parts:
The trigonometric identities you were made to memorise are the statement that obeys the exponential law. There is nothing else to them.
Only a handful of expansions do most of the real work in physics, and the thing worth noticing about each is what its first correction is for. The leading term of the sine series is the small-angle approximation everyone uses; the term after it reports the size of the error being made, and therefore says exactly when to stop using the approximation. An expansion that comes with its next term is an approximation carrying its own warranty, and that is the whole difference between a rule of thumb and a controlled statement.
The genuine surprise here comes from noticing that the exponential series never asks what kind of number its argument is. Feed it an imaginary quantity and sort the resulting terms by whether their power is even or odd, and the even ones assemble into the cosine while the odd ones assemble into the sine. Those were never three separate functions. They are one object, sorted by parity.
This collects the promise left standing at the end of the first chapter. Multiplying by the imaginary unit is a quarter turn in the plane, so the equation that made a concentration shrink away becomes, once that single factor is inserted, an equation that rotates something at fixed length instead. Quantum mechanics conserves probability because probability is the squared length of a thing being rotated, and rotations do not change lengths.
3 · Convergence, radius, and a genuine surprise
A power series is an infinite sum, and infinite sums are promises. We need a test for when the promise is kept.
3.1 · The ratio test
Suppose . There are two cases, and we take them in turn.
If , pick any with . Then beyond some index we have , hence . The tail of the series is therefore dominated term by term by the geometric series , which is finite. A series of positive terms with bounded partial sums converges, because its partial sums increase and are bounded above.
If instead the terms eventually grow, so they cannot tend to zero, and a series whose terms do not tend to zero cannot converge. Hence:
Apply this to a power series by putting . The ratio is , so if the series converges for and diverges for . That number is the radius of convergence.
Try it on two of §2's expansions. For , , so the radius is infinite. For , , so the radius is , as the grind box already showed by hand.
3.2 · The surprise
Here is a function with nothing whatsoever wrong with it:
(the series follows from with , or from the same exact finite-sum identity used in the grind box). On the real line is smooth, positive, bounded by , infinitely differentiable, and utterly featureless. It is the Lorentzian lineshape, and it does nothing interesting at or anywhere else.
Yet by the ratio test its series converges only for . At the terms are , and the "approximation" runs away to infinity no matter how many terms you take.
Something is setting a radius of , and it is not visible on the real line. Watch it happen:
3.3 · Where the boundary actually is
The series in (0.3.19) makes perfect sense with a complex variable in place of , and there the answer is immediate: blows up at , both at distance from the origin. In fact at every term of the series equals , so the partial sums are . The series does not merely fail to represent the function there. It diverges outright.
That is enough to pin the radius, given one small lemma: the set where a power series converges is a disc. If converges then its terms are bounded, say , and for any with ,
which is a convergent geometric series. So convergence at one point forces absolute convergence everywhere strictly closer to the origin. Contrapositive: divergence at forces divergence at every . The radius is exactly , and it was determined by two points that do not lie on the real line at all.
The real line does not know why its own series fails. The complex plane does. A real function can look flawless over its entire domain and still have a Taylor series that gives up at a specific radius. The radius is set by the nearest singularity in the complex plane, and that is a place your real-valued function never visits and never reports on.
This is not a curiosity. It is the reason that analyticity, poles and branch cuts become physical objects later in the book. In Chapter 5.11 a scattering amplitude is treated as an analytic function of complex energy, and then the dictionary reads:
- poles on the real axis are stable particles,
- poles just off it are resonances, with lifetimes given by the imaginary part,
- and branch cuts are thresholds where new particles can be produced.
The analytic structure in a region you cannot experimentally visit controls what you measure in the region you can. That entire subject is this footnote about , taken seriously.
An infinite sum is a promise rather than a completed act, so there has to be a test for when the promise is kept, and comparing each term with the one before it settles nearly every case that arises in practice.
Then comes the unsettling part. Take a function that is smooth, bounded, positive and completely uneventful along the whole real line, and its series still gives up at a definite distance from the origin, beyond which each additional term makes the answer worse rather than better. Nothing whatsoever happens to the function at that distance. Nothing on the real line accounts for the boundary, and no amount of looking harder along the real line ever will.
The explanation lies off the line entirely. Allow the variable to be complex and the function has two points where it blows up, both at exactly that distance from the origin, and the radius is set by whichever bad point is nearest, including bad points in directions the original problem never visits. This is the first time the description you have been using turns out to be a restricted view of something larger, and it will not be the last. It becomes physics later on: the structure of a scattering amplitude in a region no experiment can reach is what fixes the particle masses and lifetimes measured in the region experiments can.
4 · Asymptotic series — where physics really lives
Everything so far has been about series that converge. Now the uncomfortable truth. Most of the series physics actually runs on do not converge, for any value of the parameter, ever. They are still the most accurate predictions humanity has made.
That apparent contradiction is resolved by a single idea, and this section is where the chapter earns its title.
4.1 · Two ways to take a limit
A power series statement involves two variables, the argument and the number of terms , and there are two different things you can mean.
Convergent: fix , let , and demand the partial sums approach . That is what §3 was about.
Asymptotic: fix , let , and demand the error die faster than the last term you kept. Formally, as means that for each fixed ,
These are genuinely different demands, and the two limits do not commute. A series can satisfy (0.3.21) for every while diverging for every .
Put the two questions side by side. The convergent case asks "does adding terms help forever?". The asymptotic case asks "does the approximation improve as the parameter gets small?".
Physics almost always wants the second question. In physics the parameter is small and handed to you, whether it is or or the fine structure constant, and the number of terms is a choice you make.
4.2 · A completely explicit example
Consider
This integral is perfectly well defined and finite for every : the integrand is positive and bounded by , so .
We want a series in , so expand the denominator with the exact finite-sum identity from the grind box, with , and integrate term by term using (integrate by parts times):
Everything in those two lines is exact. Our next goal is to put a bound on that remainder, and since for , it obeys
The right-hand side is precisely the magnitude of the first term you dropped. That bound is beautiful, and it is the practical rule of thumb for asymptotic series: the error is no bigger than the first omitted term.
Note what has and has not been shown. Here it is a theorem, because the remainder came out as an explicit alternating integral we could bound. For a general asymptotic series it is a reliable heuristic and not a theorem, and the honest general statement is only . The bound also proves (0.3.21) immediately, since . So the series is asymptotic to .
Now apply the ratio test to . The ratio is , which tends to infinity for every . The radius of convergence is zero. The series diverges everywhere except the single point .
Both statements are true at once, and the resolution is in the shape of the terms . Multiply one term by to get the next. While that factor is less than one and the terms shrink. Once it exceeds one and they grow, forever. The series is useful up to the turning point and poisonous after it.
Grind box — optimal truncation and the accuracy floor
We want the minimising the error bound . Take logarithms and use the crude Stirling estimate, which we can derive on the spot by comparing a sum to an integral:
(The comparison is legitimate because is increasing, so the sum is trapped between and . We need only the leading terms.) Then
Differentiate with respect to , treating it as continuous:
It is a minimum because . Substituting back, at :
Check it against the figure. At this predicts stopping at about ten terms with an accuracy near . The numerically computed optimum is with error . At the predicted floor is and the actual optimum is with error . The exponential scaling is exactly right. The factor of a few comes from the subleading in Stirling that we discarded.
Remember this shape. An asymptotic series in a small parameter delivers accuracy and no better. That exponentially small leftover is not noise. It is a real physical effect that the power series is structurally incapable of seeing, as the warning box below explains.
4.3 · The punchline
Now the statement this section exists for.
The perturbation series of quantum electrodynamics is asymptotic and divergent. It predicts the electron's magnetic moment to about twelve significant figures. That is the most accurately confirmed prediction in the history of science.
The expansion parameter is the fine structure constant , and the series for the electron's anomalous magnetic moment runs in powers of . The measured value is , which is twelve significant figures.
The calculated value has been worked out through five orders in , with over twelve thousand Feynman diagrams at the fifth order. It agrees within the combined uncertainties, at the level of about one part in .
That the series diverges is not a suspicion but an argument, due to Dyson in 1952, and it is worth sketching because it is pure §3 reasoning. If the series in had any nonzero radius of convergence, it would also converge for small negative , since a radius means a disc.
But with like charges attract. The vacuum could then lower its energy without bound by spontaneously separating clumps of like charge, so there is no stable ground state and no sensible theory at all. A function cannot be analytic at a point on the boundary between "fine" and "catastrophically ill-defined". Hence the radius is zero. (The mechanism, if you want one: the number of Feynman diagrams at order grows like , exactly as in our toy integral.)
So how is anyone getting twelve digits? By the grind box. With , optimal truncation sits near terms and the accuracy floor is . Nobody is anywhere near the turning point. The terms are still falling steeply at fifth order and will keep falling for another four hundred.
Physicists are not being sloppy with a divergent series. They are using an asymptotic expansion correctly, in the regime where the error bound is the first omitted term and that term is astronomically small. Chapter 5.11 makes this precise and shows what the leftover actually is.
Everything up to here has concerned series that converge, and most of the series physics actually runs on do not converge, for any value of their parameter, ever. They remain the most accurate predictive tools humanity has built, and the contradiction is only apparent. Dissolving it takes one distinction.
There are two different things a statement about a series can mean. One fixes the parameter and asks whether adding terms forever eventually arrives at the answer. The other fixes the number of terms and asks whether the approximation improves as the parameter is made small. Physics almost always wants the second, because the small parameter is handed to you by nature while the number of terms is your choice, and a series can satisfy the second demand perfectly while failing the first everywhere.
Such a series behaves in a characteristic way. Its terms shrink for a while, reach a smallest term, and thereafter grow without limit, and while you are still in the shrinking regime the error is no larger than the first term left out. So the series has a best place to stop and a best accuracy it can reach, both calculable in advance. The perturbation series of quantum electrodynamics diverges, and it predicts the electron's magnetic moment to twelve significant figures, because nobody is remotely near the turning point and the first neglected term is minute.
5 · Dimensional analysis and orders of magnitude
The last tool in this chapter costs almost nothing and repeatedly returns answers that look like they should have required a calculation.
5.1 · Homogeneity, and why it is not a convention
Every additive term in a physical equation must have the same dimensions. The reason is not etiquette. A physical law must hold whatever units you measure in.
Suppose an equation reads where is a length and a time. Switch from metres to feet and changes by a factor of while does not. An equation true in one unit system would then be false in another, which would make the choice of units a piece of physics. It is not. Hence , always.
In mechanics we track three independent dimensions, mass , length and time , and write , , and so on.
The same argument in its stronger form is the Buckingham π idea. If a physical relation involves dimensional quantities built from independent dimensions, it is equivalent to a relation among just dimensionless combinations.
The reason is simple enough to state in one sentence: you may use of the quantities to define your units, setting them to , after which only genuinely independent numbers remain to be related. Everything below is that sentence applied.
5.2 · Demonstration: the pendulum
What can the period of a pendulum depend on? Plausibly the length (), the bob mass (), and gravity (). That is quantities and dimensions, so there is dimensionless group and the physics must be that this group is a constant.
Start with the mass. It is the only quantity carrying , so no dimensionless group can contain it, since any power of would leave an uncancelled behind. That is a conclusion before any calculation: the period does not depend on the mass. Now seek :
so for some pure number . Two lines, no differential equation, and we have the scaling and the independence from mass.
What dimensional analysis cannot give you is . It happens to be , and getting it requires actually solving the equation of motion, which we do in Worked example 1.
Worse, dimensional analysis is blind to dimensionless parameters. The amplitude is already a pure number, so nothing forbids from being a function . It is. The period of a real pendulum grows with amplitude, and the "constant" is only constant in the small-angle limit.
So the method tells you the structure and is silent about everything dimensionless. That is precisely why had to be measured rather than derived.
5.3 · The payoff: the Planck scale
Now something that looks like it should be impossible. Take the three constants that define the three great theories:
| Constant | Belongs to | Dimensions |
|---|---|---|
| Relativity | ||
| Quantum mechanics | ||
| Gravity |
Ask for a length: . Matching exponents of , , gives three linear equations:
From the first, . The third then gives . Substituting both into the second, , so and . Putting those exponents back gives the Planck length, and with it a time and a mass:
The mass looks unremarkable at first, about the mass of a flea's egg. Then convert it to an energy: , roughly times the energy the Large Hadron Collider delivers per collision.
Now the part that matters. The three equations (0.3.26) form a linear system whose coefficient matrix has determinant
A nonzero determinant means the system has exactly one solution. That is true for a length, and equally for any other target dimension. So is not a length you can build from . It is the only one.
There is no dimensionless knob to slide it, no free parameter, no alternative combination. Combine gravity with quantum mechanics and relativity, and a length appears whether you wanted it or not.
That is the content of the sentence in Chapter 7.1 saying that quantum gravity has a built-in scale. It is worth contrasting with electromagnetism, where the analogous attempt fails.
Take , and the Coulomb constant (dimensions , i.e. energy × length). From those three you can build no length at all, because they are dimensionally dependent. The combination is already a pure number, so the corresponding determinant vanishes.
That absence of an intrinsic scale is the deep reason QED behaves so differently from gravity, and it is why the gravitational coupling being dimensionful is exactly the disease diagnosed by power counting in Chapter 7.1.
5.4 · Orders of magnitude
The habit that goes with dimensional analysis is estimating first and computing later. You decide whether an effect can possibly matter before spending a week on it. Here is one example, using only what is on this page. How does gravity compare to electricity inside an atom? Both forces fall as , so the ratio is a pure number independent of separation:
Forty orders of magnitude. That is why no chemistry textbook mentions gravity. Read the other way, it is also why quantum gravity is experimentally out of reach. To make gravity comparable to the other forces you must reach the Planck energy, which is where (0.3.27) said it would be. A single ratio, computed in one line, explains both the structure of Part IV and the predicament of Part VII.
Before spending these tools, one expansion from outside physics, where the mathematics is not analogous but identical.
You have been performing first-order Taylor expansions in clinic for your entire career, under the name "the odds ratio approximates the relative risk when the outcome is rare." Let us actually do the expansion, because it tells you something the rule of thumb does not.
With event probabilities (treated) and (control),
So the two differ by exactly the factor , with no approximation yet. Now write with , and expand that factor in the small quantity using the geometric series :
The middle expression is exact. The last is the first-order expansion. Read the leading term: the fractional discrepancy between OR and RR is approximately (baseline event rate) × (RR − 1). Two consequences follow that the usual "if the outcome is under 10%" formulation cannot express.
1. The rate is only half the story. The error is controlled by the product , so a large treatment effect degrades the approximation just as fast as a common outcome does. Take a control rate of , comfortably "rare", together with . The exact factor is , so the odds ratio is against a relative risk of , an overstatement of . The rule of thumb waves this case through. The expansion does not.
2. It tells you the rate of decay, not just a threshold. Fix and walk the baseline rate upward. At the discrepancy is . At it is . At it is , meaning against .
Notice also that the first-order estimate itself degrades. It predicts , and against true values of , and . The expansion is warning you, in the only way an expansion can, that by its own higher terms have taken over and the approximation should be abandoned. A good expansion tells you when to stop trusting it. As §4 showed, that is the same service the first omitted term performs for QED.
The second everyday expansion is (0.3.9). A hazard ratio is reported as , and is why a log-hazard of is read off as "a 5% reduction". It is really , a reduction, and the error is the next term, . At the same reflex gives "50% increase" for a true , which is a increase. Same series, same failure mode, same fix: look at the next term.
You would be forgiven for assuming that a function with infinitely many derivatives is determined by them. It is not, and the counterexample is one you can check by hand.
Define for and . This function is smooth everywhere, including at the origin, where it flattens out so aggressively that it beats every polynomial. Differentiating repeatedly gives for some polynomial . That is true for , and each derivative of such an expression is again of that form by the chain and product rules. To evaluate the derivatives at the origin, use the definition:
where , and the limit vanishes because beats any power: from the exponential series with all-positive terms, , so once .
So every Taylor coefficient at the origin is zero. The Taylor series of is , which converges beautifully, everywhere and absolutely, to the function . And is not : . The series converges, and it converges to the wrong function.
Recall the warning after (0.3.6). What matters is whether , and here for every , unchanging and never small. Smoothness is not analyticity.
Now the physics, because this is not a pathology. It is a sector of reality. Perturbation theory produces a power series in a coupling . Any quantity of the form has, by the argument above, an identically zero power series at . It is invisible at every order of perturbation theory, to all orders, forever.
Such terms are exactly what the accuracy floor of §4 was made of. An asymptotic series bottoms out at because that is the size of what it cannot see.
And those terms are real.
- Quantum tunnelling amplitudes go as .
- Instantons in Yang–Mills theory carry a factor .
- The confinement of quarks and the mass gap of QCD are non-perturbative phenomena of exactly this kind.
When Chapter 6.5 says that the proton's mass does not come from the Higgs but from the strong interaction's own dynamics, it is describing physics that lives entirely in the invisible sector defined by . The most important effects in the strong force are the ones that every term of the series says are zero.
The last tool in the chapter costs almost nothing and returns answers that look as though they should have required work. It rests on one observation: the units you measure in are a choice you made, so a law of physics cannot depend on them, and that demand alone constrains what an answer may look like. Applied to a pendulum it delivers the dependence on length and on gravity and rules out any dependence on the bob's mass, before an equation of motion has been written.
Pushed harder the method produces something startling. Take the constant belonging to relativity, the one belonging to quantum mechanics and the one belonging to gravity: there is exactly one length you can build from them, not one among many. It is forced into existence, it is unimaginably small, and it is why a quantum theory of gravity arrives with a built-in scale no experiment can approach. What it can never give you is a pure number, which is why some constants of nature must be measured rather than derived.
A closing caution, since the chapter has been one long argument for expanding things. Having every derivative is not the same as being determined by them: some effects are so flat near the origin that every term of every expansion reports them as zero. They are not fictions. The mass of the proton lives there.
6 · Worked examples
Two examples, one from each half of the chapter. The first uses Taylor's theorem to explain why the harmonic oscillator keeps reappearing. The second uses the binomial series to go one term past Newton.
Take any smooth potential energy with a minimum at . What is the motion of a particle released nearby?
Expand about using (0.3.6), writing :
Now kill the terms one at a time.
The constant is irrelevant. Only differences of potential energy have physical consequences, because force is and adding a constant changes no derivative. Drop it.
The linear term vanishes, because it is a minimum. is the definition of a stationary point. This is the crucial step and it is not an approximation: at an equilibrium the first-order term is exactly absent. Equilibrium is precisely the condition that makes the quadratic term leading.
The cubic and beyond are small. The ratio of the cubic to the quadratic term is , so they are negligible provided the amplitude satisfies . That is a quantitative criterion, not a hope.
Dropping the constant, setting the linear term to zero and keeping only the quadratic term, what is left is
and Newton's second law reads . From Chapter 0.1 we know the function that is minus its own second derivative, so with
Every stable system, close enough to equilibrium, is a harmonic oscillator with spring constant . Not "can be modelled as". It is one, to leading order, with an error you can bound. The harmonic oscillator is not a special case that physics happens to like. It is the generic case, the first surviving term of the expansion of anything stable.
Concretely: the pendulum. A bob on a rigid rod of length has height below the pivot, so and the kinetic energy is . The effective mass for the coordinate is therefore . Then , which vanishes at as required, and . Feed those into the formula for and we get
There is the that §5 said dimensional analysis could not supply, and the mass has cancelled exactly as dimensional analysis promised it must.
You can reach the same place by expanding the cosine with (0.3.8), giving : a constant, then the harmonic term , then a quartic correction whose effect is to make the period depend on amplitude at relative order . The expansion cannot give that coefficient without the perturbation methods of Chapter 0.8, but it already tells you the form. The correction is quadratic in amplitude, which is why a pendulum clock keeps time provided you keep the swing small.
Where this goes. This one paragraph is why the harmonic oscillator will not leave you alone.
- Chapter 0.8, the normal modes of coupled systems.
- Chapter 4.8, the quantum oscillator and its ladder operators.
- Chapter 5.3, where a free quantum field is an infinite collection of oscillators, one per momentum mode, which is where particles come from.
- Chapter 7.4, the string's vibrational modes, which is where the graviton comes from.
Quantum field theory is built out of oscillators for exactly the reason on this page: expand any stable field configuration about its minimum, and the quadratic term is what you get first.
Chapter 0.1 expanded far enough to find . With the binomial series (0.3.12) we can now go as far as we like at no extra cost. Put into :
How good is the truncation? At the exact is , while and . The error after two terms is , and the first omitted term is . Those two agree to two figures, exactly as the remainder theory said. Relative to the kinetic energy the correction is , which is at a tenth the speed of light.
A sign that catches people out. Quantum mechanics works with momentum, not velocity, and the relativistic momentum is , not . Expanding the energy in instead requires rather than :
The correction is now negative. Both results are correct. They disagree because and are not proportional once , so "fourth order in " and "fourth order in " are different expansions. This is a standing hazard with series: the coefficients depend on what you chose to expand in, and only the physical answer is invariant.
Where the term is seen. The Hamiltonian correction is the leading relativistic contribution to atomic energy levels. In hydrogen the electron's typical speed is (Chapter 4.13), so this term is smaller than the Rydberg energy by a factor . Those are shifts of order , which together with the spin–orbit and Darwin terms make up the fine structure of the hydrogen spectrum. It was measured in the nineteenth century, before anyone knew what it was. It is the third term of a Taylor series.
7 · Your turn
Problem 1 · radius of convergence, twice
(a) Find the radius of convergence of the Maclaurin series of , first by writing the series explicitly and applying the ratio test, then explain the answer geometrically.
(b) Find the radius of convergence of , identify the function it represents, and say what happens at the two endpoints .
Solution
(a) Factor out the and use the geometric series with :
Successive terms are in the ratio , so the series converges when , i.e. : the radius is .
Geometrically: as a function of a complex variable, blows up where , that is at , both at distance from the origin. The disc of convergence is the largest disc centred at containing no singularity, and those two poles stop it at radius . On the real line the function is smooth and bounded everywhere, with a maximum of at the origin and no feature at . The boundary is invisible from the real axis, as in §3.
(b) Coefficients , so
hence . The function: from (0.3.9) with ,
which indeed has a singularity at , on the real axis this time, so no complex detective work is needed.
Endpoints. The ratio test says nothing when the limit equals , and here the two ends genuinely differ. At the series is , the harmonic series, which diverges, and that matches the fact that there. At it is , the alternating harmonic series, which converges (to ), and the function is perfectly finite there. So the boundary circle of a power series can behave differently at different points on it. The radius tells you about the interior only.
Problem 2 · the Lorentz factor to fourth order
Expand to order from the binomial series. Evaluate the truncation error at and check it against the first omitted term. At what speed does the term contribute of the kinetic energy, and where does this term show up physically?
Solution
With and argument , the coefficients from (0.3.11) are , , . Each is multiplied by , so all signs come out positive:
Numerical check at . Exact: . Truncated at : . Error . First omitted term: . The two agree to better than of themselves, so the remainder really is "the next term, near enough".
The speed. Kinetic energy is , so the ratio of the second term to the first is
So Newtonian kinetic energy is good to up to about of light speed. That is , faster than anything in the solar system and slower than everything in a particle accelerator. That single number explains why relativity went unnoticed until people started looking at light and electrons.
Physically. Rewritten in momentum, the same expansion gives the correction to the Hamiltonian (Worked example 2), which is the leading relativistic contribution to atomic fine structure, of relative size in hydrogen.
Problem 3 · how far can you push the rare-outcome approximation?
A trial reports 2-year recurrence in of controls and on treatment. Compute the relative risk and the odds ratio exactly, express the discrepancy as a percentage, and compare it with both the exact factor and the first-order estimate from the Familiar Ground box. What does the size of the second-order term tell you about reporting the odds ratio as though it were a risk ratio here?
Solution
With and :
The ratio is , i.e. the odds ratio sits below the relative risk.
Against the exact formula, with and (so ):
Exact, as it must be, since that expression involved no approximation.
Against the first-order estimate: , predicting a discrepancy where the truth is . The expansion of shows where the missing points come from: the second-order term is of the first, which is exactly .
Interpretation. Reported as an odds ratio, this treatment "reduces the odds by ". The actual reduction in risk is . At a baseline the odds ratio is not a stand-in for the risk ratio. It exaggerates the effect by a fifth of the effect size, and the exaggeration is always away from .
The expansion also tells you the fix without needing a new rule of thumb. The discrepancy scales as , so if you want it under with you need . Note finally that the first-order estimate was itself off. When the leading correction is , the next one is around –, and you are no longer in the regime where one term is enough.
Problem 4 · the Bohr radius from dimensions alone
(a) The hydrogen atom is built from , the electron mass , and the Coulomb constant (dimensions of energy × length). Find the unique length that can be built from them and evaluate it. What else can you get, and what can dimensional analysis not tell you?
(b) Show that , and alone cannot produce a length, and say why that is structurally different from the case of §5.
Solution
(a) , , . Seek :
The third gives . Substituting that into the second, , so , , and . Hence
which is the Bohr radius, correct including its numerical factor. That is a piece of luck, not a guarantee. The same three constants give exactly one energy, , and exactly one speed, .
What is fixed and what is not. Fixed: the combination, hence all the scaling. Double the electron mass and the atom shrinks by a factor of two, and that conclusion needs no quantum mechanics. Not fixed: every pure number. The ground-state energy is , which is of the natural energy scale, and dimensional analysis cannot produce that any more than it produced the pendulum's . Nor can it produce the spectrum's , since is dimensionless.
Notice also the third result. Dividing the natural speed by gives , the electron's typical speed in units of light speed. That is exactly the small parameter that made the relativistic correction of Worked example 2 of relative size .
(b) Try . The exponent matrix, with columns and rows , is
The determinant vanishes, so the system is singular: either no solution or infinitely many, and never a unique one. The kernel is the reason. The combination has zero dimensions, being the fine structure constant , so you may multiply any candidate length by any power of and get another candidate. Electromagnetism, quantum mechanics and relativity together define no length. You must import a mass, such as , before the atom has a size.
The contrast. For the determinant was ((0.3.28)): the three constants are dimensionally independent, they fix a complete system of units, and a length exists whether or not anyone wants it. Gravity therefore arrives with its own scale built in, and the dimensionful coupling is exactly what makes the perturbative expansion of gravity break down at that scale. That is the power counting of Chapter 7.1.
You have Taylor's theorem with its remainder, derived by nothing more exotic than repeated integration by parts, in both integral and Lagrange forms. That is what makes every truncation you ever take come with a bound rather than a hope.
You have the handful of expansions that carry most of physics, including the binomial series for arbitrary real exponent. And you have Euler's formula as a consequence of the exponential series rather than as a slogan: is a rotation, which is why quantum mechanics conserves probability while first-order kinetics does not conserve drug.
Four further things you now know:
- The radius of convergence is set in the complex plane, even for a real function that never misbehaves.
- Most series in physics diverge and are used anyway, correctly, by truncating near the smallest term.
- Smooth is not analytic, and an entire sector of physics hides in the difference.
- Dimensional analysis gets you the structure of an answer for free while remaining permanently blind to pure numbers.
Where this gets spent. Taylor expansion → Chapter 1.2 (varying an action is expanding it to first order in a whole function), Chapter 2.5 (the Newtonian limit of relativistic dynamics), Chapter 5.8 (every Feynman diagram is a term in a series). Euler's formula → Chapters 0.8 and 0.9 (oscillators, Fourier analysis) and then the whole of quantum mechanics from Chapter 4.2 onward. Asymptotic series → Chapter 5.11, where the divergence becomes the renormalisation group. Non-analyticity → Chapter 6.5, where instantons and confinement live in the part of the answer perturbation theory cannot see. Dimensional analysis → Chapter 7.1, where is the reason quantum gravity is hard rather than merely unfinished. And small oscillations → Chapters 0.8, 4.8, 5.3 and 7.4, because "expand about the minimum and keep the quadratic term" is, in the end, why quantum field theory is built out of harmonic oscillators.