Part 0 · The Toolkit — Chapter 0.1
What a Derivative Really Is
The single most important reframe in the book, and it happens on page one.
You have met derivatives before, so this chapter is not going to reteach you the power rule. What it is going to do is swap one definition for another.
The definition you were given at school is that a derivative is a slope. We are going to set that aside and put a different one in its place. The new one is chosen because it survives everything this book will later do to it: curved spacetime, spaces whose points are entire functions, and eventually differentiation with respect to a whole field configuration.
Here is the replacement, stated once now so you can see where we are heading. The derivative is the coefficient of the best linear approximation to a function near a point.
Slope is what that sentence means in one dimension, on a flat page. It is a special case, and it happens to be the special case that does not generalise. Learn the general version now and Chapters 0.6, 2.4, 3.2 and 5.7 become far easier than they have any right to be.
Tools you'll need · None. This is the ground floor.
1 · The question that forces the idea
Physics is almost entirely the study of things that change. The first thing we need, then, is a way of saying how fast something changes. That sounds trivial until you try to be precise about it, and this section is about where the trouble lies.
Suppose a particle's position is . Between times and it moves a distance . Divide that by the time it took and you have the average velocity over the stretch:
There is no controversy in that. The trouble starts when you ask for the velocity right now, at the single instant .
The obvious move is to set in (0.1.1). Do it and you get , which is not a number. Yet an object clearly does have a speed at each moment. Your speedometer reads something. So there is a real quantity here, and the naive algebra cannot reach it.
The way out is to refuse to set at all. We ask instead what the ratio approaches as shrinks. That evasive-sounding manoeuvre is the whole of differential calculus.
It works for a reason worth naming now. The numerator and the denominator both go to zero, and the numerator does it no more slowly than the denominator does. Their ratio therefore settles down to something finite rather than blowing up.
Almost everything physics has to say concerns things that change, so the first thing the subject needs is a way of saying how fast something is changing right now, at a single moment, rather than on average across a stretch of time. That is harder to state than it sounds. Over a stretch you can divide the distance covered by the time taken and get an honest answer, but a moment has no duration, and no distance divided by no time yields nothing at all.
The way out is to stop forcing the arithmetic. Rather than asking what the ratio equals once the interval has shrunk to nothing, we ask what it is heading towards as the interval shrinks, and that question has an answer because the top of the fraction dwindles at least as fast as the bottom does, so the ratio settles down instead of running away.
This is the first appearance of a habit that runs through the whole book. A quantity you cannot reach directly gets defined by what a sequence of quantities you can reach approaches, and such a definition is worth having only when that approach is guaranteed rather than hoped for. A speedometer is evidence that instantaneous speed is real; this chapter is the account of why it is also well defined.
2 · The definition, stated carefully
Let's write the definition down, and then spend a moment on the one clause in it that carries all the weight.
The derivative of at the point is
whenever that limit exists.
That last phrase is doing real work, and it is not a formality. It means there is a single number such that the quotient can be forced as close to as you like, just by making small enough. It means one thing more as well, and that is the part which matters. The same has to come out whether you approach from the left or from the right.
A function can be perfectly continuous and still have no derivative at a point. By continuous we mean no jumps and no gaps, so that you can draw the graph without lifting your pen.
The standard example is at . Approach from the right and the quotient is . Approach from the left and it is . Two different answers, so there is no single limit and no derivative. The graph has a corner there, and "the best straight line through here" is genuinely ambiguous.
Worse is possible. Weierstrass constructed a function that is continuous everywhere and differentiable nowhere. It is infinitely crinkled at every scale, so that zooming in never straightens it out.
Such functions are not exotic curiosities. The path of a quantum particle in the Feynman path integral (Chapter 5.6) is exactly of this type.
Here is why that should concern you on page one. Physics assumes smoothness constantly and quietly, and the places where the assumption fails are reliably the interesting places: shocks, phase transitions, singularities.
Written out formally, the definition carries a clause that is easy to read past: the derivative exists whenever the limit exists. That clause is not legal boilerplate, and it is the most informative part of the sentence. It demands a single number that the ratio approaches however the interval shrinks, and in particular the same number whether you close in from the left or from the right. Where two different answers arrive, there is no best straight line through the point and nothing for the derivative to be.
Being able to draw a curve without lifting your pen is not enough to guarantee this. A graph with a sharp corner is perfectly unbroken and still gives two answers at the corner, and worse is possible: functions that are unbroken everywhere and have a derivative nowhere, wrinkled at every magnification, so that zooming in never straightens anything out.
Those are not museum pieces. Physics assumes smoothness constantly and almost always silently, and the places where the assumption fails are reliably the places worth studying, among them shock waves, phase transitions, and the paths a quantum particle takes. It is worth knowing from the first page that smoothness is an assumption being made rather than a property of the world.
3 · The reframe: derivative as linear approximation
Our goal in this section is to get the definition into a form that mentions no graph at all. The route is two lines of algebra applied to (0.1.2), and we will take them one at a time.
Saying that the quotient tends to is the same as saying that the difference between the quotient and tends to zero. Give that difference a name, :
What we want in the end is a statement about itself rather than about a quotient, so multiply that line through by and move the terms around:
Read (0.1.4) slowly, because it is the sentence this entire book is built on. Here is what it says. Near , the function is a constant plus a linear function of the displacement. The error you make by stopping there is not merely small. It is small compared to itself.
That last clause is the sharp part. Lots of approximations have small errors. This one has an error that vanishes faster than the thing you are approximating.
And is the unique number with that property. No other choice of coefficient makes the leftover shrink faster than linearly. Uniqueness is what lets us turn the whole thing around and use it as the definition:
is the unique number such that approximates with an error that vanishes faster than .
Notice what that sentence leaves out. There is no slope in it, no graph, and nothing tying it to one dimension. That is exactly why it travels. Each of the later derivatives in this book is the same sentence with one word exchanged.
- Replace by a displacement vector and by a linear map, and you have the derivative in several variables (Chapter 0.6).
- Replace the flat displacement by a tangent vector on a curved manifold, and you have the derivative in general relativity (Chapter 3.2).
- Replace the function by a functional and the number by a whole function, and you have the variational derivative that produces the Euler–Lagrange equations (Chapter 1.2) and eventually the field equations of the Standard Model.
Every one of those is this same sentence. "Slope" gets you none of them.
3.1 · Seeing it: local flatness
The claim in (0.1.4) is one you can watch happen, and the figure below lets you do it.
The reasoning behind the figure runs like this. If the error really does die faster than , then zooming in on the graph near must make the curve and its tangent line converge. Not merely look similar. They should become indistinguishable in a way that keeps improving the harder you look.
What happens here is a deliberate replacement, and it deserves to be flagged as one. The definition given at school, that a derivative is the slope of a graph, is being set aside for one that mentions no graph at all: the derivative is the number that makes a straight-line stand-in for a function as good as a straight line can be. On a flat page in one dimension the two agree, which is why slope was ever taught, but slope is the special case that will not survive curved space, several variables, or the later chapters in which the thing being varied is a whole history.
The precision that makes the new definition work lies in how the error is described. To say the leftover is small would be worth almost nothing, since nearly any approximation has a small error when the displacement is small. The claim is far stronger: the leftover shrinks faster than the displacement it accompanies, so that measured against the displacement it disappears entirely. Exactly one coefficient achieves that, and uniqueness is what turns a good guess into a definition.
The leftover does not vanish without trace, though. Divided by the square of the displacement it settles on a definite number rather than dying, so the error is specifically quadratic, and that one observation is the door into the two chapters that follow.
4 · The rules, derived
Almost every differentiation rule you were once asked to memorise falls out of (0.1.4) in a line or two. We will take three of them here, in the order sum, product, chain.
Deriving them once is worth more than memorising them ten times. The reason is practical rather than moral. The derivations are the part that carries over into settings where the memorised forms mean nothing.
4.1 · Sum
Linear approximations add, and that is the whole of it. If and , then .
The errors still need checking, since the strong sense of §3 is what we are claiming. Each of the two error terms is smaller than , so their sum is smaller than as well. Hence , and the same argument gives .
4.2 · Product — with a picture
Let , and picture it as the area of a rectangle with sides and . That picture is exact rather than decorative, and it will hand us the answer before we do any algebra.
Nudge by . The sides grow by and , so the area grows by three pieces: a strip along the top, a strip along the side, and a tiny square in the corner where the two growths overlap.
We are after a rate of change, so divide by and see which of the three pieces survive the division.
The first two terms give . The corner term is , and since and , that is roughly , which goes to zero. The corner is second order and it dies. What is left is
The corner square is the entire reason the product rule looks the way it does.
It is also your first meeting with a move you will make hundreds of times before this book is over: second-order terms die. It is worth noticing that this is a move rather than a law of nature. In Chapter 3.4 we transport a vector around a closed loop, and the second-order terms that refuse to die there are precisely the curvature of spacetime.
4.3 · Chain rule — where the reframe pays
This is the derivation that justifies everything said in §3, so let's go slowly.
We want at . Write for the point at which the outer function gets evaluated.
Start by nudging by and asking how far moves. The linear approximation of at answers that directly:
So the inner function hands the outer one a displacement of size . Our next step is to feed that displacement into , using the linear approximation of at :
The second line put in plus something smaller than , which sweeps every error into a single bucket measured against . That is now exactly the shape of (0.1.4), so we can read off whatever multiplies :
Notice what actually happened. We composed two linear approximations, and linear maps compose by multiplication. That is the whole content of the chain rule.
The same sentence keeps working as the setting gets harder. In several variables the linear maps are matrices, and matrices compose by matrix multiplication, so the chain rule becomes , a product of Jacobians (Chapter 0.6). In general relativity it becomes the transformation law that defines a tensor (Chapter 2.4). It is the same statement each time. Try getting there from "slope times slope".
Grind box — quotient rule, and the fine print on the chain rule
Quotient. Write and combine the product rule with the chain rule applied to , whose derivative is :
There is no reason to memorise this separately.
Fine print. The chain-rule argument above has a gap. Suppose for some sequence of . That is a perfectly legal thing for to do, and a constant does it at every step. The step "" then divides by something that may vanish.
The clean repair is to define
which is continuous at precisely because is differentiable there. Then holds for all , including , and
No division by zero anywhere. This is the kind of hole that does not matter for any function you will meet in physics. It is still worth seeing once that somebody noticed it and plugged it.
Every rule you were once asked to memorise falls out of that one sentence, usually in a line or two. Deriving them once is worth more than rehearsing them ten times, because the derivations are what carries over into settings where the memorised forms mean nothing.
The product rule is the clearest case. Picture a product as the area of a rectangle and nudge both sides outwards: the area gains a strip along the top, a strip along the side, and a small square in the corner where the two growths overlap. Each strip is proportional to the nudge and survives; the corner is proportional to the nudge twice over and dies. That is the entire rule, and discarding a second-order corner is a move you will make for the rest of the book. Much later, the pieces that refuse to die when a vector is carried around a closed loop turn out to be the curvature of spacetime.
The chain rule is where the reframe pays for itself immediately. Composing two functions near a point amounts to composing their straight-line stand-ins, and straight-line maps compose by multiplication. The same sentence holds with matrices in place of numbers when there are several variables, and in general relativity it becomes the rule that defines a tensor. There is no route to any of that from slope times slope.
5 · Where comes from
The exponential is not a function anyone chose out of fondness. It is forced, and this section shows what forces it.
Begin by differentiating straight from the definition, leaving the base unspecified:
Something remarkable is visible already, before any base has been picked. Whatever is, the derivative of comes out proportional to itself. Exponentials are the functions whose rate of growth is set by their current size.
All that distinguishes one base from another is the constant of proportionality , an awkward leftover sitting out in front. It depends on the base, with and . The natural move is to ask for the base that makes the leftover disappear.
Before taking that base for granted, let's check that it exists and that only one does. For rational , , and continuity fills in the irrationals. So is a continuous, strictly increasing function of the base on , with and as . A continuous increasing function that starts at and runs to infinity takes the value exactly once. That base is what we call .
So this is the definition of : the base for which the function is its own derivative. The numerical value is a consequence rather than a starting point.
The formula you already know falls straight out of it. Setting means , which says for small , so . Put and you have the familiar .
The logarithm follows in one line, by the chain rule. Start from and differentiate both sides. That gives , and therefore .
You already use (0.1.11) daily, under a different name. A drug cleared by first-order kinetics obeys
That differential equation says exactly what (0.1.10) said. The rate of change is proportional to the amount present. Nothing else has that property, which is why the solution had to be an exponential. The same three lines describe radioactive decay, a discharging capacitor, and the washout of a tracer.
Now here is the part worth filing away. In Chapter 4.6 the fundamental law of quantum mechanics will turn out to be
Structurally, that is your clearance equation. The only new ingredient is the factor of , and in Chapter 0.4 you will see that multiplying by is a rotation by 90°.
So the difference between a drug concentration decaying away and a quantum state evolving in time comes down to this. One exponential shrinks. The other rotates. Same equation, turned through a right angle. Hold onto that, because it is not a metaphor.
The exponential is not a function anyone was fond of; it is one that was cornered into existence. Differentiate any base raised to a power and something striking appears before a base has even been chosen: the rate of growth comes out proportional to the current size, whichever base you picked. All that distinguishes one base from another is an awkward constant left sitting in front, so the natural move is to ask which base makes that constant equal to one.
That request has exactly one answer, and the answer is the number we call e. Notice the direction of the definition. The value 2.718 and so on is not the starting point but the consequence; the starting point is the demand that a quantity be its own rate of change, and the number is whatever has to be true for that demand to be met. A great many objects in this book arrive the same way, forced into existence by a requirement rather than chosen for convenience.
You already rely on the pattern. Anything cleared at a rate proportional to how much is present decays exponentially, because nothing else can. Hold onto that shape: the fundamental equation of quantum mechanics has the identical form with one extra factor in it, and that factor turns shrinking into turning. Three chapters from here you will see exactly why.
6 · Trigonometric derivatives
Everything in this section rests on one genuinely geometric input, so let's isolate that input before spending it.
In the unit circle, compare three areas that sit one inside the next: the triangle with vertices at the centre, at and at the point on the circle at angle ; the circular sector spanning the same angle; and the triangle cut off by the tangent line at , which pokes out beyond the circle. For their areas are , and , so the comparison gives .
What the derivative is going to need is a statement about , so divide the whole chain through by :
and that limit follows because , which leaves the middle quantity trapped between two things both heading for .
This is the point at which radians enter physics permanently. Measure angles in degrees instead and the limit comes out as , so every formula downstream inherits the litter.
Grind box — from the squeeze to
We need a companion limit first, for the cosine terms that are about to appear. Using ,
With both limits in hand, expand the difference quotient for using the angle-addition identity:
The same manipulation on gives . Now note what follows from doing it twice: . That is a function proportional to minus its own second derivative, which is the harmonic oscillator equation. It is about to become the most reused equation in this book (Chapters 0.8, 4.8, 5.3, 7.4).
Trigonometry needs one genuinely geometric input before any of it can be differentiated, and the input arrives by trapping a circular sector between two triangles and comparing their areas. Everything else follows from that single squeeze.
Where a choice becomes visible is in the units. The clean result, that the sine of a small angle is very nearly the angle itself, holds only when angles are measured in radians, which is the unit in which an angle is a length along a circle rather than a turn cut into a convenient number of parts. Measure in degrees and a conversion factor lodges itself in the derivative and never leaves, so every formula downstream inherits the litter. This is the first instance of something the book returns to constantly, which is that a description always rests on a choice, and that one choice among many makes the structure visible while the rest bury it.
The consequence worth carrying away is a single line. Differentiate the sine twice and you get the sine back with its sign reversed, which is to say a quantity proportional to the negative of its own second rate of change. That is the equation obeyed by anything that oscillates, and from a few chapters onward it never really leaves the book.
7 · Higher derivatives, and what they are for
Differentiate again and you get , the rate at which the rate is changing. The first three of these are so common in physics that they have names: position, velocity , and acceleration .
Newton's second law is a statement about the second derivative, and that is exactly why mechanics is hard. You are told the acceleration and asked for the position. Getting there means integrating twice, and each integration picks up a constant. The two constants encode where the thing started and how fast it was going.
Geometrically, is concavity. The sharper statement is the one the zoom figure already showed you numerically, which is that is the leading error in linearisation,
Keep going in that direction and you get Taylor's theorem. That is Chapter 0.3, and it more or less runs the rest of physics.
7.1 · Notation, and one deliberate abuse
Four notations are in common use and you will need to read all four. They are not rivals. Each one is convenient for a different job.
| Form | Written | Best for |
|---|---|---|
| Lagrange | , | Pure functions of one variable |
| Leibniz | , | Anything where you must track with respect to what |
| Newton | , | Derivatives with respect to time, universally, in mechanics |
| Operator | , | When the derivative itself is the object of study, as in QM and QFT |
is not a fraction. It is a single symbol denoting a limit of fractions, and the limit has already been taken.
Yet every physicist you will ever read writes things like and cheerfully cancels differentials. They are not being sloppy. They are being terse about something true.
The justification is (0.1.4). The derivative is a linear map. Read that way, is an input displacement and the corresponding output displacement, so is a legitimate equation between linear maps rather than an approximation. Separating variables in a differential equation works for the same reason.
The proper definitions arrive later. In Chapter 0.6 the object becomes a one-form. In Chapter 3.5 it acquires an algebra, and starts to matter.
So use the abuse, because it is sound. Just know that you are cashing a cheque written in Chapter 0.6.
Differentiating a second time gives the rate at which the rate is changing, and the usual gloss, namely how sharply a graph bends, undersells it. The sharper statement is the one you see on magnifying the graph: the second derivative is the size of the leftover from replacing a function by a straight line. Bending is what that looks like when drawn; the leading error is what it is.
This is also why mechanics is hard, in a specific and honest way. Newton's law is a statement about the second derivative, so you are handed an acceleration and asked for a position, which means accumulating twice and picking up two unknown constants. Those constants are where the object started and how fast it was going, so a law of motion does not by itself determine a motion; law and starting conditions together do.
One last thing, because it looks like cheating and is not. Physicists write derivatives as though they were fractions and cancel the pieces freely, and the licence is real: the derivative is a map from a small input displacement to a small output one, so the manipulation is an exact statement about such maps rather than an approximation. You leave with the derivative in the form that generalises, and with a reason to keep chasing the leftovers, which two chapters from here becomes a method.
8 · Worked examples
A pendulum has period . You lengthen it by 1%. What happens to the period?
You could compute twice and subtract. Don't. There is a better route, and it starts by taking the logarithm, which turns the product-and-power structure into a plain sum:
Sums are easy to differentiate, which was the point of taking the logarithm. Since , every term turns into a fractional change:
Now put the numbers in. With fixed and , we get . The pendulum runs half a percent slow.
The general rule is worth having permanently. For ,
Exponents become sensitivity coefficients. This one identity does most of the error-propagation work in experimental physics, and it is why a quantity's exponent tells you at a glance how much that quantity matters. It is the linearisation idea of §3, applied to instead of to .
In Chapter 2.5 the energy of a moving particle will turn out to be with
Nothing about that looks like the kinetic energy you know. Our goal is to see the familiar expression come out of it, so let's linearise. Set and apply the chain rule to :
At that derivative is zero, so the linear term vanishes and the leading behaviour has to be quadratic. What we need next is therefore . Differentiate again and evaluate at , where only the first factor survives, giving . Feeding that into (0.1.13) with :
Let's look at what that last line is saying. The kinetic energy of your first physics course is the quadratic term in the linearisation of relativity. It was never a fundamental quantity. It was the second-order coefficient, and the that seemed arbitrary is the sitting in (0.1.13).
This is what it feels like when a new theory contains an old one. You expand, and the old theory is the first surviving term.
9 · Your turn
Problem 1 · inverse functions
Without using the formula for , show that if is differentiable and invertible then . Then use it to get and .
Solution
Differentiate the identity with respect to using the chain rule:
Geometrically, this says that reflecting a graph in the line turns a slope into its reciprocal.
Logarithm. , , , so .
Arcsine. , , so . Since on the principal branch and ,
That is not a coincidence. It is the same square root that appears in , and it is there for the same geometric reason.
Problem 2 · the error really is quadratic
Linearise at , i.e. approximate it by . Compute the exact error for and verify numerically that while . Explain the .
Solution
marches to zero, which confirms that the linear approximation is correct in the strong sense of (0.1.4). marches to instead, and by (0.1.13) that limit must be . ✓
The moral is worth stating plainly. "The error is small" is a weak statement. "The error is with known coefficient" is a usable one, because it lets you decide in advance how many terms you need.
Problem 3 · sensitivity in practice
The escape velocity from a body of mass and radius is . Using the logarithmic-derivative rule, estimate the percentage change in if is 3% larger and is 2% larger. Which quantity is more sensitive to?
Solution
, so
The exponents are , so is equally sensitive in magnitude to both. They carry opposite signs, so the two effects largely cancel.
Note how much faster this is than evaluating twice. It also tells you something the two evaluations would not, namely why they nearly cancel.
The linearisation is valid because 3% is small. At 300% it would be worthless, and you would need the higher terms of Chapter 0.3.
Problem 4 · half-life versus time constant
For , find the time to fall to of the initial value and the time to fall to . Show the ratio is and compute it. Then: a colleague says "the drug is 63% cleared after one half-life." Where did they go wrong, and what is the correct statement?
Solution
Setting gives the time constant . Setting gives .
The ratio is , so the time constant is about 44% longer than the half-life.
The colleague has fused the two. After one time constant , the fraction remaining is , so 63.2% is gone. After one half-life, by definition, 50% is gone. The 63% figure belongs to and not to , and the two differ by that factor of .
The confusion is persistent for a reason. is the natural constant of the differential equation, being what appears in , while is the natural constant for a human reading a chart. The mathematics prefers . Our base-ten intuitions prefer halves. And is the toll for crossing between them.
You now have the derivative in the form that generalises: . That is the coefficient of the best linear approximation, with an error that dies faster than linearly.
You have the rules as consequences rather than incantations, with the chain rule understood as composition of linear maps. You have defined by the property that makes it inevitable. And you have seen that quantum time evolution is your own first-order kinetics with an in the exponent.
Where this gets spent. The linear-approximation view is cashed in four places: Chapter 0.6 for the total derivative as a matrix, Chapter 1.2 for varying a whole path instead of a number, Chapter 3.2 for tangent spaces on a curved manifold, and Chapter 5.7 for differentiating with respect to a field.
Two smaller debts come due as well. The observation that the leftover error is quadratic is the door into Chapter 0.3. And , a function proportional to minus its own second derivative, is the harmonic oscillator, which from Chapter 0.8 onward never leaves.