Part 0 · The Toolkit — Chapter 0.9
Fourier, Delta Functions, and Probability
The uncertainty principle, three parts before quantum mechanics, as a theorem about integrals.
Part 0 closes here, and it closes with the one sentence that organises everything below: the Fourier transform is a change of basis (Chapter 0.4) in an infinite-dimensional inner-product space (Chapter 0.5), and the basis is chosen because it is the one in which the derivative operator is diagonal. That is the whole subject. Every formula in this chapter is a consequence of it. If a formula ever looks arbitrary, go back to that sentence and read it again.
Two earlier results are about to meet. Chapter 0.5 proved that a Hermitian operator has an orthonormal basis of eigenvectors, and that in that basis it is nothing but a list of numbers. Chapter 0.8 found that are the eigenfunctions of .
Now put those two together. Take the eigenfunctions of the derivative and use them as a basis, and what you get is Fourier analysis. From there, every linear differential equation in this book becomes an algebra problem. Differentiation becomes multiplication. Convolution becomes multiplication. Green's functions become division.
And then the payoff, which is the reason this chapter exists. In §6 you will define the width of a function and the width of its transform, and prove that their product cannot be smaller than one half. That inequality is the Heisenberg uncertainty principle. You will have derived it with no particles, no measurement, no , and no physics of any kind. The two ingredients are Cauchy–Schwarz from Chapter 0.5 and integration by parts from Chapter 0.2. Chapter 4.9 will add exactly one thing to it: the substitution .
Tools you'll need — Chapter 0.2: integration by parts, the Gaussian integral and its completed-square cousin, and Problem 4 (the transform of a Gaussian), which §6 cashes in explicitly. Chapter 0.3: Euler's formula, and Taylor expansion to second order with control of the remainder. Chapter 0.5: inner products, Cauchy–Schwarz with its equality condition, orthonormal bases and the coefficient formula , completeness , unitarity, and Dirac notation. Chapter 0.7: the divergence theorem and the diffusion equation. Chapter 0.8: that is the eigenfunction of , the driven oscillator, and the Duhamel/integrating-factor solution. Sections 3 and 6 each cash a specific theorem from Chapters 0.2 and 0.5. Nothing here is decoration.
1 · Fourier series as an orthonormal basis
Let's begin by giving functions the same treatment vectors got in Chapter 0.5. The plan for this section is short: define an overlap between two functions, check it is a genuine inner product, and then choose a basis. Chapter 0.5's table of inner-product spaces already had the row we need, and we skipped past it at the time. It was functions on an interval, with
Nothing new is being asserted here. This is the same definition Chapter 0.5 gave, with the vectors now being functions. Take the three requirements in turn.
- Conjugate symmetry, because .
- Linearity in the second slot, because integration is linear.
- Positive definiteness, because and vanishes only for .
Chapter 0.4 warned that a space of functions was coming, and that it would be infinite-dimensional. Here it is.
Everything Chapter 0.5 proved about inner-product spaces is therefore available to us. What we do not yet have is a basis. We are about to choose one, and the choice is not a matter of taste. §3 shows that the functions below are precisely the eigenfunctions of .
1.1 · The orthogonality relation, by direct integration
Work on the interval and define the wavenumbers
These are exactly the for which returns to its starting value after one interval length. Euler from Chapter 0.3 shows it in one line: .
We want to know whether two of these waves are perpendicular, so compute their inner product. There are two cases. For the integrand is , and the integral is just the length . For , write , which is not zero, and integrate directly:
because is a non-zero integer and of an integer multiple of vanishes. (The middle step used , which is Euler's formula subtracted from its own conjugate.) Both cases together:
That is the entire technical content of Fourier series. Divide by to normalise and the functions
That is an orthonormal set in exactly Chapter 0.5's sense , verified by integration rather than assumed.
1.2 · Coefficients are inner products
Chapter 0.5 proved that in an orthonormal basis, finding coordinates requires no work at all: . Transcribe it. Suppose . Take of both sides and use (0.9.5):
Not one integral had to be solved and not one linear system inverted. The orthogonality did all the work. Now write that out and fold the into the coefficient, and we arrive at the standard form:
1.3 · Completeness — quoted, and flagged as quoted
Everything above shows that the are orthonormal. It does not show that they are enough. Whether every function of interest is actually reachable as a sum (0.9.7) is a separate question, and the property being asked about is called completeness. We are going to use it without proving it.
Here is the honest statement. For every with , the partial sums converge to in the norm of (0.9.1), meaning . This is convergence. It is the version that matters for physics, because it is the version that makes the energy bookkeeping work.
It does not imply that at every point, and §1.4 exhibits a place where that fails permanently. Pointwise convergence needs extra hypotheses, the Dirichlet conditions: piecewise smooth, finitely many extrema and jumps. And at a jump the series converges to the midpoint of the two one-sided limits, rather than to either of them.
Both facts are proved in Chapter 4.3, where the space of square-integrable functions gets its proper name, a Hilbert space, and completeness stops being an assumption. Until then we are standing on it, and you know that we are.
1.4 · A worked expansion, and an honest failure
Take , so , and expand the square wave
Since is real and odd, it is easier to work with sines directly. Write , where . That orthogonality is the same computation as (0.9.4), rewritten with Euler.
Now find the coefficients. The integrand is a product of two odd functions, hence even, so the integral over is twice the integral over , where :
The bracket is for odd and for even , so
Let's look at what that last line is saying. A discontinuous function, built entirely out of smooth ones. That should bother you slightly, and it should. No finite partial sum is discontinuous, so the discontinuity is manufactured only in the limit. The manufacturing process leaves a permanent scar.
The Gibbs phenomenon. Near the jump at , every partial sum overshoots. Not by less and less: by the same amount, forever. The partial sum peaks at rather than , and as grows the overshoot narrows and slides toward the jump but never shrinks. Since the jump has size , the overshoot is of the jump, and that number is a constant of nature for every jump discontinuity in every Fourier series. The grind box derives it.
This is a genuine failure of pointwise convergence, and it is worth naming as one rather than apologising for it. It is also not a failure of convergence: the overshoot region shrinks in width fast enough that regardless. The two notions of convergence disagree, and the one physics needs is the one that survives.
Grind box — where comes from
Take the partial sum of (0.9.10) to terms (odd only) and differentiate it. The sum of cosines telescopes into something closed:
where . The closed form follows from summing the geometric series and taking real parts. Now integrate from to get back the partial sum itself:
The first maximum is at the first zero of the integrand's numerator past the origin, i.e. . For large that is small, so throughout the range, and substituting :
The number is not elementary. It is a value of the sine integral, and it is perfectly definite. Numerically, the peak of is at , at , at , at . It is converging, and not to .
Why it matters beyond aesthetics. To band-limit a signal is to truncate its Fourier content, and that is what every detector, every lens and every digital filter does. Any time you do it, you ring at the edges by of the step. Sharp edges and finite bandwidth are incompatible, and §6 makes that a theorem.
A space whose vectors are functions was promised when vector spaces were first defined, and it arrives here fully equipped. Give two functions an overlap by multiplying one against the other and integrating, and every result about lengths, angles and perpendicularity becomes available at once, no new theorem required. What is missing is a basis, and the one chosen is the collection of pure waves that fit a whole number of times into the interval.
That the waves are mutually perpendicular takes one short integration, and it is the entire technical content of the subject. Once it is done the earlier result about perpendicular bases takes over: each coefficient is a single overlap, computed on its own without reference to the others, so breaking a function into waves costs only a sequence of separate integrals.
Two cautions belong alongside it. Perpendicularity does not establish that the waves are numerous enough to reach every function, and that they are is quoted here rather than proved. And the sense in which a series of smooth waves reproduces a function with a jump is weaker than one would like, because near the jump every partial sum overshoots, by the same nine per cent of the step, however many terms are taken. A sharp edge and a limited supply of frequencies are incompatible, which stops being an annoyance and becomes a theorem shortly.
2 · From series to transform
Fourier series need a box. Physics on the whole line does not have one, so we take and watch what survives.
2.1 · The limit, taken carefully
The allowed wavenumbers (0.9.2) are spaced by
So as the box grows, the modes crowd together, and in the limit they form a continuum. The sum over in (0.9.7) is going to become an integral over . That works only if we first arrange for the sum to look like a Riemann sum, which means every term has to carry a factor of .
Manufacturing that factor is the whole job of this subsection. Start by defining, for a function on the whole line,
and notice that the coefficient in (0.9.7) is . Our goal is to get the series into Riemann-sum form, so substitute that expression for back into it:
Look at what appeared. The factor that had to be manufactured is exactly the mode spacing (0.9.11). So the right-hand side is a Riemann sum, of the kind Chapter 0.2 defined the integral to be the limit of. Let and the sum becomes that integral:
The two boxed formulas are the Fourier transform pair. Let's compare them with Chapter 0.5's . Then (0.9.12) is the coefficient , and (0.9.14) is the expansion, with the sum over a discrete index replaced by an integral over a continuous one. It is the same equation.
What has changed is that the basis is now labelled by a real number instead of an integer. That one change is the difficulty Chapter 4.5 has to work to make legitimate. The functions do not belong to the space they are supposed to be a basis of, because . Chapter 0.5's warning box said this would happen. §5 gives it the tool it needs.
2.2 · The convention, chosen and held
The has to go somewhere and there is no canonical place to put it. We have chosen the symmetric convention: a factor on each of (0.9.12) and (0.9.14), and the sign going forward, coming back. We will hold it for the entire book.
Other texts choose differently, and this is a real and recurring source of confusion. Factors of and appear and disappear between references for no mathematical reason whatsoever. Here are the three conventions you will meet.
| Convention | Forward | Inverse | Convolution theorem |
|---|---|---|---|
| Symmetric (ours) | |||
| Asymmetric | |||
| Ordinary frequency |
The symmetric one is chosen here for one reason: it makes the transform unitary, which is the next result, and unitary is the property Chapter 4.6 needs.
One further wrinkle is worth flagging now, because it catches people. In time-dependent problems physics usually writes for the forward direction in time, which is the opposite sign from our in space. The reason is that it makes a plane wave read . That is a convention about which way waves travel rather than a different transform. We will flag it again when it first matters.
2.3 · Plancherel: the transform is a rotation
Chapter 0.5's Parseval identity said that in an orthonormal basis, the squared length of a vector is the sum of the squared magnitudes of its coordinates. We want to know what that statement becomes on the whole line, so take it into the box first and then to the limit.
In the box, expand using (0.9.7), then collapse the double sum with the orthogonality (0.9.4). Only the diagonal terms survive:
Now we want the right-hand side in terms of the transform rather than the coefficients, so substitute and manufacture the mode spacing again:
Now look at the two ends of that chain. The left-hand side of the first line never depended on at all, so it needs no limit taken, and the right-hand side has just become an integral. Reading them together, we have proved Plancherel's theorem:
Let's read that structurally rather than as a formula. It says the map preserves norms. A norm-preserving linear map on an inner-product space preserves the inner product as well, because the inner product is recoverable from the norm. To see that, expand and and solve. So the transform preserves overlaps too:
(0.9.18) is exactly Chapter 0.5's , the defining property of a unitary map. So the Fourier transform is not merely a useful formula. It is a rotation in function space: it preserves lengths, preserves angles, and carries orthonormal bases to orthonormal bases. Nothing is created or destroyed by transforming. The object is the same, viewed along different axes.
Chapter 0.5 also proved that unitary maps are exactly the ones that conserve total probability. So this one line is the reason a wavefunction in Chapter 4.6 can be moved freely between position and momentum representations without the probabilities ceasing to add to one. Physicists say the Fourier transform "conserves probability". They mean it is a rotation.
Let the interval holding the function grow without limit and watch which parts of the description survive. The permitted wavelengths crowd together as the box lengthens, the discrete list of coefficients becomes a function of a continuous label, and the sum over modes becomes an integral, provided the bookkeeping is arranged so each term carries the spacing between neighbouring modes. It is the same expansion along perpendicular directions as before, the directions now labelled by a real number rather than a whole one, and that innocuous-looking change is the one thing here that later takes real work to make legitimate.
What comes through the limit untouched is the statement that total size is preserved. The squared size of a function, added over positions, equals the squared size of its transform, added over wavelengths, which says that transforming is a rotation of the space of functions rather than a computation performed upon them. It preserves lengths, it preserves angles, and it carries perpendicular bases to perpendicular bases.
Hold that in the form stated, because rotations are precisely the maps that keep a total probability equal to one. When a physicist remarks that the transform conserves probability, the content of the remark is that it is a rotation, and that nothing about the object has been altered by transforming it. Only the axes it is being described against have moved.
3 · Why Fourier, actually: the derivative becomes multiplication
Now the reason the subject exists. Everything so far could have been said about any orthonormal basis. This section says what is special about this one.
3.1 · The derivative theorem
We want to know what the transform does to a derivative, so transform and integrate by parts, which by Chapter 0.2 is the product rule run backwards:
The boundary term is dropped, and the assumption behind dropping it should be said out loud: as . Since exactly, the oscillating factor cannot help. The decay has to come from itself.
That assumption is worth translating, because it is a physical statement in disguise. It says the disturbance is localised. The wave packet is somewhere. The field dies away. The concentration is finite. It is the same hypothesis Chapter 0.2 flagged when it derived integration by parts, and granting it here leaves
and by iterating, . Differentiation has become multiplication.
3.2 · Why: the basis diagonalises
Do not accept (0.9.20) as a computational trick. It is the spectral theorem. Chapter 0.8 observed that exponentials are the eigenfunctions of the derivative, so write that observation in the form this chapter needs:
The operator acting on the function returns the same function times the number . That is Chapter 0.5's with , and . So the Fourier basis is the eigenbasis of the derivative operator.
Which is why (0.9.20) is nothing but Chapter 0.5's observation that an operator, written in its own eigenbasis, is a list of numbers. Here the list is , indexed by a continuum rather than by an integer.
A linear differential equation is an operator equation , where is built out of and multiplication by constants. In the Fourier basis every becomes , so becomes a polynomial in . That is an ordinary function, one number for each . The equation becomes
Solving a differential equation has been reduced to dividing. All the work has moved into transforming in and transforming back. This is exactly what diagonalising a matrix does for a linear system in Chapter 0.5, and it is the same theorem: a hard coupled problem becomes a list of trivial uncoupled ones, one per basis direction.
The honest caveat: Chapter 0.5's spectral theorem was proved in finite dimensions, and the argument above applies it to an operator on a function space, where the eigenfunctions are not even in the space. That gap is real. Chapter 4.5 closes it, and the price is a genuinely more careful theory. We proceed knowing what we are standing on.
3.3 · A demonstration: the driven oscillator
Chapter 0.8 fought the damped driven oscillator with trial solutions and complex amplitudes:
Our goal is to turn that into an equation with no derivatives in it at all, so transform in with conjugate variable , using (0.9.20) twice. Each time derivative becomes a factor :
The differential equation has become an algebraic equation, for each separately, and it is solved by division:
The function is the transfer function. It also goes by response function and by susceptibility, three names for one object. Everything Chapter 0.8 extracted by hand is now readable straight off it.
- The response amplitude is , which peaks where the denominator is smallest. That peak is resonance.
- The phase lag is , which passes through at .
One more reading is available if we look closely at the peak. Near resonance the denominator is approximately , so
That is the Lorentzian of Chapter 0.8, with half-width-at-half-maximum , now derived rather than fitted. Problem 4 turns this same shape into the natural linewidth of an excited state, and the connection runs through §6.
Grind box — what "transform the equation" is really doing, and two traps
It is a change of basis, and only that. In Chapter 0.5's language, an equation in the position basis becomes, after inserting completeness, an equation between components in the basis. Since is diagonal there, the matrix equation degenerates to with no sum at all. That is one equation per mode, each solved by dividing. There is no new mathematics in §3. There is a change of coordinates.
Trap 1: initial conditions. Transforming in over the whole line quietly assumes the solution exists for all time and decays at both ends. That is the right thing to want for a steady-state response, and the wrong thing for an initial-value problem, where the system is switched on at and you are told and . For those, the correct tool is the Laplace transform. It is the same idea with and the integral run over , and that change of limits produces the initial data as boundary terms in the integration by parts instead of discarding them. Everything in this chapter has a Laplace twin.
Trap 2: can vanish. Dividing by is illegal where has a zero on the real axis, which is exactly the undamped resonance , . The division then fails because the answer genuinely does not exist: an undamped oscillator driven exactly at resonance has no bounded steady state, it grows linearly forever. The formula's failure is reporting a physical fact. When damping is present, has no real zeros and the division is safe. That is why was doing more work in Chapter 0.8 than it looked.
Any perpendicular basis would have delivered everything so far, so the case for this one has yet to be made, and it rests on a property no other basis possesses. Differentiating a pure wave returns the same wave multiplied by a number. The chosen basis is the one in which differentiation is diagonal, which is why it was chosen and the reason the subject exists.
In that description differentiation stops being differentiation. It becomes multiplication by a number differing from one wave to the next, so a differential equation, which ties the value of a function to the values at its neighbours, becomes an algebraic equation with one instance per wave, answered by dividing. The difficulty has moved into transforming in and back out, and none of it is left in the middle.
By now the manoeuvre should be recognisable, because this is the fourth performance under the fourth different name. A symmetric array became a list of numbers along perpendicular directions. Two coupled masses became two oscillators that ignore each other. An exponential turned out to be what differentiation merely rescales. And a function on the line has become a list of amplitudes, one per wavelength, none interfering with any other. Find the description in which the problem falls apart into independent pieces: one idea, and it has now paid for itself four times over.
4 · Convolution and Green's functions
Definition. The convolution of two functions is
Read it as: at each point , take a weighted average of , with the weights given by centred on and reflected. The operation is symmetric. Substituting turns (0.9.26) into , so . It is bilinear as well, which makes it a genuine product on the space of functions.
4.1 · The convolution theorem, derived
Transform (0.9.26) and split the exponential using :
The second line substituted at fixed . That is a shift, so and the limits are unchanged. The double integral now factorises, because the integrand is a product of a function of alone and a function of alone:
The stray is the price of the symmetric convention, and the table in §2.2 lists what it becomes in the others. Get it wrong and every amplitude in Part V is off by a constant factor, so it is worth one moment of care now.
Now the headline. Convolution in is multiplication in . An operation that costs a full integral at every single point becomes a pointwise product.
4.2 · Linear time-invariant systems, and Green's functions
Here is why (0.9.28) is one of the most load-bearing facts in physics. Consider any system that is linear, meaning that doubling the input doubles the output and that inputs superpose, and time-invariant, meaning that its behaviour does not depend on what o'clock it is. Feed it an input and call the output .
Let be the output produced by a single sharp unit impulse at . That output carries two names, the impulse response and the Green's function. Now put the two hypotheses to work. Time-invariance says an impulse at produces . Linearity says a general input, regarded as a dense sequence of impulses of strength , produces the superposition of the individual responses:
That is the whole content of linear response theory, and it is a convolution. Transform it with (0.9.28) and it becomes . Now set that beside (0.9.24) and read off that the transfer function is . The transfer function is the transformed Green's function. Two objects, one thing.
Chapter 0.8 already built one of these and said so. Its integrating-factor solution of (Chapter 0.8 calls the rate constant , but here is taken, so it is ) was
with for and for . The upper limit in Chapter 0.8's formula and the vanishing of for negative argument are the same statement, and that statement is causality: the response cannot precede the impulse. Duhamel's formula was a convolution all along.
Everything above generalises with replaced by a spacetime point. In Chapter 5.4 the operator is a wave operator, , and its Green's function is called the propagator , which is the amplitude for a disturbance created at to arrive at . Transforming, the propagator becomes
which is (0.9.24) with : division by the polynomial, exactly as in §3.
Now the payoff. Take a particle that propagates from to , interacts, and then propagates to . Its amplitude is a convolution of propagators in position space, because the interaction point is unknown and has to be integrated over. By (0.9.28), in momentum space that integral collapses into a product.
This is why the Feynman rules read the way they do: write down a factor for each line and each vertex and multiply them together. There is no deeper combinatorial magic in the rules. Feynman diagrams are products in momentum space because the underlying operation is convolution in position space, and the convolution theorem is (0.9.28), which you have just proved in four lines. That is a real and load-bearing insight, and it is worth knowing now rather than in Part V.
Hit a system once, sharply, and record everything it does afterwards. Provided it responds in proportion to what is done to it and behaves the same way whatever the hour, that record is the whole system, because any input can be regarded as a dense succession of sharp hits and the output is whatever survives of each, added together. The operation assembling the answer, in which one function is slid across another and the overlap recorded at every displacement, is expensive: a full integral at every point.
Written in terms of waves it is a product. Transform both functions, multiply them together, transform the result back, and an integral at every point has become one arithmetical operation at each wavelength. That is why the transform earns its place in daily practice, quite separately from why it exists in principle.
The consequence is larger than it looks. In the last part of the book the response to a disturbance created at one point of spacetime and detected at another is this same object with more indices, and the amplitude for a particle to travel, interact somewhere unspecified, and travel onward is a sliding overlap of such responses. Written in terms of wavelengths those overlaps turn into multiplications, which is exactly why the rules for Feynman diagrams tell you to write a factor for each line and multiply.
5 · The Dirac delta, handled honestly
Two debts have been quietly accumulating. §4 needed "a single sharp unit impulse" and did not define one. §2 needed a completeness relation for a continuum of basis functions and did not have one. The same object settles both debts, and that object is not a function.
That last sentence is the whole difficulty of this section, so this is the plan. First we define the thing by what it does rather than by what it is. Then we show it is the narrow-bump limit Chapter 0.2 was already drawing. Then we derive its rules, and at each step you will see why the "not a function" caveat has to be respected rather than merely noted.
5.1 · Define it by what it does
Do not ask what value takes at each point. There is no answer, and §7's warning box shows in detail why the question is malformed. Ask instead what the object does to other functions. The Dirac delta is defined by
for every sufficiently well-behaved test function . That is the definition, complete.
Now notice what kind of object that makes . It is a rule that eats a function and returns a number, and it does so linearly, since . A linear map from functions to numbers is a linear functional, which is precisely Chapter 0.5's definition of a bra. Objects defined this way are called distributions, and is the simplest non-trivial one.
Shifting the argument gives the general sifting property, by substituting :
The delta reaches into a function and pulls out its value at one point, discarding everything else. Every use of in this book is that sentence.
5.2 · It is the limit Chapter 0.2 was drawing
Take the normalised Gaussian, whose total area is by Chapter 0.2's :
This is exactly the curve in Chapter 0.2's width figure. It is the one whose measured area stayed pinned at while the peak rose as , and whose caption said that pushing to its minimum was watching the Dirac delta form. It was. Here is that claim made precise. For continuous ,
Proof. What we have to drive to zero is the difference between the two sides. Since , we may write that difference as
Fix . By continuity there is an with for . That part of the integral is therefore smaller than . Outside , the factor is at most , which goes to zero faster than any power as , so (for bounded ) that part vanishes in the limit. Hence the whole thing is below for small enough , and was arbitrary.
Two things are worth extracting from that proof, and the first is the one that will keep you out of trouble.
The limit is taken after integrating, never before. Take it before and you get pointwise, which is for and at . That object carries no information at all: it is zero almost everywhere, so any integral of it against anything is zero, and the one point where it is not zero contributes nothing. The order of the two operations is what the delta is. This is why every rule in §5.4 is proved by testing both sides against an , rather than by comparing values point by point, and it is why the delta lives inside an integral sign and nowhere else.
The Gaussian was not special. Any normalised bump that narrows will do. Chapter 0.2's grind box on differentiating under the integral sign exhibited another one, on the half-line, and called it "the Dirac delta being born". It was.
5.3 · The Fourier representation, and what it really is
The single most-used identity in Parts IV–VII is
which cannot be true as it stands, since the integrand has modulus everywhere and the integral does not converge. Here is the derivation that makes it respectable, and it uses nothing beyond Chapter 0.2. There are three steps. Insert a convergence factor with , do the now-honest Gaussian integral with , and take at the end:
The regulated integral is a normalised Gaussian of width . That is precisely (0.9.33), which tends to by §5.2. So (0.9.36) is true in the only sense in which it was ever meant: as a limit of honest integrals, evaluated under an integral sign.
The manoeuvre was to insert a regulator, compute, and then remove the regulator. It is not a dodge to be embarrassed about. It is the same move that will be called renormalisation in Chapter 5.10, met here in its harmless form.
Now let's see what (0.9.36) actually is. Chapter 0.5's completeness relation reads . In function language, with the components, that says , since acting on it has to return . So put in the Fourier basis (0.9.5), , and manufacture the mode spacing one last time:
So (0.9.36) is the completeness relation in continuous disguise. That is the honest reading, and it explains why the delta appears everywhere the Fourier transform does: every time you insert the identity in a continuous basis, a delta drops out. In Dirac notation the whole of §2 compresses to and , and (0.9.36) is .
5.4 · The rules, derived
Scaling. Claim: for . Test it against an arbitrary and substitute . If then and the limits keep their order:
If the substitution reverses the limits, and swapping them back supplies a minus sign, so the answer is . Both cases give , which is what gives. Since the two functionals agree on every test function, they are equal:
Note the consequence: is not dimensionless. If carries units of length then carries units of , so that (0.9.31) balances. Checking that in an unfamiliar formula catches an impressive number of errors.
The derivative. We would like to obey integration by parts. So define it to, discarding the boundary term (test functions vanish at infinity):
This is how every distribution is differentiated, and it is why distributions are so useful. They can always be differentiated, as many times as you like, with no smoothness required of anything. The step function has a derivative, and that derivative is . Then has a derivative of its own, and so on forever. Chapter 5.2's field equations lean on this constantly.
The three-dimensional delta, and Chapter 0.7's loose end. Define , so that . Chapter 0.7 left a warning here. The Coulomb potential satisfies everywhere the computation is legal, and yet a point charge is manifestly a source. Both halves are true, and the resolution is a distribution.
Take the two halves in turn. Direct differentiation away from the origin gives for . Now instead integrate over a ball of any radius and use the divergence theorem of Chapter 0.7, with :
An object that vanishes everywhere except one point yet integrates to over every ball containing that point is, by (0.9.31), exactly times a delta:
Chapter 0.7's loose end is tied. Now read (0.9.43) once more in the language of §4. It says that is the Green's function of the Laplacian, the response to a unit point source. Poisson's equation is therefore solved by convolution with it, and that convolution is
That is Coulomb's law, recovered as a special case of §4.2 rather than posited.
Grind box — distributions, said properly
The construction that makes all of §5 rigorous is due to Schwartz and is worth seeing once, in outline, because it explains exactly which manipulations are safe.
Test functions. Fix a space of very well-behaved functions. They are infinitely differentiable, and either compactly supported or, in the Schwartz space , decaying faster than any power along with all their derivatives. is the natural home for Fourier analysis, because (0.9.20) and its partner trade decay for smoothness in both directions, so the transform maps onto itself.
Distributions. A distribution is a continuous linear map . Ordinary functions embed: becomes . Some distributions, like , are not of that form for any , and that is the precise sense of " is not a function".
Everything is defined by moving the operation onto the test function. Derivative: . Scaling and translation: as in §5.4. Fourier transform: . Applied to , that last one gives , so , a constant. A perfectly sharp spike in has perfectly flat content in . Hold that thought for §6.
What is not defined. There is no general product of two distributions. The recipe above works because differentiating or transforming a test function gives another test function. Multiplying two distributions has no such move available, and is genuinely undefined. Not hard, not subtle, undefined. See the warning box.
The single sharp kick relied on a moment ago was never defined, and defining it honestly means giving up the idea that it is a function. Nothing assigns a value to each point of the line and behaves as required, because a function altered at one point leaves every integral where it was, and infinity is not a value anything takes. The way out is to stop asking what it is and state what it does, which is to eat a function and hand back its value at one point.
That is not a partial answer but the whole definition, and objects specified this way are perfectly legitimate. Two habits go with them. Such an object appears only under an integral sign, paired with something well behaved, and the limit producing it, in which a bump narrows while its area is held at one, must be taken after the integration and not before. Take it before and you have a quantity that is zero everywhere and infinite at a point, carrying no information.
Worth extracting too is the device that makes its most quoted representation respectable. Insert a factor rendering a divergent integral convergent, compute the honest answer, and remove the factor at the end. That is neither a dodge nor a trick peculiar to this object. It is the procedure later called renormalisation, met here in its harmless form.
6 · The bandwidth theorem
This is what the chapter was for. Here is the plan, in plain English, before we start. We give a function a width and give its transform a width, using the ordinary statistical definition of spread. Then we show the two widths cannot both be small, and we get the exact number they are bounded by.
6.1 · Defining a width
Let be normalised, , which by Plancherel (0.9.17) means as well. Then and are both non-negative and both integrate to one: they are probability densities, whether or not anybody intends them as such. So give each the standard deviation of §7:
with and .
We may assume . That is not a convenience granted to ourselves. It is a fact, and it takes two lines.
- Replacing by shifts bodily and changes only by the phase . So and are both untouched, while moves by .
- Replacing by leaves alone and shifts to . That moves by , again with untouched.
So both means can be driven to zero without changing either spread, and we do so.
6.2 · The move that makes it work
The difficulty is that is an integral over while is an integral over . They live in different places, so as things stand there is nothing to compare. Our goal is to get both of them onto the same axis, and §2 and §3 together do exactly that. Use the derivative theorem (0.9.20) first, then Plancherel (0.9.17):
Let's pause on that line, because it is the crux. The spread in has become a statement about alone: it is the size of the derivative of . In hindsight the shape of it makes sense. A function with a lot of high- content is a function that wiggles, and a function that wiggles has a big derivative. What we have now is an identity rather than an intuition. Both integrals run over , so we may compare them.
6.3 · Cauchy–Schwarz
Apply Chapter 0.5's inequality in the function space with the inner product (0.9.1), to the two vectors
Their norms are exactly the two quantities we want, by (0.9.45) and (0.9.46):
So the theorem is proved as soon as we can put a lower bound on the left-hand side. Call it (the is real, so the conjugation does nothing to it) and add to its own conjugate. The two terms assemble into a derivative by the product rule:
Now integrate by parts (Chapter 0.2). The boundary term drops because at both ends. That is a hypothesis rather than a triviality, and the grind box justifies it:
That is the whole calculation. The normalisation , which looked like tidiness, has turned into the number on the right-hand side.
Write for "multiply by " and for . Then for any ,
And , while by the same integration by parts. So (0.9.50) reads . The integration by parts and the commutator are the same computation, seen from two sides. That is exactly why Chapter 4.9 will be able to state the general uncertainty relation as , as Chapter 0.5's insight box already promised. The here is the of quantum mechanics with the not yet installed.
6.4 · The theorem
A complex number is at least as big as its real part in absolute value, so from (0.9.50),
Feed that into (0.9.48) and take the square root of both sides:
Done. Let's list what went into it, because the shortness of the list is the point.
- The definition of the transform (§2).
- Plancherel (§2).
- The derivative theorem (§3).
- Cauchy–Schwarz (Chapter 0.5).
- Integration by parts (Chapter 0.2).
- The normalisation.
There is nothing else in it.
6.5 · The Gaussian saturates it
We want to know which functions actually achieve the bound, so run the argument backwards from its two inequalities. Chapter 0.5 recorded the equality condition for Cauchy–Schwarz: equality holds exactly when the two vectors are parallel. Here that reads for some constant .
The other inequality was (0.9.51), and equality there needs , which forces to be real. Since , that in turn forces to be real. Now solve by separation (Chapter 0.8). It gives , so
The minimisers are exactly the Gaussians, and nothing else. This is stronger than the claim that Gaussians happen to be one example. The derivation ran backwards from the equality condition and produced them uniquely.
Now verify the value directly, cashing Chapter 0.2's Problem 4. Write , so and , which by Chapter 0.2's Worked example 2 has . Its transform, by with , is
We now have the same function described both ways, so the two widths can finally be multiplied together. Squaring the transform gives , and matching that against the standard Gaussian shape gives . Therefore
The bound is attained, so (0.9.52) cannot be improved. Squeeze the Gaussian and its transform broadens by exactly the reciprocal factor. The product does not move. That is a thing you should now watch happen.
Grind box — the hypotheses, said out loud
The boundary term. We assumed at . Here is why it is automatic whenever the theorem's quantities are finite. By Cauchy–Schwarz, , so is integrable. Its derivative is , whose second piece is bounded in integral by , so the derivative is integrable too. A function whose derivative is integrable has limits at . A function that is itself integrable and has limits must have those limits equal to zero. So the boundary term vanishes for exactly the functions for which the statement has content.
When the spreads are infinite. If or the inequality is vacuously true, and this is not a rare corner: any function with a jump discontinuity has (Problem 1), and any function with power-law tails has . The theorem is sharp only among functions smooth enough and localised enough to have both moments. Outside that class one uses different width measures: the width at half maximum, the width containing of the power, the reciprocal of the peak height. All of them obey inequalities of the same shape with different constants, and none of those constants is .
Why the standard deviation, then? Because it is the measure for which the constant is exactly and the minimiser is exactly the Gaussian, and because it is the one Chapter 4.9 needs. There, is a variance of measurement outcomes, and it has to be that and nothing else.
(0.9.52) is the Heisenberg uncertainty principle. It is a theorem about Fourier transforms, and it has just been proved with no physics anywhere in it. No particles, no measurement, no observer, no . Every ingredient came from Part 0.
All that quantum mechanics will add, in Chapter 4.9, is a single substitution: . That is one physical claim, the de Broglie relation, and it is the only new thing. Multiply both sides of (0.9.52) by :
The inequality was never quantum. It is a fact about waves, as true of a radio pulse or a seismic trace as of an electron, and it was known to signal engineers before it was known to physicists. The quantum part, the genuinely strange and genuinely new part, is the claim that a particle's momentum is the wavenumber of a wave. Once you accept that, the uncertainty principle is not an additional mystery. It is arithmetic you did in Part 0.
It does not say that measuring a particle's position disturbs its momentum. That story, the microscope and the photon kicking the electron, is Heisenberg's own 1927 heuristic. It is not what the theorem states, and it has misled essentially every popular account since.
Look at what was actually proved. (0.9.52) is a relation between two representations of one object. The function and its transform are the same thing written in two bases, and the theorem says their widths cannot both be small. No measurement occurs anywhere in the derivation. No observer appears. Nothing is disturbed, because nothing is done.
A short pulse is broad in frequency, and that is the end of it. A femtosecond flash of light contains a hundred nanometres of spectrum whether or not anyone looks at it (Problem 4). A bass note plucked for a tenth of a second has an intrinsically fuzzy pitch. The spread is a property of the object, not of the act of examining it.
Measurement disturbance is a real and separate phenomenon with its own theorems, and Chapter 4.9 will keep the two apart. If you carry one thing away from §6, carry this: the uncertainty principle is a statement about what a localised wave is, not about what happens when you poke it.
Whatever else a wave may be, it cannot be both brief and pure, and the exact statement of that impossibility is the strongest thing in the toolkit. Measure the spread of a function across positions in the ordinary statistical way, measure the spread of its transform across wavelengths likewise, and the product can never fall below one half. Squeezing either widens the other by the compensating factor, and the bell curve is the unique shape attaining the minimum, a conclusion the argument produces rather than assumes.
Look at what the proof needed, because the list is short and every item was in hand: the definition of the transform, the fact that transforming preserves size, the fact that differentiating has become multiplying, the overlap inequality proved when inner products were introduced and applied here to two particular functions, and one integration by parts. There is nothing else in it, and no physics anywhere: no particles, no measurement, no observer and no Planck's constant.
That inequality is the uncertainty principle, and all quantum mechanics adds to it later is one substitution: that a particle's momentum is the wavelength of a wave. It says nothing about disturbing a thing by looking at it, because no measurement occurs in the derivation and nothing is done to anything. A brief flash of light contains a wide spread of colours whether or not anybody watches.
7 · Probability, and why it is the same subject
Probability enters this chapter for a structural reason and not a curricular one: the central limit theorem is a Fourier argument, and the object that makes it one is a Fourier transform under another name.
7.1 · The apparatus, briefly
A random variable taking real values is described by a probability density with , meaning . Its cumulative distribution is .
By the fundamental theorem of calculus (Chapter 0.2), those two are related by . The density is the derivative of the CDF, and that one sentence is what makes both objects worth having. Next, the expectation of any function is
which is linear in because integration is. The mean is , the variance is , and expanding the square with linearity gives the form you actually compute with:
The standard deviation is . That is the same quantity §6 called a spread, computed from the same integral. And the Gaussian density is
whose normalising constant is Chapter 0.2's with , and whose variance is by Chapter 0.2's Worked example 2. Nothing in (0.9.58) is quoted.
7.2 · The characteristic function is a Fourier transform
Define the characteristic function of :
It is the Fourier transform of the density, up to the convention's constant and a sign. Three properties follow at once, and all three get used below.
- It always exists, because and is integrable. Moments are not like that. A moment can be infinite, and the grind box at the end of §7 turns on a case where one is.
- It determines uniquely, because the transform is invertible by (0.9.14).
- Its Taylor coefficients are the moments, because differentiating under the integral (Chapter 0.2) brings down factors of .
Writing that last property out to second order:
7.3 · Independent sums convolve, so characteristic functions multiply
Let and be independent, with densities , and set . Independence means the joint density factorises, . Then
and differentiating with respect to under the integral, using on the inner integral:
Adding independent random variables convolves their densities. That is a surprising fact with a one-line derivation, and it is the bridge we came here for, because §4's convolution theorem now applies to probability.
So transform (0.9.62), then translate the result back into characteristic functions with (0.9.59). The two stray factors of cancel against the one in (0.9.28), and what is left is
A hard operation on densities, an integral at every point, has become multiplication of two ordinary functions. This is (0.9.28) and nothing more, and it is the entire reason the central limit theorem is provable.
7.4 · Variance adds, and the
Take without loss of generality (subtract the means). Then
and the cross term factorises under independence, . Hence
Variances add. Standard deviations do not. Everything below is bookkeeping on that one line.
Take independent copies, each of variance . Their sum has variance . And since follows directly from the definition, the mean has
The standard error falls as , and the exponent is not a convention or an empirical rule. It is the square root of "variances add", which is (0.9.65), which is the vanishing of a cross term.
7.5 · The central limit theorem, sketched honestly
What follows gets the mechanism exactly right and is genuinely where the Gaussian comes from. It leaves two things unproved, and both are named below rather than hidden.
Let be independent with the same distribution, mean , finite variance . Standardise:
Write , so and . By (0.9.63), and since straight from the definition,
Take logarithms, which turns the power into a product and is the move that makes the whole thing work. Expand to second order using (0.9.60) with , then expand from Chapter 0.3:
Watch what happened, because it is the entire theorem. The quadratic term carried exactly one power of , so multiplying by left it finite and unchanged. Every higher term carried a higher power of , so multiplying by still left it going to zero. The quadratic term is the unique survivor. It survives not because we truncated anything, but because the scaling by in (0.9.67) was chosen precisely to let it survive. Hence
All that remains is to recognise the answer. By Chapter 0.2's Gaussian integral with and , the function is the characteristic function of the standard normal density , as you can check by doing that integral. So the limit is a Gaussian. And it is a Gaussian because is the fixed point of "square and rescale", which is what (0.9.68) does.
What is not proved here. Two things, and here they are.
- The remainder needs control uniform in . Doing that properly requires either a third moment or a truncation argument, and the honest general statement is Lindeberg's.
- Convergence of characteristic functions implies convergence of distributions. That is Lévy's continuity theorem, and we are quoting it.
Neither gap affects the mechanism. The Gaussian appears because the quadratic term is the only one that survives the scaling, and that part is fully derived above.
Grind box — when the central limit theorem is false
The hypothesis is doing real work. Drop it and the conclusion fails, in an instructive way.
Take the Cauchy density , whose tails fall as . That makes diverge, so there is no variance to speak of. Its characteristic function is (⚑ a contour integral, quoted here, and the only place in this book that uses complex analysis, which Chapter 5.4 builds). Then by (0.9.63) the mean of samples has
the characteristic function of a single Cauchy sample. Averaging a million Cauchy observations gives you exactly the precision of one observation. The is gone, because (0.9.65) never applied: there was no variance to add. Note also that is not differentiable at . The failure of the mean to exist is visible as a corner in the characteristic function, exactly where (0.9.60) would have read it off.
This is not a pathology invented to embarrass the theorem. Cauchy tails arise from ratios of normal variables, from resonance line shapes (the Lorentzian of (0.9.25) is a Cauchy density), and in any setting where the largest single contribution is comparable to the sum of the rest. When you see a sample mean that refuses to stabilise as data accumulate, this is the first thing to suspect.
Equation (0.9.66) is the most expensive line of algebra in clinical research, and it is worth seeing that it is a theorem rather than a statistical convention. The standard error of a mean is because variances add for independent observations (0.9.65), which is true because a cross term vanishes. That is the whole reason. No other exponent is available.
The consequences are brutal and entirely determined. A confidence interval has half-width proportional to , so halving the width of an interval costs four times the sample size, and cutting it to a third costs nine times. For a two-arm trial comparing means with two-sided and power, the required size per arm is
with and . The detectable effect enters squared, which is (0.9.66) inverted: chasing an effect half the size costs four times the patients, and an effect a quarter the size costs sixteen times. Every argument you have ever had about whether a trial is powered for a realistic effect is an argument about that exponent, and the exponent is not negotiable.
Now the bridge, and it is an identity rather than an analogy. Counting is a sum of independent yes/no events, so (0.9.65) applies to it verbatim. One event either registers or does not, with probability . By (0.9.57) its variance is , since for a variable taking only the values and . Now add independent such events. The mean count is , and the variance is for small .
So counts carry a standard deviation of and a relative precision , derived in three lines from the same cross term vanishing. That single fact sets the noise floor of every measurement in physics. It sets photon shot noise in a detector, the error bar on a bin of a collider histogram, the counting error on a radioactive sample, and the sensitivity of a gravitational-wave interferometer.
When a particle physicist demands before claiming a discovery, they are asking for an excess whose probability under the null is . The they are counting in is , from (0.9.65). When you compute the power of a trial, you use the same line with replaced by patients.
So a discovery threshold at CERN and a sample-size calculation on a protocol sheet rest on identical algebra, derived above in four lines. The difference between the two fields is which quantity is expensive to increase.
The delta is not a function, and pretending otherwise produces real damage. There is no assignment of a value to each point of the real line that does what (0.9.31) requires. Take the natural candidate, zero everywhere except at one point. It fails, because changing a function at a single point does not change any integral at all. That candidate has for every , which is not . Nor can the value at the origin be "infinity", since infinity is not a real number, and no arithmetic involving it would make the integral come out to rather than to something else.
So the delta means something only under an integral sign, paired with a test function, and its type is a linear functional rather than a function. That reads like a piece of pedantry, and it is the opposite. It is the rule that tells you which manipulations are safe, and the rest of this box is what happens to people who ignore it.
Treat it as a function anyway and you immediately generate two expressions that look harmless and are not.
- . Setting in (0.9.36) gives , a divergent integral over all wavenumbers.
- . There is no product of distributions, and would have to be "", which is the previous problem again.
Both show up in quantum field theory, unavoidably and on the first page of any real calculation. A enforcing momentum conservation gets squared when you square an amplitude, and the resulting is the volume of spacetime. It is infinite, and correctly so, because you asked for a total probability rather than for a rate per unit volume. Products of propagators at coincident points produce the same disease.
In every case the appearance of or is not a mistake. It is a signal that the question was posed in a limit that has not yet been taken carefully, and the repair is a regulator. Put the system in a box, smear the point, analytically continue the dimension. It is exactly what §5.3 did when it inserted , computed, and then removed it. Chapter 5.10 does nothing else for an entire chapter and calls it regularisation, and Chapter 5.11 is what you do with the answer afterwards.
Probability arrives at the close not as a change of subject but as the same subject in other words. Adding two independent random quantities slides one density across the other and records the overlap, which is exactly the operation of a moment ago, so in the wave description the densities multiply. Repeat the multiplication many times, rescaling as you go, and every feature of the original distribution is ground away except the quadratic one, which survives because the rescaling was chosen to let it. The bell curve is what remains once everything else has been averaged out of existence, and that is why it is everywhere.
The toolkit is finished. What you have now is a way of describing change and a way of accumulating it, the habit of expanding rather than solving, a language of spaces and the maps between them, the theorem that the right description makes a hard problem fall apart into independent pieces, a calculus for quantities defined at every point, and the one basis in which differentiation is arithmetic.
None of it is physics. All of it was assembled for a single job, and the next part begins that job by discarding the thing most people take to be the foundation of the whole subject. It throws away forces, and puts in their place one number attached to each history the world might have followed.
8 · Worked examples
Compute for , , and verify exactly.
Straight from the definition (0.9.12):
Complete the square, exactly as in Chapter 0.2:
The second term is constant in and comes out. The first is a Gaussian shifted by an imaginary amount, and shifting the contour off the real axis is the one step here that needs a word. It is legal because is analytic and decays in the strip between the two contours, so a rectangle contour has nothing left on its connecting edges. That argument is quoted here and built properly in Chapter 5.4. ⚑ Chapter 0.2's Problem 4 checked that same continuation independently, by solving a differential equation in instead. Two routes, one answer, which is how you know the continuation is sound. With Chapter 0.2's :
A Gaussian transforms to a Gaussian, of reciprocal width. A large (narrow in ) produces a small in the exponent (broad in ), and vice versa.
The widths. The densities are and . Matching each against :
Two things to notice. First, the answer is independent of . Squeeze the function and the transform spreads by precisely the reciprocal factor, which is the behaviour the figure in §6.5 measures numerically. Second, this is the equality case of (0.9.52), and §6.5 proved it is the only one. Chapter 0.2's Problem 4 got here first, by a completely different route, and said so.
Solve with all the material initially at the origin, .
Chapter 0.7 got the scaling out of this equation by dimensional analysis: has units of , so the only length available is . That argument is correct and gives no constant and no shape. Here is the actual solution.
Transform in only, holding as a parameter. By (0.9.20) applied twice, , while the time derivative passes through the -integral untouched:
The partial differential equation has become, for each separately, the first-order ordinary equation of Chapter 0.1, the one that says rate is proportional to amount. Its solution is the exponential, with a decay constant that depends on the mode:
The initial condition. By (0.9.31), the transform of is : a constant, the same at every wavenumber. A point source contains every mode equally. So
and the physics is already visible without inverting anything: each mode decays at rate , so short-wavelength structure is erased fastest, quadratically fastest. That is what smoothing is.
Transform back with (0.9.14), using Chapter 0.2's Gaussian integral with and :
Read the answer. It is a Gaussian for every , with total integral at all times (Chapter 0.2 again), which is to say that matter is conserved. Now match against to read off the width:
This closes the loop Chapter 0.7 opened. There, the scaling came from dimensional analysis, and Problem 3 of that chapter said explicitly that the method could give the scaling but never the pure number in front. Here is the number: the distribution is Gaussian, its variance is , and the constant is . Nothing was guessed.
It also explains why the answer had to be Gaussian. The solution is the density of a sum of many independent random steps, and §7.5 says such a sum is Gaussian. The heat kernel and the central limit theorem are the same statement, which is why Chapter 5.6's path integral will be able to treat diffusion and quantum amplitudes with one piece of machinery.
One last reading is available. is the Green's function of the diffusion operator, so by §4.2 the solution for any initial profile is the convolution . Diffusion is convolution with a spreading Gaussian, which is precisely why repeated diffusion steps compose by adding variances.
9 · Your turn
Problem 1 — the square pulse, the sinc, and a divergence worth understanding
Let for and otherwise. (a) Compute . (b) Verify Plancherel (0.9.17) for it, given . (c) Compute . (d) Attempt , and explain what you find in terms of (0.9.52).
Solution
(a) Directly from (0.9.12):
using the same step as (0.9.3). This is the sinc function: a central lobe of half-width , then side lobes of alternating sign whose envelope decays only as .
(b) . On the other side, substituting ,
(c) , so
(d) Now attempt the same in . The numerator is
because the cancels the exactly. The integrand no longer decays at all. It oscillates between and forever, and the integral diverges. So and . The theorem holds, in the loudest possible way.
This is not a technicality. It is the theorem telling you something true. A hard edge is infinitely sharp, and by §6.2 the spread in is the size of . For a step, is a delta, and a delta has infinite norm. A discontinuity costs infinite bandwidth.
Truncating at shows how it diverges. Then (the averages to ), so and
This is what the §6.5 figure is doing on the square-pulse preset. The readout is tracking the sampling cutoff rather than converging to anything. Double the number of grid points and rises by a factor of . The measured values are for , with ratios . The discrete pulse has its own sampling corrections on top, so the number does not match the continuum formula exactly. The point is that there is no limit for it to match. Worth knowing where a number comes from before you trust it.
A finite honest statement. of the power lies in the central lobe (numerically, from ). If you use the central-lobe half-width as the bandwidth, the product is , comfortably above . Different width measure, different constant, same conclusion. And it is the last of the power, out in the tails, that makes the variance-based measure blow up.
Problem 2 — the scaling law for the delta, in three dimensions too
From the defining property (0.9.31) alone, prove for real . Then deduce , find the analogous rule for , and use the result to evaluate for .
Solution
The scaling law. Two distributions are equal when they give the same number on every test function, so test both sides. Substitute , .
For the limits keep their order:
For the substitution sends , so the limits arrive reversed and swapping them costs a sign:
Both cases give , which is what gives. Equal on every test function, hence equal.
Consequence. Put and you get , so the delta is even. That had better be true, since (0.9.36) is manifestly even in .
Three dimensions. , and each factor contributes :
The exponent is the dimension, because the delta's units are one over a volume. In Chapter 5.8 the deltas are four-dimensional and this is how their Jacobians are tracked.
The composite argument. vanishes at , and near each zero it is approximately linear (Chapter 0.1): with . The delta only sees neighbourhoods of the zeros, so it splits into one term per zero, each rescaled by the local slope via the law just proved:
The general rule, over the simple zeros , is this argument run once. It is used on every page of relativistic kinematics, where puts a particle on its mass shell and the becomes the factor in the phase-space measure.
Problem 3 — Gaussians are closed under addition
Let and be independent. Using the convolution theorem rather than a direct integral, show that is Gaussian with variance . Then say what would have gone wrong without independence.
Solution
Step 1, the characteristic function of a Gaussian. With , Chapter 0.2's with , gives
Step 2, multiply. The density of is by (0.9.62), so by (0.9.63) the characteristic functions multiply:
Step 3, recognise it. That is Step 1 again with . Since the characteristic function determines the density uniquely (the transform is invertible, §7.2), .
Compare the effort. Convolving two Gaussians directly means completing a square inside a double exponential. In Fourier space it is adding two exponents. This is (0.9.28) earning its keep, and it is the reason the Gaussian family is closed under addition at all. The transform turns convolution into multiplication, and the exponential of a quadratic is exactly the shape that is closed under multiplication, because quadratics add.
Without independence. Step 2 is the only place it was used, and it fails completely: is not unless the joint density factorises, so the characteristic functions do not multiply. Correlated Gaussians still sum to a Gaussian (that is a separate fact about the joint normal distribution), but the variance is , from (0.9.64) with the cross term retained. Positively correlated errors add up faster than . That is why correlated measurements do not average down, and why cluster randomisation costs more patients than it looks like it should.
Problem 4 — short pulses and narrow lines
(a) From (0.9.55), show that a Gaussian pulse whose intensity has full-width-at-half-maximum has a frequency content of FWHM . Evaluate the required spectral width, in nanometres, for a pulse centred at . (b) An excited state decays with lifetime , so its amplitude behaves as for and before. Transform it, show the intensity spectrum is a Lorentzian, and find the linewidth. Evaluate for .
Solution
(a) For a Gaussian, (0.9.55) gives in standard deviations, i.e. (the pulse is a function of and its transform a function of , so nothing changes but the names). A Gaussian intensity has FWHM , from solving . So
and converting the angular frequency to ordinary frequency, :
This is the transform limit. A pulse cannot be shorter than this for its bandwidth, and a pulse that achieves the bound is called transform-limited.
Numbers. gives . Converting to wavelength via , so (Chapter 0.1's logarithmic derivative):
A pulse at must span roughly of spectrum. That is more than a tenth of its own centre wavelength, and a large fraction of the visible range. This is why ultrafast lasers need broadband gain media, and why "monochromatic ultrashort pulse" is a contradiction rather than an engineering challenge.
(b) The carrier as written, , puts the spectral line at once our forward kernel has acted. Nothing in the physics depends on that sign, so take the carrier to be instead and the line lands at . Transform with (0.9.12), integrating only over where . The exponent collects to , and the integral is elementary:
The evaluation at the upper limit vanishes because of the factor. The decay is what makes the transform exist at all.
The measured intensity is , and the modulus squared of a reciprocal is the reciprocal of the modulus squared:
That is a Lorentzian, the identical shape to the driven-oscillator response (0.9.25). The match is not a coincidence: an excited state is a damped oscillator, with . Its half-width at half maximum is , so the full width is
Numbers. For : , so , and . Against an optical transition at () that is a fractional width of . Spectacularly sharp, and yet not zero, and not zero for a reason that has nothing to do with the apparatus.
The point. The relation is usually presented as a fourth Heisenberg relation, "energy–time uncertainty", with an air of mystery. It is (0.9.52) applied in the time domain, plus . A state that does not last forever does not have a sharp energy, for the same reason a note that does not last forever does not have a sharp pitch. No measurement anywhere.
You built the Fourier basis as an orthonormal basis, with the orthogonality (0.9.4) proved by direct integration and the coefficients recognised as Chapter 0.5's inner products.
You took honestly and got the transform pair. You proved Plancherel and read it as the statement that the transform is unitary, a rotation in function space.
You showed that the basis diagonalises the derivative, which is the reason the subject exists, and turned two differential equations into division. You derived the convolution theorem and saw that a linear system's response is a Green's function convolution, hence a product of transfer functions, and hence a product of propagators.
You gave the Dirac delta a definition it can survive, derived its Fourier representation with a regulator, and recognised it as Chapter 0.5's completeness relation in continuous disguise.
You proved the bandwidth theorem from Cauchy–Schwarz, showed the Gaussian is its unique minimiser, and measured that number live. And you found the characteristic function to be a Fourier transform, which made the central limit theorem a statement about which Taylor term survives a rescaling.
Where this gets spent.
- The Fourier basis and Plancherel → Chapter 4.3 (where completeness is finally proved and the position and momentum representations become two bases for one Hilbert space), Chapter 4.6 (where unitarity is conservation of probability), Chapter 5.3 (where a quantum field is expanded in exactly (0.9.14), one harmonic oscillator per ).
- Diagonalising the derivative → Chapter 5.4, where every propagator is (0.9.24) with replaced by a four-momentum, and where momentum space is the only place the calculations are tractable.
- Convolution and Green's functions → Chapter 5.4 again: the Green's function of a wave operator is the propagator, and Feynman diagrams are products precisely because convolution in position space is multiplication in momentum space (0.9.28).
- The delta → Chapter 4.5 (continuum normalisation ), Chapter 5.3 (equal-time commutators), Chapter 5.10 (where and appear and force regularisation).
- The bandwidth theorem → Chapter 4.9, which adds and nothing else.
- Probability, variance, the CLT → Chapter 4.19 (density matrices and measurement statistics), Chapter 5.11 (where the renormalisation group is a central-limit theorem for fluctuations at successive scales, by the same "only one term survives the rescaling" argument as (0.9.69)).
And that closes Part 0. Nine chapters, one accumulating toolkit, no physics yet. Here is the whole of it in one list.
- The linear approximation (0.1).
- The integral as an accumulation, with the Gaussian done three ways (0.2).
- The series expansion with control of its error (0.3).
- The vector space and the linear map (0.4).
- The spectral theorem in an inner-product space (0.5).
- The total derivative and the Jacobian (0.6).
- The field theorems that turn interior cancellation into boundary integrals (0.7).
- The oscillator and the differential equations around it (0.8).
- And now the Fourier transform that makes all of them commute with each other.
Chapter 1.1 begins the physics. What happens there is not that you start using one of these tools. It is that the first serious question in mechanics, why do objects move along the paths they do?, needs the linear approximation, the integral, the expansion, and the variational idea all at once, in the first three pages. That is the point at which Part 0 stops looking like a long detour and starts looking like the only sensible way in.