Part IV · Quantum Mechanics — Chapter 4.3
Function Spaces: Measure, L², and Completeness
Three promises come due together: the integral is thrown away and rebuilt, the space of states becomes a space, and completeness stops being an assumption.
Chapter 4.2 ended by saying what it had bought on credit. Every time that chapter wrote "expand in a basis", it was assuming that the basis reaches everything, and every time it wrote "the state after the measurement", it was assuming that the object it produced was still a state. In finite dimensions both are theorems of Chapter 0.5. For a particle on a line the space is a space of functions, and neither one has been established at all.
So this chapter builds the space. Chapter 0.4 built the vector space and Chapter 0.5 built the operators on it. This chapter builds the space and Chapters 4.4 and 4.5 build the operators on it, in that order and for that reason, which is the shape Chapter 4.2's closing brick announced.
Three separate promises fall due here, and all three were made long before Part IV. Chapter 0.2 said the integral would be thrown away and rebuilt from scratch. Chapter 0.2 also used dominated convergence to legalise a step and said it would be proved properly here. Chapter 0.9 quoted the completeness of the Fourier basis, marked the quotation, and named this chapter as the place the mark comes off. The work is one piece, because you cannot prove that a basis reaches everything until you know what "everything" is, and you cannot say what "everything" is until the integral defining the norm behaves under limits.
Here is the route. Section 1 reopens the hole Chapter 0.2 dug in front of you, and it will be the same hole and the same function. Section 2 replaces the idea of length by the idea of measure, and quotes one construction. Section 3 builds the new integral. Section 4 proves the two theorems that make it worth having, and those two theorems are the point of the whole rebuild. Section 5 assembles the space of square-integrable functions and checks Chapter 0.5's axioms against it one by one. Section 6 is the centre of the chapter: the space is complete, and the exact sequence that had no limit in §1 acquires one. Sections 7 and 8 turn that into a working basis and show that the Fourier modes really are one.
Conventions. This chapter is real analysis and uses no complex analysis anywhere, in keeping with the rest of Part IV. The Fourier convention is Chapter 0.9's, unchanged. The inner product is linear in its second slot, as Chapter 0.5 §1.1 chose it. Two results are used without being proved here, and both are marked where they enter: the construction of Lebesgue measure in §2.3, and the proof of the completeness theorem in §6.2. A third is quoted and its mark is not this chapter's to raise: Heine–Cantor, standing since Chapter 0.2 §1.1, which §8.2 leans on and cites rather than raising again. There are no others. A chapter this heavy carrying two new marks is the claim it is making, and §9 and the closing brick say where each one is spent.
Tools you'll need — Chapter 0.2 §1.1 above all, for the Riemann integral as a limit of tagged sums, for the function , and for the sequence of functions its warning box constructed and then set aside. Also its §4.4 for differentiation under the integral sign, and the grind box there that named the hypothesis this chapter proves. Chapter 0.5 §1 for the three inner-product axioms and Cauchy–Schwarz, §2 for orthonormal bases, Parseval and the resolution of the identity, and §3.1 for the fact that a projection is the closest point in a subspace. Its §1.2 table has the row this chapter exists to make legitimate. Chapter 0.9 §1 for the Fourier modes and their orthogonality, §1.3 for the completeness assumption and its mark, §1.4 for the square wave and the Gibbs overshoot, §2.3 for Plancherel, and §5.2 for the argument about a narrowing bump that §8 reuses almost word for word. Chapter 0.3 §3 for convergence and the tests that decide it without naming a limit. Chapter 0.4 §2 for basis and dimension, which is the argument §7 has to replace. Chapter 0.8 §6.4 for the Lorentzian, which Worked example 3 turns into a state. Chapter 4.2 §8 for the three lines that force infinite dimension, and its closing brick for the division of labour between this chapter and the next. Chapter 4.1 §5 for the integral whose term-by-term evaluation §4.1 here makes airtight.
1 · Why the Riemann integral has to go
Here is the debt this section collects, and it is Chapter 0.2's. We look again at the function that chapter could not integrate, define the one piece of vocabulary the book has been using without ever writing down, and then exhibit the failure that actually matters. The failure is not that some function is not integrable. It is that the class of integrable functions is not closed under the limits we need to take, and by the end of the section you will have a sequence with nowhere to converge to.
1.1 · The function Chapter 0.2 left on the table
Chapter 0.2 defined the integral as a limit of tagged sums and then, in a warning box attached to that definition, told you the definition would not survive. The words were these: "In Chapter 4.3 we will throw away [the definition] and rebuild the integral from scratch (the Lebesgue integral), and the reason is visible already." We are going to do exactly that, so let's start by looking again at what was visible already.
The function is Chapter 0.2's , and we keep its name. It takes the value at every rational number and at every irrational one. Now recall why the definition of the integral cannot cope with it. A tagged sum lets you choose the sample point in each cell, and every cell, however short, contains rationals and irrationals both. So you may tag entirely at rationals, or entirely at irrationals, and on the two choices give
and both are available at every mesh, however fine. The definition asks for a single number that every tagging approaches, and there are two numbers that every mesh can produce. So is not Riemann integrable, and no refinement helps, because refining does not remove the choice.
Now let's be honest about how much that proves, because on its own it proves very little. A single function that cannot be integrated is an annoyance rather than a crisis. Physics is not short of integrals it declines to attempt, and a pathology you can name and step around costs nothing. If this were the whole case against the Riemann integral, the sensible response would be to note as a curiosity and carry on.
The real case is about limits, and to state it we need one word the book has been leaning on without ever defining.
1.2 · Cauchy sequences, defined here
Chapter 0.3 §3 gave you tests that decide whether a series converges. The ratio test is the one you use most, and notice what it does: it settles convergence without producing the limit. That is the whole trick, and it is worth naming, because everything in this chapter turns on the difference between a sequence that is heading somewhere and a sequence that has somewhere to go.
Work in a space with a norm, which for now means Chapter 0.5's . A sequence converges to when , and that definition names the limit. The alternative names nothing outside the sequence itself. Call a Cauchy sequence when its terms eventually huddle together:
The two conditions are not symmetric, and the asymmetry is the subject of this chapter. One direction is immediate. If then, given , take beyond which every term is within of , and the triangle inequality of Chapter 0.5 §1.4 closes it:
Every convergent sequence is therefore Cauchy. The converse is the interesting question, and the first thing to see about it is that it is not a question about the sequence at all. It is a question about the space the sequence lives in.
Take the rational numbers with the ordinary absolute value, and take the decimal truncations of . Any two terms beyond the -th differ by less than , so the sequence is Cauchy. It does not converge in , because the number it is closing in on is not a rational number. Nothing is wrong with the sequence. Something is missing from the space, and the missing thing has a name.
A space in which every Cauchy sequence converges to a point of the space is called complete. The real numbers are what you get by insisting on exactly that: you take the Cauchy sequences of rationals and declare the ones that ought to have the same limit to be the same number. Completeness enters this book at that point, as part of what a real number is rather than as a theorem about them, and Chapter 0.3's convergence tests all quietly stood on it. A test that certifies convergence without exhibiting a limit is only worth having in a space where the limit is guaranteed to be there.
That is the property we are going to need for functions, and §6 is where it is established. Keep (4.3.2) in view, because the rest of this section is one sequence of functions that satisfies it and has nowhere to land.
1.3 · Chapter 0.2's sequence, and the two failures it already exhibits
The same warning box in Chapter 0.2 that promised the rebuild also built the sequence that forces it, and we use that sequence rather than a new one. Enumerate the rationals as , which can be done, and let be the function on equal to at and everywhere else. Each differs from the zero function at finitely many points, so each is Riemann integrable, and the value is
Chapter 0.2 then observed that pointwise and stopped there. Let's carry it further, because the sequence shows two distinct defects and they need separating.
The first defect is that the class is not closed under limits. The sequence is increasing, every term is Riemann integrable, and every integral is , which is as tame as a sequence can be. Its pointwise limit is , which is not Riemann integrable at all. So the natural statement that any theory of integration wants to make,
fails here for the worst possible reason. It is not that the two sides differ. It is that the right-hand side does not refer to anything. Any convergence theorem is a statement of the form (4.3.5), so within the Riemann theory there can be no convergence theorem worth the name.
The second defect is that the norm is not a norm. Chapter 0.5 §1.2's table proposed as an inner product on functions, and said in the next paragraph that the row "is the entire reason Chapter 4.3 is possible." Test the axioms on the sequence above. For the difference is at finitely many points and elsewhere, so
and yet . That is a direct violation of Chapter 0.5's axiom (iii), the demand that for every non-zero . Without axiom (iii) there is no distance, because two different objects sit at distance zero, and Chapter 0.5's whole apparatus rests on all three axioms holding. Section 5 repairs this, and the repair is not cosmetic: it changes what a vector is.
Notice also that (4.3.6) makes Cauchy in the strongest imaginable sense, every pair being at distance exactly zero. So the sequence is Cauchy and its pointwise limit is not integrable. That sounds like the third failure, and it is not, which is worth saying out loud rather than gliding over. Distance zero from also means distance zero from the zero function, so in this norm the sequence does have a limit sitting inside the class. Chapter 0.2's sequence is too thin to show that anything is missing. We fatten it.
1.4 · The same rationals, thickened, and the hole that is left
Keep Chapter 0.2's enumeration and give each rational a little room. Around put the open interval of length , and let be the part of lying in . That is a finite union of intervals, so its indicator function
is a step function, perfectly Riemann integrable, whose integral is the total length of the intervals making up . Those totals increase with and never exceed the sum of all the lengths, which is . An increasing sequence of real numbers bounded above converges, so tends to a limit .
Now measure the distance between two terms. For the difference is the indicator of with removed, so its square is itself, and
because converges and a convergent sequence of numbers is Cauchy. So is a Cauchy sequence of Riemann-integrable functions. The question is what it converges to, and the answer is the point of the section.
There is no Riemann-integrable on with . Not one that is hard to find. Not one that exists but is badly behaved. None.
The reason is a tug-of-war between two facts about the set , meaning the union of all the , which is the part of lying in . It contains every rational of the interval, so it is dense: every subinterval of meets it. And its total length is at most , so it is small. A limit of the would have to be close to everywhere in a dense set, which lets a tagged sum at any mesh be pushed up to , and it would have to have integral at most . Those two demands cannot both be met by a function the Riemann integral can handle, and the grind box turns that sentence into a proof.
Grind box — no Riemann-integrable function is the limit, in full
Suppose is Riemann integrable on and , the integrals existing. We derive a contradiction in four steps.
1. The integral of is at most . Apply Cauchy–Schwarz to and the constant function on . Chapter 0.5 §1.4 proved that inequality from all three axioms and §1.3 has just withdrawn the third for this pairing, so the citation needs a word. Its proof survives on this pair. The opening inequality asks only that never be negative, which is still true here, and every later appeal to (iii) is a division by , which for is . So
So by (4.3.7).
2. A non-negative Riemann-integrable function with zero integral has infimum zero on every subinterval. If had on a subinterval of length , then any partition with the endpoints of among its points has lower sum at least , so .
3. On every subinterval, comes arbitrarily close to . Let be any subinterval. Since contains every rational of it meets , and is open, so contains a closed subinterval . Now are open sets whose union is , so by compactness for some , and then on for every . Hence
so . By step 2 the infimum of over is zero, so takes values as close to as you please inside , and therefore .
4. Contradiction. Take any partition of . Step 3 was proved for subintervals of the open interval , and the first and last cells are not subsets of that, so the transfer needs one word. Every cell contains a subinterval of , and a supremum over a cell is at least the supremum over a piece of it, so step 3 gives on every cell. Now fix and tag each cell at a point where , which the supremum makes possible. The tagged sum is then , and such a tagging exists at every mesh. Riemann integrability is the statement that every tagging converges to , so for every , hence . Step 1 gave .
Nothing there needed the fact that the integral is the infimum of the upper sums. Chapter 0.2 defined the integral by tagged sums and never proved that characterisation, and choosing the tags by hand is both shorter and closer to the definition the book actually has.
The two steps that did the work are worth naming. Density of forced the tagged sums up. Smallness of held the integral down. A theory of integration that could see the difference between "meets every interval" and "has length " would have no trouble here, and the Riemann integral cannot see it, because a cell width and a sample point are the only instruments it has.
So we have a Cauchy sequence of perfectly ordinary step functions with no limit in the class. This is the third failure and it is the fatal one. Chapter 4.2 §7 built time evolution as a limit, and Chapter 0.9 §1.3 asserted that partial Fourier sums converge to the function they came from. Both are statements that a Cauchy sequence has a limit, and in this class they are false.
What we need is now precise rather than vague, and it is a list of four items.
- A notion of the size of a set that can tell a dense set of total length from the whole interval. Section 2.
- An integral built on that notion, defined on a class closed under limits, and agreeing with Chapter 0.2's wherever Chapter 0.2's exists. Section 3.
- Convergence theorems that let a limit pass through the integral sign, so that (4.3.5) becomes a theorem with hypotheses instead of a wish. Section 4.
- A repair of axiom (iii), so that distance zero means equality and Chapter 0.5's geometry transfers intact. Section 5.
Section 6 then shows that the four together buy completeness, and hands back the sequence with a limit attached.
The integral you were taught works by chopping the horizontal axis into thin strips and adding up the areas of rectangles. It is a good definition and it handles every function a physics problem is likely to hand you. It has one flaw, and the flaw is not about any particular function. It is about what happens when you take a sequence of functions and ask what they are approaching.
A sequence of numbers can be seen to be settling down without anyone knowing what it is settling down to. Every term past a certain point is within a hairsbreadth of every other term: that is a self-contained test, and it needs no knowledge of the answer. Whether a sequence that passes the test actually arrives somewhere is not a fact about the sequence. It is a fact about the world the sequence lives in. Inside the fractions alone, the decimal expansion of the square root of two passes the test and arrives nowhere, because the place it was heading is not a fraction. Adding the missing destinations is what the real numbers are for.
Exactly the same thing goes wrong one level up, with functions instead of numbers. There is a sequence of the most ordinary functions imaginable, each one flat except on a few short intervals, which passes the settling-down test and has nothing to settle on. The destination it wants is a function so ragged that the strip-and-rectangle definition cannot assign it an area at all. That is not an oddity to be stepped around. Quantum mechanics is going to say that a state is a vector in a space of functions, and that the limit of a sequence of states is a state. In this space that sentence is false.
So the definition of area has to be replaced, and the replacement starts one step further back than you might expect. Before asking what the area under a curve is, ask what the size of a set of points on the line is. Get that right and everything else follows.
2 · Measure, with one thing quoted
This section replaces one word. Chapter 0.2's integral was built on the length of an interval, and length is the thing that could not see the difference between a dense set and a fat one. We write down what a general notion of size would have to satisfy, and we find that three demands fix almost everything. Then one construction is quoted rather than built, because building it is a chapter of its own and this book is not going to pretend otherwise. What the quotation buys is stated precisely, and §8 comes back to collect part of it.
2.1 · The three demands, and why one of them is countable
We want to attach a number to a set of real numbers, and the number should behave the way length behaves. Three demands do it.
- It extends length. For an interval, is the length: , and the same for the closed and half-open versions.
- It is countably additive. If are disjoint then the size of the union is the sum of the sizes.
- It is translation invariant. Sliding a set along the line does not change its size.
The first and third are obvious requests. The second is the one carrying the weight, and the word countably is doing all of the carrying, so let's see why finite additivity would not be enough. Every failure in §1 was a countable process. The set of rationals is a countable union of single points, the set was a countable union of intervals, and the sequences and were indexed by the integers. A notion of size that only adds finitely many pieces cannot say anything about any of them.
Countable additivity has one consequence we will use so often that it is worth extracting now. Suppose is an increasing family with union . Write the union as a disjoint one by peeling off shells, , and apply countable additivity to the shells. The partial sums of the resulting series are exactly the , so
This is called continuity from below, and it is the exact point at which the word "countably" gets spent. Section 4's monotone convergence theorem is (4.3.9) promoted from sets to functions, and it has no other engine.
2.2 · Which sets, and the shape the collection has to have
The demands above say nothing about which sets get a size, and §2.5 will show that not all of them can. So the collection of sets we can measure is part of the structure, and its shape is forced by the operations we intend to perform on it.
A collection of subsets of is a -algebra when it contains , is closed under taking complements, and is closed under countable unions. Countable intersections come free, because is the complement of , and set differences come free for the same reason. Those are precisely the operations a countable limit process performs, which is why the definition looks the way it does rather than some other way.
The smallest -algebra containing every open interval is called the collection of Borel sets, and it contains everything you are likely to write down: open sets, closed sets, countable unions of closed sets, countable intersections of those, and so on upward. The set of §1.4 is a countable union of open intervals, so it is Borel, and so is the set of rationals.
2.3 · The construction, quoted as a package
Everything so far is a specification. Meeting it is a genuine construction, and here is the first of the two places where this chapter takes something on trust. The construction starts by defining an outer measure, which covers a set as economically as possible with intervals and takes the cheapest cover:
That definition applies to every set whatever, which is why it cannot be the end of the story: something so generous is not going to be countably additive. Carathéodory's criterion then selects the sets on which it behaves, by keeping exactly when splits every other set additively.
We use, without proof, the following. There is a -algebra of subsets of , containing every Borel set, and a function agreeing with (4.3.10), with all of these properties.
- (a) Length. is the length of for every interval.
- (b) Countable additivity. for disjoint .
- (c) Translation invariance. .
- (d) Regularity. For and any there is an open with , and a closed with .
- (e) Completeness of the measure. If and then and .
This is Lebesgue's theorem. No chapter of this book proves it, and none will, which is why the mark is here rather than a forward pointer. The construction runs from (4.3.10) to (a) through (e) in about twenty pages and opens every course in measure theory, Royden's Real Analysis and Rudin's Real and Complex Analysis being the two standard places to read it. That is the first time this book sends you to another text for an argument. Chapter 3.6 §5.4 named a text before, for a signature convention, which is a different kind of pointer.
Say what is being bought. Properties (a), (b) and (c) are the three demands of §2.1 and nothing more. Property (e) is a convenience that costs nothing and saves a clause in every later proof. Property (d) is the one that does real work later, and it is worth watching. It says a measurable set is squeezed between an open set and a closed set that are as close to it in size as you like, which is what lets an arbitrary measurable set be traded for a finite union of intervals. Section 7.4 cashes it, making that trade once in order to show the space separable, and §8.1 then re-uses the same trade to prove that continuous functions are dense in the space of states. Neither result is separately marked, because both are this package spending itself.
2.4 · Sets of measure zero, and the phrase "almost everywhere"
Call a null set when . The most useful fact about null sets is how easy they are to come by, and the proof is a construction you have already seen in this chapter.
Every countable set is null. Let the set be , fix , and cover by an interval of length . That is a countable cover of total length , so by (4.3.10), and was arbitrary. In particular
Look at what that construction was. It is §1.4's thickening of the rationals, with the thickness sent to zero instead of held at . The same cover that produced a dense set of length produces, in the limit, the statement that the rationals have no length at all. Chapter 0.2 could not write down (4.3.11), and it is the whole reason is about to become integrable.
A property holds almost everywhere, abbreviated a.e., when the set where it fails is null. So almost everywhere, and the functions of §1.3 are all equal to each other almost everywhere. That phrase is going to reorganise the whole subject in §5, where two functions agreeing almost everywhere stop being two functions.
Restrict to and read as the chance that a uniformly distributed random number lands in . Nothing has to be adjusted for the reading to work, because Lebesgue measure on is the uniform distribution, and the correspondence is term by term.
- is the statement that some outcome occurs.
- Countable additivity is the axiom that the chance of any one of countably many mutually exclusive events is the sum of their chances. It is not an approximation to that axiom. It is that axiom.
- A null set is an event of probability zero, and "almost everywhere" is what a statistician writes as "almost surely".
Equation (4.3.11) then reads: a uniform random draw from is rational with probability exactly zero, although rationals are everywhere in the interval. You have used that fact every time you treated a continuous variable as continuous. Probability zero does not mean impossible, and it never did. What §2.5 adds is the part that has no clinical counterpart: there are subsets of to which no probability can be assigned at all, on pain of contradiction, and the axioms of probability are stated over a -algebra rather than over all subsets for exactly that reason.
2.5 · Not every set is measurable
The flag in §2.3 restricted to a collection rather than to all subsets of , and a reader is entitled to ask whether that restriction is a real one or a technician's caution. It is real, and the demonstration is short enough to run here. It is due to Vitali.
The point of running it is to make the flag honest. If could be all subsets, the quoted package would be quoting a formality. It cannot, and the three demands of §2.1 are exactly what makes it impossible.
Grind box — the Vitali set, built
Work on and declare when is rational. This is an equivalence relation, so it cuts into disjoint classes, each class being one number plus every rational, cut down to . One class is the rationals of on their own. Every other class is an irrational offset plus all the rationals. Choose one representative from each class and collect the choices into a set . That step is the axiom of choice, and it is where the whole construction lives.
Now translate around the circle. For let
which is shifted by with the overhang wrapped back to the start.
They are disjoint. If and agree modulo then is rational, so and lie in the same class, so by the choice, so .
They cover. Given , let be the representative of 's class. Then is rational and lies in , so it is either or for some , and .
The contradiction. Suppose with . Each is two translated pieces of , so by (c) and (b) of §2.3 it is measurable with . The rationals in are countable, so countable additivity applies to the whole family:
If the sum is . If the sum is . Neither is , so is not in .
Every ingredient was one of the three demands, plus the axiom of choice. Drop countable additivity for finite additivity and the argument collapses, which is one reason the countable version is not a stylistic preference.
So is a genuine restriction, and the sets it leaves out are the ones you cannot write down. Every set arising from a formula, a limit, or a physical construction is Borel, hence in , and §3 never has to think about it again.
Before you can improve on the idea of area you have to improve on the idea of length. The question is how big a set of points on a line is, when the set is not an interval and may be scattered all through one.
Three requirements settle it almost entirely. The size of an interval is its length. Sliding a set along the line leaves its size alone. And if a set is cut into a list of non-overlapping pieces, even an endless list, the sizes of the pieces add up to the size of the whole. That third requirement is the only one with any bite, and the word "endless" is where the bite is. Every difficulty in the previous section was an endless process, so a rule that only adds finitely many pieces is no use.
Meeting the three requirements takes a construction, and this chapter quotes it rather than performing it, with the quotation marked in place. What the quotation includes is worth knowing, because one clause of it is spent later on something that looks unrelated: any set that can be measured can be trapped between an open set and a closed one that are as close to it in size as desired. That clause is what will eventually let a wildly discontinuous function be traded for a smooth one.
Two consequences follow at once. A set of separate points, even endlessly many of them, has size zero, because you can cover the first with a cover of length one-half of your allowance, the second with one quarter, and so on forever, and the total is your allowance however small you made it. So the rationals, dense as they are, have no size at all. And the price of the whole scheme, which is the honest part: there are sets to which no size can be given, and the proof that they exist uses nothing but the three requirements. The restriction is real rather than a technician's fussiness.
a natural place to stop · the problem is set, and so is what "size" must mean
3 · The Lebesgue integral, built
With size in hand, the integral takes one paragraph to define and one section to check. The whole idea is a single change of direction, which §3.1 states in a sentence. Section 3.2 says which functions are eligible and proves the property that the Riemann class lacked. Section 3.3 does the construction. Section 3.5 shows the new integral agrees with the old one wherever the old one works, so that nothing in Parts 0 to III has to be relearned, and §3.6 is honest about the one thing that is lost.
3.1 · One change of direction
Here is the whole idea in a sentence. Riemann slices the domain. Lebesgue slices the range.
Chapter 0.2 chopped the -axis into cells, took the value of somewhere in each cell, and added up value times width. The new definition chops the -axis instead. For each level it asks how big the set of points where is near is, and adds up level times size.
The picture that makes it stick is counting money. Given a pile of coins, Riemann's method is to go along the pile in order and add each coin's value as you reach it. Lebesgue's method is to sort the pile into denominations first, count how many coins are in each pile, and add up denomination times count. Three terms map exactly: the value of a coin is the value of , the position of a coin in the pile is the point , and the number of coins of a given denomination is the size of the set where takes that value.
One term has no counterpart, and it is the one §1 turned on. A coin has a single definite value, so there is no cell, no width, and no freedom about where to sample. Both methods therefore return the same total on every pile, which is §3.5 in miniature, and neither of them can fail. So the picture cannot show you why one method survives where the other does not, and that claim is left where it is proved, in §4.
3.2 · Which functions, and the property the Riemann class lacked
Slicing the range means asking for the size of the set where exceeds a level, so that set had better be measurable. Call measurable when
That is the entire condition, and it is weak. Every continuous function passes, since is then open. Every step function passes. So does , for which is when , the rationals when , and empty when . Three sets, all of them in , and nothing you can write down fails.
Now the property that Riemann's class did not have, and the reason the -algebra was defined with countable operations in §2.2. Let be measurable and suppose at every . Then exactly when the values eventually stay above some level strictly greater than , and "some level", "eventually" and "stay above" are three countable operations, written below in that order:
Every set on the right is in by (4.3.12), and is closed under countable unions and intersections, so the set on the left is in too. A pointwise limit of measurable functions is measurable.
Compare that with §1.3, where a monotone limit of Riemann-integrable functions was not Riemann integrable. The defect is gone, and it is gone by construction rather than by luck: the class was defined by a condition phrased in countable operations precisely so that countable operations could not escape it.
3.3 · Simple functions, then the supremum
The construction goes in two stages, and the first stage is the one where the definition actually happens. A simple function is a finite sum of indicators of measurable sets,
and its integral is defined to be the obvious thing, level times size, added up. Since the are disjoint there is no ambiguity about what to write:
That is the range-slicing of §3.1 with only finitely many levels. Two things follow from finite additivity of and nothing else, and both are used constantly below: the integral of a sum of simple functions is the sum of the integrals, and implies .
The second stage extends this to any non-negative measurable by approaching it from below with simple functions and taking the best available answer:
a supremum that may be . Monotonicity is immediate from (4.3.16), since means every competing for also competes for , so the supremum can only grow. That one-line fact is used in almost every proof that follows.
It is worth seeing that the supremum is always attained in the limit by an explicit staircase, because §4 will apply its theorems to exactly this sequence. Cut the range into steps of height up to a ceiling of and read off which step is on:
Each is simple by (4.3.12), the sequence increases because halving the step can only refine the estimate, and at every . So every non-negative measurable function is an increasing limit of simple ones, which is the range being sliced ever more finely.
Finally, drop the requirement that be non-negative. Write with and , both non-negative and measurable, and define . Call integrable when both pieces have finite integral, which is the same as saying . For a complex-valued , integrate the real and imaginary parts separately. That is the whole definition.
3.4 · The first two answers, one of them a debt
Chapter 0.2's warning box promised that the new integral "gives into the bargain", and the promise is now one line. The function restricted to is the indicator of , which is simple, so (4.3.15) applies directly and (4.3.11) supplies the size:
The function that no partition could pin down has an integral, and getting it took one multiplication. The same applies to §1.4's thickened set: is a limit of the simple functions , and its integral will turn out to be once §4 licenses passing to the limit. That is the loose end §6 ties.
3.5 · It agrees with Riemann wherever Riemann works
If the new integral disagreed with the old one anywhere the old one was defined, every calculation in Parts 0 to III would need checking, so the following matters more than its proof suggests.
If is Riemann integrable on , then is measurable, it is Lebesgue integrable, and the two integrals are equal.
The idea is that upper and lower sums are integrals of simple functions, so the Riemann machinery is already inside the Lebesgue machinery, viewed correctly. Take partitions with mesh going to zero, each refining the last, and build the two staircases
the sums running over the cells of . These are simple functions whose Lebesgue integrals, by (4.3.15), are exactly the lower and upper Riemann sums of over . Riemann integrability says those two sequences of numbers close on each other, and that squeeze is what the grind box turns into the theorem.
Grind box — Riemann implies Lebesgue, with the same value
Shift by a constant so that , which changes both integrals by the same amount. Write and for the lower and upper Riemann sums over , and for the Riemann integral.
One line first, because Chapter 0.2 defined that integral by tagged sums and not by upper and lower ones. On a fixed partition the tagged sums come as close as you like to and to , by tagging near the infimum or near the supremum in each cell, and Chapter 0.2 §1.1 observed that every tagged sum is trapped between the two. Integrability sends every tagged sum to as the mesh shrinks, so and are squeezed onto with them. Refinement makes the first increase and the second decrease, so and .
The same refinement condition makes increasing and decreasing pointwise, so both have pointwise limits
and both are measurable by §3.2. Now three steps.
1. . Monotonicity of (4.3.16) against for every and gives . Let and run and the two ends squeeze onto .
2. almost everywhere. The function is non-negative and satisfies , so monotonicity and additivity on simple functions give , hence . For any non-negative measurable the inequality and monotonicity give Chebyshev's bound , so here every vanishes, and the union over is the set where . It is null.
3. Conclusion. Since and off a null set, almost everywhere, so is measurable by property (e) of §2.3, and its integral equals , two functions agreeing off a null set having the same integral by (4.3.15) applied to the competing simple functions.
So no arithmetic in this book changes. Every Gaussian in Chapter 0.2, every Fourier coefficient in Chapter 0.9, and every action integral in Part I means the same number it always did. What has been added is a larger class of functions and, in §4, the theorems that make the enlargement worth having.
3.6 · The one thing that is lost, said plainly
The new integral is not merely more powerful than the old one, and pretending otherwise would store up a surprise. There is a class of integrals the Riemann theory handles and the Lebesgue theory declines.
Chapter 0.9's grind box on the Gibbs constant used the sine integral, and the same function gives the standard example. The improper Riemann integral
exists, because the alternating arches cancel in a controlled way as grows. But , since the -th arch contributes about and the harmonic series diverges. By §3.3 a function is Lebesgue integrable only when , so is not Lebesgue integrable on .
The reason for the asymmetry is exactly the reason the new integral is worth having. The Lebesgue integral sorts the range before it sums, so it can never rely on the order in which cancellations arrive, and an improper Riemann integral is a statement about that order. When you meet in Part V it is a limit of proper integrals, and it has to be written as one.
The new definition of area is one idea. Instead of chopping the horizontal axis into thin strips and asking how tall the function is over each strip, chop the vertical axis into thin bands and ask how wide the set of places is where the function lands in each band. Then add up height times width as before. Counting a pile of coins by walking along it is the first method; sorting the pile into denominations and multiplying is the second. For a finite pile the two totals agree, and they agree here too, wherever the old definition worked at all.
What changes is which functions are eligible. The old definition needed the function to be reasonably well behaved along the horizontal axis. The new one needs only that each of the bands corresponds to a set whose size is defined, and the collection of such sets was built in the previous section to be closed under endless unions and intersections. That is not a coincidence. It is what makes the class of eligible functions survive limits, which is the exact property the old class lacked and the whole reason for the rebuild.
The function that took the value one at every fraction and zero elsewhere now has an area, and the calculation is a single multiplication: the height is one, the set of fractions has size zero, the product is zero. A page of frustration in the earlier chapter becomes a line.
One thing is given up, and it should be said rather than discovered later. Certain integrals converge only because positive and negative contributions arrive alternately and cancel. Sorting the values before adding them destroys that arrangement, so such integrals are not admitted by the new definition and have to be written as limits of ordinary ones. That is the price of never depending on the order of the terms, and the next section is what the price buys.
4 · Monotone and dominated convergence
This is the section the rebuild was for. Section 1 showed that the Riemann integral admits no theorem of the form "the integral of the limit is the limit of the integrals". Here are the two that the Lebesgue integral does admit. The first is proved from continuity from below and nothing else. The second follows from the first through one intermediate step, and it is the one Chapter 0.2 borrowed against. Both are stated with their hypotheses in front, because in both cases the hypothesis is what the theorem is really about.
4.1 · Monotone convergence, derived
Let be measurable and let at every . Then . The limit function is measurable by §3.2, so both sides refer to something, and that is exactly what failed in §1.3.
One direction is free. Monotonicity of (4.3.16) gives for every , and the left side increases, so its limit satisfies . The work is the other direction, and it consists of showing that beats every simple function competing in the supremum that defines .
So fix a simple with , and fix a slack factor slightly below , kept a different letter from the levels because the two are about to be multiplied together. Since climbs to , every point is eventually caught above , so the sets
Write for , the integral restricted to a set. Discarding everything outside can only lower an integral of a non-negative function, so , and the last of those is a simple function's integral, which (4.3.15) writes out explicitly:
The limit in the middle is (4.3.9), continuity from below, applied to the increasing sets inside each . That is where the countable in "countably additive" is spent, and it is the only place the proof spends anything. So for every , hence , and taking the supremum over gives . With the free direction, .
Two corollaries follow at once, and the first of them repairs an omission. Nothing so far has established that for general non-negative measurable and , since (4.3.15) only gave it for simple functions. Take the staircases and of (4.3.17), note that , and apply monotone convergence to all three sequences. Additivity for simple functions passes to the limit, so the Lebesgue integral is linear.
The second corollary is the one that gets used in anger. Let be non-negative and measurable and apply monotone convergence to the partial sums , which increase because the terms are non-negative:
No convergence hypothesis is needed anywhere in (4.3.22), since both sides are allowed to be and the theorem says they are infinite together. Sign is the only thing being asked for.
One interchange that looks like a neighbour of this one is not settled here, and it is worth saying so rather than leaving a reader to wonder. Exchanging a sum with an integral is (4.3.22). Exchanging two integrals is Fubini's theorem, which Chapter 0.2 quoted and marked when it squared the Gaussian, and which needs a measure on a product space rather than on the line. That is a second construction of the kind §2.3 already quoted once, this chapter does not perform it, and nothing below uses it. Chapter 0.2's mark stands where it is.
A debt, collected. Chapter 4.1 §5 evaluated the integral in the Planck spectrum by expanding the denominator as a geometric series and integrating term by term, and it wrote in parentheses that "the theorem that makes it airtight is monotone convergence, proved in Chapter 4.3." Here it is. For the expansion has every term positive, so (4.3.22) applies with no further hypothesis, and
The Stefan–Boltzmann constant of Chapter 4.1 rests on that interchange, and the interchange now rests on a theorem rather than on an expectation. Worked example 2 derives the that turns the middle expression into the last.
4.2 · Fatou's lemma, and what it says when it is strict
Monotone convergence needs the sequence to climb, and most sequences do not. The next result asks for nothing but non-negativity, and pays for that generosity with an inequality instead of an equality.
For non-negative measurable , . The proof is monotone convergence applied to the wrong sequence on purpose. Set . That sequence has to be measurable before monotone convergence can touch it, and §3.2 closed the class under pointwise limits rather than under infima, so here is the missing line. Writing puts each in , then is in as a countable intersection, and is the condition (4.3.12) asks for. The increase with and climb to by definition, so monotone convergence gives . Also for every , so monotonicity gives . Let run and the right side becomes .
An inequality is only interesting when you know how it fails to be an equality, and Chapter 0.2 already built the example. Its grind box on differentiating under the integral sign used
while as for every fixed . Take and you have a sequence with and . Chapter 0.2 said of it that "the mass has not vanished. It has escaped into an ever narrower, ever taller spike near the origin", and Chapter 0.9 §5.2 identified the same object as the Dirac delta being born. Fatou's lemma is the exact statement that mass can escape but cannot appear from nowhere, and the direction of the inequality is the direction mass can go.
4.3 · Dominated convergence, and Chapter 0.2's outstanding hypothesis
Now the theorem the book has been waiting on. It buys back the equality, and the price is a single function that holds the whole sequence down.
Let be measurable with pointwise almost everywhere, and suppose there is one integrable with for every . Then is integrable and , so in particular .
The proof is Fatou applied to a cleverly chosen non-negative sequence. Since as well, we have , so is non-negative and Fatou applies to it. Its pointwise lower limit is , because , so
using the linearity established in §4.1. Now is finite, which is the whole content of the hypothesis, so it may be cancelled from both sides. That leaves , and the quantity is non-negative, so it is zero.
The promise, collected verbatim. Chapter 0.2 §4.4's grind box justified swapping a derivative with an integral and wrote: "The sufficient hypothesis (dominated convergence, proved properly in Chapter 4.3) is that the -derivative of the integrand be bounded, uniformly in near the point of interest, by a single fixed integrable function." That is the theorem above, and the swap it licenses is worth stating as its own result rather than left implied.
Suppose is integrable in for each near , that exists, and that for all such , with integrable. Then
To see it, take any sequence and form the difference quotients . The mean value theorem writes each as for some between and , so for every , which is the domination. And pointwise, which is the convergence. Dominated convergence therefore gives , and the left side is by construction the difference quotient of . Every sequence gives the same limit, so the derivative exists and (4.3.26) holds.
Chapter 0.2's own case now closes. There and you were asked to accept the bound
with the right-hand side a fixed function of finite integral. That is exactly the hypothesis of (4.3.26), so every Gaussian moment Chapter 0.2 extracted by differentiating under the integral sign is now derived rather than provisional. The generating function of Chapter 0.2 §4.5, and with it the whole differentiation-with-respect-to-a-source apparatus that Part V runs on, inherits the same licence.
It is tempting to read "there is an integrable with " as a regularity condition that any reasonable sequence satisfies. Chapter 0.2's escaping spike is the counterexample and it is worth computing, because it shows the hypothesis failing by exactly the amount the conclusion fails.
Any dominating function must beat of (4.3.24) at every , so it must beat the supremum over . Substituting turns into , whose maximum over is at , so
The smallest possible dominating function is not integrable, so no dominating function exists, and dominated convergence does not apply. It is not that the theorem is silent about this sequence while the conclusion happens to hold. The conclusion is false here, and the hypothesis is the exact thing that fails.
The practical version, for a reader who will be swapping limits and integrals for the rest of the book: the danger is always mass escaping to where the integral cannot follow it, either up a narrowing spike or out to infinity. A dominating function is a fence that stops both.
Everything so far has been preparation for two statements, and these two are what the rebuilt integral is actually for. Both answer the same question: if a sequence of functions is approaching a limit, does the area under them approach the area under the limit?
The first says yes whenever the functions only ever increase, filling in from below and never retreating. Nothing else is required, and the proof leans on exactly one thing: that the size of a growing family of sets is the limit of their sizes. That was the awkward clause about endless lists in the definition of size, and this is what it was for.
The second is the one used daily. It says yes whenever the whole sequence can be held under a single fixed function of finite area, a ceiling that does not move as the sequence evolves. Without such a ceiling the statement is simply false, and the way it fails is always the same: the area runs away somewhere the limit cannot see it, up a spike that grows taller as it grows narrower, or off toward infinity in a bump that never shrinks but keeps moving. Each of those has a fixed amount of area at every stage and none in the limit. A ceiling of finite area is a fence against both.
The second statement is the one an earlier chapter borrowed and promised to repay. It was used to justify differentiating an integral by differentiating inside it, which is the trick that produced every moment of the Gaussian without solving a single new integral. The justification is now complete, and the bound it needs was written down at the time.
a natural place to stop · the integral is built; what follows is the space
5 · , and the functions that are not functions
Now the space itself. We collect the square-integrable functions, put Chapter 0.5's inner product on them, and check its three axioms one at a time. Two hold with no work at all. The third fails, and it fails on exactly the sequence Chapter 0.2 built, so the repair is forced rather than chosen. Making the repair costs something real, and §5.3 says what. Once it is made, every theorem of Chapter 0.5 §1 arrives with no proof needed, and §5.5 records the one way this space is not like .
5.1 · The set, and that it is a vector space
Fix a measurable set , which for our purposes is or an interval, and collect the measurable complex-valued functions on it whose squared modulus has finite integral:
Two things need checking before that definition means anything, and both are one line. First, the collection is closed under addition, since pointwise and the right-hand side is integrable. Second, the inner product converges, since by the arithmetic-geometric mean inequality, so the integral defining is finite for every pair in the set.
That second check is the reason the exponent is rather than anything else. The square is the power for which the inner product of two members always exists, which is why quantum mechanics lives in this space and not in its neighbours.
5.2 · The three axioms, and the one that fails
Chapter 0.5 §1.1 laid out exactly three axioms, and the whole of that chapter follows from them. Take them in order against (4.3.28).
- (i) Conjugate symmetry, . Holds, because and conjugation commutes with the integral.
- (ii) Linearity in the second slot. Holds, because the integral is linear, which §4.1 established for the Lebesgue integral.
- (iii) Positive definiteness, for . Fails.
It fails, and the counterexample is the one this chapter opened with. Chapter 0.2's is non-zero at every rational number, so it is not the zero function, and yet (4.3.18) gives . The same holds for every of §1.3.
The general statement is worth having, because it says precisely how bad the failure is. For non-negative measurable , Chebyshev's bound from the grind box of §3.5 gives , so a vanishing integral forces every one of those sets to be null, and their union is :
So the axiom does not fail badly. It fails by exactly the width of a null set, and (4.3.29) tells you what to do about it.
5.3 · The repair, and what it costs
Since the norm cannot tell two functions apart when they agree almost everywhere, stop asking it to. Declare and to be the same vector when almost everywhere. That is an equivalence relation, and the space of states is the set of equivalence classes:
Axiom (iii) now holds by construction, since says is in the class of the zero vector, which is what " in " now means. All three axioms hold, so is an inner-product space in exactly Chapter 0.5's sense, and this is the moment that chapter was pointing at when it said of its table's function row that "that row is the entire reason Chapter 4.3 is possible."
The quotient is not free, and pretending it is would set up a misunderstanding that costs a whole chapter later.
Ask what is for and the question has no answer. Every single point is a null set, so you may change at to any value you like and it is still the same vector. Pointwise values are not properties of an element of , and no theorem about can produce one.
Three consequences, all of which arrive later in Part IV.
- "The probability of finding the particle exactly at " is not merely small, it is not defined, and this is why the Born rule for a continuous variable is stated with over a region rather than with at a point.
- The Gibbs overshoot of §8.4 is a statement about the values of a function at points, and it is therefore invisible to the norm of . Both statements are true at once and they are not in conflict, because they are about different things.
- The object , which every physics text writes as though it were a vector, cannot be one: a state concentrated at a single point is the zero vector of . Chapter 4.5 says what it is instead, and Chapter 0.9 §5.3 already flagged the same difficulty for .
In exchange for those three, the space acquires a genuine geometry and, in §6, a genuine completeness. It is a good trade, and it is the only one available.
A probability density is one. Take the density of a continuous variable, change its value at finitely many points, or on any set of probability zero, and nothing observable moves. Every probability, every quantile, every moment, every survival function and every likelihood is an integral of the density, so all of them are blind to the change. The density is genuinely not a function with definite values. It is a class of functions agreeing almost everywhere, and a particular representative is chosen for convenience when one has to be written down.
The mathematics is identical, not analogous. Equation (4.3.30) is the statement that a density is well defined as an object of integration rather than as a pointwise assignment, and it is the same quotient by the same relation. A wavefunction gets the same treatment, and it is worth being careful about why. The reason is not that settles everything measurable, because it does not. The distribution of momentum is the squared modulus of the Fourier transform of , which §8.5 below builds out of the phase of as much as out of its modulus, and no amount of determines it. The reason is that is an element of (4.3.30) by construction, so altering it on a null set does not alter it at all.
One place the parallel stops, and it is worth knowing where. A density is usually chosen to be the continuous representative when one exists, and that choice is harmless because it is unique. In there need be no continuous representative at all, and §8.1 has to prove that a continuous function can at least be found nearby.
5.4 · Chapter 0.5's theorems, transferred without a word changed
Here is the return on Chapter 0.5 having been written abstractly. That chapter proved Cauchy–Schwarz from the three axioms alone, and its proof mentions no basis, no dimension and no component. Every step of it is now available here with nothing to check:
The first is Chapter 0.5's (0.5.6) with the vectors renamed, and the second is the triangle inequality that follows from it in three lines there. Both are theorems about inner-product spaces, and is one. Chapter 0.5 said this would happen, in the solution to Problem 1 of its §10, where Gram–Schmidt run on produced the Legendre polynomials: "the machinery does not care that the vectors are functions. That is the whole content of the abstraction, and it is what Chapter 4.3 will exploit." This is the exploitation, and it consists of writing nothing down.
What does not transfer is anything Chapter 0.5 proved using finite dimension, and that chapter listed the four places it used it: "in the induction, in rank–nullity, in the interchange of sums, in the claim that an injective map is surjective." Sections 6 and 7 supply replacements for as much of that as this chapter needs, and Chapters 4.4 and 4.5 handle the rest.
5.5 · Square-integrable does not mean integrable
One structural fact separates from every finite-dimensional space and from of a bounded interval, and it has physical consequences that Chapter 4.9 will need.
On a bounded interval the two conditions are nested. Cauchy–Schwarz against the constant function gives, for ,
so square-integrable implies integrable there. On the whole line the constant function is not in , the factor has nothing to converge to, and the implication fails. Both inclusions fail, in fact, and one example each settles it:
- has and . So and .
- on and elsewhere has and . So and .
The physical reading is worth having now rather than in Chapter 4.9. A wavefunction belongs to because is what normalisation means, and that guarantees nothing about , or about . Worked example 3 exhibits a perfectly normalised state whose mean position does not exist, and the state is one Chapter 0.9 already introduced under another name.
The space of states can now be assembled. Take the functions whose squared size has a finite total, and give any two of them an overlap by multiplying one against the conjugate of the other and integrating. The choice of the square rather than some other power is not arbitrary: it is the only power for which the overlap of two members is guaranteed to be finite, which is why physics lives here.
Three requirements were laid down long ago for anything calling itself an overlap. Two of them hold immediately. The third demands that only the zero vector have zero length, and it fails, on precisely the function this chapter began with. A function supported on the fractions alone has zero length without being zero.
The failure is exactly as wide as a set of size zero, so the fix is to stop distinguishing functions that agree except on such a set. Two functions differing at a scatter of isolated points are declared to be the same vector. This sounds like an evasion and is not: it is the same convention already in force for probability densities, where changing the density at a few points changes no probability, no average and no likelihood, and nobody regards the density as ill-defined on that account.
The price is worth naming, because it is charged later. A vector in this space has no value at any particular point. The question of what the wavefunction equals at one exact position has no answer, and every honest statement about a continuous coordinate is an integral over a region. In return, the overlap becomes a genuine measure of distance, and every theorem about lengths and angles proved for abstract spaces becomes available here with no work at all, because none of those proofs ever asked what the vectors were.
6 · Completeness: the theorem, and the failure it repairs
Everything built so far exists for this section. Section 1 handed you a Cauchy sequence with nowhere to go. Sections 2 to 5 enlarged the class of functions, built an integral that survives limits, and fixed the norm so that distance zero means equality. The claim now is that the enlargement was exactly the right size: not merely bigger, but big enough that no Cauchy sequence can escape it. Section 6.2 states the theorem and marks it as quoted. Section 6.3 then takes the two sequences of §1 and shows you where they land.
6.1 · What is being claimed, and what is not
Convergence in means , which written out is
and is often called convergence in mean square. Read what it says and, more importantly, what it does not. It says the total squared discrepancy shrinks to nothing. It says nothing whatever about the value of at any particular point, and §5.3 already explained why it cannot: the vectors have no pointwise values to make a statement about. Section 8.4 exhibits a sequence converging in the sense of (4.3.33) whose graphs never stop overshooting.
With that understood, here is the claim. Every sequence in that is Cauchy in the sense of (4.3.2) converges, in the sense of (4.3.33), to an element of . An inner-product space with that property is called a Hilbert space, and that is the proper name Chapter 0.9 §1.3 promised the space of square-integrable functions would eventually be given.
6.2 · Riesz–Fischer
Riesz–Fischer. is complete for every measurable , and is therefore a Hilbert space.
The proof is quoted rather than run, and the shape of it is worth having so that the quotation is not a black box. Given a Cauchy sequence, thin it to a subsequence whose consecutive gaps shrink geometrically, , which is always possible. Then apply monotone convergence to the increasing partial sums of to show that the series converges at almost every point, so the telescoping sum defines a function almost everywhere. Fatou's lemma applied to then shows both that is in and that in norm. Finally, a Cauchy sequence with a convergent subsequence converges, by the triangle inequality.
It is proved in this form in the same two measure theory texts §2.3 named, and it is short enough there to read in an afternoon.
What is being assumed, precisely. Nothing beyond §2.3's package and §4's two theorems. The sketch above uses monotone convergence, Fatou, and the definition of the integral, and it uses no new input at all. So this mark is different in kind from the one in §2.3: that one imports a construction the book never performs, and this one marks an argument the book chooses not to write out. Both are honest to declare and only one of them is a genuine debt.
Two remarks before the theorem is spent. First, the sources quoted above prove the same statement for every at once, so nothing about the exponent is doing the work here. What the exponent gives is the inner product of §5.1, and it is the combination of the two that makes the space quantum mechanics needs. Second, the proof produced convergence almost everywhere only along a subsequence, and it is worth being exact about what that does and does not mean. Section 8.4's example is not evidence that the thinning is necessary. Its partial sums converge in norm and they also converge at every point, as §8.4's own theorem and Problem 5(c) both say, so what they fail is uniform convergence, which is a different statement. Whether norm convergence drags a whole sequence to a limit almost everywhere is a question this chapter does not take up, and the subsequence is all the sketch above delivers.
6.3 · The two sequences of §1, and where they land
Now collect. Both sequences from §1 are sequences in , both are Cauchy, and both now have limits. They are different kinds of repair and it is worth taking them in order.
Chapter 0.2's sequence. The functions of §1.3 were at and elsewhere. In each of them is the zero vector, by (4.3.29) and the vanishing of . Their pointwise limit is also the zero vector, for the same reason. So the sequence converges, its limit is in the space, the limit's integral exists, and the integral is
Compare (4.3.34) with (4.3.5), the statement that failed in §1.3. It failed there because its right-hand side referred to nothing. Both sides now refer to something and the two agree. Chapter 0.2's warning box has been answered in full: the definition was thrown away, the integral was rebuilt, and came with it.
The thickened sequence. This is the one that showed a genuine hole, so it is the one worth watching close. The step functions of §1.4 increase to at every point, and is measurable because is open. Its square integral is finite, so it is a vector of , and the distance from the sequence to it is computed by (4.3.9):
So the sequence converges in , its limit is a genuine element of the space, and the limit has an integral, namely , a number at most .
Now read that against the grind box of §1.4, which is where the payoff is. That box proved that no Riemann-integrable function is an limit of the . Since is that limit, and limits in a metric space are unique up to distance zero, the box says exactly this:
The vector the sequence was reaching for is not a relabelling of anything the old theory had. It sits at positive distance from every last one of them. The hole exhibited in §1 was real, the enlargement of §2 to §5 filled it, and (4.3.36) is the receipt.
6.4 · What completeness gets spent on
Chapter 4.2 wrote its whole apparatus on the assumption that limits of states are states, and listed the assumption as owed. Here is what the theorem now pays for, named by chapter so the credit can be tracked.
- Section 7.2, immediately. A series with converges to a vector. Without completeness the coefficients would be square-summable and the sum would be nothing, which would leave "expand in a basis" meaningless.
- Chapter 4.5's spectral theorem. Every construction there produces its operator as a limit of simpler ones, and the limit has to exist in the space it started in.
- Chapter 4.6's time evolution. The operator is defined by a series, and applying it to a state produces a state only because the partial sums are Cauchy and converge.
- Chapter 4.15's perturbation series. The first correction to a state is an infinite sum over the unperturbed basis, and it names a state only because its coefficients are square-summable and §7.2 then supplies a limit.
This is the centre of the chapter. The claim is that the enlarged space has no gaps: any sequence of states that is settling down settles down onto a state, and never onto something outside the space. That statement is what an earlier chapter took entirely on credit every time it said that a limit of states is a state, and it is the property whose absence made the old integral unusable.
The theorem itself is quoted rather than proved, and the quotation is marked. The shape of the argument is short enough to describe: thin the sequence so that consecutive terms are twice as close as the last pair, add up the gaps, and the two convergence theorems of the previous section show both that the total settles at almost every point and that what it settles on belongs to the space.
The satisfying part is what happens to the two sequences from the beginning of the chapter. The first, built from the fractions one at a time, now converges, and the object it converges to is the zero vector, with an area of exactly zero. That was promised eight chapters ago and it is now delivered.
The second is the one that mattered. It converged nowhere before, and now it converges to a perfectly respectable member of the space, an object at a positive distance from every function the old theory could handle. The gap that was exhibited was real, and what fills it is genuinely new rather than a relabelling. That is what it means for the enlargement to have been exactly the right size.
a natural place to stop · the space is complete; what follows is finding a basis in it
7 · Orthonormal bases in infinite dimensions
Completeness in hand, we can finally ask what a basis is here. In finite dimensions Chapter 0.4 settled it by counting, and the count is not available. Two tools come first. Section 7.1 shows that the coefficients of any vector against any orthonormal set are square-summable, and §7.2 shows that any square-summable list of coefficients assembles into a vector, which is where §6 gets spent. What then replaces the count is four statements which look different, are all called completeness of a basis, and turn out to be the same statement. Section 7.3 proves that. Section 7.4 then shows that a basis of this space can always be listed, which is what every physics calculation has been assuming.
One older promise is being kept here too. Chapter 1.1 §4.3 observed that the configuration of a field is a function rather than a list of numbers, that the space of such functions is therefore infinite-dimensional, and that "Chapter 4.3 will take its infinite-dimensionality seriously." Taking it seriously is exactly this section, and the cost is now visible. The dimension count that settled every question in Chapter 0.4 is gone, and four separate statements have to be proved equivalent to replace it.
7.1 · Bessel's inequality
Let be an orthonormal set in , meaning , with no assumption yet that there are enough of them. For any write and look at how far is from . That combination is the projection of onto the span of the first , which by Chapter 0.5 §3.1 is the closest point there, and that is why it is the combination worth measuring against. The inequality itself uses none of it. All it uses is that a squared norm can never be negative, so expand the squared distance and let the cross terms collapse:
That is Bessel's inequality, and the derivation is Chapter 0.5's Parseval computation with the equality weakened to an inequality because the set may not be big enough. Its content is small and its consequence is not: whatever else is true, the coefficients of any against any orthonormal set are square-summable.
7.2 · Square-summable coefficients build a vector
Bessel says the coefficients are well behaved. Completeness turns that into an actual vector, and this is the first place §6 gets spent. Suppose and look at the partial sums . Orthonormality makes the distance between two of them a tail of a convergent series:
So is Cauchy, and by Riesz–Fischer it converges to some vector of . The series therefore means something, for any square-summable coefficients whatever. Without §6 the partial sums would be closing on a gap and the notation would name nothing.
One small tool is needed repeatedly below and is worth extracting. Cauchy–Schwarz gives , so the inner product is continuous in each slot. Norm convergence may therefore be taken inside any inner product, which is used at every step of §7.3.
7.3 · Four statements, and the proof that they are one
The word "complete" is used of a basis as well as of a space, and it means something different there. Worse, it means four different-looking things, and physics texts move between them without comment. Here they are, for an orthonormal set in a Hilbert space .
- (a) Spanning. The finite linear combinations of the come arbitrarily close to every vector of .
- (b) Expansion. Every equals , the series converging in norm.
- (c) Parseval. Every satisfies .
- (d) No blind spot. The only vector orthogonal to every is the zero vector.
An orthonormal set satisfying these is an orthonormal basis. The equivalence is what makes the definition checkable, because (d) is the one you can usually verify and (b) is the one you want to use.
(b) implies (a), since the partial sums of the expansion are finite combinations and they converge to .
(a) implies (d). Let be orthogonal to every , hence to every finite combination of them. By (a) there are finite combinations in norm. Continuity of the inner product from §7.2 then gives
because every term in the sequence is zero. So is the zero vector.
(d) implies (b), and this is the step that uses everything. Given , its coefficients are square-summable by Bessel, so §7.2 says exists in . Pair any against the difference , moving the limit inside by continuity, and only one term survives orthonormality:
So is orthogonal to every , and (d) forces , which is (b).
(b) implies (c), by taking the norm of the expansion. Pythagoras from Chapter 0.5 §1.3 gives , and the norm is continuous because , so letting run turns that into (c). And (c) implies (d) immediately: a vector orthogonal to every has all coefficients zero, so (c) makes its norm zero, so it is the zero vector.
The implications proved are (b) to (a) to (d) to (b), which is a closed cycle containing three of the four, and (b) to (c) to (d), which brings in the fourth. All four are therefore equivalent.
Two things are worth noticing about which chapter supplied which step. The finite-dimensional proof in Chapter 0.5 §2.3 needed none of this, because there the four statements are all immediate consequences of counting dimensions. Here the count is unavailable and completeness of the space is what replaces it, entering at §7.2 and nowhere else. That is the precise sense in which Chapter 4.2's basis expansions were on credit.
7.4 · A basis can always be listed
Everything in §7.3 assumed the orthonormal set was indexed by the integers, and that assumption needs justifying, because a Hilbert space could in principle need an uncountable basis. It cannot, for this space, and the reason is that is not too large to be approximated from a list.
A space is separable when it contains a countable dense subset. Building one for takes three steps, and the second is where §2.3's regularity clause is spent. It is spent once, here, and §8.1 re-uses the same step rather than paying for it again.
- Simple functions are dense. For in take the staircase (4.3.17). Then , which is integrable, and pointwise, so dominated convergence gives . A general complex splits into four non-negative pieces.
- Indicators are nearly intervals. For measurable of finite measure, regularity supplies an open with as small as desired. Any open subset of is a countable disjoint union of open intervals, and the reason is worth a line: take for each point the largest open interval inside containing it, two such intervals are equal or disjoint, and each one holds a rational, so there are countably many. Countable additivity then lets finitely many of them carry all but a little of the measure, and keeping those leaves a set with small, so is small.
- Count them. Perturbing the endpoints of those finitely many intervals to rationals, and the coefficients of the simple function to rationals with rational imaginary parts, changes the norm by as little as you like and leaves a countable collection.
So is separable. Now the consequence, which is three lines. Let be any orthonormal set, and note that distinct members are far apart, since orthonormality and Pythagoras give
Put an open ball of radius around each , and by (4.3.41) the balls are disjoint. A countable dense set must place at least one of its points in every one of them, so the balls can be labelled injectively by a countable set, and there are at most countably many of them. Every orthonormal set in a separable Hilbert space is countable, so every orthonormal basis of can be written as a sequence and expansions can be written as ordinary series.
Existence is settled the same way. Run Chapter 0.5 §2.1's Gram–Schmidt on a countable dense set, discarding any vector that is already in the span of its predecessors. The output is a countable orthonormal set whose finite combinations reach everything the dense set reached, which is statement (a), so it is a basis. The recipe is the one from Chapter 0.5 with no modification, which is a third instance of that chapter's machinery not caring what the vectors are.
One consequence deserves stating, with a caution attached to it. Given two separable infinite-dimensional Hilbert spaces with bases and , the map sending to is defined on everything by (b), lands everywhere by (b), and preserves norms by (c). So it is unitary in Chapter 0.5 §4.3's sense. Up to a unitary map there is only one separable infinite-dimensional Hilbert space, and every question about which one to use is a question about which basis is convenient.
Here is the caution, because that result is easy to over-spend. It depends on a choice of basis at each end and singles out no particular map, so it can never tell you that some given map is the right one. Section 8.5 needs exactly that, and it gets it from Plancherel instead. Nothing later in this chapter uses the statement just proved.
In finitely many dimensions a basis is a set of perpendicular unit vectors, and you know it is big enough by counting: three of them in three dimensions, and there is nothing more to check. Counting is not available here, so the question of whether a perpendicular set is big enough has to be answered some other way.
Four different answers are in circulation and they all get called completeness. The finite combinations reach everywhere. Every vector is the sum of its components. The squared length is the sum of the squared components. And nothing but zero is perpendicular to the whole set. They look like four separate claims, and the section proves they are one claim in four costumes, which is useful because the last is usually the only one you can check and the second is always the one you want to use.
The proof leans on the previous section in exactly one place: to know that a list of components with a finite total square actually assembles into a vector rather than pointing at a gap. That is where completeness of the space is spent, and it is the reason the two senses of the word have to be kept apart.
One more thing has to be true before any of it is usable, and it is easy to overlook. A basis here could in principle be too big to write as a list. It is not, and the argument is a pleasant one: any two distinct perpendicular unit vectors are a fixed distance apart, so small balls around them never overlap, so a countable supply of approximating points cannot serve more than countably many of them. Bases can therefore be numbered, expansions are ordinary series, and everything a physicist writes with a summation sign is legitimate.
8 · The Fourier basis is a basis
Chapter 0.9 proved that the Fourier modes are orthonormal, said in a marked box that it had not proved there were enough of them, and named this chapter. Section 8.1 shows that continuous functions are dense, using nothing new. Section 8.2 proves, constructively, that a continuous periodic function is a uniform limit of trigonometric polynomials. Section 8.3 puts the two together and the mark comes off. Section 8.4 then measures the two kinds of convergence against each other on Chapter 0.9's own square wave, and §8.5 collects the last debt of the chapter.
8.1 · Continuous functions are dense, on the regularity step §7.4 already paid for
This subsection supplies the first of the two links §8.3 needs. Every vector of sits as close as you like to a continuous function, and on to a continuous periodic one, which is what will let §8.2 work only with continuous functions and still reach the whole space.
Nothing here is quoted, and it is worth being explicit about why, because the result looks like the kind of thing that ought to need a flag. Section 2.3's package included regularity, clause (d), and §7.4 spent it. This section re-uses that step and adds one link of its own. Watch the clause do the second of the two jobs it was bought for.
The chain has three links and the first two are already built. Simple functions are dense in , by the staircase argument of §7.4. An indicator is close in to the indicator of a finite union of intervals, which is §7.4's regularity step verbatim. So it is enough to approximate the indicator of a single interval by a continuous function, and that is done by hand.
Replace the vertical sides of by straight ramps of horizontal width , giving a continuous trapezoid . The two functions agree except on the two ramps, where they differ by at most , so
which goes to zero with . Composing the three links, the continuous functions are dense in . On one extra turn of the same screw is needed for §8.2: place the ramps so that the trapezoid vanishes near both endpoints, and it extends to a continuous -periodic function on the line with no jump at the seam. So the continuous -periodic functions are dense in as well.
8.2 · Fejér's theorem, built from a kernel you have already met
Work on , which is Chapter 0.9's box with and therefore , and write Chapter 0.9's coefficients and partial sums:
Direct attack on is the route that fails, and the reason is worth naming in advance. The argument below works by writing the approximation as an integral of against a weight, and it needs that weight to be non-negative. The weight belonging to is the Dirichlet kernel, written out below when the averaging is performed, and it is a ratio of two sines, which changes sign. The route that works replaces the partial sums by their running average, which turns the weight into a square. Define the Cesàro means
To see what averaging has done, substitute (4.3.43) into (4.3.44) and pull the sum inside the integral, which is legal because it is finite. What comes out is a convolution against a single function of :
The inner sum is a geometric series, and summing it gives the Dirichlet kernel . Averaging those over is a telescoping exercise. Multiply by , which turns each into , and everything cancels except the ends, leaving . So
Now read off three properties of (4.3.46), because they are the whole theorem. It is non-negative, being a square, and that is exactly what the averaging bought. Its integral is one, since for every as only the term survives, and the average of ones is one. And it concentrates at the origin, since for the denominator is bounded below and
A non-negative bump, of unit integral, narrowing at the origin. That is Chapter 0.9 §5.2's delta-sequence, and the argument that follows is that section's proof with the Gaussian replaced by (4.3.46).
Fejér's theorem. If is continuous and -periodic then uniformly. Since , the error is
and it is split at . On the inside, is uniformly continuous, being continuous on a closed bounded interval, so can be chosen once and for all to make for every at once. That part of the integral is then below , because is non-negative with unit integral. On the outside, the bracket is at most and (4.3.47) drives to zero uniformly, so that part is below for large. The bound is uniform in throughout.
Two notes on what that proof used. The uniform continuity is Heine–Cantor, which Chapter 0.2 §1.1 quoted and marked when it proved that continuous functions are Riemann integrable; this chapter leans on that standing mark and does not raise a new one. And the theorem is constructive: it produces the approximating trigonometric polynomial explicitly, as an average of partial sums, rather than asserting one exists. Since is a finite combination of the , what the theorem says is that the trigonometric polynomials are dense, in the supremum norm, in the continuous -periodic functions, and that a member of the dense set can be written down from by averaging.
8.3 · The mark comes off
Now assemble, and note that each link was proved above and none is quoted.
- Uniform convergence implies convergence on a bounded interval, since . So by Fejér's theorem in §8.2 the trigonometric polynomials come arbitrarily close in to every continuous periodic function.
- Those are dense in , by §8.1.
- Therefore the trigonometric polynomials are dense in , which is statement (a) of §7.3 for the set with .
- Chapter 0.9 §1.1 proved by direct integration that this set is orthonormal.
By the equivalence of §7.3, all four statements hold. So the Fourier modes are an orthonormal basis of , and in particular, for every with ,
That chapter wrote, in a box headed "Quoted, not derived" and carrying the mark: "Everything above shows that the are orthonormal. It does not show that they are enough… Both facts are proved in Chapter 4.3, where the space of square-integrable functions gets its proper name, a Hilbert space, and completeness stops being an assumption. Until then we are standing on it, and you know that we are."
Both facts are now proved. The first is the left half of (4.3.49), and it came from Fejér plus the equivalence of §7.3. The second, about pointwise convergence and what happens at a jump, is §8.4 below. The space has its name, by §6.2. Nothing in Chapter 0.9 changes, and nothing in it was wrong. What changes is that its mark has been paid rather than carried.
8.4 · Two kinds of convergence, measured against each other
Equation (4.3.49) is a statement about a norm, and Chapter 0.9 §1.4 exhibited a function for which it is emphatically not a statement about points. This is the last piece of that chapter's mark, and it is also the numerical confirmation this chapter owes.
First the pointwise theorem, so that the two claims can be set beside each other. It rests on one lemma about oscillating integrals, and the lemma is worth stating because it recurs throughout Part V.
Riemann–Lebesgue lemma. If is integrable then as . The reason is that rapid oscillation averages a slowly varying weight against itself, and the proof turns that into three lines in the grind box.
With it, the classical statement about jumps follows. If is -periodic and integrable, and at the point both one-sided limits exist with bounded one-sided difference quotients, then . At a point of continuity that is , and at a jump it is the midpoint, exactly as Chapter 0.9 stated without proof.
One word about the hypothesis, because Chapter 0.9 gave it a name and this one does not match it. That chapter asked for the Dirichlet conditions, piecewise smooth with finitely many extrema and jumps, imposed on across the whole interval. The statement above asks less. It is a condition at the single point , and away from it wants nothing beyond integrability. The same conclusion from a weaker hypothesis is the stronger theorem, so the Dirichlet conditions are superseded here rather than merely met.
Grind box — Riemann–Lebesgue, and convergence to the midpoint of a jump
Riemann–Lebesgue. For the integral is computed outright,
so it holds for step functions by linearity. The regularity argument of §7.4, run with the norm in place of the norm, makes step functions dense in . So given pick a step function with , and
which is below for large .
The jump. Summing the geometric series in (4.3.43) gives with . The kernel is even and , so and therefore
Take the first. Writing turns it into . Near the numerator is at most a constant times by the bounded difference quotient, and behaves like , so is bounded near the origin and integrable on . Riemann–Lebesgue kills it, and the second term goes the same way.
Now put the two convergences on the same function and watch them separate. Chapter 0.9 §1.4's square wave on is on and on , with the series it derived there,
Write for the partial sum keeping every odd harmonic up to . Parseval, which is now a theorem rather than an assumption, turns the norm of the error into a tail of a series. Using and squaring the coefficients of (4.3.50),
the tail being estimated by an integral because the odd integers past are spaced by two. Taking the square root gives the rate this chapter promised to confirm:
Against that, set the pointwise behaviour. Chapter 0.9's grind box computed the peak of the partial sum near the jump and found it converging to rather than to . The jump has size , so the overshoot, as a fraction of the jump, is
Two numbers, both exact, both about the same sequence of functions, and they do opposite things. Equation (4.3.52) falls to zero. Equation (4.3.53) settles on a constant and stays there forever.
They are consistent, and the reconciliation is one sentence. The overshoot keeps its height and loses its width. The first peak sits at , so the whole ear is squeezed into a window shrinking like , and a fixed height on a vanishing width contributes a vanishing amount to . A norm integrates and cannot see it. A maximum does not integrate and sees nothing else. The figure below is that sentence made visible, and it is worth using the zoom control, because the ear is invisible at any fixed magnification once is large and is exactly the same height at every magnification that tracks it.
The physical reading matters more than the arithmetic, and it is the reason §5.3 warned about pointwise values in advance. Every statement quantum mechanics makes about a continuous coordinate is an integral over a region, so the convergence that has physical content is the one in (4.3.52). A claim about the value of a wavefunction at a point is not a claim can support, and the Gibbs ear is what that looks like when you insist on making one.
8.5 · Position and momentum are one space, joined by a unitary map
The last debt of the chapter is Chapter 0.9's closing brick, which listed where its results would be spent and sent the first entry here: "The Fourier basis and Plancherel → Chapter 4.3 (where completeness is finally proved and the position and momentum representations become two bases for one Hilbert space)."
Assemble what is now available, and it is one thing rather than two. Chapter 0.9 §2.3 proved Plancherel, that the transform preserves the norm, and read it back as the statement that the transform is unitary. That is the whole of what is needed here, and it is a statement about this one map. Section 7.4's result, that any two separable infinite-dimensional Hilbert spaces are unitarily equivalent, singles out no map at all and does no work in this section.
So what Plancherel says is that and are not two objects but two descriptions of one. Square-integrable functions of and square-integrable functions of are the same Hilbert space, and the transform is a unitary map between the two descriptions. Chapter 0.9's phrase two bases for one Hilbert space is the loose reading of that, and the box below says exactly how loose.
Writing it with in place, as Part IV does everywhere, the momentum-space wavefunction is rescaled so that its own integral is one, with the visible:
The normalisations agree because the change of variable contributes exactly the that the prefactor removes, which is the reason the symmetric convention of Chapter 0.9 §2.2 was chosen there. So a state has one norm, and the Born rule computes the same total probability in either description.
It is tempting to read (4.3.54) as saying that and are two orthonormal bases in the sense of §7.3. They are not, and nothing above claims they are.
Section 7.4 proved that every orthonormal basis of this space is countable, and the positions on a line are not countable. The functions are not in at all, since has infinite integral, which Chapter 0.9 §5.3 had already noticed and called a real gap. And a state concentrated at one point is the zero vector, by §5.3 here.
What is established is exactly this: in the variable and in the variable are the same Hilbert space, and (4.3.54) is a unitary map between the two descriptions. That is enough for every computation in Chapters 4.6 to 4.10, and it is what physicists mean when they say the two representations are equivalent. Chapter 4.5 closes the specific gap Chapter 0.9 named, by giving and a precise meaning through box normalisation and a limit that always works. The general theory of objects like them is Chapter 5.4's, and this chapter claims neither.
Nine chapters ago the pure waves were shown to be mutually perpendicular, and it was said plainly that being perpendicular is not the same as being numerous enough, and that the second claim was being borrowed. This section returns it.
The route has two halves. First, any state in the space can be approached by a continuous function, which re-uses the one quoted clause about squeezing sets between open and closed ones. That clause was spent a section earlier, to show the space has a countable basis at all, and this is the second job it does. Second, any continuous repeating function can be approached by a finite combination of pure waves, and the argument for that is constructive rather than abstract: average the partial sums instead of taking them, and the resulting weight becomes a positive bump of unit area narrowing onto a point. That bump is the same device used earlier to make sense of an idealised spike, reused without alteration.
Putting the halves together, the waves reach everything, so they are a basis, and the borrowing is repaid.
What remains is a warning that is also the most interesting fact in the section. The sense in which a series of smooth waves reproduces a function with a jump is not the sense you would guess. Measured by total squared discrepancy the approximation improves without limit, falling in proportion to the inverse square root of the number of terms. Measured by the worst error near the jump it does not improve at all: every partial sum overshoots by about nine per cent of the step, forever. Both are true, because the overshoot keeps its height and loses its width, and an integral cannot see a fixed height on a vanishing width. Which of the two is the physically meaningful one is not a matter of taste. Every measurable prediction is an integral over a region, so the first is the one that counts, and the second is what happens to anyone who asks the theory for the value of a wavefunction at a point.
9 · Worked examples
On let , with a real parameter. (a) Find the pointwise limit. (b) For which does ? (c) For which does in ? (d) Find the smallest possible dominating function and say for which it is integrable. (e) Compare (b) and (d) and say what the comparison establishes about dominated convergence.
(a) Fix . As soon as the point is outside the interval where lives, so from then on. The pointwise limit is at every point of , for every whatever. Notice that the parameter has not appeared yet, so any dependence on below is a statement about the integral rather than about the functions.
(b) The integral is height times width:
which tends to when , sits at for every when , and diverges when . Since , the interchange is valid exactly for .
(c) Squaring doubles the exponent and leaves the width alone:
So convergence in is a strictly stronger demand than convergence of the integrals. In the window the integrals converge to the right answer while the norms do not shrink at all, which is worth holding on to: a sequence of states can have its expectation values settle without the states themselves converging.
(d) Any dominating function must exceed every member of the family, so the smallest one is . For the functions that are non-zero at are those with , and for the largest of them wins, so is for the largest integer below , which is up to a bounded factor. Then
So an integrable dominating function exists exactly when , and for the family is bounded by a constant and the question is trivial.
(e) Compare the two answers. The interchange in (b) is valid for . A dominating function exists in (d) for . The two conditions coincide, so on this family dominated convergence is not merely sufficient but sharp. There is no for which the conclusion holds and the hypothesis fails, and none for which the hypothesis holds and the conclusion fails. That is unusual, and it is the reason this family is the right one to keep in mind: when you cannot find a dominating function it is worth suspecting that the interchange is not merely unproved but false.
Two further readings. Applying dominated convergence to instead needs to be integrable, which asks and matches (c) exactly, so the same sharpness holds for the statement. And Fatou's lemma is consistent throughout, reading , with equality for and strict inequality for , which is mass escaping up the spike.
(a) Apply Parseval to Chapter 0.9's square wave and evaluate , then . (b) Apply it to on and evaluate . (c) Use (b) to finish Chapter 4.1's integral (4.3.23). (d) Use (a) to derive the exact error of the square wave's partial sums and confirm the rate of (4.3.52).
(a) The square wave has , and by (4.3.50) its expansion in the orthogonal set , each of squared norm , has coefficients for odd and zero otherwise. Parseval, which §8.3 turned into a theorem, equates the two:
The full sum follows by splitting off the even terms, which are the full sum again with every doubled. Writing , the even part is , so and . That is the Basel sum, obtained here as a by-product of a square wave.
(b) Take on and use the complex form (4.3.43). The coefficient is the mean, . For , two integrations by parts give , and the sine part vanishes by oddness, so . Parseval in the form (4.3.49) reads
the factor counting and together. Rearranging, , so .
(c) Section 4.1 used monotone convergence to write Chapter 4.1's integral as and stopped there. With (b) in hand,
which is the number the Stefan–Boltzmann constant of Chapter 4.1 was computed from. Both steps of that computation are now theorems: the term-by-term integration is monotone convergence and the resulting sum is Parseval.
(d) The error of a partial sum is the tail of Parseval's series, so from (a),
which is exact at every and is the quantity the figure plots. Dividing by and taking the root gives the relative error
Now the asymptotics, and the estimate is worth making carefully because the approach turns out to be slow. The odd integers beyond are spaced by two and itself is odd, so they sit at the midpoints of the length-two intervals tiling . Their reciprocal squares therefore sum to , with a midpoint-rule error of order . So
The rate is , which is (4.3.52), and the deficit factor in front of it is why the constant arrives so slowly. That factor is at , so the product is still half a per cent short of there, and it is at .
The figure's first readout computes exactly and , so both numbers can be checked by moving the slider. At it reads and . At it reads and . Against , the rate is confirmed and the constant is confirmed with it, with the last three digits still on their way.
Take , with a length. (a) Check that is normalised and identify . (b) Decide whether is in . (c) Decide whether exists. (d) Say what is wrong with the argument that by symmetry. (e) Say where such a state comes from physically.
(a) The squared modulus is , whose integral is
So is a perfectly good normalised state, and is the Lorentzian of Chapter 0.8 §6.4, which Chapter 0.9's grind box on the central limit theorem identified as the Cauchy density.
(b) No. For large the function itself falls only as , so
a logarithmic divergence at both ends. This is §5.5's first bullet in the flesh: and , and there is nothing pathological about the state that produced it.
(c) No. By §3.3 a function is Lebesgue integrable only when the integral of its modulus is finite, and
So is not integrable and does not exist. The same computation with diverges worse, so does not exist either, and neither does the position uncertainty. The state is normalised and has no mean position.
(d) The symmetry argument computes something, and the something is not the integral. What it computes is
which is a limit of proper integrals over a particular family of symmetric windows. Section 3.6 is exactly about this. The Lebesgue integral sorts values before summing them and can never depend on the order in which cancellations arrive, so a quantity defined by the order of arrival is not one of its values.
To see how completely the answer belongs to the family rather than to the state, move the right edge. The antiderivative is a logarithm, so for any fixed ,
Every one of those limits is finite and no two of them agree. Taking recovers the the symmetry argument produced, and taking gives , which is for . So the family does not merely disturb the answer. It sets it.
Divergence needs more than a wider window: it needs the right edge to outrun the left one faster than any fixed ratio. The family does that, and the same antiderivative gives , which grows like and has no limit. The number is a property of the window family, not of the state, and no family is the one the state prefers.
(e) From any resonance. Chapter 0.8 §6.4 derived the Lorentzian as the response of a damped oscillator near its natural frequency and identified its width as the decay rate, and Chapter 0.9 noted that Cauchy tails "arise from ratios of normal variables, from resonance line shapes… and in any setting where the largest single contribution is comparable to the sum of the rest." An unstable state has a Lorentzian energy distribution, so its mean energy does not exist as a Lebesgue integral, and every quoted "mean energy" of such a state is a symmetric window in disguise. That is not a defect in the physics. It is a warning that the quoted number depends on a truncation, and it is exactly the kind of thing the rebuilt integral was designed to make visible.
10 · Your turn
Problem 1 — a space with a hole in it, built by hand
(a) Show that a Cauchy sequence with a convergent subsequence converges, and say where the triangle inequality enters. (b) On let be the continuous function that is for , rises linearly to across the interval of width centred at , and is thereafter. Show that is Cauchy in the norm and identify its limit. (c) Show that no continuous function is that limit, so the continuous functions with the norm form an incomplete space. (d) Show that is not Cauchy in the supremum norm, and say why that does not rescue anything. (e) Which of §7.3's four statements could fail in an incomplete inner-product space, and at which step of the proof?
Solution
(a) Let and let . Choose with for , then choose with and . For any the triangle inequality gives . The triangle inequality is what lets the subsequence act as a stepping stone, and it is the only tool used.
(b) For the two functions agree outside the wider ramp, which has width , and differ by at most inside it. So , which goes to zero, so the sequence is Cauchy. The same estimate against the step function gives , so the limit is that step function.
(c) Suppose is continuous with almost everywhere. On the two agree off a null set, and a continuous function agreeing with off a null set is identically there, since otherwise it would be non-zero on a whole interval, which has positive measure. So on and, by the same argument, on . Continuity at then demands . So the sequence is Cauchy in and converges to nothing in it.
(d) At we have while , so , which tends to as grows with fixed. So the sequence is not Cauchy in the supremum norm. That does not rescue the space, because completeness must hold for the norm you are actually using, and the norm quantum mechanics uses is the one, since it is the one that comes from an inner product and therefore carries the geometry.
(e) Statement (b), the expansion, is the one at risk, and the step is §7.2. There the partial sums were shown to be Cauchy and completeness was invoked to give them a limit. In an incomplete space the coefficients would still be square-summable by Bessel and the series would still be Cauchy, and it would converge to nothing. Statements (a), (c) and (d) can all still be stated, and the chain of implications breaks precisely at (d) implies (b).
Problem 2 — thin and small are different properties
(a) The middle-thirds Cantor set is what remains of after removing the open middle third of every interval at every stage, forever. Show that it has measure zero. (b) Now remove, at stage , an open interval of length from the centre of each of the intervals then present. Show that what remains is closed, contains no interval, and has measure . (c) Conclude that "contains no interval" and "has measure zero" are independent properties, and say which of the two Chapter 0.2's Riemann integral is sensitive to. (d) Show that the indicator of the set in (b) is not Riemann integrable, by the argument of §1.4 run in the other direction.
Solution
(a) At stage there are intervals each of length , so what remains after stages has measure . The Cantor set is contained in every one of those, so its measure is at most for every , hence zero. It is nevertheless uncountable, since the ternary expansions using only the digits and are in bijection with the binary expansions of , so "null" is much weaker than "countable".
(b) The set is an intersection of closed sets, hence closed. It contains no interval, because after stages every remaining interval has length below , so any interval inside the set would have to have length zero. The total length removed is
so the measure of what remains is . This is called a fat Cantor set.
(c) The set in (a) contains no interval and has measure zero. The set in (b) contains no interval and has measure . An interval contains an interval and has positive measure. So the two properties are independent. Upper and lower Riemann sums see only whether a set is dense in a cell or absent from it, so the Riemann integral is sensitive to the topological property and blind to the measure-theoretic one, which is precisely the defect §1.4 exploited.
(d) Let be the fat Cantor set and take any partition. Every cell contains points outside , since contains no interval, so on every cell and every lower sum is . The cells that meet have and their union covers , so their total length is at least and every upper sum is at least . The gap is at least at every mesh, so is not Riemann integrable. Its Lebesgue integral is , by (4.3.15) in one line.
Problem 3 — the three ways mass escapes
On define , then , then . (a) Show that all three tend to zero at every point and that all three have integral . (b) Which of them converge to zero in ? (c) For each, compute of the family and show it is not integrable. (d) State what Fatou's lemma says about each, and check it. (e) Say in one sentence each what physical process the three pictures correspond to.
Solution
(a) For fixed , the spike misses once , the plateau has height , and the travelling bump has left behind once . Each integral is height times width, which is , then , then .
(b) Squaring the height and keeping the width, , which diverges, , which tends to zero, and , which does not move. So only the plateau converges in , and it does so while its integral stays pinned at . Convergence in and convergence of integrals are independent conditions on , in both directions.
(c) For the spike, the largest with gives comparable to on , whose integral diverges at the origin. For the plateau, is comparable to for large , whose integral diverges at infinity. For the travelling bump, , whose integral is infinite. In each case the smallest candidate ceiling has infinite integral, so no dominating function exists and dominated convergence does not apply. It could not, since the conclusion is false in all three.
(d) Fatou says , which reads in all three cases. It holds, strictly. That is the content of the inequality: mass may disappear in the limit but may not appear, and each of these loses exactly one unit of it.
(e) The spike is a pulse of fixed energy compressed into a shorter and shorter time, which is Chapter 0.9's delta being born. The plateau is a fixed quantity spread ever more thinly, which is a wavepacket dispersing. The travelling bump is a fixed quantity leaving the region of interest without changing shape, which is a particle escaping to infinity. Chapter 4.6 has to handle exactly this for a free packet, and Chapter 4.7 for the unbound states of a genuine potential.
Problem 4 — an orthonormal set that is not a basis
(a) Prove Bessel's inequality by expanding with , and say which axiom of Chapter 0.5 §1 makes the left-hand side non-negative. (b) On show that , for ranging over the integers, is an orthonormal set. (c) Exhibit a non-zero orthogonal to every , and conclude that the set is not a basis. (d) Compute both sides of Parseval for your and say by how much the inequality of (a) is strict. (e) Identify the subspace of which is a basis.
Solution
(a) Expanding, and using to collapse the double sum,
So the partial sums of are bounded by , and being increasing they converge to something no larger. The left-hand side is non-negative by axiom (iii), positive definiteness, which is the axiom §5.3 had to repair by a quotient before any of this was available for functions.
(b) The inner product is , which is when and otherwise vanishes, since is a non-zero integer and Chapter 0.9's computation (4.3.43) applies unchanged.
(c) Take . Then for every , because is odd and therefore never zero. Yet . So statement (d) of §7.3 fails, and by the equivalence proved there all four fail.
(d) Parseval would say . The inequality of (a) is therefore strict by the entire norm of . Bessel is not close to being an equality here, which is worth noticing: an orthonormal set can miss a whole direction rather than a little of one.
(e) Each satisfies , so every finite combination does, and so does every limit of such combinations. Conversely a -periodic function on is determined by its restriction to an interval of length , on which the functions are the full Fourier system rescaled and therefore a basis by §8.3. So is an orthonormal basis of the subspace of -periodic functions, which is a proper closed subspace of .
Problem 5 — the two convergences, quantitatively
Work with the square wave and its partial sums throughout. (a) Using , write exactly, and recover . (b) The first peak of sits at and overshoots by a fixed amount. Show that the contribution of the window to tends to zero, and say which factor is responsible. (c) Section 8.4's theorem says at every . Explain in one paragraph why that does not contradict an overshoot that never shrinks. (d) A detector records a band-limited version of a step, keeping harmonics up to . State what it measures in each of the two senses, and which of the two a physical measurement returns.
Solution
(a) Parseval gives , using the total . The tail is a sum over odd integers past , spaced by two, so it is about , whence . Dividing by and taking the square root gives .
(b) On that window is bounded by a constant, since the partial sums themselves are bounded by about near the jump and is bounded by . So the contribution is at most a constant times the width of the window, which is and tends to zero. The height is not responsible for anything, because it does not change. The width is. Comparing with (a), the window contributes at the same rate as the whole error, so the ear is not a negligible part of the error. It is a fixed fraction of a total that is itself vanishing.
(c) Fix any . The overshoot happens at , which tends to zero, so for large enough the peak lies strictly between and , and the value at itself is converging to . Every point is eventually to the right of the trouble. What never happens is that the trouble goes away, because for every there is some point at which the error is about . The maximum over is taken at a moving location, and pointwise convergence says nothing about a quantity evaluated at a moving location. Convergence is pointwise but not uniform, and the Gibbs overshoot is the exact measure of the gap between those two words.
(d) In the sense it measures the step to within a relative error , which for a hundred harmonics is about six per cent and falls as more are kept. In the pointwise sense it rings at the edge by of the step no matter how many harmonics are kept, which is the ringing Chapter 0.9 §1.4 said every lens, detector and digital filter produces. A physical measurement integrates over a region, a time window or a pixel, so it returns the first. The second is what the record looks like if you plot it, and it is why sharp edges and finite bandwidth are incompatible, a statement Chapter 0.9 §6 turned into a theorem.
The Riemann integral was discarded for a reason, and the reason was exhibited rather than asserted. Chapter 0.2's is not integrable, which alone proves nothing. Chapter 0.2's own sequence is increasing, has every integral equal to zero, and has a limit outside the class, so no convergence theorem is available. And thickening that same sequence, by giving the -th rational an interval of length , produces a Cauchy sequence of step functions whose limit is at a positive distance from every Riemann-integrable function on the interval. Density let a tagged sum at every mesh be pushed up to and total length held the integral below , and those two facts cannot be reconciled by any function the old theory can handle.
What replaced it. Size, first: three demands, one of them countable, and one construction quoted. Then an integral that slices the range instead of the domain, defined on a class closed under countable limits by construction. Then the two theorems the whole rebuild existed for, one proved from continuity from below alone and the other from it by way of Fatou. Then the space: square-integrable functions with Chapter 0.5's inner product, two axioms holding for free and the third repaired by declaring functions equal almost everywhere to be one vector.
And then completeness, which was the point. Every Cauchy sequence in converges in , so the space earns the name Hilbert space that Chapter 0.9 promised it. The two sequences of §1 were handed back with limits attached. Chapter 0.2's converges to the zero vector, with as that chapter said it would. The thickened one converges to a genuinely new vector, and §1.4's grind box is the proof that it is new rather than a relabelling.
Bases, and the Fourier modes. Four statements of completeness for an orthonormal set were proved equivalent, with completeness of the space entering at exactly one step. Every orthonormal set here is countable, because two of them are always apart and a countable dense set cannot serve uncountably many disjoint balls, so "expand in a basis" means an ordinary series. Fejér's theorem then supplies the density of trigonometric polynomials constructively, out of a positive kernel of unit integral narrowing at the origin, which is Chapter 0.9 §5.2's argument reused without alteration. Chapter 0.9 §1.3's mark is paid, and it stays on the page where it was raised, naming this chapter as the payer.
The number this chapter owed. For the square wave, the relative error of the partial sums is exactly , which falls like across three decades in the figure, while the overshoot at the jump sits at of the jump and does not move. Both are measured in the figure rather than quoted, and the readout reports the two slopes as and . The reconciliation is that the ear keeps its height and loses its width.
Two marks, and that is the claim. The construction of Lebesgue measure in §2.3, quoted as a package with its five clauses written out. The proof of Riesz–Fischer in §6.2, whose shape was described and whose inputs are only §2.3 and §4. Nothing else in this chapter is asserted without being derived, and the Vitali set of §2.5 was built rather than mentioned, so the first mark is a real restriction rather than a formality. One clause of that first package, regularity, was deliberately cashed in §7.4 to make the space separable, and re-used in §8.1 to prove that continuous functions are dense. Neither result carries a separate mark of its own, precisely so that you can watch a flag do the work it claimed. This chapter also leans on one mark already standing elsewhere, Heine–Cantor from Chapter 0.2 §1.1, used once in §8.2 and cited there rather than re-raised. Two new marks for the chapter that rebuilds the integral, proves the space complete, and establishes the basis every later calculation expands in, is the honest count.
Where this gets spent. Chapter 4.4 takes the space built here and puts operators on it, which is the second instalment of the bill Chapter 4.2 named, and it needs completeness at every step for domains and for the difference between symmetric and self-adjoint. Chapter 4.5 is the third instalment, and needs it again for spectra with no eigenvectors in the space, and for the precise meaning of and that §8.5 declined to claim. Chapter 4.6 needs to map states to states, which is completeness applied to a series. Chapter 4.5 also proves the Hermite functions complete, and that proof is statement (d) of §7.3 above run on an integral that only the dominated convergence of §4.3 licenses, after which Chapter 4.8 expands in them freely. Chapter 4.9 uses Cauchy–Schwarz in the form §5.4 transferred, and needs to explain why a normalised state can have no mean position. And every chapter of Part V expands a field in modes and integrates term by term, which is monotone convergence and Parseval, both proved here. The integral, the two convergence theorems, the space, and the basis: that is the whole of what Part IV and Part V spend, and none of it is on credit any longer.