Part IV · Quantum Mechanics — Chapter 4.3

Function Spaces: Measure, L², and Completeness

Three promises come due together: the integral is thrown away and rebuilt, the space of states becomes a space, and completeness stops being an assumption.

Where we are

Chapter 4.2 ended by saying what it had bought on credit. Every time that chapter wrote "expand in a basis", it was assuming that the basis reaches everything, and every time it wrote "the state after the measurement", it was assuming that the object it produced was still a state. In finite dimensions both are theorems of Chapter 0.5. For a particle on a line the space is a space of functions, and neither one has been established at all.

So this chapter builds the space. Chapter 0.4 built the vector space and Chapter 0.5 built the operators on it. This chapter builds the space and Chapters 4.4 and 4.5 build the operators on it, in that order and for that reason, which is the shape Chapter 4.2's closing brick announced.

Three separate promises fall due here, and all three were made long before Part IV. Chapter 0.2 said the integral would be thrown away and rebuilt from scratch. Chapter 0.2 also used dominated convergence to legalise a step and said it would be proved properly here. Chapter 0.9 quoted the completeness of the Fourier basis, marked the quotation, and named this chapter as the place the mark comes off. The work is one piece, because you cannot prove that a basis reaches everything until you know what "everything" is, and you cannot say what "everything" is until the integral defining the norm behaves under limits.

Here is the route. Section 1 reopens the hole Chapter 0.2 dug in front of you, and it will be the same hole and the same function. Section 2 replaces the idea of length by the idea of measure, and quotes one construction. Section 3 builds the new integral. Section 4 proves the two theorems that make it worth having, and those two theorems are the point of the whole rebuild. Section 5 assembles the space of square-integrable functions and checks Chapter 0.5's axioms against it one by one. Section 6 is the centre of the chapter: the space is complete, and the exact sequence that had no limit in §1 acquires one. Sections 7 and 8 turn that into a working basis and show that the Fourier modes really are one.

Conventions. This chapter is real analysis and uses no complex analysis anywhere, in keeping with the rest of Part IV. The Fourier convention is Chapter 0.9's, unchanged. The inner product is linear in its second slot, as Chapter 0.5 §1.1 chose it. Two results are used without being proved here, and both are marked where they enter: the construction of Lebesgue measure in §2.3, and the proof of the completeness theorem in §6.2. A third is quoted and its mark is not this chapter's to raise: Heine–Cantor, standing since Chapter 0.2 §1.1, which §8.2 leans on and cites rather than raising again. There are no others. A chapter this heavy carrying two new marks is the claim it is making, and §9 and the closing brick say where each one is spent.

Tools you'll need  — Chapter 0.2 §1.1 above all, for the Riemann integral as a limit of tagged sums, for the function χ\chi, and for the sequence of functions its warning box constructed and then set aside. Also its §4.4 for differentiation under the integral sign, and the grind box there that named the hypothesis this chapter proves. Chapter 0.5 §1 for the three inner-product axioms and Cauchy–Schwarz, §2 for orthonormal bases, Parseval and the resolution of the identity, and §3.1 for the fact that a projection is the closest point in a subspace. Its §1.2 table has the row this chapter exists to make legitimate. Chapter 0.9 §1 for the Fourier modes and their orthogonality, §1.3 for the completeness assumption and its mark, §1.4 for the square wave and the Gibbs overshoot, §2.3 for Plancherel, and §5.2 for the argument about a narrowing bump that §8 reuses almost word for word. Chapter 0.3 §3 for convergence and the tests that decide it without naming a limit. Chapter 0.4 §2 for basis and dimension, which is the argument §7 has to replace. Chapter 0.8 §6.4 for the Lorentzian, which Worked example 3 turns into a state. Chapter 4.2 §8 for the three lines that force infinite dimension, and its closing brick for the division of labour between this chapter and the next. Chapter 4.1 §5 for the integral whose term-by-term evaluation §4.1 here makes airtight.

1 · Why the Riemann integral has to go

Here is the debt this section collects, and it is Chapter 0.2's. We look again at the function that chapter could not integrate, define the one piece of vocabulary the book has been using without ever writing down, and then exhibit the failure that actually matters. The failure is not that some function is not integrable. It is that the class of integrable functions is not closed under the limits we need to take, and by the end of the section you will have a sequence with nowhere to converge to.

1.1 · The function Chapter 0.2 left on the table

Chapter 0.2 defined the integral as a limit of tagged sums and then, in a warning box attached to that definition, told you the definition would not survive. The words were these: "In Chapter 4.3 we will throw away [the definition] and rebuild the integral from scratch (the Lebesgue integral), and the reason is visible already." We are going to do exactly that, so let's start by looking again at what was visible already.

The function is Chapter 0.2's χ\chi, and we keep its name. It takes the value 11 at every rational number and 00 at every irrational one. Now recall why the definition of the integral cannot cope with it. A tagged sum lets you choose the sample point in each cell, and every cell, however short, contains rationals and irrationals both. So you may tag entirely at rationals, or entirely at irrationals, and on [0,1][0,1] the two choices give

iχ(xi)Δxi  =  1oriχ(xi)Δxi  =  0, \sum_{i} \chi(x_i^{*})\,\Delta x_i \;=\; 1 \qquad\text{or}\qquad \sum_{i} \chi(x_i^{*})\,\Delta x_i \;=\; 0, (4.3.1)

and both are available at every mesh, however fine. The definition asks for a single number that every tagging approaches, and there are two numbers that every mesh can produce. So χ\chi is not Riemann integrable, and no refinement helps, because refining does not remove the choice.

Now let's be honest about how much that proves, because on its own it proves very little. A single function that cannot be integrated is an annoyance rather than a crisis. Physics is not short of integrals it declines to attempt, and a pathology you can name and step around costs nothing. If this were the whole case against the Riemann integral, the sensible response would be to note χ\chi as a curiosity and carry on.

The real case is about limits, and to state it we need one word the book has been leaning on without ever defining.

1.2 · Cauchy sequences, defined here

Chapter 0.3 §3 gave you tests that decide whether a series converges. The ratio test is the one you use most, and notice what it does: it settles convergence without producing the limit. That is the whole trick, and it is worth naming, because everything in this chapter turns on the difference between a sequence that is heading somewhere and a sequence that has somewhere to go.

Work in a space with a norm, which for now means Chapter 0.5's v=v,v\lVert v\rVert=\sqrt{\avg{v,v}}. A sequence v1,v2,v_1,v_2,\ldots converges to vv when vnv0\lVert v_n-v\rVert\to0, and that definition names the limit. The alternative names nothing outside the sequence itself. Call (vn)(v_n) a Cauchy sequence when its terms eventually huddle together:

for every ε>0  there is an M  with  vnvm<ε  for all n,mM. \text{for every } \varepsilon\gt0 \ \text{ there is an } M \ \text{ with } \ \lVert v_n-v_m\rVert\lt\varepsilon \ \text{ for all } n,m\ge M. (4.3.2)

The two conditions are not symmetric, and the asymmetry is the subject of this chapter. One direction is immediate. If vnvv_n\to v then, given ε\varepsilon, take MM beyond which every term is within ε/2\varepsilon/2 of vv, and the triangle inequality of Chapter 0.5 §1.4 closes it:

vnvm  =  (vnv)(vmv)    vnv+vmv  <  ε. \lVert v_n-v_m\rVert \;=\; \lVert (v_n-v)-(v_m-v)\rVert \;\le\; \lVert v_n-v\rVert+\lVert v_m-v\rVert \;\lt\; \varepsilon. (4.3.3)

Every convergent sequence is therefore Cauchy. The converse is the interesting question, and the first thing to see about it is that it is not a question about the sequence at all. It is a question about the space the sequence lives in.

Take the rational numbers with the ordinary absolute value, and take the decimal truncations 1,1.4,1.41,1.414,1,\,1.4,\,1.41,\,1.414,\,\ldots of 2\sqrt2. Any two terms beyond the MM-th differ by less than 10M10^{-M}, so the sequence is Cauchy. It does not converge in Q\mathbb{Q}, because the number it is closing in on is not a rational number. Nothing is wrong with the sequence. Something is missing from the space, and the missing thing has a name.

A space in which every Cauchy sequence converges to a point of the space is called complete. The real numbers are what you get by insisting on exactly that: you take the Cauchy sequences of rationals and declare the ones that ought to have the same limit to be the same number. Completeness enters this book at that point, as part of what a real number is rather than as a theorem about them, and Chapter 0.3's convergence tests all quietly stood on it. A test that certifies convergence without exhibiting a limit is only worth having in a space where the limit is guaranteed to be there.

That is the property we are going to need for functions, and §6 is where it is established. Keep (4.3.2) in view, because the rest of this section is one sequence of functions that satisfies it and has nowhere to land.

1.3 · Chapter 0.2's sequence, and the two failures it already exhibits

The same warning box in Chapter 0.2 that promised the rebuild also built the sequence that forces it, and we use that sequence rather than a new one. Enumerate the rationals as q1,q2,q3,q_1,q_2,q_3,\ldots, which can be done, and let fnf_n be the function on [0,1][0,1] equal to 11 at q1,,qnq_1,\ldots,q_n and 00 everywhere else. Each fnf_n differs from the zero function at finitely many points, so each is Riemann integrable, and the value is

01fn(x)dx  =  0for every n. \int_0^1 f_n(x)\,\dd x \;=\; 0 \qquad\text{for every } n. (4.3.4)

Chapter 0.2 then observed that fnχf_n\to\chi pointwise and stopped there. Let's carry it further, because the sequence shows two distinct defects and they need separating.

The first defect is that the class is not closed under limits. The sequence f1f2f3f_1\le f_2\le f_3\le\cdots is increasing, every term is Riemann integrable, and every integral is 00, which is as tame as a sequence can be. Its pointwise limit is χ\chi, which is not Riemann integrable at all. So the natural statement that any theory of integration wants to make,

limn01fn  =  01limnfn, \lim_{n\to\infty}\int_0^1 f_n \;=\; \int_0^1 \lim_{n\to\infty} f_n, (4.3.5)

fails here for the worst possible reason. It is not that the two sides differ. It is that the right-hand side does not refer to anything. Any convergence theorem is a statement of the form (4.3.5), so within the Riemann theory there can be no convergence theorem worth the name.

The second defect is that the norm is not a norm. Chapter 0.5 §1.2's table proposed f,g=fg\avg{f,g}=\int\overline{f}g as an inner product on functions, and said in the next paragraph that the row "is the entire reason Chapter 4.3 is possible." Test the axioms on the sequence above. For n>mn\gt m the difference fnfmf_n-f_m is 11 at finitely many points and 00 elsewhere, so

fnfm22  =  01fnfm2dx  =  0, \lVert f_n-f_m\rVert_2^{\,2} \;=\; \int_0^1 \abs{f_n-f_m}^{2}\,\dd x \;=\; 0, (4.3.6)

and yet fnfmf_n\ne f_m. That is a direct violation of Chapter 0.5's axiom (iii), the demand that v,v>0\avg{v,v}\gt0 for every non-zero vv. Without axiom (iii) there is no distance, because two different objects sit at distance zero, and Chapter 0.5's whole apparatus rests on all three axioms holding. Section 5 repairs this, and the repair is not cosmetic: it changes what a vector is.

Notice also that (4.3.6) makes (fn)(f_n) Cauchy in the strongest imaginable sense, every pair being at distance exactly zero. So the sequence is Cauchy and its pointwise limit is not integrable. That sounds like the third failure, and it is not, which is worth saying out loud rather than gliding over. Distance zero from f1f_1 also means distance zero from the zero function, so in this norm the sequence does have a limit sitting inside the class. Chapter 0.2's sequence is too thin to show that anything is missing. We fatten it.

1.4 · The same rationals, thickened, and the hole that is left

Keep Chapter 0.2's enumeration q1,q2,q_1,q_2,\ldots and give each rational a little room. Around qkq_k put the open interval IkI_k of length 2k22^{-k-2}, and let UnU_n be the part of I1InI_1\cup\cdots\cup I_n lying in (0,1)(0,1). That is a finite union of intervals, so its indicator function

gn(x)  =  {1,xUn,0,otherwise,01gn(x)dx  =  Un, g_n(x) \;=\; \begin{cases} 1, & x\in U_n,\\ 0, & \text{otherwise,}\end{cases} \qquad \int_0^1 g_n(x)\,\dd x \;=\; \abs{U_n}, (4.3.7)

is a step function, perfectly Riemann integrable, whose integral is the total length Un\abs{U_n} of the intervals making up UnU_n. Those totals increase with nn and never exceed the sum of all the lengths, which is k12k2=14\sum_{k\ge1}2^{-k-2}=\tfrac14. An increasing sequence of real numbers bounded above converges, so Un\abs{U_n} tends to a limit λ14\lambda\le\tfrac14.

Now measure the distance between two terms. For n>mn\gt m the difference gngmg_n-g_m is the indicator of UnU_n with UmU_m removed, so its square is itself, and

gngm22  =  UnUm  m,n  0, \lVert g_n-g_m\rVert_2^{\,2} \;=\; \abs{U_n}-\abs{U_m} \;\xrightarrow[m,n\to\infty]{}\; 0, (4.3.8)

because Un\abs{U_n} converges and a convergent sequence of numbers is Cauchy. So (gn)(g_n) is a Cauchy sequence of Riemann-integrable functions. The question is what it converges to, and the answer is the point of the section.

⚠ The claim, and why it is worth the grind box

There is no Riemann-integrable hh on [0,1][0,1] with 01gnh2dx0\int_0^1\abs{g_n-h}^{2}\dd x\to0. Not one that is hard to find. Not one that exists but is badly behaved. None.

The reason is a tug-of-war between two facts about the set UU, meaning the union of all the UnU_n, which is the part of I1I2I_1\cup I_2\cup\cdots lying in (0,1)(0,1). It contains every rational of the interval, so it is dense: every subinterval of (0,1)(0,1) meets it. And its total length is at most 14\tfrac14, so it is small. A limit of the gng_n would have to be close to 11 everywhere in a dense set, which lets a tagged sum at any mesh be pushed up to 11, and it would have to have integral at most 14\tfrac14. Those two demands cannot both be met by a function the Riemann integral can handle, and the grind box turns that sentence into a proof.

Grind box — no Riemann-integrable function is the limit, in full

Suppose hh is Riemann integrable on [0,1][0,1] and 01gnh2dx0\int_0^1\abs{g_n-h}^{2}\dd x\to0, the integrals existing. We derive a contradiction in four steps.

1. The integral of hh is at most 14\tfrac14. Apply Cauchy–Schwarz to gnh\abs{g_n-h} and the constant function 11 on [0,1][0,1]. Chapter 0.5 §1.4 proved that inequality from all three axioms and §1.3 has just withdrawn the third for this pairing, so the citation needs a word. Its proof survives on this pair. The opening inequality asks only that u,u\avg{u,u} never be negative, which is still true here, and every later appeal to (iii) is a division by v,v\avg{v,v}, which for v=1v=\mathbf 1 is 1,1=1>0\avg{\mathbf 1,\mathbf 1}=1\gt0. So

01(hgn)    01hgn    hgn212  =  hgn2    0. \left|\int_0^1 (h-g_n)\right| \;\le\; \int_0^1\abs{h-g_n} \;\le\; \lVert h-g_n\rVert_2\cdot\lVert 1\rVert_2 \;=\; \lVert h-g_n\rVert_2 \;\to\; 0 .

So 01h=limn01gn=λ14\int_0^1 h=\lim_n\int_0^1 g_n=\lambda\le\tfrac14 by (4.3.7).

2. A non-negative Riemann-integrable function with zero integral has infimum zero on every subinterval. If ϕ0\phi\ge0 had infJϕ=c>0\inf_J\phi=c\gt0 on a subinterval JJ of length \ell, then any partition with the endpoints of JJ among its points has lower sum at least c>0c\ell\gt0, so ϕc>0\int\phi\ge c\ell\gt0.

3. On every subinterval, hh comes arbitrarily close to 11. Let I(0,1)I\subseteq(0,1) be any subinterval. Since UU contains every rational of (0,1)(0,1) it meets II, and UU is open, so IUI\cap U contains a closed subinterval JJ. Now U1U2U_1\subseteq U_2\subseteq\cdots are open sets whose union is UJU\supseteq J, so by compactness JUnJ\subseteq U_n for some nn, and then gm=1g_m=1 on JJ for every mnm\ge n. Hence

J(1h)2dx  =  Jgmh2dx    01gmh2dx    0, \int_J (1-h)^{2}\,\dd x \;=\; \int_J\abs{g_m-h}^{2}\dd x \;\le\; \int_0^1\abs{g_m-h}^{2}\dd x \;\to\;0,

so J(1h)2=0\int_J(1-h)^{2}=0. By step 2 the infimum of (1h)2(1-h)^{2} over JJ is zero, so hh takes values as close to 11 as you please inside JJ, and therefore supIhsupJh1\sup_I h\ge\sup_J h\ge1.

4. Contradiction. Take any partition of [0,1][0,1]. Step 3 was proved for subintervals of the open interval (0,1)(0,1), and the first and last cells are not subsets of that, so the transfer needs one word. Every cell contains a subinterval of (0,1)(0,1), and a supremum over a cell is at least the supremum over a piece of it, so step 3 gives supCih1\sup_{C_i}h\ge1 on every cell. Now fix ε>0\varepsilon\gt0 and tag each cell at a point where h(xi)>1εh(x_i^{*})\gt1-\varepsilon, which the supremum makes possible. The tagged sum is then ih(xi)Δxi>(1ε)iΔxi=1ε\sum_i h(x_i^{*})\Delta x_i\gt(1-\varepsilon)\sum_i\Delta x_i=1-\varepsilon, and such a tagging exists at every mesh. Riemann integrability is the statement that every tagging converges to 01h\int_0^1h, so 01h1ε\int_0^1 h\ge1-\varepsilon for every ε\varepsilon, hence 01h1\int_0^1 h\ge1. Step 1 gave 01h14\int_0^1 h\le\tfrac14. \blacksquare

Nothing there needed the fact that the integral is the infimum of the upper sums. Chapter 0.2 defined the integral by tagged sums and never proved that characterisation, and choosing the tags by hand is both shorter and closer to the definition the book actually has.

The two steps that did the work are worth naming. Density of UU forced the tagged sums up. Smallness of UU held the integral down. A theory of integration that could see the difference between "meets every interval" and "has length 14\tfrac14" would have no trouble here, and the Riemann integral cannot see it, because a cell width and a sample point are the only instruments it has.

So we have a Cauchy sequence of perfectly ordinary step functions with no limit in the class. This is the third failure and it is the fatal one. Chapter 4.2 §7 built time evolution as a limit, and Chapter 0.9 §1.3 asserted that partial Fourier sums converge to the function they came from. Both are statements that a Cauchy sequence has a limit, and in this class they are false.

What we need is now precise rather than vague, and it is a list of four items.

  • A notion of the size of a set that can tell a dense set of total length 14\tfrac14 from the whole interval. Section 2.
  • An integral built on that notion, defined on a class closed under limits, and agreeing with Chapter 0.2's wherever Chapter 0.2's exists. Section 3.
  • Convergence theorems that let a limit pass through the integral sign, so that (4.3.5) becomes a theorem with hypotheses instead of a wish. Section 4.
  • A repair of axiom (iii), so that distance zero means equality and Chapter 0.5's geometry transfers intact. Section 5.

Section 6 then shows that the four together buy completeness, and hands back the sequence (gn)(g_n) with a limit attached.

In plain terms 4.3.1

The integral you were taught works by chopping the horizontal axis into thin strips and adding up the areas of rectangles. It is a good definition and it handles every function a physics problem is likely to hand you. It has one flaw, and the flaw is not about any particular function. It is about what happens when you take a sequence of functions and ask what they are approaching.

A sequence of numbers can be seen to be settling down without anyone knowing what it is settling down to. Every term past a certain point is within a hairsbreadth of every other term: that is a self-contained test, and it needs no knowledge of the answer. Whether a sequence that passes the test actually arrives somewhere is not a fact about the sequence. It is a fact about the world the sequence lives in. Inside the fractions alone, the decimal expansion of the square root of two passes the test and arrives nowhere, because the place it was heading is not a fraction. Adding the missing destinations is what the real numbers are for.

Exactly the same thing goes wrong one level up, with functions instead of numbers. There is a sequence of the most ordinary functions imaginable, each one flat except on a few short intervals, which passes the settling-down test and has nothing to settle on. The destination it wants is a function so ragged that the strip-and-rectangle definition cannot assign it an area at all. That is not an oddity to be stepped around. Quantum mechanics is going to say that a state is a vector in a space of functions, and that the limit of a sequence of states is a state. In this space that sentence is false.

So the definition of area has to be replaced, and the replacement starts one step further back than you might expect. Before asking what the area under a curve is, ask what the size of a set of points on the line is. Get that right and everything else follows.

2 · Measure, with one thing quoted

This section replaces one word. Chapter 0.2's integral was built on the length of an interval, and length is the thing that could not see the difference between a dense set and a fat one. We write down what a general notion of size would have to satisfy, and we find that three demands fix almost everything. Then one construction is quoted rather than built, because building it is a chapter of its own and this book is not going to pretend otherwise. What the quotation buys is stated precisely, and §8 comes back to collect part of it.

2.1 · The three demands, and why one of them is countable

We want to attach a number μ(A)\mu(A) to a set AA of real numbers, and the number should behave the way length behaves. Three demands do it.

  • It extends length. For an interval, μ\mu is the length: μ((a,b))=ba\mu\big((a,b)\big)=b-a, and the same for the closed and half-open versions.
  • It is countably additive. If A1,A2,A_1,A_2,\ldots are disjoint then the size of the union is the sum of the sizes.
  • It is translation invariant. Sliding a set along the line does not change its size.

The first and third are obvious requests. The second is the one carrying the weight, and the word countably is doing all of the carrying, so let's see why finite additivity would not be enough. Every failure in §1 was a countable process. The set of rationals is a countable union of single points, the set UU was a countable union of intervals, and the sequences (fn)(f_n) and (gn)(g_n) were indexed by the integers. A notion of size that only adds finitely many pieces cannot say anything about any of them.

Countable additivity has one consequence we will use so often that it is worth extracting now. Suppose A1A2A_1\subseteq A_2\subseteq\cdots is an increasing family with union AA. Write the union as a disjoint one by peeling off shells, A=A1(A2A1)(A3A2)A=A_1\cup(A_2\setminus A_1)\cup(A_3\setminus A_2)\cup\cdots, and apply countable additivity to the shells. The partial sums of the resulting series are exactly the μ(An)\mu(A_n), so

A1A2,A=nAnμ(A)  =  limnμ(An). A_1\subseteq A_2\subseteq\cdots, \quad A=\bigcup_n A_n \qquad\Longrightarrow\qquad \mu(A) \;=\; \lim_{n\to\infty}\mu(A_n). (4.3.9)

This is called continuity from below, and it is the exact point at which the word "countably" gets spent. Section 4's monotone convergence theorem is (4.3.9) promoted from sets to functions, and it has no other engine.

2.2 · Which sets, and the shape the collection has to have

The demands above say nothing about which sets get a size, and §2.5 will show that not all of them can. So the collection of sets we can measure is part of the structure, and its shape is forced by the operations we intend to perform on it.

A collection M\mathcal M of subsets of R\R is a σ\sigma-algebra when it contains R\R, is closed under taking complements, and is closed under countable unions. Countable intersections come free, because nAn\bigcap_n A_n is the complement of nAnc\bigcup_n A_n^{c}, and set differences come free for the same reason. Those are precisely the operations a countable limit process performs, which is why the definition looks the way it does rather than some other way.

The smallest σ\sigma-algebra containing every open interval is called the collection of Borel sets, and it contains everything you are likely to write down: open sets, closed sets, countable unions of closed sets, countable intersections of those, and so on upward. The set UU of §1.4 is a countable union of open intervals, so it is Borel, and so is the set of rationals.

2.3 · The construction, quoted as a package

Everything so far is a specification. Meeting it is a genuine construction, and here is the first of the two places where this chapter takes something on trust. The construction starts by defining an outer measure, which covers a set as economically as possible with intervals and takes the cheapest cover:

μ(A)  =  inf{k=1Ik  :  Ak=1Ik,  Ik open intervals}. \mu^{*}(A) \;=\; \inf\left\{\sum_{k=1}^{\infty}\abs{I_k} \;:\; A\subseteq\bigcup_{k=1}^{\infty}I_k,\ \ I_k \text{ open intervals}\right\}. (4.3.10)

That definition applies to every set whatever, which is why it cannot be the end of the story: something so generous is not going to be countably additive. Carathéodory's criterion then selects the sets on which it behaves, by keeping AA exactly when AA splits every other set additively.

⚑ Quoted, not derived — the construction of Lebesgue measure

We use, without proof, the following. There is a σ\sigma-algebra M\mathcal M of subsets of R\R, containing every Borel set, and a function μ:M[0,]\mu:\mathcal M\to[0,\infty] agreeing with (4.3.10), with all of these properties.

  • (a) Length. μ(I)\mu(I) is the length of II for every interval.
  • (b) Countable additivity. μ(kAk)=kμ(Ak)\mu\big(\bigcup_k A_k\big)=\sum_k\mu(A_k) for disjoint AkMA_k\in\mathcal M.
  • (c) Translation invariance. μ(A+t)=μ(A)\mu(A+t)=\mu(A).
  • (d) Regularity. For AMA\in\mathcal M and any ε>0\varepsilon\gt0 there is an open OAO\supseteq A with μ(OA)<ε\mu(O\setminus A)\lt\varepsilon, and a closed CAC\subseteq A with μ(AC)<ε\mu(A\setminus C)\lt\varepsilon.
  • (e) Completeness of the measure. If μ(N)=0\mu(N)=0 and BNB\subseteq N then BMB\in\mathcal M and μ(B)=0\mu(B)=0.

This is Lebesgue's theorem. No chapter of this book proves it, and none will, which is why the mark is here rather than a forward pointer. The construction runs from (4.3.10) to (a) through (e) in about twenty pages and opens every course in measure theory, Royden's Real Analysis and Rudin's Real and Complex Analysis being the two standard places to read it. That is the first time this book sends you to another text for an argument. Chapter 3.6 §5.4 named a text before, for a signature convention, which is a different kind of pointer.

Say what is being bought. Properties (a), (b) and (c) are the three demands of §2.1 and nothing more. Property (e) is a convenience that costs nothing and saves a clause in every later proof. Property (d) is the one that does real work later, and it is worth watching. It says a measurable set is squeezed between an open set and a closed set that are as close to it in size as you like, which is what lets an arbitrary measurable set be traded for a finite union of intervals. Section 7.4 cashes it, making that trade once in order to show the space separable, and §8.1 then re-uses the same trade to prove that continuous functions are dense in the space of states. Neither result is separately marked, because both are this package spending itself.

2.4 · Sets of measure zero, and the phrase "almost everywhere"

Call NN a null set when μ(N)=0\mu(N)=0. The most useful fact about null sets is how easy they are to come by, and the proof is a construction you have already seen in this chapter.

Every countable set is null. Let the set be {p1,p2,}\{p_1,p_2,\ldots\}, fix ε>0\varepsilon\gt0, and cover pkp_k by an interval of length ε2k\varepsilon2^{-k}. That is a countable cover of total length ε\varepsilon, so με\mu^{*}\le\varepsilon by (4.3.10), and ε\varepsilon was arbitrary. In particular

μ(Q[0,1])  =  0. \mu\big(\mathbb{Q}\cap[0,1]\big) \;=\; 0. (4.3.11)

Look at what that construction was. It is §1.4's thickening of the rationals, with the thickness sent to zero instead of held at 14\tfrac14. The same cover that produced a dense set of length 14\tfrac14 produces, in the limit, the statement that the rationals have no length at all. Chapter 0.2 could not write down (4.3.11), and it is the whole reason χ\chi is about to become integrable.

A property holds almost everywhere, abbreviated a.e., when the set where it fails is null. So χ=0\chi=0 almost everywhere, and the functions fnf_n of §1.3 are all equal to each other almost everywhere. That phrase is going to reorganise the whole subject in §5, where two functions agreeing almost everywhere stop being two functions.

Familiar ground — measure is probability, and the part with no clinical counterpart

Restrict μ\mu to [0,1][0,1] and read μ(A)\mu(A) as the chance that a uniformly distributed random number lands in AA. Nothing has to be adjusted for the reading to work, because Lebesgue measure on [0,1][0,1] is the uniform distribution, and the correspondence is term by term.

  • μ(whole interval)=1\mu(\text{whole interval})=1 is the statement that some outcome occurs.
  • Countable additivity is the axiom that the chance of any one of countably many mutually exclusive events is the sum of their chances. It is not an approximation to that axiom. It is that axiom.
  • A null set is an event of probability zero, and "almost everywhere" is what a statistician writes as "almost surely".

Equation (4.3.11) then reads: a uniform random draw from [0,1][0,1] is rational with probability exactly zero, although rationals are everywhere in the interval. You have used that fact every time you treated a continuous variable as continuous. Probability zero does not mean impossible, and it never did. What §2.5 adds is the part that has no clinical counterpart: there are subsets of [0,1][0,1] to which no probability can be assigned at all, on pain of contradiction, and the axioms of probability are stated over a σ\sigma-algebra rather than over all subsets for exactly that reason.

2.5 · Not every set is measurable

The flag in §2.3 restricted μ\mu to a collection M\mathcal M rather than to all subsets of R\R, and a reader is entitled to ask whether that restriction is a real one or a technician's caution. It is real, and the demonstration is short enough to run here. It is due to Vitali.

The point of running it is to make the flag honest. If M\mathcal M could be all subsets, the quoted package would be quoting a formality. It cannot, and the three demands of §2.1 are exactly what makes it impossible.

Grind box — the Vitali set, built

Work on [0,1)[0,1) and declare xyx\sim y when xyx-y is rational. This is an equivalence relation, so it cuts [0,1)[0,1) into disjoint classes, each class being one number plus every rational, cut down to [0,1)[0,1). One class is the rationals of [0,1)[0,1) on their own. Every other class is an irrational offset plus all the rationals. Choose one representative from each class and collect the choices into a set VV. That step is the axiom of choice, and it is where the whole construction lives.

Now translate VV around the circle. For qQ[0,1)q\in\mathbb{Q}\cap[0,1) let

Vq  =  {v+q:vV, v+q<1}    {v+q1:vV, v+q1}, V\oplus q \;=\; \{v+q : v\in V,\ v+q\lt1\} \;\cup\; \{v+q-1 : v\in V,\ v+q\ge1\},

which is VV shifted by qq with the overhang wrapped back to the start.

They are disjoint. If v1+q1v_1+q_1 and v2+q2v_2+q_2 agree modulo 11 then v1v2v_1-v_2 is rational, so v1v_1 and v2v_2 lie in the same class, so v1=v2v_1=v_2 by the choice, so q1=q2q_1=q_2.

They cover. Given x[0,1)x\in[0,1), let vv be the representative of xx's class. Then xvx-v is rational and lies in (1,1)(-1,1), so it is either qq or q1q-1 for some qQ[0,1)q\in\mathbb{Q}\cap[0,1), and xVqx\in V\oplus q.

The contradiction. Suppose VMV\in\mathcal M with μ(V)=m\mu(V)=m. Each VqV\oplus q is two translated pieces of VV, so by (c) and (b) of §2.3 it is measurable with μ(Vq)=m\mu(V\oplus q)=m. The rationals in [0,1)[0,1) are countable, so countable additivity applies to the whole family:

1  =  μ([0,1))  =  qQ[0,1)μ(Vq)  =  qm. 1 \;=\; \mu\big([0,1)\big) \;=\; \sum_{q\in\mathbb{Q}\cap[0,1)}\mu(V\oplus q) \;=\; \sum_{q} m .

If m=0m=0 the sum is 00. If m>0m\gt0 the sum is ++\infty. Neither is 11, so VV is not in M\mathcal M. \blacksquare

Every ingredient was one of the three demands, plus the axiom of choice. Drop countable additivity for finite additivity and the argument collapses, which is one reason the countable version is not a stylistic preference.

So M\mathcal M is a genuine restriction, and the sets it leaves out are the ones you cannot write down. Every set arising from a formula, a limit, or a physical construction is Borel, hence in M\mathcal M, and §3 never has to think about it again.

In plain terms 4.3.2

Before you can improve on the idea of area you have to improve on the idea of length. The question is how big a set of points on a line is, when the set is not an interval and may be scattered all through one.

Three requirements settle it almost entirely. The size of an interval is its length. Sliding a set along the line leaves its size alone. And if a set is cut into a list of non-overlapping pieces, even an endless list, the sizes of the pieces add up to the size of the whole. That third requirement is the only one with any bite, and the word "endless" is where the bite is. Every difficulty in the previous section was an endless process, so a rule that only adds finitely many pieces is no use.

Meeting the three requirements takes a construction, and this chapter quotes it rather than performing it, with the quotation marked in place. What the quotation includes is worth knowing, because one clause of it is spent later on something that looks unrelated: any set that can be measured can be trapped between an open set and a closed one that are as close to it in size as desired. That clause is what will eventually let a wildly discontinuous function be traded for a smooth one.

Two consequences follow at once. A set of separate points, even endlessly many of them, has size zero, because you can cover the first with a cover of length one-half of your allowance, the second with one quarter, and so on forever, and the total is your allowance however small you made it. So the rationals, dense as they are, have no size at all. And the price of the whole scheme, which is the honest part: there are sets to which no size can be given, and the proof that they exist uses nothing but the three requirements. The restriction is real rather than a technician's fussiness.

a natural place to stop  ·  the problem is set, and so is what "size" must mean

3 · The Lebesgue integral, built

With size in hand, the integral takes one paragraph to define and one section to check. The whole idea is a single change of direction, which §3.1 states in a sentence. Section 3.2 says which functions are eligible and proves the property that the Riemann class lacked. Section 3.3 does the construction. Section 3.5 shows the new integral agrees with the old one wherever the old one works, so that nothing in Parts 0 to III has to be relearned, and §3.6 is honest about the one thing that is lost.

3.1 · One change of direction

Here is the whole idea in a sentence. Riemann slices the domain. Lebesgue slices the range.

Chapter 0.2 chopped the xx-axis into cells, took the value of ff somewhere in each cell, and added up value times width. The new definition chops the yy-axis instead. For each level yy it asks how big the set of points where ff is near yy is, and adds up level times size.

The picture that makes it stick is counting money. Given a pile of coins, Riemann's method is to go along the pile in order and add each coin's value as you reach it. Lebesgue's method is to sort the pile into denominations first, count how many coins are in each pile, and add up denomination times count. Three terms map exactly: the value of a coin is the value of ff, the position of a coin in the pile is the point xx, and the number of coins of a given denomination is the size of the set where ff takes that value.

One term has no counterpart, and it is the one §1 turned on. A coin has a single definite value, so there is no cell, no width, and no freedom about where to sample. Both methods therefore return the same total on every pile, which is §3.5 in miniature, and neither of them can fail. So the picture cannot show you why one method survives where the other does not, and that claim is left where it is proved, in §4.

3.2 · Which functions, and the property the Riemann class lacked

Slicing the range means asking for the size of the set where ff exceeds a level, so that set had better be measurable. Call ff measurable when

{x:f(x)>a}    Mfor every real a. \{x : f(x)\gt a\} \;\in\; \mathcal M \qquad\text{for every real } a. (4.3.12)

That is the entire condition, and it is weak. Every continuous function passes, since {f>a}\{f\gt a\} is then open. Every step function passes. So does χ\chi, for which {χ>a}\{\chi\gt a\} is R\R when a<0a\lt0, the rationals when 0a<10\le a\lt1, and empty when a1a\ge1. Three sets, all of them in M\mathcal M, and nothing you can write down fails.

Now the property that Riemann's class did not have, and the reason the σ\sigma-algebra was defined with countable operations in §2.2. Let f1,f2,f_1,f_2,\ldots be measurable and suppose fn(x)f(x)f_n(x)\to f(x) at every xx. Then f(x)>af(x)\gt a exactly when the values fn(x)f_n(x) eventually stay above some level strictly greater than aa, and "some level", "eventually" and "stay above" are three countable operations, written below in that order:

{f>a}  =  k=1 N=1 nN {fn>a+1k}. \{f\gt a\} \;=\; \bigcup_{k=1}^{\infty}\ \bigcup_{N=1}^{\infty}\ \bigcap_{n\ge N}\ \Big\{f_n\gt a+\tfrac1k\Big\}. (4.3.13)

Every set on the right is in M\mathcal M by (4.3.12), and M\mathcal M is closed under countable unions and intersections, so the set on the left is in M\mathcal M too. A pointwise limit of measurable functions is measurable.

Compare that with §1.3, where a monotone limit of Riemann-integrable functions was not Riemann integrable. The defect is gone, and it is gone by construction rather than by luck: the class was defined by a condition phrased in countable operations precisely so that countable operations could not escape it.

3.3 · Simple functions, then the supremum

The construction goes in two stages, and the first stage is the one where the definition actually happens. A simple function is a finite sum of indicators of measurable sets,

φ  =  k=1Kck1Ek,ck0,EkM disjoint, \varphi \;=\; \sum_{k=1}^{K} c_k\,\mathbf{1}_{E_k}, \qquad c_k\ge0, \quad E_k\in\mathcal M \ \text{disjoint}, (4.3.14)

and its integral is defined to be the obvious thing, level times size, added up. Since the EkE_k are disjoint there is no ambiguity about what to write:

φdμ    k=1Kckμ(Ek). \int \varphi \,\dd\mu \;\equiv\; \sum_{k=1}^{K} c_k\,\mu(E_k). (4.3.15)

That is the range-slicing of §3.1 with only finitely many levels. Two things follow from finite additivity of μ\mu and nothing else, and both are used constantly below: the integral of a sum of simple functions is the sum of the integrals, and φψ\varphi\le\psi implies φψ\int\varphi\le\int\psi.

The second stage extends this to any non-negative measurable ff by approaching it from below with simple functions and taking the best available answer:

fdμ    sup{φdμ  :  φ simple, 0φf}, \int f \,\dd\mu \;\equiv\; \sup\left\{\int\varphi\,\dd\mu \;:\; \varphi \ \text{simple},\ 0\le\varphi\le f\right\}, (4.3.16)

a supremum that may be ++\infty. Monotonicity is immediate from (4.3.16), since fgf\le g means every φ\varphi competing for ff also competes for gg, so the supremum can only grow. That one-line fact is used in almost every proof that follows.

It is worth seeing that the supremum is always attained in the limit by an explicit staircase, because §4 will apply its theorems to exactly this sequence. Cut the range into steps of height 2n2^{-n} up to a ceiling of nn and read off which step ff is on:

φn  =  j=0n2n1j2n  1{j2nf<(j+1)2n}  +  n1{fn}. \varphi_n \;=\; \sum_{j=0}^{n2^{n}-1}\frac{j}{2^{n}}\;\mathbf{1}_{\{\,j2^{-n}\,\le\, f\,\lt\,(j+1)2^{-n}\}} \;+\; n\,\mathbf{1}_{\{f\ge n\}}. (4.3.17)

Each φn\varphi_n is simple by (4.3.12), the sequence increases because halving the step can only refine the estimate, and φn(x)f(x)\varphi_n(x)\to f(x) at every xx. So every non-negative measurable function is an increasing limit of simple ones, which is the range being sliced ever more finely.

Finally, drop the requirement that ff be non-negative. Write f=f+ff=f^{+}-f^{-} with f+=max(f,0)f^{+}=\max(f,0) and f=max(f,0)f^{-}=\max(-f,0), both non-negative and measurable, and define f=f+f\int f=\int f^{+}-\int f^{-}. Call ff integrable when both pieces have finite integral, which is the same as saying f<\int\abs f\lt\infty. For a complex-valued ff, integrate the real and imaginary parts separately. That is the whole definition.

3.4 · The first two answers, one of them a debt

Chapter 0.2's warning box promised that the new integral "gives 01χ=0\int_0^1\chi=0 into the bargain", and the promise is now one line. The function χ\chi restricted to [0,1][0,1] is the indicator of Q[0,1]\mathbb{Q}\cap[0,1], which is simple, so (4.3.15) applies directly and (4.3.11) supplies the size:

01χdμ  =  1μ(Q[0,1])  =  0. \int_0^1 \chi \,\dd\mu \;=\; 1\cdot\mu\big(\mathbb{Q}\cap[0,1]\big) \;=\; 0. (4.3.18)

The function that no partition could pin down has an integral, and getting it took one multiplication. The same applies to §1.4's thickened set: 1U\mathbf 1_U is a limit of the simple functions gng_n, and its integral will turn out to be μ(U)\mu(U) once §4 licenses passing to the limit. That is the loose end §6 ties.

3.5 · It agrees with Riemann wherever Riemann works

If the new integral disagreed with the old one anywhere the old one was defined, every calculation in Parts 0 to III would need checking, so the following matters more than its proof suggests.

If ff is Riemann integrable on [a,b][a,b], then ff is measurable, it is Lebesgue integrable, and the two integrals are equal.

The idea is that upper and lower sums are integrals of simple functions, so the Riemann machinery is already inside the Lebesgue machinery, viewed correctly. Take partitions PnP_n with mesh going to zero, each refining the last, and build the two staircases

n  =  i(infCif)1Ci,un  =  i(supCif)1Ci, \ell_n \;=\; \sum_i \Big(\inf_{C_i} f\Big)\mathbf 1_{C_i}, \qquad u_n \;=\; \sum_i \Big(\sup_{C_i} f\Big)\mathbf 1_{C_i}, (4.3.19)

the sums running over the cells CiC_i of PnP_n. These are simple functions whose Lebesgue integrals, by (4.3.15), are exactly the lower and upper Riemann sums of ff over PnP_n. Riemann integrability says those two sequences of numbers close on each other, and that squeeze is what the grind box turns into the theorem.

Grind box — Riemann implies Lebesgue, with the same value

Shift ff by a constant so that f0f\ge0, which changes both integrals by the same amount. Write LnL_n and UnU_n for the lower and upper Riemann sums over PnP_n, and RR for the Riemann integral.

One line first, because Chapter 0.2 defined that integral by tagged sums and not by upper and lower ones. On a fixed partition the tagged sums come as close as you like to LnL_n and to UnU_n, by tagging near the infimum or near the supremum in each cell, and Chapter 0.2 §1.1 observed that every tagged sum is trapped between the two. Integrability sends every tagged sum to RR as the mesh shrinks, so LnL_n and UnU_n are squeezed onto RR with them. Refinement makes the first increase and the second decrease, so LnRL_n\uparrow R and UnRU_n\downarrow R.

The same refinement condition makes n\ell_n increasing and unu_n decreasing pointwise, so both have pointwise limits

=limnn,u=limnun,    f    u, \ell=\lim_n \ell_n, \qquad u=\lim_n u_n, \qquad \ell\;\le\; f\;\le\; u,

and both are measurable by §3.2. Now three steps.

1. =R\int\ell=R. Monotonicity of (4.3.16) against num\ell_n\le\ell\le u_m for every nn and mm gives LnUmL_n\le\int\ell\le U_m. Let nn and mm run and the two ends squeeze onto RR.

2. u=u=\ell almost everywhere. The function uu-\ell is non-negative and satisfies uunnu-\ell\le u_n-\ell_n, so monotonicity and additivity on simple functions give (u)UnLn0\int(u-\ell)\le U_n-L_n\to0, hence (u)=0\int(u-\ell)=0. For any non-negative measurable ϕ\phi the inequality 1k1{ϕ>1/k}ϕ\tfrac1k\mathbf 1_{\{\phi\gt1/k\}}\le\phi and monotonicity give Chebyshev's bound μ{ϕ>1/k}kϕ\mu\{\phi\gt1/k\}\le k\int\phi, so here every μ{u>1/k}\mu\{u-\ell\gt1/k\} vanishes, and the union over kk is the set where u>u\gt\ell. It is null.

3. Conclusion. Since fu\ell\le f\le u and =u\ell=u off a null set, f=f=\ell almost everywhere, so ff is measurable by property (e) of §2.3, and its integral equals =R\int\ell=R, two functions agreeing off a null set having the same integral by (4.3.15) applied to the competing simple functions. \blacksquare

So no arithmetic in this book changes. Every Gaussian in Chapter 0.2, every Fourier coefficient in Chapter 0.9, and every action integral in Part I means the same number it always did. What has been added is a larger class of functions and, in §4, the theorems that make the enlargement worth having.

3.6 · The one thing that is lost, said plainly

The new integral is not merely more powerful than the old one, and pretending otherwise would store up a surprise. There is a class of integrals the Riemann theory handles and the Lebesgue theory declines.

⚠ A conditionally convergent integral is not a Lebesgue integral

Chapter 0.9's grind box on the Gibbs constant used the sine integral, and the same function gives the standard example. The improper Riemann integral

0sinxxdx  =  limT0Tsinxxdx \int_0^{\infty}\frac{\sin x}{x}\,\dd x \;=\; \lim_{T\to\infty}\int_0^{T}\frac{\sin x}{x}\,\dd x

exists, because the alternating arches cancel in a controlled way as TT grows. But 0sinx/xdx=\int_0^{\infty}\abs{\sin x/x}\,\dd x=\infty, since the nn-th arch contributes about 2/πn2/\pi n and the harmonic series diverges. By §3.3 a function is Lebesgue integrable only when f<\int\abs f\lt\infty, so sinx/x\sin x/x is not Lebesgue integrable on [0,)[0,\infty).

The reason for the asymmetry is exactly the reason the new integral is worth having. The Lebesgue integral sorts the range before it sums, so it can never rely on the order in which cancellations arrive, and an improper Riemann integral is a statement about that order. When you meet 0sinx/x\int_0^\infty\sin x/x in Part V it is a limit of proper integrals, and it has to be written as one.

In plain terms 4.3.3

The new definition of area is one idea. Instead of chopping the horizontal axis into thin strips and asking how tall the function is over each strip, chop the vertical axis into thin bands and ask how wide the set of places is where the function lands in each band. Then add up height times width as before. Counting a pile of coins by walking along it is the first method; sorting the pile into denominations and multiplying is the second. For a finite pile the two totals agree, and they agree here too, wherever the old definition worked at all.

What changes is which functions are eligible. The old definition needed the function to be reasonably well behaved along the horizontal axis. The new one needs only that each of the bands corresponds to a set whose size is defined, and the collection of such sets was built in the previous section to be closed under endless unions and intersections. That is not a coincidence. It is what makes the class of eligible functions survive limits, which is the exact property the old class lacked and the whole reason for the rebuild.

The function that took the value one at every fraction and zero elsewhere now has an area, and the calculation is a single multiplication: the height is one, the set of fractions has size zero, the product is zero. A page of frustration in the earlier chapter becomes a line.

One thing is given up, and it should be said rather than discovered later. Certain integrals converge only because positive and negative contributions arrive alternately and cancel. Sorting the values before adding them destroys that arrangement, so such integrals are not admitted by the new definition and have to be written as limits of ordinary ones. That is the price of never depending on the order of the terms, and the next section is what the price buys.

4 · Monotone and dominated convergence

This is the section the rebuild was for. Section 1 showed that the Riemann integral admits no theorem of the form "the integral of the limit is the limit of the integrals". Here are the two that the Lebesgue integral does admit. The first is proved from continuity from below and nothing else. The second follows from the first through one intermediate step, and it is the one Chapter 0.2 borrowed against. Both are stated with their hypotheses in front, because in both cases the hypothesis is what the theorem is really about.

4.1 · Monotone convergence, derived

Let 0f1f20\le f_1\le f_2\le\cdots be measurable and let fn(x)f(x)f_n(x)\to f(x) at every xx. Then fnf\int f_n\to\int f. The limit function is measurable by §3.2, so both sides refer to something, and that is exactly what failed in §1.3.

One direction is free. Monotonicity of (4.3.16) gives fnf\int f_n\le\int f for every nn, and the left side increases, so its limit LL satisfies LfL\le\int f. The work is the other direction, and it consists of showing that LL beats every simple function competing in the supremum that defines f\int f.

So fix a simple φ=kck1Ek\varphi=\sum_k c_k\mathbf 1_{E_k} with 0φf0\le\varphi\le f, and fix a slack factor θ\theta slightly below 11, kept a different letter from the levels ckc_k because the two are about to be multiplied together. Since fnf_n climbs to fφf\ge\varphi, every point is eventually caught above θφ\theta\varphi, so the sets

An  =  {x:fn(x)θφ(x)}satisfyA1A2,nAn  =  R. \begin{aligned} A_n \;&=\; \big\{x : f_n(x)\ge \theta\,\varphi(x)\big\} \quad\text{satisfy}\\[4pt] A_1\subseteq A_2&\subseteq\cdots, \qquad \bigcup_n A_n \;=\; \R. \end{aligned} (4.3.20)

Write Ag\int_A g for g1A\int g\,\mathbf 1_A, the integral restricted to a set. Discarding everything outside AnA_n can only lower an integral of a non-negative function, so fnAnfnθAnφ\int f_n\ge\int_{A_n}f_n\ge\theta\int_{A_n}\varphi, and the last of those is a simple function's integral, which (4.3.15) writes out explicitly:

fn    θkckμ(EkAn)  n  θkckμ(Ek)  =  θφ. \int f_n \;\ge\; \theta\sum_k c_k\,\mu(E_k\cap A_n) \;\xrightarrow[n\to\infty]{}\; \theta\sum_k c_k\,\mu(E_k) \;=\; \theta\int\varphi. (4.3.21)

The limit in the middle is (4.3.9), continuity from below, applied to the increasing sets EkAnE_k\cap A_n inside each EkE_k. That is where the countable in "countably additive" is spent, and it is the only place the proof spends anything. So LθφL\ge\theta\int\varphi for every θ<1\theta\lt1, hence LφL\ge\int\varphi, and taking the supremum over φ\varphi gives LfL\ge\int f. With the free direction, L=fL=\int f. \blacksquare

Two corollaries follow at once, and the first of them repairs an omission. Nothing so far has established that (f+g)=f+g\int(f+g)=\int f+\int g for general non-negative measurable ff and gg, since (4.3.15) only gave it for simple functions. Take the staircases φnf\varphi_n\uparrow f and ψng\psi_n\uparrow g of (4.3.17), note that φn+ψnf+g\varphi_n+\psi_n\uparrow f+g, and apply monotone convergence to all three sequences. Additivity for simple functions passes to the limit, so the Lebesgue integral is linear.

The second corollary is the one that gets used in anger. Let h1,h2,h_1,h_2,\ldots be non-negative and measurable and apply monotone convergence to the partial sums knhk\sum_{k\le n}h_k, which increase because the terms are non-negative:

k=1hkdμ  =  k=1hkdμ,hk0. \int \sum_{k=1}^{\infty} h_k \,\dd\mu \;=\; \sum_{k=1}^{\infty}\int h_k\,\dd\mu, \qquad h_k\ge0. (4.3.22)

No convergence hypothesis is needed anywhere in (4.3.22), since both sides are allowed to be ++\infty and the theorem says they are infinite together. Sign is the only thing being asked for.

One interchange that looks like a neighbour of this one is not settled here, and it is worth saying so rather than leaving a reader to wonder. Exchanging a sum with an integral is (4.3.22). Exchanging two integrals is Fubini's theorem, which Chapter 0.2 quoted and marked when it squared the Gaussian, and which needs a measure on a product space rather than on the line. That is a second construction of the kind §2.3 already quoted once, this chapter does not perform it, and nothing below uses it. Chapter 0.2's mark stands where it is.

A debt, collected. Chapter 4.1 §5 evaluated the integral in the Planck spectrum by expanding the denominator as a geometric series and integrating term by term, and it wrote in parentheses that "the theorem that makes it airtight is monotone convergence, proved in Chapter 4.3." Here it is. For y>0y\gt0 the expansion 1/(ey1)=n1eny1/(\ee^{y}-1)=\sum_{n\ge1}\ee^{-ny} has every term positive, so (4.3.22) applies with no further hypothesis, and

0y3dyey1  =  n=10y3enydy  =  n=13!n4  =  π415. \int_{0}^{\infty}\frac{y^{3}\,\dd y}{\ee^{y}-1} \;=\; \sum_{n=1}^{\infty}\int_{0}^{\infty}y^{3}\ee^{-ny}\,\dd y \;=\; \sum_{n=1}^{\infty}\frac{3!}{n^{4}} \;=\; \frac{\pi^{4}}{15}. (4.3.23)

The Stefan–Boltzmann constant of Chapter 4.1 rests on that interchange, and the interchange now rests on a theorem rather than on an expectation. Worked example 2 derives the n4=π4/90\sum n^{-4}=\pi^{4}/90 that turns the middle expression into the last.

4.2 · Fatou's lemma, and what it says when it is strict

Monotone convergence needs the sequence to climb, and most sequences do not. The next result asks for nothing but non-negativity, and pays for that generosity with an inequality instead of an equality.

For non-negative measurable fnf_n,  lim infnfnlim infnfn\ \int\liminf_n f_n\le\liminf_n\int f_n. The proof is monotone convergence applied to the wrong sequence on purpose. Set hn=infknfkh_n=\inf_{k\ge n}f_k. That sequence has to be measurable before monotone convergence can touch it, and §3.2 closed the class under pointwise limits rather than under infima, so here is the missing line. Writing {fa}=j1{f>a1/j}\{f\ge a\}=\bigcap_{j\ge1}\{f\gt a-1/j\} puts each {fka}\{f_k\ge a\} in M\mathcal M, then {hna}=kn{fka}\{h_n\ge a\}=\bigcap_{k\ge n}\{f_k\ge a\} is in M\mathcal M as a countable intersection, and {hn>a}=j1{hna+1/j}\{h_n\gt a\}=\bigcup_{j\ge1}\{h_n\ge a+1/j\} is the condition (4.3.12) asks for. The hnh_n increase with nn and climb to lim inffn\liminf f_n by definition, so monotone convergence gives hnlim inffn\int h_n\to\int\liminf f_n. Also hnfkh_n\le f_k for every knk\ge n, so monotonicity gives hninfknfk\int h_n\le\inf_{k\ge n}\int f_k. Let nn run and the right side becomes lim inffn\liminf\int f_n. \blacksquare

An inequality is only interesting when you know how it fails to be an equality, and Chapter 0.2 already built the example. Its grind box on differentiating under the integral sign used

h(x,a)  =  2xa2ex2/a2,01h(x,a)dx=1e1/a2  a0  1, h(x,a) \;=\; \frac{2x}{a^{2}}\,\ee^{-x^{2}/a^{2}}, \qquad \int_0^1 h(x,a)\,\dd x = 1-\ee^{-1/a^{2}} \;\xrightarrow[a\to0]{}\; 1, (4.3.24)

while h(x,a)0h(x,a)\to0 as a0a\to0 for every fixed x>0x\gt0. Take a=1/na=1/n and you have a sequence with lim inf=0\int\liminf=0 and lim inf=1\liminf\int=1. Chapter 0.2 said of it that "the mass has not vanished. It has escaped into an ever narrower, ever taller spike near the origin", and Chapter 0.9 §5.2 identified the same object as the Dirac delta being born. Fatou's lemma is the exact statement that mass can escape but cannot appear from nowhere, and the direction of the inequality is the direction mass can go.

4.3 · Dominated convergence, and Chapter 0.2's outstanding hypothesis

Now the theorem the book has been waiting on. It buys back the equality, and the price is a single function that holds the whole sequence down.

Let fnf_n be measurable with fnff_n\to f pointwise almost everywhere, and suppose there is one integrable gg with fng\abs{f_n}\le g for every nn. Then ff is integrable and fnf0\int\abs{f_n-f}\to0, so in particular fnf\int f_n\to\int f.

The proof is Fatou applied to a cleverly chosen non-negative sequence. Since fg\abs f\le g as well, we have fnf2g\abs{f_n-f}\le2g, so 2gfnf2g-\abs{f_n-f} is non-negative and Fatou applies to it. Its pointwise lower limit is 2g2g, because fnf0\abs{f_n-f}\to0, so

2g    lim infn(2gfnf)  =  2g    lim supnfnf, \int 2g \;\le\; \liminf_n\int\big(2g-\abs{f_n-f}\big) \;=\; \int 2g \;-\; \limsup_n\int\abs{f_n-f}, (4.3.25)

using the linearity established in §4.1. Now 2g\int 2g is finite, which is the whole content of the hypothesis, so it may be cancelled from both sides. That leaves lim supfnf0\limsup\int\abs{f_n-f}\le0, and the quantity is non-negative, so it is zero. \blacksquare

The promise, collected verbatim. Chapter 0.2 §4.4's grind box justified swapping a derivative with an integral and wrote: "The sufficient hypothesis (dominated convergence, proved properly in Chapter 4.3) is that the aa-derivative of the integrand be bounded, uniformly in aa near the point of interest, by a single fixed integrable function." That is the theorem above, and the swap it licenses is worth stating as its own result rather than left implied.

Suppose F(x,a)F(x,a) is integrable in xx for each aa near a0a_0, that aF\partial_aF exists, and that aF(x,a)G(x)\abs{\partial_aF(x,a)}\le G(x) for all such aa, with GG integrable. Then

ddaF(x,a)dx  =  Fa(x,a)dxat a=a0. \dv{}{a}\int F(x,a)\,\dd x \;=\; \int \pdv{F}{a}(x,a)\,\dd x \qquad\text{at } a=a_0. (4.3.26)

To see it, take any sequence ana0a_n\to a_0 and form the difference quotients Qn(x)=[F(x,an)F(x,a0)]/(ana0)Q_n(x)=[F(x,a_n)-F(x,a_0)]/(a_n-a_0). The mean value theorem writes each Qn(x)Q_n(x) as aF(x,ξ)\partial_aF(x,\xi) for some ξ\xi between ana_n and a0a_0, so QnG\abs{Q_n}\le G for every nn, which is the domination. And QnaF(x,a0)Q_n\to\partial_aF(x,a_0) pointwise, which is the convergence. Dominated convergence therefore gives QnaF\int Q_n\to\int\partial_aF, and the left side is by construction the difference quotient of aFa\mapsto\int F. Every sequence gives the same limit, so the derivative exists and (4.3.26) holds.

Chapter 0.2's own case now closes. There F(x,a)=eax2F(x,a)=\ee^{-ax^{2}} and you were asked to accept the bound

aeax2  =  x2eax2    x2ea0x2for aa0>0, \abs{\pdv{}{a}\ee^{-ax^{2}}} \;=\; x^{2}\ee^{-ax^{2}} \;\le\; x^{2}\ee^{-a_0x^{2}} \qquad\text{for } a\ge a_0\gt0, (4.3.27)

with the right-hand side a fixed function of finite integral. That is exactly the hypothesis of (4.3.26), so every Gaussian moment Chapter 0.2 extracted by differentiating under the integral sign is now derived rather than provisional. The generating function of Chapter 0.2 §4.5, and with it the whole differentiation-with-respect-to-a-source apparatus that Part V runs on, inherits the same licence.

⚠ The dominating function is the theorem, not a technicality

It is tempting to read "there is an integrable gg with fng\abs{f_n}\le g" as a regularity condition that any reasonable sequence satisfies. Chapter 0.2's escaping spike is the counterexample and it is worth computing, because it shows the hypothesis failing by exactly the amount the conclusion fails.

Any dominating function must beat h(x,a)h(x,a) of (4.3.24) at every aa, so it must beat the supremum over aa. Substituting u=x/au=x/a turns hh into (2u2/x)eu2(2u^{2}/x)\ee^{-u^{2}}, whose maximum over u>0u\gt0 is at u=1u=1, so

supa>0h(x,a)  =  2ex,012exdx  =  . \sup_{a\gt0} h(x,a) \;=\; \frac{2}{\ee\,x}, \qquad \int_0^1\frac{2}{\ee\,x}\,\dd x \;=\; \infty.

The smallest possible dominating function is not integrable, so no dominating function exists, and dominated convergence does not apply. It is not that the theorem is silent about this sequence while the conclusion happens to hold. The conclusion is false here, and the hypothesis is the exact thing that fails.

The practical version, for a reader who will be swapping limits and integrals for the rest of the book: the danger is always mass escaping to where the integral cannot follow it, either up a narrowing spike or out to infinity. A dominating function is a fence that stops both.

In plain terms 4.3.4

Everything so far has been preparation for two statements, and these two are what the rebuilt integral is actually for. Both answer the same question: if a sequence of functions is approaching a limit, does the area under them approach the area under the limit?

The first says yes whenever the functions only ever increase, filling in from below and never retreating. Nothing else is required, and the proof leans on exactly one thing: that the size of a growing family of sets is the limit of their sizes. That was the awkward clause about endless lists in the definition of size, and this is what it was for.

The second is the one used daily. It says yes whenever the whole sequence can be held under a single fixed function of finite area, a ceiling that does not move as the sequence evolves. Without such a ceiling the statement is simply false, and the way it fails is always the same: the area runs away somewhere the limit cannot see it, up a spike that grows taller as it grows narrower, or off toward infinity in a bump that never shrinks but keeps moving. Each of those has a fixed amount of area at every stage and none in the limit. A ceiling of finite area is a fence against both.

The second statement is the one an earlier chapter borrowed and promised to repay. It was used to justify differentiating an integral by differentiating inside it, which is the trick that produced every moment of the Gaussian without solving a single new integral. The justification is now complete, and the bound it needs was written down at the time.

a natural place to stop  ·  the integral is built; what follows is the space

5 · L2L^{2}, and the functions that are not functions

Now the space itself. We collect the square-integrable functions, put Chapter 0.5's inner product on them, and check its three axioms one at a time. Two hold with no work at all. The third fails, and it fails on exactly the sequence Chapter 0.2 built, so the repair is forced rather than chosen. Making the repair costs something real, and §5.3 says what. Once it is made, every theorem of Chapter 0.5 §1 arrives with no proof needed, and §5.5 records the one way this space is not like Rn\R^{n}.

5.1 · The set, and that it is a vector space

Fix a measurable set XX, which for our purposes is R\R or an interval, and collect the measurable complex-valued functions on it whose squared modulus has finite integral:

L2(X)  =  {f:XC measurable : Xf2dμ<},f,g  =  Xfgdμ. \begin{aligned} \mathcal L^{2}(X) \;&=\; \left\{ f: X\to\C \ \text{measurable} \ :\ \int_X\abs{f}^{2}\dd\mu\lt\infty\right\},\\[6pt] \avg{f,g} \;&=\; \int_X \overline{f}\,g\,\dd\mu. \end{aligned} (4.3.28)

Two things need checking before that definition means anything, and both are one line. First, the collection is closed under addition, since f+g22f2+2g2\abs{f+g}^{2}\le2\abs f^{2}+2\abs g^{2} pointwise and the right-hand side is integrable. Second, the inner product converges, since fg12(f2+g2)\abs{\overline fg}\le\tfrac12(\abs f^{2}+\abs g^{2}) by the arithmetic-geometric mean inequality, so the integral defining f,g\avg{f,g} is finite for every pair in the set.

That second check is the reason the exponent is 22 rather than anything else. The square is the power for which the inner product of two members always exists, which is why quantum mechanics lives in this space and not in its neighbours.

5.2 · The three axioms, and the one that fails

Chapter 0.5 §1.1 laid out exactly three axioms, and the whole of that chapter follows from them. Take them in order against (4.3.28).

  • (i) Conjugate symmetry, f,g=g,f\avg{f,g}=\overline{\avg{g,f}}. Holds, because gˉf=fˉg\overline{\int\bar gf}=\int\bar fg and conjugation commutes with the integral.
  • (ii) Linearity in the second slot. Holds, because the integral is linear, which §4.1 established for the Lebesgue integral.
  • (iii) Positive definiteness, f,f>0\avg{f,f}\gt0 for f0f\ne0. Fails.

It fails, and the counterexample is the one this chapter opened with. Chapter 0.2's χ\chi is non-zero at every rational number, so it is not the zero function, and yet (4.3.18) gives χ,χ=χ2=0\avg{\chi,\chi}=\int\abs\chi^{2}=0. The same holds for every fnf_n of §1.3.

The general statement is worth having, because it says precisely how bad the failure is. For non-negative measurable ϕ\phi, Chebyshev's bound from the grind box of §3.5 gives μ{ϕ>1/k}kϕ\mu\{\phi\gt1/k\}\le k\int\phi, so a vanishing integral forces every one of those sets to be null, and their union is {ϕ>0}\{\phi\gt0\}:

Xf2dμ  =  0f=0  almost everywhere. \int_X \abs f^{2}\,\dd\mu \;=\; 0 \qquad\Longleftrightarrow\qquad f=0 \ \text{ almost everywhere}. (4.3.29)

So the axiom does not fail badly. It fails by exactly the width of a null set, and (4.3.29) tells you what to do about it.

5.3 · The repair, and what it costs

Since the norm cannot tell two functions apart when they agree almost everywhere, stop asking it to. Declare ff and gg to be the same vector when f=gf=g almost everywhere. That is an equivalence relation, and the space of states is the set of equivalence classes:

L2(X)  =  L2(X)/{f:f=0 almost everywhere}. L^{2}(X) \;=\; \mathcal L^{2}(X)\big/\big\{f : f=0 \text{ almost everywhere}\big\}. (4.3.30)

Axiom (iii) now holds by construction, since f,f=0\avg{f,f}=0 says ff is in the class of the zero vector, which is what "f=0f=0 in L2L^{2}" now means. All three axioms hold, so L2(X)L^{2}(X) is an inner-product space in exactly Chapter 0.5's sense, and this is the moment that chapter was pointing at when it said of its table's function row that "that row is the entire reason Chapter 4.3 is possible."

The quotient is not free, and pretending it is would set up a misunderstanding that costs a whole chapter later.

⚠ A vector of L2L^{2} has no value at any point

Ask what ψ(0)\psi(0) is for ψL2(R)\psi\in L^{2}(\R) and the question has no answer. Every single point is a null set, so you may change ψ\psi at x=0x=0 to any value you like and it is still the same vector. Pointwise values are not properties of an element of L2L^{2}, and no theorem about L2L^{2} can produce one.

Three consequences, all of which arrive later in Part IV.

  • "The probability of finding the particle exactly at x0x_0" is not merely small, it is not defined, and this is why the Born rule for a continuous variable is stated with abψ2dx\int_a^b\abs\psi^{2}\dd x over a region rather than with ψ(x0)2\abs{\psi(x_0)}^{2} at a point.
  • The Gibbs overshoot of §8.4 is a statement about the values of a function at points, and it is therefore invisible to the norm of L2L^{2}. Both statements are true at once and they are not in conflict, because they are about different things.
  • The object x\ket x, which every physics text writes as though it were a vector, cannot be one: a state concentrated at a single point is the zero vector of L2L^{2}. Chapter 4.5 says what it is instead, and Chapter 0.9 §5.3 already flagged the same difficulty for eikx\ee^{\ii kx}.

In exchange for those three, the space acquires a genuine geometry and, in §6, a genuine completeness. It is a good trade, and it is the only one available.

Familiar ground — a density is already defined only up to a null set, and where the parallel stops

A probability density is one. Take the density of a continuous variable, change its value at finitely many points, or on any set of probability zero, and nothing observable moves. Every probability, every quantile, every moment, every survival function and every likelihood is an integral of the density, so all of them are blind to the change. The density is genuinely not a function with definite values. It is a class of functions agreeing almost everywhere, and a particular representative is chosen for convenience when one has to be written down.

The mathematics is identical, not analogous. Equation (4.3.30) is the statement that a density is well defined as an object of integration rather than as a pointwise assignment, and it is the same quotient by the same relation. A wavefunction gets the same treatment, and it is worth being careful about why. The reason is not that ψ2\abs\psi^{2} settles everything measurable, because it does not. The distribution of momentum is the squared modulus of the Fourier transform of ψ\psi, which §8.5 below builds out of the phase of ψ\psi as much as out of its modulus, and no amount of ψ2\abs\psi^{2} determines it. The reason is that ψ\psi is an element of (4.3.30) by construction, so altering it on a null set does not alter it at all.

One place the parallel stops, and it is worth knowing where. A density is usually chosen to be the continuous representative when one exists, and that choice is harmless because it is unique. In L2L^{2} there need be no continuous representative at all, and §8.1 has to prove that a continuous function can at least be found nearby.

5.4 · Chapter 0.5's theorems, transferred without a word changed

Here is the return on Chapter 0.5 having been written abstractly. That chapter proved Cauchy–Schwarz from the three axioms alone, and its proof mentions no basis, no dimension and no component. Every step of it is now available here with nothing to check:

Xfgdμ2    Xf2dμ Xg2dμ,f+g2    f2+g2. \abs{\int_X \overline{f}\,g\,\dd\mu}^{2} \;\le\; \int_X\abs f^{2}\dd\mu \ \int_X\abs g^{2}\dd\mu, \qquad \lVert f+g\rVert_2 \;\le\; \lVert f\rVert_2+\lVert g\rVert_2 . (4.3.31)

The first is Chapter 0.5's (0.5.6) with the vectors renamed, and the second is the triangle inequality that follows from it in three lines there. Both are theorems about inner-product spaces, and L2L^{2} is one. Chapter 0.5 said this would happen, in the solution to Problem 1 of its §10, where Gram–Schmidt run on 1,x,x2,1,x,x^{2},\ldots produced the Legendre polynomials: "the machinery does not care that the vectors are functions. That is the whole content of the abstraction, and it is what Chapter 4.3 will exploit." This is the exploitation, and it consists of writing nothing down.

What does not transfer is anything Chapter 0.5 proved using finite dimension, and that chapter listed the four places it used it: "in the induction, in rank–nullity, in the interchange of sums, in the claim that an injective map is surjective." Sections 6 and 7 supply replacements for as much of that as this chapter needs, and Chapters 4.4 and 4.5 handle the rest.

5.5 · Square-integrable does not mean integrable

One structural fact separates L2(R)L^{2}(\R) from every finite-dimensional space and from L2L^{2} of a bounded interval, and it has physical consequences that Chapter 4.9 will need.

On a bounded interval the two conditions are nested. Cauchy–Schwarz against the constant function 11 gives, for fL2[a,b]f\in L^{2}[a,b],

abfdμ    ba  f2  <  , \int_a^b \abs f\,\dd\mu \;\le\; \sqrt{b-a}\;\lVert f\rVert_2 \;\lt\;\infty, (4.3.32)

so square-integrable implies integrable there. On the whole line the constant function 11 is not in L2L^{2}, the factor ba\sqrt{b-a} has nothing to converge to, and the implication fails. Both inclusions fail, in fact, and one example each settles it:

  • f(x)=(1+x)3/4f(x)=(1+\abs x)^{-3/4} has f2=(1+x)3/2<\int\abs f^{2}=\int(1+\abs x)^{-3/2}\lt\infty and f=(1+x)3/4=\int\abs f=\int(1+\abs x)^{-3/4}=\infty. So fL2(R)f\in L^{2}(\R) and fL1(R)f\notin L^{1}(\R).
  • f(x)=x1/2f(x)=\abs x^{-1/2} on (0,1)(0,1) and 00 elsewhere has f=2\int\abs f=2 and f2=01x1dx=\int\abs f^{2}=\int_0^1 x^{-1}\dd x=\infty. So fL1(R)f\in L^{1}(\R) and fL2(R)f\notin L^{2}(\R).

The physical reading is worth having now rather than in Chapter 4.9. A wavefunction belongs to L2L^{2} because ψ2=1\int\abs\psi^{2}=1 is what normalisation means, and that guarantees nothing about ψ\int\psi, or about xψ2\int x\abs\psi^{2}. Worked example 3 exhibits a perfectly normalised state whose mean position does not exist, and the state is one Chapter 0.9 already introduced under another name.

In plain terms 4.3.5

The space of states can now be assembled. Take the functions whose squared size has a finite total, and give any two of them an overlap by multiplying one against the conjugate of the other and integrating. The choice of the square rather than some other power is not arbitrary: it is the only power for which the overlap of two members is guaranteed to be finite, which is why physics lives here.

Three requirements were laid down long ago for anything calling itself an overlap. Two of them hold immediately. The third demands that only the zero vector have zero length, and it fails, on precisely the function this chapter began with. A function supported on the fractions alone has zero length without being zero.

The failure is exactly as wide as a set of size zero, so the fix is to stop distinguishing functions that agree except on such a set. Two functions differing at a scatter of isolated points are declared to be the same vector. This sounds like an evasion and is not: it is the same convention already in force for probability densities, where changing the density at a few points changes no probability, no average and no likelihood, and nobody regards the density as ill-defined on that account.

The price is worth naming, because it is charged later. A vector in this space has no value at any particular point. The question of what the wavefunction equals at one exact position has no answer, and every honest statement about a continuous coordinate is an integral over a region. In return, the overlap becomes a genuine measure of distance, and every theorem about lengths and angles proved for abstract spaces becomes available here with no work at all, because none of those proofs ever asked what the vectors were.

6 · Completeness: the theorem, and the failure it repairs

Everything built so far exists for this section. Section 1 handed you a Cauchy sequence with nowhere to go. Sections 2 to 5 enlarged the class of functions, built an integral that survives limits, and fixed the norm so that distance zero means equality. The claim now is that the enlargement was exactly the right size: not merely bigger, but big enough that no Cauchy sequence can escape it. Section 6.2 states the theorem and marks it as quoted. Section 6.3 then takes the two sequences of §1 and shows you where they land.

6.1 · What is being claimed, and what is not

Convergence in L2L^{2} means fnf20\lVert f_n-f\rVert_2\to0, which written out is

Xfnf2dμ  n  0, \int_X \abs{f_n-f}^{2}\,\dd\mu \;\xrightarrow[n\to\infty]{}\; 0, (4.3.33)

and is often called convergence in mean square. Read what it says and, more importantly, what it does not. It says the total squared discrepancy shrinks to nothing. It says nothing whatever about the value of fnf_n at any particular point, and §5.3 already explained why it cannot: the vectors have no pointwise values to make a statement about. Section 8.4 exhibits a sequence converging in the sense of (4.3.33) whose graphs never stop overshooting.

With that understood, here is the claim. Every sequence in L2(X)L^{2}(X) that is Cauchy in the sense of (4.3.2) converges, in the sense of (4.3.33), to an element of L2(X)L^{2}(X). An inner-product space with that property is called a Hilbert space, and that is the proper name Chapter 0.9 §1.3 promised the space of square-integrable functions would eventually be given.

6.2 · Riesz–Fischer

⚑ Quoted, not derived — the completeness of L2L^{2}

Riesz–Fischer. L2(X)L^{2}(X) is complete for every measurable XX, and is therefore a Hilbert space.

The proof is quoted rather than run, and the shape of it is worth having so that the quotation is not a black box. Given a Cauchy sequence, thin it to a subsequence whose consecutive gaps shrink geometrically, fnk+1fnk2<2k\lVert f_{n_{k+1}}-f_{n_k}\rVert_2\lt2^{-k}, which is always possible. Then apply monotone convergence to the increasing partial sums of kfnk+1fnk\sum_k\abs{f_{n_{k+1}}-f_{n_k}} to show that the series converges at almost every point, so the telescoping sum fn1+k(fnk+1fnk)f_{n_1}+\sum_k(f_{n_{k+1}}-f_{n_k}) defines a function ff almost everywhere. Fatou's lemma applied to ffnk2\abs{f-f_{n_k}}^{2} then shows both that ff is in L2L^{2} and that fnkff_{n_k}\to f in norm. Finally, a Cauchy sequence with a convergent subsequence converges, by the triangle inequality.

It is proved in this form in the same two measure theory texts §2.3 named, and it is short enough there to read in an afternoon.

What is being assumed, precisely. Nothing beyond §2.3's package and §4's two theorems. The sketch above uses monotone convergence, Fatou, and the definition of the integral, and it uses no new input at all. So this mark is different in kind from the one in §2.3: that one imports a construction the book never performs, and this one marks an argument the book chooses not to write out. Both are honest to declare and only one of them is a genuine debt.

Two remarks before the theorem is spent. First, the sources quoted above prove the same statement for every LpL^{p} at once, so nothing about the exponent 22 is doing the work here. What the exponent 22 gives is the inner product of §5.1, and it is the combination of the two that makes L2L^{2} the space quantum mechanics needs. Second, the proof produced convergence almost everywhere only along a subsequence, and it is worth being exact about what that does and does not mean. Section 8.4's example is not evidence that the thinning is necessary. Its partial sums converge in norm and they also converge at every point, as §8.4's own theorem and Problem 5(c) both say, so what they fail is uniform convergence, which is a different statement. Whether norm convergence drags a whole sequence to a limit almost everywhere is a question this chapter does not take up, and the subsequence is all the sketch above delivers.

6.3 · The two sequences of §1, and where they land

Now collect. Both sequences from §1 are sequences in L2[0,1]L^{2}[0,1], both are Cauchy, and both now have limits. They are different kinds of repair and it is worth taking them in order.

Chapter 0.2's sequence. The functions fnf_n of §1.3 were 11 at q1,,qnq_1,\ldots,q_n and 00 elsewhere. In L2L^{2} each of them is the zero vector, by (4.3.29) and the vanishing of fn2\lVert f_n\rVert_2. Their pointwise limit χ\chi is also the zero vector, for the same reason. So the sequence converges, its limit is in the space, the limit's integral exists, and the integral is

01χdμ  =  0  =  limn01fndμ. \int_0^1 \chi\,\dd\mu \;=\; 0 \;=\; \lim_{n\to\infty}\int_0^1 f_n\,\dd\mu . (4.3.34)

Compare (4.3.34) with (4.3.5), the statement that failed in §1.3. It failed there because its right-hand side referred to nothing. Both sides now refer to something and the two agree. Chapter 0.2's warning box has been answered in full: the definition was thrown away, the integral was rebuilt, and 01χ=0\int_0^1\chi=0 came with it.

The thickened sequence. This is the one that showed a genuine hole, so it is the one worth watching close. The step functions gn=1Ung_n=\mathbf 1_{U_n} of §1.4 increase to 1U\mathbf 1_U at every point, and 1U\mathbf 1_U is measurable because UU is open. Its square integral is finite, so it is a vector of L2[0,1]L^{2}[0,1], and the distance from the sequence to it is computed by (4.3.9):

gn1U22  =  μ(UUn)  =  μ(U)μ(Un)  n  0. \lVert g_n-\mathbf 1_U\rVert_2^{\,2} \;=\; \mu(U\setminus U_n) \;=\; \mu(U)-\mu(U_n) \;\xrightarrow[n\to\infty]{}\; 0 . (4.3.35)

So the sequence converges in L2L^{2}, its limit is a genuine element of the space, and the limit has an integral, namely 011U=μ(U)\int_0^1\mathbf 1_U=\mu(U), a number at most 14\tfrac14.

Now read that against the grind box of §1.4, which is where the payoff is. That box proved that no Riemann-integrable function is an L2L^{2} limit of the gng_n. Since 1U\mathbf 1_U is that limit, and limits in a metric space are unique up to distance zero, the box says exactly this:

1Uh2  >  0for every Riemann-integrable h on [0,1]. \lVert \mathbf 1_U - h\rVert_2 \;\gt\; 0 \qquad\text{for every Riemann-integrable } h \text{ on } [0,1]. (4.3.36)

The vector the sequence was reaching for is not a relabelling of anything the old theory had. It sits at positive distance from every last one of them. The hole exhibited in §1 was real, the enlargement of §2 to §5 filled it, and (4.3.36) is the receipt.

6.4 · What completeness gets spent on

Chapter 4.2 wrote its whole apparatus on the assumption that limits of states are states, and listed the assumption as owed. Here is what the theorem now pays for, named by chapter so the credit can be tracked.

  • Section 7.2, immediately. A series ncnen\sum_n c_ne_n with ncn2<\sum_n\abs{c_n}^{2}\lt\infty converges to a vector. Without completeness the coefficients would be square-summable and the sum would be nothing, which would leave "expand in a basis" meaningless.
  • Chapter 4.5's spectral theorem. Every construction there produces its operator as a limit of simpler ones, and the limit has to exist in the space it started in.
  • Chapter 4.6's time evolution. The operator eiH^t/\ee^{-\ii\hat Ht/\hbar} is defined by a series, and applying it to a state produces a state only because the partial sums are Cauchy and converge.
  • Chapter 4.15's perturbation series. The first correction to a state is an infinite sum over the unperturbed basis, and it names a state only because its coefficients are square-summable and §7.2 then supplies a limit.
In plain terms 4.3.6

This is the centre of the chapter. The claim is that the enlarged space has no gaps: any sequence of states that is settling down settles down onto a state, and never onto something outside the space. That statement is what an earlier chapter took entirely on credit every time it said that a limit of states is a state, and it is the property whose absence made the old integral unusable.

The theorem itself is quoted rather than proved, and the quotation is marked. The shape of the argument is short enough to describe: thin the sequence so that consecutive terms are twice as close as the last pair, add up the gaps, and the two convergence theorems of the previous section show both that the total settles at almost every point and that what it settles on belongs to the space.

The satisfying part is what happens to the two sequences from the beginning of the chapter. The first, built from the fractions one at a time, now converges, and the object it converges to is the zero vector, with an area of exactly zero. That was promised eight chapters ago and it is now delivered.

The second is the one that mattered. It converged nowhere before, and now it converges to a perfectly respectable member of the space, an object at a positive distance from every function the old theory could handle. The gap that was exhibited was real, and what fills it is genuinely new rather than a relabelling. That is what it means for the enlargement to have been exactly the right size.

a natural place to stop  ·  the space is complete; what follows is finding a basis in it

7 · Orthonormal bases in infinite dimensions

Completeness in hand, we can finally ask what a basis is here. In finite dimensions Chapter 0.4 settled it by counting, and the count is not available. Two tools come first. Section 7.1 shows that the coefficients of any vector against any orthonormal set are square-summable, and §7.2 shows that any square-summable list of coefficients assembles into a vector, which is where §6 gets spent. What then replaces the count is four statements which look different, are all called completeness of a basis, and turn out to be the same statement. Section 7.3 proves that. Section 7.4 then shows that a basis of this space can always be listed, which is what every physics calculation has been assuming.

One older promise is being kept here too. Chapter 1.1 §4.3 observed that the configuration of a field is a function rather than a list of numbers, that the space of such functions is therefore infinite-dimensional, and that "Chapter 4.3 will take its infinite-dimensionality seriously." Taking it seriously is exactly this section, and the cost is now visible. The dimension count that settled every question in Chapter 0.4 is gone, and four separate statements have to be proved equivalent to replace it.

7.1 · Bessel's inequality

Let e1,e2,e_1,e_2,\ldots be an orthonormal set in L2L^{2}, meaning em,en=δmn\avg{e_m,e_n}=\delta_{mn}, with no assumption yet that there are enough of them. For any ff write cn=en,fc_n=\avg{e_n,f} and look at how far ff is from nNcnen\sum_{n\le N}c_ne_n. That combination is the projection of ff onto the span of the first NN, which by Chapter 0.5 §3.1 is the closest point there, and that is why it is the combination worth measuring against. The inequality itself uses none of it. All it uses is that a squared norm can never be negative, so expand the squared distance and let the cross terms collapse:

0    fnNcnen2  =  f2nNcn2n=1cn2    f2. 0 \;\le\; \Big\lVert f-\sum_{n\le N}c_n e_n\Big\rVert^{2} \;=\; \lVert f\rVert^{2}-\sum_{n\le N}\abs{c_n}^{2} \qquad\Longrightarrow\qquad \sum_{n=1}^{\infty}\abs{c_n}^{2} \;\le\; \lVert f\rVert^{2}. (4.3.37)

That is Bessel's inequality, and the derivation is Chapter 0.5's Parseval computation with the equality weakened to an inequality because the set may not be big enough. Its content is small and its consequence is not: whatever else is true, the coefficients of any ff against any orthonormal set are square-summable.

7.2 · Square-summable coefficients build a vector

Bessel says the coefficients are well behaved. Completeness turns that into an actual vector, and this is the first place §6 gets spent. Suppose ncn2<\sum_n\abs{c_n}^{2}\lt\infty and look at the partial sums SN=nNcnenS_N=\sum_{n\le N}c_ne_n. Orthonormality makes the distance between two of them a tail of a convergent series:

SNSM2  =  M<nNcn2  M,N  0. \lVert S_N-S_M\rVert^{2} \;=\; \sum_{M\lt n\le N}\abs{c_n}^{2} \;\xrightarrow[M,N\to\infty]{}\; 0 . (4.3.38)

So (SN)(S_N) is Cauchy, and by Riesz–Fischer it converges to some vector of L2L^{2}. The series ncnen\sum_n c_ne_n therefore means something, for any square-summable coefficients whatever. Without §6 the partial sums would be closing on a gap and the notation would name nothing.

One small tool is needed repeatedly below and is worth extracting. Cauchy–Schwarz gives u,vu,v=u,vvuvv\abs{\avg{u,v}-\avg{u,v'}}=\abs{\avg{u,v-v'}}\le\lVert u\rVert\lVert v-v'\rVert, so the inner product is continuous in each slot. Norm convergence may therefore be taken inside any inner product, which is used at every step of §7.3.

7.3 · Four statements, and the proof that they are one

The word "complete" is used of a basis as well as of a space, and it means something different there. Worse, it means four different-looking things, and physics texts move between them without comment. Here they are, for an orthonormal set {en}\{e_n\} in a Hilbert space HH.

  • (a) Spanning. The finite linear combinations of the ene_n come arbitrarily close to every vector of HH.
  • (b) Expansion. Every ff equals nen,fen\sum_n\avg{e_n,f}e_n, the series converging in norm.
  • (c) Parseval. Every ff satisfies f2=nen,f2\lVert f\rVert^{2}=\sum_n\abs{\avg{e_n,f}}^{2}.
  • (d) No blind spot. The only vector orthogonal to every ene_n is the zero vector.

An orthonormal set satisfying these is an orthonormal basis. The equivalence is what makes the definition checkable, because (d) is the one you can usually verify and (b) is the one you want to use.

(b) implies (a), since the partial sums of the expansion are finite combinations and they converge to ff.

(a) implies (d). Let ff be orthogonal to every ene_n, hence to every finite combination of them. By (a) there are finite combinations gkfg_k\to f in norm. Continuity of the inner product from §7.2 then gives

f2  =  f,f  =  limkf,gk  =  0, \lVert f\rVert^{2} \;=\; \avg{f,f} \;=\; \lim_{k\to\infty}\avg{f,g_k} \;=\; 0, (4.3.39)

because every term in the sequence is zero. So ff is the zero vector.

(d) implies (b), and this is the step that uses everything. Given ff, its coefficients cn=en,fc_n=\avg{e_n,f} are square-summable by Bessel, so §7.2 says g=ncneng=\sum_nc_ne_n exists in HH. Pair any eme_m against the difference fgf-g, moving the limit inside by continuity, and only one term survives orthonormality:

em,fg  =  em,fncnem,en  =  cmcm  =  0. \avg{e_m,\,f-g} \;=\; \avg{e_m,f}-\sum_{n}c_n\avg{e_m,e_n} \;=\; c_m-c_m \;=\; 0. (4.3.40)

So fgf-g is orthogonal to every eme_m, and (d) forces f=gf=g, which is (b).

(b) implies (c), by taking the norm of the expansion. Pythagoras from Chapter 0.5 §1.3 gives SN2=nNcn2\lVert S_N\rVert^{2}=\sum_{n\le N}\abs{c_n}^{2}, and the norm is continuous because uvuv\big|\lVert u\rVert-\lVert v\rVert\big|\le\lVert u-v\rVert, so letting NN run turns that into (c). And (c) implies (d) immediately: a vector orthogonal to every ene_n has all coefficients zero, so (c) makes its norm zero, so it is the zero vector.

The implications proved are (b) to (a) to (d) to (b), which is a closed cycle containing three of the four, and (b) to (c) to (d), which brings in the fourth. All four are therefore equivalent. \blacksquare

Two things are worth noticing about which chapter supplied which step. The finite-dimensional proof in Chapter 0.5 §2.3 needed none of this, because there the four statements are all immediate consequences of counting dimensions. Here the count is unavailable and completeness of the space is what replaces it, entering at §7.2 and nowhere else. That is the precise sense in which Chapter 4.2's basis expansions were on credit.

7.4 · A basis can always be listed

Everything in §7.3 assumed the orthonormal set was indexed by the integers, and that assumption needs justifying, because a Hilbert space could in principle need an uncountable basis. It cannot, for this space, and the reason is that L2L^{2} is not too large to be approximated from a list.

A space is separable when it contains a countable dense subset. Building one for L2(R)L^{2}(\R) takes three steps, and the second is where §2.3's regularity clause is spent. It is spent once, here, and §8.1 re-uses the same step rather than paying for it again.

  • Simple functions are dense. For f0f\ge0 in L2L^{2} take the staircase (4.3.17). Then fφn2f2\abs{f-\varphi_n}^{2}\le\abs f^{2}, which is integrable, and fφn20\abs{f-\varphi_n}^{2}\to0 pointwise, so dominated convergence gives fφn20\lVert f-\varphi_n\rVert_2\to0. A general complex ff splits into four non-negative pieces.
  • Indicators are nearly intervals. For measurable EE of finite measure, regularity supplies an open OEO\supseteq E with μ(OE)\mu(O\setminus E) as small as desired. Any open subset of R\R is a countable disjoint union of open intervals, and the reason is worth a line: take for each point the largest open interval inside OO containing it, two such intervals are equal or disjoint, and each one holds a rational, so there are countably many. Countable additivity then lets finitely many of them carry all but a little of the measure, and keeping those leaves a set AA with μ(EA)\mu(E\triangle A) small, so 1E1A22=μ(EA)\lVert\mathbf 1_E-\mathbf 1_A\rVert_2^{2}=\mu(E\triangle A) is small.
  • Count them. Perturbing the endpoints of those finitely many intervals to rationals, and the coefficients of the simple function to rationals with rational imaginary parts, changes the norm by as little as you like and leaves a countable collection.

So L2(R)L^{2}(\R) is separable. Now the consequence, which is three lines. Let {eα}\{e_\alpha\} be any orthonormal set, and note that distinct members are far apart, since orthonormality and Pythagoras give

eαeβ2  =  eα2+eβ2  =  2,αβ. \lVert e_\alpha-e_\beta\rVert^{2} \;=\; \lVert e_\alpha\rVert^{2}+\lVert e_\beta\rVert^{2} \;=\; 2, \qquad \alpha\ne\beta. (4.3.41)

Put an open ball of radius 122\tfrac12\sqrt2 around each eαe_\alpha, and by (4.3.41) the balls are disjoint. A countable dense set must place at least one of its points in every one of them, so the balls can be labelled injectively by a countable set, and there are at most countably many of them. Every orthonormal set in a separable Hilbert space is countable, so every orthonormal basis of L2L^{2} can be written as a sequence e1,e2,e3,e_1,e_2,e_3,\ldots and expansions can be written as ordinary series.

Existence is settled the same way. Run Chapter 0.5 §2.1's Gram–Schmidt on a countable dense set, discarding any vector that is already in the span of its predecessors. The output is a countable orthonormal set whose finite combinations reach everything the dense set reached, which is statement (a), so it is a basis. The recipe is the one from Chapter 0.5 with no modification, which is a third instance of that chapter's machinery not caring what the vectors are.

One consequence deserves stating, with a caution attached to it. Given two separable infinite-dimensional Hilbert spaces with bases {en}\{e_n\} and {en}\{e_n'\}, the map sending cnen\sum c_ne_n to cnen\sum c_ne_n' is defined on everything by (b), lands everywhere by (b), and preserves norms by (c). So it is unitary in Chapter 0.5 §4.3's sense. Up to a unitary map there is only one separable infinite-dimensional Hilbert space, and every question about which one to use is a question about which basis is convenient.

Here is the caution, because that result is easy to over-spend. It depends on a choice of basis at each end and singles out no particular map, so it can never tell you that some given map is the right one. Section 8.5 needs exactly that, and it gets it from Plancherel instead. Nothing later in this chapter uses the statement just proved.

In plain terms 4.3.7

In finitely many dimensions a basis is a set of perpendicular unit vectors, and you know it is big enough by counting: three of them in three dimensions, and there is nothing more to check. Counting is not available here, so the question of whether a perpendicular set is big enough has to be answered some other way.

Four different answers are in circulation and they all get called completeness. The finite combinations reach everywhere. Every vector is the sum of its components. The squared length is the sum of the squared components. And nothing but zero is perpendicular to the whole set. They look like four separate claims, and the section proves they are one claim in four costumes, which is useful because the last is usually the only one you can check and the second is always the one you want to use.

The proof leans on the previous section in exactly one place: to know that a list of components with a finite total square actually assembles into a vector rather than pointing at a gap. That is where completeness of the space is spent, and it is the reason the two senses of the word have to be kept apart.

One more thing has to be true before any of it is usable, and it is easy to overlook. A basis here could in principle be too big to write as a list. It is not, and the argument is a pleasant one: any two distinct perpendicular unit vectors are a fixed distance apart, so small balls around them never overlap, so a countable supply of approximating points cannot serve more than countably many of them. Bases can therefore be numbered, expansions are ordinary series, and everything a physicist writes with a summation sign is legitimate.

8 · The Fourier basis is a basis

Chapter 0.9 proved that the Fourier modes are orthonormal, said in a marked box that it had not proved there were enough of them, and named this chapter. Section 8.1 shows that continuous functions are dense, using nothing new. Section 8.2 proves, constructively, that a continuous periodic function is a uniform limit of trigonometric polynomials. Section 8.3 puts the two together and the mark comes off. Section 8.4 then measures the two kinds of convergence against each other on Chapter 0.9's own square wave, and §8.5 collects the last debt of the chapter.

8.1 · Continuous functions are dense, on the regularity step §7.4 already paid for

This subsection supplies the first of the two links §8.3 needs. Every vector of L2L^{2} sits as close as you like to a continuous function, and on [π,π][-\pi,\pi] to a continuous periodic one, which is what will let §8.2 work only with continuous functions and still reach the whole space.

Nothing here is quoted, and it is worth being explicit about why, because the result looks like the kind of thing that ought to need a flag. Section 2.3's package included regularity, clause (d), and §7.4 spent it. This section re-uses that step and adds one link of its own. Watch the clause do the second of the two jobs it was bought for.

The chain has three links and the first two are already built. Simple functions are dense in L2L^{2}, by the staircase argument of §7.4. An indicator 1E\mathbf 1_E is close in L2L^{2} to the indicator of a finite union of intervals, which is §7.4's regularity step verbatim. So it is enough to approximate the indicator of a single interval by a continuous function, and that is done by hand.

Replace the vertical sides of 1[a,b]\mathbf 1_{[a,b]} by straight ramps of horizontal width δ\delta, giving a continuous trapezoid TδT_\delta. The two functions agree except on the two ramps, where they differ by at most 11, so

1[a,b]Tδ22  =  ramps1[a,b]Tδ2dμ    2δ, \lVert \mathbf 1_{[a,b]}-T_\delta\rVert_2^{\,2} \;=\; \int_{\text{ramps}}\abs{\mathbf 1_{[a,b]}-T_\delta}^{2}\,\dd\mu \;\le\; 2\delta, (4.3.42)

which goes to zero with δ\delta. Composing the three links, the continuous functions are dense in L2L^{2}. On [π,π][-\pi,\pi] one extra turn of the same screw is needed for §8.2: place the ramps so that the trapezoid vanishes near both endpoints, and it extends to a continuous 2π2\pi-periodic function on the line with no jump at the seam. So the continuous 2π2\pi-periodic functions are dense in L2[π,π]L^{2}[-\pi,\pi] as well.

8.2 · Fejér's theorem, built from a kernel you have already met

Work on [π,π][-\pi,\pi], which is Chapter 0.9's box with L=2πL=2\pi and therefore kn=nk_n=n, and write Chapter 0.9's coefficients and partial sums:

an=12πππf(y)einydy,SNf(x)=nNaneinx. a_n = \frac{1}{2\pi}\int_{-\pi}^{\pi}f(y)\,\ee^{-\ii ny}\,\dd y, \qquad S_N f(x) = \sum_{\abs n\le N} a_n\,\ee^{\ii nx}. (4.3.43)

Direct attack on SNfS_Nf is the route that fails, and the reason is worth naming in advance. The argument below works by writing the approximation as an integral of ff against a weight, and it needs that weight to be non-negative. The weight belonging to SNfS_Nf is the Dirichlet kernel, written out below when the averaging is performed, and it is a ratio of two sines, which changes sign. The route that works replaces the partial sums by their running average, which turns the weight into a square. Define the Cesàro means

σNf  =  1N+1M=0NSMf. \sigma_N f \;=\; \frac{1}{N+1}\sum_{M=0}^{N} S_M f. (4.3.44)

To see what averaging has done, substitute (4.3.43) into (4.3.44) and pull the sum inside the integral, which is legal because it is finite. What comes out is a convolution against a single function of t=xyt=x-y:

σNf(x)  =  12πππf(xt)KN(t)dt,KN(t)  =  1N+1M=0N nMeint. \sigma_N f(x) \;=\; \frac{1}{2\pi}\int_{-\pi}^{\pi} f(x-t)\,K_N(t)\,\dd t, \qquad K_N(t) \;=\; \frac{1}{N+1}\sum_{M=0}^{N}\ \sum_{\abs n\le M}\ee^{\ii nt}. (4.3.45)

The inner sum is a geometric series, and summing it gives the Dirichlet kernel DM(t)=sin((M+12)t)/sin(t/2)D_M(t)=\sin\big((M+\tfrac12)t\big)/\sin(t/2). Averaging those over MM is a telescoping exercise. Multiply by 2sin2(t/2)2\sin^{2}(t/2), which turns each DMD_M into 2sin(t/2)sin((M+12)t)=cos(Mt)cos((M+1)t)2\sin(t/2)\sin\big((M+\tfrac12)t\big)=\cos(Mt)-\cos\big((M+1)t\big), and everything cancels except the ends, leaving 1cos((N+1)t)=2sin2((N+1)t/2)1-\cos\big((N+1)t\big)=2\sin^{2}\big((N+1)t/2\big). So

KN(t)  =  1N+1(sin((N+1)t/2)sin(t/2))2. K_N(t) \;=\; \frac{1}{N+1}\left(\frac{\sin\big((N+1)t/2\big)}{\sin(t/2)}\right)^{2}. (4.3.46)

Now read off three properties of (4.3.46), because they are the whole theorem. It is non-negative, being a square, and that is exactly what the averaging bought. Its integral is one, since 12πππDM=1\frac{1}{2\pi}\int_{-\pi}^{\pi}D_M=1 for every MM as only the n=0n=0 term survives, and the average of ones is one. And it concentrates at the origin, since for δtπ\delta\le\abs t\le\pi the denominator is bounded below and

0    KN(t)    1(N+1)sin2(δ/2)  N  0uniformly on δtπ. 0 \;\le\; K_N(t) \;\le\; \frac{1}{(N+1)\sin^{2}(\delta/2)} \;\xrightarrow[N\to\infty]{}\; 0 \quad\text{uniformly on } \delta\le\abs t\le\pi. (4.3.47)

A non-negative bump, of unit integral, narrowing at the origin. That is Chapter 0.9 §5.2's delta-sequence, and the argument that follows is that section's proof with the Gaussian replaced by (4.3.46).

Fejér's theorem. If ff is continuous and 2π2\pi-periodic then σNff\sigma_Nf\to f uniformly. Since KN/2π=1\int K_N/2\pi=1, the error is

σNf(x)f(x)  =  12πππ[f(xt)f(x)]KN(t)dt, \sigma_Nf(x)-f(x) \;=\; \frac{1}{2\pi}\int_{-\pi}^{\pi}\big[f(x-t)-f(x)\big]\,K_N(t)\,\dd t, (4.3.48)

and it is split at t=δ\abs t=\delta. On the inside, ff is uniformly continuous, being continuous on a closed bounded interval, so δ\delta can be chosen once and for all to make f(xt)f(x)<ε\abs{f(x-t)-f(x)}\lt\varepsilon for every xx at once. That part of the integral is then below ε\varepsilon, because KNK_N is non-negative with unit integral. On the outside, the bracket is at most 2maxf2\max\abs f and (4.3.47) drives KNK_N to zero uniformly, so that part is below ε\varepsilon for NN large. The bound is uniform in xx throughout. \blacksquare

Two notes on what that proof used. The uniform continuity is Heine–Cantor, which Chapter 0.2 §1.1 quoted and marked when it proved that continuous functions are Riemann integrable; this chapter leans on that standing mark and does not raise a new one. And the theorem is constructive: it produces the approximating trigonometric polynomial explicitly, as an average of partial sums, rather than asserting one exists. Since σNf\sigma_Nf is a finite combination of the einx\ee^{\ii nx}, what the theorem says is that the trigonometric polynomials are dense, in the supremum norm, in the continuous 2π2\pi-periodic functions, and that a member of the dense set can be written down from ff by averaging.

8.3 · The mark comes off

Now assemble, and note that each link was proved above and none is quoted.

  • Uniform convergence implies L2L^{2} convergence on a bounded interval, since g22πmaxg\lVert g\rVert_2\le\sqrt{2\pi}\max\abs g. So by Fejér's theorem in §8.2 the trigonometric polynomials come arbitrarily close in L2L^{2} to every continuous periodic function.
  • Those are dense in L2[π,π]L^{2}[-\pi,\pi], by §8.1.
  • Therefore the trigonometric polynomials are dense in L2[π,π]L^{2}[-\pi,\pi], which is statement (a) of §7.3 for the set {en}\{e_n\} with en(x)=einx/2πe_n(x)=\ee^{\ii nx}/\sqrt{2\pi}.
  • Chapter 0.9 §1.1 proved by direct integration that this set is orthonormal.

By the equivalence of §7.3, all four statements hold. So the Fourier modes are an orthonormal basis of L2[π,π]L^{2}[-\pi,\pi], and in particular, for every ff with f2<\int\abs f^{2}\lt\infty,

fSNf2  N  0and12πππf2dx  =  n=an2. \lVert f-S_Nf\rVert_2 \;\xrightarrow[N\to\infty]{}\; 0 \qquad\text{and}\qquad \frac{1}{2\pi}\int_{-\pi}^{\pi}\abs f^{2}\dd x \;=\; \sum_{n=-\infty}^{\infty}\abs{a_n}^{2}. (4.3.49)
Chapter 0.9 §1.3's mark is paid, not carried

That chapter wrote, in a box headed "Quoted, not derived" and carrying the mark: "Everything above shows that the ene_n are orthonormal. It does not show that they are enough… Both facts are proved in Chapter 4.3, where the space of square-integrable functions gets its proper name, a Hilbert space, and completeness stops being an assumption. Until then we are standing on it, and you know that we are."

Both facts are now proved. The first is the left half of (4.3.49), and it came from Fejér plus the equivalence of §7.3. The second, about pointwise convergence and what happens at a jump, is §8.4 below. The space has its name, by §6.2. Nothing in Chapter 0.9 changes, and nothing in it was wrong. What changes is that its mark has been paid rather than carried.

8.4 · Two kinds of convergence, measured against each other

Equation (4.3.49) is a statement about a norm, and Chapter 0.9 §1.4 exhibited a function for which it is emphatically not a statement about points. This is the last piece of that chapter's mark, and it is also the numerical confirmation this chapter owes.

First the pointwise theorem, so that the two claims can be set beside each other. It rests on one lemma about oscillating integrals, and the lemma is worth stating because it recurs throughout Part V.

Riemann–Lebesgue lemma. If ϕ\phi is integrable then ϕ(t)sin(λt)dt0\int\phi(t)\sin(\lambda t)\,\dd t\to0 as λ\lambda\to\infty. The reason is that rapid oscillation averages a slowly varying weight against itself, and the proof turns that into three lines in the grind box.

With it, the classical statement about jumps follows. If ff is 2π2\pi-periodic and integrable, and at the point xx both one-sided limits exist with bounded one-sided difference quotients, then SNf(x)12[f(x+)+f(x)]S_Nf(x)\to\tfrac12\big[f(x^{+})+f(x^{-})\big]. At a point of continuity that is f(x)f(x), and at a jump it is the midpoint, exactly as Chapter 0.9 stated without proof.

One word about the hypothesis, because Chapter 0.9 gave it a name and this one does not match it. That chapter asked for the Dirichlet conditions, piecewise smooth with finitely many extrema and jumps, imposed on ff across the whole interval. The statement above asks less. It is a condition at the single point xx, and away from xx it wants nothing beyond integrability. The same conclusion from a weaker hypothesis is the stronger theorem, so the Dirichlet conditions are superseded here rather than merely met.

Grind box — Riemann–Lebesgue, and convergence to the midpoint of a jump

Riemann–Lebesgue. For ϕ=1[a,b]\phi=\mathbf 1_{[a,b]} the integral is computed outright,

absin(λt)dt  =  cosλacosλbλ    2λ    0, \left|\int_a^b\sin(\lambda t)\,\dd t\right| \;=\; \left|\frac{\cos\lambda a-\cos\lambda b}{\lambda}\right| \;\le\; \frac{2}{\lambda} \;\to\;0,

so it holds for step functions by linearity. The regularity argument of §7.4, run with the L1L^{1} norm in place of the L2L^{2} norm, makes step functions dense in L1L^{1}. So given ε\varepsilon pick a step function ss with ϕs1<ε\lVert\phi-s\rVert_1\lt\varepsilon, and

ϕsinλt    ϕs  +  ssinλt  <  ε+Cλ, \left|\int\phi\sin\lambda t\right| \;\le\; \int\abs{\phi-s} \;+\; \left|\int s\sin\lambda t\right| \;\lt\; \varepsilon + \frac{C}{\lambda},

which is below 2ε2\varepsilon for large λ\lambda. \blacksquare

The jump. Summing the geometric series in (4.3.43) gives SNf(x)=12πππf(xt)DN(t)dtS_Nf(x)=\frac{1}{2\pi}\int_{-\pi}^{\pi}f(x-t)D_N(t)\dd t with DN(t)=sin((N+12)t)/sin(t/2)D_N(t)=\sin\big((N+\frac12)t\big)/\sin(t/2). The kernel is even and 12πππDN=1\frac{1}{2\pi}\int_{-\pi}^{\pi}D_N=1, so 12π0πDN=12\frac{1}{2\pi}\int_0^{\pi}D_N=\frac12 and therefore

SNf(x)f(x+)+f(x)2  =  12π0π ⁣[f(xt)f(x)]DN(t)dt+  12π0π ⁣[f(x+t)f(x+)]DN(t)dt. \begin{aligned} S_Nf(x)-\frac{f(x^{+})+f(x^{-})}{2} \;=\;& \frac{1}{2\pi}\int_0^{\pi}\!\big[f(x-t)-f(x^{-})\big]D_N(t)\dd t\\[4pt] +\;& \frac{1}{2\pi}\int_0^{\pi}\!\big[f(x+t)-f(x^{+})\big]D_N(t)\dd t. \end{aligned}

Take the first. Writing ϕ(t)=[f(xt)f(x)]/sin(t/2)\phi(t)=\big[f(x-t)-f(x^{-})\big]/\sin(t/2) turns it into 12π0πϕ(t)sin((N+12)t)dt\frac{1}{2\pi}\int_0^{\pi}\phi(t)\sin\big((N+\frac12)t\big)\dd t. Near t=0t=0 the numerator is at most a constant times tt by the bounded difference quotient, and sin(t/2)\sin(t/2) behaves like t/2t/2, so ϕ\phi is bounded near the origin and integrable on [0,π][0,\pi]. Riemann–Lebesgue kills it, and the second term goes the same way. \blacksquare

Now put the two convergences on the same function and watch them separate. Chapter 0.9 §1.4's square wave on [π,π][-\pi,\pi] is f=+1f=+1 on (0,π)(0,\pi) and f=1f=-1 on (π,0)(-\pi,0), with the series it derived there,

f(x)  =  4π(sinx+sin3x3+sin5x5+),f22=ππ1dx=2π. f(x) \;=\; \frac4\pi\left(\sin x+\frac{\sin3x}{3}+\frac{\sin5x}{5}+\cdots\right), \qquad \lVert f\rVert_2^{\,2} = \int_{-\pi}^{\pi}1\,\dd x = 2\pi . (4.3.50)

Write SNS_N for the partial sum keeping every odd harmonic up to NN. Parseval, which is now a theorem rather than an assumption, turns the norm of the error into a tail of a series. Using ππsin2nxdx=π\int_{-\pi}^{\pi}\sin^{2}nx\,\dd x=\pi and squaring the coefficients of (4.3.50),

fSN22f22  =  8π2n>Nn odd1n2    8π212N  =  4π2N, \frac{\lVert f-S_N\rVert_2^{\,2}}{\lVert f\rVert_2^{\,2}} \;=\; \frac{8}{\pi^{2}}\sum_{\substack{n\gt N\\ n\ \mathrm{odd}}}\frac{1}{n^{2}} \;\approx\; \frac{8}{\pi^{2}}\cdot\frac{1}{2N} \;=\; \frac{4}{\pi^{2}N}, (4.3.51)

the tail being estimated by an integral because the odd integers past NN are spaced by two. Taking the square root gives the rate this chapter promised to confirm:

fSN2f2    2πN1/2  =  0.63662  N1/2. \frac{\lVert f-S_N\rVert_2}{\lVert f\rVert_2} \;\approx\; \frac{2}{\pi}\,N^{-1/2} \;=\; 0.63662\;N^{-1/2}. (4.3.52)

Against that, set the pointwise behaviour. Chapter 0.9's grind box computed the peak of the partial sum near the jump and found it converging to 2πSi(π)=1.178980\tfrac2\pi\mathrm{Si}(\pi)=1.178980 rather than to 11. The jump has size 22, so the overshoot, as a fraction of the jump, is

maxxSN(x)12  N  1π0πsinttdt12=  0.58948990.5  =  0.0894899. \begin{aligned} \frac{\max_x S_N(x)-1}{2} \;&\xrightarrow[N\to\infty]{}\; \frac{1}{\pi}\int_0^{\pi}\frac{\sin t}{t}\,\dd t-\frac12\\[4pt] &=\; 0.5894899-0.5 \;=\; 0.0894899 . \end{aligned} (4.3.53)

Two numbers, both exact, both about the same sequence of functions, and they do opposite things. Equation (4.3.52) falls to zero. Equation (4.3.53) settles on a constant and stays there forever.

They are consistent, and the reconciliation is one sentence. The overshoot keeps its height and loses its width. The first peak sits at x=π/(N+1)x=\pi/(N+1), so the whole ear is squeezed into a window shrinking like 1/N1/N, and a fixed height on a vanishing width contributes a vanishing amount to fSN2\int\abs{f-S_N}^{2}. A norm integrates and cannot see it. A maximum does not integrate and sees nothing else. The figure below is that sentence made visible, and it is worth using the zoom control, because the ear is invisible at any fixed magnification once NN is large and is exactly the same height at every magnification that tracks it.

N = 17
L2: || f - S_N ||_2 / || f ||_2 = 0.149976 times sqrt(N) = 0.618366 against 2/pi = 0.636620 (exact, from the coefficient tail)
pointwise: peak of S_N = 1.180010 overshoot = 0.090005 of the jump against (1/pi)Si(pi) - 1/2 = 0.089490 (measured on the plotted curve)
slopes over the last decade of N: L2 error -0.4981 overshoot -0.0001 -- one converges like N^(-1/2), the other does not converge at all
One sequence, two convergences, and they are not the same convergence. Top: the square wave of (4.3.50) in grey, its partial sum SNS_N in blue, and the limiting overshoot level 1.1789801.178980 as a dashed amber rule. Bottom: both errors against NN, on logarithmic axes over three decades. The blue curve is fSN2/f2\lVert f-S_N\rVert_2/\lVert f\rVert_2, computed from the exact coefficient tail of (4.3.51), with (4.3.52) dashed beside it. The amber curve is the overshoot as a fraction of the jump, found by scanning the plotted partial sum and refining the peak parabolically, with no use of (4.3.53) anywhere in the computation. The two curves are the point of the figure. One has slope 12-\tfrac12 across three decades and the other has slope zero, and the third readout measures both slopes over the last decade rather than asserting them. Now press "track the jump". The window follows π/(N+1)\pi/(N+1), so the first arch always fills the frame, and the ear stands at exactly the same height at every setting of the slider while the width of the window it occupies falls by a factor of 501501, the half-width being 7π/(N+1)7\pi/(N+1) and the slider running from N=1N=1 to N=1001N=1001. That is the whole reconciliation: convergence in the norm of L2L^{2} is compatible with a permanent, unshrinking error in the values, provided the region carrying it shrinks. Chapter 0.9 said the two notions disagree. This is by how much.

The physical reading matters more than the arithmetic, and it is the reason §5.3 warned about pointwise values in advance. Every statement quantum mechanics makes about a continuous coordinate is an integral over a region, so the convergence that has physical content is the one in (4.3.52). A claim about the value of a wavefunction at a point is not a claim L2L^{2} can support, and the Gibbs ear is what that looks like when you insist on making one.

8.5 · Position and momentum are one space, joined by a unitary map

The last debt of the chapter is Chapter 0.9's closing brick, which listed where its results would be spent and sent the first entry here: "The Fourier basis and Plancherel → Chapter 4.3 (where completeness is finally proved and the position and momentum representations become two bases for one Hilbert space)."

Assemble what is now available, and it is one thing rather than two. Chapter 0.9 §2.3 proved Plancherel, that the transform f~(k)=12πfeikxdx\tilde f(k)=\frac{1}{\sqrt{2\pi}}\int f\,\ee^{-\ii kx}\dd x preserves the norm, and read it back as the statement that the transform is unitary. That is the whole of what is needed here, and it is a statement about this one map. Section 7.4's result, that any two separable infinite-dimensional Hilbert spaces are unitarily equivalent, singles out no map at all and does no work in this section.

So what Plancherel says is that ψ\psi and ψ~\tilde\psi are not two objects but two descriptions of one. Square-integrable functions of xx and square-integrable functions of kk are the same Hilbert space, and the transform is a unitary map between the two descriptions. Chapter 0.9's phrase two bases for one Hilbert space is the loose reading of that, and the box below says exactly how loose.

Writing it with \hbar in place, as Part IV does everywhere, the momentum-space wavefunction is ψ~\tilde\psi rescaled so that its own integral is one, with the \sqrt\hbar visible:

ψ^(p)  =  1ψ~ ⁣(p)  =  12πψ(x)eipx/dx,ψ^(p)2dp  =  ψ(x)2dx. \begin{aligned} \hat\psi(p) \;&=\; \frac{1}{\sqrt{\hbar}}\,\tilde\psi\!\left(\frac{p}{\hbar}\right) \;=\; \frac{1}{\sqrt{2\pi\hbar}}\int_{-\infty}^{\infty}\psi(x)\,\ee^{-\ii px/\hbar}\,\dd x,\\[6pt] \int\abs{\hat\psi(p)}^{2}\dd p \;&=\; \int\abs{\psi(x)}^{2}\dd x . \end{aligned} (4.3.54)

The normalisations agree because the change of variable p=kp=\hbar k contributes exactly the \hbar that the prefactor removes, which is the reason the symmetric convention of Chapter 0.9 §2.2 was chosen there. So a state has one norm, and the Born rule computes the same total probability in either description.

⚠ What this does not establish, said before you assume it

It is tempting to read (4.3.54) as saying that {x}\{\ket x\} and {p}\{\ket p\} are two orthonormal bases in the sense of §7.3. They are not, and nothing above claims they are.

Section 7.4 proved that every orthonormal basis of this space is countable, and the positions on a line are not countable. The functions eikx\ee^{\ii kx} are not in L2(R)L^{2}(\R) at all, since eikx2=1\abs{\ee^{\ii kx}}^{2}=1 has infinite integral, which Chapter 0.9 §5.3 had already noticed and called a real gap. And a state concentrated at one point is the zero vector, by §5.3 here.

What is established is exactly this: L2(R)L^{2}(\R) in the variable xx and L2(R)L^{2}(\R) in the variable pp are the same Hilbert space, and (4.3.54) is a unitary map between the two descriptions. That is enough for every computation in Chapters 4.6 to 4.10, and it is what physicists mean when they say the two representations are equivalent. Chapter 4.5 closes the specific gap Chapter 0.9 named, by giving x\ket x and p\ket p a precise meaning through box normalisation and a limit that always works. The general theory of objects like them is Chapter 5.4's, and this chapter claims neither.

In plain terms 4.3.8

Nine chapters ago the pure waves were shown to be mutually perpendicular, and it was said plainly that being perpendicular is not the same as being numerous enough, and that the second claim was being borrowed. This section returns it.

The route has two halves. First, any state in the space can be approached by a continuous function, which re-uses the one quoted clause about squeezing sets between open and closed ones. That clause was spent a section earlier, to show the space has a countable basis at all, and this is the second job it does. Second, any continuous repeating function can be approached by a finite combination of pure waves, and the argument for that is constructive rather than abstract: average the partial sums instead of taking them, and the resulting weight becomes a positive bump of unit area narrowing onto a point. That bump is the same device used earlier to make sense of an idealised spike, reused without alteration.

Putting the halves together, the waves reach everything, so they are a basis, and the borrowing is repaid.

What remains is a warning that is also the most interesting fact in the section. The sense in which a series of smooth waves reproduces a function with a jump is not the sense you would guess. Measured by total squared discrepancy the approximation improves without limit, falling in proportion to the inverse square root of the number of terms. Measured by the worst error near the jump it does not improve at all: every partial sum overshoots by about nine per cent of the step, forever. Both are true, because the overshoot keeps its height and loses its width, and an integral cannot see a fixed height on a vanishing width. Which of the two is the physically meaningful one is not a matter of taste. Every measurable prediction is an integral over a region, so the first is the one that counts, and the second is what happens to anyone who asks the theory for the value of a wavefunction at a point.

9 · Worked examples

Worked example 1 — one family of spikes, and the exact reach of each theorem

On (0,1)(0,1) let fn=nα1(0,1/n)f_n=n^{\alpha}\mathbf 1_{(0,1/n)}, with α\alpha a real parameter. (a) Find the pointwise limit. (b) For which α\alpha does fnlimfn\int f_n\to\int\lim f_n? (c) For which α\alpha does fn0f_n\to0 in L2L^{2}? (d) Find the smallest possible dominating function and say for which α\alpha it is integrable. (e) Compare (b) and (d) and say what the comparison establishes about dominated convergence.

(a) Fix x(0,1)x\in(0,1). As soon as n>1/xn\gt1/x the point xx is outside the interval where fnf_n lives, so fn(x)=0f_n(x)=0 from then on. The pointwise limit is 00 at every point of (0,1)(0,1), for every α\alpha whatever. Notice that the parameter has not appeared yet, so any dependence on α\alpha below is a statement about the integral rather than about the functions.

(b) The integral is height times width:

01fn  =  nα1n  =  nα1, \int_0^1 f_n \;=\; n^{\alpha}\cdot\frac1n \;=\; n^{\alpha-1},

which tends to 00 when α<1\alpha\lt1, sits at 11 for every nn when α=1\alpha=1, and diverges when α>1\alpha\gt1. Since limfn=0\int\lim f_n=0, the interchange is valid exactly for α<1\alpha\lt1.

(c) Squaring doubles the exponent and leaves the width alone:

fn22  =  n2α1n  =  n2α1    0α<12. \lVert f_n\rVert_2^{\,2} \;=\; n^{2\alpha}\cdot\frac1n \;=\; n^{2\alpha-1} \;\longrightarrow\;0 \quad\Longleftrightarrow\quad \alpha\lt\tfrac12.

So convergence in L2L^{2} is a strictly stronger demand than convergence of the integrals. In the window 12α<1\tfrac12\le\alpha\lt1 the integrals converge to the right answer while the L2L^{2} norms do not shrink at all, which is worth holding on to: a sequence of states can have its expectation values settle without the states themselves converging.

(d) Any dominating function must exceed every member of the family, so the smallest one is F(x)=supnfn(x)F(x)=\sup_n f_n(x). For x(0,1)x\in(0,1) the functions that are non-zero at xx are those with n<1/xn\lt1/x, and for α>0\alpha\gt0 the largest of them wins, so F(x)F(x) is nαn^{\alpha} for the largest integer nn below 1/x1/x, which is xαx^{-\alpha} up to a bounded factor. Then

01xαdx  =  11αfor α<1,01xαdx  =  for α1. \int_0^1 x^{-\alpha}\,\dd x \;=\; \frac{1}{1-\alpha} \quad\text{for }\alpha\lt1, \qquad \int_0^1 x^{-\alpha}\,\dd x \;=\;\infty \quad\text{for }\alpha\ge1 .

So an integrable dominating function exists exactly when α<1\alpha\lt1, and for α0\alpha\le0 the family is bounded by a constant and the question is trivial.

(e) Compare the two answers. The interchange in (b) is valid for α<1\alpha\lt1. A dominating function exists in (d) for α<1\alpha\lt1. The two conditions coincide, so on this family dominated convergence is not merely sufficient but sharp. There is no α\alpha for which the conclusion holds and the hypothesis fails, and none for which the hypothesis holds and the conclusion fails. That is unusual, and it is the reason this family is the right one to keep in mind: when you cannot find a dominating function it is worth suspecting that the interchange is not merely unproved but false.

Two further readings. Applying dominated convergence to fn2\abs{f_n}^{2} instead needs supnfn2x2α\sup_n f_n^{2}\approx x^{-2\alpha} to be integrable, which asks α<12\alpha\lt\tfrac12 and matches (c) exactly, so the same sharpness holds for the L2L^{2} statement. And Fatou's lemma is consistent throughout, reading 0=lim inffnlim inffn=limnα10=\int\liminf f_n\le\liminf\int f_n=\lim n^{\alpha-1}, with equality for α<1\alpha\lt1 and strict inequality for α1\alpha\ge1, which is mass escaping up the spike.

Worked example 2 — Parseval, twice, and two sums Part IV has already used

(a) Apply Parseval to Chapter 0.9's square wave and evaluate n oddn2\sum_{n\ \mathrm{odd}}n^{-2}, then n1n2\sum_{n\ge1}n^{-2}. (b) Apply it to f(x)=x2f(x)=x^{2} on [π,π][-\pi,\pi] and evaluate n1n4\sum_{n\ge1}n^{-4}. (c) Use (b) to finish Chapter 4.1's integral (4.3.23). (d) Use (a) to derive the exact L2L^{2} error of the square wave's partial sums and confirm the rate of (4.3.52).

(a) The square wave has f22=ππ1dx=2π\lVert f\rVert_2^{2}=\int_{-\pi}^{\pi}1\,\dd x=2\pi, and by (4.3.50) its expansion in the orthogonal set {sinnx}\{\sin nx\}, each of squared norm π\pi, has coefficients 4/πn4/\pi n for odd nn and zero otherwise. Parseval, which §8.3 turned into a theorem, equates the two:

2π  =  n odd(4πn)2π  =  16πn odd1n2n odd1n2  =  π28. 2\pi \;=\; \sum_{n\ \mathrm{odd}}\left(\frac{4}{\pi n}\right)^{2}\pi \;=\; \frac{16}{\pi}\sum_{n\ \mathrm{odd}}\frac{1}{n^{2}} \qquad\Longrightarrow\qquad \sum_{n\ \mathrm{odd}}\frac{1}{n^{2}} \;=\; \frac{\pi^{2}}{8}.

The full sum follows by splitting off the even terms, which are the full sum again with every nn doubled. Writing S=n1n2S=\sum_{n\ge1}n^{-2}, the even part is (2m)2=S/4\sum(2m)^{-2}=S/4, so S=π2/8+S/4S=\pi^{2}/8+S/4 and S=π2/6S=\pi^{2}/6. That is the Basel sum, obtained here as a by-product of a square wave.

(b) Take f(x)=x2f(x)=x^{2} on [π,π][-\pi,\pi] and use the complex form (4.3.43). The n=0n=0 coefficient is the mean, a0=12πx2=π2/3a_0=\frac{1}{2\pi}\int x^{2}=\pi^{2}/3. For n0n\ne0, two integrations by parts give ππx2cosnxdx=4π(1)n/n2\int_{-\pi}^{\pi}x^{2}\cos nx\,\dd x=4\pi(-1)^{n}/n^{2}, and the sine part vanishes by oddness, so an=2(1)n/n2a_n=2(-1)^{n}/n^{2}. Parseval in the form (4.3.49) reads

12πππx4dx  =  π45  =  π49  +  2n14n4, \frac{1}{2\pi}\int_{-\pi}^{\pi}x^{4}\,\dd x \;=\; \frac{\pi^{4}}{5} \;=\; \frac{\pi^{4}}{9} \;+\; 2\sum_{n\ge1}\frac{4}{n^{4}},

the factor 22 counting nn and n-n together. Rearranging, 8n4=π4(1519)=445π48\sum n^{-4}=\pi^{4}(\tfrac15-\tfrac19)=\tfrac{4}{45}\pi^{4}, so n1n4=π4/90\sum_{n\ge1}n^{-4}=\pi^{4}/90.

(c) Section 4.1 used monotone convergence to write Chapter 4.1's integral as 6n1n46\sum_{n\ge1}n^{-4} and stopped there. With (b) in hand,

0y3dyey1  =  6π490  =  π415  =  6.49394, \int_{0}^{\infty}\frac{y^{3}\,\dd y}{\ee^{y}-1} \;=\; 6\cdot\frac{\pi^{4}}{90} \;=\; \frac{\pi^{4}}{15} \;=\; 6.49394,

which is the number the Stefan–Boltzmann constant of Chapter 4.1 was computed from. Both steps of that computation are now theorems: the term-by-term integration is monotone convergence and the resulting sum is Parseval.

(d) The error of a partial sum is the tail of Parseval's series, so from (a),

fSN22  =  16πn>Nn odd1n2  =  2π16πnNn odd1n2, \lVert f-S_N\rVert_2^{\,2} \;=\; \frac{16}{\pi}\sum_{\substack{n\gt N\\ n\ \mathrm{odd}}}\frac{1}{n^{2}} \;=\; 2\pi-\frac{16}{\pi}\sum_{\substack{n\le N\\ n\ \mathrm{odd}}}\frac{1}{n^{2}},

which is exact at every NN and is the quantity the figure plots. Dividing by f22=2π\lVert f\rVert_2^{2}=2\pi and taking the root gives the relative error

E(N)  =  (18π2nNn odd1n2)1/2. E(N) \;=\; \left(1-\frac{8}{\pi^{2}}\sum_{\substack{n\le N\\ n\ \mathrm{odd}}}\frac{1}{n^{2}}\right)^{1/2}.

Now the asymptotics, and the estimate is worth making carefully because the approach turns out to be slow. The odd integers beyond NN are spaced by two and NN itself is odd, so they sit at the midpoints of the length-two intervals tiling (N+1,)(N+1,\infty). Their reciprocal squares therefore sum to 12N+1x2dx=12(N+1)\tfrac12\int_{N+1}^{\infty}x^{-2}\dd x=\tfrac{1}{2(N+1)}, with a midpoint-rule error of order N3N^{-3}. So

E(N)2  =  4π2(N+1)  +  O(N3),E(N)N  =  2πNN+1  +  O(N2). \begin{aligned} E(N)^{2} \;&=\; \frac{4}{\pi^{2}(N+1)} \;+\; O(N^{-3}),\\[4pt] E(N)\,\sqrt N \;&=\; \frac{2}{\pi}\sqrt{\frac{N}{N+1}} \;+\; O(N^{-2}). \end{aligned}

The rate is 2/πN1/22/\pi\cdot N^{-1/2}, which is (4.3.52), and the deficit factor N/(N+1)\sqrt{N/(N+1)} in front of it is why the constant arrives so slowly. That factor is 0.995090.99509 at N=101N=101, so the product is still half a per cent short of 2/π2/\pi there, and it is 0.999500.99950 at N=1001N=1001.

The figure's first readout computes exactly E(N)E(N) and E(N)NE(N)\sqrt N, so both numbers can be checked by moving the slider. At N=101N=101 it reads E=0.063034E=0.063034 and EN=0.633481E\sqrt N=0.633481. At N=1001N=1001 it reads E=0.020112E=0.020112 and EN=0.636302E\sqrt N=0.636302. Against 2/π=0.6366202/\pi=0.636620, the rate is confirmed and the constant is confirmed with it, with the last three digits still on their way.

Worked example 3 — a normalised state whose mean position does not exist

Take ψ(x)=a/π(a2+x2)1/2\psi(x)=\sqrt{a/\pi}\,(a^{2}+x^{2})^{-1/2}, with aa a length. (a) Check that ψ\psi is normalised and identify ψ2\abs\psi^{2}. (b) Decide whether ψ\psi is in L1(R)L^{1}(\R). (c) Decide whether x\avg x exists. (d) Say what is wrong with the argument that x=0\avg x=0 by symmetry. (e) Say where such a state comes from physically.

(a) The squared modulus is ψ(x)2=aπ(a2+x2)\abs{\psi(x)}^{2}=\dfrac{a}{\pi(a^{2}+x^{2})}, whose integral is

adxπ(a2+x2)  =  1π[arctanxa]  =  1π(π2+π2)  =  1. \int_{-\infty}^{\infty}\frac{a\,\dd x}{\pi(a^{2}+x^{2})} \;=\; \frac{1}{\pi}\Big[\arctan\frac{x}{a}\Big]_{-\infty}^{\infty} \;=\; \frac{1}{\pi}\left(\frac\pi2+\frac\pi2\right) \;=\;1 .

So ψ\psi is a perfectly good normalised state, and ψ2\abs\psi^{2} is the Lorentzian of Chapter 0.8 §6.4, which Chapter 0.9's grind box on the central limit theorem identified as the Cauchy density.

(b) No. For large x\abs x the function ψ\psi itself falls only as 1/x1/\abs x, so

ψdx  =  aπdxa2+x2  =  , \int_{-\infty}^{\infty}\abs{\psi}\,\dd x \;=\; \sqrt{\frac a\pi}\int\frac{\dd x}{\sqrt{a^{2}+x^{2}}} \;=\;\infty,

a logarithmic divergence at both ends. This is §5.5's first bullet in the flesh: ψL2\psi\in L^{2} and ψL1\psi\notin L^{1}, and there is nothing pathological about the state that produced it.

(c) No. By §3.3 a function is Lebesgue integrable only when the integral of its modulus is finite, and

xψ(x)2dx  =  2aπ0xdxa2+x2  =  aπ[ln(a2+x2)]0  =  . \int_{-\infty}^{\infty}\abs{x}\,\abs{\psi(x)}^{2}\,\dd x \;=\; \frac{2a}{\pi}\int_0^{\infty}\frac{x\,\dd x}{a^{2}+x^{2}} \;=\; \frac{a}{\pi}\Big[\ln(a^{2}+x^{2})\Big]_0^{\infty} \;=\;\infty .

So xψ2x\abs\psi^{2} is not integrable and x\avg x does not exist. The same computation with x2x^{2} diverges worse, so x2\avg{x^{2}} does not exist either, and neither does the position uncertainty. The state is normalised and has no mean position.

(d) The symmetry argument computes something, and the something is not the integral. What it computes is

limRRRxψ(x)2dx  =  0, \lim_{R\to\infty}\int_{-R}^{R}x\,\abs{\psi(x)}^{2}\,\dd x \;=\; 0,

which is a limit of proper integrals over a particular family of symmetric windows. Section 3.6 is exactly about this. The Lebesgue integral sorts values before summing them and can never depend on the order in which cancellations arrive, so a quantity defined by the order of arrival is not one of its values.

To see how completely the answer belongs to the family rather than to the state, move the right edge. The antiderivative is a logarithm, so for any fixed c>0c\gt0,

RcRaxdxπ(a2+x2)  =  a2πlna2+c2R2a2+R2  R  aπlnc. \int_{-R}^{cR}\frac{a\,x\,\dd x}{\pi(a^{2}+x^{2})} \;=\; \frac{a}{2\pi}\ln\frac{a^{2}+c^{2}R^{2}}{a^{2}+R^{2}} \;\xrightarrow[R\to\infty]{}\; \frac{a}{\pi}\ln c .

Every one of those limits is finite and no two of them agree. Taking c=1c=1 recovers the 00 the symmetry argument produced, and taking c=2c=2 gives aln2/πa\ln2/\pi, which is 0.2206360.220636 for a=1a=1. So the family does not merely disturb the answer. It sets it.

Divergence needs more than a wider window: it needs the right edge to outrun the left one faster than any fixed ratio. The family [R,R2][-R,R^{2}] does that, and the same antiderivative gives a2πln[(a2+R4)/(a2+R2)]\tfrac{a}{2\pi}\ln\big[(a^{2}+R^{4})/(a^{2}+R^{2})\big], which grows like aπlnR\tfrac a\pi\ln R and has no limit. The number 00 is a property of the window family, not of the state, and no family is the one the state prefers.

(e) From any resonance. Chapter 0.8 §6.4 derived the Lorentzian as the response of a damped oscillator near its natural frequency and identified its width as the decay rate, and Chapter 0.9 noted that Cauchy tails "arise from ratios of normal variables, from resonance line shapes… and in any setting where the largest single contribution is comparable to the sum of the rest." An unstable state has a Lorentzian energy distribution, so its mean energy does not exist as a Lebesgue integral, and every quoted "mean energy" of such a state is a symmetric window in disguise. That is not a defect in the physics. It is a warning that the quoted number depends on a truncation, and it is exactly the kind of thing the rebuilt integral was designed to make visible.

10 · Your turn

Problem 1 — a space with a hole in it, built by hand

(a) Show that a Cauchy sequence with a convergent subsequence converges, and say where the triangle inequality enters. (b) On [0,1][0,1] let hnh_n be the continuous function that is 00 for x1212nx\le\half-\frac{1}{2n}, rises linearly to 11 across the interval of width 1/n1/n centred at 12\half, and is 11 thereafter. Show that (hn)(h_n) is Cauchy in the L2L^{2} norm and identify its L2L^{2} limit. (c) Show that no continuous function is that limit, so the continuous functions with the L2L^{2} norm form an incomplete space. (d) Show that (hn)(h_n) is not Cauchy in the supremum norm, and say why that does not rescue anything. (e) Which of §7.3's four statements could fail in an incomplete inner-product space, and at which step of the proof?

Solution

(a) Let vnkvv_{n_k}\to v and let ε>0\varepsilon\gt0. Choose MM with vnvm<ε/2\lVert v_n-v_m\rVert\lt\varepsilon/2 for n,mMn,m\ge M, then choose kk with nkMn_k\ge M and vnkv<ε/2\lVert v_{n_k}-v\rVert\lt\varepsilon/2. For any nMn\ge M the triangle inequality gives vnvvnvnk+vnkv<ε\lVert v_n-v\rVert\le\lVert v_n-v_{n_k}\rVert+\lVert v_{n_k}-v\rVert\lt\varepsilon. The triangle inequality is what lets the subsequence act as a stepping stone, and it is the only tool used.

(b) For m<nm\lt n the two functions agree outside the wider ramp, which has width 1/m1/m, and differ by at most 11 inside it. So hnhm221/m\lVert h_n-h_m\rVert_2^{2}\le1/m, which goes to zero, so the sequence is Cauchy. The same estimate against the step function 1[1/2,1]\mathbf 1_{[1/2,1]} gives hn1[1/2,1]221/n0\lVert h_n-\mathbf 1_{[1/2,1]}\rVert_2^{2}\le1/n\to0, so the L2L^{2} limit is that step function.

(c) Suppose gg is continuous with g=1[1/2,1]g=\mathbf 1_{[1/2,1]} almost everywhere. On (0,12)(0,\half) the two agree off a null set, and a continuous function agreeing with 00 off a null set is identically 00 there, since otherwise it would be non-zero on a whole interval, which has positive measure. So g0g\equiv0 on (0,12)(0,\half) and, by the same argument, g1g\equiv1 on (12,1)(\half,1). Continuity at 12\half then demands 0=10=1. So the sequence is Cauchy in C[0,1]C[0,1] and converges to nothing in it.

(d) At x=12+12nx=\half+\frac{1}{2n} we have hn=1h_n=1 while hm=12+m2nh_m=\half+\frac{m}{2n}, so hnhm12m2n\lVert h_n-h_m\rVert_\infty\ge\half-\frac{m}{2n}, which tends to 12\half as nn grows with mm fixed. So the sequence is not Cauchy in the supremum norm. That does not rescue the space, because completeness must hold for the norm you are actually using, and the norm quantum mechanics uses is the L2L^{2} one, since it is the one that comes from an inner product and therefore carries the geometry.

(e) Statement (b), the expansion, is the one at risk, and the step is §7.2. There the partial sums were shown to be Cauchy and completeness was invoked to give them a limit. In an incomplete space the coefficients would still be square-summable by Bessel and the series would still be Cauchy, and it would converge to nothing. Statements (a), (c) and (d) can all still be stated, and the chain of implications breaks precisely at (d) implies (b).

Problem 2 — thin and small are different properties

(a) The middle-thirds Cantor set is what remains of [0,1][0,1] after removing the open middle third of every interval at every stage, forever. Show that it has measure zero. (b) Now remove, at stage kk, an open interval of length 4k4^{-k} from the centre of each of the 2k12^{k-1} intervals then present. Show that what remains is closed, contains no interval, and has measure 12\half. (c) Conclude that "contains no interval" and "has measure zero" are independent properties, and say which of the two Chapter 0.2's Riemann integral is sensitive to. (d) Show that the indicator of the set in (b) is not Riemann integrable, by the argument of §1.4 run in the other direction.

Solution

(a) At stage kk there are 2k2^{k} intervals each of length 3k3^{-k}, so what remains after kk stages has measure (2/3)k(2/3)^{k}. The Cantor set is contained in every one of those, so its measure is at most (2/3)k(2/3)^{k} for every kk, hence zero. It is nevertheless uncountable, since the ternary expansions using only the digits 00 and 22 are in bijection with the binary expansions of [0,1][0,1], so "null" is much weaker than "countable".

(b) The set is an intersection of closed sets, hence closed. It contains no interval, because after kk stages every remaining interval has length below 2k2^{-k}, so any interval inside the set would have to have length zero. The total length removed is

k=12k14k  =  12k=12k  =  12, \sum_{k=1}^{\infty} 2^{k-1}\cdot 4^{-k} \;=\; \frac12\sum_{k=1}^{\infty}2^{-k} \;=\; \frac12,

so the measure of what remains is 112=121-\half=\half. This is called a fat Cantor set.

(c) The set in (a) contains no interval and has measure zero. The set in (b) contains no interval and has measure 12\half. An interval contains an interval and has positive measure. So the two properties are independent. Upper and lower Riemann sums see only whether a set is dense in a cell or absent from it, so the Riemann integral is sensitive to the topological property and blind to the measure-theoretic one, which is precisely the defect §1.4 exploited.

(d) Let KK be the fat Cantor set and take any partition. Every cell contains points outside KK, since KK contains no interval, so inf1K=0\inf\mathbf 1_K=0 on every cell and every lower sum is 00. The cells that meet KK have sup1K=1\sup\mathbf 1_K=1 and their union covers KK, so their total length is at least μ(K)=12\mu(K)=\half and every upper sum is at least 12\half. The gap is at least 12\half at every mesh, so 1K\mathbf 1_K is not Riemann integrable. Its Lebesgue integral is μ(K)=12\mu(K)=\half, by (4.3.15) in one line.

Problem 3 — the three ways mass escapes

On R\R define fn=n1(0,1/n)f_n=n\,\mathbf 1_{(0,1/n)}, then gn=1n1(0,n)g_n=\tfrac1n\mathbf 1_{(0,n)}, then hn=1(n,n+1)h_n=\mathbf 1_{(n,n+1)}. (a) Show that all three tend to zero at every point and that all three have integral 11. (b) Which of them converge to zero in L2(R)L^{2}(\R)? (c) For each, compute supn\sup_n of the family and show it is not integrable. (d) State what Fatou's lemma says about each, and check it. (e) Say in one sentence each what physical process the three pictures correspond to.

Solution

(a) For fixed xx, the spike fnf_n misses xx once n>1/xn\gt1/x, the plateau gng_n has height 1/n01/n\to0, and the travelling bump hnh_n has left xx behind once n>xn\gt x. Each integral is height times width, which is n1nn\cdot\frac1n, then 1nn\frac1n\cdot n, then 111\cdot1.

(b) Squaring the height and keeping the width, fn22=n\lVert f_n\rVert_2^{2}=n, which diverges, gn22=1/n\lVert g_n\rVert_2^{2}=1/n, which tends to zero, and hn22=1\lVert h_n\rVert_2^{2}=1, which does not move. So only the plateau converges in L2L^{2}, and it does so while its integral stays pinned at 11. Convergence in L2L^{2} and convergence of integrals are independent conditions on R\R, in both directions.

(c) For the spike, the largest nn with 1/n>x1/n\gt x gives supnfn(x)\sup_n f_n(x) comparable to 1/x1/x on (0,1)(0,1), whose integral diverges at the origin. For the plateau, supngn(x)\sup_n g_n(x) is comparable to 1/x1/x for large xx, whose integral diverges at infinity. For the travelling bump, supnhn=1(1,)\sup_n h_n=\mathbf 1_{(1,\infty)}, whose integral is infinite. In each case the smallest candidate ceiling has infinite integral, so no dominating function exists and dominated convergence does not apply. It could not, since the conclusion is false in all three.

(d) Fatou says lim inflim inf\int\liminf\le\liminf\int, which reads 010\le1 in all three cases. It holds, strictly. That is the content of the inequality: mass may disappear in the limit but may not appear, and each of these loses exactly one unit of it.

(e) The spike is a pulse of fixed energy compressed into a shorter and shorter time, which is Chapter 0.9's delta being born. The plateau is a fixed quantity spread ever more thinly, which is a wavepacket dispersing. The travelling bump is a fixed quantity leaving the region of interest without changing shape, which is a particle escaping to infinity. Chapter 4.6 has to handle exactly this for a free packet, and Chapter 4.7 for the unbound states of a genuine potential.

Problem 4 — an orthonormal set that is not a basis

(a) Prove Bessel's inequality by expanding fnNcnen2\lVert f-\sum_{n\le N}c_ne_n\rVert^{2} with cn=en,fc_n=\avg{e_n,f}, and say which axiom of Chapter 0.5 §1 makes the left-hand side non-negative. (b) On [π,π][-\pi,\pi] show that un(x)=e2inx/2πu_n(x)=\ee^{2\ii nx}/\sqrt{2\pi}, for nn ranging over the integers, is an orthonormal set. (c) Exhibit a non-zero fL2[π,π]f\in L^{2}[-\pi,\pi] orthogonal to every unu_n, and conclude that the set is not a basis. (d) Compute both sides of Parseval for your ff and say by how much the inequality of (a) is strict. (e) Identify the subspace of which {un}\{u_n\} is a basis.

Solution

(a) Expanding, and using em,en=δmn\avg{e_m,e_n}=\delta_{mn} to collapse the double sum,

0    fnNcnen2=f2nNcˉncnnNcncˉn+nNcn2=f2nNcn2. 0\;\le\;\Big\lVert f-\sum_{n\le N}c_ne_n\Big\rVert^{2} = \lVert f\rVert^{2}-\sum_{n\le N}\bar c_nc_n-\sum_{n\le N}c_n\bar c_n+\sum_{n\le N}\abs{c_n}^{2} = \lVert f\rVert^{2}-\sum_{n\le N}\abs{c_n}^{2}.

So the partial sums of cn2\sum\abs{c_n}^{2} are bounded by f2\lVert f\rVert^{2}, and being increasing they converge to something no larger. The left-hand side is non-negative by axiom (iii), positive definiteness, which is the axiom §5.3 had to repair by a quotient before any of this was available for functions.

(b) The inner product is 12πππe2i(nm)xdx\frac{1}{2\pi}\int_{-\pi}^{\pi}\ee^{2\ii(n-m)x}\dd x, which is 11 when n=mn=m and otherwise vanishes, since 2(nm)2(n-m) is a non-zero integer and Chapter 0.9's computation (4.3.43) applies unchanged.

(c) Take f(x)=eixf(x)=\ee^{\ii x}. Then un,f=12πππei(12n)xdx=0\avg{u_n,f}=\frac{1}{\sqrt{2\pi}}\int_{-\pi}^{\pi}\ee^{\ii(1-2n)x}\dd x=0 for every nn, because 12n1-2n is odd and therefore never zero. Yet f2=2π0\lVert f\rVert^{2}=2\pi\ne0. So statement (d) of §7.3 fails, and by the equivalence proved there all four fail.

(d) Parseval would say 2π=nun,f2=02\pi=\sum_n\abs{\avg{u_n,f}}^{2}=0. The inequality of (a) is therefore strict by the entire norm of ff. Bessel is not close to being an equality here, which is worth noticing: an orthonormal set can miss a whole direction rather than a little of one.

(e) Each unu_n satisfies un(x+π)=un(x)u_n(x+\pi)=u_n(x), so every finite combination does, and so does every L2L^{2} limit of such combinations. Conversely a π\pi-periodic function on [π,π][-\pi,\pi] is determined by its restriction to an interval of length π\pi, on which the functions e2inx\ee^{2\ii nx} are the full Fourier system rescaled and therefore a basis by §8.3. So {un}\{u_n\} is an orthonormal basis of the subspace of π\pi-periodic functions, which is a proper closed subspace of L2[π,π]L^{2}[-\pi,\pi].

Problem 5 — the two convergences, quantitatively

Work with the square wave and its partial sums SNS_N throughout. (a) Using n oddn2=π2/8\sum_{n\ \mathrm{odd}}n^{-2}=\pi^{2}/8, write fSN22\lVert f-S_N\rVert_2^{2} exactly, and recover fSN2/f2(2/π)N1/2\lVert f-S_N\rVert_2/\lVert f\rVert_2\approx(2/\pi)N^{-1/2}. (b) The first peak of SNS_N sits at x=π/(N+1)x^{*}=\pi/(N+1) and overshoots by a fixed amount. Show that the contribution of the window x<3x\abs x\lt3x^{*} to fSN2\int\abs{f-S_N}^{2} tends to zero, and say which factor is responsible. (c) Section 8.4's theorem says SN(x)f(x)S_N(x)\to f(x) at every x0x\ne0. Explain in one paragraph why that does not contradict an overshoot that never shrinks. (d) A detector records a band-limited version of a step, keeping harmonics up to NN. State what it measures in each of the two senses, and which of the two a physical measurement returns.

Solution

(a) Parseval gives fSN22=16πn>N, oddn2=2π16πnN, oddn2\lVert f-S_N\rVert_2^{2}=\frac{16}{\pi}\sum_{n\gt N,\ \mathrm{odd}}n^{-2} =2\pi-\frac{16}{\pi}\sum_{n\le N,\ \mathrm{odd}}n^{-2}, using the total π2/8\pi^{2}/8. The tail is a sum over odd integers past NN, spaced by two, so it is about 12Nx2dx=1/2N\tfrac12\int_N^\infty x^{-2}\dd x=1/2N, whence fSN228/πN\lVert f-S_N\rVert_2^{2}\approx8/\pi N. Dividing by f22=2π\lVert f\rVert_2^{2}=2\pi and taking the square root gives 2/πN2/\pi\sqrt N.

(b) On that window fSN\abs{f-S_N} is bounded by a constant, since the partial sums themselves are bounded by about 1.181.18 near the jump and ff is bounded by 11. So the contribution is at most a constant times the width of the window, which is 6π/(N+1)6\pi/(N+1) and tends to zero. The height is not responsible for anything, because it does not change. The width is. Comparing with (a), the window contributes at the same rate 1/N1/N as the whole error, so the ear is not a negligible part of the error. It is a fixed fraction of a total that is itself vanishing.

(c) Fix any x0x\ne0. The overshoot happens at x=π/(N+1)x^{*}=\pi/(N+1), which tends to zero, so for NN large enough the peak lies strictly between 00 and xx, and the value at xx itself is converging to f(x)f(x). Every point is eventually to the right of the trouble. What never happens is that the trouble goes away, because for every NN there is some point at which the error is about 0.1790.179. The maximum over xx is taken at a moving location, and pointwise convergence says nothing about a quantity evaluated at a moving location. Convergence is pointwise but not uniform, and the Gibbs overshoot is the exact measure of the gap between those two words.

(d) In the L2L^{2} sense it measures the step to within a relative error 2/πN2/\pi\sqrt N, which for a hundred harmonics is about six per cent and falls as more are kept. In the pointwise sense it rings at the edge by 8.95%8.95\% of the step no matter how many harmonics are kept, which is the ringing Chapter 0.9 §1.4 said every lens, detector and digital filter produces. A physical measurement integrates over a region, a time window or a pixel, so it returns the first. The second is what the record looks like if you plot it, and it is why sharp edges and finite bandwidth are incompatible, a statement Chapter 0.9 §6 turned into a theorem.

The brick you just laid — the space quantum mechanics lives in

The Riemann integral was discarded for a reason, and the reason was exhibited rather than asserted. Chapter 0.2's χ\chi is not integrable, which alone proves nothing. Chapter 0.2's own sequence fnf_n is increasing, has every integral equal to zero, and has a limit outside the class, so no convergence theorem is available. And thickening that same sequence, by giving the kk-th rational an interval of length 2k22^{-k-2}, produces a Cauchy sequence of step functions whose limit is at a positive distance from every Riemann-integrable function on the interval. Density let a tagged sum at every mesh be pushed up to 11 and total length held the integral below 14\tfrac14, and those two facts cannot be reconciled by any function the old theory can handle.

What replaced it. Size, first: three demands, one of them countable, and one construction quoted. Then an integral that slices the range instead of the domain, defined on a class closed under countable limits by construction. Then the two theorems the whole rebuild existed for, one proved from continuity from below alone and the other from it by way of Fatou. Then the space: square-integrable functions with Chapter 0.5's inner product, two axioms holding for free and the third repaired by declaring functions equal almost everywhere to be one vector.

And then completeness, which was the point. Every Cauchy sequence in L2L^{2} converges in L2L^{2}, so the space earns the name Hilbert space that Chapter 0.9 promised it. The two sequences of §1 were handed back with limits attached. Chapter 0.2's converges to the zero vector, with 01χ=0\int_0^1\chi=0 as that chapter said it would. The thickened one converges to a genuinely new vector, and §1.4's grind box is the proof that it is new rather than a relabelling.

Bases, and the Fourier modes. Four statements of completeness for an orthonormal set were proved equivalent, with completeness of the space entering at exactly one step. Every orthonormal set here is countable, because two of them are always 2\sqrt2 apart and a countable dense set cannot serve uncountably many disjoint balls, so "expand in a basis" means an ordinary series. Fejér's theorem then supplies the density of trigonometric polynomials constructively, out of a positive kernel of unit integral narrowing at the origin, which is Chapter 0.9 §5.2's argument reused without alteration. Chapter 0.9 §1.3's mark is paid, and it stays on the page where it was raised, naming this chapter as the payer.

The number this chapter owed. For the square wave, the relative L2L^{2} error of the partial sums is exactly (18π2nN, oddn2)1/2\big(1-\tfrac{8}{\pi^{2}}\sum_{n\le N,\ \mathrm{odd}}n^{-2}\big)^{1/2}, which falls like 0.63662N1/20.63662\,N^{-1/2} across three decades in the figure, while the overshoot at the jump sits at 0.08948990.0894899 of the jump and does not move. Both are measured in the figure rather than quoted, and the readout reports the two slopes as 12-\tfrac12 and 00. The reconciliation is that the ear keeps its height and loses its width.

Two marks, and that is the claim. The construction of Lebesgue measure in §2.3, quoted as a package with its five clauses written out. The proof of Riesz–Fischer in §6.2, whose shape was described and whose inputs are only §2.3 and §4. Nothing else in this chapter is asserted without being derived, and the Vitali set of §2.5 was built rather than mentioned, so the first mark is a real restriction rather than a formality. One clause of that first package, regularity, was deliberately cashed in §7.4 to make the space separable, and re-used in §8.1 to prove that continuous functions are dense. Neither result carries a separate mark of its own, precisely so that you can watch a flag do the work it claimed. This chapter also leans on one mark already standing elsewhere, Heine–Cantor from Chapter 0.2 §1.1, used once in §8.2 and cited there rather than re-raised. Two new marks for the chapter that rebuilds the integral, proves the space complete, and establishes the basis every later calculation expands in, is the honest count.

Where this gets spent. Chapter 4.4 takes the space built here and puts operators on it, which is the second instalment of the bill Chapter 4.2 named, and it needs completeness at every step for domains and for the difference between symmetric and self-adjoint. Chapter 4.5 is the third instalment, and needs it again for spectra with no eigenvectors in the space, and for the precise meaning of x\ket x and p\ket p that §8.5 declined to claim. Chapter 4.6 needs eiH^t/\ee^{-\ii\hat Ht/\hbar} to map states to states, which is completeness applied to a series. Chapter 4.5 also proves the Hermite functions complete, and that proof is statement (d) of §7.3 above run on an integral that only the dominated convergence of §4.3 licenses, after which Chapter 4.8 expands in them freely. Chapter 4.9 uses Cauchy–Schwarz in the form §5.4 transferred, and needs L2⊄L1L^{2}\not\subset L^{1} to explain why a normalised state can have no mean position. And every chapter of Part V expands a field in modes and integrates term by term, which is monotone convergence and Parseval, both proved here. The integral, the two convergence theorems, the space, and the basis: that is the whole of what Part IV and Part V spend, and none of it is on credit any longer.