Part 0 · The Toolkit — Chapter 0.5
Inner Products, Eigenvectors, and the Spectral Theorem
Quantum mechanics, four parts early, with the physics stripped off.
Here is the stake, stated as plainly as it can be stated. Every postulate of quantum mechanics is a statement about Hermitian operators on an inner-product space. Not modelled by. Not analogous to. Is. Measurement outcomes are eigenvalues. States are unit vectors. Time evolution is a unitary map. Compatible observables are commuting operators. Take the physics away and what is left is a chapter of linear algebra, and it is this one.
So we are going to prove that chapter of linear algebra now, four parts early, with no physics attached whatsoever. No wavefunctions, no , no measurement. By the end you will have proved, with no mention of physics anywhere in the argument, that a certain class of operators has only real eigenvalues, and that eigenvectors belonging to different eigenvalues are exactly perpendicular. Chapter 4.2 announces that measurement results are real numbers and that distinct outcomes are perfectly distinguishable. When it does, you will already have the proof. That makes Chapter 4.2 a translation exercise: a table of renamings, not a new subject.
That is the whole design of this book compressed into one chapter. The mathematics is not scaffolding erected to support the physics. It is the physics, wearing different words.
Tools you'll need — Chapter 0.1: nothing beyond fluency. Chapter 0.3: Euler's formula , and power series. Chapter 0.4: vector spaces, basis and dimension, linear maps and their matrices, change of basis, determinant, trace, and the fact that over every square matrix has at least one eigenvalue. Everything here is built on those.
1 · Inner products
A vector space as Chapter 0.4 left it is a bleak place. You can add vectors and scale them, and that is all. There is no length, no angle, no notion of two vectors pointing in different directions. Every geometric word you know is missing.
An inner product is the single extra structure that puts all of them back at once. Let's see what it has to look like, and why it has no freedom in the matter.
1.1 · Why the axioms look the way they do
You already have an inner product in , namely the dot product . Its essential property is that is the squared length: a non-negative number that vanishes only for the zero vector.
Now try to carry that same formula over to unchanged. It breaks immediately. Take in and add up the squares of the components:
and yet . A non-zero vector with zero length is useless to us. We would lose the ability to say that two states are different, and telling states apart is the one job a length has to do.
So the formula needs repair, and there is exactly one repair available. Replace by , which is non-negative for every complex number and zero only when . That forces the complex inner product to conjugate one of its two arguments. Everything below is bookkeeping around that single fact.
An inner product on a complex vector space is a map , written , satisfying three axioms.
Three remarks follow, in increasing order of importance.
Axiom (iii) is not quite well posed until (i) is used. Writing "" presumes that is real, and that is a consequence rather than an assumption. Set in (i) and you get . A complex number equal to its own conjugate is real. So the axioms are consistent, and we derived the fact rather than assuming it.
Linearity in one slot forces conjugate-linearity in the other. We only demanded linearity in the second argument. The first is then determined, and conjugate symmetry shows it in one line:
Pull a scalar out of the second slot and it comes out unchanged. Pull it out of the first and it comes out conjugated. This asymmetry is real and you have to keep track of it. It is also the reason the notation of §2 exists.
The convention. Which slot is the linear one is pure convention, and the two communities chose differently. Mathematicians make the first slot linear. Physicists make the second. We have chosen the physics convention above and will keep it for the whole book. The reason is Chapter 4.2, where becomes a probability amplitude, read right-to-left as "amplitude to find in state ". The linear slot had better be the one holding the state that is actually evolving. If you open a mathematics text and every formula appears to have its conjugates on the wrong side, this is why.
1.2 · The standard examples
| Space | Inner product | Where it turns up |
|---|---|---|
| Ordinary geometry; least squares (§3) | ||
| Spin, qubits, any finite quantum system (4.12) | ||
| , weighted | , | Weighted regression; fitting |
| Functions on | Fourier series (0.9); wavefunctions (4.3) | |
| matrices | Density matrices (4.19) |
The fourth row is the one to test, because it is the one that will matter most later. Check it against the axioms and you will find nothing new is needed. Conjugate symmetry holds because . Linearity in the second slot holds because integration is linear. Positive definiteness holds because , with equality only if .
That row is the entire reason Chapter 4.3 is possible. A space of functions is an inner-product space, and every theorem of this chapter will try to move there. Some of them survive the trip. Some need repair, and the warning at the end of §6 says which ones and why.
1.3 · Length, distance, and the angle we cannot yet define
An inner product is all we need in order to say how long a vector is and how far apart two vectors are. So define the norm and the distance:
Both are well defined. We showed above that is real and non-negative, so the square root is a real number. Call and orthogonal when . That relation is symmetric, since as well, so it does not matter which vector you name first.
Now let's see what the three axioms hand us for free. Expand the squared norm of a sum, using nothing but the axioms, and the generalised Pythagoras drops out:
The reason is that , and the middle two terms are conjugates of each other, so they add to twice the real part. When the cross term dies and . That is Pythagoras, in any dimension, over , derived in one line from three axioms.
Now the angle. In you were told that , and we would like to define the angle that way in general. A definition is only legal if the right-hand side lands in , and nothing we have written down so far guarantees that it does. Proving it is the next item, and it is much more than a technicality.
1.4 · The Cauchy–Schwarz inequality
The claim is that for all in an inner-product space,
with equality exactly when and are parallel.
The proof rests on one idea. Positivity of a norm is an infinite supply of inequalities, and we get to choose which one to use. Whatever we pick, axiom (iii) guarantees that
That statement is true for every , which is a great deal of information. To get at it we have to see the explicitly, so expand the right-hand side, remembering that the first slot conjugates by (0.5.3) and the second does not:
where the two middle terms combined because .
Now let's spend the freedom we have been saving. This quantity is non-negative for every , so it is non-negative for the that makes it smallest, and that is the choice worth making. Assume first that . (If then both sides of (0.5.6) are zero and there is nothing to prove.) Then take
With that choice, , which is real and non-negative, and comes out to the same thing. Substituting both into (0.5.8), the two copies partially cancel:
Multiply through by and you have (0.5.6). Equality holds precisely when the quantity in (0.5.7) is zero, which is to say when . That is exactly the statement that the two vectors are parallel.
So the angle is now legal. Divide (0.5.6) by and it says , so we may define
Grind box — the triangle inequality, and what was
The triangle inequality falls out. Start from (0.5.5) and bound the cross term. Two facts do it, applied in order: , and then Cauchy–Schwarz.
Taking square roots gives . Every metric-space fact you have ever used about these spaces descends from this line, and this line descends from (0.5.6).
What the magic was. It looks pulled from a hat. It is not. Minimise (0.5.8) honestly and it appears on its own. Write with real, so that is real. Then the expression is a real quadratic in :
and a real quadratic is minimised where its derivative vanishes, so set that derivative to zero:
which reproduces (0.5.9) exactly. And the minimising vector is precisely the orthogonal projection of onto the line through , which is the subject of §3. So Cauchy–Schwarz is the statement that a projection is never longer than the thing being projected. Everything in this chapter is one geometric picture seen from different angles.
An alternative proof, for the file. If then . Choose to be minus the argument of , which rotates that product onto the positive real axis, and the real part becomes , which gives directly. The general case follows by rescaling. Same idea, with phase doing the work instead of magnitude.
Chapter 4.9 will prove that for any two observables and measured on a state ,
and the proof there is (0.5.6) applied to a particular pair of vectors built out of the state and the two observables, namely
followed by one line splitting into its real and imaginary parts. That is the entire derivation. Nothing is added in Chapter 4.9 except the physical meaning of the symbols.
Let's say what that means without softening it. The Heisenberg uncertainty principle is the single most quoted statement in twentieth-century physics, the one that gets prose written about the limits of human knowledge. It is Cauchy–Schwarz. You have just proved it. The only thing still missing is permission to say what the letters mean.
Until now the space has had no geometry in it whatever: no lengths, no angles, no way of saying that two directions differ rather than merely being labelled differently. An inner product restores all of that in one stroke, and the question it answers is how much of one vector lies along another. Over the complex numbers the definition is obliged to conjugate one of its two arguments, and that is forced rather than chosen, because without the conjugation a perfectly respectable non-zero vector can be assigned zero length, and the ability to tell two states apart would be lost on the first page.
Out of three short axioms comes an inequality stating that the overlap between two vectors can never exceed the product of their lengths, with equality only when the two are parallel. It has the air of housekeeping. What it actually does is make the word angle legal in a setting where no angles had been defined, and it will do considerably more than that.
Worth saying plainly now, because it is later dressed up beyond recognition. That inequality, applied to two particular vectors assembled from a state and a pair of measurements, is the uncertainty principle. The most quoted sentence of twentieth-century physics is a statement about overlaps, and nothing is added to it afterwards except an interpretation of the letters.
2 · Orthonormal bases
Chapter 0.4 showed that any basis lets you write every vector uniquely in coordinates. It also showed that finding those coordinates means solving a linear system, which is work. With an inner product in hand, one class of basis makes that work vanish altogether.
A set is orthonormal if
Orthonormal sets are automatically linearly independent, so in dimension any of them form a basis. Here is the one-line reason. Suppose , and take of both sides. Linearity in the second slot gives for every , so every coefficient vanishes.
That same computation is the payoff, so let's run it again on a vector we actually care about. Write any and hit it with :
The double sum has collapsed to a single term, and the -th coefficient has fallen out on its own with nothing solved. Renaming the index, that gives the expansion we will use for the rest of the book:
Coordinates have become inner products. There is no system to solve and no matrix to invert. To find the -th coordinate you take one inner product. This is the single practical reason orthonormal bases dominate physics.
Notice also that the physics convention is already paying for itself. Because the second slot is the linear one, is the coefficient exactly, with no stray conjugate attached to it.
Two more consequences follow, and both come from the same move. Substitute (0.5.14) into an inner product, then use (0.5.12) to collapse the resulting double sum:
The second of those is Parseval's identity: the squared length of a vector is the sum of the squared magnitudes of its coordinates. The words "in an orthonormal basis" are load-bearing here, and the identity is false without them.
It is worth knowing now where Parseval gets spent, because it gets spent twice. In Chapter 0.9 the become the Fourier modes, and Parseval becomes the statement that total energy equals the sum of the energies in each mode. In Chapter 4.2 the become probabilities, and Parseval becomes the statement that they add to one.
2.1 · Gram–Schmidt: orthonormal bases always exist
The formulae above are worthless if orthonormal bases are rare. They are not rare, and what proves it is a recipe rather than an existence argument. Given any basis , the following procedure manufactures an orthonormal one. Set
running in order. For the sum is empty, so . In words: take the next vector, subtract off everything it has in common with the directions already fixed, and normalise what is left.
Grind box — why Gram–Schmidt works, line by line
Three things must be checked: that we never divide by zero, that the output is orthonormal, and that it spans the same space. All three come from one induction hypothesis:
: is orthonormal and .
Base case. because it belongs to a basis, so and is a unit vector spanning the same line. holds.
No division by zero. Assume . If were , then by (0.5.16) would lie in . That contradicts the linear independence of the 's. So and is defined.
Orthogonality. For any , take the inner product of (0.5.16) with and use :
Exactly one term of the sum survives, and it is exactly the term that cancels. Dividing by preserves this, so for all , and by construction.
Same span. From (0.5.16), is a combination of and earlier 's, hence of . Rearranged, the same equation says is a combination of . So each set spans the other's span, holds, and induction completes the proof at .
Two consequences worth naming. First, every finite-dimensional inner-product space has an orthonormal basis, so the theorems below never have to assume one exists. Second, Gram–Schmidt is triangular: involves only . Write that relationship as a matrix statement and you get , with having orthonormal columns and upper triangular. That is the QR decomposition, and it is how a computer actually solves the least-squares problem of §3.
2.2 · Dirac notation, introduced where it is harmless
We now install the notation that Chapter 4.2 will use for everything. It is standard to meet it in the middle of learning quantum mechanics, where it acquires an entirely undeserved mystique. It is a piece of pure linear algebra and there is nothing in it. Here it is.
Write a vector as , a ket. Now fix a vector and consider the map that sends any to the number . By axiom (ii) that map is linear in , so it is a linear functional, an element of the space of linear maps . (That space is itself a vector space, since sums and scalar multiples of linear maps are linear.) Call the functional , a bra. Then:
The bracket closes, and a closed bracket is a number.
Now leave the bracket open instead, and see what kind of object comes back. Define to be the map that acts on by applying first and using the resulting number to scale :
Is that a linear map? Check it: is linear in because is, so yes. So the outer product is an operator, while the inner product is a number. In coordinates with respect to an orthonormal basis, the distinction is one you already know from Chapter 0.4:
Row times column is a number. Column times row is a matrix. That is the whole content.
Dirac notation is a device for making the type of an object visible in its shape. You can tell at a glance whether an expression is a scalar or an operator, without tracking any dimensions. Note also, from (0.5.3), that the map is conjugate-linear: . The conjugates in (0.5.19) are that fact written out in coordinates.
2.3 · Completeness: the resolution of the identity
Now rewrite the expansion (0.5.14) in the new notation. Its coefficient is , so
The regrouping in the second step is legal by (0.5.18). Since this holds for every , the operator in brackets acts as the identity on every vector. And two linear maps that agree on every vector are equal, so we may drop the from both sides:
This is the completeness relation, also called the resolution of the identity. It says the basis is not missing any directions. Reconstruct any vector from its components and you get the vector back, with nothing left over.
It is used constantly, and always in the same way. You insert into the middle of an expression in the form (0.5.21), and a hard object turns into a sum of easy ones. Chapter 4.2 does this on nearly every page. So does Chapter 0.9, where the sum becomes an integral over a continuum of modes.
Not all descriptions cost the same, and the difference is large enough to change what is worth attempting. In a general basis, finding the coordinates of a vector means solving a system of equations. Choose the basis so that its members are mutually perpendicular and each of unit length, and the work disappears: every coordinate is a single overlap, computed on its own, without reference to any of the others. Nothing has been approximated. This is why perpendicular bases dominate physics so thoroughly that people forget the other kind exists.
Two consequences arrive immediately. The squared length of a vector becomes the sum of the squared sizes of its coordinates, which later says that the total energy of a vibrating string is the sum of the energies in its modes, and later still that a set of probabilities adds to one. And the claim that the basis has missed no direction can be written as an operation which, inserted anywhere in an expression, changes nothing, so that a difficult object breaks into a sum of easy ones without being altered.
None of it would matter if such bases were rare. They are not. Take any basis at all, subtract from each new vector everything it shares with the directions already settled, and rescale what survives. The procedure never fails, so no theorem ahead has to assume its raw material exists.
3 · Projection, and why least squares is a right angle
Let be a subspace, meaning a subset closed under addition and scalar multiplication. Such a set is a vector space in its own right, and it inherits 's inner product unchanged. By Gram–Schmidt it has an orthonormal basis , and that basis is all we need in order to define the orthogonal projection onto :
Compare that with (0.5.21). It is the same construction, truncated. We keep only the directions inside and throw the rest away.
Three properties follow, and all three are derived rather than assumed.
lands in and fixes . is visibly a combination of the , so . And if then (0.5.14) applied inside says . Put those two together. For any at all, the vector lies in , so applying to it again changes nothing, . In symbols,
Projecting something that has already been projected does nothing. That is what "projection" means, and (0.5.23) is the algebraic form of it.
is symmetric under the inner product. To see it, expand both sides with (0.5.22) and watch them meet in the middle:
The first equality uses conjugate-linearity in slot 1, in the form . In §4 this property will be written .
The residual is orthogonal to . The residual is what the projection threw away, so take its inner product with each basis vector of in turn. For each ,
so for every . That is enough to give entirely, because an inner product with any combination of the is the corresponding combination of zeros.
3.1 · Projection is the closest point — proved
Now the theorem that makes projection useful rather than merely tidy.
Claim: for every , , with equality only for .
We want to compare an arbitrary against , so the move is to write the vector in terms of . Add and subtract it:
The first piece is orthogonal to by (0.5.25). The second lies in , because both and do. So the two pieces are orthogonal to each other, which means Pythagoras (0.5.5) applies with no cross term at all:
Let's look at what that last line is actually saying. The first term on the right does not depend on in any way. The second is non-negative, and it vanishes only when . So the minimum is attained there and nowhere else.
Read (0.5.27) once more. It says "nearest point" and "perpendicular residual" are the same condition. Not similar, not related: the same. Every optimisation problem that turns out to have a linear-algebra answer is this theorem in disguise.
Fit , with the outcomes, the design matrix, and the coefficients to be chosen. Minimising the residual sum of squares means minimising , which is a squared distance in with the ordinary inner product.
As ranges over all of , the vector ranges over exactly the column space of , the subspace spanned by the predictors. So the problem is this: find the point of that subspace closest to . By the theorem just proved, the answer is the orthogonal projection, and it is characterised by the residual being perpendicular to the subspace. Perpendicular to the subspace means perpendicular to every column of , and writing that down gives
Those are the normal equations, and the word "normal" in their name has always meant "perpendicular". Solving them gives and with . That is the "hat matrix" of every regression output you have ever read, and it satisfies and because it is the projection of (0.5.22).
One more line pays a dividend. If contains an intercept column, the constant vector lies in the column space, so (0.5.27) applied to gives
The ANOVA decomposition is Pythagoras' theorem. And is the squared cosine of the angle between the centred outcome vector and its projection. Regression is trigonometry in dimensions.
Suppose you are stranded outside some subspace and want the point of it nearest to you. Keep the parts of your vector lying along that subspace's own perpendicular directions and discard everything else, and what you kept is the nearest point, uniquely so. The proof is Pythagoras used once. What is left over is perpendicular to the subspace, the error committed by choosing any other point lies inside the subspace, and two perpendicular pieces add as squares, so wandering away can only add.
The sentence to carry off is that nearest point and perpendicular leftover are not two conditions that happen to agree. They are one condition. Every optimisation problem whose answer turns out to be linear algebra is this fact wearing a costume.
The most familiar costume is ordinary least squares. Fitting a model means choosing the combination of predictors closest to the observed outcomes; the combinations you can reach form a subspace; the fitted values are therefore the projection of the data onto it. The equations defining the fit assert that the residual is perpendicular to every predictor, which is why the word normal in their name has always meant perpendicular. The decomposition of variance printed at the foot of every regression table is Pythagoras with the right angle at the fitted values.
4 · The adjoint
Every operator on an inner-product space has a partner. What defines the partner is not a formula in coordinates but a piece of behaviour: how it interacts with the inner product.
Definition. The adjoint of a linear map is the linear map satisfying
Read it as a rule for moving an operator across the comma, and note that doing so costs you a dagger. No basis appears anywhere in (0.5.28), and that is the point. The adjoint is a property of and the inner product, not of any coordinate system.
A definition by behaviour owes us two checks: that such a map exists at all, and that there is only one of them. Both are short.
It is unique. Suppose and both satisfy (0.5.28). Then for every and . We get to choose , so choose . This gives , so for every , so .
It exists. Fix an orthonormal basis and define . This is linear in , because the inner product is linear in slot 2. Now expand via (0.5.14) and compute the left-hand side of the definition:
The right-hand side of that chain is , which is what (0.5.28) asked for. Existence and uniqueness settled, in five lines.
4.1 · In an orthonormal basis it is the conjugate transpose
Chapter 0.4 defined the matrix of by . Take of that and use (0.5.12), and you get a formula worth remembering: . A matrix element is an inner product.
Now compute the matrix of the same way. Use (0.5.28) to move the operator across the comma, then conjugate symmetry to turn the result around:
So is the conjugate transpose: transpose the matrix and conjugate the entries. That familiar recipe is not the definition. It is what the definition (0.5.28) looks like once you commit to an orthonormal basis. Note the hypothesis, because it is an easy one to lose. In a non-orthonormal basis (0.5.30) is flatly false.
Three algebraic rules follow immediately from (0.5.28), each by moving operators across the comma one at a time:
The last one is worth doing out loud, since it is the one people misremember: , and uniqueness does the rest. The order reverses, exactly as it does for the transpose and the inverse. One more rule gets used constantly below and is derived the same way: .
4.2 · The three species of operator
| Name | Condition | What it is for |
|---|---|---|
| Hermitian (self-adjoint) | Observables. §6 is about these | |
| Unitary | Symmetries, time evolution, change of orthonormal basis | |
| Normal | The largest class that can be diagonalised orthonormally |
Both Hermitian and unitary operators are normal. For a Hermitian operator the reason is that . For a unitary one it takes an extra step: implies, in finite dimensions, that is invertible with , and therefore as well.
For a Hermitian matrix, (0.5.30) says . Read that entry by entry. The diagonal entries are real, and entries across the diagonal are conjugates of each other. A real Hermitian matrix is nothing but a symmetric matrix.
4.3 · Unitary maps preserve all geometry
This is a two-line calculation with large consequences. We want to know what does to an inner product, so move across the comma using its defining property and watch it cancel. If then for any :
The inner product is untouched, and everything geometric follows from that one line. Setting gives , so lengths are preserved. Feeding both facts into (0.5.11) leaves unchanged, so angles are preserved too. And taking from an orthonormal basis shows is again orthonormal, so unitary maps are exactly the maps that carry orthonormal bases to orthonormal bases.
A real unitary matrix satisfies , which is precisely an orthogonal matrix, meaning a rotation or a reflection. Unitary is the complex analogue of rotation. Hold that next to the last section of Chapter 0.3, where turned out to be a rotation of the plane. The connection is not an analogy, and §7 will make it an identity.
Here is where this goes. In Chapter 4.6 the state of a quantum system evolves by , and must be unitary because is a total probability and probabilities have to keep summing to one. Equation (0.5.32) is that requirement written in linear algebra, and it is the reason time evolution in quantum mechanics is a rotation rather than a flow that stretches.
Every map acquires a partner the moment its space has an inner product, and the partner is specified by a piece of behaviour rather than by a recipe. It is the map you substitute when you want to shift an operator from one side of an overlap to the other. No basis is mentioned anywhere in that description, which is exactly what makes it worth having. The familiar instruction to transpose an array and conjugate its entries is not the definition; it is what the definition looks like after you commit to a perpendicular basis, and in any other basis it is false.
Three kinds of map are then singled out by how each sits beside its partner. One kind is its own partner. A second has a partner that undoes it. A third merely commutes with its partner, and is the widest family for which the next section's conclusions survive at all.
The second kind earns a word now. A map whose partner undoes it leaves every overlap exactly as it found it, so it alters no length and no angle, and it carries perpendicular bases to perpendicular bases. It is the complex counterpart of a rigid rotation. That is why evolution in time, in any theory where a total probability must stay equal to one, has no choice about what kind of map it is.
5 · Eigenvalues and eigenvectors
A linear map generally does two things at once to a vector. It turns the vector, and it rescales it. Some directions are special, and they are the ones where only the second thing happens.
Definition. A non-zero is an eigenvector of with eigenvalue if
Geometrically, spans a direction the map does not turn. It only stretches it, by the factor . A negative flips it, and a complex means, as we will see in a moment, that no such real direction exists at all. The requirement is essential, since holds for every and would make the definition vacuous.
To find these directions we need something we can solve, so rewrite (0.5.33) as . A non-zero solution exists exactly when fails to be invertible, which by Chapter 0.4 happens exactly when its determinant vanishes. So the eigenvalues are the roots of the characteristic polynomial:
Expanding the determinant of an matrix produces a polynomial of degree in , with leading term . Which numbers you are allowed to use now matters enormously. Over a polynomial need not have any roots at all, since has none, and correspondingly a real matrix can have no real eigenvectors whatsoever. A rotation of the plane turns every direction, so none is preserved. Over this cannot happen. As recorded in Chapter 0.4, every non-constant polynomial has a complex root, so
That single sentence is why quantum mechanics is built over and not , and it is the hypothesis that makes §6 work.
Two pieces of vocabulary before we go on. The set of eigenvectors for a given , together with , forms the eigenspace , a subspace of dimension at least one. An eigenvalue whose eigenspace has dimension greater than one is called degenerate.
5.1 · Seeing it
Before proving anything, look at the phenomenon. The figure below applies a real matrix to every vector on the unit circle. The image is an ellipse. Take that as observed for the moment, since §6 will prove it, and the proof turns out to be the spectral theorem itself. A rotating input vector is drawn together with its image, and the eigen-directions are marked, meaning the angles at which input and image become parallel.
The figure's whole behaviour comes out of one small computation, and it is worth doing so that the sliders stop being magic. Write the family it uses as a symmetric matrix plus an antisymmetric one:
The eigenvalues are the roots of the quadratic (0.5.34), so what decides whether they are real is the discriminant. Form it:
At the discriminant is a sum of squares. It is therefore non-negative, and the eigenvalues of a real symmetric matrix are always real. That is a preview of §6(a).
Turning on eats into the discriminant, and the eigenvalues stay real only until , which is exactly the squared half-gap between the eigenvalues of the symmetric part. In the figure , , , so the symmetric matrix has eigenvalues and and a half-gap of . The collapse therefore happens at , which is what the button does.
Turning and rescaling are what a map does to a vector at once, and the directions in which the turning stops are where the map is at its most legible. Locating them comes down to asking when the map, with a multiple of the identity subtracted from it, stops being invertible, since only then can something non-zero be sent to nothing. That condition is the vanishing of a determinant, which converts the search into finding the roots of a polynomial.
What the polynomial returns depends on which numbers you permit, exactly as the previous chapter warned. It depends on the map too, in a way better watched than read about. Begin with a symmetric object and its two special directions stand at right angles. Add an antisymmetric piece and they lean towards one another, meeting and merging at a definite amount of skew, beyond which they are gone altogether and nothing is left unturned.
That behaviour is what the remainder of the chapter accounts for. Perpendicular special directions are not a general property of maps and may never be assumed; they belong to symmetric ones specifically, and the symmetry does every bit of the work. Notice which way the dependence runs, because an enormous amount will later be resting on it.
6 · The spectral theorem
This is the centre of the chapter and of Part 0. Everything before it was preparation, and a great deal after it is consequence. The theorem comes in three parts, and we take them in order.
6.1 · (a) Hermitian ⇒ every eigenvalue is real
Let and with . Since is Hermitian, the defining property (0.5.28) reads . We want a statement about on its own, so apply that with and evaluate each side using (0.5.33):
The first step used conjugate-linearity in slot 1, and the last used linearity in slot 2. Since , axiom (iii) gives , so we may divide by it:
Three lines. The entire asymmetry between the two slots of the inner product has been converted into a statement that a number is real.
6.2 · (b) Hermitian ⇒ eigenvectors with distinct eigenvalues are orthogonal
Let and with , both real by (a). We want to show that , and the route there is to compute one quantity in two different ways. Take :
The last step is where reality is spent, since it needs from (a). The two expressions have to agree, so subtract one from the other:
since . Different eigenvalues force exactly perpendicular eigenvectors, not merely independent ones. Nothing was assumed about except .
6.3 · (c) The theorem
Let be a Hermitian operator on a finite-dimensional complex inner-product space . Then has an orthonormal basis consisting of eigenvectors of , with real eigenvalues , and
Before the proof, let's look at what that last equation says. Compare it with completeness (0.5.21). The identity is , and is the same sum with the terms weighted by the eigenvalues. An operator that would be complicated in a random basis is, in its own eigenbasis, nothing but a list of numbers attached to a list of perpendicular directions.
Grant the orthonormal eigenbasis for a moment, and the formula is immediate. Both sides are linear maps, so it is enough to check that they agree on every basis vector, and they do:
By Chapter 0.4, two linear maps agreeing on a basis are equal. So the work is entirely in producing the basis, and producing it is what the grind box does.
Grind box — the induction proof of the spectral theorem
We induct on . The claim to be proved at each stage is: every Hermitian operator on an -dimensional complex inner-product space has an orthonormal basis of eigenvectors.
Base case . Pick any unit vector . It spans , so is a multiple of it, . Done, and is real by part (a).
Step 1, get one eigenvector. Let and assume the claim in dimension . Because is complex, (0.5.35) supplies an eigenvalue and an eigenvector, which we normalise to a unit vector . By (a), .
Step 2, pass to the orthogonal complement. The plan is to peel off the line through and apply the induction hypothesis to whatever is left, so define
It is a subspace, because is linear. It inherits the inner product of and every axiom still holds on a subset, so is a complex inner-product space in its own right. We claim . To see it, split any using the projection of §3 onto the line through :
and the second bracket lies in by (0.5.25). So if is any basis of , then spans . It is also independent: applying to a vanishing combination kills every term and leaves , after which the independence of the forces every . So it is a basis, and since dimension is well defined (Chapter 0.4), .
Step 3, maps into . This is the step that uses Hermiticity, and it is the hinge of the whole proof. For ,
using in the first step and in the third. So , which is to say the restriction is a genuine operator on . Without Hermiticity this fails. A general operator does not preserve the orthogonal complement of one of its eigenvectors, and the induction dies right here. That is precisely why the figure in §5 loses its perpendicular eigen-directions the moment you break the symmetry.
Step 4, the restriction is still Hermitian. For , the defining property is inherited verbatim, because it is a statement about inner products of vectors that all lie in :
Step 5, induct. By the induction hypothesis has an orthonormal basis of eigenvectors of . But acts as does, so these are eigenvectors of . Each lies in , hence is orthogonal to , and they are orthonormal among themselves. So is an orthonormal set of eigenvectors in an -dimensional space, which makes it a basis.
What was used. Exactly three things: that guarantees an eigenvalue (0.5.35), that Hermiticity makes the orthogonal complement invariant, and that the property survives restriction. Remove any one of the three and the theorem is false.
6.4 · Degeneracy, and the honest statement
Nothing above assumed the eigenvalues were distinct, so let's ask what happens when they are not. If is degenerate, the induction hands you several basis vectors with the same eigenvalue, and any orthonormal basis of the eigenspace will do. The choice is not unique, and no theorem prefers one. What is unique is the eigenspace itself.
That suggests writing the decomposition in terms of eigenspaces rather than eigenvectors, since only the eigenspaces are forced on us. Group the terms by distinct eigenvalue and write for the projection onto :
The middle identity holds because eigenvectors from different eigenspaces are orthogonal by (b). The right-hand one is completeness (0.5.21), regrouped. This is the basis-independent form, and it is the one that survives to infinite dimensions in Chapter 4.5.
Finally, the matrix version. Assemble the orthonormal eigenvectors as the columns of a matrix . Orthonormality of the columns is exactly the statement , so is unitary, and for every column reads . We want by itself, so multiply on the right by :
In the language of Chapter 0.4 this is a change of basis, and the change of basis is unitary, which is to say a rotation. Every Hermitian operator is a diagonal one, seen from a rotated angle.
Two dividends fall out at once, using cyclicity of the trace from Chapter 0.4:
6.5 · The real case, and the ellipse we borrowed
A real symmetric matrix is Hermitian, so (a) gives it real eigenvalues. For a real the matrix is real and singular, so it has a real null vector. And Gram–Schmidt inside each eigenspace uses only real arithmetic. Put those together and a real symmetric matrix has an orthonormal basis of real eigenvectors, so that (0.5.44) becomes with real orthogonal. That is the classical principal axes theorem.
That closes the loan taken out in §5. Let be real, invertible, acting on the unit circle. A point of the image is with , which is the same as saying . Squaring that condition turns it into a quadratic form set equal to one:
Now we are in a position to use the theorem, because is exactly the kind of matrix it governs. Note that , and that for . So is real symmetric with strictly positive eigenvalues , each . Substituting turns (0.5.46) into , an ellipse with semi-axes along the eigenvectors of . The image of a circle under any invertible linear map is an ellipse, and the reason is the theorem you just proved.
One detail is worth extracting, because the figure in §5 shows it. The ellipse's axes are the eigenvectors of , not of . If is symmetric then has the same eigenvectors and the two coincide, which is why the eigen-lines lie exactly along the ellipse in the symmetric preset. Break the symmetry and they part company, visibly.
Read the three results back, replacing "Hermitian operator" with "observable" and "eigenvalue" with "possible measurement outcome". Do not accept the replacement yet. Just look at the sentences.
(a) Eigenvalues are real becomes: measurement outcomes are real numbers. An apparatus reads out a number on a dial, and it cannot read grams. In Chapter 4.2 this is presented as a requirement on the theory. It is not a requirement. It is (0.5.39).
(b) Eigenvectors are orthogonal becomes: distinct outcomes are perfectly distinguishable. Two states with different measured values have zero overlap, so the probability of confusing one for the other, which will turn out to be , is exactly zero. Not small. Zero.
(c) The eigenvectors form a complete orthonormal basis becomes: any state whatsoever can be written as a superposition of outcomes, with . And by Parseval (0.5.15), so the probabilities add to one automatically.
These are three of the postulates of quantum mechanics as they are usually taught, and they are not postulates. They are theorems about Hermitian matrices, proved above, in a chapter containing no physics. Chapter 4.2 will do exactly one thing that this chapter did not. It will assert that a physical state is a unit vector and a physical observable is a Hermitian operator. That is one postulate, not four. Everything else is renaming.
Most matrices are not diagonalisable at all. The spectral theorem's hypothesis is doing heavy work, and it is easy to miss how heavy. Consider
The only eigenvalue is , repeated twice. Its eigenvectors solve , i.e. , which forces . So the eigenspace is the single line spanned by . A two-dimensional space, and only a one-dimensional supply of eigenvectors. There is no basis of eigenvectors, no diagonal form, and no spectral decomposition.
Such a matrix is called defective, and it is not exotic. The skew slider in the figure above manufactures one at . Note that , so the theorem is not violated. It is inapplicable, which is a different thing. When a physics text says "expand in eigenstates", it is quietly leaning on Hermiticity, every time.
In infinite dimensions the theorem needs genuine repair. Everything proved in this chapter is a finite-dimensional theorem, and two of the steps fail outright in a function space. Both failures are worth knowing rather than filing away, because you will meet them.
First, "Hermitian" and "self-adjoint" come apart. An unbounded operator carries a domain, the subset of functions it is allowed to act on, and carries its own, generally larger. The symmetric condition on a domain is strictly weaker than genuine self-adjointness, and the difference is not pedantry. The momentum operator on the half-line is symmetric and admits no self-adjoint extension whatsoever.
Second, eigenvectors can fail to exist inside the space at all. On the whole line, the momentum eigenfunctions solve the eigenvalue equation but are not square-integrable, so they are not vectors of the Hilbert space. The spectrum becomes continuous, the sum becomes an integral over projection-valued measures, and completeness must be re-proved rather than assumed. Chapter 4.5 pays this bill in full. Until then, everything here is exact and everything here is finite-dimensional.
At the centre of the toolkit sit three statements, none of them long to prove. A map that is its own partner has multipliers that are real numbers. Its special directions belonging to different multipliers are exactly perpendicular, not merely independent. And there are enough of those directions to describe every vector in the space, with nothing left over.
Taken together they say that such a map, viewed in the right description, is a list of numbers attached to a list of mutually perpendicular directions, each one acting alone and none of them speaking to any other. This is the move most of the book is built from, and it deserves its name: the problem has fallen apart into independent pieces. Coupled oscillators, the modes of a plucked string, the components of a wave, the energy levels of an atom — every one of those is this, and the repetition is not laziness. It is one theorem being spent over and over.
Nothing in the argument is physics. Yet read the three again with the words changed and they say that measurement outcomes are real numbers, that distinct outcomes are perfectly distinguishable rather than merely different, and that any state can be written as a combination of outcomes whose weights add to one. Those are usually offered as postulates about nature. They are theorems about symmetric arrays.
7 · Functions of operators
You know what means for a number. What could mean for an operator? The spectral theorem answers it at once. If acts on each by multiplying by , then whatever is, it ought to act on by multiplying by . So define
for Hermitian with spectral decomposition , and any function defined on the eigenvalues.
There is one question to settle before we trust that definition. Does the answer depend on which orthonormal basis we happened to pick inside a degenerate eigenspace? In the projection form (0.5.43) the definition reads , and the are unique, so the answer is no. Equation (0.5.47) is well posed.
It had better also agree with the answer you would get by substituting into a power series, and it does. The key is that orthonormality makes the cross terms vanish, which is easiest to see by squaring and watching the double sum collapse:
By the same collapse, inductively, for every . It holds for as well, since by completeness. So every power behaves, and that is enough to push a whole series through. For any whose series converges at every eigenvalue,
Swapping the two sums is legitimate because the sum over is finite. So (0.5.47) and "plug the matrix into the series" agree wherever both make sense. And (0.5.47) keeps working for functions with no power series at all, such as on a positive operator.
7.1 · The exponential of times a Hermitian operator
Claim: if is Hermitian then is unitary.
Unitarity is a statement about the adjoint, so the first thing we need is the adjoint. Take it of (0.5.47) with , using and from (0.5.31):
The middle step is where reality is spent. It uses , and holds only because is Hermitian. Now we have both factors of , so multiply them together, collapsing with orthonormality exactly as in (0.5.48):
So is unitary.
The converse holds as well: every unitary operator is for some Hermitian , and the grind box proves it. Put the two halves together and Hermitian operators and unitary operators are the same information, related by an exponential. That is exactly how real numbers and points on the unit circle are related by in Chapter 0.3. This is that statement with the number replaced by an operator.
Grind box — the converse: every unitary is
We need a spectral theorem for unitary operators. The induction of §6 goes through with two substitutions, and here they are.
Eigenvalues have modulus one. If with then by (0.5.32), , so and we may write with real.
The orthogonal complement is invariant. In finite dimensions , and gives since . So for ,
which says . The restriction preserves inner products on and maps into . Being injective on a finite-dimensional space it is onto, hence unitary on .
Induct exactly as before, using (0.5.35) to start. The result is an orthonormal basis with , so
That is Hermitian because the are real and each is self-adjoint. One caveat is worth recording. The are only defined modulo , so is not unique, and the exponential map is many-to-one, exactly as is on .
In Chapter 4.6 the Hamiltonian is Hermitian because energy is an observable, and time evolution is
That operator is unitary by the theorem just proved, and unitary means probability is conserved by (0.5.32). So the sentence "energy is observable, therefore probability is conserved" is not a physical argument at all. It is (0.5.51). Differentiate the exponential and you get , so the Schrödinger equation is the statement that a Hermitian operator generates a unitary flow.
This pairing of a Hermitian generator with a unitary group element is the finite-dimensional shadow of the Lie algebra and Lie group relationship of Chapter 6.1, where maps the algebra, a vector space of generators, to the group, a curved space of transformations. Every symmetry in the Standard Model is an instance. You have just met the whole idea in two dimensions, with no differential geometry anywhere.
Once an object has been broken into independent directions you can do arithmetic on it one direction at a time, and that is the whole of what it means to take a function of a map. Multiply each direction's number by whatever the function does to that number, and leave the directions themselves alone. This agrees with the obvious alternative of substituting the map into a power series, because perpendicularity kills every cross term, and it goes on working for functions that have no series at all.
The case that matters is the exponential of the imaginary unit times a self-partnered map. Each of the real numbers becomes a phase of unit length, so the result preserves every overlap and belongs to the rotation-like family met earlier. The converse holds as well, so the two families carry the same information and are related by exponentiation, precisely as the real numbers are related to the points of a circle.
That equivalence is more consequential than its length suggests. It is the reason a quantity being observable forces the flow it generates to conserve probability, so that an argument people offer as physics is a line of algebra. It is also the first sighting of the pairing between a generator and the family of transformations it builds, which is how every symmetry in fundamental physics is eventually described.
8 · Commuting operators and simultaneous diagonalisation
Define the commutator . It measures the failure of two operators to be interchangeable, and it is zero exactly when order does not matter. The theorem below is the reason that quantity is worth a name.
Theorem. Two Hermitian operators and on a finite-dimensional complex inner-product space admit a common orthonormal basis of eigenvectors if and only if .
8.1 · Forward direction (easy)
Suppose is an orthonormal basis with and . We want to compare with , so apply each of them to a basis vector and see what comes out:
and because they are numbers. So and agree on a basis, hence are equal (Chapter 0.4), so .
8.2 · Reverse direction (the real content)
Here the difficulty is degeneracy. If had distinct eigenvalues the argument would be three lines. It generally does not, and the honest proof must build the common basis inside each eigenspace.
Grind box — the reverse direction, degeneracy handled
Assume , , and .
Step 1, decompose by . By the spectral theorem, is the orthogonal direct sum of the eigenspaces of for its distinct eigenvalues :
with the orthogonality being part (b) of §6 and the spanning being part (c).
Step 2, preserves each eigenspace. This is the only place the hypothesis is used. Let , so . We want to know where sits, so feed it to and use the commutation to move past :
so is again an eigenvector of with eigenvalue , or else zero, which lies in too. Hence .
Step 3, the restriction is Hermitian. This runs exactly as in Step 4 of the spectral-theorem proof. For , holds because it holds in all of , and both vectors lie in the subspace. So is a Hermitian operator on the finite-dimensional inner-product space .
Step 4, apply the spectral theorem inside the eigenspace. Therefore has an orthonormal basis of eigenvectors of . And every vector in is automatically an eigenvector of , with eigenvalue , because that is what the eigenspace is. So this basis consists of simultaneous eigenvectors of both operators. This is exactly where degeneracy is handled. Within a degenerate eigenspace of , the operator chooses the basis that could not.
Step 5, assemble. Take the union of these bases over . Vectors from the same eigenspace are orthonormal by construction, and vectors from different eigenspaces are orthogonal by Step 1. The total count is , so the union is an orthonormal basis of consisting of simultaneous eigenvectors of and .
Remark. Run the same argument with three or more mutually commuting Hermitian operators and it produces a basis of simultaneous eigenvectors for all of them. Apply Step 2 to each in turn, refining the decomposition as far as the operators allow. Whether that refinement reaches eigenspaces of dimension one depends on the set: it does exactly when the operators are enough to tell every basis vector apart by its eigenvalues, and a set that is not enough still yields a common eigenbasis while leaving some joint eigenspace of dimension two or more. Chapter 4.9 §3 turns that distinction into a criterion.
The theorem promises that a common eigenbasis exists. It does not promise that the one your solver just handed you is it, and the figure below is that distinction, driven to the point where it bites.
An orthonormal basis of simultaneous eigenvectors is a way of labelling every state by a list of numbers, one eigenvalue from each operator. Chapter 4.11 calls such a list the quantum numbers of the state, and calls the operators a complete set of commuting observables. Complete means two things: mutually commuting, and enough of them that every eigenspace has been cut down to a single line, so the labels determine the state uniquely. When you write the hydrogen states as you are naming the eigenvalues of three commuting operators. And the reason the labels are consistent, which is to say the reason a state can have a definite and a definite at the same time, is Step 4 above, and nothing else.
Run it backwards and you get the other half. If there is no common eigenbasis, so there is no state in which both observables are sharp. That is the qualitative content of Chapter 4.9. The quantitative version is the Cauchy–Schwarz bound from §1, whose right-hand side is non-zero exactly when the operators fail to commute. The two great mysteries of quantum mechanics, why can some quantities be known together and others not?, are one theorem about matrices, proved above, with an "if and only if" in the middle.
You have run principal component analysis on expression data, on imaging features, on multiplexed assay panels. Here is what was actually happening, and why it could not have failed.
Centre the data matrix ( samples features) and form the covariance matrix , whose entries are . Covariance is symmetric in its two arguments, so . A real covariance matrix is Hermitian by construction, not by assumption and not by luck. Section 6 therefore guarantees, as a theorem:
- The eigenvalues are real. They are variances, so this had better be true. It is not a numerical accident, and it does not depend on the data.
- The principal components are mutually orthogonal. Uncorrelated components are not something PCA arranges by cleverness. They come out perpendicular because §6(b) says eigenvectors of a symmetric matrix with distinct eigenvalues have no choice.
- They span the whole feature space. Completeness (0.5.21) says no direction of variation is missed, so reconstructing from all components returns your data exactly.
More is true, and it is worth being precise about the extra. is not merely Hermitian. It is positive semi-definite, and one line shows it for any direction :
because it is a sum of squares, and that sum is exactly the sample variance of the data projected onto . Apply this to a unit eigenvector and it gives . So the eigenvalues are not just real, they are non-negative, and each eigenvalue is the variance along its own component. Positive semi-definiteness is strictly stronger than Hermiticity: a Hermitian matrix may have negative eigenvalues, and a covariance matrix may not.
Two further standard facts now cost one line each. Write a unit vector as in the eigenbasis with . The spectral decomposition then gives , a weighted average of the eigenvalues, and a weighted average is maximised by putting all the weight on the largest. The first principal component is the top eigenvector because a weighted average is largest when it is not an average. And by (0.5.45) the total variance is , which is why "proportion of variance explained" is and why those fractions sum to one.
Now the point worth carrying away. PCA works for exactly the same reason quantum measurement outcomes are real numbers. One theorem, §6, applied twice. Once to a covariance matrix, where the real eigenvalues are variances and the orthogonal eigenvectors are uncorrelated components. Once to a Hamiltonian, where the real eigenvalues are energies and the orthogonal eigenvectors are distinguishable states. It is not an analogy and not a coincidence of formalism. It is one theorem with two costumes, and you have been using the harder-looking costume for years without noticing.
Two maps of the self-partnered kind fall apart into the same independent directions when, and only when, the order of applying them makes no difference. Half of that is easy: if both are lists of numbers over one shared set of directions, then numbers commute and so must the maps. The other half carries the content. Wherever the first map cannot distinguish between several directions, having given them all the same number, the second goes inside that ambiguity and chooses, and between them they produce one set of directions suiting both.
This is how a physical state comes to be labelled by a handful of numbers, one from each map, the labels fixing the state completely once there are enough maps to leave no ambiguity. The reason such labels can coexist, which is to say the reason a state may have a definite value of one quantity and simultaneously of another, is the argument above and nothing besides.
Run it backwards for the other half. Where two maps fail to commute there is no shared set of directions, so no state exists in which both quantities are sharp, and the size of the failure is what the overlap inequality from the opening of the chapter measures. The two things people find most mysterious about quantum mechanics are one theorem about matrices with an if and only if in it.
9 · Worked examples
Take the Hermitian matrix
It is real and symmetric, hence Hermitian, so §6 applies before we compute anything. The eigenvalues will be real and the eigenvectors perpendicular. Let us watch it happen.
Eigenvalues. From (0.5.34),
and that is a difference of two squares, so take square roots of both sides:
Real, as promised. And note what has happened here. A matrix with equal diagonal entries, two "levels" of the same value , has eigenvalues that are split apart by the off-diagonal coupling, into and , separated by .
Eigenvectors. For , solve :
For the matrix becomes , forcing , so . Both have been normalised by dividing by .
Orthogonality, checked. ✓ Perpendicular, exactly as §6(b) requires. We never used that fact to find them, so this is a genuine test rather than a restatement.
Spectral decomposition. Build the projections as column times row:
Sanity checks: ✓ (completeness), ✓ (orthogonality), ✓ (idempotence).
Reconstruction. Now rebuild from its spectrum alone, which is the claim of §6.3 made concrete:
The matrix has been completely dissolved into two numbers and two perpendicular directions, and reassembled from them.
Why this example and not another. The eigenvalues are a level splitting. Two states that would be degenerate at energy are pushed apart by the coupling , and the new eigenstates are the symmetric and antisymmetric combinations rather than either original state. This matrix is the most reused object in all of quantum mechanics. It is the ammonia molecule, whose two mirror-image configurations split by and give the ammonia maser. It is neutrino oscillation, where the states of definite mass are not the states of definite flavour and the mismatch makes them convert into one another. It is the qubit, where and are the two states you engineer and control. Chapter 4.2 will do all three, and the linear algebra will already be finished.
Let
which is real symmetric, hence Hermitian. Compute for real using (0.5.47) rather than by summing a series.
Step 1, the spectrum. , so . The eigenvectors are the same as in Worked example 1. (Set and there and compare.) That gives the spectral decomposition
Step 2, apply the function. The operator in the exponent is , whose eigenvalues are on the same eigenvectors. So with ,
Step 3, regroup using Euler. Write from Chapter 0.3 and collect the 's:
using completeness for the first bracket and Step 1 for the second. Written out as a matrix, that is
What this is. Compare it with . The structure is identical, and this is Euler's formula with a matrix in the exponent. The reason it works is that , so plays the role that plays for numbers, except that it squares to rather than . The eigenvalues are what carry the alternating signs. Any Hermitian with gives by the identical three steps.
Checks. Unitarity: , and multiplying gives ✓, as §7 promised. Determinant: ✓.
Where it goes. Unitary matrices of determinant form the group , and the operator above is a member of it. In Chapter 4.12, is precisely the operator that rotates a spin- state about the axis by angle . That is the same computation as this one, with a general unit vector in place of the -direction. The factor of in the exponent is why a spin must be turned through to return to itself, and you can already see it coming: at the formula above gives , not . Chapter 6.1 explains what kind of object does that.
10 · Your turn
Problem 1 · Gram–Schmidt, twice
(a) Apply (0.5.16) to the basis , , of with the ordinary dot product. Verify the result is orthonormal.
(b) Now do the same in a space with no arrows in it. Take the polynomials on with , and orthonormalise them.
Solution
(a) , so .
Next, , so
and normalising that vector gives the second basis element:
Finally and , so subtracting both components off leaves
which normalises to
Check. , and , and . Each has norm by construction. ✓
(b) , so .
, because the integrand is odd. So is already orthogonal to the constants and nothing needs subtracting. With , normalising gives
For : by oddness again, while . Subtracting that one component,
and the norm of what is left is
so that dividing through by it gives
Those are the first three Legendre polynomials, normalised: , , , each divided by . They were not invented. They are what you get by running Gram–Schmidt on , and there is nothing else they could have been. The identical procedure with a weight on the whole line produces the Hermite polynomials, which are the quantum harmonic oscillator states of Chapter 4.8. The machinery does not care that the vectors are functions. That is the whole content of the abstraction, and it is what Chapter 4.3 will exploit.
Problem 2 · the spectrum of a unitary operator
Prove that every eigenvalue of a unitary operator has . Then interpret it: what does this say geometrically, and what would go wrong physically if it were false?
Solution
Let with . Apply (0.5.32) with both arguments equal to :
Since , axiom (iii) gives and we may divide, leaving . So and for some real .
Geometry. A unitary map preserves all lengths, so along any direction it preserves it must stretch by a factor of modulus one, which is to say not stretch at all but only rotate the phase. A unitary operator has no expanding or contracting directions anywhere. In its eigenbasis it is , a separate rotation on each of perpendicular axes. Contrast a Hermitian operator, which in its eigenbasis is with the real: it only stretches and never rotates. The two classes are exactly complementary, which is the content of .
Physically. If some , a state along that eigenvector would grow in norm under repeated application of . Since is total probability, that state would end up with probability greater than one, meaning an outcome that occurs more often than always. If the state would fade away and probability would leak out of the universe. Chapter 4.6's insistence on unitarity is exactly the demand that neither happens, and by the result above there is no third option, since unitary forces automatically.
Problem 3 · symmetry is what makes the right angle
(a) Diagonalise : find the eigenvalues, normalise the eigenvectors, and verify they are orthogonal.
(b) Now break the symmetry with , which has the same symmetric part. Find its eigenvalues and eigen-directions and compute the angle between them. Comment on what happened.
Solution
(a) gives , so and . For : forces , giving , at . For : , giving , at .
Orthogonality: ✓. The two eigen-directions are exactly apart, as §6(b) guarantees for any real symmetric matrix with distinct eigenvalues.
(b) and , so the quadratic formula gives
Real, but that is luck rather than a theorem. The discriminant came out positive here and will not in general. For , the first row of gives
a direction at . For the sign flips and the direction is at . Their dot product, with , is , so they are not perpendicular, and the angle between them follows from dividing by the two lengths:
Comment. The right angle is gone: became . Nothing was done to the symmetric part of the matrix, since is an antisymmetric addition, and yet the eigen-directions tilted toward each other by each. This is the figure of §5 with different numbers, and by (0.5.37) the eigenvalues would have gone complex once the antisymmetric coefficient exceeded half the eigenvalue gap of , i.e. . Orthogonality of eigenvectors is not a property of matrices. It is a property of Hermitian matrices, and it is lost the moment the hypothesis is.
Problem 4 · non-degenerate commuting operators share everything
Let and be Hermitian with , and suppose is non-degenerate: every eigenvalue of has a one-dimensional eigenspace. Show that every eigenvector of is automatically an eigenvector of . Then explain why the general proof of §8 needed to be longer than this one.
Solution
Let with . We want to know where sends , so push through and use to move inside:
So lies in the eigenspace . By hypothesis is one-dimensional and contains , so and therefore for some scalar . Hence is an eigenvector of . (If then , which is still an eigenvalue, since the zero vector is excluded as an eigenvector, not as an eigenvalue.) The eigenvectors of already form an orthonormal basis, so that basis diagonalises too.
Why the general case is harder. The step " is one-dimensional, therefore is a multiple of " is exactly what fails under degeneracy. If , all we learn is that maps that three-dimensional space into itself. The vector could be any vector in it, and a generic will not be an eigenvector of . The repair is the whole content of the grind box in §8: restricted to is still Hermitian, so the spectral theorem applies inside and selects the right basis there. Degeneracy does not break the theorem. It means alone no longer specifies the basis and needs 's help to finish the job.
The physics of that last sentence, for later: a degenerate energy level is one where the Hamiltonian alone does not tell you which state you are in, and you must measure a second commuting observable to say. That is exactly why hydrogen states need three labels rather than one, and why "lifting a degeneracy" with a magnetic field, say, is such a common experimental move. Chapter 4.15.
You have the inner product from three axioms, with the conjugate forced by the demand that no non-zero vector have zero length, and every geometric notion rebuilt from it: length, distance, orthogonality, angle. You have Cauchy–Schwarz, proved by optimising a manifestly non-negative quantity, which is what makes "angle" definable at all. You have orthonormal bases, in which coordinates are inner products. You have Gram–Schmidt, which shows they always exist, and completeness, , in Dirac notation introduced where it is nothing but bookkeeping about rows and columns. You have orthogonal projection as the nearest-point map, and with it the fact that least-squares regression is a right angle and ANOVA is Pythagoras. You have the adjoint defined without reference to any basis, and shown to be the conjugate transpose in an orthonormal one. And you have the spectral theorem: a Hermitian operator has real eigenvalues, perpendicular eigenvectors, and a complete orthonormal eigenbasis, so that it dissolves into . Functions of operators, the Hermitian–unitary correspondence , and simultaneous diagonalisation of commuting operators all came out of it as corollaries.
Where this gets spent. Cauchy–Schwarz goes to Chapter 4.9, where it becomes the uncertainty principle with nothing added but interpretation. Orthonormal bases and completeness go to Chapter 4.2, for superposition and probability amplitudes, and to Chapter 0.9, where Fourier analysis turns out to be an orthonormal expansion in a function space and Parseval becomes conservation of energy across modes. Adjoint and Hermiticity go to Chapter 4.2, where "an observable is a Hermitian operator" is the one postulate that has to be made. The spectral theorem goes to Chapter 4.2 for the measurement postulates, which are these theorems renamed, to Chapters 4.7 and 4.8, where solving a system means diagonalising its Hamiltonian, to Chapter 0.8, where normal modes of coupled oscillators are eigenvectors of a symmetric matrix, and to every PCA you will ever run. goes to Chapter 4.6, where a Hermitian Hamiltonian generates unitary time evolution and probability is conserved as a matter of algebra, and to Chapter 6.1, where the same pairing becomes the Lie algebra of a Lie group. Commuting observables go to Chapters 4.9 and 4.11, where they become the complete set of quantum numbers that labels every state of every atom.
One thing remains outstanding, and it is worth naming so that you notice when it is paid. Every proof in this chapter used finite dimension: in the induction, in rank–nullity, in the interchange of sums, in the claim that an injective map is surjective. Quantum mechanics happens in infinite dimensions. Chapters 4.4 and 4.5 are where the bill comes due.