Part IV · Quantum Mechanics — Chapter 4.2
The Linear Algebra of Quantum States
A table of renamings, seven postulates, and nothing else.
Chapter 4.1 ended with two statements. What has to change is what a state is, and the mathematics for that change was built nine chapters ago. This chapter makes good on both. It is the chapter Part 0 was a down-payment on. The claim it has to honour was written out in the opening paragraph of Chapter 0.5 in these words: "every postulate of quantum mechanics is a statement about Hermitian operators on an inner-product space. Not modelled by, not analogous to — is." That same chapter then told you what this one would look like: "Chapter 4.2 becomes a translation exercise: a table of renamings, not a new subject."
So the spine of what follows is a two-column table. On the left, a theorem of Chapter 0.5, with the number of the equation that states it. On the right, a sentence about measurement. The rows are not analogies, and the right-hand column is not an interpretation of the left. Each row is one statement written twice, in two vocabularies. Section 2 builds the table, and the rest of the chapter refers back to it.
What is genuinely new is short. That shortness is the reason this chapter is written differently from every chapter before it. Part I cornered the action principle out of Newton. Part II cornered Lorentz invariance out of two facts about light. Part III cornered the field equations out of free fall and one identity. Quantum mechanics cannot be cornered. The honest response is to say so out loud at each of the seven places where something is asserted rather than derived. Each of those seven gets its own box, before it is used, and each carries a ⚑ as well. The box says what is being claimed. The mark says it was not earned. This chapter therefore has the largest flag count in the book, and that is the point rather than a defect.
One of the seven matters more than the others. The Born rule, which is §5, is the first thing in twenty-nine chapters that is posited rather than cornered, and it is still open. It is not smuggled in as a definition. It is not derived from a plausible-sounding axiom, and it is not established by the theorem that comes closest. Section 5 says that plainly, at the moment it enters, and Chapter 4.20 returns to what it costs.
Conventions. SI throughout, with written explicitly everywhere. Part IV does not use natural units and does not adopt them until Chapter 5.1. Dirac notation is exactly as Chapter 0.5 §2.2 installed it: hats on operators, no hats on eigenvalues or classical quantities, for the identity. The inner product is linear in its second slot and conjugate-linear in its first, as Chapter 0.5 §1.1 chose and said it was choosing. Section 2.1 below restates why, because the Born rule is built on it. Everything proved in Chapter 0.5 is a finite-dimensional theorem. Where a statement here will need repair in infinite dimensions, the repair is named by chapter, and you are told which of the two mathematics chapters does it.
Tools you'll need — Chapter 0.5, all of it, and this chapter spends nothing else of comparable weight. Section 1 has the inner-product axioms and the choice of linear slot. Section 1.4 has Cauchy–Schwarz, used in §8 here and cashed as the uncertainty principle in Chapter 4.9. Section 2 has orthonormal bases, Parseval and the resolution of the identity. Section 3 has orthogonal projection and . Section 4 has the adjoint and the three species of operator. Section 6 has the spectral theorem in all three of its parts. Section 7 has functions of operators and . Section 8 has commuting operators, simultaneous diagonalisation and its figure fig-joint. And its Worked example 1 has a matrix that §10 here uses three times. Chapter 0.4 §2 for basis and dimension, which §9 needs to build the tensor product. Its §6.1 for the cyclic property of the trace, which §8 turns into a proof that this subject cannot be finite-dimensional. Its §7 for the one theorem Part 0 imported rather than proved. Chapter 1.3 §6 for Poisson brackets and the fundamental brackets , and §7 for the statement that a conserved quantity is the generator of its symmetry. Sections 7 and 8 here repeat both with operators. Chapter 4.1 §6.5 for and , and §6.6 for the tension this chapter is here to resolve. Chapter 2.5 §4.3 for , which §10.3 expands to get the neutrino oscillation length. Chapter 0.3 for Euler's formula, used in §10 without comment.
1 · The claim, stated before it is earned
Announce the destination. This section does no mathematics. It states exactly what the chapter will assert and exactly what it will not. It gives the total number of assumptions quantum mechanics requires, and it sets out the two rules by which the rest of Part IV should be read. Doing this first costs two pages. What it buys is the ability to check, at every later line, whether something was derived or put in by hand.
1.1 · What Chapter 0.5 claimed
Chapter 0.5 proved a sequence of theorems about operators on a complex inner-product space, with no physics anywhere in the argument. Then, in an insight box at the end of its §6, it read three of them back with two words changed:
"(a) Eigenvalues are real becomes: measurement outcomes are real numbers. … (b) Eigenvectors are orthogonal becomes: distinct outcomes are perfectly distinguishable. … (c) The eigenvectors form a complete orthonormal basis becomes: any state whatsoever can be written as a superposition of outcomes… These are three of the postulates of quantum mechanics as they are usually taught, and they are not postulates. They are theorems about Hermitian matrices, proved above, in a chapter containing no physics. Chapter 4.2 will do exactly one thing that this chapter did not: assert that a physical state is a unit vector and a physical observable is a Hermitian operator. That is one postulate, not four. Everything else is renaming."
That is the claim. This chapter either honours it, or the design of the book fails retrospectively. So the chapter is organised to make the claim checkable. Section 2 lays out the renamings as a table, and every later section either points at a row of that table or announces an assumption in a box. If you find a place where something has been slipped in without one or the other, the claim has failed there.
1.2 · The complete bill, said out loud
Here is the number, because a subject that has to assume things should be willing to say how many. Quantum mechanics as this book develops it rests on eight postulates and one experimental measurement. Seven of the eight arrive in this chapter. The eighth is the symmetrisation of identical-particle states, which arrives in Chapter 4.18. The measurement is that the electron carries half-integer angular momentum, which is Stern–Gerlach's, quoted in Chapter 4.12. Nothing else in Part IV is asserted without being derived. Chapter 4.20 §9 puts the whole list back on one page, with an honest statement of what each item does and does not settle.
| What is asserted | Where | |
|---|---|---|
| P1 | A pure state is a unit vector, defined only up to an overall phase | §3 |
| P2 | An observable is a Hermitian operator | §4 |
| P3 | The Born rule. The probability of outcome is | §5 |
| P4 | The state update. After outcome the state is | §6 |
| P5 | The generator of time evolution is the energy | §7 |
| P6 | Canonical quantisation. | §8 |
| P7 | A composite system's space is the tensor product of the parts' | §9 |
| P8 | Identical-particle states are totally symmetric or totally antisymmetric | Chapter 4.18 |
| E1 | Measured, not postulated: the electron carries | Chapter 4.12 |
Now set that against the list of things which are not on it. Every one of them is a theorem of Chapter 0.5 that you have already proved:
- that measured values are real numbers (0.5 §6.1).
- that distinct measured values belong to exactly perpendicular states (0.5 §6.2).
- that every state can be resolved into the possible outcomes of any observable (0.5 §6.3).
- that the probabilities of the outcomes add to one (0.5 §2, Parseval).
- that two quantities can be sharp together exactly when their operators commute (0.5 §8).
- that a norm-preserving evolution preserves every amplitude, so that conserving total probability and being unitary are the same condition (0.5 §4.3).
- that a Hermitian generator produces exactly such an evolution, and every such evolution has one (0.5 §7.1 and its grind box).
Seven theorems against seven assertions. The seven theorems are the ones usually presented as the mysterious part. That contrast is the payoff for nine chapters of mathematics, and it is the thing this chapter exists to make visible.
1.3 · Two rules for reading the rest of Part IV
Every postulate is announced in a box, before it is used, and carries a ⚑. Part III never had to do this, because general relativity was cornered rather than posited. There you had the equivalence principle plus one identity, and the field equations follow. Quantum mechanics has no such route. The response is not to hide that fact but to timestamp it. A box means this is being asserted. The ⚑ means the book uses this and does not derive it. Both, every time, in place.
Say each time whether a step is a derivation or an identification. These are different things, and Part IV moves between them from sentence to sentence. That is unitary when is Hermitian is algebra, proved in Chapter 0.5 §7.1. That the operator in that exponent is the energy is a physical identification with experiments behind it. Chapter 4.1 already ran this distinction once. The form of the cavity spectrum was derived and its normalisation was matched, and the chapter said so rather than letting the two blur. You have been trained for twenty-nine chapters to ask which is which, and this is the part where the answer changes most often.
Something changes in this chapter, and it is worth naming before it happens rather than afterwards. Everything so far has been cornered. Geometry was not chosen; it was what survived once you insisted that no observer's description could be privileged. Even the field equations of gravity were forced, in the sense that the alternatives were eliminated one by one until a single form was left standing. What arrives now cannot be got that way, and no honest telling pretends otherwise.
So the method changes. Certain statements are going to be put down as assumptions, each in its own box, each marked, each stated before anything leans on it. That is not a weakness in the presentation. It is the difference between a subject that was derived and a subject that was discovered, and confusing the two is how people come away believing quantum mechanics was deduced from something more familiar. It was not. It was guessed, tested for a century, and never yet found wrong.
What makes the chapter worth reading anyway is how little has to be assumed. The mathematics is finished — it was finished nine chapters ago, in a chapter with no physics in it at all. What gets added here is short enough to list on one page, and the rest of the part is that short list being spent.
2 · The table of renamings
Announce the destination. We set out, in one place, every result of Chapter 0.5 that this part of the book uses, beside the physical statement it becomes. Nothing in this section is assumed and nothing is proved. The proofs are all in Chapter 0.5, and each one is cited by equation. The purpose is to make the size of what has to be added visible before any of it is added.
2.1 · One convention, restated because everything rests on it
Chapter 0.5 §1.1 defined the inner product to be linear in its second argument and conjugate-linear in its first, and gave the reason in advance: "in Chapter 4.2 the object becomes a probability amplitude, read right-to-left as 'amplitude to find in state ', and the linear slot had better be the one holding the state that is actually evolving." That reason can now be stated properly rather than promised.
The number is going to be an amplitude, meaning a complex number whose squared modulus is a probability. The state that is evolving, being prepared, being superposed, is , and superposition is a linear operation on it. Split a beam into two paths and recombine it, and what comes out is . The amplitude to find that in must be the sum of the two amplitudes, or interference could not be computed at all. So the slot holding has to be the linear one.
The conjugation then lands on , and that is exactly what makes , and hence . Read that last equality in words. The probability of finding in equals the probability of finding in , which is a symmetry a probability had better have.
Mathematics texts make the other choice, and then every formula below acquires its conjugates on the other side. Nothing physical depends on which convention you pick. What does depend on it is arithmetic. Mixing the two conventions inside one calculation is the single most common way to produce a wrong sign.
2.2 · The table
Read it once from left to right and once from right to left. Left to right it is a dictionary. Right to left it is a list of physical facts, with the place each one was proved.
| Chapter 0.5 proved | Where | Which says, about a physical system |
|---|---|---|
| A complex inner-product space: vectors add and scale, and is linear in | §1 | Superposition. If two states are possible, so is any combination of them, and is the amplitude to find in |
| for | §1.1, axiom (iii) | Two different states can always be told apart: no non-zero state has zero length |
| Cauchy–Schwarz, | §1.4 | No probability exceeds one — and, in Chapter 4.9, the uncertainty relation, with nothing added but the meaning of the letters |
| Orthonormal basis , | §2 | A complete set of mutually exclusive outcomes of one measurement |
| §2 | Every state is a superposition of the outcomes, with the coefficients read off as single overlaps | |
| Parseval, | §2 | The probabilities add to one, and they add to one automatically rather than by a normalisation imposed afterwards |
| Resolution of the identity, | §2.3 | The set of outcomes has missed nothing: insert it anywhere and a hard amplitude becomes a sum of easy ones |
| Orthogonal projection , with and | §3 | An outcome. Idempotence becomes the statement that repeating a measurement immediately returns the same value |
| The adjoint, | §4 | The rule for moving an operator across an amplitude — used on every page of Part IV |
| Hermitian, | §4.2 | An observable |
| Unitary, , preserves every inner product | §4.3 | Evolution, and every symmetry: exactly the maps under which total probability is conserved |
| Eigenvalue , eigenvector , | §5 | A measured value, and the state that gives it with certainty |
| Over every operator has an eigenvalue | §5, from 0.4 §7 | Every observable has possible measured values. Over this is false, and that is why the space is complex |
| Hermitian every eigenvalue is real | §6.1 | Measured values are real numbers. A dial cannot read |
| Hermitian eigenvectors with different eigenvalues are orthogonal | §6.2 | Distinct outcomes are perfectly distinguishable. The chance of confusing one for another is zero, not small |
| The spectral theorem: an orthonormal eigenbasis exists | §6.3 | Any state at all is a superposition of the outcomes of any observable |
| , with and | §6.4 | An observable is its list of possible values together with the outcomes they belong to, and nothing else |
| Degeneracy: the eigenspace is unique, a basis inside it is not | §6.4 | Two states can share a measured value, and then that measurement does not determine the state |
| §7 | Functions of observables, including the one that generates motion | |
| Hermitian unitary, and conversely | §7.1 | Observables generate symmetries and evolution, and every symmetry has an observable behind it |
| a common orthonormal eigenbasis exists | §8 | Compatible measurements, and the quantum numbers that label a state |
| no common eigenbasis | §8 | There is no state in which both quantities are sharp — the qualitative half of Chapter 4.9 |
2.3 · What has been added so far, exactly
Nothing. Not one line of the table is an assumption about the world. Every left-hand entry is a theorem with a proof you have read, and every right-hand entry is that theorem with two or three words changed. Chapter 0.5 §6.5 put the count in its own words: "That is one postulate, not four. Everything else is renaming". The table is that sentence, itemised.
The physics enters at exactly seven places, and the sections that follow take them one at a time.
- Two of the seven, §3 and §4, are the pair Chapter 0.5 named: a state is a unit vector, an observable is a Hermitian operator.
- Two more, §5 and §6, are the ones Chapter 0.5 never had to mention, because a chapter of linear algebra has no reason to talk about probability.
- The last three, §§7–9, are dynamics, the relation between position and momentum, and how two systems combine.
That is the whole of it.
If you have met a list of "the postulates of quantum mechanics" before, compare it with §1.2's. The usual list runs to five or six items. It includes, as postulates, the statements that measurement outcomes are real, that the eigenstates of an observable form a complete orthonormal set, and that probabilities add to one. None of those three is on §1.2's list, because none of them is an assumption. They are 0.5 §6.1, 0.5 §6.3 and Parseval, and they were proved four parts ago in a chapter with no physics in it.
Notice which way round that leaves things. The items a reader finds strange are the theorems: real outcomes, perfect distinguishability, superposition, probabilities that sum to one automatically, and the impossibility of knowing two incompatible quantities together. The items that are genuinely assumed are dull-looking by comparison: a state is a unit vector, an observable is a Hermitian operator, evolution is generated by the energy. So the strangeness and the assumption sit in different places. The whole reason Chapter 0.5 was written where it was, four parts before it was needed, was to make that visible rather than assertable.
One item on the usual list survives intact, and §5 is about it.
Here is the debt being settled. Nine chapters ago a chapter of pure linear algebra was written with the announcement that it was quantum mechanics with the physics taken out, and that when the physics came back the work would already be done. What that amounts to is a table with two columns. On the left sits a theorem about self-partnered maps on a space with an inner product, proved with no physics in the room. On the right sits a sentence about measurement. The rows are not analogies and the right-hand column is not an interpretation of the left; each row is one statement appearing twice, in two vocabularies.
Read the table and notice how little is added. Measured values are real because the multipliers of such a map are real. Different outcomes are never mistaken for one another because the special directions belonging to different multipliers are exactly perpendicular. Any state at all is a combination of outcomes because there are enough of those directions to span everything. Weights adding to one is the statement that the squared length of a vector is the sum of the squared sizes of its coordinates.
What has to be supplied on top is small enough to count: seven assertions in this chapter, one more in the last chapter of this part, and one experimental fact. Everything else in the table was earned long ago.
3 · States are unit rays
Announce the destination. The first assertion. A state is a unit vector, and two unit vectors differing by an overall complex factor of modulus one describe the same state. The first clause is bookkeeping, and we show why. The second clause is physics, and we show what it forbids and what it permits. The section ends with a parameter count that turns the abstraction into something you can hold: a two-state system has two real parameters, not four.
Asserted. The physical states of a system are in one-to-one correspondence with the unit vectors of a complex inner-product space. One understanding comes with that: and for real are the same state, not two states that agree. Such an equivalence class is called a ray.
Not derived, and nothing here derives it. The vector-space structure is the whole content. It says that if and are possible states, then so is for any complex making the result a unit vector. That is the superposition principle, and it is an assertion about nature. Chapter 4.1 §6.6 described two descriptions of light, each supported by measurements the other cannot account for, and said the tension is resolved by changing what a state is. This is the change.
What is being deferred. Which inner-product space is not specified here, and cannot be. For a spin it is . For a particle on a line it is a space of functions that Chapter 4.3 builds and Chapters 4.4 and 4.5 supply operators for. Every theorem quoted below is Chapter 0.5's finite-dimensional one, and every use of it in infinite dimensions is a promissory note on those two chapters.
3.1 · Why the length is fixed at one, and why that is not a restriction
The reason to normalise is that the numbers we are about to call probabilities are squared overlaps, and Parseval says their sum is . So expand any state in an orthonormal basis , which Chapter 0.5 §2 permits:
Setting makes the right-hand sum equal to one, which is what a complete list of probabilities has to do. Notice the direction of the logic. Normalisation is chosen so that Parseval's identity reads as a statement about probability. Probability is not being normalised afterwards, by hand.
Scaling by any non-zero complex scales every by , and every by . So it changes no ratio and describes nothing new. A state is therefore a direction, and the unit-length convention picks one representative from each direction.
3.2 · The overall phase, and why it is invisible
Fixing the length leaves one freedom untouched. If then as well, since . So unit length does not pick a unique vector. P1 asserts that this leftover freedom is physically empty. The reason can be given completely only once §5 has supplied the Born rule. But the computation itself is one line, and it does not depend on anything §5 adds beyond the form of the expression, so we can do it now. For any projection ,
because the first slot conjugates (Chapter 0.5 §1.1) and . Every probability the theory produces is of this form, so every one of them is untouched. The same one-line argument kills the phase in every expectation value, since by the identical cancellation. There is no measurement in the theory that can see .
Verified. Taking the Hermitian matrix and the state , symbolic algebra gives the three probabilities , , both with and without a factor , identically in , and they sum to .
3.3 · The relative phase, which is everything
Now the contrast that makes §3.2 content rather than triviality. Multiply one term of a superposition by a phase and nothing cancels. Take a two-state system, the two states being the two paths of an interferometer, and compare
Both are unit vectors. Now we want the number a detector actually reports, so project onto the recombined state , which is what a detector at one output port of an interferometer responds to. The amplitude is , and squaring its modulus gives
which runs from at to at . That is a fringe, and it is what an interferometer measures. So a phase applied to every component is unobservable, and a phase applied to one component is the difference between full transmission and complete extinction.
This is not a subtlety to be filed away. It is the whole reason the theory needs complex numbers, and the whole content of the two-slit experiment. Chapter 4.1 §6.6 insisted that light's wave behaviour was untouched by the evidence for parcels of energy, and (4.2.4) is where the wave behaviour now lives: in the relative phase between the terms of a superposition.
3.4 · How many parameters a state actually has
Count, because the count is the cleanest way to feel the difference between a vector and a ray. A unit vector in has real components, subject to one real constraint , so the unit sphere has real dimensions. Identifying vectors that differ by a phase removes one more. So
For that is two, not four. Two real parameters is a sphere, and we can see the sphere directly by writing the general unit vector as
which uses the phase freedom to make the first coefficient real and non-negative, and that exhausts the freedom. Every state of a two-level system is (4.2.6) for exactly one , and are polar coordinates on a sphere. That sphere is the Bloch sphere. It is not a picture drawn for intuition. It is the state space, exactly, and §12's Problem 1 shows that the three expectation values are the Cartesian coordinates of the corresponding point. The sphere is therefore something you can measure.
For the count gives four real parameters, and there is no comparably nice picture. The two-level case is special, and its picture does not generalise. Better to say that out loud than to let the sphere become a mental model of quantum mechanics in general.
The first assertion is about what a state is, and it has two clauses that do different work. A state is a direction in the space, and its length is fixed at one. Fixing the length is not a restriction on nature but a choice of bookkeeping, because the weights are going to be squared lengths and they have to add to one; doubling every component would double every weight and describe nothing new.
The second clause is stranger and more consequential. Multiplying the whole state by a phase — a complex number of size one, applied to every component at once — changes nothing that can ever be measured, so states differing only by such a factor are the same state rather than two states that happen to agree. What is emphatically not invisible is a phase applied to one component and not another. That relative phase is the entire difference between two beams that reinforce and two that cancel, and it is what a wave description was carrying all along.
So the correct object is not a vector but a direction with its overall phase forgotten, and the counting bears this out: a two-state system has two real parameters to its name rather than four, which is why every account of one draws a sphere.
a natural place to stop · the dictionary is laid out; what follows is one postulate at a time
4 · Observables are Hermitian operators
Announce the destination. The second assertion, and the one Chapter 0.5 named in advance as "the one postulate that has to be made". We state it, then collect the three theorems it immediately buys without re-proving any of them. Then we read the spectral decomposition as a physical statement about what an observable is. Then we use Chapter 0.5 §8 to define a term the rest of Part IV lives on: a quantum number. One warning is attached to the postulate itself and cannot wait, because the word "Hermitian" is not strong enough for the spaces Chapter 4.3 builds.
Asserted. Every measurable quantity of a physical system is represented by a Hermitian operator on the state space, and the possible results of measuring it are the eigenvalues of and nothing else.
Not derived. This is Chapter 0.5's "an observable is a Hermitian operator is the one postulate that has to be made", collected. Notice what is being assumed and what is not. It is not being assumed that outcomes are real, or that different outcomes are distinguishable, or that the possible outcomes span the state space. Those are consequences, proved in Chapter 0.5 §6 with no physics in the argument, and §4.1 below reads them off.
⚠ The word will be sharpened, and the difference is physical. In a finite-dimensional space, Hermitian () is the right condition and nothing more need be said. In a space of functions it is not enough. An operator like cannot act on every vector of the space, because differentiating a square-integrable function need not give a square-integrable function. So it carries a domain, and its adjoint carries a domain of its own, fixed by the definition and with no reason to be the same one. The condition holding on a domain is called symmetric, and that is strictly weaker than self-adjoint, which requires the domains to match as well. Chapter 4.4 §4 makes the correction, and P2 should be read as saying self-adjoint from that point on.
Here is why that is not pedantry. Chapter 4.4 §5 shows that on the half-line is symmetric and has no self-adjoint extension whatever. So "the momentum of a particle confined to a half-line" is not an observable at all. A physical conclusion, from a domain.
4.1 · Three theorems, collected rather than re-proved
With P2 in place, the following are facts about nature that were established in a chapter containing no physics. Each is cited at the section that states it, and none is proved again here. Proving them again would be a misreading of what this chapter is.
Measured values are real. Chapter 0.5 §6.1 showed that a Hermitian operator's eigenvalues satisfy . A dial reads a real number, and now it has to.
Distinct measured values belong to orthogonal states. Chapter 0.5 §6.2 showed that eigenvectors with satisfy . Section 5 will make the probability of finding one where the other is, so that probability is exactly zero. Not small, and not zero to within experimental error. Two states with different definite values of the same observable are perfectly distinguishable in a single measurement.
Every state is a superposition of outcomes. Chapter 0.5 §6.3 produced an orthonormal basis of eigenvectors. So for any state and any observable whatever, with by Parseval. The probabilities of the outcomes add to one automatically, and no separate assumption is needed to make them.
4.2 · What an observable is, read off the spectral decomposition
We want the form of the decomposition that will still be true when the space becomes infinite dimensional, so group Chapter 0.5 §6.4's decomposition by distinct eigenvalue rather than by basis vector. That is the form Part IV uses throughout:
with the real and distinct. Let's read (4.2.7) as a definition of the object rather than as a factorisation of it. An observable is a list of real numbers, together with a set of mutually orthogonal outcomes that exhaust the space. It is a labelled partition of the state space into perpendicular pieces. The operator is not something extra that the list has. The list is the operator. Two observables are the same observable exactly when they have the same list, which is why in practice one specifies a measurement by saying what its possible answers are and which states give each of them with certainty.
The degeneracy clause carries physics. If has rank greater than one, several independent states share the value , and Chapter 0.5 §6.4 was careful that the eigenspace is unique but a basis inside it is not, and no theorem prefers one. Physically: that measurement, on its own, does not determine the state. Something else has to be measured, and the next subsection says exactly what "something else" is allowed to be.
4.3 · Compatible observables, and the definition of a quantum number
Chapter 0.5 §8 proved that two Hermitian operators admit a common orthonormal eigenbasis if and only if they commute. Rename both halves.
If there are states in which both quantities have definite values at once, and enough of them to span the space. Call two such observables compatible. Chapter 0.5 §8.2's Step 4 is where the content sits. Inside a degenerate eigenspace of , where has said all it can say, the restriction of is still Hermitian. So the spectral theorem applies inside that eigenspace, and chooses the basis could not.
If there is no common eigenbasis, so there is no state at all in which both quantities are sharp. That is the qualitative content of the uncertainty principle. Chapter 4.9 supplies the quantitative version by applying Cauchy–Schwarz to two particular vectors. That is a calculation Chapter 0.5 §1.4 already set up, and it adds nothing but the meaning of the symbols.
Now the definition this chapter owes the rest of the part. A set of mutually commuting observables is a complete set of commuting observables when their common eigenspaces are all one-dimensional. Said the other way round, the list of eigenvalues determines the state uniquely up to phase. That list is called the quantum numbers of the state. Writing a hydrogen state as is naming the eigenvalues of three commuting operators and nothing more. Chapter 0.5's insight box promised the term to Chapter 4.11. The term is defined here instead, where the machinery arrives. Chapter 4.9 §3 asks how one knows a set is complete, and Chapters 4.11 and 4.13 spend it on angular momentum and on hydrogen.
One caution, and Chapter 0.5's figure fig-joint is where to see it rather than have it described again. The theorem promises that a common eigenbasis exists. It does not promise that the basis a calculation happens to produce is that one. Drive the commutator of two operators to exactly zero, and a standard eigensolver will still hand back states in which the second observable has a spread. Inside a degenerate eigenspace it returned an eigenbasis of the first, rather than the one the second prefers. Go and move the sliders. The point is worth the minute, and it is exactly what Chapter 4.13's degeneracies will need.
4.4 · Which operators are observables, and which are not
P2 says every observable is Hermitian. It does not say every Hermitian operator corresponds to something an apparatus can measure, and that converse is genuinely open. For a system with superselection rules, or for gauge-dependent quantities, some Hermitian operators correspond to nothing measurable. This book does not need the converse and does not assert it. What it does need is the direction P2 states, plus one immediate negative consequence worth recording now.
A unitary operator is not, in general, an observable. Chapter 0.5's Problem 2 showed that its eigenvalues have modulus one, so they are complex unless they are , and P2 requires real outcomes. Unitary operators are how the state changes. Hermitian operators are what can be read. Chapter 0.5 §7.1's correspondence says the two classes carry the same information, and §7 below is the section that spends it.
The second assertion is the one the whole design was aimed at. A measurable quantity is represented by a map that is its own partner, and that single sentence is the last thing needing to be assumed about measurement outcomes. Everything usually presented alongside it as a separate postulate is a theorem already proved: that the readings are real numbers, that distinct readings are perfectly distinguishable rather than merely different, that any state can be resolved into the possible readings, that the weights add to one.
One honest qualification belongs here rather than later. The word used is Hermitian, and in a space of functions that word is not quite strong enough — the stronger property involves specifying which functions the map is allowed to act on, and the difference is physical rather than pedantic. A particle confined to a half-line turns out to have no momentum observable at all. The repair has its own chapter.
The other dividend is labels. Two quantities can be sharp at once exactly when the order of the two maps does not matter, and when there are enough such quantities to leave no ambiguity, the list of their values names the state uniquely. Those numbers are what atoms are catalogued by.
5 · The Born rule
Announce the destination. This is the section the book's honesty rests on. We state the rule connecting a state to the probabilities of outcomes. We say without hedging that it is not derivable, and that nothing in this book or outside it has derived it. We deal properly with the theorem that comes closest. Then we derive from the rule everything that is derivable from it: that probabilities are non-negative and sum to one, that overall phase is invisible, and that the average and the spread of a measurement are the two short expressions everyone uses.
5.1 · Why nothing so far has produced a probability
Look at what P1 and P2 have supplied. There is a state, which is a unit ray. There is an observable, which by (4.2.7) is a set of real values with orthogonal outcomes summing to the identity. The state resolves into those outcomes, , since . The pieces are mutually orthogonal, so Chapter 0.5 §1.3's Pythagoras applies with no cross terms, and their squared lengths add to one:
the cross terms vanishing because by Hermiticity and orthogonality of the projections. So there is a list of non-negative numbers, one per possible outcome, adding to one, sitting in the formalism already.
And that is as far as mathematics goes. A list of non-negative numbers summing to one is consistent with being a probability distribution. It is not a demonstration that it is one. Nothing in Chapter 0.5, and nothing in P1 or P2, connects any number in the theory to the relative frequency of an outcome in a long run of experiments. That connection has to be asserted, and here it is.
Asserted. A system in state , measured for the observable , yields the value with probability
When is non-degenerate this is , the squared modulus of a single amplitude. When it is degenerate it is the sum of the squared moduli over any orthonormal basis of the eigenspace, and the projection form is what makes that independent of which basis you pick.
This is the first thing in this book that is posited rather than cornered, and it is permanently open. Every earlier ⚑ in twenty-nine chapters marked something established elsewhere that this book chose not to establish, whether a theorem of mathematics, a measurement, or a result belonging to a subject not built here. P3 is none of those. It cannot be deduced from P1 and P2. It does not follow from the linearity of the theory. And no argument in physics has established it: the literature contains many attempted derivations, and each assumes something equivalent to what it concludes. The mark on this box is therefore of a different kind from most in this book. Elsewhere ⚑ means this book chose not to prove this. Here it means nobody has proved this. Chapter 4.20 §9 returns to what that costs, and states precisely which part of the problem decoherence addresses and which part it does not touch.
⚑ Gleason's theorem, with its hypotheses, because it is the thing that comes closest. Statement: work on a complex Hilbert space of dimension at least three. Take any function assigning a number in to every orthogonal projection, subject to and for any countable family of mutually orthogonal projections. Then has the form for a unique positive operator of unit trace. For a pure state that is exactly P3. It is a genuine theorem, and it is not proved here.
What it does and does not do: look at how much its hypotheses already assume. They assume that probabilities are assigned to projections. They assume those probabilities are additive over orthogonal projections. And they assume the dimension is at least three. That last condition is not a technicality. The two-dimensional case is a genuine counterexample, and spin- is two-dimensional. Assume all of that and the squared length is forced. But assuming all of that is assuming most of what one wanted explained. It narrows the freedom inside a framework rather than deriving the framework. It is not a derivation of the Born rule, and this book does not present it as one.
5.2 · What follows immediately, and what does not
The probabilities are legitimate. Each is a squared norm, hence real and non-negative, and (4.2.8) says they add to one. Neither fact was assumed in P3. Both are Chapter 0.5's, and P3 was written in the form precisely so that they would be automatic.
The overall phase is invisible. Proved in (4.2.2), and it is worth noticing that §3.2's computation was written before P3 arrived and needed only the form . P1 and P3 fit together exactly. Had the rule been anything not quadratic in , P1's identification of with would have been inconsistent.
What does not follow: which outcome occurs. P3 gives the distribution and says nothing about the individual run. This is not a gap the chapter will close.
5.3 · The expectation value and the variance, derived
These are the two quantities actually reported by an experiment, and both come out of P3 in two lines each. We want the mean of the measured value over many repetitions. By P3 that mean is the weighted average . So substitute P3 into it, and then run (4.2.7) backwards:
The middle step used , which is and from Chapter 0.5 §3. The last step is the spectral decomposition read from right to left. So the ubiquitous formula is derived from P3. It is not a separate definition of what "average" means.
The spread comes the same way. The variance of the measured value is . We want that in operator form, so expand the square, then use (which is Chapter 0.5 §7's function-of-an-operator with ) together with :
The second form is the one Chapter 4.9 needs, because it exhibits as the length of a vector, and Cauchy–Schwarz is a statement about lengths. That is the whole reason the uncertainty principle is one line there rather than a new idea.
Verified. With and , computing the two projections and summing gives , against , and against . Exact rationals both ways.
5.4 · Amplitudes are not probabilities
A superposition is not a mixture. The state is not a description of a system that is really in or really in with equal chance and we do not know which. Those two situations give identical predictions for a measurement in the basis, where both give and . They give different predictions for every other measurement, and that is what makes the difference a fact rather than a preference.
Take the observable whose outcomes are . For the superposition, : the outcome occurs every time. For the mixture, the answer is . Certainty against a coin flip. The arithmetic of the two differs because in the first case the amplitudes are added and then squared, and in the second the squares are added. Amplitudes are complex numbers, so adding them first allows cancellation that adding squares can never produce.
This is why the theory needs a vector space over and not a probability distribution over a list of possibilities, and it is the single most common place for a reader to install a picture that will have to be removed later. The formal object describing a genuine mixture exists, and it is the density operator. It is built in Chapter 4.19, and Chapter 4.20 spends it on decoherence, which is where this distinction does its work.
This is where the book stops cornering and starts assuming, and the moment deserves to be marked rather than slipped past. Everything so far has been forced: the action principle out of Newton, the geometry of spacetime out of two facts about light, the field equations out of free fall and one identity. The rule connecting a state to the odds of each outcome is not like that. It is put in by hand, because nothing available implies it and nothing since has managed to.
The rule itself is short. Project the state onto the directions belonging to a given reading, and the squared length of what survives is the probability of getting that reading. Non-negative because it is a squared length; adding to one because the projections reconstruct the whole vector with nothing left over; unaffected by an overall phase because that phase has size one and squaring removes it.
There is a celebrated theorem which comes close to deriving it, and its hypotheses are worth naming because they are where the content hides: it assumes that probabilities are already assigned to projections, additively, and it needs at least three dimensions. Assume that much and the squared length is forced. Which is to say the theorem explains the formula given the framework, and the framework is most of what one wanted explained.
6 · Measurement, and the state after it
Announce the destination. A fourth assertion is needed, and most books fold it into the third. We keep them apart, because they are logically independent, and because the separation is what lets Chapter 4.20 say precisely which of the two decoherence addresses. We then derive the one consequence that makes the postulate testable, which is that an immediately repeated measurement returns the same value with certainty. Next we set the algebra beside the conditioning you do every day, and say exactly where the two part company. We end by naming the one imported theorem that all four postulates rest on.
6.1 · Why P3 is not enough
P3 says what the odds are. It says nothing whatever about the system afterwards, and something has to be said, because measurements are performed in sequence. An apparatus is calibrated by measuring the same thing twice. A Stern–Gerlach beam is sent through a second magnet. A qubit is read out and then used. Without a rule for the post-measurement state, the theory cannot predict the second measurement at all.
Asserted. If a measurement of on the state yields the value , then immediately afterwards the state of the system is
Not derived, and independent of P3. One could consistently write down a theory with P3 and a different update rule, or with P3 and no update rule at all for measurements never repeated. The division of labour matters: P3 is about frequencies in an ensemble. P4 is about an individual system after an individual outcome. Most treatments merge them into one "measurement postulate", and thereby lose the ability to say which half is in difficulty. Chapter 4.20 §9 needs the separation. Decoherence gives a good account of why the alternatives stop interfering, which is a statement about P3's probabilities becoming classical. It gives no account at all of why one of them happens, and that is P4's clause.
The denominator is not a second assumption. is exactly the square root of P3's probability, so the rescaling is forced by P1's requirement that a state be a unit vector. And the update is undefined precisely when , which is when the outcome has probability zero and therefore does not occur.
6.2 · Repeatability, derived
The postulate has one immediate consequence that is a genuine prediction, and it costs three lines. Measure again straight away, on the updated state . By P3 the probability of getting is , and
using from (4.2.7). So the probability is for and otherwise: an immediately repeated measurement returns the same value with certainty. That is a fact about laboratories, and it is what makes the word "measurement" mean anything. It comes from , which is Chapter 0.5 §3's idempotence, proved there as the algebraic statement that projecting twice is projecting once.
Notice also which vector is when is non-degenerate. Then has rank one, so , and the update gives up to a phase, which by P1 is . The state after the measurement is the eigenstate, and every memory of except the fact of the outcome is gone. When is degenerate the update keeps more. It keeps the component of inside the eigenspace, direction and all. That distinction is not decorative. It is what makes a measurement of a degenerate observable weaker than a measurement of a complete set, and it is the reason Chapter 4.13's hydrogen states need three labels.
The algebra of P4 is one you perform daily. Given a joint distribution and an observation, you restrict to the cases consistent with what was seen and divide by the total weight of what survived, so that the remaining probabilities add to one again: . Restrict, then renormalise. P4 is the same two operations in the same order. Here is the restriction, and is the weight of what survived. The correspondence is exact, symbol for symbol, including the fact that both are undefined when the observed event had probability zero. When you compute a positive predictive value from a prior and a test result, you are doing P4 with a diagonal projection.
Where it stops, and this is the whole of the difference. Conditioning restricts a probability. P4 restricts an amplitude. Probabilities are non-negative, so discarding branches can only remove weight, and a branch that survives conditioning contributes what it contributes. Amplitudes are complex, so two surviving branches can cancel. That is why the ordinary law of total probability fails here. For a classical variable always holds. Worked example 1 exhibits a case where the quantum analogue gives if the intermediate quantity is measured and exactly if it is not. Nothing about the second run differs. The first apparatus was switched on, and that changed the answer from certainty to a coin flip. No amount of care with priors reproduces that, because the object being conditioned is not a probability.
The second difference is more subtle, and worth naming because it is the one people carry the wrong intuition about. Conditioning tells you something about a variable that already had a value. P4 makes no such claim, and §5.4 showed why: a superposition is not an unknown value. So the update is a change of state and not merely a change of information, and every attempt to read it as the latter runs into Chapter 4.20's inequalities.
6.3 · What is not settled, named
P3 and P4 together let you predict every laboratory result in this book. They do not explain anything about the process they describe. Three questions are left completely open, and it is better to list them than to let them accumulate as unease.
What counts as a measurement? P4 refers to "a measurement" as a primitive. The apparatus is itself made of atoms and should be describable by P1 and P5 like anything else, in which case nothing in the joint evolution ever produces a single outcome. This is the measurement problem and this book does not solve it.
Why one outcome rather than another? P3 supplies the distribution and P4 the conditional state. Neither says what selects the individual result.
When does the update happen? P4 says "immediately afterwards", which is not a time. Chapter 4.20 §9 states which of these decoherence answers, essentially the first and only in part, and which it leaves untouched.
6.4 · What all four postulates rest on
One structural remark, and it is owed. Every measurement postulate above is a statement about the projections of an observable, and those exist because of Chapter 0.5's spectral theorem. Look at Step 1 of that theorem's proof: "Because is complex, [Chapter 0.4] supplies an eigenvalue." Chapter 0.4 §7 supplied it from the fundamental theorem of algebra, which says that every non-constant complex polynomial has a root. That chapter flagged the theorem as the one result Part 0 imported rather than proved ⚑.
So the whole apparatus of measurement in this book descends from one imported theorem. That is a small debt, and it is worth knowing exactly where it sits, because it is not a hidden one. Without it, an observable might have no eigenvalues at all, and there would be no possible measured values to assign probabilities to. Chapter 0.4 §7 already made the point in those words. The debt is paid in Chapter 5.4, where complex analysis is built and the fundamental theorem of algebra becomes a three-line corollary of Liouville's theorem. Until then it is quoted, and it is the only thing in this chapter that is quoted rather than either derived or postulated.
A second assertion is needed and most accounts fold it into the first, which costs them the ability to say later which of the two is in trouble. The first says what the odds are. The second says what is left afterwards: whatever part of the state pointed along the directions belonging to the reading obtained, rescaled back to length one. The rest is gone.
The mathematics of that is something done every day. Restrict attention to the cases consistent with what was observed, then renormalise so the surviving weights add to one again — this is conditioning, and the algebra is identical, line for line. The break comes from what is being restricted. Here it is amplitudes rather than probabilities, and amplitudes can cancel, so a component that survives conditioning may still be annulled by another that also survived. Send a beam through a filter, and whether a second filter passes anything depends on whether the first one was looked at.
What none of this settles is why anything is left at all — why one reading occurs rather than the others, and what physical process the rescaling describes. That question is not answered in this chapter or in this book, and the last chapter of the part says exactly which part of it decoherence addresses and which part it leaves alone.
a natural place to stop · states, observables and measurement are in place; what follows is motion
7 · Evolution has a Hermitian generator
Announce the destination. Here almost everything is derived, and the assertion is a single clause. In four steps we show that time evolution must be unitary, that a unitary flow must be the exponential of a Hermitian operator, and that the operator therefore exists and has the dimensions of an energy. Then we state the one thing that does not follow, which is that it is the energy. Differentiating the result gives the Schrödinger equation, which Chapter 4.6 takes seriously. The section closes by collecting Chapter 1.3's promise that observables generate symmetries, with operators in place of functions and every word unchanged.
7.1 · Step one: evolution preserves the norm
Write for the map taking the state at time to the state at time , and require two things of it.
The first is linearity. If a system can be prepared in or , P1 says it can be prepared in any superposition, and evolution must take that superposition to the corresponding superposition of the evolved states. Otherwise the two beams of an interferometer, prepared together and travelling together, could not be treated as they are.
The second is conservation of total probability. At every time, the outcomes of any measurement must have probabilities summing to one, which by §5.2 is the statement . Put the two together, and is a linear map preserving the norm of every state.
7.2 · Step two: norm-preserving and linear means unitary
Preserving lengths sounds weaker than preserving inner products. It is not. The missing inner-product information can be extracted from lengths alone, and that extraction is worth doing on the page, because it is three lines and it is the hinge of the section. Apply Chapter 0.5 §1.3's expansion to , which is by linearity, and use three times:
so . That is the real part. Now we want the imaginary part, so run the same line with replaced by . The second slot is the linear one, so , and , which gives . Both parts agree, so
The equivalence on the right is Chapter 0.5 §4.3 read backwards. We have by the definition of the adjoint, and an operator whose inner products with everything agree with the identity's is the identity. So time evolution is unitary, and this is a consequence rather than an assumption. Chapter 0.5 §4.3 said the same thing forward, in the words " must be unitary because is a total probability". Here is the argument it was pointing at.
7.3 · Step three: a unitary flow has a Hermitian generator
Now bring in time. Take a system whose physical arrangement is not itself changing, with no field being switched on and no apparatus being moved. For such a system, evolving for and then for must be the same as evolving for , since there is nothing to distinguish the intermediate instant. So
We want a differential equation rather than a functional one, so differentiate (4.2.14) with respect to at . The left side gives and the right gives , so writing ,
the solution being the matrix exponential , which converges for every operator and is the unique solution with . Once is identified below as with Hermitian, that same exponential is Chapter 0.5 §7's function of an operator with .
What kind of operator is ? Unitarity is the condition we have not yet spent, so differentiate at , using :
so is anti-Hermitian. Then is Hermitian, because using from Chapter 0.5 §4.1. Every anti-Hermitian operator is times a Hermitian one and conversely, which is the same correspondence as Chapter 0.5 §7.1's , seen at the level of generators rather than group elements.
Now the dimensions. has the dimensions of one over time, since has to be dimensionless. Multiply and divide by the constant with the dimensions of action that Chapter 4.1 produced, writing . Then is Hermitian, and it has the dimensions of , which is energy. Nothing has been assumed to get here. The existence of a Hermitian operator with the dimensions of energy generating the evolution is forced by unitarity, the group law and dimensional analysis.
7.4 · Step four: the one thing that does not follow
Asserted. The Hermitian operator of (4.2.17) is the observable representing the system's energy. It is written and called the Hamiltonian, and time evolution is
Exactly half of this was derived. That evolution has the form for some Hermitian of energy dimensions is §7.3, with no physics beyond linearity and conservation of probability. That is the energy is the other half. The energy here means the same quantity a calorimeter measures, the same quantity Chapter 1.3 built as the Legendre transform of the Lagrangian, the same quantity Chapter 4.1's counted in parcels. Saying is that quantity is an identification with experiments behind it, and it is the postulate. Chapter 4.6 §2 states the sign convention loudly and shows what the identification buys. The rigorous infinite-dimensional version of §7.3, where "differentiate the group law" needs care, is Stone's theorem, quoted in Chapter 4.5 §9.
The hypothesis that was used, named. (4.2.14) assumed the physical arrangement does not change during the evolution. Think of a system being driven, such as an atom in a laser pulse or a spin in a swept magnetic field. There depends on time, the group law fails, and is wrong. Chapter 4.17 §3 handles that case, where the exponential is replaced by a time-ordered series.
Differentiating (4.2.17) with and multiplying by gives the equation the next chapter is named after:
Chapter 0.5's closing insight box named this exact equation as where its would be spent, and said the sentence "energy is observable, therefore probability is conserved" is not a physical argument but a line of algebra. That is now literally true. (4.2.13) is the algebra, and P5 is the only physics in the paragraph.
7.5 · Observables generate symmetries: Chapter 1.3, with operators
Nothing in §7.3 used the fact that the parameter was time. Any Hermitian generates a one-parameter family of unitaries , and the statement that this family is a symmetry of the system is the statement that it leaves alone. So let's work out what such a family does to an arbitrary observable, to first order in . Expand the exponential, keeping terms to first order:
the two cross terms combining into the commutator because by definition. Subtracting leaves the change itself, which is the object we are after:
Set that beside Chapter 1.3 §7, which proved for classical mechanics that the change in under the flow generated by is . The two are the same statement under the single replacement
which is precisely the substitution Chapter 1.3 §6.4 announced in advance. Taking and turns (4.2.20) into the equation of motion for an observable, , which is Chapter 1.3's with the same replacement.
Now read the same line the other way. The condition says two things at once: that is conserved, and that is unchanged by the symmetry generates. A conserved quantity is the generator of its symmetry, in operator form, which is what Chapter 1.3 said this chapter would repeat word for word.
Two things are being collected there, and it is worth separating them. The algebraic statement (4.2.20) is derived, from §7.3 and nothing else. The claim that (4.2.21) is the right correspondence for every pair of classical observables is a different animal. It is not derived, it is not assumed here either, and it is false. Section 8 postulates it for the single pair it needs, and Chapter 4.10 §8 proves that it cannot be extended consistently to all of them.
A two-compartment pharmacokinetic model, or a two-state Markov model of treatment response, is a vector of amounts obeying , with a real matrix whose columns sum to zero. Its solution is , a matrix exponential, evaluated exactly as Chapter 0.5 §7 evaluates one, by diagonalising and exponentiating the eigenvalues. Everything structural in §7.3 is already in that calculation: a one-parameter family, a group law , a generator read off as the derivative at zero, and a conserved total because the columns sum to zero. (4.2.18) is that equation with replaced by , and nothing else about the machinery changes.
Where the two part company is exactly the , and the consequence is total. A rate matrix has eigenvalues with non-positive real part, and real ones for the two-state case here. So has factors . The compartments relax, the transients die, and the system settles into a steady state that does not remember how it started. An anti-Hermitian generator has purely imaginary eigenvalues, so the factors are , of modulus one at every time. Nothing decays, nothing relaxes, and there is no steady state at all. The components hold their sizes forever, and only their relative phases move.
That is why §10's two-state system oscillates undamped instead of equilibrating at fifty-fifty as the corresponding Markov chain does, and it is the entire content of requiring probability to be conserved rather than merely bounded. Suppose you want relaxation back, and real atoms do relax. It has to come from coupling to something else that has been left out of the state. That is Chapter 4.20's subject, and it is not available by adjusting .
Now motion, and here almost everything is derived rather than assumed. Total probability must remain one for as long as the system exists, so whatever moves a state forward in time cannot change its length, and a map that preserves all lengths preserves all overlaps and is a rotation of the space. That much is not a physical hypothesis; it is the bookkeeping of the state postulate followed to its conclusion.
Then the correspondence from the toolkit chapter takes over. Rotations of this kind are exactly the exponentials of self-partnered maps, so there is a generator, and it is an observable. The whole question is which one, and the answer is the only genuinely new thing here: the generator is the energy. Nothing forces that. It is an identification with an experiment behind it, in exactly the sense that the constant relating colour to energy was an identification in the last chapter, and it is worth marking as such because it is the single input.
Differentiate the rotation and an equation of motion appears, which the next chapter takes seriously. The pattern underneath is one already met in the classical setting: a conserved quantity and the motion it generates are one object, and putting operators in place of functions repeats every word.
8 · Canonical quantisation
Announce the destination. Position and momentum have to be related to each other somehow, and the classical theory already supplies the relation. Chapter 1.3 computed the Poisson bracket of a coordinate with its conjugate momentum and got one. We postulate the operator version, under §7.5's substitution, for that pair alone. Then we spend three lines proving something that decides the shape of the next two chapters. No finite-dimensional space can carry that relation. Quantum mechanics is infinite-dimensional before any physics is done, and you can prove it here with Chapter 0.4's trace.
8.1 · The classical relation, and the operator version
Chapter 1.3 §6 built the Poisson bracket and computed the fundamental brackets
which say that momentum generates translation of position, and that neither generates any motion of its own kind. Section 7.5 showed that the operator analogue of a Poisson bracket is , derived for a generator acting on an observable. Applying it to (4.2.22) gives the relation this part runs on. It is an assumption, and here is its box.
Asserted. The observables representing the Cartesian coordinates of a particle and their conjugate momenta satisfy
Not derived. Chapter 1.3 §6.4 announced the substitution and named Chapter 4.9 as where it would be taken seriously. The commutator itself is postulated here, because Chapters 4.6, 4.8 and 4.11 all need it before Chapter 4.9 arrives. What Chapter 4.10 §8 supplies is the sharper and more interesting statement, and that statement is negative. The substitution cannot be extended consistently to every classical observable at once. Assign operators to all polynomials in and obeying it, and you obtain a contradiction. So P6 is a postulate about this one pair and not a general dictionary, and it should not be read as one.
What it does not fix. P6 says nothing about which space, which functions, or what and look like. Chapter 4.6 supplies the standard realisation acting on functions of position, and §8.4 below says why the realisation cannot be dodged. Notice also that the relation is dimensionally forced once it is assumed to be a multiple of the identity. The bracket has the dimensions of position times momentum, which is action. Chapter 4.1 §5.7 already observed that the new constant has exactly those dimensions, and that it is exactly the phase-space unit classical statistical mechanics was missing.
8.2 · Three lines that decide the next two chapters
Suppose the state space were finite-dimensional, of dimension , so that and are matrices. Take the trace of both sides of P6's first relation. Chapter 0.4 §6.1 proved that the trace is cyclic, , for any two square matrices of matching size. Hence the trace of any commutator vanishes:
Now take the trace of the right-hand side. The identity matrix in dimension has ones on its diagonal, so
Two numbers that must be equal are and . There are no matrices satisfying P6, for any whatever.
Stop and look at what that argument is. It uses one fact about the trace, proved in Chapter 0.4 in a chapter about determinants and change of basis, with no analysis, no limits and no physics. It is complete. And it says that the state space of a particle with a position cannot be for any . The reason is not that is too coarse an approximation, and not that a continuum is more elegant. It is that the arithmetic is impossible. The infinite-dimensionality of quantum mechanics is forced by one postulate and one line of linear algebra, before any physics has been done at all.
8.3 · How badly finite matrices fail, quantified
"Impossible" invites the response that some large matrix must come close, and the question deserves a number rather than a reassurance. It does not come close, and the bound is exact.
Use the inner product on matrices from Chapter 0.5 §1.2's table, , whose norm is . Apply Chapter 0.5 §1.4's Cauchy–Schwarz inequality in that space, to and any matrix :
Now we want to apply that to the miss itself, so set , the amount by which a candidate pair fails. By (4.2.23) its trace is , so , and (4.2.25) gives , that is
for every pair of matrices and every . Read the right-hand equality. The error is at least as large as the thing being reproduced. In relative terms the miss is at least , at every size, so a finite matrix pair cannot get the canonical commutator even half right.
Verified numerically, and the bound is attained. Minimising directly over all complex pairs by Nelder–Mead from twelve random starts bottoms out at for , so the bound is right and it is attained. Problem 3 identifies the equality case, and it is worth having in advance. Equality needs parallel to , which forces . So the best a finite-dimensional model can do is to commute, that is to reproduce none of the relation at all, and every pair that makes the commutator non-zero does strictly worse. For the natural truncation with a nearest-neighbour , the trace of the commutator is to machine precision at and , exactly as (4.2.23) requires.
8.4 · Why spin is not a counterexample, and what has to happen next
Here is the objection to raise. Chapter 4.11 will describe electron spin with matrices, and that is finite-dimensional. It is, and no contradiction arises, because a spin has no position operator. The three spin components satisfy , whose right-hand side is not a multiple of the identity. Its trace therefore can vanish, and does, since each is traceless. Section 8.2 rules out finite dimensions only for a system carrying a canonically conjugate pair. That it rules them out so cleanly while leaving spin alone is a check on the argument rather than a limitation of it.
What §8.2 forces, then, is this. The space of states for a particle has to be infinite-dimensional, and every theorem quoted in this chapter was proved in Chapter 0.5 in finite dimensions. Chapter 0.5 was explicit about the four places its proofs used that hypothesis, naming them as "in the induction, in rank–nullity, in the interchange of sums, in the claim that an injective map is surjective", and it said the bill would come due. It comes due in two instalments, and the division is worth stating precisely rather than gesturing at:
- Chapter 4.3 builds the space. The Riemann integral is discarded and rebuilt as the Lebesgue integral. That is what makes complete, so that a limit of states is a state, and what makes the Fourier modes genuinely an orthonormal basis rather than an assumption. Everything in §§1–3 of this chapter that used "an orthonormal basis exists" is on credit until then.
- Chapters 4.4 and 4.5 build the operators on it. Unbounded operators, domains, the difference between symmetric and self-adjoint that P2's box already warned about, spectra with no eigenvectors, the spectral theorem in the form that survives, and the meaning of and . Everything in §§4–7 that used "the spectral theorem" is on credit until then.
That shape is not new. Chapter 0.4 built the space and Chapter 0.5 built the operators on it. Chapter 4.3 builds the space and Chapters 4.4 and 4.5 build the operators on it, in the same order and for the same reason.
Position and momentum need a relation to each other, and the classical theory supplies one already: the bracket that measured how a quantity changes under the flow another one generates gave a particular answer for those two. The assertion here is that the same relation holds with the bracket replaced by the failure of two maps to commute, times a constant. That is a real assumption, and whether it can be extended to every quantity at once is a question with a sharp and negative answer several chapters ahead.
What the assertion costs is immediate and enormous, and the argument is three lines long. Adding up the diagonal entries of a product does not care about the order of multiplication, so the diagonal sum of the failure-to-commute is zero for any two square arrays whatever. The relation demands that it equal a fixed non-zero number on every diagonal entry, whose sum is therefore not zero. No finite list of numbers can do this. Not approximately either: the best possible attempt at any size misses by as much as the target itself.
So the space cannot be finite-dimensional, and this is settled before any physics is done. Two chapters of mathematics follow, and they were not chosen for thoroughness. They were forced here.
9 · Two systems: the tensor product
Announce the destination. Everything so far has described one system. Two systems require a rule for combining their state spaces, and there are only two candidates: the dimensions add, or the dimensions multiply. We build the object in which they multiply out of Chapter 0.4's basis and dimension, and it turns out to be a basis of pairs and nothing more. We postulate that it is the right one, and then we count. The counting is what makes entanglement a fact about arithmetic rather than a mystery, and it is the whole of what Chapter 4.19 needs from this chapter.
9.1 · A basis of pairs
Let system have state space with orthonormal basis and system have with . Whatever the joint space is, it must contain a state for each way of specifying both systems separately: in and in . Write that state , abbreviated , and there are such pairs.
Now use P1. If those states are all possible states of the pair, then by the superposition principle so is every complex combination of them, and the joint space contains their span. Declare them to be an orthonormal basis of that span. That declaration is the definition of the tensor product , and by Chapter 0.4 §2.2's definition of dimension it gives
The inner product comes with it, and nothing new has to be chosen. Declare on the basis and extend by conjugate-linearity in the first slot and linearity in the second, which is Chapter 0.5 §1's construction unchanged. For a general pair of states and this gives
the second identity following by expanding both sides on the basis and collapsing with the deltas. Amplitudes for independent systems multiply, which is what they had better do.
Asserted. The state space of a system composed of two parts is the tensor product of the parts' state spaces, so its dimension is the product of theirs. Observables belonging to part alone act as , and any two such observables commute with any two belonging to alone.
Not derived. There is a competing rule, which says that describing two systems means listing two descriptions, so that the dimensions add. That is what classical mechanics does with configuration spaces, and nothing in P1 to P6 rules it out. Which one nature uses is a physical question, and its answer is the source of every effect in Chapters 4.18 to 4.20.
What is not being claimed. Nothing here says a state of the whole is built from states of the parts, and §9.2 shows almost none of them are. Nothing here handles identical parts either, where a further restriction applies. That restriction is P8, the symmetrisation postulate, and it is stated in Chapter 4.18 §3.
9.2 · Entanglement, as a dimension count
Ask how many of the dimensions are reached by states of the form . Count parameters. Choosing takes complex numbers and takes , but rescaling by and by gives the same product, so one complex parameter is shared. The set of product vectors therefore has complex dimension , inside a space of complex dimension . For two two-level systems that is inside . For two ten-level systems it is inside .
So almost every state of a composite system is not a pair of states of its parts. Such states are called entangled, and the word names a counting fact rather than an influence: the joint space has more directions in it than the product construction reaches.
One explicit case, checkable by hand in three lines. Take two two-level systems and the state . If it were , then matching coefficients requires
From either , which contradicts , or , which contradicts . No solution exists. The state assigns no state at all to either half on its own, and the question "what is the first system's state?" has no answer of the kind P1 provides. That is why Chapter 4.19 has to build a different object, the density operator, to answer it.
9.3 · The size of the space, with a number
Iterating (4.2.27) over two-level systems gives a state space of dimension , so specifying a general state takes complex amplitudes, while specifying a state of each system separately takes real angles by (4.2.5). At :
The observable universe contains of order atoms. Three hundred two-level systems means three hundred atoms, which is a small molecule. They have more amplitudes in their joint description than there are atoms available to record them, and the excess is entirely the gap between and , compounded three hundred times. That gap is what makes simulating quantum systems on ordinary computers hard, what makes building quantum ones interesting, and what Chapter 4.20 measures with Bell's inequality. It follows from P7 and nothing else.
Put two systems side by side and ask what describes the pair. The answer is the one place where quantum mechanics departs from ordinary intuition by an amount that can be written as a number. Classically, describing two things means describing each and listing both, so the descriptions add. Here a basis for the pair is a list of pairs of basis states, one drawn from each, so the dimensions multiply.
That difference is entirely responsible for the phenomenon everyone finds strange. States describing each system separately and then pairing them do exist, but they form a vanishingly thin subset — counted properly, of dimension roughly the sum where the whole space has dimension the product. Almost every state of the pair is therefore not of that form, which means it assigns no state at all to either half on its own. A concrete two-by-two example can be checked by hand in three lines: four conditions on four numbers, and no solution.
The practical scale of this is worth a number. Three hundred two-state systems require more amplitudes to specify than there are atoms in the observable universe, while a description of each one separately needs six hundred angles. That gap is where the interest in building such machines comes from, and the last chapter of this part is about what fills it.
a natural place to stop · the framework is complete; what follows is three systems worked in full
10 · Three two-state systems, in full
Announce the destination. Chapter 0.5's Worked example 1 diagonalised one Hermitian matrix completely and then said what it was: "It is the ammonia molecule… it is neutrino oscillation… it is the qubit. Chapter 4.2 will do all three, and the linear algebra will already be finished." Here it is done. We solve the matrix once with §7's evolution operator, and obtain a single formula. Then we read that formula three times, in three sets of units: picoseconds, kilometres, nanoseconds. The linear algebra takes half a page because it was finished nine chapters ago. The rest is identification and arithmetic.
10.1 · One matrix, solved once
Two states, called and , whatever they turn out to be physically. The most general Hermitian operator on their span has four real parameters. One of them, the phase of the off-diagonal entry, can be absorbed into the definition of , so take it real. Write the remaining three as a mean, a difference and a coupling:
Chapter 0.5's Worked example 1 is the case . The eigenvalues come from the characteristic polynomial, , giving
Real, as Chapter 0.5 §6.1 requires, and the level splitting is . Notice the shape. The coupling pushes the two levels apart whatever the diagonal difference is, and the splitting is never smaller than . Two levels that would cross as passes through zero instead approach to and separate again. That is an avoided crossing, and the phenomenon is entirely (4.2.32).
Define the mixing angle by
which is legitimate because . Then the normalised eigenvectors are
orthogonal by inspection, as Chapter 0.5 §6.2 guarantees without being consulted. The grind box verifies them and runs §7's evolution operator on the state that starts as . The result is one formula, and it is the whole of this section:
with exactly. Four features are worth naming before the physics starts, because each is one of this chapter's postulates doing visible work.
The mean energy does not appear. It cancelled, because it contributes an overall factor to the state. That is a global phase, which P1 says is not there. Only energy differences are observable, and that is P1 rather than a separate principle.
The oscillation never stops and never damps. By §7's familiar box, the eigenvalues of the generator are purely imaginary, so nothing decays. A classical two-state rate process would relax to fifty-fifty and stay there.
Depth and speed move in opposite directions. Increasing the detuning at fixed coupling makes smaller, which is a shallower oscillation, and larger, which is a faster one.
And at short times the two effects cancel exactly. Expanding (4.2.35) for small , the contributes and the prefactor contributes , so
and the leading term does not contain at all. However far off resonance the system is driven, the transition probability leaves zero along the same parabola. Detuning changes when the curve turns over, not how fast it starts. That is the seed of the transition-rate formula Chapter 4.17 derives, and the figure below makes it visible.
Grind box — the eigenvectors, the evolution, and
Step 1 · check the eigenvectors. Subtract , which shifts both eigenvalues and moves no eigenvector. Using (4.2.33),
the last equality being the two identities and , both instances of and . The same computation with gives .
Step 2 · invert. (4.2.34) is a rotation by , so its inverse is a rotation by :
Step 3 · evolve. By §7, each eigenstate picks up its own phase and nothing else:
Step 4 · project. By P3 the amplitude to find is , and Step 2 gives , :
using from Euler's formula and . Squaring the modulus kills both the global phase and the , leaving (4.2.35).
Verified symbolically. Forming directly from (4.2.31) and taking the modulus squared of its entries returns for the off-diagonal transition, and gives identically in . Not to some precision. Identically. The series expansion of the same expression returns (4.2.36) term by term.
Everything from here to the end of §10 is (4.2.35) with numbers put in. The numbers are measurements and are quoted, not derived. The formula is derived, and it is the same one in all three cases.
Ammonia. The inversion line of in the rotational state is at , which is the line the first maser ran on in 1954. The corresponding splitting for the non-rotating molecule is . The barrier to inversion is about .
Neutrinos. with consistent with . Then with . Baselines and energies as quoted per experiment.
Qubit. A superconducting two-level circuit with level splitting , driven at a Rabi frequency . These are representative numbers for a transmon, chosen because they are round.
The two-state description itself is a model in every case.
- Ammonia is a molecule with many vibrational and rotational levels, and the two-state treatment is the restriction to the lowest inversion doublet.
- The neutrino has three flavours, and the two-flavour treatment is exact only when one mass splitting dominates.
- The qubit is a weakly anharmonic ladder truncated to its bottom two rungs.
Each restriction is a good approximation for the reason its own field gives, and none of them is being derived here.
10.2 · The ammonia molecule
Ammonia is a nitrogen atom and three hydrogens arranged as a shallow pyramid. The nitrogen can sit above the plane of the hydrogens or below it, and these two arrangements are mirror images with identical energy. Call them and . Two facts fix (4.2.31) completely.
The detuning is zero, by symmetry. Reflecting the molecule exchanges and and changes no energy, so the two diagonal entries of are equal and . This is not an approximation or a convenient choice. It is a symmetry of the Coulomb interaction between the same four nuclei and the same ten electrons. Hence by (4.2.33), and , and the oscillation goes all the way from one configuration to the other and back.
The coupling is not zero. Classically the nitrogen cannot get from one side to the other, because passing through the plane of the hydrogens costs about , which is , far more than the molecule has. The off-diagonal entry of a Hermitian operator connecting and is under no such prohibition. Nothing in P1 to P5 requires a state to get from one place to another through the intervening places, and Chapter 4.7 computes for a barrier. Here it is read off a measurement.
So the energy eigenstates are the symmetric and antisymmetric combinations, , split by , and
We want a number out of that rather than a shape, so put the measured into it. The splitting is
and the time for the molecule to turn itself inside out and back is , with the halfway point, nitrogen fully on the other side, at . Three consequences follow, and each of them is a postulate of this chapter made visible.
A molecule in its ground state has no shape. The stationary states are , not or . A state of definite shape is a superposition of the two energy levels, and therefore not stationary. This is §4.3 in the most concrete form available. Shape and energy are represented by operators that do not commute, so there is no state with both sharp, and the molecule in its lowest energy level is not in either configuration.
The splitting is tiny because the barrier is large. The ratio of barrier to splitting is . That the coupling is suppressed by three orders of magnitude rather than being zero is the quantitative content of tunnelling, and Chapter 4.10's WKB approximation computes the suppression as an exponential in the barrier's width and height.
The frequency lands in the microwave band, which is why the maser came first. The wavelength is . Chapter 4.1's Worked example 3 derived the ratio of stimulated to spontaneous emission for an atom in thermal radiation as . At and , , and the ratio is
so stimulated emission beats spontaneous emission by a factor of at room temperature at this frequency, against for an optical transition. Now sort the two energy eigenstates with an inhomogeneous electric field, which works because have opposite parity and so respond oppositely. Feed the upper one into a cavity tuned to , and the third process Einstein was forced to invent in Chapter 4.1 §4.5 amplifies. That is the ammonia maser, built in 1954, and every number in its design is in this paragraph.
10.3 · Neutrino oscillation
The same matrix, and the identification is where all the work is. A neutrino is produced by a weak interaction in a definite flavour, with a muon or with an electron, and it is detected the same way. But flavour is not what propagation cares about. Propagation is generated by , whose eigenstates are the states of definite mass.
Call the flavour states and , and the mass states and . The assertion tested by every oscillation experiment is that these are two different orthonormal bases of the same two-dimensional space, related by (4.2.34) for some angle . Nothing about that is unusual. It is Chapter 0.4's change of basis, and §4.3 of this chapter says the two observables do not commute.
Now the energies, and this is the one step needing Part II. A neutrino of definite mass and momentum has, by Chapter 2.5 §4.3, energy . Neutrinos in these experiments carry energies in the MeV to GeV range, while their masses are below an electronvolt, so and the square root can be expanded. We want the difference of two nearly equal energies, and the leading terms will cancel, which is exactly why expanding is the right move:
using from Chapter 0.3. The common cancels in the difference, and with for the beam energy,
That is the level splitting in (4.2.35). One more identification is needed. A neutrino travels at essentially , so the proper substitution for the elapsed time is , where is the distance from source to detector. That distance is the quantity an experiment actually controls. Substituting both into (4.2.35):
which is the formula every oscillation experiment is analysed with, obtained from (4.2.35) by two substitutions and no new physics. Now convert the phase to the units experiments use, with in , in kilometres and in GeV. Using ,
the familiar coefficient, derived rather than quoted. Next we want the distance over which the pattern repeats. That is the oscillation length, the analogue of ammonia's , and it is the making the phase advance by :
Numbers, for the T2K experiment. A muon-neutrino beam is made at Tokai and detected at Kamioka, away, with the beam tuned to . With the phase is , which is , within three per cent of , the first oscillation maximum. So
for maximal mixing: essentially every muon neutrino in the beam has become something else by the time it arrives. The oscillation length is , so the first minimum sits at against a baseline of . Equivalently, is exactly the first minimum for a beam energy of . The beam energy was chosen to put the detector at the minimum, and (4.2.42) is how it was chosen.
The same formula with a small mixing angle gives the other kind of experiment. At the Daya Bay reactor, electron antineutrinos of about are counted at . The phase is , so with the predicted disappearance is . That is a deficit of about eight per cent, which is what is measured, and it is how is known.
One honesty note about the derivation. Treating the two mass states as plane waves with a common momentum, and setting , is a shortcut. A neutrino is produced as a localised wave packet, the two mass components travel at slightly different speeds, and the careful treatment follows the packets and asks when they still overlap at the detector. It gives (4.2.42) in every regime these experiments operate in, and the corrections are suppressed by the ratio of the packet width to . The shortcut is standard, and it is a shortcut. The machinery to do it properly is Chapter 4.6's wave packets.
10.4 · The qubit
The third reading inverts the relationship between the formula and the apparatus. In ammonia the parameters are whatever chemistry supplies. In a neutrino they are whatever the mass matrix supplies. In an engineered two-level system both are set by the designer, and (4.2.35) becomes a specification rather than a prediction.
Take a superconducting circuit with two levels split by , with . That is a splitting of , which is twelve times at the such circuits run at. So Chapter 0.6's Boltzmann factor leaves the system in its lower level rather than thermally stirred.
Now drive it with a microwave field at exactly . Chapter 4.17 shows that in a frame rotating with the drive the problem becomes (4.2.31) with and , where is proportional to the drive amplitude. That reduction is Chapter 4.17's and is not derived here. With it granted, (4.2.35) reads
the Rabi formula. With the full period is , so a pulse lasting
takes to with probability . That is a NOT gate, performed by leaving the drive on for exactly half an oscillation. Half that, , gives probability and produces the operator whose square is NOT. Worked example 3 builds it and shows that no classical stochastic process has such a square root. Detuning the drive off resonance is , and (4.2.35) says the gate then fails to reach probability no matter how long it is left on, with the shortfall . That is where the tolerance on a control line comes from.
Set the three side by side. Same matrix, same formula, three sets of units:
| are | Splitting | Full period of | ||
|---|---|---|---|---|
| Ammonia | nitrogen above / below the hydrogens | , by mirror symmetry | ||
| Neutrino | the two detectable flavours | set by the mass matrix | of flight | |
| Qubit | the two engineered levels | drive detuning, chosen |
The splittings span a factor of , from down to , and the periods span exactly the same factor the other way, because (4.2.35) makes period and splitting reciprocal and nothing else enters. The linear algebra is one matrix that Chapter 0.5 diagonalised before any of this was mentioned.
10.5 · One computed evolution, read in three sets of units
The figure below integrates (4.2.18) numerically for (4.2.31) and plots what comes out. Nothing in it uses (4.2.35). The closed form appears only in the readout, as the thing being checked against. Two of the readouts are the numerical confirmation this chapter owes.
The point of this section is that there is one calculation, and it is finished. Take two states of nearly the same energy with something connecting them, which is the smallest interesting arrangement there is, and solve it once. What comes out is a single expression: the chance of finding the system in the other state swings back and forth forever, at a rate set by the separation of the two energy levels, with a depth set by how evenly the connection mixes them.
Now read the answer three times. In a molecule of ammonia the two states are the nitrogen atom sitting on either side of its three hydrogens, the connection is its ability to pass through them, and the swing takes forty-two trillionths of a second — which corresponds to a microwave line that was used to build the first device of its kind. In a neutrino the two states are the two identities it can be detected with, the connection is the mismatch between those and the two definite masses, and the swing takes six hundred kilometres of flight. In an engineered two-level circuit both numbers are chosen by the designer, and half a swing at twenty nanoseconds is what turns one state into the other.
Three subjects, three sets of units, one matrix. That is what the chapter has been claiming.
11 · Worked examples
A beam of spin- particles is prepared with a definite value of , passed through a second apparatus measuring , and then through a third measuring again. Using only P3 and P4: (a) compute the probability of the third magnet reporting spin-down. (b) Compute the same probability with the middle magnet removed. (c) Compute it with the middle magnet present but not read, so that the two -paths are recombined coherently. (d) Say which postulate each answer used, and what the comparison establishes.
Set up the two-dimensional space with an orthonormal basis. The observable has, by Chapter 0.5's Worked example 1 with , the eigenstates
which are orthogonal, as §4.1 requires without being asked. All four overlaps between the two bases have modulus .
(a) Measure, keep the beam, measure again. Starting from , P3 gives , and P4 replaces the state by . Then P3 again: . The two are independent because P4 wiped out all memory of the first state, so
(b) Middle magnet removed. Nothing happens between the two measurements, so the state is still and P3 gives , exactly. The two eigenstates of are orthogonal, and §4.1's second theorem says the confusion probability is zero rather than small.
(c) Middle magnet present, nothing read, beams recombined. Now no measurement occurred, so P4 does not apply and P3 is used once at the end. The state is unchanged and can be rewritten using the resolution of the identity in the basis, which is Chapter 0.5 §2.3's inserted and nothing more:
The two amplitudes are equal in magnitude and opposite in sign, and they cancel exactly. So the answer is , agreeing with (b) as it must, since inserting changes nothing.
(d) What the three answers establish. Compare (a) with (c). The same apparatus is in the beam line in both, the same two paths are travelled, and the same final measurement is made. In (c) the two amplitudes are added and then squared, giving . In (a) they are squared and then added, giving , of which one branch was kept, so . The only difference in the beam line is that in (a) the intermediate value was read. Keeping one branch afterwards is what turns the into , and the reading is what destroyed the cancellation.
Certainty becomes a coin flip because somebody looked. That is P4 doing work that P3 alone cannot do. P3 assigns probabilities to the outcomes of a measurement, and it is P4 that destroys the coherence between the branches so that the third magnet sees no interference. It also shows why §6.2's familiar box insisted that conditioning is not the whole story. The classical law of total probability would give whether anyone looked or not, and the measured answer without looking is .
A three-level system has the observable
(a) Find the eigenvalues and the projections. (b) Apply P3 and check the probabilities sum to one. (c) Compute and two ways. (d) Apply P4 and say precisely what the state becomes. (e) Find a second observable that completes the set, and give the quantum numbers of the three joint eigenstates.
(a) is real symmetric, hence Hermitian. Its top-left block is Chapter 0.5's with eigenvalues and eigenvectors . The third basis vector is an eigenvector with eigenvalue . So the distinct eigenvalues are , twice degenerate, and , once. Building as the sum of over an orthonormal basis of each eigenspace:
Checks: ✓, ✓, each is idempotent and Hermitian ✓, and ✓, which is (4.2.7).
(b) and , so
summing to . Note what would have gone wrong with the careless version of P3. Writing "the probability is " and picking the eigenvector alone gives , and picking alone gives . Neither of those is the probability of the outcome , which is their sum . The projection form is written the way it is precisely so that a degenerate eigenvalue is handled correctly and independently of which basis was chosen inside the eigenspace.
(c) From P3, . From (4.2.9), ✓. For the spread, because both eigenvalues square to , so and . Directly, ✓.
(d) If is obtained, P4 gives
This is the point of the example. The state after the measurement is not and not . It is the particular direction inside the eigenspace that already pointed along, rescaled. A degenerate measurement removes the component outside the eigenspace and touches nothing inside it, which is exactly §6.2's remark that such a measurement is weaker than a complete one.
(e) Any commuting with and distinguishing the two directions inside 's eigenspace will do. Take . Its commutator with is the zero matrix by direct multiplication, since is a multiple of the identity on the block where is non-diagonal. Inside the eigenspace of , the operator has eigenvalue on and on , so it chooses the basis could not. That is Chapter 0.5 §8.2's Step 4, in three dimensions. The joint labels are
each one-dimensional, so is a complete set of commuting observables and those pairs are the quantum numbers of §4.3. Measuring on gives and , summing to one, and after it the state carries a complete label.
Section 10.4 said that half a Rabi period turns into with certainty, and that a quarter period produces an operator whose square is that gate. (a) Build the quarter-period operator from §7 and check it is unitary. (b) Square it. (c) Show that the classical analogue does not exist, meaning a two-state stochastic process whose square is a certain flip. (d) Say which postulate the difference belongs to.
(a) The resonant drive of §10.4 is (4.2.31) with and , so and by (4.2.17) the evolution operator is . Chapter 0.5's Worked example 2 evaluated exactly this exponential using and got with . The quarter period is , giving
Unitary, since ✓, as §7.2 guarantees for any with Hermitian. Applied to it gives , whose two probabilities are and .
(b) Squaring,
That is the NOT gate multiplied by , and by P1 an overall factor of modulus one is not there. So applied twice takes to with probability . It is a genuine square root of NOT, and it is what the pulse of §10.4 performs after .
(c) Now the classical version. A stochastic process on two states that treats them symmetrically is described by the matrix with , meaning "stay with probability , flip with probability ". Two applications give
and requiring to be the certain flip means requiring the diagonal entry to vanish. It is a sum of two squares of real numbers, so it vanishes only if both vanish, which needs and at once. Its minimum over all real is , attained at . No classical two-state process, applied twice, produces a certain flip. The best possible leaves a half chance of being where it started.
(d) The difference is P1's, not P3's. The quantum operator succeeds because the two intermediate amplitudes are and , and on the second application the contributions to "stay" are and , which cancel. Cancellation requires the entries to be complex numbers of opposite sign, which requires the state to be a vector in a complex space rather than a list of probabilities. Section 5.4 named that as the single place a reader is most likely to install the wrong picture. Half-way through a NOT gate the qubit is not "probably still ". It is in a definite state whose phase is what makes the second half of the gate work.
12 · Your turn
Problem 1 — rays, not vectors, and the sphere that follows
(a) Show that the set of states of an -level system has real parameters, and say which constraint removes each of the two. (b) For , show that every state can be written as (4.2.6) for exactly one pair with and , treating and as the two exceptions and saying what happens there. (c) Compute in the state (4.2.6), with , , , and identify what the three numbers are. (d) Deduce that two states are the same state exactly when those three expectation values agree, and say why that makes the Bloch sphere a measurement rather than a picture.
Solution
(a) A vector in is real numbers. The condition is one real equation, leaving . P1 identifies with , and is one real parameter, leaving . Normalisation removes one, and the global phase removes the other.
(b) Write . Choose the phase freedom to make real and non-negative, which fixes uniquely unless . Then for a unique , since runs monotonically from to on that range, and , so with unique modulo . The exceptions are the poles. At we have and is undefined, and at we have and is the leftover phase, which P1 discards. Both are single states, and they are exactly the two points where a polar coordinate system on a sphere degenerates. The parameter count is right, and the coordinates are the ones misbehaving.
(c) With , and (4.2.9),
and the same computation with gives with the sign convention above, which is . So the three expectation values are the Cartesian coordinates of the point at polar angle and azimuth on the unit sphere. Check the length: , so every state sits exactly on the sphere and none inside it. (States inside it exist and are mixtures, which is Chapter 4.19's density operator.)
(d) The map from to the point is a bijection onto the sphere by (b) and (c), so equal expectation values force equal and hence the same state. Each of the three numbers is the average of a measurement, and Chapter 4.12 measures exactly these with three Stern–Gerlach magnets in three orientations. So the position of a state on the sphere is determined by an experiment, and the sphere is where the states are rather than a way of imagining them. Note that this fails for . The parameter count gives four real parameters, there is no two-sphere, and no comparable picture exists.
Problem 2 — how the ammonia maser sorts its molecules
Put the molecule of §10.2 in a uniform electric field . The two configurations and have electric dipole moments of magnitude pointing in opposite directions along the symmetry axis, so the field shifts their energies by : the matrix is (4.2.31) with and unchanged. (a) Write down the two energies as functions of and sketch their behaviour at small and large field. (b) Find the field at which the two régimes cross over, and evaluate it for . (c) Explain, from (a), why passing a beam through an inhomogeneous field separates the two energy eigenstates, and which way each is deflected. (d) What happens to the inversion oscillation of (4.2.37) as the field is raised?
Solution
(a) Straight from (4.2.32),
At small field, expand: , so the shift is quadratic in with coefficient . At large field the square root is dominated by and the shift is linear, , which is the classical answer for a dipole in a field. The two levels never cross. They repel to a minimum separation of at zero field, which is the avoided crossing of §10.1.
(b) The crossover is where the two terms under the root are equal, . With ,
At the splitting has grown by only . At it has grown by a factor of , and the response is essentially linear.
(c) An inhomogeneous field exerts a force on a molecule in the corresponding state. Since increases with field strength and decreases, the upper state is pushed towards weak field and the lower state towards strong field. They go opposite ways, and a quadrupole field with a minimum on the axis focuses the upper state down the axis while throwing the lower state out. That is the maser's state selector, and what emerges is a beam of molecules in the upper energy state only. That is a population inversion, which Chapter 4.1's Worked example 3 proved cannot occur in thermal equilibrium at any temperature, and which therefore has to be manufactured.
(d) By (4.2.35) the oscillation amplitude is , which falls as the field rises, while the frequency rises. At the amplitude goes as and the molecule is essentially frozen into whichever configuration it started in, because the field has made the two configurations energetically distinguishable by far more than the tunnelling can bridge. A strong field switches the inversion off, and this is the same mechanism by which a large molecule's environment freezes it into one enantiomer.
Problem 3 — how badly a finite matrix misses the canonical commutator
Section 8.3 asserted a bound. Prove it and interpret it. (a) Using the inner product on complex matrices, which is Chapter 0.5 §1.2's fifth row, verify the three axioms hold, and compute . (b) Prove for every , and state which theorem of Chapter 0.5 you used. (c) Deduce for all matrices , and say what the right-hand side is the norm of. (d) Determine exactly when equality holds, using Chapter 0.5 §1.4's equality condition, and say what that means about how close a finite-dimensional model can get. Check your answer against the pair , .
Solution
(a) Conjugate symmetry: , since the trace of the conjugate transpose is the conjugate of the trace. Linearity in the second slot is linearity of the trace. Positive definiteness: , vanishing only if every entry does. And , so .
(b) , so Cauchy–Schwarz (Chapter 0.5 §1.4) applied in this inner-product space gives . Take square roots. The point worth noticing is that Chapter 0.5 proved Cauchy–Schwarz from the three axioms alone and never mentioned what the vectors were, so it applies verbatim to a space whose vectors are matrices.
(c) Put . By cyclicity of the trace (Chapter 0.4 §6.1) the commutator is traceless, so and . Then (b) gives , that is . The right-hand side is exactly : the error is at least as big as the target.
(d) Chapter 0.5 §1.4 proved that Cauchy–Schwarz is an equality exactly when the two vectors are parallel, so equality here needs for some complex . Taking the trace, , so and
The bound is attained exactly when the two matrices commute, which is to say when they reproduce none of the canonical relation whatever. Every pair that actually tries, by having a non-zero commutator, does strictly worse. Check it on the suggested pair. We have and , so , which has half of it exactly right. Then with , a relative error of against the commuting pair's . Getting one diagonal entry exactly right cost more than giving up did.
So there is no sequence of finite-dimensional models converging to the canonical commutator, not even slowly: the relative error is bounded below by at every dimension, and the minimum is achieved by doing nothing. Numerical minimisation over all complex pairs from twelve random starts returns at , which is the bound and its equality case, found without being told about either.
Problem 4 — design a neutrino experiment
The other mass splitting is with . Reactors emit electron antineutrinos with energies of a few MeV; take . (a) Using (4.2.43), find the baseline at which the disappearance is greatest, and the oscillation length. (b) The KamLAND experiment used reactors at an average of about . Compute the survival probability there at and say where on the oscillation curve that sits. (c) Daya Bay, measuring the other angle, is at . Show that this is the right baseline for and the wrong one for , with numbers. (d) A colleague proposes measuring with a detector from a reactor, arguing that being closer means more neutrinos. Say quantitatively why this fails.
Solution
(a) Maximum disappearance is the first maximum of , at phase :
with in GeV. The oscillation length from (4.2.44) is , twice the first maximum as it must be.
(b) At the phase is , so and
That is past the first minimum and on the way back up the second oscillation. Now notice what this implies about the analysis. The phase depends on , the reactor spectrum spans roughly to , and at the phase varies from to across that band, which is more than a full oscillation. The measured survival is therefore an average over the spectrum and is closer to . What makes the measurement powerful is that the shape of the distortion as a function of fixes far better than any single rate could.
(c) For at , the first maximum is at , so sits at phase , giving . That is near enough to the maximum, and §10.3 computed the resulting deficit. For at the same baseline the phase is , so and the disappearance from that term is , invisible under the systematics. The two splittings differ by a factor of , so the two baselines differ by the same factor, and one experiment cannot do both.
(d) The flux does rise, as , so at instead of there are times as many events. But the signal is not the number of events. It is the fraction missing, and at the phase is , so and the disappearance is . That is five parts in a million, against a reactor flux normalisation known to a per cent at best. Increasing the count rate by improves the statistical error by , while the effect being sought has shrunk by . The trade is losing by a factor of about , and that is before the systematic error, which does not improve with the count rate at all and which is already far larger than the signal. Oscillation experiments are placed by (4.2.44) and not by flux.
Problem 5 — entanglement, by counting
(a) Show that the four-dimensional space of two two-level systems contains states that are not of the form , by exhibiting one and proving no factorisation exists. (b) A general two-qubit state is . Show that it factorises if and only if , and interpret that as one complex condition on four complex numbers. (c) Confirm §9.2's parameter count for using (b), remembering P1. (d) For , compute the probabilities of the four outcomes of measuring on both systems, and then the two outcomes of measuring on the first system alone. Say what is strange about the answer and which chapter handles it.
Solution
(a) Take and suppose . Expanding by (4.2.28) and matching the four coefficients gives (4.2.29): and . From , either , which contradicts , or , which contradicts . No factorisation exists.
(b) If the state factorises then and , so the condition is necessary. Conversely suppose . If , set , , , . Then ✓, ✓, ✓, and by hypothesis ✓. If the same argument runs from whichever coefficient is non-zero. So factorising is exactly the vanishing of the determinant . That is one complex equation on four complex numbers, which is a two-real-dimensional condition, and it confirms that product states are a thin set rather than a large one.
(c) Product states satisfy one complex equation inside , so they form a set of complex dimension with , exactly §9.2's count. Now count them as states. Imposing and discarding the global phase removes two real dimensions, by Problem 1(a). So the whole space has real parameters and the product states have , which is also : two real parameters for each qubit's own Bloch sphere, as it must be.
(d) The joint measurement projects onto the four basis states, and by P3 the probabilities are : and . Measuring only the first system means measuring , whose eigenvalue has the two-dimensional eigenspace spanned by . By P3 in its projection form, and likewise .
What is strange is the combination. Each system on its own is completely unpredictable, a fair coin, and yet the two always agree. There is no state of the first system alone that reproduces this. Assigning it or contradicts the observed randomness, and assigning it any superposition contradicts (a). The description "each has a definite but unknown value, correlated at preparation" reproduces these particular numbers and is ruled out by measuring other observables, which is the content of Bell's inequality. Chapter 4.19 builds the density operator to answer "what is the first system's state", and Chapter 4.20 measures the inequality.
The table was the chapter. Twenty-two rows, every left-hand entry a theorem of Chapter 0.5 cited at the section that states it, every right-hand entry that theorem with two or three words changed. Real measured values are 0.5 §6.1. Perfectly distinguishable outcomes are §6.2, and the confusion probability is exactly zero rather than small. Any state being a superposition of the outcomes of any observable is §6.3. Probabilities adding to one is Parseval, §2. Compatible measurements and the quantum numbers that label a state are §8. Symmetries and evolution generated by observables are §7.1. None of that was assumed here and none of it was re-proved here, which was the instruction.
Seven assertions, boxed and flagged, and they are the whole of the addition.
- A state is a unit ray, so that normalisation makes Parseval read as probability and the overall phase drops out of every prediction by one line of conjugation. The relative phase between terms is the opposite case. It produces , and it is the entire wave behaviour of Chapter 4.1 §6.6, now living somewhere definite.
- An observable is Hermitian, which is the only postulate the measurement outcomes need. Chapter 4.4 §4 will sharpen it to self-adjoint for a physical reason: momentum on a half-line has no self-adjoint extension and is therefore not an observable.
- The Born rule, which is not derivable, is still open, and was marked as such at the moment it entered rather than afterwards. Gleason's theorem was quoted with its hypotheses and explicitly not presented as a derivation.
- The state update is a separate postulate, and Worked example 1 shows why the separation is not bookkeeping. Three magnets give when the middle one is read, and exactly when it is not.
- The generator of time evolution is the energy, which is the only new content in §7. Unitarity followed from probability conservation in three lines, and the existence of a Hermitian generator followed from unitarity in three more.
- The canonical commutator is postulated here rather than in Chapter 4.9, because Chapters 4.6, 4.8 and 4.11 all need it first.
- And a composite system's space is the tensor product, so dimensions multiply.
Three lines that settle the shape of the book. The trace of a commutator is zero by cyclicity, from Chapter 0.4 §6.1. The trace of is . So no finite matrices satisfy . Cauchy–Schwarz in the matrix inner product sharpens that: at every dimension, with equality exactly for commuting pairs. So a finite model's best strategy is to reproduce none of the relation, and direct numerical minimisation confirms the ratio at . Quantum mechanics is infinite-dimensional before a single physical question has been asked, and the proof needs nothing but a trace.
And one matrix did all three systems. With having mean , difference and coupling , the splitting is and , with absent because it is a global phase. Read once: ammonia, by mirror symmetry, splitting , inversion period , wavelength , stimulated emission favoured over spontaneous by at room temperature. That is the maser, with every design number on the page. Read again with and expanded from Chapter 2.5's dispersion relation: the oscillation formula , its coefficient derived rather than quoted, and T2K's at sitting at against an oscillation length of . Read a third time with both parameters chosen by an engineer: a NOT gate and a square root of it, which Worked example 3 shows no classical two-state process possesses. The figure integrates the equation numerically and holds to about through the whole cycle at every setting, moves not at all under a full turn of the global phase, and shows every detuned curve leaving the origin along the same parabola , which is the fact Chapter 4.17's transition rate is built on.
Where this gets spent. Ten flags. Seven are the postulates. One is Gleason with its hypotheses. One is the fundamental theorem of algebra inherited from Chapter 0.4 §7, which is the single thing this chapter quotes rather than derives or postulates, and the thing the whole measurement apparatus descends from, paid off in Chapter 5.4. And one is the measured parameters of §10. That is the largest count in the book and it is correct, because this is the only chapter whose subject is the postulates.
What is owed is now precise rather than vague. Chapter 0.5 named the four places its proofs used finite dimension, calling them "in the induction, in rank–nullity, in the interchange of sums, in the claim that an injective map is surjective", and §8 showed that infinite dimension is forced. The bill is paid in three instalments, and the shape repeats one you have already seen: Chapter 0.4 built the space and Chapter 0.5 built the operators on it. Chapter 4.3 builds the space and Chapters 4.4 and 4.5 build the operators on it.
Chapter 4.3 discards the Riemann integral, builds the Lebesgue integral, proves complete so that a limit of states is a state, and shows the Fourier modes really are an orthonormal basis. Everything here that said "expand in a basis" is on credit until then. Chapter 4.4 supplies domains and the difference between symmetric and self-adjoint that P2's box warned about. Chapter 4.5 supplies spectra with no eigenvectors in the space, the spectral theorem in the form that survives, Stone's theorem for §7.3, and the meaning of and . Everything here that said "the spectral theorem" is on credit until then.
After that, the list gets spent one chapter at a time.
- Chapter 4.6 turns (4.2.18) into a differential equation and solves it.
- Chapter 4.8 diagonalises the oscillator, which is §4.2's "solving a system means diagonalising its Hamiltonian".
- Chapter 4.9 applies Cauchy–Schwarz to and and gets the uncertainty principle with nothing added, and Chapter 4.10 proves that P6 cannot be extended to every observable.
- Chapter 4.11 spends §4.3's quantum numbers on angular momentum.
- Chapter 4.20 returns to P3, P4 and P7 together, which is where the three of them turn out to have been the interesting ones all along.