Part IV · Quantum Mechanics — Chapter 4.1
What Classical Physics Cannot Do
Four failures, made quantitative, so that when quantisation arrives it arrives forced rather than proposed.
Every part of this book so far has been built by cornering. Chapter 1.1 cornered the action principle out of Newton. Part II cornered Lorentz invariance out of two postulates about light that a nineteenth-century physicist would have accepted. Part III cornered the field equations out of the equivalence principle and one identity. Quantum mechanics cannot be cornered that way. So this chapter exists to make sure that when the postulates arrive in Chapter 4.2 you already know exactly which measurements they were invented to account for, and exactly how badly the alternative fails.
Four failures, then. They are not chosen for variety. They are chosen because each one is a number.
- A classical calculation of the light inside a hot box gives infinity.
- A classical account of the photoelectric effect predicts a delay of seconds, where nanoseconds are measured.
- A classical account of X-ray scattering predicts no wavelength shift at all, where a shift of several per cent is routine.
- A classical hydrogen atom lives for sixteen picoseconds.
A subject is replaced when its predictions are wrong by factors like these, not when its picture becomes unfashionable. Each of these numbers is one you can produce yourself before the section ends.
The centre of gravity is the first of the four. Sections 1 to 5 derive the spectrum of cavity radiation, and the route we take is Einstein's: three rates, a balance condition, and Chapter 0.6's Boltzmann weights. We take that route because it needs no statistics of light, and there is no statistics of light in this book until Chapter 4.18. The classical answer is derived first and in full, and not as throat-clearing. It is used as a boundary condition inside the derivation that replaces it. Sections 6 and 7 then take the three failures that are about matter meeting light, and §7 ends by naming what has to be given up.
Conventions. SI throughout, with , and written explicitly. The symbol does not appear until §6.5, where it is defined at the point of use. Frequency means ordinary frequency in hertz, not angular frequency, because every spectroscopic measurement quoted here is reported that way. We write where the wave four-vector of Chapter 2.5 is used. The symbol means spectral energy density: energy per unit volume per unit frequency interval, so that is an energy density and is the total energy per unit volume. Since the 2019 SI redefinition, , , and are exact by definition. So every constant this chapter computes from them is exact arithmetic, and any disagreement with a measurement is a real disagreement.
Tools you'll need — Chapter 0.6 Worked example 1 above all: the Boltzmann distribution derived from maximum entropy and two Lagrange multipliers, with identified rather than assumed. Sections 3 and 4 both run on it and nothing else in statistical mechanics is used or needed. Chapter 0.8 §7.6 for the standing waves of a field in a box and §7.2 for the fact that the modes are independent oscillators. Section 2 here is that argument in three dimensions. Also §4.3 for the harmonic oscillator's energy as a sum of two squares, which §3.1 needs. Chapter 0.2 §5 for when an improper integral converges, which is the whole content of §3.3, and §4 for interchanging a sum and an integral. Chapter 0.3 §5 for dimensional analysis, used in §5.7 to say what is, and for , which §7.3 needs. Chapter 0.9 §1 for Fourier series and Parseval's identity, used once in §5.6's grind box to evaluate a sum. Also §7.5 for the Cauchy distribution, whose divergence has the same structure as §3.3's. Chapter 2.5 §7.1 for the wave four-vector and its null condition, §4.3 for and , and Worked example 1, the Compton calculation, which §6.4 collects unchanged. Chapter 2.1 §3 for the electromagnetic wave equation, the dispersion relation and the transversality of light, which supplies §2.4's factor of two. Chapter 3.9 §1.2 for the microwave background at , which Worked example 1 takes apart as a blackbody.
1 · What a classical prediction of the blackbody spectrum actually is
Announce the destination. Before anything can be predicted, there has to be something whose value does not depend on unrecorded details of the apparatus. We show that the light inside a hot cavity is such a thing: its spectrum is a universal function of frequency and temperature, identical for every cavity of every material and shape. The argument costs one page and it uses the second law of thermodynamics once.
1.1 · The object is a cavity, not a surface
A hot poker glows, and the colour of the glow depends on the poker. Iron, carbon and tungsten at the same temperature do not look the same, because each emits and absorbs different wavelengths with different efficiency, and what leaves the surface is the product of the two. So the radiation from a surface is not a universal object, and there is nothing clean to predict about it.
Change the object. Take a closed cavity with opaque walls, hold the whole thing at temperature , and wait until nothing further changes. The radiation inside is then in thermal equilibrium with the walls, meaning that every part of the wall absorbs exactly as much as it emits, at every wavelength. If that were not so, the part in question would be heating or cooling, and the state would not be steady.
Let be the energy per unit volume carried by radiation with frequency between and . To measure it, drill a hole small enough that letting radiation out does not disturb the equilibrium inside, and analyse what comes through. That is what a laboratory blackbody is, and it is why the spectrum of a furnace with a peephole is reproducible when the spectrum of a glowing surface is not.
1.2 · Kirchhoff's argument, in full
Here is the claim to be established. The function depends on and on , and on nothing else at all. Not on the material of the walls, not on the shape of the cavity, not on its volume. The argument is one Kirchhoff gave in 1859, and it is worth having complete, because it is what makes the problem well posed and because it uses exactly one thing this book has not built.
The argument below uses the statement that heat does not pass spontaneously from a colder body to a hotter one, with no other change occurring. This book has not derived the second law and will not. Chapter 0.6's Worked example 1 built the Boltzmann distribution from maximum entropy, and maximum entropy is a statement about inference rather than about time. Getting from there to the irreversibility of macroscopic heat flow requires an argument about initial conditions that is not attempted anywhere here.
What is being assumed, precisely. Two bodies at the same temperature, connected only to each other and otherwise isolated, cannot spontaneously come to different temperatures. That is weaker than the full second law and is all §1.2 needs.
Take two cavities, and , with walls of different materials and of different shapes, both held at the same temperature . Now the move. Join them by a short pipe and place in the pipe a filter that transmits radiation in a narrow band of frequencies about and reflects everything else perfectly. Such a filter exchanges no net energy with the radiation, because what it does not transmit it returns.
Suppose . Then more energy per unit time crosses the filter from to than from to . The reason is that the flux through an aperture is proportional to the energy density on that side, and to nothing else the filter is sensitive to. So warms and cools, and two bodies that started at the same temperature have spontaneously arrived at different temperatures with nothing being done to either. That is forbidden. The same argument with and exchanged rules out . Hence
and since and were arbitrary cavities, is a universal function. Let's note where each hypothesis entered. The filter's narrowness is what gives the conclusion frequency by frequency rather than merely for the total. Its perfect reflection outside the band is what stops it acting as a reservoir of its own.
1.3 · What is therefore to be predicted
The problem is now completely specified, which is a rare and valuable state of affairs. There is one function of two variables, it belongs to no particular material, and any theory of light and heat must produce it. Two things follow immediately and are worth writing down before any calculation, because they are what makes the calculation of §2 legitimate.
The cavity may be chosen for convenience. Since the answer does not depend on shape or material, take a cube with perfectly conducting walls. That is not an approximation to a real furnace. It is a different cavity with the same .
The answer is an energy density, not an energy. carries no dependence on volume, so a cavity twice the size holds twice the energy at every frequency. This is what lets §2 count modes per unit volume and stop worrying about the box.
A kiln with a small hole in it glows, and the colour of the glow depends on how hot the kiln is and on nothing else at all — not the clay, not the shape, not the polish of the inside. That claim sounds too strong to be free, so the argument is worth having: it converts a question about pottery into one with exactly one correct answer.
Set two hot boxes side by side at the same temperature and join them through a pipe carrying a filter that passes one narrow band of colour and nothing else. If the two boxes held different amounts of that colour, more of it would cross one way than the other, one box would grow hotter than the other with nothing whatever being done to either, and heat would have flowed from cold to hot on its own. Thermodynamics forbids that outright. So the two amounts must match, band by band, and since the boxes were arbitrary, the amount of each colour inside any hot cavity is one universal function of colour and temperature.
That is what makes the problem worth attacking rather than merely interesting. Nothing is left to argue about — no material, no geometry, no surface finish — and any theory of light and heat whatsoever has to deliver that single function. The rest of this chapter is what the classical theory delivers when it is asked.
2 · Counting modes in a box
Announce the destination. We want the number of independent oscillations of the electromagnetic field per unit volume with frequency between and . Four steps: identify the standing waves, notice that their wavevectors form a lattice, count lattice points inside a ball, and change variable from wavenumber to frequency. The answer is and every step is on this page, because this count is used three more times in the book and once here to produce a divergence.
2.1 · A field in a box is a set of independent oscillators
Chapter 0.8 §7.6 took a chain of masses and springs, let the spacing go to zero, and obtained a field obeying the wave equation. Its normal modes are the standing waves , and each one of them is an independent harmonic oscillator with its own frequency, its own amplitude and its own conserved energy. Chapter 2.1 §2 derived from Maxwell's equations that each Cartesian component of in vacuum obeys exactly that wave equation, with . So the argument transfers with one change: three dimensions instead of one.
In a cubical cavity of side with perfectly conducting walls, the tangential electric field must vanish on every wall. A standing wave
satisfies that condition exactly when are positive integers, since the sine vanishes at and at precisely then. Write the wavevector components as . The wave equation then forces , which is Chapter 2.1's dispersion relation, so
The reason to want the count is that each such mode is an independent oscillator. So the total energy in the cavity is a sum over modes of whatever energy each one carries, and §3 will supply the energy of one.
2.2 · The allowed wavevectors form a lattice
Read (4.1.3) as a statement about geometry rather than about waves. The allowed are the points of a cubic lattice of spacing , restricted to the octant where all three components are positive. Each lattice point sits at the corner of a small cube of volume , and each such cube contains exactly one lattice point when the points are assigned to cubes consistently.
So counting modes with below some value becomes counting lattice points inside a sphere of radius , restricted to one octant. Suppose is large compared with the lattice spacing . That is the same as saying the wavelength is small compared with the box, which for visible light in a laboratory cavity means a ratio of about . Then the count is the volume available divided by the volume per point:
The is the octant and the is the ball. Nothing else has happened.
2.3 · Two polarisations
Each allowed carries not one mode but two, and the reason is a fact about Maxwell's equations rather than about counting. In vacuum , and for a wave with wavevector that condition reads : the field is transverse, as Chapter 2.1 §2 derived. The vectors perpendicular to a given form a two-dimensional space, so there are two independent polarisations for each and no more. Hence
2.4 · From wavenumber to frequency, and the density of modes
We want the count per unit frequency. So substitute from (4.1.3) before differentiating. That way the chain rule is done once and visibly, rather than twice and silently:
Differentiate with respect to to get the number of modes in a thin band, and divide by the volume because §1.3 established that the answer is a density:
modes per unit volume per unit frequency. Check the dimensions before going on: , so is a pure number per unit volume, as a mode count must be. The box has vanished from the answer, exactly as Kirchhoff's argument said it had to.
2.5 · What was thrown away, and how big it is
Two approximations were made in (4.1.4) and both deserve a number rather than a reassurance, because a divergence is about to be blamed on this formula.
The boundary of the ball. Replacing a lattice count by a volume miscounts the points near the spherical surface and near the three coordinate planes. Both errors are surface effects, so they scale as the square of the radius against the volume's cube, and the relative error therefore falls as . Counting the lattice points exactly for spheres of radius lattice spacings gives a ratio to (4.1.4) of at and at . That is to within a percent in the coefficient, which is a correction with a coefficient of order one, exactly as predicted. For a cavity of holding visible light, and the correction is at the level of .
The choice of boundary condition. Perfectly conducting walls gave sines. Periodic boundary conditions give complex exponentials instead, with and of either sign. The lattice spacing doubles, which divides the count by , but the region is now the whole ball rather than one octant, which multiplies it by . The two changes cancel exactly and (4.1.7) is unchanged. It had better be, since §1.2 proved that the walls cannot matter.
What this closes. The mode density (4.1.7) is the object the book has referred to and never built. It is used again in Chapter 4.18 for occupation numbers and in Part V for field quantisation, and it is derived here once, completely, so that it never has to be quoted.
The way to count what a box of light can do is one you already own from the chapter on oscillators. A field in a box is a collection of standing waves, each with its own frequency, each swinging independently of every other, and that decomposition is the same trick that turned two masses on three springs into two separate motions and then turned a chain of masses into a continuous string. Nothing new is needed for three dimensions: there are three whole numbers to choose instead of one, because the wave has to fit a whole number of half-wavelengths across the box in each direction.
Counting them is then geometry rather than physics. Each choice of three whole numbers is a point on a lattice, one to each little cube, so the number of standing waves with pitch below some limit is the number of lattice points inside a ball, which is the volume of that ball divided by the volume of one cube. Ball volumes grow as the cube of the radius, so the running total grows as the cube of the frequency, and the number of new modes appearing in each thin slice of frequency therefore grows as the square. That one fact is what makes the classical prediction diverge, and it comes from nothing deeper than the surface of a sphere growing as the square of its radius.
3 · Rayleigh–Jeans, and a divergent integral
Announce the destination. We give each mode of §2 the energy classical statistical mechanics assigns it, multiply, and integrate. The multiplication takes one line. The integration does not converge, and the whole of §§4 and 5 exists to repair that. Along the way we set out the two experimental facts any correct formula must reproduce, so that when one is produced you can score it against targets that were fixed in advance.
3.1 · Equipartition, derived rather than quoted
Chapter 0.6's Worked example 1 maximised the entropy subject to normalisation and a fixed mean energy, and obtained with and . We want the mean energy of one oscillator, so we need that result with the sum over states replaced by an integral. The reason for the replacement is that a classical oscillator's state is a continuous pair rather than a label from a list.
Making that replacement is a step worth naming as the place where classical statistical mechanics is quietly incomplete. A sum over discrete states becomes an integral over phase space only once you decide how much phase-space area counts as one state, and classical mechanics offers no such unit. Write the unknown unit as , with the dimensions of times , so that
Nothing below depends on . The quantity we want is , which is the second constraint of 0.6's Worked example 1 restated, and a constant multiplying adds a -independent constant to . Differentiation kills that constant. Keep the symbol anyway. Section 5.7 shows that the new constant has exactly the dimensions of the missing unit, and the fact that classical physics had a hole of precisely this shape is not a coincidence.
A mode is a harmonic oscillator, so by Chapter 0.8 §4.3 its energy is a sum of two squares,
Both integrals in (4.1.8) then factorise, and each is a Gaussian. Rather than evaluate them, notice that the only thing needed is how each depends on : in , substitute , which pulls outside and leaves a -independent integral. So each quadratic term contributes a factor to , and
Now that we have as a function of , the mean energy is one differentiation away, so differentiate:
per mode. Half of that comes from the kinetic term and half from the potential term, which is the general statement of equipartition: for every coordinate the energy is quadratic in. The result is independent of , of and therefore of , and that independence is the whole problem. A mode oscillating at is assigned exactly as much energy as one at .
3.2 · The Rayleigh–Jeans law
Energy density equals modes per unit volume per unit frequency, times energy per mode. Multiply (4.1.7) by (4.1.11):
This is the Rayleigh–Jeans law, and everything in it was derived: the mode count from Maxwell's equations and a box, the energy per mode from maximum entropy. There is no fitted parameter anywhere. It is also, at low frequency, right. Right to the precision of the best measurements, in fact, which is what §5 turns into a boundary condition. Radio astronomers use it daily to convert a measured intensity into a brightness temperature, and Worked example 1 computes where it stops being safe.
3.3 · The integral, and the number
The total energy per unit volume is the area under (4.1.12):
Chapter 0.2 §5 classified exactly this: converges only for , and here . So the failure is not a large number, an inaccuracy or a régime of validity. It is a divergence, and the phrase ultraviolet catastrophe is a literal description of where it lives. It lives in the high-frequency tail, which is exactly the region where the mode count grows.
Say the number, because "infinite" is easier to shrug off than an arithmetic comparison. Cut the integral off at a frequency and it evaluates to . At :
| Cut off at | Rayleigh–Jeans energy density | The measured total, for comparison |
|---|---|---|
| (far infrared) | ||
| (violet) | ||
| (soft X-ray) |
By the time the integral has been carried to violet light it has overshot the measured total by a factor of , and every further decade of frequency multiplies the answer by again with no end. Taken at its word, (4.1.13) says that a sealed box of air at room temperature contains unbounded electromagnetic energy, and that opening it would produce an unbounded flash of X-rays. Nobody has observed that.
One clarification, because the mistake is common. The divergence is not caused by the box being idealised, by the walls being perfect conductors, or by the failure of the lattice count near the origin. Those are all effects at low , and §2.5 showed they are relative corrections of order that shrink as grows. Everything that goes wrong goes wrong in the tail, and every correction available is at the other end.
The structure of this failure is one you have met in a different subject. A Cauchy distribution has a perfectly sensible density everywhere, a finite value at every point, a clear peak and a well-defined median. That is Chapter 0.9 §7.5, and it is the same Lorentzian shape as Chapter 0.8 §6.4's resonance. It has no mean, because does not converge, and the mathematics of that failure is identical to (4.1.13)'s: a density falling off too slowly for a moment to exist. In both cases every individual value is finite, every finite sample gives a finite answer, and the quantity you actually wanted does not exist. It is why a sample mean of Cauchy-distributed data does not settle as grows: each new decade of range contributes as much as all the previous ones. The cavity is the more violent case, since there each decade contributes a thousand times the last rather than the same again. But the diagnosis is identical, and it points at the same place, which is the tail.
Where the analogy stops, and it stops somewhere important. A Cauchy distribution's missing mean is a fact about a model, and the model may be the wrong one for the data. Here the divergent quantity is a measured energy with a known value. So this is not a case of a statistic being ill-defined. It is a case of a correct-looking calculation returning for a number that has been weighed. And the repair is not analogous either. One does not truncate (4.1.13), or take its median, or robustify it. The occupancy of the high-frequency modes has to be wrong, which means (4.1.11) is wrong, which means the classical statistical mechanics of an oscillator is wrong.
3.4 · Two measured facts, fixed as targets before the derivation
Set the scoreboard before playing. Two features of real cavity radiation were established experimentally in the 1890s, and they are stated here so that §5's formula can be tested against numbers it was not built from.
Wien's displacement law (1893). The wavelength at which peaks, measured per unit wavelength, satisfies with . Hotter bodies peak bluer, in exact inverse proportion. Separately, and this is the form §5.5 produces directly, the peak measured per unit frequency satisfies . These two statements are not the same statement and their peaks are at different places. Problem 2 works out why.
The Stefan–Boltzmann law (1879, 1884). The total power radiated per unit area from a hole in a cavity is with . The fourth power was Stefan's measurement and Boltzmann's thermodynamic argument. The value of was measured and, before 1900, could not be computed from anything.
Neither is derived in this book, and both are used below only as targets. Note also that the Rayleigh–Jeans law is consistent with neither: it has no peak to displace, and its total is infinite rather than proportional to .
One conversion is needed to use the second of these, and it is geometry rather than physics: the flux out of a small hole is , derived in the grind box. The factor of four is the average of over an outward hemisphere, weighted by solid angle.
Grind box A — why the flux through a hole is times the energy density
Inside the cavity the radiation is isotropic, so the energy density travelling into solid angle about any direction is . That energy moves at speed . Through a hole of area in the wall, the component of the flow along the outward normal is , and only directions in the outward hemisphere contribute.
The integral gives . The integral is with . Hence
So , and a formula for delivers .
Now put one shared quantity of heat into each of those modes. That is what the classical theory says to do, and it is not a guess: the same maximum-entropy calculation that produced the Boltzmann weights hands back the same average energy for every mode whose energy is quadratic in its amplitude, whatever its frequency, because rescaling an integration variable does not care what it is rescaling. Multiply the count of modes by the energy of each, and the prediction is finished in a single line.
What comes out is a spectrum with no peak. It rises as the square of the frequency and never turns over, so the total energy in the box is the area under a curve that climbs forever, and that area is infinite. Not large — infinite, and infinite in a way no care with the constants repairs, because the divergence lives in the tail where the modes are and nothing done at low frequency can reach it. Taken at its word, a cavity at room temperature holds unbounded energy and would empty itself into the ultraviolet the moment anyone opened it.
The failure has a shape worth recognising, since it is the same one that makes a Lorentzian have no mean: a density falling off too slowly to be summed, with everything going wrong out where nobody was looking.
4 · Einstein's A and B coefficients
Announce the destination. We put atoms in the cavity and require that the radiation and the atoms be in equilibrium with each other. Writing every process that changes the atomic populations at a rate proportional to a population, imposing balance, and feeding in Chapter 0.6's Boltzmann weights produces an expression for . The first attempt, with the two processes anybody would write down, fails. And it fails in a way that forces a third process nobody had seen. That is the argument's best moment, and it is why this route is worth taking rather than a shorter one.
This route is longer than the usual textbook one, and the reason for taking it belongs in front of it. The short derivation runs on Bose–Einstein statistics, which requires the symmetrisation of identical-particle states, which requires the notion of an identical particle. That last item is the final chapter of this part, so the short road would make the argument circular. Einstein's road uses nothing but rate bookkeeping and the Boltzmann distribution you already own, and it produces a physical prediction on the way that the short road hides.
4.1 · The setting, and the one identification being made
Put into the cavity a dilute gas of atoms, each of which has two energy levels , and let and be the numbers of atoms in each. Dilute means the atoms interact with the radiation and not with each other, so nothing but radiation moves atoms between levels. Write the level spacing as
where is the frequency of the radiation this transition exchanges energy with, and is a constant.
This is an identification, not a derivation, and it is worth being precise about what is being assumed and what is not. Two things are already established. First, by §1.2 the radiation's spectrum is universal. So a relation between the atom's level spacing and the frequency it couples to cannot depend on which atom is used, because it is a property of the field. Second, by Chapter 0.3 §5's dimensional analysis, no product of powers of and has the dimensions of energy: , which carries no mass whatever the exponents. Some new constant is therefore unavoidable.
What is assumed is that one constant suffices, meaning that the ratio is the same number at every frequency rather than a function of frequency. That assumption is tested three separate times below. Sections 5.5 and 5.6 check it against two measurements the formula was not fitted to, and §6 measures the same constant twice more in experiments involving no cavity at all. If it were false, those checks would disagree.
4.2 · Detailed balance, stated as the assumption it is
The assumption. In equilibrium, each microscopic process is balanced by its own reverse, separately. It is not merely that the total rate into each level equals the total rate out.
Why that is stronger than a steady state. A steady state requires only . With two levels the two conditions coincide, since the only way in to level is the only way out of level . With three or more levels they come apart, and a system can sit in a perfectly steady state with a constant net current circulating round a cycle . Detailed balance forbids that circulation. It is an extra hypothesis about equilibrium and it is not derivable from stationarity.
What justifies it here. It follows from time-reversal invariance of the underlying microscopic dynamics, which electromagnetism has. But this book has not developed the machinery to make that implication precise, so detailed balance is quoted. Its consequences are checked against measurement in §5.
You use detailed balance whenever you write a dissociation constant. A ligand binding a receptor at rate and unbinding at rate sits at equilibrium when those two rates are equal, and the ratio then follows with no further information about the mechanism. The algebra below is the same algebra: one process, its own reverse, and a ratio of rate constants determined by requiring balance. Einstein's is and his is , with the ligand replaced by radiation, and the Boltzmann factor plays the part plays in the thermodynamic version of the same statement.
Where it stops, and the failure is one you have seen argued about. The identity holds only for a genuine equilibrium. Take a three-state system: receptor, ligand-bound, internalised, back to receptor. It can reach a perfectly steady state in which concentrations do not change while a net flux runs round the cycle, sustained by whatever is driving it. That state is stationary and it does not satisfy detailed balance. So the pairwise ratios of rate constants are not fixed by the concentrations, and quoting a for any one step from steady-state data is then wrong. The same distinction is exactly why §4.2 has to state detailed balance as an assumption rather than deducing it from . Two levels hide the difference. Three do not.
4.3 · The populations, from Chapter 0.6
Before any rates, the ratio of populations is fixed. Chapter 0.6's Worked example 1 gives with , and since the numbers of atoms in each level are proportional to those probabilities,
Two remarks before this gets used. Suppose the levels were degenerate, with and states of the same energy. Every state still carries the same Boltzmann weight, so the ratio would be , and every conclusion below acquires a in a predictable place. Section 4.6 says where. Take throughout.
And note the shape of (4.1.15) against something you compute routinely. The fraction of atoms in the upper level is , which is a logistic function of . It is the same expression, symbol for symbol, as a two-outcome logistic model with linear predictor . The Boltzmann ratio is the odds of finding an atom excited, and is the log-odds. That is not an analogy. Both come from normalising two exponential weights. The identity is useful because it means the intuition you have for how fast an odds ratio collapses as the linear predictor grows is exactly the right intuition for how fast the upper level empties as the frequency rises.
4.4 · Two processes, and why they do not close
Now the rates. The atoms are in a bath of radiation of spectral density at the transition frequency. Consider processes whose rate is proportional to the population available to undergo them, which is what "rate constant" means and what mass action assumes.
Absorption. An atom in level absorbs and goes to level . Nothing happens without radiation present, so the rate is proportional to as well as to :
Spontaneous emission. An excited atom drops to level , emitting. An excited atom in the dark still decays, which is what a fluorescence lifetime is. So this rate is proportional to and to nothing else:
That is everything a two-process model can offer: one rate up, one rate down. So impose detailed balance on just these two, and see what spectrum comes out:
Now test it, and the test is the cheapest one available: hold fixed and let grow without bound. The right-hand side of (4.1.18) approaches the finite number . But at fixed and large enough we are in the régime , where §3.2's Rayleigh–Jeans law holds and is measurably correct, and it says , which grows without bound. Note carefully how little of Rayleigh–Jeans is being used: only that as , not the coefficient. Even that weak statement is enough.
So the two-process scheme is not merely inaccurate. It is inconsistent, for every choice of and . Something is missing, and the failure says what kind of thing. The emission rate must be able to grow without bound as does, so there must be an emission process whose rate depends on .
4.5 · The third process, forced
Ask what can be added. The scheme's ground rules are that every rate is proportional to a population and, if it requires radiation, to . Level has one population and there is one field density, so the only term available that is not already present is
radiation already present makes an excited atom drop and emit. This is stimulated emission, and the point to hold onto is that it was not added to the model to explain anything. It was the only repair the model's own rules permitted, and it is required by consistency alone. Einstein wrote it down in 1917. No experiment had suggested it, and none would show it directly for decades.
Count what has just happened. Two rate constants, a balance condition, and Boltzmann weights went in. No new physics, no new constant, no new postulate. And the arithmetic refuses to close. The refusal is not a numerical disagreement that a better value of could fix. The left side of (4.1.18) is unbounded in and the right side is bounded, so no choice of constants works at all.
The only available repair introduces a physical process that nobody had observed, and the argument then determines its rate constant exactly, in §4.6. Stimulated emission is the mechanism of every laser and every maser ever built, and it arrived as the thing that had to exist for a consistency check to pass. Worked example 3 computes how strongly it competes with spontaneous emission under ordinary conditions, and the answer explains why building a laser took forty years after the prediction.
4.6 · The balance, solved, and the first coefficient relation
With all three processes, detailed balance reads
We want , so collect its terms on one side and divide:
The population ratio is the one piece of physics still missing from that expression, so bring it in now. Substituting (4.1.15) gives , and hence
Everything about the spectrum is now in two ratios of rate constants, and two conditions fix them. Here is the first, and it is one line.
Let at fixed . The exponential tends to , so the denominator of (4.1.22) tends to and tends to . But must grow without bound, by the argument of §4.4. A ratio grows without bound only if its denominator goes to zero. Hence
Absorption and stimulated emission have exactly the same rate constant. Nothing was fitted, and one limit did the whole job. (With degeneracies the same line gives , which is the promised predictable place.) Write for the common value, and (4.1.22) collapses to
Notice that this also closes §4.1's identification, which was left standing as an assumption. is a property of the atom and carries no temperature, so the entire temperature dependence of at fixed sits in that exponent. Two atoms coupling to the same with different level spacings would give the one universal two different temperature dependences, and §1.2 forbids that. So is a function of alone, and the same function for every atom. That means the only thing §4.1 assumed beyond what is now proved is that the function is a straight line through the origin. That is a narrow assumption rather than a broad one, and the three checks in §5.5, §5.6 and §6 are tests of the linearity alone.
Recap, since ten lines have gone by. Three rate constants went in, one of them belonging to a process introduced because the other two could not balance. Detailed balance plus Boltzmann weights gave (4.1.22). One limit then gave . What is left is the complete shape of the spectrum, with a single unknown constant standing in front of it. That constant is a property of the atom, and the next section shows it cannot be.
Einstein's route to the answer is the one taken here, and the reason is worth stating. It needs no statistics of light at all, only the bookkeeping of atoms trading energy with radiation, and that matters because the statistics of light cannot honestly be written down until the end of this part. Take atoms with two levels sitting in the radiation and write every process that could change how many sit in each, at a rate proportional to how many are available to undergo it. Two are obvious: an atom absorbs and goes up, an atom drops on its own and emits.
The argument will not close with only those two, and that is the moment worth stopping on. Demand that the populations hold steady with each process exactly balancing its own reverse, feed in the ratio of populations the Boltzmann weights supply, and then drive the temperature up without limit. The radiation must grow without limit too; the two-process bookkeeping says it cannot, no matter what values the two rates are given.
The repair is forced rather than chosen. There has to be a third process in which radiation already present makes an excited atom drop and emit, at a rate proportional to how much radiation is present. Nobody had seen such a thing. The algebra insisted on it, and it is the mechanism of every laser since.
5 · The Planck spectrum, and what is
Announce the destination. One constant is left in (4.1.24), and one condition remains unused: at low frequency the answer must reduce to the Rayleigh–Jeans law, which is right there. Imposing it fixes the constant and finishes the spectrum. We then differentiate the result and integrate it, and compare both against the two numbers §3.4 fixed in advance. Finally we say, carefully, what has and has not been established, and what is at this stage of the book.
5.1 · Rayleigh–Jeans as a boundary condition
Write and take in (4.1.24). Since , the leading behaviour is
That must agree with (4.1.12), which was derived from Maxwell's equations and maximum entropy with nothing fitted, and which is measurably correct at low frequency. Setting the two equal,
The factor cancels from both sides. It had to cancel, since is a ratio of rate constants and cannot depend on the temperature of a cavity the atom happens to be sitting in. That cancellation is a check, not a step. What remains is
and note what has just been established: the rate at which an excited atom decays in the dark is not independent of the rate at which it absorbs. The two are locked together, by a factor built from , and the transition frequency and containing nothing about the atom at all. An atom that absorbs strongly must fluoresce quickly, and the says that a transition at twice the frequency decays eight times faster for the same absorption strength.
5.2 · The Planck spectrum
Substitute (4.1.27) into (4.1.24):
This is the Planck spectrum. Read it as the mode count of §2 times an average energy per mode:
The mode count is untouched. It came from Maxwell's equations and a box, and neither of those is in question. What has changed is the energy per mode, and the way it has changed is the whole content. The quantity has been replaced by something that equals when and falls off exponentially when . High-frequency modes are not sharing the heat. That is what cures §3.3's divergence, and it is what equipartition got wrong.
5.3 · What this argument did, and what it did not do
Be exact about the logical shape here, because it is unusual and a careful reader will notice.
What was derived. That must have the form came from detailed balance plus Boltzmann weights plus the existence of a third process, and nothing else. That is a genuine derivation of the shape, and it is the part that could not have been guessed. Similarly, came from a limit and is exact.
What was supplied from outside. The constant in front came from requiring agreement with the Rayleigh–Jeans law, in the régime where that law is correct. That is the very law being replaced. So the classical theory was not merely refuted here. It was used, as a boundary condition, to fix two numbers in a formula whose form came from elsewhere. That is a legitimate move, and it is what matching asymptotics always does. But it is not the same as deriving a law from nothing, and the chapter would be dishonest to let it read as though it were.
What repairs it. Chapter 4.18 §5 derives (4.1.28) a second time and by a completely different route. That route counts the states of the field with Bose–Einstein statistics, which needs the symmetrisation postulate and therefore cannot be done here. The two routes share no step. When the second one arrives, the constant in front is computed rather than matched, and the present argument stops being a fit and becomes one of two independent confirmations. Until then, the honest statement is this. The form is derived, the normalisation is matched, and three checks follow immediately that were not used in the matching: Wien's 1896 distribution law as a limit (§5.4), the displacement constant (§5.5), and Stefan's (§5.6).
5.4 · The two limits, and the two empirical laws they reproduce
Low frequency, . By construction and (4.1.28) becomes , the Rayleigh–Jeans law. There is no information in that. It is the boundary condition being satisfied.
High frequency, . Now and
This is Wien's distribution law, the formula fitted to the high-frequency measurements in 1896 and known to fail in the infrared. It was not used as an input anywhere above, so this is the first real prediction. The same expression reproduces both of the empirical laws that were known before 1900, each in its own régime, and interpolates smoothly between them. How badly each fails outside its own régime is arithmetic:
| Rayleigh–Jeans Planck | Wien Planck | |
|---|---|---|
The two ratios are and , both read straight off (4.1.28). The crossover sits at of order one, so the peak of the spectrum is exactly where neither classical formula works. That is why the two camps could each claim their own half of the data, and why neither could claim the middle.
5.5 · Wien's displacement, with the arithmetic
The first thing §3.4 fixed as a target. Differentiate (4.1.28) with respect to at fixed and set the result to zero. Substituting makes this a question about the single function , because everything else is a constant multiplier. So the answer will be a pure number, and will be proportional to before any arithmetic is done. Using the quotient rule,
The denominator is finite and non-zero for , so set the numerator to zero. Divide by , which is legitimate since at a maximum:
This is transcendental and has no closed-form root, so it has to be solved numerically. The method is Chapter 0.3's first-order Taylor expansion run backwards. Near any point, . The step that would send to zero is therefore , and iterating that step is Newton's method. With and , starting from :
Three steps, and the root is . Hence , or
which is . Compare §3.4's measured . Nothing was fitted to this. The only number put into the formula was matched at low frequency, and the peak sits at , where that low-frequency law is wrong by a factor of .
Problem 2 does the same calculation per unit wavelength, gets the root of , and recovers Wien's constant against the measured . It also settles the trap that .
5.6 · Stefan–Boltzmann, with the arithmetic
The second target. Integrate (4.1.28) over all frequencies. The substitution is the one to make, because it removes every trace of from the integrand and leaves it in the prefactor, where its power can be read off. With and ,
The has appeared before the integral has been touched, which is the point of doing the substitution first: the fourth power comes from and nothing else, and would be there whatever the remaining integral turned out to be. That integral is a pure number, and the grind box evaluates it as . So
What §3.4 asked for was the flux out of a hole rather than the energy density inside, and Grind box A supplies the conversion . So
Put the numbers in, piece by piece, so the arithmetic is checkable. We have , then , then , and finally . The numerator is and the denominator is , so
Against §3.4's measured . Every digit shown agrees. The reason it is that good rather than merely good is that since 2019 , and are exact by definition, so the modern "measured" is this arithmetic. Before 2019 the comparison was a genuine confrontation between an experiment and a formula, and it agreed to the precision of the experiment. Two measurements, both fixed in §3.4 before the derivation ran, both recovered from one formula containing one constant.
Grind box B — the integral , and from Parseval
Step 1. Expand the denominator. For write and expand the geometric series in , which converges because :
Step 2. Interchange sum and integral. Every term is non-negative, so the partial sums increase monotonically and the interchange is the case Chapter 0.2 §4 discussed. (The theorem that makes it airtight is monotone convergence, proved in Chapter 4.3.)
the inner integral being by three integrations by parts, or by differentiating three times with respect to .
Step 3. The sum, from Parseval. Take on and expand it in Chapter 0.9 §1's orthonormal basis . The sine coefficients vanish because is even, and
the second by two integrations by parts, the boundary terms surviving only in the second one because and . Parseval's identity says . That is Chapter 0.5 §2, transferred to this basis by the completeness that Chapter 0.9 §1.3 flagged as quoted and Chapter 4.3 proves. Here . So
Rearranged, , so . Therefore
(Checked numerically: the same integral evaluated by quadrature gives .)
5.7 · What is, at this stage
Say plainly what has and has not been claimed about the new constant, because everything later in this part depends on not over-reading it now.
Its dimensions. From (4.1.14), is an energy divided by a frequency, so . Those are the dimensions of action, which is energy times time, and which is also momentum times length and angular momentum. Chapter 1.2 built a whole subject out of a quantity with those dimensions, and Chapter 1.3 gave phase space its area. So the coincidence is not one. It is the subject of Chapter 4.10.
Its value. Fitted. The formula's shape was derived and its normalisation matched, and the one remaining number is whatever makes (4.1.28) agree with the measured spectrum: . Since 2019 that value is exact by definition, because the kilogram is now defined by fixing it. But that is a statement about the SI, not about physics, and in 1900 it was a fit to a curve.
What it is not, yet. Nothing above says that energy comes in indivisible units, that the oscillator's energy levels are discrete, or that light is made of particles. All that has been shown is that the average energy of a cavity mode is rather than . Planck himself resisted the stronger readings for years and was right to want more evidence. Section 6 supplies it, from two experiments that involve no cavity.
One structural remark, and it is the reason to have kept in §3.1. The classical partition function (4.1.8) needed a unit of phase-space area and classical mechanics had none to offer. The constant that has just appeared has exactly the dimensions of phase-space area, meaning momentum times length, which is Chapter 1.3's . It is the missing unit, and Chapter 4.10 makes the statement precise by showing that a classical orbit enclosing area in phase space corresponds to about quantum states.
Two unknown ratios survive the balancing, and two demands fix them. Requiring the answer to keep growing as the temperature rises makes absorption and the new third process exactly equally strong. Requiring it to reproduce the classical law at low frequency, where that law is right and measurably so, fixes the other. What is left is a formula with one new constant in it, and the honest description is not that a law was derived from nothing: the shape came from the balancing, and the numbers in front came from a limit supplied by the theory being replaced. A second, independent route arrives near the end of this part, and it is that route which makes the first more than a fit.
The new constant is at this stage nothing but a fitted number carrying the units of energy multiplied by time. Leave it that way. What it does is set, for each temperature, a frequency above which the shared quantity of heat is no longer enough to stir a mode at all, so those modes sit the whole thing out. That is what cures the divergence: the modes still pile up as the square of the frequency, but their occupancy dies faster than any power.
Two measurements the formula was never fitted to — where the peak sits, and how the total grows with temperature — come out right, the second of them to nine figures.
6 · Light carries momentum in parcels: the photoelectric effect and Compton
Announce the destination. Two experiments, neither involving a cavity. The first measures a second time, from a straight line whose slope has nothing to do with §5, and shows that light delivers energy in amounts fixed by frequency. The second shows that the parcel carries momentum too, and does so by a calculation Chapter 2.5 has already done in full. Together they promote the relation , which Chapter 2.5 could only note as a structural resemblance, into a law.
6.1 · The classical prediction, stated before it is contradicted
Shine light on a clean metal surface in vacuum and electrons come off. Classically light is a wave carrying energy density spread continuously over the surface, and an electron bound in the metal is a charge in that oscillating field, driven and gradually shaken loose. Three predictions follow, and it is worth writing them down before the data, because each is definite.
(i) The maximum kinetic energy should rise with intensity. A stronger field shakes harder, so an electron freed by a bright beam should leave faster than one freed by a dim beam of the same colour.
(ii) There should be no threshold frequency. Any colour should work eventually: energy accumulates, and enough of it accumulated is enough of it accumulated.
(iii) There should be a measurable delay at low intensity. Estimate it. A beam of intensity falling on an atom of radius delivers power to that atom. With , which is bright, the area is and the power is . To accumulate the needed to free an electron from sodium takes
That estimate is generous to the classical picture in every direction. It assumes the atom collects the whole beam falling on its geometric cross-section and loses none of it. Reduce the intensity by a factor of a million, which is still an easily visible beam, and the delay becomes , or four and a half months.
6.2 · What is measured
None of the following is derived here. All of it was established between 1902 and 1916, most decisively by Millikan, who set out to disprove Einstein's 1905 relation and measured it instead.
(i) The maximum kinetic energy of the emitted electrons is completely independent of intensity. Increasing the brightness increases the number of electrons in exact proportion and does not change their energies at all.
(ii) There is a sharp threshold frequency , characteristic of the metal. Below it no electrons are emitted at any intensity whatever, and above it emission begins immediately.
(iii) No delay is observed. Emission follows illumination within the resolution of the measurement, which by the 1920s was nanoseconds. That is ten orders of magnitude away from (4.1.39).
(iv) The maximum kinetic energy is a linear function of frequency. And here is the part that matters: the slope of that line is the same for every metal. Only the intercept changes. Work functions used below, in electronvolts: caesium , sodium , zinc , platinum .
Every one of (i), (ii) and (iii) contradicts a prediction of §6.1, and (iv) is not a prediction of §6.1 at all. A wave picture supplies no reason for a universal slope, because it supplies no constant with the right dimensions.
6.3 · , and a second measurement of
The reading that accounts for all four facts at once is that light of frequency delivers its energy to one electron in an indivisible amount , and it is the same , on the hypothesis being tested. Then an electron that absorbs one such amount and pays the minimum energy required to escape the metal emerges with at most
Check each fact against it.
- Intensity sets how many amounts arrive per second and not how large each is, so is independent of intensity. That is fact (i).
- Below no single amount is large enough, and no waiting helps, because energy is not being accumulated. Those are facts (ii) and (iii).
- And (4.1.40) is a straight line in with slope and intercept , so the slope is universal and the intercept is the material's. That is fact (iv).
Measuring the slope measures . In practice one measures the stopping potential , the retarding voltage that just prevents any electron from reaching the collector, so that and
Take the two extreme mercury lines from a sodium surface. At , and . At , and . The slope is
so . That is §5.7's value to six figures, obtained from a photocell and a voltmeter with no furnace, no spectrometer and no thermodynamics anywhere in the apparatus. Two measurements of the same constant by methods sharing no equipment is what turns §4.1's identification from an assumption into a result. Problem 3 does the full five-point fit.
6.4 · Compton scattering, collected from Chapter 2.5
Compton, in 1923, sent molybdenum X-rays of wavelength into graphite and measured the wavelength of the scattered radiation as a function of scattering angle. He found a shift to longer wavelength, growing with angle, reaching about in backscatter, and independent of the incident wavelength and of the scattering material. The measurements are quoted. The calculation below is not.
Chapter 2.5's Worked example 1 did this calculation in full and this chapter reuses it unchanged, exactly as that chapter said it would. Here is the setup it used. A particle of energy and momentum , massless so that by 2.5 §4.3, strikes an electron at rest. Conservation of four-momentum is imposed. The unknown electron four-momentum is isolated on one side and squared, which replaces it by and removes the recoil direction from the problem entirely. The result is
Two things about that formula matter here and did not matter in Chapter 2.5.
The one quoted input was , and it is no longer quoted. Chapter 2.5 flagged as an import and named this chapter as where it would be argued for. It has now been argued for twice, in §5 and §6.3, so the flag is discharged and the Compton formula stands on this book's own results.
The independence of is the whole evidential weight. A classical wave shakes an electron at the driving frequency, and an oscillating charge re-radiates at the frequency it is oscillating at. So the scattered light has the same wavelength as the incident light, and the predicted shift is exactly zero, at every angle, for every material. The measured shift is not small. At on X-rays it is , a change, and at it is , or . This is not a discrepancy in a coefficient. It is a measured effect where the classical prediction is zero, and it is reproduced to the precision of the experiment by treating the encounter as a two-body collision with the four-momentum bookkeeping of Part II applied without modification.
6.5 · , promoted from a suspicion to a law
Chapter 2.5 §7.1 ended with a deliberately unfinished observation. It had built the wave four-vector and proved it a four-vector by counting crests. It then noted that for light in vacuum , so . Setting that beside from its §4.3, it remarked that the two objects have the same transformation law, the same null condition and the same relation between their time and space components. It closed with "Chapter 4.1 supplies the missing constant." Here it is.
A plane light wave travelling along the unit vector has, from Maxwell's equations and Chapter 2.5 §4.3, energy and momentum in the ratio with along . So whatever amount of energy the wave delivers, it delivers with it four-momentum
The two are proportional, with the same unit vector, because both are built from the same propagation direction, so and the entire question is what is. Sections 5 and 6.3 established that the energy is delivered in amounts , so . Give that combination its own name, since every equation in the rest of this part carries it:
which is a change of units and nothing more. Here goes with and goes with , so that . Then
Read off the two components separately, because each is a statement that was measured independently. The time component is , which is §6.3's photoelectric line. The space component is , so . That is precisely the momentum Chapter 2.5's Compton calculation assigned the incident radiation, and the calculation reproduced the measurement. So each half of (4.1.46) has its own experiment, and the two experiments are unrelated. That is what promotes the relation from a structural suspicion about two null four-vectors into a law.
One consequence is worth having now, since Part V leans on it. Because (4.1.46) is an equation between four-vectors, it holds in every inertial frame with the same : Lorentz-transform both sides and the relation is unchanged. So is a Lorentz invariant, which a relation between a frame-dependent energy and a frame-dependent frequency had no obligation to be.
6.6 · What has not been shown
None of §6 shows that light "is a stream of particles", and reading it that way will make the next ten chapters incomprehensible.
What has been shown is that light exchanges energy and momentum with matter in amounts and , indivisibly. Everything that made light a wave is untouched.
- The same X-rays that scatter off an electron like billiard balls are diffracted by a crystal according to Bragg's condition, which is an interference condition and needs a phase.
- The same light that ejects electrons one at a time from a photocathode produces two-slit interference fringes whose spacing tracks its wavelength.
- And the derivation of in §2, which is an essential ingredient of the Planck spectrum, is a count of standing waves and would be meaningless for a gas of classical particles.
So the position at the end of this chapter is that two descriptions are each supported by measurements the other cannot account for. That is not resolved by declaring light to be "both", nor by saying it is "sometimes one and sometimes the other". It is resolved in Chapter 4.2 by changing what a state is, so that the two descriptions stop being alternatives. That change is the whole content of the postulates. Hold the tension rather than dissolving it prematurely.
Two experiments now, of quite different kinds, saying different things. The first is about arrival. Shine light on a metal and electrons come off with a maximum energy that depends on the colour and not at all on the brightness, and below a threshold colour nothing comes off at all however long anyone waits. A wave delivering its energy smoothly and evenly cannot do that, and the classical estimate of how long one atom would need to gather enough energy from a dim beam runs to seconds, against the nanoseconds measured. Energy arrives in lumps whose size is set by the frequency, through the same constant that turned up in the cavity.
The second experiment is about collisions. Bounce X-rays off electrons and the wavelength shifts by an amount fixed by the scattering angle and by nothing else — not by the incoming wavelength, not by the target material. A wave shaking an electron makes it re-radiate at the driving frequency and produces no shift whatever, so that is not a small discrepancy but a contradiction. Treat the encounter instead as two particles colliding, with the relativistic bookkeeping of the second part used unchanged and nothing added, and the measured shift falls out exactly.
So the lump carries momentum as well as energy, in the proportion the wave description had already quietly suggested, and a suspicion recorded two parts ago becomes a law.
7 · Spectral lines, and why an orbiting electron is a catastrophe
Announce the destination. The last two failures are about atoms, and they are of a different kind from the first two. Those were numerical: a formula gave the wrong number, or infinity. These are structural: classical mechanics has no way to produce the form of what is observed, and the one classical model of an atom that reproduces the right energy scale destroys itself in sixteen picoseconds. We do the integral and give the number.
7.1 · The lines, and the arithmetic that fits them
An electric discharge through hydrogen emits light at sharp, isolated wavelengths and nowhere in between. The four visible lines and one ultraviolet, in vacuum, in nanometres: , , , , . Quoted, not derived. The same goes for the measured ionisation energy of that §7.4 scores its calculation against. Every element behaves the same way, each with its own set of lines, which is why a spectrum identifies a substance.
Balmer noticed in 1885 that those five numbers are fitted by one formula. Take the reciprocal of each wavelength and divide by for in turn:
| / nm | / | ratio / | ||
|---|---|---|---|---|
The five ratios agree to nine parts in a million. That is not a fit with five parameters. It is one number, , reproducing five measurements. Rydberg then generalised it, and the generalisation predicts series nobody had looked for:
Setting predicts an ultraviolet series starting at , found by Lyman in 1906. Setting predicts an infrared series starting at , found by Paschen in 1908. So (4.1.47) is not a curve-fit but a formula with predictive content, and it is left entirely unexplained here. Chapter 4.13 derives it, including the value of , and that derivation is one of the three or four things quantum mechanics is believed for.
7.2 · Nothing classical produces whole numbers
Look at what (4.1.47) is made of: two integers and one constant. Now ask where a classical atom could get an integer from.
A charge moving in a Coulomb field has, by Chapter 1.3's Kepler problem, a continuous family of bound orbits: pick any energy and any angular momentum consistent with it, and there is an orbit. The orbital frequency is a continuous function of the energy, and a charge circling at frequency radiates at and its harmonics, by Chapter 2.1's wave equation with a periodic source. So a gas of classical atoms with energies spread over any range at all should emit a continuum. What is measured is a set of isolated lines whose positions are fixed to six figures and are the same for every hydrogen atom in the universe.
This is a different kind of failure from §3.3's, and worth naming as such. There, a calculation produced a number that was wrong. Here, the classical framework has no room for the observed structure: there is no continuous parameter to tune, no coefficient to adjust, and no limit in which integers emerge from (4.1.47)'s left-hand side. Integers appear in classical physics in exactly one place, as mode numbers of a wave confined to a region, which is §2.1's . Following that hint is what the rest of Part IV does, with Chapter 4.13 supplying the two integers in (4.1.47) specifically.
7.3 · Larmor's formula, and the size of an atom
The formula. A non-relativistic point charge with acceleration radiates energy at the rate . This is a consequence of Maxwell's equations, derived by computing the retarded fields of an accelerated charge and integrating the Poynting flux over a distant sphere. That derivation is not in this book, so the formula is quoted. It is the same physics that makes a radio transmitter work, and it is not in any doubt.
This book's rule on constants is to use the measured combination rather than a product of separately measured pieces. So the Coulomb constant is written as , which is Chapter 0.3 §5's symbol and value, with the measured quantity. It is not reassembled from a tabulated and a tabulated . Since , Larmor's formula becomes .
The length. Atoms are about across, from the density of solids, from the viscosity of gases and from X-ray crystallography. Take the electron's orbital radius as , which is the measured value for hydrogen.
7.4 · The integral, and sixteen picoseconds
Model the atom as an electron in a circular orbit of radius about a proton. Two things follow before any radiation is considered, from Newton's second law and the Coulomb force alone. Balancing the centripetal acceleration against the Coulomb attraction gives , so . That makes the kinetic energy , and the total energy
Pause on the first of these before continuing, because it is the one thing the classical model gets right. At the measured radius, , against a measured ionisation energy for hydrogen of . The classical picture has the energy scale of an atom right to a part in a thousand. Everything else about it is about to fail completely. That combination of right scale and wrong structure is what makes the failure interesting rather than merely wrong.
Now let it radiate. Substituting the acceleration into Larmor's formula,
At this is , which for one atom is enormous: it is of binding energy divided by . That crude estimate is the right order and the exact answer is three times smaller, because the power grows as the orbit shrinks.
Do it properly. The orbit's energy falls at the rate it radiates, so . Differentiate (4.1.48) with respect to and use the chain rule, since is what is changing:
That is separable, in the sense of Chapter 0.8 §2.1. So put every on one side and every on the other, and integrate from the initial radius to zero:
Both sides carry a minus sign and a factor of three, and both of those cancel. Solving that last equality for the time gives
The numbers, piece by piece. We have , then , then , and finally . The numerator is and the denominator is , so
Two derived checks on that number, both exact rather than numerical. The crude estimate equals , which is exactly , so the ratio is three, as the arithmetic above showed. And the number of orbits completed is with . Substituting turns that into an elementary integral and gives exactly , which with is
So the classical hydrogen atom completes about two hundred thousand orbits and then does not exist, sixteen picoseconds after it is assembled. And the light it emits on the way is a chirp: the orbital frequency starts at , in the ultraviolet, and rises without bound as , sweeping continuously through every frequency above that. Compare with §7.1's five sharp lines whose positions are stable to six figures.
Both features are wrong, and they are wrong in different directions. The atom is not merely unstable on some long timescale. It is unstable on a timescale shorter than the period of any molecular vibration and shorter than the fastest chemical process anyone can name, so that no sample of hydrogen could ever be assembled. And the spectrum is not merely shifted. It is continuous where the measurement is discrete. Nothing about the model can be adjusted to repair either, because both come from the same two things: that a bound charge can have any energy, and that an accelerating charge radiates.
7.5 · The four failures, collected
Line them up with their numbers, because the list is the chapter's product.
| What classical physics predicts | What is measured | Discrepancy |
|---|---|---|
| Cavity energy density | , finite | Infinite against at |
| Photoemission delay at , rising with intensity, no threshold | Delay under , independent of intensity, sharp threshold | in time, and two qualitative facts reversed |
| Compton shift at every angle | at , independent of and material | A measured effect against a prediction of zero |
| Continuous emission, and the atom collapses in | Five sharp lines fitted by two integers, and atoms are stable | Structural: no continuous parameter can produce integers |
Notice what the four have in common, because it is the thing Chapter 4.2 changes. In every case the classical theory assigns a continuous range of values to something the measurement finds restricted. The list is the energy of a cavity mode, the energy transferred to one electron, the momentum carried by a light beam, and the energy of a bound electron. And in every case the restriction is set by one constant with the dimensions of action, measured twice here by unrelated means and agreeing.
The last two failures are about atoms, and they are worse than the first two, because they are not quantitative errors but structural ones. Heated gases emit at sharp isolated frequencies, and for hydrogen those frequencies fit one arithmetic formula built out of two whole numbers with a single constant in front, to a part in a hundred thousand across the whole visible series. Nothing in classical physics produces whole numbers. A charge orbiting freely can circle at any radius at all, so it can emit at any frequency at all, and a continuous smear is precisely what is not seen.
Worse still, it cannot orbit. An accelerating charge radiates, which is the same theory that explained the radio transmitter and is not optional, so an orbiting electron bleeds energy continuously, spirals in, and reaches the nucleus. Doing that integral gives sixteen picoseconds, after two hundred thousand turns, with the emitted frequency sliding upward the entire way. Every atom in the universe should have collapsed long before anyone could look at one.
So the object needs replacing rather than adjusting, and what replaces it is a state that is a direction in an abstract space instead of a position and a velocity. The mathematics for that was built nine chapters ago and has been waiting since. The next chapter puts the physics back into it.
8 · Worked examples
Chapter 3.9 §1.2 quoted the microwave background as thermal radiation at , isotropic to one part in , and used its isotropy as the evidence for the cosmological principle. It said nothing about the word "thermal". Using (4.1.28) and nothing else. (a) Where does it peak? (b) How much energy does it carry? (c) How many photons per cubic centimetre, and what is the mean energy of one? (d) Up to what frequency may a radio astronomer use the Rayleigh–Jeans law instead?
(a) The peak. Straight from (4.1.34):
whose wavelength is . That is why the instruments that measure it are millimetre-wave telescopes, and why it was found with a horn antenna rather than a photographic plate.
(b) The energy. From (4.1.36), with :
which in more digestible units is , or about a quarter of an electronvolt per cubic centimetre, everywhere in the universe. The flux crossing any surface one way is .
(c) How many photons. This needs a different integral, since a photon count weights each mode by the number of energy parcels in it rather than by their energy. So divide the integrand of (4.1.35) by before integrating. The same substitution gives
Grind box B's method applies verbatim with replaced by : expand the denominator, interchange, and use , giving . Unlike this sum has no closed form, because the Parseval trick produces even powers only. So evaluate it numerically. The first six terms give and the tail beyond is bounded by , and carrying it properly gives . Hence the integral is . With ,
Four hundred microwave-background photons in every cubic centimetre here, and in every cubic centimetre between here and the edge of the observable universe. (The room you are in also holds its own thermal radiation at , and since that is times as many photons, but only inside the room.) The mean energy of one background photon is the ratio of (b) to (c), and the temperature cancels out of it as a matter of algebra:
Note that this is not . The mean photon energy and the peak of the energy spectrum are different questions about the same distribution and they have different answers, which is the same distinction Problem 2 turns on.
(d) Where Rayleigh–Jeans is safe. From §5.4 the ratio of the Rayleigh–Jeans value to the true one is . Setting that to and solving numerically gives . Setting it to gives . Converting with :
So a radio astronomer working below a gigahertz may use freely, with an error under one per cent, and quote a brightness temperature, meaning an intensity expressed as the temperature a blackbody would need to produce it. Above about that convention starts to lie, and at the peak the Rayleigh–Jeans value is too big by a factor of . Solving gives , so the error reaches a factor of two at . The ultraviolet catastrophe is not a remote pathology: for the most-studied radiation field in astronomy it is a factor of two before the spectrum has even reached its peak.
Using (4.1.43) with Compton's molybdenum line, . (a) Tabulate the shift, the scattered photon's energy and the electron's recoil energy at four angles. (b) Explain why the effect is invisible with visible light. (c) State what wavelength is needed for a one per cent shift, and what that implies about when the experiment became possible.
(a) The incident photon energy is . The shift is with , and the electron takes because the collision is elastic in the four-momentum sense. Energy is conserved, and the electron's share is whatever the photon lost.
| / pm | / pm | shift | / keV | / keV | ||
|---|---|---|---|---|---|---|
A wavelength change is comfortably resolved by a crystal spectrometer, and the electron energies are in the range an ionisation chamber measures. Note that even in backscatter the photon keeps of its energy: the electron is of rest energy and a photon barely moves it, which is why the shift is a small fraction here and why it is not for gamma rays.
(b) Why not with visible light. The shift (4.1.43) is an absolute length, the same at whatever the incident wavelength. What a spectrometer resolves is a relative change. For green light at ,
which is five parts in a million. That is below what a nineteenth-century spectroscope could see, and swamped in any case by the thermal motion of the scatterers. The effect was always there. It was invisible because the natural yardstick is a picometre and visible wavelengths are half a million of them.
(c) The threshold of visibility. A one per cent shift at needs , an energy of . That is a soft X-ray. So the experiment became possible exactly when X-ray sources and crystal spectrometers did: von Laue's diffraction in 1912, the Braggs' spectrometer in 1913, Compton's measurement in 1923. The delay between the physics being true and the physics being visible was set by one length, and Chapter 0.3 §5's dimensional argument already said that is the only length that can be built from , and .
Section 4.5 forced stimulated emission into existence. Using only (4.1.23), (4.1.27) and (4.1.28). (a) Find the ratio of stimulated to spontaneous emission rates for an atom sitting in thermal radiation. (b) Evaluate it for a red laser transition at room temperature and for a microwave transition at room temperature. (c) Say what would have to be true for stimulated emission to dominate, and show it cannot happen in equilibrium.
(a) The two emission rates from (4.1.19) and (4.1.17) are and , so their ratio is , and both and every property of the atom cancel. Substituting (4.1.28) for and (4.1.27) for :
Everything has cancelled but one exponential. The ratio is a property of the radiation, not of the atom. That is exactly what one should expect from §5.1, where turned out to contain nothing atomic.
(b) Two numbers. For the helium–neon laser line at , , and at , , so and
At optical frequencies stimulated emission is, in ordinary surroundings, utterly negligible: an excited atom in a warm room fluoresces spontaneously and the third process Einstein was forced to invent contributes one part in . Now take a microwave transition at , the frequency of a domestic oven. There , so , and since we may use :
Stimulated emission dominates spontaneous emission by three orders of magnitude at microwave frequencies and is suppressed by thirty-three orders of magnitude at optical ones, and the whole of that swing is the in (4.1.27) combined with the Boltzmann factor. That is the quantitative reason the maser was built in 1953 and the laser not until 1960: at microwave frequencies you are working with a process that thermal radiation already favours, and at optical frequencies you are fighting a factor of .
(c) What amplification requires. A beam grows only if stimulated emission outruns absorption. Those two rates are and , with the same in both, by (4.1.23). So the condition is
But (4.1.15) says , which is less than for every positive and approaches only as . No system in thermal equilibrium can amplify light at any temperature whatever. Amplification requires a population inversion, which is by definition not an equilibrium state and has to be pumped and maintained. The equality , which cost one line in §4.6, is what makes that statement exact rather than approximate: if absorption and stimulated emission had different rate constants, a clever choice of materials might have amplified light in equilibrium, and the balance argument forbids exactly that.
9 · Your turn
Problem 1 · is the catastrophe an accident of three dimensions?
Repeat §2's mode count in spatial dimensions. The volume of a ball of radius in dimensions is with , , . (a) Show that the mode density per unit volume is for a constant you should give, taking one polarisation for simplicity so that the answer is clean. (b) Write the Rayleigh–Jeans law in dimensions and determine for which the total energy diverges. (c) Do the same for the Planck spectrum and find the analogue of the Stefan–Boltzmann law. (d) State in one sentence what (b) and (c) together say about the diagnosis in §3.3.
Solution
(a) The allowed still form a lattice of spacing restricted to the positive orthant, which is now a fraction of the ball rather than . So , and with ,
so . Check : , giving , which is (4.1.7) with the factor of two for polarisation removed. ✓ Notice the bookkeeping: the orthant's cancels the inside , and the lattice spacing's cancels the , so the only surviving geometry is . That is the same kind of cancellation §2.5 found between conducting and periodic boundary conditions.
(b) Equipartition is dimension-independent, since (4.1.11) counted quadratic terms rather than directions. So and
diverges for every , since the exponent is never less than . The ultraviolet catastrophe happens in every dimension. It is not a feature of three.
(c) The Planck form replaces by , so
and the remaining integral converges for every , since near the integrand behaves as , integrable for , and at large it is killed exponentially. Grind box B's expansion gives it as . So : the fourth power of the Stefan–Boltzmann law is , and the is the dimension of space. In the law would be , and in it would be with .
(d) The divergence comes from a mode count that grows without bound together with an energy per mode that does not fall, and only the first of those is dimension-dependent. So the diagnosis in §3.3 was correctly aimed at equipartition rather than at the counting.
Problem 2 · Wien's displacement law per unit wavelength, and a trap
The spectrum measured per unit wavelength, , is defined by , since the energy in a band must not depend on which variable labels the band. (a) Convert (4.1.28) to . (b) Maximise it and show the condition is with . Solve by Newton's method and obtain Wien's constant . (c) Compute and explain, in terms of (a), why it is not and why no error has been made. (d) The Sun's surface is at . Where does it peak in wavelength, and where in frequency?
Solution
(a) With we have , so . Substituting into (4.1.28),
(b) Put , so and . The same quotient-rule computation as (4.1.31), with replaced by , gives , that is . Newton from with , :
and the root is . Hence
against §3.4's measured . ✓ Note is the whole of the constant content. The transcendental root supplies the rest.
(c) . Not , and nothing is wrong. The two peaks are the maxima of two different functions: and differ by the Jacobian factor , which is not constant, so it shifts where the maximum sits. Asking "at what wavelength does a blackbody peak" is therefore an incomplete question. One must say per unit what. The total energy is the same either way, since that is an integral rather than a maximum, and integrals do transform correctly under a change of variable.
(d) , green-blue, near the middle of the visible band. And , whose wavelength is , in the near infrared. Both statements are true about the same Sun. The familiar claim that the Sun peaks where human vision peaks is a statement about and is not true of .
Problem 3 · measure and a work function from five data points
A sodium photocathode is illuminated in turn by five mercury lines and the stopping potential measured. This is the same quoted class of measurement as §6.2's. The data are idealised, with noise removed so that the arithmetic is checkable:
| / nm | |||||
|---|---|---|---|---|---|
| / V |
(a) Convert to frequencies and verify that is linear in by checking that the slope between successive points is constant. (b) Extract and the work function . (c) Predict the threshold wavelength and check it against the value quoted in §6.2. (d) A student argues that the linearity could be a coincidence of this metal. What one further measurement settles it, and what would it show?
Solution
(a) gives, in units of : , , , , . Take differences of adjacent columns, both of and of , and divide. In units of the four ratios are
each numerator being in volts and each denominator in units of . Four independent slopes agreeing to about one part in , the drift in the last one being the rounding of the tabulated voltages. That is the content of the experiment: not the value of any single point, but that the points lie on a line.
(b) The slope is by (4.1.41), and the end-to-end value is , so
For , use any point: at , , so .
(c) , which is green. So a sodium photocathode is blind to red and orange light of any brightness whatever. This agrees with §6.2's quoted work function, as it must, since that is where the value came from.
(d) Repeat the whole experiment with a different cathode: zinc, say, with . The prediction is that the line moves down by at every frequency while its slope is unchanged to within the experimental error, and that the threshold moves from to , in the ultraviolet. A slope that came out different for zinc would mean is a property of the metal, which would destroy the whole reading. A slope that matches means the constant belongs to the light. This is exactly the measurement Millikan made, and it is why the photoelectric effect, rather than the blackbody spectrum, is what persuaded people.
Problem 4 · run the argument without stimulated emission
Suppose Einstein had not been forced to add the third process, and instead had allowed the two remaining coefficients to be arbitrary functions of frequency. (a) Show that the two-process balance (4.1.18) can be made to agree with the measured spectrum at high frequency, and identify what would have to be. (b) Show that the resulting law then fails at low frequency, and compute by what factor it is wrong at . (c) Deduce that the existence of stimulated emission is equivalent to the statement that as at fixed , and say in one sentence which measurement establishes that. (d) Suppose instead that stimulated emission exists but with . What does the spectrum become, and at what temperature does it become absurd?
Solution
(a) (4.1.18) gives , which is exactly the shape of Wien's distribution law (4.1.30). Matching gives , which is the same ratio §5.1 found, since (4.1.30) is the limit of (4.1.28) and the two agree there. So the two-process scheme reproduces the high-frequency data perfectly, which is why Wien's law survived from 1896 to 1900.
(b) The ratio of Wien's law to the truth is , from §5.4. At that is : the two-process theory predicts about a tenth of the energy actually present, and the shortfall grows without bound as since . This is the discrepancy Rubens and Kurlbaum measured in the far infrared in 1900, and it is what prompted Planck to look for an interpolation.
(c) From (4.1.22), is bounded as unless the denominator vanishes there, and the denominator is : unbounded growth requires , and is the existence of stimulated emission. The two statements are therefore equivalent. What establishes the unbounded growth is the Rayleigh–Jeans law, itself derived in §3.2 and confirmed by every low-frequency measurement, which is the same infrared data of (b).
(d) With , (4.1.22) gives
which is finite and positive for all and tends to as , so it never reproduces the Rayleigh–Jeans growth. The absurdity is not a divergence but a ceiling: the energy density of cavity radiation would saturate, and a furnace at would hold no more low-frequency radiation than one at . The failure sets in as soon as falls below about , since that is where the true starts growing like while this one flattens. At room temperature , so the whole of the radio, microwave and far-infrared spectrum is in the failing régime.
Problem 5 · how bad is the catastrophe, in kilograms
Take a cubic metre of empty space at and integrate the Rayleigh–Jeans law up to the frequency at which one parcel of light would carry the rest energy of an electron, . (This is not a principled cut-off. It is a frequency you can name.) (a) Compute . (b) Compute the energy. (c) Express it as a mass by and compare with the true answer from (4.1.36). (d) State precisely why choosing a more sensible cut-off would not rescue the classical theory.
Solution
(a) .
(b) From §3.3, . With , and :
(c) Dividing by , that is , or twenty-seven grams of mass-energy in every cubic metre of room-temperature space, about two per cent of the mass of the air that would occupy it and present in a vacuum just the same. The true answer, , is smaller by a factor of .
(d) Because the answer scales as and the true total is finite, any cut-off yields agreement only at one particular value, and that value would have to depend on , since matching requires . So a cut-off is not a missing detail of the apparatus but a temperature-dependent parameter with no physical interpretation, chosen after the fact to reproduce a measurement. That is what (4.1.28) replaces: not a sharp cut-off but a smooth suppression whose scale is derived and is proportional to for the reason the fudge would have required.
The problem was made well posed before it was attacked. Kirchhoff's filter argument, run on two cavities joined through a narrow-band filter, shows that a spectrum different in the two would move heat from cold to hot. So the spectrum inside any hot cavity is one universal function of frequency and temperature, independent of material, shape and volume. That is what licenses §2's choice of a cube with conducting walls, and it is why the answer is an energy density with no box in it.
The mode count, built and not quoted. Standing waves in a cube put the allowed wavevectors on a lattice of spacing in one octant. Counting lattice points inside a ball gives , transversality doubles it, and turns it into per unit volume. The neglected surface terms are of relative order , which for a centimetre cavity and visible light is , and periodic boundary conditions give the same answer because two factors of cancel. Equipartition then follows from Chapter 0.6's Boltzmann distribution in three lines, since each quadratic term pulls a out of a Gaussian, so that and . Multiplying gives , whose integral is and does not exist. Carried to violet light at it has already overshot the measured total by , and it gains per decade forever.
Stimulated emission was forced, not proposed. Two rate processes, detailed balance and Boltzmann weights give , which is bounded as . The true is not, by the Rayleigh–Jeans law that had to be derived first. No values of and repair that, and the only term the model's own rules permit is an emission rate proportional to . Requiring the repaired balance to survive then gives in one line, and requiring it to reproduce Rayleigh–Jeans at low frequency gives . So , with no quantum statistics used anywhere, which was the point. The shape is derived and the normalisation is matched against the theory being replaced. Chapter 4.18 §5 does it again from the other end and closes that loop.
Scored against targets fixed in advance. Differentiating gives , whose root makes . Integrating gives and hence . Both were quoted in §3.4 before the derivation ran, and both come out of one formula containing one constant. The same formula reproduces Wien's 1896 law as its high-frequency limit, which was never used as an input.
And the constant is measured again, twice, without a cavity. A photocell gives , a line whose slope is the same for every metal, and five mercury lines on sodium return . Compton's X-rays give at , independent of wavelength and material, against a classical prediction of exactly zero. That calculation is Chapter 2.5's Worked example 1, unchanged, with its one flagged import now discharged. Setting the null four-vectors of Chapter 2.5 §7.1 and §4.3 side by side then gives , whose two components are the two experiments. Finally the atom: a classical electron at the measured radius has the right binding energy to a part in a thousand, radiates by Larmor's formula, and reaches the nucleus in after turns, emitting a rising chirp where five sharp lines fitted by two integers are observed.
Where this gets spent. Four failures with four numbers, and one thing common to all of them: a quantity classical physics lets vary continuously is found restricted, and the scale of the restriction is one constant with the dimensions of action, measured twice by unrelated means and agreeing to six figures. That is what Part IV now has to account for, and it cannot be done by adjusting a coefficient. Section 7.2 showed there is no coefficient to adjust. What has to change is what a state is. Chapter 4.2 makes that change, and the surprise is how little new mathematics it needs: Chapter 0.5 built Hermitian operators, orthonormal eigenbases, the spectral theorem and simultaneous diagonalisation on a complex vector space, and 4.2 is a two-column table with 0.5's theorems on the left and their physical readings on the right, plus a short list of postulates that are announced as postulates and flagged as such. Everything in this chapter that looked like a separate mystery becomes one statement in that language: parcels of energy, a constant with the dimensions of action, integers in a spectrum. Nothing here is repealed: the Planck spectrum is re-derived in 4.18 from the other side, is the relation Part V quantises, is the density of states every later calculation counts with, and returns in Chapter 4.17 as the equality of two matrix elements that Hermiticity makes automatic. Part IV begins where Chapter 0.5 left off, and it begins with the postulates named out loud.