Part VI · Heterogeneity, evolution, and metastatic biology · Chapter 26
A framework for heterogeneity
Four different observations share one word, and most of what they measure is not yet actionable.
1 · Taxonomy, interpatient, intratumoral, temporal, and intermetastatic
Heterogeneity names four different observations. They are distinguished by what was compared with what.
Interpatient heterogeneity compares one patient's tumour with another's. It is the observation that produced subtype classification, and it needs one sample per patient.
Intratumoral heterogeneity compares regions of a single lesion. It needs several samples from that lesion, taken at one time.
Temporal heterogeneity compares one lesion with itself at a later date. It needs samples separated in time. It cannot be observed at all from a single specimen.
Intermetastatic heterogeneity compares separate deposits in one patient. It needs samples from more than one site, taken close enough together that elapsed time is not the explanation.
The sampling requirement is the useful part of this taxonomy. A study that took one biopsy per patient can report only interpatient heterogeneity, whatever its title says. Multiregion sequencing of primary tumours reports the second1. Paired primary and metastatic sequencing reports a mixture of the third and the fourth2. Sampling several deposits at autopsy reports the fourth with the third largely controlled3.
Two of these are confounded in almost every clinical dataset. A metastasis biopsied at progression differs from the primary for three reasons at once. Different cells were sampled, time has passed, and treatment intervened. A paired design cannot apportion the difference between them. Receptor discordance and conversion is about how to reason when it cannot.
2 · Genetic, phenotypic, and microenvironmental heterogeneity as separate axes
Heterogeneity is measured on at least three axes, and they are not one measurement rescaled.
Genetic heterogeneity is variation in DNA between cells. Phenotypic heterogeneity is variation in what cells are doing, measured as protein, transcript or state. Microenvironmental heterogeneity is variation in what surrounds the cells, including immune infiltrate, stroma and vasculature.
The axes would be uninteresting if they tracked one another. They do not.
The clearest demonstration is an integrated spatial study of 280 microdissected tumour regions from 33 breast tumours4. Proteomic heterogeneity was scored across regions, and in a subset of tumours the same regions were exome sequenced. Proteomic heterogeneity did not correspond to genomic heterogeneity within tumours, with correlation coefficients below 0.18 and no significant association. It did correspond to differences in the microenvironment.
The same study produced a dissociation worth stating plainly. Receptor-based heterogeneity, scored from ER, PR and HER2 expression patterns across regions, was lower in grade 3 tumours than in grade 2 tumours. Proteomic heterogeneity moved the other way and rose in grade 3. Receptor heterogeneity correlated positively with disease-free survival. Proteomic heterogeneity correlated negatively with it.
The cohort is 33 tumours and the survival analysis is correlational, so the outcome direction should be held loosely. The dissociation itself is the durable finding. Two heterogeneity measures made on the same tissue moved in opposite directions with grade.
That result forbids a common shortcut. A tumour described as heterogeneous is not one fact with several available assays. It is several different facts, and choosing an assay commits you to one of them.
3 · Models of evolution, linear, branched, neutral, and punctuated
Four models are in circulation. Each makes a different claim about where diversity comes from.
Linear evolution describes sequential replacement. Each new clone is fitter and sweeps the previous one away, so the tumour at any moment is approximately one population. Branched evolution describes divergence from a common ancestor with the branches persisting, so the tumour is several populations at once.
Neutral evolution makes a stronger claim. It holds that most observed diversity is not under selection and accumulates as passenger mutations during growth. Fitting bulk mutant allele frequencies to the power law this predicts, one analysis found the neutral model fitted 323 of 904 tumours across 14 cancer types5. Where it fits, subclonal structure carries no information about which clone will matter later.
Punctuated evolution holds that most genomic change happens in short early bursts rather than gradually. Single-nucleus sequencing of 1,000 cells from tumours in 12 patients with triple-negative disease identified one to three major clonal subpopulations per tumour sharing a common lineage6. The copy number patterns were difficult to explain by gradual accumulation, and the authors concluded that most aberrations were acquired early and followed by stable clonal expansion.
These are not competing universal laws. Multiregion sequencing of 303 samples from 50 primary breast cancers found that the extent of subclonal diversification varied between cases, and found no strict temporal order for the common driver genes1. Events in PIK3CA, TP53, PTEN, BRCA2 and MYC occurred early in some tumours and late in others.
The model inferred also depends on the assay. A bulk allele-frequency spectrum can support neutrality. Single-cell copy number in the same disease supports punctuation. Those two measurements are made on different quantities at different resolutions, so a disagreement between them is not necessarily a disagreement about the tumour.
4 · Fitness landscapes, bottlenecks, and selection under therapy
Fitness is not a property a clone carries. It is a relationship between a clone and an environment, and it changes when the environment does. A subclone with no advantage under one regimen can be the only surviving population under the next. Nothing about that clone changed.
Treatment is therefore better read as a change of landscape than as a reduction in tumour burden. What the change selects for is predictable from what the drug removes. That argument is developed agent class by agent class in Temporal heterogeneity and clonal evolution.
Bottlenecks are the other half of the structure. A bottleneck is any event that reduces the population to a small number of survivors. The transition from in situ to invasive growth is one. Metastatic seeding is another. Each line of therapy is another.
The important observation is that a bottleneck does not end evolution. Sequencing of 299 samples from 170 patients with locally relapsed or metastatic breast cancer showed that the clones seeding metastasis disseminate late from the primary and then continue to acquire mutations2. Most distant metastases carried driver mutations absent from the primary, drawn from a wider set of cancer genes than the early drivers. The population narrows and then diversifies again.
One finding should temper the intuition that more diversity is always worse. In a pan-cancer analysis of 1,165 exomes across 12 cancer types, 86% of tumours contained at least two detectable clones7. Mortality risk rose where more than two clones coexisted in the sample, with a hazard ratio of 1.49 against samples with fewer. Risk then fell again where more than four coexisted. The relationship was not monotonic.
Two readings of that non-monotonicity are available and neither can be dismissed. Genomic instability may generate both the variation selection acts on and a cost the cells must carry. Alternatively, clone counts estimated from a single bulk exome may simply be noisy at the upper end. The measurement question is taken up next.
5 · The measurement problem, every assay reports a different kind of heterogeneity
This section is the discipline the rest of this part depends on.
Any measurement of heterogeneity fixes four things. It fixes a unit of observation, which may be a cell, a microscope field, a region, a lesion or a patient. It fixes a quantity, such as copy number, mutation, protein or transcriptional state. It fixes a sampling scheme. It fixes a threshold that converts a continuous measure into a category.
Change any one of the four and the number changes. Four worked examples make the point better than a general statement does.
The unit of observation. The clone counts above were estimated from one bulk exome per tumour7. A clone below the detection frequency is not counted, and a clone confined to a region that was not sampled is not counted either. The resulting number is a property of the specimen as much as of the disease.
The scale. Quantitative immunofluorescence for CD3, CD8 and CD20 across 93 samples from 31 resected primary breast carcinomas partitioned the variance explicitly8. Between 66% and 69% of the variance sat between fields of view within a single section. Between 30% and 33% sat between biopsies taken from different regions of the same tumour. Differences between sections were negligible. For these markers most of the variability is at a scale finer than any clinical sampling scheme addresses.
The threshold. The ER boundary is set at 1% of tumour nuclei staining, with a separate reporting category for 1% to 10% because benefit at that level is poorly characterised9. In a multi-institutional study of HER2 immunohistochemistry, overall agreement between pathologists on cases scored 0 was 25%10. A heterogeneity statistic built on a category boundary inherits that boundary's reproducibility.
The assay. Proteomic and genomic heterogeneity measured on the same microdissected regions did not correlate4.
Two published heterogeneity frequencies are comparable only if the unit of observation, the quantity, the sampling scheme and the threshold all match. They usually do not. A range quoted across studies is often a range of methods rather than a range of tumours. The same problem recurs for HER2 in Reported frequency and its dependence on detection method and sampling and for receptor conversion in Receptor discordance and conversion, arriving each time through a different door.
6 · When heterogeneity is a curiosity and when it changes a decision
Most measured heterogeneity is not actionable. This chapter should say so before the rest of this part describes how much of it can now be measured.
A heterogeneity measurement changes a decision only when three conditions hold together. The measurement has to be reproducible enough to act on. A different action has to be available. And there has to be evidence that the different action helps in the group the measurement defines.
Almost everything in the current literature fails the third condition. A good deal of it also fails the first.
Three things pass, and the list is deliberately short.
A receptor result that differs between primary and metastasis passes, because it opens or closes access to a treatment class. In a prospective study of 121 women biopsied at suspected metastatic recurrence, the treating oncologist changed the plan in 14% of cases, with a 95% confidence interval of 8.4% to 21.5%11. Receptor discordance and conversion sets out when that change is justified.
An emergent resistance alteration detected before radiographic progression passes, on the strength of one trial. In SERENA-6, switching the endocrine partner at the emergence of an ESR1 mutation in circulating tumour DNA gave a median progression-free survival of 16.8 months against 9.2 months for continuing12. That is temporal heterogeneity acted on prospectively, and it is developed in Acting on evolution, intermittent dosing, adaptive therapy, pre-emptive switching.
HER2 heterogeneity in the neoadjuvant setting passes a weaker version of the test. It predicts pathologic complete response, which was 0% in heterogeneous against 55% in homogeneous tumours in one prospective study13. It calibrates expectation before treatment starts. It does not select therapy, and What to do with a heterogeneous HER2 report in clinic today is explicit that treatment should not be reduced on the basis of it.
Everything else on the current list is prognostic at best. Clone counts, proteomic heterogeneity scores, imaging-derived heterogeneity indices and spatial immune metrics are each associated with outcome in at least one series. None has been shown to change what should be done. Association with outcome and utility for a decision are different claims, and the framework that separates them is set out in Analytical validity, clinical validity, and clinical utility defined.
Ask of any heterogeneity result what decision it would change. If the answer is none, it is a research measurement, and recording it as one is honest rather than dismissive.
Do not escalate treatment because a tumour is described as heterogeneous. No prospective trial has tested escalation selected on a heterogeneity measurement. The one prospective dataset in HER2-positive disease showed the heterogeneous group responding worse rather than differently13.
Do act on a receptor result that has changed, subject to the reasoning in Rebiopsy, when, which lesion, and when the result should change management. That is the one form of heterogeneity with an established and available action attached.
Treat a single-sample heterogeneity statistic as a property of the specimen. Two cores from the same tumour can return different values without the tumour having changed at all.
The gap between what can be measured and what can be acted on is the honest state of this field. Spatial heterogeneity, HER2 heterogeneity, Receptor discordance and conversion and Temporal heterogeneity and clonal evolution describe the measurements in detail. This section is the test to hold each of them against.
References
- Yates LR, Gerstung M, Knappskog S, et al. Subclonal diversification of primary breast cancer revealed by multiregion sequencing. Nat Med 2015 21:751-759. PMID 26099045
- Yates LR, Knappskog S, Wedge D, et al. Genomic evolution of breast cancer metastasis and relapse. Cancer Cell 2017 32:169-184. PMID 28810143
- Avigdor BE, Cimino-Mathews A, DeMarzo AM, et al. Mutational profiles of breast cancer metastases from a rapid autopsy series reveal multiple evolutionary trajectories. JCI Insight 2017 2:e96896. PMID 29263308
- Mardamshina M, Karagach S, Mohan V, et al. Integrated spatial proteomic analysis of breast cancer heterogeneity unravels cancer cell phenotypic plasticity. Nat Commun 2025 16:10482. PMID 41290667
- Williams MJ, Werner B, Barnes CP, Graham TA, Sottoriva A. Identification of neutral tumor evolution across cancer types. Nat Genet 2016 48:238-244. PMID 26780609
- Gao R, Davis A, McDonald TO, et al. Punctuated copy number evolution and clonal stasis in triple-negative breast cancer. Nat Genet 2016 48:1119-1130. PMID 27526321
- Andor N, Graham TA, Jansen M, et al. Pan-cancer analysis of the extent and consequences of intratumor heterogeneity. Nat Med 2016 22:105-113. PMID 26618723
- Mani NL, Schalper KA, Hatzis C, et al. Quantitative assessment of the spatial heterogeneity of tumor-infiltrating lymphocytes in breast cancer. Breast Cancer Res 2016 18:78. PMID 27473061
- Allison KH, Hammond MEH, Dowsett M, et al. Estrogen and progesterone receptor testing in breast cancer: ASCO/CAP guideline update. J Clin Oncol 2020 38:1346-1366. PMID 31928404
- Robbins CJ, Fernandez AI, Han G, et al. Multi-institutional assessment of pathologist scoring HER2 immunohistochemistry. Mod Pathol 2023 36:100032. PMID 36788069
- Amir E, Miller N, Geddie W, et al. Prospective study evaluating the impact of tissue confirmation of metastatic disease in patients with breast cancer. J Clin Oncol 2012 30:587-592. PMID 22124102
- Turner NC, Mayer EL, Park YH, et al. Switching to camizestrant at ESR1 mutation emergence before disease progression during first-line treatment of hormone receptor-positive advanced breast cancer (SERENA-6). Lancet Oncol 2026. PMID 42442380
- Metzger Filho O, Viale G, Stein S, et al. Impact of HER2 heterogeneity on treatment response of early-stage HER2-positive breast cancer: phase II neoadjuvant clinical trial of T-DM1 combined with pertuzumab. Cancer Discov 2021 11:2474-2487. PMID 33941592