The Reference-Class Problem
A fifty-five-year-old man is diagnosed with a cancer. What is the probability he survives five years? The literature offers a five-year survival rate of 40% for this cancer. But he is also a non-smoker, physically fit, diagnosed early, with a particular genetic profile, treated at a specialist centre. Each of these defines a class with a different survival rate.
Which is his probability?
This is the reference-class problem. It was introduced by Venn and named by Reichenbach, and it is the most practically consequential difficulty in the philosophy of probability — it arises whenever a statistic is applied to an individual, which is to say in medicine, law, insurance, forecasting, and machine learning.
The problem
Every individual belongs to indefinitely many classes. Frequencies differ across them. A frequency is defined only relative to a class, so unless a class is singled out, no probability attaches to the individual at all.
The natural fixes fail in instructive ways.
Take the narrowest class. More specific information is better, so use the smallest class containing the individual. But the narrowest class is the singleton — the class containing only him — whose survival frequency is or depending on what actually happens. Taken seriously, this yields trivial probabilities for everything, which is the single-case problem in its sharpest form.
Take the narrowest class for which we have reliable statistics. Reichenbach's proposal, and what practitioners actually do. It makes the probability depend on what data happen to have been collected, so the patient's probability of survival changes when a new study is published, without anything about him changing. That is a strange consequence for a supposedly objective quantity. The proposal is best understood as a rule for estimating a probability rather than as an account of what the probability is.
Take the class of all and only causally relevant factors. Attractive, and it is roughly right — but it requires prior causal knowledge, so probability can no longer be used to analyse causation without circularity. It also does not terminate: the set of causally relevant factors, fully specified, again approaches the singleton.
Homogeneity. Salmon's proposal: use a class that is objectively homogeneous, admitting no statistically relevant partition. If no property divides the class into subclasses with different frequencies, it is the right class. This is the most principled answer, and its difficulty is that in a deterministic or near-deterministic world the only objectively homogeneous classes are trivial ones — every apparent homogeneity dissolves under finer description.
Whose problem is it?
The problem's severity varies sharply with the interpretation, which is one reason it is a good diagnostic.
Frequentism — fatal. A frequency is by definition relative to a class, so without a selection rule the theory does not assign probabilities to individuals at all. Von Mises accepted this and declared single-case probability meaningless.
Propensity — dissolved in principle, returning as epistemology. The chance belongs to the concrete setup with all its properties, so there is no class to choose. But we must still estimate the propensity from data about other cases, and choosing which other cases are relevantly similar is the reference-class problem in epistemic dress.
Subjectivism — not a problem about probability but about evidence. The agent has a credence in this man's survival; the various statistics are all evidence bearing on it; the question is how to weigh them, which is ordinary Bayesian updating on a conjunction of facts about him. This is arguably subjectivism's single best argument, and it explains why the problem feels intractable to objectivists and merely difficult to Bayesians.
Best-system chance — the system's own generalisations determine which properties are chance-relevant, so the reference class is fixed by the laws rather than by our description. This is a real advantage, though it means the correct reference class is known only when the best system is.
Hájek has argued the problem is genuinely universal — that even subjectivists face it in choosing which evidence to conditionalize on, and that it therefore afflicts every interpretation. The response is that for subjectivists it is not a problem of indeterminacy but of inference: all the evidence is admissible, and the framework says to condition on all of it.
In practice
The problem is not academic, and three domains show what is at stake.
Law. Statistical evidence about a defendant's group is admitted with great caution, and the reference-class problem is part of why. The L'Heureux-Dubé and blue-bus cases raise the question directly: if 80% of buses in town are blue, does that establish that the bus that hit the plaintiff was probably blue? Most jurisdictions say naked statistical evidence is insufficient for liability even at high probability, which suggests the legal standard demands evidence about the individual case rather than about a class containing it — a distinction the likelihood framework partly captures.
Actuarial and algorithmic prediction. Insurance rating and risk scoring are exercises in reference-class selection, and the choice of variables is simultaneously a statistical and a normative decision. Excluding a predictive variable on fairness grounds is a choice of reference class made for non-epistemic reasons — and the resulting probabilities are genuinely different, not merely differently reported.
Medicine. Personalised medicine is the systematic narrowing of reference classes, and it runs into the statistical wall the problem predicts: the narrower the class, the more relevant it is and the less data there is. The trade-off between relevance and sample size is the practical form of the philosophical problem, and it has no clean solution.
What the problem shows
Two conclusions are reasonably secure.
Probabilities attach to descriptions, not to bare individuals. "The probability that he survives" is elliptical; the complete claim is "the probability that a person so described survives". This is the same structure as the model-relativity of every probability claim, arriving in a form where the suppressed clause is a class rather than a measure.
Relevance is a causal notion, not a statistical one. The properties that should determine the reference class are those that make a causal difference, and identifying them requires causal knowledge. A purely statistical criterion cannot distinguish a confounder from a coincidence, which is Simpson's paradox again: whether to partition a population is a causal question.
That second point suggests the problem is not fully solvable within probability theory, and should not be expected to be. With a causal model in hand, the reference class is determined — condition on the causal parents, not on everything correlated — and without one, no statistical rule suffices.
Where this sits
The reference-class problem is the practical face of the single-case problem, and it is the strongest argument against pure frequency accounts of probability. It also explains why the chance/credence distinction does real work: the question "what is his probability of survival?" is ambiguous between a request for an objective chance, which requires a determinate reference class or a propensity, and a request for a reasonable credence, which requires weighing all the available evidence.
The next page turns from the application of probability to individuals to its relation to ordinary categorical belief: the lottery and preface paradoxes.