Frequentism
Frequentism identifies probability with relative frequency: the probability of heads is the proportion of heads among tosses. It is the interpretation with the strongest claim to empiricist respectability, since frequencies are observable and probabilities on this view are not mysterious extra facts but summaries of what actually happens. It dominated statistical practice through the twentieth century, and the vocabulary of "sampling", "populations", and "long runs" that pervades applied science is its residue.
It is also the interpretation with the most thoroughly catalogued problems. This page distinguishes its two main forms, sets out the machinery von Mises added to make the second work, and then works through the objections — of which the reference-class problem and the single case are the ones that have never been answered.
Finite frequentism
The simplest version: the probability of an attribute in a finite reference class is the actual relative frequency of that attribute in that class. The probability that a randomly chosen resident of a country is left-handed just is the proportion of left-handers.
Its merits are real. Probabilities so defined are perfectly ascertainable — count — and they satisfy the axioms automatically, since relative frequencies in a finite set are a normalised measure. Nothing metaphysical is posited. For actuarial and demographic work this is very nearly what practitioners mean.
The objections are equally clear.
- The problem of the single case. A coin tossed exactly once and landing heads has, on this view, probability of landing heads. A coin never tossed at all has no probability whatsoever, since the reference class is empty and the ratio undefined. Both verdicts are wrong.
- Probabilities become artifacts of sample size. If a coin is tossed times and lands heads times, its probability of heads is . Any deviation of frequency from chance is definitionally impossible, so the notion of a biased sample — indispensable in statistics — cannot be stated. This is the most damaging point: the theory makes it incoherent to say that a fair coin happened to land heads seven times, which is exactly the kind of thing statistics exists to say.
- Only rational values are possible. In a class of members, every probability is a multiple of . But a physical half-life gives a decay probability that is irrational, and no finite class can realise it.
- Reference class. Every individual belongs to many classes with different frequencies, and finite frequentism supplies no rule for choosing.
Limiting frequentism
The natural repair is to move to infinite sequences. Let the probability be the limit of relative frequency as the sequence is extended:
This removes the sample-size artifact — a run of seven heads is now a fluctuation that washes out — and permits irrational values. Venn proposed it; von Mises gave it its rigorous form.
Von Mises saw that a limit alone is insufficient, because a sequence can have the right limiting frequency while being patently non-random. The sequence HTHTHTHT… has a limiting frequency of but is perfectly predictable, and no gambler would call it a chance process. He therefore required a collective: an infinite sequence satisfying two axioms.
- Convergence — the limiting relative frequency exists.
- Randomness (Principle of the Excluded Gambling System) — the limit is invariant under admissible place selection. Choosing a subsequence by a rule that does not look at the outcome being selected — every third trial, every trial following a head — must yield the same limit. No gambling system can improve one's expected return.
The second axiom is what distinguishes probability from mere statistics, and formalising it caused decades of difficulty. If every subsequence must have the same limit, no sequence qualifies, since one can always select the subsequence of all heads. Restricting to effectively computable place selections is the successful repair, due to Church and Wald: this makes collectives exist, and the resulting theory turns out to be closely related to algorithmic randomness, where Martin-Löf randomness supplies the modern and more satisfactory treatment of what it is for an individual sequence to be random.
The objections
The single case, again
Limiting frequentism does not solve the single-case problem; it makes it structural. A probability is a property of a sequence, and a single trial is not a sequence. "The probability that this nucleus decays in the next hour" is, strictly, ill formed — one can only speak of the frequency in an infinite class of similar nuclei.
Von Mises accepted this bullet explicitly, holding that single-case probability talk is simply meaningless and that the demand for it reflects a confusion. That is a coherent position, but it is in serious tension with physics, where the Born rule is naturally read as giving the chance of this measurement yielding a particular result, and with every practical use of probability in medicine, law, and decision-making.
The reference-class problem
Even granting classes rather than individuals, which class? Every event falls under many descriptions, and frequencies differ across them. The theory offers no principled selection, and the natural fix — take the narrowest class for which we have statistics — collapses, because the narrowest class containing a given individual is the singleton, whose frequency is or .
Reichenbach proposed the "narrowest class for which reliable statistics exist", which is explicitly a compromise with epistemic practicality rather than a metaphysical answer, and makes probability depend on what data we happen to have. That is a strange consequence for a theory whose selling point was objectivity.
Sequences that do not exist
The limit is taken over an infinite sequence, but actual sequences are finite. So the probability must be a claim about a hypothetical infinite extension: what the frequency would converge to if the coin were tossed forever.
This is a substantial retreat. The counterfactual is not obviously well defined — coins wear out, and a coin tossed forever is not this coin — and if the answer is that the limit is whatever the coin's disposition would produce, then the theory has covertly become a propensity theory with an extra step. Hypothetical frequentism also forfeits the empiricist advantage that motivated the position: hypothetical infinite frequencies are no more observable than propensities.
Order dependence
Limits of relative frequency depend on the order of the sequence. Any sequence with limiting frequency can be rearranged to have limiting frequency , or no limit at all, without changing which outcomes occur. If probability is a limit, then it is a property not of the collection of outcomes but of an arbitrary ordering imposed on them — which is hard to square with the claim that it is an objective feature of the world.
The law of large numbers does not help
It is tempting to think the strong law vindicates frequentism by proving that frequencies converge to probabilities. It does not, and the reason is the central methodological point of this section.
The theorem says: given a probability measure under which the trials are i.i.d. with , the set of infinite sequences whose limiting frequency equals has -measure . Every element of that statement presupposes . The theorem cannot define as the limit, because it needs to state what has measure one. It is also weaker than it looks in the relevant respect: the exceptional set is not empty, merely null, and it contains sequences with every other limiting frequency and sequences with no limit at all. So convergence is not guaranteed — it is merely assigned probability one, by the very measure whose meaning was at issue.
What the theorem does establish is a consistency result: if probabilities are understood some other way, then frequencies are excellent evidence about them, and an interpretation on which they were not would be in trouble. That is a genuine constraint on rival interpretations, and it is why frequency data bear on chance under every account. It is not a definition.
What frequentism gets right
The objections are severe enough that few philosophers now hold the view as a definition of probability. But two of its commitments are correct and are retained by its rivals.
- Frequencies are the primary evidence for chances. However chance is understood, statistical data are how we find out about it. Any interpretation must make frequency evidential, and the law of large numbers explains why.
- Chance constrains frequency. A theory on which chances and frequencies could come systematically apart — where a coin of chance regularly landed heads of the time in long runs — would not be a theory of chance at all. The connection is not identity, but it is not accidental either, and spelling it out is exactly what the fit condition does in the best-system account.
The modern descendant of frequentism is not a definitional thesis but a methodological one: error-statistical or classical statistics, which uses sampling distributions to control long-run error rates without claiming that probability means frequency. That is a defensible position and it is not touched by the arguments above.
Assessment
| Criterion | Verdict |
|---|---|
| Admissibility | passes for finite; limiting frequencies need not be countably additive |
| Ascertainability | best of any interpretation for finite classes; poor for hypothetical limits |
| Applicability | strong for mass phenomena, weak elsewhere |
| Single case | fails — by design |
| Explains the calculus | partly — additivity is automatic, countable additivity is not |
| Guidance | weak — needs a further principle to connect a class frequency to this case |
Where this sits
Frequentism is the clearest case of an interpretation that answers the metaphysical question of the introduction by deflating it: there is no probabilistic fact beyond the pattern of actual outcomes. Its failures are all versions of one complaint — that a pattern in a collection cannot be a property of a member — and the two interpretations that follow are the two ways of responding.
Propensity theories keep objectivity and locate the probability in the individual setup, paying for it with a primitive disposition. The best-system account keeps the Humean insistence that chance supervenes on the pattern of actual events, but abandons the identification of chance with frequency in favour of chance as a term in the best summary of that pattern — which is designed to preserve frequentism's metaphysical economy while escaping its technical failures.