Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

A History of the Philosophy of Probability

Probability has a short history and a strange one. The mathematics arrived late — there is no ancient theory of probability, though dice are as old as civilisation — and when it came, it came with two faces at once. From the beginning the calculus was applied both to physical setups (dice, mortality, errors of measurement) and to degrees of belief (the credibility of testimony, the reasonableness of a wager), and the question of which use was fundamental was not raised for two centuries.

This page traces the arguments still in play. The historical narrative is a means to that end rather than the object.


Before 1654

Games of chance are ancient; a mathematics of chance is not. The standard explanations for the delay — absence of good notation, fatalism, the low status of gambling, poorly standardised dice — are each partial, and the puzzle remains genuinely open.

There was, however, a rich pre-mathematical tradition. Medieval jurists and theologians developed sophisticated qualitative accounts of probability as approvability — the credibility of an opinion, its endorsement by authorities. This is closer to evidential probability than to chance, and the modern quantitative notion partly displaced and partly absorbed it.

1654–1750: the calculus emerges

Pascal and Fermat (1654), corresponding about the division of stakes in an interrupted game, produced the first systematic treatment. Their solution introduced expectation as the central notion — the value of a position is the probability-weighted average of outcomes — and expectation, not probability, remained primary for a century.

Huygens (1657) wrote the first published textbook. Pascal's Wager, in the Pensées, is the first decision-theoretic argument, and it made probability a tool of philosophy immediately; it is treated in pragmatic arguments.

Jacob Bernoulli's Ars Conjectandi (posthumous, 1713) is the first great philosophical work on the subject. It contains the first law of large numbers, proving that observed frequencies converge in probability to the underlying ratio — and Bernoulli's motivation was explicitly epistemological. He wanted to justify inferring an unknown ratio from observed frequencies, which is the inverse problem, and he saw clearly that his theorem addressed it only indirectly. That gap — between "frequencies converge to the probability" and "the observed frequency tells us the probability" — is the one frequentism and Bayesianism have been arguing about ever since.

Bernoulli also distinguished aleatory probability (chance in the setup) from epistemic probability (degree of certainty), which is the chance/credence distinction in its first clear statement.

Bayes (posthumous, 1763) solved a version of the inverse problem, giving a rule for the probability of a hypothesis given evidence. His use of a uniform prior over the unknown parameter was defended by an ingenious argument about a ball rolled on a table — an early recognition that the prior needs justification rather than stipulation.

1750–1900: classical confidence and its collapse

Laplace dominates. His Théorie analytique (1812) and Essai philosophique (1814) gave the classical interpretation its canonical form and applied probability across astronomy, jurisprudence, and demography. Laplace was a determinist: his demon knows the complete state and assigns no intermediate probabilities, so probability is entirely a measure of human ignorance. The classical interpretation is therefore epistemic in its founder's hands, which is often forgotten.

The nineteenth century then dismantled the classical framework from two directions.

The frequency turn. Ellis, Cournot, and above all Venn (The Logic of Chance, 1866) argued that probability is relative frequency in a series, and that the Principle of Indifference is worthless — it manufactures numbers from ignorance. Boole pressed the circularity charge against equipossibility.

The statistical turn in physics. Maxwell and Boltzmann brought probability into fundamental science. Boltzmann's statistical account of entropy, and the reversibility and recurrence objections of Loschmidt and Zermelo, made the interpretation of physical probability urgent for the first time.

Peirce deserves mention as an early propensity theorist, holding that a die's probability is a disposition — a "would-be" — and anticipating the twentieth-century position by fifty years.

1900–1950: the interpretations crystallise

The classical consensus having collapsed, the modern positions were staked out within four decades.

  • Keynes, A Treatise on Probability (1921) — the logical interpretation: probability as an objective logical relation between propositions, sometimes merely ordinal.
  • von Mises, Grundlagen (1919) — rigorous frequentism via collectives and the axiom of randomness.
  • Ramsey, "Truth and Probability" (1926) — the subjectivist foundation: degrees of belief measured by betting behaviour, constrained by coherence, with the first Dutch-book argument and the first representation theorem. Written against Keynes, and decisive.
  • de Finetti (1930s) — "probability does not exist"; the representation theorem for exchangeable sequences; the rejection of countable additivity.
  • Kolmogorov, Grundbegriffe (1933) — the measure-theoretic axiomatisation. Its importance for philosophy is precisely its neutrality: it settles the calculus and leaves the interpretation entirely open, which is why the disputes above continue unaffected by it.
  • Popper (1930s–50s) — falsificationism, and later the propensity interpretation, introduced to make sense of single-case quantum probabilities.
  • Reichenbach (1930s–50s) — frequentism, the pragmatic vindication of induction, and the common-cause principle.
  • Cox (1946) — the derivation of the probability axioms from structural desiderata on plausible reasoning.
  • Carnap, Logical Foundations (1950) — the most systematic inductive logic, and the one whose failure was most instructive.

1950–present

The Bayesian revival. Savage's Foundations of Statistics (1954) gave the definitive representation theorem. Jeffrey's The Logic of Decision (1965) generalised conditionalization to uncertain evidence. Bayesian methods spread through statistics, philosophy of science, and eventually machine learning.

Lewis. "A Subjectivist's Guide to Objective Chance" (1980) introduced the Principal Principle and reframed the entire chance/credence question. The best-system account, and Lewis's own discovery of the Big Bad Bug, set the agenda for the following decades.

Negative results. Goodman's grue (1955), Putnam's diagonal argument, and the collapse of Carnap's programme established that no purely formal inductive logic is possible. Bell's theorem (1964) and its experimental confirmation showed that a philosophical principle about probability and causation could be empirically refuted.

Causal modelling. Suppes, then Spirtes–Glymour–Scheines and Pearl (1980s–2000s) developed the graphical framework, giving a rigorous account of the relation between probability and causation after a century of failed reductive analyses.

Accuracy-first epistemology. Joyce (1998) replaced pragmatic Dutch-book arguments with accuracy-dominance arguments, and the programme has since been extended to conditionalization and the Principal Principle.

Self-location. Elga's Sleeping Beauty paper (2000) and Bostrom's Anthropic Bias (2002) opened the self-locating problem, which remains the clearest unresolved gap in the Bayesian framework.

The shape of the history

Three observations worth carrying away.

The two faces were there from the start. Bernoulli distinguished aleatory from epistemic probability in 1713, and the field has been trying to relate them ever since. The chance/credence distinction is not a modern refinement but the subject's founding structure.

Formalisation settled less than expected. Kolmogorov's axiomatisation was decisive mathematically and inert philosophically. Every interpretive dispute alive in 1930 is alive now, because the axioms are neutral among them.

The most durable results are negative. Grue, the failure of Carnap's programme, Bell's theorem, the Big Bad Bug, the pooling impossibility results, Simpson's paradox, and Bertrand's paradox are what the field has established most securely. That is not a poor showing: knowing precisely which inferences fail, and why, is the main protection against the invalid ones catalogued on the method page.