Models, Measures, and Applications
Every interpretation of probability takes for granted that some particular probability model is the right one for the situation at hand. The frequentist counts outcomes in a class, the Bayesian assigns credences over a partition of hypotheses, the propensity theorist attributes a tendency to a setup — and each has already assumed a sample space, an event structure, and often a parameterisation. None of these is given by the world. They are chosen, and rival choices are formally impeccable and mutually inconsistent.
This is the application question of the introduction, and it is the most easily overlooked of the three because it looks like a technicality. It is not. It is where the Principle of Indifference fails, where "at random" turns out to be undefined, and where the gap between a mathematical structure and a physical situation has to be closed by something other than mathematics.
A measure is not found, it is built
The Kolmogorov axioms define a probability space as a triple . Applying probability to anything real requires supplying all three, and each involves a decision that is not forced.
- The sample space — what are the possible outcomes? A coin toss has two outcomes if we care only about the face, three if we admit landing on edge, and continuum-many if we individuate by angle of rest. The choice is a decision about how finely to describe the world, and probabilities are relative to it.
- The event algebra — which sets of outcomes count as events? For finite this is usually the full power set and the choice is invisible. For continuous it cannot be, since no translation-invariant countably additive measure is defined on every subset of the reals. Some questions are not merely unanswered but unaskable, and which ones depends on the algebra.
- The measure — which assignment? Even with and fixed, infinitely many measures satisfy the axioms. Selecting one is the whole substance of applying probability, and the axioms are silent on it.
The upshot is that a probability claim about the world is always model-relative, and disputes that appear to be about probabilities are often disputes about models. Recognising this dissolves some puzzles outright and sharpens others into genuine problems.
Bertrand's paradox
The classic demonstration. Take an equilateral triangle inscribed in a circle and draw a chord "at random". What is the probability that the chord is longer than a side of the triangle?
Three natural methods give three different answers, and each is a legitimate construction:
| Method | Randomisation | Answer |
|---|---|---|
| Random endpoints | fix one endpoint, choose the other uniformly on the circumference | |
| Random radius | choose a radius, then a point uniformly along it, and take the perpendicular chord | |
| Random midpoint | choose the chord's midpoint uniformly over the disc |
Nothing is wrong with any calculation. What is wrong is the question: "at random" does not specify a measure, and the three methods impose different ones. The uniform distribution on endpoints is not the uniform distribution on midpoints, because the map between the two parameterisations is not measure-preserving. Once a physical procedure for producing chords is specified — spinning a pointer, rolling a straw onto a circle, sighting through a slit — the answer is determined, and different procedures genuinely give different answers.
Two lessons, and they point in opposite directions.
- Against the Principle of Indifference. "Assign equal probability to possibilities among which you have no reason to discriminate" is not a well-defined instruction, because possibilities can be redescribed. Indifference over one parameterisation is not indifference over another, and the principle by itself does not say which to use. This is why the classical and logical interpretations cannot be self-standing.
- In defence of the axioms. Bertrand's paradox is not an inconsistency in probability theory. It is a demonstration that an under-described problem has no answer. The mathematics behaves impeccably: three different measures, three different values. The failure is in the passage from an English sentence to a model.
Jaynes proposed a partial rescue: require the solution to be invariant under the transformations the problem statement does not fix — for a chord problem, scaling, rotation, and translation of the circle — and the radius method is uniquely selected. The maximum-entropy and transformation-group approaches generalise this. Whether it succeeds is contested; the invariance requirement is powerful where a genuine symmetry group is available and silent where it is not, and critics observe that the choice of which invariances to demand reintroduces the original arbitrariness one level up.
The same problem, less obviously
Bertrand is a toy. The same structure appears wherever a probability is asserted, and the following are all instances of it.
- Continuous parameters in science. A prior "uniform over the possible values" of a physical constant is not invariant under reparameterisation: uniform in a length is not uniform in its logarithm or its reciprocal. Every choice of an uninformative prior is a choice of parameterisation, which is why invariance arguments — Jeffreys priors, maximum entropy — occupy so much space in objective Bayesian work.
- The Doomsday argument and self-location. Reasoning about one's position in a sequence of observers requires a measure over observers, and the argument's force depends entirely on which reference class and which sampling assumption are adopted. This is Bertrand with humans in place of chords.
- Statistical mechanics. The standard measure on phase space is Liouville measure, and it is doing genuine work: the Past Hypothesis is a claim about a low-entropy macrostate, and "most" microstates within it behave thermodynamically only relative to that measure. Choosing it is justified by its dynamical invariance rather than by indifference, which is exactly the Jaynesian move made rigorous.
- Fine-tuning. Claims that life-permitting constants occupy a "tiny fraction" of parameter space presuppose a measure over that space, often over an unbounded range where no normalisable uniform measure exists. The teleological argument inherits this problem in full, and it is one of the standard replies to it.
Conditioning is model-relative too
The application question has a second face, less often noticed, that concerns not the initial model but conditioning within it. In elementary settings is well defined whenever , and no interpretive issue arises. When the ratio is undefined, and yet such conditions are asserted constantly — "given that the point lies on this great circle", "given that the parameter equals exactly this value".
The rigorous treatment defines conditional expectation with respect to a -algebra rather than an event, as a Radon–Nikodym derivative, and the resulting object is unique only up to null sets. The philosophical consequence is the Borel–Kolmogorov paradox: conditioning on the same null event, approached through different families of conditions, yields different distributions. Conditioning on a great circle of a sphere gives a uniform distribution if the circle is regarded as a line of longitude and a non-uniform one if it is regarded as a limit of latitudes.
So "the probability given that " is not well defined by the event alone. It requires a specification of how the conditioning event is embedded in a family — which is to say, a model of the measurement or limiting procedure. Once again the mathematics is exactly right and the English is underdetermined.
What closes the gap?
If the axioms do not select a model, what does? Four kinds of consideration are actually used, and it is worth being explicit that all four are extra-mathematical.
- Physical symmetry. If the setup is invariant under a group of transformations and the quantity of interest is too, the measure should be invariant as well. This is the most respectable ground and the one Jaynes exploited. It applies only when there is a genuine physical symmetry, not merely a descriptive one.
- The generating procedure. Specifying how outcomes are actually produced — the mechanism, the sampling protocol, the experimental design — usually determines the model uniquely. This is why Bertrand's paradox evaporates for any concrete chord-drawing device, and why stopping rules matter in statistics.
- Empirical adequacy. A model can be tested. If the chosen measure implies frequencies that do not obtain, it is the wrong model, and this is the main corrective in practice.
- Robustness. Where no unique model is defensible, one can ask whether the conclusion survives across the plausible range. If it does, the underdetermination is harmless; if the conclusion flips, the argument was resting on the modelling choice rather than on the evidence. This is a discipline the anthropic and fine-tuning literature would benefit from more than it practises it.
Idealisation and the limits of the formalism
A final constraint on any application. Probability models are idealisations, and the idealisation sometimes carries the argument.
Infinite sequences of trials do not occur, but limiting frequentism requires them. Perfectly repeatable setups do not exist, but propensity talk assumes them. Agents with logically closed, real-valued credences over a complete algebra of propositions do not exist either, and this idealisation is the source of the problem of old evidence and of every difficulty about logical omniscience in Bayesian epistemology. In each case the formal object is well behaved and the physical or psychological system it represents is not.
None of this is an objection to using models — the same could be said of frictionless planes — but it does bear on what conclusions the models can support. An argument that depends essentially on a feature present only in the idealisation, and absent in every real instance, has not established what it appears to. Whether the infinite limits in frequentism and the ideal agents in Bayesianism are harmless idealisations or load-bearing fictions is a question each of the following pages has to answer for itself.
Where this leaves us
The formal calculus is neutral not only between interpretations but between models, and the second neutrality has to be resolved before the first even arises. Assertions of the form "the probability is " are elliptical for "on this model, the probability is ", and a great many disputes — Bertrand, indifference, fine-tuning, self-location, conditioning on null events — turn out to be disputes about the suppressed clause.
This completes the spine. The three questions are in place, the chance/credence distinction organises the field, the interpretations are mapped, and the model-relativity of every probability claim is on the table. The interpretations folder takes the positions in turn, beginning with the classical account and the Principle of Indifference whose failure this page has already prepared.