Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Priors and Indifference

Probabilism fixes the structure of credence and conditionalization fixes its dynamics. Neither fixes its content. Given a prior, the Bayesian machine grinds out posteriors; but the machine has to be started, and nothing in the framework says where. This is the problem of the priors, and it is the point at which every dispute in Bayesian epistemology ultimately arrives.

The problem is not merely that priors are unspecified. It is that the posterior after any finite body of evidence depends on them, sometimes decisively, so that a disagreement about priors is a disagreement about what the evidence shows. If priors are unconstrained, Bayesian confirmation is a bookkeeping device rather than a theory of evidence.


The problem

Two agents observe the same data and reach opposite conclusions, each impeccably coherent, each having conditionalized correctly. One assigned the hypothesis a prior of , the other . Neither has made an error by the standards of subjective Bayesianism.

The standard reply is washing out: as evidence accumulates, agents with different priors converge, so the priors' influence is transient. The Blackwell–Dubins merging theorem makes this precise. Three qualifications matter, and together they take most of the force out of the reply.

  • Convergence requires priors to be mutually absolutely continuous — to agree on which events have probability zero. Where they disagree, no amount of evidence brings them together.
  • It is asymptotic. Nothing is guaranteed at any finite stage, and the rate can be arbitrarily slow. Since all actual inquiry is finite, this is a statement about a limit no one reaches.
  • It concerns agreement, not truth. Two agents can converge on a falsehood if both assign the truth a sufficiently low prior.

So washing out shows that priors do not matter in the limit for agents who already agree about what is possible. It does not show that any finite body of evidence compels a particular credence.

Indifference and its failure

The oldest proposal is the Principle of Indifference: absent any reason to discriminate, assign equal credence. Its failures are set out under classical probability, and they are decisive as a general rule — the wine–water paradox, the cube factory, and Bertrand's paradox all exhibit the same defect, that indifference is not invariant under reparameterisation.

The discrete version fails too, and it is worth restating here because it is the form that arises in practice. Given a coin of unknown bias, indifference over outcomes gives to heads; indifference over hypotheses about the bias gives a different answer depending on how finely the hypotheses are individuated. Ignorance does not come with a canonical partition, and the principle needs one.

Invariance

The serious modern repair replaces "no reason to discriminate" with an explicit symmetry requirement: the prior should be invariant under transformations that leave the problem unchanged.

This is a genuine improvement, because a symmetry claim is checkable and can be wrong, whereas a claim about the absence of reasons is neither. The standard results:

  • Location parameters. If the problem is invariant under , the prior must be uniform (improper on the whole line).
  • Scale parameters. If invariant under — as for a quantity with no natural unit — the prior must be , equivalently uniform in . This is why "uniform between and " is the wrong prior for a scale quantity, and it dissolves the cube-factory paradox: side, area, and volume are related by a scale transformation, and the scale-invariant prior gives consistent answers across all three.
  • Jeffreys priors generalise this using the Fisher information, giving a prior invariant under arbitrary smooth reparameterisation.

Maximum entropy extends the approach to cases with constraints: among distributions consistent with what is known, choose the one maximising entropy. With no constraints this returns the uniform distribution; with a known mean, the exponential; with mean and variance on the line, the Gaussian.

The limits are real and should not be glossed. Invariance determines a prior only when a symmetry group is available, and many problems have none — there is no natural group acting on the space of scientific theories. Maximum entropy in the continuous case is not reparameterisation-invariant unless defined relative to a reference measure, and specifying that measure is the original problem restated. And which invariance to demand can itself be contested, as the competing resolutions of Bertrand's paradox illustrate.

Improper priors

Scale and location invariance yield improper priors: diverges, so these are not probability distributions at all. They are used anyway, on the grounds that the posterior is often proper even when the prior is not.

The practice is defensible but hazardous. Improper priors can yield improper posteriors, in which case the inference is meaningless; they violate probabilism outright, so an agent with an improper prior is not a Bayesian agent in the strict sense; and they can produce marginalisation paradoxes, where two valid routes to the same posterior disagree. They are best regarded as convenient limits of proper priors rather than as representations of ignorance.

Regularity and dogmatism

Two failure modes bracket the space of admissible priors.

Dogmatism: assigning or to a contingent proposition makes it unrevisable, since conditionalization cannot move a credence away from the extremes. Cromwell's rule counsels against it. But strict regularity — no contingent proposition gets — is unattainable in continuous spaces, where almost every proposition must receive zero.

The practical concern is subtler than the formal one. A prior of is technically regular and effectively dogmatic: no feasible evidence will raise it to significance. Formal regularity is thus neither necessary nor sufficient for open-mindedness, and the real desideratum — that priors be responsive to attainable evidence — has no crisp formulation.

Simplicity and the structure of the prior

Priors are not only about single hypotheses; they encode a shape over hypothesis space, and that shape carries substantive commitments.

The most important is simplicity. Any Bayesian account of why simpler theories are preferable must locate that preference in the prior — assigning higher probability to simpler hypotheses — or in the likelihoods, since nothing else is available. The first option makes simplicity a brute prior commitment and raises the question why the world should be expected to be simple. There is also a partial structural answer: a simple theory, having fewer adjustable parameters, spreads its probability over fewer possible data sets, so it assigns higher probability to the data it does predict and is more sharply confirmed when they obtain. This is the Bayesian reconstruction of Occam's razor, and it explains part of the preference without a brute posit.

Two well-known priors formalise the idea. The Solomonoff prior weights each hypothesis by where is its Kolmogorov complexity, giving a universal prior with attractive convergence properties — at the cost of being uncomputable and dependent on the choice of universal machine. Minimum description length is the practical descendant.

Goodman's grue problem is the standing obstacle to all of this: simplicity is language-relative, and a language with gerrymandered primitives makes gerrymandered hypotheses simple. The same difficulty sank the logical interpretation, and it is not solved by moving the burden from a confirmation function to a prior.

Where this leaves the dispute

The positions on priors correspond exactly to the positions on interpretation:

PositionConstraint on priorsCost
Strict subjectivismcoherence onlypermissiveness; confirmation is agent-relative
Objective Bayesianismcoherence + calibration + equivocationequivocation inherits the indifference paradoxes
Logical probabilityuniquely determined by the languagefailed; language dependence, -continuum
Practical Bayesianismconventional priors, checked by sensitivity analysisadmits the problem is unsolved and manages it

The last row is not a philosophical position but it is what working practice looks like, and it embodies a defensible methodological attitude: choose a prior, be explicit about it, and check whether the conclusions survive across the plausible range. Where they do, the underdetermination is harmless; where they do not, the argument was resting on the prior rather than on the evidence. Applied to the fine-tuning and anthropic literatures, this test is considerably more demanding than it might appear.

Where this sits

The problem of the priors is the unsolved core of Bayesian epistemology, and its irreducibility is the strongest argument for permissivism — not that any prior is as good as any other, but that no principle has been found that narrows the field to one.

It is also why the Principal Principle matters so much. If objective chances exist and rationally constrain credence, then at least some priors are fixed by the world rather than chosen by the agent, and the arbitrariness is bounded. That is the subject of the next folder.

The remaining option is to deny that credences must be precise at all, and to represent ignorance by a set of probability functions rather than a single one — the subject of imprecise probability.