Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Bayesian Confirmation

Bayesian confirmation theory analyses evidential support in probabilistic terms. Evidence confirms hypothesis when it raises its probability:

and disconfirms it when the inequality reverses. The account is simple, general, applies to any proposition whatever, and explains a great deal of scientific practice. It is the dominant theory of confirmation, and this page sets out its machinery, what it explains, and where it strains.

The first thing to fix is a distinction that resolves a surprising amount of confusion. Incremental confirmation is the raising of probability; absolute confirmation is a hypothesis's having high probability. Evidence can incrementally confirm a hypothesis that remains very improbable — a positive test result for a rare disease raises the probability from to , which is confirmation in the incremental sense and leaves the disease unlikely. Most paradoxes of confirmation trade on sliding between the two.


The machinery

Bayes' theorem is a trivial consequence of the axioms:

The three inputs each carry a name and a philosophical weight. is the prior; is the likelihood — how expected the evidence is if the hypothesis is true; is the expectedness, computed as .

The ratio form makes the dynamics vivid:

Confirmation occurs exactly when — when the hypothesis makes the evidence more likely than it was unconditionally. Three consequences follow immediately, and they are the theory's main explanatory successes:

  • Surprising evidence confirms more. The smaller , the larger the ratio. Evidence that would be astonishing otherwise, but is expected given , confirms strongly. This is why the 1919 eclipse observations mattered so much for general relativity.
  • Entailed evidence confirms. If entails then , so any with confirms . Successful prediction is confirmation.
  • Refutation is absolute. If entails and is false, . Falsification is the limiting case of disconfirmation rather than a separate logic — which is a point in the theory's favour against Popper.

Measures of confirmation

That confirms is one thing; how much is another, and here the theory is less unified than it appears. Several measures are in use, all agreeing on the direction of confirmation and disagreeing on its magnitude and on comparisons:

MeasureDefinition
Difference
Ratio
Likelihood ratio
Normalised difference

These are not equivalent, and they give different verdicts about which of two pieces of evidence confirms more. The choice is not merely technical: arguments in the philosophy of science — about whether diverse evidence confirms better, about the ravens, about old evidence — can come out differently depending on the measure, and papers sometimes reach opposite conclusions because they silently adopt different ones. The likelihood ratio has the best claim to measure evidential strength as opposed to resulting belief, and it is the subject of the next page.

What the theory explains

The ravens paradox. Hempel observed that "all ravens are black" is logically equivalent to "all non-black things are non-ravens", so by the equivalence condition a white shoe confirms the raven hypothesis. This seems absurd.

The Bayesian treatment is one of the theory's best results: the shoe does confirm the hypothesis, but only by an utterly negligible amount. Since non-black things vastly outnumber ravens, observing a non-black non-raven eliminates a tiny fraction of the ways the hypothesis could fail, while observing a black raven eliminates a much larger fraction. The paradox dissolves into a quantitative point — the intuition that shoes are irrelevant is an approximation of "almost irrelevant" — and the theory explains why the amount is negligible rather than merely asserting it. Note that the result depends on background assumptions about relative class sizes, so it is not purely formal.

Diverse evidence. Varied evidence confirms better than repetitive evidence, because after several similar observations the next one is already expected — is high — so the confirmation ratio is near . Diverse evidence retains low expectedness.

Severe tests. A test is severe when the evidence would probably not have obtained if the hypothesis were false, that is when is small. This makes the likelihood ratio large, and it reconstructs in probabilistic terms a notion the error-statistical tradition takes as basic.

The problems

Irrelevant conjunction

If confirms , then also confirms for any irrelevant — the conjunction of relativity with the claim that the moon is made of cheese is confirmed by the eclipse data. Since entails just as does, the likelihood is unchanged.

The standard reply is that the conjunction is confirmed less, on most measures, and that the tacked-on conjunct receives no confirmation itself. This blunts the objection without wholly removing it, and the details depend on which measure of confirmation is adopted — an instance of the general point above.

Duhem–Quine

Hypotheses do not entail predictions alone; they need auxiliary assumptions about instruments, background conditions, and approximations. When a prediction fails, logic tells us only that the conjunction is false, not which conjunct to blame.

Bayesianism handles this better than falsificationism, and this is a genuine strength. Disconfirmation distributes over the conjuncts in proportion to their prior probabilities and the likelihoods, so a well-established auxiliary absorbs little blame while a shaky one absorbs much. This explains why anomalies sometimes overthrow theories and sometimes indict the apparatus, and it does so quantitatively rather than by appeal to scientific judgement.

The cost is familiar: the distribution of blame depends on the priors over auxiliaries, which are not themselves determined. The framework says what follows from a distribution of confidence without saying which distribution is right.

The problem of the priors

The standing difficulty, treated in full under priors and indifference. depends on , and if priors are unconstrained then so is confirmation. Two scientists can agree on all the evidence and all the likelihoods and disagree about the posterior. Washing-out results help only asymptotically and only for agents who already agree about what is possible.

Likelihoods for catch-alls

Computing requires — the probability of the evidence given that the hypothesis is false. But is not a definite hypothesis; it is the disjunction of every alternative, including those nobody has conceived. Assigning it a likelihood requires assessing theories that do not exist.

In practice one compares specific rivals rather than a hypothesis against its negation, which is the comparative approach the likelihood framework makes explicit. But this means confirmation is always relative to a considered set of alternatives, and the introduction of a new theory can retrospectively change how well-confirmed an old one was — connecting to the new-hypothesis problem.

Old evidence

If is already known, and : known evidence confirms nothing. Yet the perihelion of Mercury, known since 1859, was among the strongest evidence for general relativity in 1915. This is Glymour's problem, and it gets its own page.

Where this sits

Bayesian confirmation is the most successful general account of evidential support available, and its successes are real: the ravens, diverse evidence, severity, and Duhem–Quine all receive quantitative treatments that no rival framework matches. It also unifies confirmation and refutation in a single formalism.

Its limitations are all versions of one point. The theory tells us how evidence shifts credence and is silent on where credence starts and on which hypotheses are in play. That leaves the problem of induction untouched, since a coherent grue-agent updates just as correctly, and it makes confirmation relative to a prior and to a hypothesis space.

The next page isolates the component of the framework that is least prior-dependent, and that some hold to be the whole of evidence: the likelihood.