Likelihood and Evidential Support
In Bayes' theorem, , the prior is the contested ingredient and the likelihood is not. Two scientists who disagree profoundly about how plausible a hypothesis is can often agree exactly on how probable the data would be if it were true, because the likelihood is typically fixed by the hypothesis itself — a statistical model specifies it.
This suggests isolating the likelihood as the carrier of evidential import, distinguishing what the data say from what one should believe after taking priors into account. The resulting position — likelihoodism — is a distinct third option alongside Bayesianism and frequentism, and the notions it introduces are used well beyond it.
The law of likelihood
Hacking's formulation:
Evidence favours over if and only if , and the degree of favouring is the ratio .
Three features distinguish this from Bayesian confirmation.
Evidence is comparative. There is no such thing as evidence supporting a hypothesis simpliciter — only support for one hypothesis against a specified rival. This avoids the catch-all problem: one never needs , only the likelihoods of definite alternatives.
No priors are required. The likelihood ratio is computable from the hypotheses and the data alone.
It measures evidence, not belief. A likelihood ratio of favouring a hypothesis with a prior of leaves it improbable. Likelihoodists take this as a feature: the data strongly favour the hypothesis, and it remains implausible, and these are different claims that should not be collapsed. Royall's illustration is the standard one — a positive result on a highly reliable test for a very rare disease is strong evidence for the disease while leaving it unlikely that the patient has it. Conflating the two is the base-rate fallacy in one direction and the prosecutor's fallacy in the other.
Royall's conventional benchmarks — ratios of and as "moderately strong" and "strong" — are calibrated by the observation that a ratio of can arise by chance with probability at most .
Bayes factors
The Bayesian counterpart. For composite hypotheses with parameters, the Bayes factor is the ratio of marginal likelihoods, each averaged over the parameter's prior:
It is exactly the factor by which the prior odds are multiplied to give the posterior odds, so it isolates the contribution of the data to the change in belief.
The Bayes factor has an important structural property, and it is the most satisfying formal account of Occam's razor available. A hypothesis with many free parameters spreads its likelihood thinly across the many data sets it could accommodate; a simpler hypothesis concentrates it. When the data fall where the simple hypothesis predicts, it wins the comparison decisively. Simplicity is thereby rewarded automatically, with no separate simplicity prior — an "Ockham factor" emerging from the arithmetic. This is a genuine explanation of why flexible theories are penalised, and it is the strongest Bayesian response to the worry that simplicity preferences are brute.
The catch is that the averaging requires priors over parameters, so Bayes factors are not prior-free. They can be highly sensitive to the parameter prior, and are undefined for improper priors — one of the sharpest practical arguments against improper priors, since the resulting factor depends on an arbitrary normalising constant.
The likelihood principle
The strong claim: all the evidential meaning of the data is contained in the likelihood function. Two experiments yielding proportional likelihood functions provide the same evidence about the hypotheses, whatever their designs.
This follows, by the Birnbaum argument, from two premises most statisticians find compelling in isolation — sufficiency (only sufficient statistics matter) and conditionality (if a coin flip decides which of two experiments to run, only the one actually run is relevant). That two innocuous principles entail a controversial one is what makes the result interesting.
Its most striking consequence concerns stopping rules. Consider two experiments: one fixes 12 coin tosses in advance and observes 9 heads; the other tosses until 3 tails appear, which happens on the 12th toss. The data are identical, and the likelihood functions are proportional — the binomial and negative binomial differ only by a combinatorial constant not involving the parameter. So on the likelihood principle the evidence about the coin's bias is identical.
Classical significance testing disagrees, and gives different -values, because the sampling distributions differ: what would have happened under other possible outcomes depends on the stopping rule. So the experimenter's intentions affect the classical verdict while leaving the likelihood untouched.
Which side is right is a genuine philosophical dispute, not a technicality. Likelihoodists and Bayesians find it absurd that intentions should matter; error statisticians reply that the stopping rule affects the procedure's error properties, and that a researcher who samples until reaching significance will reach it eventually — the practice now called -hacking. That is a real problem the likelihood principle appears to be indifferent to.
The Bayesian response is that optional stopping does not mislead a Bayesian in the same way: a sceptical prior is not overturned by data mined in this fashion, since the likelihood ratio must actually be large. But the intuition that design matters is not fully accommodated, and the disagreement remains live.
The error-statistical alternative
Mayo's severity account is the principal rival, and it takes seriously exactly what likelihoodism sets aside.
Data provide good evidence for to the extent that has passed a severe test — one that would very probably have produced a worse fit with if were false.
Severity is a property of the testing procedure, not merely of the data, so it depends on the sampling distribution, the stopping rule, and the space of outcomes that might have occurred. This is what makes it able to condemn -hacking directly: a result obtained by optional stopping has not passed a severe test, because the procedure would very probably have produced such a result even if the hypothesis were false.
The frameworks disagree about what evidence is. For the likelihoodist it is a relation between data and hypotheses; for the error statistician it is a property of an inferential procedure and its error probabilities. The disagreement over stopping rules is the visible symptom of that deeper difference, and neither side has a knock-down argument.
Limitations
The catch-all again. Likelihoodism avoids needing only by always comparing specified rivals. But the true hypothesis may not be among those considered, and the framework then reports which of several false hypotheses the data favour — a fact of limited interest. Comparative support is silent on absolute adequacy.
Composite hypotheses. "The coin is biased" does not fix a likelihood; only a specific bias does. Handling composite hypotheses requires either maximising over the parameter (giving complex hypotheses an unfair advantage) or averaging with a prior (abandoning prior-freedom).
It gives no guidance for action. Since likelihoodism deliberately declines to say what to believe, it must be supplemented for decision-making — and the natural supplement is a prior, which returns us to Bayesianism.
Misleading evidence is possible. Data can favour a false hypothesis. Likelihoodists accept this and quantify it: the probability of obtaining a likelihood ratio of or more in favour of a false hypothesis is at most . This universal bound is a genuinely attractive result, and it is what licenses the conventional benchmarks above.
Where this sits
The likelihood is the least contested part of the Bayesian apparatus, and separating it from the prior clarifies what is really in dispute in most arguments about evidence. When two parties disagree about whether some finding supports a hypothesis, they are usually disagreeing about priors, about which alternatives are on the table, or about whether the procedure was severe — rarely about the likelihood itself.
The distinction between evidence and belief also does useful work elsewhere in these notes. The Bayesian treatment of testimony for miracles turns precisely on it: strong testimony can constitute powerful evidence for an event that remains improbable given a low enough prior, which is Hume's point restated in likelihood terms and is why the debate is genuinely about priors rather than about testimony.
The next page takes up the case where the framework appears to break down entirely — evidence that is already known, and hypotheses that did not exist when the evidence was collected: old evidence and new theories.