Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

The Philosophy of Bell's Theorem

Bell's theorem is the closest thing physics has to an experimental metaphysics: a proof that no theory of a certain very natural kind — one in which the outcomes of spacelike-separated measurements are fixed by locally-carried information — can reproduce the statistical predictions of quantum mechanics, together with laboratory tests that decide the matter in QM's favour. Its philosophical importance is that the kind of theory it excludes is exactly the kind common sense expects the world to be: one in which distant events are correlated only because of what they share in their common past, and in which measurement merely reveals pre-existing values. This remark reads the Entanglement, EPR, and Bell's Theorem page philosophically. It does so in two movements. First, it works through the derivation of the CHSH inequality and its quantum violation in enough detail that every assumption is visible on the page — because the philosophy is entirely a dispute about which assumption to blame. Second, it dismantles those assumptions one by one, catalogues the experimental loopholes and their closure, and lays out the menu of metaphysical options the theorem leaves open: give up locality, give up definite values, give up free choice, or give up single outcomes.

The guiding tension is set by the two historical bookends. Einstein, Podolsky, and Rosen argued in 1935 that quantum mechanics must be incomplete — that its entangled correlations could only be explained by "elements of reality" the wavefunction fails to mention, on pain of a spooky action at a distance no respectable physics should tolerate. Their argument was a reductio aimed at forcing hidden variables. Bell's 1964 reply turned the reductio into a test: he showed that the very hidden-variable theories EPR demanded make predictions that differ numerically from quantum mechanics, so the question "is the world locally explicable?" is not a matter of taste but of experiment — and experiment has answered no.


Part I — The derivation, made explicit

1. The experimental scenario

Fix the setup that all the mathematics refers to. A source emits pairs of systems; one member goes to Alice, the other to Bob, and the two wings are far enough apart that a measurement in one can be completed before any light-speed signal from the other could arrive (they are spacelike separated). Each party has two available measurement settings (two fixed apparatus configurations, e.g. two directions), and on each run freely selects one of them and records one of two outcomes, :

  • Alice's two settings are labelled and . On a given run she selects one of them; write for her choice and for the outcome she records (the outcome depends on ).
  • Bob's two settings are labelled and ; his choice is and his outcome is .

Here are fixed labels for the four possible settings, not variables; are the choices that range over them. Over many runs one estimates the four correlation functions

the expectation of the product of the two outcomes. is perfect correlation, perfect anticorrelation, no correlation. Everything below is a constraint on the four numbers — nothing else about the systems is used. That austerity is what makes the theorem so powerful: it is device-independent, indifferent to what the systems are or how they are measured.

graph LR
    A["Alice<br/>chooses a or a′<br/>outcome A = ±1"]
    S(("Source<br/>entangled pair"))
    B["Bob<br/>chooses b or b′<br/>outcome B = ±1"]
    A ---|"◀ system 1"| S
    S ---|"system 2 ▶"| B

The two wings are spacelike separated: each measurement completes before any light-speed signal from the other could arrive. Alice and Bob each freely pick one of two settings and record a outcome; the correlations are reconstructed only afterward, by comparing the two records over an ordinary classical channel. The schematic is deliberately abstract — §4 anchors the same skeleton to concrete hardware.

Agreed in advance vs. chosen at runtime. A protocol subtlety that the theorem hinges on: what the two parties fix beforehand and what they must not.

  • Fixed at design time (shared): the menu of two settings each party will choose between ( for Alice; for Bob), and a common reference frame — since depends on the relative angle between settings, the wings must share a definition of "angle zero," or the correlators are undefined.
  • Chosen at runtime (independent, never coordinated): which of the two settings is used on each run. These picks must be made randomly and independently, ideally so late that Alice's and Bob's choices are themselves spacelike separated. Pre-agreeing a schedule of choices, or signalling them across during the run, would reopen the locality and free-choice loopholes the test exists to close — which is why real experiments drive the settings with independent fast random-number generators (or starlight) at each station.

The correlators are then assembled after the runs, by sorting the paired records according to which setting combination occurred.

2. What "local hidden variables" means, precisely

The class of theories Bell excludes is defined by a single factorization condition, and naming its ingredients is the whole game philosophically. A local hidden-variable (LHV) model posits:

  1. A complete state — the "hidden variable," possibly a whole list of them — carried by the pair and fixed at the source. It ranges over some space with a probability density , . ( supplements the quantum state ; it need not replace it.)
  2. Local response probabilities and : the chance of each local outcome depends only on the local setting and on . Read as the probability of Alice getting outcome , given that she used setting and the hidden state was .
  3. Factorizability — the mathematical heart of "locality":

Read (3) slowly: given the complete state , the joint distribution of the two outcomes factorizes, so that once is fixed there is no residual correlation between the wings, and neither outcome's statistics depend on the distant setting. This is Bell's condition of local causality: screens off the two spacelike events from each other. It quietly bundles two independent assumptions (§10 pulls them apart) and presupposes a third — that the choice of settings is statistically independent of , i.e. (measurement independence, §11). The correlation a model predicts is the average over the hidden state:

Because any stochastic model can be simulated by averaging deterministic ones (absorb the local coin-flips into extra components of — Fine's theorem), we lose no generality by taking the responses to be definite functions , so , . Determinism is thus not an extra assumption; it is a free consequence of factorizability — a point that will matter a great deal in §12.

3. Deriving the CHSH inequality

We derive the bound rigorously from the three assumptions isolated in §2, keeping the model stochastic throughout so that the exact steps at which Parameter Independence (PI) and Outcome Independence (OI) are invoked are visible on the page. (Both are defined in full in §9: PI = a wing's outcome statistics are independent of the distant setting; OI = the two outcomes are uncorrelated once and both settings are fixed. Their conjunction is the factorizability (3) of §2.)

Notation. Fix the hidden state . Let be the joint outcome distribution, , with local marginals and . Define the local mean outcomes and the correlator at fixed

Step 1 — factor the correlator. Apply the two conditions of §2 in turn. First split the joint distribution:

← Outcome Independence used here. OI, , is precisely what lets the joint expectation split into a product of local means. Without it and everything below collapses.

Next remove the distant setting from each factor:

← Parameter Independence used here. PI, (and symmetrically for ), strips Bob's setting from Alice's response and vice versa, so each mean depends on its own setting only. This is what will allow a single value to be reused across both of Alice's settings in Step 3.

Step 2 — abbreviate. At fixed write

Step 3 — the algebraic core. Form the fixed- CHSH combination (Clauser–Horne–Shimony–Holt, 1969) and factor — the factorization is legitimate only because PI made the same numbers in the and terms:

With and the elementary identity for real ,

When the responses are deterministic — the case §2 reduces to via Fine's theorem — , so exactly one of vanishes and the other is , sharpening this to .

Step 4 — average over the hidden state. The measured correlator is . Measurement independence lets the same weight serve all four setting pairs, so they combine under one integral:

← Measurement Independence used here. MI, , is what allows the four correlators — each in principle weighted by its own — to be collected under the single average . Without it the terms cannot be added at fixed .

The triangle inequality with and then yields the CHSH inequality

Every local hidden-variable theory obeys . Tracing the logic back locates each premise exactly: OI turned the joint correlator into a product (Step 1); PI made each factor depend on its own setting only (Step 1), which licensed the factorization (Step 3); and measurement independence let the four terms share one average (Step 4). No physics entered — only these three premises and the two-valuedness of the outcomes. Drop any one and the bound is lost, which is precisely why a measured violation forces the rejection of at least one of them (§6, §9, §11).

The original 1964 inequality

Bell's 1964 argument predates CHSH and is less general in one respect — it needs an idealized perfect-anticorrelation premise — but it deserves the same rigorous treatment, because it turns on a different assumption and ties directly to the EPR step of §7.

Premises. Keep deterministic local responses (here denote generic measurement directions, placeholders for any of the settings below; is Alice's response function, is Bob's) — the deterministic form of PI: each is a function of its own setting and only; with determinism, OI is automatic — and measurement independence , and adjoin the empirical

  • (PA) Perfect anticorrelation at equal settings. Along the same direction , the two wings always disagree: for every — the singlet's signature.

Step 0 — PA forces on aligned settings. Since and , an average equal to its minimum pins the integrand:

so for -almost-every . This lets a single family carry both wings:

← Perfect anticorrelation used here. PA (equivalently, locality the observed — the EPR step of §7) is the sole idealization the 1964 form pays; CHSH (§3) never invokes it. This is exactly the "perfect-correlation premise" that is never exactly realizable in the lab.

Step 1 — three settings. For any three directions , use to write

hence, from ,

Step 2 — bound. The prefactor has modulus , and the bracket is nonnegative, (since ). Therefore

Step 3 — re-express via . By Step 0, , which yields the original Bell inequality

Quantum violation. With and coplanar settings , one has and , so the inequality demands , i.e. false. Quantum mechanics breaks the 1964 bound just as it breaks CHSH.

Comparison with CHSH. The two are the same discovery in different dress. The 1964 inequality buys vividness — via PA it forges the direct EPR-style link between perfect correlation and predetermined values (§7) — at the cost of a premise, exactly, that no real apparatus meets (finite detector efficiency and alignment). CHSH discards PA, needing only PI, OI, and MI, and is therefore the robust, lab-ready form every modern experiment reports as .

4. The quantum violation

Quantum mechanics predicts correlators that break the bound. Take the spin singlet and let each party measure spin along a direction in a plane, at angles . The quantum expectation is

Choose the settings at successive : , , , . Then each correlator has magnitude , and

The prediction exceeds the local bound of by a wide, experimentally comfortable margin. The correlations are too strong for any common cause carried from the source to explain.

The angles are not unique. Nothing privileges the particular numbers ; three separate freedoms remain.

  • Global orientation is free. Since depends only on the difference of angles (the singlet is rotationally invariant), rotating all four settings by a common offset changes nothing. Fixing is mere convention — starting at and shifting the rest along works identically.
  • Only the maximum is special. Equal spacing is singled out solely because it saturates Tsirelson's bound (§5); up to the global rotation above and relabelling/reflection it is the unique configuration that does so.
  • You need only beat , not reach . Take four equally spaced settings at a variable step , i.e. . Then

which exceeds the local bound of for every step from up to , and peaks at exactly at . So an entire continuum of angle choices refutes local hidden variables; the textbook set merely does it by the largest, most noise-tolerant margin. (For photons every physical angle is halved — see the factor-of-2 note below — so the same state-space spacing appears as on the polarizers.)

A concrete realization: from spins to photons

The machinery is easier to trust once anchored to real hardware. Two systems do the pedagogical work — one cleanest to reason about, one that actual experiments use. Both share the same skeleton sketched in §1: a central source emits an entangled pair, one member to each wing, where a freely-chosen setting yields a outcome.

The table below is the dictionary: it fixes what each abstract placeholder of §1§2 becomes in each realization.

Abstract element (§1–§2)Spin-½ / Stern–GerlachPhoton / SPDC
Sourcesinglet molecule dissociatingnonlinear crystal (down-conversion)
Systems 1 & 2 (the emitted pair)two neutral spin-½ atomstwo photons
Quantum state of the pairsinglet Bell state
Setting labels (Alice)magnet angles polarizer angles
Setting labels (Bob)magnet angles polarizer angles
Choice which magnet orientation is usedwhich polarizer orientation is used
Outcome deflected up () / down ()parallel channel () / orthogonal ()
Correlator

One placeholder is deliberately left unfilled: the hidden variable has no quantum referent in either column. It is exactly the extra ingredient a local model would have to add on top of the quantum state to explain the correlations classically — and the content of the theorem is that the experimental leaves no room for it. In the quantum description the state (or ) is the whole story; there is nothing else the pair "carries."

The conceptual version: spin-½ and Stern–Gerlach. The derivation above already is this version. A spin-0 source emits two spin-½ particles in the singlet — concretely, a singlet diatomic molecule dissociating into two neutral spin-½ atoms (silver or an alkali, whose moment is a single unpaired electron spin), the source EPR–Bohm imagined; not a spin-0 pion, whose weak decay yields an undetectable neutrino and a parity-fixed helicity rather than a rotatable singlet. Alice and Bob each own a Stern–Gerlach magnet oriented at an angle in a plane, and the particle is deflected up () or down () — a position on a screen, which is exactly the kind of definite binary record the argument needs. (Neutral atoms are essential: free electrons, being charged, cannot be Stern–Gerlach-analyzed — the Lorentz force blurs the spin splitting.) The law is , and the successive- settings of §4 give . Simplest algebra, most vivid outcomes.

The experimental version: polarization-entangled photons. Essentially every real Bell test — Aspect (1982), Weihs (1998), the 2015 loophole-free photon experiments — uses light. A nonlinear crystal undergoing spontaneous parametric down-conversion (SPDC) splits one pump photon into an entangled pair, e.g. . Each party sends their photon through a polarizer (or a polarizing beam-splitter with a detector on each port) rotated to angle : transmission in the parallel channel is , the orthogonal channel . The correlation is

and the settings give .

The factor of 2 to watch for. The photon angles () are half the spin angles (). This is not arbitrary: polarization is a spin-1 degree of freedom, so rotating a polarizer by a physical angle rotates the quantum state by on the Poincaré sphere — hence for photons against for spin-½.

Why photons in practice. They are cheap to entangle (SPDC), travel kilometres through fiber or free space with little decoherence, and allow setting switches faster than light can cross the apparatus — closing the locality loophole (§12). Their historic weakness was detector efficiency (the detection loophole), which is why the loophole-free tests used either near-unit-efficiency photodetectors (NIST, Vienna 2015) or matter qubits — the Delft 2015 experiment entangled electron spins in nitrogen-vacancy centres in diamond via photon interference. For teaching: reason with spins, connect to the lab with photons.

5. Tsirelson's bound: quantum mechanics is nonlocal but not maximally so

Quantum mechanics violates CHSH, but only up to a ceiling. Promote the outcomes to -valued observables (Hermitian, , and since Alice's and Bob's operators act on different factors). For the operator a short computation gives the sum-of-squares identity

Each commutator of two observables has norm , so , whence

This is Tsirelson's bound (1980), saturated by the singlet with the settings of §4. Its philosophical interest is sharp: the purely relativistic constraint of no-signalling (§10) by itself permits as large as — the hypothetical Popescu–Rohrlich (PR) box. So nature is nonlocal, yet strictly less nonlocal than causality alone would allow. Why it stops exactly at has become a research programme in the reconstruction of quantum theory from information-theoretic principles (e.g. information causality, Pawłowski et al. 2009), and a clue that the Hilbert-space formalism encodes a principle we have not yet named.


Part II — The philosophical analysis

6. What, exactly, has been proved?

State the logic carefully, because loose statements of "what Bell proved" cause most of the confusion. The demonstrable core is a conditional:

Experiment finds . By modus tollens, at least one antecedent is false. The theorem itself is a piece of mathematics and is not in dispute; all the philosophy lives in the choice of which antecedent to reject — and each choice is a different picture of the world. The remainder of this remark is a tour of those antecedents and the positions that deny each.

The full ledger of premises — gathered from §2, §3, and §9 into one place — is:

AssumptionFormal statementMeaningNeeded byRejected by
Realism / hidden variables with , fixing the outcome responses outcomes trace to a state carried from the source (locality is not yet assumed — that is the PI/OI rows)both formsCopenhagen, QBism (§7, §13)
Parameter Independence (PI)local marginal is independent of the distant settingbothBohmian mechanics (§9, §13)
Outcome Independence (OI)no residual outcome–outcome correlation once is fixedbothcollapse / orthodox QM (§9, §13)
PI OI = FactorizabilityBell local causality: screens off the wingsboth— (the conjunction)
Measurement Independence (MI)settings are chosen independently of bothsuperdeterminism, retrocausality (§11)
Two-valued outcomeseach run yields a single definite resultbothmany-worlds (§13)
Perfect anticorrelation (PA) for all singlet certainty on aligned settings1964 form only— (CHSH needs it not)

Two entries deserve emphasis. Determinism is absent from the list: it is not assumed but derived from factorizability via Fine's theorem (§2), so "give up determinism" is not among the escape routes. And PA is required by the 1964 inequality but not by CHSH (§3) — which is exactly why every modern, loophole-tolerant experiment reports the CHSH quantity . Everything below dismantles these premises one at a time.

7. The realism assumption and a common oversimplification

Textbooks routinely say Bell refuted "local realism," presenting realism (or "hidden variables," or "definiteness") and locality as two separable premises, either of which might be dropped — and inviting the comfortable conclusion that we may keep locality by merely abandoning naïve realism about unmeasured quantities. This gloss is at best incomplete, and Bell himself rejected it.

The subtlety is that determinism/definiteness is not an independent postulate of the argument — it can be derived. Bell's own two-part reasoning (following EPR) runs:

  1. EPR step. For the singlet, whenever Alice and Bob measure along the same axis they get perfectly anticorrelated results with certainty. If the world is local, Alice's distant choice cannot influence Bob's system; so Bob's definite outcome must have been fixed in advance by something carried locally — a hidden variable . Locality + perfect correlations predetermined values. Determinism is a conclusion here, not an assumption.
  2. Bell step. Those predetermined local values obey the inequality (§3).
  3. Experiment. The inequality is violated.

On this reading the only premise standing between "local" and the false inequality is locality itself (given the empirical perfect correlations), so the honest conclusion is not "local realism is false" but "locality is false, full stop." This is the position of Bell's mature essay La nouvelle cuisine and of its contemporary defenders (Maudlin, Norsen): the world is not locally causal, and no retreat to anti-realism rescues locality.

The opposing camp notes that the CHSH derivation of §3 needs only factorizability and never invokes the perfect correlations that power the EPR step; factorizability can be motivated as "locality a separate assumption that fixes outcome probabilities" (a mild realism). On this view "realism" is a genuine, droppable premise, and interpretations that deny observer-independent outcomes (§14) exploit exactly that. Both readings are internally coherent; the disagreement is about which minimal set of premises best regiments the physics. The safe, neutral statement — the one to actually remember — is:

The world cannot be both local and such that measurements reveal pre-existing, locally-determined values. At least one of those must go.

8. Locality, decomposed: Bell locality as a screening-off condition

To see what "locality" is doing, connect factorization (3) to a principle from the philosophy of causation. Reichenbach's common cause principle says that a correlation between two events that do not cause one another must be due to a common cause in their shared past that screens off the correlation — renders the events statistically independent once the common cause is specified. Bell's factorizability is Reichenbach's screening-off, applied to the complete common cause lying in the past light cone of both measurements. The violation of CHSH therefore says:

  • Either there is no common cause that screens off the correlations — so Reichenbach's principle fails for quantum systems (Van Fraassen's and Cartwright's reading), and we must accept brute spacelike correlations with no local explanation,
  • Or there is a genuine non-local influence — a direct causal link between the wings, an "action at a distance" of the sort EPR found intolerable.

Either horn abandons the classical picture of a world knit together only through its past. This is the deepest content of the theorem: the common-cause structure of spacetime — the assumption that spacelike correlations bottom out in the past — does not survive quantum mechanics.

9. Jarrett–Shimony: parameter independence vs. outcome independence

The single most illuminating philosophical move is to split factorizability into two logically independent conditions (Jarrett 1984; Shimony). Factorization (3) is equivalent to the conjunction of:

  • Parameter Independence (PI) — Alice's outcome statistics do not depend on Bob's setting: . (Jarrett's "locality.")
  • Outcome Independence (OI) — given the settings and , Alice's outcome is statistically independent of Bob's outcome: . (Jarrett's "completeness.")

Now the crucial fact: quantum mechanics violates OI but satisfies PI. The entangled correlations couple the two outcomes (learning Bob's result changes the distribution of Alice's), yet neither party's marginal statistics depend on the distant party's setting choice. This asymmetry is the linchpin of the whole philosophy of the subject:

Violated by QM?Consequence if violated
Parameter IndependenceNoWould permit superluminal signalling — controllable, usable communication
Outcome IndependenceYesOnly uncontrollable correlation — no signalling

Whose OI, whose PI? "QM violates OI, satisfies PI" is the orthodox reading, where the quantum state is the complete state (). Bell refutes only the conjunction, so which conjunct fails is interpretation-relative: a deterministic hidden-variable theory (Bohmian mechanics) instead satisfies OI and violates PI — with the outcomes fixed as definite functions of , nothing is left to correlate once and the settings are given (OI holds trivially), and the whole nonlocality lands on PI (the distant setting moves the local outcome). So OI is "the culprit" only in PI-respecting theories; the two camps are tabulated in §13.

Because PI survives, the correlations cannot be used to send anything: Alice, by her choice of setting, cannot alter what Bob sees. Shimony christened the residual, PI- respecting nonlocality "passion at a distance" as opposed to "action at a distance." The world is nonlocal in its correlations but not in any signal — which is exactly the loophole through which quantum theory and relativity coexist.

PI and OI presuppose realism. A structural point worth stating explicitly: both conditions are defined relative to a complete state — they are constraints on — so the very whose existence is the realism premise (§6) is logically prior to them. Two consequences follow. First, one cannot "reject realism yet accept both PI and OI": holding both is just factorizability, which forces and is refuted; and rejecting the value-framework wholesale (strong Copenhagen, or Everett denying single outcomes) leaves PI/OI with no referent — they become inapplicable, not merely false. Second, there is a thin middle road — "reject only the extra hidden variables and take " (orthodox QM): then PI/OI are definable, and orthodox QM keeps PI ( no-signalling) while dropping OI (the entanglement of is the OI violation). So the escape "no hidden variables" always resolves into either keep PI, drop OI or dissolve the framework — never keep both.

10. No-signalling and the "peaceful coexistence" with relativity

That PI holds is the theorem of no-signalling: Bob's reduced state is completely unaffected by Alice's choice of measurement, so no marginal — nothing Bob can access locally — carries information about Alice's setting. The nonlocal correlations become visible only after the two records are brought together over an ordinary, luminally-limited classical channel. Hence:

  • At the operational/relativistic level there is genuine peace: no experiment exploiting entanglement can send a superluminal message, so special relativity's prohibition on faster-than-light signalling is untouched. (This is the sense in which the block-universe causal structure of relativity survives.)
  • At the ontological level the peace is uneasy. A dynamical collapse correlated across a spacelike interval has no Lorentz-invariant "order" — which measurement happened "first" is frame-dependent — so any story in which one outcome brings about the other seems to need a preferred foliation of spacetime that relativity denies. Bohmian mechanics (§13) bites this bullet with an explicit preferred frame; others deny there is any bringing-about to order. The tension is not a paradox but a cost that every interpretation must pay somewhere.

The moral is a careful distinction philosophy of physics insists upon: signal locality (no superluminal communication) is empirically secure; Bell locality / local causality (no superluminal influence of any kind) is refuted. The two come apart precisely because of the PI/OI split.

11. The measurement-independence assumption: superdeterminism and retrocausality

Buried in the derivation is a premise so natural it is easy to miss: — the hidden state is statistically independent of which settings get chosen. This is measurement independence (also "free choice," "no-conspiracy," "-independence"). It is often called the free-will assumption, because it is what licenses treating the experimenters' setting choices as freely (or at least genuinely randomly) made rather than pre-arranged to match — though no metaphysical libertarian free will is needed, only effective statistical independence. Deny it and the derivation collapses: if can be correlated with the future settings, a purely local model can fake any correlation whatever, because the source "knows" in advance how the detectors will be set.

Two very different metaphysics exploit this:

  • Superdeterminism ('t Hooft; Hossenfelder). A single deterministic law fixes both the hidden variables and the experimenters' setting choices from common initial conditions, so the required correlation between and is built into the initial state of the universe. Locality is saved — at the price of denying that the settings are freely (or even effectively randomly) chosen. Critics regard this as a conspiracy that would undermine the very possibility of experimental science (if setting choices are secretly correlated with what is measured, no controlled experiment is trustworthy); defenders reply that "conspiracy" begs the question about what independence nature owes us.
  • Retrocausality (transactional and two-state-vector interpretations; Price). The measurement setting influences the hidden state backward along the light cone, so the correlation is causal but time-symmetric rather than superluminal. This can preserve a form of locality-in-spacetime by trading it for causation running both temporal directions — attractive to those who take the time-symmetry of microphysics seriously, unsettling to those who take the asymmetry of causation as basic.

Because measurement independence is an assumption, it cannot be proved, only made implausible: experiments have pushed the setting choices to sources maximally decorrelated from any plausible common past — human free choices (the BIG Bell Test, 2018) and the light of distant quasars emitted billions of years ago (cosmic Bell tests, 2018) — shrinking the space in which a non-conspiratorial correlation could hide. The connection to the metaphysics of agency is direct, and links this remark to the free-will debate: what the experimenter's "free choice" of setting is becomes a load-bearing physical assumption.

12. The loopholes: how an experiment can fail to close the argument

A raw violation of refutes local hidden variables only if the experiment actually instantiates the derivation's premises. Each way it might fail to is a loophole — a local-realist escape route kept open by an experimental imperfection. The history of the subject is the history of closing them.

  • Locality (communication) loophole. If the two measurements are not spacelike separated, a subluminal influence — the setting or outcome at one wing propagating to the other — could coordinate the results locally. Closed by choosing settings randomly and fast and placing the wings far enough apart that no light-speed signal can cross during a measurement (Aspect's switching, 1982; Weihs et al. with truly random, spacelike-separated choices, 1998).
  • Detection (fair-sampling) loophole. If detectors miss most pairs, the detected sub-ensemble may violate CHSH while the full ensemble obeys it — a local model can exploit low efficiency by letting decide whether a particle is detected. Closing it requires detection efficiency above a threshold (the Eberhard bound, for optimized states), reached with trapped ions, atoms, superconducting qubits, and eventually high-efficiency photodetectors.
  • Freedom-of-choice loophole. The formal face of measurement independence (§11): if setting choices share a common cause with , all bets are off. Mitigated, never fully closed, by decorrelating the choices as far into the independent past as possible (quasar light, human choices).
  • Memory loophole. If trials are not independent and identically distributed, a local model with memory of past settings could bias the statistics; handled by proper hypothesis testing that assumes no i.i.d. structure (Gill).
  • Coincidence-time / collapse-locality loopholes. Subtler timing- and collapse-model escapes, addressed by fixed time-windows and by pushing any putative collapse to spacelike separation.

The loophole-free experiments of 2015 — Hensen et al. (Delft, entangled electron spins in nitrogen-vacancy centres via event-ready entanglement swapping), Giustina et al. (Vienna) and Shalm et al. (NIST), both with high-efficiency photons — closed the locality and detection loopholes simultaneously for the first time. The 2022 Nobel Prize to Clauser, Aspect, and Zeilinger recognized this arc. What remains is only the freedom-of-choice/superdeterminism loophole, which is not so much an experimental defect as a metaphysical stance about whether nature grants us independent choices at all (§11). Within any framework that grants ordinary statistical independence of setting choices, local hidden-variable theories are dead.

13. The menu of interpretations, sorted by which premise they reject

Every serious interpretation of quantum mechanics can be located by which antecedent of the Bell conditional (§6) it sacrifices. This is the most useful map the theorem provides, and it connects directly to the measurement problem. The four "give up X" options of the introduction correspond one-to-one to the premises isolated in §6–§9:

Premise sacrificedInterpretation(s)What it keepsWhat it costs
Locality — Bell local causality (§8, §10)Bohmian mechanics; spontaneous collapse (GRW, CSL)determinism, definite outcomes, a single worldavowedly nonlocal dynamics; a preferred foliation straining Lorentz invariance
Outcome definiteness — "realism," pre-existing values (§7)Copenhagen, QBism, relational QMlocality (PI intact, §9)no observer-independent values; the wavefunction becomes knowledge/belief, not a thing
Single outcomesMany-worlds (Everett)locality and determinisma branching-multiverse ontology; recovering the Born-rule probabilities
Measurement independence — "free choice" (§11)Superdeterminism; retrocausalitylocality and definitenessdenies free / effectively-random settings — a cosmic "conspiracy," or backward causation

A few nuances the grid flattens. Bohmian mechanics wears its nonlocality on its sleeve — its guidance equation makes each particle's motion depend instantaneously on the distant particle's position — and its virtue is precisely that honesty: it shows the deterministic, realist completion EPR wanted is possible, but only if nonlocal. Many-worlds claims to buy locality back: with every outcome realized on a separate branch there is no single result- pair for an influence to correlate, and the branching spreads no faster than light via decoherence. QBism dissolves the nonlocality into information updating rather than physical influence, reading as an agent's belief state.

Refining the "Locality" row: PI vs. OI. The first row is deliberately coarse. By the Jarrett–Shimony split (§9), Bell locality is the conjunction PI OI, and violating the inequality means dropping at least one half — but which half is interpretation-dependent, and it divides the nonlocal theories into two camps:

Nonlocal theoryOIPIWhy
Deterministic hidden variables (Bohmian mechanics)satisfied (trivially)violatedwith definite values, once and both settings are fixed the outcomes are determined, so no residual outcome–outcome correlation remains to break (OI holds); the nonlocality must live in PI — the distant setting genuinely affects the local outcome
Stochastic collapse / orthodox QM (GRW, CSL, Copenhagen)violatedsatisfiedthe local marginals are independent of the distant setting (no-signalling PI), yet learning the distant outcome shifts the local distribution

Both camps are nonlocal, but they break different halves: Bohm sacrifices Parameter Independence; collapse/orthodox theories sacrifice Outcome Independence. Crucially, both preserve signal-locality (§10) — Bohm's PI-violation never surfaces in the statistics under quantum equilibrium, so it cannot be used to send a message. One caveat: PI and OI presuppose a joint distribution — i.e. that there are outcomes to correlate — so the split applies to the "Locality" row only; the outcome-definiteness and single-outcome rows reject that framework upstream.

No option is free. The theorem does not select an interpretation; it taxes every one, and the philosophy of quantum mechanics is largely the accountancy of these taxes.

14. Beyond CHSH: sharper and broader forms

Two extensions deepen the philosophical picture without changing its moral.

  • Nonlocality without inequalities. For three or more particles, the Greenberger–Horne–Zeilinger (GHZ) state yields an all-or-nothing contradiction: local hidden variables must assign values that satisfy a set of equations with no solution, so a single ideal run — not a statistical margin — refutes local realism. Hardy's two-particle construction achieves something similar for a fraction of runs. These show the Bell phenomenon is not an artifact of statistics but a flat logical incompatibility.
  • Contextuality (Kochen–Specker). Bell nonlocality is a special case of a more general impossibility: no assignment of pre-existing values to all observables can be non-contextual (independent of which compatible observables are co-measured). Kochen–Specker needs no spatial separation at all — it is a constraint on hidden variables within a single system. Locality is then seen as spatial non- contextuality, and Bell's theorem as its most physically dramatic instance.

15. What the theorem does and does not license

A closing ledger, to guard against the two opposite over-readings.

It does establish:

  • The world is not locally causal: spacelike correlations cannot, in general, be explained by a common cause in the shared past screening off the wings (§8). This is a fact about nature, not about quantum formalism — any successor theory must reproduce it.
  • Therefore one must give up at least one of: locality, outcome-definiteness, single outcomes, or measurement independence (§13). There is no locally-causal, single-world, definite-outcome, free-choice theory of our world.

It does not establish:

  • That superluminal signalling is possible — no-signalling holds, and relativity's operational prohibition is safe (§10).
  • That "realism" simpliciter is refuted — whether realism is even a separable premise is contested (§7); a nonlocal realist theory (Bohm) fits all the data.
  • That the correlations are "spooky" in any usable sense — the nonlocality is passion, not action (§9): visible only after classical comparison.

The positive payoff is not merely negative metaphysics. Because a CHSH violation certifies that no local, predetermined mechanism produced the data, it underwrites device-independent cryptography and certified randomness: the same feature that embarrassed Einstein is now a resource that lets one trust a random number without trusting the device that made it. That is the fitting last word on a theorem that turned a thought-experiment about incompleteness into working technology — and a piece of metaphysics into a laboratory fact.

See also