Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

The Philosophy of Bell's Theorem

Bell's theorem is the closest thing physics has to an experimental metaphysics: a proof that no theory of a certain very natural kind — one in which the outcomes of spacelike-separated measurements are fixed by locally-carried information — can reproduce the statistical predictions of quantum mechanics, together with laboratory tests that decide the matter in QM's favour. Its philosophical importance is that the kind of theory it excludes is exactly the kind common sense expects the world to be: one in which distant events are correlated only because of what they share in their common past, and in which measurement merely reveals pre-existing values. This remark reads the Entanglement, EPR, and Bell's Theorem page philosophically. It does so in two movements. First, it works through the derivation of the CHSH inequality, its quantum violation, and the experimental machinery that puts both to the test, in enough detail that every assumption is visible on the page — because the philosophy is entirely a dispute about which assumption to blame. Second, it dismantles those assumptions one by one, catalogues the experimental loopholes and their closure, and lays out the menu of metaphysical options the theorem leaves open: give up locality, give up definite values, give up free choice, or give up single outcomes.

The guiding tension is set by the two historical bookends. Einstein, Podolsky, and Rosen argued in 1935 that quantum mechanics must be incomplete — that its entangled correlations could only be explained by "elements of reality" the wavefunction fails to mention, on pain of a spooky action at a distance no respectable physics should tolerate. Their argument was a reductio aimed at forcing hidden variables. Bell's 1964 reply turned the reductio into a test: he showed that the very hidden-variable theories EPR demanded make predictions that differ numerically from quantum mechanics, so the question "is the world locally explicable?" is not a matter of taste but of experiment — and experiment has answered no.


Part I — The theorem and the experiment

1. The experimental scenario

Fix the setup that all the mathematics refers to. A source emits pairs of systems; one member goes to Alice, the other to Bob, and the two wings are far enough apart that a measurement in one can be completed before any light-speed signal from the other could arrive (they are spacelike separated). Each party has two available measurement settings (two fixed apparatus configurations, e.g. two directions), and on each run selects one of them and records one of two outcomes, (why two settings is the minimum that can prove anything is shown in §2):

  • Alice's two settings are labelled and . On a given run exactly one of them is in force; write for the setting used and for the outcome recorded (the outcome depends on ).
  • Bob's two settings are labelled and ; the setting used is and his outcome is .

Here are fixed labels for the four possible settings, not variables; range over them. How a run comes to use one setting rather than another — whether by free choice, a random generator, or a fixed schedule — is a question of experimental protocol, and is deferred entirely to §6. The derivation in §3 never asks. Over many runs one estimates the four correlation functions

the expectation of the product of the two outcomes. is perfect correlation, perfect anticorrelation, no correlation. Everything below is a constraint on the four numbers — nothing else about the systems is used. That austerity is what makes the theorem so powerful: it is device-independent, indifferent to what the systems are or how they are measured.

graph LR
    A["Alice<br/>setting a or a′<br/>outcome A = ±1"]
    S(("Source<br/>entangled pair"))
    B["Bob<br/>setting b or b′<br/>outcome B = ±1"]
    A ---|"◀ system 1"| S
    S ---|"system 2 ▶"| B

The two wings are spacelike separated: each measurement completes before any light-speed signal from the other could arrive. On each run each wing has one of its two settings in force and records a outcome; the correlations are reconstructed only afterward, by comparing the two records over an ordinary classical channel. The schematic is deliberately abstract — it fixes the objects the derivation quantifies over and nothing more. §5 anchors the same skeleton to concrete hardware, and §6 supplies the protocol that decides which setting a given run uses.

2. What "local hidden variables" means, precisely

The class of theories Bell excludes is defined by a single factorization condition, and naming its ingredients is the whole game philosophically. A local hidden-variable (LHV) model posits:

  1. A complete state — the "hidden variable," possibly a whole list of them — carried by the pair and fixed at the source. It ranges over some space with a probability density , . ( supplements the quantum state ; it need not replace it.)
  2. Local response probabilities and : the chance of each local outcome depends only on the local setting and on . Read as the probability of Alice getting outcome , given that she used setting and the hidden state was .
  3. Factorizability — the mathematical heart of "locality":

Read (3) slowly: given the complete state , the joint distribution of the two outcomes factorizes, so that once is fixed there is no residual correlation between the wings, and neither outcome's statistics depend on the distant setting. This is Bell's condition of local causality: screens off the two spacelike events from each other. It quietly bundles two independent assumptions (§11 pulls them apart) and presupposes a third — that the choice of settings is statistically independent of , i.e. (measurement independence, §13). The correlation a model predicts is the average over the hidden state:

Because any stochastic model can be simulated by averaging deterministic ones (absorb the local coin-flips into extra components of — Fine's theorem), we lose no generality by taking the responses to be definite functions , so , . Determinism is thus not an extra assumption; it is a free consequence of factorizability — a point that will matter a great deal in §8–§9.

Why two settings per wing? The scenario of §1 stipulates alternatives at each station, and it is worth seeing at once that this is not a bookkeeping convenience but the precondition for there being anything to prove. With only one setting per wing, every observed joint distribution admits an LHV model: take drawn from that very distribution and have each wing report its own component. The same holds if only one party has alternatives (draw from its marginal, then each independently from ). So no experiment with fewer than two settings on each side can refute local hidden variables, however strange its correlations look — is the minimum.

What conditions (1)–(3) really commit one to is that a single fixes the responses to alternatives that are never jointly measured: on any given run Alice uses or , yet the model must say what she would have got either way. Equivalently, an LHV model entails a joint distribution over the four counterfactual values whose pairwise marginals are the four measurable correlators. Fine's theorem makes the correspondence exact: in this scenario CHSH holds if and only if such a joint distribution exists. A violation is therefore precisely the statement that the four correlators cannot be glued into one consistent classical ledger — and there is nothing to glue until there are alternatives.

3. Deriving the CHSH inequality

We derive the bound rigorously from the three assumptions isolated in §2, keeping the model stochastic throughout so that the exact steps at which Parameter Independence (PI) and Outcome Independence (OI) are invoked are visible on the page. (Both are defined in full in §11: PI = a wing's outcome statistics are independent of the distant setting; OI = the two outcomes are uncorrelated once and both settings are fixed. Their conjunction is the factorizability (3) of §2.)

Notation. Fix the hidden state . Let be the joint outcome distribution, , with local marginals and . Define the local mean outcomes and the correlator at fixed

Step 1 — factor the correlator. Apply the two conditions of §2 in turn. First split the joint distribution:

← Outcome Independence used here. OI, , is precisely what lets the joint expectation split into a product of local means. Without it and everything below collapses.

Next remove the distant setting from each factor:

← Parameter Independence used here. PI, (and symmetrically for ), strips Bob's setting from Alice's response and vice versa, so each mean depends on its own setting only. This is what will allow a single value to be reused across both of Alice's settings in Step 3.

Step 2 — abbreviate. At fixed write

Step 3 — the algebraic core. Form the fixed- CHSH combination (Clauser–Horne–Shimony–Holt, 1969) and factor — the factorization is legitimate only because PI made the same numbers in the and terms:

With and the elementary identity for real ,

When the responses are deterministic — the case §2 reduces to via Fine's theorem — , so exactly one of vanishes and the other is , sharpening this to .

Step 4 — average over the hidden state. The measured correlator is . Measurement independence lets the same weight serve all four setting pairs, so they combine under one integral:

← Measurement Independence used here. MI, , is what allows the four correlators — each in principle weighted by its own — to be collected under the single average . Without it the terms cannot be added at fixed .

The triangle inequality with and then yields the CHSH inequality

Every local hidden-variable theory obeys . Tracing the logic back locates each premise exactly: OI turned the joint correlator into a product (Step 1); PI made each factor depend on its own setting only (Step 1), which licensed the factorization (Step 3); and measurement independence let the four terms share one average (Step 4). No physics entered — only these three premises and the two-valuedness of the outcomes. Drop any one and the bound is lost, which is precisely why a measured violation forces the rejection of at least one of them (§8, §11, §13).

Nothing here varies, and nothing is chosen. It is worth pausing on what the derivation did not use. There are no runs, no time-ordering, and no agents anywhere in it: are four points in the domain of the response functions, and evaluates all four at the same simultaneously. What an LHV model must supply is a total function on that domain — a value for even as the model is also asked for — which is the counterfactual commitment of §2, not an experimental procedure. So read strictly, the theorem is atemporal algebra: any pair of functions into , averaged against a single measure , satisfies .

Where, then, do "vary," "randomize," and "choose" come from? Three levels are worth keeping apart, because conflating them is the source of most confusion about what Bell tests are doing.

  1. The algebra (this section). Settings are free variables and is one fixed measure. Nothing varies, nothing happens, nobody chooses.
  2. The theory class (§2, §8). Measurement independence is the stipulation that a model's genuinely does not depend on . This is a substantive restriction on which models count as local hidden-variable models — but it is still a claim about theories, with no experimenter anywhere in it.
  3. The experiment (§6, §13–§14). Only here do choosing, randomizing, and fast spacelike switching appear, and they have exactly one job: to warrant identifying four measured sub-ensemble averages with four integrals against a common . Randomization is evidential hygiene for level 3 and no part of levels 1 or 2.

The original 1964 inequality

Bell's 1964 argument predates CHSH and is less general in one respect — it needs an idealized perfect-anticorrelation premise — but it deserves the same rigorous treatment, because it turns on a different assumption and ties directly to the EPR step of §9.

Premises. Keep deterministic local responses (here denote generic measurement directions, placeholders for any of the settings below; is Alice's response function, is Bob's) — the deterministic form of PI: each is a function of its own setting and only; with determinism, OI is automatic — and measurement independence , and adjoin the empirical

  • (PA) Perfect anticorrelation at equal settings. Along the same direction , the two wings always disagree: for every — the singlet's signature.

Step 0 — PA forces on aligned settings. Since and , an average equal to its minimum pins the integrand:

so for -almost-every . This lets a single family carry both wings:

← Perfect anticorrelation used here. PA (equivalently, locality the observed — the EPR step of §9) is the sole idealization the 1964 form pays; CHSH (§3) never invokes it. This is exactly the "perfect-correlation premise" that is never exactly realizable in the lab.

Step 1 — three settings. For any three directions , use to write

hence, from ,

Step 2 — bound. The prefactor has modulus , and the bracket is nonnegative, (since ). Therefore

Step 3 — re-express via . By Step 0, , which yields the original Bell inequality

Quantum violation. With and coplanar settings , one has and , so the inequality demands , i.e. false. Quantum mechanics breaks the 1964 bound just as it breaks CHSH.

Comparison with CHSH. The two are the same discovery in different dress. The 1964 inequality buys vividness — via PA it forges the direct EPR-style link between perfect correlation and predetermined values (§9) — at the cost of a premise, exactly, that no real apparatus meets (finite detector efficiency and alignment). CHSH discards PA, needing only PI, OI, and MI, and is therefore the robust, lab-ready form every modern experiment reports as .

4. The quantum violation

Quantum mechanics predicts correlators that break the bound. Take the spin singlet and let each party measure spin along a direction in a plane, at angles . The quantum expectation is

Choose the settings at successive : , , , . Then each correlator has magnitude , and

shared 0° reference 45° 45° 45° a = 0° b = 45° a′ = 90° b′ = 135° ■ Alice (a, a′) ■ Bob (b, b′)

The four CHSH axes in the shared reference frame of §1: Alice's (blue) and Bob's (red). Every adjacent pair sits apart; Alice's two settings are apart, Bob's two are apart, interleaved at — the arrangement that reaches . (These are the spin-½ magnet angles; for polarization-entangled photons every angle is halved — the factor-of-2 note below — giving on the polarizers.)

source particle 1 particle 2 θA +1 −1 Alice θB +1 −1 Bob spacelike separated

The same experiment as a physical scene: the source in the middle emits the pair to two spatially separated stations. Each analyzer is a disk (a polarizer, or a Stern–Gerlach magnet) set to its own angle measured from a shared vertical reference, and it sorts its particle into a or detector. Only the relative angle enters the correlation, ; the CHSH choice realizes as the and of the compass above.

The prediction exceeds the local bound of by a wide, experimentally comfortable margin. The correlations are too strong for any common cause carried from the source to explain.

The angles are not unique. Nothing privileges the particular numbers ; three separate freedoms remain.

  • Global orientation is free. Since depends only on the difference of angles (the singlet is rotationally invariant), rotating all four settings by a common offset changes nothing. Fixing is mere convention — starting at and shifting the rest along works identically.
  • Only the maximum is special. Equal spacing is singled out solely because it saturates Tsirelson's bound (§7); up to the global rotation above and relabelling/reflection it is the unique configuration that does so.
  • You need only beat , not reach . Take four equally spaced settings at a variable step , i.e. . Then

which exceeds the local bound of for every step strictly between and , and peaks at exactly at . (Both endpoints sit on the bound, not above it: is the degenerate case where all four settings coincide, and by construction.) So an entire continuum of angle choices refutes local hidden variables; the textbook set merely does it by the largest, most noise-tolerant margin. (For photons every physical angle is halved — see the factor-of-2 note below — so the same state-space spacing appears as on the polarizers.)

5. Realizing the experiment: spins and photons

The machinery is easier to trust once anchored to real hardware. Two systems do the pedagogical work — one cleanest to reason about, one that actual experiments use. Both share the same skeleton sketched in §1: a central source emits an entangled pair, one member to each wing, where a freely-chosen setting yields a outcome.

The table below is the dictionary: it fixes what each abstract placeholder of §1§2 becomes in each realization.

Abstract element (§1–§2)Spin-½ / Stern–GerlachPhoton / SPDC
Sourcesinglet molecule dissociatingnonlinear crystal (down-conversion)
Systems 1 & 2 (the emitted pair)two neutral spin-½ atomstwo photons
Quantum state of the pairsinglet Bell state
Setting labels (Alice)magnet angles polarizer angles
Setting labels (Bob)magnet angles polarizer angles
Choice which magnet orientation is usedwhich polarizer orientation is used
Outcome deflected up () / down ()parallel channel () / orthogonal ()
Correlator

One placeholder is deliberately left unfilled: the hidden variable has no quantum referent in either column. It is exactly the extra ingredient a local model would have to add on top of the quantum state to explain the correlations classically — and the content of the theorem is that the experimental leaves no room for it. In the quantum description the state (or ) is the whole story; there is nothing else the pair "carries."

The conceptual version: spin-½ and Stern–Gerlach. The derivation above already is this version. A spin-0 source emits two spin-½ particles in the singlet — concretely, a singlet diatomic molecule dissociating into two neutral spin-½ atoms (silver or an alkali, whose moment is a single unpaired electron spin), the source EPR–Bohm imagined; not a spin-0 pion, whose weak decay yields an undetectable neutrino and a parity-fixed helicity rather than a rotatable singlet. Alice and Bob each own a Stern–Gerlach magnet oriented at an angle in a plane, and the particle is deflected up () or down () — a position on a screen, which is exactly the kind of definite binary record the argument needs. (Neutral atoms are essential: free electrons, being charged, cannot be Stern–Gerlach-analyzed — the Lorentz force blurs the spin splitting.) The law is , and the successive- settings of §4 give . Simplest algebra, most vivid outcomes.

The experimental version: polarization-entangled photons. Essentially every real Bell test — Aspect (1982), Weihs (1998), the 2015 loophole-free photon experiments — uses light. A nonlinear crystal undergoing spontaneous parametric down-conversion (SPDC) splits one pump photon into an entangled pair, e.g. . Each party sends their photon through a polarizer (or a polarizing beam-splitter with a detector on each port) rotated to angle : transmission in the parallel channel is , the orthogonal channel . The correlation is

and the settings give .

The factor of 2 to watch for. The photon angles () are half the spin angles (). This is not arbitrary: polarization is a spin-1 degree of freedom, so rotating a polarizer by a physical angle rotates the quantum state by on the Poincaré sphere — hence for photons against for spin-½.

Why photons in practice. They are cheap to entangle (SPDC), travel kilometres through fiber or free space with little decoherence, and allow setting switches faster than light can cross the apparatus — closing the locality loophole (§14). Their historic weakness was detector efficiency (the detection loophole), which is why the loophole-free tests used either near-unit-efficiency photodetectors (NIST, Vienna 2015) or matter qubits — the Delft 2015 experiment entangled electron spins in nitrogen-vacancy centres in diamond via photon interference. For teaching: reason with spins, connect to the lab with photons.

6. How a Bell test is actually run: from protocol to

Everything up to here treated the settings as free variables (§1, §3). An actual apparatus has to do something on each run, and what it does turns out to be constrained in ways the algebra never hinted at. This section supplies the protocol the derivation deliberately omitted.

Agreed in advance vs. chosen at runtime. The first division to get right is what the two parties fix beforehand and what they must not.

  • Fixed at design time (shared): the menu of two settings each party will choose between ( for Alice; for Bob), and a common reference frame — since depends only on the relative angle between settings, the wings must share a definition of "angle zero," or the relative angles go uncontrolled. (Not undefined: without a shared frame you still measure four perfectly well-defined correlators — you simply no longer know which four, and cannot tune them to the CHSH-optimal configuration.) This particular pre-agreement is harmless, because a shared frame is fixed independently of and so leaves measurement independence untouched (§13).
  • Chosen at runtime (independent, never coordinated): which of the two settings is used on each run. These picks must be made randomly and independently, ideally so late that Alice's and Bob's choices are themselves spacelike separated. Pre-agreeing a schedule of choices, or signalling them across during the run, would reopen the locality and free-choice loopholes the test exists to close — which is why real experiments drive the settings with independent fast random-number generators (or starlight) at each station. Just how badly a pre-agreed schedule fails is worked out below, and again in §13.

The correlators are then assembled after the runs, by sorting the paired records according to which setting combination occurred. Neither nor even a single correlator is read off a dial — both are statistics reconstructed from raw detector counts in three layers.

1. Primitive data — coincidence counts. Each run records only which setting was active on each side and which detector fired: a setting pair chosen randomly, and an outcome pair . Tallying many runs gives the coincidence counts for the four outcome combinations . In a two-channel analyzer (a polarizing beam-splitter, or a Stern–Gerlach magnet with a detector on each port) the values are literally separate detectors, so these are just click totals; matching Alice's click to Bob's uses time-tagging within a coincidence window — a shared down-conversion pair arrives within nanoseconds, the crack the coincidence-time loophole exploits (§14).

2. Each correlator is (agree − disagree)/total. For a fixed setting pair the correlator is estimated by the sample mean of the product :

Being a ratio, is insensitive to the overall pair-production rate and — in a two-channel setup — to symmetric losses; it is the empirical stand-in for the quantum expectation .

3. Combine four disjoint sub-ensembles into . The four setting pairs are measured on different runs — one cannot measure and on the same pair, since Alice used a different setting — so

is assembled from four disjoint sub-ensembles. Combining them into one bound requires that the four sub-ensembles share the same effective hidden-state distribution — an assumption about unobservables that identical source preparation cannot certify. It has two teeth: measurement independence (§13) — that the carried on each run is uncorrelated with the setting pair chosen, which a superdeterministic model violates while keeping a perfectly identical source; and fair sampling (§14) — that the detected sub-ensemble faithfully represents the prepared one, which fails if detection efficiency depends on . Preparing the particles identically and interleaving the four settings randomly (to defeat drift) is necessary but not sufficient; the residual gap is exactly what the loopholes below quantify.

Why not measure the four correlators in blocks? One might ask: why switch settings randomly run-to-run, rather than measure all runs first, then all , and sum at the end? If nature is quantum-mechanical and non-conspiratorial, both protocols return the same — which is precisely why the question feels innocent. But predicting the same number is not the same as carrying the same evidential weight, and here the two come apart completely: a Bell test is not a measurement of a quantum prediction, it is an attempt to exclude a class of rival theories, and blocking readmits the entire class.

The foundational reason to interleave is therefore specific to Bell tests: each setting must be chosen unpredictably, at the last moment, and spacelike-separated from the far wing. Any pre-committed schedule — block or round-robin — is fixed at both stations from the outset, which puts the distant setting into the common past, where a perfectly local device can simply read it. The lookup-table construction of §13 shows how total the damage is: given a known schedule and a shared random seed, a local model reproduces any no-signalling correlation, up to . Note that no superdeterminism is needed for this — plain foreknowledge of the schedule suffices — and the freedom-of-choice (§13), locality, and memory (§14) loopholes all reopen at once. Fresh per-run randomness is also what the martingale statistics below require, since those bounds assume each setting is unpredictable given the past.

There is a second, non-foundational reason that is just ordinary experimental hygiene shared with every comparative measurement: real sources and detectors drift (temperature, alignment, efficiency), and blocking correlates that drift with the setting choice, so the four 's sample different distributions. This is the standard randomization-against-drift principle (cf. lock-in detection, ABBA metrology), not anything peculiar to Bell — its only Bell-specific sting is that here the drift can mimic the quantum signal itself rather than merely blur it. Interleaving makes all four correlators sample the same time-averaged source.

The two reasons cash out differently, and it is worth not overstating the first. Against drift, interleaving genuinely enforces what was previously an assumption: all four correlators demonstrably sample the same time-averaged source. Against measurement dependence it enforces nothing — MI cannot be established by any experiment (§13). What randomization achieves is narrower but still decisive: it closes off every route by which a locally accessible past could have carried the setting, leaving only a conspiracy arranged in the deep common past — which is exactly the residue the cosmic Bell tests then push back billions of years.

A worked case: the blocked protocol. Make it concrete. Run the four setting pairs as four consecutive blocks of runs each — , then , then , then — changing an angle only three times in the entire session, on a source nobody otherwise touches. (The even milder version, a single switch mid-session, is broken in exactly the same way; it simply cannot deliver on its own, since CHSH needs all four correlators.) Five things break at once:

  1. The setting becomes a function of the run index. "Block fixes " is a public, deterministic rule, and both stations count runs against a shared clock. Alice's apparatus can therefore compute , and the lookup-table attack of §13 applies verbatim: a local model with a shared seed reproduces anything up to . (What the hidden variable actually is in that model is spelled out just below.)
  2. "The same ensemble" is not on offer. Each pair is measured once and destroyed — the point of step 3 above — so each block consists of different pairs, drawn later. "Same ensemble" is silently the claim , and no amount of not-touching-the-source certifies it.
  3. Drift becomes perfectly aliased onto the setting. Block runs at time . Any monotone wander in pump power, crystal temperature, alignment, or detector efficiency is now exactly correlated with the setting pair — the one correlation MI forbids. If efficiency also depends on , this by itself can manufacture from a strictly local source.
  4. The -value stops being computable. Martingale bounds require each setting to be unpredictable given the past; a block schedule is the most predictable schedule there is, so the statistics the loophole-free experiments rely on simply do not apply.
  5. The choice sits in the common past. You picked each block's angles before those pairs were even emitted, putting the choice event inside the past light cone of both measurements — the exact opposite of switching while the particles are in flight.

What plays the role of here? Worth pinning down, because the hidden variable is emphatically not where the cheating happens. Take : a single number drawn uniformly on at the source and delivered to both wings, fresh each run and independent across runs — the most innocent classical common cause a local realist could ask for. On a run in the block the singlet's target joint distribution at is and . Both devices partition into those four intervals in the same agreed order, see which interval contains , and each reports its own component of the resulting pair. Same , same partition, so the same — distributed exactly as the singlet, giving block by block and .

So is a fair coin, carrying nothing. What transports the illicit information is the response function — a standing lookup table installed at design time and indexed by run number — not the state the pair carries. This is also why the blame is a bookkeeping choice (§13): leave the table in the device's program and is genuinely independent of the settings, so MI holds and PI fails, since Alice's response is explicitly ; absorb the table and run index into instead and the response becomes formally , so PI holds while now determines the settings and MI fails. Same machine, same fraud, two descriptions.

And the sting is that none of this looks like failure. If nature is quantum, non-conspiratorial, and the rig is stable, this protocol still returns . The defect is not that the number comes out wrong; it is that the right number has stopped excluding anything — which is precisely the difference between predicting a value and testing a theory.

What "confirming " means. Each is a ratio of near-Poissonian counts, so carries a statistical error ; one accumulates coincidences until the observed stands many standard deviations above . The loophole-free experiments avoid even the i.i.d. assumption, reporting a -value under the local-realist null hypothesis via martingale-based tests (Gill) — Delft 2015 with , Vienna/NIST with . (The single-channel analyzers of the earliest tests lack a "" port and instead use the Clauser–Horne form with a fair-sampling assumption, since superseded by two-channel detection above the Eberhard efficiency , §14.)

7. Tsirelson's bound: quantum mechanics is nonlocal but not maximally so

Quantum mechanics violates CHSH, but only up to a ceiling. Promote the outcomes to -valued observables (Hermitian, , and since Alice's and Bob's operators act on different factors). For the operator a short computation gives the sum-of-squares identity

(Write ; the two squares contribute , and the cross terms collapse to . The familiar textbook form differs only because it places the minus sign on the term instead.)

Each commutator of two observables has norm , so , whence

This is Tsirelson's bound (1980), saturated by the singlet with the settings of §4. Its philosophical interest is sharp: the purely relativistic constraint of no-signalling (§12) by itself permits as large as — the hypothetical Popescu–Rohrlich (PR) box.

What a PR box is (Popescu & Rohrlich 1994). One box per wing, each taking a bit in and returning a bit out: inputs , outputs , obeying with each output marginally uniform. The wings therefore agree on three of the four input pairs and disagree on the fourth, which makes every CHSH correlator with exactly the signs that add — so , the algebraic ceiling, since each of the four terms is bounded by . Yet it cannot signal: each output is a fair coin whatever the distant input, so the correlation surfaces only when the two records are later compared. No quantum state reproduces it, and no one has built one; it is a conceptual device for asking why nature stops where it does.

The three sets therefore nest strictly,

so nature is nonlocal, yet strictly less nonlocal than causality alone would allow — and no-signalling by itself does not single out quantum mechanics. Why it stops exactly at has become a research programme in the reconstruction of quantum theory from information-theoretic principles (e.g. information causality, Pawłowski et al. 2009, which recovers precisely the Tsirelson bound; van Dam's result that PR boxes would collapse communication complexity to a single bit is a second reason to doubt they exist), and a clue that the Hilbert-space formalism encodes a principle we have not yet named.


Part II — The philosophical analysis

8. What, exactly, has been proved?

State the logic carefully, because loose statements of "what Bell proved" cause most of the confusion. The demonstrable core is a conditional:

Experiment finds . By modus tollens, at least one antecedent is false. The theorem itself is a piece of mathematics and is not in dispute; all the philosophy lives in the choice of which antecedent to reject — and each choice is a different picture of the world. The remainder of this remark is a tour of those antecedents and the positions that deny each.

The full ledger of premises — gathered from §2, §3, and §11 into one place — is:

AssumptionFormal statementMeaningNeeded byRejected by
Realism / hidden variables with , fixing the outcome responses outcomes trace to a state carried from the source (locality is not yet assumed — that is the PI/OI rows)both formsCopenhagen, QBism (§9, §15)
Parameter Independence (PI)local marginal is independent of the distant settingbothBohmian mechanics (§11, §15)
Outcome Independence (OI)no residual outcome–outcome correlation once is fixedbothcollapse / orthodox QM (§11, §15)
PI OI = FactorizabilityBell local causality: screens off the wingsboth— (the conjunction)
Measurement Independence (MI)settings are chosen independently of bothsuperdeterminism, retrocausality (§13)
Single definite outcomeeach run yields exactly one and one there is one fact of the matter per run for to average overbothmany-worlds (§15)
Perfect anticorrelation (PA) for all singlet certainty on aligned settings1964 form only— (CHSH needs it not)

Two entries deserve emphasis. Determinism is absent from the list: it is not assumed but derived from factorizability via Fine's theorem (§2), so "give up determinism" is not among the escape routes. And PA is required by the 1964 inequality but not by CHSH (§3) — which is exactly why every modern, loophole-tolerant experiment reports the CHSH quantity . Everything below dismantles these premises one at a time.

9. The realism assumption and a common oversimplification

Textbooks routinely say Bell refuted "local realism," presenting realism (or "hidden variables," or "definiteness") and locality as two separable premises, either of which might be dropped — and inviting the comfortable conclusion that we may keep locality by merely abandoning naïve realism about unmeasured quantities. This gloss is at best incomplete, and Bell himself rejected it.

The subtlety is that determinism/definiteness is not an independent postulate of the argument — it can be derived. Bell's own two-part reasoning (following EPR) runs:

  1. EPR step. For the singlet, whenever Alice and Bob measure along the same axis they get perfectly anticorrelated results with certainty. If the world is local, Alice's distant choice cannot influence Bob's system; so Bob's definite outcome must have been fixed in advance by something carried locally — a hidden variable . Locality + perfect correlations predetermined values. Determinism is a conclusion here, not an assumption.
  2. Bell step. Those predetermined local values obey the inequality (§3).
  3. Experiment. The inequality is violated.

On this reading the only premise standing between "local" and the false inequality is locality itself (given the empirical perfect correlations), so the honest conclusion is not "local realism is false" but "locality is false, full stop." This is the position of Bell's mature essay La nouvelle cuisine and of its contemporary defenders (Maudlin, Norsen): the world is not locally causal, and no retreat to anti-realism rescues locality.

One premise this reading still needs. "Locality is the only premise" is shorthand for "the only premise besides measurement independence." The EPR step reasons counterfactually — had Alice measured along a different axis, Bob's system would have been unaffected — which presupposes that the axis could vary independently of (§13). No reading of Bell's theorem, however Bell-orthodox, escapes that assumption.

The opposing camp notes that the CHSH derivation of §3 needs only factorizability and never invokes the perfect correlations that power the EPR step; factorizability can be motivated as "locality a separate assumption that fixes outcome probabilities" (a mild realism). On this view "realism" is a genuine, droppable premise, and interpretations that deny observer-independent outcomes (§16) exploit exactly that. Both readings are internally coherent; the disagreement is about which minimal set of premises best regiments the physics. The safe, neutral statement — the one to actually remember — is:

The world cannot be both local and such that measurements reveal pre-existing, locally-determined values. At least one of those must go.

10. Locality, decomposed: Bell locality as a screening-off condition

To see what "locality" is doing, connect factorization (3) to a principle from the philosophy of causation. Reichenbach's common cause principle says that a correlation between two events that do not cause one another must be due to a common cause in their shared past that screens off the correlation — renders the events statistically independent once the common cause is specified. Bell's factorizability is Reichenbach's screening-off, applied to the complete common cause lying in the past light cone of both measurements.

Two qualifications the slogan hides. First, factorizability is strictly stronger than Reichenbach's principle as usually stated. Reichenbach demands a screener for each correlation taken singly; Bell demands that one and the same screen off all four setting pairs at once — a common common cause. The difference is not academic: Hofer-Szabó, Rédei and Szabó showed that correlation-by-correlation common causes for the quantum correlations do exist, so what the experiments refute is the demand for a single, shared screener, not Reichenbach's principle in its weakest form. Second, deriving screening-off from the causal picture also requires measurement independence (§13); without it is not purely a common cause of the outcomes but is itself correlated with the setting choices, and the Reichenbachian reading of factorizability lapses.

With those in hand, the violation of CHSH says:

  • Either there is no common cause that screens off the correlations — so Reichenbach's principle fails for quantum systems (Van Fraassen's and Cartwright's reading), and we must accept brute spacelike correlations with no local explanation,
  • Or there is a genuine non-local influence — a direct causal link between the wings, an "action at a distance" of the sort EPR found intolerable.

Either horn abandons the classical picture of a world knit together only through its past (and both presuppose measurement independence — denying that is a third way out, §13). This is the deepest content of the theorem: the common-cause structure of spacetime — the assumption that spacelike correlations bottom out in a single screener in the shared past — does not survive quantum mechanics.

11. Jarrett–Shimony: parameter independence vs. outcome independence

The single most illuminating philosophical move is to split factorizability into two logically independent conditions (Jarrett 1984; Shimony). Factorization (3) is equivalent to the conjunction of:

  • Parameter Independence (PI) — Alice's outcome statistics do not depend on Bob's setting: . (Jarrett's "locality.")
  • Outcome Independence (OI) — given the settings and , Alice's outcome is statistically independent of Bob's outcome: . (Jarrett's "completeness.")

Now the crucial fact: quantum mechanics violates OI but satisfies PI. The entangled correlations couple the two outcomes (learning Bob's result changes the distribution of Alice's), yet neither party's marginal statistics depend on the distant party's setting choice. This asymmetry is the linchpin of the whole philosophy of the subject:

Violated by QM?Consequence if violated
Parameter IndependenceNoWould permit superluminal signalling — controllable, usable communication
Outcome IndependenceYesOnly uncontrollable correlation — no signalling

Whose OI, whose PI? "QM violates OI, satisfies PI" is the orthodox reading, where the quantum state is the complete state (). Bell refutes only the conjunction, so which conjunct fails is interpretation-relative: a deterministic hidden-variable theory (Bohmian mechanics) instead satisfies OI and violates PI — with the outcomes fixed as definite functions of , nothing is left to correlate once and the settings are given (OI holds trivially), and the whole nonlocality lands on PI (the distant setting moves the local outcome). So OI is "the culprit" only in PI-respecting theories; the two camps are tabulated in §15.

Because PI survives, the correlations cannot be used to send anything: Alice, by her choice of setting, cannot alter what Bob sees. Shimony christened the residual, PI- respecting nonlocality "passion at a distance" as opposed to "action at a distance." The world is nonlocal in its correlations but not in any signal — which is exactly the loophole through which quantum theory and relativity coexist.

PI and OI presuppose realism. A structural point worth stating explicitly: both conditions are defined relative to a complete state — they are constraints on — so the very whose existence is the realism premise (§8) is logically prior to them. Two consequences follow. First, one cannot "reject realism yet accept both PI and OI": holding both is just factorizability, which forces and is refuted; and rejecting the value-framework wholesale (strong Copenhagen, or Everett denying single outcomes) leaves PI/OI with no referent — they become inapplicable, not merely false. Second, there is a thin middle road — "reject only the extra hidden variables and take " (orthodox QM): then PI/OI are definable, and orthodox QM keeps PI ( no-signalling) while dropping OI (the entanglement of is the OI violation). So the escape "no hidden variables" always resolves into either keep PI, drop OI or dissolve the framework — never keep both.

Is the split legitimate? Bell himself, and after him Norsen and Maudlin, resisted the Jarrett reading — objecting to the labels, not the algebra. Calling PI "locality" and OI "completeness" insinuates that only PI is a genuine causal-locality condition and that violating OI is a merely formal shortcoming of the description. But OI follows from local causality just as directly: if really is a complete specification of the shared past, then a residual dependence of Alice's outcome on Bob's spacelike-separated outcome is a nonlocal influence, full stop. On this view Bell locality is one condition, not two, and "QM violates only OI" understates the damage — what fails is local causality, and the PI/OI bookkeeping merely records that the failure happens to be uncontrollable. The split is still the right tool for classifying interpretations (§15) and for seeing why relativity is not immediately contradicted (§12); it should be read as a partition of consequences, not as the discovery that locality was two independent assumptions all along.

12. No-signalling and the "peaceful coexistence" with relativity

That PI holds is the theorem of no-signalling: Bob's reduced state is completely unaffected by Alice's choice of measurement, so no marginal — nothing Bob can access locally — carries information about Alice's setting. The nonlocal correlations become visible only after the two records are brought together over an ordinary, luminally-limited classical channel. Hence:

  • At the operational/relativistic level there is genuine peace: no experiment exploiting entanglement can send a superluminal message, so special relativity's prohibition on faster-than-light signalling is untouched. (This is the sense in which the block-universe causal structure of relativity survives.)
  • At the ontological level the peace is uneasy. A dynamical collapse correlated across a spacelike interval has no Lorentz-invariant "order" — which measurement happened "first" is frame-dependent — so any story in which one outcome brings about the other seems to need a preferred foliation of spacetime that relativity denies. Bohmian mechanics (§15) bites this bullet with an explicit preferred frame; others deny there is any bringing-about to order. The tension is not a paradox but a cost that every interpretation must pay somewhere.

The mapping between the two conditions of §11 and the two grades of locality is exact:

Condition (§11)Grade of localityRequired by SR?Status in QM
Parameter Independencesignal locality — no controllable superluminal communicationYes — violating it would let a setting choice send a message, and in some frame that message runs backward in time (causal loop)holds (no-signalling)
Outcome IndependenceBell locality's surplus — no spacelike outcome–outcome influence beyond common causesNo — its failure carries no signal (outcomes are uncontrollable), so SR's operational prohibition is untouchedfails (this is the entanglement)
PI OI (factorizability)full Bell locality / local causalitystronger than SR demandsfails

The middle row is the whole story of the coexistence: quantum mechanics gives up only the grade of locality that special relativity never asked for. In relativistic QFT the surviving PI reappears as microcausality — spacelike-separated local observables commute, for spacelike — the field-theoretic encoding of no-signalling. (For a bosonic field this is the familiar ; fermionic fields anticommute at spacelike separation, and it is precisely because physical observables are built from them in even numbers that the observables still commute. The condition is on the observable algebra, not on the fields.)

The moral is a careful distinction philosophy of physics insists upon: signal locality (no superluminal communication) is empirically secure; Bell locality / local causality (no superluminal influence of any kind) is refuted. The two come apart precisely because of the PI/OI split.

13. The measurement-independence assumption: superdeterminism and retrocausality

Buried in the derivation is a premise so natural it is easy to miss: — the hidden state is statistically independent of which settings get chosen. This is measurement independence (also "free choice," "no-conspiracy," "-independence"). It is often called the free-will assumption, because it is what licenses treating the experimenters' setting choices as freely (or at least genuinely randomly) made rather than pre-arranged to match — though no metaphysical libertarian free will is needed, only effective statistical independence. Deny it and the derivation collapses: if can be correlated with the future settings, a purely local model can fake any correlation whatever, because the source "knows" in advance how the detectors will be set.

Free choice secures MI; it is not the same as MI. Separate the two jobs the settings do. Having alternatives at each wing is what gives the theorem its grip at all (§2); merely measuring then asks no more than that the ensemble take all four setting pairs — a fixed round-robin suffices. MI is the further condition that which pair a run uses carries no information about . Free, random, ideally spacelike-separated choice is how one secures MI against a predictable schedule — but it is a guarantor, not a constituent: where MI already holds as a fact, and the far setting is not locally knowable, the choosing adds nothing. Call an ensemble strong-ideal when four conditions hold together:

  1. Stationary source — every pair drawn from one fixed , unchanging across the run sequence;
  2. Fair detection — the detected sub-ensemble faithfully represents the prepared one, i.e. efficiency does not depend on (§14);
  3. MI as a fact is genuinely independent of the settings, , holding of its own accord rather than by enforcement;
  4. No local access to the distant setting — neither wing's apparatus can compute or receive the far setting before it records its outcome.

Against a strong-ideal ensemble the deterministic round-robin already yields the full bound, and no choosing is needed. Condition 4 is why the list is not three, and the reason is worth spelling out, because a pre-agreed schedule does not merely weaken the test — it makes the test vacuous.

The lookup-table attack. Fix any target correlation — the singlet's, or even a PR box with . Give each station (i) the public schedule, from which it reads the run index and hence both settings , and (ii) a shared seed drawn uniformly on at the source, independent of everything else. On run each device computes the same distribution and samples it by inverting the same CDF at the same ; Alice reports her component, Bob reports his. The emitted pairs are then distributed exactly as — with no signal travelling anywhere, and with the seed strictly independent of the settings. A deterministic schedule therefore certifies nothing whatever: it can fake not only but the full no-signalling polytope.

Which premise died? That turns out to be a bookkeeping choice, and seeing why is the point. Keep the schedule outside : then MI holds — the seed really is independent of the settings — and it is PI that fails, since Alice's response is an explicit function of . Now absorb the schedule into , the more natural move since it does lie in the common past: PI and OI both hold formally, because is recoverable from and conditioning on it adds no information — but now determines the settings, so MI fails outright. The physical fact is the same under both descriptions: the far setting was fixed in the common past, and is therefore locally available. Condition 3 does not exclude this, because a public schedule is not correlated with — it is merely known.

The moral is that a real test cannot simply declare MI and run a round-robin. Settings must be generated fresh, late, and unpredictably — from events no local device could have held in advance — which is exactly what fast physical RNGs, quasar light, and human bits supply (§13, §14). This is also what §6's block-versus-interleave note is tracking. Choice earns its keep whenever the ensemble is merely weak-ideal — perfect source and detectors, but conditions 3 and 4 left open — where it secures MI against the one thing a perfect i.i.d. source does not exclude: a superdeterministic common past correlating with a foreseeable schedule. ("i.i.d. source" constrains each stream alone, not the joint dependence of on the settings.)

Why only the setting needs randomizing. This invites an obvious worry. Only the angle is chosen afresh each run; everything else about the apparatus — the source design, the crystal, the detectors, the shared reference frame, even which two angles sit on each menu — is fixed in advance, in the common past, exactly where the lookup-table attack lives. Do those variables not carry the same disease?

They do not, and the reason is sharp. MI is required only of variables that vary run to run and select which correlator a run feeds. Anything held constant across the four sub-ensembles may be correlated with as strongly as you like, because the derivation simply runs conditionally on it: if is such a constant, MI-within- gives , and since never changes, . Conditioning on constants is free. A universe that correlated with the brand of your photodiodes, the lab temperature, or the menu itself would therefore buy a local model nothing — within any fixed menu the four correlators must still glue together (§2). The operative word is constant, not pre-agreed: a design variable is harmless exactly when its distribution is the same across all four sub-ensembles. Anything that shifts in step with the setting — a detector whose efficiency changes as the polarizer rotates, an analyzer that warms as it switches — is no longer a constant but an effective setting-correlated variable, and reintroduces precisely the dependence MI forbids. (Hence the experimental discipline of routing both settings through the same detectors rather than dedicating hardware to each.)

Two things follow, and they are worth having explicitly.

  • The residual selectors just are the loopholes. Settings are not the only run-to-run variable deciding which runs enter which correlator. Whether a pair is detected is a second — the fair-sampling/detection loophole — and which detections are counted as one coincidence is a third — the coincidence-time loophole. The list in §14 is, in effect, the enumeration of the run-to-run selectors that have not been randomized away.
  • It sharpens what superdeterminism has to claim. "Everything is determined" is by itself harmless: a deterministic universe that fixes the hardware, the menu, and even the RNG seeds is perfectly compatible with . The conspiracy must be aimed specifically at the correlation between and the setting selector. That specificity is what critics mean by contrivance — and it is also why the loophole can be quantified at all (below), since there is one well-defined correlation to measure rather than a diffuse "everything affects everything."

Two very different metaphysics exploit this:

  • Superdeterminism ('t Hooft; Hossenfelder). A single deterministic law fixes both the hidden variables and the experimenters' setting choices from common initial conditions, so the required correlation between and is built into the initial state of the universe. Locality is saved — at the price of denying that the settings are freely (or even effectively randomly) chosen. Critics regard this as a conspiracy that would undermine the very possibility of experimental science (if setting choices are secretly correlated with what is measured, no controlled experiment is trustworthy); defenders reply that "conspiracy" begs the question about what independence nature owes us.
  • Retrocausality (transactional and two-state-vector interpretations; Price). The measurement setting influences the hidden state backward along the light cone, so the correlation is causal but time-symmetric rather than superluminal. This can preserve a form of locality-in-spacetime by trading it for causation running both temporal directions — attractive to those who take the time-symmetry of microphysics seriously, unsettling to those who take the asymmetry of causation as basic.

Where the superdeterminist's lookup table lives. The construction in the box above puts the table in the apparatus, but a superdeterminist need not — and standardly does not — put it there. It can sit in the experimenter. A human deliberating over which angle to set is a physical system like any other, its state fixed by the same initial conditions that fix ; on that view the deliberation simply is the table, and no hardware needs to consult anything. This is why the BIG Bell Test's hundred thousand volunteers strengthen the case without closing it: nothing about a choice being felt as free makes it statistically independent of anything. Two consequences, cutting in opposite directions.

  • It costs far more than the apparatus version. The lookup-table attack is devastating precisely because it is cheap: the schedule is a public record sitting in the common past, and reading it demands no fine-tuning whatever. Relocated into a nervous system the correlation cannot be read — nobody, the experimenter included, knows the setting before the choice — so it must instead be pre-established in the initial conditions and preserved intact through the entire causal history of a brain. The move keeps the attack's logical sufficiency and throws away its mechanism.
  • But the conspiracy must be multiply realized. This is the real force of the experimental programme, and it is stronger than "we picked a setting source that is hard to conspire with." Modern tests drive the settings from causally independent systems: thermal noise in a diode, human volunteers, and light from two unrelated quasars emitted billions of years ago. A superdeterministic correlation with would have to hold simultaneously in all of them — the same -dependence realized in neurons, in semiconductor junctions, and in the emission statistics of two high-redshift objects — forcing the arrangement back to at least the earlier quasar's emission epoch. It is not that any one source is unconspirable; it is that the conspiracy must be present in every one of them at once.

So the link to the metaphysics of agency is structural, not decorative. Conway and Kochen's Free Will Theorem states one direction of the dependence sharply: if the experimenters' choices are not functions of the information available in their past light cones, then neither are the particles' responses — indeterminism at the bottom follows from freedom at the top. The converse does not hold, and the slogan invites the error: rejecting measurement independence only blocks the inference from to "local causality fails," which leaves a local model permitted rather than established.

Because measurement independence is an assumption, it cannot be proved, only made implausible: experiments have pushed the setting choices to sources maximally decorrelated from any plausible common past — human free choices (the BIG Bell Test, 2018) and the light of distant quasars emitted billions of years ago (cosmic Bell tests, 2018) — shrinking the space in which a non-conspiratorial correlation could hide.

It is a dial, not a switch. How much independence must fail is quantifiable, and the answer is sobering: a local deterministic model can reproduce the singlet correlations exactly while relaxing measurement independence by only a fraction of a bit of mutual information between and the setting pair (Hall 2010, 2011), and above a modest threshold of setting knowledge every no-signalling correlation — up to the PR box — becomes locally simulable (Barrett & Gisin 2011). The rhetorical charge of "conspiracy" therefore overstates what the loophole costs in information: the required correlation is small, not cosmically fine-tuned. What it costs instead is the reliability of randomized experiment in general — which is why most physicists judge the price prohibitive even though the quantitative demand is mild.

The connection to the metaphysics of agency is direct, and links this remark to the free-will debate: what the experimenter's "free choice" of setting is becomes a load-bearing physical assumption.

14. The loopholes: how an experiment can fail to close the argument

A raw violation of refutes local hidden variables only if the experiment actually instantiates the derivation's premises. Each way it might fail to is a loophole — a local-realist escape route kept open by an experimental imperfection. The history of the subject is the history of closing them.

  • Locality (communication) loophole. If the two wings are not spacelike separated during the critical interval, a subluminal influence — Alice's setting or outcome propagating to Bob — could coordinate the results locally. Closing it imposes a strict timing budget: at each wing the setting choice, the measurement, and the recording of the outcome must all finish within the time light needs to cross the wing separation . Aspect (1982) first switched settings during the photons' flight with acousto-optic modulators; Weihs et al. (1998) used independent physical random-number generators and fast electro-optic modulators across m (a s budget); the Delft 2015 test separated its two nitrogen-vacancy centres by km.
  • Detection (fair-sampling) loophole. If detectors register only a minority of pairs, the detected sub-ensemble can violate CHSH while the full prepared ensemble obeys it. The explicit local strategy: let decide not only the outcomes but whether the particle fires at all for a given setting, so the detected sample is precisely the biased sub-ensemble that fakes a violation — the concrete failure of the "detected prepared" assumption flagged in §6. Closing it needs the total system efficiency above threshold: for maximally entangled states, relaxable to using non-maximally entangled states (Eberhard 1993) at the cost of a smaller violation — the trick the 2015 photon experiments exploited. Trapped ions, atoms, superconducting qubits, and NV centres reach near-unit efficiency naturally; only photons had to fight for it.
  • Freedom-of-choice loophole. The experimental face of measurement independence (§13): if the setting choices share any common cause with , the four sub-ensembles need not share a -distribution and a local model fakes the violation. It cannot be closed, only made implausible by decorrelating the choices from any plausible common past — the cosmic Bell tests drove the settings with light from Milky-Way stars (2017) and then high-redshift quasars (2018), forcing any conspiracy to have been arranged billions of years ago, while the BIG Bell Test (2018) used random bits from human volunteers. Each shrinks, but cannot eliminate, the hiding space.
  • Memory loophole. If trials are not independent and identically distributed — the apparatus "remembers" past settings and outcomes — a local model could bias the running statistics. Defused by hypothesis tests valid without the i.i.d. assumption (martingale / Gill bounds), which is exactly why the loophole-free experiments quote a -value rather than a Gaussian error bar (§6).
  • Coincidence-time loophole. When a coincidence is defined by a window measured relative to the detections themselves, a local model can shift detection times as a function of to manufacture spurious coincidences. Closed by fixed, externally-clocked time slots (or a pulsed source) in place of self-referential windows.
  • Collapse-locality loophole (Kent). The premise that each measurement is complete — irreversibly recorded — while the wings are still spacelike separated. Were collapse deferred to some later stage (amplification, or even a conscious observer), the correlated events would not truly be spacelike. Addressed by pushing the irreversible amplification and recording out to spacelike separation.

The loophole-free experiments of 2015 — Hensen et al. (Delft, entangled electron spins in nitrogen-vacancy centres via event-ready entanglement swapping), Giustina et al. (Vienna) and Shalm et al. (NIST), both with high-efficiency photons — closed the locality and detection loopholes simultaneously for the first time. The 2022 Nobel Prize to Clauser, Aspect, and Zeilinger recognized this arc. What remains is only the freedom-of-choice/superdeterminism loophole, which is not so much an experimental defect as a metaphysical stance about whether nature grants us independent choices at all (§13). Within any framework that grants ordinary statistical independence of setting choices, local hidden-variable theories are dead.

15. The menu of interpretations, sorted by which premise they reject

Every serious interpretation of quantum mechanics can be located by which antecedent of the Bell conditional (§8) it sacrifices. This is the most useful map the theorem provides, and it connects directly to the measurement problem. The four "give up X" options of the introduction correspond one-to-one to the premises isolated in §8–§11:

Premise sacrificedInterpretation(s)What it keepsWhat it costs
Locality — Bell local causality (§10, §12)Bohmian mechanics; spontaneous collapse (GRW, CSL)definite outcomes and a single world (and determinism, but only for Bohm — GRW/CSL are irreducibly stochastic)avowedly nonlocal dynamics; a preferred foliation straining Lorentz invariance
Outcome definiteness — "realism," pre-existing values (§9)Copenhagen, QBism, relational QMlocality (PI intact, §11)no observer-independent values; the wavefunction becomes knowledge/belief, not a thing
Single outcomesMany-worlds (Everett)locality and determinisma branching-multiverse ontology; recovering the Born-rule probabilities
Measurement independence — "free choice" (§13)Superdeterminism; retrocausalitylocality and definitenessdenies free / effectively-random settings — a cosmic "conspiracy," or backward causation

A few nuances the grid flattens. Bohmian mechanics wears its nonlocality on its sleeve — its guidance equation makes each particle's motion depend instantaneously on the distant particle's position — and its virtue is precisely that honesty: it shows the deterministic, realist completion EPR wanted is possible, but only if nonlocal. Many-worlds claims to buy locality back: with every outcome realized on a separate branch there is no single result- pair for an influence to correlate, and the branching spreads no faster than light via decoherence. QBism dissolves the nonlocality into information updating rather than physical influence, reading as an agent's belief state.

Two entries in the "costs" column are softer than the grid suggests. The preferred foliation is not compulsory: Tumulka's relativistic flash-GRW model (2006) is a fully Lorentz-invariant collapse theory reproducing the Bell correlations without any distinguished frame, and Bohmian programmes have been formulated with the foliation determined by the wave function rather than posited as extra structure (Dürr, Goldstein, Münch-Berndl & Zanghì) — Lorentz invariance is a severe strain, not a proven impossibility. And Many-worlds' claim to locality is itself contested: recovering the appearance of a definite correlated pair still requires an account of when and how Alice's and Bob's branches "join up," and critics argue the locality is bought by relocating the problem into the branch-identification story rather than dissolving it.

Refining the "Locality" row: PI vs. OI. The first row is deliberately coarse. By the Jarrett–Shimony split (§11), Bell locality is the conjunction PI OI, and violating the inequality means dropping at least one half — but which half is interpretation-dependent, and it divides the nonlocal theories into two camps:

Nonlocal theoryOIPIWhy
Deterministic hidden variables (Bohmian mechanics)satisfied (trivially)violatedwith definite values, once and both settings are fixed the outcomes are determined, so no residual outcome–outcome correlation remains to break (OI holds); the nonlocality must live in PI — the distant setting genuinely affects the local outcome
Stochastic collapse / orthodox QM (GRW, CSL, Copenhagen)violatedsatisfiedthe local marginals are independent of the distant setting (no-signalling PI), yet learning the distant outcome shifts the local distribution

Both camps are nonlocal, but they break different halves: Bohm sacrifices Parameter Independence; collapse/orthodox theories sacrifice Outcome Independence. Crucially, both preserve signal-locality (§12) — Bohm's PI-violation never surfaces in the statistics under quantum equilibrium, so it cannot be used to send a message. One caveat: PI and OI presuppose a joint distribution — i.e. that there are outcomes to correlate — so the split applies to the "Locality" row only; the outcome-definiteness and single-outcome rows reject that framework upstream.

No option is free. The theorem does not select an interpretation; it taxes every one, and the philosophy of quantum mechanics is largely the accountancy of these taxes.

16. Beyond CHSH: sharper and broader forms

Three extensions deepen the philosophical picture without changing its moral.

  • Nonlocality without inequalities. For three or more particles, the Greenberger–Horne–Zeilinger (GHZ) state yields an all-or-nothing contradiction: local hidden variables must assign values that satisfy a set of equations with no solution, so a single ideal run — not a statistical margin — refutes local realism. Hardy's two-particle construction achieves something similar for a fraction of runs. These show the Bell phenomenon is not an artifact of statistics but a flat logical incompatibility.
  • Contextuality (Kochen–Specker). Bell nonlocality is a special case of a more general impossibility: no assignment of pre-existing values to all observables can be non-contextual (independent of which compatible observables are co-measured). Kochen–Specker needs no spatial separation at all — it is a constraint on hidden variables within a single system. Locality is then seen as spatial non- contextuality, and Bell's theorem as its most physically dramatic instance. (One caveat: Kochen–Specker requires a Hilbert space of dimension — for a single qubit a non-contextual value assignment does exist, as Bell noted in 1966. Bell's theorem gets its bite from separation instead, which is why it works for two qubits.)
  • A hierarchy, not a dichotomy. "Entangled" and "Bell-nonlocal" are not synonyms. Entanglement, EPR-steering, and Bell nonlocality form three strictly nested classes (Wiseman, Jones & Doherty 2007): every Bell-nonlocal state is steerable, every steerable state is entangled, and both inclusions are strict — Werner (1989) exhibited entangled states that admit an explicit local hidden-variable model for all projective measurements, and so violate no Bell inequality at all. The metaphysical conclusions of §8 therefore attach to the top of the hierarchy only; entanglement as such does not refute local causality.

17. What the theorem does and does not license

A closing ledger, to guard against the two opposite over-readings.

It does establish:

  • The world is not locally causal: spacelike correlations cannot, in general, be explained by a common cause in the shared past screening off the wings (§10). This is a fact about nature, not about quantum formalism — any successor theory must reproduce it.
  • Therefore one must give up at least one of: locality, outcome-definiteness, single outcomes, or measurement independence (§15). There is no locally-causal, single-world, definite-outcome, free-choice theory of our world.

It does not establish:

  • That superluminal signalling is possible — no-signalling holds, and relativity's operational prohibition is safe (§12).
  • That "realism" simpliciter is refuted — whether realism is even a separable premise is contested (§9); a nonlocal realist theory (Bohm) fits all the data.
  • That the correlations are "spooky" in any usable sense — the nonlocality is passion, not action (§11): visible only after classical comparison.

The positive payoff is not merely negative metaphysics. Because a CHSH violation certifies that no local, predetermined mechanism produced the data, it underwrites device-independent cryptography and certified randomness: the same feature that embarrassed Einstein is now a resource that lets one trust a random number without trusting the device that made it. That is the fitting last word on a theorem that turned a thought-experiment about incompleteness into working technology — and a piece of metaphysics into a laboratory fact.

See also