Helen Brand had just learned she could not get Miles Bron convicted of murdering her twin sister, Andi. Not enough evidence, not in time. She went feral instead, and somewhere in the wreckage she found the sample of Klear, Miles’s unstable prototype fuel, and threw it into the fire. The island mansion, Glass Onion, was made almost entirely of glass. Near enough heat, it started cracking, small sounds first, then the panels gave out. The Mona Lisa hung in the middle of it, on loan from the Louvre. It was the one object in that house Miles had not bought, only borrowed. He stood beside it at parties and let people draw their own conclusion about which one of them would be remembered. “Your fuel of the future just barbecued the most famous painting in the world,” Helen told him, watching it go.

That is the ending of Glass Onion: A Knives Out Mystery. The fire is not the twist. Andi’s death, Helen’s rampage, none of it happens if a courtroom had gotten one thing right, years earlier. Miles Bron was accused, back then, of stealing an idea that was never his.

Andi Brand had an idea. She sketched it once on a napkin at a bar called the Glass Onion, its logo pressed faint into the paper under her pen. That idea became Alpha, the technology company Miles Bron built into the billions. He pushed her out, and the friends who owed him their careers rose in that courtroom and swore, without blinking, that his version of events was the truth. Andi searched everywhere for the napkin and could not find it. By the time the case went to court, all she had was her own word, thin against a story everyone else had already agreed to tell. She lost.

Years after the case ended, Andi found the original napkin. She emailed the friend group a photo of it, proof she had found it, and warned them what was coming. Miles killed her before she could follow through, and staged it to look like she had done it herself. Helen, trying to find out what had happened to her sister, later found the same napkin hidden in Miles’s house, and held it up in front of everyone at the party. Miles burned it a moment later, a small fire that Helen would answer with a much larger one before the story was over.

That napkin was the proof the trial itself never had. Only one question decided who owned Alpha: did Miles Bron produce the design himself. If he did, the idea was his. If he did not, it belonged to whoever thought of it, Andi. The trial was about deciding whether Miles was innocent of stealing Andi’s idea, given the evidence. Statisticians call an assumption like that the null hypothesis, innocent until proven guilty, the default that stands until something specific knocks it down. The complementary statement, that he was not innocent and Andi’s version was true, is called the alternative hypothesis.

The true state of the world is fixed, but it usually sits out of reach, a gap between what is true and what evidence can show. Statisticians run into that gap constantly. What arrives instead is sample data, a partial, secondhand account standing in for a population nobody gets to inspect in full. A test never touches the truth. It only ever touches what got produced as evidence for it. Everyone without direct access to the truth had only that stand-in to work with. The evidence used in that courtroom was verbal, Miles’s word against Andi’s, backed by his friends’ testimony. A stronger form of evidence existed to unseat the null, the very napkin Andi would not find again for years.

Whatever that truth was, however far out of reach it sat, the friend group facing it only ever had two doors: turn on Miles’s story, or let it stand. Statisticians call the first door reject the null and the second fail to reject it, and no test, no jury, ever gets a third option, never gets to say maybe, or close enough. The friend group took the second door. They let the story stand. Even another trial, had one come, would have found nothing new to rule on without the napkin, unless the friends had changed their story.

That is a test too, weighing what you know against what you would rather believe, and I have run it myself more times than I can count. The same test shows up in my own body too. Persistent soreness the morning after a hard workout can mean two things: an ordinary training adaptation, muscle fibers doing the repair work they are supposed to do, or something further along than a normal ache. Set alpha at five percent. I only stop training and see a physiotherapist when a presentation like that would show up in fewer than five out of a hundred ordinary recovery cycles, genuinely atypical rather than an ordinary ache. Set it lower, at one percent, and I train through almost everything, until a minor strain turns into a real injury. Set it higher, at twenty percent, and I rest at every twinge, and lose the consistency the training was for in the first place. The same fuzzier test runs quietly on a meditation cushion. A stretch of restlessness or low mood during a sit is either ordinary noise passing through, or a signal that something off the cushion, sleep, stress, diet, needs attention.

Exactly How Wrong

The truth has two possible states: either Alpha was Miles’ idea, or it was not. The courtroom’s verdict also has two possible states: it ruled for Miles, or it ruled for Andi. Crossing those two, what was true and what the courtroom decided, produces four possible outcomes. The table below shows all four. Two of them match the truth. Two of them do not.

Court rules for Miles(Fail to Reject) Court rules for Andi(Reject)
Null: Alpha is Miles’ idea Correct call Type I error: the rightful owner loses anyway
Alternative: Alpha is not Miles’ idea Type II error: what the courtroom delivered Correct call

One correct call is the null holding up: whatever evidence was brought against it did not meet the standard needed to overturn it, whether because none was offered or because it was successfully challenged as inadmissible or insufficient. The other correct call is the null getting overturned, and that only happens when evidence does meet that standard. Rejecting the null requires evidence that meets it. Failing to reject it only means nothing has met it yet.

The other two boxes are errors. Wrongly rejecting the null when it is true, ruling for Andi when Alpha really was Miles’ idea, is a Type I error, a false positive. The null was false, Alpha was Andi’s, and the court still failed to reject it, ruling for Miles anyway. That is a Type II error, a false negative.

The court never got the chance to correct that error, because no new trial ever reopened the question. A verdict only reports one thing: reject, or fail to reject. It never reports which of the four boxes was the real one. The film lets its audience see which box was real: Andi’s sketch, Miles’s lie, both shown on screen. A real test does not get that privilege. It only ever touches the sample, never the population, the same gap named earlier. Miles made sure that sample said something false.

The court’s actual verdict, fail to reject, left Miles’s story unrefuted, nothing more. Statisticians call what was missing power.

Neyman and Pearson’s framework works in two steps: fix the Type I error at a threshold, alpha, and then, among every test that keeps its Type I error at or below that threshold, pick the one with the highest power. Consider a friend group that would never turn on Miles, no matter what evidence appeared. Power is the probability of rejecting the null when the null is false, the flip side of Type I error. Both depend on whether the test ever rejects at all. Since this group never rejects, both are zero: their Type I error is zero, which stays within any alpha above zero, but their power is zero too, so this is not an optimal test. Now consider the opposite friend group, one that would always turn on Miles regardless of evidence. This group exceeds the threshold instead, committing an inflated Type I error, since they would reject even when Miles was innocent.

A p-value is the probability of evidence at least this extreme, if the null were true, the same number this article already earned in full. The rejection rule is fixed alongside it: reject the null when the p-value is at or below alpha, fail to reject it otherwise. Both the threshold and the rule have to be set before the evidence arrives, not after.

Imagine, for a moment, a different piece of evidence surfacing in that courtroom, one that worked out to a p-value of 0.08, against an alpha fixed in advance at 0.05. The rule says fail to reject: the evidence falls short, and Miles’s story stands. Cheating looks like raising alpha to 0.10 after already seeing that 0.08, so the same evidence suddenly counts as enough. That is not a stricter test producing a different answer. It is the same weak evidence, with the threshold lowered to meet it after the fact.

A threshold does not promise a person, or a test, will never be wrong. It promises to know, ahead of time, exactly how wrong it is willing to be, before the null is ever rejected.

Four Was Never Enough to Ask

Picture ten friends instead of four, called to the stand one at a time, each answering the same question without hearing what the others said. If the friend group was really as divided as it claimed, no more likely to side with Miles than with Andi, that is the same as ten fair coins. The question becomes how many would need to land on Miles before the pattern stopped looking like chance. Nine out of ten gets you close to the standard cutoff, one in twenty, but not exactly: it works out to 0.0107, not 0.05. Eight out of ten gets you 0.0547, closer to 0.05, but now past it. No number of friends out of ten lands exactly on 0.05. Whoever designs the test has to pick, in advance, which of the two nearest options to use: the stricter rule, requiring nine of ten to agree, conservative at 0.0107, or the looser rule, requiring eight of ten, slightly inflated at 0.0547. Those two numbers straddle 0.05, one just under, one just over.

There were not ten. There were four, Claire, Lionel, Birdie, and Duke, and four is not a rounding error off of ten, it is a different problem. The only outcome that could ever look suspicious with four friends is complete unanimity, all four backing Miles, and the odds of that from four honest coins is 1/16 = 0.0625. That is the most extreme thing four friends could have done, and it still would not have cleared 0.05. There was no version of four honest testimonies that would have forced a rejection. The friend group was not just small enough to look plausible. It was small enough that no evidence it could have produced, short of one of them breaking ranks, was ever going to count as proof of anything.

Even 0.0625 assumes something that was not true: four independent opinions. Claire did not check her own memory without knowing what Lionel planned to say. Duke was not weighing Andi’s claim on its own merits, he was thinking about his paycheck. Miles paid all four of them together, as one group. So what looked, from the math, like four honest people happening to agree was not four separate opinions. It was one decision, repeated four times by four different people, so it would look like agreement instead of one person talking through four mouths. A test needs its trials to be independent for its numbers to mean what they claim to mean. Without that independence, there was no coin being flipped at all. There was only one person’s answer, said four times.

The level of significance, 0.05 fixed in advance, states how much error a test is willing to accept, decided before a single friend takes the stand. The Type I error rate is different, whatever number the actual rejection rule produces once the arithmetic is done, 0.0107 for reject at nine of ten, 0.0625 for reject only at complete unanimity of four. Neither number lands exactly on 0.05. A friend group of ten or a friend group of four, no matter which, cannot produce a rule that does: every rule either size can produce sits to one side of 0.05 or the other, never exactly on it. Statisticians call the number that gets chosen alpha, lowercase, the nearest a real test can get to 0.05 without going over it.

The threshold, alpha, and Alpha the company share only a name, a coincidence the film never intended. But once the number decides what counts as enough evidence, the coincidence starts pulling weight anyway. Whatever evidence rises against a null only wins if it clears alpha, whatever alpha happens to be attached to. In this courtroom, alpha was the number the friend group’s testimony needed to clear before the company named Alpha could change hands. Nothing they said came near it. Only the napkin might have, and it surfaced years too late to be entered as anything at all.

A Closing Reflection. Somewhere this week, a call will get made before the truth behind it is fully known. Notice it arriving, not what it decides.

  1. The moment right before you let something stand: what you notice in your own body when you choose the easier read, a loosening in the shoulders, a held breath let out, the beat before you change the subject.
  2. The moment right before you turn on something instead: what tips you, a tone in someone’s voice, a detail that will not sit still, the second doubt turns into a decision.
  3. A story sitting untouched in your life right now, one you have not looked at closely in a while: not the story itself, just where you feel it when it comes up, a tightness, a quickness to change the subject, a readiness with the same old answer.

A verdict never announces which of the four boxes it landed in. That gets found out later, if it gets found out at all, by whoever is still around to notice the napkin when it finally surfaces. There is a version of you that never turns on anything, never risks being the one who was wrong, and it costs something too, quietly, in every true thing it never gets around to catching. That version keeps its own small fire lit in private. Somewhere, it is still waiting to become the larger one.

Where This Framework Came From

Jerzy Neyman and Egon Pearson built this framework together, over more than a decade of joint work, deciding how a test should report what it found. The null hypothesis, the alternative needed to unseat it, alpha fixed before a verdict, the power a test gets once alpha is set, the false positive, the false negative: all of it came out of that work. It is still the framework this piece runs on, and the one science still runs on too.

Intellectual Honesty Note. Glass Onion: A Knives Out Mystery (2022, directed by Rian Johnson) is a real film; the account of Miles Bron, Andi Brand, and the Disruptors given here follows its actual plot, though it is a fictional story, not a documented real-world event. The Mona Lisa burned in the film’s ending was a museum-grade replica, not the real painting, which remained in the Louvre throughout; the production reportedly documented its destruction on camera, a standard practice when a prop replica is convincing enough to otherwise risk passing as a forgery later. Null hypothesis testing, the alternative hypothesis, alpha, power, and the four-outcome table crossing a null’s truth against a test’s decision, along with the terms Type I and Type II error, are standard statistical vocabulary, not invented for this piece. The shared name between Miles Bron’s company and the statistical alpha is a coincidence of the film’s own naming, not a pun the film intended. The connection this piece draws between them, alpha as the threshold the evidence had to clear and Alpha as the company that threshold decided, is this piece’s own reading of the coincidence, not a claim about what the film meant by it. The friend group who testified against Andi in the original lawsuit numbers four in the film: Claire Debella, Lionel Toussaint, Birdie Jay, and Duke Cody. The comparison to ten witnesses is a constructed illustration built for this piece to show how a small, discrete sample limits which significance levels a test can reach, not part of the film’s own account.

References

Johnson, R. (Director). (2022). Glass Onion: A Knives Out Mystery [Film]. Netflix.

Neyman, J., & Pearson, E. S. (1933). On the problem of the most efficient tests of statistical hypotheses. Philosophical Transactions of the Royal Society of London, Series A, 231, 289-337.