Helen Brand had just learned she could not get Miles Bron convicted of murdering her twin sister, Andi. Not enough evidence, not in time. She went feral instead, and somewhere in the wreckage she found the sample of Klear, Miles’s unstable prototype fuel, and threw it into the fire. The island mansion, Glass Onion, was made almost entirely of glass. Near enough heat, it started cracking, small sounds first, then the panels gave out. The Mona Lisa hung in the middle of it, on loan from the Louvre. It was the one object in that house Miles had not bought, only borrowed. He stood beside it at parties and let people draw their own conclusion about which one of them would be remembered. “Your fuel of the future just barbecued the most famous painting in the world,” Helen told him, watching it go.

That is the ending of Glass Onion: A Knives Out Mystery. The fire is not the twist. Andi’s death, Helen’s rampage, none of it happens if a courtroom had gotten one thing right, years earlier. Miles Bron was accused of stealing an idea.

The idea was Andi Brand’s. She sketched it once on a napkin at a bar called the Glass Onion, its logo pressed faint into the paper under her pen. That idea became Alpha, the technology company Miles Bron built into the billions. He pushed her out, and the friends who owed him their careers rose in that courtroom and swore, without blinking, that his version of events was the truth. Andi searched everywhere for the napkin and could not find it. By the time the case went to court, all she had was her own word, thin against a story everyone else had already agreed to tell. She lost.

Years after the case ended, Andi found the original napkin. She emailed the friend group a photo of it and warned them what was coming. Miles killed her before she could follow through, and staged it to look like she had done it herself. Helen, trying to find out what had happened to her sister, later found the same napkin hidden in Miles’s house, and held it up in front of everyone at the party. Miles burned it a moment later, a small fire that Helen would answer with a much larger one before the story was over.

That napkin was the proof the courtroom never had. The trial was about deciding whether Miles was innocent of stealing Andi’s idea, given the evidence. Statisticians call an assumption like that the null hypothesis, the default that stands until evidence overturns it: innocent until proven guilty in a courtroom, a taster who is only guessing in the lady tasting tea. The burden of knocking it down fell on Andi, not Miles: his story held unless she could prove otherwise. The complementary statement, that he was not innocent and Andi’s version was true, is called the alternative hypothesis.

The true state of the world is fixed, but it usually sits out of reach, a gap between what is true and what evidence can show. What arrives instead is sample data, a partial, secondhand account standing in for a population nobody gets to inspect in full. The evidence used in that courtroom was verbal, Miles’s word against Andi’s, backed by his friends’ testimony. A stronger form of evidence existed to unseat the null, the very napkin Andi would not find again for years.

Whatever that truth was, however far out of reach it sat, the friend group facing it only ever had two doors: turn on Miles’s story, or let it stand. Statisticians call the first door reject the null and the second fail to reject it. No test and no jury ever gets a third door, a maybe, or a close enough. Rejecting the null takes evidence strong enough to overturn it. Failing to reject only means that the evidence provided is not that strong yet. The friend group took the second door. They let the story stand. Even another trial, had one come, would have found nothing new to rule on without the napkin, unless the friends had changed their story.

You probably have a story standing the way Miles’s did. Maybe it is the one about why a friendship ended, the version you settled on years ago and stopped checking. Nobody proved it. Nobody produced the napkin. It holds because it was the default, and the evidence against it never got strong enough, or never got looked for. Letting a story stand is not the same as finding it true.

Exactly How Wrong

The truth has two possible states: either Alpha was Miles’s idea, or it was not. The courtroom’s verdict went through one of the two doors: it ruled for Miles, or it ruled for Andi. Crossing those two, what was true and what the courtroom decided, produces four possible outcomes. The table below shows all four. Two of them match the truth. Two of them do not.

Court rules for Miles
(Fail to Reject)
Court rules for Andi
(Reject)
Null: Alpha is Miles’s idea Correct call Type I error: the rightful owner loses anyway
Alternative: Alpha is not Miles’s idea Type II error: what the courtroom delivered Correct call

One correct call is the null holding up: the evidence brought against it did not meet the standard needed to overturn it. The other correct call is the null getting overturned, which happens only when the evidence does meet that standard.

Two ways to be wrong. The other two boxes are the errors. Rejecting the null when it is true is a Type I error, a false positive: ruling for Andi when the idea really was Miles’s. Failing to reject the null when it is false is a Type II error, a false negative: ruling for Miles when the idea was Andi’s.

A verdict only reports one thing: reject, or fail to reject. It never reports which of the four boxes was the real one. The film lets its audience see which box was real: Andi’s sketch, Miles’s lie, both shown on screen. A real test does not get that privilege. It only ever touches the sample, never the population, the same gap named earlier. The court’s actual verdict, fail to reject the null, left Miles’s story unrefuted, nothing more.

Deciding how wrong to be. Naming the two errors is one thing. Building a test that holds both in check is another, and that is the problem Jerzy Neyman and Egon Pearson set out to solve.

A framework is a decision structure built on top of a mental model, an internal picture of how something behaves. The model explains the mechanism. The framework says what to do about it, and it does so the same way each time the problem returns. Those are the tests that tell the two apart. Neyman and Pearson’s framework works in two steps. First, fix the maximum tolerable Type I error at a threshold called the level of significance, denoted by alpha. Then, among every test that keeps its Type I error at or below alpha, pick the one with the highest power. Power is the probability of rejecting the null when the null is false, the complement of a Type II error: one minus the chance of letting a false story stand. The order matters, and two friend groups show why. One would never turn on Miles, whatever the evidence. It never rejects, so its Type I error is zero and its power is zero too: perfectly safe, and perfectly useless. The other would always turn on Miles, whatever the evidence. Its power is one, the highest it can be, and its Type I error is one too: it rules for Andi even when the idea was Miles’s. Power is worth nothing when it comes from rejecting all the time. That is why the framework controls the false positives first, capping Type I error at alpha, and only then reaches for the most power it can get.

A threshold cannot keep a test from being wrong. What it promises is exactly how wrong the test will allow itself to be, decided before any verdict.

One null has carried the piece this far: Alpha is Miles’s idea, and the burden to overturn it stayed with Andi. The strongest proof she could have brought was the napkin, and she did not have it. What the court had instead was the friends, all four of them, who swore the idea was Miles’s.

Four Was Never Enough to Ask

Picture ten friends for a moment instead of four, each called to the stand alone, none of them hearing what the others said. Each one answers the same either-or question and nothing more: does the idea belong to Miles or to Andi. No reasons, no account, no follow-up, just the name. Count the friends who say Andi, and that count is the whole of the evidence the test will weigh. Not what any of them argued, not how certain they sounded, only how many said her name. A real courtroom would never reduce testimony to that, where one careful account can outweigh a roomful of nodding.

The null and the alternative from the courtroom carry over, restated as a count. The null is that the friends are no more likely to name Andi than Miles. The alternative is that they lean toward Andi. Neyman and Pearson built their method on having both, because the two steps that follow, fixing an error rate and then choosing a rule, cannot even be stated without the alternative. The alternative is what tells you which tallies would count in Andi’s favor.

Before any friend speaks, the court has to decide what a win for Andi looks like: how many of the ten naming her would be enough to hand her the company. That decision is the rejection rule, and like the bar for passing a bill or carrying an election, it is set before the votes come in, not after. Because the alternative points to Andi, only a lean toward her can count as a win, so the number has to be high. Five of ten is an even split, evidence for neither. Anything below five leans toward Miles. The rejection rule belongs somewhere in the top half. The question is where.

Start with the strictest rules and work down: rule for Andi only if all ten name her, or if at least nine do, or at least eight, or at least seven. Each one gets the same question. Suppose the friends were really a 50/50 split, as likely to say Andi as Miles. How often would a split like that hand up that many Andi-votes or more, by luck alone? That probability is the rule’s Type I error rate.

Candidate Rule Type I Error Rate
Rule for Andi only if all ten friends name her 0.0010
Rule for Andi if at least nine friends name her 0.0107
Rule for Andi if at least eight friends name her 0.0547
Rule for Andi if at least seven friends name her 0.1719

Before reading which of these a statistician would keep, put the line where instinct puts it. Six of ten naming Andi is already a majority, and most people would stop there and call the company hers. But a 50/50 split produces six Andi-votes or more about 38 times in 100, by luck alone. Set the bar at six, and even a fair coin with no lean at all would clear it more than a third of the time, ruling for Andi when nothing pointed to her. That is why six is not on the list, and why the real question is how far above a bare majority the bar has to sit.

This is the first step of Neyman and Pearson’s method: control the Type I error. Fix alpha, the largest Type I error rate you are willing to accept, and set it at 0.05. Alpha is the limit you allow yourself. What a given rule delivers against that limit is its size, and every candidate in the table above has one, its entry in the Type I Error Rate column. A rule stays in the running only if its size falls at or below alpha, with no rounding and no benefit of the doubt. Two of the four qualify: at least ten at 0.0010, and at least nine at 0.0107. At least eight, at 0.0547, is over the line, barely, but barely over is still over. At least seven, at 0.1719, is nowhere close. So the bar cannot fall below nine of the ten naming Andi, three votes past the simple majority of six that intuitively felt like enough.

Among the rules that clear that first step, the second keeps the most powerful. Power is the chance the verdict rules for Andi when the idea truly is hers. To put a number on it, suppose that in that case each friend names her nine times in ten. Here is the power of each rule near the line, beside the Type I error rate from before.

Candidate Rule Type I Error Rate Power
At least eight name Andi 0.0547 0.93
At least nine name Andi 0.0107 0.74
At least ten name Andi 0.0010 0.35

So the second step of the Neyman-Pearson framework picks the optimal test from the rules that survived the first. At 0.05 only two survived, at least nine and at least ten, and of those the more powerful is at least nine, 0.74 against 0.35. Its rejection rule: reject only if at least nine of the ten name Andi. Now double the tolerance to 0.10. At least eight, at 0.0547, clears the higher bar, and it is the most powerful of the survivors, so the optimal rule drops from nine to eight and power climbs from 0.74 to 0.93. Tolerating twice the false-positive risk buys something real here, a lower bar to clear and a better chance of ruling for Andi when the idea for the company truly was hers.

In the courtroom there were not ten friends. There were four: Claire, Lionel, Birdie, and Duke. Build the same kind of rule for four, and the strictest one possible is unanimity, all four naming Andi. Its Type I error rate is 0.0625. Held against the same 0.05 threshold, 0.0625 is over, the way 0.0547 was over for the ten. Unanimity is the most any rejection rule could demand, and unanimity already fails the first step. With four friends there is no rule that keeps the Type I error at or below 0.05, so no rule runs at all. Four testimonies were never enough to count as evidence against the null.

This is what a test with no power looks like: the friend group that would never turn on Miles, now made real. The verdict defaults to fail to reject, however the four friends vote. All five tallies, from none naming Andi to all four, land on Miles. A test that never rejects has a size of zero, and a power of zero with it. For the ten, the rule that ran was at least nine, at 0.0107. For the four, the 0.0625 belongs to a rule that was thrown out.

Loosen alpha to 0.10 and a single rule squeaks in: unanimity, at 0.0625, now falls under the line, while the next rule down, at least three of four, is 0.3125 and stays out. One rule qualifies, so it wins by default, no power step required.

And in the film the four never came down as four anyway. They came down as one. Miles paid the group together, and they answered together: Claire and Lionel matched their stories before taking the stand, and Duke was weighing his paycheck, not the claim. All four named Miles, none named Andi. The 0.0625 assumed something that was not true of them, four friends answering independently, each on their own. That was the mental model under the whole test: a count where every voice is its own. What the arithmetic counted as four voices for Miles was one decision spoken four times, staged to look like agreement. When the responses are not independent, the count is not four pieces of evidence. It is one piece wearing four faces.

So at the 0.05 cutoff, the court could not rule for Andi. It was correct the way a coin with one face is correct. Nothing the four could have said would have changed the verdict, so their testimony was never weighed at all. The verdict came from how small the group was. For the friends to have carried Andi’s case, there would have had to be many more of them, with most of them naming her. Clearing them from the stand would have left Miles’s story with nothing holding it up, and still no proof the idea was hers.

The threshold alpha and the company Alpha share nothing but a name, whether or not the film meant it. Once alpha sets what counts as enough, though, the coincidence starts to carry weight. A rule wins its verdict only when its size falls at or below alpha, and here alpha was the bar the friends’ testimony had to clear before the company named Alpha could change hands. Their testimony never came near it. Only the napkin might have, and it surfaced too late to be entered at all.

A Closing Reflection. Somewhere this week, a call will get made before the truth behind it is fully known. Notice it arriving, not what it decides.

  1. The moment right before you let something stand: what you notice in your own body when you choose the easier read, a loosening in the shoulders, a held breath let out, the beat before you change the subject.
  2. The moment right before you turn on something instead: what tips you, a tone in someone’s voice, a detail that will not sit still, the second doubt turns into a decision.
  3. A story sitting untouched in your life right now, one you have not looked at closely in a while: not the story itself, just where you feel it when it comes up, a tightness, a quickness to change the subject, a readiness with the same old answer.

A verdict never announces which of the four boxes it landed in. That gets found out later, if it gets found out at all, by whoever is still around to notice the napkin when it finally surfaces. There is a version of you that never turns on anything, never risks being the one who was wrong, and it costs something too, quietly, in every true thing it never gets around to catching. That version keeps its own small fire lit in private. Somewhere, it is still waiting to become the larger one.

Where This Framework Came From

Jerzy Neyman and Egon Pearson built this framework together, over more than a decade of joint work, deciding how a test should report what it found. The term null hypothesis is Fisher’s, from the same tea test. The alternative needed to unseat it, alpha fixed before a verdict, the power a test gets once alpha is set, the false positive, the false negative: all of that came out of their work. It is still the framework this piece runs on, and the one science still runs on too.

Intellectual Honesty Note. Glass Onion: A Knives Out Mystery (2022, directed by Rian Johnson) is a real film telling a fictional story. The account of Miles Bron, Andi Brand, and the friend group given here follows its plot. The film presents the painting burned in its ending as the real Mona Lisa, on loan from the Louvre. The production burned a replica, and the real painting never left the Louvre. Null hypothesis testing, the alternative hypothesis, alpha, power, and the four-outcome table crossing a null’s truth against a test’s decision, along with the terms Type I and Type II error, are standard statistical vocabulary, not invented for this piece. The shared name between Miles Bron’s company and the statistical alpha may or may not be a pun the film intended. The connection this piece draws between them, alpha as the threshold the evidence had to clear and Alpha as the company that threshold decided, is this piece’s own reading of the coincidence, not a claim about what the film meant by it. The friend group who testified against Andi in the original lawsuit numbers four in the film: Claire Debella, Lionel Toussaint, Birdie Jay, and Duke Cody. The comparison to ten witnesses is an illustration built for this piece, not part of the film. It treats the friends’ testimony as the evidence for whether the idea was Andi’s, to show how a small, discrete sample limits which significance levels a test can reach. In the illustration each friend answers a bare yes or no on ownership, so a single count can serve as the whole evidence. A real trial turns on the substance of testimony, not a tally, and the film’s case was not decided this way.

References

Johnson, R. (Director). (2022). Glass Onion: A Knives Out Mystery [Film]. Netflix.

Neyman, J., & Pearson, E. S. (1933). On the problem of the most efficient tests of statistical hypotheses. Philosophical Transactions of the Royal Society of London, Series A, 231, 289-337.