“Roosevelt was my closest friend.” Taft said brokenly, then wept.
It was April 25, 1912. A journalist had found him slumped in his train car after a brutal day of campaigning in Boston, against Theodore Roosevelt, the man who had picked him as his own successor four years earlier. Everyone watching that race knew how this kind of story ends. A friendship breaks, a party splits, somebody loses, in this case badly enough that Woodrow Wilson won the whole election. That is the version of betrayal most people know how to recognize.
There is a quieter version, one that never resolves, no tears anyone thought to record, no losing side, still running inside a method used in every introductory statistics class. Somewhere in that course, you probably ran into Fisher, Neyman, and Pearson, even if you never clocked how much they shaped the field.
Jerzy Neyman stood up to read a paper, and by the time he sat back down, a friendship was coming apart in the room with him. It was March 28, 1935, at the Royal Statistical Society. Neyman’s paper, on the design of agricultural experiments, questioned work belonging to the department Ronald A. Fisher had spent fourteen years building at Rothamsted. Fisher rose during the discussion and told the room it was “clear to everyone present that Dr. Neyman has been somewhat unwise in his choice of topics,” then said flatly, “Dr. Neyman has misunderstood the intention… of the z test and of the Latin Square.” Egon Pearson, Neyman’s collaborator, stood up next. He did not defend the paper’s mathematics. He defended the right to have written it at all, telling Fisher directly that he ought to “beg leave to question the wisdom of accusing a fellow worker of incompetence without… showing that he had succeeded in mastering the argument” himself. The minutes recorded the exchange. Nobody in that room mistook it for an ordinary disagreement.
A few years earlier, William Sealy Gosset, who published his own statistical work under the name “Student” and counted both men as friends, wrote to Fisher trying to arrange a research visit for Neyman at Rothamsted. He called Neyman “fonder of algebra than correlation tables,” and “the only person except yourself I have heard talk about maximum likelihood as if he enjoyed it.” Gosset was not being polite. That was the highest compliment he had, one craftsman recognizing another across a shared obsession. Whatever broke in that room in 1935 had been, only a short while before, warm enough that a mutual friend wanted the two of them working side by side.
If I were Fisher, I would feel blindsided, mostly. I had spent at least a decade building that department. The challenge to it came from someone I had been warm with, someone a mutual friend once thought I should be working alongside, not defending myself against. That is a strange kind of vertigo: the relationship changes shape in public, in real time, and I only find out by watching myself react to it. But the disagreement itself is not the sharper sting. It is Pearson standing up right after, not even engaging the argument, going straight for “you don’t get to talk to him that way.” That is worse than losing on the merits. I came in defending a decade of work. I left having been quietly told, before I had a second to compose myself, that I had behaved badly while doing it.
Almost everyone has felt some version of this: watching someone they trained, mentored, or simply assumed was still on their side take a public shot at the thing they built, and realizing the friendship shifted before they ever got a vote. There is no private version of that conversation. People find out it changed by watching it break in front of witnesses.
However, if I were Neyman, I would feel exposed first. Standing up to read a paper is the vulnerable position. I was the one putting a claim in front of the room, and Fisher was not just another audience member. He ran the field. He does not answer with a counter-argument. First I am told I was unwise to have chosen this topic at all, before I have defended a word of it. Then I am told I misunderstood the very tests I was correcting. Neither line touches my math. One says I should not have spoken. The other says I did not understand what I was speaking about. That is not “you’re wrong.” That is “you didn’t do the reading,” delivered by the one person whose opinion of my competence mattered, in front of everyone whose opinion mattered too.
That kind of dismissal carries a particular unfairness, worse than losing an argument on the merits. If Fisher had said, “here’s where your math breaks,” I could answer it. “You misunderstood the logic” forecloses the conversation before it starts. It is a status move, not a substantive one, and I have no comeback that does not sound defensive. I built this paper because I saw a real gap in a framework I respected. I did not expect to be told I had not earned the right to see it. Then Pearson stands up for me. Relief, probably, but a complicated kind, because now the room has also watched me need rescuing. I did not get to win the argument on my own. Someone else had to win the right for me to have made it at all. I thought I had earned something with Fisher. Gosset thought enough of what I understood to try to put us in a room together. Whatever that was, it did not survive me having a good idea he had not had first.
Roosevelt and Taft’s rift had an ending anyone could point to: Wilson took the election. Taft finished behind even Roosevelt, third in a race he had entered as the sitting president. Fisher and Neyman’s rift never found that kind of ending. Neither man ever conceded the argument. Neither framework lost. The two got run together anyway, the same method holding both without either man’s signature on it. The historian of science Gerd Gigerenzer later traced the stitching back to its two separate sources. Most people who have ever read the phrase “statistically significant” in a headline, watched a poll get called the moment it crossed some invisible threshold, or run the test themselves without asking where the threshold came from, have never heard either statistician’s name, or known there was ever a rupture behind the method they were using. It reads as one settled thing: state a hypothesis, run the numbers, compare the result to a threshold, decide. It is not one settled thing. It is a truce between two camps who never agreed with each other, papered over so completely that the seam does not show anymore. Seeing that seam means taking the two camps apart first, on their own terms, starting with what each one built.
What Framework Each Camp Built
Strip away the personalities and both camps start from the same five ingredients: a null hypothesis, a population standing behind whatever sample got collected, a test statistic computed from that sample, a threshold the statistic gets measured against, and a decision at the end, some verdict about the null. Say those five words to Fisher and to Neyman and Pearson, and neither man objects to the list.
Take the real afternoon behind what statisticians still call the lady tasting tea, at Rothamsted Experimental Station sometime in the 1920s. Fisher’s colleague Muriel Bristol claimed she could tell, by taste alone, whether the tea went into the cup before the milk or after. Fisher designed a test on the spot: eight cups, four poured tea first and four poured milk first, presented to her in a random order she could not predict. She sorted all eight correctly. That one experiment carries all five ingredients at once. The null hypothesis: Bristol is only guessing. The population: every way eight cups could come out sorted if guessing were all that was happening.
Shrunk down to a guessable size, that population is easier to see whole. Suppose Bristol was only given four cups, two poured tea-first and two poured milk-first. Guessing at random only means choosing which two cups she calls tea-first. Whichever two are left over become milk-first on their own, no separate guess required.
Say the true tea-first cups are {2,3}. Bristol guesses {2,4} instead: cup 2 matches, cup 4 does not, one of her two answers landed right. That single correct answer also decides the automatic milk-first side: cup 1 gets labeled milk-first correctly, cup 3 gets mislabeled milk-first when it was tea-first, so two of the four cups end up correctly sorted in total. The table below adds a column for that, how many cups each of the six possible arrangements gets right, against that same true answer, {2,3}.
| Bristol calls tea-first | Milk-first (automatic) | Correct, against true tea-first {2,3} |
|---|---|---|
| {1,2} | {3,4} | 1 of 2 tea-first (2 correct of 4 total) |
| {1,3} | {2,4} | 1 of 2 tea-first (2 correct of 4 total) |
| {1,4} | {2,3} | 0 of 2 tea-first (0 correct of 4 total) |
| {2,3} | {1,4} | 2 of 2 tea-first (4 correct of 4 total) |
| {2,4} | {1,3} | 1 of 2 tea-first (2 correct of 4 total) |
| {3,4} | {1,2} | 1 of 2 tea-first (2 correct of 4 total) |
Six guesses, six equally likely outcomes, and the third column shows only one of them gets every cup right. Fisher’s real design runs the same logic on a longer list: choosing four cups out of eight to call tea-first, instead of two out of four, gives seventy equally likely guesses instead of six, with still only one perfect match. That full list of seventy possible guesses is the population. Only one of those seventy guesses matches the true split, which puts a pure guesser’s odds of sorting all eight cups right at 1/70. There are also five possible outcomes, depending on how many of the four true tea-first cups she matches: four matched, eight cups correct in total; three matched, six cups correct; two matched, four cups correct; one matched, two cups correct; or none matched, zero cups correct.
The test statistic: the number she got right, eight out of eight. The threshold: some line for how rare a perfect sort would have to be before guessing stops being a believable story. The decision: whether this result counts as evidence against the guessing story. What each camp built on top of that same list turned into a separate practice of its own, both still wearing the same five ingredients.
Fisher called his significance testing. In his own words, “every experiment may be said to exist only to give the facts a chance of disproving the null hypothesis.” That 1/70 is the p-value: the probability, computed by assuming the null hypothesis is true, of a result at least this extreme. Fisher computed that same probability for whichever of the five outcomes came in. Getting three or more of the four tea-first cups right, for instance, would happen by pure guessing close to one time in four (17/70), not remotely rare. What he refused to do was name, once and for all, the precise point on that scale where a result stops looking like luck.
One scientist might draw the line there: only four correct answers, a perfect sort, count as real evidence against guessing.
Fisher’s framework also emphasized that a test can reject the null, but it was never built to prove the null true, only to leave it standing, unrefuted, if the evidence falls short. No threshold fixed in advance, and no verdict that counts as proof: those two choices together are the feature Fisher would later go to war to defend.
Neyman and Pearson called theirs hypothesis testing. They set out to answer a question Fisher’s framework left open: once a test is built, how often will it be wrong, and in which direction? Their answer was to name an alternative hypothesis before ever touching the data, that Bristol can really tell the difference and is not just guessing, and to fix two things in advance of pouring a single cup.
The significance level, alpha, is the chance of rejecting the guessing story when Bristol was only guessing, a false positive, fixed by Neyman and Pearson before any data comes in. Say they set it at five percent, the number that would later become standard. Before Bristol touches a single cup, they have to decide which results will count as rejecting the guessing story. Reject only on a perfect sort, four correct: that region’s true rate, what statisticians call its size, is 1/70, about 1.4 percent, safely under five percent. Widen the region to include three correct as well as four, and the size jumps to 17/70, close to one time in four, nearly five times over budget. So the rejection region ends up holding one outcome only, four correct and nothing else, because that is the largest region that still stays under the five percent alpha they agreed to in advance.
That one-outcome region is the real difference from Fisher. Fisher looks at whatever result happens, say four correct or three or two, and judges each one fresh. Neyman and Pearson commit to that single region before Bristol pours a cup, and once it is set there is no more judging left to do: her result either lands inside it or it does not.
A second number follows from that same region: the power, how often it would catch Bristol if she really can tell the difference, correctly rejecting the guessing story when it deserves to be rejected. The opposite mistake runs the other direction: failing to reject the guessing story when Bristol could tell the difference, calling a real taster lucky and clearing her as a guesser instead. Both numbers, size and power, are properties of the test itself, true before a single cup gets poured and just as true after, run again and again on a thousand different tasters. This is a behavioristic idea, closer to industrial quality control, accepting or rejecting a batch, than to a single scientist puzzling over a single dataset.
Same eight cups, same five ingredients, each side after something else.
Fisher wanted a probability statement about this one afternoon, Bristol’s actual eight cups, sorted the one time she sorted them. Neyman and Pearson were willing to give that up entirely, in exchange for a guarantee about how often the method as a whole would mislead someone, over a thousand different tasters in one afternoon or a single taster in a thousand different afternoons. Fisher’s real objection ran deeper than the math, and it had two parts. First, naming an alternative hypothesis in advance, the way power requires, asks a scientist to already know the answer to the very question the experiment exists to find out. Second, the whole framework asks a scientist puzzling over one dataset to act like a factory inspector clearing a batch, the same comparison this piece made earlier. Fisher called that reframing of an ordinary test of significance into “some kind of acceptance procedure” a mistake that, in his own words, “originated in several misapprehensions and has led, apparently, to several more.” A guarantee about a method’s long run behavior, however well earned, has nothing to say about Bristol, this taster, this afternoon, the only one that actually happened.
The Seam Nobody Named
What gets taught today, a p-value reported next to a pre-fixed significance level of 0.05, is neither Fisher’s method nor Neyman and Pearson’s. It is the stitching Gigerenzer traced back at the start of this piece: statisticians and historians of the field call it hybridization, a synthesis built by textbook authors and working scientists over decades, not by either camp, and not something either man would have signed his name to.
Fisher never stopped being blunt about it in print, dismissing what he called “mad Neymanians” for treating a scientific question like a factory inspection. Neyman was just as blunt back, at one point calling fiducial inference, Fisher’s own attempt to wring a probability statement out of a single sample without Neyman and Pearson’s machinery, “simply non-existent.” The personal hostility outlived both of them, and the hybrid got taught anyway. It explains why the test can feel airtight in a textbook and slippery the moment it meets a real decision, because it is quietly being asked to run two frameworks at once.
A Closing Reflection. Some line in your life got drawn before any data came in. Somewhere else, nothing got drawn at all, just your own fresh read of the one case standing in front of you right now. Both are probably running today, on different decisions, in the same hour.
- A line drawn before you needed it, applied the same way every time since: a spending cap picked before a trip and followed no matter what turns up once you are there, a screen-time limit that ends the day at the same minute whether it was a good day or a hard one, a household rule enforced the same way whether the moment it interrupts is boring or the best one you have had all week.
- A call made fresh each time, with nothing fixed to check it against in advance: how much to say to a friend who just told you something hard, how firm to be with a kid having a bad day, whether an apology sounds real enough to accept this time. Nobody wrote the rule down beforehand. You are reading this one case, and only this one, before you decide.
- A moment today when you catch the tell of which one you are running: the flat, pre-decided feel of an answer arriving before the question finishes, versus the slower drag of weighing this one case, right now, in front of you. Notice which shows up first, before you decide what it means.
Neither the line drawn in advance nor the read made fresh is the safer one. A line fixed before any data arrives survives the days when judgment would waver, and misses the one case in front of it that never looked like the others. A read made fresh every time can see that one case clearly, and still leaves nothing behind for the next person who asks how it was decided. Most days run on some working truce between the two, the same seam this whole piece went looking for, unsigned, unannounced, holding anyway.
Where This Practice Came From
Taft’s remark, “Roosevelt was my closest friend,” and the circumstance of his weeping in his train car on April 25, 1912, are documented in Pringle (1939), pp. 781-782. The March 1935 Royal Statistical Society meeting, the Fisher and Pearson exchange recorded in its minutes, and Gosset’s earlier letter to Fisher are real and documented, drawn from Louçã (2008), which cites the original 1935 discussion minutes and Joan Fisher Box’s 1978 biography of her father directly. The Lady Tasting Tea experiment, eight cups, four poured each way, presented in random order, with a one in seventy chance of a perfect sort by guessing alone, is Fisher’s own account of a real event at Rothamsted, given in his own book, not the six-pair version later used for teaching by Lindley (1993), which is a pedagogical simplification and is kept separate here on purpose. The line “some kind of acceptance procedure… originated in several misapprehensions and has led, apparently, to several more” is Fisher’s own, from the opening of Fisher (1955), his direct attack on the Neyman-Pearson framework published sixteen years after the Rothamsted rupture.
Intellectual Honesty Note. The lifelong personal hostility between Fisher and Neyman, including the specific terms quoted above, is documented in the historical record of the dispute, not invented for effect; it coexisted with real, sourced warmth between them earlier in their careers, and neither fact cancels the other. Framework is standard terminology in the statistics literature for both Fisherian significance testing and the Neyman-Pearson approach; Fisher himself used it this way in 1955, and Gigerenzer et al. (1989) and Louçã (2008) both frame the 1930s dispute the same way. Hybridization is Gigerenzer et al.’s own term for the synthesis taught today, not this piece’s invention.
References
Fisher, R.A. (1935). The Design of Experiments. Oliver and Boyd.
Fisher, R.A. (1935). Contribution to the discussion of J. Neyman, “Statistical problems in agricultural experimentation.” Journal of the Royal Statistical Society Supplement, 2, 154-157.
Fisher, R.A. (1955). “Statistical Methods and Scientific Induction.” Journal of the Royal Statistical Society, Series B, 17(1), 69-78.
Fisher, J. (1978). R.A. Fisher: The Life of a Scientist. Wiley.
Gigerenzer, G., Swijtink, Z., Porter, T., Daston, L., Beatty, J., & Krüger, L. (1989). The Empire of Chance: How Probability Changed Science and Everyday Life. Cambridge University Press.
Gosset, W.S. (1970). Letters from W.S. Gosset to R.A. Fisher, 1915-1936. Issued for private circulation. Dublin: Arthur Guinness.
Lindley, D.V. (1993). “The Analysis of Experimental Data: The Appreciation of Tea and Wine.” Teaching Statistics, 15(1), 22-25.
Louçã, F. (2008). “The Widest Cleft in Statistics: How and Why Fisher Opposed Neyman and Pearson.” ISEG/UTL Working Paper 02/2008/DE/UECE.
Pringle, H.F. (1939). The Life and Times of William Howard Taft: A Biography. New York: Farrar & Rinehart.