“Roosevelt was my closest friend.” Taft said brokenly, then wept.
It was April 25, 1912. A journalist had found him slumped in his train car after a brutal day of campaigning in Boston, against Theodore Roosevelt, the man who had picked him as his own successor four years earlier. Everyone watching that race knew how this kind of story ends. A friendship breaks, a party splits, somebody loses, in this case badly enough that Woodrow Wilson won the whole election. That is the version of betrayal most people know how to recognize.
There is a quieter version: no resolution, no recorded tears, no losing side, just a method still running in nearly every introductory statistics class. Anyone who sat through that course met Fisher, Neyman, and Pearson, whether or not they clocked how much of it those three wrote.
Jerzy Neyman stood up to read a paper, and by the time he sat back down, a friendship was coming apart in the room with him. It was March 28, 1935, at the Royal Statistical Society. The paper he presented, on the design of agricultural experiments and written with two colleagues, questioned work belonging to the department Ronald A. Fisher had spent fourteen years building at Rothamsted. Fisher rose during the discussion and told the room it was “clear to everyone present that Dr. Neyman has been somewhat unwise in his choice of topics,” then said flatly, “Dr. Neyman has misunderstood the intention… of the z test and of the Latin Square.” Egon Pearson, Neyman’s collaborator, stood up next. He did not defend the paper’s mathematics. He defended the right to have written it at all, telling the room he would “beg leave to question the wisdom of accusing a fellow worker of incompetence without… showing that he had succeeded in mastering the argument.” The minutes recorded the exchange.
Years earlier, William Sealy Gosset, who published his own statistical work under the name “Student” and counted both men as friends, wrote to Fisher trying to arrange a research visit for Neyman at Rothamsted. He called Neyman “fonder of algebra than correlation tables,” and “the only person except yourself I have heard talk about maximum likelihood as if he enjoyed it.” That was the highest compliment he had, one craftsman recognizing another across a shared obsession. Whatever broke in that room in 1935 had been, years before, warm enough that a mutual friend wanted the two of them working side by side.
If I were Fisher, I would feel blindsided, mostly. I had spent at least a decade building that department. The challenge to it came from someone I had been warm with, someone a mutual friend once thought I should be working alongside, not defending myself against. But the disagreement itself is not the sharper sting. It is Pearson standing up right after, not even engaging the argument, going straight for how I had spoken to him. That is worse than losing on the merits. I left having been quietly told, before I had a second to compose myself, that I had behaved badly while defending my own work.
Almost everyone has felt some version of this: watching someone they trained, mentored, or simply assumed was still on their side take a public shot at the thing they built, and realizing the friendship shifted before they ever got a vote. There is no private version of that conversation. People find out it changed by watching it break in front of witnesses.
However, if I were Neyman, I would feel exposed first. Standing up to read a paper is the vulnerable position. I was the one putting a claim in front of the room, and Fisher was not just another audience member. He ran the field. First I am told I was unwise to have chosen this topic at all, before I have defended a word of it. Then I am told I misunderstood the very tests I was correcting. Neither line touches my math. One says I should not have spoken. The other says I did not understand what I was speaking about. Together they say I did not do the reading, and they say it through the one person whose opinion of my competence mattered, in front of everyone whose opinion mattered too.
It is a status move. If Fisher had said, “here is where your math breaks,” I could answer it. “You misunderstood” leaves me no comeback that does not sound defensive. Then Pearson stands up for me. Relief, probably, but a complicated kind, because now the room has also watched me need rescuing. Gosset once thought enough of what I understood to try to put us in a room together. Whatever that was, it did not survive me having a good idea Fisher had not had first.
Roosevelt and Taft’s rift had an ending anyone could point to: Wilson took the election. Taft finished behind even Roosevelt, third in a race he had entered as the sitting president. Fisher and Neyman’s rift never found that kind of ending. Neither man ever conceded the argument. Neither framework lost. The two got run together anyway, the same method holding both without either man’s signature on it. The historian of science Gerd Gigerenzer later traced the seam back to its two separate sources.
You have probably used this method without either name attached. Maybe you read “statistically significant” in a headline, or watched a poll get called the moment it crossed some invisible threshold. Maybe you ran the test yourself and never asked where the threshold came from. Each time, a number cleared a line and you were told something was real. Whether that is a fair reading depends on which question the number was answering. Fisher asked how surprising this one result would be if nothing were going on. Neyman and Pearson asked for a rule that, used again and again, would mislead you only so often. The method you ran no longer tells you which one it answered.
It reads as one settled thing: state a hypothesis, run the numbers, compare the result to a threshold, decide. It is not one settled thing. It is a truce between two camps who never agreed with each other, stitched together so completely that the seam does not show anymore.
What Each Camp Built
Strip away the rupture in that room and both camps still share four pieces. Both start with a null hypothesis. Both fix in advance a summary measure drawn from the sample, either the test statistic itself or a function of it, its form set before any data are collected. Both then take a real sample and compute that number, its observed value. And both end on a decision about the null, made by comparing that value to some cutoff. This much, neither camp disputes. Agreement ends there.
Where the seam would later run is easiest to see on a road. Picture a police officer with a radar gun on a road with a sixty-five mile an hour limit. Three drivers get clocked over the course of the day: one at eighty, one at eighty-four, one at one hundred. Fisher and Neyman–Pearson report different things about those same three readings, and the difference is the whole argument.
Fisher called his significance testing. The vocabulary, null hypothesis and p-value, is worked out cup by cup in a separate piece on the lady tasting tea, the founding story most people know his method by. In his own words, “every experiment may be said to exist only in order to give the facts a chance of disproving the null hypothesis.” Fisher’s method writes a report that gets worse the faster the driver was going. Eighty in a sixty-five zone reads as mildly over. Eighty-four reads as a firmer flag. One hundred reads as reckless, wildly over the line. Every new reading earns its own description, a faster one always earning a harsher note. What he refused to do was name, once and for all, the exact speed where “a little fast” turns into “reckless.”
Fisher also emphasized that a test can reject the null, but it was never built to prove the null true, only to leave it standing, unrefuted, if the evidence falls short. It also never named a specific alternative, no second hypothesis waiting to be tested against the null, only the null itself and the data’s willingness to contradict it. Together with the threshold he refused to fix, these were the features Fisher would later go to war to defend.
Neyman–Pearson called theirs hypothesis testing. They set out to answer a question Fisher left open: once a test is built, how often will it be wrong, and in which direction? They built machinery to answer it: the null and alternative hypotheses, Type I and Type II errors, and a test’s level of significance, size and power. All of it is unpacked in the courtroom case built on Glass Onion, where a stolen idea is tried on the word of four friends. Only the verdict this machinery produces belongs here. Fix a citation threshold, the rejection region, in advance, and all three readings land the same way. Whether a driver gets caught at eighty, eighty-four, or one hundred, Neyman–Pearson report the same thing every time: reject the null, issue the citation. The verdict was decided before any of them was ever clocked, not by how fast they were going.
Put the same three readings in the units an introductory course uses, where the number is a test statistic, written z. The null here is that a driver is keeping pace with traffic. On this road, traffic averages the limit, sixty-five. Some drivers go faster and some slower, with a spread of roughly six and a half. The two close drivers, eighty and eighty-four, land at z = 2.3 and 2.9. Fisher separates even those: a driver keeping pace with traffic reads eighty or faster about 1 time in 93, and eighty-four or faster about 1 in 540. The reckless one, one hundred, sits so far out it would turn up about 1 time in 25 million. Neyman–Pearson fix their cutoff beforehand at the 0.05 level of significance, which puts the rejection region at z > 1.645. On this road that is near seventy-six, eleven over the limit. Every reading clears it, so all three draw the identical citation, whether the radar gun read eighty or one hundred.
One Test, Two Frameworks
Fisher and Neyman–Pearson run the same test on the same reading, but answer two different questions with it. The grey spine down the middle is the skeleton both camps agree on. The colored cells are where they part.
Fisher
Significance testing
the statistic is a propertyof the sample
Neyman–Pearson
Hypothesis testing
size and power are propertiesof the test
Set the null, H₀.
Set the null H₀, and name the alternative Hₐ.
Fix the test statistic T and its distribution under H₀. Its form is set before any data.
Nothing is fixed in advance.
No line drawn. No alternative to weigh.
Fix the level α and the rejection region R, before a single measurement.
Draw a real sample. Compute T.
Report the p-value: how rare is this reading under H₀? The report sharpens as the reading climbs.
Is T inside R? A yes or no, settled before the data arrived.
Reject H₀ if p is small. Judged fresh each time. The null is left standing, never proved.
If T is in R, reject H₀ and accept Hₐ. The verdict was set before the data.
Inductive inference: a probability statement about this one sample.
Inductive behavior: a guarantee about how often the method misleads, over the long run.
Same test, same four pieces, each side after something else.
The divide begins with the population each camp imagines behind that one reading. Fisher’s is an infinite hypothetical population, and the reading is a single extraction from it. Neyman–Pearson’s is repeated sampling, the same test taken again and again, without end. Both built a mental model, a working picture of a mechanism. Neyman and Pearson put a framework on theirs. Fisher, on purpose, did not.
Fisher wanted a probability statement about this one sample, the actual reading in front of him, taken the one time it was taken. Neyman–Pearson were willing to give that up entirely, in exchange for a guarantee about how often the method as a whole would mislead someone, over many different samples in the long run. That trade is what makes theirs a framework in the strict sense, a decision structure run the same way on every new case. Fisher’s real objection ran deeper than the math, and it had two parts. First, power, the chance a test catches a real effect when there is one, requires naming an alternative hypothesis in advance. Naming it in advance asks a scientist to already know the answer to the very question the experiment exists to find out. Second, the whole framework asks a scientist puzzling over one dataset to act like a factory inspector clearing a batch. Fisher called that reframing of an ordinary test of significance into “some kind of acceptance procedure” a mistake that “originated in several misapprehensions and has led, apparently, to several more.” A guarantee about a method’s long run behavior, however well earned, has nothing to say about this sample, this dataset, the only one that happened.
The Seam Nobody Named
What gets taught today, a p-value reported next to a pre-fixed significance level of 0.05, is neither Fisher’s method nor Neyman–Pearson’s. It is the seam Gigerenzer traced back at the start of this piece. Statisticians and historians of the field call it hybridization: a synthesis built by textbook authors and working scientists over decades, not by either camp. Neither man would have signed his name to it.
Fisher never stopped being blunt about it in print, dismissing what he called “mad Neymanians” for treating a scientific question like a factory inspection. Neyman was just as blunt back. Fisher had tried to wring a probability statement out of a single sample without Neyman–Pearson’s machinery, a method he called fiducial inference, and Neyman called it “simply non-existent.” The hostility ran for decades, and the hybrid got taught anyway. So a study reports p = 0.03 beside the 0.05 line, and two readers take away different things. One hears a verdict: reject, move on. The other hears a strength of evidence, about 1 in 33. Each is reading half of the hybrid.
A Closing Reflection. A speed limit does not care how fast you were going, over is over. A judgment made new each time cares about nothing else. You run both of these right now, in your own life. Some decisions follow a standard set before anything happens, no exceptions. Others get sized up as they come, nothing settled beforehand to check them against.
- A line drawn before you needed it, applied the same way every time: the wallet staying shut at the register because the trip’s cap is spent, the screen going dark at the fixed minute, good day or hard one, the house rule enforced mid-laugh.
- A call made new each time, a beat of quiet before you answer, eyes on the other person’s face while you weigh it. How much to say to a friend who just told you something hard. How firm to be with a kid having a bad day. Whether an apology sounds real enough to accept this time.
- A moment today when you catch the tell of which one you are running: the flat, pre-decided feel of an answer arriving before the question finishes, versus the slower drag of weighing this one case, right now, in front of you. Notice which shows up first, before you decide what it means.
Neither the line drawn in advance nor the call made new is the safer one. The first survives the days when judgment would waver, and misses the one case in front of it that never looked like the others, the citation that reads the same at eighty as at one hundred. The second can see that one case clearly, and still leaves nothing behind for the next person who asks how it was decided. Taft at least got an ending, something to weep over in a train car. Most days get no ending, just a working truce between the two, unsigned and unannounced, stitched so close the seam does not show.
Where This Practice Came From
Taft’s remark, “Roosevelt was my closest friend,” and the circumstance of his weeping in his train car on April 25, 1912, are documented in Pringle (1939), pp. 781-782. The March 1935 Royal Statistical Society meeting, the Fisher and Pearson exchange recorded in its minutes, and Gosset’s earlier letter to Fisher are real and documented, drawn from Louçã (2008), which cites the original 1935 discussion minutes and Joan Fisher Box’s 1978 biography of her father directly. The lady tasting tea, the story behind Fisher’s vocabulary here, is his own, given in The Design of Experiments (1935b). What really happened that afternoon, and how much of the famous ending holds up, is taken up in full in the cup-by-cup retelling of the tea test. The z-test comparison, Fisher’s shrinking p-value against Neyman–Pearson’s fixed report, along with the z = 2.3 and z = 2.9 example itself, is drawn from Louçã (2008), who in turn credits Berger (2003) for the example. The Neyman–Pearson machinery this piece only budgets for, the alternative hypothesis, the error rate fixed in advance, and the four-outcome table it produces, is worked out in the Glass Onion courtroom piece. The line “some kind of acceptance procedure… originated in several misapprehensions and has led, apparently, to several more” is Fisher’s own, from the opening of Fisher (1955), his direct attack on the Neyman–Pearson framework published twenty years after the 1935 meeting. Fisher’s “mad Neymanians” and Neyman’s “simply non-existent” are quoted from Louçã (2008).
Intellectual Honesty Note. The personal hostility between Fisher and Neyman, which ran for decades, is documented in the historical record of the dispute, including the terms quoted above. It coexisted with real, sourced warmth between them earlier in their careers. Neither fact cancels the other. The two “if I were” passages are dramatized. They imagine each man’s side of the 1935 meeting from the documented exchange, not from either man’s own account. Framework is standard terminology in the statistics literature for both Fisherian significance testing and the Neyman–Pearson approach. Fisher himself used it this way in 1955, and Gigerenzer et al. (1989) and Louçã (2008) both frame the 1930s dispute in those terms. Hybridization is Gigerenzer et al.’s own term for the synthesis taught today, not this piece’s invention. The p-values quoted for the radar example are computed one-sided, because that example counts only a reading over the limit, not one under it. The speeds, and the spread of about six and a half miles an hour, are invented to fit the sourced z-values. The same z-values, given two-sided as the example usually is, would read as 1 in 47 and 1 in 270.
References
Berger, J. (2003). “Could Fisher, Jeffreys and Neyman Have Agreed on Testing?” Statistical Science, 18(1), 1-32.
Fisher, R.A. (1935a). Contribution to the discussion of J. Neyman, “Statistical problems in agricultural experimentation.” Journal of the Royal Statistical Society Supplement, 2, 154-157.
Fisher, R.A. (1935b). The Design of Experiments. Oliver and Boyd.
Fisher, R.A. (1955). “Statistical Methods and Scientific Induction.” Journal of the Royal Statistical Society, Series B, 17(1), 69-78.
Fisher, J. (1978). R.A. Fisher: The Life of a Scientist. Wiley.
Gigerenzer, G., Swijtink, Z., Porter, T., Daston, L., Beatty, J., & Krüger, L. (1989). The Empire of Chance: How Probability Changed Science and Everyday Life. Cambridge University Press.
Gosset, W.S. (1970). Letters from W.S. Gosset to R.A. Fisher, 1915-1936. Issued for private circulation. Dublin: Arthur Guinness.
Louçã, F. (2008). “The Widest Cleft in Statistics: How and Why Fisher Opposed Neyman and Pearson.” ISEG/UTL Working Paper 02/2008/DE/UECE.
Neyman, J., with the cooperation of Iwaszkiewicz, K., & Kołodziejczyk, St. (1935). “Statistical Problems in Agricultural Experimentation.” Journal of the Royal Statistical Society Supplement, 2(2), 107-154 (discussion 154-180).
Pringle, H.F. (1939). The Life and Times of William Howard Taft: A Biography. New York: Farrar & Rinehart.