<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>History on statistical.systems</title>
    <link>https://statistical.systems/tags/history/</link>
    <description>Recent content in History on statistical.systems</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en-us</language>
    <lastBuildDate>Mon, 17 Aug 2026 23:59:00 -0400</lastBuildDate><atom:link href="https://statistical.systems/tags/history/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>The Seam Does Not Show Anymore</title>
      <link>https://statistical.systems/blog/the_seam_does_not_show_anymore/</link>
      <pubDate>Mon, 17 Aug 2026 23:59:00 -0400</pubDate>
      
      <guid>https://statistical.systems/blog/the_seam_does_not_show_anymore/</guid>
      <description>A room at the Royal Statistical Society in 1935, a friendship breaking apart in front of everyone, and two statisticians whose falling out never got the ending anyone expects.</description>
      <content:encoded><![CDATA[<p>&ldquo;Roosevelt was my closest friend.&rdquo; Taft said brokenly, then wept.</p>
<p>It was April 25, 1912. A journalist had found him slumped in his train car after a brutal day of campaigning in Boston, against Theodore Roosevelt, the man who had picked him as his own successor four years earlier. Everyone watching that race knew how this kind of story ends. A friendship breaks, a party splits, somebody loses, in this case badly enough that Woodrow Wilson won the whole election. That is the version of betrayal most people know how to recognize.</p>
<p>There is a quieter version: no resolution, no recorded tears, no losing side, just a method still running in every introductory statistics class. Somewhere in that course, you probably ran into Fisher, Neyman, and Pearson, even if you never clocked how much they shaped the field.</p>
<p>Jerzy Neyman stood up to read a paper, and by the time he sat back down, a friendship was coming apart in the room with him. It was March 28, 1935, at the Royal Statistical Society. Neyman&rsquo;s paper, on the design of agricultural experiments, questioned work belonging to the department Ronald A. Fisher had spent fourteen years building at Rothamsted. Fisher rose during the discussion and told the room it was &ldquo;clear to everyone present that Dr. Neyman has been somewhat unwise in his choice of topics,&rdquo; then said flatly, &ldquo;Dr. Neyman has misunderstood the intention&hellip; of the z test and of the Latin Square.&rdquo; Egon Pearson, Neyman&rsquo;s collaborator, stood up next. He did not defend the paper&rsquo;s mathematics. He defended the right to have written it at all, telling Fisher directly that he ought to &ldquo;beg leave to question the wisdom of accusing a fellow worker of incompetence without&hellip; showing that he had succeeded in mastering the argument&rdquo; himself. The minutes recorded the exchange. Nobody in that room mistook it for an ordinary disagreement.</p>
<p>A few years earlier, William Sealy Gosset, who published his own statistical work under the name &ldquo;Student&rdquo; and counted both men as friends, wrote to Fisher trying to arrange a research visit for Neyman at Rothamsted. He called Neyman &ldquo;fonder of algebra than correlation tables,&rdquo; and &ldquo;the only person except yourself I have heard talk about maximum likelihood as if he enjoyed it.&rdquo; Gosset was not being polite. That was the highest compliment he had, one craftsman recognizing another across a shared obsession. Whatever broke in that room in 1935 had been, only a short while before, warm enough that a mutual friend wanted the two of them working side by side.</p>
<p>If I were Fisher, I would feel blindsided, mostly. I had spent at least a decade building that department. The challenge to it came from someone I had been warm with, someone a mutual friend once thought I should be working alongside, not defending myself against. That is a strange kind of vertigo: the relationship changes shape in public, in real time, and I only find out by watching myself react to it. But the disagreement itself is not the sharper sting. It is Pearson standing up right after, not even engaging the argument, going straight for &ldquo;you don&rsquo;t get to talk to him that way.&rdquo; That is worse than losing on the merits. I came in defending a decade of work. I left having been quietly told, before I had a second to compose myself, that I had behaved badly while doing it.</p>
<p>Almost everyone has felt some version of this: watching someone they trained, mentored, or simply assumed was still on their side take a public shot at the thing they built, and realizing the friendship shifted before they ever got a vote. There is no private version of that conversation. People find out it changed by watching it break in front of witnesses.</p>
<p>However, if I were Neyman, I would feel exposed first. Standing up to read a paper is the vulnerable position. I was the one putting a claim in front of the room, and Fisher was not just another audience member. He ran the field. He does not answer with a counter-argument. First I am told I was unwise to have chosen this topic at all, before I have defended a word of it. Then I am told I misunderstood the very tests I was correcting. Neither line touches my math. One says I should not have spoken. The other says I did not understand what I was speaking about. That is not &ldquo;you&rsquo;re wrong.&rdquo; That is &ldquo;you didn&rsquo;t do the reading,&rdquo; delivered by the one person whose opinion of my competence mattered, in front of everyone whose opinion mattered too.</p>
<p>That kind of dismissal carries a particular unfairness, worse than losing an argument on the merits. If Fisher had said, &ldquo;here&rsquo;s where your math breaks,&rdquo; I could answer it. &ldquo;You misunderstood the logic&rdquo; forecloses the conversation before it starts. It is a status move, not a substantive one, and I have no comeback that does not sound defensive. I built this paper because I saw a real gap in a framework I respected. I did not expect to be told I had not earned the right to see it. Then Pearson stands up for me. Relief, probably, but a complicated kind, because now the room has also watched me need rescuing. I did not get to win the argument on my own. Someone else had to win the right for me to have made it at all. I thought I had earned something with Fisher. Gosset thought enough of what I understood to try to put us in a room together. Whatever that was, it did not survive me having a good idea he had not had first.</p>
<p>Roosevelt and Taft&rsquo;s rift had an ending anyone could point to: Wilson took the election. Taft finished behind even Roosevelt, third in a race he had entered as the sitting president. Fisher and Neyman&rsquo;s rift never found that kind of ending. Neither man ever conceded the argument. Neither framework lost. The two got run together anyway, the same method holding both without either man&rsquo;s signature on it. The historian of science Gerd Gigerenzer later traced the seam back to its two separate sources. Most people who have ever read the phrase &ldquo;statistically significant&rdquo; in a headline, watched a poll get called the moment it crossed some invisible threshold, or run the test themselves without asking where the threshold came from, have never heard either statistician&rsquo;s name, or known there was ever a rupture behind the method they were using. It reads as one settled thing: state a hypothesis, run the numbers, compare the result to a threshold, decide. It is not one settled thing. It is a truce between two camps who never agreed with each other, stitched together so completely that the seam does not show anymore.</p>
<h2 id="what-framework-each-camp-built">What Framework Each Camp Built</h2>
<p>Strip away the personalities and both camps start from the same five ingredients: a null hypothesis, a population standing behind whatever sample got collected, a test statistic computed from that sample, a threshold the statistic gets measured against, and a decision at the end, some verdict about the null. Say those five words to Fisher and to Neyman and Pearson, and neither man objects to the list. The design both camps fought over, the eight-cup test built for the lady tasting tea, belongs to <a href="/blog/what_it_would_take_in_cups/">this article</a>. What belongs here is only what the two men built on top of any shared design, and the cleanest way to see it is a case with no story attached at all.</p>
<p>Picture a police officer with a radar gun on a road with a sixty-five mile an hour limit. One driver gets clocked at seventy. A second driver, caught later the same day, gets clocked at ninety-six. Fisher and Neyman-Pearson report two different things about those same two readings, and the difference is the whole argument.</p>
<p><strong>Fisher called his significance testing.</strong> In his own words, &ldquo;every experiment may be said to exist only in order to give the facts a chance of disproving the null hypothesis.&rdquo; Fisher&rsquo;s method writes a report that gets worse the faster the driver was going. Seventy in a sixty-five zone reads as a routine note. Ninety-six reads as reckless, clearly over the line. Every fresh reading earns its own fresh description, a faster reading always earning a harsher one. What he refused to do was name, once and for all, the exact speed where &ldquo;a little fast&rdquo; turns into &ldquo;reckless.&rdquo;</p>
<p>Fisher&rsquo;s framework also emphasized that a test can reject the null, but it was never built to prove the null true, only to leave it standing, unrefuted, if the evidence falls short. No threshold fixed in advance, and no verdict that counts as proof: those two choices together are the feature Fisher would later go to war to defend.</p>
<p><strong>Neyman and Pearson called theirs hypothesis testing.</strong> They set out to answer a question Fisher&rsquo;s framework left open: once a test is built, how often will it be wrong, and in which direction? The machinery they built to answer it, an alternative hypothesis named before the data ever arrives and an error rate budgeted on purpose, gets its own full treatment in <a href="/blog/one_napkin_two_fires/">this article</a>. Only the verdict that budget produces belongs here. Fix the citation threshold at the speed limit itself, sixty-five, and the two readings from before land the same way. Whether a driver gets caught at seventy or ninety-six, Neyman and Pearson report the same thing either time: reject the null, issue the citation. The verdict was decided before either driver was ever clocked, not by how fast they were going.</p>
<p>That fixed report is the real difference from Fisher. Fisher looks at whatever speed happens and judges it fresh, a faster reading earning a harsher description every time. Neyman and Pearson commit to a threshold before any car passes, and once it is set there is no more judging left to do: a driver is either over the limit or not, and the citation reads the same whether they barely qualified or blew past it by thirty miles an hour.</p>
<p>A second number follows from that same budget: power, how often the test catches a real effect instead of clearing it as noise. Size and power are both properties of the test itself, true before a single measurement is taken and just as true after, run again and again over many samples. This is a behavioristic idea, closer to industrial quality control, accepting or rejecting a batch, than to a single scientist puzzling over a single dataset.</p>
<p>Same test, same five ingredients, each side after something else.</p>
<p>Fisher wanted a probability statement about this one sample, the actual reading in front of him, taken the one time it was taken. Neyman and Pearson were willing to give that up entirely, in exchange for a guarantee about how often the method as a whole would mislead someone, over many different samples in the long run. Fisher&rsquo;s real objection ran deeper than the math, and it had two parts. First, naming an alternative hypothesis in advance, the way power requires, asks a scientist to already know the answer to the very question the experiment exists to find out. Second, the whole framework asks a scientist puzzling over one dataset to act like a factory inspector clearing a batch, the same comparison this piece made earlier. Fisher called that reframing of an ordinary test of significance into &ldquo;some kind of acceptance procedure&rdquo; a mistake that &ldquo;originated in several misapprehensions and has led, apparently, to several more.&rdquo; A guarantee about a method&rsquo;s long run behavior, however well earned, has nothing to say about this sample, this dataset, the only one that happened.</p>
<h2 id="the-seam-nobody-named">The Seam Nobody Named</h2>
<p>What gets taught today, a p-value reported next to a pre-fixed significance level of 0.05, is neither Fisher&rsquo;s method nor Neyman and Pearson&rsquo;s. It is the seam Gigerenzer traced back at the start of this piece: statisticians and historians of the field call it hybridization, a synthesis built by textbook authors and working scientists over decades, not by either camp, and not something either man would have signed his name to.</p>
<p>Fisher never stopped being blunt about it in print, dismissing what he called &ldquo;mad Neymanians&rdquo; for treating a scientific question like a factory inspection. Neyman was just as blunt back, at one point calling fiducial inference, Fisher&rsquo;s own attempt to wring a probability statement out of a single sample without Neyman and Pearson&rsquo;s machinery, &ldquo;simply non-existent.&rdquo; The personal hostility outlived both of them, and the hybrid got taught anyway. It explains why the test can feel airtight in a textbook and slippery the moment it meets a real decision, because it is quietly being asked to run two frameworks at once.</p>
<blockquote>
<p><strong>A Closing Reflection.</strong> <em>A speed limit does not care how fast you were going, over is over. A judgment made new each time cares about nothing else. You run both of these right now, in your own life. Some decisions follow a standard set before anything happens, no exceptions. Others get sized up as they come, nothing settled beforehand to check them against. Both are running today, on different choices, in the same hour.</em></p>
<ol>
<li>A line drawn before you needed it, applied the same way every time since, no pause before it lands: a spending cap picked before a trip and followed no matter what turns up once you are there, a screen-time limit that ends the day at the same minute whether it was a good day or a hard one, a household rule enforced the same way whether the moment it interrupts is boring or the best one you have had all week.</li>
<li>A call made fresh each time, a beat of quiet before you answer, eyes on the other person&rsquo;s face while you weigh it: how much to say to a friend who just told you something hard, how firm to be with a kid having a bad day, whether an apology sounds real enough to accept this time. Nobody wrote the rule down beforehand. You are reading this one case, and only this one, before you decide.</li>
<li>A moment today when you catch the tell of which one you are running: the flat, pre-decided feel of an answer arriving before the question finishes, versus the slower drag of weighing this one case, right now, in front of you. Notice which shows up first, before you decide what it means.</li>
</ol>
<p><em>Neither the line drawn in advance nor the call made fresh is the safer one. The first survives the days when judgment would waver, and misses the one case in front of it that never looked like the others, the citation that reads the same at seventy as at ninety-six. The second can see that one case clearly, and still leaves nothing behind for the next person who asks how it was decided. Most days run on some working truce between the two, the same seam this whole piece went looking for, unsigned, unannounced, holding anyway.</em></p></blockquote>
<h2 id="where-this-practice-came-from">Where This Practice Came From</h2>
<p>Taft&rsquo;s remark, &ldquo;Roosevelt was my closest friend,&rdquo; and the circumstance of his weeping in his train car on April 25, 1912, are documented in Pringle (1939), pp. 781-782. The March 1935 Royal Statistical Society meeting, the Fisher and Pearson exchange recorded in its minutes, and Gosset&rsquo;s earlier letter to Fisher are real and documented, drawn from Louçã (2008), which cites the original 1935 discussion minutes and Joan Fisher Box&rsquo;s 1978 biography of her father directly. The lady tasting tea, the design both camps are illustrated against, is Fisher&rsquo;s own, given in <em>The Design of Experiments</em> (1935); what really happened the afternoon behind it, and how much of the famous ending holds up, is a question <a href="/blog/what_it_would_take_in_cups/">What It Would Take, In Cups</a> takes up in full. The z-test comparison, Fisher&rsquo;s shrinking p-value against Neyman and Pearson&rsquo;s fixed report, along with the z = 2.3 and z = 2.9 example itself, is drawn from Louçã (2008), who in turn credits Berger (2003) for the example. The Neyman-Pearson machinery this piece only budgets for, the alternative hypothesis, the error rate fixed in advance, and the four-outcome table it produces, is worked out in full, against a different case, in <a href="/blog/one_napkin_two_fires/">One Napkin, Two Fires</a>. The line &ldquo;some kind of acceptance procedure&hellip; originated in several misapprehensions and has led, apparently, to several more&rdquo; is Fisher&rsquo;s own, from the opening of Fisher (1955), his direct attack on the Neyman-Pearson framework published sixteen years after the Rothamsted rupture.</p>
<p><strong>Intellectual Honesty Note.</strong> The lifelong personal hostility between Fisher and Neyman, including the specific terms quoted above, is documented in the historical record of the dispute, not invented for effect; it coexisted with real, sourced warmth between them earlier in their careers, and neither fact cancels the other. Framework is standard terminology in the statistics literature for both Fisherian significance testing and the Neyman-Pearson approach; Fisher himself used it this way in 1955, and Gigerenzer et al. (1989) and Louçã (2008) both frame the 1930s dispute the same way. Hybridization is Gigerenzer et al.&rsquo;s own term for the synthesis taught today, not this piece&rsquo;s invention.</p>
<h2 id="references">References</h2>
<p>Berger, J. (2003). &ldquo;Could Fisher, Jeffreys and Neyman Have Agreed on Testing?&rdquo; <em>Statistical Science</em>, 18(1), 1-32.</p>
<p>Fisher, R.A. (1935). <em>The Design of Experiments</em>. Oliver and Boyd.</p>
<p>Fisher, R.A. (1935). Contribution to the discussion of J. Neyman, &ldquo;Statistical problems in agricultural experimentation.&rdquo; <em>Journal of the Royal Statistical Society Supplement</em>, 2, 154-157.</p>
<p>Fisher, R.A. (1955). &ldquo;Statistical Methods and Scientific Induction.&rdquo; <em>Journal of the Royal Statistical Society, Series B</em>, 17(1), 69-78.</p>
<p>Fisher, J. (1978). <em>R.A. Fisher: The Life of a Scientist</em>. Wiley.</p>
<p>Gigerenzer, G., Swijtink, Z., Porter, T., Daston, L., Beatty, J., &amp; Krüger, L. (1989). <em>The Empire of Chance: How Probability Changed Science and Everyday Life</em>. Cambridge University Press.</p>
<p>Gosset, W.S. (1970). <em>Letters from W.S. Gosset to R.A. Fisher, 1915-1936</em>. Issued for private circulation. Dublin: Arthur Guinness.</p>
<p>Louçã, F. (2008). &ldquo;The Widest Cleft in Statistics: How and Why Fisher Opposed Neyman and Pearson.&rdquo; ISEG/UTL Working Paper 02/2008/DE/UECE.</p>
<p>Pringle, H.F. (1939). <em>The Life and Times of William Howard Taft: A Biography</em>. New York: Farrar &amp; Rinehart.</p>
]]></content:encoded>
    </item>
    
  </channel>
</rss>
