One afternoon in the early 1920s, at Rothamsted, an agricultural research station in Hertfordshire, R.A. Fisher drew a cup of tea from the urn and offered it to the woman beside him. She declined it. Muriel Bristol, a phycologist, told him she preferred a cup with the milk poured in first, and that she could tell the difference by taste. Not guess at it. Tell.

“Nonsense,” Fisher said. “Surely it makes no difference.” The polite thing at that point is to say how interesting and let it go. Bristol maintained, with emphasis, that of course it did.

Then a voice from just behind them said, “Let’s test her.” It belonged to William Roach, a biochemist at the same station.

They ran it there and then. Cups were poured at the trestle table, Roach handing them round and writing down what she said. Bristol picked out enough of them correctly, and Roach exulted.

That should have been the end of it. A claim was made, a test was run, and the skeptic was corrected on the spot.

Fisher was not satisfied. What Roach had thrown together had come out in her favor without being able to account for it, and Fisher was already turning over what it would have taken to do it properly. How many cups. Whether they should be paired. In what order they should be presented. What to do about variation in temperature and sweetness. What could be concluded from a perfect score, and what from a score with a single error in it. The answer was years away, and it sits in The Design of Experiments in 1935. Eight cups, four poured tea first and four poured milk first, presented in a random order. The lady is told in advance what the test will consist of, and that there are four of each kind. Her task is to divide the eight into two sets of four.

The design is documented. The afternoon is only remembered. The eight cups are Fisher’s own, published under his name. He never named the lady, never said where or when, and never reported how she did.

Nearly all of that comes from one place. Joan Fisher Box, his daughter and his biographer, is the source for the station, the urn, the trestle table, and Bristol by name. Roach at her elbow with the pencil is Box as read by John Richardson, who in 2021 set the surviving accounts against the records. She has Bristol getting more than enough of the cups right to prove her case, which is Roach’s verdict, spoken by the man who had called for the test and would marry her not long afterward. Box then adds that her personal triumph was never recorded.

The familiar ending, in which Bristol sorts all eight correctly, is in neither Box nor Fisher. It comes from David Salsburg, who heard it in the late 1960s from Hugh Fairfield Smith, a statistician who said he had been there. Smith never wrote it down. Salsburg published it in 2001. Forty-five years from the afternoon to the telling, thirty-five more from the telling to the page. In that version it is a Cambridge tea party, the people at the table are university dons and their wives, and the woman has no name at all.

Then there is the arithmetic, which needs one borrowed assumption to run at all. No source records how many cups Bristol was handed. Box does not say, and Fisher’s eight is a number from a book published more than a decade later. Richardson runs the eight through Box’s account anyway, because it is the only structure on offer: four poured each way, sorted into two piles of four. Under that structure a single error cannot happen. Getting one cup wrong means getting a second one wrong, one from each pile. Roach’s verdict was that she got more than enough right, which is not a claim that she got all of them. If she fell short of all of them, she made two errors at least. Eight cups with two errors is the kind of afternoon guessing produces about one time in four. Nothing Roach said rules that afternoon out.

Maurice Kendall, in an obituary published the year after Fisher died, reports him saying that he never performed the tea-tasting experiment at all. Box’s version is what keeps that from meaning the afternoon never happened. On her telling it was Roach who ran the cups, with Fisher standing beside the table watching it go by. The design can be checked and the result cannot. Hold the perfect score loosely, and hold the eight cups tight.

Most claims like Bristol’s never get that far. My husband can taste aspartame. Hand him the diet version of anything, no label and no warning, and he will tell you before the second sip. He can also name the singer before the first line of a song is over, Kate Bush against Joanna Newsom, two voices I would swear were the same one. He has had to say so more than once, and for years my answer was a softer version of Fisher’s. It is the same drink, I would tell him. It is the same voice.

I was wrong in the way Fisher was wrong. Aspartame is a different molecule and always was, the way milk poured first makes a different cup and always did. The chemistry cares about the order. Milk poured into tea that is close to boiling meets that heat at the point where it lands, and the whey proteins denature on contact. That is where the faint scalded note comes from. Tea poured over milk brings the milk up gradually and never reaches the same peak. The proteins survive it, and the cup comes out rounder. Eighty years later the Royal Society of Chemistry issued guidance on brewing a proper cup and recommended milk first.

When I said it was the same drink, I was not saying anything about the drink. I was saying that I could not taste the difference, and treating that as the end of the matter. The correction would have cost me one evening. Four glasses of the diet version, four of the original, poured out of his sight, and by the end of it his claim would have had a way to come out right. I design experiments for a living. Roach called for a test the first time the question came up. Fisher, unsatisfied by a result he could not check, spent years building one he could. I have done neither. That the two drinks differ is settled chemistry. Whether he can tell them apart is a claim about him, and it has been sitting at my own table for years, untested.

What Fisher Had to Decide First

The version Fisher settled on meant answering three questions in advance, and every test since has answered the same three, whatever it is testing.

What gets believed for free. The starting assumption is that she is only guessing, also called the null hypothesis. That assumption is the only reason there is a test at all. Standing there with the cooling cup, neither of them could have said what it would look like for Bristol to be right. The guessing story is what supplies the missing piece. Someone who cannot taste at all still sorts the cups, and still gets a share of them right, with a probability that can be worked out in advance. Being able to taste the order of milk and tea now means something specific: improving that chance by more than luck can account for. What Fisher built was not only a way to check her claim. It was the first version of her claim precise enough to check.

Everything that could have happened. If the taster were guessing, there is a fixed list of ways the test could come out, and the list can be written down. Choosing four cups out of eight to call tea first gives seventy possible guesses, every one of them equally likely, and only one of them entirely right.

Seventy is too many to look at whole, so shrink it. Suppose there were only four cups, two poured each way. Guessing means choosing which two she calls tea first, and the other two become milk first on their own, with no separate guess required. Say the true tea-first cups are the second and third. Here is the entire list of what she could have said.

She calls tea first Milk first, by default Correct out of four
1, 2 3, 4 2
1, 3 2, 4 2
1, 4 2, 3 0
2, 3 1, 4 4
2, 4 1, 3 2
3, 4 1, 2 2

Six possibilities, all equally likely if she is guessing, and exactly one of them perfect. That is the shape of the real thing, only smaller. Fisher’s eight cups run the same logic out to seventy possibilities with, again, just one perfect sort.

The shrunken version also shows why four cups would not have been enough to run. One perfect sort in six means a pure guesser gets it exactly right about seventeen percent of the time. A test that small is not a neutral test. She can sort every cup correctly and still be told that one guesser in six manages that much, because the doubt is the thing the design protects. It is a dismissal with a procedure attached. Eight cups is where that stops being true. A perfect sort is then one chance in seventy, which is past what anyone will hand to luck.

How rare the count would be under the default. The count is the number she gets right, a single number standing in for the whole test. Eight out of eight. That puts a pure guesser’s odds at 1 in 70, or about 1.4 percent. That number, the probability of a result at least this extreme when the guessing story is true, is the p-value.

What the p-value does not come with is a verdict. It says how surprising eight out of eight would be if she were guessing. It does not say how surprising is surprising enough, and no amount of arithmetic decides that. Fisher left it there deliberately, and spent the rest of his career defending the choice. He would compute the number for whatever came in and read it fresh, and he refused to name, once and for all, the point where luck stops being a believable story. His own words on it, twenty years later, in Statistical Methods and Scientific Inference: “no scientific worker has a fixed level of significance at which from year to year, and in all circumstances, he rejects hypotheses; he rather gives his mind to each particular case in the light of his evidence and his ideas.” Where that line comes from, and what happens to a test once somebody insists on fixing it in advance, is a fight of its own.

There is a second thing the p-value does not do, and a perfect score is what hides it. Eight out of eight makes it hard to go on believing she was only guessing. Six out of eight would not have made it true. It would have left the test silent, and a silent test says nothing at all about whether Muriel Bristol could tell the difference. He was explicit about the asymmetry in the same chapter of The Design of Experiments that lays out the eight cups. The null hypothesis “is never proved or established, but is possibly disproved, in the course of experimentation.”

The ordinary version of this mistake files a search that found nothing as proof there was nothing to find. Absence of evidence, taken for evidence of absence. My ear does it every time a record comes on. Finding nothing is what an ear like mine does, and it is not the same as there being nothing there.

Why This Story Keeps Getting Told

Statistics does not have many founding stories, and this is the one it kept. I did not know how many times it had been told until I started writing this piece. Box has a version. Salsburg has a version and a book title. Senn worked through the chemistry, Richardson wrote his to sort the others out, Schwab and Starbuck have one from 2024, and the textbook versions are past counting. Every one of those writers got here before I did. The honest question is why write it again.

A claim meets skepticism, gets tested, and comes out vindicated. That pattern is satisfying no matter what sits inside it: tea, a coin flip, a psychic, a suspect in a lineup. A claim is doubted. Someone devises a way to check it. The claim holds. A reader does not need to care about hypothesis testing to feel the pull of that arc. Every retelling of this story is really a retelling of that trope, with statistics standing in as the excuse.

Almost none of the retellings say who she was before that afternoon. Muriel Bristol held a doctorate, a DSc awarded in 1919, and was on the scientific staff at Rothamsted. She had been publishing on algae for years, and a Danish algologist who had read two of her papers would go on to name an entire genus, Muriella, after her. The tellings bring her in as a lady at a tea party with a charming claim, the charm hers and the reasoning his. Many never name her at all. Whatever else was true of the woman holding the cup, her whole training was in noticing small differences and being believed about them. That thread is sitting right there, and almost nobody pulls it.

And almost none of them sit with what her claim actually was. The retellings that reach for the statistics treat the tasting itself as a pretext, something to get past on the way to the design. But a claim like hers is a specific kind of claim: knowledge that lives in performance rather than explanation. Nobody argues their way to knowing whether Bristol could really taste the difference. It can only be tested. That is also true of a husband who can name a singer before the first line finishes, or catch aspartame before the second sip. Neither of us can explain how he knows. The claim only shows up as something that gets tested, or something that gets filed away.

That leaves two things this version can still add: a scientist flattened into a charming guest, and a kind of knowing that keeps getting tested or dismissed without ever being named. The first pulled me toward writing this. The second is why I did not stop once the trope delivered its ending.

A Closing Reflection. Fisher’s first answer took about one second. The test at the table took an afternoon. The one that could settle it took him years. Both of those are still available, every time somebody tells you something about their own experience that does not line up with yours.

  1. A claim someone close to you keeps making about what they can taste, hear, feel, or tell coming, that you have quietly filed as exaggeration without ever once building a way to find out.
  2. Something you can tell that other people around you cannot, that you stopped saying out loud because of the way it landed the last time you did.
  3. A disagreement where you notice you are gathering reasons rather than looking, and where you could not now say what would have changed your mind an hour ago.

Nobody at that research station was obliged to take her seriously. He believed the chemistry was on his side, the room certainly was, and the sentence he reached for first is the sentence almost everyone reaches for first. What he did instead was cheap. Eight cups, one pot, an order she could not see, and it gives a claim something a dismissal never does, which is a way to turn out to be right. Somewhere this week a claim is going to land in front of you that makes no sense given what you already know. Notice how fast the filing happens. Then notice what it would take, in cups, to find out.

Where This Came From

Fisher’s own chapter 2 of The Design of Experiments opens on an unnamed lady who “declares that by tasting a cup of tea made with milk she can discriminate whether the milk or the tea infusion was first added to the cup,” gives no place, no date and no outcome, and leaves the whole thing reading as hypothetical. The Rothamsted afternoon is Box’s addition to it. The line on the null hypothesis quoted in the piece is from that same chapter. The passage on fixed significance levels is from Statistical Methods and Scientific Inference (1956).

Fisher’s reply, “Nonsense, surely it makes no difference,” is quoted by Lindley (1993) from Box. Lindley marks it there as reported rather than transcribed, and uses it in a parenthesis about Fisher’s state of mind rather than in his narration of the event. The popular retellings all paraphrase it differently. No contemporaneous record of the exchange exists. Box was born in 1933, after the afternoon she describes. She does not name her informant for the tea scene. Her preface credits the recollections of William Roach, who worked in the insecticides and antiseptics laboratory at the same station. Bristol had died in 1949, so the account most likely reaches us through the one person at the table with a stake in how it came out. Senn (2012) gives a different date and place, 1950 in Bristol. This piece follows Richardson, the later account, which cites the General Register Office directly. No source fixes the date of the afternoon itself. Box places it soon after Fisher came to Rothamsted in 1919 and shortly before the 1923 marriage. Richardson reads that as the early 1920s, which is where this piece puts it. Richardson also notes that afternoon tea at Rothamsted was a routine staff event rather than a gathering, which is why no crowd appears here. Box calls Bristol an algologist. Phycologist is the same discipline under its more current name. The genus named for her, Muriella, was given by the Danish algologist Johannes Boye Petersen in 1932.

Most retellings merge two events that Box keeps apart: the improvised test at the tea table, and the design questions it left Fisher turning over afterward. The eight cups being “years away” means only that they were published in 1935. Box has him turning the questions over the same afternoon. No source records when he arrived at the design itself. The arithmetic run against Box in this piece is Richardson’s. George Box (1976) gives a different wording of Roach’s verdict, that she made nearly every choice correctly rather than proving her case outright. That puts the afternoon at odds a guesser reaches about one time in four, p = 0.243 in Richardson’s own calculation. Richardson’s own conclusion is blunter than this piece’s: on the record as it stands, Fisher was probably right.

What Richardson records as explicitly Smith’s are the claim of having been present, the identification of Fisher as the man with the Vandyke beard, and the perfect score. The Cambridge staging itself is Salsburg’s narration, which is why this piece says “in that version” rather than putting Cambridge in Smith’s mouth. Salsburg gives no evidence for his late 1920s dating or for Cambridge. Fisher held no post at Cambridge between 1913 and 1943. The one Cambridge appearance Richardson finds on record is a DSc, approved by correspondence in February 1926 and conferred in person on 12 March. That is spring, not the summer afternoon Salsburg’s account needs. Smith was born in 1904. He studied at Edinburgh from 1922 to 1926, then at Cornell from 1926 to 1928. That does not sit comfortably with his having been in the room for Salsburg’s own late 1920s dating, though no source rules it out directly. The gaps this piece gives, forty-five years and thirty-five, are this piece’s own arithmetic, run from Box’s early 1920s dating for the afternoon rather than from Salsburg’s late 1920s one. Richardson gives forty years and another thirty, computed from Salsburg’s own dating. Box is the better supported of the two accounts on every point. This piece follows it. Schwab and Starbuck (2024) is a recent example of the merge, running Roach’s line straight into the eight-cup design and reporting the perfect score, citing Fisher and Salsburg but not Box.

The Royal Society of Chemistry’s guidance was issued in 2003, which is where the “eighty years later” comes from. Senn (2012) works through the same chemistry alongside the statistics.

Intellectual Honesty Note. Muriel Bristol was a real scientist at Rothamsted, and the disagreement over the order of milk and tea is a documented episode, not staged for this piece. Fisher’s reply carries the opening scene and is the hinge the rest of the essay turns on, which is the most weight it can be asked to carry, and it is reported speech rather than a transcript. The opening should be read that way. The piece does not assert that Roach had a stake in the outcome, because Box gives him no motive. It gives the marriage and leaves the reader to draw it, and that inference is mine. Calling the table test thrown together draws on Box’s word extempore, and no source records what was agreed before the pouring started. The taste descriptors, the faint scalded note and the rounder cup, are this piece’s own rendering of the chemistry and not anything a source reports. The four-cup table is a shrunken version built for this piece, not Fisher’s. The personal material is real: my husband can tell an aspartame-sweetened drink from the original, and he has told me more than once that Kate Bush and Joanna Newsom sound nothing alike. I have never run the eight-glass version on him. No test was conducted for this piece.

References

Box, G. E. P. (1976). Science and Statistics. Journal of the American Statistical Association, 71(356), 791-799.

Box, J. F. (1978). R. A. Fisher: The Life of a Scientist. Wiley.

Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd.

Fisher, R. A. (1956). Statistical Methods and Scientific Inference. Oliver and Boyd.

Kendall, M. G. (1963). Ronald Aylmer Fisher, 1890-1962. Biometrika, 50(1-2), 1-15.

Lindley, D. V. (1993). The Analysis of Experimental Data: The Appreciation of Tea and Wine. Teaching Statistics, 15(1), 22-25.

Richardson, J. T. E. (2021). A closer look at the lady tasting tea. Significance, 18(5), 34-37.

Royal Society of Chemistry. (2003). How to Make a Perfect Cup of Tea [Press release].

Salsburg, D. (2001). The Lady Tasting Tea: How Statistics Revolutionized Science in the Twentieth Century. W. H. Freeman.

Schwab, A., & Starbuck, W. H. (2024). How Muriel’s tea stained management research through statistical significance tests. Manuscript.

Senn, S. (2012). Tea for three: Of infusions and inferences and milk in first. Significance, 9(6), 30-33.