R.A. Fisher drew a cup of tea from the urn and offered it to the woman beside him. She declined it. She told him she preferred a cup with the milk poured in first, and that she could tell the difference by taste. Not guess at it. Tell.

The woman was Muriel Bristol, a phycologist. The urn stood at Rothamsted, an agricultural research station in Hertfordshire, set out for the staff’s afternoon tea on some day in the early 1920s.

“Nonsense,” Fisher said. “Surely it makes no difference.” The polite thing at that point is to say how interesting and let it go. Bristol maintained, with emphasis, that of course it did.

Then a voice from just behind them said, “Let’s test her.” It belonged to William Roach, a biochemist at the same station.

They ran the test there and then. Cups were poured at the trestle table, Roach handing them round and writing down what she said. Bristol picked out enough of them correctly, and Roach exulted.

That should have been the end of it. A claim was made, a test was run, and the skeptic was corrected on the spot.

Fisher was not satisfied. The test Roach had thrown together came out in her favor, but it could not rule out a lucky run of guesses. Fisher was already turning over what it would have taken to rule that out. How many cups. Whether they should be paired. In what order they should be presented. What to do about variation in temperature and sweetness. What could be concluded from a perfect score, and what from a score with a single error in it. The answer was years away, and it was published in The Design of Experiments in 1935. Eight cups, four poured tea first and four poured milk first, presented in a random order. The lady is told in advance what the test will consist of, and that there are four of each kind. Her task is to divide the eight into two sets of four.

The design is documented. The tasting is only remembered. The eight cups are Fisher’s own, published under his name. He never named the lady, never said where or when, and never reported how many she got right.

Bristol’s name, the place and the rough date all come from one source: Joan Fisher Box, Fisher’s daughter and biographer. Box has Bristol getting more than enough of the cups right to prove her case. That is Roach’s verdict, spoken by the man who had called for the test and would marry Bristol not long afterward. Her personal triumph, Box adds, was never recorded.

The familiar ending, in which Bristol sorts all eight correctly, is in neither Box nor Fisher. David Salsburg heard it in the late 1960s from Hugh Fairfield Smith, a statistician who said he had been there, and published it in 2001. Forty-five years from the tasting to the telling, thirty-five more from the telling to the page. In that version it is a Cambridge tea party, and the woman has no name at all.

Then there is the arithmetic. No source records how many cups Bristol was handed. John Richardson, who in 2021 set the surviving accounts against the records, runs Fisher’s eight through Box’s account anyway, because it is the only structure on offer: four poured each way, sorted into two piles of four. Under that structure, getting one cup wrong means getting a second one wrong, one from each pile. “More than enough” is not a claim that she got all of them. If she missed any, she missed two at least. Six out of eight is the kind of score that guessing produces about one time in four. Nothing Roach said rules that score out.

Maurice Kendall reports Fisher saying he never performed the tea-tasting experiment at all. On Box’s telling, that is true. Roach ran the cups, and Fisher stood beside the table and watched. The design can be checked and the result cannot. Hold the perfect score loosely, and hold the eight cups tight.

Most claims like Bristol’s never get tested. My husband claims he can taste aspartame. Hand him the diet version of anything, no label and no warning, and he will tell you before the second sip. He can also name the singer before the first line of a song is over, Kate Bush against Joanna Newsom, two voices I would swear were the same one. He has had to say so more than once, and my answer was a softer version of Fisher’s. It is the same drink, I would tell him. It is the same voice.

I was wrong in the way Fisher was wrong. Aspartame is a different molecule and always was, the way milk poured first makes a different cup and always did. The chemistry cares about the order. Milk poured into tea that is close to boiling meets that heat at the point where it lands, and the whey proteins denature on contact. That is where the faint scalded note comes from. Tea poured over milk brings the milk up gradually and never reaches the same peak. The proteins survive it, and the cup comes out rounder. Eighty years after the tasting, the Royal Society of Chemistry issued guidance on brewing a proper cup and recommended milk first.

When I said it was the same drink, I was not saying anything about the drink. I was saying that I could not taste the difference, and treating that as the end of the matter. The correction would have been easy. I could pour one can of the original into four tiny cups and one can of the diet version into four more, out of his sight, and give his claim a way to come out right. I design experiments for a living. Roach called for a test the first time the question came up. Fisher, unsatisfied by a result he could not check, spent years building one he could. I have done neither. That the two drinks differ is settled chemistry. Whether my husband can tell them apart is a claim about him, untested.

Most people run into a claim like his. A patient tells a doctor that this headache is not like the others. A mechanic hears a rattle the owner swears was always there. Some of those claims are right and some are not. A dismissal treats them all alike, and a test is what sorts them.

What Fisher Had to Decide First

The version Fisher settled on meant answering three questions in advance, and every significance test since has answered the same three, whatever it is testing.

What gets believed for free. The starting assumption is that she is guessing, also called the null hypothesis. Without that assumption there would be nothing to test against. Standing there with the cooling cup, neither Fisher nor Roach could have said what it would look like for Bristol to be right. The guessing story is what supplies the missing piece. Someone who cannot tell the difference still sorts the cups, and still gets a share of them right, with a probability that can be worked out in advance. That guesser is a mental model, a small-scale copy of how something works, and every count that follows is measured against it. Being able to taste the order of milk and tea now means something specific: improving that chance by more than luck can account for. What Fisher built was the first version of her claim precise enough to check.

Everything that could have happened. If the taster were guessing, there is a fixed list of ways the test could come out, and the list can be written down. Choosing four cups out of eight to call tea first gives seventy possible guesses, every one of them equally likely, and only one of them entirely right.

Seventy is too many to look at whole, so shrink it. Suppose there were only four cups, two poured each way. Guessing means choosing which two she calls tea first, and the other two become milk first on their own, with no separate guess required. Say the true tea-first cups are the second and third. Here is the entire list of what she could have said.

She calls tea first Milk first, by default Correct out of four
1, 2 3, 4 2
1, 3 2, 4 2
1, 4 2, 3 0
2, 3 1, 4 4
2, 4 1, 3 2
3, 4 1, 2 2

Six possibilities, all equally likely if she is guessing, and exactly one of them perfect. This is the real test, only smaller. Fisher’s eight cups run the same logic out to seventy possibilities with, again, just one perfect sort.

The shrunken version also shows why four cups would not have been enough to run. One perfect sort in six means a pure guesser gets it exactly right about 17 percent of the time. A test that small leans toward the doubter. She can sort every cup correctly and still be told that one guesser in six manages that much. A test like that is a dismissal with a procedure attached. Eight cups is where that stops being true. A perfect sort is then one chance in seventy, rare enough that a guesser would almost never manage it.

How rare the count would be under the default. The count is the number she gets right, a single number standing in for the whole test. Eight out of eight. That puts a pure guesser’s odds at 1 in 70, or about 1.4 percent. That number, the probability of a result at least this extreme when the guessing story is true, is the p-value.

What the p-value does not come with is a verdict. It says how surprising eight out of eight would be if she were guessing. It does not say how surprising is surprising enough, and no amount of arithmetic decides that. Fisher left it there deliberately, and spent the rest of his career defending the choice. He would compute the number for whatever came in and read it fresh, and he refused to name, once and for all, the point where luck stops being a believable story. His own words on it, twenty years later, in Statistical Methods and Scientific Inference: “no scientific worker has a fixed level of significance at which from year to year, and in all circumstances, he rejects hypotheses; he rather gives his mind to each particular case in the light of his evidence and his ideas.” A level of significance is that threshold: the point, set before the data comes in, that a p-value has to clear before the evidence against the null counts as sufficient.

There is a second thing the p-value does not do, and a perfect score is what hides it. Eight out of eight makes it hard to go on believing she was only guessing. Six out of eight would not have settled whether she was guessing. It would have left the test silent, and a silent test says nothing at all about whether Muriel Bristol could tell the difference. Fisher was explicit about the asymmetry in the same chapter of The Design of Experiments that lays out the eight cups. The null hypothesis “is never proved or established, but is possibly disproved, in the course of experimentation.”

Reading a silent test as a verdict is the ordinary mistake. A search that found nothing gets filed as proof there was nothing to find. Absence of evidence, taken for evidence of absence.

Why This Story Keeps Getting Told

Statistics does not have many founding stories, and the eight cups are the one it kept. I did not know how many times it had been told until I started writing this piece. Box has a version. Salsburg has a version and a book title. Senn worked through the chemistry, Richardson wrote his to sort the others out, Schwab and Starbuck have one from 2024, and the textbook versions are past counting. Every one of those writers got here before I did. So why pour the cups again?

A claim meets skepticism, gets tested, and comes out vindicated. That pattern is satisfying no matter what sits inside it: tea, a coin flip, a psychic, a suspect in a lineup. A claim is doubted. Someone devises a way to check it. The claim holds. Most tellings of this story are that trope, with statistics standing in as the excuse.

Almost none of the retellings say who she was before she declined Fisher’s cup. Muriel Bristol held a doctorate, a DSc awarded in 1919, and was on the scientific staff at Rothamsted. She had been publishing on algae for years, and a Danish algologist who had read two of her papers would go on to name an entire genus, Muriella, after her. The tellings bring her in as a lady at a tea party with a charming claim, the charm hers and the reasoning Fisher’s. Many never name her at all. Whatever else was true of the woman holding the cup, her whole training was in noticing small differences and being believed about them. That thread is sitting right there, and almost nobody pulls it.

And almost none of them stay with what her claim was. The ones that reach for the statistics treat the tasting itself as a pretext, something to get past on the way to the design. But a claim like hers is a specific kind of claim: knowledge that lives in performance rather than explanation. Nobody argues their way to knowing whether Bristol could taste the difference. It can only be tested. The same goes for my husband and his aspartame. Neither my husband nor I can explain how he knows. The claim only shows up as something that gets tested, or something that gets filed away.

A Closing Reflection. Fisher’s first answer took about one second. The test at the table took an afternoon. The one that could settle it took him years. All three are still available, every time somebody tells you something about their own experience that does not line up with yours.

  1. The half-smile you give when someone close to you says they can taste, hear, or feel a difference you cannot.
  2. The pause before you mention something you can tell and the people around you cannot, and how your voice drops when you finally do.
  3. The moment in a disagreement when you stop listening and start collecting reasons, and the tightness in your jaw that tells you it happened.

Nobody at that research station was obliged to take Bristol seriously. Fisher believed the chemistry was on his side, and he said so bluntly. Most people reach for the same certainty, just more quietly. What Fisher did instead was simple. Eight cups, one pot, an order she could not see. That design gives a claim something a dismissal never does: a way to turn out to be right. Somewhere this week a claim is going to land in front of you that makes no sense given what you already know. Notice how fast the dismissal happens. Then notice what it would take, in cups, to find out.

Where This Came From

Fisher’s own Chapter 2 of The Design of Experiments opens on an unnamed lady who “declares that by tasting a cup of tea made with milk she can discriminate whether the milk or the tea infusion was first added to the cup,” gives no place, no date and no outcome, and leaves the whole thing reading as hypothetical. The Rothamsted tea table is Box’s addition to it. The line on the null hypothesis quoted in the piece is from that same chapter. “Absence of evidence, taken for evidence of absence” adapts the saying “absence of evidence is not evidence of absence,” usually attributed to the astronomer Martin Rees and made famous by Carl Sagan. The passage on fixed significance levels is from Statistical Methods and Scientific Inference (1956).

Fisher’s reply, “Nonsense, surely it makes no difference,” is quoted by Lindley (1993) from Box. Lindley marks it there as reported rather than transcribed, and uses it in a parenthesis about Fisher’s state of mind rather than in his narration of the event. The popular retellings all paraphrase it differently. No contemporaneous record of the exchange exists.

Box is the source for the station, the urn, the trestle table, and Bristol by name. Roach handing the cups round and writing down her answers is Box as read by Richardson. Kendall’s report of Fisher disowning the experiment is from his obituary of Fisher, published in 1963, the year after Fisher died. Box was born in 1933, after the tasting she describes. She does not name her informant for the tea scene. Her preface credits the recollections of William Roach, who worked in the insecticides and antiseptics laboratory at the same station. Bristol had died in 1949, so the account most likely reaches us through the one person at the table with a stake in how it came out. Senn (2012) gives a different date and place, 1950 in Bristol. This piece follows Richardson, the later account, which cites the General Register Office directly. No source fixes the date of the tasting itself. Box places it soon after Fisher came to Rothamsted in 1919 and shortly before the 1923 marriage. Richardson reads that as the early 1920s, which is where this piece puts it. Richardson also notes that afternoon tea at Rothamsted was a routine staff event rather than a gathering, which is why no crowd appears here. Box calls Bristol an algologist. Phycologist is the same discipline under its more current name. The genus named for her, Muriella, was given by the Danish algologist Johannes Boye Petersen in 1932.

Most retellings merge two events that Box keeps apart: the improvised test at the tea table, and the design questions it left Fisher turning over afterward. The eight cups being “years away” means only that they were published in 1935. Box has him turning the questions over the same afternoon. No source records when he arrived at the design itself. The arithmetic run against Box in this piece is Richardson’s. George Box (1976) gives a different wording of Roach’s verdict, that she made nearly every choice correctly rather than proving her case outright. That puts Bristol’s score at the same odds a guesser reaches about one time in four, p = 0.243 in Richardson’s own calculation. Richardson’s own conclusion is blunter than this piece’s: on the record as it stands, Fisher was probably right.

What Richardson records as explicitly Smith’s are the claim of having been present, the identification of Fisher as the man with the Vandyke beard, and the perfect score. The Cambridge staging itself is Salsburg’s narration, which is why this piece says “in that version” rather than putting Cambridge in Smith’s mouth. Salsburg gives no evidence for his late 1920s dating or for Cambridge. Fisher held no post at Cambridge between 1913 and 1943. The one Cambridge appearance Richardson finds on record is a DSc, approved by correspondence in February 1926 and conferred in person on 12 March. That is spring, not the summer afternoon Salsburg’s account needs. Smith was born in 1904. He studied at Edinburgh from 1922 to 1926, then at Cornell from 1926 to 1928. That does not sit comfortably with his having been in the room for Salsburg’s own late 1920s dating, though no source rules it out directly. The gaps this piece gives, forty-five years and thirty-five, are this piece’s own arithmetic, run from Box’s early 1920s dating for the tasting rather than from Salsburg’s late 1920s one. Richardson gives forty years and another thirty, computed from Salsburg’s own dating. Box is the better supported of the two accounts on every point. This piece follows it. Schwab and Starbuck (2024) is a recent example of the merge, running Roach’s line straight into the eight-cup design and reporting the perfect score, citing Fisher and Salsburg but not Box.

The Royal Society of Chemistry’s guidance was issued in 2003, which is where the “eighty years” comes from. Senn (2012) works through the same chemistry alongside the statistics.

Intellectual Honesty Note. Muriel Bristol was a real scientist at Rothamsted, and the disagreement over the order of milk and tea is a documented episode, not staged for this piece. Fisher’s reply is the hinge the rest of the essay turns on, which is the most weight it can be asked to carry, and it is reported speech rather than a transcript. The exchange at the urn should be read that way. The piece does not assert that Roach had a stake in the outcome, because Box gives him no motive. It gives the marriage and leaves the reader to draw it, and that inference is mine. Calling the table test thrown together draws on Box’s word extempore, and no source records what was agreed before the pouring started. The taste descriptors, the faint scalded note and the rounder cup, are this piece’s own rendering of the chemistry and not anything a source reports. The four-cup table is a shrunken version built for this piece, not Fisher’s. The personal material is real: my husband claims he can tell an aspartame-sweetened drink from the original, and he has told me more than once that Kate Bush and Joanna Newsom sound nothing alike. I have never run the eight-cup version on him. No test was conducted for this piece.

References

Box, G. E. P. (1976). Science and Statistics. Journal of the American Statistical Association, 71(356), 791-799.

Box, J. F. (1978). R. A. Fisher: The Life of a Scientist. Wiley.

Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd.

Fisher, R. A. (1956). Statistical Methods and Scientific Inference. Oliver and Boyd.

Kendall, M. G. (1963). Ronald Aylmer Fisher, 1890-1962. Biometrika, 50(1-2), 1-15.

Lindley, D. V. (1993). The Analysis of Experimental Data: The Appreciation of Tea and Wine. Teaching Statistics, 15(1), 22-25.

Richardson, J. T. E. (2021). A closer look at the lady tasting tea. Significance, 18(5), 34-37.

Royal Society of Chemistry. (2003). How to Make a Perfect Cup of Tea [Press release].

Salsburg, D. (2001). The Lady Tasting Tea: How Statistics Revolutionized Science in the Twentieth Century. W. H. Freeman.

Schwab, A., & Starbuck, W. H. (2024). How Muriel’s tea stained management research through statistical significance tests. Manuscript.

Senn, S. (2012). Tea for three: Of infusions and inferences and milk in first. Significance, 9(6), 30-33.