.
Picking up where I left off in a 2023 post, I will (finally!) return to Gardiner and Zaharos’s discussion of sensitivity in epistemology and its connection to my notion of severity. But before turning to Parts II (and III), I’d better reblog Part I. Here it is:
I’ve been reading an illuminating paper by Georgi Gardiner and Brian Zaharatos (Gardiner and Zaharatos, 2022; hereafter, G & Z), “The safe, the sensitive and the severely tested,” that forges links between contemporary epistemology and my severe testing account. It’s part of a collection published in Synthese on “Recent issues in Philosophy of Statistics”. Gardiner and Zaharatos were among the 15 faculty who attended the 2019 summer seminar in philstat that I ran (with Aris Spanos). The authors courageously jump over some high hurdles separating the two projects (whether a palisade or a ha ha–see G & Z) and manage to bring them into close connection. The traditional epistemologist is largely focused on an analytic task of defining what is meant by knowledge (generally restricted to low-level perceptual claims, or claims about single events) whereas the severe tester is keen to articulate when scientific hypotheses are well or poorly warranted by data. Still, while severity grows out of statistical testing, I intend for the account to hold for any case of error-prone inference. So it should stand up to the examples with which one meets in the jungles of epistemology. For all of the examples I’ve seen so far, it does. I will admit, the epistemologists have storehouses of thorny examples, many of which I’ll come back to. This will be part 1 of two, possible even three, posts on the topic; revisions to this part will be indicated with ii, iii, etc., and no I haven’t used the chatbot or anything in writing this.
I won’t dwell over many differences in goals and language, but focus on the points of contrast that best reveal what severity has to offer the epistemologist. The epistemological notion that is closest in spirit to severity, G & Z propose, is what epistemologists call “sensitivity,”
Mayo has independently developed a sensitivity condition without drawing on the resources of contemporary epistemological theory. She has developed a sensitivity account, without perceiving herself as such. (G & Z, 19)
So am I like the Molière of sensitivity? [1] In fact, severe testing quite consciously gives an account of inference that is sensitive to erroneously inferring (warranting or believing) claims. It is only because statistical methods deliberately supply such tools that I, as a philosopher, look to them in the first place.
It is noteworthy that the authors link severity with “non-formal epistemology” rather than formal epistemology. I think this seems right. Formal epistemology is largely Bayesian or at least “probabilist”, in the sense of using probability to capture degrees of belief, plausibility, or support in hypotheses or claims. However, non-formal epistemologists regularly slip into probabilist talk in speaking of statistical evidence and inference, and this creates obstacles for their own project. Notably, they are all too happy with crude induction: from k% of A’s have been observed to be B’s, to inferring the probability that a specific A is a B equals k. Fallacies of probabilistic instantiation, reference class problems, and lack of randomness loom large, but do not get attention. Following a second big assumption—that a claim’s being probable (in some sense) warrants inferring it—non-formal epistemologists set about to find principles to block such an apparent warrant. But I’m getting ahead of myself. I return to this at the end of the current post.
Severity.
Informal. Here’s an informal take on my notion of severity, weak and a strong. We start with a minimal requirement: we deny there is evidence for claim C if little if anything has been done that could have found flaws in C (weak severity). Claim C is warranted (by data x) just to the extent it has been subjected to, and passes, a test that probably would have found flaws in C, if they are present. This probability is the severity with which C has passed the test. Claims that pass with high severity are said to be warranted (strong severity). The severity accorded a claim is automatically deemed low if it’s impossible to assess the relevant error probabilities, even approximately.[2]
The authors are right to suggest there are echoes of “sensitivity” in epistemology. Very roughly, sensitivity requires, for any claim C that is judged or believed to be true, that if C were false, it would (probably?) not be believed, or judged or the like. However, they also aver that severity faces, or appears to face, a problem thought to bedevil sensitivity in epistemology. If correct, it suggests that skeptical possibilities result in even well-tested claims failing the minimal requirement for evidence. I will show how the severe tester debunks the allegation that sensitivity (in epistemology) is said to face in Part 2.
A full discussion of severe testing may be found in my Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (CUP 2022): STINT. All 16 “tours” of the book may be found, in proof form, in these excerpts on this blog.
Severity in error statistics.
The informal description of severity that I gave above is liable to be misunderstood without first understanding a little bit of how they arise in error statistical methods such as statistical significance tests—even though my severity idea results in reformulating these tools. If severity is misunderstood, it will be misapplied when call on for the epistemological project. So here are some elements from error statistics: In formal error statistics, probability arises to assess and control the probabilty a method leads to misinterpreting data. These are the method’s error probabilities. While, technically, error probabilities of a method allude to its behavior in (actual or hypothetical) repeated use, they may also serve to capture the capabilities of methods to avoid error in the case at hand. This is what allows moving from error probabilities (of a method) to specific warrants (to applications of that method). Or so I have been arguing for some time
Pre-data, a method such as a hypothesis test is specified so as to ensure that true claims are passed or inferred, and false claims are not passed or are rejected—although these claims are qualified probabilistically. For example, in a statistical significance test, the claims of interest might be: the data (say from a randomized control trial) are evidence a given treatment is beneficial for a given disease, or the data fail to provide evidence of benefit (at least in the experimental population). The standard type I error here is to infer evidence of benefit when it is absent (i.e., when the observed effect is merely due to background variability or “noise”); while the type II error is to fail to infer benefit when it exists. The probability a method M “would have found” flaws in claim C refers to the probability method M would have rejected C, computed under various statistical hypotheses. This “computed under…” phrase may be unfamiliar to you, but it’s important. It is not a conditional probability, and does not use a prior probability. I’ve written so much on this blog about statistical tests, that it’s best to just direct you to search this blog. Can it be used in informal epistemology? I say yes.
One need not consider the case a “test” in any official way (e.g., the claim C need not be prespecified). Claim C can be an estimation, prediction, perceptual judgment or other. You don’t have to call the inference from data x to claim C a test. But keeping to testing language underscores that we are interested in how well probed claims are—at least when we’re doing an analysis as we are. (In fact, I dub the severe tester’s use of probability “probativism,” in contrast to “probabilism” and “performance”. (These are briefly discussed in G & Z.) The main thing is that there is a method M that moves from data and background to a claim C (which may be a denial of some claim).
Severity is post-data. The severity notion I develop is a post-data measure. For example, the null hypothesis H0 might assert no benefit, whereas alternatives under H1 assert positive magnitudes of benefit. We rarely want to merely infer evidence of some benefit, but rather how much of a benefit. G & Z illustrate with a figure relating effect size to severity. Once a test result is at hand, and an inference reached with data x, severe testers evaluate how well (or poorly) warranted various claims are. Any inference reached is accompanied with at least one poorly warranted claim (in relation to the relevant errors).
Levels. The severity assessment is always at a “level”—analogous to a significance level or confidence level. The focus is on the level attained post-data. This depends on how well probed the claim actually is with the given data x and method M. In informal settings—the ones we are usually faced with—these error probabilities remain qualitative and so does the level of warrant: poorly probed, reasonably well probed, extremely well probed, etc.. Even in formal statistical contexts, a precise error probability is rarely required. Epistemologists like to talk about beliefs, especially a “subject S believes that P” (for a proposition P). You’re free to do so, although I won’t, except when quoting others. What level of severity to require for an indication, strong evidence, knowledge (or warrant) will vary with the context—but I will not have to specify thresholds here. A method that errs over 50% of the time (in relation to a type of claim) is unreliable.
The kinds of claims that arise in epistemology are dichotomous: rather than consider magnitudes of effects, the possibilities are generally exhausted by C and its denial. That simplifies things.
Counterfactuals in error statistics.
My severity requirement alludes to counterfactuals but, unlike their typical treatment by epistemologists, I don’t cash them out in terms of possible worlds, be they close or distant. In formal statistics, the needed counterfactuals stem from the sampling distribution of a statistic (“computed under” various assumptions about the world giving rise to the data). The statistic, T, called the test statistic, measures the accordance or fit between data x and claim H. The larger T is, the more improbable x is, when computed under H, (and the more probable under ~H). If the data are sufficiently distant from what would be expected under H, the method M outputs ~H. So the probability attaches to the method.
Auditing. A crucial part of the severity requirement is checking the assumptions underlying these claims, which I call auditing. An application of an adequate statistical method may be said to supply nominally good error probabilities, but they may actually be poor if they do not stand up to testing by the required audit. Audits themselves occur at two levels: the primary inference to C, and the secondary scrutiny of assumptions (typically, in statistics, of a model for the data generation).
A passage from Gardiner and Zaharatos (G & Z):
Strong severity aims to characterise the epistemic value of good tests. A good test is good because were H false, the test would have detected it. For observed data e to support a hypothesis H, on Mayo’s view, it does not suffice for e to fit H. In addition, e’s fitting H must be a good test of H. A test is good if H were false, the data wouldn’t fit H. (ibid. 4)
I especially like the first sentence of this passage, because it is often supposed that good error probabilities are for shop-keeping and acceptance sampling in industry. Among the ways I’d wish to qualify the claims in this passage:
Remember the method used by Scott Harkonen? (Mayo 2020; and these two blogposts (Oct 9, 2013 and Feb 1, 2020). Failing to find statistical significance on 10 different endpoints he tries and tries again until finding a subgroup of patients with a sufficiently high number showing benefit from the treatment (for a serious lung disease). Let HPD be the post-data hypothesis Harkonen formulates based on the unblinded data. HPD: the treatment benefits patients with such and such characteristics. Look at what happens in relation to the specially generated null hypothesis: ~HPD: the treatment does not benefit these patients. The data x “do not fit” the post-data null hypothesis ~HPD. The only reason Harkonen selected this postdoc subgroup is that he could declare that the data do not fit the post-designated null hypothesis ~HPD. Using a likelihood ratio as a measure of fit, the data do not fit ~HPD since Pr(x;~HPD) < Pr(x;HPD). The biased selection would not show up in the likelihood ratio, if the associated error probability was not considered. What resources does the epistemologist have to pick up on the fact that the method had high error probabilities—in our sense? I’m not sure, but that’s what I’d like to supply them.
Sensitivity in Epistemology.
In one place, Gardiner and Zaharatos define sensitivity as
S’s true judgement that p is sensitive iff in the nearest possible worlds in which p is not true, S does not judge that p. (G & Z, 13)
I’m not sure if they intended to write “true judgment” here. They drop the possible worlds in the following:
Sensitivity of belief. S’s belief that p is sensitive iff if p were false, S would not believe that p (ibid.)
A judgement that p is sensitive iff were p false, the agent would not have judged that p. This ‘judgement’ might be a legal verdict, scientific conclusion, formal finding, news report or similar. The agent might be a group or community. For some such judgements, an individual’s believing p is not a central or necessary condition that p (ibid.)
That’s good because we want to drop the “subject (or knower) S” from the notion. But we need to qualify with an error probability. An article on sensitivity in legal epistemology, by Enoch et al. (2012, 204), defines sensitivity using probability rather than possible worlds:
Sensitivity: S’s belief that p is sensitive =df. Had it not been the case that p, S would (most probably) not have believed that p.
Nudging their definition closer to severity, one might try:
Claim C, which passes test M, is warranted (with severity?) iff were C false, C would (most probably) not have passed or passed so well, with method M.
I insert “?” because it is at most something that might be tried. A main problem is that I don’t know how probability is being used here. Enoch et al, I take it, construe it as probabilifying the claim C itself, perhaps with a posterior probability. My position would side with Peirce
if universes were as plenty as blackberries, if we could put a quantity of them in a bag, shake them well up, draw out a sample and examine them to see what proportion of them had one arrangement and what proportion another. (2.684)
It still seems odd to my ears to call the claim (belief or judgment) sensitive, rather than the method that outputs the claim. I also worry that insufficient attention is paid to ensuring that true claims pass.
Some Classic Examples.
What allows viewing sensitivity and severity as in the same spirit is considering how they are used in debunking some classic examples. So let’s turn to that.
Lottery paradox. The lottery paradox, for which my friend Henry Kyburg is famous, goes like this:
The evidence, x, is that a person A has bought a ticket in a fair lottery with only a one in a million chance of having her ticket drawn as the winner. While the probability that “A will not win” is high, it does not pass with severity with this evidence. because even if A’s ticket is a winner, there is no chance of finding this out. The paradox that Kyburg was on about is that if it is inferred, for each ticket-holder, that they will not win, then we would infer no one will win, contradicting the supposition that there will be a winner. (Kyburg’s solution denies we should conjoin highly probable claims to infer A1 will not win, A2 will not win etc.) Here, the assumption is that it is a fair lottery, so that each ticket has the same probability of being selected.
The severe tester denies it is warranted to infer A will not win. The improbability of winning is given in the description of the lottery, so nothing has been done to distinguish a winning ticket from a losing ticket. I discuss this in Mayo 1996. (I will want to come back to this case in a later post.)
The authors, G & Z discuss a few other popular examples. Take the example of “prisoners”.
Prisoner (“guilt by association”).
Security footage reveals that ninety-nine prisoners together attack a guard. One prisoner refuses to participate. Prison officials decide that since for each prisoner it is 99% probable they are guilty, they have adequate evidence to successfully prosecute individual prisoners for assault. They charge Ryan, an arbitrarily selected prisoner in the yard, with assault. A guilty verdict is returned. Given the evidence, it is highly probable that Ryan rioted. But convicting Ryan on this evidence seems epistemically inappropriate. (G & Z, 11)
G & Z rightly appeal to severity (or sensitivity) to show the epistemic inappropriateness, but what about the inappropriateness of supposing that “for each prisoner it is 99% probable they are guilty” and “Given the evidence, it is highly probable that Ryan rioted”? You could say that the probability that a randomly selected prisoner rioted is .95, but this does not mean that Ryan, the one selected, has a .95 probability of having gatecrashed (whatever this might mean). Either Ryan is guilty or he isn’t. Moreover, The fact that we can randomly select prisoners does not mean they each had an equal probability of rioting.
Even if one wants to arrive at a method for assigning epistemic warrant to specific claims, based on proportions (it need not be a probability), the problem of the reference class must be dealt with, if the method is to be decently reliable. (Principles of indifference do not suffice.) (I need to return to this type of “naked statistics” in a later post.)
Putting aside this issue, G & Z are right in handling this example: nothing has been done to distinguish Ryan’s guilt from his innocence.
Let me end this post with their excellent sum-up unifying sensitivity and severity:
Unification. The parallel between sensitivity and severe testing is apparent. Sensitivity is not a matter of how probable the claim is given the evidence. A judgement can have very high evidential probability, and yet be insensitive. This is exemplified by the lottery, prisoner, and sex crime examples. Instead sensitivity asks ‘were the claim false, would this falsity be detectable?’ … Severe testing likewise focuses on this subjunctive question: If the claim were wrong, would the fit between the favoured hypothesis and the data be notably weaker? And has anything been done so that were the hypothesis false, the data collected would indicate this falsity? In cases like Prisoner and Lottery, the answer is resoundingly no to both questions. (G & Z, 14)
Gardiner and Zaharatos have managed to put a new spin on severity. Their paper–which I highly recommend–has encouraged me to revisit* the jungles of classic epistemology, at least where severity has something to say. But I don’t expect to start speaking like a native. Stay tuned for part 2. In the mean time, share your thoughts in the comments to this blog.
*Or, more correctly, visit for the first time.
References
Enoch D., Spectre, L. & Fisher, T. (2012). Statistical evidence, sensitivity, and the legal value of knowledge. Philosophy & Public Affairs, 40(3), 197–224.
Gardiner, G., & Zaharatos, B. (2022). The safe, the sensitive, and the severely tested: a unified account. Synthese: An International Journal for Epistemology, Methodology and Philosophy of Science, 200(5). [See class on 4/26 from Mayo’s 2023 graduate seminar (syllabus here).]
Mayo, D. G. (2020). P-Values on Trial: Selective Reporting of (Best Practice Guides Against) Selective Reporting. Harvard Data Science Review 2.1.
Mayo, D. G. (2018). Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars, Cambridge: Cambridge University Press. [See this post of excepts for proofs.]
Mayo, D. G. (1996). Error and the Growth of Experimental Knowledge, Chicago: Chicago University Press, 1996.
Peirce, C. S. (1931-35). Collected papers. Vols. 1-6. Edited by C. Hartshorne and P. Weiss. Cambridge: Harvard University Press.
ENDNOTES
[1] I refer to Molière (1890), for those of you who had to read Le Bourgois Gentilhomme in high school:
My faith! For more than forty years I have been speaking prose while knowing nothing of it, and I am the most obliged person in the world to you for telling me so.
[2] See Statistical Inference as Severe Testing (SIST 2018, 18). Gambits that result in low severity or the inability to assess severity—variants of cherry-picking, P-hacking, data-dredging, and optional stopping– are examples of biasing selection effects (92). Gellerization is another term I’ve used.
2024-25 Cruise
Our second stop in 2025 on the leisurely tour of SIST is Excursion 4 Tour II which you can read here. This criticism of statistical significance tests continues to be controversial, but it shouldn’t be. One should not suppose that quantities measuring different things ought to be equal. At the bottom you will see links to posts discussing this issue, each with a large number of comments. The comments from readers are of interest!
getting beyond…
Excerpt from Excursion 4 Tour II*4.4 Do P-Values Exaggerate the Evidence?
“Significance levels overstate the evidence against the null hypothesis,” is a line you may often hear. Your first question is:
What do you mean by overstating the evidence against a hypothesis?
Several (honest) answers are possible. Here is one possibility:
What I mean is that when I put a lump of prior weight π0 of 1/2 on a point null H0 (or a very small interval around it), the P-value is smaller than my Bayesian posterior probability on H0.
More generally, the “P-values exaggerate” criticism typically boils down to showing that if inference is appraised via one of the probabilisms – Bayesian posteriors, Bayes factors, or likelihood ratios – the evidence against the null (or against the null and in favor of some alternative) isn’t as big as 1 − P.
You might react by observing that: (a) P-values are not intended as posteriors in H0 (or Bayes ratios, likelihood ratios) but rather are used to determine if there’s an indication of discrepancy from, or inconsistency with, H0. This might only mean it’s worth getting more data to probe for a real effect. It’s not a degree of belief or comparative strength of support to walk away with. (b) Thus there’s no reason to suppose a P-value should match numbers computed in very different accounts, that differ among themselves, and are measuring entirely different things. Stephen Senn gives an analogy with “height and stones”:
. . . [S]ome Bayesians in criticizing P-values seem to think that it is appropriate to use a threshold for significance of 0.95 of the probability of the alternative hypothesis being true. This makes no more sense than, in moving from a minimum height standard (say) for recruiting police officers to a minimum weight standard, declaring that since it was previously 6 foot it must now be 6 stone. (Senn 2001b, p. 202)
To top off your rejoinder, you might ask: (c) Why assume that “the” or even “a” correct measure of evidence (relevant for scrutinizing the P-value) is one of the probabilist ones?
All such retorts are valid, and we’ll want to explore how they play out here. Yet, I want to push beyond them. Let’s be open to the possibility that evidential measures from very different accounts can be used to scrutinize each other.
Getting Beyond “I’m Rubber and You’re Glue”. The danger in critiquing statistical method X from the standpoint of the goals and measures of a distinct school Y, is that of falling into begging the question. If the P-value is exaggerating evidence against a null, meaning it seems too small from the perspective of school Y, then Y’s numbers are too big, or just irrelevant, from the perspective of school X. Whatever you say about me bounces off and sticks to you. This is a genuine worry, but it’ s not fatal. The goal of this journey is to identify minimal theses about “ bad evidence, no test (BENT)” that enable some degree of scrutiny of any statistical inference account – at least on the meta-level. Why assume all schools of statistical inference embrace the minimum severity principle? I don’t, and they don’t. But by identifying when methods violate severity, we can pull back the veil on at least one source of disagreement behind the battles.
Thus, in tackling this latest canard, let’ s resist depicting the critics as committing a gross blunder of confusing a P-value with a posterior probability in a null. We resist, as well, merely denying we care about their measure of support. I say we should look at exactly what the critics are on about. When we do, we will have gleaned some short-cuts for grasping a plethora of critical debates. We may even wind up with new respect for what a P-value, the least popular girl in the class, really does.
To visit the core arguments, we travel to 1987 to papers by J. Berger and Sellke, and Casella and R. Berger. These, in turn, are based on a handful of older ones (Cox 1977, E, L, & S 1963, Pratt 1965), and current discussions invariably revert back to them. Our struggles through quicksand of Excursion 3, Tour II, are about to pay large dividends.
This excerpt comes from Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (Mayo, CUP 2018).
Readers can find blogposts that trace out the discussion of this topic, as I was developing it, along with comments. The following 2 are central:
(7/14) “P-values overstate the evidence against the null”: legit or fallacious? (revised) 71 comments
(7/23) Continued:”P-values overstate the evidence against the null”: legit or fallacious? 39 comments
Where you are in the journey.
2024-2025 Cruise
Our first stop in 2025 on the leisurely tour of SIST is Excursion 4 Tour I which you can read here. I hope that this will give you the chutzpah to push back in 2025, if you hear that objectivity in science is just a myth. This leisurely tour may be a bit more leisurely than I intended, but this is philosophy, so slow blogging is best. (Plus, we’ve had some poor sailing weather). Please use the comments to share thoughts.
.
Tour I The Myth of “The Myth of Objectivity”*
Objectivity in statistics, as in science more generally, is a matter of both aims and methods. Objective science, in our view, aims to find out what is the case as regards aspects of the world [that hold] independently of our beliefs, biases and interests; thus objective methods aim for the critical control of inferences and hypotheses, constraining them by evidence and checks of error. (Cox and Mayo 2010, p. 276) [i]
Whenever you come up against blanket slogans such as “no methods are objective” or “all methods are equally objective and subjective” it is a good guess that the problem is being trivialized into oblivion. Yes, there are judgments, disagreements, and values in any human activity, which alone makes it too trivial an observation to distinguish among very different ways that threats of bias and unwarranted inferences may be controlled. Is the objectivity–subjectivity distinction really toothless, as many will have you believe? I say no. I know it’s a meme promulgated by statistical high priests, but you agreed, did you not, to use a bit of chutzpah on this excursion? Besides, cavalier attitudes toward objectivity are at odds with even more widely endorsed grass roots movements to promote replication, reproducibility, and to come clean on a number of sources behind illicit results: multiple testing, cherry picking, failed assumptions, researcher latitude, publication bias and so on. The moves to take back science are rooted in the supposition that we can more objectively scrutinize results – even if it’s only to point out those that are BENT. The fact that these terms are used equivocally should not be taken as grounds to oust them but rather to engage in the difficult work of identifying what there is in “objectivity” that we won’t give up, and shouldn’t.
The Key Is Getting Pushback! While knowledge gaps leave plenty of room for biases, arbitrariness, and wishful thinking, we regularly come up against data that thwart our expectations and disagree with the predictions we try to foist upon the world. We get pushback! This supplies objective constraints on which our critical capacity is built. Our ability to recognize when data fail to match anticipations affords the opportunity to systematically improve our orientation. Explicit attention needs to be paid to communicating results to set the stage for others to check, debate, and extend the inferences reached. Which conclusions are likely to stand up? Where do the weakest parts remain? Don’t let anyone say you can’t hold them to an objective account.
Excursion 2, Tour II led us from a Popperian tribe to a workable demarcation for scientific inquiry. That will serve as our guide now for scrutinizing the myth of the myth of objectivity. First, good sciences put claims to the test of refutation, and must be able to embark on an inquiry to pin down the sources of any apparent effects. Second, refuted claims aren’t held on to in the face of anomalies and failed replications; they are treated as refuted in further work (at least provisionally); well-corroborated claims are used to build on theory or method: science is not just stamp collecting. The good scientist deliberately arranges inquiries so as to capitalize on pushback, on effects that will not go away, on strategies to get errors to ramify quickly and force us to pay attention to them. The ability to register how hunting, optional stopping, and cherry picking alter their error-probing capacities is a crucial part of a method’s objectivity. In statistical design, day-to-day tricks of the trade to combat bias are consciously amplified and made systematic. It is not because of a “disinterested stance” that we invent such methods; it is that we, quite competitively and self-interestedly, want our theories to succeed in the market place of ideas.
Admittedly, that desire won’t suffice to incentivize objective scrutiny if you can do just as well producing junk. Successful scrutiny is very different from success at grants, getting publications and honors. That is why the reward structure of science is so often blamed nowadays. New incentives, gold stars and badges for sharing data and for resisting the urge to cut corners are being adopted in some fields. Fortunately, for me, our travels will bypass lands of policy recommendations, where I have no special expertise. I will stop at the perimeters of scrutiny of methods which at least provide us citizen scientists armor against being misled. Still, if the allure of carrots has grown stronger than the sticks, we need stronger sticks.
Problems of objectivity in statistical inference are deeply intertwined with a jungle of philosophical problems, in particular with questions about what objectivity demands, and disagreements about “objective versus subjective” probability. On to the jungle!
[i] Mayo and Cox (2010), “Objectivity and Conditionality in Frequentist Inference”, is the paper that led me to the critical analysis of Birnbaum on the Likelihood Principle. How could I write on “conditionality” if it leads to renouncing error probabilities? I asked David Cox. We agreed that it did not.
From Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars* (Mayo 2018, CUP)
To see where you are in the book, check the full Itinerary here.If you want to follow us, write to jemille6@vt.edu, for a clean copy of the readings.
.
Remember that old Woody Allen movie, “Midnight in Paris,” where the main character (I forget who plays it, I saw it on a plane), a writer finishing a novel, steps into a cab that mysteriously picks him up at midnight and transports him back in time where he gets to run his work by such famous authors as Hemingway and Virginia Wolf? (It was a new movie when I began the blog in 2011.) He is wowed when his work earns their approval and he comes back each night in the same mysterious cab…Well, ever since I began this blog in 2011, I imagine being picked up in a mysterious taxi at midnight on New Year’s Eve, and lo and behold, find myself in the 1960s New York City, in the company of Allan Birnbaum who is is looking deeply contemplative, perhaps studying his 1962 paper…Birnbaum reveals some new and surprising twists this year! [i]
(The pic on the left is the only blurry image I have of the club I’m taken to.) It has been a decade since I published my article in Statistical Science (“On the Birnbaum Argument for the Strong Likelihood Principle”), which includes commentaries by A. P. David, Michael Evans, Martin and Liu, D. A. S. Fraser, Jan Hannig, and Jan Bjornstad. David Cox, who very sadly did in January 2022, is the one who encouraged me to write and publish it. Not only does the (Strong) Likelihood Principle (LP or SLP) remain at the heart of many of the criticisms of Neyman-Pearson (N-P) statistics and of error statistics in general, but a decade after my 2014 paper, it is more central than ever–even if it is often unrecognized.
OUR EXCHANGE:
ERROR STATISTICIAN: It’s wonderful to meet you Professor Birnbaum; I’ve always been extremely impressed with the important impact your work has had on philosophical foundations of statistics. I happen to have published on your famous argument about the likelihood principle (LP). (whispers: I can’t believe this!)
BIRNBAUM: Ultimately you know I rejected the LP as failing to control the error probabilities needed for my Confidence concept. But you know all this, I’ve read it in your book: Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (STINT, 2018, CUP).
ERROR STATISTICIAN: You’ve read my book? Wow! Then you know I don’t think your argument shows that the LP follows from such frequentist concepts as sufficiency S and the weak conditionality principle WLP. I don’t rehearse my argument there, but I first found the problem in 2006, when I was writing something on “conditioning” with David Cox. [ii] Sorry,…I know it’s famous…
BIRNBAUM: Well, I shall happily invite you to take any case that violates the LP and allow me to demonstrate that the frequentist is led to inconsistency, provided she also wishes to adhere to the WLP and sufficiency (although less than S is needed).
ERROR STATISTICIAN: Well I show that no contradiction follows from holding WCP and S, while denying the LP.
BIRNBAUM: Well, well, well: I’ll bet you a bottle of Elba Grease champagne that I can demonstrate it!
ERROR STATISTICAL PHILOSOPHER: It is a great drink, I must admit that: I love lemons.
BIRNBAUM: OK. (A waiter brings a bottle, they each pour a glass and resume talking). Whoever wins this little argument pays for this whole bottle of vintage Ebar or Elbow or whatever it is Grease.
.
ERROR STATISTICAL PHILOSOPHER: I really don’t mind paying for the bottle.
BIRNBAUM: Good, you will have to. Take any LP violation. Let x’ be 2-standard deviation difference from the null (asserting μ = 0) in testing a normal mean from the fixed sample size experiment E’, say n = 100; and let x” be a 2-standard deviation difference from an optional stopping experiment E”, which happens to stop at 100. Do you agree that:
(0) For a frequentist, outcome x’ from E’ (fixed sample size) is NOT evidentially equivalent to x” from E” (optional stopping that stops at n)
ERROR STATISTICAL PHILOSOPHER: Yes, that’s a clear case where we reject the strong LP, and it makes perfect sense to distinguish their corresponding p-values (which we can write as p’ and p”, respectively). The searching in the optional stopping experiment makes the p-value quite a bit higher than with the fixed sample size. For n = 100, data x’ yields p’= ~.05; while p” is ~.3. Clearly, p’ is not equal to p”, I don’t see how you can make them equal.
BIRNBAUM: Suppose you’ve observed x”, a 2-standard deviation difference from an optional stopping experiment E”, that finally stops at n=100. You admit, do you not, that this outcome could have occurred as a result of a different experiment? It could have been that a fair coin was flipped where it is agreed that heads instructs you to perform E’ (fixed sample size experiment, with n = 100) and tails instructs you to perform the optional stopping experiment E”, stopping as soon as you obtain a 2-standard deviation difference, and you happened to get tails, and performed the experiment E”, which happened to stop with n =100.
ERROR STATISTICAL PHILOSOPHER: Well, that is not how x” was obtained, but ok, it could have occurred that way.
BIRNBAUM: Good. Then you must grant further that your result could have come from a special experiment I have dreamt up, call it a BB-experiment. In a BB-experiment, if the outcome from the experiment you actually performed has an outcome with a proportional likelihood to one in some other experiment not performed, E’, then we say that your result has an “LP pair”. For any violation of the strong LP, the outcome observed, let it be x”, has an “LP pair”, call it x’, in some other experiment E’. In that case, a BB-experiment stipulates that you are to report x” as if you had determined whether to run E’ or E” by flipping a fair coin.
(They fill their glasses again)
ERROR STATISTICAL PHILOSOPHER: You’re saying that if my outcome from trying and trying again, that is, optional stopping experiment E”, with an “LP pair” in the fixed sample size experiment I did not perform, then I am to report x” as if the determination to run E” was by flipping a fair coin (which decides between E’ and E”)?
BIRNBAUM: Yes, and one more thing. If your outcome had actually come from the fixed sample size experiment E’, it too would have an “LP pair” in the experiment you did not perform, E”. Whether you actually observed x” from E”, or x’ from E’, you are to report it as x” from E”.
ERROR STATISTICAL PHILOSOPHER: So let’s see if I understand a Birnbaum BB-experiment: whether my observed 2-standard deviation difference came from E’ or E” (with sample size n) the result is reported as x’, as if it came from E’ (fixed sample size), and as a result of this strange type of a mixture experiment.
BIRNBAUM: Yes, or equivalently you could just report x*: my result is a 2-standard deviation difference and it could have come from either E’ (fixed sampling, n= 100) or E” (optional stopping, which happens to stop at the 100th trial). That’s how I sometimes formulate a BB-experiment.
ERROR STATISTICAL PHILOSOPHER: You’re saying in effect that if my result has an LP pair in the experiment not performed, I should act as if I accept the strong LP and just report it’s likelihood; so if the likelihoods are proportional in the two experiments (both testing the same mean), the outcomes are evidentially equivalent.
BIRNBAUM: Well, but since the BB- experiment is an imagined “mixture” it is a single experiment, so really you only need to apply the weak LP which frequentists accept. Yes? (The weak LP is the same as the sufficiency principle).
ERROR STATISTICAL PHILOSOPHER: But what is the sampling distribution in this imaginary BB- experiment? Suppose I have Birnbaumized my experimental result, just as you describe, and observed a 2-standard deviation difference from optional stopping experiment E”. How do I calculate the p-value within a Birnbaumized experiment?
BIRNBAUM: I don’t think anyone has ever called it that.
ERROR STATISTICAL PHILOSOPHER: I just wanted to have a shorthand for the operation you are describing, there’s no need to use it, if you’d rather I not. So how do I calculate the p-value within a BB-experiment?
BIRNBAUM: You would report the overall p-value, which would be the average over the sampling distributions: (p’ + p”)/2
Say p’ is ~.05, and p” is ~.3; whatever they are, we know they are different, that’s what makes this a violation of the strong LP (given in premise (0)).
ERROR STATISTICAL PHILOSOPHER: So you’re saying that if I observe a 2-standard deviation difference from E’, I do not report the associated p-value p’, but instead I am to report the average p-value, averaging over some other experiment E” that could have given rise to an outcome with a proportional likelihood to the one I observed, even though I didn’t obtain it this way?
BIRNBAUM: I’m saying that you have to grant that x’ from a fixed sample size experiment E’ could have been generated through a BB-experiment.
My this drink is sour!
ERROR STATISTICAL PHILOSOPHER: Yes, I love pure lemon.
BIRNBAUM: Perhaps you’re in want of a gene; never mind.
I’m saying you have to grant that x’ from a fixed sample size experiment E’ could have been generated through a BB-experiment. If you are to interpret your experiment as if you are within the rules of a BB experiment, then x’ is evidentially equivalent to x” (is equivalent to x*). This is premise (1).
ERROR STATISTICAL PHILOSOPHER: But the result would be that the p-value associated with x’ (fixed sample size) is reported to be larger than it actually is (.05), because I’d be averaging over fixed and optional stopping experiments; while observing x” (optional stopping) is reported to be smaller than it is–in both cases because of an experiment I did not perform.
BIRNBAUM: Yes, the BB-experiment computes the P-value in an unconditional manner: it takes the convex combination over the 2 ways the result could have come about.
ERROR STATISTICAL PHILOSOPHER: this is just a matter of your definitions, it is an analytical or mathematical result, so long as we grant being within your BB experiment.
BIRNBAUM: True, (1) plays the role of the sufficiency assumption, but one need not even appeal to sufficiency, it is just a matter of mathematical equivalence.
By the way, I am focusing just on LP violations, therefore, the outcome, by definition, has an LP pair. In other cases, where there is no LP pair, you just report things as usual.
ERROR STATISTICAL PHILOSOPHER: OK, but p’ still differs from p”; so I still don’t how I’m forced to infer the strong LP which identifies the two. In short, I don’t see the contradiction with my rejecting the strong LP in premise (0). (Also we should come back to the “other cases” at some point….)
BIRNBAUM: Wait! Don’t be so impatient; I’m about to get to step (2). Here, let’s toast to the new year: “To Elbar Grease!”
ERROR STATISTICAL PHILOSOPHER: To Elbar Grease!
BIRNBAUM: So far all of this was step (1).
ERROR STATISTICAL PHILOSOPHER: : Oy, what is step 2?
BIRNBAUM: STEP 2 is this: Surely, you agree, that once you know from which experiment the observed 2-standard deviation difference actually came, you ought to report the p-value corresponding to that experiment. You ought NOT to report the average (p’ + p”)/2 as you were instructed to do in the BB experiment.
This gives us premise (2a):
(2a) outcome x”, once it is known that it came from E”, should NOT be analyzed as in a BB- experiment where p-values are averaged. The report should instead use the sampling distribution of the optional stopping test E”, yielding the p-value, p” (~.37). In fact, .37 is the value you give in STINT p. 44 (imagining the experimenter keeps taking 10 more).
ERROR STATISTICAL PHILOSOPHER: So, having first insisted I imagine myself in a Birnbaumized, I mean a BB-experiment, and report an average p-value, I’m now to return to my senses and “condition” in order to get back to the only place I ever wanted to be, i.e., back to where I was to begin with?
BIRNBAUM: Yes, at least if you hold to the weak conditionality principle WCP (of D. R. Cox)—surely you agree to this.
(2b) Likewise, if you knew the 2-standard deviation difference came from E’, then
x’ should NOT be deemed evidentially equivalent to x” (as in the BB experiment), the report should instead use the sampling distribution of fixed test E’, (.05).
ERROR STATISTICAL PHILOSOPHER: So, having first insisted I consider myself in a BB-experiment, in which I report the average p-value, I’m now to return to my senses and allow that if I know the result came from optional stopping, E”, I should “condition” on and report p”.
BIRNBAUM: Yes. There was no need to repeat the whole spiel.
ERROR STATISTICAL PHILOSOPHER: I just wanted to be clear I understood you. Of course, all of this assumes the model is correct or adequate to begin with.
BIRNBAUM: Yes, the LP (or SLP, to indicate it’s the strong LP) is a principle for parametric inference within a given model. So you arrive at (2a) and (2b), yes?
ERROR STATISTICAL PHILOSOPHER: OK, but it might be noted that unlike premise (1), premises (2a) and (2b) are not given by definition, they concern an evidential standpoint about how one ought to interpret a result once you know which experiment it came from. In particular, premises (2a) and (2b) say I should condition and use the sampling distribution of the experiment known to have been actually performed, when interpreting the result.
BIRNBAUM: Yes, and isn’t this weak conditionality principle WCP one that you happily accept?
ERROR STATISTICAL PHILOSOPHER: Well the WCP originally refers to actual mixtures, where one flipped a coin to determine if E’ or E” is performed, whereas, you’re requiring I consider an imaginary Birnbaum mixture experiment, where the choice of the experiment not performed will vary depending on the outcome that needs an LP pair; and I cannot even determine what this might be until after I’ve observed the result that would violate the LP? I don’t know what the sample size will be ahead of time.
BIRNBAUM: Sure, but you admit that your observed x” could have come about through a BB-experiment, and that’s all I need. Notice
(1), (2a) and (2b) yield the strong LP!
Outcome x” from E”(optional stopping that stops at n) is evidentially equivalent to x’ from E’ (fixed sample size n).
ERROR STATISTICAL PHILOSOPHER: Clever, but your “proof” is obviously unsound; and before I demonstrate this, notice that the conclusion, were it to follow, asserts p’ = p”, (e.g., .05 = .3!), even though it is unquestioned that p’ is not equal to p”, that is because we must start with an LP violation (premise (0)).
BIRNBAUM: Yes, it is puzzling, but where have I gone wrong?
(The waiter comes by and fills their glasses; they are so deeply engrossed in thought they do not even notice him.)
ERROR STATISTICAL PHILOSOPHER: There are many routes to explaining a fallacious argument. The one I find most satisfactory is in Mayo (2014). But, given we’ve been partying, here’s a very simple one. What is required for STEP 1 to hold, is the denial of what’s needed for STEP 2 to hold:
Step 1 requires us to analyze results in accordance with a BB- experiment. If we do so, true enough we get:
premise (1): outcome x” (in a BB experiment) is evidentially equivalent to outcome x’ (in a BB experiment):
That is because in either case, the p-value would be (p’ + p”)/2
Step 2 now insists that we should NOT calculate evidential import as if we were in a BB- experiment. Instead we should consider the experiment from which the data actually came, E’ or E”:
premise (2a): outcome x” (in a BB experiment) is/should be evidentially equivalent to x” from E” (optional stopping that stops at n): its p-value should be p”.
premise (2b): outcome x’ (within in a BB experiment) is/should be evidentially equivalent to x’ from E’ (fixed sample size): its p-value should be p’.
If (1) is true, then (2a) and (2b) must be false!
If (1) is true and we keep fixed the stipulation of a BB experiment (which we must to apply step 2), then (2a) is asserting:
The average p-value (p’ + p”)/2 = p’ which is false.
Likewise if (1) is true, then (2b) is asserting:
the average p-value (p’ + p”)/2 = p” which is false
Alternatively, we can see what goes wrong by realizing:
If (2a) and (2b) are true, then premise (1) must be false.
In short your famous argument requires us to assess evidence in a given experiment in two contradictory ways: as if we are within a BB- experiment (and report the average p-value) and also that we are not, but rather should report the actual p-value.
I can render it as formally valid, but then its premises can never all be true; alternatively, I can get the premises to come out true, but then the conclusion is false—so it is invalid. In no way does it show the frequentist is open to contradiction (by dint of accepting S, WCP, and denying the LP).
BIRNBAUM: Yet some people still think it is a breakthrough. I never agreed to go as far as Jimmy Savage wanted me too, namely, to be a Bayesian….
ERROR STATISTICAL PHILOSOPHER: I’ve come to see that clarifying the entire argument turns on defining the WCP. Have you seen my 2014 paper in Statistical Science? The key difference is that in (2014), the WCP is stated as an equivalence, as you intended. Cox’s WCP, many claim, was not an equivalence, going in 2 directions. Slides from a presentation may be found on this blogpost.
BIRNBAUM: Yes, the “monster of the LP” arises from viewing WCP as an equivalence, instead of going in one direction (from mixtures to the known result).
ERROR STATISTICAL PHILOSOPHER: In my 2014 paper (unlike my earlier treatments) I too construe WCP as giving an “equivalence” but there is an equivocation that invalidates the purported move to the LP.
On the one hand, it’s true that if z is known (and known for example to have come from optional stopping), it’s irrelevant that it could have resulted from either fixed sample testing or optional stopping.
But it does not follow that if z is known, it’s irrelevant whether it resulted from fixed sample testing or optional stopping. It’s the slippery slide into this second statement–which surely sounds the same as the first–that makes your argument such a brain buster. (Mayo 2014)
BIRNBAUM: Yes I have seen your 2014 paper! Your Rejoinder to some of the critics is gutsy, to say the least. I’ve also seen the slides on your blog.
ERROR STATISTICAL PHILOSOPHER: Thank you, I’m amazed you follow my blog! I haven’t kept it up that much lately; blogs have fallen out of fashion.
BIRNBAUM: As has inferential statistics it seems–it’s all AI/ML. But I have to admit that CHAT GPT illuminates at least part of your argument as to why my reasoning was flawed.
ERROR STATISTICAL PHILOSOPHER: I never thought to check CHAT GPT on my paper, that’s amazing.
BIRNBAUM: Here is what I found on the Chatbot:
CHAT GPT
Birnbaum’s Argument and the Likelihood Principle
In his 1962 paper, Birnbaum argued that if frequentists accept two principles—sufficiency and weak conditionality—they are logically compelled to accept the likelihood principle. The likelihood principle states that all the evidence in data is contained in the likelihood function, meaning that the sampling distribution (and hence frequentist error probabilities) is irrelevant to evidential assessment….
Error Statistician’s Dilemma
If Birnbaum’s argument is correct, then frequentist methods (which rely on error probabilities) would be rendered irrelevant for assessing evidence. This would make it difficult for frequentists to defend their approach as coherent, particularly in the face of Bayesian methods that naturally adhere to the likelihood principle.
However, Deborah Mayo, in her 2014 work, critiques Birnbaum’s argument, exposing a logical flaw in his alleged proof.
BIRNBAUM: The bot does not get your argument right. The who experience has encouraged me to write the first draft of a completely revised paper, reflecting a large advance in my thinking on this. It’s not quite ready to share….
ERROR STATISTICAL PHILOSOPHER: Wow! I’d love to read it…have you identified the problem? In your last couple of papers, you suggest you’d discovered the flaw in your argument for the LP. Am I right? Even in the discussion of your (1962) paper, you seemed to agree with Pratt that WCP can’t do the job you intend. I just want to know, and won’t share your answer with anyone….
(She notices Birnbaum is holding a paper on long legal-sized yellow sheets filled with tiny hand-written comments, covering both sides.)
Sudden interruption by the waiter:
WAITER: Who gets the tab?
BIRNBAUM: I do. To Elbar Grease! To Severe Testing!
Happy New Year!
BIRNBAUM (looking wistful): Savage, you know, never got off my case about remaining at “the half-way house” of likelihood, and not going full Bayesian. Then I wrote the review about the Confidence Concept as the one rock on a shifting scene… Pratt thought the argument should instead appeal to a Censoring Principle (basically, it doesn’t matter if your instrument cannot measure beyond k units if the measurement you’re making is under k units.)
ERROR STATISTICAL PHILOSOPHER: Yes, but who says frequentist error statisticians deny the Censoring Principle? So back to my question,…you did uncover the flaw in your argument, yes?
WAITER: We’re closing now; shall I call a Taxi?
BIRNBAUM: Yes, yes!
ERROR STATISTICAL PHILOSOPHER: ‘Yes’, you discovered the flaw in the argument, or ‘yes’ to the taxi?
MANAGER: We’re closing now; I’m sorry you must leave.
ERROR STATISTICAL PHILOSOPHER: We’re leaving I just need him to clarify his answer….
BIRNBAUM: I predict that 2025 will be the year that people will finally take seriously your paper from a decade ago!
ERROR STATISTICAL PHILOSOPHER: I’ll drink to that!
Suddenly a large group of people bustle past the manager…it’s all chaos.
Prof. Birnbaum…? Allan? Where did he go? (oy, not again!)
Link to complete discussion:
Mayo, Deborah G. On the Birnbaum Argument for the Strong Likelihood Principle (with discussion & rejoinder).Statistical Science 29 (2014), no. 2, 227-266.
[i] Many links on the strong likelihood principle (LP or SLP) and Birnbaum may be found by searching this blog. Good sources for where to start as well as classic background papers may be found in this blogpost. A link to slides and video of a very introductory presentation of my argument from the 2021 Phil Stat Forum is here.
January 7: “Putting the Brakes on the Breakthrough: On the Birnbaum Argument for the Strong Likelihood Principle” (D.Mayo)
[ii] In 2023 I wrote a paper on Cox’s statistical philosophy. Sadly he died in 2022. (The first David R. Cox Foundations of Statistics Prize, currently given by the ASA on even-numbered years, was awarded to Nancy Reid at the JSM 2023.)
.
I took a side trip to David Cox’s famous “weighing machine” example” a month ago, an example thought to have caused “a subtle earthquake” in foundations of statistics, because knew we’d be coming back to it at the end of December when we revisit the (strong) Likelihood Principle [SLP]. It’s been a decade since I published my Statistical Science article on this, Mayo (2014), which includes several commentators, but the issue is still mired in controversy. It’s generally dismissed as an annoying, mind-bending puzzle on which those in statistical foundations tend to hold absurdly strong opinions. Mostly it has been ignored. Yet I sense that 2025 is the year that people will return to it again, given some recent and soon to be published items. This post gives some background, and collects the essential links that you would need if you want to delve into it. Many readers know that each year I return to the issue on New Year’s Eve…. But that’s tomorrow.
By the way, this is not part of our lesurely tour of SIST. In fact, the argument is not even in SIST, although the SLP (or LP) arises a lot. But if you want to go off the beaten track with me to the SLP conundrum, here’s your opportunity.
What’s it all about? An essential component of inference based on familiar frequentist notions: p-values, significance and confidence levels, is the relevant sampling distribution (hence the term sampling theory, or my preferred error statistics, as we get error probabilities from the sampling distribution). This feature results in violations of a principle known as the strong likelihood principle (SLP). To state the SLP roughly, it asserts that all the evidential import in the data (for parametric inference within a model) resides in the likelihoods. If accepted, it would render error probabilities irrelevant post data.
SLP (We often drop the “strong” and just call it the LP. The “weak” LP just boils down to sufficiency)
For any two experiments E1 and E2 with different probability models f1, f2, but with the same unknown parameter θ, if outcomes x and y (from E1 and E2 respectively) determine the same (i.e., proportional) likelihood function (f1(x; θ) = cf2(*y; θ) for all θ), then x and y are inferentially equivalent (for an inference about θ*).
(What differentiates the weak and the strong LP is that the weak refers to a single experiment.)
Violation of SLP:
Whenever outcomes x and y from experiments E1 and E2 with different probability models f1, f2, but with the same unknown parameter θ, and f1(x; θ) = cf2(*y; θ) for all θ, and yet outcomes x and y have different implications for an inference about θ*.
For an example of a SLP violation, E1 might be sampling from a Normal distribution with a fixed sample size n, and E2 the corresponding experiment that uses an optional stopping rule: keep sampling until you obtain a result 2 standard deviations away from a null hypothesis that θ = 0 (and for simplicity, a known standard deviation). When you do, stop and reject the point null (in 2-sided testing).
The SLP tells us (in relation to the optional stopping rule) that once you have observed a 2-standard deviation result, there should be no evidential difference between its having arisen from experiment E1, where n was fixed, say, at 100, and experiment E2 where the stopping rule happens to stop at n = 100. For the error statistician, by contrast, there is a difference, and this constitutes a violation of the SLP.
———————-
Now for the surprising part: In Cox’s weighing machine example, recall, a coin is flipped to decide which of two experiments to perform. David Cox (1958) proposes something called the Weak Conditionality Principle (WCP) to restrict the space of relevant repetitions for frequentist inference. The WCP says that once it is known which Ei produced the measurement, the assessment should be in terms of the properties of the particular Ei. Nothing could be more obvious.
The surprising upshot of Allan Birnbaum’s (1962) argument is that the SLP appears to follow from applying the WCP in the case of mixture experiments, and so uncontroversial a principle as sufficiency (SP)–although even that has been shown to be optional to the argument, strictly speaking. Were this true, it would preclude the use of sampling distributions. J. Savage calls Birnbaum’s argument “a landmark in statistics” (see [i]).
Although his argument purports that [(WCP and SP) entails SLP], in fact data may violate the SLP while holding both the WCP and SP. Such cases also directly refute [WCP entails SLP].
Binge reading the Likelihood Principle.
If you’re keen to binge read the SLP–a way to break holiday/winter break doldrums–or if it comes up during 2025, I’ve pasted most of the early historical sources below. The argument is simple; showing what’s wrong with it took a long time.
My earliest treatment, via counterexample, is in Mayo (2010)–in an appendix to a paper I wrote with David Cox on objectivity and conditionality in frequentist inference. But the treatment in the appendix doesn’t go far enough, so if you’re interested, it’s best to just check out Mayo (2014) in Statistical Science.[ii] An intermediate paper Mayo (2013) corresponds to a talk I presented at the JSM in 2013.
Interested readers may search this blog for quite a lot of discussion of the SLP including “U-Phils” (discussions by readers) (e.g., here, and here), and amusing notes (e.g., Don’t Birnbaumize that experiment my friend.
This conundrum is relevant to the very notion of “evidence”, blithely taken for granted in both statistics and philosophy. [iii] There’s no statistics involved, just logic and language.My 2014 paper shows the logical problem, but I still think that it will take an astute philosopher of language to adequately classify the linguistic fallacy being committed.
To have a list for binging, I’ve grouped some key readings below.
Classic Birnbaum Papers:
Note to Reader: If you look at the (1962) “discussion”, you can already see Birnbaum backtracking a bit, in response to Pratt’s comments.
Some additional early discussion papers:
Durbin:
There’s also a good discussion in Cox and Hinkley 1974.
Evans, Fraser, and Monette:
Kalbfleisch:
My discussions (also noted above):
[ii] The link Mayo (2014) includes comments on my paper by Bjornstad, Dawid, Evans, Fraser, Hannig, and Martin and Liu, and my rejoinder.
[iii] In Birnbaum’s argument, he introduces an informal, and rather vague, notion of the “evidence (or evidential meaning) of an outcome z from experiment E”. He writes it: Ev(E,z).
In my formulation of the argument, I introduce a new symbol ⇒ to represent a function from a given experiment-outcome pair, (E,z) to a generic inference implication. It (hopefully) lets us be clearer than does Ev.
(E,z) ⇒ InfrE(z) is to be read “the inference implication from outcome z in experiment E” (according to whatever inference type/school is being discussed).
If E is within error statistics, for example, it is necessary to know the relevant sampling distribution associated with a statistic. If it is within a Bayesian account, a relevant prior would be needed.
[iv] I’ve blogged these links in the past; please let me know if any links are broken.
2024 Cruise
We are now at stop 3 on our December leisurely cruise through SIST: Excursion 3 Tour III. I am pasting the slides and video from this session during the LSE Research Seminars in 2020 (from which this cruise derives). (Remember it was early pandemic, and we weren’t so adept with zooming.) The Higgs discussion clarifies (and defends) a somewhat controversial interpretation of p-values. (If you’re interested in the Higgs discovery, there’s a lot more on this blog you can find with the search. Ben Recht recently blogged that the Higgs discovery did not take place. HEP physicists roundly responded. I would omit the section on “capability and severity” were I to write a second edition, while keeping the duality of tests and CIs. Share your remarks in the comments.
.
III. (November 2024) Deeper Concepts: Confidence Intervals and Tests: Higgs’ Discovery:
Reading: SIST: Excursion 3 Tour III
Interested in joining us? Please email Jean Miller (jemille6@vt.edu), with your info, and she will send you a clean copy of the monthly materials.
General Info Items: -References: Captain’s Bibliography
–Souvenirs Meeting 3: N (Rule of Thumb for SEV), O (Interpreting Probable Flukes), L (Beyond Incompatibilist Tunnels), M (Quicksand Takeaway]
-Summaries of 16 Tours (abstracts & keywords)
–Excerpts & Mementos on Error Statistics Philosophy Blog
Slides & Video Links from Meeting 3 of the LSE 2020 seminar:Slides: (PDF)
Video:
https://errorstatistics.com/wp-content/uploads/2024/11/lecture_3_ph500_trimmed-1.mp4
2024 Cruise
Welcome to the December leisurely cruise:Wherever we are sailing, assume that it’s warm. This is an overview of our first set of readings for December from my Statistical Inference as Severe Testing: How to get beyond the statistics wars (CUP 2018): [SIST]–Excursion 3 Tour II–(although I already snuck in one of the examples from 3.4, Cox’s weighing machine). This leisurely cruise is intended to take a whole month to cover one week of readings from my 2020 LSE Seminars, except for December and January which double up.
What do you think of “3.6 Hocus-Pocus: P-values Are Not Error probabilities, Are Not Even Frequentist”? This section refers to Jim Berger’s attempted unification of Jeffreys, Neyman and Fisher in 2003. The unification considers testing 2 simple hypotheses using a random sample from a Normal distribution, computing their two P-values, rejecting whichever gets a smaller P-value, and then computing its posterior probability, assuming each gets a prior of .5. This he calls the “Bayesian error probability”. The result violates what he calls the “frequentist principle”. According to Berger Neyman criticized p-values for violating the frequentist principle (SIST p. 186).
Some snapshots from Excursion 3 tour II.
Excursion 3 Tour II: It’s The Methods, Stupid
Tour II disentangles a jungle of conceptual issues at the heart of today’s statistics wars. The first stop (3.4) unearths the basis for a number of howlers and chestnuts thought to be licensed by Fisherian or N-P tests.* In each exhibit, we study the basis for the joke. Together, they show: the need for an adequate test statistic, the difference between implicationary (i assumptions) and actual assumptions, and the fact that tail areas serve to raise, and not lower, the bar for rejecting a null hypothesis. (Additional howlers occur in Excursion 3 Tour III)
recommended: medium to heavy shovel
Stop (3.5) pulls back the curtain on the view that Fisher and N-P tests form an incompatible hybrid. Incompatibilist tribes retain caricatures of F & N-P tests, and rob each from notions they need (e.g., power and alternatives for F, P-values & post-data error probabilities for N-P). Those who allege that Fisherian P-values are not error probabilities often mean simply that Fisher wanted an evidential not a performance interpretation. This is a philosophical not a mathematical claim. N-P and Fisher tended to use P-values in both ways. It’s time to get beyond incompatibilism. Even if we couldn’t point to quotes and applications that break out of the strict “evidential versus behavioral” split, we should be the ones to interpret the methods for inference, and supply the statistical philosophy that directs their right use.” (p. 181)
strongly recommended: light to medium shovel, thick-skinned jacket
In (3.6) we slip into the jungle. Critics argue that P-values are for evidence, unlike error probabilities, but then aver P-values aren’t good measures of evidence either, since they disagree with probabilist measures: likelihood ratios, Bayes Factors or posteriors. A famous peace-treaty between Fisher, Jeffreys & Bayes promises a unification. A bit of magic ensues! The meaning of error probability changes into a type of Bayesian posterior probability. It’s then possible to say ordinary frequentist error probabilities (e.g., type I & II error probabilities) aren’t error probabilities. We get beyond this marshy swamp by introducing subscripts 1 and 2. Whatever you think of the two concepts, they are very different. This recognition suffices to get you out of quicksand.
required: easily removed shoes, stiff walking stick (review Souvenir M on day of departure)
*Several of these may be found in searching for “Saturday night comedy” on this blog. In SIST, however I trace out the basis for the jokes.
selected key terms and ideas
Howlers and chestnuts of statistical tests
armchair science
Jeffreys tail area criticism
Limb sawing logic
Two machines with different precisions
Weak conditionality principle (WCP)
Conditioning (see WCP)
Likelihood principle
Long run performance vs probabilism
Alphas and p’s
Fisher as behaviorist
Hypothetical long-runs
Freudian metaphor for significance tests
Pearson, on cases where there’s no repetition
Armour-piercing naval shell
Error probability1 and error probability 2Incompatibilist philosophy (F and N-P must remain separate)
Test statistic requirements (p. 159)
Please send me your questions, other key terms to add, and any typos you find, in the comments.
2024 Cruise
.
We’re stopping briefly to consider one of the “chestnuts” in the exhibits of “chestnuts and howlers” in Excursion 3 (Tour II) of my book Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (SIST). It is now 66 years since Cox gave his famous weighing machine example in Sir David Cox (1958)[1]. It’s still relevant So, let’s go back to it, with an excerpt from SIST (pp. 170-173).
Exhibit (vi): Two Measuring Instruments of Different Precisions. Did you hear about the frequentist who, knowing she used a scale that’s right only half the time, claimed her method of weighing is right 75% of the time?
She says, “I flipped a coin to decide whether to use a scale that’s right 100% of the time, or one that’s right only half the time, so, overall, I’m right 75% of the time.” (She wants credit because she could have used a better scale, even knowing she used a lousy one.)
Basis for the joke: An N-P test bases error probability on all possible outcomes or measurements that could have occurred in repetitions, but did not.
As with many infamous pathological examples, often presented as knockdown criticisms of all of frequentist statistics, this was invented by a frequentist, Cox (1958). It was a way to highlight what could go wrong in the case at hand, if one embraced an unthinking behavioral-performance view. Yes, error probabilities are taken over hypothetical repetitions of a process, but not just any repetitions will do. Here’s the statistical formulation.
We flip a fair coin to decide which of two instruments, E1 or E2, to use in observing a Normally distributed random sample Z to make inferences about mean θ. E1 has variance of 1, while that of E2 is 106. Any randomizing device used to choose which instrument to use will do, so long as it is irrelevant to θ. This is called a mixture experiment. The full data would report both the result of the coin flip and the measurement made with that instrument. We can write the report as having two parts: First, which experiment was run and second the measurement: (Ei, z), i = 1 or 2.
In testing a null hypothesis such as θ = 0, the same z measurement would correspond to a much smaller P-value were it to have come from E1 rather than from E2: denote them as p1(z) and p2(z), respectively. The overall significance level of the mixture: [p1(z) + p2(z)]/2, would give a misleading report of the precision of the actual experimental measurement. The claim is that N-P statistics would report the average P-value rather than the one corresponding to the scale you actually used! These are often called the unconditional and the conditional test, respectively. The claim is that the frequentist statistician must use the unconditional test.
Suppose that we know we have observed a measurement from E2 with its much larger variance:
The unconditional test says that we can assign this a higher level of significance than we ordinarily do, because if we were to repeat the experiment, we might sample some quite different distribution. But this fact seems irrelevant to the interpretation of an observation which we know came from a distribution [with the larger variance]. (Cox 1958, p. 361)
Once it is known which Ei has produced z, the P-value or other inferential assessment should be made with reference to the experiment actually run. As we say in Cox and Mayo (2010):
The point essentially is that the marginal distribution of a P-value averaged over the two possible configurations is misleading for a particular set of data. It would mean that an individual fortunate in obtaining the use of a precise instrument in effect sacrifices some of that information in order to rescue an investigator who has been unfortunate enough to have the randomizer choose a far less precise tool. From the perspective of interpreting the specific data that are actually available, this makes no sense. (p. 296)
To scotch his famous example, Cox (1958) introduces a principle: weak conditionality.
Weak Conditionality Principle (WCP): If a mixture experiment (of the aforementioned type) is performed, then, if it is known which experiment produced the data, inferences about θ are appropriately drawn in terms of the sampling behavior in the experiment known to have been performed (Cox and Mayo 2010, p. 296).
It is called weak conditionality because there are more general principles of conditioning that go beyond the special case of mixtures of measuring instruments.
While conditioning on the instrument actually used seems obviously correct, nothing precludes the N-P theory from choosing the procedure “which is best on the average over both experiments” (Lehmann and Romano 2005, p. 394), and it’s even possible that the average or unconditional power is better than the conditional. In the case of such a conflict, Lehmann says relevant conditioning takes precedence over average power (1993b).He allows that in some cases of acceptance sampling, the average behavior may be relevant, but in scientific contexts the conditional result would be the appropriate one (see Lehmann 1993b, p. 1246). Context matters. Did Neyman and Pearson ever weigh in on this? Not to my knowledge, but I’m sure they’d concur with N-P tribe leader Lehmann. Admittedly, if your goal in life is to attain a precise α level, then when discrete distributions preclude this, a solution would be to flip a coin to decide the borderline cases! (See also Example 4.6, Cox and Hinkley 1974, pp. 95–6; Birnbaum 1962, p. 491.)
Is There a Catch?
The “two measuring instruments” example occupies a famous spot in the pantheon of statistical foundations, regarded by some as causing “a subtle earthquake” in statistical foundations. Analogous examples are made out in terms of confidence interval estimation methods (Tour III, Exhibit (viii)). It is a warning to the most behavioristic accounts of testing from which we have already distinguished the present approach. Yet justification for the conditioning (WCP) is fully within the frequentist error statistical philosophy, for contexts of scientific inference. There is no suggestion, for example, that only the particular data set be considered. That would entail abandoning the sampling distribution as the basis for inference, and with it the severity goal. Yet we are told that “there is a catch” and that WCP leads to the Likelihood Principle (LP)!
It is not uncommon to see statistics texts argue that in frequentist theory one is faced with the following dilemma: either to deny the appropriateness of conditioning on the precision of the tool chosen by the toss of a coin, or else to embrace the strong likelihood principle, which entails that frequentist sampling distributions are irrelevant to inference once the data are obtained. This is a false dilemma. Conditioning is warranted to achieve objective frequentist goals, and the [weak] conditionality principle coupled with sufficiency does not entail the strong likelihood principle. The ‘dilemma’ argument is therefore an illusion. (Cox and Mayo 2010, p. 298)
There is a large literature surrounding the argument for the Likelihood Principle, made famous by Birnbaum (1962). Birnbaum hankered for something in between radical behaviorism and throwing error probabilities out the window. Yet he himself had apparently proved there is no middle ground (if you accept WCP)! Even people who thought there was something fishy about Birnbaum’s “proof” were discomfited by the lack of resolution to the paradox. It is time for post-LP philosophies of inference. So long as the Birnbaum argument, which Savage and many others deemed important enough to dub a “breakthrough in statistics,” went unanswered, the frequentist was thought to be boxed into the pathological examples. She is not.
In fact, I show there is a flaw in his venerable argument (Mayo 2010b, 2013a, 2014b). That’s a relief. Now some of you will howl, “Mayo, not everyone agrees with your disproof! Some say the issue is not settled.” Fine, please explain where my refutation breaks down. It’s an ideal brainbuster to work on along the promenade after a long day’s tour. Don’t be dismayed by the fact that it has been accepted for so long. But I won’t revisit it here.
From Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (Mayo 2018, CUP).
Excursion 3 Tour II, pp. 170-173.
If you’re keen to follow our abbreviated cruise, write to Jean Miller (jemille6@vt.edu) and she will send you the final pages of the monthly readings.
Note to the Reader:
Textbooks should not call a claim a theorem if it’s not a theorem, i.e., if there isn’t a proof of it (within the relevant formal system). Yet you will find many statistics texts, and numerous discussion articles, that blithely repeat that the (strong) Likelihood Principle is a theorem, shown to follow if you accept the (WCP) which frequentist error statisticians do.[2] I argue that Allan Birnbaum’s (1962) alleged proof is circular. So, in in 2025, when you find a text that claims the LP is a theorem, provable from the (WEP), please let me know.
If statistical inference follows Bayesian posterior probabilism, the LP follows easily. It’s shown in just a couple of pages of Excursion 1 Tour II (45-6). All the excitement is whether the frequentist (error statistician) is bound to hold it. If she is, then error probabilities become irrelevant to the evidential import of data (once the data are given), at least when making parametric inferences within a statistical model.
The LP was a main topic for the first few years of this blog. That’s because I was still refining an earlier disproof from Mayo (2010), based on giving a counterexample. I later saw the need for a deeper argument which I give in Mayo (2014) in Statistical Science.[3] (There, among other subtleties, the WCP is put as a logical equivalence as intended.)
“It was the adoption of an unqualified equivalence formulation of conditionality, and related concepts, which led, in my 1962 paper, to the monster of the likelihood axiom,” (Birnbaum 1975, 263).
If you’re keen to try your hand at the arguments (Birnbaum’s or mine), you might start with a summary post (based on slides) here, or an intermediate paper Mayo (2013) that I presented at the JSM. It is not included in SIST. It’s a brainbuster, though, I warn you. There’s no real mathematics or statistics involved, it’s pure logic. But it’s very circuitous, which is why the supposed “proof” has stuck around as long as it has.But I’ve always thought that clarifying it fully demanded the expertise of a philosopher of language, but I haven’t found one yet.
[1] Cox 1958 has a different variant of the chestnut.
[2] Note sufficiency is not really needed in the “proof”.
[3] The discussion includes commentaries by Dawid, Evans, Martin and Liu, Hannig, and Bjørnstad–some of whom are very unhappy with me. But I’m given the final word in the rejoinder.
References (outside of the excerpt; for refs within SIST, please see SIST):
Birnbaum, A. (1962), “On the Foundations of Statistical Inference“, Journal of the American Statistical Association 57(298), 269-306.
Birnbaum, A. (1975). Comments on Paper by J. D. Kalbfleisch. Biometrika, 62 (2), 262–264.
Cox, D. R. (1958), “Some problems connected with statistical inference“, The Annals of Mathematical Statistics, 29, 357-372.
Mayo, D. G. (2013) “Presented Version: On the Birnbaum Argument for the Strong Likelihood Principle”, in JSM Proceedings, Section on Bayesian Statistical Science. Alexandria, VA: American Statistical Association: 440-453.
Mayo, D. G. (2014). Mayo paper: “On the Birnbaum Argument for the Strong Likelihood Principle,” Paper with discussion and Mayo rejoinder: Statistical Science29(2) pp. 227-239, 261-266.
November Cruise
Although the numbers used in the introductory example are fine, I’m unhappy with it and seek a replacement–ideally with the same or similar numbers. It is assumed that there is a concern both with inferring larger, as well as smaller, discrepancies than warranted. Actions taken if too high a temperature is inferred would be deleterious. But, given the presentation, the more “serious” error would be failing to report an increase, calling for H0: μ ≥ 150 as the null. But the focus on one-sided positive discrepancies is used through the book, so I wanted to keep to that. I needed a one-sided test with a null value other than 0, and saw an example like this in a book. I think it was ecology. Of course, the example is purely for a simple. numerical illustration. Fortunately, the severity analysis gives the same interpretation of the data regardless of how the test and alternative hypotheses are specified. Still, I’m calling for reader replacements, a suitable reward to be ascertained.
Exhibit (i) N-P Methods as Severe Tests: First Look (Water Plant Accident)
There’s been an accident at a water plant where our ship is docked, and the cooling system had to be repaired. It is meant to ensure that the mean temperature of discharged water stays below the temperature that threatens the ecosystem, perhaps not much beyond 150 degrees Fahrenheit. There were 100 water measurements taken at randomly selected times and the sample mean x computed, each with a known standard deviation σ = 10. When the cooling system is effective, each measurement is like observing X ~ N(150, 102). Because of this variability, we expect different 100-fold water samples to lead to different values of X, but we can deduce its distribution. If each X ~N(μ = 150, 102) then X is also Normal with μ = 150, but the standard deviation of X is only σ/√n= 10/√100 = 1. So X ~ N(μ = 150, 1).
It is the distribution of X that is the relevant sampling distribution here. Because it’s a large random sample, the sampling distribution of X is Normal or approximately so, thanks to the Central Limit Theorem. Note the mean of the sampling distribution of X is the same as the underlying mean, both are μ. The frequency link was created by randomly selecting the sample, and we assume for the moment it was successful. Suppose they are testing:
H0: μ ≤ 150 vs. H1: μ > 150.
The test rule for α = 0.025 is:
Reject H0: iff X > 150 + cασ/√100 = 150 + 1.96(1)=151.96,
since cα = 1.96.
For simplicity, let’s go to the 2-standard error cut-off for rejection:
Reject H0 (infer there’s an indication that μ > 150) iff X ≥ 152.
The test statistic d(x) is a standard Normal variable: Z = √100( X – 150)/10 = X – 150 which, for x = 152 is 2. The area to the right of 2 under the standard Normal is around 0.025.
Now we begin to move beyond the strict N-P interpretation. Say x is just significant at the 0.025 level (x = 152). What warrants taking the data as indicating μ > 150 is not that they’d rarely be wrong in repeated trials on cooling systems by acting this way–even though that’s true. There’s a good indication that it’s not in compliance right now. Why? The severity rationale: Were the mean temperature no higher than 150, then over 97% of the time their method would have resulted in a lower mean temperature than observed. Were it clearly in the safe zone, say μ = 149 degrees, a lower observed mean would be even more probable. Thus, x = 152 indicates some positive discrepancy from H0 (though we don’t consider it rejected by a single outcome). They’re going to take another round of measurements before acting. In the context of a policy action, to which this indication might lead, some type of loss function would be introduced. We’re just considering the evidence, based on these measurements; all for illustrative purposes.
Severity Function:I will abbreviate “the severity with which claim C passes test T with data x“:
SEV(test T, outcome x, claim C).
Reject/Do Not Reject: will be interpreted inferentially, in this case as an indication or evidence of the presence or absence of discrepancies of interest.
Let us suppose we are interested in assessing the severity of C: μ > 153. I imagine this would be a full-on emergency for the ecosystem!
Reject H0. Suppose the observed mean is x = 152, just at the cut-off for rejecting H0:
d(x0) = √100(152 – 150)/10 = 2.
The data reject H0 at level 0.025. We want to compute
SEV(T, x = 152, C: μ > 153).
We may say: “the data accord with C: μ > 153,” that is, severity condition (S-1) is satisfied; but severity requires there to be at least a reasonable probability of a worse fit with C if C is false (S-2). Here, “worse fit with C” means x ≤ 152 (i.e., d(x0) ≤ 2). Given it’s continuous, as with all the following examples, < or ≤ give the same result. The context indicates which is more useful. This probability must be high for C to pass severely; if it’s low, it’s BENT.
We need Pr(X ≤ 152; μ > 153 is false). To say μ > 153 is false is to say μ ≤ 153. So we want Pr(X ≤ 152; μ ≤ 153). But we need only evaluate severity at the point μ = 153, because this probability is even greater for μ < 153:
Pr(X ≤ 152; μ = 153) = Pr(Z ≤ -1) = 0.16.
Here, Z = √100(152 – 153)/10 = -1. Thus SEV(T, x = 152, C: μ > 153) = 0.16. Very low. Our minimal severity principle blocks μ > 153 because it’s fairly probable (84% of the time) that the test would yield an even larger mean temperature than we got, if the water samples came from a body of water whose mean temperature is 153. Table 3.1 gives the severity values associated with different claims, given x = 152. Call tests of this form T+
In each case, we are making inferences of form: μ > μ1 = 150 + γ, for different values of γ. To merely infer μ > 150 , the severity is 0.97 since Pr(X ≤ 152; μ = 150) = Pr(Z ≤ 2) = 0.97. While the data give an indication of non-compliance, μ > 150, to infer C: μ > 153 would be making mountains out of molehills. In this case, the observed difference just hit the cut-off for rejection. N-P tests leave things at that coarse level in computing power and the probability of a Type II error, but severity will take into account the actual outcome. Table 3.2 gives the severity values associated with different claims, given x = 153.
If “the major criticism of the Neyman-Pearson frequentist approach” is that it fails to provide “error probabilities fully varying with the data” as J. Berger alleges, (2003, p.6) then, we’ve answered the major criticism.
Non-rejection. Now suppose x = 151, so the test does not reject H0. The standard formulation of N-P (as well as Fisherian) tests stops there. But we want to be alert to a fallacious interpretation of a “negative” result: inferring there’s no positive discrepancy from μ = 150. No (statistical) evidence of non-compliance isn’t evidence of compliance, here’s why. We have (S-1): the data “accord with” H0, but what if the test had little capacity to have alerted us to discrepancies from 150? The alert comes by way of “a worse fit” with H0–namely, a mean x > 151. Condition (S-2) requires us to consider Pr(X > 151; μ = 150), which is only 0.16. To get this, standardize X to obtain a standard Normal variate: Z = √100(151 – 150)/10 = 1; and Pr(X > 151; μ = 150) = 0.16. Thus, SEV(T+, x = 151, C: μ ≤ 150) = low(0.16). Table 3.3 gives the severity values associated with different inferences of form: μ ≤ μ1= 150 + γ, given x* = 151.
Can they at least say that x = 151 is a good indication that μ ≤ 150.5? No, SEV(T+, x = 151, C: μ ≤ 150.5) ≅ 0.3, [Z = 151 – 150.5 = 0.5]. But x = 151 is a good indication that μ ≤ 152 and μ ≤ 153 (with severity indications of 0.84 and 0.97, respectively).
You might say, assessing severity is no different from what we would do with a judicious use of existing error probabilities. That’s what the severe tester says. Formally speaking, it may be seen merely as a good rule of thumb to avoid fallacious interpretations. What’s new is the statistical philosophy behind it. We no longer seek either probabilism or performance, but rather using relevant error probabilities to assess and control severity.5
5Initial developments of the severity idea were Mayo (1983, 1988, 1991, 1996). In Mayo and Spanos (2006, 2011), it was developed much further.
NOTE: I will set out some quiz examples of severity in the next week for practice.
*There is a typo in the book here, it has “-” rather than “>”
You can find the beginning of this section (3.2), the development of N-P tests, in this post.
To read further, see Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (Mayo 2018, CUP).
Where you are in the journey:
Excursion 3: Statistical Tests and Scientific Inference
Tour I Ingenious and Severe Tests 119
3.1 Statistical Inference and Sexy Science: The 1919
Eclipse Test 121
3.2 N-P Tests: An Episode in Anglo-Polish Collaboration 131
YOU exhibit (i) N-P Methods as Severe Tests: First Look (Water Plant Accident)
3.3 How to Do All N-P Tests Do (and more) While
a Member of the Fisherian Tribe 146
Interested in joining us? Please email Jean Miller (jemille6@vt.edu), with your info, and she will send you a clean copy of the monthly materials.
Neyman & Pearson
November Cruise: 3.2
This second of November’s stops in the leisurely cruise of SIST aligns well with my recent Neyman Seminar at Berkeley. Egon Pearson’s description of the three steps in formulating tests is too rarely recognized today. Note especially the order of the steps. Share queries and thoughts in the comments.
3.2 N-P Tests: An Episode in Anglo-Polish Collaboration*
We proceed by setting up a specific hypothesis to test, H0 in Neyman’s and my terminology, the null hypothesis in R. A. Fisher’s . . . in choosing the test, we take into account alternatives to H0 which we believe possible or at any rate consider it most important to be on the look out for . . .Three steps in constructing the test may be defined:
Step 1. We must first specify the set of results . . .
Step 2. We then divide this set by a system of ordered boundaries . . .such that as we pass across one boundary and proceed to the next, we come to a class of results which makes us more and more inclined, on the information available, to reject the hypothesis tested in favour of alternatives which differ from it by increasing amounts.
Step 3. We then, if possible, associate with each contour level the chance that, if H0 is true, a result will occur in random sampling lying beyond that level . . .
In our first papers [in 1928] we suggested that the likelihood ratio criterion, λ, was a very useful one . . . Thus Step 2 proceeded Step 3. In later papers [1933–1938] we started with a fixed value for the chance, ε, of Step 3 . . . However, although the mathematical procedure may put Step 3 before 2, we cannot put this into operation before we have decided, under Step 2, on the guiding principle to be used in choosing the contour system. That is why I have numbered the steps in this order. (Egon Pearson 1947, p. 173)
In addition to Pearson’s 1947 paper, the museum follows his account in “The Neyman–Pearson Story: 1926–34” (Pearson 1970). The subtitle is “Historical Sidelights on an Episode in Anglo-Polish Collaboration”!
We meet Jerzy Neyman at the point he’s sent to have his work sized up by Karl Pearson at University College in 1925/26. Neyman wasn’t that impressed:
Neyman found . . . [K.]Pearson himself surprisingly ignorant of modern mathematics. (The fact that Pearson did not understand the difference between independence and lack of correlation led to a misunderstanding that nearly terminated Neyman’s stay . . .) (Lehmann 1988, p. 2)
Thus, instead of spending his second fellowship year in London, Neyman goes to Paris where his wife Olga (“Lola”) is pursuing a career in art, and where he could attend lectures in mathematics by Lebesque and Borel. “[W]ere it not for Egon Pearson [whom I had briefly met while in London], I would have probably drifted to my earlier passion for [pure mathematics]” (Neyman quoted in Lehmann 1988, p. 3).
What pulled him back to statistics was Egon Pearson’s letter in 1926. E. Pearson had been “suddenly smitten” with doubt about the justification of tests then in use, and he needed someone with a stronger mathematical background to pursue his concerns. Neyman had just returned from his fellowship years to a hectic and difficult life in Warsaw, working multiple jobs in applied statistics.
[H]is financial situation was always precarious. The bright spot in this difficult period was his work with the younger Pearson. Trying to find a unifying, logical basis which would lead systematically to the various statistical tests that had been proposed by Student and Fisher was a ‘big problem’ of the kind for which he had hoped . . . (ibid., p. 3)
….Interim pages 131-6 (in proofs) are here.
Historical Sidelight. Except for short visits and holidays, their work proceeded by mail. When Pearson visited Neyman in 1929, he was shocked at the conditions in which Neyman and other academics lived and worked in Poland. Numerous letters from Neyman describe the precarious position in his statistics lab: “You may have heard that we have in Poland a terrific crisis in everything” [1931] (C. Reid 1998, p. 99). In 1932, “I simply cannot work; the crisis and the struggle for existence takes all my time and energy” (Lehmann 2011, p. 40). Yet he managed to produce quite a lot. While at the start, the initiative for the joint work was from Pearson, it soon turned in the other direction with Neyman leading the way.
By comparison, Egon Pearson’s greatest troubles at the time were personal: He had fallen in love “at first sight” with a woman engaged to his cousin George Sharpe, and she with him. She returned the ring the very next day, but Egon still gave his cousin two years to win her back (C. Reid 1998, p. 86). In 1929, buoyed by his work with Neyman, Egon finally declares his love and they are set to be married, but he let himself be intimidated by his father, Karl, deciding “that I could not go against my family’s opinion that I had stolen my cousin’s fiancée . . . at any rate my courage failed” (ibid., p. 94). Whenever Pearson says he was “suddenly smitten” with doubts about the justification of tests while gazing on the fruit station that his cousin directed, I can’t help thinking he’s also referring to this woman (ibid., p. 60). He was lovelorn for years, but refused to tell Neyman what was bothering him.
…..….Interim pages 137-9 are here.
Performance versus Severity Construals of Tests
“The work [of N-P] quite literally transformed mathematical statistics” (C. Reid 1998, p. 104). The idea that appraising statistical methods revolves around optimality (of some sort) goes viral. Some compared it “to the effect of the theory of relativity upon physics” (ibid.). Even when the optimal tests were absent, the optimal properties served as benchmarks against which the performance of methods could be gauged. They had established a new pattern for appraising methods, paving the way for Abraham Wald’s decision theory, and the seminal texts by Lehmann and others. The rigorous program overshadowed the more informal Fisherian tests. This came to irk Fisher. Famous feuds between Fisher and Neyman erupted as to whose paradigm would reign supreme. Those who sided with Fisher erected examples to show that tests could satisfy predesignated criteria and long-run error control while leading to counterintuitive tests in specific cases. That was Barnard’s point on the eclipse experiments (Section 3.1): no one would consider the class of repetitions as referring to the hoped-for 12 photos, when in fact only some smaller number were usable. We’ll meet up with other classic chestnuts as we proceed.
N-P tests began to be couched as formal mapping rules taking data into “reject H0” or “do not reject H0” so as to ensure the probabilities of erroneous rejection and erroneous acceptance are controlled at small values, independent of the true hypothesis and regardless of prior probabilities of parameters. Lost in this behavioristic formulation was how the test criteria naturally grew out of the requirements of probative tests, rather than good long-run performance. Pearson underscores this in his paper (1947) in the epigraph of Section 3.2: Step 2 comes before Step 3. You must first have a sensible distance measure. Since tests that pass muster on performance grounds can simultaneously serve as probative tests, the severe tester breaks out of the behavioristic prison. Neither Neyman nor Pearson, in their applied work, was wedded to it. Where performance and probativeness conflict, probativeness takes precedent. Two decades after Fisher allegedly threw Neyman’s wood models to the floor (Section 5.8), Pearson (1955) tells Fisher: “From the start we shared Professor Fisher’s view that in scientific enquiry, a statistical test is ‘a means of learning’” (p. 206):
. . . it was not till after the main lines of this theory had taken shape with its necessary formalization in terms of critical regions, the class of admissible hypotheses, the two sources of error, the power function, etc., that the fact that there was a remarkable parallelism of ideas in the field of acceptance sampling became apparent. Abraham Wald’s contributions to decision theory of ten to fifteen years later were perhaps strongly influenced by acceptance sampling problems, but that is another story. (ibid., pp. 204–5)
In fact, the tests as developed by Neyman–Pearson began as an attempt to obtain tests that Fisher deemed intuitively plausible, and this goal is easily interpreted as that of computing and controlling the severity with which claims are inferred. Not only did Fisher reply encouragingly to Neyman’s letters during the development of their results, it was Fisher who first informed Neyman of the split of K. Pearson’s duties between himself and Egon, opening up the possibility of Neyman’s leaving his difficult life in Poland and gaining a position at University College in London. Guess what else? Fisher was a referee for the all-important N–P 1933 paper, and approved of it.
To Neyman it has always been a source of satisfaction and amusement that his and Egon’s fundamental paper was presented to the Royal Society by Karl Pearson, who was hostile and skeptical of its contents, and favorably refereed by the formidable Fisher, who was later to be highly critical of much of the Neyman–Pearson theory. (C. Reid 1998, p. 103)
…To read further, see Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars* (Mayo 2018, CUP).
Where you are in the journey:
Excursion 3: Statistical Tests and Scientific Inference
Tour I Ingenious and Severe Tests 119
3.1 Statistical Inference and Sexy Science: The 1919
Eclipse Test 121
3.2 N-P Tests: An Episode in Anglo-Polish Collaboration 131
YOU
3.3 How to Do All N-P Tests Do (and more) While
a Member of the Fisherian Tribe 146
Interested in joining us? Please email Jean Miller (jemille6@vt.edu), with your info, and she will send you a clean copy of the monthly materials.
November Cruise
This first excerpt for November is really just the preface to 3.1. Remember, our abbreviated cruise this fall is based on my LSE Seminars in 2020, and since there are only 5, I had to cut. So those seminars skipped 3.1 on the eclipse tests of GTR. But I want to share snippets from 3.1 with current readers, along with reflections in the comments. (I promise, I’ve even numbered them below)
Excursion 3 Statistical Tests and Scientific Inference
Tour I Ingenious and Severe Tests
[T]he impressive thing about [the 1919 tests of Einstein’s theory of gravity] is the risk involved in a prediction of this kind. If observation shows that the predicted effect is definitely absent, then the theory is simply refuted.The theory is incompatible with certain possible results of observation – in fact with results which everybody before Einstein would have expected. This is quite different from the situation I have previously described, [where] . . . it was practically impossible to describe any human behavior that might not be claimed to be a verification of these [psychological] theories. (Popper 1962, p. 36)
The 1919 eclipse experiments opened Popper’ s eyes to what made Einstein’ s theory so different from other revolutionary theories of the day: Einstein was prepared to subject his theory to risky tests.[1] Einstein was eager to galvanize scientists to test his theory of gravity, knowing the solar eclipse was coming up on May 29, 1919. Leading the expedition to test GTR was a perfect opportunity for Sir Arthur Eddington, a devout follower of Einstein as well as a devout Quaker and conscientious objector. Fearing “ a scandal if one of its young stars went to jail as a conscientious objector,” officials at Cambridge argued that Eddington couldn’ t very well be allowed to go off to war when the country needed him to prepare the journey to test Einstein’ s predicted light deflection (Kaku 2005, p. 113).
The museum ramps up from Popper through a gallery on “ Data Analysis in the 1919 Eclipse” (Section 3.1) which then leads to the main gallery on origins of statistical tests (Section 3.2). Here’ s our Museum Guide:
According to Einstein’ s theory of gravitation, to an observer on earth, light passing near the sun is deflected by an angle, λ , reaching its maximum of 1.75″ for light just grazing the sun, but the light deflection would be undetectable on earth with the instruments available in 1919. Although the light deflection of stars near the sun (approximately1 second of arc) would be detectable, the sun’ s glare renders such stars invisible, save during a total eclipse, which “ by strange good fortune” would occur on May 29, 1919 (Eddington [1920] 1987, p. 113).
There were three hypotheses for which “ it was especially desired to discriminate between” (Dyson et al. 1920 p. 291). Each is a statement about a parameter, the deflection of light at the limb of the sun (in arc seconds): λ = 0″ (no deflection), λ = 0.87″ (Newton), λ = 1.75″ (Einstein). The Newtonian predicted deflection stems from assuming light has mass and follows Newton’ s Law of Gravity. The difference in statistical prediction masks the deep theoretical differences in how each explains gravitational phenomena. Newtonian gravitation describes a force of attraction between two bodies; while for Einstein gravitational effects are actually the result of the curvature of spacetime. A gravitating body like the sun distorts its surrounding spacetime, and other bodies are reacting to those distortions.
Where Are Some of the Members of Our Statistical Cast of Characters in 1919? In 1919, Fisher had just accepted a job as a statistician at Rothamsted Experimental Station. He preferred this temporary slot to a more secure offer by Karl Pearson (KP), which had so many strings attached – requiring KP to approve everything Fisher taught or published – that Joan Fisher Box writes: After years during which Fisher “ had been rather consistently snubbed” by KP, “It seemed that the lover was at last to be admitted to his lady’ s court – on conditions that he first submit to castration” (J. Box 1978, p. 61). Fisher had already challenged the old guard. Whereas KP, after working on the problem for over 20 years, had only approximated “the first two moments of the sample correlation coefficient; Fisher derived the relevant distribution, not just the first two moments” in 1915 (Spanos 2013a). Unable to fight in WWI due to poor eyesight, Fisher felt that becoming a subsistence farmer during the war, making food coupons unnecessary, was the best way for him to exercise his patriotic duty.
In 1919, Neyman is living a hardscrabble life in a land alternately part of Russia or Poland, while the civil war between Reds and Whites is raging. “It was in the course of selling matches for food” (C. Reid 1998, p. 31) that Neyman was first imprisoned (for a few days) in 1919. Describing life amongst “roaming bands of anarchists, epidemics” (ibid., p. 32), Neyman tells us,“existence” was the primary concern (ibid., p. 31). With little academic work in statistics, and “ since no one in Poland was able to gauge the importance of his statistical work (he was ‘sui generis,’ as he later described himself)” (Lehmann 1994, p. 398), Polish authorities sent him to University College in London in 1925/1926 to get the great Karl Pearson’ s assessment. Neyman and E. Pearson begin work together in 1926. Egon Pearson, son of Karl, gets his B.A. in 1919, and begins studies at Cambridge the next year, including a course by Eddington on the theory of errors. Egon is shy and intimidated, reticent and diffi dent, living in the shadow of his eminent father, whom he gradually starts to question after Fisher’ s criticisms. He describes the psychological crisis he’ s going through at the time Neyman arrives in London: “ I was torn between conflicting emotions: a. finding it difficult to understand R.A.F., b. hating [Fisher] for his attacks on my paternal ‘ god,’ c. realizing that in some things at least he was right” (C. Reid 1998, p. 56). As far as appearances amongst the statistical cast: there are the two Pearsons: tall, Edwardian, genteel; there’ s hardscrabble Neyman with his strong Polish accent and small, toothbrush mustache; and Fisher: short, bearded, very thick glasses, pipe, and eight children. Let’ s go back to 1919, which saw Albert Einstein go from being a little known German scientist to becoming an international celebrity.
As noted at the start, my LSE sectures skipped Section 3.1, but there are things in it I’d like to talk to readers about.
3.1 Statistical Inference and Sexy Science: The 1919 Eclipse Test
p. 121 …….I get the impression that statisticians consider there to be a world of difference between statistical inference and appraising large-scale theories in “glamorous” or “sexy science.” The way it actually unfolds, which may not be what you find in philosophical accounts of theory change, revolves around local data analysis and statistical inference. Even large-scale, sexy theories are made to connect with actual data only by intermediate hypotheses and models. To falsify, or even provide anomalies, for a large-scale theory like Newton’s, we saw, is to infer “falsifying hypotheses,” which are statistical in nature….
p. 122 …There are two key stages of inquiry corresponding to two questions within the broad umbrella of auditing an inquiry:
(i) is there a deflection effect of the amount predicted by Einstein as against Newton (the “Einstein effect”)?
(ii) is it attributable to the sun’s gravitational field as described in Einstein’s hypothesis?
A distinct third question, “higher” in our hierarchy, in the sense of being more theoretical and more general, is: is GTR an adequate account of gravity as a whole?…… Comment 3.1.1.
p. 123…The problem in (i) is reduced to a statistical one: the observed mean deflections (from sets of photographs) are Normally distributed around the predicted mean deflection .
The proper way to frame this as a statistical test is to choose one of the values as H0 and define composite H1 to include alternative values of interest. For instance, the Newtonian “half deflection” can specify H0: μ ≤ 0.87 , and the H1: μ > 0.87 includes the Einsteinian value of 1.75
p. 124…A text by Ghosh et al. (2010, p. 48) presents the Eddington results as a twosided Normal test of Normal test of H0: μ = 1.75 (the Einstein value) vs. H1:≠ 1.75, with a lump of prior probability given to the point null. If any theoretical prediction were to get a lump at this stage, it is Newton’s. …Comment 3.1.2
p. 125 …Some Popperian Confusions About Falsification and Severe Tests
Popper lauds GTR as sticking its neck out, bravely being ready to admit its falsity were the deflection effect not found (1962, pp. 36-7). Even if no deflection effect had been found in the 1919 experiments, it would have been blamed on the sheer difficulty in discerning so small an effect. This would have been entirely correct. Yet many Popperians, perhaps Popper himself, get this wrong. Listen to Popperian Meehl:
[T]he stipulation beforehand that one will be pleased about substantive theory when the numerical results come out as forecast, but will not necessarily abandon it when they do not, seems on the face of it to be about as blatant a violation of the Popperian commandment as you could commit. For the investigator, in a way, is doing … what astrologers and Marxists and psychoanalysts allegedly do, playing ‘heads I win, tails you lose.’ (Meehl 1978, p. 821)
There is a confusion here, and it’s rather common. …Here’s how the severity requirement cashes this out…Comment 3.1.3
p. 127… Big Picture Inference: Can Other Hypotheses Explain the Observed Deflection?
Even to the extent that they had found a deflection effect, it would have been fallacious to infer the effect “attributable to the sun’s gravitational field.” The question (ii) must be tackled: A statistical effect is not a substantive effect. Addressing the causal attribution demands the use of the eclipse data as well as considerable background information. Here we’re in the land of “big picture” inference: the inference is “given everything we know”. In this sense, the observed effect is used and is “non-novel” (in the use-novel sense). Once the deflection effect was known, imprecise as it was, it had to be used. Deliberately seeking a way to explain the eclipse effect while saving Newton’s Law of Gravity from falsification isn’t the slightest bit pejorative – so long as each conjecture is subject to severe test. …
It’s Not How Plausible, but How Well Probed…
p. 129…Souvenir I: So What Is a Statistical Test, Really?
So what’s in a statistical test? First there is a question or problem, a piece of which is to be considered statistically, either because of a planned experimental design, or by embedding it in a formal statistical model. There are (A) hypotheses, and a set of possible outcomes or data; (B) a measure of accordance or discordance, fit, or misfit, between possible answers (hypotheses) and data; and (C) an appraisal of a relevant distribution associated with . Since we want to tell what’s true about tests now in existence, we need an apparatus to capture them, while also offering latitude to diverge from their straight and narrow paths. Comment 3.1.3
…To read further, see Tour I Ex3 TI (full proofs) of Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (CUP, 2018)
Where you are in the journey:
Excursion 3: Statistical Tests and Scientific Inference
Tour I Ingenious and Severe Tests 119
YOU
3.1 Statistical Inference and Sexy Science: The 1919
Eclipse Test 121
3.2 N-P Tests: An Episode in Anglo-Polish Collaboration 131
3.3 How to Do All N-P Tests D (and more) While
a Member of the Fisherian Tribe 146
.
We continue our leisurely tour of Statistical Inference as Severe Testing [SIST] (Mayo 2018, CUP) with Excursion 3. This is based on my 5 seminars at the London School of Economics in 2020; I include slides and video for those who are interested. (use the comments for questions)
November’s Leisurely Tour: N-P and Fisherian Tests, Severe Testing
Reading:SIST: Excursion 3 Tour I (focus on pages up to p. 152): 3.1, 3.2, 3.3
Optional: Excursion 2 Tour II pp. 92-100 (Sections 2.4-2.7)
Quick refresher on means, variance, standard deviations, the Normal distribution, standard normal
Slides & Video Links for November (from my LSE Seminar)Slides:Meeting #2 main slides (PDF)
Supplemental slides (Likelihoodist vs. Significance Tester w/ Bernoulli Trials) (PDF)
Video
Interested in joining us? Please email Jean Miller (jemille6@vt.edu), with your info, and she will send you a clean copy of the monthly materials.
-References: Captain’s Bibliography
–Souvenirs: October Meeting : A-D; November Meeting: (E) An Array of Questions, Problems, Models, (I) So What Is a Statistical Test, Really?, (J) UMP Tests, (K) Probativism
[Souvenirs from optional pages–they’re free: (F) Getting Free of Popperian Constraints on Language, (G) The Current State of Play in Psychology, (H) Solving Induction Is Showing Methods with Error Control]
.
In this post, I consider the questions posed for my (October 9) Neyman Seminar by Philip Stark, Distinguished Professor of Statistics at UC Berkeley. We didn’t directly deal with them during the panel discussion following my talk, and I find some of them a bit surprising. (Other panelist’s questions are here).
Philip Stark asks:
When and how did Statistics lose its way and become (largely) a mechanical way to bless results rather than a serious attempt to avoid fooling ourselves and others?
These are important and highly provocative questions! To a large extent, Stark and other statisticians would be the ones to address them. As an outsider, and as a philosopher of science, I will merely analyze these questions. and in so doing raise some questions about them. That’s Part I of this post. In Part II, I will list some of Stark’s replies to #5 in his (2018) joint paper with Andrea Saltelli “Cargo-cult statistics and scientific crisis”. (The full paper is relevant for #1-4 as well.)
Part I. Some Questions Provoked by Stark’s Questions
There may be a blurring these days between blaming methods as incapable of performing their job, as opposed to their misuse and abuse. It seems clear that Stark would not be asking in Question 5 “What can academic statisticians do to help get the train back on the tracks?” if he thought statistical methods themselves were corrupt. As I read Stark, he means not that the methods themselves merely provide holy water, but that, driven by perverse incentives, celebrity culture, researcher flexibility and the like, many(?) researchers are led to misuse them so that, in effect, they serve merely as holy water to bless results. If so, it’s not that Statistics lost its way, but that many statistical inquiries are unsound or unscientific. This makes his position (as I understand it) importantly different from what I took Ben Recht to be claiming in his recent blogpost, which I discuss and reply to in an earlier blogpost. [i]
For Stark, academic statisticians can help get the train back on the tracks by solving some of the problems he lists of statistics instruction (Question 4, whose parts I label (a)-(c)), although (d) and (e) point more to weak sciences, flexible methods, perverse incentives and moral failings.
When understanding, care, and honesty become valued less than novelty, visibility, scale, funding, and salary, science is at risk. … bad science outcompetes better science. (Stark and Saltelli, 2018)
Bad science violates my minimal severity requirement. (We don’t have evidence for a claim if little if anything has been done to probe the ways it can be wrong.) Stark’s work, to his credit, has used error statistical methods to weed out and advance severe error probes of weak and insevere statistical inferences. But here are some questions on his questions:
Question 1. Do statisticians worry that declaring “Statistics has become corrupt” will be taken as further grist for the mills of the movement to abandon significance, and even, in some quarters, to downplay instruction in statistical inference methods? While Stark may have all of Statistics in mind, my talk focused on error statistical tests, and they are the methods most often blamed for cookbook, sciency statistics, so I keep to them. Are statisticians concerned about how it might sound to a graduate student to say: “Statistics is (largely) corrupt: would you like to study for a Ph.D in Statistics?” Might overly harsh self-criticism, especially (ironiclly) of methods designed for self-criticism, weaken those methods in the meta-statistical competition between rival schools or philosophies of statistics?
I begin my book, Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars CUP, 2018, 3)[i]:
It is easy to lie with statistics. Or so the cliché goes. It is also very difficult to uncover these lies without statistical methods – at least of the right kind. Self-correcting statistical methods are needed, and, with minimal technical fanfare, that’s what I aim to illuminate.[ii]
Self-critical methods can also be ,inadvertently, self-destructive, if criticism isn’t handled in a constructive manner in appraising competing methods. Here’s what I mean:
Question 2. Isn’t it possible that the more self-critical and honest methods, in their willingness to concede defects due to perverse incentives, may lose out in the battle with less self-critical and less self-effacing methodologies? Brad Efron (1998) is right to say the frequentist (error statistician) is the pessimist, who worries that “if anything can go wrong it will,” while the Bayesian optimistically assumes if anything can go right it will (Efron 1998, p. 99). Suppose error statisticians hold themselves to a uniquely high standard, deeply skeptical of findings that lack valid error probabilities. They scrutinize models, acknowledging that multiple testing, optional stopping, data-dredging and other biasing selection effects can readily undermine the integrity of these probabilities. Suppose that when statisticians from other schools criticize their tests for allowing illicit p-value, the error statistician confesses: “Yes, we are largely corrupt,” while Bayesians or other statisticians reply: “I agree, you are corrupt, but we’re not!”
How do statisticians ensure that methods that open themselves to criticism for severe testing violations are not replaced by tools that are less capable of error control? I asked a question in my Neyman seminar as to whether accounts of evidence that are insensitive to error probabilities somehow escape the consequences of biasing selection effects. We did not discuss it, but my answer was: not for a severe tester. Good science requires being able to apply severity at the meta-level–in scrutinizing inferences and arguments about which methods to use for given problems. The same kind of social and ethical conflicts that Stark so aptly uncovers are operative here. Here, I think Statistics might be ideally positioned to call them out, although they rarely do. [iii]
Question 3. How, specifically, might statistical instructors tackle the issues raised in Stark’s Question 4? (See Part II.) Question 4(c) notes the need for a correct elucidation of the meaning of concepts. However, it seems to me that the kind of strictly correct, but shallow and unenthusiastic, recitals of definitions of things like p-values that we often see scarcely help. I say that a genuine appreciation of their value would also be required. If one adopts, as a default, a pessimistic standpoint, it can be an obstacle to communicating what error-statistical thinking is all about. Another obstacle to a non-equivocal interpretation, even if only unconscious, is being wedded to a notion of evidence and inference at odds with the one underlying statistical tests.
In Regina Nuzzo’s “Tips for communicating p-values”(2018) she says
One little-known requirement… is that all analyses and results be presented, no matter the outcome. Yes, these seem like strange, nit-picky rules, but they’re part of the deal when using p-values. (Nuzzo, 2018)
It’s a cool paper, (discussed in my blogpost), but are these “strange, nit-picky rules”? Are other methods for the job of distinguishing genuine free spurious statistical effects free of such rules? Effective communication requires something like an “enthusiastic grasp” of what the tools can accomplish, and how violating the “strange” rules permit being fooled by randomness. Statisticians might be better able to describe what I’m after. The type of passion-driven statistics is found in Stark’s own work.
What do readers think? I now turn to some of Stark’s suggestions in Stark and Saltelli (2018).
Part II. What Can Statisticians Do? Some suggestions from Stark and Saltelli (2018)
- Statisticians can help with important, controversial issues with immediate consequences for society. We can help fight power asymmetries in the use of evidence. We can stand up for the responsible use of statistics, even when that means taking personal risks.
- We should be vocally critical of cargo-cult statistics, including where study design is ignored, where p-values, confidence intervals and posterior distributions are misused… We should be critical even when the abuses involve politically charged issues, such as the social cost of climate change. …
- We can insist that “service” courses foster statistical thinking, deep understanding, and appropriate scepticism, rather than promulgating cargo-cult statistics. We can help empower individuals to appraise quantitative information critically – to be informed, effective citizens of the world. …
- When we appraise each other’s work in academia, we can ignore impact factors, citation counts, and the like: they do not measure importance, correctness, or quality. We can pay attention to the work itself, rather than the masthead of the journal in which it appeared, the press coverage it received, or the funding that supported it. We can insist on evidence that the work is correct – on reproducibility and replicability – rather than pretend that editors and referees can reliably vet research by proxy when the requisite evidence was not even submitted for scrutiny.
- We can decline to referee manuscripts that do not include enough information to tell whether they are correct. We can commit to working reproducibly, to publishing code and data, and generally to contributing to the intellectual commons.
- And we can be of service. Direct involvement of statisticians on the side of citizens in societal and environmental problems can help earn the justified trust of society. …
These are laudable goals in sync with that of severity. The authors are to be credited for how they have implemented them in their work. Please share your reactions to them, as well as the 4 questions I raise, in the comments.
[i] I don’t claim to be clear on Recht’s position because I’m unsure of Recht’s reply to the queries I raised on his posts. But it became clear in our (blog) discussion that he regards statistical tests, by which he means statistical significance tests, as serving important regulatory control of error rates. (See my earlier blogpost.)
[ii] After the Feynman quote about bending over backwards.
[iii] Stark is an exception. An example is his comment on my editorial “The statistics wars and intellectual conflicts of interest.”
Ship Statinfasst
We are starting on Tour II of Excursion 1 (4th stop). The 3rd stop is in an earlier blog post. As I promised, this cruise of SIST is leisurely. I have not yet shared new reflections in the comments–but I will!
Where YOU are in the journey:
.
1.4 The Law of Likelihood and Error Statistics
If you want to understand what’s true about statistical inference, you should begin with what has long been a holy grail–to use probability to arrive at a type of logic of evidential support–and in the first instance you should look not at full-blown Bayesian probabilism, but at comparative accounts that sidestep prior probabilities in hypotheses. An intuitively plausible logic of comparative support was given by the philosopher Ian Hacking (1965)–the Law of Likelihood. Fortunately, the Museum of Statistics is organized by theme, and the Law of Likelihood and the related Likelihood Principle is a big one.
Law of Likelihood (LL):Data x are better evidence for hypothesis H1 than for H0 if x is more probable under H1 than under H0: Pr(x; H1) > Pr(x; H0) that is, the likelihood ratio LR of H1 over H0 exceeds 1.
H0 andH1 are statistical hypotheses that assign probabilities to the values of the random variable X. A fixed value of X is written x0, but we often want to generalize about this value, in which case, following others, I use x. The likelihood of the hypothesis H, given data x, is the probability of observing x, under the assumption that H is true or adequate in some sense. Typically, the ratio of the likelihood of H1 over H0 also supplies the quantitative measure of comparative support. Note when Xis continuous, the probability is assigned over a small interval around X to avoid probability 0.
Does the Law of Likelihood Obey the Minimal Requirement for Severity?
Likelihoods are vital to all statistical accounts, but they are often misunderstood because the data are fixed and the hypothesis varies. Likelihoods of hypotheses should not be confused with their probabilities. Two ways to see this. First, suppose you discover all of the stocks in Pickrite’s promotional letter went up in value (x)–all winners. A hypothesis H to explain this is that their method always succeeds in picking winners. H entails x, so the likelihood of H given x is 1. Yet we wouldn’t say H is therefore highly probable, especially without reason to put to rest that they culled the winners post hoc. For a second way, at any time, the same phenomenon may be perfectly predicted or explained by two rival theories; so both theories are equally likely on the data, even though they cannot both be true.
Suppose Bristol-Roach, in our Bernoulli tea tasting example, got two correct guesses followed by one failure. The observed data can be represented as x0 =<1,1,0>. Let the hypotheses be different values for θ, the probability of success on each independent trial. The likelihood of the hypothesis H0 : θ = 0.5, given x0, which we may write as Lik(0.5), equals (½)(½)(½) = 1/8. Strictly speaking, we should write Lik(θ;x0), because it’s always computed given data x0; I will do so later on. The likelihood of the hypothesis θ = 0.2 is Lik(0.2)= (0.2)(0.2)(0.8) = 0.032. In general, the likelihood in the case of Bernoulli independent and identically distributed trials takes the form: Lik(θ)= θs(1- θ)f, 0< θ<1, where s is the number of successes and f the number of failures. Infinitely many values for θ between 0 and 1 yield positive likelihoods; clearly, then likelihoods do not sum to 1, or any number in particular. Likelihoods do not obey the probability calculus.
The Law of Likelihood (LL) will immediately be seen to fail on our minimal severity requirement – at least if it is taken as an account of inference. Why? There is no onus on the Likelihoodist to predesignate the rival hypotheses – you are free to search, hunt, and post-designate a more likely, or even maximally likely, rival to a test hypothesis H0
Consider the hypothesis that θ = 1 on trials one and two and 0 on trial three. That makes the probability of x maximal. For another example, hypothesize that the observed pattern would always recur in three-trials of the experiment (I. J. Good said in his cryptoanalysis work these were called “kinkera”). Hunting for an impressive fit, or trying and trying again, one is sure to find a rival hypothesis H1 much better “supported” than H0 even when H0 is true. As George Barnard puts it, “there always is such a rival hypothesis, viz. that things just had to turn out the way they actually did” (1972, p. 129).
Note that for any outcome of n Bernoulli trials, the likelihood of H0 : θ = 0.5 is (0.5)n, so is quite small. The likelihood ratio (LR) of a best-supported alternative compared to H0 would be quite high. Since one could always erect such an alternative,
() Pr(LR in favor of H1 over H0; H0*) = maximal.
Thus the LL permits BENT evidence. The severity for H1 is minimal, though the particular H1 is not formulated until the data are in hand.I call such maximally fitting, but minimally severely tested, hypotheses Gellerized, since Uri Geller was apt to erect a way to explain his results in ESP trials. Our Texas sharpshooter is analogous because he can always draw a circle around a cluster of bullet holes, or around each single hole. One needn’t go to such an extreme rival, but it suffices to show that the LL does not control the probability of erroneous interpretations.
What do we do to compute ()? We look beyond the specific observed data to the behavior of the general rule or method, here the LL. The output is always a comparison of likelihoods. We observe one outcome, but we can consider that for any outcome, unless it makes H0 maximally likely, we can find an H1 that is more likely. This lets us compute the relevant properties of the method: its inability to block erroneous interpretations of data. As always, a severity assessment is one level removed: you give me the rule, and I consider its latitude for erroneous outputs. We’re actually looking at the probability distribution of the rule, over outcomes in the sample space. This distribution is called a sampling distribution.* It’s not a very apt term, but nothing has arisen to replace it. For those who embrace the LL, once the data are given, it’s irrelevant what other outcomes could have been observed but were not. Likelihoodists say that such considerations make sense only if the concern is the performance of a rule over repetitions, but not for inference from the data. Likelihoodists hold to “the irrelevance of the sample space” (once the data are given). This is the key contrast between accounts based on error probabilities (error statistical) and logics of statistical inference.
To continue reading Excursion 1 Tour II, go here.
Reader questions are invited in the comments.
Interested in joining us? Please email Jean Miller (jemille6@vt.edu), with your info, and she will send you a clean copy of the monthly materials starting November.
__________
This excerpt comes from Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (Mayo, CUP 2018).
Blurbs of all 16 Tours can be found here.
Giordano, Snow, Yu, Stark, Recht
My Neyman Seminar in the Statistics Department at Berkeley was followed by a lively panel discussion including 4 Berkeley faculty, orchestrated by Ryan Giordano (Dept of Statistics):
I want to share the many fascinating (and difficult) questions put forward by panelists earlier on the day of my talk. Of course, once the live discussion began it took on a life of its own, and we rarely looked back at this list. I’m only now returning to many of those that we didn’t cover–and I’m interested in reader comments. I’m extremely grateful to the organizers, panelists, and audience for creating such a uniquely enriching and provocative exchange of ideas–one that is sure to be continued!
Philip Stark:
Snow Zhang
Bin Yu
Ben Recht
I’ve had a few ‘blogologues’ with Recht recently on his blog and mine, e.g., here, here, and here.
Slides from my presentation are in my previous post. Please share your thoughts, and proposed replies, in the comments.
.
There was a very valuable panel discussion after my October 9 Neyman Seminar in the Statistics Department at UC Berkeley. I want to respond to many of the questions put forward by the participants (Ben Recht, Philip Stark, Bin Yu, Snow Zhang) that we did not address during that panel. Slides from my presentation, “Severity as a basic concept of philosophy of statistics” are at the end of this post (but with none of the animations). I begin in this post by responding to Ben Recht, a professor of Artificial Intelligence and Computer Science at Berkeley, and his recent blogpost, What is Statistics’ Purpose? On severe testing, regulation, and butter passing, on my talk. I will consider: (1) A complex or leading question; (2) Why I chose to focus about Neyman’s philosophy of statistics and (3) What the “100 years of fighting and browbeating” were/are all about.
(1) A complex or leading question.
A question Recht submitted to the panel is this:
Even if the epistemological value of statistical tests is highly questionable, is it reasonable to use statistical tests as benchmarks for regulatory approval (say for drugs or policy)?
Where has it been shown that the epistemological value of statistical tests is highly questionable? I follow Recht in using “statistical tests” to refer to statistical significance tests. One can affirm the value of statistical tests for regulatory approval, as he does, while denying the allegations some have voiced that tests don’t also have epistemological value. The criticisms of statistical tests run to type. Either they are based on:
(A) misinterpretations or misuses of tests, or
(B) assuming notions of evidence and inference that are at odds with those underlying statistical tests.
If Recht knows of others that don’t fit under these umbrellas, I’d be interested to hear. The best known examples of (A) are: interpreting p-values as posterior probabilities, unwarranted moves from statistical to substantive claims and magnitude errors, taking no evidence against H0 as evidence for it, and illicit error probabilities due to biasing selection effects (e.g., multiple testing, optional stopping, cherry picking, outcome switching, etc.) For the main examples of (B), see slides 58-60 of my talk.
Recht is right to stress that “statistics asks many different kinds of questions”, and that statistical (significance) tests are only a small part of a rich methodology I dub error statistics (itself a proper subset of statistics), but this does not show they lack epistemological value for the problems they are intended to address. Severity makes this explicit.
The fallacy of statistical affirming the consequent
I realize that there are misuses of statistical tests that, in some circles, are so baked-in that they are thought to actually be licensed by tests, notably the view that rejecting a statistical test or null hypothesis H0 warrants a substantive research claim H*. In places, Recht seems to associate my notion of severe testing with this illicit notion, often associated with something called NHST, but this is precisely what severe testing denies. He refers to Paul Meehl:
In Meehlian language, the verisimilitude of claims can be tested by experiment, and we want to demonstrate a “damned strange coincidence” to convince ourselves that our claim is causally associated with our experiment. For falsificationists, the more surprising the experimental outcome, the more we are assured our claim resembles the truth.
It is important to emphasize that the only claim severely tested by dint of statistically significant results is the denial of the test or null hypothesis H0. A claim C is severely tested only by passing a test that C would probably have failed if it is false. That statistically significant results would be very surprising (i.e., very improbable) under the assumption that H0, say that there’s no positive effect, does not warrant with severity the truthlikeness of some alternative claim C (however one likes to define versimilitude) even if C “explains” or entails the effect. A falsificationist would not be “assured our claim resembles the truth” simply because there’s evidence of a genuine effect, not due to chance. Finding a genuine discrepancy from H0 only gives evidence of an incompatibility with or discrepancy from H0—of the particular sort the test is probing.
Recht has a series of interesting blogposts on Meehl, so I might note that when Meehl and the “damned strange coincidence” comes up in my Statistical Inference as Severe Testing: How to get beyond the statistics wars [SIST] (CUP, 2018), I emphasize the difference between ruling out coincidence and finding evidence for research hypothesis H*.
For the corroboration to be strong, we have to have ‘Popperian risk’, … ‘severe test’ [as in Mayo], or what philosopher Wesley Salmon called a highly improbable coincidence [“damn strange coincidence”]. (Meehl and Waller 2002, p. 284)
Yet we mustn’t blur an argument from coincidence merely to a real effect, and one that underwrites arguing from coincidence to research hypothesis H*.
…Meehl’s critiques [of NHST] rarely mention the methodological falsificationism of Neyman and Pearson. Why is the field that cares about power—which is defined in terms of N-P tests—so hung up on simple significance tests?…With N-P tests, the statistical alternative to the null hypothesis is made explicit: the null and alternative exhaust the possibilities. There can be no illicit jumping of levels from statistical to causal (from H1 to H*) Fisher didn’t allow the illicit move either, but he was less explicit. (SIST 95-6)
That’s one big reason Neyman-Pearson (N-P) tests are so relevant for today’s hand-wringing about illicit statistical affirming the consequent.
Is Recht suggesting that the logic of statistical significance tests purports to warrant evidence for the truth of a substantive alternative H*? Something often called NHST is portrayed as warranting the illicit inference from statistical to substantive claims, and that is why, in my discussion with Gelman, I say we should move away from “NHST”–a term which was never an official designation in statistical testing. Also, NHST is often described as using the point nil null. The severe tester (in sync with both Neyman and with Cox) would consider one-sided tests, or two one-sided tests with an adjustment for selection.
We are not merely interested in inferring the existence of a discrepancy; we generally wish to infer those discrepancies (or population effective sizes) that are well or poorly tested. The severe tester considers a number of alternatives to the reference hypothesis H0, and reports severity curves (as in slide #62. ) Note that severity decreases as the test’s power at the corresponding alternative increases.[1]
(2) On why I chose to speak about Neyman.
In answer to Recht’s question as to why I focus on Neyman (also Fisher, Cox, Lehmann, Pearson) rather than some other great statisticians let me give 4 main reasons:
(1) I am giving a Neyman seminar,
(2) Neyman’s (and Pearson’s) development of Fisherian tests avoid the central fallacy that Recht is on about (moving from rejecting H0 to inferring a substantive claim H*), and
(3) Neyman’s construal of tests in terms of error statistical performance provides a crucial strategy that takes us beyond the traditional problems of induction, and also beyond Popper (although Popper could have used N-P tests to flesh out his notion of “methodological falsification”). Recht is in favor of the performance, acceptance-sampling philosophy. I don’t think it goes far enough for using tests to make inferences with stringency. A fourth reason that I didn’t take up in my talk:
(4) Neyman’s applied papers provide important insights about the value of statistical models for finding out true things, even though the models themselves are at best approximations. They still enable adequate error probability control. I especially like his discussion of a conjecture and refutation exercise of a model for pest control (SIST, exhibit xii, p. 299 in Excursion 4 Tour iv).
Formal statistical tests give us error probabilities defined in terms of the sampling distribution of a test statistic. The more Fisherian construal focuses on the attained p-value: “pobs is the probability that we would mistakenly declare there to be evidence against H0, were we to regard the data under analysis as just decisive against H0.” (Cox and Hinkley 1974, 66.) (See slide #32.) Neyman-Pearson view a statistical test as a rule that maps observed values of an appropriate test statistic d(x) into either “reject H0” or “do not reject H0” in such a way that there is a low probability of erroneously rejecting H0and a much higher probability of correctly rejecting H0. “Reject H0” and “fail to reject H0” are generally interpreted as x is evidence against H0, or x fails to provide evidence against H0 (which is not the same as evidence for H0.)
Contrary to what is often supposed, N-P testers also advocate reporting attained p-values post-data. Erich Lehmann, Neyman’s first Ph.D student at Berkeley makes this clear. See my slides (#30-31) from Lehmann’s (1993) “The Fisher, Neyman-Pearson Theories of Testing Hypotheses: One theory or two?”.
So error statistical tests supply error probabilities associated with test outputs. Granted, there is a missing premise that moves from them to any inference, claim, decision or any other output of a test. That is what the severity concept and severity requirements supply for contexts where the goal is finding something out, solving a statistical problem or critically scrutinizing a conjectured solution to a problem. That is why severity is a basic concept for philosophy of statistics—as in my title.
(3) What are the 100 years of fighting and browbeating all about?
According to Recht:
If, after 100 years of fighting and browbeating, we see that statistical testing consistently fails to be severe testing, then it’s pretty silly to keep teaching our students that statistical tests are severe tests. (Recht post)
I am not sure how Recht is viewing the “100 years of fighting and browbeating”? Does he think they are over the well-known fallacy of moving of statistical to substantive significance–affirming the consequent discussed in (1) above? The only claim inferred with severity from a rejection of H0 is its denial (in relation to the test statistic). The fact that Meehl was wrestling with fallacious uses of tests in psychology does not alter what tests actually do, nor provide grounds to suppose that statisticians have been teaching their students that a statistically significant effect automatically warrants a substantive claim H. The error probabilities of tests do not apply to H, unless it is tantamount to the denial of H0. So the “100 years of fighting and browbeating” is not over whether statistical tests supply ways to move directly from statistical to substantive claims, from correlational to causal claims or the like. Those have been well-known fallacies for donkey’s years.
Perhaps Recht means that there has been 100 years of fighting and browbeating over whether statistical tests can supply tests with good error probabilities (which is all they strictly claim to do), but that is not so. We know they can supply them. (Neyman also developed confidence interval estimation as inverses to tests, with corresponding coverage probabilities.) The 100 years of fighting is whether error probabilities matter for statistical inference, given they do not supply degrees of support, belief, or probability of statistical hypotheses. The 100 years of fighting, in other words, is over (frequentist) error statistical performance versus Bayesian (or other) probabilisms. It is conceptual and philosophical. Recht does not say anything about this here, and I’d still like to know what he thinks. Notice that Bayesian confirmation does permit moving from rejecting H0 to a substantive H insofar as H** receives a Bayes boost (is made more probable). The evidence from the data for updating, or for Bayes factors, is in the likelihood ratio (see slide #38)[ the likelihood principle]. This is at odds with error probabilities.
I might note that no formal statistical methods make use of the term I dub “severity”, although I thank Meehl for mentioning me (in the same breath as Popper and Salmon!) Severity leads to reformulating tests so as to avoid classic fallacies of rejection and non-rejection. I observe in SIST:
none of the [existing] formal notions directly give severity assessments. There isn’t even a statistical school or tribe that has explicitly endorsed this goal. I find this perplexing. That will not preclude our immersion into the mindset of a futuristic tribe whose members use error probabilities for assessing severity; it’s just the ticket for our task: understanding and getting beyond the statistics wars. We may call this tribe the severe testers. (SIST, 9)
So what of Recht’s question of the purpose of statistical tests?
The purposes of statistics, of course, are enormously broad, and I will leave that to statisticians. But the purposes of statistical significance tests are, first and foremost, as Benjamini puts it, to supply our “first line of defense against being fooled by randomness” (2016, p. 1). If an observed effect is explainable as due to chance variability, then we’d very probably fail to reliably generate statistically significant results. It is by dint of non-significant results that failed replication is identified, and, notice, critics of tests presuppose the use of tests for this critical role. (The Replication Paradox.)
By dint of statistically significant results, on the other hand, tests uncover discordancies and inconsistencies between data and a reference or test hypothesis H0. This is the essence of model criticism. In the severe tester’s formulation, attained p-values may be used to indicate the extent of discrepancies that are well or poorly warranted. This is conjecture and refutation, followed by new conjecture which is then put to the test (combined with background theory), and the error probing occurs all along the path from data collection, modeling, and inference. (I see this as akin to Bin Yu’s veridical data science, which I’m only recently learning about.)
The introspection Recht calls for at the end of his post, as I see it, is an invitation to the philosophical and conceptual considerations that I discuss in my talk (probabilism, performance, and probativism).
For statisticians, this means there’s a need for introspection about what the field is for. …Statistics asks many different kinds of questions, but we confuse our students because the methods often look the same. Are we trying to quantify the verisimilitude of a theory or assertion? Or are we trying to quantify the error in a measurement to aid decision making? We need to speak with clarity about this. …until we disambiguate use, we will have more dumb arguments about what p-value thresholds mean for the replicability of science.
Whether one wants to call it measuring plausibility, probability, support, or verisimilitude, these would all fall under what I call probabilisms. Highly probable (however one measures it) differs from highly well probed, and distinct tools are needed for these different goals. Moreover, good performance is necessary but not sufficient for severe testing. (This was the distinction between Fisher and Neyman that Lehmann draws at the end of his paper.) In addition to clarifying use, probabilists need to clarify their chosen measures. At present, there are subjective, objective, empirical, pragmatic and many systems within each, with little agreement on which to use or how to interpret them. How the goals of AI/ML and data science more generally fit in, is yet a distinct question.
Recht’s final point leads me to add one more remark. He says:
So I’ll close a word to the folks who like to philosophize about statistics, whether they be philosophers, statisticians, or bloggers. We need less focus on Popper’s modus tollens and more on his piecemeal social engineering. (Recht post)
The severe tester is happy to promote Popper’s piecemeal engineering for problems of policy reform: it is precisely akin to what is right-headed in his view of solving problems piecemeal. The need for democratic checks, and considering how any one side on controversial policies may be wrong, leading to unintended consequences, is crucial. That was the gist of my (2021) editorial in Conservation Biology “The statistics wars and intellectual conflicts of interest”, for the policy of abandoning significance and p-value thresholds. The editorial is here.
I’m grateful to statistician Philip Stark, also a panelist at my Berkeley talk, for his published comment (in Conservation Biology) on my editorial. In his view:
I also agree with Prof. Mayo’s thesis that abandoning P-values exacerbates moral hazard for journal editors, although there has always been moral hazard in the gatekeeping function. Absent any objective assessment of the agreement between the data and competing theories, publication decisions may be even more subject to cronyism, “taste,” confirmation bias, etc.
Throwing away P-values because many practitioners don’t know how to use them is like banning scalpels because most people don’t know how to perform surgery. Those who would perform surgery should be trained in the proper use of scalpels, and those who would use statistics should be trained in the proper use of P-values. (Stark 2022, 1)
I ended my editorial as follows:
The key function of statistical tests is to constrain the human tendency to selectively favor views they believe in. There are ample forums for debating statistical methodologies. There is no call for executive directors or journal editors to place a thumb on the scale. Whether in dealing with environmental policy advocates, drug lobbyists, or avid calls to expel statistical significance tests, a strong belief in the efficacy of an intervention is distinct from its having been well tested. Applied science will be well served by editorial policies that uphold that distinction.
Notes
[1] For example, in a one-sided test (of the mean): H0: μ < μ0 vs H1: μ > μ0 , if the power of the test to detect μ’ is high, (i.e., POW(μ’) is high) then a just statistically significant result is poor evidence that μ > μ’: the severity associated with inferring μ > μ’ is low.
My slides from the Neyman Seminar are below (pdf):
Severity as a basic concept in philosophy of statistics
Third Stop
Readers: With this third stop we’ve covered Tour 1 of Excursion 1. My slides from the first LSE meeting in 2020 which dealt with elements of Excursion 1 can be found at the end of this post. There’s also a video giving an overall intro to SIST, Excursion 1. It’s noteworthy to consider just how much things seem to have changed in just the past few years. Or have they? What would the view from the hot-air balloon look like now? I will try to address this in the comments.
The Current State of Play in Statistical Foundations: A View From a Hot-Air Balloon (1.3)
.
How can a discipline, central to science and to critical thinking, have two methodologies, two logics, two approaches that frequently give substantively different answers to the same problems? … Is complacency in the face of contradiction acceptable for a central discipline of science? (Donald Fraser 2011, p. 329)
We [statisticians] are not blameless … we have not made a concerted professional effort to provide the scientific world with a unified testing methodology. (J. Berger 2003, p. 4)
From the aerial perspective of a hot-air balloon, we may see contemporary statistics as a place of happy multiplicity: the wealth of computational ability allows for the application of countless methods, with little handwringing about foundations. Doesn’t this show we may have reached “the end of statistical foundations”? One might have thought so. Yet, descending close to a marshy wetland, and especially scratching a bit below the surface, reveals unease on all sides. The false dilemma between probabilism and long-run performance lets us get a handle on it. In fact, the Bayesian versus frequentist dispute arises as a dispute between probabilism and performance. This gets to my second reason for why the time is right to jump back into these debates: the “statistics wars” present new twists and turns. Rival tribes are more likely to live closer and in mixed neighborhoods since around the turn of the century. Yet, to the beginning student, it can appear as a jungle.
Statistics Debates: Bayesian versus Frequentist
These days there is less distance between Bayesians and frequentists, especially with the rise of objective [default] Bayesianism, and we may even be heading toward a coalition government. (Efron 2013, p. 145)
A central way to formally capture probabilism is by means of the formula for conditional probability, where Pr(x) > 0:
Since Pr(H and x) = Pr(x|H)Pr(H) and Pr(x) = Pr(x|H)Pr(H) + Pr(x|~H)Pr(~H), we get:
where ~H is the denial of H. It would be cashed out in terms of all rivals to H within a frame of reference. Some call it Bayes’ Rule or inverse probability. Leaving probability uninterpreted for now, if the data are very improbable given H, then our probability in H after seeing x, the posterior probability Pr(H|x), may be lower than the probability in H prior to x, the prior prob- ability Pr(H). Bayes’ Theorem is just a theorem stemming from the definition of conditional probability; it is only when statistical inference is thought to be encompassed by it that it becomes a statistical philosophy. Using Bayes’ Theorem doesn’t make you a Bayesian.
Larry Wasserman, a statistician and master of brevity, boils it down to a contrast of goals. According to him (2012b):
The Goal of Frequentist Inference: Construct procedure with frequentist guarantees [i.e., low error rates].
The Goal of Bayesian Inference: Quantify and manipulate your degrees of beliefs. In other words, Bayesian inference is the Analysis of Beliefs.
At times he suggests we use B(H) for belief and F(H) for frequencies. The distinctions in goals are too crude, but they give a feel for what is often regarded as the Bayesian-frequentist controversy. However, they present us with the false dilemma (performance or probabilism) I’ve said we need to get beyond.
Today’s Bayesian–frequentist debates clearly differ from those of some years ago. In fact, many of the same discussants, who only a decade ago were arguing for the irreconcilability of frequentist P-values and Bayesian measures, are now smoking the peace pipe, calling for ways to unify and marry the two. I want to show you what really drew me back into the Bayesian–frequentist debates sometime around 2000. If you lean over the edge of the gondola, you can hear some Bayesian family feuds starting around then or a bit after. Principles that had long been part of the Bayesian hard core are being questioned or even abandoned by members of the Bayesian family. Suddenly sparks are flying, mostly kept shrouded within Bayesian walls, but nothing can long be kept secret even there. Spontaneous combustion looms. Hard core subjectivists are accusing the increasingly popular “objective (non-subjective)” and “reference” Bayesians of practicing in bad faith; the new frequentist–Bayesian unificationists are taking pains to show they are not subjective; and some are calling the new Bayesian kids on the block “pseudo Bayesian.” Then there are the Bayesians camping somewhere in the middle (or perhaps out in left field) who, though they still use the Bayesian umbrella, are flatly denying the very idea that Bayesian updating fits anything they actually do in statistics. Obeisance to Bayesian reasoning remains, but on some kind of a priori philosophical grounds. Let’s start with the unifications.
While subjective Bayesianism offers an algorithm for coherently updating prior degrees of belief in possible hypotheses H1, H2, …, Hn, these unifications fall under the umbrella of non-subjective Bayesian paradigms. Here the prior probabilities in hypotheses are not taken to express degrees of belief but are given by various formal assignments, ideally to have minimal impact on the posterior probability. I will call such Bayesian priors default. Advocates of unifications are keen to show that (i) default Bayesian methods have good performance in a long series of repetitions – so probabilism may yield performance; or alternatively, (ii) frequentist quantities are similar to Bayesian ones (at least in certain cases) – so performance may yield probabilist numbers. Why is this not bliss? Why are so many from all sides dissatisfied?
True blue subjective Bayesians are understandably unhappy with non- subjective priors. Rather than quantify prior beliefs, non-subjective priors are viewed as primitives or conventions for obtaining posterior probabilities. Take Jay Kadane (2008):
The growth in use and popularity of Bayesian methods has stunned many of us who were involved in exploring their implications decades ago. The result … is that there are users of these methods who do not understand the philosophical basis of the methods they are using, and hence may misinterpret or badly use the results … No doubt helping people to use Bayesian methods more appropriately is an important task of our time. (p. 457, emphasis added)
I have some sympathy here: Many modern Bayesians aren’t aware of the traditional philosophy behind the methods they’re buying into. Yet there is not just one philosophical basis for a given set of methods. This takes us to one of the most dramatic shifts in contemporary statistical foundations. It had long been assumed that only subjective or personalistic Bayesianism had a shot at providing genuine philosophical foundations, but you’ll notice that groups holding this position, while they still dot the landscape in 2018, have been gradually shrinking. Some Bayesians have come to question whether the wide- spread use of methods under the Bayesian umbrella, however useful, indicates support for subjective Bayesianism as a foundation.
Marriages of Convenience?
The current frequentist–Bayesian unifications are often marriages of convenience; statisticians rationalize them less on philosophical than on practical grounds. For one thing, some are concerned that methodological conflicts are bad for the profession. For another, frequentist tribes, contrary to expectation, have not disappeared. Ensuring that accounts can control their error probabilities remains a desideratum that scientists are unwilling to forgo. Frequentists have an incentive to marry as well. Lacking a suitable epistemic interpretation of error probabilities – significance levels, power, and confidence levels – frequentists are constantly put on the defensive. Jim Berger (2003) proposes a construal of significance tests on which the tribes of Fisher, Jeffreys, and Neyman could agree, yet none of the chiefs of those tribes concur (Mayo 2003b). The success stories are based on agreements on numbers that are not obviously true to any of the three philosophies. Beneath the surface – while it’s not often said in polite company – the most serious disputes live on. I plan to lay them bare.
If it’s assumed an evidential assessment of hypothesis H should take the form of a posterior probability of H – a form of probabilism – then P-values and confidence levels are applicable only through misinterpretation and mistranslation. Resigned to live with P-values, some are keen to show that construing them as posterior probabilities is not so bad (e.g., Greenland and Poole 2013). Others focus on long-run error control, but cede territory wherein probability captures the epistemological ground of statistical inference. Why assume significance levels and confidence levels lack an authentic epistemological function? I say they do: to secure and evaluate how well probed and how severely tested claims are.
Eclecticism and Ecumenism
If you look carefully between dense forest trees, you can distinguish unification country from lands of eclecticism (Cox 1978) and ecumenism (Box 1983), where tools first constructed by rival tribes are separate, and more or less equal (for different aims). Current-day eclecticisms have a long history – the dabbling in tools from competing statistical tribes has not been thought to pose serious challenges. For example, frequentist methods have long been employed to check or calibrate Bayesian methods (e.g., Box 1983); you might test your statistical model using a simple significance test, say, and then proceed to Bayesian updating. Others suggest scrutinizing a posterior probability or a likelihood ratio from an error probability standpoint. What this boils down to will depend on the notion of probability used. If a procedure frequently gives high probability for claim C even if C is false, severe testers deny convincing evidence has been provided, and never mind about the meaning of probability. One argument is that throwing different methods at a problem is all to the good, that it increases the chances that at least one will get it right. This may be so, provided one understands how to interpret competing answers. Using multiple methods is valuable when a shortcoming of one is rescued by a strength in another. For example, when randomized studies are used to expose the failure to replicate observational studies, there is a presumption that the former is capable of discerning problems with the latter. But what happens if one procedure fosters a goal that is not recognized or is even opposed by another? Members of rival tribes are free to sneak ammunition from a rival’s arsenal – but what if at the same time they denounce the rival method as useless or ineffective?
Decoupling.On the horizon is the idea that statistical methods may be decoupled from the philosophies in which they are traditionally couched. In an attempted meeting of the minds (Bayesian and error statistical), Andrew Gelman and Cosma Shalizi (2013) claim that “implicit in the best Bayesian practice is a stance that has much in common with the error-statistical approach of Mayo” (p. 10). In particular, Bayesian model checking, they say, uses statistics to satisfy Popperian criteria for severe tests. The idea of error statistical foundations for Bayesian tools is not as preposterous as it may seem. The concept of severe testing is sufficiently general to apply to any of the methods now in use. On the face of it, any inference, whether to the adequacy of a model or to a posterior probability, can be said to be warranted just to the extent that it has withstood severe testing. Where this will land us is still futuristic.
Why Our Journey?
We have all, or nearly all, moved past these old [Bayesian-frequentist] debates, yet our textbook explanations have not caught up with the eclecticism of statistical practice. (Kass 2011, p. 1)
When Kass proffers “a philosophy that matches contemporary attitudes,” he finds resistance to his big tent. Being hesitant to reopen wounds from old battles does not heal them. Distilling them in inoffensive terms just leads to the marshy swamp. Textbooks can’t “catch-up” by soft-peddling competing statistical accounts. They show up in the current problems of scientific integrity, irreproducibility, questionable research practices, and in the swirl of methodological reforms and guidelines that spin their way down from journals and reports.
From an elevated altitude we see how it occurs. Once high-profile failures of replication spread to biomedicine, and other “hard” sciences, the problem took on a new seriousness. Where does the new scrutiny look? By and large, it collects from the earlier social science “significance test controversy” and the traditional philosophies coupled to Bayesian and frequentist accounts, along with the newer Bayesian–frequentist unifications we just surveyed. This jungle has never been disentangled. No wonder leading reforms and semi-popular guidebooks contain misleading views about all these tools. No wonder we see the same fallacies that earlier reforms were designed to avoid, and even brand new ones. Let me be clear, I’m not speaking about flat-out howlers such as interpreting a P-value as a posterior probability. By and large, they are more subtle; you’ll want to reach your own position on them. It’s not a matter of switching your tribe, but excavating the roots of tribal warfare. To tell what’s true about them. I don’t mean understand them at the socio-psychological levels, although there’s a good story there (and I’ll leak some of the juicy parts during our travels).
How can we make progress when it is difficult even to tell what is true about the different methods of statistics? We must start afresh, taking responsibility to offer a new standpoint from which to interpret the cluster of tools around which there has been so much controversy. Only then can we alter and extend their limits. I admit that the statistical philosophy that girds our explorations is not out there ready-made; if it was, there would be no need for our holiday cruise. While there are plenty of giant shoulders on which we stand, we won’t be restricted by the pronouncements of any of the high and low priests, as sagacious as many of their words have been. In fact, we’ll brazenly question some of their most entrenched mantras. Grab on to the gondola, our balloon’s about to land.
In Tour II, I’ll give you a glimpse of the core behind statistics battles, with a firm promise to retrace the steps more slowly in later trips.
Please share your constructive thoughts and queries in the comments.
FOR ALL OF TOUR I: SIST Excursion 1 Tour I
THE FULL ITINERARY: Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars: SIST Itinerary
I wrote detailed notes/outline of Excursion 1 with definitions of some key terms.
REFERENCES:
Berger, J. (2003). ‘Could Fisher, Jeffreys and Neyman Have Agreed on Testing?’ and ‘Rejoinder’, Statistical Science 18(1), 1–12; 28–32.
Box, G. (1983). ‘An Apology for Ecumenism in Statistics’, in Box, G., Leonard, T., and Wu, D. (eds.), Scientific Inference, Data Analysis, and Robustness, New York:
Academic Press, 51–84.
Cox, D. (1978). ‘Foundations of Statistical Inference: The Case for Eclecticism’, Australian Journal of Statistics 20(1), 43–59.
Efron, B. (2013). ‘A 250-Year Argument: Belief, Behavior, and the Bootstrap’, Bulletin of the American Mathematical Society 50(1), 126–46.
Fraser, D. (2011). ‘Is Bayes Posterior Just Quick and Dirty Confidence?’ and ‘Rejoinder’, Statistical Science 26(3), 299–316; 329–31.
Gelman, A. and Shalizi, C. (2013). ‘Philosophy and the Practice of Bayesian Statistics’ and ‘Rejoinder’, British Journal of Mathematical and Statistical Psychology 66(1), 8–38; 76–80.
Greenland, S. and Poole, C. (2013). ‘Living with P Values: Resurrecting a Bayesian Perspective on Frequentist Statistics’ and ‘Rejoinder: Living with Statistics in Observational Research’, Epidemiology 24(1), 62–8; 73–8.
Kadane, J. (2008). ‘Comment on Article by Gelman’, Bayesian Analysis 3(3), 455–8.
Kass, R. (2011). ‘Statistical Inference: The Big Picture (with discussion and rejoinder)’, Statistical Science 26(1), 1–20.
Mayo, D. (2003b). ‘Could Fisher, Jeffreys and Neyman Have Agreed on Testing? Commentary on J. Berger’s Fisher Address’, Statistical Science 18, 19–24.
Wasserman, L. (2012b). ‘What is Bayesian/Frequentist Inference?’, Blogpost on normaldeviate.wordpress.com (11/7/2012).
Below are my slides from the first meeting. Slides: (PDF)
Introductory Video: Excursion 1
For the first meeting of the LSE 500, we used the introductory video to a 2019 Summer Seminar that I ran with Prof. Aris Spanos for college professors. It covers the same material we covered in Meeting 1 at the 2020 LSE seminar and is a useful introduction to SIST and overview of Excursion 1. (If you are curious about the Summer Seminar we ran in 2019, you can find the slides and videos from it on The Summer Seminar Blog.)
The first 38 minutes of the video below comprise an introduction and discussion of Excursion 1 Tour I. At minute 39, we shift to discussing/reviewing Tour II of Excursion 1.
2 notes for watching the video: There is a volume control on the little movie box (left-hand side) to control volume. Secondly, it’s suggested to watch the video in full screen mode to help with buffering, if that is a problem for you.
.
Readers: I gave the Neyman Seminar at Berkeley last Wednesday, October 9, and had been so busy preparing it that I did not update my leisurely cruise for October. This is the second stop. I will shortly post remarks on the the panel discussion that followed my Neyman talk (with panelists, Ben Recht, Philip Stark, Bin Yu, and Snow Zhang), which was quite illuminating.
“I shall be concerned with the foundations of the subject. But in case it should be thought that this means I am not here strongly concerned with practical applications, let me say right away that confusion about the foundations of the subject is responsible, in my opinion, for much of the misuse of the statistics that one meets in fields of application such as medicine, psychology, sociology, economics, and so forth”. (George Barnard 1985, p. 2)
While statistical science (as with other sciences) generally goes about its business without attending to its own foundations, implicit in every statistical methodology are core ideas that direct its principles, methods, and interpretations. I will call this its statistical philosophy. To tell what’s true about statistical inference, understanding the associated philosophy (or philosophies) is essential. Discussions of statistical foundations tend to focus on how to interpret probability, and much less on the overarching question of how probability ought to be used in inference. Assumptions about the latter lurk implicitly behind debates, but rarely get the limelight. If we put the spotlight on them, we see that there are two main philosophies about the roles of probability in statistical inference: We may dub them performance (in the long run) and probabilism.
The performance philosophy sees the key function of statistical method as controlling the relative frequency of erroneous inferences in the long run of applications. For example, a frequentist statistical test, in its naked form, can be seen as a rule: whenever your outcome exceeds some value (say, X > x), reject a hypothesis H0 and infer H1. The value of the rule, according to its performance-oriented defenders, is that it can ensure that, regardless of which hypothesis is true, there is both a low probability of erroneously rejecting H0 (rejecting H0 when it is true) as well as erroneously accepting H0 (failing to reject H*0 when it is false).
The second philosophy, probabilism, views probability as a way to assign degrees of belief, support, or plausibility to hypotheses. Many keep to a comparative report, for example that H0 is more believable than is H1 given data x; others strive to say H0 is less believable given data x than before, and offer a quantitative report of the difference.
What happened to the goal of scrutinizing BENT science by the severity criterion? [See 1.1] Neither “probabilism” nor “performance” directly captures that demand. To take these goals at face value, it’s easy to see why they come up short. Potti and Nevins’ strong belief in the reliability of their prediction model for cancer therapy scarcely made up for the shoddy testing. Neither is good long-run performance a sufficient condition. Most obviously, there may be no long-run repetitions, and our interest in science is often just the particular statistical inference before us. Crude long-run requirements may be met by silly methods. Most importantly, good performance alone fails to get at why methods work when they do; namely – I claim – to let us assess and control the stringency of tests. This is the key to answering a burning question that has caused major headaches in statistical foundations: why should a low relative frequency of error matter to the appraisal of the inference at hand? It is not probabilism or performance we seek to quantify, but probativeness.
I do not mean to disparage the long-run performance goal – there are plenty of tasks in inquiry where performance is absolutely key. Examples are screening in high-throughput data analysis, and methods for deciding which of tens of millions of collisions in high-energy physics to capture and analyze. New applications of machine learning may lead some to say that only low rates of prediction or classification errors matter. Even with prediction, “black-box” modeling, and non-probabilistic inquiries, there is concern with solving a problem. We want to know if a good job has been done in the case at hand.
Severity (Strong): Argument from Coincidence
The weakest version of the severity requirement (Section 1.1), in the sense of easiest to justify, is negative, warning us when BENT data are at hand, and a surprising amount of mileage may be had from that negative principle alone. It is when we recognize how poorly certain claims are warranted that we get ideas for improved inquiries. In fact, if you wish to stop at the negative requirement, you can still go pretty far along with me. I also advocate the positive counterpart:
Severity (strong): We have evidence for a claim C just to the extent it survives a stringent scrutiny. If C passes a test that was highly capable of finding flaws or discrepancies from C, and yet none or few are found, then the passing result, x, is evidence for C.
One way this can be achieved is by an argument from coincidence. The most vivid cases occur outside formal statistics.
Some of my strongest examples tend to revolve around my weight. Before leaving the USA for the UK, I record my weight on two scales at home, one digital, one not, and the big medical scale at my doctor’s office. Suppose they are well calibrated and nearly identical in their readings, and they also all pick up on the extra 3 pounds when I’m weighed carrying three copies of my 1-pound book, Error and the Growth of Experimental Knowledge (EGEK). Returning from the UK, to my astonishment, not one but all three scales show anywhere from a 4–5 pound gain. There’s no difference when I place the three books on the scales, so I must conclude, unfortunately, that I’ve gained around 4 pounds. Even for me, that’s a lot. I’ve surely falsified the supposition that I lost weight! From this informal example, we may make two rather obvious points that will serve for less obvious cases. First, there’s the idea I call lift-off.
Lift-off: An overall inference can be more reliable and precise than its premises individually.
Each scale, by itself, has some possibility of error, and limited precision. But the fact that all of them have me at an over 4-pound gain, while none show any difference in the weights of EGEK, pretty well seals it. Were one scale off balance, it would be discovered by another, and would show up in the weighing of books. They cannot all be systematically misleading just when it comes to objects of unknown weight, can they? Rejecting a conspiracy of the scales, I conclude I’ve gained weight, at least 4 pounds. We may call this an argument from coincidence, and by its means we can attain lift-off. Lift-off runs directly counter to a seemingly obvious claim of drag-down.
Drag-down: An overall inference is only as reliable/precise as is its weakest premise.
The drag-down assumption is common among empiricist philosophers: As they like to say, “It’s turtles all the way down.” Sometimes our inferences do stand as a kind of tower built on linked stones – if even one stone fails they all come tumbling down. Call that a linked argument.
Our most prized scientific inferences would be in a very bad way if piling on assumptions invariably leads to weakened conclusions. Fortunately we also can build what may be called convergent arguments, where lift-off is attained. This seemingly banal point suffices to combat some of the most well entrenched skepticisms in philosophy of science. And statistics happens to be the science par excellence for demonstrating lift-off!
Now consider what justifies my weight conclusion, based, as we are supposing it is, on a strong argument from coincidence. No one would say: “I can be assured that by following such a procedure, in the long run I would rarely report weight gains erroneously, but I can tell nothing from these readings about my weight now.” To justify my conclusion by long-run performance would be absurd. Instead we say that the procedure had enormous capacity to reveal if any of the scales were wrong, and from this I argue about the source of the readings: H: I’ve gained weight. Simple as that. It would be a preposterous coincidence if none of the scales registered even slight weight shifts when weighing objects of known weight, and yet were systematically misleading when applied to my weight. You see where I’m going with this. This is the key – granted with a homely example – that can fill a very important gap in frequentist foundations: Just because an account is touted as having a long-run rationale, it does not mean it lacks a short run rationale, or even one relevant for the particular case at hand. Nor is it merely the improbability of all the results were H false; it is rather like denying an evil demon has read my mind just in the cases where I do not know the weight of an object, and deliberately deceived me. The argument to “weight gain” is an example of an argument from coincidence to the absence of an error, what I call:
Arguing from Error: There is evidence an error is absent to the extent that a procedure with a very high capability of signaling the error, if and only if it is present, nevertheless detects no error.
I am using “signaling” and “detecting” synonymously: It is important to keep in mind that we don’t know if the test output is correct, only that it gives a signal or alert, like sounding a bell. Methods that enable strong arguments to the absence (or presence) of an error I call strong error probes. Our ability to develop strong arguments from coincidence, I will argue, is the basis for solving the “problem of induction.”
Glaring Demonstrations of Deception
Intelligence is indicated by a capacity for deliberate deviousness. Such deviousness becomes self-conscious in inquiry: An example is the use of a placebo to find out what it would be like if the drug has no effect. What impressed me the most in my first statistics class was the demonstration of how apparently impressive results are readily produced when nothing’s going on, i.e., “by chance alone.” Once you see how it is done, and done easily, there is no going back. The toy hypotheses used in statistical testing are nearly always overly simple as scientific hypotheses. But when it comes to framing rather blatant deceptions, they are just the ticket!
When Fisher offered Muriel Bristol-Roach a cup of tea back in the 1920s, she refused it because he had put the milk in first. What difference could it make? Her husband and Fisher thought it would be fun to put her to the test (1935a). Say she doesn’t claim to get it right all the time but does claim that she has some genuine discerning ability. Suppose Fisher subjects her to 16 trials and she gets 9 of them right. Should I be impressed or not? By a simple experiment of randomly assigning milk first/tea first Fisher sought to answer this stringently. But don’t be fooled: a great deal of work goes into controlling biases and confounders before the experimental design can work. The main point just now is this: so long as lacking ability is sufficiently like the canonical “coin tossing” (Bernoulli) model (with the probability of success at each trial of 0.5), we can learn from the test procedure. In the Bernoulli model, we record success or failure, assume a fixed probability of success θ on each trial, and that trials are independent. If the probability of getting even more successes than she got, merely by guessing, is fairly high, there’s little indication of special tasting ability. The probability of at least 9 of 16 successes, even if θ = 0.5, is 0.4. To abbreviate, Pr(at least 9 of 16 successes; H0: θ = 0.5) = 0.4. This is the P-value of the observed difference; an unimpressive 0.4. You’d expect as many or even more “successes” 40% of the time merely by guessing. It’s also the significance level attained by the result. (I often use P-value as it’s shorter.) Muriel Bristol-Roach pledges that if her performance may be regarded as scarcely better than guessing, then she hasn’t shown her ability. Typically, a small value such as 0.05, 0.025, or 0.01 is required.
Such artificial and simplistic statistical hypotheses play valuable roles at stages of inquiry where what is needed are blatant standards of “nothing’s going on.” There is no presumption of a metaphysical chance agency, just that there is expected variability – otherwise one test would suffice – and that probability models from games of chance can be used to distinguish genuine from spurious effects. Although the goal of inquiry is to find things out, the hypotheses erected to this end are generally approximations and may be deliberately false. To present statistical hypotheses as identical to substantive scientific claims is to mischaracterize them. We want to tell what’s true about statistical inference. Among the most notable of these truths is:
P-values can be readily invalidated due to how the data (or hypotheses!) are generated or selected for testing.
If you fool around with the results afterwards, reporting only successful guesses, your report will be invalid. You may claim it’s very difficult to get such an impressive result due to chance, when in fact it’s very easy to do so, with selective reporting. Another way to put this: your computed P-value is small, but the actual P-value is high! Concern with spurious findings, while an ancient problem, is considered sufficiently serious to have motivated the American Statistical Association to issue a guide on how not to interpret P-values (Wasserstein and Lazar 2016); hereafter, ASA 2016 Guide. It may seem that if a statistical account is free to ignore such fooling around then the problem disappears! It doesn’t.
Incidentally, Bristol-Roach got all the cases correct, and thereby taught her husband a lesson about putting her claims to the test.
skips p. 18 on Peirce
Texas Marksman
Take an even simpler and more blatant argument of deception. It is my favorite: the Texas Marksman. A Texan wants to demonstrate his shooting prowess. He shoots all his bullets any old way into the side of a barn and then paints a bull’s-eye in spots where the bullet holes are clustered. This fails utterly to severely test his marksmanship ability. When some visitors come to town and notice the incredible number of bull’s-eyes, they ask to meet this marksman and are introduced to a little kid. How’d you do so well, they ask? Easy, I just drew the bull’s-eye around the most tightly clustered shots. There is impressive “agreement” with shooting ability, he might even compute how improbably so many bull’s-eyes would occur by chance. Yet his ability to shoot was not tested in the least by this little exercise. There’s a real effect all right, but it’s not caused by his marksmanship! It serves as a potent analogy for a cluster of formal statistical fallacies from data-dependent findings of “exceptional” patterns.
The term “apophenia” refers to a tendency to zero in on an apparent regularity or cluster within a vast sea of data and claim a genuine regularity. One of our fundamental problems (and skills) is that we’re apopheniacs. Some investment funds, none that we actually know, are alleged to produce several portfolios by random selection of stocks and send out only the one that did best. Call it the Pickrite method. They want you to infer that it would be a preposterous coincidence to get so great a portfolio if the Pickrite method were like guessing. So their methods are genuinely wonderful, or so you are to infer. If this had been their only portfolio, the probability of doing so well by luck is low. But the probability of at least one of many portfolios doing so well (even if each is generated by chance) is high, if not guaranteed.
Let’s review the rogues’ gallery of glaring arguments from deception. The lady tasting tea showed how a statistical model of “no effect” could be used to amplify our ordinary capacities to discern if something really unusual is going on. The P-value is the probability of at least as high a success rate as observed, assuming the test or null hypothesis, the probability of success is 0.5. Since even more successes than she got is fairly frequent through guessing alone (the P-value is moderate), there’s poor evidence of a genuine ability. The Playfair and Texas sharpshooter examples, while quasi-formal or informal, demonstrate how to invalidate reports of significant effects. They show how gambits of post-data adjustments or selection can render a method highly capable of spewing out impressive looking fits even when it’s just random noise.
We appeal to the same statistical reasoning to show the problematic cases as to show genuine arguments from coincidence.
So am I proposing that a key role for statistical inference is to identify ways to spot egregious deceptions (BENT cases) and create strong arguments from coincidence? Yes, I am.
Skips “Spurious P-values and Auditing” (p. 20) up to Souvenir A (p. 21)
Souvenir A: Postcard to Send
The gift shop has a postcard listing the four slogans from the start of this Tour. Much of today’s handwringing about statistical inference is unified by a call to block these fallacies. In some realms, trafficking in too-easy claims for evidence, if not criminal offenses, are “bad statistics”; in others, notably some social sciences, they are accepted cavalierly – much to the despair of panels on research integrity. We are more sophisticated than ever about the ways researchers can repress unwanted, and magnify wanted, results. Fraud-busting is everywhere, and the most important grain of truth is this: all the fraud-busting is based on error statistical reasoning (if only on the meta-level). The minimal requirement to avoid BENT isn’t met. It’s hard to see how one can grant the criticisms while denying the critical logic.
We should oust mechanical, recipe-like uses of statistical methods that have long been lampooned, and are doubtless made easier by Big Data mining. They should be supplemented with tools to report magnitudes of effects that have and have not been warranted with severity. But simple significance tests have their uses, and shouldn’t be ousted simply because some people are liable to violate Fisher’s warning and report isolated results. They should be seen as a part of a conglomeration of error statistical tools for distinguishing genuine and spurious effects. They offer assets that are essential to our task: they have the means by which to register formally the fallacies in the postcard list. The failed statistical assumptions, the selection effects from trying and trying again, all alter a test’s error-probing capacities. This sets off important alarm bells, and we want to hear them. Don’t throw out the error-control baby with the bad statistics bathwater.
The slogans about lying with statistics? View them, not as a litany of embarrassments, but as announcing what any responsible method must register, if not control or avoid. Criticisms of statistical tests, where valid, boil down to problems with the critical alert function. Far from the high capacity to warn, “Curb your enthusiasm!” as correct uses of tests do, there are practices that make sending out spurious enthusiasm as easy as pie. This is a failure for sure, but don’t trade them in for methods that cannot detect failure at all. If you’re shopping for a statistical account, or appraising a statistical reform, your number one question should be: does it embody trigger warnings of spurious effects? Of bias? Of cherry picking and multiple tries? If the response is: “No problem; if you use our method, those practices require no change in statistical assessment!” all I can say is, if it sounds too good to be true, you might wish to hold off buying it.
Skips remainder of section 1.2 (bott p. 22- middle p. 23).
NOTES:
2 This is the traditional use of “bias” as a systematic error. Ioannidis (2005) alludes to biasing as behaviors that result in a reported significance level differing from the value it actually has or ought to have (e.g., post-data endpoints, selective reporting). I will call those biasing selection effects.
FOR ALL OF TOUR I: SIST Excursion 1 Tour I
THE FULL ITINERARY: Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars: SIST Itinerary
Ship Statinfasst
Excerpt from excursion 1 Tour I: Beyond Probabilism and Performance: Severity Requirement (1.1)
NOTE: The following is an excerpt from my existing book: Statistical Inference as Severe Testing: How to get beyond the statistics wars (CUP, 2018). For any new reflections or corrections, I will use the comments. The initial announcement is here.
I’m talking about a specific, extra type of integrity that is [beyond] not lying, but bending over backwards to show how you’re maybe wrong, that you ought to have when acting as a scientist. (Feynman 1974/1985, p. 387)
It is easy to lie with statistics. Or so the cliché goes. It is also very difficult to uncover these lies without statistical methods – at least of the right kind. Self- correcting statistical methods are needed, and, with minimal technical fanfare, that’s what I aim to illuminate. Since Darrell Huff wrote How to Lie with Statistics in 1954, ways of lying with statistics are so well worn as to have emerged in reverberating slogans:
Exposés of fallacies and foibles ranging from professional manuals and task forces to more popularized debunking treatises are legion. New evidence has piled up showing lack of replication and all manner of selection and publication biases. Even expanded “evidence-based” practices, whose very rationale is to emulate experimental controls, are not immune from allegations of illicit cherry picking, significance seeking, P-hacking, and assorted modes of extra- ordinary rendition of data. Attempts to restore credibility have gone far beyond the cottage industries of just a few years ago, to entirely new research programs: statistical fraud-busting, statistical forensics, technical activism, and widespread reproducibility studies. There are proposed methodological reforms – many are generally welcome (preregistration of experiments, transparency about data collection, discouraging mechanical uses of statistics), some are quite radical. If we are to appraise these evidence policy reforms, a much better grasp of some central statistical problems is needed.
Getting Philosophical
Are philosophies about science, evidence, and inference relevant here? Because the problems involve questions about uncertain evidence, probabilistic models, science, and pseudoscience – all of which are intertwined with technical statistical concepts and presuppositions – they certainly ought to be. Even in an open-access world in which we have become increasingly fearless about taking on scientific complexities, a certain trepidation and groupthink take over when it comes to philosophically tinged notions such as inductive reasoning, objectivity, rationality, and science versus pseudoscience. The general area of philosophy that deals with knowledge, evidence, inference, and rationality is called epistemology. The epistemological standpoints of leaders, be they philosophers or scientists, are too readily taken as canon by others. We want to understand what’s true about some of the popular memes: “All models are false,” “Everything is equally subjective and objective,” “P-values exaggerate evidence,” and “[M]ost published research findings are false” (Ioannidis 2005) – at least if you publish a single statistically significant result after data finagling. (Do people do that? Shame on them.) Yet R. A. Fisher, founder of modern statistical tests, denied that an isolated statistically significant result counts.
[W]e need, not an isolated record, but a reliable method of procedure. In relation to the test of significance, we may say that a phenomenon is experimentally demonstrable when we know how to conduct an experiment which will rarely fail to give us a statistically significant result. (Fisher 1935b/1947, p. 14)
Satisfying this requirement depends on the proper use of background knowledge and deliberate design and modeling.
This opening excursion will launch us into the main themes we will encounter. You mustn’t suppose, by its title, that I will be talking about how to tell the truth using statistics. Although I expect to make some progress there, my goal is to tell what’s true about statistical methods themselves! There are so many misrepresentations of those methods that telling what is true about them is no mean feat. It may be thought that the basic statistical concepts are well understood. But I show that this is simply not true.
Nor can you just open a statistical text or advice manual for the goal at hand. The issues run deeper. Here’s where I come in. Having long had one foot in philosophy of science and the other in foundations of statistics, I will zero in on the central philosophical issues that lie below the surface of today’s raging debates. “Getting philosophical” is not about articulating rarified concepts divorced from statistical practice. It is to provide tools to avoid obfuscating the terms and issues being bandied about. Readers should be empowered to understand the core presuppositions on which rival positions are based – and on which they depend.
Do I hear a protest? “There is nothing philosophical about our criticism of statistical significance tests (someone might say). The problem is that a small P-value is invariably, and erroneously, interpreted as giving a small probability to the null hypothesis.” Really? P-values are not intended to be used this way; presupposing they ought to be so interpreted grows out of a specific conception of the role of probability in statistical inference. That conception is philosophical. Methods characterized through the lens of over-simple epistemological orthodoxies are methods misapplied and mischaracterized. This may lead one to lie, however unwittingly, about the nature and goals of statistical inference, when what we want is to tell what’s true about them.
1.1 Severity Requirement: Bad Evidence, No Test (BENT)Fisher observed long ago, “[t]he political principle that anything can be proved by statistics arises from the practice of presenting only a selected subset of the data available” (Fisher 1955, p. 75). If you report results selectively, it becomes easy to prejudge hypotheses: yes, the data may accord amazingly well with a hypothesis H, but such a method is practically guaranteed to issue so good a fit even if H is false and not warranted by the evidence. If it is predetermined that a way will be found to either obtain or interpret data as evidence for H, then data are not being taken seriously in appraising H. H is essentially immune to having its flaws uncovered by the data. H might be said to have “passed” the test, but it is a test that lacks stringency or severity. Everyone understands that this is bad evidence, or no test at all. I call this the severity requirement. In its weakest form it supplies a minimal requirement for evidence:
Severity Requirement (weak): One does not have evidence for a claim if nothing has been done to rule out ways the claim may be false. If data x agree with a claim C but the method used is practically guaranteed to find such agreement, and had little or no capability of finding flaws with C even if they exist, then we have bad evidence, no test (BENT).
The “practically guaranteed” acknowledges that even if the method had some slim chance of producing a disagreement when C is false, we still regard the evidence as lousy. Little if anything has been done to rule out erroneous construals of data. We’ll need many different ways to state this minimal principle of evidence, depending on context….
skips bottom of p. 5-bottom of p. 6
Do We Always Want to Find Things Out?
The severity requirement gives a minimal principle based on the fact that highly insevere tests yield bad evidence, no tests (BENT). We can all agree on this much, I think. We will explore how much mileage we can get from it. It applies at a number of junctures in collecting and modeling data, in linking data to statistical inference, and to substantive questions and claims. This will be our linchpin for understanding what’s true about statistical inference. In addition to our minimal principle for evidence, one more thing is needed, at least during the time we are engaged in this project: the goal of finding things out.
The desire to find things out is an obvious goal; yet most of the time it is not what drives us. We typically may be uninterested in, if not quite resistant to, finding flaws or incongruencies with ideas we like. Often it is entirely proper to gather information to make your case, and ignore anything that fails to support it. Only if you really desire to find out something, or to challenge so-and-so’s (“trust me”) assurances, will you be prepared to stick your (or their) neck out to conduct a genuine “conjecture and refutation” exercise. Because you want to learn, you will be prepared to risk the possibility that the conjecture is found flawed.
We hear that “motivated reasoning has interacted with tribalism and new media technologies since the 1990s in unfortunate ways” (Haidt and Iyer 2016). Not only do we see things through the tunnel of our tribe, social media and web searches enable us to live in the echo chamber of our tribe more than ever. We might think we’re trying to find things out but we’re not. Since craving truth is rare (unless your life depends on it) and the “perverse incentives” of publishing novel results so shiny, the wise will invite methods that make uncovering errors and biases as quick and painless as possible. Methods of inference that fail to satisfy the minimal severity requirement fail us in an essential way.
With the rise of Big Data, data analytics, machine learning, and bioinformatics, statistics has been undergoing a good deal of introspection. Exciting results are often being turned out by researchers without a traditional statistics background; biostatistician Jeff Leek (2016) explains: “There is a structural reason for this: data was sparse when they were trained and there wasn’t any reason for them to learn statistics.” The problem goes beyond turf battles. It’s discovering that many data analytic applications are missing key ingredients of statistical thinking. Brown and Kass (2009) crystalize its essence. “Statistical thinking uses probabilistic descriptions of variability in (1) inductive reasoning and (2) analysis of procedures for data collection, prediction, and scientific inference” (p. 107). A word on each.
(1) Types of statistical inference are too varied to neatly encompass. Typically we employ data to learn something about the process or mechanism producing the data. The claims inferred are not specific events, but statistical generalizations, parameters in theories and models, causal claims, and general predictions. Statistical inference goes beyond the data – by definition that makes it an inductive inference. The risk of error is to be expected. There is no need to be reckless. The secret is controlling and learning from error. Ideally we take precautions in advance: pre-data, we devise methods that make it hard for claims to pass muster unless they are approximately true or adequately solve our problem. With data in hand, post-data, we scrutinize what, if anything, can be inferred.
What’s the essence of analyzing procedures in (2)? Brown and Kass don’t specifically say, but the gist can be gleaned from what vexes them; namely, ad hoc data analytic algorithms where researchers “have done nothing to indicate that it performs well” (p. 107). Minimally, statistical thinking means never ignoring the fact that there are alternative methods: Why is this one a good tool for the job? Statistical thinking requires stepping back and examining a method’s capabilities, whether it’s designing or choosing a method, or scrutinizing the results.
A Philosophical Excursion
Taking the severity principle then, along with the aim that we desire to find things out without being obstructed in this goal, let’s set sail on a philosophical excursion to illuminate statistical inference. Envision yourself embarking on a special interest cruise featuring “exceptional itineraries to popular destinations worldwide as well as unique routes” (Smithsonian Journeys). What our cruise lacks in glamour will be more than made up for in our ability to travel back in time to hear what Fisher, Neyman, Pearson, Popper, Savage, and many others were saying and thinking, and then zoom forward to current debates. There will be exhibits, a blend of statistics, philosophy, and history, and even a bit of theater. Our standpoint will be pragmatic in this sense: my interest is not in some ideal form of knowledge or rational agency, no omniscience or God’s-eye view – although we’ll start and end surveying the landscape from a hot-air balloon. I’m interested in the problem of how we get the kind of knowledge we do manage to obtain – and how we can get more of it. Statistical methods should not be seen as tools for what philosophers call “rational reconstruction” of a piece of reasoning. Rather, they are forward-looking tools to find something out faster and more efficiently, and to discriminate how good or poor a job others have done.
The job of the philosopher is to clarify but also to provoke reflection and scrutiny precisely in those areas that go unchallenged in ordinary practice. My focus will be on the issues having the most influence, and being most liable to obfuscation. Fortunately, that doesn’t require an abundance of technicalities, but you can opt out of any daytrip that appears too technical: an idea not caught in one place should be illuminated in another. Our philosophical excursion may well land us in positions that are provocative to all existing sides of the debate about probability and statistics in scientific inquiry.
Methodology and Meta-methodology
We are studying statistical methods from various schools. What shall we call methods for doing so? Borrowing a term from philosophy of science, we may call it our meta-methodology – it’s one level removed.1 To put my cards on the table: A severity scrutiny is going to be a key method of our meta-methodology. It is fairly obvious that we want to scrutinize how capable a statistical method is at detecting and avoiding erroneous interpretations of data. So when it comes to the role of probability as a pedagogical tool for our purposes, severity – its assessment and control – will be at the center. The term “severity” is Popper’s, though he never adequately defined it. It’s not part of any statistical methodology as of yet. Viewing statistical inference as severe testing lets us stand one level removed from existing accounts, where the air is a bit clearer.
Our intuitive, minimal, requirement for evidence connects readily to formal statistics. The probabilities that a statistical method lands in erroneous interpretations of data are often called its error probabilities. So an account that revolves around control of error probabilities I call an error statistical account. But “error probability” has been used in different ways. Most familiar are those in relation to hypotheses tests (Type I and II errors), significance levels, confidence levels, and power – all of which we will explore in detail. It has occasionally been used in relation to the proportion of false hypotheses among those now in circulation, which is different. For now it suffices to say that none of the formal notions directly give severity assessments. There isn’t even a statistical school or tribe that has explicitly endorsed this goal. I find this perplexing. That will not preclude our immersion into the mindset of a futuristic tribe whose members use error probabilities for assessing severity; it’s just the ticket for our task: understanding and getting beyond the statistics wars. We may call this tribe the severe testers.
We can keep to testing language. See it as part of the meta-language we use to talk about formal statistical methods, where the latter include estimation, exploration, prediction, and data analysis. I will use the term “hypothesis,” or just “claim,” for any conjecture we wish to entertain; it need not be one set out in advance of data. Even predesignating hypotheses, by the way, doesn’t preclude bias: that view is a holdover from a crude empiricism that assumes data are unproblematically “given,” rather than selected and interpreted. Conversely, using the same data to arrive at and test a claim can, in some cases, be accomplished with stringency.
As we embark on statistical foundations, we must avoid blurring formal terms such as probability and likelihood with their ordinary English meanings. Actually, “probability” comes from the Latin probare, meaning to try, test, or prove. “Proof” in “The proof is in the pudding” refers to how you put some- thing to the test. You must show or demonstrate, not just believe strongly. Ironically, using probability this way would bring it very close to the idea of measuring well-testedness (or how well shown). But it’s not our current, informal English sense of probability, as varied as that can be. To see this, consider “improbable.” Calling a claim improbable, in ordinary English, can mean a host of things: I bet it’s not so; all things considered, given what I know, it’s implausible; and other things besides. Describing a claim as poorly tested generally means something quite different: little has been done to probe whether the claim holds or not, the method used was highly unreliable, or things of that nature. In short, our informal notion of poorly tested comes rather close to the lack of severity in statistics. There’s a difference between finding H poorly tested by data x, and finding x renders H improbable – in any of the many senses the latter takes on. The existence of a Higgs particle was thought to be probable if not necessary before it was regarded as well tested around 2012. Physicists had to show or demonstrate its existence for it to be well tested. It follows that you are free to pursue our testing goal without implying there are no other statistical goals. One other thing on language: I will have to retain the terms currently used in exploring them. That doesn’t mean I’m in favor of them; in fact, I will jettison some of them by the end of the journey.
To sum up this first tour so far, statistical inference uses data to reach claims about aspects of processes and mechanisms producing them, accompanied by an assessment of the properties of the inference methods: their capabilities to control and alert us to erroneous interpretations. We need to report if the method has satisfied the most minimal requirement for solving such a problem. Has anything been tested with a modicum of severity, or not? The severe tester also requires reporting of what has been poorly probed, and highlights the need to “bend over backwards,” as Feynman puts it, to admit where weaknesses lie. In formal statistical testing, the crude dichotomy of “pass/fail” or “significant or not” will scarcely do. We must determine the magnitudes (and directions) of any statistical discrepancies warranted, and the limits to any substantive claims you may be entitled to infer from the statistical ones. Using just our minimal principle of evidence, and a sturdy pair of shoes, join me on a tour of statistical inference, back to the leading museums of statistics, and forward to current offshoots and statistical tribes.
.
Why We Must Get Beyond the Statistics Wars
Some readers may be surprised to learn that the field of statistics, arid and staid as it seems, has a fascinating and colorful history of philosophical debate, marked by unusual heights of passion, personality, and controversy for at least a century. Others know them all too well and regard supporting any one side largely as proselytizing. I’ve heard some refer to statistical debates as “theological.” I do not want to rehash the “statistics wars” that have raged in every decade, although the significance test controversy is still hotly debated among practitioners, and even though each generation fights these wars anew – with task forces set up to stem reflexive, recipe-like statistics that have long been deplored.
The time is ripe for a fair-minded engagement in the debates about statistical foundations; more than that, it is becoming of pressing importance. Not only because
nor because
– as important as those facets are – but because what is at stake is a critical standpoint that we may be in danger of losing. Without it, we forfeit the ability to communicate with, and hold accountable, the “experts,” the agencies, the quants, and all those data handlers increasingly exerting power over our lives. Understanding the nature and basis of statistical inference must not be considered as all about mathematical details; it is at the heart of what it means to reason scientifically and with integrity about any field whatever. Robert Kass (2011) puts it this way:
We care about our philosophy of statistics, first and foremost, because statistical inference sheds light on an important part of human existence, inductive reasoning, and we want to understand it. (p. 19)
Isolating out a particular conception of statistical inference as severe testing is a way of telling what’s true about the statistics wars, and getting beyond them.
Chutzpah, No Proselytizing
Our task is twofold: not only must we analyze statistical methods; we must also scrutinize the jousting on various sides of the debates. Our meta-level standpoint will let us rise above much of the cacophony; but the excursion will involve a dose of chutzpah that is out of the ordinary in professional discussions. You will need to critically evaluate the texts and the teams of critics, including brilliant leaders, high priests, maybe even royalty. Are they asking the most unbiased questions in examining methods, or are they like admen touting their brand, dragging out howlers to make their favorite method look good? (I am not sparing any of the statistical tribes here.) There are those who are earnest but brainwashed, or are stuck holding banners from an earlier battle now over; some are wedded to what they’ve learned, to what’s in fashion, to what pays the rent. Some are so jaundiced about the abuses of statistics as to wonder at my admittedly herculean task. I have a considerable degree of sympathy with them. But, I do not sympathize with those who ask: “why bother to clarify statistical concepts if they are invariably misinterpreted?” and then proceed to misinterpret them. Anyone is free to dismiss statistical notions as irrelevant to them, but then why set out a shingle as a “statistical reformer”? You may even be shilling for one of the proffered reforms, thinking it the road to restoring credibility, when it will do nothing of the kind.
You might say, since rival statistical methods turn on issues of philosophy and on rival conceptions of scientific learning, that it’s impossible to say anything “true” about them. You just did. It’s precisely these interpretative and philosophical issues that I plan to discuss. Understanding the issues is different from settling them, but it’s of value nonetheless. Although statistical disagreements involve philosophy, statistical practitioners and not philosophers are the ones leading today’s discussions of foundations. Is it possible to pursue our task in a way that will be seen as neither too philosophical nor not philosophical enough? Too statistical or not statistically sophisticated enough? Probably not, I expect grievances from both sides.
Finally, I will not be proselytizing for a given statistical school, so you can relax. Frankly, they all have shortcomings, insofar as one can even glean a clear statement of a given statistical “school.” What we have is more like a jumble with tribal members often speaking right past each other. View the severity requirement as a heuristic tool for telling what’s true about statistical controversies. Whether you resist some of the ports of call we arrive at is unimportant; it suffices that visiting them provides a key to unlock current mysteries that are leaving many consumers and students of statistics in the dark about a crucial portion of science.
NOTE:
1 This contrasts with the use of “metaresearch” to describe work on methodological reforms by non-philosophers. This is not to say they don’t tread on philosophical territory often: they do.
FOR ALL OF TOUR I: SIST Excursion 1 Tour I
THE FULL ITINERARY: Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars: SIST Itinerary
Ship Statinfasst
We’re embarking on a leisurely cruise through the highlights of Statistical Inference as Severe Testing [SIST]: How to Get Beyond the Statistics Wars (CUP 2018) this fall (Oct-Jan), following the 5 seminars I led for a 2020 London School of Economics (LSE) Graduate Research Seminar. It was run entirely online due to Covid (as were the workshops that followed). In this new, relaxed (self-paced) journey, excursions that had been covered in a week, will be spread out over a month [i] and I’ll be posting abbreviated excerpts on this blog a few times a month. Look for the posts marked with the picture of ship StatInfAsSt. [ii]
We set sail September 30, but to get a head start, I am giving the author’s proofs of what I plan for the October tours below. A link to the materials from the LSE Graduate Seminar, along with videos of the sessions, all on phil-stat-wars.com, is at the end of this post.
Interested in joining us? Please email Jean Miller (jemille6@vt.edu), with your info, and she will send you a clean copy of the monthly materials.
In the initial writing of this book, I benefitted enormously from the comments on early drafts by readers of this blog. This fall, I will share some updates, including remarks on what I would change. [iii] I’m hoping for some great discussion in the comment sections on this blog!
I. (October 2024) Introduction: Controversies in Phil Stat:
Readings: SIST: Preface, Excursion 1 (click on links below for proofs)
Preface
Excursion 1 Tour I
Excursion 1 Tour II
Notes/Outline of Excursion 1 Postcard
Information Items for SIST
-Captain’s (2023) Library: Bibliography–Souvenirs (here)
-Summaries of 16 Tours (abstracts & keywords)
Please share this announcement with anyone you think might want to follow us. Here’s a link to the LSE Graduate Research Seminar:
LSE Research Seminar PH500 21 May – 25 June, 2020
[i] Except that I will sneak in an additional meeting during the last couple of months. There was a fifth, bonus meeting in the LSE seminar, with
David Hand–an optional excursion.
[ii] I expect to find errors in this first announcement, and will mark updates with v1, 2 etc.
[iii] If I were writing a new edition, for example.
.
Below is an email exchange that Andrew Gelman posted on this day 5 years ago on his blog, Statistical Modeling, Causal Inference, and Social Science. (You can find the original exchange, with its 130 comments, here.) Note: “Me” refers to Gelman. I will share my current reflections in the comments.
Exchange with Deborah Mayo on abandoning statistical significancePosted on September 11, 2019 9:52 AM by AndrewThe philosopher wrote:
The big move in the statistics wars these days is to fight irreplication by making it harder to reject, and find evidence against, a null hypothesis.
Mayo is referring to, among other things, the proposal to “redefine statistical significance” as p less than 0.005. My colleagues and I do not actually like that idea, so I responded to Mayo as follows:
I don’t know what the big moves are, but my own perspective, and I think that of the three authors of the recent article being discussed, is that we should not be “rejecting” at all, that we should move beyond the idea that the purpose of statistics is to reject the null hypothesis of zero effect and zero systematic error.
I don’t want to ban speech, and I don’t think the authors of that article do, either. I’m on record that I’d like to see everything published, including Bem’s ESP ~~paper~~ data and various other silly research. My problem is with the idea that rejecting the null hypothesis tells us anything useful.
Mayo replied:
I just don’t see that you can really mean to say that nothing is learned from finding low-p values, especially if it’s not an isolated case but time and again. We may know a hypothesis/model is strictly false, but we do not yet know in which way we will find violations. Otherwise we could never learn from data. As a falsificationist, you must think we find things out from discovering our theory clashes with the facts–enough even to direct a change in your model. Even though inferences are strictly fallible, we may argue from coincidence to a genuine anomaly & even to pinpointing the source of the misfit.So I’m puzzled.
I hope that “only” will be added to the statement in the editorial to the ASA collection. Doesn’t the ASA worry that the whole effort might otherwise be discredited as anti-science?
My response:
The problem with null hypothesis significance testing is that rejection of straw-man hypothesis B is used as evidence in favor of preferred alternative A. This is a disaster. See here.
Then Mayo:
I know all this. I’ve been writing about it for donkey’s years. But that’s a testing fallacy. N-P and Fisher couldn’t have been clearer. That does not mean we learn nothing from a correct use of tests. N-P tests have a statistical alternative and at most one learns, say, about a discrepancy from a hypothesized value. If a double blind RCT clinical trial repeatedly shows statistically significant (small p-value) increase in cancer risks among exposed, will you deny that’s evidence?
Me:
I don’t care about the people, Neyman, Fisher, and Pearson. I care about what researchers do. They do something called NHST, and it’s a disaster, and I’m glad that Greenland and others are writing papers pointing this out.
Mayo:
We’ve been saying this for years and years. Are you saying you would no longer falsify models because some people will move from falsifying a model to their favorite alternative theory that fits the data? That’s crazy. You don’t give up on correct logic because some people use illogic. The clinical trials I’m speaking about do not commit those crimes. would you really be willing to say that they’re all bunk because some psychology researchers do erroneous experiments and make inferences to claims where we don’t even know we’re measuring the intended phenomenon?
Ironically, by the way, the Greenland argument only weakens the possibility of finding failed replications.
Me:
I pretty much said it all here.
I don’t think clinical trials are all bunk. I think that existing methods, NHST included, can be adapted to useful purposes at times. But I think the principles underlying these methods don’t correspond to the scientific questions of interest, and I think there are lots of ways to do better.
Mayo:
And I’ve said it all many times in great detail. I say drop NHST. It was never part of any official methodology. That is no justification for endorsing official policy that denies we can learn from statistically significant effects in controlled clinical trials among other legitimate probes. Why not punish the wrong-doers rather than all of science that uses statistical falsification?
Would critics of statistical significance tests use a drug that resulted in statistically significant increased risks in patients time and again? Would they recommend it to members of their family? If the answer to these questions is “no”, then they cannot at the same time deny that anything can be learned from finding statistical significance.
Me:
In those cases where NHST works, I think other methods work better. To me, the main value of significance testing is: (a) when the test doesn’t reject, that tells you your data are too noisy to reject the null model, and so it’s good to know that, and (b) in some cases as a convenient shorthand for a more thorough analysis, and (3) for finding flaws in models that we are interested in (as in chapter 6 of BDA). I would not use significance testing to evaluate a drug, or to prove that some psychological manipulation has a nonzero effect, or whatever, and those are the sorts of examples that keep coming up.
In answer to your previous email, I don’t want to punish anyone, I just think statistical significance is a bad idea and I think we’d all be better off without it. In your example of a drug, the key phrase is “time and again.” No statistical significance is needed here.
Mayo:
One or two times would be enough if they were well controlled. And the ONLY reason they have meaning even if it were time and time again is because they are well controlled. I’m totally puzzled as to how you can falsify models using p-values & deny p-value reasoning.
As I discuss through my book, Statistical Inference as Severe Testing, the most important role of the severity requirement is to block claims—precisely the kinds of claims that get support under other methods be they likelihood or Bayesian.
Stop using NHST—there’s speech ban I can agree with. In many cases the best way to evaluate a drug is via controlled trials. I think you forget that for me, since any claim must be well probed to be warranted, estimations can still be viewed as tests.
I will stop trading in biotechs if the rule to just report observed effects gets passed and the responsibility that went with claiming a genuinely statistically significant effect goes by the board.That said, it’s fun to be talking with you again.
Me:
I’m interested in falsifying real models, not straw-man nulls of zero effect. Regarding your example of the new drug: yes, it can be solved using confidence intervals, or z-scores, or estimates and standard errors, or p-values, or Bayesian methods, or just about anything, if the evidence is strong enough. I agree there are simple problems for which many methods work, including p-values when properly interpreted. But I don’t see the point of using hypothesis testing in those situations either—it seems to make much more sense to treat them as estimation problems: how effective is the drug, ideally for each person or else just estimate the average effect if you’re ok fitting that simpler model.
I can blog our exchange if you’d like.
And so I did.
Please be polite in any comments. Thank you.
I am posting this with Gelman’s approval. You might find it interesting to check out some of the 130 comments on his blog here. I invite you to share reflections in the comments to this post.
.
Georgi Georgiev
In online experimentation, a.k.a. online A/B testing, one is primarily interested in estimating if and how different user experiences affect key business metrics such as average revenue per user. A trivial example would be to determine if a given change to the purchase flow of an e-commerce website is positive or negative as measured by average revenue per user, and by how much. An online controlled experiment would be conducted with actual users assigned randomly to either the currently implemented experience or the changed one.
Despite excellent motivation, good alignment of interests, and a growing body of knowledge, unbiased estimates from several sources show the median true effect in online experiments to be approximately zero [1]. Half of the proposed changes to business websites and mobile apps would have had no effect or a detrimental effect on the respective business, had they been permanently released to all end users. The effects of most of the rest are measured in single digit percentages, except for a long and thin positive tail. Such a median value of estimated effect sizes gives rise to the need to statistically discern true from false null hypotheses while the prevalence of small effect sizes necessitates relatively high-powered experiments.
Due to the competitive nature of business enterprises, there is a constant push to improve one’s experimentation program. In such an environment, it did not take long to feel ripples from the calls to abandon statistical significance and the “Moving to a world beyond p < .05” special issue of The American Statistician. However, major moves away from p-values and frequentist (error) statistics were ongoing in online A/B testing for several years prior.
Online A/B testing falls for the Bayesian allureA noticeable shift from frequentist to Bayesian approaches occurred rapidly in 2015 and 2016 at which time two of the three most used A/B testing software providers shifted their statistical engines from simple fixed-sample frequentist tests to Bayesian approaches.
At that time, the major motivating factors behind the move away from statistical significance were, primarily:
The allure of ‘Stopping rules do not matter’ was inspired by a realization that stopping rules do matter if one cares to discern between false and true effects. This happened quickly in online experimentation due to the near real-time nature of the data. When you look at results that show your business losing (or failing to capture) 100’s of thousands of dollars per day / month, next to which there is a glowing sign saying ‘99% confidence’ (or even ‘100% confidence’), or an equivalently low p-value, it is easy to stop a test early. Do this a couple of times and one starts to realize that looking at statistics computed under a fixed-sample assumption hour-by-hour or on a daily basis is a really bad way to figure out what has a positive impact. One of the earliest short studies of the issue included the pointed quote: “A/B testing with repeated tests is totally legitimate, if your long-run ambition is to be wrong 100% of the time.” [2]. A less technical reason many were brought to similar realizations is that examinations of the expected combined impact of the tested changes turned out to be far below what was reported by the accounting team after the fact.
Instead of realizing the core issue with such a flawed process of optional stopping are the broken error guarantees, several leading software companies turned to measures which did away with such guarantees, namely different flavors of Bayesian inference and estimation [3].
The belief that Bayesian accounts of probability are better aligned with the objectives of online experimentation compared to frequentist ones relies on a combination of severe twisting of the meaning of words, oversimplification, and mischaracterization of frequentist inference. The major culprit in my view was a naïve mixing of decision-theoretic approaches and Bayesian methods. To avoid a major tangent in the article, those interested in an overview of the Bayesian v Frequentist debate in online experimentation should refer to reference [4].
Developments since the call to abandon statistical significanceIn the years since the 2019 special issue [of The American Statistician], the main concern raised about p-values has changed slightly with a focus on insisting on how one is bound to misinterpret a p-value. These critiques unsurprisingly include issues of inferring a non-significant result to mean there is no real effect, of interpreting the observed p-value as the probability of the null hypothesis being true or false, and other valid concerns. Prominently featured is the argument that proposed Bayesian alternatives such as Bayes factors and posterior odds ratios are somehow easier to grasp, despite their objectively higher complexity.
The debate in the online experimentation community is therefore no different than the broader scientific discussion on the issue. It is, however, notable that what has primarily been put forward by some of the highest authorities in the business consists of simplistic explanations and appeals to the ‘intuitiveness’ of Bayesian probability.
How intuitive is Bayesian probability, really?As a prime example, this “definition” comes from the product documentation of Google Optimize – a widely used A/B testing tool between 2016 and late 2023:
“Probability to be best tells you which variant is likely to be the best performing overall. It is exactly what it sounds like — no extra interpretation needed!”
The only clarification to the above statement comes in the form of an answer to the question “What is “probability to beat baseline”? Is that the same as confidence?”, with the answer reading: “probability to beat baseline is exactly what it sounds like: the probability that a variant is going to perform better than the original”.
The above is the whole definition users of the software were supposed to be satisfied with in regard to the Bayesian measures of probability they were presented. I believe this betrays entitlement and arrogance on the part of Bayesians not typical elsewhere.
Skeptical of such appeals to intuitiveness, I conducted a small poll among practitioners to see if they really understand probability in Bayesian terms. See [5] for the poll and its results, as well as a broader discussion on the topic, but I believe it is important to summarize it here.
The question was to imagine a test with a true null so no treatment was administered to either group and how a ‘probability’ measure would change from day one with 1000 users per group to day ten with 10,000 users per group. No definition of ‘probability’ was given on purpose. The possible answers were that such a ‘probability’ would either ‘Increase substantially’, ‘Decrease substantially’, or ‘Remain roughly the same as on day one’.
The third option is what is most likely to happen with most Bayesian software used in the field, including the most popular one (Google Optimize). In all of them the posterior odds would remain roughly unchanged. Yet, that option was only chosen by less than a third of respondents. Two thirds answered ‘Decrease substantially’ which would make sense in a logical construct in which the null hypothesis is either true or false (a frequentist view) and reflects the behavior of a consistent estimator of such a ‘probability’.
While the poll had just 61 respondents, they have all self-identified within the higher brackets of online A/B testing practitioners. By no means conclusive on its own, the poll remains a rare attempt to quantify the merit of the claim that Bayesian probability is more intuitive than frequentist probability. I would challenge all those critical of p-values and frequentist probability to conduct better polls and show if it is indeed the case that other frameworks are more intuitive.
Other recent developmentsA relatively new angle of attack against p-values that has gained some traction in the last couple of years focuses on replicability as well as false positive risk (FPR) in line with Ioannidis (2005) [6] and Benjamin (2017) [7]. As framed by Colquhoun (2017) [8]: “We wish to answer this question: If you observe a ‘significant’ p-value after doing a single unbiased experiment, what is the probability that your result is a false positive?”. The inability of the p-value to give an answer to the above question is pointed out as a deficiency in need of addressing. The proposed way to solve it is by supplementing or outright replacing the reporting of p-values with reports of FPR, a.k.a. false positive probability.
A good example of the argument for false positive risk can be found in a paper by Kohavi, Deng & Vermeer (2022) [9]. Its motivating example is an egregious misuse of statistical inference and estimation deserving of every critique imaginable. The paper contains several good points to that effect. However, it also features the following suggestions:
p-value thresholds in online experimentationIt should be noted that there is no consensus threshold in the industry as a whole. Different companies or departments might impose their own thresholds or stick to textbook alpha of 0.05 for lack of a deeper understanding. More advanced teams may choose the threshold for each A/B test or each type of A/B tests performed such that it reflects the potential impact of the decision(s) to which it is relevant using some kind of a decision framework. A sample of thresholds used is available in [1].
In business experiments the risks and rewards associated with any experiment are often quantifiable. One can therefore arrive at significance thresholds and sample sizes which result in (roughly) optimal balance between risk and reward. For example, a company might employ a less strict threshold of 0.1 or 0.05 for mundane tests as part of regular quality assurance, whereas high stakes experiments might be subject to much lower thresholds in terms of both statistical significance and type II errors.
False positive riskThe main issue with using false positive risk is that it cannot be objectively computed for any single online experiment. It also does not surface any test-specific information which is not already contained in the p-value, but just augments the p-value with data from a set of other experiments with always questionable relevance. It relies on the assumption that the experiment at hand is drawn randomly from a sample much like a set of previously observed experiments, but that is not at all what happens in practice.
A further major issue is that the formula typically put forth for the calculation of FPR does not compute what the FPR concept is defined as, namely the probability that a statistically significant result is a false positive. This is something I’ve examined in much detail in [10].
Where is online experimentation heading to?Despite examples like the above-mentioned work, in the past several years there has been no noticeable overall shift away from p-values and towards alternative measures of evidence, or even toward preferring lower p-value thresholds on Bayesian grounds. Critiques of p-values and proposals of alternatives remain a side topic in most of the published research as it continues to focus on improving the efficiency of existing methods, the removal of sources of bias, as well as dealing with violations of standard model assumptions in different scenarios.
For what that’s worth, experimentation programs at high profile corporations continue to share mostly experiments conducted in a typically frequentist fashion. I’m aware of some experimentation programs which exhibit methodological eclecticism by offering users the ability to conduct both frequentist and Bayesian tests, with the caveat that most seem to use uniform or noninformative priors.
At the high level, to the extent to which the field had swayed Bayesian in the mid-to-late 2010s, it seems it may have lately been headed back to error-statistical territory. The adoption of Bayesian methods since 2019 is either mostly unchanged or somewhat on the decline. The decline I believe is partly due to the discontinuation of Google Optimize in late 2023. It also has to do with the rapidly increasing popularity of frequentist sequential testing such as group sequential tests and methods of so-called ‘Always valid inference’ [11][12]. Many vendors have added such methods in just the past couple of years as there is now a wide recognition that optional stopping is an issue for the trustworthiness of test outcomes. Yet, reliable numbers on how many tests are conducted using frequentist methods vs Bayesian ones are near-impossible to come by so estimations come by proxy through the rough numbers of companies using particular vendors, methods used in publicly shared tests in various cases, methods discussed in research papers, etc.
Share your reflections and questions in the comments on this post.
References:
[1] Georgiev, G. (2022) What Can Be Learned From 1,001 A/B Tests?, https://blog.analytics-toolkit.com/2022/what-can-be-learned-from-1001-a-b-tests/
[2] Downey A. (2011) Repeated tests: how bad can it be? https://allendowney.blogspot.com/2011/10/repeated-tests-how-bad-can-it-be.html
[3] Deng A., Lu J., Chen S. (2016) Continuous Monitoring of A/B Tests without Pain: Optional Stopping in Bayesian Testing; https://doi.org/10.1109/DSAA.2016.33
[4] Georgiev, G. (2020) Frequentist vs Bayesian Inference, https://blog.analytics-toolkit.com/2020/frequentist-vs-bayesian-inference/
[5] Georgiev, G. (2020) Bayesian Probability and Nonsensical Bayesian Statistics in A/B Testing, https://blog.analytics-toolkit.com/2020/bayesian-probability-and-nonsensical-bayesian-statistics-in-a-b-testing/
[6] Ioannidis J.P.A (2005) Why Most Published Research Findings Are False; https://doi.org/10.1371/journal.pmed.0020124
[7] Benjamin, D.J., et al. (2018) Redefine statistical significance; https://doi.org/10.1038/s41562-017-0189-z
[8] Colquhoun, D. (2017) The reproducibility of research and the misinterpretation of p-values. Royal Society Open Science (4). https://doi.org/10.1098/rsos.171085
[9] Kohavi, R., Deng, A., Vermeer, L. (2022) A/B Testing Intuition Busters, KDD ’22: Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp.3168–3177; https://doi.org/10.1145/3534678.3539160
[10] Georgiev, G. (2023) False Positive Risk in A/B Testing, https://blog.analytics-toolkit.com/2023/false-positive-risk-in-a-b-testing/
[11] Johari, R., Pekelis, L., Walsh, D. J. (2015) Always valid inference: Bringing sequential analysis to A/B testing, arXiv preprint arXiv:1512.04922
[12] Johari R., Koomen P., Pekelis L., Walsch D. (2017) “Peeking at A/B Tests: Why it matters, and what to do about it” https://doi.org/10.1145/3097983.3097992
About Georgi Z. Georgiev:
Author of “Statistical methods in online A/B testing”, founder of Analytics-Toolkit.com, statistics Instructor at CXL Institute
.
A topic that came up in some comments recently reflects a recent tendency to divorce statistical inference (bad) from statistical thinking (good), and it deserves the spotlight of a post. I always alert authors of papers that come up on this blog, inviting them to comment, and one from Christopher Tong (reacting to a comment on Ron Kenett) concerns this dichotomy.
Response by Christopher Tong to D. Mayo’s July 14 comment
TONG: In responding to Prof. Kenett, Prof. Mayo states: “we should reject the supposed dichotomy between ‘statistical method and statistical thinking’ which unfortunately gives rise to such titles as ‘Statistical inference enables bad science, statistical thinking enables good science,’ in the special TAS 2019 issue. This is nonsense.” [Mayo July 14 comment here.]
I am the author of the paper whose title she attacks as “nonsense”. If she had read my paper she would know that, like Kenett, I am advocating placing statistical thinking at the center of statistical teaching and practice. The dichotomy that she thinks is false exists in much of actual teaching and practice, and is one that it seems both Kenett and I are trying to undo. The title of my paper reflects the real (not the ideal) situation, and if that’s “nonsense”, then (and I would agree) much of statistical teaching and practice is nonsense. Finally I note that mine is one of only two papers in the 2019 special issue that even contains the phrase “statistical thinking” in the title. I strongly recommend the other one, which offers a concrete solution to how the “integrating” that Kenett speaks of can be done in statistics education.
The views expressed are my own.
Response by Mayo to Tong:
MAYO: Thank you for your comment. My thinking was that it would be good to alert the authors of the papers Lakens discusses, and I’m glad that you have. I have read your paper, and, as much as your highly provocative title earns rewards in contexts such as the special issue in which it appears, in my thinking, it does an enormous disservice to statistical inference as “enabling bad science”. Your paper itself—which reviews many right-headed contributions—shows that the very insights and tools that your “good statistical thinking” requires are themselves at the foundations of frequentist error statistical methodology and depend upon statistical inference methods, formal and informal. The formal tools were developed as deliberate idealizations by the founders as exemplars—to check and improve our ordinary (pre-statistics) statistical thinking. Grasping the brilliance of how this works demands a clear understanding of the mathematical and conceptual tools.
I wrote a book Statistical Inferisence as Severe Testing: How to Get Beyond the Statistics Wars (CUP, 2018): SIST. You can find all 16 “tours” on this blog (in final draft form) in this post: Blurbs of 16 Tours: Statistical Inference as Severe Testing; How to Get Beyond the Statistics Wars (SIST).
The notion of statistical inference developed there is very different from your sterile depiction. You talk as if practicing scientists in fields that employ statistical method gain their first exposure to statistics at the point that they are doing applied research. This should not be true. High school students, if they are to be critical consumers of the policies and decisions that will affect them in their lives—let alone conduct research—should study statistical method, including experimental design. Nor can statistical researchers without a VERY clear understanding of statistical concepts and computations assume they only need to think about the domain field, and assume good science will emerge. They should have a deep grasp of the formal methods that others will use to check their models and results. Statistical significance tests and tests of statistical hypotheses more generally, are intimately connected to experimental design, as Fisher emphasized.
I worry that your paper warns these students off, claiming it will only endanger their ability to do good science. What a relief for the students! This is one of their hardest courses, and now they can point to an important journal that has an article that warns us NOT to study statistical inference.
Statistical significance tests are just one small part of statistical science, but they are piecemeal methods and cannot all be learned at one time. Fisher wrote a book, Statistical Methods and Scientific Inference; the integration of the two was there from the start.
Testing statistical assumptions is a crucial part of error statistical methods. You mention Box, but he is talking about Bayesian vs frequentist methods. Box considered that Bayesian inference gives the formal, deductive part of inference which, in his view, could enter only after the creative, inductive work of arriving at and testing a model, which he claimed requires statistical significance tests:
[S]ome check is needed on [the brain’s] pattern seeking ability, for common experience shows that some pattern or other can be seen in almost any set of data or facts. This is the object of diagnostic checks and tests of fit which, I will argue, require frequentist theory significance tests for their formal justification. (Box 1983, 57)
Yet you say “formal, probability-based statistical inference should play no role in most scientific research, which is inherently exploratory, requiring flexible methods of analysis that inherently risk overfitting”. Box disagrees, saying we need checks on such risks, and statistical significance tests provides that. Eye-balling the data won’t suffice. (I say this after having worked with Aris Spanos, an expert on testing model assumptions.) Whenever we use data to solve statistical problems we are doing statistical inference: this goes beyond the data, and thus it is inductive or ampliative. (A paper I wrote with David Cox in 2006 is called: “Frequentist Statistics as a Theory of Inductive Inference”.)
There is no suggestion whatever that the significance test would typically be the only analysis reported. In fact, a fundamental tenet of the conception of inductive learning most at home with the frequentist philosophy is that inductive inference requires building up incisive arguments and inferences by putting together several different piece-meal results. Although the complexity of the story makes it more difficult to set out neatly, as, for example, if a single algorithm is thought to capture the whole of inductive inference, the payoff is an account that approaches the kind of full-bodied arguments that scientists build up in order to obtain reliable knowledge and understanding of a field. (Mayo & Cox 2006, 82.)
Perhaps some pedagogical treatments of statistical inference methods are overly formal, allowing students to just use computers to get the answer. Maybe that’s what’s behind your saying that there’s a divorce between statistical inference (bad) and statistical thinking (good). I say that computing solutions by hand provides a much deeper understanding of methods, and of where our intuitive thinking about probability and statistical inference is often badly wrong. It seems you’re missing that the key rationale for using deliberately idealized models in statistics is in order to learn from data how they fail and how to improve them. Used correctly, they serve as references for severe testing.
Of course, as you stress, “exploratory” inquiry and model building require a data dependence that would not be kosher in a predesignated “confirmatory inquiry”. But even in exploratory inquiry, we can use data both to build and severely probe such questions as whether a given method or model ought to be modified, whether it will serve to find out what we want to know, despite approximations, etc. Moreover, in exploratory inference, there are still statistical assumptions that ought to, and can, be checked by methods with different assumptions, and triangulating results. Fisher, Neyman and many, many others gave us mathematics to show how various designs (e.g., randomizations), and remodeling of data allow “subtracting out” or compensating misspecifications. Contemporary methods go further, but puzzlingly, you reject all such “technical fixes”.
I’m inclined to think that John Byrd had it right (in his comment on Kenett—who I do not think shares your view of statistical inference as enabling bad science):
So, I say that the reasoning underlying these [data science] approaches was given to us by Fisher, Neyman, Pearson, Deming, Cohen, Cox and others from our past. If “data science” becomes ignorant of what statistics can teach us, they will end up re-inventing these same concepts that guide error control, sampling issues, etc. Then we all get to watch a younger generation think they invented such concepts. (Link to J. Byrd July 14 full comment)
Another response to Kenett that seems right-headed is that of Christian Hennig:
Then, on the other side, statistical methodology is to quite some extent a formalisation of principles of statistical thinking, and if we want to analyse formally the implications of our thinking (and more broadly how to do it best), we generate “statistical methodology” by modelling situations (probability models) and decision making (statistical methods, model-based or not). “Errors” and “error probabilities” are then relevant again in the sense that statistical thinking can be criticised by saying, “if you apply “statistical thinking principle”, i.e., method A in artificial situation B in which we know (as we can “control” the truth when assuming models) that we should arrive at conclusion C, in fact you will quite likely arrive at conclusion D which is opposite to C”, then we have learnt something about how statistical thinking can be led astray.
I don’t really think this kind of reasoning can be easily replaced, and I claim that quality statistical thinking needs to be informed by such knowledge. (Link to C. Hennig’s July 17 comment)
There are several other comments on Kenett’s post both before and after Tong’s that might interest y. I invite your thoughts in the comments.
References
Box, G. (1983). An Apology for Ecumenism in Statistics, in Box, G., Leonard, T., and Wu, D. (eds.), Scientific Inference, Data Analysis, and Robustness, New York: Academic Press, 51–84.
Tong, C. (2019). Statistical Inference Enables Bad Science; Statistical Thinking Enables Good Science. The American Statistician, 73(sup1), 246–261. https://doi.org/10.1080/00031305.2018.1518264
.
Professor Andrew GelmanHiggins Professor of Statistics
Professor of Political Science
Director of the Applied Statistics Center
Columbia University
(Trying to) clear up a misunderstanding about decision analysis and significance testing
Background
In our 2019 article, Abandon Statistical Significance, Blake McShane, David Gal, Christian Robert, Jennifer Tackett, and I talk about three scenarios: summarizing research, scientific publication, and decision making.
In making our recommendations, we’re not saying it will be easy; we’re just saying that screening based on statistical significance has lots of problems. P-values and related measures are not useless—there can be value in saying that an estimate is only 1 standard error away from 0 and so it is consistent with the null hypothesis, or that an estimate is 10 standard errors from zero and so the null can be rejected, or than an estimate is 2 standard errors from zero, which is something that we would not usually see if the null hypothesis were true. Comparison to a null model can be a useful statistical tool, in its place. The problem we see with “statistical significance” is when this tool is used as a dominant or default or master paradigm:
We have no desire to “ban” p-values or other purely statistical measures. Rather, we believe that such measures should not be thresholded and that, thresholded or not, they should not take priority over the currently subordinate factors. We also argue that it seldom makes sense to calibrate evidence as a function of p-values or other purely statistical measures.
For summarizing research, we recommend the acceptance of uncertainty and that “the p-value be demoted from its threshold screening role and instead, treated continuously, be considered along with the currently subordinate factors [e.g., related prior evidence, plausibility of mechanism, study design and data quality, real world costs and benefits, novelty of finding, and other factors that vary by research domain], as just one among many pieces of evidence.”
For publication, we again recommend against using statistical-significance-based thresholding or screening; as we pointed out, even those journals that screen, implicitly or explicitly, on p-values or other thresholds do not automatically publish every submission that contains a statistically-significant finding: these journals necessarily use other criteria to decide what to publish, and so we’re just saying to use these other criteria without applying a significance screen.
Finally, for decision making more generally, we recommend cost-benefit analysis accounting for experimental uncertainty. Or, when there are no clear costs and benefits, we recommend that decisions be made based on all the available results and “working directly with hypotheses of interest rather than reasoning indirectly from a null model.”
A point of dispute or confusion
Recently we had an online discussion regarding our recommendations for decision analysis. A recent guest post [by Daniël Lakens] on Deborah Mayo’s blog expressed the view that our decision rule will be in practice equivalent to a p-value. This claim is not correct—or, to put it another way, given any dataset summarized by a single point estimate, a decision threshold can be mapped to a p-value or a Bayes factor or a z-score or any other monotonic summary of that point estimate, but then the cutoff p-value (or whatever) will depend on the data, so it’s not a p-value (or whatever) threshold in the usual sense.
The applied point is that our recommended approach depends not on type I and II errors (or, for that matter, on type M and S errors) but rather on estimates of costs and benefits of the decision being considered, where costs and benefits are measured in terms of dollars, lives, or some other external units.
This is not to say that type I, II, M, and S errors are irrelevant! They can be useful concepts in helping to understand and evaluate statistical procedures. We just don’t want to use them as decision thresholds.
I think some of the confusion in our online discussion on decision thresholds came because there was confusion about what sort of decisions we were talking about. So I thought it would help to explain using a simple hypothetical example. We also presented this example on our blog where there is some discussion in comments. I thank Daniël Lakens for providing the comment that motivated this post and Deborah Mayo for the opportunity to add to this discussion.
Working through an example
Consider the following decision problem. Your company is considering adopting an innovation that would cost $10,000 and that has uncertain benefit. Your prior on the benefit is normal(0, $100,000). In order to inform your decision you conduct an experiment that gives you an unbiased estimate of the benefit with some standard error, s. To give some notation here, let C be the cost ($10,000 in this case), let sigma_0 be the prior sd ($100,000 in this case), and let theta be the benefit, so that theta – C is the net gain, and you get an estimate theta_hat ~ normal(theta, s) with prior theta ~ normal(0, sigma_0). Further assume that sigma_0 is a small amount of money relative to the size of your company, so that your utility is a linear function of the net gain.
Given the estimate theta_hat, what should your decision be? There are lots of rules you could use here, but I’m gonna go Bayesian, given that all the costs, benefits, and uncertainties are defined up front.
Given the data, the posterior mean of theta is (theta_hat/s^2)/(1/s^2 + 1/sigma_0^2), so the expected net benefit of adopting the innovation is (theta_hat/s^2)/(1/s^2 + 1/sigma_0^2) – C, and you should adopt the innovation if (theta_hat/s^2)/(1/s^2 + 1/sigma_0^2) > C, that is, if theta_hat > C(1 + (s/sigma_0)^2). We’re assuming C=10.
Here are a couple examples: If s = $10,000, then s/sigma_0 = 0.1 and the rule is to adopt the innovation if theta_hat > $10,100. If s = $100,000, then the rule is to adopt if theta_hat > $20,000.
Another possible decision rule is based on statistical significance: adopt the innovation if the estimate is statistically significantly greater than zero, given some p-value threshold, which in this normal model is equivalent to adopting if the z-score—the estimate divided by its standard error—exceeds some value. In the above notation, the rule would be to accept if theta_hat > A*s, where A is some pre-chosen value. A commonly-used threshold is A=2, but sometimes people want to be conservative and set A to some larger value such as 3, while other times people will set A to a lower value such as 1.5 so as not to stifle innovation.
The Bayesian rule and the statistical significance rule are different! Not just in their motivations, but also in their mathematical forms. The Bayesian rule is to accept the innovation if theta_hat > C(1 + (s/sigma_0)^2); the statistical significance rule is to accept the innovation if theta_hat > A*s. The classical rule has two parameters: the threshold z-score A and the standard error s. The Bayesian rule has three parameters: the cost of the innovation C, the prior sd sigma_0 of the benefit, and the standard error s.
You might try to wrestle the Bayesian and statistical significance rules into agreement by setting the threshold z-score A to an appropriate value, but the resulting rules will still be different functions of s. If you have a Bayesian rule given some values of C and sigma_0, there is no significance threshold A which will give you the equivalent rule.
You might think that all this could be saved if the Bayesian inference used a flat prior, but, no, not at all! A flat prior here corresponds to sigma_0 = infinity, which in turn implies the decision rule to adopt the innovation if theta_hat > C. This rule depends entirely on the point estimate, not on the standard error at all! For it to be a p-value threshold, the decision would have to depend on theta_hat/s, which it doesn’t.
This is all basic decision theory (modulo any algebra or calculation or conceptual errors I’ve made in my above quick derivation), and I hadn’t thought much about it recently until I came across a comment from psychology researcher Daniël Lakens on Deborah Mayo’s blog where he pointed to this quote from a paper I’d coauthored: “Instead, to the extent that decisions do need to be made about which lines of research to pursue further, we recommend making such decisions using a model of the distribution of effect sizes and variation, thus working directly with hypotheses of interest rather than reasoning indirectly from a null model,” and responded:
I [Lakens] think if I would see the authors do this in practice, any decision they make ends up being a dichotomous claim based on a critical value, and I will be able to recompute that critical value into a p-value. I am happy to be proven wrong, and all the authors need to do is to link me to a real life example where this is not the case.
The request for a real-life example was easy to satisfy; I provided some in my follow-up comment, pointing to chapter 9 of Bayesian Data Analysis, third edition, and this article about decision making for radon measurement and remediation where we went through all aspects of uncertainty and utility quantification in detail. Providing references was no big deal; it can be hard to find things in an unfamiliar literature, and I was happy to help Lakens by answering his question by supplying examples.
The more interesting question to me was how it was that someone who’d thought about decision analysis and hypothesis testing could have thought that “any decision [we] make ends up being a dichotomous claim based on a critical value, and I [Lakens] will be able to recompute that critical value into a p-value.” If a data-based decision is dichotomous (in the above example, adopt the proposed innovation or stick with the status quo), it can be said to be based on a critical value, but, as the above example shows, that critical value cannot in general be recomputed into a p-value—unless you allow the p-value threshold to itself depend on various features of the decision problem, the prior distribution, and the data (in the above example, C, sigma_0, and s), in which case this is not a p-value threshold in the usual sense in that it cannot be defined before seeing the data.
After some reflection, though, I think I understand what Lakens was talking about, and it’s good to have a chance to clarify the point.
In my decision problem as stated above, there are two decision options:
(i) Adopt the innovation, or
(ii) Stick with the status quo,
and the relative utility of option (i) compared to option (ii) is theta-C. Very direct; as I said, it’s standard decision theory from the 1940s through today.
But here’s another way to frame the problem. You take the same two decision options, but the goal is not to maximize dollars but to make the correct decision. If you take option (i), there is a loss L1 if the innovation was not beneficial (that is, if theta<0), and if you take option (ii), there is a loss L2 if theta>0. (In our problem with a continuous prior distribution, we can ignore the zero-probability event that theta=0 exactly. You could include that possibility and it wouldn’t change the basic features of the problem.) These are the costs of classical type I and II errors, and the relative utility of option (1) compared to option (ii) is – L11_{theta<0} + L21_{theta>0}.
If you then attack the problem from a Bayesian perspective, the relative expected utility of option (1) compared to option (ii) is L2Pr(theta>0|data) – L1Pr(theta<0|data) = (L1 + L2) * Pr(theta>0|data) – L1, and this is positive if Pr(theta>0|data) > L1/(L1+L2).
Let’s quickly check this calculation. If L1=L2, then the two sorts of losses are equivalent, so it makes sense that you’d adopt the innovation if and only if it is more likely than not to have a positive benefit. If L1>L2, then the loss of mistakenly adapting the innovation is higher than mistakenly sticking with the status quo, so there’s a higher threshold, etc.
For the problem above, Pr(theta>0|data) = Phi((theta_hat/s^2)/sqrt(1/s^2 + 1/sigma_0^2)), where Phi is the cumulative normal distribution function, and so the Bayesian rule is to adopt the innovation if Phi((theta_hat/s^2)/sqrt(1/s^2 + 1/sigma_0^2)) > L1/(L1+L2), which maps to adopting if theta_hat > s^2sqrt(1/s^2 + 1/sigma_0^2)(Phi^{-1}(L1/(L1+L2)). Again, the threshold depends on s, hence there is no critical p-value corresponding to this decision. Under a flat prior, though, with sigma_0=infinity, the decision rule is to adopt the innovation if theta_hat/s > L1/(L1+L2), which would be a p-value threshold. So maybe that’s what Lakens was thinking about.
From a decision-analytic perspective, I don’t think that this second formulation, in which the losses depend only on whether the optimal decision was chosen, makes sense. In any medical, business, or policy example I’ve seen, the benefit of the decision should depend on the actual costs and benefits of the decision options—indeed, that’s what McShane and I and the others were thinking when we wrote:
For regulatory, policy, and business decisions, cost-benefit calculations seem clearly superior to acontextual statistical thresholds. Specifically, and as noted, such thresholds implicitly express a particular tradeoff between Type I and Type II error, but in reality this tradeoff should depend on the costs, benefits, and probabilities of all outcomes.
I can’t be sure, but I’m guessing that what Lakens was talking about in his above-quoted comment was not a business or policy decision (whether to adopt an innovation of uncertain benefit relative to its cost) but rather a scientific decision of whether to accept or reject a hypothesis. All I can say there is that I don’t see decision analysis as being particularly relevant here because I don’t see the need for accepting or rejecting a hypothesis or scientific model. Here’s what we wrote in our paper:
While we see the intuitive appeal of using p-value or other statistical thresholds as a screening device to decide what avenues (e.g., ideas, drugs, or genes) to pursue further, this approach fundamentally does not make efficient use of data: there is in general no connection between a p-value—a probability based on a particular null model—and either the potential gains from pursuing a potential research lead or the predictive probability that the lead in question will ultimately be successful. Instead, to the extent that decisions do need to be made about which lines of research to pursue further, we recommend making such decisions using a model of the distribution of effect sizes and variation, thus working directly with hypotheses of interest rather than reasoning indirectly from a null model.
The general point here is that there is no general equivalence between data-informed cost-benefit analysis and statistical significance thresholds or Neyman-Pearson hypothesis testing more generally. If you really want to use a p-value threshold, you can reverse-engineer a Bayesian decision analysis that will get there, but you have to use this weird utility function where a gain of a billion dollars (or the discovery of a huge effect) is as good as a gain of 50 cents (or the discovery of a tiny effect which might not even hold up in the future).
Some reasonable arguments in favor of statistical significance
That all said, I can see potential good reasons for using p-values and statistical significance to make decisions:
– It’s a standard approach, so by doing this you can communicate with others using this set of methods, and your results will be comparable to other results obtained in this way.
– Bayesian inference and, even more so, Bayesian decision analysis, has lots of knobs you can turn, lots of ways for a researcher to mess up or cheat.
– Bayesian inference can be done in so many different ways. Even setting aside problems with researcher degrees of freedom, the advice to follow Bayesian methods does not in general provide much guidance to an applied researcher: there are so many Bayesian methods out there!
– If you’re already familiar with classical approaches, it might not be worth the effort to retool to learn other methods.
– Bayesian methods that have been effective for me in applications in pharmacology, political science, etc., might not be so appropriate in some areas of psychology.
– Our methods might work for us, but we actually could do just as well by judicious use of significance testing and p-values.
– Just cos significance testing has problems, it doesn’t mean that any particular specified alternative, Bayesian or otherwise, will perform better.
Ultimately, despite these issues, I still will recommend abandoning statistical significance, but I recognize that the methods used by myself and my collaborators are far from perfect. Other people can have good reasons for not following our advice. Also, as a researcher in statistical theory and methods, I remain interested in statistical significance and p-values, if for no other reason than that it behooves us to understand these methods that so many people remain committed to.
The point of this post is not to argue to abandon statistical significance—for that, I recommend our article!—but rather to highlight the fundamental differences between Bayesian decision analysis and statistical significance thresholds. They’re not just two ways of doing the same thing, and that’s important!
Real-world decisions are typically more complicated
To keep things simple, I’ve focused the above discussion on a single binary decision: go with the innovation or stick with the status quo. Blake McShane points out that “more typically, however, decisions are not binary; the utility function is unknown or different stakeholders disagree about it, etc.; the data generating process is unknown, etc. I would think all of these would accentuate differences between the decision theoretic and significance testing approaches.” Blake also points out that adopting a decision theoretic point of view does not require adopting a Bayesian approach to statistics. I find Bayesian inference to be a convenient framework here, but the key point is that if the decision depends on real-world costs and benefits—that is, if the utility or equivalent depends on theta, not just on theta>0 or theta!=0 or whatever—then decisions do not in general map to statistical-significance thresholds.
Sander Greenland also pointed out that what I’m calling the classical or significance-testing framework itself comes in many forms and has a long history. The present post is not intended to be any summary or tutorial of that approach. Rather, I’m focusing on the lack of equivalence between cost-benefit decision rules and significance-testing decision rules.
We’re incoherent too
Yes, our approach has holes! Our models are always under construction and we’re always checking them. A model check can be seen as a sort of hypothesis check, where if the data shows important features that are not predicted by the model, we need to alter or expand the model. Our predictions are stochastic, and so any such evaluation must be probabilistic. So . . . when we do the necessary step of model checking, we’re doing something like significance testing, asking whether the data differ from the model in important ways, more than could be explained by chance alone. This is a point that Deborah Mayo and others have made, and which we have discussed too. To loop back to “Abandon statistical significance,” there’s an awkwardness here because we are using some aspects of statistical significance when we do model checking. I’ll just say that we’re not using significance as a threshold; we’re using it as one piece of information in our decision making. Still, I wanted to acknowledge the tension that remains in our workflow.
Note from Mayo: This discussion began in the comments to Lakens’ guest post, part 1 and part 2 . We welcome your questions and remarks in the comments to this blogpost.
.
Aris Spanos
Wilson Schmidt Professor of Economics
Department of Economics
Virginia Tech
The following guest post (link to PDF of this post) was written as a comment to Mayo’s recent post: “Abandon Statistical Significance and Bayesian Epistemology: some troubles in philosophy v3“.
On Frequentist Testing: revisiting widely held confusions and misinterpretations
After reading chapter 13.2 of the 2022 book Fundamentals of Bayesian Epistemology 2: Arguments, Challenges, Alternatives, by Michael G. Titelbaum, I decided to write a few comments relating to his discussion in an attempt to delineate certain key concepts in frequentist testing with a view to shed light on several long-standing confusions and misinterpretations of these testing procedures. The key concepts include ‘what is a frequentist test’, ‘what is a test statistic and how it is chosen’, and ‘how the hypotheses of interest are framed’.
The first thing that needs to be brought out immediately is that frequentist testing is framed in terms of applied mathematics and its proper framing is of utmost importance in forefending confusions and misinterpretations. Regrettably, philosophers of science tend to undervalue that formalism as cumbersome and needless, which rings hollow when one compares the statistical framing with that of first order logic and other parts of epistemology!
To understand the frequentist testing one needs to place it in its proper context which is a particular statistical model comprising several probabilistic assumptions whose validity of the particular data is of paramount importance. The simplest statistical model is a simple Bernoulli model denoted by:
Xk ~BerlD(θ, θ(1- θ)), 0 < θ <1, xk=0, 1, k=1, 2, —n, …, (1)
where ‘BerIID’ stands for Bernoulli (Ber), Independent and Identically Distributed (IID). The underlying random variable X takes only two values, X =1, say Head (H) and X =0 Tails (T), with P(X =1)= θ and P(X =0)=1− θ. In the context of the statistical model in (1), the relevant data x0:=(x1, x2, …, xn) are viewed as a single realization of the sample X:=( X1, X2 , …, Xn). Note that a random variable is denoted by a capital letter (Xk) and the corresponding observation by a small letter (xk); this minor notational point will avert numerous confusions in practice!
The first issue that arises in practice is how to frame the hypothesis of interest in frequentist testing. This arises because there are a number of confusions between R.A. Fisher’s (1922) framing of the null hypothesis (H0) and the Neyman-Pearson (N-P) (1933) framing that includes both the null (H0) and the alternative (H1) hypothesis. The truth is that the objective is identical for both framings: learn from data x0:=(x1, x2, …, xn) about the ‘true’ value, say θ of the unknown parameter θ. In light of that, the entire parameter space (0, 1) is relevant for frequentist testing since theoretically θ** can take any one of the values in this interval. Hence, the framing of the hypotheses needs to cover the whole of the parameter space. This was first stated in the Neyman-Pearson (N-P) lemma that provided the cornerstone of frequentist testing and included Fisher’s framing as a special case; see Note 1 below.
In summary, the proper way to specify the hypotheses of interest in frequentist testing is to framed them in terms of the model’s unknown parameter(s) θ, and ensure that they constitute a partition of the parameter space. For the statistical model in (1), partitioning can take various forms, including:
H0: θ ≤ θ0 vs. H1: θ > θ0, where θ0 =.5. (2)
What is a frequentist test? It is not just a statistic, say whose distribution in this case is Binomial (Bin):
derived by assuming the validity of the statistical model assumptions ‘BerIID’. Any attempt to use (3) to define any error probabilities, including the p-value (Titelbaum (2022), p. 465), is improper and will give rise to the wrong inference.
A frequentist test comprises two equally important components. For the hypotheses in (2), the optimal N-P test takes the form:
In frequentist testing, the test statistic is always a distance function d(X) framed in terms of a statistic and a frequentist test also includes a rejection region whose choices are neither arbitrary or whimsical. Indeed, the choice of d(X) and C1(α) is interrelated and needs to satisfy two conditions. The first is that the distribution of d(X) can be evaluated under both H0 and H1, and the second is that it has to give rise to an optimal test in terms of learning from data about θ, and there are good choices for d(X) and C1(α) and bad ones, which are evaluated in terms of their capacity to approximate θ** framed in terms of their pre-data type I and II error probabilities.
In the case of the test in (4), the N-P optimal theory of testing renders it Uniformly Most Powerful (UMP) whose relevant sampling distributions are:
where ‘≈’ indicates an approximation of the scaled Binomial by the Normal distribution; see Spanos (2019).
Example 1. Consider particular data x0 representing 17 Hs out of n=20 flips (Titelbaum, 2022, p. 465):
HHHTHHHHHTHHHHTHHHHH (6)
In light of the fact that this p-value is based on n=20 observations, this indicates a clear departure from θ0 =.5. It is important to emphasize that the p-value in (8) is NOT the probability of the particular configuration in (6) occurring as a realization of the sample!
The question that arises at this stage is ‘what about the validity of the probabilistic assumptions comprising the invoked statistical model in (1) for the particular data x0?*’ If any of these assumptions are invalid for x0 the above inference results are unreliable! The Bernoulli assumption is innocuous since any data with two outcomes can always be framed with such a distribution. The IID assumptions, however, are often invalid with real data. When IID is invalid, the assumed distributions under both H0 and H*1 in (5) are invalid, inducing sizeable discrepancies between the actual error probabilities and the nominal ones derived by assuming the validity of IID! This renders any evaluations based on their tail areas highly misleading; see Spanos (2019), ch. 15. Hence, in practice one needs to test the IID assumptions before using x0 in the context of the statistical model in question to draw inferences about *θ.
Example 2. Let us return to the particular data x0 representing 17 Hs out of n=20 flips, and ask the question: Are the IID assumptions valid for data x0? Although there are many ways to test the IID assumptions (Spanos, 2019, ch.15), a particularly simple misspecification test is the runs test. A ‘run’ is a segment of the sequence of outcomes consisting of adjacent identical elements which are followed and proceeded by a different symbol. For the observed sequence in (6), the sequence of 7 runs is shown below:
The runs test compares the actual number of runs R with the number of expected runs E(R)— assuming that the sample is IID process — to construct the runs test:
Note that the above runs test can be extended to a more sophisticated test which accounts, not only for the number of runs, but also their different lengths, e.g. run 1 has length 3, run 2 has length 1 and run 3 has length 5.
In this case the sample size n =20, is rather small, but we can apply the test anyway for illustration purposes:
where the p-value indicates departures from the IID assumptions at any significance level α ≥ .032. One can dispute this particular threshold α but in light of the small sample size n=20, this threshold ensures that the test has sufficient power to detect departures from IID. On the issue of why the p-value should always be one-sided is because the data render irrelevant one of the two tails post-data (when dR(x0) is revealed). This is the key difference between the p-value and type I and II error probabilities that are pre-data; see Spanos (2019). ch. 13.
It is important to emphasize that the runs test in (10) is probing the validity of the IID assumptions underlying the invoked statistical model in (1), and thus it poses very different questions to the data when compared to the N-P test in (5) which assumes the validity of the assumptions and probes for θ! Indeed, misspecification testing should not be framed in terms of the parameter(s) of the underlying statistical model because it probes outside the boundaries of the given model, as opposed to N-P testing that probes within its boundaries*. In practice, misspecification testing predates N-P testing to secure the validity of the invoked statistical model and thus the reliability of the ensuing inference; see Mayo and Spanos (2004).
More broadly, the accept/reject H0 results and small/large p-values do not provide evidence for or against particular hypotheses since the sample size n in conjunction with the pre-specified α play a crucial role in transforming such results into evidence using their post-data severity evaluation that outputs the warranted discrepancy from the null value; see Mayo and Spanos (2006). For a given α, ignoring the sample size n is likely to give rise to the fallacies of acceptance and rejection. This is due to the inherent trade-off between the type I and II error probabilities, which implies that for a given α the power of the test increases with n. The post-data severity evaluation provides an evidential account of the accept/reject H0 results by taking fully into account the sample size n; see Mayo (2018), Spanos (2023).
Note 1. The widely held impression that Fisher’s significance testing and N-P testing are two very different approaches is just another misconstrual of frequentist testing. They are not so different! Stating just a point null hypothesis H0: θ = θ0 and a threshold α, Fisher brings into play the type I and II error probabilities (and power) indirectly into his significance testing. Don’t take my word for this claim, read Fisher (1935), pp. 21-22 describing how the power of the test increases with the sample size n but calling it the ‘sensitivity’ of a test. How could one explain the acerbic Fisher vs. Neyman-Pearson exchanges? They were talking passed each other since the type I and II error probabilities are pre-data — framing the capacity of the test —, and Fisher’s p-value is post-data evaluation indicating potential departures from H0 in light of the observed test statistic. Fisher can get away without specifying an alternative hypothesis H1 since the sign of the observed test statistic indicates the direction of departure, eliminating one of the two tails; see Spanos (2019), ch. 13. It should be noted that Fisher constructed his numerous test statistics using intuition, but they turned out to define optimal frequentist tests when supplemented by an appropriate rejection region, and Neyman and Pearson (1933) give him credit for that.
Note 2. For further discussions on frequentist testing, its misinterpretations and misuses, including the base-rate fallacy, the Jeffreys-Lindley Paradox, Akaike type selection criteria, etc., see the following papers:
References
[1] Fisher, R.A. (1922) “On the mathematical foundations of theoretical statistics”, Philosophical Transactions of the Royal Society A, 222: 309-368.
[2] Fisher, R.A. (1935) The Design of Experiments, Oliver and Boyd, Edinburgh.
[3] Mayo, Deborah G. (2018) Statistical inference as severe testing: How to get beyond the statistics wars, Cambridge University Press.
[4] Mayo, D.G. and A. Spanos (2004) “Methodology in Practice: Statistical Misspecification Testing”, Philosophy of Science, 71: 1007-1025.
[5] Mayo, D.G. and A. Spanos. (2006) “Severe Testing as a Basic Concept in a Neyman-Pearson Philosophy of Induction”, The British Journal for the Philosophy of Science, 57: 323-357.
[6] Neyman, J. and E.S. Pearson (1933) “On the problem of the most efficient tests of statistical hypotheses”, Philosophical Transactions of the Royal Society, A, 231, 289-337.
[7] Spanos, Aris (2019) Introduction to Probability Theory and Statistical Inference: Empirical Modeling with Observational Data, 2nd edition, Cambridge University Press, Cambridge.
[8] Spanos, Aris (2023) “Revisiting the Large n (Sample Size) Problem: How to Avert Spurious Significance Results.” Stats 6(4): 1323-1338.
.
Professor Yudi Pawitan
Department of Medical Epidemiology and Biostatistics
Karolinska Institutet, Stockholm, Sweden
[An earlier guest post on this topic by Y. Pawitan is Jan 10, 2022: Yudi Pawitan: Behavioral aspects in the statistical significance war-game]
Behavioral aspects in the statistical significance war-game
I remember with fondness the good old days when the only ‘statistical war’-game was fought between the Bayesian and the frequentist. It was simpler and the participants were for the most part collegial. Moreover, there was a feeling that it was a philosophical debate. Even though the Bayesian-frequentist war is not fully settled, we can see areas of consensus, for example in objective Bayesianism or in conditional inference. However, on the P-value and statistical significance front, the war looks less simple since it is about statistical praxis; it is no longer Bayesian vs frequentist, with no consensus in sight and with wide implications affecting the day-to-day use of statistics.
Typically, a persistent controversy between otherwise sensible and knowledgeable people might indicate we are missing some common perspectives or the big picture. In complex issues, there can be genuinely distinct aspects about which different parties disagree and, at some point, agree to disagree. I am not sure we have reached that point yet, with each side still working to persuade the other side about the faults of their position. For now, I can still concur with Mayo (2021)’s appeal that at least the umpires – reviewers and journals editors – recognize (a) the issue at hand and (b) that genuine debates are still ongoing, so it is not yet time to take sides.
I have previously described my disagreement with the ideas of banning the P-value or just its threshold, and retiring statistical significance and categorical significance statements (Pawitan, 2020). To summarize briefly,
However, no matter how much we have talked, debated and argued, it seems the disagreement persists. So, instead of repeating or expanding the arguments, here I would like to discuss where genuine disagreements can occur and be accepted. In game-theoretic or behavior-economic analyses, it is accepted that rational-intelligent individuals can act differently, thus disagree, reflecting different personal preferences or utility functions. In this game-theoretic framework, the differing parties accept each other’s position, and they feel no need to persuade and change each other’s opinion.
So let’s start by assuming we are all rational-intelligent players: we fully understand the correct meaning and usage of the P-value in particular and statistical inference in general. Excluding deliberate frauds, most objections to the P-value or its threshold seem to refer to at least three concerns:
Let’s start with the last concern first. Since the P-value threshold controls the false-positive rate under the null, we must suppose that the concern regarding false positives is either due to a belief that (a) the reported P-value does not represent the true level of uncertainty, or (b) that the standard threshold – such as 0.05 – is too large. The first issue will arise, for instance, when one reports winners (i.e., selective inference) in a multiple testing situation without properly accounting for the multiplicity. In principle, this problem can be cured by more rigorous inference procedures, for example, using the false discovery rate methods.
Regarding (b), reducing the threshold will increase the false-negative rate, so it’s not cost-free. It seems to me the attitude is: don’t bother trying to balance these two errors, let’s just drop the P-value or its threshold. This reflects a preference that is perhaps amenable to further theoretical analysis and discussion, for example in relation to replicability/validation, but in any case, it will not affect the other two concerns.
The first two concerns are different: They reflect some degree of distrust of non-experts and the gullible public. Although I share these concerns, there is a genuine difference in where I put them on my own utility scale relative to the advantages of having the P-value. Furthermore, in the game theory for a social setting, we talk about a personal preference and a social preference. On any single issue, these can be distinct or may also coincide. For instance, personally I would never consider abortion, but I will not impose my personal preference on other people, so in my social preference, abortion is acceptable. However, somebody else might not only reject abortion for herself, but also wants to live in a society that does not allow abortion, so would militate for its ban. Even in liberal countries, where individual preferences/liberty are supposed to be paramount, there are many issues where you might want to project your personal preference as the social preference: vaccination, addictive drugs from marijuana to cocaine, pornography, prostitution, open-carry firearms, gambling, death penalty, euthanasia, etc. These social issues are typically solved by democratic means, directly in referenda or indirectly by decisions of elected representatives. In either case, there is a mechanism – such as voting – and an authority that can impose the agreed decision as a social contract to the whole society.
What kind of social solution is suitable for something like the P-value war? It is indeed a challenging problem, since we have (i) no boundary that defines the legitimate stake-holders (academic statisticians? +applied statisticians? +chartered statisticians? +statistically literate scientists?+…?), (ii) no formal mechanism to express and combine preferences, and (iii) no real authority to impose any agreed decision. As in society in general, social norms that are not formally democratically controlled are dictated by culture. But how cultures evolve and which social rules get adopted are not predictable; in particular they may not be decided by the majority. They may well depend on a small number of influencers, in our case perhaps top-ranked-journal editors, or top-ranked statisticians or scientists. Nassim Taleb (2020) highlighted how social changes can be driven by a small intolerant/loud minority in the face of a tolerant/quiet majority. For instance, the few editors of the journal Basic and Applied Social Psychology banned the P-value and statistical inference, and the numerous authors must acquiesce regardless of their personal views.
I have never seen any rigorous opinion poll on the use of P-value. An informal poll (n=303) done during a public debate on the P-value at the National Institute of Statistical Sciences (Oct 2020) showed that a clear majority (55%) of the audience would use the P-value alone vs 28% both the P-value and Bayes Factor vs 8% the Bayes Factor alone vs 13% Neither (private communication with JL Rosenberger who did the poll; see the transcript in https://errorstatistics.com/2020/12/13/the-statistics-debate-niss-in-transcript-form-question-1/) Formal professional bodies such as the American Statistical Association (ASA) or the Royal Statistical Society (RSS) could perhaps run such a poll. They will of course still face the boundary problem I mention above, as their members do not represent all users of statistics, but it will be a start. A rigorous poll would be useful, so we can judge the extent of the division within our profession. One may argue strongly that science is not a democratic enterprise: 1000 dissenting but wrong votes cannot beat a single correct vote. But on an issue with no definite right-wrong answer, such as the use of the P-value, a large support for banning it or its threshold should encourage all of us to come to a workable consensus. But a small support – please do not ask for a threshold! – should give the intolerant/loud minority pause for thought.
References
Luo, J. et al. (2007) Oral use of Swedish moist snuff (snus) and risk for cancer of the mouth, lung, and pancreas in male construction workers: a retrospective cohort study. Lancet, 369: 2015–20.
Mayo, D. (2021) The statistics wars and intellectual conflicts of interest. Conservation Biology. https://conbio.onlinelibrary.wiley.com/doi/full/10.1111/cobi.13861
Pawitan, Y. (2020). Defending the P-value. https://arxiv.org/abs/2009.02099
Pawitan, Y. and Lee, Y. (2024). Philosophies, Puzzles and Paradoxes. Boca Raton: Chapman and Hall/CRC Press.
Taleb, N. N. (2020). Skin in the Game: Hidden Asymmetries in Daily Life. New York: Random House.
.
Has the “abandon significance” movement in statistics trickled down into philosophy of science? A little bit. Nowadays (since the late 1990’s [i]), probabilistic inference and confirmation enter in philosophy by way of fields dubbed formal epistemology and Bayesian epistemology. These fields, as I see them, are essentially ways to do analytic epistemology using probability. Given its goals, I do not criticize the best known current text in Bayesian Epistemology with that title, Titelbaum 2022, for not engaging in foundational problems of Bayesian practice, be it subjective, non-subjective (conventional), empirical or what some call “pragmatic” Bayesianism. The text focuses on probability as subjective degree of belief. I have employed chapters from it in my own seminars in spring 2023 to explain some Bayesian puzzles such as the tacking paradox. But I am troubled with some of the examples Titelbaum uses in criticizing statistical significance tests. I only came across them while flipping through some later chapters of the text while observing a session of my colleague Rohan Sud’s course on Bayesian Epistemology this spring. It was not a topic of his seminar.
1. A test of statistical significance. What is it? First of all, it’s a test of a statistical hypothesis. There is a set of possible outcomes, a sample space, and a statisticatest hypothesis H0 which assigns probabilities (or densities) to outcomes modeled in terms of the distribution of a random variable X. There is a test statistic d(X), a function of data x and H0, such that the larger the observed d(x) the more improbable x is, computed according to H0. H0 provides this probability assignment by hypothesizing the value of an unknown parameter θ in a statistical model M. θ is viewed as a fixed, but unknown quantity, except for special cases where θ may itself be regarded as a random variable that takes on values with given probabilities.
The test is a rule that maps observed values of d(x) into either “reject H0” or “do not reject H0” in such a way that there is a low probability of erroneously rejecting H0 and a much higher probability of correctly rejecting H0. “Reject H0” and “fail to reject H0” are generally interpreted as x is evidence against H0, or x fails to provide evidence against H0, (which is not the same as evidence for H0.) Titelbaum focuses on simple Fiasherian tests, so I will too.
Coin tossing (Bernouilli) trials. For instance, in a random sample X of n coin tosses (X = X1, X2,…Xn), H0 might assert that θ, the probability of heads on each, trial is .5. A test statistic d(X) would be the difference between the observed proportion of heads and .5 (in units of the standard error SE). If we observe 60% heads in 100 trials, the test statistic is .6 – .5 in SE units, which yields 2. (Under H0, the SE is .05.) The probability d(X) exceeds 2 is ~ .02. This is the p-value associated with the result for testing H0.The test infers: there is evidence of inconsistency (or discordance) with H0 at statistical significance level .02. The rationale is that 98% of the time, we’d observe a smaller proportion of heads than we did, under the assumption that H0. (Note: .98 is 1 – the p-value.)
Until we check the assumptions of the model, I would say we merely have an indication of evidence. And, even if they check out, as Fisher repeatedly emphasized, isolated significant results do not suffice: “we may say that a phenomenon is experimentally demonstrable when we know how to conduct an experiment which will rarely fail to give us a statistically significant result (Fisher 1947, p. 14).
Testing for soccer. Treating it as such leads to absurd results. “Consider the hypothesis that John plays soccer. Conditional on this hypothesis, it’s unlikely that he plays goalie….But now suppose we observe John playing goalie” where it is stipulated “that goalies are rare soccer players….Yet it would be a mistake to conclude that John doesn’t play soccer” (Titelbaum 463). Yes, but Titelbaum is mistaken to think significance tests license such an inference. [ii] Since he has told us that goalies are rare soccer players (I don’t have a clue about sports), this leads to a logical contradiction. [Premises: (Gj & Sj), (Gj → Sj); conclusion: ~Sj, where G and S are the predicates “is a goalie”, “is a soccer player”, respectively, and j is the name John.] There are no p-values here, and terrible error probabilities. The probability that an inference to “John is not a soccer player” is erroneous is 1. (Also, all of us non-goalies, are erroneously inferred to be soccer players.)[iii]
Titelbaum cites Dickson and Baird (2011, 219-220) as the source of the example. Since I don’t think in terms of sports, note that on their construal of statistical significance tests, observing that a person is a Nobel prize winner would license inferring she was not a person, since Nobel prize winners are rare. This makes no sense. John being a soccer player is not a statistical hypothesis assigning probabilities to outcomes, and even if it were a statistical H, the probability of data x given H is not its p-value. I turn to this.
It’s still a mistake with statistical hypotheses: What if we have a genuine statistical hypothesis, such as the coin tossing case:s on each trial is .5? Every random sample of 100 toss, x, leads a very low probability under H0, (θ = .5) in the Bernouilli model M. Thus, on Titelbaum’s definition of p-value, all outcomes lead to low p-values, and all outcomes would equally reject H0, even if true. But it’s incorrect to view the probability of x given H0 as the p-value.
Anyone who reads this blog knows I don’t consider rudimentary statistical significance tests an adequate account of statistical inference. Here I keep to a rudimentary Fisherian view, because that is what Titelbaum does, and that suffices to block his criticisms [iv]. When Titelbaum says “unlikely events occur, and occur all the time” (463), he thinks he’s objecting to Fisherian tests, but Fisher would agree (remember to construe Titelbaum’s “likely: as “probable”). What’s infrequent is for test statistic d(x) to land in the rejection region computed under H0. (It can be made frequent only by violating assumptions of the model M).
3. Diagnostic screening. Titelbaum’s next criticism (464) is better known than the soccer case. Here we are given frequentist “prior” probabilities for a disease or abnormality.
Suppose a randomly selected member of the population receives a positive test result for a particular disease. This result may receive a very low p-value relative to the null hypothesis that the individual lacks the disease. But if the frequency of the disease in the general population is even lower, it would be a mistake to reject the null.
Do you agree? Is it is a mistake to take a positive test result as evidence for a disease, where a positive result is very rare under no disease? If it is a mistake, then it’s even more of a mistake to take a negative result as evidence for disease. So if we required a high posterior probability of disease given a positive result, the test would have no chance of detecting disease even if present.
Doctor: Your positive result is not evidence of disease, given how rare the disease is.
Patient: Does your screening have any chance of finding evidence of disease even if present?
If a high posterior is required for evidence of disease, the doctor would have to say no.
To have some numbers, assume the probability of a positive result among those with no disease is very low, say .01, while the probability of a positive result among those with the disease is high (the typical criticism sets it at 1). If it is stipulated that the probability of the disease is sufficiently small, say .001, the probability of disease can still be very small, in this case .1. The posterior probability is correct, but the test has not done its job: to discern evidence of the rare disease.
Titelbaum’s diagnostic screening criticism is based on assuming that taking a positive result + as evidence for abnormality or disease requires Pr(disease|+) to be high. I don’t think most Bayesians would agree with this, given that Pr(+|disease) is 100 times Pr(+|no disease). Bayesians generally allow that evidence confirms a claim (to some degree) if its posterior probability has increased from its prior. Here the Bayes factor is 100.
The example is crudely simplistic, but it’s the one we’re given. For frequentists to infer evidence of disease is not to infer the the disease is definitely present: there is a statistical claim, with given error probabilities. With diagnostic screening, the error statistical report is generally in terms of sensitivity and specificity. In this example, it is also stipulated that there are prior frequentist probabilities or prior prevalences. If so, my frequentist doctor reports that while the positive result is evidence of the presence of abnormality, there is a very low frequentist probability of disease among those testing positive, here .1.
In would be unusual for a diagnostic test of disease to infer its presence or absence, rather than to infer evidence of some sign of abnormality warranting a call-back (e.g., in breast-cancer screening). There would also typically be a report of degrees and type of abnormality. My doctor would also add that the proportion with disease among callbacks is low, especially among women with characteristics q, r, s, etc.or further investigation. The question of what action to take is distinct: with luggage, ringing the alarm just leads to rummaging through your bag.
The base rate fallacy. Titelbaum (464) says, “Admittedly the framing of [the medical screening] example brings prior probabilities into the discussion, which the frequentist is keen to avoid. But ignoring such considerations encourages us to commit the Base Rate Fallacy” (464). He assumes, in other words, that since frequentists do not assign probabilities to hypotheses where doing so is illegitimate, that they will “abstain” from assigning them even when proper frequentist probabilities are given!
It is not that the frequentist is keen to avoid prior probabilities of events with given frequentist probabilities. She is keen to avoid assigning probabilities to hypotheses unless they can be construed as random variables with frequentist distributions. Examples such as diagnostic screening with given probabilities are just ordinary conditional probability examples. Many would deny there’s anything Bayesian about it, unless any use of conditional probability is to be called Bayesian.
By and large, by the way, Bayesians share the frequentist perspective that parameters in a model are constants (except for random effects). The difference is that Bayesians are prepared to assign probabilities to constants by allowing probability to be degree of belief assignments.
Frequentists do not abstain from assigning probabilities to claims when legitimate frequentist priors are given, but she will critically evaluate how they are arrived at. Where do we get relevant base rates? Which reference class should we use?
Probabilistic instantiation fallacy. Thus far I have criticized Titelbaum’s diagnostic screening example for assuming evidence of H requires the posterior probability of H to be high. I denies this, even where priors are given. For one thing, it led to never having evidence for a rare disease, even if present. Even hypotheses that are probable in some sense, need not have been well probed by the data x. I now go further. What about the assumed prior probability? In the diagnostic screening example, it is assumed if an individual, say Jill, is randomly selected from a population where 1 out of 1000 have a property (e.g., disease) then the probability that Jill has the property is .001. But this is a fallacy, akin to assuming that a randomly selected .95 confidence interval estimate has a probability of .95 of covering the true parameter value. [Granted, applying uniform priors can yield .95 as the probability the particular estimate is true.]
It’s important to realize that probability, in randomly selecting members of a population, refers to the selecting process. Even if we imagine Jill was randomly selected from a population where the abnormality is very rare, it does not follow that the probability Jill has it is low. Supposing it is, is to commit the fallacy of probabilistic instantiation.
I have discussed this elsewhere [v]. Peter Achinstein (2010, 187), who considers himself an objective Bayesian epistemologist, says it would be a fallacy if he were a frequentist:
My response to the probabilistic fallacy charge is to say that it would be true if the probabilities in question were construed as relative frequencies. However, … I am concerned with epistemic probability.
Achinstein is prepared to assign objective epistemic probabilities this if “the only thing known” is that Jill was randomly selected from a population where 1 in 1000 have the disease. The problems in supposing we get knowledge from ignorance, indifference, or uninformative priors are very. I’m not sure what Titelbaum’s position is on them. His diagnostic screening example was intended to give frequentist probabilities.
I am not saying frequentists can never assign probabilities to Jill having a disease, or other specific events. With disease, it will involve a combination of genetic, environmental, and several other background variables to assess the relevant reference class in which to place Jill. Perhaps the new AI/ML techniques will provide them. When they do, the question of whether and how to employ them in evaluating evidence of hypotheses will still be a separate issue.
Please share corrections, reactions, and questions in the comments. Libnks to the previous 5 guestposts are especially welcome.
Notes:
[i] In an earlier period, as in the 80s, philosophers of science engaged regularly with statistics.
[ii] Using “likely” both as probability of events and when referring to the technical term of a hypothesis being “likely” is a slippery business. For example, many rival hypotheses can all be maximally likely, in the technical sense.
[iii] Here, H0 is: John is a soccer player; the test rule is: observe goalie, infer not soccer; observe not goalie, infer soccer. If you have another construal, let me know.
[iv] Although I prefer an account closer to Neyman and Pearson (N-P), with an explicit alternative, there are important roles for Fisherian tests where the choice is left to specifying the test statistic d(X); notably, testing assumptions of the model. Fisher gives a demanding set of requirements for an adequate test statistic (see SIST). For a broad class of tests, as David Cox shows, Fisher’s tests lead to the same place as N-P.
[v] E.g., Mayo 1997, 2005, 2010, 2018. See also Spanos 2010. In the Achinstein discussion, a student is randomly selected from a population where college-readiness is rare. High test scores still yield a low posterior probability for being classified as ready.
References
.
John Park, MD
Medical Director of Radiation Oncology
North Kansas City Hospital
Clinical Assistant Professor
Univ. Of Missouri-Kansas City
[An earlier post by J. Park on this topic: Jan 17, 2022: John Park: Poisoned Priors: Will You Drink from This Well? (Guest Post)]
Abandoning P-values and Embracing Artificial Intelligence in Medicine
The move to abandon P-values that started 5 years ago was, as we say in medicine, merely a symptom of a deeper more sinister diagnosis. Within medicine, the diagnosis was a lack of statistical and philosophical knowledge. Specifically, this presented as an uncritical move towards Bayesianism away from frequentist methods, that went essentially unchallenged. The debate between frequentists and Bayesians, though longstanding, was little known inside oncology. Out of concern, I sought a collaboration with Prof. Mayo, which culminated into a lecture given at the 2021 American Society of Radiation Oncology meeting. The lecture included not only representatives from frequentist and Bayesian statistics, but another interesting guest that was flying under the radar in my field at that time… artificial intelligence (AI).
Fast forward 3 years from that meeting: AI and Machine Learning (ML) have taken medicine by storm with the rise of large language models being adapted to specific diseases, automated reading of images, and even contouring of cancers in radiation oncology. Given there was no firm statistical foundation, it is not surprising medicine has given way to AI/ML data analysis essentially without challenge once again.
The current state is like the wild, wild, west with academic medicine publishing AI/ML papers at an incredible rate in all the top oncology journals. The New England Journal of Medicine (NEJM) even created its own spin off journal called NEJM AI. Looking at these papers there is no uniform consensus on how to validate big data results. Severity is severely (pun intended) lacking. This is further compounded by the fact that traditional statistics, no matter what the flavor, were looking for causal inferences, however for AI methodologists, some say it is not their concern and many ML algorithms do not try to quantify error rates (Watson, Synthese, 2023). The usual concerns about normality, independence, and identical distributions do not apply because frequently there are no strong assumptions about the data. Along with the fact that many algorithms are black boxes with proprietary rights and even if we could see the algorithm, we could not understand it.
Further compounding the issue there are no universal algorithms being used, whereas, in most medical research regression models, Kaplan-Meier analysis, etc. are standardized. A snapshot of 3 articles from top oncology journals display the disparate methodologies used.
| Study | Algorithm | Training Methodologies | | Bladder Cancer Lymph Node Detection (Wu, Lancet Oncol, 2023) | High-Resolution Net (HRNet) neural network using a novel architecture that ran multi-resolution streams in parallel | Dynamic balanced sampling scheme with hard negative mining approach, F2 Score, Brier Score, ROC Curve | | Multiple Myeloma Prognostication (Maura, JCO, 2024) | 3 algorithm approach using multivariate cox-proportional-hazard with regularization, random survival forest, and neural networks | 5-fold cross validation (repeated x 10), Harrell’s and Uno’s Concordances, Integrated Brier Score, Negative Binomial Log-likelihood | | Lung Cancer Toxicity Prediction (Ladbury, IJROBP, 2023) | 6 algorithm Interpretable ML approach using logistic regression, naive Bayes, k-nearest neighbors, support vector machine, random forest, and extreme gradient boosting | Synthetic Minority Oversampling Technique (SMOTE) TomekLinks method, the SMOTE-Edited Nearest Neighbors method, tabular generative adversarial networks, 10-fold cross validation to maximize AUC, Shapley additive explanation (SHAP) framework |
The question now is are we able to severely test AI/ML studies? First, we can subject the AI algorithms to prospective clinical trials and subject them to rigorous testing in this manner. The table below is list of trials from a paper my colleagues and I wrote discussing the integration of AI/ML into oncology trials (Kang, Sem Rad Onc, 2023).
Although, in my humble opinion, this is the best way test the AI/ML methods, it will not be feasible to test the multitude of algorithms that way. A good start to this problem is the formation of AI/ML guidelines such as TRIPOD+AI that provide authors a checklist to “to promote the complete, accurate, and transparent reporting of studies that develop a prediction model or evaluate its performance” (Collins, BMJ, 2024). Given the multiple different types of data analysis, there are also different guidelines for when AI reads medical imaging vs analyzing data points per the table below (Collins, BMJ, 2024).
An area that needs much more growth is on what constitutes severe testing for AI/ML. The stakes are much higher for cancer patients than say for shopping preferences when looking at A/B testing. Trying to make the models more transparent with explainable AI (XAI), runs into the problems of ambiguous fidelity, lack of severe testing, and an emphasis on product over process (Watson, Synthese, 2023). The other question is when are AI/ML methods actually needed? In a study comparing statistical methods (Cox model and accelerated failure time model with log-logistic distribution) with ML methods (random survival forest and PC-Hazard deep neural network), the ML methods needed 2-3 times the sample size to achieve the same performance (Infante, Stats in Medicine, 2023), asking the question how many studies actually need the power of AI/ML, as opposed to being the hot topic that is publishable? I do not have any answers, save for prospectively testing these methods in clinical trials, but before the answers come these questions must be raised. The aforementioned guidelines are a start, but it is hard to see, or at least it is not explicit, how effective they are at safeguarding against bias, irreproducibility, or distinguishing signal from noise. These were the same warnings given out five years ago when P-values and statistical significance were called to be abandoned. My hope is that they will be heeded.
References
Infante, G., Miceli, R., Ambrogi, F. Sample size and predictive performance of machine learning methods with survival data: A simulation study. Statistics in Medicine. Volume 42, Issue 30, p5657-5675 (2023).
Kang, J., Chowdhry, A.K., Pugh, S.L., Park, J.H. Integrating Artificial Intelligence and Machine Learning Into Cancer Clinical Trials. Seminars in Radiation Oncology. Volume 33, Issue 4, p386-394 (2023).
Ladbury, C., Li, R., Danesharasteh, A., et al. Explainable Artificial Intelligence to Identify Dosimetric Predictors of Toxicity in Patients with Locally Advanced Non-Small Cell Lung Cancer: A Secondary Analysis of RTOG 0617. International Journal of Radiation OncologyBiologyPhysics. Volume 117, Issue 5, p1287-1296 (2023).
Maura, F., Rajanna, A.R., Ziccheduu, B., et al. Genomic Classification and Individualized Prognosis in Multiple Myeloma. Journal of Clinical Oncology. Volume 42, Number 11 (2024).
Watson, D.S. Conceptual challenges for interpretable machine learning. Synthese. Volume 200, Article 65, (2022).
Wu. S., Guibin, H., Xu, A. et al. Artificial intelligence-based model for lymph node metastases detection on whole slide images in bladder cancer: a retrospective, multicentre, diagnostic study. Lancet Oncology Volume 24, Issue 4, p360-37, (2023).
.
Ron S. Kenett
Chairman of the KPA Group;
Senior Research Fellow, the Samuel Neaman Institute, Technion, Haifa;
Chairman, Data Science Society, Israel
What’s happening in statistical practice since the “abandon statistical significance” call
This is a retrospective view from experience gained by applying statistics to a wide range of problems, with an emphasis on the past few years. The post is kept at a general level in order to provide a bird’s eye view of the points being made.
An important impact on the current practice of statistics is the merging of empirical predictive analytics with probability-based modelling. In many applications, but not in all, one has access to massive data coming from sensor technologies and unstructured formats such as text and images. In these cases, methods to fit and assess a model are based on splitting the data into a training set and a validation set. The training data is used to fit a model, the validation set, to evaluate it. If the fit of a model to the training data is very high but the fit to the validation set is low, we experience what is labeled “overfitting”. Overfitting reduces the ability to generalize the model to future implementations and implies poor predictive performance. This approach is different from the classical statistical analysis relying on probability models and hypothesis testing.
R. A. Fisher, in his fundamental paper “On the Mathematical Foundations of Theoretical Statistics”, stated that “the object of statistical method is the reduction of data.” He then identified “three problems which arise in the reduction of data.”: Problem 1: Specification—choosing the right mathematical model for a population. Problem 2: Estimation—methods to calculate estimates, from a sample, and Problem 3: Distribution—properties of estimators derived from samples. The predictive analytic methods outlined in the first paragraph change all this. The Specification, Estimation, Distribution trio is replaced, with computer intensive methods including cross validation, bootstrapping and simulations. The models used in this context are supervised or unsupervised, with transfer learning, active learning labelling, zero shot and few shot learning.
In this context of new problems and new methods we still face the fundamental issue of collecting data with relevance to the problem at hand, what Colin Mallows called the “zeroth problem.” For example, splitting of the data into training and validation sets must be consistent with the data generation process. Some principles for achieving this are formulated in an approach titled “befitting cross validation” (BCV), see Kenett et al (2022, https://xwdeng80.github.io/BCV2022.pdf ).
The performance of a model applied to a validation set is determining goodness of fit. In classification problems one computes misclassification errors, lift, ROC and confusion matrices.
The combination of computer intensive methods and classical statistical methods (Bayesian or frequentists) is a challenge for current work in data analysis. In many companies, Statistics groups have been replaced by Data Science groups. Some universities do not even involve statisticians in their data science programs.
With this perspective, let me share my views on what’s happening in statistical practice since the “abandon significance” call five years ago.
The few years considered here start at the Bethesda ASA symposium on statistical inference (SSI) on October 2017. The event followed the 2016 ASA “p-value statement” and fed a large special issue of the American Statistician. I summarized what happened there in a blog titled “to p or not to p”. See: https://blogisbis.wordpress.com/2017/10/24/to-p-or-not-to-p-my-thoughts-on-the-asa-symposium-on-statistical-inference/
Following SSI, Mayo’s error statistics philosophy blog provided a platform to discuss controversies around hypothesis testing and p-values as presented in Bethesda and beyond. It also provided some clarity into what was official ASA policy and what was not. The blog gathered perspectives under an umbrella labelled: “The statistics wars”. At some point I mentioned, in a comment, that these debates seem to be localised to specific groups and that most users of statistics ignored them. I also noted that the statistics wars appeared to have no impact on statistical analysis software platforms or on statistics curriculum in academia and elsewhere.
Today, I believe that these comments still hold and that the cargo cult application of statistics, presented in Stark and Saltelli, is prevalent. Moreover, the statistics war discussions in the literature, that reached a peak five years ago, seem waning. As Shakespeare wrote: “much ado about nothing.”
In contrast, there are many advances in data analysis. Some examples include, assessing selection bias (Benjamini, 2019), evaluating fairness in analytic models (Plecko and Bareinboim, 2023) and considering information quality (https://sites.google.com/site/datainfoq).
Moreover, it seems that the p-value discussions, started by the ASA over five years ago, did not strengthen the general position of statistics or statisticians. These discussions coincided with a particularly delicate period where other disciplines got deeply involved in data analysis and modelling. The result is that statistics is at a crossroad. Some thoughts on how to address current challenges in meeting this crossroad were presented in this seminar. See also here. A summary is listed below:
In summary, the positioning of analytics is at a peak. Statisticians should leverage this opportunity by clarifying the unique selling points of statistics. This was the context of the conference “On the foundations of applied statistics” held at the Samuel Neaman Institute, Technion, Israel in April 2024. Presenters addressed various aspects of applied statistics including philosophical methods, historical examples and designing experiments for generalizability of findings. For slides and a recording see https://neaman.org.il/en/On-the-foundations-of-applied-statistics. In the conference, Daniel Lakens proposed to have a Vatican like event where a multiperspective group sits down to map applied statistics foundations, until white smoke is observed. Why not…
.
Professor Daniël Lakens
Human Technology Interaction
Eindhoven University of Technology
[Some earlier posts by D. Lakens on this topic are at the end of this post]*
This continues Part 1:4: Most do not offer any alternative at allAt this point, it might be worthwhile to point out that most of the contributions to the special issue do not discuss alternative approaches to p < .05 at all. They discuss general problems with low quality research (Kmetz, 2019), the importance of improving quality control (D. W. Hubbard & Carriquiry, 2019), results blind reviewing (Locascio, 2019), or the role of subjective judgment (Brownstein et al., 2019). There are historical perspectives on how we got to this point (Kennedy-Shaffer, 2019), ideas about how science should work instead, many stressing the importance of replication studies (R. Hubbard et al., 2019; Tong, 2019). Note that Trafimow both recommends replication as an alternative (Trafimow, 2019), but also co-authors a paper stating we should not expect findings to replicate (Amrhein et al., 2019), thereby directly contradicting himself within the same special issue. Others propose not simply giving up on p-values, but on generalizable knowledge (Amrhein et al., 2019). The suggestion is to only report descriptive statistics.
5: Why has Nothing Changed?It is worth reflecting on why we have had these discussions about the use of p-values approximately every decade for the last century, while nothing changes. I will make 3 personal and largely unsubstantiated observations that you are free to disagree with.
As far as I noticed, not a single contribution to the special issue engaged with philosophy of science. This is almost unbelievable, as the choice for a statistical method is based on the aims of science (Laudan, 1986), and what the aims of science are is per definition a philosophical question. One would have expected every contribution to start with three paragraphs about philosophy of science, and only then, after having explained what their thoughts about the aims of science are, continue with their proposed recommendations. It is worthwhile to acknowledge that the practice of p < .05 is perfectly defensible from a methodological falsificationist approach to knowledge generation (Uygun Tunç et al., 2023). It is fine to criticize methodological falsificationism, but one can hardly blame methodological falsificationists for using the method that is coherent with their aims, and the claims they want to make. One is also free to align one’s recommendations with a different philosophy of science, but then it has to be coherent and well-developed.
Second, there were very few real-life examples of how one should analyze data in practice. In our own contribution to the special issue (Dongen et al., 2019) we analyzed the same dataset in four different ways, and all approaches lead to the same conclusion. Of course different approaches might yield different results in some applied contexts, and in some contexts it makes no sense to ask certain statistical questions (including those answered by p < .05). But there might be little use in recommendations that are not related to an applied context. Some solutions work in some contexts, but not in others. It requires a careful understanding of the research context to know which recommendations might work in some research areas. Without this understanding, statisticians might make proposals no one needs. It reminds me of a statistician who presented their work, and at the end of their talk asked: “so, does anyone have a dataset where they would need the method I have developed?”
Third, there was by and large a lack of engagement with other perspectives. Researchers propose their own views, but there is no honest discussion about relative strengths and weaknesses of what they propose. The statistics community has a reward structure that is similar to most scientific fields, where provocative and attention grabbing titles lead to a lot of citations. Actually solving problems is not strongly rewarded (with some exceptions, of course). Reaching consensus, or examining if proposals actually work in practice, is a lot of work and does not get any statistician tenure. The field is not an exception, but it helps to explain why discussions about p-values are never ending, I think.
References(References below are for both parts )
Amrhein, V., Trafimow, D., & Greenland, S. (2019). Inferential Statistics as Descriptive Statistics: There Is No Replication Crisis if We Don’t Expect Replication. The American Statistician, 73(sup1), 262–270. https://doi.org/10.1080/00031305.2018.1543137
Anderson, A. A. (2019). Assessing Statistical Results: Magnitude, Precision, and Model Uncertainty. The American Statistician, 73(sup1), 118–121. https://doi.org/10.1080/00031305.2018.1537889
Benjamin, D. J., & Berger, J. O. (2019). Three Recommendations for Improving the Use of p-Values. The American Statistician, 73(sup1), 186–191. https://doi.org/10.1080/00031305.2018.1543135
Betensky, R. A. (2019). The p-Value Requires Context, Not a Threshold. The American Statistician, 73(sup1), 115–117. https://doi.org/10.1080/00031305.2018.1529624
Blume, J. D., Greevy, R. A., Welty, V. F., Smith, J. R., & Dupont, W. D. (2019). An Introduction to Second-Generation p-Values. The American Statistician, 73(sup1), 157–167. https://doi.org/10.1080/00031305.2018.1537893
Brownstein, N. C., Louis, T. A., O’Hagan, A., & Pendergast, J. (2019). The Role of Expert Judgment in Statistical Inference and Evidence-Based Decision-Making. The American Statistician, 73(sup1), 56–68. https://doi.org/10.1080/00031305.2018.1529623
Calin-Jageman, R. J., & Cumming, G. (2019). The New Statistics for Better Science: Ask How Much, How Uncertain, and What Else Is Known. The American Statistician, 73(sup1), 271–280. https://doi.org/10.1080/00031305.2018.1518266
Cohen, J. (1994). The earth is round (p < .05). American Psychologist, 49(12), 997–1003. https://doi.org/10.1037/0003-066X.49.12.997
Cohen, J. (1995). The earth is round ( p < .05): Rejoinder. American Psychologist, 50(12), 1103. http://dx.doi.org/10.1037/0003-066X.50.12.1103
Colquhoun, D. (2019). The False Positive Risk: A Proposal Concerning What to Do About p-Values. The American Statistician, 73(sup1), 192–201. https://doi.org/10.1080/00031305.2018.1529622
Dongen, N. N. N. van, Doorn, J. B. van, Gronau, Q. F., Ravenzwaaij, D. van, Hoekstra, R., Haucke, M. N., Lakens, D., Hennig, C., Morey, R. D., Homer, S., Gelman, A., Sprenger, J., & Wagenmakers, E.-J. (2019). Multiple Perspectives on Inference for Two Simple Statistical Scenarios. The American Statistician, 73(sup1), 328–339. https://doi.org/10.1080/00031305.2019.1565553
Fricker, R. D., Burke, K., Han, X., & Woodall, W. H. (2019). Assessing the Statistical Analyses Used in Basic and Applied Social Psychology After Their p-Value Ban. The American Statistician, 73(sup1), 374–384. https://doi.org/10.1080/00031305.2018.1537892
Gannon, M. A., de Bragança Pereira, C. A., & Polpo, A. (2019). Blending Bayesian and Classical Tools to Define Optimal Sample-Size-Dependent Significance Levels. The American Statistician, 73(sup1), 213–222. https://doi.org/10.1080/00031305.2018.1518268
Goodman, W. M., Spruill, S. E., & Komaroff, E. (2019). A Proposed Hybrid Effect Size Plus p-Value Criterion: Empirical Evidence Supporting its Use. The American Statistician, 73(sup1), 168–185. https://doi.org/10.1080/00031305.2018.1564697
Greenland, S. (2019). Valid P-Values Behave Exactly as They Should: Some Misleading Criticisms of P-Values and Their Resolution With S-Values. The American Statistician, 73(sup1), 106–114. https://doi.org/10.1080/00031305.2018.1529625
Hodges, J. L., & Lehmann, E. L. (1954). Testing the Approximate Validity of Statistical Hypotheses. Journal of the Royal Statistical Society. Series B (Methodological), 16(2), 261–268. https://doi.org/10.1111/j.2517-6161.1954.tb00169.x
Hubbard, D. W., & Carriquiry, A. L. (2019). Quality Control for Scientific Research: Addressing Reproducibility, Responsiveness, and Relevance. The American Statistician, 73(sup1), 46–55. https://doi.org/10.1080/00031305.2018.1543138
Hubbard, R., Haig, B. D., & Parsa, R. A. (2019). The Limited Role of Formal Statistical Inference in Scientific Inference. The American Statistician, 73(sup1), 91–98. https://doi.org/10.1080/00031305.2018.1464947
Kennedy-Shaffer, L. (2019). Before p < 0.05 to Beyond p < 0.05: Using History to Contextualize p-Values and Significance Testing. The American Statistician, 73(sup1), 82–90. https://doi.org/10.1080/00031305.2018.1537891
Kmetz, J. L. (2019). Correcting Corrupt Research: Recommendations for the Profession to Stop Misuse of p-Values. The American Statistician, 73(sup1), 36–45. https://doi.org/10.1080/00031305.2018.1518271
Krueger, J. I., & Heck, P. R. (2019). Putting the P-Value in its Place. The American Statistician, 73(sup1), 122–128. https://doi.org/10.1080/00031305.2018.1470033
Lakens, D., & Delacre, M. (2020). Equivalence Testing and the Second Generation P-Value. Meta-Psychology, 4, 1–11. https://doi.org/10.15626/MP.2018.933
Lakens, D., Adolfi, F.G., Albers, C.J. et al. Justify your alpha. Nat Hum Behav 2, 168–171 (2018). https://doi.org/10.1038/s41562-018-0311-x
Laudan, L. (1986). Science and Values: The Aims of Science and Their Role in Scientific Debate.
Locascio, J. J. (2019). The Impact of Results Blind Science Publishing on Statistical Consultation and Collaboration. The American Statistician, 73(sup1), 346–351. https://doi.org/10.1080/00031305.2018.1505658
Maier, M., & Lakens, D. (2022). Justify Your Alpha: A Primer on Two Practical Approaches. Advances in Methods and Practices in Psychological Science, 5(2), 25152459221080396. https://doi.org/10.1177/25152459221080396
Manski, C. F. (2019). Treatment Choice With Trial Data: Statistical Decision Theory Should Supplant Hypothesis Testing. The American Statistician, 73(sup1), 296–304. https://doi.org/10.1080/00031305.2018.1513377
Matthews, R. A. J. (2019). Moving Towards the Post p < 0.05 Era via the Analysis of Credibility. The American Statistician, 73(sup1), 202–212. https://doi.org/10.1080/00031305.2018.1543136
Mazzolari, R., Porcelli, S., Bishop, D. J., & Lakens, D. (2022). Myths and methodologies: The use of equivalence and non-inferiority tests for interventional studies in exercise physiology and sport science. Experimental Physiology, 107(3), 201–212. https://doi.org/10.1113/EP090171
McShane, B. B., Gal, D., Gelman, A., Robert, C., & Tackett, J. L. (2019). Abandon Statistical Significance. The American Statistician, 73(sup1), 235–245. https://doi.org/10.1080/00031305.2018.1527253
Murphy, K. R., & Myors, B. (1999). Testing the hypothesis that treatments have negligible effects: Minimum-effect tests in the general linear model. Journal of Applied Psychology, 84(2), 234–248. https://doi.org/10.1037/0021-9010.84.2.234
Pogrow, S. (2019). How Effect Size (Practical Significance) Misleads Clinical Practice: The Case for Switching to Practical Benefit to Assess Applied Research Findings. The American Statistician, 73(sup1), 223–234. https://doi.org/10.1080/00031305.2018.1549101
Tong, C. (2019). Statistical Inference Enables Bad Science; Statistical Thinking Enables Good Science. The American Statistician, 73(sup1), 246–261. https://doi.org/10.1080/00031305.2018.1518264
Trafimow, D. (2019). Five Nonobvious Changes in Editorial Practice for Editors and Reviewers to Consider When Evaluating Submissions in a Post p < 0.05 Universe. The American Statistician, 73(sup1), 340–345. https://doi.org/10.1080/00031305.2018.1537888
Uygun Tunç, D., Tunç, M. N., & Lakens, D. (2023). The epistemic and pragmatic function of dichotomous claims based on statistical hypothesis tests. Theory & Psychology, 09593543231160112. https://doi.org/10.1177/09593543231160112
Wasserstein, R. L., Schirm, A. L., & Lazar, N. A. (2019). Moving to a World Beyond “p < 0.05.” The American Statistician, 73(sup1), 1–19. https://doi.org/10.1080/00031305.2019.1583913
Wong, T. K., Kiers, H., & Tendeiro, J. (2022). On the Potential Mismatch Between the Function of the Bayes Factor and Researchers’ Expectations. Collabra: Psychology, 8(1). https://doi.org/10.1525/collabra.36357
*Other related guestposts by Lakens on this blog:
June 8, 2016 “So you banned p-values, how’s that working out for you?” D. Lakens exposes the consequences of a puzzling “ban” on statistical inference
Jan 5, 2022: Lakens (Guest Post): Averting journal editors from making fools of themselves.
June 30, 2022: D. Lakens responds to confidence interval crusading journal editors]
.
Professor Daniël Lakens
Human Technology Interaction
Eindhoven University of Technology
*[Some earlier posts by D. Lakens on this topic are listed at the end of part 2, forthcoming this week]
How were we supposed to move beyond p < .05, and why didn’t we?
It has been 5 years since the special issue “Moving to a world beyond p < .05” came out (Wasserstein et al., 2019). I might be the only person in the world who has read all 43 contributions to this special issue. [In part 1] I will provide a summary of what the articles proposed we should do instead of p < .05, and [in part 2] offer some reflections on why they did not lead to any noticeable change.
1: Test Range predictionsPerhaps surprisingly the most common recommendation is to continue to make dichotomous claims based on whether a statistic falls below a critical score – exactly as we now do with p-values. However, instead of computing test scores for null hypothesis significance tests, the recommendation is to compute a test statistic for interval hypothesis tests. This recommendation is 70 years old (Hodges & Lehmann, 1954), and it solves most of the criticisms people have on the current use of p-values. However, it is challenging to specify a range of values to test against, and this is why adoption of range predictions is incredibly slow. It requires a large time investment to figure out which effect sizes we would meaningfully predict, and most research areas have not invested that time yet.
Anderson discusses the importance of testing against effects that are considered practically significant, and to take the precision into account (Anderson, 2019). He interprets confidence intervals as tests, and therefore his contribution boils down to a reminder that range predictions are an improvement to NHST. If you can specify a smallest effect of interest, and have an idea of how precise you want the estimate to be, Anderson proposes you should design studies that return confidence intervals that yield informative results with respect to an effect size you have determined would be practically significant.
Betensky also proposes testing range predictions (Betensky, 2019), and specifically minimum effect tests. She writes “The p-value and sample size jointly yield 95% confidence bounds for the effect of interest, which can be compared to the predetermined meaningful effect size to make inferences about the true effect.” She proposes: “The principle: Reject the null in favor of a meaningful effect if and only if the lower 95% confidence bound exceeds the smallest effect size considered meaningful. Thus, rejecting the null means we can be 95% confident that the true effect size is at least as large as the size considered to be clinically meaningful.” This is a test, that yields a dichotomous claim, based on a p-value. We are not moving beyond p < .05 in this proposal, but we are testing against meaningful effects, instead of a null hypothesis of 0.
Pogrow suggests to focus on practical benefit, which he proposes is evaluated by comparing the unadjusted actual performance of an experimental group to an existing benchmark (Pogrow, 2019). The recommendation is not very good – in essence, the author suggests to ignore all statistical inferences. He acknowledges that this leaves open the question when a difference is ‘large enough’. He writes: “At the very least “oomph” is a clearly noticeable improvement or other benefit that does not require precise statistical criteria to discern.” This will not do (my general writing advice is that if you cannot define what is within quotation marks, you typically need to think more about what you are proposing). The correct way to achieve this proposal would be a test (even if the Type 1 error rate is set to a higher level than 5% based on a cost-benefit analysis (Maier & Lakens, 2022)). Nevertheless, in essence the paper recommends comparing effects against a meaningful effect size.
Goodman, Spruill and Komaroff (2019) also propose a test of a range prediction, and call it a minimum effect size plus p-value (MESP) approach. They write “Here, rejecting the null also requires the difference of the observed statistic from the exact null to be meaningfully large or practically significant, in the researcher’s judgment and experience.” Again, this still is a test with a dichotomous conclusion based on a p-value – and in essence simply a minimal effect test (Mazzolari et al., 2022; Murphy & Myors, 1999).
Blume and colleagues (2019) state that p-values might be criticized in the literature, but “having a gross indicator for when a set of data are sufficient to separate signal from noise is not a bad idea” (p. 157). They propose ‘Second Generation P-Values’, which is an interval null hypothesis test. As we have pointed out elsewhere, their idea is conceptually extremely similar to traditional equivalence tests (Lakens & Delacre, 2020), and equivalence tests are a superior solution. Regardless, this proposal again leads to dichotomous decisions based on whether a confidence interval falls within a range prediction, or not.
Greenland (2019) also suggests “computing P-values for contextually important alternatives to null (nil or “no effect”) hypotheses, such as minimal important differences”, in a piece in which he criticizes unwarranted criticism of p-values. He does respond against the dichotomization of p-values, but mainly in observational studies: “It may be argued that exceptions to (4) [the dichotomization of p-values] arise, for example when the central purpose of an analysis is to reach a decision based solely on comparing p to an α, especially when that decision rule has been given explicit justification including both false-acceptance (Type-II) error costs under relevant alternatives and false-rejection (Type-I) error costs (Lakens et al. Citation2018), as in quality-control applications. Such thoroughly justified applications are, however, uncommon in observational research settings.” This is perhaps similar to Cohen (1994) who criticized the use of NHST, only to accept that it might play a role in well-controlled experiments (Cohen, 1995).
Calin-Jageman and Cumming stress an estimation approach (Calin-Jageman & Cumming, 2019), but when a researcher wants to make a decision (such as whether or not a prediction is supported or not) they write “Another frequent concern is that scientists need to make clear Yes/No decisions (e.g. Does this drug work? Is this project worth funding?). No problem! Focusing on effect sizes and uncertainty does not preclude making decisions—in fact, it makes it easier because one can easily test a result against any desired standard of evidence. For example, suppose you know that a drug improves outcomes by 10% with a 95% confidence interval from 2% up to 18%. If the standard of evidence required is at least a 1% increase in outcomes, the drug would be considered suitable (because a 1% increase is not within the range of the confidence interval).” In other words, the authors propose testing range predictions, such as minimal effect tests, based on the confidence interval.
2: Complement p with something elseSo far, we have seen that 7 of the articles either directly propose the use of p-values for interval hypothesis tests, or make a proposal that is strongly in line with it. Several other papers support the continued use of dichotomous claims, but recommend adding other statistics. This is commonly already required in reporting guidelines, which typically require researchers to add effect sizes and confidence intervals. But the special issue saw some additional proposals.
For example Colquhoun foresees the continued use of p-values, and suggests complementing it with additional statistics. Colquhoun (2019) writes: “It is suggested that p-values and confidence intervals should continue to be given, but that they should be supplemented by a single additional number that conveys the strength of the evidence better than the p-value. This number could be the minimum FPR (that calculated on the assumption of a prior probability of 0.5, the largest value that can be assumed in the absence of hard prior data).”
Matthews (2019) suggests the use of Bayesian credible intervals instead of p-values (called AnCred), but the proposed alternative is still a testing procedure that leads to dichotomous conclusions, just based on a different critical value, which is now determined with a Bayesian flavor. As Matthews writes: “Perhaps the most obvious potential criticism of AnCred is that the concept of credibility simply replaces statistical significance as a means of dichotomizing findings. However, it should be stressed that dichotomization per se has never been the problem with NHST; it is the actions that flow from it.”
Benjamin and Berger (Benjamin & Berger, 2019) state that “In statistical practice, perhaps the single biggest problem with p-values is that they are often misinterpreted in ways that lead to overstating the evidence against the null hypothesis.” They propose to complement p-values with a Bayes factor bound to communicate how strong the evidence in the data is. Their long term plan is to make researchers more comfortable with Bayesian approaches. They hint at being ok with completely replacing p-values in the future, but when pressed, I doubt they would support giving up error control when making claims. As another contribution in the special issue shows, removing statistical error control will lead to more researchers claiming their predictions are supported than is warranted based on commonly used error rates (Fricker et al., 2019). Given the widespread misuse of Bayes factors (Wong et al., 2022), anyone proposing an alternative to p-values will also need to engage with the problems that will emerge if this recommendation would be adopted at the same scale as p-values are currently adopted.
Krueger and Heck (2019) write “As experimentalists, we are reluctant to relinquish dichotomous decision-making entirely” (p. 125) and provide a defense of dichotomous claims in science. The authors do not really provide a suggestion for an alternative, beyond a generic statement that “We join those who recommend researchers use a toolbox of statistical techniques, employ good judgment, and keep an eye on developments in statistical and data science.”
3: Use Statistical Decision TheoryStatistical decision theory is an important upgrade to Neyman-Pearson hypothesis testing, and yet, Neyman-Pearson hypothesis tests remain the dominant approach to p-values in science. The reason statistical decision theory is not adopted is because we don’t know how. In statistical decision theory we still make dichotomous claims based on p-values, but the alpha level is no longer set to 5%, and the Type 2 error is no longer a minimum of 20%, but error rates are specified based on a cost benefit analysis (Maier & Lakens, 2022). Researchers need to be able to quantify the costs and benefits of doing their study – and they can’t. This is already difficult enough in applied research, but only becomes more difficult in theoretical research. Several authors in the special issue propose to use statistical decision theory, but no one explains how we can achieve it.
Manski (2019) provides a nice overview of statistical decision theory, and suggests to use it instead of NHST. I would love to see this used more, but personally, I would not know how to implement it in practice in my research.
Gannon et al. (2019) join the choir of people who say we need to keep using p-values, but stress the need to choose alpha levels in a smarter way. They write “This article argues that researchers do not need to completely abandon the p-value, the best-known significance index, but should instead stop using significance levels that do not depend on sample sizes.” This is a solid recommendation, but it is a recommendation to justify alpha levels (Lakens et al., 2018), not a suggestion to move beyond p-values. It uses the simplest form of statistical decision theory, where we try to minimize errors, but consider all errors equally costly.
Despite the misleading title of ‘Abandon Statistical Significance’, McShane and colleagues (McShane et al., 2019) also propose the use of statistical decision theory to make dichotomous claims. They write: “While we see the intuitive appeal of using p-value or other statistical thresholds as a screening device to decide what avenues (e.g., ideas, drugs, or genes) to pursue further, this approach fundamentally does not make efficient use of data: there is in general no connection between a p-value—a probability based on a particular null model—and either the potential gains from pursuing a potential research lead or the predictive probability that the lead in question will ultimately be successful. Instead, to the extent that decisions do need to be made about which lines of research to pursue further, we recommend making such decisions using a model of the distribution of effect sizes and variation, thus working directly with hypotheses of interest rather than reasoning indirectly from a null model.” This proposal would still be based on p-values (or a comparable critical value) and dichotomous decisions.
To be continued in Part 2 (to be posted later this week)…
.
Professor Christian Hennig
Department of Statistical Sciences “Paolo Fortunati”
University of Bologna
[An earlier post by C. Hennig on this topic: Jan 9, 2022: The ASA controversy on P-values as an illustration of the difficulty of statistics]
Statistical tests in five random research papers of 2024, and related thoughts on the “don’t say significant” initiative
This text follows an invitation to write on “abandon statistical significance 5 years on”, so I decided to do a tiny bit of empirical research. I had a look at five new papers listed on May 17 on the “Research Articles” site of Scientific Reports. I chose the most recent five papers when I looked without being selective. As I “sampled” papers for a general impression, I don’t want this to be a criticism of particular papers or authors, however in the interest of transparency, the doi addresses of the papers are:
https://doi.org/10.1038/s41598-024-62172-2
https://doi.org/10.1038/s41598-024-61552-y
https://doi.org/10.1038/s41598-024-62074-3
https://doi.org/10.1038/s41598-024-59702-3
https://doi.org/10.1038/s41598-024-62350-2
Four of the papers contain statistical tests. None of the papers contains any of the methods that were proposed in the statistical literature as alternatives to tests and p-values such as s-values (Cole et al. 2020), second generation p-values (Blume et al. 2019), e-values (Grünwald et al. 2023), relevance (Stahel 2021), or any Bayesian analysis. Severity (Mayo 2018) does not feature either. There is no trace of any influence of the “abandon statistical significance” discussion. The papers have a substantial amount of material and key conclusions that don’t rely on statistical tests, on which I will not comment here.
Obviously this is not a representative sample of what goes on in science (for sure looking at only one journal provides a very narrow view); I was just curious to see an interdisciplinary mix of recent papers in a reasonably well reputed journal. Despite this, it suggests (together with some even less formal further looking around) that significance testing has for sure not been abandoned, but still happens all over the place.
How do I feel about this? As probably almost all statisticians do, I think that statistical tests are endemically misinterpreted and misused. Although I don’t agree with demands for tests to be abandoned (more on this later), I hope that the controversy about tests causes at least some researchers to question what they are doing, and to understand a bit better when to use them or not, and how to interpret the results. This may indeed happen in some quarters, but I suspect that it concerns a small minority.
Looking at the four papers containing tests, there are some issues. One of them only reports whether p<0.05 or not, but not the precise p-value. This bugs me; it is of great informative value whether p is close to 0.05 (not a very strong indication of anything) or rather 0.0001; binary thinking based on just distinguishing “significant” from “insignificant” was one of the major issues with tests highlighted by Wasserstein et al. (2019), and I agree that (reasonably) precise p-values should be shown.
Another paper states that the null hypothesis is true in case of non-rejection (“there is no difference”; “data are normal” – no they aren’t! -when tested and not rejected). I find this problematic, but given that the general tone of the paper conveys that results refer to the specific experiment and authors avoid overgeneralising claims and don’t seem to imply that their results are the final word on the matter, this seems rather harmless in the specific case.
The same paper wrongly replaces a paired t-test by an unpaired Wilcoxon because of “non-normality”, ignoring dependence between observations on the same individual observed in two conditions. A data plot indicates, however, that the highly significant test result would also have been significant with a correct paired test.
One paper has a rather nonsensical description of a post hoc power analysis clearly revealing a lack of insight. Another one seems to imply that running a Kolmogorov-Smirnov test amounts to just plotting two distribution functions together without computing a p-value. The one paper that doesn’t run a formal statistical test claims anyway that a certain result is “significant”. This could have been formally tested, though the test would have been non-standard and not so simple. A statistical plot illustrates the result. I believe that a formal test would have confirmed significance due to a large sample size, and it would have been informative to actually run the test.
All four papers with tests run several tests. Two of them don’t account for or even mention multiple testing. This is particularly problematic where it is only indicated whether “p<0.05”, because multiple testing corrections would require p-values much smaller than 0.05, although the “p<0.05 paper” actually uses a multiple testing correction in one place, if not in another, and not in a way that the reader could see how this plays out. I don’t think that running corrections for multiple testing is mandatory (this depends on what kind of conclusion is drawn), but if authors don’t demonstrate any awareness for the issue, this raises suspicion. Fortunately, almost all given p-values are either seriously small or comfortably bigger than 0.05, so that I hardly expect conclusions to be affected by multiple testing – unless there actually was more testing that was not reported.
On the positive side, authors were interested in effect sizes (which are not confused with small p-values) and (mostly) gave intervals quantifying uncertainty. Also, often they plot the data in a way that the reader can see how the distributions look like and how big differences from what is expected under the null hypotheses actually are. One paper shows plots diagnosing model assumptions stating “the residuals are distributed uniformly around the zero line, indicating the
adequacy of the model”, but the corresponding plot indicates clear heteroscedasticity violating a model assumption.
My overall impression is somewhat ambivalent. There are misunderstandings and misinterpretations galore, but I don’t have the impression that any of them leads to a grossly wrong or misleading assessment of the subject matter. Of course I can’t rule out selective reporting or even computation errors or fake data, but none of the papers seems to hint at such a thing.
I’m fine with running a test as a routine device to check whether what was observed is compatible with meaningless random variation, e.g., “The concentration of sAA were similar between pasture and paddock (46.25 +/- 19.44 and 47.52 +/- 13.11 U/L, respectively; p = 0.7742) (…) BChol was found to be higher during pasture stable and lower in a statistically significant way when horses moved to paddock (12.44 +/- 6.30 and 5.58 +/- 2.39 mU/mL, respectively; p = 0.0068)” (Bazzano et al., 2024) with accompanying boxplots. The p-values here convey a message that is very relevant to the aim of the paper. If it is not possible to tell apart a difference from a model for meaningless random variation, it can for sure not be used as indication of anything meaningful.
I don’t actually think that this message could have been conveyed any better with any of the above mentioned “alternatives to p-values”, at least given that information to assess effect sizes is presented as well. The word “significant” doesn’t do any harm here as far as I’m concerned.
Scepticism regarding the generalisability of such results is always justified anyway, and for sure a p-value in isolation (and be it with an interval of effect sizes) doesn’t make a convincing discovery. There are always various ways to raise doubts. Ultimately I think that in order to achieve the status of a properly reliable scientific discovery, a result needs to be confirmed from different angles and by different authors. A single study with limited scope is generally not enough. If this were generally accepted, quite a bit of the worst trouble with overinterpretation of statistical tests would be out of the window.
Opponents of statistical tests could argue that in these papers there is a problem with binary thinking, which is encouraged by the binary logic of Neyman-Pearson type tests. Furthermore the list of issues with testing alone in the small set of papers considered here is rather long indeed.
I respond that (a) the tests serve a purpose, and (b) I don’t see how these problems will be solved using any other statistical method in the place of the tests (or just leaving them out). Most of the problems (e.g., confusion about paired observations; not being able to spot a model assumption violation) are on a level that will create issues with any statistical approach. Binary thinking will be brought in whenever an author wants to say that an observed effect is either meaningful or not, regardless of what method is used (a posterior probability for example can be thresholded just as well). Binary thinking (in the sense that “data show that either null hypothesis or alternative is true”) is inappropriate, but there is sometimes a need for binary decisions (e.g., should we use a method based on a normality assumption?), and language is essentially discrete, so any interpretation of a numerical result in words will imply some thresholding, if not necessarily transparently.
Statistical tests are the formalisation of an elementary intuition, namely that the observation of an event that is very unlikely under a certain probability model (which may involve fixed parameter values) indicates evidence against that model. This is probably the closest thing to falsification that we can have for data that comes with random variation and models that allow for the possibility of any outcome. As such, the idea behind tests looks very simple. Admittedly, fleshing it out and getting it into use in science comes with difficulties and complexities. In particular, often the probability for any specific result is low; at least for continuous distributions the probability for any precise observation is 0. Obviously just on this basis it wouldn’t make sense to say that there is evidence against the model whatever we observe. Instead, “statistical falsification” will require the specification of a rejection region (or rather regions at various levels, implying the definition of the p-value as “borderline rejection level”) in advance. There is more than one possible way to do this. Neyman and Pearson proposed to choose a test so that the probability to reject the null hypothesis is maximised under a certain alternative hypothesis of interest. Not taking for granted that reality follows either the null or the alternative hypothesis, such a test can be interpreted saying that the test statistic and nominal alternative hypothesis suggest a certain direction of deviation from the null hypothesis in case of rejection, such as generating larger values on average. In some situations tests may be preferred that work well distinguishing bigger sets of null and alternative hypotheses against each other rather than being optimal for specific ones. There are further subtleties, for example the issue of multiple testing, i.e., the potentially large probability of finding a meaningless significance if enough tests are run.
Most proposed alternatives to statistical tests are however even more complex (this may be controversial, but I won’t elaborate on the complexities of any specific alternative here), and several of them require understanding tests first, which is already hard. The superficial simplicity of tests is a blessing as well as a curse, as people all over the place without much statistical insight feel encouraged (or even forced) to use them. I suspect that some users of tests don’t even care about understanding what they are doing; following the “ritual” and getting published is enough. Other users may feel that they understand the basic idea, but may not know about the subtleties.
I actually think that coming up with alternative ways to formalise the evidence in the data is laudable. Many of these approaches have advantages that are potentially useful in certain situations, and I don’t think that people who use them should be pushed back into using standard tests (in many cases doing both, running a test and on top of it giving an e-value, relevance value or similar, can be informative). The scope of a p-value is limited, and there is more information in the data regarding hypotheses of interest than can be captured in a single number. So complementing a p-value with other information is a good thing, highlighting certain issues of p-values is helpful as well, but I am pretty convinced that replacing p-values with alternative single numbers will not improve matters.
This seems to be an instance of “solutionism”; the hope that the problems with statistics in practice can be solved by new methodology without addressing the underlying lack of statistical competence. Statisticians tend to get more credit (and higher level publications) by inventing new methodology than by increasing the understanding of existing ones, and the potential to come up with something that will later be used by generations of researchers is a strong incentive (even though this hope will be disappointed more often than not).
The aim of any initiative to improve the use of statistics should be to improve understanding. Pointing out misunderstandings and misinterpretations is necessary and worthwhile (reviewers should have asked for precise p-values instead of just stating “p<0.05”, and of course they should not have let authors get away with using a two-sample Wilcoxon for paired observations or a wrong interpretation of model diagnostics). I doubt that understanding can be improved by pushing people to use more complex methodology with which little experience exists. Changing methods or philosophy will not solve what really is the problem here. Wasserstein et al. 2019 and the papers in the Special Issue introduced by that editorial have some good things to say on how to improve data analysis in science (some of these papers explicitly advocate retaining statistical tests). Unfortunately the authors of the editorial chose to emphasise the red herring of “abandoning significance” over the more helpful aspects of their initiative.
References
Blume, J. D., Greevy, R. A., Welty, V. F., Smith, J. R., & Dupont, W. D. (2019). An Introduction to Second-Generation p-Values. The American Statistician, 73(sup1), 157-167. https://doi.org/10.1080/00031305.2018.1537893
Cole S. R., Edwards, J. K. and Greenland, S. (2021) Surprise! American Journal of Epidemiology 190(2), 191-193 https://doi.org/10.1093/aje/kwaa136
Grünwald, P., de Heide, R., & Koolen, W. M. (2023). Safe testing. arxiv, https://doi.org/10.48550/arXiv.1906.07801. To appear as discussion paper in Journal of the Royal Statistical Society
Mayo, D. (2010) Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars, Cambridge University Press.
Stahel, W. A. (2021) New relevance and significance measures to replace
p-values. PLoS ONE 16(6): e0252991. https://doi.org/10.1371/journal.pone.0252991
Wasserstein, R., Schirm, A. and Lazar, N. (2019) Moving to a World Beyond ‘p < 0.05’, The American Statistician 73(S1), 1-19: Editorial. https://doi.org/10.1080/00031305.2019.1583913
.
Professor Andrea Saltelli
UPF Barcelona School of Management, Barcelona, Spain, Centre for the Study of the Sciences and the Humanities, University of Bergen, Bergen, Norway
[An earlier post by A. Saltelli on this topic: Nov 22, 2019: A. Saltelli (Guest post): What can we learn from the debate on statistical significance?]
Analytic flexibility: a badly kept secret?
In a previous post in this blog I expressed concern about a loss of trust that could incur the activity of scientific quantification – as practiced in several discipline – unless some technical and normative element of crisis could be managed. The piece warned that the phenomenon could lead to “a decline of public trust in the findings of science”. Five years and one pandemic later, we may wonder if the danger has indeed materialized.
Looking at a broader time horizon, two decades have elapsed from the onset of the reproducibility crisis, with Ioannidis 2005 categorical “Why Most Published Research Findings Are False” all the way to the large reproducibility experiments of present times, with “many researchers using the same data”.
My personal reading is that the last five years have witnessed an acceleration of the crisis leading to a progressive recalibration of what science can do and not do in a context characterized by increased polarization. We inhabit now an increasingly sophisticated landscape where the epistemic authority of science is appropriated via fact signalling – “a practice where the stylistic tropes of logical thinking, scientific research, or data analysis is worn like a costume to bolster a sense of moral righteousness and certitude” – by practically all actors, from purported fact-checkers to corporate interests to governments, without forgetting scientists themselves.
Statisticians and econometricians have continued in recent times to shed light on the so-called analytical flexibility of statistical work which has variously been described as a garden of forking paths or a universe of uncertainty hiding in plain sight.
Modellers reached a similar conclusion four decades ago – when hydrologists provocatively wrote that with some ‘judiciously fiddling’ any conclusion could be reached using a mathematical model. This became substance for satire when Douglas Adam, the lucky author of the Hitchhiker’s Guide to the Galaxy, created as a character in one of his novels the model developer that could link any desired conclusions to the available observations via a “plausible series of logical-sounding steps”.
A few practitioners have realized that as an alternative, or a complement, to tens of different teams trying to replicate the same data (seventy-three in Breznau et. al) one may try to anticipate what would result from such an experiment by propagating the uncertainty or ambiguity implicit in the garden of forking paths by simulation, i.e. by running iteratively the analysis with different combinations of assumptions about data and models. Some authors have called this with some imagination (and a catchy name) multiverse analysis.
With my collaborators, we call this with the admittedly less catchy ‘modelling of the modelling process’, or simply global sensitivity analysis. These approaches compare favourably with the alternative, I daresay; they rely on decades of experience in global sensitivity analysis, augmented by sensitivity auditing, an approach that aims to explore the entire model generating process, inclusive of motivation and biases of the developers, expectations and purposes of the recipient of the analysis and so on. Sensitivity auditing is recommended in guidelines of European institutions and academia and examples of sensitivity auditing and of modelling of the modelling process can be found in a recent book on the policy of modelling published by Oxford University Press. On a personal note, sensitivity auditing is inspired by a philosophical orientation known as post-normal science that helped me in my own research.
An important element of these approaches is that they do not eschew engaging with sociology of quantification. As noted in my previous post, our crises are also epistemological and philosophical, and we need more than ever disciplines to work together abandoning the tragic divide among different families of science. The reproducibility crisis has important political and legal ramifications, that feed into our discussions about the facticity of facts and about how science is mobilized to support policy decisions. A comparison among different communities on the merit of the different ways to look at our quantification would be enlightening. Statisticians are now convinced that “it is easy to miss or downplay philosophical presuppositions, especially if one has a strong interest in endorsing the policy upshot” (Mayo 2021), and that quantifications, including statistical ones, need to be looked at with a double lens, technical as well as normative, as recommended by sociologists of quantification.
If we go back to the question posed above, about what is happening with public trust in science five years after the high point of the statistical wars on retiring significance, we must say that the rapid clock of policy has consigned us with a world where not just some, but possibly any issue involving society, technology and policy, has become intensely and increasingly polarized. As a result, even the interpretation of where we are now with public trust in science is a matter of conflict – the success and failures to address the pandemic being a case in point. For example, when it comes to successes and failure of science, scientific journals whose business model depends of ‘science solving problems’ may be led to reject works that point in the opposite direction.
My vision, that cannot be neutral for all that we discussed thus far, is that a new covenant between science and society is needed, an enterprise where statisticians can offer a special contribution due to the history of their discipline being rich in contributions from sociology and philosophy – more so I daresay than in other communities producing numbers and algorithms, but this is also up for debate.
I’m reblogging reader commentaries on my editorial, “The statistics wars and intellectual conflicts of interest“. 3 are published in Conservation Biology; a 4th, by Lakens, is in the Journal of the International Society of Physiotherapy. This post was first published on May 15, 2022. Thus, “soon to be” refers to the past. Share your remarks in the comments.
There are 3 commentaries soon to be published in Conservation Biology o also published in Conservation Biology.
.
Professor Philip B. Stark
Department of Statistics
University of California, Berkeley
You can read a draft of Philip Stark’s commentary here
.
Professor Christian Hennig
Department of Statistical Sciences “Paolo Fortunati”
University of Bologna.
Here is a draft of Christian Hennig’s commentary
Kent Staley and John Park , who each wrote individual commentaries for the blog, joined forces to write a joint commentary for the journal. You can read the draft here.
.
Kent W. Staley
Professor, Coordinator of Graduate Studies
Department of Philosophy
Saint Louis University
.
John Park, MD
Radiation Oncologist
Kansas City VA Medical Center
A commentary by Lakens will appear in the psychology journal that adopted the “abandon significance” recommendation discussed in my editorial. (See this post for a link.) Others may appear as letters or part of longer papers elsewhere. I will update this blog once I have that information. I’m very impressed with these interesting, interdisciplinary efforts, and grateful to all the authors for pursuing them.
All of the initial blog commentaries on Mayo’s (2021) editorial (up through Jan 18, 2022) are below*
Schachtman
Park
Dennis
StarkStaley
Pawitan
Hennig
Ionides and Ritov
Haig
Lakens
*After David Cox dies on January 18, 2022, I posted some memorial items. Two additional commentaries posted later are by Paul Daniels and Yu-li Ko.
For background: The slides and video from our January 11, 2022 Forum, with presentations by Y. Benjamini, D. Hand and I, which grew out of my editorial, can be found here:
January 11 Forum: “Statistical Significance Test Anxiety” : Benjamini, Mayo, Hand
.
Before posting new reflections on where we are 5 years after the ASA P-value controversy–both my own and readers’–I will reblog some reader commentaries from 2022 in connection with my (2022) editorial in Conservation Biology: “The Statistical Wars and Intellectual Conflicts of Interest”. First, here are excerpts from my editorial:
The Statistics Wars and Intellectual Conflicts of Interest
How should journal editors react to heated disagreements about statistical significance tests in applied fields, such as conservation science, where statistical inferences often are the basis for controversial policy decisions? They should avoid taking sides. They should also avoid obeisance to calls for author guidelines to reflect a particular statistical philosophy or standpoint. The question is how to prevent the misuse of statistical methods without selectively favoring one side.
The statistical‐significance‐test controversies are well known in conservation science. In a forum revolving around Murtaugh’s (2014) “In Defense of P values,” Murtaugh argues, correctly, that most criticisms of statistical significance tests “stem from misunderstandings or incorrect interpretations, rather than from intrinsic shortcomings of the P value” (p. 611). However, underlying those criticisms, and especially proposed reforms, are often controversial philosophical presuppositions about the proper uses of probability in uncertain inference. Should probability be used to assess a method’s probability of avoiding erroneous interpretations of data (i.e., error probabilities) or to measure comparative degrees of belief or support? Wars between frequentists and Bayesians continue to simmer in calls for reform.
Consider how, in commenting on Murtaugh (2014), Burnham and Anderson (2014 : 627) aver that “P‐values are not proper evidence as they violate the likelihood principle (Royall, 1997).” This presupposes that statistical methods ought to obey the likelihood principle (LP), a long‐standing point of controversy in the statistics wars. The LP says that all the evidence is contained in a ratio of likelihoods (Berger & Wolpert, 1988). Because this is to condition on the particular sample data, there is no consideration of outcomes other than those observed and thus no consideration of error probabilities. One should not write this off because it seems technical: methods that obey the LP fail to directly register gambits that alter their capability to probe error. Whatever one’s view, a criticism based on presupposing the irrelevance of error probabilities is radically different from one that points to misuses of tests for their intended purpose—to assess and control error probabilities.
Error control is nullified by biasing selection effects: cherry‐picking, multiple testing, data dredging, and flexible stopping rules. The resulting (nominal) p values are not legitimate p values. In conservation science and elsewhere, such misuses can result from a publish‐or‐perish mentality and experimenter’s flexibility (Fidler et al., 2017). These led to calls for preregistration of hypotheses and stopping rules–one of the most effective ways to promote replication (Simmons et al., 2012). However, data dredging can also occur with likelihood ratios, Bayes factors, and Bayesian updating, but the direct grounds to criticize inferences as flouting error probability control is lost. This conflicts with a central motivation for using p values as a “first line of defense against being fooled by randomness” (Benjamini, 2016). The introduction of prior probabilities (subjective, default, or empirical)–which may also be data dependent–offers further flexibility.
Signs that one is going beyond merely enforcing proper use of statistical significance tests are that the proposed reform is either the subject of heated controversy or is based on presupposing a philosophy at odds with that of statistical significance testing. It is easy to miss or downplay philosophical presuppositions, especially if one has a strong interest in endorsing the policy upshot: to abandon statistical significance. Having the power to enforce such a policy, however, can create a conflict of interest (COI). Unlike a typical COI, this one is intellectual and could threaten the intended goals of integrity, reproducibility, and transparency in science.
If the reward structure is seducing even researchers who are aware of the pitfalls of capitalizing on selection biases, then one is dealing with a highly susceptible group. For a journal or organization to take sides in these long-standing controversies—or even to appear to do so—encourages groupthink and discourages practitioners from arriving at their own reflective conclusions about methods.
The American Statistical Association (ASA) Board appointed a President’s Task Force on Statistical Significance and Replicability in 2019 that was put in the odd position of needing to “address concerns that a 2019 editorial [by the ASA’s executive director (Wasserstein et al., 2019)] might be mistakenly interpreted as official ASA policy” (Benjamini et al., 2021)—as if the editorial continues the 2016 ASA Statement on p-values (Wasserstein & Lazar, 2016). That policy statement merely warns against well‐known fallacies in using p values. But Wasserstein et al. (2019) claim it “stopped just short of recommending that declarations of ‘statistical significance’ be abandoned” and announce taking that step. They call on practitioners not to use the phrase statistical significance and to avoid p value thresholds. Call this the no‐threshold view. The 2016 statement was largely uncontroversial; the 2019 editorial was anything but. The President’s Task Force should be commended for working to resolve the confusion (Kafadar, 2019). Their report concludes: “P-values are valid statistical measures that provide convenient conventions for communicating the uncertainty inherent in quantitative results” (Benjamini et al., 2021). A disclaimer that Wasserstein et al., 2019 was not ASA policy would have avoided both the confusion and the slight to opposing views within the Association.
The no‐threshold view has consequences (likely unintended). Statistical significance tests arise “to test the conformity of the particular data under analysis with [a statistical hypothesis] H0 in some respect to be specified” (Mayo & Cox, 2006: 81). There is a function D of the data, the test statistic, such that the larger its value (d), the more inconsistent are the data with H0. The p value is the probability the test would have given rise to a result more discordant from H0 than d is were the results due to background or chance variability (as described in H0). In computing p, hypothesis H0 is assumed merely for drawing out its probabilistic implications. If even larger differences than d are frequently brought about by chance alone (p is not small), the data are not evidence of inconsistency with H0. Requiring a low pvalue before inferring inconsistency with H0 controls the probability of a type I error (i.e., erroneously finding evidence against H0).
…
Whether interpreting a simple Fisherian or an N‐P test, avoiding fallacies calls for considering one or more discrepancies from the null hypothesis under test. Consider testing a normal mean H0: μ ≤ μ0 versus H1: μ > μ0. If the test would fairly probably have resulted in a smaller p value than observed, if μ = μ1 were true (where μ1 = μ0 + γ, for γ > 0), then the data provide poor evidence that μ exceeds μ1. It would be unwarranted to infer evidence of μ > μ1. Tests do not need to be abandoned when the fallacy is easily avoided by computing p values for one or two additional benchmarks (Burgman, 2005; Hand, 2021; Mayo, 2018; Mayo & Spanos, 2006).
The same is true for avoiding fallacious interpretations of nonsignificant results. These are often of concern in conservation, especially when interpreted as no risks exist. In fact, the test may have had a low probability to detect risks. But nonsignificant results are not uninformative. If the test very probably would have resulted in a more statistically significant result were there a meaningful effect, say μ > μ1 (where μ1 = μ0 + γ, for γ > 0), then the data are evidence that μ < μ1. (This is not to infer μ ≤ μ0.) “Such an assessment is more relevant to specific data than is the notion of power” (Mayo & Cox, 2006: 89). This also matches inferring that μ is less than the upper bound of the corresponding confidence interval (at the associated confidence level) or a severity assessment (Mayo, 2018). Others advance equivalence tests (Lakens, 2017; Wellek, 2017). An N‐P test tells one to specify H0 so that the type I error is the more serious (considering costs); that alone can alleviate problems in the examples critics adduce (H0would be that the risk exists).
Many think the no‐threshold view merely insists that the attained p value be reported. But leading N‐P theorists already recommend reporting p, which “gives an idea of how strongly the data contradict the hypothesis…[and] enables others to reach a verdict based on the significance level of their choice” (Lehmann & Romano, 2005: 63−64). What the no‐threshold view does, if taken strictly, is preclude testing. If one cannot say ahead of time about any result that it will not be allowed to count in favor of a claim, then one does not test that claim. There is no test or falsification, even of the statistical variety. What is the point of insisting on replication if at no stage can one say the effect failed to replicate? One may argue for approaches other than tests, but it is unwarranted to claim by fiat that tests do not provide evidence. (For a discussion of rival views of evidence in ecology, see Taper & Lele, 2004.)
Many sign on to the no‐threshold view thinking it blocks perverse incentives to data dredge, multiple test, and p hack when confronted with a large, statistically nonsignificant p value. Carefully considered, the reverse seems true. Even without the word significance, researchers could not present a large (nonsignificant) p value as indicating a genuine effect. It would be nonsensical to say that even though more extreme results would frequently occur by random variability alone that their data are evidence of a genuine effect. The researcher would still need a small p value, which is to operate with a threshold. However, it would be harder to hold data dredgers culpable for reporting a nominally small p value obtained through data dredging. What distinguishes nominal p values from actual ones is that they fail to meet a prespecified error probability threshold.
…
While it is well known that stopping when the data look good inflates the type I error probability, a strict Bayesian is not required to adjust for interim checking because the posterior probability is unaltered. Advocates of Bayesian clinical trials are in a quandary because “The [regulatory] requirement of Type I error control for Bayesian [trials] causes them to lose many of their philosophical advantages, such as compliance with the likelihood principle” (Ryan etal., 2020: 7).
It may be retorted that implausible inferences will indirectly be blocked by appropriate prior degrees of belief (informative priors), but this misses the crucial point. The key function of statistical tests is to constrain the human tendency to selectively favor views they believe in. There are ample forums for debating statistical methodologies. There is no call for executive directors or journal editors to place a thumb on the scale. Whether in dealing with environmental policy advocates, drug lobbyists, or avid calls to expel statistical significance tests, a strong belief in the efficacy of an intervention is distinct from its having been well tested. Applied science will be well served by editorial policies that uphold that distinction.
During the pandemic, I ran a phil stat research seminar seminar (for the London School of Economics), a series of forums, and, finally, 4 workshops: “Phil Stat Wars and their Casualties” (phil-stat-wars.com), as our planned in-person events kept getting delayed. You can find videos for all these at phil-stat-wars.com.
les stats, c’est moi
This is the last of the selected posts I will reblog from 5 years ago on the 2019 statistical significance controversy. The original post, published on this blog on December 13, 2019, had 85 comments, so you might find them of interest. I invite readers to share their thoughts as to where the field is now, in relation to that episode, and to alternatives being used as replacements for statistical significance tests. Use the comments and send me guest posts.
When it comes to the statistics wars, leaders of rival tribes sometimes sound as if they believed “les stats, c’est moi”. [1]. So, rather than say they would like to supplement some well-known tenets (e.g., “a statistically significant effect may not be substantively important”) with a new rule that advances their particular preferred language or statistical philosophy, they may simply blurt out: “we take that step here!” followed by whatever rule of language or statistical philosophy they happen to prefer (as if they have just added the new rule to the existing, uncontested tenets). Karan Kefadar, in her last official (December) report as President of the American Statistical Association (ASA), expresses her determination to call out this problem at the ASA itself. (She raised it first in her June article, discussed in my last post.)
One final challenge, which I hope to address in my final month as ASA president, concerns issues of significance, multiplicity, and reproducibility. In 2016, the ASA published a statement that simply reiterated what p-values are and are not. It did not recommend specific approaches, other than “good statistical practice … principles of good study design and conduct, a variety of numerical and graphical summaries of data, understanding of the phenomenon under study, interpretation of results in context, complete reporting and proper logical and quantitative understanding of what data summaries mean.”
The guest editors of the March 2019 supplement to The American Statistician went further, writing: “The ASA Statement on P-Values and Statistical Significance stopped just short of recommending that declarations of ‘statistical significance’ be abandoned. We take that step here. … [I]t is time to stop using the term ‘statistically significant’ entirely.”
Many of you have written of instances in which authors and journal editors—and even some ASA members—have mistakenly assumed this editorial represented ASA policy. The mistake is understandable: The editorial was co-authored by an official of the ASA. In fact, the ASA does not endorse any article, by any author, in any journal—even an article written by a member of its own staff in a journal the ASA publishes. (Kafadar, December President’s Corner)
Yet Wasserstein et al. 2019 describes itself as a continuation of the ASA 2016 Statement on P-values, which I abbreviate as ASA I. (Wasserstein is the Executive Director of the ASA.) It describes itself as merely recording the decision to “take that step here”, and add one more “don’t” to ASA I. As part of this new “don’t,” it also stipulates that we should not consider “at all” whether pre-designated P-value thresholds are met. (It also restates four of the six principles in ASA I so as to be considerably stronger than those in ASA I. I argue, in fact, the resulting principles are inconsistent with principles 1 and 4 of ASA I. See my post from June 17, 2019.) Since it describes itself as a continuation of the ASA policy in ASA I, and that description survived peer review at the journal TAS, readers presume that’s what it is; absent any disclaimer to the contrary, that conception (or misconception) remains operative.
There really is no other way to read the claim in the Wasserstein et al. March 2019 editorial: “The ASA Statement on P-Values and Statistical Significance stopped just short of recommending that declarations of ‘statistical significance’ be abandoned.[2] We take that step here.” Had the authors viewed their follow-up as anything but a continuation of ASA I, they would have said something like: “Our own recommendation is to go much further than ASA I. We suggest that all branches of science stop using the term ‘statistically significant’ entirely.” They do not say that. What they say is written from the perspective of “Les stats, c’est moi”.
The 2019 P-value Project II
Kafadar deserves a great deal of credit for providing some needed qualification in her December note. However, there needs to be a disclaimer by ASA as regards what it calls its P-value Project. The P-value project, started in 2014, refers to the overall ASA campaign to provide guides for the correct use and interpretation of P-values and statistical significance, and journal editors and societies are to consider revising their instructions to authors taking into account its guidelines. ASA I was distilled from many meetings and discussions from representatives in statistics. The only difference in today’s P-value Project is that both ASA I and the 2019 editorial by Wasserstein et al. are to form the new ASA guidelines–even if the latter is not to be regarded as a continuation of ASA I (in accord with Kafadar’s qualification). I will refer to it as the 2019 ASA P-value Project II.(note) Wasserstein et al. 2019 is a piece of the P-value project, and the authors thank the ASA for its support of this Project at the end of the article. [4] [5]
Of Policies and Working Groups
Kafadar continues:
Even our own ASA members are asking each other, “What do we tell our collaborators when they ask us what they should do about statistical hypothesis tests and p-values?” Should the ASA have a policy on hypothesis testing or on using “statistical significance”?
Allow me to weigh in here: No, no it should not. At one time I would have said yes, but no more. I can hear the policy now (sounding much like Wasserstein et al. 2019, only written in stone): “Don’t say, never say, or if you really feel you must say significance, and are prepared to thoroughly justify such a “thoughtless” term, then you may only say “significance level p” where p is continuous, and never rounded up or cut off, ever. But never, ever use the “ant” ending: significant. You can’t, can’t, can’t say results are statistically significant (at level p). The only exception would be if you’re giving the history of statistics. (3)
Why can’t the ASA merely provide a bipartisan forum for discussion of the multitude of models, methods, aims, goals, and philosophies of its members? Wasserstein et al. 2019 admits there is no agreement, and that there might never be. Spare us another document whose implication is: we need not test, and cannot falsify claims, even statistically (since that is the consequence of no thresholds). I realize that Kafadar is calling for a serious statement–one that counters the impression of the Wasserstein et al. opinion.
To address these issues, I hope to establish a working group that will prepare a thoughtful and concise piece reflecting “good statistical practice,” without leaving the impression that p-values and hypothesis tests—and, perhaps by extension as many have inferred, statistical methods generally—have no role in “good statistical practice.” …The ASA should develop—and publicize—a properly endorsed statement on these issues that will guide good practice.
Be careful what you wish for. I give major plaudits to Kafadar for pressing hard to see that alternative views are respected, and to counter the popular but terrible arguments of the form: since these methods are misused, they should be banished, and replaced with methods advocated by group Z (even if the credentials of Z’s methods haven’t been scrutinized!) We have already seen in 2019 the extensive politicization and sensationalizing of bandwagons in statistics. (See my editorial P-value Thresholds: Forfeit at your Peril.) The average ASA member, who doesn’t happen to be a thought leader or member of a politically correct statistical-philosophical tribe, is in great danger of being muffled entirely. There’s already a loss of trust. We already know, under the motto that “a crisis should never be wasted”, that many leaders of statistical tribes view the crisis of replication as an opportunity to sell alternative methods they have long been promoting. Rather than the properly endorsed, truly representative, statement that Kafadar seeks, we may get dictates from those who are quite convinced that they know best: “les stats, c’est moi”.
APPENDIX. How a Working Group on P-values and Significance Testing Could Work
I see one way that a working group could actually work. The 2016 ASA statement, ASA I, had a principle, it was #4. You don’t hear about it in the 2019 follow-up. It is that “P-values and related statistics” cannot be correctly interpreted without knowing how many hypotheses were tested, how data were specified and results selected for inference. Notice the qualification “and related statistics”. The presumption is that some methods don’t require that information! That information is necessary only if one is out to control the error probabilities associated with an inference.
Here’s my idea: Have the group consist of those who work in areas where statistical inferences depend on controlling error probabilities (I call such methods error statistical). They would be involved in current uses and developments of statistical significance testing and the much larger (frequentist) error statistical methodology within which it forms just a part. They would be familiar with, and some would be involved in developing, the latest error statistical tools, including tests and confidence distributions, P-values with high dimensional data, current problems of adjusting for multiple testing, and of testing statistical model assumptions, and they would be capable of different aspects of comparative statistical methods (Bayesian and error statistical). They would present their findings and recommendations, and responses sought.
The need for the kind of forum I’m envisioning is so pressing, that it should not be contingent on being created by any outside association. It should emerge spontaneously in 2020. We take that step here.
Please share your comments in the comments.
[1] This is a pun on “l’état, c’est moi” (“I am the state”, Louis XIV.) I thank Glenn Shafer for the appropriate French spelling for my pun. (Thanks to S. Senn for noticing I was missing the X in Louis XIV.)
[2] They are referring to the last section of ASA I on “other measures of evidence”. Indeed, that section suggests an endorsement of an assortment of alternative measures of evidence including Bayes factors, likelihood ratios and others. There is no attention to whether any of these methods accomplish the key task of the statistical significance test–to distinguish genuine from spurious effects. For a fuller explanation of this last section, please see my post from June 17, 2019 and November 14, 2019. And, obviously, check the last section of ASA I.
Shortly after the 2019 editorial appeared, I queried Wasserstein as to the relationship between it and ASA I. It was never clarified. I hope now that it will be. At the same time I informed him of what appeared to me to be slips in expressing principles of ASA I, and I offered friendly amendments (see my post from June 17, 2019).
[3] If you’re giving the history of statistics, you can speak of those bad, bad men–dichotomaniacs, Neyman and Pearson– who, following Fisher, divided results into significant and non-significant discrepancies (introduced the alternative hypotheses, type I and II errors, power and optimal tests) and thereby tried to reduce all of statistics to acceptance sampling, engineering, and 5-year plans in Russia–as Fisher (1955) himself said (after the professional break with Neyman in 1935). Never mind that Neyman developed confidence intervals at the same time, 1930. For a full discussion of the history of the Fisher-Neyman (and related) wars, please see my Statistical Inference as severe Testing: How to Get Beyond the Statistics Wars (CUP, 2018).
[4] I was just sent this podcast and interview of Ron Wasserstein, so I’m adding it as a footnote. There, Wasserstein et al. 2019 is clearly described as the ASA’s “further guidance”, and Wasserstein takes no exception to it. The interviewer says:
“But it would seem as though Ron’s work has only just begun. The ASA has just published further guidance in the most recent edition of The American Statistician, which is open access and written for non-statisticians. The guidance is intended to go further and argues for an end to the concept of statistical significance and towards a model which the ASA have coined their ATOM Principle: Accept uncertainty, Thoughtful, Open and Modest.”
https://www.howresearchers.com/wp-content/uploads/2019/05/hrcw-transcript-episode-2.pdf
[5]Nathan Schachtman, in a new post just added to his law blog on this very topic, displays a letter from the ASA acknowledging that a journal has revised its guidelines taking into account both ASA I and the 2019 Wasserstein et al. editorial. I had seen this letter, in relation to the NEJM, but it’s hard to know what to make of it. I haven’t seen others acknowledging other journals, and there have been around 7 at this point. I may just be out of the loop.
Selected blog posts on ASA I and the Wasserstein et al. 2019 editorial:
.
I continue my selective 5-year review of some of the posts revolving around the statistical significance test controversy from 2019. This post was first published on the blog on November 14, 2019. I feared then that many of the howlers of statistical significance tests would be further etched in granite after the ASA’s P-value project, and in many quarters this is, unfortunately, true. One that I’ve noticed quite a lot is the (false) supposition that negative results are uninformative. Some fields, notably psychology, keep to a version of simple Fisherian tests, ignoring Neyman-Pearson (N-P) tests (never minding that Jacob Cohen was a psychologist who gave us “power analysis”). (See note [1]) For N-P, “it is immaterial which of the two alternatives…is labelled the hypothesis tested” (Neyman 1950, 259). Failing to find evidence of a genuine effect, coupled with a test’s having high capability to detect meaningful effects, warrants inferring the absence of meaningful effects. Even with the simple Fisherian test, failing to reject H0 is informative. Null results figure importantly throughout science, such as when the ether was falsified by Michelson-Morley, and in directing attention away from unproductive theory development.
Please share your comments on this blogpost.
Everything is impeach and remove these days! Should that hold also for the concept of statistical significance and P-value thresholds? There’s an active campaign that says yes, but I aver it is doing more harm than good. In my last post, I said I would count the ways it is detrimental until I became “too disconsolate to continue”. There I showed why the new movement, launched by Executive Director of the ASA (American Statistical Association), Ronald Wasserstein (in what I dub ASA II(note)), is self-defeating: it instantiates and encourages the human-all-too-human tendency to exploit researcher flexibility, rewards, and openings for bias in research (F, R & B Hypothesis). That was reason #1. Just reviewing it already fills me with such dismay, that I fear I will become too disconsolate to continue before even getting to reason #2. So let me just quickly jot down reasons #2, 3, 4, and 5 (without full arguments) before I expire.
[I thought that with my book Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (2018, CUP), that I had said pretty much all I cared to say on this topic (and by and large, this is true), but almost as soon as it appeared in print just around a year ago, things got very strange.]
But wait. Someone might object that I’m the one doing more harm than good by linking the ASA (The American Statistical Association) to Wasserstein’s campaign to get publishers, journalists, authors and the general public to buy into the recommendations of ASA II(note). “Shhhhh!” some counsel, “don’t give it more attention; we want people to look away”. Nothing to see? I don’t think so. I will discuss this point in this post in PART II, as soon as I sketch my list of reasons #2-5.
Before starting, let me remind readers that what I abbreviate as ASA II(note) only refers to those portions of the 2019 editorial by Wasserstein, Schirm, and Lazar that allude to their general recommendations, not their summaries of contributed papers in the issue of TAS.
PART I
2 Decriminalize theft to end robbery. The key arguments for impeaching and removing statistical significance levels and P-value thresholds commit fallacies of the “cut-off your nose to spite your face” variety. For example, we should ban P-value thresholds because they cause biased selection and data dredging. Discard P-value thresholds and P-hacking disappears! Or so it is argued. Faced with unwelcome nonsignificant results, eager researchers are still led to massage, spin, and data dredge–only now it is much harder to directly hold them accountable. For the argument, see my (“P-value Thresholds: Forfeit at your Peril“, 2019).
3 Straw men and women fallacies. ASA I and II(note) do more harm than good by presenting oversimple caricatures of the tests. Even ASA I excludes a consideration of alternatives, error probabilities and power[1]. At the same time, it will contrast these threadbare “nil null” hypothesis tests with confidence intervals (CIs)–never minding that the latter employs alternatives. No wonder CIs look better, but such a test is unfair. (Neyman developed confidence intervals as inversions of tests at the same time he was developing hypotheses tests with alternatives in 1930. Using only significance tests, you could recover the lower (and upper) 1-α CI bounds if you wanted, by asking for the hypotheses that the data are statistically significantly greater (smaller) than, at level α, using the usual 2-sided computation).
In ASA II(note), we learn that “no p-value can reveal the …presence…of an association or effect” (at odds with principle 1 of ASA I). That could be true only in the sense that no formal statistical quantity alone could reveal the presence of an association. But in a realistic setting, small p-values surely do reveal the presence of effects. Yes, there are assumptions, but significance tests are prime tools to probe them. We hear of “the seductive certainty falsely promised by statistical significance”, and are told that “a declaration of statistical significance is the antithesis of thoughtfulness”. (How an account that never issues an inference without an associated error probability can be promising certainty is unexplained. On the second allegation, thresholds are rendered meaningful by choosing them to reflect background information and a host of theoretical and epistemic considerations,.) The requirement in philosophy of a reasonably generous interpretation of what your criticizing isn’t a call for being kind or gentle, it’s that otherwise your criticism is guilty of straw men (and women) fallacies, and thus fails.
4 Alternatives to significance testing are given a pass.You will not find any appraisal of the alternative methods recommended to replace significance tests for their intended tasks. Although many of the “alternative measures of evidence” listed in ASA I and II(note): Likelihood ratios, Bayes factors (subjective, default, empirical), posterior predictive values (in diagnostic screening) have been critically evaluated by leading statisticians, no word of criticism is heard here. Here’s an exercise: run down the list of 6 “principles” of ASA I, applying them to any of the alternative measures of evidence on offer. Take, for example, Bayes factors. I claim that they do worse than do significance tests, even without modifications.[2]
5 Assumes probabilism. Any fair (non question-begging) comparison of statistical methods should recognize different roles probability may play in inference. The role of probability in inference by way of statistical falsification is quite different from using probability to quantify degrees of confirmation, support, plausibility or belief in a statistical hypothesis or model–or comparative measures of these. I abbreviate the former as error statistical methods, the latter, as variants on probabilism. Use whatever terms you like. Statistical significance tests are part of a panoply of methods where probability arises to assess and control misleading interpretations of data.
Error probabilities quantify the capabilities of a method to detect the ways a claim (hypothesis, model or other) may be false, or specifiably flawed. The basic principle of testing is minimalist: there is evidence for a claim only to the extent it has been subjected to, and passes, a test that had at least a reasonable probability of having discerned how the claim may be false. (For a more detailed exposition, see Mayo 2018, or excerpts from SIST on this blog).
Reason #5, then, is that “measures of evidence” in both ASA I and II(note) beg this key question (about the role of probability in statistical inference) in favor of probabilisms–usually comparative as with Bayes factors. If the recommendation in ASA II(note) to remove statistical thresholds is taken seriously, there are no tests and no statistical falsification. Recall what Ioannidis said in objecting to “don’t say significance”, cited in my last post:
Potential for falsification is a prerequisite for science. Fields that obstinately resist refutation can hide behind the abolition of statistical significance but risk becoming self-ostracized from the remit of science. (Ioannidis 2019)
“Self-ostracizing” is a great term. ASA should ostracize self-ostracizing. This takes me back to the question I promised to come back to: is it a mistake to see the ASA as entangled in the campaign to ban use of the “S-word”, and kill P-value thresholds?
PART II (to read this part, please go to the original post)
[1] “To keep the statement reasonably simple, we did not address alternative hypotheses, error types, or power (among other things)” (ASA I)
[2] The ASA 2016 Guide’s Six Principles
[3] I am grateful to Ron Wasserstein for inviting me to be a “philosophical observer” at the 2015 meeting to draft ASA I.
2019 Blog posts on ASA II(note):
On ASA I:
REFERENCES:
Ioannidis J. (2019). The importance of predefined rules and prespecified statistical analyses: do not abandon significance. JAMA 321:2067‐2068. (pdf)
Ionides, E., Giessing, A., Ritov, Y. & Page, S. (2017). Response to the ASA’s Statement on p-Values: Context, Process, and Purpose, The American Statistician, 71:1, 88-89. (pdf)
Mayo, (2018). Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars, SIST (2018, CUP).
Mayo, D. G. (2019), P‐value thresholds: Forfeit at your peril. Eur J Clin Invest, 49: e13170. (pdf) doi:10.1111/eci.13170
Neyman, J. (1950), First course in Probability and Statistics, NY: Henry Holt.
Wasserstein, R. & Lazar, N. (2016), The ASA’s Statement on p-Values: Context, Process, and Purpose”. Volume 70, 2016 – Issue 2.
Wasserstein, R., Schirm, A. and Lazar, N. (2019) “Moving to a World Beyond ‘p < 0.05’”, The American Statistician 73(S1): 1-19: Editorial. (ASA II(note))(pdf)
I continue my 5-year review of some highlights from the “abandon significance” movement from 2019. This post was first published on this blog on November 30, 2019, It was based on a call by then American Statistical Association President, Karen Kafadar, which sparked a counter-movement. I will soon begin sharing a few invited guest posts reflecting on current thinking either on the episode or on statistical methodology more generally. I may continue to post such reflections over the summer, as they come in, so let me know if you’d like to contribute something. Share your thoughts in the comments.
Mayo writing to Kafadar
I never met Karen Kafadar, the 2019 President of the American Statistical Association (ASA), but the other day I wrote to her in response to a call in her extremely interesting June 2019 President’s Corner: “Statistics and Unintended Consequences“:
I only recently came across her call, and I will share my letter below. First, here are some excerpts from her June President’s Corner (her December report is due any day).
Recently, at chapter meetings, conferences, and other events, I’ve had the good fortune to meet many of our members, many of whom feel queasy about the effects of differing views on p-values expressed in the March 2019 supplement of The American Statistician (TAS). The guest editors— Ronald Wasserstein, Allen Schirm, and Nicole Lazar—introduced the ASA Statement on P-Values (2016) by stating the obvious: “Let us be clear. Nothing in the ASA statement is new.” Indeed, the six principles are well-known to statisticians.The guest editors continued, “We hoped that a statement from the world’s largest professional association of statisticians would open a fresh discussion and draw renewed and vigorous attention to changing the practice of science with regards to the use of statistical inference.”…
Wait a minute. I’m confused about who is speaking. The statements “Let us be clear…” and “We hoped that a statement from the world’s largest professional association…” come from the 2016 ASA Statement on P-values. I abbreviate this as ASA I (Wasserstein and Lazar 2016). The March 2019 editorial that Kafadar says is making many members “feel queasy,” is the update (Wasserstein, Schirm, and Lazar 2019). I abbreviate it as ASA II [i].(note)
A healthy debate about statistical approaches can lead to better methods. But, just as Wilks and his colleagues discovered, unintended consequences may have arisen: Nonstatisticians (the target of the issue) may be confused about what to do. Worse, “by breaking free from the bonds of statistical significance” as the editors suggest and several authors urge, researchers may read the call to “abandon statistical significance” as “abandon statistical methods altogether.” …
But we may need more. How exactly are researchers supposed to implement this “new concept” of statistical thinking? Without specifics, questions such as “Why is getting rid of p-values so hard?” may lead some of our scientific colleagues to hear the message as, “Abandon p-values”—despite the guest editors’ statement: “We are not recommending that the calculation and use of continuous p-values be discontinued.”
Brad Efron once said, “Those who ignore statistics are condemned to re-invent it.” In his commentary (“It’s not the p-value’s fault”) following the 2016 ASA Statement on P-Values, Yoav Benjamini wrote, “The ASA Board statement about the p-values may be read as discouraging the use of p-values because they can be misused, while the other approaches offered there might be misused in much the same way.” Indeed, p-values (and all statistical methods in general) can be misused. (So may cars and computers and cell phones and alcohol. Even words in the English language get misused!) But banishing them will not prevent misuse; analysts will simply find other ways to document a point—perhaps better ways, but perhaps less reliable ones. And, as Benjamini further writes, p-values have stood the test of time in part because they offer “a first line of defense against being fooled by randomness, separating signal from noise, because the models it requires are simpler than any other statistical tool needs”—especially now that Efron’s bootstrap has become a familiar tool in all branches of science for characterizing uncertainty in statistical estimates.[Benjamini is commenting on ASA I.]
… It is reassuring that “Nature is not seeking to change how it considers statistical evaluation of papers at this time,” but this line is buried in its March 20 editorial, titled “It’s Time to Talk About Ditching Statistical Significance.” Which sentence do you think will be more memorable? We can wait to see if other journals follow BASP’s lead and then respond. But then we’re back to “reactive” versus “proactive” mode (see February’s column), which is how we got here in the first place.
… Indeed, the ASA has a professional responsibility to ensure good science is conducted—and statistical inference is an essential part of good science. Given the confusion in the scientific community (to which the ASA’s peer-reviewed 2019 TAS supplement may have unintentionally contributed), we cannot afford to sit back. After all, that’s what started us down the “abuse of p-values” path.
Is it unintentional? [ii]
…Tukey wrote years ago about Bayesian methods: “It is relatively clear that discarding Bayesian techniques would be a real mistake; trying to use them everywhere, however, would in my judgment, be a considerably greater mistake.” In the present context, perhaps he might have said: “It is relatively clear that trusting or dismissing results based on a single p-value would be a real mistake; discarding p-values entirely, however, would in my judgment, be a considerably greater mistake.” We should take responsibility for the situation in which we find ourselves today (and during the past decades) to ensure that our well-researched and theoretically sound statistical methodology is neither abused nor dismissed categorically. I welcome your suggestions for how we can communicate the importance of statistical inference and the proper interpretation of p-values to our scientific partners and science journal editors in a way they will understand and appreciate and can use with confidence and comfort—before they change their policies and abandon statistics altogether. Please send me your ideas!
You can read the full June President’s Corner.
On Fri, Nov 8, 2019 at 2:09 PM Deborah Mayo mayod@vt.edu wrote:
Dear Professor Kafadar;
Your article in the President’s Corner of the ASA for June 2019 was sent to me by someone who had read my “P-value Thresholds: Forfeit at your Peril” editorial, invited by John Ioannidis. I find your sentiment welcome and I’m responding to your call for suggestions.
For starters, when representatives of the ASA issue articles criticizing P-values and significance tests, recommending their supplementation or replacement by others, three very simple principles should be followed:
Here’s what I recommend ASA do now in order to correct the distorted picture that is now widespread and growing: Run a conference akin to the one Wasserstein ran on “A World Beyond ‘P < 0.05′” except that it would be on evaluating some competing methods for statistical inference: Comparative Methods of Statistical Inference: Problems and Prospects.
The workshop would consist of serious critical discussions on Bayes Factors, confidence intervals[iii], Likelihoodist methods, other Bayesian approaches (subjective, default non-subjective, empirical), particularly in relation to today’s replication crisis. …
Growth of the use of these alternative methods have been sufficiently widespread to have garnered discussions on well-known problems….The conference I’m describing will easily attract the leading statisticians in the world. …
Sincerely,
D. Mayo
Please share your comments on this blogpost.
[i] My reference to ASA II(note) refers just to the portion of the editorial encompassing their general recommendations: don’t say significance or significant, oust P-value thresholds. (It mostly encompasses the first 10 pages.) It begins with a review of 4 of the 6 principles from ASA I, even though they are stated in more extreme terms than in ASA I. (As I point out in my blogpost, the result is to give us principles that are in tension with the original 6.) Note my new qualification in [ii]*
[ii]As soon as I saw the 2019 document, I queried Wasserstein as to the relationship between ASA I and II(note). It was never clarified. I hope now that it will be, with some kind of disclaimer. That will help, but merely noting that it never came to a Board vote will not quell the confusion now rattling some ASA members. The ASA’s P-value campaign to editors to revise their author guidelines asks them to take account of both ASA I and II(note). In carrying out the P-value campaign, at which he is highly effective, Ron Wasserstein obviously wears his Executive Director’s hat. See The ASA’s P-value Project: Why It’s Doing More Harm than Good. So, until some kind of clarification is issued by the ASA, I’ve hit upon this solution.
The ASA P-value Project existed before the 2016 ASA I. The only difference in today’s P-value Project–since the March 20, 2019 editorial by Wasserstein et al– is that the ASA Executive Director (in talks, presentations, correspondence) recommends ASA I and the general stipulations of ASA II(note)–even though that’s not a policy document. I will now call it the 2019 ASA P-value Project II. It also includes the rather stronger principles in ASA II(note). Even many who entirely agree with the “don’t say significance” and “don’t use P-value thresholds” recommendations have concurred with my “friendly amendments” to ASA II(note) (including, for example, Greenland, Hurlbert, and others). See my post from June 17, 2019.
You merely have to look at the comments to that blog. If Wasserstein would make those slight revisions, the 2019 P-value Project II wouldn’t contain the inconsistencies, or at least “tensions” that it now does, assuming that it retains ASA I. The 2019 ASA P-value Project II sanctions making the recommendations in ASA II(note), even though ASA II(note) is not an ASA policy statement.
However, I don’t see that those made queasy by ASA II(note) would be any less upset with the reality of the ASA P-value Project II.
[iii]Confidence intervals (CIs) clearly aren’t “alternative measures of evidence” in relation to statistical significance tests. The same man, Neyman, developed tests (with Pearson) and CIs, even earlier ~1930. They were developed as duals, or inversions, of tests. Yet the advocates of CIs–the CI Crusaders, S. Hurlbert calls them–are some of today’s harshest and most ungenerous critics of tests. For these crusaders, it has to be “CIs only”. Supplementing p-values with CIs isn’t good enough. Now look what’s happened to CIS in the latest guidelines of the NEJM. You can readily find them searching NEJM on this blog. (My own favored measure, severity, improves on CIs, moves away from the fixed confidence level, and provides a different assessment corresponding to each point in the CI.
*Or is it not obvious? I think it is, because he is invited and speaks, writes, and corresponds in that capacity.
Wasserstein, R. & Lazar, N. (2016) [ASA I], The ASA’s Statement on p-Values: Context, Process, and Purpose”. Volume 70, 2016 – Issue 2.
Wasserstein, R., Schirm, A. and Lazar, N. (2019) [ASA II(note)] “Moving to a World Beyond ‘p < 0.05’”, The American Statistician 73(S1): 1-19: Editorial.(pdf)
Related posts on ASA II(note):
Related book (excerpts from posts on this blog are collected here)
Mayo, (2018). Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars, SIST (2018, CUP).
.
Soon after the Wasserstein et al (2019) “don’t say significance” editorial, John Ioannidis invited Andrew Gelman and I to write editorials from our different perspectives on an associated editorial that Nature invited. It was written by Amrhein, Greenland and McShane (AGM, 2019). Prior to the publication of AGM 2019, people were given the opportunity to add their names to the Nature article.
A campaign followed that aimed at the collection of signatures in what was called a ‘petition’ on the widely popular blogsite of Andrew Gelman. Ultimately, 854 scientists signed the petition and the list of their names was published along with commentary. (Hardwicke and Ioannidis, 2019, p. 2)
Tom Hardwicke and John Ioannidis (2019) took advantage of the opportunity “to perform a survey of the signatories to understand how and why they signed the endorsement” (ibid.). This post, reblogged from September 25 2019, includes all 3 articles: the survey by Hardwicke and Ioannidis, and the editorials by Gelman and I. They appeared in the European Journal of Clinical Investigations (2019). I’m still interested in reader responses (in the comments) to the question I pose.
The October 2019 issue of the European Journal of Clinical Investigations came out today. It includes the PERSPECTIVE article by Tom Hardwicke and John Ioannidis, an invited editorial by Gelman and one by me:
Petitions in scientific argumentation: Dissecting the request to retire statistical significance, by Tom Hardwicke and John Ioannidis
When we make recommendations for scientific practice, we are (at best) acting as social scientists, by Andrew Gelman
P-value thresholds: Forfeit at your peril, by Deborah Mayo
I blogged excerpts from my preprint, and some related posts, here.
All agree to the disagreement on the statistical and metastatistical issues:
Despite the admitted disparate views, ASA representatives come out, in 2019, forcefully on the side of: Don’t use P-value thresholds (“at all”) in interpreting data, and Never describe results as attaining “statistical significance at level p”. Should the ASA, as an umbrella group, be striving to provide a relatively neutral forum for open, pressure-free, discussion of different methods–their pros and cons? This is a leading question, true. As an outsider, I’m interested to know what both insiders and outsiders think.[i]
Links to ASA I and IInote:
Wasserstein, and N. Lazar. 2016 ASA Statement on P-Values and Statistical Significance (ASA I).
Wasserstein, R., Schirm A. and N. Lazar. “Moving to a world beyond ‘p< 0.05‘” (ASA II)note
This is the guest post by Bran Haig on July 12, 2019 in response to the “abandon statistical significance” editorial in The American Statistician (TAS) by Wasserstein, Schirm, and Lazar (WSL 2019). In the post it is referred to as ASAII with a note added once we learned that it is actually not a continuation of the 2016 ASA policy statement. (I decided to leave it that way, as otherwise the context seems lost. But in the title to this post, I refer to the journal TAS.) Brian lists some of the benefits that were to result from abandoning statistical significance. I welcome your constructive thoughts in the comments.
Brian Haig, Professor Emeritus
Department of Psychology
University of Canterbury
Christchurch, New Zealand
The American Statistical Association’s (ASA)(note) recent effort to advise the statistical and scientific communities on how they should think about statistics in research is ambitious in scope. It is concerned with an initial attempt to depict what empirical research might look like in “a world beyond p<0.05” (The American Statistician, 2019, 73, S1,1-401). Quite surprisingly, the main recommendation of the lead editorial article in the Special Issue of The American Statistician devoted to this topic (Wasserstein, Schirm, & Lazar, 2019; hereafter, ASA II(note)) is that “it is time to stop using the term ‘statistically significant’ entirely”. (p.2) ASA II(note) acknowledges the controversial nature of this directive and anticipates that it will be subject to critical examination. Indeed, in a recent post, Deborah Mayo began her evaluation of ASA II(note) by making constructive amendments (reblogged) to three recommendations that appear early in the document (‘Error Statistics Philosophy’, June 17, 2019). These amendments have received numerous endorsements, and I record mine here. In this short commentary, I briefly state a number of general reservations that I have about ASA II(note).
1. The proposal that we should stop using the expression “statistical significance” is given a weak justification
ASA II(note) proposes a superficial linguistic reform that is unlikely to overcome the widespread misconceptions and misuse of the concept of significance testing. A more reasonable, and common-sense, strategy would be to diagnose the reasons for the misconceptions and misuse and take ameliorative action through the provision of better statistics education, much as ASA I did with p values. Interestingly, ASA II(note) references Mayo’s recent book, Statistical Inference as Severe Testing (2018), when mentioning the “statistics wars”. However, it refrains from considering the fact that her error-statistical perspective provides an informed justification for continuing to use tests of significance, along with the expression, “statistically significant”. Further, ASA II(note) reports cases where some of the Special Issue authors thought that use of a p-value threshold might be acceptable. However, it makes no effort to consider how these cases might challenge their main recommendation.
2. The claimed benefits of abandoning talk of statistical significance are hopeful conjectures.
ASA II(note) makes a number of claims about the benefits that it thinks will follow from abandoning talk of statistical significance. It says,“researchers will see their results more easily replicated – and, even when not, they will better understand why”. “[We] will begin to see fewer false alarms [and] fewer overlooked discoveries …”. And, “As ‘statistical significance’ is used less, statistical thinking will be used more.” (p.1) I do not believe that any of these claims are likely to follow from retirement of the expression, “statistical significance”. Unfortunately, no justification is provided for the plausibility of any of the alleged benefits. To take two of these claims: First, removal of the common expression, “significance testing” will make little difference to the success rate of replications. It is well known that successful replications depend on a number of important factors, including research design, data quality, effect size, and study power, along with the multiple criteria often invoked in ascertaining replication success. Second, it is just implausible to suggest that refraining from talk about statistical significance will appreciably help overcome mechanical decision-making in statistical practice, and lead to a greater engagement with statistical thinking. Such an outcome will require, among other things, the implementation of science education reforms that centre on the conceptual foundations of statistical inference.
3. ASA II’s(note) main recommendation is not a majority view.
ASA II(note) bases its main recommendation to stop using the language of “statistical significance” in good part on its review of the articles in the Special Issue. However, an inspection of the Special Issue reveals that this recommendation is at variance with the views of many of the 40-odd articles it contains. Those articles range widely over topics covered, and attitudes to, the usefulness of tests of significance. By my reckoning, only two of the articles advocate banning talk of significance testing. To be fair, ASA II(note) acknowledges the diversity of views held about the nature of tests of significance. However, I think that this diversity should have prompted it to take proper account of the fact that its recommendation is only one of a number of alternative views about significance testing. At the very least, ASA II(note) should have tempered its strong recommendation not to speak of statistical significance any more.
4.The claim for continuity between ASA I and ASA II(note) is misleading. There is no evidence in ASA I (Wasserstein & Lazar, 2016) for the assertion made in ASA II(note) that the earlier document stopped just short of recommending that claims of “statistical significance” should be eliminated. In fact, ASA II(note) marks a clear departure from ASA I, which was essentially concerned with how to better understand and use p-values. There is nothing in the earlier document to suggest that abandoning talk of statistical significance might be the next appropriate step forward in the ASA’s efforts to guide our statistical thinking.
5. Nothing is said about scientific method, and little is said about science.
The announcement of the ASA’s 2017 Symposium on Statistical Inference stated that the Symposium would “focus on specific approaches for advancing scientific methods in the 21st century”. However, the Symposium, and the resulting Special Issue of The American Statistician, showed little interest in matters to do with scientific method. This is regrettable because the myriad insights about scientific inquiry contained in contemporary scientific methodology have the potential to greatly enrich statistical science. The post-p< 0.05 world depicted by ASA II(note) is not an informed scientific world. It is an important truism that statistical inference plays a major role in scientific reasoning. However, for this role to be properly conveyed, ASA II(note) would have to employ an informative conception of the nature of scientific inquiry.
6. Scientists who speak of statistical significance do embrace uncertainty. I think that it is uncharitable, indeed incorrect, of ASA II(note) to depict many researchers who use the language of significance testing as being engaged in a quest for certainty. John Dewey, Charles Peirce, and Karl Popper taught us quite some time ago that we are fallible, error-prone creatures, and that we must embrace uncertainty. Further, despite their limitations, our science education efforts frequently instruct learners to think of uncertainty as an appropriate epistemic attitude to hold in science. This fact, combined with the oft-made claim that statistics employs ideas about probability in order to quantify uncertainty, requires from ASA II(note) a factually-based justification for its claim that many scientists who employ tests of significance do so in a quest for certainty.
Under the heading, “Has the American Statistical Association Gone Post-Modern?”, the legal scholar, Nathan Schachtman, recently stated:
The ASA may claim to be agnostic in the face of the contradictory recommendations, but there is one thing we know for sure: over-reaching litigants and their expert witnesses will exploit the real or apparent chaos in the ASA’s approach. The lack of coherent, consistent guidance will launch a thousand litigation ships, with no epistemic compass.(‘Schachtman’s Law’, March 24, 2019)
I suggest that, with appropriate adjustment, the same can fairly be said about researchers and statisticians, who might look to ASA II(note) as an informative guide to a better understanding of tests of significance, and the many misconceptions about them that need to be corrected.
References
Haig, B. D. (2019). Stats: Don’t retire significance testing. Nature, 569, 487.
Mayo, D. G. (2019). The 2019 ASA Guide to P-values and Statistical Significance: Don’t Say What You Don’t Mean (Some Recommendations)(ii),blog post on Error Statistics Philosophy Blog, June 17, 2019.
Mayo, D. G. (2018). Statistical inference as severe testing: How to get beyond the statistics wars. New York, NY: Cambridge University Press.
Wasserstein, R. L., & Lazar, N. A. (2016). The ASA’s statement on p-values: Context, process, and purpose. The American Statistician, 70, 129-133.
Wasserstein, R. L., Schirm. A. L., & Lazar, N. A. (2019). Editorial: Moving to a world beyond “p<0.05”. The American Statistician, 73, S1, 1-19.
In a July 19, 2019 post I discussed The New England Journal of Medicine’s response to Wasserstein’s (2019) call for journals to change their guidelines in reaction to the “abandon significance” drive. The NEJM said “no thanks” [A]. However confidence intervals CIs got hurt in the mix. In this reblog, I kept the reference to “ASA II” with a note, because that best conveys the context of the discussion at the time. Switching it to WSL (2019) just didn’t read right. I invite your comments.
The New England Journal of Medicine NEJM announced new guidelines for authors for statistical reporting yesterday. The ASA describes the change as “in response to the ASA Statement on P-values and Statistical Significance and subsequent The American Statistician* special issue on statistical inference” (ASA I and II,(note) in my abbreviation). If so, it seems to have backfired. I don’t know all the differences in the new guidelines, but those explicitly noted appear to me to move in the reverse direction from where the ASA I and II(note) guidelines were heading.
The most notable point is that the NEJM highlights the need for error control, especially for constraining the Type I error probability, and pays a lot of attention to adjusting P-values for multiple testing and post hoc subgroups. ASA I included an important principle (#4) that P-values are altered and may be invalidated by multiple testing, but they do not call for adjustments for multiplicity, nor do I find a discussion of Type I or II error probabilities in the ASA documents. NEJM gives strict requirements for controlling family-wise error rate or false discovery rates (understood as the Benjamini and Hochberg frequentist adjustments). They do not go along with the ASA II(note) call for ousting thresholds, ending the use of the words “significance/significant”, or banning “p ≤ 0.05”. In the associated article, we read:
“Clinicians and regulatory agencies must make decisions about which treatment to use or to allow to be marketed, and P values interpreted by reliably calculated thresholds subjected to appropriate adjustments have a role in those decisions”.
When it comes to confidence intervals, the recommendations of ASA II(note), to the extent they were influential on the NEJM, seem to have had the opposite effect to what was intended–or is this really what they wanted?
Significance levels and P-values, in other words, are terms to be reserved for contexts in which their error statistical meaning is legitimate. This is a key strong point of the NEJM guidelines. Confidence levels, for the NEJM, lose their error statistical or “coverage probability” meaning, unless they follow the adjustments that legitimate P-values call for. But they must be accompanied by a sign that warns the reader the intervals were not adjusted for multiple testing and thus “the inferences drawn may not be reproducible.” The P-value, but not the confidence interval, remains an inferential tool with control of error probabilities. Now CIs are inversions of tests, and strictly speaking should also have error control. Authors may be allowed to forfeit this, but then CIs can’t replace significance tests and their use may even (inadvertently, perhaps) signal lack of error control. (In my view, that is not a good thing.) Here are some excerpts:
For all studies:
Significance tests should be accompanied by confidence intervals for estimated effect sizes, measures of association, or other parameters of interest. The confidence intervals should be adjusted to match any adjustment made to significance levels in the corresponding test.
For clinical trials:
Original and final protocols and statistical analysis plans (SAPs) should be submitted along with the manuscript, as well as a table of amendments made to the protocol and SAP indicating the date of the change and its content.
The analyses of the primary outcome in manuscripts reporting results of clinical trials should match the analyses prespecified in the original protocol, except in unusual circumstances. Analyses that do not conform to the protocol should be justified in the Methods section of the manuscript. …
When comparing outcomes in two or more groups in confirmatory analyses, investigators should use the testing procedures specified in the protocol and SAP to control overall type I error — for example, Bonferroni adjustments or prespecified hierarchical procedures. P values adjusted for multiplicity should be reported when appropriate and labeled as such in the manuscript. In hierarchical testing procedures, P values should be reported only until the last comparison for which the P value was statistically significant. P values for the first nonsignificant comparison and for all comparisons thereafter should not be reported. For prespecified exploratory analyses, investigators should use methods for controlling false discovery rate described in the SAP — for example, Benjamini–Hochberg procedures.
When no method to adjust for multiplicity of inferences or controlling false discovery rate was specified in the protocol or SAP of a clinical trial, the report of all secondary and exploratory endpoints should be limited to point estimates of treatment effects with 95% confidence intervals. In such cases, the Methods section should note that the widths of the intervals have not been adjusted for multiplicity and that the inferences drawn may not be reproducible. No P values should be reported for these analyses.
As noted earlier, since P-values would be invalidated in such cases, it’s entirely right not to give them. CIs are permitted, yes, but are required to sport an alert warning that, even though multiple testing was done, the intervals were not adjusted for this and therefore “the inferences drawn may not be reproducible.” In short their coverage probability justification goes by the board.
I wonder if practitioners can opt out of this weakening of CIs, and declare in advance that they are members of a subset of CI users who will only report confidence levels with a valid error statistical meaning, dual to statistical hypothesis tests.
The NEJM guidelines continue:
…When the SAP prespecifies an analysis of certain subgroups, that analysis should conform to the method described in the SAP. If the study team believes a post hoc analysis of subgroups is important, the rationale for conducting that analysis should be stated. Post hoc analyses should be clearly labeled as post hoc in the manuscript.
Forest plots are often used to present results from an analysis of the consistency of a treatment effect across subgroups of factors of interest. …A list of P values for treatment by subgroup interactions is subject to the problems of multiplicity and has limited value for inference. Therefore, in most cases, no P values for interaction should be provided in the forest plots.
If significance tests of safety outcomes (when not primary outcomes) are reported along with the treatment-specific estimates, no adjustment for multiplicity is necessary. Because information contained in the safety endpoints may signal problems within specific organ classes, the editors believe that the type I error rates larger than 0.05 are acceptable. Editors may request that P values be reported for comparisons of the frequency of adverse events among treatment groups, regardless of whether such comparisons were prespecified in the SAP.
When possible, the editors prefer that absolute event counts or rates be reported before relative risks or hazard ratios. The goal is to provide the reader with both the actual event frequency and the relative frequency. Odds ratios should be avoided, as they may overestimate the relative risks in many settings and be misinterpreted.
Authors should provide a flow diagram in CONSORT format. The editors also encourage authors to submit all the relevant information included in the CONSORT checklist. …The CONSORT statement, checklist, and flow diagram are available on the CONSORT
Detailed instructions to ensure that observational studies retain control of error rates are given.
In the associated article:
P values indicate how incompatible the observed data may be with a null hypothesis; “P<0.05” implies that a treatment effect or exposure association larger than that observed would occur less than 5% of the time under a null hypothesis of no effect or association and assuming no confounding. Concluding that the null hypothesis is false when in fact it is true (a type I error in statistical terms) has a likelihood of less than 5%. [i]…
The use of P values to summarize evidence in a study requires, on the one hand, thresholds that have a strong theoretical and empirical justification and, on the other hand, proper attention to the error that can result from uncritical interpretation of multiple inferences.5 This inflation due to multiple comparisons can also occur when comparisons have been conducted by investigators but are not reported in a manuscript. A large array of methods to adjust for multiple comparisons is available and can be used to control the type I error probability in an analysis when specified in the design of a study.6,7 Finally, the notion that a treatment is effective for a particular outcome if P<0.05 and ineffective if that threshold is not reached is a reductionist view of medicine that does not always reflect reality. [ii]
… A well-designed randomized or observational study will have a primary hypothesis and a prespecified method of analysis, and the significance level from that analysis is a reliable indicator of the extent to which the observed data contradict a null hypothesis of no association between an intervention or an exposure and a response. Clinicians and regulatory agencies must make decisions about which treatment to use or to allow to be marketed, and P values interpreted by reliably calculated thresholds subjected to appropriate adjustments have a role in those decisions.
Finally, the current guidelines are limited to studies with a traditional frequentist design and analysis, since that matches the large majority of manuscripts submitted to the Journal. We do not mean to imply that these are the only acceptable designs and analyses. The Journal has published many studies with Bayesian designs and analyses8-10 and expects to see more such trials in the future. When appropriate, our guidelines will be expanded to include best practices for reporting trials with Bayesian and other designs.
What do you think?
The author guidelines:
https://www.nejm.org/author-center/new-manuscripts
The associated article:
https://www.nejm.org/doi/full/10.1056/NEJMe1906559
*I meant to thank Nathan Schachtman for notifying me and sending links; also Stuart Hurlbert.
[i] It would be better, it seems to me, if the term “likelihood” was used only for its technical meaning in a document like this.
[ii] I don’t see it as a matter of “reductionism” but simply a matter of the properties of the test and the discrepancies of interest in the context at hand.
[A] A self-published book on this episode, by Donald Macnaughton, came out in 2021: The War on Statistical Significance: The American Statistician vs. the New England Journal of Medicine.
.
On June 1, 2019, I posted portions of an article [i],“There is Still a Place for Significance Testing in Clinical Trials,” in Clinical Trials responding to the 2019 call to abandon significance. I reblog it here. While very short, it effectively responds to the 2019 movement (by some) to abandon the concept of statistical significance [ii]. I have recently been involved in researching drug trials for a condition of a family member, and I can say that I’m extremely grateful that they are still reporting error statistical assessments of new treatments, and using carefully designed statistical significance tests with thresholds. Without them, I think we’d be lost in a sea of potential treatments and clinical trials. Please share any of your own experiences in the comments. The emphasis in this excerpt is mine:
Much hand-wringing has been stimulated by the reflection that reports of clinical studies often misinterpret and misrepresent the findings of the statistical analyses. Recent proposals to address these concerns have included abandoning p-values and much of the traditional classical approach to statistical inference, or dropping the concept of statistical significance while still allowing some place for p-values. How should we in the clinical trials community respond to these concerns? Responses may vary from bemusement, pity for our colleagues working in the wilderness outside the relatively protected environment of clinical trials, to unease about the implications for those of us engaged in clinical trials….
However, we should not be shy about asserting the unique role that clinical trials play in scientific research. A clinical trial is a much safer context within which to carry out a statistical test than most other settings. Properly designed and executed clinical trials have opportunities and safeguards that other types of research do not typically possess, such as protocolisation of study design; scientific review prior to commencement; prospective data collection; trial registration; specification of outcomes of interest including, importantly, a primary outcome; and others. For randomised trials, there is even more protection of scientific validity provided by the randomisation of the interventions being compared. It would be a mistake to allow the tail to wag the dog by being overly influenced by flawed statistical inferences that commonly occur in less carefully planned settings….
Furthermore, the research question addressed by clinical trials (comparing alternative strategies) fits well with such an approach and the corresponding decision-making settings (e.g. regulatory agencies, data and safety monitoring committees and clinical guideline bodies) are often ones within which statistical experts are available to guide interpretation. The carefully designed clinical trial based on a traditional statistical testing framework has served as the benchmark for many decades. It enjoys broad support in both the academic and policy communities. There is no competing paradigm that has to date achieved such broad support. The proposals for abandoning p-values altogether often suggest adopting the exclusive use of Bayesian methods. For these proposals to be convincing, it is essential their presumed superior attributes be demonstrated without sacrificing the clear merits of the traditional framework. Many of us have dabbled with Bayesian approaches and find them to be useful for certain aspects of clinical trial design and analysis, but still tend to default to the conventional approach notwithstanding its limitations. While attractive in principle, the reality of regularly using Bayesian approaches on important clinical trials has been substantially less appealing – hence their lack of widespread uptake.
The issues that have led to the criticisms of conventional statistical testing are of much greater concern where statistical inferences are derived from observational data. … Even when the study is appropriately designed, there is also a common converse misinterpretation of statistical tests whereby the investigator incorrectly infers and reports that a non-significant finding conclusively demonstrates no effect. However, it is important to recognise that an appropriately designed and powered clinical trial enables the investigators to potentially conclude there is ‘no meaningful effect’ for the principal analysis.[iii] More generally, these problems are largely due to the fact that many individuals who perform statistical analyses are not sufficiently trained in statistics. It is naive to suggest that banning statistical testing and replacing it with greater use of confidence intervals, or Bayesian methods, or whatever, will resolve any of these widespread interpretive problems. Even the more modest proposal of dropping the concept of ‘statistical significance’ when conducting statistical tests could make things worse. By removing the prespecified significance level, typically 5%, interpretation could become completely arbitrary. It will also not stop data-dredging, selective reporting, or the numerous other ways in which data analytic strategies can result in grossly misleading conclusions.
These considerations notwithstanding, the field of clinical trials is in rapid evolution and it is entirely possible and appropriate that the statistical framework used for their evaluation must also change. However, such evolution should emerge from careful methodological research and open-minded, self-critical enquiry. We earnestly hope that Clinical Trials will continue to be seen as a natural academic home for exploration and debate about alternative statistical frameworks for making inferences from clinical trials. The Editors welcome articles that evaluate or debate the merits of such alternative paradigms along with the conventional one within the context of clinical trials. Especially welcome are exemplar trial articles and those which are illustrated using practical examples from clinical trials that permit a realistic evaluation of the strengths and weaknesses of the approach.
You can read the full article here.
Please share your comments.
[i] Jonathan A Cook, Dean A Fergusson, Ian Ford , Mithat Gonen, Jonathan Kimmelman, Edward L Korn and Colin B Begg (2019). “There is still a place for significance testing in clinical trials”, Clinical Trials 2019, Vol. 16(3) 223–224.
[ii] PBack in 2019, was trying to find an apt acronym. I played with the idea of calling those driven to Stop Error Statistical Tests “Obsessed”. I thank Nathan Schachtman for sending me the article.
[iii] It’s disappointing how many critics of tests seem unaware of this simple power analysis point, and how it avoids egregious fallacies of non-rejection, or moderate P-value. It precisely follows simple significance test reasoning. The severity account that I favor gives a more custom-tailored approach that is sensitive to the actual outcome. (See, for example, Excursion 5 of Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (2018, CUP).
[iv] Bayes factors, like other comparative measures, are not “tests”, and do not falsify (even statistically). They can only say one hypothesis or model is better than a selected other hypothesis or model, based on some^ selected criteria. They can both (all) be improbable, unlikely, or terribly tested. One can always add a “falsification rule”, but it must be shown that the resulting test avoids frequently passing/failing claims erroneously.
^The Anti-Testers would have to say “arbitrary criterion”, to be consistent with their considering any P-value “arbitrary”, and denying that a statistically significant difference, reaching any P-value, indicates a genuine difference from a reference hypothesis.
.
It was 3 months before I decided to write a blogpost in response to Wasserstein, Schirm and Lazar (2019)’s editorial in The American Statistician in which they recommend that the concept of “statistical significance” be abandoned, hereafter, WSL 2019. (I titled it “Don’t Say What You don’t Mean”.) In that June 17, 2019 blogpost, pasted below, I proposed 3 “friendly amendments” to the language of that document. (There are 97 comments on that post!) The problem is that WSL 2019 presents several of the 6 principles from ASA I (the 2016 ASA statement on Statistical Significance) in a far stronger fashion so as to be inconsistent or at least in tension with some of them. I didn’t think they really meant what they said. I discussed these amendments with Ron Wasserstein, Executive Director of the ASA at the time. Had these friendly amendments been carried out, the document would not have caused as much of a problem, and people might focus more on the positive recommendations it includes about scientific integrity. The proposed ban on a key concept of statistics would still be problematic, resulting in the 2019 ASA President’s Task Force, but it would have helped the document. At the time, it was still not known whether WSL 2019 was intended as a continuation of the 2016 ASA policy document [ASA I]. That explains why I first referred to WSL 2019 in this blogpost as ASA II. Once it was revealed that it was not official policy at all (many months later), but only the recommendations of the 3 authors, I placed a “note” after each mention of ASA II. But given it caused sufficient confusion as to result in the then ASA president (Karen Kafadar) appointing an ASA Task Force on Statistical Significance and Replicability in 2019 (see here and here), and later, a disclaimer by the authors, in this reblog I refer to it as WSL 2019. You can search this blog for other posts on the 2019 Task Force: their report is here, and the disclaimer here.
“The 2019 Guide to P-values and Statistical significance: Don’t Say What You don’t Mean” (June 17, 2019)
Some have asked me why I haven’t blogged on the recent follow-up to the ASA Statement on P-Values and Statistical Significance (Wasserstein and Lazar 2016)–hereafter, ASA I. They’re referring to the editorial by Wasserstein, R., Schirm, A. and Lazar, N. (2019)–hereafter, [WSL 2019]–opening a special on-line issue of over 40 contributions responding to the call to describe “a world beyond P < 0.05”.[1] Am I falling down on the job? Not really. All of the issues are thoroughly visited in my Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars, SIST (2018, CUP). I invite interested readers to join me on the statistical cruise therein.[2] As the [WSL 2019]. authors observe: “At times in this editorial and the papers you’ll hear deep dissonance, the echoes of ‘statistics wars’ still simmering today (Mayo 2018)”. True, and reluctance to reopen old wounds has only allowed them to fester. However, I will admit, that when new attempts at reforms are put forward, a philosopher of science who has written on the statistics wars ought to weigh in on the specific prescriptions/proscriptions, especially when a jumble of fuzzy conceptual issues are interwoven through a cacophony of competing reforms. (My published comment on ASA I, “Don’t Throw Out the Error Control Baby With the Bad Statistics Bathwater” is here.)
So I should say something. But the task is delicate. And painful. Very. I should start by asking: What is it (i.e., what is it actually saying)? Then I can offer some constructive suggestions.
The Invitation to Broader Consideration and Debate
The papers in this issue propose many new ideas, ideas that in our determination as editors merited publication to enable broader consideration and debate. The ideas in this editorial are likewise open to debate. ([WSL 2019] p. 1)
The questions around reform need consideration and debate. (p. 9)
Excellent! A broad, open, critical debate is sorely needed. Still, we can only debate something when there is a degree of clarity as to what “it” is. I will be very happy to post reader’s meanderings on [WSL 2019]. (~1000 words) if you send them to me.
My focus here is just on the intended positions of the ASA [or WSL 2019], not the summaries of articles. This comprises around the first 10 pages. Even from just the first few pages the reader is met with some noteworthy declarations:
Don’t conclude anything about scientific or practical importance based on statistical significance (or lack thereof). (p. 1)
No p-value can reveal the plausibility, presence, truth, or importance of an association or effect. (p.2)
A declaration of statistical significance is the antithesis of thoughtfulness. (p. 4)
Whether a p-value passes any arbitrary threshold should not be considered at all when deciding which results to present or highlight. (p. 2, my emphasis)
It is time to stop using the term “statistically significant” entirely. Nor should variants such as “significantly different,” “p < 0.05,” and “nonsignificant” survive. (p.2)
“Statistically significant”– don’t say it and don’t use it. (p. 2)
(Wow!)
I am very sympathetic with the concerns about rigid cut-offs, and fallacies of moving from statistical significance to substantive scientific claims. I feel as if I’ve just written a whole book on it! I say, on p. 10 of SIST:
In formal statistical testing, the crude dichotomy of “pass/fail” or “significant or not” will scarcely do. We must determine the magnitudes (and directions) of any statistical discrepancies warranted, and the limits to any substantive claims you may be entitled to infer from the statistical ones.
Since [WSL (2019)] will still use P-values, you’re bound to wonder why a user wouldn’t just report “the difference is statistically significant at the P-value attained”. (The probability of observing even larger differences, under the assumption of chance variability alone is p.) Confidence intervals (CIs) are already routinely given alongside P-values. So there is clearly more to the current movement than meets the eye. But for now I’m just trying to decipher what the ASA position is.
What’s the Relationship Between ASA I and [WSL 2019]?
I assume, for this post, that [WSL 2019] is intended to be an extension of ASA I. In that case, it would subsume the 6 principles of ASA I. There is evidence for this. For one thing, it begins by sketching a “sampling” of “don’ts” from ASA I, for those who are new to the debate. Secondly, it recommends that ASA I be widely disseminated. But some Principles (1, 4) are apparently missing[3], and others are rephrased in ways that alter the initial meanings. Do they really mean these declarations as written? Let us try to take them at their word.
But right away we are struck with a conflict with Principle 1 of ASA I–which happens to be the only positive principle given. (See Note 5 for the six Principles of ASA I.)
Principle 1. P-values can indicate how incompatible the data are with a specified statistical model.
A p-value provides one approach to summarizing the incompatibility between a particular set of data and a proposed model for the data. The most common context is a model, constructed under a set of assumptions, together with a so-called “null hypothesis.” Often the null hypothesis postulates the absence of an effect, such as no difference between two groups, or the absence of a relationship between a factor and an outcome. The smaller the p-value, the greater the statistical incompatibility of the data with the null hypothesis, if the underlying assumptions used to calculate the p-value hold. This incompatibility can be interpreted as casting doubt on or providing evidence against the null hypothesis or the underlying assumptions.” (ASA I p. 131)
However, an indication of how incompatible data are with a claim of the absence of a relationship between a factor and an outcome would be an indication of the presence of the relationship; and providing evidence against a claim of no difference between two groups would often be of scientific or practical importance.
So, Principle 1 (from ASA I) doesn’t appear to square with the first bulleted item I listed (from [WSL 2019]:
(1) “Don’t conclude anything about scientific or practical importance based on statistical significance (or lack thereof)” [WSL 2019].
Either modify (1) or erase Principle 1. But if you erase all thresholds for finding incompatibility (whether using P-values or other measures), there are no tests, and no falsifications, even of the statistical kind.
My understanding (from Ron Wasserstein) is that this bullet is intended to correspond to Principle 5 in ASA I – that P-values do not give population effect sizes. But it is now saying something stronger (at least to my ears and to everyone else I’ve asked). Do the authors mean to be saying that nothing (of scientific or practical importance) can be learned from statistical significance tests? I think not.
So, my first recommendation is:
Replace (1) with:
“Don’t conclude anything about the scientific or practical importance of the (population) effect size based only on statistical significance (or lack thereof).”
Either that, or simply stick to Principle 5 from ASA I : “A p-value, or statistical significance[4], does not measure the size of an effect or the importance of a result.” (p. 132) This statement is, strictly speaking, a tautology, true by the definitions of terms: probability isn’t itself a measure of the size of a (population) effect. However, you can use statistically significant differences to infer what the data indicate about the size of the (population) effect.[4]
My second friendly amendment concerns the second bulleted item:
(2) No p-value can reveal the plausibility, presence, truth, or importance of an association or effect. (p. 2)
Focus just on “presence”. From this assertion it would seem to follow that no P-values[5], however small, even from well-controlled trials, can reveal the presence of an association or effect–and that is too strong. Again, we get a conflict with Principle 1 from ASA I. But I’m guessing, for now, the authors do not intend to say this. If you don’t mean it, don’t say it.
So, my second recommendation is to replace (2) with:
“No p-value by itself can reveal the plausibility, presence, truth, or importance of an association or effect.
Without this friendly amendment, [WSL 2019] is at loggerheads with ASA I, and they should not be advocating those 6 principles without changing either or both. Without this or a similar modification, moreover, the ability of any other statistical quantity or evidential measure is likewise unable to reveal these things. Or so many would argue. These modest revisions might prevent some readers stopping after the first few pages, and that would be a shame, as they would miss the many right-headed insights about linking statistical and scientific inference.
This leads to my third bulleted item from [WSL 2019]:
(3) A declaration of statistical significance is the antithesis of thoughtfulness… it ignores what previous studies have contributed to our knowledge. (p. 4)
Surely the authors do not mean to say that anyone who asserts the observed difference is statistically significant at level p has her hands tied and invariably ignores all previous studies, background information and theories in planning and reaching conclusions, decisions, proposed solutions to problems. I’m totally on board with the importance of backgrounds, and multiple steps relating data to scientific claims and problems. Here’s what I say in SIST:
The error statistician begins with a substantive problem or question. She jumps in and out of piecemeal statistical tests both formal and quasi-formal.The pieces are integrated in building up arguments from coincidence, informing background theory, self-correcting via blatant deceptions, in an iterative movement. The inference is qualified by using error probabilities to determine not “ how probable,” but rather, “ how well-probed” claims are, and what has been poorly probed. (SIST, p. 162)
But good inquiry is piecemeal: There is no reason to suppose one does everything at once in inquiry, and it seems clear from the [WSL 2019] guide that the authors agree. Since I don’t think they literally mean (3), why say it?
Practitioners who use these methods in medicine and elsewhere have detailed protocols for how background knowledge is employed in designing, running, and interpreting tests. When medical researchers specify primary outcomes, for just one example, it’s very explicitly with due regard for the mechanism of drug action. It’s intended as the most direct way to pick up on the drug’s mechanism. Finding incompatibility using P-values, inherits the meaning already attached to a sensible test hypothesis. That valid P-values require context is presupposed by the very important Principle 4 of ASA I (see note (3).
As lawyer Nathan Schachtman observes, in a recent conversation on [WSL (2019)].
By the time a phase III clinical trial is being reviewed for approval, there is a mountain of data on pharmacology, pharmacokinetics, mechanism, target organ, etc. If Wasserstein wants to suggest that there are some people who misuse or misinterpret p-values, fine. The principle of charity requires that we give a more sympathetic reading to the broad field of users of statistical significance testing. (Schachtman 2019)
Now it is possible the authors are saying a reported P-value can never be thoughtful because thoughtfulness requires that a statistical measure, at any stage of probing, incorporate everything we know (SIST dubs this “big picture” inference.) Do we want that? Or maybe (3) is their way of saying a statistical measure must incorporate background beliefs in the manner of Bayesian degree-of-belief (?) priors. Many would beg to differ, including some leading Bayesians. Andrew Gelman (2012) has suggested that ‘Bayesians Want Everybody Else to be Non-Bayesian’:
Bayesian inference proceeds by taking the likelihoods from different data sources and then combining them with a prior (or, more generally, a hierarchical model). The likelihood is key. . . No funny stuff, no posterior distributions, just the likelihood. . . I don’t want everybody coming to me with their posterior distribution – I’d just have to divide away their prior distributions before getting to my own analysis. (ibid., p. 54)
So, my third recommendation is to replace (3) with (something like):
“failing to report anything beyond a declaration of statistical significance is the antithesis of thoughtfulness.”
There’s much else that bears critical analysis and debate in [WSL (2019)]; I’ll come back to it. I hope to hear from the authors of [WSL (2019)] about my very slight, constructive amendments (to avoid a conflict with Principle 1).
Meanwhile, I fear we will see court cases piling up denying that anyone can be found culpable for abusing p-values and significance tests, since the ASA declared that all p-values are arbitrary, and whether predesignated thresholds are honored or breached should not be considered at all. (This was already happening based on ASA I.)[6]
Please share your thoughts and any errors in the comments, I will indicate later drafts of this post with (i), (ii),…Do send me other articles you find discussing this. Version (ii) of this post begins a list:
Nathan Schachtman (2019): Has the ASA Gone Post-Modern?
Cook et al.,(2019) There is Still Place for Significance Testing in Clinical Trials
NEJM Manuscript & Statistical Guidelines 2019Harrington, New Guidelines for Statistical Reporting in the Journal NEJM 2019
References:
Gelman, A. (2012) “Ethics and the Statistical Use of Prior Information”. http://www.stat.columbia.edu/~gelman/research/published/ChanceEthics5.pdf
Mayo, D. (2016). “Don’t Throw out the Error Control Baby with the Bad Statistics Bathwater: A Commentary” on R. Wasserstein and N. Lazar: “The ASA’s Statement on P-values: Context, Process, and Purpose”, The American Statistician 70(2).
Mayo, D. (2018). Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars. Cambridge: Cambridge University Press.
Schachtman, N. (2019). (private communication)
Wasserstein, R. and Lazar, N. (2016). “The ASA’s Statement on P-values: Context, Process and Purpose”, (and supplemental materials), The American Statistician 70(2), 129–33. (ASA I)
Wasserstein, R., Schirm, A. and Lazar, N. (2019) Editorial: “Moving to a World Beyond ‘p < 0.05’”, The American Statistician 73(S1): 1-19.
NOTES
[1] I gave an invited paper at the conference (“A world Beyond…”) out of which the idea for this volume grew. I was in a session with a few other exiles to describe the contexts where statistical significance tests are of value. I was too much involved in completing my book to write up my paper for this volume, nor did others in our small group. Links are here to: my slides and Yoav Benjamini’s slides. I did post notes to journalists on the Amrhein article here.
[2] Excerpts and mementos from SIST are here.
..
Five years ago on this day, a news correspondent at NPR, Richard Harris, published this article, the same day as Wasserstein et al., (2019). Moving to a world beyond “p < 0.05”. TAS, and Amrhein et al., (2019), Comment: Retire statistical significance. Nature. I was one of several people Harris interviewed for his article. He starts by talking of flip-flops regarding the healthfulness of eggs.
Statisticians say it may not be wise to put all your eggs in the significance basket.
A recent study that questioned the healthfulness of eggs raised a perpetual question: Why do studies, as has been the case with health research involving eggs, so often flip-flop from one answer to another?
The truth isn’t changing all the time. But one reason for the fluctuations is that scientists have a hard time handling the uncertainty that’s inherent in all studies. There’s a new push to address this shortcoming in a widely used – and abused – scientific method.
Scientists and statisticians are putting forth a bold idea: Ban the very concept of “statistical significance.”
We hear that phrase all the time in relation to scientific studies. Critics, who are numerous, say that declaring a result to be statistically significant or not essentially forces complicated questions to be answered as true or false.
“The world is much more uncertain than that,” says Nicole Lazar, a professor of statistics at the University of Georgia. She is involved in the latest push to ban the use of the term “statistical significance.”
An entire issue of the journal The American Statistician is devoted to this question, with 43 articles and a 17,500-word editorial that Lazar co-authored.
Some of the scientists involved in that effort also wrote a more digestible commentary that appears in Thursday’s issue of Nature. More than 850 scientists and statisticians told the Naturecommentary authors they want to endorse this idea.
In the early 20th century, the father of statistics, R.A. Fisher, developed a test of significance. It involves a variable called the p-value, that he intended to be a guide for judging results.
Over the years, scientists have warped that idea beyond all recognition. They’ve created an arbitrary threshold for the p-value, typically 0.05, and they use that to declare whether a scientific result is significant or not.
This shortcut often determines whether studies get published or not, whether scientists get promoted and who gets grant funding.
“It’s really gotten stretched all out of proportion,” says Ron Wasserstein, the executive director of the American Statistical Association. He’s been advocating this change for years and he’s not alone.
“Failure to make these changes are really now starting to have a sustained negative impact on how science is conducted,” he says. “It’s time to start making the changes. It’s time to move on.”
There are many downsides to this, he says. One is that scientists have been known to massage their data to make their results hit this magic threshold. Arguably worse, scientists often find that they can’t publish their interesting (if somewhat ambiguous) results if they aren’t statistically significant. But that information is actually still useful, and advocates say it’s wasteful simply to throw it away.
There are some prominent voices in the world of statistics who reject the call to abolish the term “statistical significance.”
“Nature ought to invite somebody to bring out the weakness and dangers of some of these recommendations,” says Deborah Mayo, a philosopher of science at Virginia Tech.
“Banning the word ‘significance’ may well free researchers from being held accountable when they downplay negative results” and otherwise manipulate their findings, she notes.
“We should be very wary of giving up on something that allows us to hold researchers accountable.”
Her desire to keep “statistical significance” is deeply embedded.
Scientists – like the rest of us – are far more likely to believe that a result is true if it’s statistically significant. Still, Blake McShane, a statistician at the Kellogg School of Management at Northwestern University, says we put far too much faith in the concept.
“All statistics naturally bounce around quite a lot from study to study to study,” McShane says. That’s because there’s lots of variation from one group of people to another, and also because subtle differences in approach can lead to different conclusions.
So, he says, we shouldn’t be at all surprised if a result that’s statistically significant in one study doesn’t meet that threshold in the next.
McShane, who co-authored the Nature commentary, says this phenomenon also partly explains why studies done in one lab are frequently not reproduced in other labs. This is sometimes referred to as the “reproducibility crisis,” when in fact, the apparent conflict between studies may be an artifact of relying on the concept of statistical significance.
But despite these flaws, science embraces statistical significance because it’s a shortcut that provides at least some insight into the strength of an observation.
Journals are reluctant to abandon the concept. “Nature is not seeking to change how it considers statistical analysis in evaluation of papers at this time,” the journal noted in an editorial that accompanies the commentary.
Veronique Kiermer, publisher and executive editor of the PLOS journals, bemoans the overreliance on statistical significance, but says her journals don’t have the leverage to force a change.
“The problem is that the practice is so engrained in the research community,” she writes in an email, “that change needs to start there, when hypotheses are formulated, experiments designed and analyzed, and when researchers decide whether to write up and publish their work.”
One problem is what would scientists use instead of statistical significance. The advocates for change say the community can still use the p-value test, but as part of a broader approach to measuring uncertainty.
A bit more humility would also be in order, these advocates for change say.
“Uncertainty is present always,” Wasserstein says. “That’s part of science. So rather than trying to dance around it, we [should] accept it.”
That goes a bit against human nature. After all, we want answers, not more questions.
But McShane says arriving at a yes/no answer about whether to eat eggs is too simplistic. If we step beyond that, we can ask more important questions. How big is the risk? How likely is it to be real? What are the costs and benefits to an individual?
Lazar has an even more extreme view. She says when she hears about individual studies, like the egg one, her statistical intuition leads her to shrug: “I don’t even pay attention to it anymore.”
You can reach NPR Science Correspondent Richard Harris at rharris@npr.org.View this story on npr.org
Use the comments to share your first thoughts and reflections on consequences, positive and negative, of the movement to abandon significance over the past 5 years. I may collect second thoughts for a separate blogpost.
..
In my last post, I sketched some first remarks I would have made had I been able to travel to London to fulfill my invitation to speak at a Royal Society conference, March 4 and 5, 2024, on “the promises and pitfalls of preregistration.” This is a continuation. It’s a welcome consequence of today’s statistical crisis of replication that some social sciences are taking a page from medical trials and calling for preregistration of sampling protocols and full reporting. In 2018, Brian Nosek and others wrote of the “Preregistration Revolution”, as part of open science initiatives.
The main sources of failed replication are not mysterious: data-dredging, multiple testing, outcome-switching, cherry-picking, optional stopping and a host of related “biasing selection effects” can practically guarantee an impressive-looking effect, even if it is spurious. The inferred effect, H, agrees with the data but the test H has passed lacks stringency or severity. However, as I noted, I would be keen to distinguish right off pejorative from non-pejorative data dredging because one of the main arguments I hear against preregistration sets sail by presenting us with cases where trenchant searching is a mark of good science and the route to discovery. One need not always get new data to severely rule out errors of relevance (although what counts as “new” is often unclear). Think of trying and trying again to find a DNA match, the source of a faster than speed of light anomaly,[1] or the location of your keys (always the last place you look!). Calls for preregistration to block data-dredging—where they matter—are calls for severe testing, notably where the goal is avoiding being fooled by random chance. Since statistical significance tests are key tools for this end, preregistration (or, in some cases, compensation by error probability adjustments) is valued to avoid systematically misleading results (violating David Cox’s weak repeated sampling principle). In medicine, and other fields, perverse incentives to generate and present data so as to selectively advantage desirable inferences have led to elaborate protocols “on best practice in trial reporting, which are endorsed by 585 academic journals” (Goldacre et al., 2019, p. 2), as well as methods for p-value adjustments (Benjamini, 2020).
Now I enter the second portion of my thoughts on the topic.
However, there are rivals to error statistical methods that hold principles of evidence where evidence is not altered by altered error probabilities. On one such principle, all the evidence is via likelihood ratios (LR) of hypotheses:
Pr(x0;H1)/Pr(x0;H0)
where Pr(x0;H1) is the probability of x0 under hypothesis H1, and Pr(x0;H0) the probability of x0 under hypothesis H0. This is often called the (strong) likelihood principle.[2]
There is a lot of confusion about likelihoods. With likelihoods, the data x0 are fixed, the hypotheses vary. Often, “likelihood” is used interchangeably with “probability”, but this leads to trouble when we’re keen to talk about the formal concept of likelihood. A hypothesis that perfectly fits the data has a likelihood equal to 1; so, since many rival hypotheses can fit x0, it’s clear that likelihoods do not obey the probability calculus.[3]
Of pertinence to the issue of preregistration, if all the evidence is in the likelihood ratio, then error probabilities drop out once the data are in hand. As subjective Bayesian, Dennis Lindley observed long ago:
Sampling distributions, significance levels, power, all depend on something more [than the likelihood function]–something that is irrelevant in Bayesian inference—namely the sample space. (Lindley 1971, 436)
What he means is that once the data are observed, other outcomes that could have occurred are irrelevant. Yet error probabilities consider outcomes other than x0. The error statistician cannot assess the evidence in x0 without knowing how the method of sampling would have behaved under different outcomes (i.e., the sampling distribution). Ignoring this error probabilistic behavior, she knows she can be fooled. Finding the probability of x0 under hypothesis H0 is low, one can easily construct an alternative H1 that fits the data swimmingly in order to get a high likelihood ratio in favor of H1. As statistician George Barnard puts it, “there always is such a rival hypothesis viz., that things just had to turn out the way they actually did” (Barnard 1972, p. 129). The probability of finding some better fitting alternative or other can be high, or guaranteed, even when H0 correctly describes the data generation. This probability enters in assessing how well an inferred claim has been severely probed.
Preregistration and controversies about error probabilities
A classic problematic example is when a researcher, failing to find a benefit in a (double-blind) randomized control trial on a medical treatment, searches the unblinded data until finding a subgroup where those treated do better than the controls in some way or other. They might search for patterns among patients with different characteristics (age, sex, employment, education, medical conditions diseases, etc.), and next try different proxy variables to use in measuring benefit (“outcome switching”). Let H1PD be the “post data” hypothesis arrived at from the subgroup search and outcome tinkering. The data x supports H1PD better than H0PD that there is 0 benefit. We know this because the subgroup has been deliberately selected such that those with the treatment do better than the untreated group by a chosen amount. The likelihood ratio or Bayes factor is
Pr(x|H1PD)/Pr(x|H0PD).
This way of proceeding has a high probability of issuing in a report of drug benefit H1 (in some subgroup or other), even if no benefit exists (i.e., even if the null or test hypothesis H0 is true). The researcher has drawn that line around the post-data subgroup just like the Texas marksman. It is even worse if the researcher reports this as a result of the original double-blind trial!
Nevertheless, if all of the evidence for a statistical inference is in the likelihood ratio, then alterations to error probabilities do not alter the import of the evidence. Thus, there is a tension between popular calls for preregistration—arguably, one of the most promising ways to boost replication—and accounts that downplay error probabilities: Bayes Factors, Bayesian posteriors, likelihood ratios.
Bayesian analysis does not base decisions on error control. Indeed, Bayesian analysis does not use sampling distributions. …As Bayesian analysis ignores counterfactual error rates, it cannot control them. (Kruschke and Liddell 2017, 13, 15)
So, the long-standing controversies about error probabilities, I would expect, would be central in a conference on preregistration in statistics. I am interested to learn if it was, and what was said.
The constructive upshot of the replication crisis–for most
(Frequentist) error statistical methods are often put on the defensive as to why control of error probabilities matters to the inference at hand. Error control, according to some statistical schools, is only of concern to ensure low error rates in some long run series of applications. Members of these schools say, we are happy to have methods with good long-run “operating properties” at the design stage, but once the data are in hand, error probabilities drop out. It should be clear from the replication crisis that what bothers us about pejorative data dredging is not about long-runs–even though they do damage the reliability of performance. It is that a poor job has been done in the case at hand in distinguishing genuine from spurious effects. Little has been done to mitigate and prevent known ways to blow up the probability of false positive results.
By and large, the replication crisis has had the constructive upshot of raising the consciousness of researchers. We have replication research and, as with this conference, a focus on preregistration and registered reports. Well-known statistical critics from psychology, Joseph Simmons, Leif Nelson, and Uri Simonsohn, place at the top of their list of requirements the need to block flexible or “optional” stopping: “Researchers often decide when to stop data collection on the basis of interim data analysis … many believe this practice exerts no more than a trivial influence on false-positive rates” (Simmons et al. 2011, p. 1361). “Contradicting this intuition” they show the probability of erroneous rejections balloons.
Consider the often discussed example of optional stopping in two-sided testing of a 0 versus a non-zero Normal mean (H0: μ = μ0 vs. H1: μ > μ1): (known σ) “[I]f an experimenter uses this procedure, then with probability 1 he will eventually reject any sharp null hypothesis, even though it be true.” (Edwards, Lindman, and Savage 1963, 239) While both Bayesians and non-Bayesians employ such adaptive or sequential trials, the error statistician must take account of, and perhaps adjust for, the stopping rule. But Edwards, Lindman, and Savage aver “the import of the sequence of n data actually observed will be exactly the same as it would be had you planned to take exactly n observations in the first place. (ibid., pp. 238-9)
The likelihood principle emphasized in Bayesian statistics implies, among other things, that the rules governing when data collection stops are irrelevant to data interpretation. (Edwards, Lindman, and Savage 1963, p. 193)
[The same stopping procedure can be used to ensure that μ0 is always excluded from the corresponding confidence interval.] Authors of the Likelihood Principle, Jim Berger and Robert Wolpert, remark: “It seems very strange that a frequentist could not analyze a given set of data…if the stopping rule is not given….Data should be able to speak for itself” (Berger and Wolpert 1988, 78).
The question that arises is this: if a method’s error probabilities do not enter in appraising evidence, it is unclear how to use registered reports of what would alter error probabilities in scrutinizing an inference. (Recall my description in the previous post of what the critical reader of a preregistered report might consider.) How would a Bayesian, for example, assuming they accept the likelihood principle do so?[4] They can, of course, report that violations of preregistration are problematic for any reported p-values (or other error probability notions, type 1 and 2 errors, confidence levels). They might go on to show, or purport to show, that this would not be problematic for their preferred account. If so, does it follow we can skip the preregistration revolution and adopt their methods? I assume the preregistration conference took advantage of the opportunity to discuss this. (I will report back once I find out.)
Severity on the meta-level
Here things become especially tricky. If one purports to show that biasing selection effects make no difference to a given methodology, it will not do to employ an analysis that has no antenna for picking up on how error probabilities are altered by biasing selection effects. Right? It wouldn’t suffice to declare, for example, that wearing likelihood principle glasses, the likelihood ratio or Bayes factor is insensitive to biasing selection effects that alter error rates. After all, the severity principle also applies on the meta-level:
Severity requirement (minimal). If an inquiry had little or no capability of unearthing the falsity of a claim inferred, then the claim is unwarranted by the data from that inquiry.
As always, the severity assessment naturally takes into account the particular error of inference that is relevant in the context of the inference or claim. Since meta-statistics should be one-level removed, one needs to remove likelihood principle glasses and ask what the effects of biasing selection effects would be on an error statistical account of evidence. This is not always done, however, and some might assume that they need not worry about predesignation if they don’t compute p-values. But they do have to worry (at least in the contexts I am identifying as pejorative selection effects).
Now error statistical assessments of rivals to error statistical methods have often been done, with varying degrees of success. They are routinely carried out by regulatory agencies in examining Bayesian adaptive or sequential trials in exploratory inquiry[5]. Bayesian trialists Ryan et al. (2020) in radiation oncology give an interesting objection to such frequentist scrutiny. While they admit “the type I error was inflated in their Bayesian adaptive designs,
The requirement of type I error control for Bayesian adaptive designs causes them to lose many of their philosophical advantages, such as compliance with the likelihood principle, and creates a design that is inherently frequentist. (Ryan et al., 2020, p. 7)
I think it is a hybrid Bayesian-frequentist assessment. Regulatory agencies perform a kind of hybrid Bayesian-frequentist computation of a type I error probability by “determining how frequently the Bayesian design incorrectly declares a treatment to be effective or superior when it is assumed that there is truly no difference” (ibid., p. 3). The idea is to simulate many thousands of trials and find the proportion that result in assigning a posterior probability of .9 (or other threshold) to H, a treatment is effective, assuming H is false. To correct for the raised frequency of reaching such thresholds in sequential trials, researchers are required to raise them.
To be clear, I am not saying error statistical methods could be replaced by such attempted calibrations, even if successful. They cannot. My point is only that an inquiry keen to generalize and to distinguish real from spurious should not feel free to ignore preregistration protocols in testing, even if they ditch p-values–at least not until they severely checked the consequences, wearing error statistical spectacles. This brings me to a closing remark about the links between the general debate about evidence in statistics and in philosophy.
A word about the philosophical advantages
The reason the likelihood principle is viewed as having philosophical imprimatur connects the controversy about preregistration to a controversy in philosophy between Popper’s account of falsification and confirmation theories. To logical empiricist philosophers, the holy grail was to find a logic of inductive inference that is purely formal in the same sense as deductive logic. All you would need to consider were the statements of evidence and hypotheses to arrive at inferences. Evidence “e” was the unproblematic, empirical, ground statement; confirmation logics defined various C-functions between e and h: C(h, e).[6] Popper challenged the idea that data were given, recognizing that we always observe through the lens of a perspective, an interest, a theory, a framework.[7] Nowadays, few hold to the idea that “data” are value-free, offering a relatively unproblematic basis for inference, but the ideal has not lost its attractiveness in many quarters.
I will undoubtedly return to make corrections to this continuation. If I do, I will indicate the version # in the title.
References can be found (with links) on the “Captain’s Biblio” here.
You can find all of the 16 “tours” in my book, Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (CUP 2018) on this blogpost. (It includes the entire manuscript in “proof” form, which is quite readable.)
All topics are also discussed on this blog, with very useful comments from readers.
________
[1] The Opera experiment (link).
[2] If you are interested, you can search this blog for an enormous amount of material on the likelihood principle.
[3] The law of likelihood (Hacking 1965), which is weaker than the likelihood principle, says that x0 comparatively favors H1 over H0 if Pr(x0;H1) exceeds Pr(x0;H0). Hacking rejects it in 1980 along with the declaration that “there is no such thing as a logic of induction”.
[4] Objective or conventional Bayesians accept technical violations of the likelihood principle enabling them to achieve matching with frequentist error probabilities, but as I explain in SIST (2018), one feels their hearts aren’t in it. Empirical Bayesians like Efron reject the likelihood principle which is why Lindley says there’s “no one less Bayesian than an empirical Bayesian” (1969, p. 421). A ‘falsificationist Bayesian’, as Andrew Gelman calls himself, employs error statistics in model checking, at least he did a decade ago. Please see links in SIST (2018), and in searching this blog.
[5] I don’t think they are permitted in confirmatory trials.
[6] Of course, they were never able to find an adequate inductive logic. Carnap constructed a “continuum” of inductive logics. Harold Jeffreys developed an early “objective” Bayesian account in statistics, later built on my many others. Objective Bayesians haven’t settled on an adequate system. A few are: reference, frequentist matching, invariance.
[7] Nor can you just look at data x for Popper; you’d need to know if it resulted from a ‘sincere effort’ to falsify the claim inferred. He never fully fleshed out his philosophy of severe testing. He wrote to me once that he regretted not learning more about formal statistical testing.
.
I had been invited to speak at a Royal Society meeting, held March 4 and 5, 2024, on “the promises and pitfalls of preregistration”—a topic in which I’m keenly interested. The meeting was organized by Dr Tom Hardwicke, Professor Marcus Munafò, Dr Sophia Crüwell, Professor Dorothy Bishop FRS FMedSci, and Professor Eric-Jan Wagenmakers. Unfortunately, I was unable to travel to London, so I had to decline attending a few months ago. But, I thought I might jot down some remarks here.
The flyer defines preregistration as “publicly declaring study plans before collecting or analyzing data”. I regard it as a welcome consequence of today’s statistical crisis of replication that some social sciences are taking a page from medical trials and calling for preregistration of sampling protocols and full reporting.[1] The major source of failed replication stems from the ease of obtaining impressive looking findings by data-dredging, multiple testing, outcome-switching, cherry-picking, optional stopping and a host of related “selection effects”. Such gambits may practically guarantee an impressive-looking effect, even if it is spurious. The inferred effect, H, agrees with the data but the test H has passed lacks stringency or severity. Such agreements, Popper (1983) might say, are “too cheap to be worth having”. Little if anything has been done to avoid the key flaw of concern: being fooled by chance. The key promise of preregistration of protocols (along with full reporting and replication checks) is to help block the biasing selection effects known to permit insevere tests.
Pejorative vs non-pejorative cases.
However, there are also cases of data dredging and multiple testing that satisfy severe testing requirements. Consider searching a full database for a DNA match with a criminal’s DNA, where we suppose the probabilities of false negatives and false positives are both extremely low. Since the probability is very high of a mismatch with person i, if i were not the criminal (and a nonmatch virtually excludes the person), a match with i warrants inferring that i is the criminal. Unlike examples of pejorative searching, where the concern is mistakenly inferring an effect is genuine, here there is a known effect or specific event, the criminal’s DNA, and stringent procedures are used to track down the source. Ruling out random chance is very different from explaining a known effect.
Nor need the data dredged hypotheses be prespecified to be tested with severity by data—even where those data were “used” to arrive at or select the hypothesis inferred. An example that is actually quite routine (although it sparked some controversy) is from the data analysis of the 1919 eclipse data. The same data were used to arrive at, as well as to test, the source of one set of blurred eclipse data during the tests of the Einstein deflection effect. The culprit–a distorted telescope mirror due to the sun’s heat–was not predesignated, but it was severely tested (i.e., it passed with severity).
This reminds me to make a remark on language: That H was severely tested means H passes a test with (high) severity. That is, it was subjected to, and passes, a test that it probably would have failed, just in case H is false.
It is the severity of a test, or lack of it, that distinguishes whether a data-dredged claim is warranted by data.
Had I been fortunate enough to fulfill the high honor of being invited to speak at this RS meeting, I think I would have started with these points. I would be keen to distinguish right off pejorative from non-pejorative data dredging. That is because one of the alleged “pitfalls” of preregistration is that it can discourage the great discoveries of science. Didn’t all these great scientists reach important discoveries by trenchantly exploring the data for patterns? Sure, but whether they stringently tested those patterns is a distinct question. (Note too that when data-driven discoveries are tested on brand new data, it is not considered data-dredging. The new data were not used in arriving at the hypothesis being scrutinized.) Moreover, frequentist error statisticians are sometimes wrongly criticized for requiring adjustments for selection in cases where no adjustment is needed or makes sense. The case of DNA matching is an example.[2] So I would want to clear the air right away that calls for preregistration to block data-dredging—where they matter—are calls for severe testing, notably where there’s an interest in avoiding being fooled by random chance.
The fact that there are ways to compensate for selection effects does not detract from preregistration’s legitimate rationale. Quite the opposite. It only underscores the importance of recognizing how selection effects can alter the reliability of inferences. Needing (or striving) to compensate, entails that they cannot just be ignored. With today’s use of big data, and AI and machine learning methods, “post data selection” is a research area of its own.[3]
Note that multiplicity and “trying and trying again” can refer to the hypothesis that is dredged to accord with given data, or it may refer to trying and trying again to find data that accord with a fixed hypothesis. (It can refer to other things besides.) There can be just as much latitude for bias in selecting what is reported as “the data” for testing a predesignated hypothesis as there is in selecting a hypothesis that agrees with data.
Biasing selection effects.In my view, pejorative selection or biasing selection effects occur when data or hypotheses are selected, generated or interpreted in such a way as to result in failing the severity requirement.
Under failing the severity requirement I include, not only reporting as severely tested a claim that actually passes with very low severity, but the inability to assess how severely a claim has been tested, even approximately. For example, the FDA allows adjusting for selection only among prespecified hypotheses or endpoints. Open-ended data torturing precludes such adjustments.[4]
Prior to adoption of preregistered endpoints, the FDA reports, it was not uncommon for a clinical trialist, failing to find a statistically significant treatment benefit on prespecified endpoints, to ransack the unblinded data, cherry-pick a subgroup where treateds do better in some respect than controls, and report it as evidence of treatment benefit. Drawing a line around treateds who happen to show some beneficial effect is akin to the Texas Marksman circling a cluster of closely placed bullet holes, and regarding it as evidence of his marksmanship. The accordance between data and hypothesis is due to the biasing selection effects, not the truth of the hypotheses (about the drug’s benefit or his marksmanship).
Preregistration and error probabilities.
An adequate statistical account must be able to pick up on how biasing selection effects alter the capabilities of a method to assess and control erroneous interpretations of data. Moreover, these altered capabilities must show up in the evidential assessment—even if it is only to declare the inference unwarranted due to data torturing. They show up in a severity assessment of the inference, whether quantitative, semi-quantitative, or qualitative. In a formal statistical context, they may be provided by using the relevant sampling distribution.
Consider what the critical reader of a preregistered report might do, whether pre-data or post-data. She looks, in effect, at the number of chances the researchers give themselves to find apparent effects by chance alone. She asks, what’s the probability that one or another hypothesis, stopping point, choice of grouping variables, ways to interpret a measurement, and so on, could lead (or have led) to a false positive—even without a formal error probability computation? If it’s fairly high, she denies there is evidence for the effect. The onus is on those claiming to provide evidence to show they have worked to avoid known traps that blow up the probability of false positives and make it all too easy to mistake chance variability as real. Thus, the rationale for preregistration goes hand in hand with that of controlling error probabilities. There is a tension, therefore, between popular calls for preregistration and statistical accounts that downplay error probabilities. That’s what I would talk about next.
To be continued. Please share your thoughts in the comments. If I find any of the contributors’ talks are available, I will link to them.
I will update and make corrections, indicating the version.
[1] Anyone who has ever read this blog knows I’ve talked quite a lot about this topic over many years (e.g., Mayo 2018, CUP), and have exchanged ideas on the topic with many others. I will refrain from references here, but please search the blog. Links to most of my published papers are at https://errorstatistics.com/mayo-publications/.
[2] Dawid’s (2000) comment on Lindley is here.
[3] AI/ML prediction models might compensate for using the “same” data by cross validation and data splitting, at least with IID; but these fields also face reproducibility and replicability crises.
[4] Usually the primary endpoint must be found statistically significant before secondary endpoints are considered.
17 Feb 1890-29 July 1962
Today is R.A. Fisher’s birthday! I am reblogging what I call the “Triad”–an exchange between Fisher, Neyman and Pearson (N-P) published 20 years after the Fisher-Neyman break-up. While my favorite is still the reply by E.S. Pearson, which alone should have shattered Fisher’s allegations that N-P “reinterpret” tests of significance as “some kind of acceptance procedure”, all three are chock full of gems for different reasons. They are short and worth rereading. Neyman’s article pulls back the cover on what is really behind Fisher’s over-the-top polemics, what with Russian 5-year plans and commercialism in the U.S. Not only is Fisher jealous that N-P tests came to overshadow “his” tests, he is furious at Neyman for driving home the fact that Fisher’s fiducial approach had been shown to be inconsistent (by others). The flaw is illustrated by Neyman in his portion of the triad. Details may be found in my book, SIST (2018) especially pp 388-392 linked to here. It speaks to a common fallacy seen every day in interpreting confidence intervals. As for Neyman’s “behaviorism”, Pearson’s last sentence is revealing.
HAPPY BIRTHDAY R.A. FISHER!
“Statistical Methods and Scientific Induction“
by Sir Ronald Fisher (1955)
SUMMARY
The attempt to reinterpret the common tests of significance used in scientific research as though they constituted some kind of acceptance procedure and led to “decisions” in Wald’s sense, originated in several misapprehensions and has led, apparently, to several more.
The three phrases examined here, with a view to elucidating they fallacies they embody, are:
Mathematicians without personal contact with the Natural Sciences have often been misled by such phrases. The errors to which they lead are not only numerical.
To continue reading Fisher’s paper.
“Note on an Article by Sir Ronald Fisher“
by Jerzy Neyman (1956)
Neyman
Summary
(1) FISHER’S allegation that, contrary to some passages in the introduction and on the cover of the book by Wald, this book does not really deal with experimental design is unfounded. In actual fact, the book is permeated with problems of experimentation. (2) Without consideration of hypotheses alternative to the one under test and without the study of probabilities of the two kinds, no purely probabilistic theory of tests is possible.
(3) The conceptual fallacy of the notion of fiducial distribution rests upon the lack of recognition that valid probability statements about random variables usually cease to be valid if the random variables are replaced by their particular values. The notorious multitude of “paradoxes” of fiducial theory is a consequence of this oversight. (4) The idea of a “cost function for faulty judgments” appears to be due to Laplace, followed by Gauss.
E.S. Pearson
“Statistical Concepts in Their Relation to Reality“.
by E.S. Pearson (1955)
Controversies in the field of mathematical statistics seem largely to have arisen because statisticians have been unable to agree upon how theory is to provide, in terms of probability statements, the numerical measures most helpful to those who have to draw conclusions from observational data. We are concerned here with the ways in which mathematical theory may be put, as it were, into gear with the common processes of rational thought, and there seems no reason to suppose that there is one best way in which this can be done. If, therefore, Sir Ronald Fisher recapitulates and enlarges on his views upon statistical methods and scientific induction we can all only be grateful, but when he takes this opportunity to criticize the work of others through misapprehension of their views as he has done in his recent contribution to this Journal (Fisher 1955 “Statistical Methods and Scientific Induction” ), it is impossible to leave him altogether unanswered.
In the first place it seems unfortunate that much of Fisher’s criticism of Neyman and Pearson’s approach to the testing of statistical hypotheses should be built upon a “penetrating observation” ascribed to Professor G.A. Barnard, the assumption involved in which happens to be historically incorrect. There was no question of a difference in point of view having “originated” when Neyman “reinterpreted” Fisher’s early work on tests of significance “in terms of that technological and commercial apparatus which is known as an acceptance procedure”. There was no sudden descent upon British soil of Russian ideas regarding the function of science in relation to technology and to five-year plans. It was really much simpler–or worse. The original heresy, as we shall see, was a Pearson one!…
Use this link to continue reading, “Statistical Concepts in Their Relation to Reality“.
I will be giving an online talk on Friday, Feb 2, 4:30-5:45 NYC time, at a conference you can watch on zoom this week (Jan 30-Feb 2): Is Philosophy Useful for Science, and/or Vice Versa? It’s taking place in-person and online at Chapman University. My talk is: “The importance of philosophy of science for Statistical Science and vice versa”. I’ll touch on a current paper I’m writing that (finally) gets back to “Bayesian conceptions of severity”, (in contrast to error statistical severity) as begun on the post on Van Dongen, Springer, and Wagenmaker (2022).
I’ll put my slides up here on Friday. Here are my slides from the conference.
You can view the conference presentations live on Zoom here: https://chapman.zoom.us/j/92313633595
Today looks to be mostly philosophy in mathematics, tomorrow, biology and psychology, and Friday, physics and a couple on statistics. Here is the full Program: Philosophy Science Conference Program 01.24.2024
$8,765.53 X 2
I want to extend my warmest thanks to all who became Friends of David R. Cox in 2022. Your generous donations to the David R. Cox Foundations of Statistics Award are honoring the contributions of David R. Cox, and promoting the importance of statistical foundations in the American Statistical Association (ASA):
Karim Abadir, Heather Battey, Yoav Benjamini, Stuart Bevan, Alex Blocker, John Bibby, Lynne Billard, Sheila M. Bird, William Browning, John Byrd, Nancy Cartwright, Michael P. Cohen, Noel Cressie, Robert Crouchley, Gary R. Cutter, Anthony C. Davison, Bianca De Stavola, Edgar, Dobriban, Christl Donnelly, Vern Farewell, Samuel Fletcher, David Firth, David Freeborn, Andrew Gelman, David J. Hand, Sylvia Halasz, Frank E. Harrell, Maria-Eglee Perez Hernandez, Klaus Hinkelmann, Michelle Jackson, Patricia A. Jacobs, Harold Jaffe, Bimal Jain, Christiana Kartsonaki, Robert and Loretta Kass, Jesse Krijthe, Daniel Lakens, Ji-Hyun Lee, Lei Liu, Francisco Louzada, Donald Macnaughton, Giovanni M. Marchetti, Kanti V. Mardia, Peter McCullagh, Xiao-Li Meng, Jean Miller, Georges A. A Monette, Pavlos Msaouel, David Oakes, David Oliver, Yusuke Ono, John Park, Jose G. Ramirez, Nancy Reid, James L. Rosenberger, Richard J. Samworth, Stephen J. Senn, Dylan S. Small, David M. Smith, Aris Spanos, Alex Sutherland, John Tomenson, Tengyao Wang, Ronald L. Wasserstein, Gideon Weiss, Manyu Wong, Henry L. Wyneken, Henry Wynn
We exceeded our increased goal (from $5,000-$7,500) last year, raising $8,765.53, all of which is being matched! *
*During the matching period, anyone who donated $50 became a “friend” of David Cox. (It is ordinarily $100.)
This award honors the contributions of David R. Cox to the foundations of statistical inference, experimental design, and data analysis. It was established in 2022 to promote research and teaching that illuminates conceptual, theoretical, philosophical, and historical perspectives on statistical science and to advance understanding of comparative approaches to the interpretation and communication of statistical data. The honoree will receive a $2,000 honorarium and up to $1,000 toward travel expenses to present a lecture at JSM.
The inaugural winner was Nancy Reid. More information on the David R. Cox Foundations of Statistics Award, and forms for nominating someone, can be found at the ASA website on this page.
Many thanks to Amanda Malloy and Ron Wasserstein for all of their help in managing the matching program.
The ASA secure donation link can be found here.
.
For three of the last four years, it was not feasible to actually revisit that spot in the road, looking to get into a strange-looking taxi, to head to “Midnight With Birnbaum”. Even last year was iffy. But this year I will, and I’m about to leave at 9pm. (The pic on the left is the only blurry image I have of the club I’m taken to.) My book Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (CUP, 2018) doesn’t include the argument from my article in Statistical Science (“On the Birnbaum Argument for the Strong Likelihood Principle”), but you can read it at that link along with commentaries by A. P. David, Michael Evans, Martin and Liu, D. A. S. Fraser, Jan Hannig, and Jan Bjornstad. David Cox, who very sadly did in January 2022, is the one who encouraged me to write and publish it. (The first David R. Cox Foundations of Statistics Prize will be awarded at the JSM 2023.) Not only does the (Strong) Likelihood Principle (LP or SLP) remain at the heart of many of the criticisms of Neyman-Pearson (N-P) statistics and of error statistics in general, but a decade after my 2014 paper, it is more central than ever–even if it is often unrecognized.
As Birnbaum emphasized, the “confidence concept” is the “one rock in a shifting scene” of statistical foundations, insofar as there’s interest in controlling the frequency of erroneous interpretations of data. (See my rejoinder to commentators.) Birnbaum bemoaned the lack of an explicit evidential interpretation of N-P methods. I purport to give one in SIST 2018 based on severe testing. Anyway, let’s see what happens this year when the event ought to be in full swing. Happy New Year!
BACKGROUND
You know how in that Woody Allen movie, “Midnight in Paris,” the main character (I forget who plays it, I saw it on a plane) is a writer finishing a novel, and he steps into a cab that mysteriously picks him up at midnight and transports him back in time where he gets to run his work by such famous authors as Hemingway and Virginia Wolf? (It was a new movie when I began the blog in 2011.) He is wowed when his work earns their approval and he comes back each night in the same mysterious cab…Well, imagine an error statistical philosopher is picked up in a mysterious taxi at midnight on New Year’s Eve and lo and behold, finds herself in the company of Allan Birnbaum.[i]
OUR EXCHANGE:
ERROR STATISTICIAN: It’s wonderful to meet you Professor Birnbaum; I’ve always been extremely impressed with the important impact your work has had on philosophical foundations of statistics. I happen to have published on your famous argument about the likelihood principle (LP). (whispers: I can’t believe this!)
BIRNBAUM: Ultimately you know I rejected the LP as failing to control the error probabilities needed for my Confidence concept. But you know all this, I’ve read it in your book: Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (STINT, 2018, CUP).
ERROR STATISTICIAN: You’ve read my book? Wow! Then you know I don’t think your argument shows that the LP follows from such frequentist concepts as sufficiency S and the weak conditionality principle WLP. I don’t rehearse my argument there, but I first found the problem in 2006, when I was writing something on “conditioning” with David Cox. [ii] Sorry,…I know it’s famous…
BIRNBAUM: Well, I shall happily invite you to take any case that violates the LP and allow me to demonstrate that the frequentist is led to inconsistency, provided she also wishes to adhere to the WLP and sufficiency (although less than S is needed).
ERROR STATISTICIAN: Well I show that no contradiction follows from holding WCP and S, while denying the LP.
BIRNBAUM: Well, well, well: I’ll bet you a bottle of Elba Grease champagne that I can demonstrate it!
ERROR STATISTICAL PHILOSOPHER: It is a great drink, I must admit that: I love lemons.
BIRNBAUM: OK. (A waiter brings a bottle, they each pour a glass and resume talking). Whoever wins this little argument pays for this whole bottle of vintage Ebar or Elbow or whatever it is Grease.
.
ERROR STATISTICAL PHILOSOPHER: I really don’t mind paying for the bottle.
BIRNBAUM: Good, you will have to. Take any LP violation. Let x’ be 2-standard deviation difference from the null (asserting μ = 0) in testing a normal mean from the fixed sample size experiment E’, say n = 100; and let x” be a 2-standard deviation difference from an optional stopping experiment E”, which happens to stop at 100. Do you agree that:
(0) For a frequentist, outcome x’ from E’ (fixed sample size) is NOT evidentially equivalent to x” from E” (optional stopping that stops at n)
ERROR STATISTICAL PHILOSOPHER: Yes, that’s a clear case where we reject the strong LP, and it makes perfect sense to distinguish their corresponding p-values (which we can write as p’ and p”, respectively). The searching in the optional stopping experiment makes the p-value quite a bit higher than with the fixed sample size. For n = 100, data x’ yields p’= ~.05; while p” is ~.3. Clearly, p’ is not equal to p”, I don’t see how you can make them equal.
BIRNBAUM: Suppose you’ve observed x”, a 2-standard deviation difference from an optional stopping experiment E”, that finally stops at n=100. You admit, do you not, that this outcome could have occurred as a result of a different experiment? It could have been that a fair coin was flipped where it is agreed that heads instructs you to perform E’ (fixed sample size experiment, with n = 100) and tails instructs you to perform the optional stopping experiment E”, stopping as soon as you obtain a 2-standard deviation difference, and you happened to get tails, and performed the experiment E”, which happened to stop with n =100.
ERROR STATISTICAL PHILOSOPHER: Well, that is not how x” was obtained, but ok, it could have occurred that way.
BIRNBAUM: Good. Then you must grant further that your result could have come from a special experiment I have dreamt up, call it a BB-experiment. In a BB-experiment, if the outcome from the experiment you actually performed has an outcome with a proportional likelihood to one in some other experiment not performed, E’, then we say that your result has an “LP pair”. For any violation of the strong LP, the outcome observed, let it be x”, has an “LP pair”, call it x’, in some other experiment E’. In that case, a BB-experiment stipulates that you are to report x” as if you had determined whether to run E’ or E” by flipping a fair coin.
(They fill their glasses again)
ERROR STATISTICAL PHILOSOPHER: You’re saying that if my outcome from trying and trying again, that is, optional stopping experiment E”, with an “LP pair” in the fixed sample size experiment I did not perform, then I am to report x” as if the determination to run E” was by flipping a fair coin (which decides between E’ and E”)?
BIRNBAUM: Yes, and one more thing. If your outcome had actually come from the fixed sample size experiment E’, it too would have an “LP pair” in the experiment you did not perform, E”. Whether you actually observed x” from E”, or x’ from E’, you are to report it as x” from E”.
ERROR STATISTICAL PHILOSOPHER: So let’s see if I understand a Birnbaum BB-experiment: whether my observed 2-standard deviation difference came from E’ or E” (with sample size n) the result is reported as x’, as if it came from E’ (fixed sample size), and as a result of this strange type of a mixture experiment.
BIRNBAUM: Yes, or equivalently you could just report x*: my result is a 2-standard deviation difference and it could have come from either E’ (fixed sampling, n= 100) or E” (optional stopping, which happens to stop at the 100th trial). That’s how I sometimes formulate a BB-experiment.
ERROR STATISTICAL PHILOSOPHER: You’re saying in effect that if my result has an LP pair in the experiment not performed, I should act as if I accept the strong LP and just report it’s likelihood; so if the likelihoods are proportional in the two experiments (both testing the same mean), the outcomes are evidentially equivalent.
BIRNBAUM: Well, but since the BB- experiment is an imagined “mixture” it is a single experiment, so really you only need to apply the weak LP which frequentists accept. Yes? (The weak LP is the same as the sufficiency principle).
ERROR STATISTICAL PHILOSOPHER: But what is the sampling distribution in this imaginary BB- experiment? Suppose I have Birnbaumized my experimental result, just as you describe, and observed a 2-standard deviation difference from optional stopping experiment E”. How do I calculate the p-value within a Birnbaumized experiment?
BIRNBAUM: I don’t think anyone has ever called it that.
ERROR STATISTICAL PHILOSOPHER: I just wanted to have a shorthand for the operation you are describing, there’s no need to use it, if you’d rather I not. So how do I calculate the p-value within a BB-experiment?
BIRNBAUM: You would report the overall p-value, which would be the average over the sampling distributions: (p’ + p”)/2
Say p’ is ~.05, and p” is ~.3; whatever they are, we know they are different, that’s what makes this a violation of the strong LP (given in premise (0)).
ERROR STATISTICAL PHILOSOPHER: So you’re saying that if I observe a 2-standard deviation difference from E’, I do not report the associated p-value p’, but instead I am to report the average p-value, averaging over some other experiment E” that could have given rise to an outcome with a proportional likelihood to the one I observed, even though I didn’t obtain it this way?
BIRNBAUM: I’m saying that you have to grant that x’ from a fixed sample size experiment E’ could have been generated through a BB-experiment.
My this drink is sour!
ERROR STATISTICAL PHILOSOPHER: Yes, I love pure lemon.
BIRNBAUM: Perhaps you’re in want of a gene; never mind.
I’m saying you have to grant that x’ from a fixed sample size experiment E’ could have been generated through a BB-experiment. If you are to interpret your experiment as if you are within the rules of a BB experiment, then x’ is evidentially equivalent to x” (is equivalent to x*). This is premise (1).
ERROR STATISTICAL PHILOSOPHER: But the result would be that the p-value associated with x’ (fixed sample size) is reported to be larger than it actually is (.05), because I’d be averaging over fixed and optional stopping experiments; while observing x” (optional stopping) is reported to be smaller than it is–in both cases because of an experiment I did not perform.
BIRNBAUM: Yes, the BB-experiment computes the P-value in an unconditional manner: it takes the convex combination over the 2 ways the result could have come about.
ERROR STATISTICAL PHILOSOPHER: this is just a matter of your definitions, it is an analytical or mathematical result, so long as we grant being within your BB experiment.
BIRNBAUM: True, (1) plays the role of the sufficiency assumption, but one need not even appeal to sufficiency, it is just a matter of mathematical equivalence.
By the way, I am focusing just on LP violations, therefore, the outcome, by definition, has an LP pair. In other cases, where there is no LP pair, you just report things as usual.
ERROR STATISTICAL PHILOSOPHER: OK, but p’ still differs from p”; so I still don’t how I’m forced to infer the strong LP which identifies the two. In short, I don’t see the contradiction with my rejecting the strong LP in premise (0). (Also we should come back to the “other cases” at some point….)
BIRNBAUM: Wait! Don’t be so impatient; I’m about to get to step (2). Here, let’s toast to the new year: “To Elbar Grease!”
ERROR STATISTICAL PHILOSOPHER: To Elbar Grease!
BIRNBAUM: So far all of this was step (1).
ERROR STATISTICAL PHILOSOPHER: : Oy, what is step 2?
BIRNBAUM: STEP 2 is this: Surely, you agree, that once you know from which experiment the observed 2-standard deviation difference actually came, you ought to report the p-value corresponding to that experiment. You ought NOT to report the average (p’ + p”)/2 as you were instructed to do in the BB experiment.
This gives us premise (2a):
(2a) outcome x”, once it is known that it came from E”, should NOT be analyzed as in a BB- experiment where p-values are averaged. The report should instead use the sampling distribution of the optional stopping test E”, yielding the p-value, p” (~.37). In fact, .37 is the value you give in STINT p. 44 (imagining the experimenter keeps taking 10 more).
ERROR STATISTICAL PHILOSOPHER: So, having first insisted I imagine myself in a Birnbaumized, I mean a BB-experiment, and report an average p-value, I’m now to return to my senses and “condition” in order to get back to the only place I ever wanted to be, i.e., back to where I was to begin with?
BIRNBAUM: Yes, at least if you hold to the weak conditionality principle WCP (of D. R. Cox)—surely you agree to this.
(2b) Likewise, if you knew the 2-standard deviation difference came from E’, then
x’ should NOT be deemed evidentially equivalent to x” (as in the BB experiment), the report should instead use the sampling distribution of fixed test E’, (.05).
ERROR STATISTICAL PHILOSOPHER: So, having first insisted I consider myself in a BB-experiment, in which I report the average p-value, I’m now to return to my senses and allow that if I know the result came from optional stopping, E”, I should “condition” on and report p”.
BIRNBAUM: Yes. There was no need to repeat the whole spiel.
ERROR STATISTICAL PHILOSOPHER: I just wanted to be clear I understood you. Of course, all of this assumes the model is correct or adequate to begin with.
BIRNBAUM: Yes, the LP (or SLP, to indicate it’s the strong LP) is a principle for parametric inference within a given model. So you arrive at (2a) and (2b), yes?
ERROR STATISTICAL PHILOSOPHER: OK, but it might be noted that unlike premise (1), premises (2a) and (2b) are not given by definition, they concern an evidential standpoint about how one ought to interpret a result once you know which experiment it came from. In particular, premises (2a) and (2b) say I should condition and use the sampling distribution of the experiment known to have been actually performed, when interpreting the result.
BIRNBAUM: Yes, and isn’t this weak conditionality principle WCP one that you happily accept?
ERROR STATISTICAL PHILOSOPHER: Well the WCP originally refers to actual mixtures, where one flipped a coin to determine if E’ or E” is performed, whereas, you’re requiring I consider an imaginary Birnbaum mixture experiment, where the choice of the experiment not performed will vary depending on the outcome that needs an LP pair; and I cannot even determine what this might be until after I’ve observed the result that would violate the LP? I don’t know what the sample size will be ahead of time.
BIRNBAUM: Sure, but you admit that your observed x” could have come about through a BB-experiment, and that’s all I need. Notice
(1), (2a) and (2b) yield the strong LP!
Outcome x” from E”(optional stopping that stops at n) is evidentially equivalent to x’ from E’ (fixed sample size n).
ERROR STATISTICAL PHILOSOPHER: Clever, but your “proof” is obviously unsound; and before I demonstrate this, notice that the conclusion, were it to follow, asserts p’ = p”, (e.g., .05 = .3!), even though it is unquestioned that p’ is not equal to p”, that is because we must start with an LP violation (premise (0)).
BIRNBAUM: Yes, it is puzzling, but where have I gone wrong?
(The waiter comes by and fills their glasses; they are so deeply engrossed in thought they do not even notice him.)
ERROR STATISTICAL PHILOSOPHER: There are many routes to explaining a fallacious argument. The one I find most satisfactory is in Mayo (2014). But, given we’ve been partying, here’s a very simple one. What is required for STEP 1 to hold, is the denial of what’s needed for STEP 2 to hold:
Step 1 requires us to analyze results in accordance with a BB- experiment. If we do so, true enough we get:
premise (1): outcome x” (in a BB experiment) is evidentially equivalent to outcome x’ (in a BB experiment):
That is because in either case, the p-value would be (p’ + p”)/2
Step 2 now insists that we should NOT calculate evidential import as if we were in a BB- experiment. Instead we should consider the experiment from which the data actually came, E’ or E”:
premise (2a): outcome x” (in a BB experiment) is/should be evidentially equivalent to x” from E” (optional stopping that stops at n): its p-value should be p”.
premise (2b): outcome x’ (within in a BB experiment) is/should be evidentially equivalent to x’ from E’ (fixed sample size): its p-value should be p’.
If (1) is true, then (2a) and (2b) must be false!
If (1) is true and we keep fixed the stipulation of a BB experiment (which we must to apply step 2), then (2a) is asserting:
The average p-value (p’ + p”)/2 = p’ which is false.
Likewise if (1) is true, then (2b) is asserting:
the average p-value (p’ + p”)/2 = p” which is false
Alternatively, we can see what goes wrong by realizing:
If (2a) and (2b) are true, then premise (1) must be false.
In short your famous argument requires us to assess evidence in a given experiment in two contradictory ways: as if we are within a BB- experiment (and report the average p-value) and also that we are not, but rather should report the actual p-value.
I can render it as formally valid, but then its premises can never all be true; alternatively, I can get the premises to come out true, but then the conclusion is false—so it is invalid. In no way does it show the frequentist is open to contradiction (by dint of accepting S, WCP, and denying the LP).
BIRNBAUM: Yet some people still think it is a breakthrough. I never agreed to go as far as Jimmy Savage wanted me too, namely, to be a Bayesian….
ERROR STATISTICAL PHILOSOPHER: My 2014 paper gives a much clearer exposition of what goes wrong in your argument than I did in the discussion from 2010. There were still several gaps, and lack of a clear articulation of the WCP. In fact, I’ve come to see that clarifying the entire argument turns on defining the WCP. Have you seen my 2014 paper in Statistical Science? The key difference is that in (2014), the WCP is stated as an equivalence, as you intended. Cox’s WCP, many claim, was not an equivalence, going in 2 directions. Slides from a presentation may be found on this blogpost.
Birnbaum: Yes, the “monster of the LP” arises from viewing WCP as an equivalence, instead of going in one direction (from mixtures to the known result).
ERROR STATISTICAL PHILOSOPHER: In my 2014 paper (unlike my earlier treatments) I too construe WCP as giving an “equivalence” but there is an equivocation that invalidates the purported move to the LP.
On the one hand, it’s true that if z is known (and known for example to have come from optional stopping), it’s irrelevant that it could have resulted from either fixed sample testing or optional stopping.
But it does not follow that if z is known, it’s irrelevant whether it resulted from fixed sample testing or optional stopping. It’s the slippery slide into this second statement–which surely sounds the same as the first–that makes your argument such a brain buster.
BIRNBAUM: Yes I have seen your 2014 paper! Your Rejoinder to some of the critics is gutsy, to say the least. I’ve also seen the slides on your blog.
ERROR STATISTICAL PHILOSOPHER: Thank you, I’m amazed you follow my blog! But look I must get your answer to a question before you leave this year.
Sudden interruption by the waiter who, very wisely, is wearing an N95 mask:
WAITER: Who gets the tab? We’re closing a bit early due to Covid.
BIRNBAUM: I do. To Elbar Grease! And to your (still) new book SIST! I have a list of comments and questions right here…
ERROR STATISTICAL PHILOSOPHER: I believe you handed them to me last year, but when I returned, they were nowhere to be found. I won’t let them disappear this year. (She takes a long legal-sized yellow sheet from Birnbaum, noticing it is filled with tiny hand-written comments, covering both sides.)
BIRNBAUM: To Elbar Grease! To Severe Testing! Happy New Year!
ERROR STATISTICAL PHILOSOPHER: I have one quick question, Professor Birnbaum, and I swear that whatever you say will be just between us, I won’t tell a soul. In your last couple of papers, you suggest you’d discovered the flaw in your argument for the LP. Am I right? Even in the discussion of your (1962) paper, you seemed to agree with Pratt that WCP can’t do the job you intend.
BIRNBAUM: Savage, you know, never got off my case about remaining at “the half-way house” of likelihood, and not going full Bayesian. Then I wrote the review about the Confidence Concept as the one rock on a shifting scene… Pratt thought the argument should instead appeal to a Censoring Principle (basically, it doesn’t matter if your instrument cannot measure beyond k units if the measurement you’re making is under k units.)
ERROR STATISTICAL PHILOSOPHER: Yes, but who says frequentist error statisticians deny the Censoring Principle? So back to my question, you disappeared before answering last year…I just want to know…you did see the flaw, yes?
WAITER: We’re closing now; shall I call Remote Taxi?
BIRNBAUM: Yes, yes!
ERROR STATISTICAL PHILOSOPHER: ‘Yes’, you discovered the flaw in the argument, or ‘yes’ to the taxi?
MANAGER: We’re closing now; I’m sorry you must leave.
ERROR STATISTICAL PHILOSOPHER: We’re leaving I just need him to clarify his answer….
BIRNBAUM: I predict that 2024 will be the year that people will finally take seriously your paper from a decade ago!
ERROR STATISTICAL PHILOSOPHER: I’ll drink to that!
Suddenly a large group of people bustle past the manager…it’s all chaos.
Prof. Birnbaum…? Allan? Where did he go? (oy, not again!)
Link to complete discussion:
Mayo, Deborah G. On the Birnbaum Argument for the Strong Likelihood Principle (with discussion & rejoinder).Statistical Science 29 (2014), no. 2, 227-266.
[i] Many links on the strong likelihood principle (LP or SLP) and Birnbaum may be found by searching this blog. Good sources for where to start as well as classic background papers may be found in this blogpost. A link to slides and video of a very introductory presentation of my argument from the 2021 Phil Stat Forum is here.
January 7: “Putting the Brakes on the Breakthrough: On the Birnbaum Argument for the Strong Likelihood Principle” (D.Mayo)
[ii] I recently wrote a paper on Cox’s statistical philosophy. Sadly he died in 2022. By the way, Ronald Giere gave me numerous original papers of yours. They’re in files in my attic library. Some are in mimeo, others typed…I mean, obviously for that time that’s what they’d be…now of course, oh never mind, sorry.
.
If you read my 2023 paper on Cox’s philosophy of statistics, you’ll have come across Cox’s famous “weighing machine” example, which is thought to have caused “a subtle earthquake” in foundations of statistics. If you’re curious as to why that is, you’ll be interested to know that each year, on New Year’s Eve, I return to the conundrum. This post gives some background, and collects the essential links.
An essential component of inference based on familiar frequentist notions: p-values, significance and confidence levels, is the relevant sampling distribution (hence the term sampling theory, or my preferred error statistics, as we get error probabilities from the sampling distribution). This feature results in violations of a principle known as the strong likelihood principle (SLP). To state the SLP roughly, it asserts that all the evidential import in the data (for parametric inference within a model) resides in the likelihoods. If accepted, it would render error probabilities irrelevant post data.
SLP (We often drop the “strong” and just call it the LP. The “weak” LP just boils down to sufficiency)
For any two experiments E1 and E2 with different probability models f1, f2, but with the same unknown parameter θ, if outcomes x and y (from E1 and E2 respectively) determine the same (i.e., proportional) likelihood function (f1(x; θ) = cf2(*y; θ) for all θ), then x and y are inferentially equivalent (for an inference about θ*).
(What differentiates the weak and the strong LP is that the weak refers to a single experiment.)
Violation of SLP:
Whenever outcomes x and y from experiments E1 and E2 with different probability models f1, f2, but with the same unknown parameter θ, and f1(x; θ) = cf2(*y; θ) for all θ, and yet outcomes x and y have different implications for an inference about θ*.
For an example of a SLP violation, E1 might be sampling from a Normal distribution with a fixed sample size n, and E2 the corresponding experiment that uses an optional stopping rule: keep sampling until you obtain a result 2 standard deviations away from a null hypothesis that θ = 0 (and for simplicity, a known standard deviation). When you do, stop and reject the point null (in 2-sided testing).
The SLP tells us (in relation to the optional stopping rule) that once you have observed a 2-standard deviation result, there should be no evidential difference between its having arisen from experiment E1, where n was fixed, say, at 100, and experiment E2 where the stopping rule happens to stop at n = 100. For the error statistician, by contrast, there is a difference, and this constitutes a violation of the SLP.
———————-
Now for the surprising part: In Cox’s weighing machine example, discussed on this post, a coin is flipped to decide which of two experiments to perform? David Cox (1958) proposes something called the Weak Conditionality Principle (WCP) to restrict the space of relevant repetitions for frequentist inference. The WCP says that once it is known which Ei produced the measurement, the assessment should be in terms of the properties of the particular Ei. Nothing could be more obvious.
The surprising upshot of Allan Birnbaum’s (1962) argument is that the SLP appears to follow from applying the WCP in the case of mixture experiments, and so uncontroversial a principle as sufficiency (SP)–although even that has been shown to be optional, strictly speaking. But this would preclude the use of sampling distributions. J. Savage calls Birnbaum’s argument “a landmark in statistics” (see [i]).
Although his argument purports that [(WCP and SP) entails SLP], we will see that data may violate the SLP while holding both the WCP and SP. Such cases also directly refute [WCP entails SLP].
Binge reading the Likelihood Principle.
If you’re keen to binge read the SLP–a way to break holiday/winter break/pandemic doldrums– I’ve pasted most of the early historical sources below. The argument is simple; showing what’s wrong with it took a long time.
My earliest treatment, via counterexample, in Mayo (2010). A deeper argument is in Mayo (2014) in Statistical Science.[ii] An intermediate paper Mayo (2013) corresponds to a talk I presented at the JSM in 2013.
Interested readers may search this blog for quite a lot of discussion of the SLP including “U-Phils” (discussions by readers) (e.g., here, and here), and amusing notes (e.g., Don’t Birnbaumize that experiment my friend, and Midnight with Birnbaum).
This conundrum is relevant to the very notion of “evidence”, blithely taken for granted in both statistics and philosophy. [iii] I call upon astute philosophers of language to further articulate the problem, and my argument, for an informal philosophical audience.
To have a list for binging, I’ve grouped some key readings below [iv].
Classic Birnbaum Papers:
Note to Reader: If you look at the (1962) “discussion”, you can already see Birnbaum backtracking a bit, in response to Pratt’s comments.
Some additional early discussion papers:
Durbin:
There’s also a good discussion in Cox and Hinkley 1974.
Evans, Fraser, and Monette:
Kalbfleisch:
My discussions (also noted above):
[ii] The link includes comments on my paper by Bjornstad, Dawid, Evans, Fraser, Hannig, and Martin and Liu, and my rejoinder.
[iii] In Birnbaum’s argument, he introduces an informal, and rather vague, notion of the “evidence (or evidential meaning) of an outcome z from experiment E”. He writes it: Ev(E,z).
In my formulation of the argument, I introduce a new symbol ⇒ to represent a function from a given experiment-outcome pair, (E,z) to a generic inference implication. It (hopefully) lets us be clearer than does Ev.
(E,z) ⇒ InfrE(z) is to be read “the inference implication from outcome z in experiment E” (according to whatever inference type/school is being discussed).
If E is within error statistics, for example, it is necessary to know the relevant sampling distribution associated with a statistic. If it is within a Bayesian account, a relevant prior would be needed.
[iv] I’ve blogged these links in the past; please let me know if any links are broken.
On November 14, I gave a talk at the Seminar in Advanced Research Methods for the Department of Psychology, Princeton University.
“Statistical Inference as Severe Testing: Beyond Probabilism and Performance”The video of my talk is below along with the slides. It reminds me to return to a paper, half-written, replying to a paper on “A Bayesian Perspective on Severity” (van Dongen, Sprenger, Wagenmakers (2022). These authors claim that Bayesians can satisfy severity “regardless of whether the test has been conducted in a severe or less severe fashion”, but what they mean is that data can be much more probable on hypothesis H1 than on H0 –the Bayes factor can be high. However, “severity” can be satisfied in their comparative (subjective) Bayesian sense even for claims that are poorly probed in the error statistical sense (slides 55-6). Share your comments.
ABSTRACT: I develop a statistical philosophy in which error probabilities of methods may be used to evaluate and control the stringency or severity of tests. A claim is severely tested to the extent it has been subjected to and passes a test that probably would have found flaws, were they present. The severe-testing requirement leads to reformulating statistical significance tests to avoid familiar criticisms and abuses. While high-profile failures of replication in the social and biological sciences stem from biasing selection effects—data dredging, multiple testing, optional stopping—some reforms and proposed alternatives to statistical significance tests conflict with the error control that is required to satisfy severity. I discuss recent arguments to redefine, abandon, or replace statistical significance.
Below is a video Princeton recorded of the talk. My slides are below that.
RECORDING OF TALK:
https://videopress.com/v/tAgbbutz?resizeToParent=true&cover=true&preloadContent=metadata&useAverageColor=true“Statistical Inference as Severe Testing: Beyond Probabilism and Performance” at Princeton, Nov. 14MY SLIDES:
https://www.slideshare.net/slideshow/embed_code/key/JdVlC889AUoz9F?hostedIn=slideshare&page=upload
The Statistics Wars and Their Casualties WorkshopIt’s been 1 year (December 8, 2022) since our workshop, The Statistics Wars and Their Casualties! There were four sessions, held over 4 days. Below are the videos and slides from all four sessions of the Workshop. The first two sessions were held on September 22 & 23, 2022. Session 1 speakers were: Deborah Mayo (Virginia Tech), Richard Morey (Cardiff University), Stephen Senn (Edinburgh, Scotland). Session 2 speakers were: Daniël Lakens (Eindhoven University of Technology), Christian Hennig (University of Bologna), Yoav Benjamini (Tel Aviv University). The last two sessions were held on December 1 and 8. Session 3 speakers were: Daniele Fanelli (London School of Economics and Political Science), Stephan Guttinger (University of Exeter), and David Hand (Imperial College London). Session 4 speakers were: Jon Williamson (University of Kent), Margherita Harris (London School of Economics and Political Science), Aris Spanos (Virginia Tech), and Uri Simonsohn (Esade Ramon Llull University).
Abstracts can be found here and the schedule here. Some participant related publications are on this page.
The Statistics Wars and Their Casualties Workshop blog can be found here.
.SESSION 1
Brief Intro to Session 1 by David Hand (Imperial College)
https://videopress.com/v/WKkpy3Pl?resizeToParent=true&cover=true&posterUrl=https%3A%2F%2Fvideos.files.wordpress.com%2FWKkpy3Pl%2Fhand-intro_mp4_std.original.jpg&preloadContent=metadata&useAverageColor=trueDeborah Mayo (Virginia Tech):
The Statistics Wars and Their Casualties
https://videopress.com/v/kkNDwWO7?resizeToParent=true&cover=true&preloadContent=metadata&useAverageColor=trueRichard Morey (Cardiff University)
Bayes factors, p values, and the replication crisis
https://videopress.com/v/c2WiYGBZ?resizeToParent=true&cover=true&posterUrl=https%3A%2F%2Fvideos.files.wordpress.com%2Fc2WiYGBZ%2Fmorey-presentation-x.5-2_mp4_std.original.jpg&preloadContent=metadata&useAverageColor=trueSlide show is posted on his webpage here.
Stephen Senn (Edinburgh)
The replication crisis: are P-values the problem and are Bayes factors the solution?
https://videopress.com/v/0TyMHldF?resizeToParent=true&cover=true&posterUrl=https%3A%2F%2Fvideos.files.wordpress.com%2F0TyMHldF%2Fsenn-presentation-x.5-3_mp4_std.original.jpg&preloadContent=metadata&useAverageColor=trueSession 1 Discussion
https://videopress.com/v/DdNEl7bW?resizeToParent=true&cover=true&preloadContent=metadata&useAverageColor=trueSESSION 2
[Brief Intro to Session 2 by Stephen Senn (Edinburgh)]
Daniël Lakens (Eindhoven University of Technology)
The role of background assumptions in severity appraisal
https://videopress.com/v/IazBYWAI?resizeToParent=true&cover=true&preloadContent=metadata&useAverageColor=trueChristian Hennig (University of Bologna)
On the interpretation of the mathematical characteristics of statistical tests
https://videopress.com/v/iWZ7YhQt?resizeToParent=true&cover=true&preloadContent=metadata&useAverageColor=trueYoav Benjamini (Tel Aviv University)
The two statistical cornerstones of replicability: addressing selective inference and irrelevant variability
https://videopress.com/v/2nW1C8Oi?resizeToParent=true&cover=true&posterUrl=https%3A%2F%2Fvideos.files.wordpress.com%2F2nW1C8Oi%2Fbenjamini-presentation_mp4_std.original.jpg&preloadContent=metadata&useAverageColor=trueSession 2 Discussion
https://videopress.com/v/DI6Ugocg?resizeToParent=true&cover=true&preloadContent=metadata&useAverageColor=trueBelow are the videos and slides from the 7 talks from Session 3 and Session 4 of our workshop The Statistics Wars and Their Casualties held on December 1 & 8, 2022.
SESSION 3
Recap of recaps summary of Sessions 1 & 2:
https://videopress.com/v/ym9Q8hKl?resizeToParent=true&cover=true&posterUrl=https%3A%2F%2Fvideos.files.wordpress.com%2Fym9Q8hKl%2Frecaps-of-recaps-1_mp4_std.original.jpg&preloadContent=metadata&useAverageColor=trueIntroduction to Session: Daniël Lakens (Eindhoven University of Technology)
https://videopress.com/v/AebkX0UP?resizeToParent=true&cover=true&posterUrl=https%3A%2F%2Fvideos.files.wordpress.com%2FAebkX0UP%2Fintroduction-to-session_mp4_std.original.jpg&preloadContent=metadata&useAverageColor=trueDaniele Fanelli (London School of Economics and Political Science)
The neglected importance of complexity in statistics and Metascience
Stephan Guttinger (University of Exeter)
What are questionable research practices?
https://videopress.com/v/wmMxqj8v?resizeToParent=true&cover=true&preloadContent=metadata&useAverageColor=trueDavid Hand (Imperial College London)
What’s the question?
https://videopress.com/v/GZ3B1BnN?resizeToParent=true&cover=true&preloadContent=metadata&useAverageColor=trueDiscussion (Session 3): (a) Panel discussion of speakers; (b) general audience discussion; (c) “Where do we go from here (Part i)” participant discussion.
SESSION 4
Introduction to Session 4: Deborah Mayo (Virginia Tech)
Jon Williamson (University of Kent)
Causal inference is not statistical inference
Margherita Harris (London School of Economics and Political Science)
On Severity, the Weight of Evidence, and the Relationship Between the Two
https://videopress.com/v/ruv6LGet?resizeToParent=true&cover=true&preloadContent=metadata&useAverageColor=trueAris Spanos (Virginia Tech)
Revisiting the Two Cultures in Statistical Modeling and Inference as they relate to the Statistics Wars and Their Potential Casualties
Uri Simonsohn (Esade Ramon Llull University)
Mathematically Elegant Answers to Research Questions No One is Asking (meta-analysis, random effects models, and Bayes factors)
Where Should Stat Activists Go From Here? Deborah Mayo (Virginia Tech):
Discussion: (a) Panel discussions; (b) General audience discussion; (c) “Where do we go from here (Part ii)” participants and audience.
The Statistics Wars and Their Casualties Workshop blog can be found here.
.
After some wrestling with the Zenodo system of uploading, my paper “Sir David Cox’s Statistical Philosophy and its Relevance to Today’s Statistical Controversies” is now published (open access) in the JSM 2023 Proceedings(link).
Abstract
I discuss Sir David Cox’s views of the nature and importance of statistical foundations and their relevance to today’s controversies about statistical inference, particularly in using statistical significance tests. A central theme in Cox’s statistical philosophy is the importance of calibrating methods by considering their behavior in (actual or hypothetical) repeated sampling. Two key questions are open to philosophical controversy:
How can the frequentist calibration be used as an evidential or inferential assessment? How can we ensure that the hypothetical long-run used in calibration is relevant to the specific data?
I will discuss the answers that emerge from Cox’s work and our jointly written papers, Mayo and Cox (2006) and Cox and Mayo (2010) on statistical significance testing, objectivity in statistics, and conditioning.
Key Words: calibration, conditioning, Sir David R. Cox, statistical foundations, statistical philosophy, statistical significance tests
Cox tended to take Fisher’s side in the Fisher-Neyman wars. However, when we worked on Mayo and Cox 2006, he was open-minded enough to read some applied papers of Neyman that deviated from the usual caricatures. This is reflected somewhat in Mayo and Cox 2006, and in a footnote in his 2006 Principles of Statistical Inference (CUP), which I only stumbled across a few years ago:
Section 3.4. The contrast made here between the calculation of p-values as measures of evidence of consistency and the more decision-focused emphasis on accepting and rejecting hypotheses might be taken as one characteristic difference between the Fisherian and the Neyman-Pearson formulations of statistical theory. While this is in some respects the case, the actual practice in specific applications as between Fisher and Neyman was almost the reverse. Neyman often in effect reported p-values whereas some of Fisher’s use of tests in applications was much more dichotomous. For a discussion of the notion of severity of tests, and the circumstances when consistency with H0 might be taken as positive support for H0, see Mayo (1996). (43-4)
Please share your questions and comments in the Comments to this blog post.
Reference (with link):
Mayo, D.G. (2023). Sir David Cox’s Statistical Philosophy and its Relevance to Today’s Statistical Controversies. JSM 2023 Proceedings, DOI: https://zenodo.org/records/10028243.
I. My first post-pandemic movie. Wearing our Barbie shirts, purchased for the occasion, my friend Billie and I* went to see the movie Barbie the other day (open caption—a great idea!).[0] It was quite funny and clever, surprisingly introspective, and self-critical—even though I think it tried a tad bit too hard to remind us it […]
Today is Sir David Cox’s birthday. He would have been 99 today. 2023 marks the first year that the David R. Cox Award in Foundations of Statistics will be given at the upcoming Joint Statistical Meetings (JSM) in Toronto. For information on the Award, see this post. I’m excited to announce the inaugural winner, Nancy […]
A fraudster’s “first time”: I was alone in my tastefully furnished office at the University. . . . I opened the file with the data that I had entered and changed an unexpected 2 into a 4; then, a little further along, I changed a 3 into a 5. . . . When the results […]
A fraudster’s “first time”: I was alone in my tastefully furnished office at the University. . . . I opened the file with the data that I had entered and changed an unexpected 2 into a 4; then, a little further along, I changed a 3 into a 5. . . . When the results […]
I’ve been reading an illuminating paper by Georgi Gardiner and Brian Zaharatos (Gardiner and Zaharatos, 2022; hereafter, G & Z), “The safe, the sensitive and the severely tested,” that forges links between contemporary epistemology and my severe testing account. It’s part of a collection published in Synthese on “Recent issues in Philosophy of Statistics”. Gardiner […]
Link to announcement on ASA website.
First Winner.
Nancy Reid
University of Toronto
For contributions to the foundations of statistics that significantly advanced the frontiers of statistics and for insight that transformed understanding of parametric statistical inference, Nancy Reid is the inaugural recipient of the David R. Cox Foundations of Statistics Award, presented by the American Statistical Association (ASA). Reid will formally receive the award and deliver a lecture at the Joint Statistical Meetings in Toronto in August.
Reid, University Professor of Statistical Sciences at the University of Toronto, co-authored with David Cox an influential 1987 J. Roy. Statist. Soc. B discussion paper entitled “Parameter orthogonality and approximate conditional inference.” With this paper, and subsequent work with Cox and others, Reid has made major contributions to higher-order inference and various aspects of conditioning.
In addition to her work in foundational areas of statistics, Reid has successfully pursued numerous other lines of research, contributing to experimental design, nonparametric statistics, robust statistics, and comparisons and contradictions between Bayesian and frequentist inference.
The David R. Cox Foundations of Statistics Award was created in 2022 through an endowment created by Deborah G. Mayo, Professor Emerita of Philosophy at Virginia Tech. The ASA presents the award in odd-numbered years. The recipient receives a $2000 honorarium and is invited to give a lecture at the Joint Statistical Meetings.See announcement on ASA website.
Anyone interested in supporting the ASA’s effort to increase the size of the award through donations that may be matched by the donor and friends of David Cox is encouraged to contact Ron Wasserstein, Executive Director, ron@amstat.org.
About the AwardThis award honors the contributions of David R. Cox to the foundations of statistical inference, experimental design, and data analysis. It was established in 2022 to promote research and teaching that illuminates conceptual, theoretical, philosophical, and historical perspectives on statistical science and to advance understanding of comparative approaches to the interpretation and communication of statistical data.
The honoree will receive a $2,000 honorarium and up to $1,000 toward travel expenses to present a lecture at JSM.
Selection CriteriaThe award will be for a paper, monograph, book, or cumulative research. Anyone who has made noteworthy contributions to statistics in the spirit of Cox’s contributions as outlined above may be nominated.
Award Recipient ResponsibilitiesThe award recipient is responsible for providing a current photograph and general personal information the year the award is presented. The American Statistical Association uses this information to publicize the award and prepare the check and certificate.
NominationsNominations are due by December 1 and require the following:
QuestionsPlease contact the committee chair.
Note: The David Cox Foundations of Statistics Award will be given biennially to start, but will be given annually when funding is sufficient. Interested in contributing to this award? Contact ASA Director of Development Amanda Malloy.
This post is open for comments and questions by all zoom and class attendees on the presentations by Aris Spanos, Richard Morey or Deborah Mayo regarding the last three sessions of Mayo’s Phil 6014 PhilStat Seminar.
Slides for Session 9 on Testing Assumptions of Statistical Models and Misspecification testing (A. Spanos’ are here).
Slides from Session 10 on Bayes Factors (R. Morey’s slides can be found here). R. Morey also has a blog post with more details on his view at this link.)
Mayo’s slides are up on the syllabus which is here.
Please use the “Leave a comment” link below.
We had a good group zooming into the first half of my seminar on March 1. I’m grateful to them for their interest. They (and anyone else who cares to) are invited to post questions for me, or other thoughts, using the comments to this post. Any new people who want to observe the March 15 session (on statistical debates in particle physics) should write to me. March 22 and 29 will have Aris Spanos and Richard Morey as guest speakers, respectively. The syllabus is here, and the questions/exercises over spring break are here.
The reading from this session is from Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (Mayo, CUP, 2018)
D. Mayo
Tour I Ingenious and Severe Tests
[T]he impressive thing about [the 1919 tests of Einstein’s theory of gravity] is the risk involved in a prediction of this kind. If observation shows that the predicted effect is definitely absent, then the theory is simply refuted.The theory is incompatible with certain possible results of observation – in fact with results which everybody before Einstein would have expected. This is quite different from the situation I have previously described, [where] . . . it was practically impossible to describe any human behavior that might not be claimed to be a verification of these [psychological] theories. (Popper 1962, p. 36)
Mayo 2018, CUP
The 1919 eclipse experiments opened Popper’ s eyes to what made Einstein’ s theory so different from other revolutionary theories of the day: Einstein was prepared to subject his theory to risky tests.[1] Einstein was eager to galvanize scientists to test his theory of gravity, knowing the solar eclipse was coming up on May 29, 1919. Leading the expedition to test GTR was a perfect opportunity for Sir Arthur Eddington, a devout follower of Einstein as well as a devout Quaker and conscientious objector. Fearing “ a scandal if one of its young stars went to jail as a conscientious objector,” officials at Cambridge argued that Eddington couldn’ t very well be allowed to go off to war when the country needed him to prepare the journey to test Einstein’ s predicted light deflection (Kaku 2005, p. 113).
The museum ramps up from Popper through a gallery on “ Data Analysis in the 1919 Eclipse” (Section 3.1) which then leads to the main gallery on origins of statistical tests (Section 3.2). Here’ s our Museum Guide:
According to Einstein’ s theory of gravitation, to an observer on earth, light passing near the sun is deflected by an angle, λ , reaching its maximum of 1.75″ for light just grazing the sun, but the light deflection would be undetectable on earth with the instruments available in 1919. Although the light deflection of stars near the sun (approximately1 second of arc) would be detectable, the sun’ s glare renders such stars invisible, save during a total eclipse, which “ by strange good fortune” would occur on May 29, 1919 (Eddington [1920] 1987, p. 113).
There were three hypotheses for which “ it was especially desired to discriminate between” (Dyson et al. 1920 p. 291). Each is a statement about a parameter, the deflection of light at the limb of the sun (in arc seconds): λ = 0″ (no deflection), λ = 0.87″ (Newton), λ = 1.75″ (Einstein). The Newtonian predicted deflection stems from assuming light has mass and follows Newton’ s Law of Gravity. The difference in statistical prediction masks the deep theoretical differences in how each explains gravitational phenomena. Newtonian gravitation describes a force of attraction between two bodies; while for Einstein gravitational effects are actually the result of the curvature of spacetime. A gravitating body like the sun distorts its surrounding spacetime, and other bodies are reacting to those distortions.
Where Are Some of the Members of Our Statistical Cast of Characters in 1919? In 1919, Fisher had just accepted a job as a statistician at Rothamsted Experimental Station. He preferred this temporary slot to a more secure offer by Karl Pearson (KP), which had so many strings attached – requiring KP to approve everything Fisher taught or published – that Joan Fisher Box writes: After years during which Fisher “ had been rather consistently snubbed” by KP, “It seemed that the lover was at last to be admitted to his lady’ s court – on conditions that he first submit to castration” (J. Box 1978, p. 61). Fisher had already challenged the old guard. Whereas KP, after working on the problem for over 20 years, had only approximated “the first two moments of the sample correlation coefficient; Fisher derived the relevant distribution, not just the first two moments” in 1915 (Spanos 2013a). Unable to fight in WWI due to poor eyesight, Fisher felt that becoming a subsistence farmer during the war, making food coupons unnecessary, was the best way for him to exercise his patriotic duty.
In 1919, Neyman is living a hardscrabble life in a land alternately part of Russia or Poland, while the civil war between Reds and Whites is raging. “It was in the course of selling matches for food” (C. Reid 1998, p. 31) that Neyman was first imprisoned (for a few days) in 1919. Describing life amongst “roaming bands of anarchists, epidemics” (ibid., p. 32), Neyman tells us,“existence” was the primary concern (ibid., p. 31). With little academic work in statistics, and “ since no one in Poland was able to gauge the importance of his statistical work (he was ‘sui generis,’ as he later described himself)” (Lehmann 1994, p. 398), Polish authorities sent him to University College in London in 1925/1926 to get the great Karl Pearson’ s assessment. Neyman and E. Pearson begin work together in 1926. Egon Pearson, son of Karl, gets his B.A. in 1919, and begins studies at Cambridge the next year, including a course by Eddington on the theory of errors. Egon is shy and intimidated, reticent and diffi dent, living in the shadow of his eminent father, whom he gradually starts to question after Fisher’ s criticisms. He describes the psychological crisis he’ s going through at the time Neyman arrives in London: “ I was torn between conflicting emotions: a. finding it difficult to understand R.A.F., b. hating [Fisher] for his attacks on my paternal ‘ god,’ c. realizing that in some things at least he was right” (C. Reid 1998, p. 56). As far as appearances amongst the statistical cast: there are the two Pearsons: tall, Edwardian, genteel; there’ s hardscrabble Neyman with his strong Polish accent and small, toothbrush mustache; and Fisher: short, bearded, very thick glasses, pipe, and eight children. Let’ s go back to 1919, which saw Albert Einstein go from being a little known German scientist to becoming an international celebrity.
…To read further, see Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (CUP, 2018)
Where you are in the journey:
Excursion 3: Statistical Tests and Scientific Inference
Tour I Ingenious and Severe Tests 119
YOU
3.1 Statistical Inference and Sexy Science: The 1919
Eclipse Test 121
3.2 N-P Tests: An Episode in Anglo-Polish Collaboration 131
3.3 How to Do All N-P Tests D (and more) While
a Member of the Fisherian Tribe 146
(we only covered portions of this)
17 Feb 1890-29 July 1962
Today is R.A. Fisher’s birthday! I am reblogging what I call the “Triad”–an exchange between Fisher, Neyman and Pearson (N-P) published 20 years after the Fisher-Neyman break-up. My seminar on PhilStat is studying these this week, so it’s timely. While my favorite is still the reply by E.S. Pearson, which alone should have shattered Fisher’s allegations that N-P “reinterpret” tests of significance as “some kind of acceptance procedure”, all three are chock full of gems for different reasons. They are short and worth rereading. Neyman’s article pulls back the cover on what is really behind Fisher’s over-the-top polemics, what with Russian 5-year plans and commercialism in the U.S. Not only is Fisher jealous that N-P tests came to overshadow “his” tests, he is furious at Neyman for driving home the fact that Fisher’s fiducial approach had been shown to be inconsistent (by others). The flaw is illustrated by Neyman in his portion of the triad. I discuss this briefly in my Philosophy of Science Association paper from a few months ago (slides are here*).Further details may be found in my book, SIST (2018) especially pp 388-392 linked to here. It speaks to a common fallacy seen every day in interpreting confidence intervals. As for Neyman’s “behaviorism”, Pearson’s last sentence is revealing.
HAPPY BIRTHDAY R.A. FISHER!
*Slides from Glymour and J. Berger’s presentations are also there.
“Statistical Methods and Scientific Induction“
by Sir Ronald Fisher (1955)
SUMMARY
The attempt to reinterpret the common tests of significance used in scientific research as though they constituted some kind of acceptance procedure and led to “decisions” in Wald’s sense, originated in several misapprehensions and has led, apparently, to several more.
The three phrases examined here, with a view to elucidating they fallacies they embody, are:
Mathematicians without personal contact with the Natural Sciences have often been misled by such phrases. The errors to which they lead are not only numerical.
To continue reading Fisher’s paper.
“Note on an Article by Sir Ronald Fisher“
by Jerzy Neyman (1956)
Neyman
Summary
(1) FISHER’S allegation that, contrary to some passages in the introduction and on the cover of the book by Wald, this book does not really deal with experimental design is unfounded. In actual fact, the book is permeated with problems of experimentation. (2) Without consideration of hypotheses alternative to the one under test and without the study of probabilities of the two kinds, no purely probabilistic theory of tests is possible.
(3) The conceptual fallacy of the notion of fiducial distribution rests upon the lack of recognition that valid probability statements about random variables usually cease to be valid if the random variables are replaced by their particular values. The notorious multitude of “paradoxes” of fiducial theory is a consequence of this oversight. (4) The idea of a “cost function for faulty judgments” appears to be due to Laplace, followed by Gauss.
E.S. Pearson
“Statistical Concepts in Their Relation to Reality“.
by E.S. Pearson (1955)
Controversies in the field of mathematical statistics seem largely to have arisen because statisticians have been unable to agree upon how theory is to provide, in terms of probability statements, the numerical measures most helpful to those who have to draw conclusions from observational data. We are concerned here with the ways in which mathematical theory may be put, as it were, into gear with the common processes of rational thought, and there seems no reason to suppose that there is one best way in which this can be done. If, therefore, Sir Ronald Fisher recapitulates and enlarges on his views upon statistical methods and scientific induction we can all only be grateful, but when he takes this opportunity to criticize the work of others through misapprehension of their views as he has done in his recent contribution to this Journal (Fisher 1955 “Statistical Methods and Scientific Induction” ), it is impossible to leave him altogether unanswered.
In the first place it seems unfortunate that much of Fisher’s criticism of Neyman and Pearson’s approach to the testing of statistical hypotheses should be built upon a “penetrating observation” ascribed to Professor G.A. Barnard, the assumption involved in which happens to be historically incorrect. There was no question of a difference in point of view having “originated” when Neyman “reinterpreted” Fisher’s early work on tests of significance “in terms of that technological and commercial apparatus which is known as an acceptance procedure”. There was no sudden descent upon British soil of Russian ideas regarding the function of science in relation to technology and to five-year plans. It was really much simpler–or worse. The original heresy, as we shall see, was a Pearson one!…
Use this link to continue reading, “Statistical Concepts in Their Relation to Reality“.
My Phil Stat seminar has been meeting for 4 weeks now, and we’re soon to experiment with a small group of outside participants zooming in (write to us, if you are interested in joining us). I’ve been so busy with the seminar that I haven’t blogged. Have you been following? All the materials are on a continually updated syllabus on this blog (SYLLABUS). We’re up to Excursion 2, Tour II.
Last week, we did something unusual: we read from Popper’s Conjectures and Refutations. I wanted to do this because scientists often appeal to distorted and unsophisticated accounts of Popper, especially in discussing falsification, and what demarcates good science from poor science. While I don’t think Popper made good on his most winning slogans, he gives us many seminal launching-off points for improved accounts of falsification, induction, corroboration, and demarcation.
Do people still assume EI is “rational”? Good science can’t be demarcated from poor, questionable, fringe science and the like by its empirical method, if that method is understood as enumerative induction (EI), says Popper, rightly. While it comes in many forms, EI (taken up in Ex 2 Tour I) infers from observed instances (or frequencies) of A’s that are B’s to inferring claims like: the next A will be a B, or most A’s are B’s, or k% of A’s are B’s or even that the probability an A is B is k. Such a method is unreliable, so we shouldn’t be keen to justify it. It permits inferring poorly probed claims and violates the minimal requirement for evidence (weak severity).
Yet we are familiar with claims from epistemologists and others that some version of (EI) is a “rational” method. (It is the basis for famous quandries in legal reasoning). An additional stipulation is generally something like: “nothing else is known,” (which itself is knowing something else), but even that does not help. Neither do claims about indifference or uninformativeness. The philosopher Carnap called EI “the straight rule” and tried for many years to justify it–unsuccessfully. Lack of randomness, biasing selection effects in generating the data and in the choice of reference classes are key issues. Although the data in EI may be seen as relative frequencies, it is very different from frequentist statistics (See SIST, 110-11 on Neyman (1955): “Statistics as the Frequentist Theory of Induction”.)
Popper also rejected the empiricist assumption that observations are known relatively unproblematically. If they are at the “foundation,” it is only because there are apt methods for testing their validity. In fact, we dub claims observable because or to the extent that they are open to stringent checks. (Popper: “anyone who has learned the relevant technique can test it” (1959, p. 99).) Accounts of hypothesis appraisal that start with “evidence x,” as in confirmation logics, vastly oversimplify how data enters in learning.
Demarcation and Investigating Bad Science. Popper’s right that if using enumerative induction (EI) makes you scientific then anyone from an astrologer to one who blithely moves from observed associations to full blown theories is scientific. Yet Popper’s criterion of testability and falsifiability – as it is typically understood – may be nearly as bad. It is both too strong and too weak. Any crazy theory found false would be scientific, and our most impressive theories are not deductively falsifiable. The only theories that deductively prohibit observations are of the sort one mainly finds in philosophy books: All swans are white is falsified by a single non-white swan. There are some statistical claims and contexts, I argue, where it’s possible to achieve deductive falsification: claims such as, these data are independent and identically distributed (IID). Going beyond a mere denial to reliably replacing them, of course, requires more work.
However, interesting claims about mechanisms and causal generalizations require numerous assumptions (substantive and statistical) and are rarely open to deductive falsification. Their tests can be reconstructed as deductively valid, but in order to warrant the premises requires evidence-transcending (ampliative) inferences. So there’s a “whiff of induction” even in Popper (as some of his critics claim), even though not of the crude (EI) sort. (Note Popper’s claim about when a statistical hypothesis is falsified below.)
“The Demise of the Demarcation Problem”. Forty years ago, Larry Laudan’s famous (1983) paper declared the demarcation problem taboo. This is a highly unsatisfactory situation for philosophers of science wishing to grapple with today’s statistical replication crisis. Laudan and I generally see eye to eye, so perhaps our disagreement here is just semantics. I share his view that what really matters is determining if a hypothesis is warranted or not, rather than whether the theory is “scientific,” but surely Popper didn’t mean logical falsifiability sufficed. Popper is clear that many unscientific theories (e.g., Marxism, astrology) are falsifiable. It’s clinging to falsified theories that leads to unscientific practices. It’s trying and trying again in the face of unwelcome results, cherry-picking cases that support preferred hypotheses, and all the rest of the biases that make it easy to find apparent support for poorly probed claims.
Following Laudan, philosophers tend to shy away from saying anything general about science versus pseudoscience – the predominant view is that there is no such thing. One gets the impression that the demarcation task is being left to committees investigating allegations of poor science or fraud. They are forced to articulate what to count as fraud, as bad statistics, or as mere questionable research practices (QRPs). People’s careers depend on their ruling: they have “skin in the game,” as Nassim Nicholas Taleb might say (2018).
Free of the qualms that give philosophers of science cold feet, the committee investigating fraudster Diederik Stapel advance some obvious, yet crucially important rules with Popperian echoes:
One of the most fundamental rules of scientific research is that an investigation must be designed in such a way that facts that might refute the research hypotheses are given at least an equal chance of emerging as do facts that confirm the research hypotheses. (Levelt Committee, Noort Committee, and Drenth Committee 2012).
This is the gist of our minimal requirement for evidence (weak severity principle). To scrutinize the scientific credentials of an inquiry is to determine if there was a serious attempt to detect and report mistaken interpretations of data.
Demarcating Inquiries (4 requirements). However, I say Popper confuses things by making it sound as if he’s asking: When is a theory unscientific? What he is actually asking or should be asking is: When is an inquiry into a theory, or an appraisal of claim H, unscientific? We want to distinguish meritorious modes of inquiry from those that are BENT. Despite being logically falsifiable, theories can be rendered immune from falsification by means of cavalier methods for their testing. Some areas have so much noise and/or flexibility that they can’t or won’t distinguish warranted from unwarranted explanations of failed predictions. It does not suffice– for an inquiry to be scientific– that there is criticism of methods and models. The criticism must be constrained by what’s actually responsible for any alleged problems. It may be correct to criticize an inference to a hypothesis H, but it may be for the wrong reason. For instance, the problem might be traced to H’s improbability when in fact the flaw is due to lack of error control due to data-dredging, optional stopping, and P-hacking.
A scientific inquiry or test must be able:
(a) to block inferences that fail the minimal requirement for severity
(b) to embark on a reliable probe to pinpoint blame for anomalies
(c) (from (a)) to directly pick up on altered error probing capacities due to biasing selection effects, optional stopping, cherry picking, data-dredging etc.
(d) (from (b)) to test and falsify claims.
So we get four requirements for an inquiry to be scientific.
Methodological probability. A valuable idea to take from Popper is that probability in learning attaches to a method: it is methodological probability. An error probability is a special case of a methodological probability.
Popper wrote to me expressing regret that he didn’t learn more statistics, but he referred to Fisher, Neiman and Pearson, and also Pierce in explaining when a statistical hypothesis is to count as falsified. Although extremely rare events may occur, Popper notes:
such occurrences would not be physical effects, because, on account of their immense improbability, they are not reproducible at will … If, however, we find reproducible deviations from a macro effect .. . deduced from a probability estimate … then we must assume that the probability estimate is falsified. (Popper 1959, p. 203)
In the same vein, we heard Fisher deny that an “isolated record” of statistically significant results suffices to warrant a reproducible or genuine effect (Fisher 1935a, p. 14). Even where a scientific hypothesis is thought to be deterministic, inaccuracies and knowledge gaps involve error-laden predictions; so our methodological rules typically involve inferring a statistical hypothesis. Popper calls it a falsifying hypothesis. It’s a hypothesis inferred in order to falsify some other claim. A first step is often to infer an anomaly is real, by falsifying a “due to chance” hypothesis. That is the role of statistical significance tests.
Insofar as we falsify general scientific claims, we are all methodological falsificationists. Some people say, “I know my models are false, so I’m done with the job of falsifying before I even begin.” Really? That’s not falsifying. Let’s look at your method: always infer that H is false, or fails to solve its intended problem. Then you’re bound to infer this even when this is erroneous. (Were H a null hypothesis of “no effect” you’d always be inferring the effect is genuine.) Your method fails the minimal severity requirement.
PHIL 6014 (crn: 20919): Spring 2023
Philosophy of Inductive-Statistical Inference
(This is an IN-PERSON class)Wed 4:00-6:30 pm, McBryde 22
There may be opportunities for zooming half-way through the semester.
Syllabus: First Installment
D. Mayo (2018) Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (CUP, 2018): SIST
Several articles will be provided.
| Date | Themes/readings | | 1. 1/18 | Introduction to the Course: How to tell what’s true about statistical inference Reading: SIST: Preface, Excursion 1 Tour I 1.1-1.3, 9-29MISC: NOTES on Excursion 1, Souvenir A Postcard to Send, SIST: Abstracts & Keywords | | 2. 1/25 (R-1) | Error Probing Tools vs Comparative Evidence: Likelihood & ProbabilityWhat counts as cheating?Intro to Logic: arguments validity & soundnessReading: SIST: Excursion 1 Tour II 1.4-1.5, 30-55(R-1) Questions for Session 2: (PDF) MISC: Souvenir B Likelihood versus Error Statistical, Souvenir C A Severe Tester’s Translation Guide, Souvenir D Why We Are So New, | | 3. 2/1 | PhilStat & Formal EpistemologyThe Traditional Problem of InductionIs Probability a Good Measure of Confirmation? Tacking ParadoxReading: SIST: Excursion 2, Tour I: 2.1-2.2, 59-74 Optional: Hawthorne and Fitelson (2004)MISC: Excursion 2 Tour I Blurb & notes | | 4. 2/8 & 5. 2/15 | Falsification, Science vs Pseudoscience, InductionStatistical Crises of Replication in Psychology & other sciences Reading for 2/8: Popper, Ch 1 from Conjectures and Refutations*, Popper TestReading for 2/15: SIST: Excursion 2, Tour II: 2.3-2.7, pages TBA Popper, severity and novelty, array of problems and modelsFallacies of rejection, Duhem’s problem; solving induction nowMISC: SIST Souvenirs (F) Getting Free of Popperian Constraints on Language, (G) The Current State of Play in Psychology, (H) Solving Induction Is Showing Methods with Error ControlExcursioon 2 Tour II Blurb & notes | | Fisher Birthday: February 17: Celebration of N-F wars | | 6. 2/22 & 7. 3/1 | *Ingenious and Severe Tests: Fisher, Neyman-Pearson, Cox: Concepts of TestsReading SIST: Excursion 3 Tour I: 3.1-3.3, 119-163 (trade-offs 328-330)the 1919 eclipse tests; Fisherian and N-P Tests; Frequentist principle of evidence: FEV Apps for statistical testing MISC: Excursion 3 Tour I Blurb & notes |
SPRING BREAK Statistical Exercises While Sunning (March 4-12)
The following is very tentative, and will depend on student interests.
| 8. 3/15 Assign 2 | Confidence & Fiducial Intervals and Deeper Concepts: Higgs Discovery | | 9. 3/22 | Objectivity in Science: Objectivity in Error Statistics & Bayesian Philosophies | | 10. 3/29 Short essay | Bayes factors and Bayes/Fisher Disagreement | | 11. 4/5 | Biasing Selection Effects, P-Hacking, Data Dredging etc. | | 12. 4/12 Assign 3 | Negative Results: Power vs Severity | | 13. 4/19 | Should Statistical Significance Tests be Abandoned, Retired, or Replaced? | | 14. 4/26 | Severity, Sensitivity, Safety: PhilStat and Classical Epistemology | | 15. 5/3 | Current Reforms and Stat Activism: Practicing Our Skills | | Final Paper |
Ship StatInfasst (Statistical Inference as Severe Testing: SIST) will set sail on Wednesday January 18 when I begin a weekly seminar on the Philosophy of Inductive-statistical inference. I’m planning to write a new edition and/or companion to SIST (Mayo 2018, CUP), so it will be good to retrace the journey. I’m not requiring a statistics or philosophy background. All materials will be on this blog, and around halfway through there may be an opportunity to zoom, if there’s interest.
One of the central roles I proposed for “stat activists” (after our recent workshop, The Statistics Wars and Their Casualties) is to critically scrutinize mistaken claims about leading statistical methods–especially when such claims are put forward as permissible viewpoints to help “the people” assess methods in an unbiased manner. The first act of 2023 under this umbrella concerns an article put forward as “statistics for the people” in a journal of radiation oncology. We are talking here about recommendations for analyzing data for treating cancer! Put forward as a fair-minded, or at least an informative, comparison of Bayesian vs frequentist methods, I find it to be little more than an advertisement for subjective Bayesian methods in favor of a caricature of frequentist error statistical methods. The journal’s “statistics for the people” section would benefit from a full-blown article on frequentist error statistical methods–not just the letter of ours they recently published–but I’m grateful to Chowdhry and other colleagues who joined me in this effort. You will find our letter below, followed by the authors’ response. You can also find a link to their original “statistics for the people” article in the references. Let me admit right off that my criticisms are a bit stronger than my co-authors.
Two quick additional things that I would wish to tell the authors in relation to their paper and response are:
I would never have come across an article in radiation oncology, if it were not for exchanges between members of a session I was in on “why we disagree” in statistical analysis in that field. I hereby invite all readers and the nearly 1000 registrants from our workshop to alert us throughout the year of interesting items under any of the stat activist banner.
Our letter: Bayesian Versus Frequentist Statistics: In Regard to Fornacon-Wood et al. (PDF of letter)
To the Editor:
We appreciate the authors bringing attention to controversies surrounding the use of Bayesian and frequentist statistics.1 [PDF of paper] There are many benefits to frequentist statistics and disadvantages of Bayesian statistics which were not discussed in the referenced article. We write this accompanying letter to aim for a more balanced presentation of Bayesian and frequentist statistics.
With frequentist statistical significance tests, we can learn whether the data indicate there is a genuine effect or difference in a statistical analysis, as they have the ability to control type I and type II error probabilities.2 Posteriors and Bayes factors do not ensure that the method rarely reports one treatment is better or worse than the other erroneously. A well-known threat to reliable results stems from the ease of using high powered methods to data-dredge and try to hunt for impressive-looking results that fail to replicate with new data. However, the Bayesian assessment is not altered by things like stopping rules-at least not without violating inference by Bayes theorem.3 The frequentist account,4 by contrast, is required to take account of such selection effects in reporting error probabilities. Another caution for those unfamiliar with practical Bayesian research is that estimation of a prior distribution is nontrivial. The priors they discuss are subjective degrees of belief, but there is considerable disagreement about which beliefs are warranted, even among experts. Furthermore, should conclusions differ if the prior is chosen by a radiation oncologist or a surgeon?5 These considerations are some of the reasons why most phase 3 studies in oncology rely on frequentist designs. The article equates frequentist methods with simple null hypothesis testing without alternatives, thereby overlooking hypothesis testing methods that control both type I and II errors. The frequentist takes account of type II errors and the corresponding notion of power. If a test has high power to detect a meaningful effect size, then failing to detect a statistically significant difference is evidence against a meaningful effect. Therefore, a value that is not small is informative.
The authors write that frequentist methods do not use background information, but this is to ignore the field of experimental design and all of the work that goes into specifying the test (eg, sample size, statistical power) and critically evaluating the connection between statistical and substantive results. An effect that corresponds to a clinically meaningful effect, or effect sizes well warranted from previous studies, would clearly influence the design.
Although their article engenders important discussion, these differences between frequentist and Bayesian methods may help readers understand why so many researchers around the world still prefer the frequentist approach.
References
Fornacon-Wood Reply: In Reply to Chowdhry et al. (PDF of letter)
To the Editor:
We thank the authors for their response to our “statistics for the people” article that aimed to introduce perhaps unfamiliar readers to Bayesian statistics and some potential advantages of their use. We agree that frequentist statistics are a useful and widespread statistical analytical approach, and we are not aiming to revisit the frequentist versus Bayesian arguments that have been well articulated in the literature. However, there are a couple of points we would like to make.
First, we acknowledge that the majority of phase 3 studies use frequentist designs, and this has the advantage of facilitating meta-analyses using established techniques. However, we would argue that the reason such frequentist designs are so prevalent is likely to have as much to do with convention (from funders/regulators as well as from researchers themselves), the relative exposure of the 2 approaches in educational materials, and the historic difficulties in calculating Bayesian posteriors as it does with the arguments the authors make.
Second, although we agree with Chowdhry et al that there are many challenges associated with the estimation of prior probability distributions, we note that similar arguments apply to effect size estimation, which they cite as a strength of the Neyman-Pearson/null hypothesis significance testing approach (ie, the use of power calculations to limit the risk of type II errors). We would also re-enforce the point we make in the article about the importance of testing the influence of the prior (represented as the divergent beliefs of the hypothetical radiation oncologist and surgeon in the communication by Chowdhry et al) in the analysis results. If the data are strong enough, the posterior distributions will be in close enough agreement to convince both parties. As we noted, it is also possible to undertake Bayesian analyses without prior information, using an uninformative prior, in which case the analysis is driven directly by the data, as for a frequentist calculation. As an aside, there is continued debate about the relative merits and deficiencies of the different frequentist approaches to significance testing, particularly around the widespread use of the hybrid Neyman-Pearson/null hypothesis significance testing approach.
Please share your constructive remarks in the comments to this post.
https://doi.org/10.1016/j.ijrobp.2022.08.034
.
For the last three years, unlike the previous 10 years that I’ve been blogging, it was not feasible to actually revisit that spot in the road, looking to get into a strange-looking taxi, to head to “Midnight With Birnbaum”. But this year I will, and I’m about to leave at 10pm. (The pic on the left is the only blurry image I have of the club I’m taken to.) My book Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (CUP, 2018) doesn’t include the argument from my article in Statistical Science (“On the Birnbaum Argument for the Strong Likelihood Principle”), but you can read it at that link along with commentaries by A. P. David, Michael Evans, Martin and Liu, D. A. S. Fraser, Jan Hannig, and Jan Bjornstad. David Cox, who very sadly did in January 2022, is the one who encouraged me to write and publish it. (The first David R. Cox Foundations of Statistics Prize will be awarded at the JSM 2023.) The (Strong) Likelihood Principle (LP or SLP) remains at the heart of many of the criticisms of Neyman-Pearson (N-P) statistics and of error statistics in general.
As Birnbaum emphasized, the “confidence concept” is the “one rock in a shifting scene” of statistical foundations, insofar as there’s interest in controlling the frequency of erroneous interpretations of data. (See my rejoinder to commentators.) Birnbaum bemoaned the lack of an explicit evidential interpretation of N-P methods. I purport to give one in SIST 2018. Anyway, let’s see what happens this year. Happy New Year!
BACKGROUND
You know how in that Woody Allen movie, “Midnight in Paris,” the main character (I forget who plays it, I saw it on a plane) is a writer finishing a novel, and he steps into a cab that mysteriously picks him up at midnight and transports him back in time where he gets to run his work by such famous authors as Hemingway and Virginia Wolf? (It was a new movie when I began the blog in 2011.) He is wowed when his work earns their approval and he comes back each night in the same mysterious cab…Well, imagine an error statistical philosopher is picked up in a mysterious taxi at midnight on New Year’s Eve and lo and behold, finds herself in the company of Allan Birnbaum.[i]
OUR EXCHANGE:
ERROR STATISTICIAN: It’s wonderful to meet you Professor Birnbaum; I’ve always been extremely impressed with the important impact your work has had on philosophical foundations of statistics. I happen to have published on your famous argument about the likelihood principle (LP). (whispers: I can’t believe this!)
BIRNBAUM: Ultimately you know I rejected the LP as failing to control the error probabilities needed for my Confidence concept. But you know all this, I’ve read it in your book: Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (STINT, 2018, CUP).
ERROR STATISTICIAN: You’ve read my book? Wow! Then you know I don’t think your argument shows that the LP follows from such frequentist concepts as sufficiency S and the weak conditionality principle WLP. I don’t rehearse my argument there, but I first found the problem in 2006, when I was writing something on “conditioning” with David Cox. [ii] Sorry,…I know it’s famous…
BIRNBAUM: Well, I shall happily invite you to take any case that violates the LP and allow me to demonstrate that the frequentist is led to inconsistency, provided she also wishes to adhere to the WLP and sufficiency (although less than S is needed).
ERROR STATISTICIAN: Well I show that no contradiction follows from holding WCP and S, while denying the LP.
BIRNBAUM: Well, well, well: I’ll bet you a bottle of Elba Grease champagne that I can demonstrate it!
ERROR STATISTICAL PHILOSOPHER: It is a great drink, I must admit that: I love lemons.
BIRNBAUM: OK. (A waiter brings a bottle, they each pour a glass and resume talking). Whoever wins this little argument pays for this whole bottle of vintage Ebar or Elbow or whatever it is Grease.
.
ERROR STATISTICAL PHILOSOPHER: I really don’t mind paying for the bottle.
BIRNBAUM: Good, you will have to. Take any LP violation. Let x’ be 2-standard deviation difference from the null (asserting μ = 0) in testing a normal mean from the fixed sample size experiment E’, say n = 100; and let x” be a 2-standard deviation difference from an optional stopping experiment E”, which happens to stop at 100. Do you agree that:
(0) For a frequentist, outcome x’ from E’ (fixed sample size) is NOT evidentially equivalent to x” from E” (optional stopping that stops at n)
ERROR STATISTICAL PHILOSOPHER: Yes, that’s a clear case where we reject the strong LP, and it makes perfect sense to distinguish their corresponding p-values (which we can write as p’ and p”, respectively). The searching in the optional stopping experiment makes the p-value quite a bit higher than with the fixed sample size. For n = 100, data x’ yields p’= ~.05; while p” is ~.3. Clearly, p’ is not equal to p”, I don’t see how you can make them equal.
BIRNBAUM: Suppose you’ve observed x”, a 2-standard deviation difference from an optional stopping experiment E”, that finally stops at n=100. You admit, do you not, that this outcome could have occurred as a result of a different experiment? It could have been that a fair coin was flipped where it is agreed that heads instructs you to perform E’ (fixed sample size experiment, with n = 100) and tails instructs you to perform the optional stopping experiment E”, stopping as soon as you obtain a 2-standard deviation difference, and you happened to get tails, and performed the experiment E”, which happened to stop with n =100.
ERROR STATISTICAL PHILOSOPHER: Well, that is not how x” was obtained, but ok, it could have occurred that way.
BIRNBAUM: Good. Then you must grant further that your result could have come from a special experiment I have dreamt up, call it a BB-experiment. In a BB-experiment, if the outcome from the experiment you actually performed has an outcome with a proportional likelihood to one in some other experiment not performed, E’, then we say that your result has an “LP pair”. For any violation of the strong LP, the outcome observed, let it be x”, has an “LP pair”, call it x’, in some other experiment E’. In that case, a BB-experiment stipulates that you are to report x” as if you had determined whether to run E’ or E” by flipping a fair coin.
(They fill their glasses again)
ERROR STATISTICAL PHILOSOPHER: You’re saying that if my outcome from trying and trying again, that is, optional stopping experiment E”, with an “LP pair” in the fixed sample size experiment I did not perform, then I am to report x” as if the determination to run E” was by flipping a fair coin (which decides between E’ and E”)?
BIRNBAUM: Yes, and one more thing. If your outcome had actually come from the fixed sample size experiment E’, it too would have an “LP pair” in the experiment you did not perform, E”. Whether you actually observed x” from E”, or x’ from E’, you are to report it as x” from E”.
ERROR STATISTICAL PHILOSOPHER: So let’s see if I understand a Birnbaum BB-experiment: whether my observed 2-standard deviation difference came from E’ or E” (with sample size n) the result is reported as x’, as if it came from E’ (fixed sample size), and as a result of this strange type of a mixture experiment.
BIRNBAUM: Yes, or equivalently you could just report x*: my result is a 2-standard deviation difference and it could have come from either E’ (fixed sampling, n= 100) or E” (optional stopping, which happens to stop at the 100th trial). That’s how I sometimes formulate a BB-experiment.
ERROR STATISTICAL PHILOSOPHER: You’re saying in effect that if my result has an LP pair in the experiment not performed, I should act as if I accept the strong LP and just report it’s likelihood; so if the likelihoods are proportional in the two experiments (both testing the same mean), the outcomes are evidentially equivalent.
BIRNBAUM: Well, but since the BB- experiment is an imagined “mixture” it is a single experiment, so really you only need to apply the weak LP which frequentists accept. Yes? (The weak LP is the same as the sufficiency principle).
ERROR STATISTICAL PHILOSOPHER: But what is the sampling distribution in this imaginary BB- experiment? Suppose I have Birnbaumized my experimental result, just as you describe, and observed a 2-standard deviation difference from optional stopping experiment E”. How do I calculate the p-value within a Birnbaumized experiment?
BIRNBAUM: I don’t think anyone has ever called it that.
ERROR STATISTICAL PHILOSOPHER: I just wanted to have a shorthand for the operation you are describing, there’s no need to use it, if you’d rather I not. So how do I calculate the p-value within a BB-experiment?
BIRNBAUM: You would report the overall p-value, which would be the average over the sampling distributions: (p’ + p”)/2
Say p’ is ~.05, and p” is ~.3; whatever they are, we know they are different, that’s what makes this a violation of the strong LP (given in premise (0)).
ERROR STATISTICAL PHILOSOPHER: So you’re saying that if I observe a 2-standard deviation difference from E’, I do not report the associated p-value p’, but instead I am to report the average p-value, averaging over some other experiment E” that could have given rise to an outcome with a proportional likelihood to the one I observed, even though I didn’t obtain it this way?
BIRNBAUM: I’m saying that you have to grant that x’ from a fixed sample size experiment E’ could have been generated through a BB-experiment.
My this drink is sour!
ERROR STATISTICAL PHILOSOPHER: Yes, I love pure lemon.
BIRNBAUM: Perhaps you’re in want of a gene; never mind.
I’m saying you have to grant that x’ from a fixed sample size experiment E’ could have been generated through a BB-experiment. If you are to interpret your experiment as if you are within the rules of a BB experiment, then x’ is evidentially equivalent to x” (is equivalent to x*). This is premise (1).
ERROR STATISTICAL PHILOSOPHER: But the result would be that the p-value associated with x’ (fixed sample size) is reported to be larger than it actually is (.05), because I’d be averaging over fixed and optional stopping experiments; while observing x” (optional stopping) is reported to be smaller than it is–in both cases because of an experiment I did not perform.
BIRNBAUM: Yes, the BB-experiment computes the P-value in an unconditional manner: it takes the convex combination over the 2 ways the result could have come about.
ERROR STATISTICAL PHILOSOPHER: this is just a matter of your definitions, it is an analytical or mathematical result, so long as we grant being within your BB experiment.
BIRNBAUM: True, (1) plays the role of the sufficiency assumption, but one need not even appeal to sufficiency, it is just a matter of mathematical equivalence.
By the way, I am focusing just on LP violations, therefore, the outcome, by definition, has an LP pair. In other cases, where there is no LP pair, you just report things as usual.
ERROR STATISTICAL PHILOSOPHER: OK, but p’ still differs from p”; so I still don’t how I’m forced to infer the strong LP which identifies the two. In short, I don’t see the contradiction with my rejecting the strong LP in premise (0). (Also we should come back to the “other cases” at some point….)
BIRNBAUM: Wait! Don’t be so impatient; I’m about to get to step (2). Here, let’s toast to the new year: “To Elbar Grease!”
ERROR STATISTICAL PHILOSOPHER: To Elbar Grease!
BIRNBAUM: So far all of this was step (1).
ERROR STATISTICAL PHILOSOPHER: : Oy, what is step 2?
BIRNBAUM: STEP 2 is this: Surely, you agree, that once you know from which experiment the observed 2-standard deviation difference actually came, you ought to report the p-value corresponding to that experiment. You ought NOT to report the average (p’ + p”)/2 as you were instructed to do in the BB experiment.
This gives us premise (2a):
(2a) outcome x”, once it is known that it came from E”, should NOT be analyzed as in a BB- experiment where p-values are averaged. The report should instead use the sampling distribution of the optional stopping test E”, yielding the p-value, p” (~.37). In fact, .37 is the value you give in STINT p. 44 (imagining the experimenter keeps taking 10 more).
ERROR STATISTICAL PHILOSOPHER: So, having first insisted I imagine myself in a Birnbaumized, I mean a BB-experiment, and report an average p-value, I’m now to return to my senses and “condition” in order to get back to the only place I ever wanted to be, i.e., back to where I was to begin with?
BIRNBAUM: Yes, at least if you hold to the weak conditionality principle WCP (of D. R. Cox)—surely you agree to this.
(2b) Likewise, if you knew the 2-standard deviation difference came from E’, then
x’ should NOT be deemed evidentially equivalent to x” (as in the BB experiment), the report should instead use the sampling distribution of fixed test E’, (.05).
ERROR STATISTICAL PHILOSOPHER: So, having first insisted I consider myself in a BB-experiment, in which I report the average p-value, I’m now to return to my senses and allow that if I know the result came from optional stopping, E”, I should “condition” on and report p”.
BIRNBAUM: Yes. There was no need to repeat the whole spiel.
ERROR STATISTICAL PHILOSOPHER: I just wanted to be clear I understood you. Of course, all of this assumes the model is correct or adequate to begin with.
BIRNBAUM: Yes, the LP (or SLP, to indicate it’s the strong LP) is a principle for parametric inference within a given model. So you arrive at (2a) and (2b), yes?
ERROR STATISTICAL PHILOSOPHER: OK, but it might be noted that unlike premise (1), premises (2a) and (2b) are not given by definition, they concern an evidential standpoint about how one ought to interpret a result once you know which experiment it came from. In particular, premises (2a) and (2b) say I should condition and use the sampling distribution of the experiment known to have been actually performed, when interpreting the result.
BIRNBAUM: Yes, and isn’t this weak conditionality principle WCP one that you happily accept?
ERROR STATISTICAL PHILOSOPHER: Well the WCP originally refers to actual mixtures, where one flipped a coin to determine if E’ or E” is performed, whereas, you’re requiring I consider an imaginary Birnbaum mixture experiment, where the choice of the experiment not performed will vary depending on the outcome that needs an LP pair; and I cannot even determine what this might be until after I’ve observed the result that would violate the LP? I don’t know what the sample size will be ahead of time.
BIRNBAUM: Sure, but you admit that your observed x” could have come about through a BB-experiment, and that’s all I need. Notice
(1), (2a) and (2b) yield the strong LP!
Outcome x” from E”(optional stopping that stops at n) is evidentially equivalent to x’ from E’ (fixed sample size n).
ERROR STATISTICAL PHILOSOPHER: Clever, but your “proof” is obviously unsound; and before I demonstrate this, notice that the conclusion, were it to follow, asserts p’ = p”, (e.g., .05 = .3!), even though it is unquestioned that p’ is not equal to p”, that is because we must start with an LP violation (premise (0)).
BIRNBAUM: Yes, it is puzzling, but where have I gone wrong?
(The waiter comes by and fills their glasses; they are so deeply engrossed in thought they do not even notice him.)
ERROR STATISTICAL PHILOSOPHER: There are many routes to explaining a fallacious argument. Here’s one. What is required for STEP 1 to hold, is the denial of what’s needed for STEP 2 to hold:
Step 1 requires us to analyze results in accordance with a BB- experiment. If we do so, true enough we get:
premise (1): outcome x” (in a BB experiment) is evidentially equivalent to outcome x’ (in a BB experiment):
That is because in either case, the p-value would be (p’ + p”)/2
Step 2 now insists that we should NOT calculate evidential import as if we were in a BB- experiment. Instead we should consider the experiment from which the data actually came, E’ or E”:
premise (2a): outcome x” (in a BB experiment) is/should be evidentially equivalent to x” from E” (optional stopping that stops at n): its p-value should be p”.
premise (2b): outcome x’ (within in a BB experiment) is/should be evidentially equivalent to x’ from E’ (fixed sample size): its p-value should be p’.
If (1) is true, then (2a) and (2b) must be false!
If (1) is true and we keep fixed the stipulation of a BB experiment (which we must to apply step 2), then (2a) is asserting:
The average p-value (p’ + p”)/2 = p’ which is false.
Likewise if (1) is true, then (2b) is asserting:
the average p-value (p’ + p”)/2 = p” which is false
Alternatively, we can see what goes wrong by realizing:
If (2a) and (2b) are true, then premise (1) must be false.
In short your famous argument requires us to assess evidence in a given experiment in two contradictory ways: as if we are within a BB- experiment (and report the average p-value) and also that we are not, but rather should report the actual p-value.
I can render it as formally valid, but then its premises can never all be true; alternatively, I can get the premises to come out true, but then the conclusion is false—so it is invalid. In no way does it show the frequentist is open to contradiction (by dint of accepting S, WCP, and denying the LP).
BIRNBAUM: Yet some people still think it is a breakthrough. I never agreed to go as far as Jimmy Savage wanted me too, namely, to be a Bayesian….
ERROR STATISTICAL PHILOSOPHER: I have a much clearer exposition of what goes wrong in your argument than I did in the discussion from 2010. There were still several gaps, and lack of a clear articulation of the WCP. In fact, I’ve come to see that clarifying the entire argument turns on defining the WCP. Have you seen my 2014 paper in Statistical Science? The key difference is that in (2014), the WCP is stated as an equivalence, as you intended. Cox’s WCP, many claim, was not an equivalence, going in 2 directions. Slides from a presentation may be found on this blogpost.
Birnbaum: Yes, the “monster of the LP” arises from viewing WCP as an equivalence, instead of going in one direction (from mixtures to the known result).
ERROR STATISTICAL PHILOSOPHER: In my 2014 paper (unlike my earlier treatments) I too construe WCP as giving an “equivalence” but there is an equivocation that invalidates the purported move to the LP.
On the one hand, it’s true that if z is known (and known for example to have come from optional stopping), it’s irrelevant that it could have resulted from either fixed sample testing or optional stopping.
But it does not follow that if z is known, it’s irrelevant whether it resulted from fixed sample testing or optional stopping. It’s the slippery slide into this second statement–which surely sounds the same as the first–that makes your argument such a brain buster.
BIRNBAUM: Yes I have seen your 2014 paper! Your Rejoinder to some of the critics is gutsy, to say the least. I’ve also seen the slides on your blog.
ERROR STATISTICAL PHILOSOPHER: Thank you, I’m amazed you follow my blog! But look I must get your answer to a question before you leave this year.
Sudden interruption by the waiter who, very wisely, is wearing an N95 mask:
WAITER: Who gets the tab? We’re closing a bit early due to Covid.
BIRNBAUM: I do. To Elbar Grease! And to your (still) new book SIST! I have a list of comments and questions right here.
ERROR STATISTICAL PHILOSOPHER: Let me see, I’d love to read your questions and comments. (She takes a long legal-sized yellow sheet from Birnbaum, noticing it is filled with tiny hand-written comments, covering both sides.)
BIRNBAUM: To Elbar Grease! To Severe Testing! Happy New Year!
ERROR STATISTICAL PHILOSOPHER: I have one quick question, Professor Birnbaum, and I swear that whatever you say will be just between us, I won’t tell a soul. In your last couple of papers, you suggest you’d discovered the flaw in your argument for the LP. Am I right? Even in the discussion of your (1962) paper, you seemed to agree with Pratt that WCP can’t do the job you intend.
BIRNBAUM: Savage, you know, never got off my case about remaining at “the half-way house” of likelihood, and not going full Bayesian. Then I wrote the review about the Confidence Concept as the one rock on a shifting scene… Pratt thought the argument should instead appeal to a Censoring Principle (basically, it doesn’t matter if your instrument cannot measure beyond k units if the measurement you’re making is under k units.)
ERROR STATISTICAL PHILOSOPHER: Yes, but who says frequentist error statisticians deny the Censoring Principle? So back to my question, you disappeared before answering last year…I just want to know…you did see the flaw, yes?
WAITER: We’re closing now; shall I call Remote Taxi?
BIRNBAUM: Yes, yes!
ERROR STATISTICAL PHILOSOPHER: ‘Yes’, you discovered the flaw in the argument, or ‘yes’ to the taxi?
MANAGER: We’re closing now; I’m sorry you must leave.
ERROR STATISTICAL PHILOSOPHER: We’re leaving I just need him to clarify his answer….
Large group of people bustle past, mostly unmasked.
Prof. Birnbaum…? Allan? Where did he go? (oy, not again!)
Link to complete discussion:
Mayo, Deborah G. On the Birnbaum Argument for the Strong Likelihood Principle (with discussion & rejoinder).Statistical Science 29 (2014), no. 2, 227-266.
[i] Many links on the strong likelihood principle (LP or SLP) and Birnbaum may be found by searching this blog. Good sources for where to start as well as classic background papers may be found in this blogpost. A link to slides and video of a very introductory presentation of my argument from the 2021 Phil Stat Forum is here.
January 7: “Putting the Brakes on the Breakthrough: On the Birnbaum Argument for the Strong Likelihood Principle” (D.Mayo)
[ii] By the way, Ronald Giere gave me numerous original papers of yours. They’re in files in my attic library. Some are in mimeo, others typed…I mean, obviously for that time that’s what they’d be…now of course, oh never mind, sorry.
Below are the videos and slides from the 7 talks from Session 3 and Session 4 of our workshop The Statistics Wars and Their Casualties held on December 1 & 8, 2022. Session 3 speakers were: Daniele Fanelli (London School of Economics and Political Science), Stephan Guttinger (University of Exeter), and David Hand (Imperial College London). Session 4 speakers were: Jon Williamson (University of Kent), Margherita Harris (London School of Economics and Political Science), Aris Spanos (Virginia Tech), and Uri Simonsohn (Esade Ramon Llull University). Abstracts can be found here. In addition to the talks, you’ll find (1) a Recap of recaps at the beginning of Session 3 that provides a summary of Sessions 1 & 2, and (2) Mayo’s (5 minute) introduction to the final discussion: “Where do we go from here (Part ii)”at the end of Session 4.
The videos & slides from Sessions 1 & 2 can be found on this post.
Readers are welcome to use the comments section on the PhilStatWars.com workshop blog post here to make constructive comments or to ask questions of the speakers. If you’re asking a question, indicate to which speaker(s) it is directed. We will leave it to speakers to respond. Thank you!
SESSION 3
Recap of recaps summary of Sessions 1 & 2:
Introduction to Session: Daniël Lakens (Eindhoven University of Technology)
Daniele Fanelli (London School of Economics and Political Science)
The neglected importance of complexity in statistics and Metascience
Stephan Guttinger (University of Exeter)
What are questionable research practices?
David Hand (Imperial College London)
What’s the question?
Discussion (Session 3): (a) Panel discussion of speakers; (b) general audience discussion; (c) “Where do we go from here (Part i)” participant discussion.
SESSION 4
Introduction to Session 4: Deborah Mayo (Virginia Tech)
Jon Williamson (University of Kent)
Causal inference is not statistical inference
Margherita Harris (London School of Economics and Political Science)
On Severity, the Weight of Evidence, and the Relationship Between the Two
Aris Spanos (Virginia Tech)
Revisiting the Two Cultures in Statistical Modeling and Inference as they relate to the Statistics Wars and Their Potential Casualties
Uri Simonsohn (Esade Ramon Llull University)
Mathematically Elegant Answers to Research Questions No One is Asking (meta-analysis, random effects models, and Bayes factors)
Where Should Stat Activists Go From Here? Deborah Mayo (Virginia Tech):
Discussion: (a) Panel discussions; (b) General audience discussion; (c) “Where do we go from here (Part ii)” participants and audience.
.
Below are slides from 4 of the talks given in our Philosophy of Science Association (PSA) session from last month: the PSA 22 Symposium: Multiplicity, Data-Dredging, and Error Control. It was held in Pittsburgh on November 13, 2022. I will write some reflections in the “comments” to this post. I invite your constructive comments there as well.
SYMPOSIUM ABSTRACT: High powered methods, the big data revolution, and the crisis of replication in medicine and social sciences have prompted new reflections and debates in both statistics and philosophy about the role of traditional statistical methodology in current science. Experts do not agree on how to improve reliability, and these disagreements reflect philosophical battles–old and new– about the nature of inductive-statistical evidence and the roles of probability in statistical inference. We consider three central questions:
•How should we cope with the fact that data-driven processes, multiplicity and selection effects can invalidate a method’s control of error probabilities?
•Can we use the same data to search non-experimental data for causal relationships and also to reliably test them?
•Can a method’s error probabilities both control a method’s performance as well as give a relevant epistemological assessment of what can be learned from data?
As reforms to methodology are being debated, constructed or (in some cases) abandoned, the time is ripe to bring the perspectives of philosophers of science (Glymour, Mayo, Mayo-Wilson) and statisticians (Berger, Thornton) to reflect on these questions.
.
Deborah Mayo (Philosophy, Virginia Tech)
Error Control and Severity
ABSTRACT: I put forward a general principle for evidence: an error-prone claim C is warranted to the extent it has been subjected to, and passes, an analysis that very probably would have found evidence of flaws in C just if they are present. This probability is the severity with which C has passed the test. When a test’s error probabilities quantify the capacity of tests to probe errors in C, I argue, they can be used to assess what has been learned from the data about C. A claim can be probable or even known to be true, yet poorly probed by the data and model at hand. The severe testing account leads to a reformulation of statistical significance tests: Moving away from a binary interpretation, we test several discrepancies from any reference hypothesis and report those well or poorly warranted. A probative test will generally involve combining several subsidiary tests, deliberately designed to unearth different flaws. The approach relates to confidence interval estimation, but, like confidence distributions (CD) (Thornton), a series of different confidence levels is considered. A 95% confidence interval method, say using the mean M of a random sample to estimate the population mean μ of a Normal distribution, will cover the true, but unknown, value of μ 95% of the time in a hypothetical series of applications. However, we cannot take .95 as the probability that a particular interval estimate (a ≤ μ ≤ b) is correct—at least not without a prior probability to μ. In the severity interpretation I propose, we can nevertheless give an inferential construal post-data, while still regarding μ as fixed. For example, there is good evidence μ ≥ a (the lower estimation limit) because if μ < a, then with high probability .95 (or .975 if viewed as one-sided) we would have observed a smaller value of M than we did. Likewise for inferring μ ≤ b. To understand a method’s capability to probe flaws in the case at hand, we cannot just consider the observed data, unlike in strict Bayesian accounts. We need to consider what the method would have inferred if other data had been observed. For each point μ’ in the interval, we assess how severely the claim μ > μ’ has been probed. I apply the severity account to the problems discussed by earlier speakers in our session. The problem with multiple testing (and selective reporting) when attempting to distinguish genuine effects from noise, is not merely that it would, if regularly applied, lead to inferences that were often wrong. Rather, it renders the method incapable, or practically so, of probing the relevant mistaken inference in the case at hand. In other cases, by contrast, (e.g., DNA matching) the searching can increase the test’s probative capacity. In this way the severe testing account can explain competing intuitions about multiplicity and data-dredging, while blocking inferences based on problematic data-dredging.
.
Suzanne Thornton (Statistics, Swarthmore College)
The Duality of Parameters and the Duality of Probability
ABSTRACT: Under any inferential paradigm, statistical inference is connected to the logic of probability. Well-known debates among these various paradigms emerge from conflicting views on the notion of probability. One dominant view understands the logic of probability as a representation of variability (frequentism), and another prominent view understands probability as a measurement of belief (Bayesianism). The first camp generally describes model parameters as fixed values, whereas the second camp views parameters as random. Just as calibration (Reid and Cox 2015, “On Some Principles of Statistical Inference,” International Statistical Review 83(2), 293-308)–the behavior of a procedure under hypothetical repetition–bypasses the need for different versions of probability, I propose that an inferential approach based on confidence distributions (CD), which I will explain, bypasses the analogous conflicting perspectives on parameters. Frequentist inference is connected to the logic of probability through the notion of empirical randomness. Sample estimates are useful only insofar as one has a sense of the extent to which the estimator may vary from one random sample to another. The bounds of a confidence interval are thus particular observations of a random variable, where the randomness is inherited by the random sampling of the data. For example, 95% confidence intervals for parameter θ can be calculated for any random sample from a Normal N(θ, 1) distribution. With repeated sampling, approximately 95% of these intervals are guaranteed to yield an interval covering the fixed value of θ. Bayesian inference produces a probability distribution for the different values of a particular parameter. However, the quality of this distribution is difficult to assess without invoking an appeal to the notion of repeated performance. For data observed from a N(θ, 1) distribution to generate a credible interval for θ requires an assumption about the plausibility of different possible values of θ, that is, one must assume a prior. However, depending on the context – is θ the recovery time for a newly created drug? or is θ the recovery time for a new version of an older drug? – there may or may not be an informed choice for the prior. Without appealing to the long-run performance of the interval, how is one to judge a 95% credible interval [a, b] versus another 95% interval [a’, b’] based on the same data but a different prior? In contrast to a posterior distribution, a CD is not a probabilistic statement about the parameter, rather it is a data-dependent estimate for a fixed parameter for which a particular behavioral property holds. The Normal distribution itself, centered around the observed average of the data (e.g. average recovery times), can be a CD for θ. It can give any level of confidence. Such estimators can be derived through Bayesian or frequentist inductive procedures, and any CD, regardless of how it is obtained, guarantees performance of the estimator under replication for a fixed target, while simultaneously producing a random estimate for the possible values of θ.
.
Clark Glymour (Philosophy, Carnegie Mellon University)
Good Data-Dredging**
ABSTRACT: “Data dredging”–searching non experimental data for causal and other relationships and taking that same data to be evidence for those relationships–was historically common in the natural sciences–the works of Kepler, Cannizzaro and Mendeleev are examples. Nowadays, “data dredging”–using data to bring hypotheses into consideration and regarding that same data as evidence bearing on their truth or falsity–is widely denounced by both philosophical and statistical methodologists. Notwithstanding, “data dredging” is routinely practiced in the human sciences using “traditional” methods–various forms of regression for example. The main thesis of my talk is that, in the spirit and letter of Mayo’s and Spanos’ notion of severe testing, modern computational algorithms that search data for causal relations severely test their resulting models in the process of “constructing” them. My claim is that in many investigations, principled computerized search is invaluable for reliable, generalizable, informative, scientific inquiry. The possible failures of traditional search methods for causal relations, multiple regression for example, are easily demonstrated by simulation in cases where even the earliest consistent graphical model search algorithms succeed. In real scientific cases in which the number of variables is large in comparison to the sample size, principled search algorithms can be indispensable. I illustrate the first claim with a simple linear model, and the second claim with an application of the oldest correct graphical model search, the PC algorithm, to genomic data followed by experimental tests of the search results. The latter example, due to Steckhoven et al. (“Causal Stability Ranking,” Bioinformatics, 28 (21), 2819-2823) involves identification of (some of the) genes responsible for bolting in A. thaliana from among more than 19,000 coding genes using as data the gene expressions and time to bolting from only 47 plants. I will also discuss Fast Causal Inference (FCI) which gives asymptotically correct results even in the presence of confounders. These and other examples raise a number of issues about using multiple hypothesis tests in strategies for severe testing, notably, the interpretation of standard errors and confidence levels as error probabilities when the structures assumed in parameter estimation are uncertain. Commonly used regression methods, I will argue, are bad data dredging methods that do not severely, or appropriately, test their results. I argue that various traditional and proposed methodological norms, including pre-specification of experimental outcomes and error probabilities for regression estimates of causal effects, are unnecessary or illusory in application. Statistics wants a number, or at least an interval, to express a normative virtue, the value of data as evidence for a hypothesis, how well the data pushes us toward the true or away from the false. Good when you can get it, but there are many circumstances where you have evidence but there is no number or interval to express it other than phony numbers with no logical connection with truth guidance. Kepler, Darwin, Cannizarro, Mendeleev had no such numbers, but they severely tested their claims by combining data dredging with severe testing.
.
James Berger (Statistics, Duke University)
Comparing Frequentists and Bayesian Control of Multiple Testing
ABSTRACT: A problem that is common to many sciences is that of having to deal with a multiplicity of statistical inferences. For instance, in GWAS (Genome Wide Association Studies), an experiment might consider 20 diseases and 100,000 genes, and conduct statistical tests of the 20×100,000=2,000,000 null hypotheses that a specific disease is associated with a specific gene. The issue is that selective reporting of only the ‘highly significant’ results could lead to many claimed disease/gene associations that turn out to be false, simply because of statistical randomness. In 2007, the seriousness of this problem was recognized in GWAS and extremely stringent standards were employed to resolve it. Indeed, it was recommended that tests for association should be conducted at an error probability of 5 x 10—7. Particle physicists similarly learned that a discovery would be reliably replicated only if the p-value of the relevant test was less than 5.7 x 10—7. This was because they had to account for a huge number of multiplicities in their analyses. Other sciences have continuing issues with multiplicity. In the Social Sciences, p-hacking and data dredging are common, which involve multiple analyses of data. Stopping rules in social sciences are often ignored, even though it has been known since 1933 that, if one keeps collecting data and computing the p-value, one is guaranteed to obtain a p-value less than 0.05 (or, indeed, any specified value), even if the null hypothesis is true. In medical studies that occur with strong oversight (e.g., by the FDA), control for multiplicity is mandated. There is also typically a large amount of replication, resulting in meta-analysis. But there are many situations where multiplicity is not handled well, such as subgroup analysis: one first tests for an overall treatment effect in the population; failing to find that, one tests for an effect among men or among women; failing to find that, one tests for an effect among old men or young men, or among old women or young women; …. I will argue that there is a single method that can address any such problems of multiplicity: Bayesian analysis, with the multiplicity being addressed through choice of prior probabilities of hypotheses. In GWAS, scientists assessed the chance of a disease/gene association to be 1/100,000, meaning that each null hypothesis of no association would be assigned a prior probability of 1-1/100,000. Only tests yielding p-values less than 5 x 10—7 would be able to overcome this strong initial belief in no association. In subgroup analysis, the set of possible subgroups under consideration can be expressed as a tree, with probabilities being assigned to differing branches of the tree to deal with the multiplicity. There are, of course, also frequentist error approaches (such as Bonferroni and FDR) for handling multiplicity of statistical inferences; indeed, these are much more familiar than the Bayesian approach. These are, however, targeted solutions for specific classes of problems and are not easily generalizable to new problems.
Thursday, December 8 will be the Final Session (Session 4) of my workshop, The Statistics Wars and Their Casualties. There will be 4 new speakers. It’s not too late to register: registration form
It’s not too late to register for Sessions #3 and #4 of our online Workshop. There will be 7 new (live) speakers and, for the the first time ever, the (short) movie; “The Recap of recaps” will be shown at the start of session #3. registration form
The Statistics Wars and Their Casualties 1 December and 8 December 2022 Sessions #3 and #4 15:00-18:15 pm London Time/10:00am-1:15pm EST ONLINE (London School of Economics, CPNSS) registration form For slides and videos of Sessions #1 and #2: see the workshop page 1 December Session 3 (Moderator: Daniël Lakens, Eindhoven University of Technology) OPENING “What Happened […]
Stephen SennConsultant StatisticianEdinburgh, Scotland A Diet of Terms A large university is interested in investigating the effects on the students of the diet provided in the university dining halls and any sex difference in these effects. Various types of data are gathered. In particular, the weight of each student at the time of his arrival […]
Some claim that no one attends Sunday morning (9am) sessions at the Philosophy of Science Association. But if you’re attending the PSA (in Pittsburgh), we hope you’ll falsify this supposition and come to hear us (Mayo, Thornton, Glymour, Mayo-Wilson, Berger) wrestle with some rival views on the trenchant problems of multiplicity, data-dredging, and error control. […]
From what standpoint should we approach the statistics wars? That’s the question from which I launched my presentation at the Statistics Wars and Their Casualties workshop (phil-stat-wars.com). In my view, it should be, not from the standpoint of technical disputes, but from the non-technical standpoint of the skeptical consumer of statistics (see my slides here). […]
I will be writing some reflections on our two workshop sessions on this blog soon, but for now, here are just the slides I used on Thursday, 22 September. If you wish to ask a question of any of the speakers, use the blogpost at phil-stat-wars.com. The slides from the other speakers will also be […]
The Statistics Wars and Their Casualties Final Schedule for September 22 & 23 (Workshop Sessions 1 & 2) Session 1: September 22 Moderator: David Hand (Imperial College London) 3:00-3:10 (10:00-10:10): Deborah Mayo, Opening Remarks and Thanks 3:10-3:15 (10:10-10:15) Chair introduction to the session 3:15-3:50 (10:15-10:50): Deborah Mayo (Virginia Tech) The Statistics Wars and Their Causalities (Abstract) 3:50-4:25 (10:50-11:25): Richard Morey (Cardiff University) Bayes factors, p values, […]
You can still register: https://phil-stat-wars.com/2022/09/19/22-23-september-workshop-schedule-the-statistics-wars-and-their-casualties/
The Statistics Wars and Their Casualties 22-23 September 2022 15:00-18:00 pm London Time ONLINE (London School of Economics, CPNSS) To register for the workshop, please fill out the registration form here. For schedules and updated details, please see the workshop webpage: phil-stat-wars.com. These will be sessions 1 & 2, there will be two more online […]
Thanks to CUP, the electronic version of my book, Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (2018), is available for free for one more week (through August 31) at this link: https://www.cambridge.org/core/books/statistical-inference-as-severe-testing/D9DF409EF568090F3F60407FF2B973B2 Blurbs of the 16 tours in the book may be found here: blurbs of the 16 tours.
This is my third and final post marking Egon Pearson’s birthday (Aug. 11). The focus is his little-known paper: “Statistical Concepts in Their Relation to Reality” (Pearson 1955). I’ve linked to it several times over the years, but always find a new gem or two, despite its being so short. E. Pearson rejected some of […]
Continuing with posts on E.S. Pearson in marking his birthday, I reblog this guest post by Aris Spanos. Egon Pearson’s Neglected Contributions to Statistics by Aris Spanos Egon Pearson (11 August 1895 – 12 June 1980), is widely known today for his contribution in recasting of Fisher’s significance testing into the Neyman-Pearson (1933) theory of […]
This is a belated birthday post for E.S. Pearson (11 August 1895-12 June, 1980)–one of my statistical heroes. It’s basically a post from 2012 which concerns an issue of interpretation (long-run performance vs probativeness) that’s badly confused these days. Yes, I know I’ve been neglecting this blog as of late, because I’m busy planning our […]
CUP will make the electronic version of my book, Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (2018), available to access for free from August 1-31 at this link: https://www.cambridge.org/core/books/statistical-inference-as-severe-testing/D9DF409EF568090F3F60407FF2B973B2 However, they will confirm the link closer to August, so check this blog on Aug 1 for any update, if you’re […]
The Statistics Wars and Their Casualties 22-23 September 2022 15:00-18:00 pm London Time ONLINE To register/receive notification of updates for the workshop, please fill out the registration/notification form here. These will be sessions 1 & 2, there will be two more The future online sessions (3 & 4) to be announced. Yoav Benjamini (Tel Aviv […]
HAPPY BIRTHDAY SIR DAVID COX! Today is David Cox’s birthday, he would have been 98 years old today. Below is a remembrance I contributed to Significance when he died, with a link to others in that same issue. “In celebrating Cox’s immense contributions, we should recognise how much there is yet to learn from him” […]
Today marks a decade since the discovery on July 4, 2012 of evidence for a Higgs particle based on a “5 sigma observed effect”. CERN celebrated with a scientific symposium (webcast here). The observed effect refers to the number of excess events of a given type that are “observed” in comparison to the number that […]
In what began as a guest commentary on my 2021 editorial in Conservation Biology, Daniël Lakens recently published a response to a recommendation against using null hypothesis significance tests by journal editors from the International Society of Physiotherapy Journal. Here are some excerpts from his full article, replies (‘response to Lakens‘), links and a few […]
The full, searchable SCOTUS dissent: https://www.politico.com/news/2022/06/24/read-supreme-court-dissent-opinion-on-roe-v-wade-pdf-00042264
Someone sent me an email the other day telling me that a disclaimer had been added to the editorial written by the ASA Executive Director and 2 co-authors (Wasserstein et al., 2019) (“Moving to a world beyond ‘p < 0.05′”). It reads: The editorial was written by the three editors acting as individuals and […]
I’ll be speaking at this conference in Philly tomorrow. My slides are also below. PDF of my slides: Statistical “Reforms”: Fixing Science or Threats to Replication and Falsification.
Prof. Deborah Mayo, Emerita Department of Philosophy Virginia Tech Prof. David Hand Department of Mathematics Statistics Imperial College London Statistical significance and its critics: practicing damaging science, or damaging scientific practice? (Synthese) [pdf of full paper.] Abstract: While the common procedure of statistical significance testing and its accompanying concept of p-values have long been […]
I had been posting commentaries daily from January 6, 2022 (on my editorial “The Statistics Wars and Intellectual conflicts of Interest”, Conservation Biology) until Sir David Cox died on January 18, at which point I switched to some memorial items. These two commentaries from what Daniell calls my ‘birthday festschrift’ were left out, and I […]
There are 3 commentaries soon to be published in Conservation Biology on my editorial, “The statistics wars and intellectual conflicts of interest” also published in Conservation Biology. Professor Philip B. Stark Department of Statistics University of California, Berkeley You can read a draft of Philip Stark’s commentary here Professor Christian Hennig Department […]
You will often hear that if you reach a just statistically significant result “and the discovery study is underpowered, the observed effects are expected to be inflated” (Ioannidis 2008, p. 64), or “exaggerated” (Gelman and Carlin 2014). This connects to what I’m referring to as the second set of concerns about statistical significance tests, power […]
The most surprising discovery about today’s statistics wars is that some who set out shingles as “statistical reformers” themselves are guilty of misdefining some of the basic concepts of error statistical tests—notably power. (See my recent post on power howlers.) A major purpose of my Statistical Inference as Severe Testing: How to Get Beyond the […]
Today is Jerzy Neyman’s birthday (April 16, 1894 – August 5, 1981). I’m reposting a link to a quirky, but fascinating, paper of his that explains one of the most misunderstood of his positions–what he was opposed to in opposing the “inferential theory”. The paper, fro 60 years ago,Neyman, J. (1962), ‘Two Breakthroughs in the […]
Suppose you are reading about a statistically significant result x that just reaches a threshold p-value α from a test T+ of the mean of a Normal distribution H0: µ ≤ 0 against H1: µ > 0 with n iid samples, and (for simplicity) known σ. The test “rejects” H0 at this level & infers evidence of a […]
One does not have evidence for a claim if little if anything has been done to rule out ways the claim may be false. The claim may be said to “pass” the test, but it’s one that utterly lacks stringency or severity. On the basis of this very simple principle, I build a notion of […]
The Statistics Wars and Their Casualties Postponed to 22-23 September 2022 London School of Economics (CPNSS) Yoav Benjamini (Tel Aviv University), Alexander Bird (University of Cambridge), Mark Burgman (Imperial College London), Daniele Fanelli (London School of Economics and Political Science), Roman Frigg (London School of Economics and Political Science), Stephen Guettinger (London School of Economics and […]
I’ve been reading about the artificial intelligence/machine learning (AI/ML) wars revolving around the use of so-called “black-box” algorithms–too complex for humans, even their inventors, to understand. Such algorithms are increasingly used to make decisions that affect you, but if you can’t understand, or aren’t told, why a machine predicted your graduate-school readiness, or which drug […]
PSA2022: Call for Contributed Papers https://psa2022.dryfta.com/ Twenty-Eighth Biennial Meeting of the Philosophy of Science AssociationNovember 10 – November 13, 2022Pittsburgh, Pennsylvania Submissions open on March 9, 2022 for contributed papers to be presented at the PSA2022 meeting in Pittsburgh, Pennsylvania, on November 10-13, 2022. The deadline for submitting a paper is 11:59 PM Pacific Standard […]
Here are all the slides along with the video from the 11 January Phil Stat Forum with speakers: Deborah G. Mayo, Yoav Benjamini and moderator/discussant David Hand. Y. Benjamini’s slides: “The ASA president Task Force Statement on Statistical Significance and Replicability“ SLIDE SHOW: Mayo slides are from the Editorial* in Conservation Biology: […]
Continuing with posts in recognition of R.A. Fisher’s birthday, I reblog (with a few new comments) one from a few years ago on a topic that had previously not been discussed on this blog: Fisher’s fiducial probability [Neyman and Pearson] “began an influential collaboration initially designed primarily, it would seem to clarify Fisher’s writing. This […]
In recognition of Fisher’s birthday (Feb 17), I reblog what I call the “Triad”–an exchange between Fisher, Neyman and Pearson (N-P) a full 20 years after the Fisher-Neyman break-up–adding a few new introductory remarks here. While my favorite is still the reply by E.S. Pearson, which alone should have shattered Fisher’s allegations that N-P “reinterpret” […]
Today is R.A. Fisher’s birthday. I’ll reblog some Fisherian items this week with a few new remarks. This paper comes just before the conflicts with Neyman and Pearson (N-P) erupted. Fisher links his tests and sufficiency, to the Neyman and Pearson lemma in terms of power. It’s as if we may see Fisher and N-P […]
Karen Kafadar, Yoav Benjamini, and Donald Macnaughton will be in a session: Should Science Abandon Statistical Significance? Friday, Feb 18 from 2-2:45 PM (EST) at the AAAS 2022 annual meeting. The general program is here. To register*, go to this page. Synopsis The concept of statistical significance is central in scientific research. However, the concept […]
Here are my slides on my Editorial in Conservation Biology: “The Statistics Wars and Intellectual Conflicts of Interest” Mayo (2021) presented at the 11 January Phil Stat Forum with speakers: Deborah G. Mayo and Yoav Benjamini and moderator David Hand. (Benjamini’s slides & full Video to come shortly) SLIDE SHOW: […]
Yesterday’s event video recording is available at:
https://www.youtube.com/watch?v=2mWYbcVflyE&t=10s
European Network for Business and Industrial Statistics (ENBIS) Webinar:
Statistical Significance and p-values
Thursday 3 Feb 2022, 14:00 → 15:30 Europe/Amsterdam (CET); 08:00-09:30 am (EST)
ENBIS will dedicate this webinar to the memory of Sir David Cox, who sadly passed away in January 2022.
This looks interesting! A bit early for me, but I plan to attend. I’m very glad they will dedicate it to Cox, and happy they haven’t abandoned the use of “statistical significance” in their title! You can find the announcement & link to register here: https://conferences.enbis.org/event/22/
In June 2011, Sir David Cox agreed to a very informal ‘interview’ on the topics of the 2010 workshop that I co-ran at the London School of Economics (CPNSS), Statistical Science and Philosophy of Science, where he was a speaker. Soon after I began taping, Cox stopped me in order to show me how to do a proper interview. He proceeded to ask me questions, beginning with:
COX: Deborah, in some fields foundations do not seem very important, but we both think foundations of statistical inference are important; why do you think that is?
MAYO: I think because they ask about fundamental questions of evidence, inference, and probability. I don’t think that foundations of different fields are all alike; because in statistics we’re so intimately connected to the scientific interest in learning about the world, we invariably cross into philosophical questions about empirical knowledge and inductive inference.
So we continued with this game, and never got back to my intended plan. But I managed to ask some of the intended questions of him while in the midst of his showing me how to do it. Here’s the link:
“A Statistical Scientist Meets a Philosopher of Science: A Conversation between Sir David Cox and Deborah Mayo”
Hinkley, Reid & Cox
Here’s an in-depth interview of Sir David Cox by Nancy Reid that brings out a rare, intellectual understanding and appreciation of some of Cox’s work. Only someone truly in the know could have managed to elicit these fascinating reflections. The interview was in Oct 1993, published in 1994.
Nancy Reid (1994). A Conversation with Sir David Cox, Statistical Science 9(3): 439-455.
Sir David Cox: July 15, 1924-Jan 18, 2022
The original Statistics Views interview is here:
“I would like to think of myself as a scientist, who happens largely to specialise in the use of statistics”– An interview with Sir David Cox
FEATURES
- Author: Statistics Views
- Date: 24 Jan 2014
Sir David Cox is arguably one of the world’s leading living statisticians. He has made pioneering and important contributions to numerous areas of statistics and applied probability over the years, of which perhaps the best known is the proportional hazards model, which is widely used in the analysis of survival data. The Cox point process was named after him.
Sir David studied mathematics at St John’s College, Cambridge and obtained his PhD from the University of Leeds in 1949. He was employed from 1944 to 1946 at the Royal Aircraft Establishment, from 1946 to 1950 at the Wool Industries Research Association in Leeds, and from 1950 to 1955 worked at the Statistical Laboratory at the University of Cambridge. From 1956 to 1966 he was Reader and then Professor of Statistics at Birkbeck College, London. In 1966, he took up the Chair position in Statistics at Imperial College Londonwhere he later became Head of the Department of Mathematics for a period. In 1988 he became Warden of Nuffield College and was a member of the Department of Statistics at Oxford University. He formally retired from these positions in 1994 but continues to work in Oxford.
Sir David has received numerous awards and honours over the years. He has been awarded the Guy Medals in Silver (1961) and Gold (1973) by the Royal Statistical Society. He was elected Fellow of the Royal Society of London in 1973, was knighted in 1985 and became an Honorary Fellow of the British Academy in 2000. He is a Foreign Associate of the US National Academy of Sciences and a foreign member of the Royal Danish Academy of Sciences and Letters. In 1990 he won the Kettering Prize and Gold Medal for Cancer Research for “the development of the Proportional Hazard Regression Model” and 2010 he was awarded the Copley Medal by the Royal Society.
He has supervised and collaborated with many students over the years, many of whom are now successful in statistics in their own right such as David Hinkley and Past President of the Royal Statistical Society, Valerie Isham. Sir David has served as President of theBernoulli Society, Royal Statistical Society, and the International Statistical Institute.
This year, Sir David is to turn 90*. Here Statistics Views talks to Sir David about his prestigious career in statistics, working with the late Professor Lindley, his thoughts on Jeffreys and Fisher, being President of the Royal Statistical Society during the Thatcher Years, Big Data and the best time of day to think of statistical methods.
1. With an educational background in mathematics at St Johns College, Cambridge and the University of Leeds, when and how did you first become aware of statistics as a discipline?
I was studying at Cambridge during the Second World War and after two years, one was sent either into the Forces or into some kind of military research establishment. There were very few statisticians then, although it was realised there was a need for statisticians. It was assumed that anybody who was doing reasonably well at mathematics could pick up statistics in a week or so! So, aged 20, I went to the Royal Aircraft Establishment in Farnborough, which is enormous and still there to this day if in a different form, and I worked in the Department of Structural and Mechanical Engineering, doing statistical work. So statistics was forced upon me, so to speak, as was the case for many mathematicians at the time because, aside from UCL, there had been very little teaching of statistics in British universities before the Second World War. Afterwards, it all started to expand.
2. From 1944 to 1946 you worked at the Royal Aircraft Establishment and then from 1946 to 1950 at the Wool Industries Research Association in Leeds. Did statistics have any role to play in your first roles out of university?
Totally. In Leeds, it was largely statistics but also to some extent, applied mathematics because there were all sorts of problems connected with the wool and textile industry in terms of the physics, chemistry and biology of the wool and some of these problems were mathematical but the great majority had a statistical component to them. That experience was not totally uncommon at the time and many who became academic statisticians had, in fact, spent several years working in a research institute first.
3. From 1950 to 1955, you worked at the Statistical Laboratory at Cambridge and would have been there at the same time as Fisher and Jeffreys. The late Professor Dennis Lindley, who was also there at that time, told me that the best people working on statistics were not in the statistics department at that time. What are your memories when you look back on that time and what do you feel were your main achievements?
Lindley was exactly right about Jeffreys and Fisher. They were two great scientists outside statistics – Jeffreys founded modern geophysics and Fisher was a major figure in genetics. Dennis was a contemporary and very impressive and effective. We were colleagues for five years and our children even played together.
The first lectures on statistics I attended as a student consisted of a short course by Harold Jeffreys who had at the time a massive reputation as virtually the inventor of modern geophysics. His Theory of Probability, published first as a monograph in physics was and remains of great importance but, amongst other things, his nervousness limited the appeal of his lectures, to put it gently. I met him personally a couple of times – he was friendly but uncommunicative. When I was later at the Statistical Laboratory in Cambridge, relations between the Director, Dr Wishart and R.A. Fisher had been at a very low ebb for 20 years and contact between the Lab and Fisher was minimal. I hear him speak on three of four occasions, interesting if often rambunctious occasions. To some, Fisher showed great generosity but not to the Statistics Lab, which was sad in view of the towering importance of his work.
“To some, Fisher showed great generosity but not to the Statistics Lab, which was sad in view of the towering importance of his work.”
4. You have also taught at many institutions over the years including Princeton, Berkeley, Cambridge, Birkbeck College and Imperial College London before joining Nuffield College here at Oxford. Over the years, how did the teaching of statistics evolve and adapt to meet the changing needs of students?
As I said, when I was a student, there was very little teaching of statistics in British universities. It has evolved over the years and was first primarily a postgraduate subject, taken after reading mathematics if you wished to be a scientific statistician, rather than an economic statistician. You took at least a diploma, or a one-year MA or a doctorate. Then statistics came into mathematics degrees, partly to make them more appealing to a wider audience and that has changed, so nowadays, most statisticians start fairly intensively in an undergraduate course, which has some advantages and some disadvantages.
5. How did your teaching and research motivated and influenced each other? Did you get research ideas from statistics and incorporate them into your teaching?
Much of my research has come from talking to scientists. Sometimes ideas come from lecturing because the way to really understand a subject is to give a course of lectures on the subject and sometimes that throws up more theoretical issues that you might not otherwise have been thought of. The overwhelming majority of my work comes either directly or indirectly from some physical biological or medical problem, but in many different ways – casual conversation sometimes.
6. You have taught many who have gone onto make their own important contributions towards statistics such as David Hinkley whom is now renowned for his work on bootstrap methods and Valerie Isham who recently served as the President of the Royal Statistical Society. The late Professor Dennis Lindley told me that “One of the joys of life is teaching a really good graduate.” Would you be in agreement?
I would say that one of the joys of life is learning from a good graduate. The first duty of a doctoral student is clearly is to educate their supervisor which my own doctoral students have down. Hopefully, they’ve learnt a bit from me occasionally! I am absolutely certain that I learnt a lot from Valerie, for instance, as we’ve worked together on and off for around forty years. Having such students is fantastic. I have been fortunate and happy as at Birkbeck, I had largely evening students. They were highly motivated and very able. Many of the graduate students at Imperial came from other places, or were international students of high standard. Also rather importantly, my students have been very nice people!
7. You are best known for your innovative work on the proportional hazards model, which is now widely used in the analysis of survival data. What research led to this discovery? What set you on the right path?
Two different things – first of all, I had been interested in reliability in an industrial context since I worked in the textile industry and to some extent, when I was at the Royal Aircraft Establishment, when strength of materials was important. I had a long interest in testing the strength and reliability which is also related to looking at the duration of life. Then the more specific thing was that at least four or five people from different areas in the US and the UK said that they had a certain kind of data with people’s survival times under various treatments and all sorts of further aspects with regards to the patient but they did not know how to analyse this data. The work led to one paper but the reason it is so popular is totally accidental. Other people wrote easily useful software in which to implement the method which is not my speciality at all. I had software to implement it but it was not suitable for general use. In a sense, it became almost too easy and so people just started to use the method because it was painless! The proportion of my life that I spent working on the proportional hazards model is, in fact, very small. I had an idea of how to solve it but I could not complete the argument and so it took me about four years on and off, often thinking about it before I went to bed.
(Editor’s note: I tell Sir David that I now had a picture in my head of him pacing the house in his pyjamas at four o’clock in the morning with a hot chocolate in one hand, thinking statistical thoughts and he laughs).
Not quite! It was right before going to bed. There is a well-established literature in mathematics that people who thought about a problem and do not know how to solve it, go to bed thinking about it and wake up the next morning with a solution. It’s not easily explicable but if you’re wide awake, you perhaps argue down the conventional lines of argument but what you need to do is something a bit crazy which you’re more likely to do if you’re half-awake or asleep. Presumably that’s the explanation!
“The proportion of my life that I spent working on the proportional hazards model is, in fact, very small. I had an idea of how to solve it but I could not complete the argument and so it took me about four years on and off…”
8. The getstats campaign by the Royal Statistical Society focuses on improving the public’s understanding of statistics in every-day life. Would you have any advice for them and what areas should they focus on that you feel there should be more awareness of in statistics?
Of course, to some extent, the notion that some very simple and non-technical ideas about collecting data and analysing it are taught to children is very good but then at the other end, there are people who are highly educated but have no sense of statistical arguments, such as many lawyers and senior civil servants. The RSS has done an excellent job in trying to interest MPs in statistical ideas. Both these extremes are important. Sending a very general message to people as far as possible helps, but also sending very focussed messages to key groups of people is more important in the short term. You do see on TV, for instance, that basic principles are being ignored in collecting and analysing evidence. Of course, it’s easy for me to stand on the sidelines and criticise.
9. You have served as the President for several societies over the years including the Royal Statistical Society, the Bernoulli Society and the International Statistical Institute. What are your memories of your time at the RSS for instance and how you helped the society adapt to the changing needs of the statistical community?
It was a bit different in my time. I was the President of the RSS at the time when Margaret Thatcher was PM and massacring the civil service and in particular, the government’s statistical service and there was a lot of activity going on about that. But it was done more by going to see people, talking to them and trying to influence them than writing them formal letters. While, of course, openness is a good thing, it is not always the best way to get results. People can take up inflexible attitudes but if you talk to them quietly in private, they are perhaps then more open to new ideas.
10. You have received numerous awards from the Guy Medal both in Silver and Gold to the Marvin Zelen Leadership Award. Is there a particular award that you were most proud of being awarded?
It would be the Copley Medal from the Royal Society as it was for general science. It is very nice to receive these awards, of course, perhaps particularly because they represent the fact that your friends have put in efforts on your behalf. Therefore, what I really value is not the award but the appreciation of friends and colleagues. That is what is important but the award and degrees are certainly an honour. If you overvalue an award, that can be dangerous.
11. You have written many papers and books. What are the ones that you are most proud of?
The one I’m going to write next, of course! I have flitted about all sorts of different topics different fields of application, different parts of the subject, and so on. Really, I don’t look back very much.
At the moment, I have just finished a book with a colleague called Case Control Studies which is mainly about epidemiological investigation.
“…what I really value is not the award but the appreciation of friends and colleagues.”
12. What has been the best book on statistics that you have ever read?
I honestly don’t know. The position of books is interesting as when I first started in my career, there were hardly any books at all that were treating statistics in a modern way. Then they very slowly began to appear and now there is a flood of them. The standard on the whole published now is very high but there is too much to keep up with!
13. What has been the most exciting development that you have worked on in statistics during your career?
I’m not sure about exciting (!) but one of the most demanding was being involved in the issues about Bovine TB in badgers, which went on for about ten years. It involved a great deal of work, which was very interesting and instructive in all sorts of ways, and not just in statistics.
I’ve been involved in other government-based topics, such as the group which made the first predictions for the AIDS epidemic, which was also very interesting.
14. At the recent Future of Statistical Sciences workshop, there was much talk about Big Data and a concern that many ‘hot areas’ such as big data/data analytics, which have close connections with statistics and the statistical sciences, are being monopolised by computer scientists and/or engineers. What do statisticians need to do to ensure their work and their profession gets noticed?
Do better quality work, which I don’t mean as a criticism as to what is done at the moment but rather, do high quality work that is important in some sense, either intellectually or practically in particular fields. Part of the problem is that relatively speaking, there are not that many statisticians who are trained to the level needed.
15. What do you think the most important recent developments in the field have been? What do you think will be the most exciting and productive areas of research in statistics during the next few years?
The most immediately important is as you said – Big Data, which will bring forward new ideas but it does not mean that old ideas from the more traditional part of the subject is useless. It is the most obvious and biggest challenge.
Ideally, we should be looking at very important practical problems in a different number of fields and see some sort of common element and build the ideas that are necessary in order to tackle any issues that arise. You should not tackle just one issue successfully but tackle a collection of issues – the Big Data aspect is one common theme undoubtedly. It goes beyond statistics – to what extent Big Data can replace small, carefully planned investigations which are much more sharply focussed on a very specific issue.
My intrinsic feeling is that more fundamental progress is more likely to be made by very focused, relatively small scale, intensive investigations than collecting millions of bits of information on millions of people, for example. It’s possible to collect such large data now, but it depends on the quality, which may be very high or not, and if it is not, what do you do about it?
16. Do you think over the years too much research has focussed on less important areas of statistics? Should the gap between research and applications get reduced? How so and by whom?
In British statistics at the moment, the gap between theory and applications is difficult. Theory has almost disappeared. Almost everyone is working on applications. The issue is whether this has gone a bit too far. Everyone has to find their best way of working in principle but if you are a theoretician, then to have really serious contact with applications is for most people, extremely fruitful and indeed almost essential. Some individuals will think that is not true and that it may be better that they sit at their desk and think great thoughts, so to speak! That is another way of working but the danger then is that the great thoughts may have no bearing on the real world. But for most people, it is the interplay which is crucial. Maybe I am not imaginative enough to just sit there and think to myself of abstract problems which are really important enough to spend time on! Others are undoubtedly much better at that which may be their better way of thinking.
“My intrinsic feeling is that more fundamental progress is more likely to be made by very focused, relatively small scale, intensive investigations than collecting millions of bits of information on millions of people, for example. It’s possible to collect such large data now, but it depends on the quality, which may be very high or not, and if it is not, what do you do about it?”
17. What do you see as the greatest challenges facing the profession of statisticians in the coming years?
I know the term ‘the profession of statistics’ is widely used but I am not that keen on it. I would like to think of myself as a scientist, who happens largely to specialise in the use of statistics. That is a question of words to some extent. One answer would be to that the challenge, preferably for an academic statistician, is to be involved in several fields of application in a non-trivial sense and combine the stimulus and the contribution you can make that way with theoretical contributions that those contacts will suggest. As I said before, I don’t think you can lay down a rule as to how what is most productive for everyone.
18. Are there people or events that have been influential in your career? Also, given that you are one of the most well respected statisticians of your generation and many statisticians look up to you, whose work do you admire (it can be someone working now, or someone whose work you admired greatly earlier on in your career?).
The person who influenced me by far the greatest was Henry Daniels. I went to work with him at the Wool Research Association and then he went to Cambridge and from there, Birmingham. He was both a very clever mathematician and a very good statistician. He was also actually a very skilful experimental physicist, which is interesting. At the Wool Research Association, he was a statistician but he also ran a measurement lab (what they called a fibre-measurement lab where he developed all sorts of clever measurement techniques).
Maurice Bartlett, who was at Manchester, UCL and then here in Oxford was another major influence and then in the background were people like R.A. Fisher and Jeffreys. I met Jeffreys a few times and went to his lectures – although he wrote beautifully, his lectures were really rather impossible, which was sad.
Otherwise, I have learnt from almost everybody that I’ve had contact with and I certainly include students amongst them.
19. If you had not got involved in the field of statistics, what do you think you would have done? (Is there another field that you could have seen yourself making an impact on?)
I thought I would go into either theoretical physics or pure mathematics but I’m very glad I didn’t. I’m not clever enough for either of those fields. They are both fascinating subjects but statistics is a much more easily satisfying life because there are so many different directions in which to go. Whereas in pure mathematics, you are possibly doing things that only two other people in the world may understand and that requires a certain austerity of spirit in order to do that, which I do not possess! I also find quantum mechanics absolutely fascinating but I am not original enough to do striking things in that field.
*He turned 90 in July 2014.
Please share your comments.
Sir David Cox speaking at the RSS meeting in a session: “Significance Tests: Rethinking the Controversy” on 5 September 2018.
We were part of a session:
Keynote 4 – Significance Tests: Rethinking the Controversy Assembly Room
Speakers:
Sir David Cox, Nuffield College, Oxford
Deborah Mayo, Virginia Tech
Richard Morey, Cardiff University
Aris Spanos, Virginia Tech
All 4 talks are on this post:
RSS 2018 – Significance Tests: Rethinking the Controversy
It was the same day and conference that my book, Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (2018, CUP) first made its physical appearance:
Blurb for session:
Intermingled in today’s statistical controversies are some long-standing, but unresolved, disagreements on the nature and principles of statistical methods and the roles for probability in statistical inference and modelling. In reaction to the so-called “replication crisis” in the sciences, some reformers suggest significance tests as a major culprit. To understand the ramifications of the proposed reforms, there is a pressing need for a deeper understanding of the source of the problems in the sciences and a balanced critique of the alternative methods being proposed to supplant significance tests. In this session speakers offer perspectives on significance tests from statistical science, econometrics, experimental psychology and philosophy of science. There will be also be panel discussion.
.
Nathan Schachtman, Esq., J.D.
Legal Counsel for Scientific Challenges
Of Significance, Error, Confidence, and Confusion – In the Law and In Statistical Practice
The metaphor of law as an “empty vessel” is frequently invoked to describe the law generally, as well as pejoratively to describe lawyers. The metaphor rings true at least in describing how the factual content of legal judgments comes from outside the law. In many varieties of litigation, not only the facts and data, but the scientific and statistical inferences must be added to the “empty vessel” to obtain a correct and meaningful outcome.
Once upon a time, the expertise component of legal judgments came from so-called expert witnesses, who were free to opine about the claims of causality solely by showing that they had more expertise than the lay jurors. In Pennsylvania, for instance, the standard to qualify witnesses to give “expert opinions” was to show that they had “a reasonable pretense to expertise on the subject.”
In the 19th and the first half of the 20th century, causal claims, whether of personal injuries, discrimination, or whatever, virtually always turned on a conception of causation as necessary and sufficient to bring about the alleged harm. In discrimination claims, plaintiffs pointed to the “inexorable zero,” in cases in which no Black citizen was ever seated on a grand jury, in a particular county, since the demise of Reconstruction. In health claims, the mode of reasoning usually followed something like Koch’s postulates.
The second half of the 20th century was marked by the rise of stochastic models in our understanding of the world. The consequence is that statistical inference made its way into the empty vessel. The rapid introduction of statistical thinking into the law did not always go well. In a seminal discrimination case, Casteneda v. Partida, 430 U.S. 432 (1977), in an opinion by Associate Justice Blackmun, the court calculated a binomial probability for observing the sample result (rather than a result at least as extreme as such a result), and mislabeled the measurement “standard deviations” rather than standard errors:
“As a general rule for such large samples, if the difference between the expected value and the observed number is greater than two or three standard deviations, then the hypothesis that the jury drawing was random would be suspect to a social scientist. The 11-year data here reflect a difference between the expected and observed number of Mexican-Americans of approximately 29 standard deviations. A detailed calculation reveals that the likelihood that such a substantial departure from the expected value would occur by chance is less than I in 10140.” Id. at 430 U.S. 482, 496 n.17 (1977). Justice Blackmun was graduated from Harvard College, summa cum laude, with a major in mathematics.
Despite the extreme statistical disparity in the 11-year run of grand juries, Justice Blackmun’s opinion provoked a robust rejoinder, not only on the statistical analysis, but on the Court’s failure to account for obvious omitted confounding variables in its simplistic analysis. And then there were the inconvenient facts that Mr. Partida was a rapist, indicted by a grand jury (50% with “Hispanic” names), which was appointed by jury commissioners (3/5 Hispanic). Partida was convicted by a petit jury (7/12 Hispanic), in front a trial judge who was Hispanic, and he was denied a writ of habeas court by Judge Garza, who went on to be a member of the Court of Appeals. In any event, Justice Blackmun’s dictum about “two or three” standard deviations soon shaped the outcome of many thousands of discrimination cases, and was translated into a necessary p-value of 5%.
Beginning in the early 1960s, statistical inference became an important feature of tort cases that involved claims based upon epidemiologic evidence. In such health-effects litigation, the judicial handling of concepts such as p-values and confidence intervals often went off the rails. In 1989, the United States Court of Appeals for the Fifth Circuit resolved an appeal involving expert witnesses who relied upon epidemiologic studies by concluding that it did not have to resolve questions of bias and confounding because the studies relied upon had presented their results with confidence intervals.[1] Judges and expert witnesses persistently interpreted single confidence intervals from one study as having a 95 percent probability of containing the actual parameter.[2] Similarly, many courts and counsel committed the transposition fallacy in interpreting p-values as posterior probabilities for the null hypothesis.[3]
Against this backdrop of mistaken and misrepresented interpretation of p-values, the American Statistical Association’s p-value statement was a helpful and understandable restatement of basic principles.[4] Within a few weeks, however, citations to the p-value Statement started to show up in the briefs and examinations of expert witnesses, to support contentions that p-values (or any procedure to evaluate random error) were unimportant, and should be disregarded.[5]
In 2019, Ronald Wasserstein, the ASA executive director, along with two other authors wrote an editorial, which explicitly called for the abandonment of using “statistical significance.”[6] Although the piece was labeled “editorial,” the journal provided no disclaimer that Wasserstein was not speaking ex cathedra.
The absence of a disclaimer provoked a great deal of confusion. Indeed, Brian Turran, the editor of Significance, published jointly by the ASA and the Royal Statistical Society, wrote an editorial interpreting the Wasserstein editorial as an official ASA “recommendation.” Turran ultimately retracted his interpretation, but only in response to a pointed letter to the editor.[7] Turran adverted to a misleading press release from the ASA as the source of his confusion. Inquiring minds might wonder why the ASA allowed such a press release to go out.
In addition to press releases, some people in the ASA started to send emails to journal editors, to nudge them to abandon statistical significance testing on the basis of what seemed like an ASA recommendation. For the most part, this campaign was unsuccessful in the major biomedical journals.[8]
While this controversy was unfolding, then President Karen Kafadar of the ASA stepped into the breach to state definitively that the Executive Director was not speaking for the ASA.[9] In November 2019, the ASA board of directors approved a motion to create a “Task Force on Statistical Significance and Replicability.”[8] Its charge was “to develop thoughtful principles and practices that the ASA can endorse and share with scientists and journal editors. The task force will be appointed by the ASA President with advice and participation from the ASA Board.”
Professor Mayo’s editorial has done the world of statistics, as well as the legal world of judges, lawyers, and legal scholars, a service in calling attention to the peculiar intellectual conflicts of interest that played a role in the editorial excesses of some of the ASA’s leadership. From a lawyer’s perspective, it is clear that courts have been misled, and distracted by, some of the ASA officials who seem to have worked to undermine a consensus position paper on p-values.[10]
Curiously, the task force’s report did not find a home in any of the ASA’s several scholarly publications. Instead “The ASA President’s Task Force Statement on Statistical Significance and Replicability”[11] appeared in the The Annals of Applied Statistics, where it is accompanied by an editorial by ASA former President Karen Kafadar.[12] In November 2021, the ASA’s official “magazine,” Chance, also published the Task Force’s Statement.[13]
Judges and litigants who must navigate claims of statistical inference need guidance on the standard of care scientists and statisticians should use in evaluating such claims. Although the Taskforce did not elaborate, it advanced five basic propositions, which had been obscured by many of the recent glosses on the ASA 2016 p-value statement, and the 2019 editorial discussed above:
Although the Task Force’s Statement will not end the debate or the “wars,” it will go a long way to correct the contentions made in court about the insignificance of significance testing, while giving courts a truer sense of the professional standard of care with respect to statistical inference in evaluating claims of health effects.
REFERENCES
[1] Brock v. Merrill Dow Pharmaceuticals, Inc., 874 F.2d 307, 311-12 (5th Cir. 1989).
[2] Richard W. Clapp & David Ozonoff, “Environment and Health: Vital Intersection or Contested Territory?” 30 Am. J. L. & Med. 189, 210 (2004) (“Thus, a RR [relative risk] of 1.8 with a confidence interval of 1.3 to 2.9 could very likely represent a true RR of greater than 2.0, and as high as 2.9 in 95 out of 100 repeated trials.”) (Both authors testify for claimants cases involving alleged environmental and occupational harms.); Schachtman, “Confidence in Intervals and Diffidence in the Courts” (Mar. 4, 2012) (collecting numerous examples of judicial offenders).
[3] See, e.g., In re Ephedra Prods. Liab. Litig., 393 F.Supp. 2d 181, 191, 193 (S.D.N.Y. 2005) (Rakoff, J.) (credulously accepting counsel’s argument that the use of a critical value of less than 5% of significance probability increased the “more likely than not” burden of proof upon a civil litigant). The decision has been criticized in the scholarly literature, but it is still widely cited without acknowledging its error. See Michael O. Finkelstein, Basic Concepts of Probability and Statistics in the Law 65 (2009).
[4] Ronald L. Wasserstein & Nicole A. Lazar, “The ASA’s Statement on p-Values: Context, Process, and Purpose,” 70 The Am. Statistician 129 (2016); see “The American Statistical Association’s Statement on and of Significance” (March 17, 2016). The commentary beyond the “bold faced” principles was at times less helpful in suggesting that there was something inherently inadequate in using p-values. With the benefit of hindsight, this commentary appears to represent editorizing by the authors, and not the sense of the expert committee that agreed to the six principles.
[5] Schachtman, “The American Statistical Association Statement on Significance Testing Goes to Court, Part I” (Nov. 13, 2018), “Part II” (Mar. 7, 2019).
[6] Ronald L. Wasserstein, Allen L. Schirm, and Nicole A. Lazar, “Editorial: Moving to a World Beyond ‘p < 0.05’,” 73 Am. Statistician S1, S2 (2019); see Schachtman,“Has the American Statistical Association Gone Post-Modern?” (Mar. 24, 2019).
[7] Brian Tarran, “THE S WORD … and what to do about it,” Significance (Aug. 2019); Donald Macnaughton, “Who Said What,” Significance 47 (Oct. 2019).
[8] See, e.g., David Harrington, Ralph B. D’Agostino, Sr., Constantine Gatsonis, Joseph W. Hogan, David J. Hunter, Sharon-Lise T. Normand, Jeffrey M. Drazen, and Mary Beth Hamel, “New Guidelines for Statistical Reporting in the Journal,” 381 New Engl. J. Med. 285 (2019); Jonathan A. Cook, Dean A. Fergusson, Ian Ford, Mithat Gonen, Jonathan Kimmelman, Edward L. Korn, and Colin B. Begg, “There is still a place for significance testing in clinical trials,” 16 Clin. Trials 223 (2019).
[9] Karen Kafadar, “The Year in Review … And More to Come,” AmStat News 3 (Dec. 2019); see also Kafadar, “Statistics & Unintended Consequences,” AmStat News 3,4 (June 2019).
[10] Deborah Mayo, “The statistics wars and intellectual conflicts of interest,” 36 Conservation Biology (2022) (in-press, online Dec. 2021).
[11] Yoav Benjamini, Richard D. DeVeaux, Bradly Efron, Scott Evans, Mark Glickman, Barry Braubard, Xuming He, Xiao Li Meng, Nancy Reid, Stephen M. Stigler, Stephen B. Vardeman, Christopher K. Wikle, Tommy Wright, Linda J. Young, and Karen Kafadar, “The ASA President’s Task Force Statement on Statistical Significance and Replicability,” 15 Annals of Applied Statistics (2021) (in press)
[12] Karen Kafadar, “Editorial: Statistical Significance, P-Values, and Replicability,” 15 Annals of Applied Statistics (2021).
[13] Yoav Benjamini, Richard D. De Veaux, Bradley Efron, Scott Evans, Mark Glickman, Barry I. Graubard, Xuming He, Xiao-Li Meng, Nancy M. Reid, Stephen M. Stigler, Stephen B. Vardeman, Christopher K. Wikle, Tommy Wright, Linda J. Young & Karen Kafadar, “ASA President’s Task Force Statement on Statistical Significance and Replicability,” 34 Chance 10 (2021).
Previous commentaries on my editorial (more to come*)
Park
Dennis
StarkStaley
Pawitan
Hennig
Ionides and Ritov
Haig
Lakens
*Let me know if you wish to write one
.
John Park, MDRadiation Oncologist
Kansas City VA Medical Center
Poisoned Priors: Will You Drink from This Well?
As an oncologist, specializing in the field of radiation oncology, “The Statistics Wars and Intellectual Conflicts of Interest”, as Prof. Mayo’s recent editorial is titled, is one of practical importance to me and my patients (Mayo, 2021). Some are flirting with Bayesian statistics to move on from statistical significance testing and the use of P-values. In fact, what many consider the world’s preeminent cancer center, MD Anderson, has a strong Bayesian group that completed 2 early phase Bayesian studies in radiation oncology that have been published in the most prestigious cancer journal —The Journal of Clinical Oncology (Liao et al., 2018 and Lin et al, 2020). This brings about the hotly contested issue of subjective priors and much ado has been written about the ability to overcome this problem. Specifically in medicine, one thinks about Spiegelhalter’s classic 1994 paper mentioning reference, clinical, skeptical, or enthusiastic priors who also uses an example from radiation oncology (Spiegelhalter et al., 1994) to make his case. This is nice and all in theory, but what if there is ample evidence that the subject matter experts have major conflicts of interests (COIs) and biases so that their priors cannot be trusted? A debate raging in oncology, is whether non-invasive radiation therapy is as good as invasive surgery for early stage lung cancer patients. This is a not a trivial question as postoperative morbidity from surgery can range from 19-50% and 90-day mortality anywhere from 0–5% (Chang et al., 2021). Radiation therapy is highly attractive as there are numerous reports hinting at equal efficacy with far less morbidity. Unfortunately, 4 major clinical trials were unable to accrue patients for this important question. Why could they not enroll patients you ask? Long story short, if a patient is referred to radiation oncology and treated with radiation, the surgeon loses out on the revenue, and vice versa. Dr. David Jones, a surgeon at Memorial Sloan Kettering, notes there was no “equipoise among enrolling investigators and medical specialties… Although the reasons are multiple… I believe the primary reason is financial” (Jones, 2015). I am not skirting responsibility for my field’s biases. Dr. Hanbo Chen, a radiation oncologist, notes in his meta-analysis of multiple publications looking at surgery vs radiation that overall survival was associated with the specialty of the first author who published the article (Chen et al, 2018). Perhaps the pen is mightier than the scalpel!
Currently, there is one surgery vs radiation trial that is accruing well, the VALOR study, a Veterans Affairs (VA) only trial. Although only 9 VA medical centers were involved in 2020, it had enrolled more participants than all previous major (phase 3) trials combined (Moghanaki and Hagan, 2020). I do not believe it is too bold to say that a major portion of this success is due to the fact there are no financial incentives for the surgeons or radiation therapists at the VA (i.e. VA physicians are salaried and do not receive payment per patient).
Here are some clear examples of what I call “poisoned priors” due to COIs. Whether financial or for prestige (would you want to be known as the inferior treatment modality for one of the most common cancers?), the COIs loom large. Many of the specialists in question are highly biased, with exposed COIs. Are we to trust priors constructed from them? Will the errors really be contained within the posteriors from these biased priors? In order to overcome this, you say that you want to use an uninformative or weakly informative prior as a statistical method to judge incoming data? Then what’s the point of having prior knowledge, in this case the surgeons’ and radiation oncologists’ priors who are the subject matter experts, if you are not willing to use them? Indeed as Prof. Mayo notes “It may be retorted that implausible inferences will indirectly be blocked by appropriate prior degrees of belief (informative priors), but this misses the crucial point. The key function of statistical tests is to constrain the human tendency to selectively favor views they believe” (Mayo, 2021). If this statement holds for appropriate prior degrees of belief, how much more is it relevant when we can show that those involved have inappropriate prior degrees belief?
These types of poisoned priors are ubiquitous in medicine and must be taken into account — we haven’t even dealt with “Big Pharma” (and don’t get me started)! We must not give up the apparatus of the phase 3 randomized trial, with its randomization, blinding, multiplicity control, and preregistered statistical thresholds for type I and II error control, which is the best form of severe testing we have for our patients.
References
All commentaries on Mayo (2021) editorial until Jan 31, 2022 (more to come*)
Schachtman
Park
Dennis
StarkStaley
Pawitan
Hennig
Ionides and Ritov
Haig
Lakens
*Let me know if you wish to write one
.
Brian Dennis
Professor Emeritus
Dept Fish and Wildlife Sciences,
Dept Mathematics and Statistical Science
University of Idaho
Journal Editors Be Warned: Statistics Won’t Be Contained
I heartily second Professor Mayo’s call, in a recent issue of Conservation Biology, for science journals to tread lightly on prescribing statistical methods (Mayo 2021). Such prescriptions are not likely to be constructive; the issues involved are too vast.
The science of ecology has long relied on innovative statistical thinking. Fisher himself, inventor of P values and a considerable portion of other statistical methods used by generations of ecologists, helped ecologists quantify patterns of biodiversity (Fisher et al. 1943) and understand how genetics and evolution were connected (Fisher 1930). G. E. Hutchinson, the “founder of modern ecology” (and my professional grandfather), early on helped build the tradition of heavy consumption of mathematics and statistics in ecological research (Slack 2010). Investigators in the early days of the subfield of conservation biology, saw the need for stochastic approaches to modeling rare or colonizing populations and for assessing extinction jeopardy (MacArthur and Wilson 1967, Leigh 1981, Lande and Orzack 1988, Dennis 1989, Dennis et al. 1991). Data arising from modern molecular genetics are now a huge cornerstone of conservation, and analyzing such data well often requires considerable statistical sophistication. Other data in ecology are highly nonstandard and require custom made generalized linear models, generalized additive models, integrated models, state space models, structural equation models, spatial capture-recapture models… an ever-expanding list. Nonstandard data, and the very theories of ecology themselves, require the modern ecologist to master an extensive statistical arsenal (Ellison and Dennis 2010).
Lack of replicability has long been acknowledged in ecology, as ecological systems are severely heterogeneous. Ecologists turned heavily to hierarchical models of various sorts to better capture a fuller picture of the sources of variability in data. The likelihoods involved, for all but the usual normal-based random effects models, are wicked multiple integrals that for many years defied computation. The Bayesian revolution swept through ecology after the discovery in statistics that the posterior distributions for such models could be simulated with MCMC algorithms, bypassing the need to calculate likelihood functions. Most ecologists I talked to had little patience for the philosophical-scientific issues involved in the Bayesian/frequentist choice but rather were enthralled with the quantum leap in complexity and realism in models that could be handled with these Bayesian methods. Frequentist methods were late to the party, but the development of algorithms for likelihood maximization such as data cloning (Lele et al. 2007, Lele et al. 2010) have now given investigators a real choice between frequentist or Bayesian inference for hierarchical models. The philosophical issues can no longer be ignored; the choice between frequentist and Bayesian approaches has consequential differences in the types of conclusions to be drawn from data (Mayo 2018, Lele 2020a,b).
It is no wonder that ecologists have long indulged in substantial introspection and questioning of statistical practice. Single papers, single papers with commentary, forums in journals, whole journal issues, and even whole journals are devoted to expounding on and debating statistical methods in ecology. The “null hypothesis” as an ecological-scientific tool rated an entire issue of The American Naturalist (November 1983).
In a contemporary example, Frontiers in Ecology and Evolution devoted a “research topic” featuring papers on “evidence statistics.” The evidence project seeks to extend Richard Royall’s (1997) ideas about evidence to statistical cases with unknown parameters and misspecified models and to endow the approach with a frequentist error structure useful for pre-data design and post-data inference (Dennis et al. 2019, Taper et al. 2021). The extension is accomplished with “evidence functions” (Lele 2004). The main structural departure from Neyman-Pearson (NP) hypothesis testing or Fisherian significance testing is that the concepts of evidence and frequentist error are separated.
The quality of inferences should increase as the amount of data available. This presents problems for NP hypothesis testing if inferences are bound to error rates, as Type I error rates (alpha) are constant regardless of sample size. On the other hand, with evidence functions, both error rates (probabilities of misleading evidence, analogous to alpha and beta in NP testing) approach zero asymptotically as sample size increases, even when models are misspecified. Results thus far suggest that differences of consistent model selection indexes (such as SIC, a.k.a. BIC) retain properties of evidence functions. AIC differences by contrast have error properties similar to NP hypothesis testing (one of the probabilities of misleading evidence does not go to zero but rather becomes constant, similar to alpha in NP hypothesis testing). Evidence functions are for comparing two models; evidence functions are point estimates of the differences of discrepancies of two models from the true data generating mechanism. Interval estimates for evidence can be produced with valid coverage properties, even when models are misspecified.
An argument against an evidence-error project is the Likelihood Principle (LP), the concept that experiment outcomes giving equal likelihood to a parameter value must be considered equal evidence for that value. The concept requires, for instance, that 7 successes out of 20 Bernoulli trials is the same evidence for a particular value of the success probability regardless of whether the experiment was a binomial experiment (number of trials fixed) or as a negative binomial experiment (trials occur until 7 successes are attained). Mayo (2018) provides an entertaining takedown of the LP on philosophical-scientific grounds. Statistically, the variances of those two success probability estimates would be different between the two experiment designs, and so any assessment of long-run error rates must depend on design as well. Similarly, to consider error rates for evidence functions, the LP must necessarily be left behind.
Journal editors can best help ecology by facilitating, promoting, and encouraging such discourse. Prescribing some fixed statistical approach (as agriculture journals once did for multiple comparisons) in the instructions to authors is likely to be ill-informed and harmful to scientific progress. The statistical landscape is growing and changing rapidly, and how statistical approaches can contribute to a particular science is best left to practitioners to sort out on the journal pages.
References
All commentaries on Mayo (2021) editorial until Jan 31, 2022 (more to come*)
Schachtman
Park
Dennis
StarkStaley
Pawitan
Hennig
Ionides and Ritov
Haig
Lakens
*Let me know if you wish to write one