.
A topic that came up in some comments recently reflects a recent tendency to divorce statistical inference (bad) from statistical thinking (good), and it deserves the spotlight of a post. I always alert authors of papers that come up on this blog, inviting them to comment, and one from Christopher Tong (reacting to a comment on Ron Kenett) concerns this dichotomy.
Response by Christopher Tong to D. Mayo’s July 14 comment
TONG: In responding to Prof. Kenett, Prof. Mayo states: “we should reject the supposed dichotomy between ‘statistical method and statistical thinking’ which unfortunately gives rise to such titles as ‘Statistical inference enables bad science, statistical thinking enables good science,’ in the special TAS 2019 issue. This is nonsense.” [Mayo July 14 comment here.]
I am the author of the paper whose title she attacks as “nonsense”. If she had read my paper she would know that, like Kenett, I am advocating placing statistical thinking at the center of statistical teaching and practice. The dichotomy that she thinks is false exists in much of actual teaching and practice, and is one that it seems both Kenett and I are trying to undo. The title of my paper reflects the real (not the ideal) situation, and if that’s “nonsense”, then (and I would agree) much of statistical teaching and practice is nonsense. Finally I note that mine is one of only two papers in the 2019 special issue that even contains the phrase “statistical thinking” in the title. I strongly recommend the other one, which offers a concrete solution to how the “integrating” that Kenett speaks of can be done in statistics education.
The views expressed are my own.
Response by Mayo to Tong:
MAYO: Thank you for your comment. My thinking was that it would be good to alert the authors of the papers Lakens discusses, and I’m glad that you have. I have read your paper, and, as much as your highly provocative title earns rewards in contexts such as the special issue in which it appears, in my thinking, it does an enormous disservice to statistical inference as “enabling bad science”. Your paper itself—which reviews many right-headed contributions—shows that the very insights and tools that your “good statistical thinking” requires are themselves at the foundations of frequentist error statistical methodology and depend upon statistical inference methods, formal and informal. The formal tools were developed as deliberate idealizations by the founders as exemplars—to check and improve our ordinary (pre-statistics) statistical thinking. Grasping the brilliance of how this works demands a clear understanding of the mathematical and conceptual tools.
I wrote a book Statistical Inferisence as Severe Testing: How to Get Beyond the Statistics Wars (CUP, 2018): SIST. You can find all 16 “tours” on this blog (in final draft form) in this post: Blurbs of 16 Tours: Statistical Inference as Severe Testing; How to Get Beyond the Statistics Wars (SIST).
The notion of statistical inference developed there is very different from your sterile depiction. You talk as if practicing scientists in fields that employ statistical method gain their first exposure to statistics at the point that they are doing applied research. This should not be true. High school students, if they are to be critical consumers of the policies and decisions that will affect them in their lives—let alone conduct research—should study statistical method, including experimental design. Nor can statistical researchers without a VERY clear understanding of statistical concepts and computations assume they only need to think about the domain field, and assume good science will emerge. They should have a deep grasp of the formal methods that others will use to check their models and results. Statistical significance tests and tests of statistical hypotheses more generally, are intimately connected to experimental design, as Fisher emphasized.
I worry that your paper warns these students off, claiming it will only endanger their ability to do good science. What a relief for the students! This is one of their hardest courses, and now they can point to an important journal that has an article that warns us NOT to study statistical inference.
Statistical significance tests are just one small part of statistical science, but they are piecemeal methods and cannot all be learned at one time. Fisher wrote a book, Statistical Methods and Scientific Inference; the integration of the two was there from the start.
Testing statistical assumptions is a crucial part of error statistical methods. You mention Box, but he is talking about Bayesian vs frequentist methods. Box considered that Bayesian inference gives the formal, deductive part of inference which, in his view, could enter only after the creative, inductive work of arriving at and testing a model, which he claimed requires statistical significance tests:
[S]ome check is needed on [the brain’s] pattern seeking ability, for common experience shows that some pattern or other can be seen in almost any set of data or facts. This is the object of diagnostic checks and tests of fit which, I will argue, require frequentist theory significance tests for their formal justification. (Box 1983, 57)
Yet you say “formal, probability-based statistical inference should play no role in most scientific research, which is inherently exploratory, requiring flexible methods of analysis that inherently risk overfitting”. Box disagrees, saying we need checks on such risks, and statistical significance tests provides that. Eye-balling the data won’t suffice. (I say this after having worked with Aris Spanos, an expert on testing model assumptions.) Whenever we use data to solve statistical problems we are doing statistical inference: this goes beyond the data, and thus it is inductive or ampliative. (A paper I wrote with David Cox in 2006 is called: “Frequentist Statistics as a Theory of Inductive Inference”.)
There is no suggestion whatever that the significance test would typically be the only analysis reported. In fact, a fundamental tenet of the conception of inductive learning most at home with the frequentist philosophy is that inductive inference requires building up incisive arguments and inferences by putting together several different piece-meal results. Although the complexity of the story makes it more difficult to set out neatly, as, for example, if a single algorithm is thought to capture the whole of inductive inference, the payoff is an account that approaches the kind of full-bodied arguments that scientists build up in order to obtain reliable knowledge and understanding of a field. (Mayo & Cox 2006, 82.)
Perhaps some pedagogical treatments of statistical inference methods are overly formal, allowing students to just use computers to get the answer. Maybe that’s what’s behind your saying that there’s a divorce between statistical inference (bad) and statistical thinking (good). I say that computing solutions by hand provides a much deeper understanding of methods, and of where our intuitive thinking about probability and statistical inference is often badly wrong. It seems you’re missing that the key rationale for using deliberately idealized models in statistics is in order to learn from data how they fail and how to improve them. Used correctly, they serve as references for severe testing.
Of course, as you stress, “exploratory” inquiry and model building require a data dependence that would not be kosher in a predesignated “confirmatory inquiry”. But even in exploratory inquiry, we can use data both to build and severely probe such questions as whether a given method or model ought to be modified, whether it will serve to find out what we want to know, despite approximations, etc. Moreover, in exploratory inference, there are still statistical assumptions that ought to, and can, be checked by methods with different assumptions, and triangulating results. Fisher, Neyman and many, many others gave us mathematics to show how various designs (e.g., randomizations), and remodeling of data allow “subtracting out” or compensating misspecifications. Contemporary methods go further, but puzzlingly, you reject all such “technical fixes”.
I’m inclined to think that John Byrd had it right (in his comment on Kenett—who I do not think shares your view of statistical inference as enabling bad science):
So, I say that the reasoning underlying these [data science] approaches was given to us by Fisher, Neyman, Pearson, Deming, Cohen, Cox and others from our past. If “data science” becomes ignorant of what statistics can teach us, they will end up re-inventing these same concepts that guide error control, sampling issues, etc. Then we all get to watch a younger generation think they invented such concepts. (Link to J. Byrd July 14 full comment)
Another response to Kenett that seems right-headed is that of Christian Hennig:
Then, on the other side, statistical methodology is to quite some extent a formalisation of principles of statistical thinking, and if we want to analyse formally the implications of our thinking (and more broadly how to do it best), we generate “statistical methodology” by modelling situations (probability models) and decision making (statistical methods, model-based or not). “Errors” and “error probabilities” are then relevant again in the sense that statistical thinking can be criticised by saying, “if you apply “statistical thinking principle”, i.e., method A in artificial situation B in which we know (as we can “control” the truth when assuming models) that we should arrive at conclusion C, in fact you will quite likely arrive at conclusion D which is opposite to C”, then we have learnt something about how statistical thinking can be led astray.
I don’t really think this kind of reasoning can be easily replaced, and I claim that quality statistical thinking needs to be informed by such knowledge. (Link to C. Hennig’s July 17 comment)
There are several other comments on Kenett’s post both before and after Tong’s that might interest y. I invite your thoughts in the comments.
References
Box, G. (1983). An Apology for Ecumenism in Statistics, in Box, G., Leonard, T., and Wu, D. (eds.), Scientific Inference, Data Analysis, and Robustness, New York: Academic Press, 51–84.
Tong, C. (2019). Statistical Inference Enables Bad Science; Statistical Thinking Enables Good Science. The American Statistician, 73(sup1), 246–261. https://doi.org/10.1080/00031305.2018.1518264