.
Aris Spanos
Wilson Schmidt Professor of Economics
Department of Economics
Virginia Tech
The following guest post (link to PDF of this post) was written as a comment to Mayo’s recent post: “Abandon Statistical Significance and Bayesian Epistemology: some troubles in philosophy v3“.
On Frequentist Testing: revisiting widely held confusions and misinterpretations
After reading chapter 13.2 of the 2022 book Fundamentals of Bayesian Epistemology 2: Arguments, Challenges, Alternatives, by Michael G. Titelbaum, I decided to write a few comments relating to his discussion in an attempt to delineate certain key concepts in frequentist testing with a view to shed light on several long-standing confusions and misinterpretations of these testing procedures. The key concepts include ‘what is a frequentist test’, ‘what is a test statistic and how it is chosen’, and ‘how the hypotheses of interest are framed’.
The first thing that needs to be brought out immediately is that frequentist testing is framed in terms of applied mathematics and its proper framing is of utmost importance in forefending confusions and misinterpretations. Regrettably, philosophers of science tend to undervalue that formalism as cumbersome and needless, which rings hollow when one compares the statistical framing with that of first order logic and other parts of epistemology!
To understand the frequentist testing one needs to place it in its proper context which is a particular statistical model comprising several probabilistic assumptions whose validity of the particular data is of paramount importance. The simplest statistical model is a simple Bernoulli model denoted by:
Xk ~BerlD(θ, θ(1- θ)), 0 < θ <1, xk=0, 1, k=1, 2, —n, …, (1)
where ‘BerIID’ stands for Bernoulli (Ber), Independent and Identically Distributed (IID). The underlying random variable X takes only two values, X =1, say Head (H) and X =0 Tails (T), with P(X =1)= θ and P(X =0)=1− θ. In the context of the statistical model in (1), the relevant data x0:=(x1, x2, …, xn) are viewed as a single realization of the sample X:=( X1, X2 , …, Xn). Note that a random variable is denoted by a capital letter (Xk) and the corresponding observation by a small letter (xk); this minor notational point will avert numerous confusions in practice!
The first issue that arises in practice is how to frame the hypothesis of interest in frequentist testing. This arises because there are a number of confusions between R.A. Fisher’s (1922) framing of the null hypothesis (H0) and the Neyman-Pearson (N-P) (1933) framing that includes both the null (H0) and the alternative (H1) hypothesis. The truth is that the objective is identical for both framings: learn from data x0:=(x1, x2, …, xn) about the ‘true’ value, say θ of the unknown parameter θ. In light of that, the entire parameter space (0, 1) is relevant for frequentist testing since theoretically θ** can take any one of the values in this interval. Hence, the framing of the hypotheses needs to cover the whole of the parameter space. This was first stated in the Neyman-Pearson (N-P) lemma that provided the cornerstone of frequentist testing and included Fisher’s framing as a special case; see Note 1 below.
In summary, the proper way to specify the hypotheses of interest in frequentist testing is to framed them in terms of the model’s unknown parameter(s) θ, and ensure that they constitute a partition of the parameter space. For the statistical model in (1), partitioning can take various forms, including:
H0: θ ≤ θ0 vs. H1: θ > θ0, where θ0 =.5. (2)
What is a frequentist test? It is not just a statistic, say whose distribution in this case is Binomial (Bin):
derived by assuming the validity of the statistical model assumptions ‘BerIID’. Any attempt to use (3) to define any error probabilities, including the p-value (Titelbaum (2022), p. 465), is improper and will give rise to the wrong inference.
A frequentist test comprises two equally important components. For the hypotheses in (2), the optimal N-P test takes the form:
In frequentist testing, the test statistic is always a distance function d(X) framed in terms of a statistic and a frequentist test also includes a rejection region whose choices are neither arbitrary or whimsical. Indeed, the choice of d(X) and C1(α) is interrelated and needs to satisfy two conditions. The first is that the distribution of d(X) can be evaluated under both H0 and H1, and the second is that it has to give rise to an optimal test in terms of learning from data about θ, and there are good choices for d(X) and C1(α) and bad ones, which are evaluated in terms of their capacity to approximate θ** framed in terms of their pre-data type I and II error probabilities.
In the case of the test in (4), the N-P optimal theory of testing renders it Uniformly Most Powerful (UMP) whose relevant sampling distributions are:
where ‘≈’ indicates an approximation of the scaled Binomial by the Normal distribution; see Spanos (2019).
Example 1. Consider particular data x0 representing 17 Hs out of n=20 flips (Titelbaum, 2022, p. 465):
HHHTHHHHHTHHHHTHHHHH (6)
In light of the fact that this p-value is based on n=20 observations, this indicates a clear departure from θ0 =.5. It is important to emphasize that the p-value in (8) is NOT the probability of the particular configuration in (6) occurring as a realization of the sample!
The question that arises at this stage is ‘what about the validity of the probabilistic assumptions comprising the invoked statistical model in (1) for the particular data x0?*’ If any of these assumptions are invalid for x0 the above inference results are unreliable! The Bernoulli assumption is innocuous since any data with two outcomes can always be framed with such a distribution. The IID assumptions, however, are often invalid with real data. When IID is invalid, the assumed distributions under both H0 and H*1 in (5) are invalid, inducing sizeable discrepancies between the actual error probabilities and the nominal ones derived by assuming the validity of IID! This renders any evaluations based on their tail areas highly misleading; see Spanos (2019), ch. 15. Hence, in practice one needs to test the IID assumptions before using x0 in the context of the statistical model in question to draw inferences about *θ.
Example 2. Let us return to the particular data x0 representing 17 Hs out of n=20 flips, and ask the question: Are the IID assumptions valid for data x0? Although there are many ways to test the IID assumptions (Spanos, 2019, ch.15), a particularly simple misspecification test is the runs test. A ‘run’ is a segment of the sequence of outcomes consisting of adjacent identical elements which are followed and proceeded by a different symbol. For the observed sequence in (6), the sequence of 7 runs is shown below:
The runs test compares the actual number of runs R with the number of expected runs E(R)— assuming that the sample is IID process — to construct the runs test:
Note that the above runs test can be extended to a more sophisticated test which accounts, not only for the number of runs, but also their different lengths, e.g. run 1 has length 3, run 2 has length 1 and run 3 has length 5.
In this case the sample size n =20, is rather small, but we can apply the test anyway for illustration purposes:
where the p-value indicates departures from the IID assumptions at any significance level α ≥ .032. One can dispute this particular threshold α but in light of the small sample size n=20, this threshold ensures that the test has sufficient power to detect departures from IID. On the issue of why the p-value should always be one-sided is because the data render irrelevant one of the two tails post-data (when dR(x0) is revealed). This is the key difference between the p-value and type I and II error probabilities that are pre-data; see Spanos (2019). ch. 13.
It is important to emphasize that the runs test in (10) is probing the validity of the IID assumptions underlying the invoked statistical model in (1), and thus it poses very different questions to the data when compared to the N-P test in (5) which assumes the validity of the assumptions and probes for θ! Indeed, misspecification testing should not be framed in terms of the parameter(s) of the underlying statistical model because it probes outside the boundaries of the given model, as opposed to N-P testing that probes within its boundaries*. In practice, misspecification testing predates N-P testing to secure the validity of the invoked statistical model and thus the reliability of the ensuing inference; see Mayo and Spanos (2004).
More broadly, the accept/reject H0 results and small/large p-values do not provide evidence for or against particular hypotheses since the sample size n in conjunction with the pre-specified α play a crucial role in transforming such results into evidence using their post-data severity evaluation that outputs the warranted discrepancy from the null value; see Mayo and Spanos (2006). For a given α, ignoring the sample size n is likely to give rise to the fallacies of acceptance and rejection. This is due to the inherent trade-off between the type I and II error probabilities, which implies that for a given α the power of the test increases with n. The post-data severity evaluation provides an evidential account of the accept/reject H0 results by taking fully into account the sample size n; see Mayo (2018), Spanos (2023).
Note 1. The widely held impression that Fisher’s significance testing and N-P testing are two very different approaches is just another misconstrual of frequentist testing. They are not so different! Stating just a point null hypothesis H0: θ = θ0 and a threshold α, Fisher brings into play the type I and II error probabilities (and power) indirectly into his significance testing. Don’t take my word for this claim, read Fisher (1935), pp. 21-22 describing how the power of the test increases with the sample size n but calling it the ‘sensitivity’ of a test. How could one explain the acerbic Fisher vs. Neyman-Pearson exchanges? They were talking passed each other since the type I and II error probabilities are pre-data — framing the capacity of the test —, and Fisher’s p-value is post-data evaluation indicating potential departures from H0 in light of the observed test statistic. Fisher can get away without specifying an alternative hypothesis H1 since the sign of the observed test statistic indicates the direction of departure, eliminating one of the two tails; see Spanos (2019), ch. 13. It should be noted that Fisher constructed his numerous test statistics using intuition, but they turned out to define optimal frequentist tests when supplemented by an appropriate rejection region, and Neyman and Pearson (1933) give him credit for that.
Note 2. For further discussions on frequentist testing, its misinterpretations and misuses, including the base-rate fallacy, the Jeffreys-Lindley Paradox, Akaike type selection criteria, etc., see the following papers:
References
[1] Fisher, R.A. (1922) “On the mathematical foundations of theoretical statistics”, Philosophical Transactions of the Royal Society A, 222: 309-368.
[2] Fisher, R.A. (1935) The Design of Experiments, Oliver and Boyd, Edinburgh.
[3] Mayo, Deborah G. (2018) Statistical inference as severe testing: How to get beyond the statistics wars, Cambridge University Press.
[4] Mayo, D.G. and A. Spanos (2004) “Methodology in Practice: Statistical Misspecification Testing”, Philosophy of Science, 71: 1007-1025.
[5] Mayo, D.G. and A. Spanos. (2006) “Severe Testing as a Basic Concept in a Neyman-Pearson Philosophy of Induction”, The British Journal for the Philosophy of Science, 57: 323-357.
[6] Neyman, J. and E.S. Pearson (1933) “On the problem of the most efficient tests of statistical hypotheses”, Philosophical Transactions of the Royal Society, A, 231, 289-337.
[7] Spanos, Aris (2019) Introduction to Probability Theory and Statistical Inference: Empirical Modeling with Observational Data, 2nd edition, Cambridge University Press, Cambridge.
[8] Spanos, Aris (2023) “Revisiting the Large n (Sample Size) Problem: How to Avert Spurious Significance Results.” Stats 6(4): 1323-1338.