Happy belated birthday Sir David Cox

15 July 1924-18 January 2022

Last week, July 15, was Sir David Cox’s birthday. [1]  It was 23 years ago that I first got to know Cox after I (boldly) invited him to be in a session I was organizing on philosophy of statistics for  the Second Erich L. Lehmann Symposium held in May 19–22, 2004; Rice University, Texas. I invited him by email, which seemed too informal back in 2023. To my surprise he said yes. Reasons for my surprise were, for one thing, the conference was in the United States while he was at Oxford. For another, Erich Lehmann had been a prominent student of Jerzy Neyman at Berkeley and had developed statistical significance testing in the Neymanian tradition that Cox wasn’t too fond of. Readers of this blog will recall how Fisher (1955) criticized Neyman for converting “his” significance tests into “acceptance procedures” more suitable for technology than science:

This difference in point of view originated when Neyman, thinking that he was correcting and improving my own early work on tests of significance, as a means to the ‘improvement of natural knowledge’, in fact reinterpreted them in terms of that technological and commercial apparatus which is known as an acceptance procedure.

Another reason for my surprise of course is that philosophy of statistics is rarely at the forefront of statistical conferences, even though Lehmann himself had encouraged me to organize the session. As it turned out, however, Cox was genuinely interested in reconciling the approaches of Fisher and Neyman, and he welcomed foundational reflection that might advance a conception of “frequentist statistics as a theory of inductive inference,” the title of our joint paper published in the Lehmann conference proceedings in 2006. Unless noted, all citations in the following are to Mayo and Cox 2006. Our follow-up collaboration was “Objectivity and Conditionality in Frequentist Inference” (Cox and Mayo, 2010).

In the preface to his 2006 book, Principles of Statistical Inference, Cox discusses the importance of statistical foundations:

Without some systematic structure statistical methods for the analysis of data become a collection of tricks that are hard to assimilate and interrelate to one another.

Calibration 

Our joint 2006 paper began “with the core elements of significance testing in a version very strongly related to but in some respect different from both Fisherian and Neyman-Pearson approaches…” (80). Statistical significance tests, as the statistician Allan Birnbaum aptly put it, are a small part of a rich set of “techniques for systematically appraising and bounding the probabilities (under respective hypotheses) of seriously misleading interpretations of data” (Birnbaum 1970, 1033). These are the method’s error probabilities and they are the basis for the calibration of frequentist or error statistical methods.

The importance of calibrating methods–that is, considering how they would behave in (actual or hypothetical) repeated sampling– is a central theme in Cox’s statistical philosophy. In his view “it seems clear that any proposed method of analysis that in repeated application would mostly give misleading answers is fatally flawed” (Cox 2006, 198). Cox dubbed this the Weak Repeated Sampling Principle. Cox and Hinkley (1974) defined it this way: “[W]e should not follow procedures which for some possible parameter values would give, in hypothetical repetitions, misleading conclusions most of the time” (45–6). Fifty percent gives a very minimal threshold.

Two questions that arise remain open to philosophical controversy:

  • How can the frequentist calibration be used as an evidential or inferential assessment (epistemological use)?
  • How can we ensure: “that the hypothetical long run used in calibration is relevant to the specific data” (Cox 2006, 198)?

The first question leads to philosophical issues for a frequentist because it is generally thought that the best, if not the only, way to use probability for an epistemological assessment is for it to supply measures of degrees of belief, support, or plausibility (absolute or comparative). We may call this probabilism. While Neyman’s behavioristic view emphasized the value of good long-run performance, error probabilities can also serve to assess what can be learned from data by evaluating how well probed specific inferences are. The second question leads to two philosophical conundrums: first how to explain when and why selection effects should alter the inferential assessment, and second, how to consider the relevant sample space without leading to the unique case, which would preclude error probabilities.

Statistical Significance Tests

If 𝐻 is a statistical hypothesis, then usually no outcome strictly contradicts it. Nor would we want to regard data as inconsistent with 𝐻 merely because they are highly improbable under H because “all individual outcomes described in detail may have very small probabilities. Rather, the issue is whether the possibly anomalous outcome represents some systematic and reproducible effect” (80). It will sometimes be claimed that a “no effect” null hypothesis is always false, but this confuses the fact that it is an idealized claim with what it is being used to express, to wit: the effect is of the sort readily produced by chance or  background variability. This is scarcely always false! Here is where statistical significance tests enter.

We have empirical data y viewed as observed values of a random variable Y whose probability distribution, defined by a statistical model, is regarded as an abstract and idealized representation of the underlying data-generating process. Data y are used to learn about the probability distribution of Y, by testing various statistical hypotheses. Neyman and Pearson called the main reference hypothesis the test hypothesis, while Fisher called it the null hypothesis, denoted by H0. A common null hypothesis H0, asserts that an experimental intervention has “no effect” or produces “no difference”.

The immediate objective is to test the conformity of the particular data under analysis with H0 in some respect to be specified. To do this we find a function t = t(y) of the data, to be called the test statistic, such that

  • the larger the value of t the more inconsistent are the data with H0;
  • the corresponding random variable T = t(Y) has a (numerically) known probability distribution when H0 is true. (81)

These two requirements for sensible test statistics are routinely glossed over in popular presentations of tests, yet they are what enable statistical tests to serve the crucial roles of testing. Not just any “statistical summary of the data” can serve this role. The first requirement is essentially that the test statistic actually track the hypothesis H0, generally given in terms of a value of a parameter: The larger the value of t, the more improbable the data under the assumption that H0 adequately captures the relevant feature of the data generation.

The second requirement is what enables computing the p-value corresponding to a value of t: p-value = Pr(T > t; H0) “regarded as a measure of concordance with H0 in the respect tested”  (81). A p-value is the probability that the test would have given rise to a result more incompatible with H0 than y is, were the results due to background or chance variability, as described in H0. It is a counterfactual claim. Small p-values (e.g., .05, .01, .005) indicate inconsistency with H0 in the respect being probed by the test. But we can distinguish two main rationales: inductive behavior and inductive inference.

Inductive Behavior vs. Inductive Inference

The first rationale is good performance: “we may give any particular value 𝑝, say, the following hypothetical interpretation: suppose that we were to treat the data as just decisive evidence against H0. Then in hypothetical repetitions H0 would be rejected in a long-run proportion 𝑝 of the cases in which it is actually true” (81-82).

In this strict behavioristic construal often associated with Neyman–which, incidentally, Egon Pearson (1955) disliked–a rejection corresponds to taking some decision or action. It could be as inferential as declaring evidence of a discrepancy from 𝐻0, or as decision-theoretic as approving a drug. Although Cox often present the low error-rate rationale of tests, he avers that, at least in scientific contexts, this is solely to convey the meaning of terms in a testable or (“operational”) manner. It is not to be applied literally.

[T]here is a distinction between the Neyman–Pearson formulation of testing regarded as clarifying the meaning of statistical significance via hypothetical repetitions and that same theory regarded as in effect an instruction on how to implement the ideas by choosing a suitable α in advance and reaching different decisions accordingly. The interpretation to be attached to accepting or rejecting a hypothesis is strongly context-dependent . . . (Cox 2006, 36)4

In Mayo (2018), I dubbed this Cox’s “meaning vs. application” distinction. A main goal in Mayo & Cox 2006 was to identify the application of error probabilities to arriving at an inferential interpretation of statistical results.

Frequentist Principle of Evidence: FEV

As a starting point, we identified a general principle that we dubbed the Frequentist Principle of Evidence, FEV:

FEV(i): y is … evidence against H0 [or] evidence of discrepancy from H0, if and only if, [were H0 adequate [3] then, with high probability, this would have resulted in a less discordant result than is exemplified by y. (Mayo & Cox 2006, 82) 

The term “discrepancy” here refers to the parametric, not an observed, discordancy.

Statistical significance test reasoning

Is akin to ordinary informal reasoning when we are keen to avoid being “fooled by randomness” (Benjamini 2016). The larger the p-value, the more easily our results can be generated by H0 and thus the less evidence of a genuine discrepancy. “Because there was a high probability (1 − 𝑝) that a less significant result would have occurred were 𝐻0 true, we may justify taking low 𝑝-values, properly computed, as evidence against 𝐻0” (81). The stipulation that p-values be “properly computed” is all important. If, for example, data have been selectively reported to ensure a low nominal p-value, then the reasoning is illicit.

The significance test is a measuring device for accordance with a specified hypothesis calibrated …by its performance in repeated applications, …we employ the performance features to make inferences about aspects of the particular thing that is measured, aspects that the measuring tool is appropriately capable of revealing. (84)

While FEV is set out in relation to a reference hypothesis H0, it rarely suffices to consider only the attained p-value. Instead, results should be interpreted by considering several discrepancies from H0, and using FEV to report how well or poorly tested they are with data y.  Consider the context of what Cox calls embedded hypotheses. Here we have exhaustive parametric hypotheses governed by a parameter θ, such as the mean μ. A typical one-sided test is H0: μ = μ0 vs. H1: μ > μ0. (While Cox preferred writing the test this way, the same test is obtained if it is framed as H0: μ < μ0 vs. H1: μ > μ0 .)[2]  To interpret evidence against a given H0, we apply FEV to several different null hypotheses, each of form H0: μ = μ’ where μ’ = μ0 + δ, δ >0. This allows determining if the data warrant inferring μ > μ’. Doing so gets around a weakness of p-values: failing to inform about the magnitude of discrepancies that are warranted. Our construal automatically blocks erroneously interpreting statistically significant results as indicating magnitudes of departures or discrepancies that are unwarranted. Note too that the FEV assessment accords with the corresponding SEV assessment for μ > μ’.

FEV (ii)

We need another principle in dealing with results that are statistically insignificant, or correspond to what Cox calls a “modest” or “moderate” p-value—namely one that is not small, say greater than .1. They are often imbued with two very different false interpretations: one is that (a) non-significance indicates the truth of the null, the other is that (b) non-significance is entirely uninformative. A nonsignificant result can be used to set an upper bound μ”:  μ  < μ” is warranted if a more significant result would have occurred, were μ as great as μ”. You can read our 2006 paper here.

Happy Belated Birthday David Cox!

[1] Some of Sir David Cox’s honors and awards are: Guy Medal (Silver, 1961) (Gold, 1973); Kettering Prize and Gold Medal for Cancer Research for the development of the Proportional Hazard Regression Model (1990); Knighted by Queen Elizabeth II (1985); George Box Medal (2005); Copley Medal (2010); International Prize in Statistics (2016). See “Remembering Sir David Cox: 1924-2022”, in Significance: Firth, Reid, Mayo, Battey (2022). Portions of these general reflections are from Mayo 2023.

[2] Cox recommended viewing two-sided tests as combining two one-sided tests, doubling the p-value for a selection effect (Cox and Hinkley 1974, 79), at least so long as one is interested in the direction of the effect. This underscores the difference from the familiar Bayesian treatment of point null hypotheses.

[3] The adequacy of H0 means it is adequate as an approximate description of the data generating mechanism, in the manner of interest.

Share your thoughts in the comments to this post.

REFERENCES

Benjamini, Y. (2016). It’s not the P-values’ fault. Comment on Wasserstein and Lazar (2016), The American Statistician, 73(1), supplemental material (online).

Birnbaum, A. (1970). Statistical methods in scientific inference. Nature, 225, 1033.

Cox, D. R. (2006). Principles of Statistical Inference. Cambridge: Cambridge University Press.

Cox, D. R. and Hinkley, D. V. (1974). Theoretical Statistics. Chapman and Hall, London.

Cox, D. R. and Mayo D. G. (2010). Objectivity and Conditionality in Frequentist Inference. In Mayo and Spanos (eds) 2010 (pp. 276-304).

Remembering Sir David Cox, 1924–2022. Significance (2022), 19: 30-37.

Fisher (1955), “Statistical Methods and Scientific Induction”. https://errorstatistics.com/wp-content/uploads/2021/02/fisher_1955-statmethssci-induct.pdf

Mayo, D. G. (2018). Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars. Cambridge: Cambridge University Press.

Mayo, D. G. (2022). A Remembrance of Sir David Cox: ‘”In celebrating Cox’s immense contributions, we should recognise how much there is yet to learn from him”. Significance. (April 2022: 36). [Link to all 4 remembrances.]

Mayo, D.G. (2023). Sir David Cox’s Statistical Philosophy and its Relevance to Today’s Statistical Controversies.  JSM 2023 Proceedings, DOI: https://zenodo.org/records/10028243.

Mayo, D. G. and Cox, D. R. (2006). Frequentist Statistics as a Theory of Inductive Inference. In Rojo, J. (Ed.) Optimality: The Second Erich L. Lehmann Symposium (pp. 77-97). Lecture Notes-Monograph series, IMS, Vol. 49 [Reprinted in Mayo and Spanos 2010.]

Mayo, D.G. and Spanos, A. (eds) (2010). Error and Inference: Recent Exchanges on Experimental Reasoning, Reliability and the Objectivity and Rationality of Science. Cambridge: Cambridge University Press.

Pearson, E. S. (1955). Statistical concepts in their relation to reality. Journal of the Royal Statistical Society B, 17, 204–207.

 

Categories: Sir David Cox | Leave a comment

Post navigation

I welcome constructive comments that are of relevance to the post and the discussion, and discourage detours into irrelevant topics, however interesting, or unconstructive declarations that "you (or they) are just all wrong". If you want to correct or remove a comment, send me an e-mail. If readers have already replied to the comment, you may be asked to replace it to retain comprehension.

Blog at WordPress.com.