Happy belated birthday Sir David Cox

15 July 1924-18 January 2022

Last week, July 15, was Sir David Cox’s birthday. [1]  It was 23 years ago that I first got to know Cox after I (boldly) invited him to be in a session I was organizing on philosophy of statistics for  the Second Erich L. Lehmann Symposium held in May 19–22, 2004; Rice University, Texas. I invited him by email, which seemed too informal back in 2023. To my surprise he said yes. Reasons for my surprise were, for one thing, the conference was in the United States while he was at Oxford. For another, Erich Lehmann had been a prominent student of Jerzy Neyman at Berkeley and had developed statistical significance testing in the Neymanian tradition that Cox wasn’t too fond of. Readers of this blog will recall how Fisher (1955) criticized Neyman for converting “his” significance tests into “acceptance procedures” more suitable for technology than science:

This difference in point of view originated when Neyman, thinking that he was correcting and improving my own early work on tests of significance, as a means to the ‘improvement of natural knowledge’, in fact reinterpreted them in terms of that technological and commercial apparatus which is known as an acceptance procedure.

Another reason for my surprise of course is that philosophy of statistics is rarely at the forefront of statistical conferences, even though Lehmann himself had encouraged me to organize the session. As it turned out, however, Cox was genuinely interested in reconciling the approaches of Fisher and Neyman, and he welcomed foundational reflection that might advance a conception of “frequentist statistics as a theory of inductive inference,” the title of our joint paper published in the Lehmann conference proceedings in 2006. Unless noted, all citations in the following are to Mayo and Cox 2006. Our follow-up collaboration was “Objectivity and Conditionality in Frequentist Inference” (Cox and Mayo, 2010).

In the preface to his 2006 book, Principles of Statistical Inference, Cox discusses the importance of statistical foundations:

Without some systematic structure statistical methods for the analysis of data become a collection of tricks that are hard to assimilate and interrelate to one another.

Calibration 

Our joint 2006 paper began “with the core elements of significance testing in a version very strongly related to but in some respect different from both Fisherian and Neyman-Pearson approaches…” (80). Statistical significance tests, as the statistician Allan Birnbaum aptly put it, are a small part of a rich set of “techniques for systematically appraising and bounding the probabilities (under respective hypotheses) of seriously misleading interpretations of data” (Birnbaum 1970, 1033). These are the method’s error probabilities and they are the basis for the calibration of frequentist or error statistical methods.

The importance of calibrating methods–that is, considering how they would behave in (actual or hypothetical) repeated sampling– is a central theme in Cox’s statistical philosophy. In his view “it seems clear that any proposed method of analysis that in repeated application would mostly give misleading answers is fatally flawed” (Cox 2006, 198). Cox dubbed this the Weak Repeated Sampling Principle. Cox and Hinkley (1974) defined it this way: “[W]e should not follow procedures which for some possible parameter values would give, in hypothetical repetitions, misleading conclusions most of the time” (45–6). Fifty percent gives a very minimal threshold.

Two questions that arise remain open to philosophical controversy:

  • How can the frequentist calibration be used as an evidential or inferential assessment (epistemological use)?
  • How can we ensure: “that the hypothetical long run used in calibration is relevant to the specific data” (Cox 2006, 198)?

The first question leads to philosophical issues for a frequentist because it is generally thought that the best, if not the only, way to use probability for an epistemological assessment is for it to supply measures of degrees of belief, support, or plausibility (absolute or comparative). We may call this probabilism. While Neyman’s behavioristic view emphasized the value of good long-run performance, error probabilities can also serve to assess what can be learned from data by evaluating how well probed specific inferences are. The second question leads to two philosophical conundrums: first how to explain when and why selection effects should alter the inferential assessment, and second, how to consider the relevant sample space without leading to the unique case, which would preclude error probabilities.

Statistical Significance Tests

If 𝐻 is a statistical hypothesis, then usually no outcome strictly contradicts it. Nor would we want to regard data as inconsistent with 𝐻 merely because they are highly improbable under H because “all individual outcomes described in detail may have very small probabilities. Rather, the issue is whether the possibly anomalous outcome represents some systematic and reproducible effect” (80). It will sometimes be claimed that a “no effect” null hypothesis is always false, but this confuses the fact that it is an idealized claim with what it is being used to express, to wit: the effect is of the sort readily produced by chance or  background variability. This is scarcely always false! Here is where statistical significance tests enter.

We have empirical data y viewed as observed values of a random variable Y whose probability distribution, defined by a statistical model, is regarded as an abstract and idealized representation of the underlying data-generating process. Data y are used to learn about the probability distribution of Y, by testing various statistical hypotheses. Neyman and Pearson called the main reference hypothesis the test hypothesis, while Fisher called it the null hypothesis, denoted by H0. A common null hypothesis H0, asserts that an experimental intervention has “no effect” or produces “no difference”.

The immediate objective is to test the conformity of the particular data under analysis with H0 in some respect to be specified. To do this we find a function t = t(y) of the data, to be called the test statistic, such that

  • the larger the value of t the more inconsistent are the data with H0;
  • the corresponding random variable T = t(Y) has a (numerically) known probability distribution when H0 is true. (81)

These two requirements for sensible test statistics are routinely glossed over in popular presentations of tests, yet they are what enable statistical tests to serve the crucial roles of testing. Not just any “statistical summary of the data” can serve this role. The first requirement is essentially that the test statistic actually track the hypothesis H0, generally given in terms of a value of a parameter: The larger the value of t, the more improbable the data under the assumption that H0 adequately captures the relevant feature of the data generation.

The second requirement is what enables computing the p-value corresponding to a value of t: p-value = Pr(T > t; H0) “regarded as a measure of concordance with H0 in the respect tested”  (81). A p-value is the probability that the test would have given rise to a result more incompatible with H0 than y is, were the results due to background or chance variability, as described in H0. It is a counterfactual claim. Small p-values (e.g., .05, .01, .005) indicate inconsistency with H0 in the respect being probed by the test. But we can distinguish two main rationales: inductive behavior and inductive inference.

Inductive Behavior vs. Inductive Inference

The first rationale is good performance: “we may give any particular value 𝑝, say, the following hypothetical interpretation: suppose that we were to treat the data as just decisive evidence against H0. Then in hypothetical repetitions H0 would be rejected in a long-run proportion 𝑝 of the cases in which it is actually true” (81-82).

In this strict behavioristic construal often associated with Neyman–which, incidentally, Egon Pearson (1955) disliked–a rejection corresponds to taking some decision or action. It could be as inferential as declaring evidence of a discrepancy from 𝐻0, or as decision-theoretic as approving a drug. Although Cox often present the low error-rate rationale of tests, he avers that, at least in scientific contexts, this is solely to convey the meaning of terms in a testable or (“operational”) manner. It is not to be applied literally.

[T]here is a distinction between the Neyman–Pearson formulation of testing regarded as clarifying the meaning of statistical significance via hypothetical repetitions and that same theory regarded as in effect an instruction on how to implement the ideas by choosing a suitable α in advance and reaching different decisions accordingly. The interpretation to be attached to accepting or rejecting a hypothesis is strongly context-dependent . . . (Cox 2006, 36)4

In Mayo (2018), I dubbed this Cox’s “meaning vs. application” distinction. A main goal in Mayo & Cox 2006 was to identify the application of error probabilities to arriving at an inferential interpretation of statistical results.

Frequentist Principle of Evidence: FEV

As a starting point, we identified a general principle that we dubbed the Frequentist Principle of Evidence, FEV:

FEV(i): y is … evidence against H0 [or] evidence of discrepancy from H0, if and only if, [were H0 adequate [3] then, with high probability, this would have resulted in a less discordant result than is exemplified by y. (Mayo & Cox 2006, 82) 

The term “discrepancy” here refers to the parametric, not an observed, discordancy.

Statistical significance test reasoning

Is akin to ordinary informal reasoning when we are keen to avoid being “fooled by randomness” (Benjamini 2016). The larger the p-value, the more easily our results can be generated by H0 and thus the less evidence of a genuine discrepancy. “Because there was a high probability (1 − 𝑝) that a less significant result would have occurred were 𝐻0 true, we may justify taking low 𝑝-values, properly computed, as evidence against 𝐻0” (81). The stipulation that p-values be “properly computed” is all important. If, for example, data have been selectively reported to ensure a low nominal p-value, then the reasoning is illicit.

The significance test is a measuring device for accordance with a specified hypothesis calibrated …by its performance in repeated applications, …we employ the performance features to make inferences about aspects of the particular thing that is measured, aspects that the measuring tool is appropriately capable of revealing. (84)

While FEV is set out in relation to a reference hypothesis H0, it rarely suffices to consider only the attained p-value. Instead, results should be interpreted by considering several discrepancies from H0, and using FEV to report how well or poorly tested they are with data y.  Consider the context of what Cox calls embedded hypotheses. Here we have exhaustive parametric hypotheses governed by a parameter θ, such as the mean μ. A typical one-sided test is H0: μ = μ0 vs. H1: μ > μ0. (While Cox preferred writing the test this way, the same test is obtained if it is framed as H0: μ < μ0 vs. H1: μ > μ0 .)[2]  To interpret evidence against a given H0, we apply FEV to several different null hypotheses, each of form H0: μ = μ’ where μ’ = μ0 + δ, δ >0. This allows determining if the data warrant inferring μ > μ’. Doing so gets around a weakness of p-values: failing to inform about the magnitude of discrepancies that are warranted. Our construal automatically blocks erroneously interpreting statistically significant results as indicating magnitudes of departures or discrepancies that are unwarranted. Note too that the FEV assessment accords with the corresponding SEV assessment for μ > μ’.

FEV (ii)

We need another principle in dealing with results that are statistically insignificant, or correspond to what Cox calls a “modest” or “moderate” p-value—namely one that is not small, say greater than .1. They are often imbued with two very different false interpretations: one is that (a) non-significance indicates the truth of the null, the other is that (b) non-significance is entirely uninformative. A nonsignificant result can be used to set an upper bound μ”:  μ  < μ” is warranted if a more significant result would have occurred, were μ as great as μ”. You can read our 2006 paper here.

Happy Belated Birthday David Cox!

[1] Some of Sir David Cox’s honors and awards are: Guy Medal (Silver, 1961) (Gold, 1973); Kettering Prize and Gold Medal for Cancer Research for the development of the Proportional Hazard Regression Model (1990); Knighted by Queen Elizabeth II (1985); George Box Medal (2005); Copley Medal (2010); International Prize in Statistics (2016). See “Remembering Sir David Cox: 1924-2022”, in Significance: Firth, Reid, Mayo, Battey (2022). Portions of these general reflections are from Mayo 2023.

[2] Cox recommended viewing two-sided tests as combining two one-sided tests, doubling the p-value for a selection effect (Cox and Hinkley 1974, 79), at least so long as one is interested in the direction of the effect. This underscores the difference from the familiar Bayesian treatment of point null hypotheses.

[3] The adequacy of H0 means it is adequate as an approximate description of the data generating mechanism, in the manner of interest.

Share your thoughts in the comments to this post.

REFERENCES

Benjamini, Y. (2016). It’s not the P-values’ fault. Comment on Wasserstein and Lazar (2016), The American Statistician, 73(1), supplemental material (online).

Birnbaum, A. (1970). Statistical methods in scientific inference. Nature, 225, 1033.

Cox, D. R. (2006). Principles of Statistical Inference. Cambridge: Cambridge University Press.

Cox, D. R. and Hinkley, D. V. (1974). Theoretical Statistics. Chapman and Hall, London.

Cox, D. R. and Mayo D. G. (2010). Objectivity and Conditionality in Frequentist Inference. In Mayo and Spanos (eds) 2010 (pp. 276-304).

Remembering Sir David Cox, 1924–2022. Significance (2022), 19: 30-37.

Fisher (1955), “Statistical Methods and Scientific Induction”. https://errorstatistics.com/wp-content/uploads/2021/02/fisher_1955-statmethssci-induct.pdf

Mayo, D. G. (2018). Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars. Cambridge: Cambridge University Press.

Mayo, D. G. (2022). A Remembrance of Sir David Cox: ‘”In celebrating Cox’s immense contributions, we should recognise how much there is yet to learn from him”. Significance. (April 2022: 36). [Link to all 4 remembrances.]

Mayo, D.G. (2023). Sir David Cox’s Statistical Philosophy and its Relevance to Today’s Statistical Controversies.  JSM 2023 Proceedings, DOI: https://zenodo.org/records/10028243.

Mayo, D. G. and Cox, D. R. (2006). Frequentist Statistics as a Theory of Inductive Inference. In Rojo, J. (Ed.) Optimality: The Second Erich L. Lehmann Symposium (pp. 77-97). Lecture Notes-Monograph series, IMS, Vol. 49 [Reprinted in Mayo and Spanos 2010.]

Mayo, D.G. and Spanos, A. (eds) (2010). Error and Inference: Recent Exchanges on Experimental Reasoning, Reliability and the Objectivity and Rationality of Science. Cambridge: Cambridge University Press.

Pearson, E. S. (1955). Statistical concepts in their relation to reality. Journal of the Royal Statistical Society B, 17, 204–207.

 

Categories: Sir David Cox | 1 Comment

Announcement: CFP Synthese Topical Collection:  Severity and Learning from Error

.

I hope that many readers of this blog will consider contributing to this!

ANNOUNCEMENT SEV26

Synthese Topical Collection CFP:  Severity and Learning from Error

This Topical Collection examines how inquiry learns from error by focusing on a basic principle of evidence in science, statistics, medicine, law, epistemology, and day-to-day learning: a claim is not well-tested, known or epistemically warranted, if it is based on a method that makes it easy to accept, conclude or infer the claim, even if it is false. Such a claim may accord well with the data, but it has not passed a stringent or severe test. While this overarching intuition is widely shared, the problem of how to understand or satisfy it remains unsolved. C. S. Peirce emphasizes randomization and (what is now called) pre-designation to achieve self-correcting methods. Popper viewed severity in terms of satisfying novel predictive success and surviving stringent attempts at falsification. Deborah Mayo (1996, 2018) combines elements from Popper and Peirce with the use of error probabilities from statistical methods: proposed solutions to problems earn warrant by surviving probes that were capable of showing them wrong or inadequate. This Topical Collection takes “severity” to be a broad meta-level concept according to which a claim – whether a report of a perception, a prediction, a hypothesis, or part of a model – is assessed according to whether, and how readily, its errors and inadequacies would have been found, if present.

Several questions arise: What errors matter for a given aim? What would it take for a method to be capable of detecting them? How in actual practice can inquirers show they have engaged in responsible error probing when there are no formal probability models? Addressing questions like these is of urgent importance today as we face high-powered methods that make it easy to find impressive looking effects that are spurious and non-replicating, or to arrive at well-fitting models that do not predict well, do not replicate, or do not provide substantive scientific understanding. These questions arise in debates about methodological shifts in AI/ML, randomized clinical trials, legal evidence, climate modeling, statistical inference, and error-prone inference in general. We seek to bring these metascience debates into direct contact and to ask what is often left hidden: What errors are now being controlled, and which have quietly dropped out of view? By bringing together philosophers, statisticians, and scientists, we aim to develop a shared set of problems and tools with a forward-looking goal: to shape emerging practices, rather than merely react to them with retrospective commentary.

We welcome submissions on any topic that broadly relates to severity or learning from error. We invite contributions that develop, apply or challenge severity-based reasoning, or that develop alternative approaches, Bayesian, frequentist, machine-learning and other, which engage the same underlying concern: how inquiry learns from error, and how claims earn warrant by surviving probes that were capable of showing them wrong or inadequate. We encourage contributions that explore connections between concepts of severity in different fields. Notably, the concepts of sensitivity and safety in contemporary epistemology can be understood through the lens of severity, and both are redolent of stability in AI. We also welcome discussions of how contemporary manifestations of severity interrelate with the traditional notions of severity from Popper and Peirce, and how concepts of severity may help in tackling fundamental problems of induction, falsification, underdetermination, and realism in philosophy of science.

The collection is partly motivated by the thirtieth anniversary of Deborah Mayo’s (1996, Chicago) Error and the Growth of Experimental Knowledge (Lakatos Prize 1998) and the development of its account of severe testing.

Appropriate Topics for Submission include, among others:

Severity and philosophy of statistics

  • Do recent controversies about the uses of error probabilities in statistics (and metastatistics) present a challenge to severity-based reasoning?
  • Do the new fields of post-selection inference (in AI and other disciplines) allow for error control despite data-driven constructions? Or do they shift attention to different errors?
  • How does severity link to such notions as calibration, security, and stability, and statistical techniques that promote such notions as robustness analyses, and multiverse analyses?

Severity and philosophy of science

  • What does it mean for a method, or for science itself, to be self-correcting or error-correcting? Does it fit best with a pragmatist philosophy?
  • How does severe probing take place in the historical sciences, e.g., climate science, geology? Can claims be well probed without being replicable?
  • Rather than probing for falsity, how can we severely probe if a model is adequate for a purpose or problem of interest?

Severity and contemporary epistemology

  • Can a useful cross-cutting epistemology that links science, statistics, and applied epistemology be built around the concept of severity?
  • Do features of severity (e.g., auditing of assumptions) point to ways to avoid problems of sensitivity and safety in epistemology?
  • Does requiring severity explain why legal epistemology resists mere base-rates and “naked statistics”? Does it solve proof paradoxes in legal epistemology?

Tracking shifts in error control

  • How does AI/ML shift from modeling data-generating mechanisms in statistics to optimizing predictive performance in machine learning.
  • How do changing guidelines for RCTs shift trials from probing biological mechanisms to predicting average treatment effects over a population?
  • What are the social, epistemic, ethical, and political consequences of shifting regimens of error control?

The value of probing error

  • How can adversarial collaborations and stress-testing advance science?
  • How can error repertoires be built and effectively employed to facilitate severity in measurement and experiment?
  • How does learning from error enter outside science (e.g., in art, architecture and life drawing)?

Submissions via: https://www.editorialmanager.com/synt/default.aspx

Under the drop-down menu, select Severity and Learning from Error.

Submitted papers will undergo the usual Synthese review process.

For further information, please contact the guest editors:

mayod@vt.edu, wendyparker@vt.edu, D.Lakens@tue.nl, staleykw@gmail.com.

The deadline for submissions is the 15th of December, 2026 (with possible short extensions). Use the comments or write to me with your ideas and questions with the subject: SEV26. The website announcement is here: https://link.springer.com/collections/ebjdhfadcd

Categories: Error Statistics, SEV 26, severity | Leave a comment

‘Low power’ and an all too standard error (continuation of “don’t turn power on its head”)

.

“In my opinion, a great deal of confusion about statistics can be traced to the fact that the point estimate is seen as being the be all and end all, the expression of uncertainty being forgotten….to provide a point estimate without also providing a standard error is, indeed, an all too standard error.”

Stephen Senn: “Error point: the importance of knowing how much you don’t know”

 

In my previous blogpost, (“How not to turn power on its head”), I argued, in relation to a one-sided test of mean μ (e.g., H0: µ  0 vs H1: µ > 0 with known SE):

If POW(μ′) is high (e.g., over .5), then a just significant result is poor evidence that μ > μ′; while if POW(μ′) is low (e.g., less than .2), it is good evidence that μ > μ′ where μ′ is a value greater than 0 (provided assumptions for these claims hold approximately).

Continue reading

Categories: power, reforming the reformers | 2 Comments

How not to turn power on its head

.

In giving some informal remarks about power at a seminar a couple of weeks ago, I proposed that the tendency to turn the notion of power on its head might be avoided by imagining we need to define a test’s error probabilities in terms of its power alone. We can refer to the power against the null hypothesis, rather than alluding to a type 1 error probability, for example. What do I mean by turning power on its head? I mean, at least here, supposing that a test provides poor evidence of discrepancies that the test has low power to detect.  Continue reading

Categories: power | 3 Comments

Error and the Growth of Experimental Knowledge cover: 30 years ago

30 years ago today, Chicago Press sent me a draft version of this cover for Error and the Growth of Experimental Knowledge for my approval (except the fuchsia and mustard in “ERROR” were switched). At first I thought it was so cartoony that it might be an April 1 joke! I had sent them a picture I drew (now in the preface), but they didn’t think that worked for a cover. They were right. It’s a fabulous cover!

To access EGEK.

Categories: Error and the Growth of Experimental Knowledge | 4 Comments

Comments on “The ASA p-value statement 10 years on” (ii)

.

Given how much I’ve blogged about the 2016 ASA p-value statement, the 2019 Executive Editor’s editorial in The American Statistician (TAS), the 2020 ASA (President’s) Task Force, and the various casualties of the related teeth pulling, I thought I should say something about the recent article by Robert Matthews in Significance (March 2026): “The ASA p-value statement 10 years on: An event of statistical significance?” He begins: “Ten years ago this month, the American Statistical Association (ASA) took the unprecedented step of issuing a statement on one of the most controversial issues in statistics: the use and abuse of p-values.” The Statement is here, 2016 ASA Statement on P-Values and Statistical Significance [1]. The Executive director of the ASA, Ronald Wasserstein, invited me to be a ”philosophical observer” at the meeting which gave rise to the 2016 statement. Although the 2016 ASA statement wasn’t radically controversial, at least as compared to the 2019 Executive Editor’s editorial, which I’ll get to in a minute, it was met with critical reactions on all sides. Stephen Senn provides a figure displaying relationships between reactions. Here’s how Matthews’ article begins: Continue reading

Categories: abandon statistical significance, ASA Task Force on Significance and Replicability, P-values, significance tests, stat wars and their casualties | 26 Comments

Power and Severity with nonsignificant results: more power puzzles? (ii)

The concept of a test’s power, originating in Neyman-Pearson’s early work, by and large, is a pre-data concept for purposes of specifying a test (notably, determining worthwhile sample size), and choosing between tests. In some papers, however, Neyman lists a third goal for power: to interpret test results post data much in the spirit of what is often called “power analysis”. This is to determine the discrepancy from a null hypothesis that may be ruled out, given nonsignificant results. One example is in a paper “The Problem of Inductive Inference” (Neyman 1955)–already a surprising title for behaviorist Neyman. The reason I’m bringing this up is that it has direct bearing on some of today’s most puzzling (and problematic) post-data uses of power. Interestingly, in that 1955 paper, Neyman is talking to none other than the logical positivist philosopher of confirmation, Rudof Carnap:

I am concerned with the term “degree of confirmation” introduced by Carnap.  …We have seen that the application of the locally best one-sided test to the data … failed to reject the hypothesis [that the n observations come from a source in which the null hypothesis is true].  The question is: does this result “confirm” the hypothesis that H0 is true of the particular data set? (Neyman, pp 40-41).

Neyman continues: Continue reading

Categories: Neyman's Nursery, power analysis | Tags: , , , | Leave a comment

Continuing the blizzard of 26 power puzzles

 

.The mayor of NYC offered $30 an hour to help shovel the ~ 30 inches of snow that fell last Sunday and Monday. From what I hear, it was a very effective program. Here’s a little power puzzle to very easily shovel through [1]

Suppose you are reading about a result x  that is just statistically significant at level α (i.e., P-value = α) in a one-sided test T+ of the mean of a Normal distribution with n iid samples, and (for simplicity) known σ:   H0: µ ≤  0 against H1: µ >  0. I have heard some people say:

A. If the test’s power to detect alternative µ’ is very low, then the just statistically significant x is poor evidence of a discrepancy (from the null) corresponding to µ’.  (i.e., there’s poor evidence that  µ > µ’ ). I am keeping symbols as simple as possible. *See point on language in notes.

They will generally also hold that if POW(µ’) is reasonably high (at least .5), then the inference to µ > µ’ is warranted, or at least not problematic.

I have heard other people say:

B. If the test’s power to detect alternative µ’ is very low, then the just statistically significant x is good evidence of a discrepancy (from the null) corresponding to µ’ (i.e., there’s good evidence that  µ > µ’).

They will generally also hold that if POW(µ’) is reasonably high (at least .5), then the inference to µ > µ’ is unwarranted.

Which is correct, from the perspective of the (error statistical) philosophy, within which power and associated tests are defined? Continue reading

Categories: blizzard of 26 power puzzles, power, reforming the reformers | 1 Comment

A Blizzard of Power Puzzles Replicate in Meta-Research

.

I often say that the most misunderstood concept in error statistics is power. One week ago, stuck in the blizzard of 2026 in NYC —exciting, if also a bit unnerving, with airports closed for two and a half days and no certainty of when I might fly out—I began collecting the many power howlers I’ve discussed in the past, because some of them are being replicated in todays meta-research about replication failure! Apparently, mistakes about statistical concepts replicate quite reliably—even when statistically significant effects do not. Others I find in medical reports of clinical trials of treatments I’m trying to evaluate in real life! Here’s one variant: A statistically significant result in a clinical trial with fairly high (e.g.,  .8) power to detect an impressive improvement δ’ is taken as good evidence of its impressive improvement δ’. Often the high power of .8 is even used as a (posterior) probability of the hypothesis of improvement being δ’. [0] If these do not immediately strike you as fallacious, compare:

  • If the house is fully ablaze, then very probably the fire alarm goes off.
  • If the fire alarm goes off, then very probably the house is fully ablaze.

The first bullet is saying the fire alarm has high power to detect the house being fully ablaze. It does not mean the converse in the second bullet. Continue reading

Categories: blizzard of 26, power, SIST, statistical significance tests | Tags: , , | 11 Comments

Leisurely Cruise February 2026: power, shpower, positive predictive value

2025-6 Leisurely Cruise

The following is the February stop of our leisurely cruise (meeting 6 from my 2020 Seminar at the LSE). There was a guest speaker, Professor David Hand. Slides and videos are below. Ship StatInfasSt may head back to port or continue for an additional stop or two, if there is interest. Although I often say on this blog that the classical notion of power, as defined by Neyman and Pearson, is one of the most misunderstood notions in stat foundations. I did not know, in writing SIST, just how ingrained those misconceptions would become. I’ll write more on this in my next post. (The following is from SIST pp. 354-356, the pages are provided below)

Shpower and Retrospective Power Analysis

It’s unusual to hear books condemn an approach in a hush-hush sort of way without explaining what’s so bad about it. This is the case with something called post hoc power analysis, practiced by some who live on the outskirts of Power Peninsula. Psst, don’t go there. We hear “there’s a sinister side to statistical power, … I’m referring to post hoc power” (Cumming 2012, pp. 340-1), also called observed power and retrospective (retro) power. I will be calling it shpower analysis. It distorts the logic of ordinary power analysis (from insignificant results). The “post hoc” part comes in because it’s based on the observed results. The trouble is that ordinary power analysis is also post-data. The criticisms are often wrongly taken to reject both. Continue reading

Categories: 2025-2026 Leisurely Cruise, power | Leave a comment

Severe testing of deep learning models of cognition (ii)

.

From time to time I hear of an application of the severe testing philosophy in intriguing ways in fields I know very little about. An example is a recent article by cognitive psychologist Jeffrey Bowers and colleagues (2023): “On the importance of severely testing deep learning models of cognition” (abstract below). Because deep neural networks (DNNs)–advanced machine learning models–seem to recognize images of objects at a similar or even better rate than humans, many researchers suppose DNNs learn to recognize objects in a way similar to humans. However, Bowers and colleagues argue that, on closer inspection, the evidence is remarkably weak, and “in order to address this problem, we argue that the philosophy of severe testing is needed”.

The problem is this. Deep learning models, after all, consist of millions of (largely uninterpretable) parameters. Without understanding how the black box model moves from inputs to outputs, it’s easy to see why observed correlations can easily occur even where the DNN output is due to a variety of factors other than using a similar mechanism as the human visual system. From the standpoint of severe testing, this is a familiar mistake. For data to provide evidence for a claim, it does not suffice that the claim agrees with data, the method must have been capable of revealing the claim to be false, (just) if it is. Here the type of claim of interest is that a given algorithmic model uses similar features or mechanisms as humans to categorize images.[1] The problem isn’t the engineering one of getting more accurate algorithmic models, the problem is inferring claim C: DNNs mimic human cognition in some sense (they focus on vision), even though C has not been well probed. Continue reading

Categories: severity and deep learning models | 5 Comments

(JAN #2) Leisurely cruise January 2026: Excursion 4 Tour II: 4.4 “Do P-Values Exaggerate the Evidence?”

2026-26 Cruise

Our second stop in 2026 on the leisurely tour of SIST is Excursion 4 Tour II which you can read here. This criticism of statistical significance tests takes a number of forms. Here I consider the best known.  The bottom line is that one should not suppose that quantities measuring different things ought to be equal. At the bottom you will see links to posts discussing this issue, each with a large number of comments. The comments from readers are of interest! We will have a zoom meeting Fri Jan 23 11AM ET on these last two posts.*If you want to join us, contact us.

getting beyond…

Excerpt from Excursion 4 Tour II*

4.4 Do P-Values Exaggerate the Evidence? Continue reading

Categories: 2026 Leisurely Cruise, frequentist/Bayesian, P-values | Leave a comment

(JAN #1) Leisurely Cruise January 2026: Excursion 4 Tour I: The Myth of “The Myth of Objectivity” (Mayo 2018, CUP)

2025-26 Cruise

Our first stop in 2026 on the leisurely tour of SIST is Excursion 4 Tour I which you can read here. I hope that this will give you the chutzpah to push back in 2026, if you hear that objectivity in science is just a myth. This leisurely tour may be a bit more leisurely than I intended, but this is philosophy, so slow blogging is best. (Plus, we’ve had some poor sailing weather). Please use the comments to share thoughts.

.

Tour I The Myth of “The Myth of Objectivity”*

Objectivity in statistics, as in science more generally, is a matter of both aims and methods. Objective science, in our view, aims to find out what is the case as regards aspects of the world [that hold] independently of our beliefs, biases and interests; thus objective methods aim for the critical control of inferences and hypotheses, constraining them by evidence and checks of error. (Cox and Mayo 2010, p. 276) [i]

Continue reading

Categories: 2026 Leisurely Cruise, objectivity, Statistical Inference as Severe Testing | Leave a comment

Midnight With Birnbaum: Happy New Year 2026!

.

Anyone here remember that old Woody Allen movie, “Midnight in Paris,” where the main character (I forget who plays it, I saw it on a plane), a writer finishing a novel, steps into a cab that mysteriously picks him up at midnight and transports him back in time where he gets to run his work by such famous authors as Hemingway and Virginia Wolf?  (It was a new movie when I began the blog in 2011.) He is wowed when his work earns their approval and he comes back each night in the same mysterious cab…Well, ever since I began this blog in 2011, I imagine being picked up in a mysterious taxi at midnight on New Year’s Eve, and lo and behold, find myself in the 1960s New York City, in the company of Allan Birnbaum who is is looking deeply contemplative, perhaps studying his 1962 paper…Birnbaum reveals some new and surprising twists this year! [i] 

(The pic on the left is the only blurry image I have of the club I’m taken to.) It has been a decade since  I published my article in Statistical Science (“On the Birnbaum Argument for the Strong Likelihood Principle”), which includes  commentaries by A. P. David, Michael Evans, Martin and Liu, D. A. S. Fraser, Jan Hannig, and Jan Bjornstad. David Cox, who very sadly did in January 2022, is the one who encouraged me to write and publish it. Not only does the (Strong) Likelihood Principle (LP or SLP) remain at the heart of many of the criticisms of Neyman-Pearson (N-P) statistics and of error statistics in general, but a decade after my 2014 paper, it is more central than ever–even if it is often unrecognized.

OUR EXCHANGE:

ERROR STATISTICIAN: It’s wonderful to meet you Professor Birnbaum; I’ve always been extremely impressed with the important impact your work has had on philosophical foundations of statistics.  I happen to have published on your famous argument about the likelihood principle (LP).  (whispers: I can’t believe this!) Continue reading

Categories: Birnbaum, CHAT GPT, Likelihood Principle, Sir David Cox | Leave a comment

For those who want to binge read the (Strong) Likelihood Principle in 2025

.

David Cox’s famous “weighing machine” example” from my last post is thought to have caused “a subtle earthquake” in foundations of statistics. It’s been 11 years since I published my Statistical Science article on this, Mayo (2014), which includes several commentators, but the issue is still mired in controversy. It’s generally dismissed as an annoying, mind-bending puzzle on which those in statistical foundations tend to hold absurdly strong opinions. Mostly it has been ignored. Yet I sense that 2026 is the year that people will return to it again. It’s at least touched upon in Roderick Little’s new book (pic below). This post gives some background, and collects the essential links that you would need if you want to delve into it. Many readers know that each year I return to the issue on New Year’s Eve…. But that’s tomorrow.

By the way, this is not part of our lesurely tour of SIST. In fact, the argument is not even in SIST, although the SLP (or LP) arises a lot. But if you want to go off the beaten track with me to the SLP conundrum, here’s your opportunity. Continue reading

Categories: 11 years ago, Likelihood Principle | Leave a comment

67 Years of Cox’s (1958) Chestnut: Excerpt from Excursion 3 Tour II

2025-26 Cruise

.

We’re stopping to consider one of the “chestnuts” in the exhibits of “chestnuts and howlers” in Excursion 3 (Tour II) of Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (SIST 2018). It is now 67 years since Cox gave his famous weighing machine example in Sir David Cox (1958)[1]. It will play a vital role in our discussion of the (strong) Likelihood Principle later this week. The excerpt is from SIST (pp. 170-173).

Exhibit (vi): Two Measuring Instruments of Different Precisions. Did you hear about the frequentist who, knowing she used a scale that’s right only half the time, claimed her method of weighing is right 75% of the time? 

She says, “I flipped a coin to decide whether to use a scale that’s right 100% of the time, or one that’s right only half the time, so, overall, I’m right 75% of the time.” (She wants credit because she could have used a better scale, even knowing she used a lousy one.)

Basis for the joke: An N-P test bases error probability on all possible outcomes or measurements that could have occurred in repetitions, but did not. Continue reading

Categories: 2025 leisurely cruise, Birnbaum, Likelihood Principle | Leave a comment

(DEC #2) December Leisurely Tour Meeting 3: SIST Excursion 3 Tour III

2025-26 Cruise

We are now at the second stop on our December leisurely cruise through SIST: Excursion 3 Tour III. I am pasting the slides and video from this session during the LSE Research Seminars in 2020 (from which this cruise derives). (Remember it was early pandemic, and we weren’t so adept with zooming.)  The Higgs discussion clarifies (and defends) a somewhat controversial interpretation of p-values. (If you’re interested in the Higgs discovery, there’s a lot more on this blog you can find with the search. I am not sure if I would include the section on “capability and severity” were I to write a second edition, though I would keep the duality of tests and CIs. My goal was to expose a fallacy that is even more common nowadays, but I would have placed a revised version later in the book. Share your remarks in the comments.

.

 

 

 

 

 

III. Deeper Concepts: Confidence Intervals and Tests: Higgs’ Discovery: Continue reading

Categories: 2025 leisurely cruise, confidence intervals and tests | Leave a comment

December leisurely cruise “It’s the Methods, Stupid!” Excursion 3 Tour II (3.4-3.6)

2025-26 Cruise

Welcome to the December leisurely cruise:
Wherever we are sailing, assume that it’s warm, warm, warm (not like today in NYC). This is an overview of our first set of readings for December from my Statistical Inference as Severe Testing: How to get beyond the statistics wars (CUP 2018): [SIST]–Excursion 3 Tour II. This leisurely cruise, participants know, is intended to take a whole month to cover one week of readings from my 2020 LSE Seminars, except for December and January which double up. 

What do you think of  “3.6 Hocus-Pocus: P-values Are Not Error probabilities, Are Not Even Frequentist”? This section refers to Jim Berger’s famous attempted unification of Jeffreys, Neyman and Fisher in 2003. The unification considers testing 2 simple hypotheses using a random sample from a Normal distribution, computing their two P-values, rejecting whichever gets a smaller P-value, and then computing its posterior probability, assuming each gets a prior of .5. This becomes what he calls the “Bayesian error probability” upon which he defines “the frequentist principle”. On Berger’s reading of an important paper* by Neyman (1977), Neyman criticized p-values for violating the frequentist principle (SIST p. 186). *The paper is “frequentist probability and frequentist statistics”. Remember that links to readings outside SIST are at the Captains biblio on the top left of the blog. Share your thoughts in the comments.

Some snapshots from Excursion 3 tour II.

Continue reading

Categories: 2025 leisurely cruise | Leave a comment

Modest replication probabilities of p-values–desirable, not regrettable: a note from Stephen Senn

.

You will often hear—especially in discussions about the “replication crisis”—that statistical significance tests exaggerate evidence. Significance testing, we hear, inflates effect sizes, inflates power, inflates the probability of a real effect, or inflates the probability of replication, and thereby misleads scientists.

If you look closely, you’ll find the charges are based on concepts and philosophical frameworks foreign to both Fisherian and Neyman–Pearson hypothesis testing. Nearly all have been discussed on this blog or in SIST (Mayo 2018), but new variations have cropped up. The emphasis that some are now placing on how biased selection effects invalidate error probabilities is welcome, but I say that the recommendations for reinterpreting quantities such as p-values and power introduce radical distortions of error statistical inferences. Before diving into the modern incarnations of the charges it’s worth recalling Stephen Senn’s response to Stephen Goodman’s attempt to convert p-values into replication probabilities nearly 20 years ago (“A Comment on Replication, P-values and Evidence,” Statistics in Medicine). I first blogged it in 2012, here. Below I am pasting some excerpts from Senn’s letter (but readers interested in the topic should look at all of it), because Senn’s clarity cuts straight through many of today’s misunderstandings. 

.

Continue reading

Categories: 13 years ago, p-values exaggerate, replication research, S. Senn | Tags: , , , | 8 Comments

First Look at N-P Methods as Severe Tests: Water plant accident [Exhibit (i) from Excursion 3]

November Cruise

The example I use here to illustrate formal severity comes in for criticism  in a paper to which I reply in a 2025 BJPS paper linked to here. Use the comments for queries.

Exhibit (i) N-P Methods as Severe Tests: First Look (Water Plant Accident) 

There’s been an accident at a water plant where our ship is docked, and the cooling system had to be repaired.  It is meant to ensure that the mean temperature of discharged water stays below the temperature that threatens the ecosystem, perhaps not much beyond 150 degrees Fahrenheit. There were 100 water measurements taken at randomly selected times and the sample mean x computed, each with a known standard deviation σ = 10.  When the cooling system is effective, each measurement is like observing X ~ N(150, 102). Because of this variability, we expect different 100-fold water samples to lead to different values of X, but we can deduce its distribution. If each X ~N(μ = 150, 102) then X is also Normal with μ = 150, but the standard deviation of X is only σ/√n = 10/√100 = 1. So X ~ N(μ = 150, 1). Continue reading

Categories: 2025 leisurely cruise, severe tests, severity function, water plant accident | Leave a comment

Blog at WordPress.com.