Preregistration has a socio-epistemological and a logical rationale

.

I will use this banner for posts that seem relevant for our Synthese Topical Collection on Severity and learning from error (CFP here). Many of the issues in today’s meta-methodology interconnect with philosophy of statistics and epistemology, and I am keen to highlight posts that touch on this. Consider preregistration. It’s a welcome consequence of today’s statistical crisis of replication that some social sciences are taking a page from medical trials and calling for preregistration of sampling protocols and full reporting. In 2018, Brian Nosek and others wrote of the “Preregistration Revolution”, as part of open science initiatives. The topic was the focus of a 2024 conference in London, which I was unable to attend, but for which I wrote these two posts here and here.

Predesignation and Philosophy of Science. Daniel Lakens [1] and colleagues (2019, 2024a,b) have long argued, correctly in my judgment, that scientific practices such as preregistration (of hypotheses and planned analysis) should be justified based on principles grounded in a philosophy of science: “Preregistration is a methodological procedure that, given a specific philosophy of science (i.e., Popper’s methodological falsificationism), improves one part of the research process – the evaluation of the test severity of hypothesis tests. Both preregistration [2] and Registered Reports [3] are aligned with methodological falsificationism” (Lakens et al. 2024b, 5). Statistical significance tests, for example, formalize tests that take observed departures from a reference or null hypothesis H0 as (statistically) falsifying H0 (inferring H1) only if the method very probably would have resulted in smaller departures from H0 than observed, were H0 true or adequate. This probability is the severity with which H1 has passed the test.

I love Benjamini’s remark that statistical significance tests are intended “as a first line of defense against being fooled by randomness” (Benjamini 2016,1). If an observed effect is actually spurious (as described in hypothesis H0), we want the probability of inferring it is genuine to be low. That is given by the p-value in a properly designed test. So the severity with which H1 may be inferred is high, 1- p. However, data dredging, multiple testing and cherry picking can result in frequently inferring there is a genuine effect erroneously. The nominal p-value may be low, but the actual error probability associated with such an inference is high. Unsurprisingly, such data dredged effects often disappear in attempted replications with stricter protocols, and the random variation goes a different way. Methods where inferential assessments require knowing the relevant error probabilities of the method producing x may be called ‘error statistical methods’ (for example, significance levels, type 1 and type 2 errors, confidence levels).[4]

However, as occurred to me in 1991, there are many cases where severity is satisfied despite data-dredging, double counting For some examples:  Consider searching a full database for a DNA match with a criminal’s DNA, where we suppose the probabilities of false negatives and false positives are both extremely low. Since the probability is very high of a mismatch with person i, if i were not the criminal (and a nonmatch virtually excludes the person), a match with i warrants inferring that i is the criminal. Nor need the data dredged hypotheses be prespecified to be tested with severity by data—even where those data were “used” to arrive at or select the hypothesis inferred. Another favorite example is from the data analysis of the 1919 eclipse data. The same data were used to arrive at, as well as to test, the source of one set of blurred eclipse data during the tests of the Einstein deflection effect.[5]

“Metascientific research has supported the prediction that due to more transparent reporting, preregistration increases the ability of peers to evaluate the severity of tests compared to non-preregistered studies” (Lakens et al. 2024b, 10).  So the first interesting thesis I extract from their discussion is, roughly, that important reforms (prespecification, registered reports) get their rationale within a philosophy of science (falsification, severe testing) and a philosophy of statistics (error statistical).

From understanding that the rationale of predesignation is severity, we understand when data dredging, double counting, adaptive trials, etc. are kosher: namely, when severity is satisfied.[7]

It’s socio-epistemological and logical. The second thesis is that there is an important “social epistemological” basis for preregistration and registered reports in significance testing. It points to one of the ways that inquiry and auditing can be arranged in practice to satisfy the formal criteria of tests.  A severe testing account, in my view, was never intended to just give a definition. Methodological, social and institutional arrangements are part of making severity operative in practice. I have lately been describing it as a kind of anti-deceit epistemology. Methodological and even institutional arrangements are part of critically auditing claims in practice. Lakens et al. (2024b, 2) put it this way: “the current approach to science is to implement methodological procedures that allow peers to transparently evaluate whether decisions introduce bias. Preregistration is such a methodological procedure and, when implemented well, makes it possible to criticize decisions related to the statistical analyses perceived to reduce the severity of tests”.

But we should be careful with sliding into the view that there is “no logical basis to treat a predesignated and a post-designated  hypothesis that is created before or after looking at data any differently” unless one means only to convey the position of purely formal logics of confirmation. That would be so under a philosophy of statistics where inferential appraisal is insensitive to the error probability of the procedure, such as those that follow the (strong) likelihood principle.  By contrast, the logic of error statistical methods breaks down with biasing selection effects. Their error-statistical guarantees no longer hold.

When the error statistic logic breaks down. Consider, for example, optional stopping. Holders of the likelihood principle aver that “the import of the sequence of n data actually observed will be exactly the same as it would be had you planned to take exactly n observations in the first place” (Edwards et al.1963, 238–39). According to the Bayesian Eric-Jan Wagenmakers, (2007, 785), “if the sampling plan is ignored, the researcher is able to always reject the null hypothesis, even if it is true. This example is sometimes used to argue that any statistical framework should somehow take the sampling plan into account. Some people feel that ‘optional stopping’ amounts to cheating . . ..This feeling is, however, contradicted by a mathematical analysis.” The trouble is that the mathematical analysis presupposes a measure of evidence (the Bayes factor) that is insensitive to error probabilities. While the Bayesian assessment of the evidence remains the same, the probability of erroneously attaining such an assessment grows. “The likelihood principle implies. . . the irrelevance of predesignation, of whether an hypothesis was thought of beforehand or was introduced to explain known effects” (Rosenkrantz 1977,122). Since the import of the evidence is through the likelihood ratio, the evidence for H1 is the same whether H1 is constructed or predesignated. But it is not the same evidence for an error statistician. [6]

Thus, there appears to be a tension between popular calls for preregistration—arguably, one of the most promising ways to boost replication—and accounts where error probabilities do not enter in interpreting results: Bayes Factors, Bayesian posteriors, likelihood ratios. It may be argued, however, that even those who follow the likelihood principle care about the expected error probabilities (or operating characteristics) of a procedure at the planning stages, before the data are collected. As the likelihoodist Richard Royall (2000b, 776) remarks: “The probabilities of weak and of misleading evidence are certainly relevant to the planning of studies. . .. But after the experiment is completed, these probabilities are not relevant for interpreting the results.” It would still, of course, be relevant for them to point out if a given significance test had an invalid p-value. [6] But the implications for the post-data critique of the inference is unclear.

Please share your thoughts and queries in the comments to this post.

NOTES

[1] Lakens is one of the guest editors for the Topical Collection, along with Kent Staley, Wendy Parker and I.

[2] “We define preregistration as a complete description of all information related to planned analyses (including the experimental design, measures, data preprocessing, and when statistical tests will corroborate or falsify predictions) that is demonstrably created without access to the data that will be analysed” (Lakens et al 2024, 2).

[3] A registered report goes further in requiring a detailed study proposal to be reviewed, followed by a commitment to publish the results before data are collected. Lakens et al (2024, 1-2) explain this concept in more detail as well as trace its evolution.

[4] Because error probabilities are based on the sampling distribution of the test statistic, these methods are often called “sampling theory methods”, but “error statistics” seems more apt.

[5] I will be like Lakatos and put interesting material in the footnotes. The recognition around 1990 that data-dredging, double use of data, violations of “novel prediction,” need not violate severity, in contrast to the reigning vs Popper-Lakatos philosophy is what led to my putting forth my severity notion.

I might mention a (real) example that first convinced me that [use-novelty] is not necessary for a good test. Here evidence was used to construct as well as to test a hypothesis of the form

H(x): x dented my 1976 Camaro.

The procedure was to hunt for a car whose tail fin perfectly matched the dent in my Camaro’s fender to construct a hypothesis about the likely make of the car that dented it. … it is practically impossible for the dent to have the features it has unless it was created by a specific type of car tail fin. (Mayo 1996, 276)

For a more recent discussion, please see Excursion 1 tour II in this excerpt of my book Mayo (2018) and my paper Mayo (2025).

[6] It is often argued that the error probability report is only concerned with long-run performance, not the evidence at hand. For the severe tester, what bothers us about pejorative data dredging is not about long-runs–even though they do damage the reliability of performance. It is that a poor job has been done in the case at hand in distinguishing genuine from spurious effects. More generally, as a method’s error probability changes, so does its capability of mitigating and prevent known ways to arrive at systematically misleading results.

REFERENCES

Benjamini, Y. (2016). It’s not the P-values’ Fault: Supplement to Wasserstein and Lazar, “The ASA statement on p-values: Context, process, and purpose”. The American Statistician 70, doi.org/10.1080/00031305.2016.1154108.

Edwards, W., Lindman, H. and Savage, L. (1963). Bayesian statistical inference for psychological research. Psychological Review 70: 193–242.

Lakens, D. (2019). The value of preregistration for Psychological Science: A conceptual Analysis. Japanese Psychological Review 62: 221–30.

Lakens, D. (2024a). When and how to deviate from a preregistration. Collabra Psychology 10(1), Article 117094.

Lakens, D., Mesquida, C., Rasti, S., & Ditroilo, M. (2024b). The benefits of preregistration and registered reports. Evidence-Based Toxicology 2(1).

Mayo, D. G. (1996). Error and the Growth of Experimental Knowledge (EGEK). Chicago: University of Chicago Press.

Mayo, D. G. (2018). Statistical inference as severe testing: How to get beyond the statistics wars, Cambridge: Cambridge University Press, 2018.

Mayo, D. G. (2025). Severe testing: Error statistics versus Bayes factor tests. British Journal for the Philosophy of Science.

Nosek, B., Ebersole, C., DeHaven, A., & Mellor, D. (2018). The preregistration revolution. Proc. Natl. Acad. Sci. 115(11) 2600-2606.

Rosenkrantz, R. (1977). Inference, method, and decision: Towards a Bayesian philosophy of science. D. Reidel.

Royall, R. (2000a,b). On the probability of observing misleading statistical evidence, and Rejoinder. Journal of the American Statistical Association, 95, pp. 760–68, 773–80.

Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p values. Psychonomic Bulletin and Review 14: 779–804.

Categories: predesignation, preregistration, SEV26 | Leave a comment

Post navigation

I welcome constructive comments that are of relevance to the post and the discussion, and discourage detours into irrelevant topics, however interesting, or unconstructive declarations that "you (or they) are just all wrong". If you want to correct or remove a comment, send me an e-mail. If readers have already replied to the comment, you may be asked to replace it to retain comprehension.

Blog at WordPress.com.