preregistration

Preregistration has a socio-epistemological and a logical rationale

.

I will use this banner for posts that seem relevant for our Synthese Topical Collection on Severity and learning from error (CFP here). Many of the issues in today’s meta-methodology interconnect with philosophy of statistics and epistemology, and I am keen to highlight posts that touch on this. Consider preregistration. It’s a welcome consequence of today’s statistical crisis of replication that some social sciences are taking a page from medical trials and calling for preregistration of sampling protocols and full reporting. In 2018, Brian Nosek and others wrote of the “Preregistration Revolution”, as part of open science initiatives. The topic was the focus of a 2024 conference in London, which I was unable to attend, but for which I wrote these two posts here and here.

Predesignation and Philosophy of Science. Daniel Lakens [1] and colleagues (2019, 2024a,b) have long argued, correctly in my judgment, that scientific practices such as preregistration (of hypotheses and planned analysis) should be justified based on principles grounded in a philosophy of science: “Preregistration is a methodological procedure that, given a specific philosophy of science (i.e., Popper’s methodological falsificationism), improves one part of the research process – the evaluation of the test severity of hypothesis tests. Both preregistration [2] and Registered Reports [3] are aligned with methodological falsificationism” (Lakens et al. 2024b, 5). Statistical significance tests, for example, formalize tests that take observed departures from a reference or null hypothesis H0 as (statistically) falsifying H0 (inferring H1) only if the method very probably would have resulted in smaller departures from H0 than observed, were H0 true or adequate. This probability is the severity with which H1 has passed the test.

I love Benjamini’s remark that statistical significance tests are intended “as a first line of defense against being fooled by randomness” (Benjamini 2016,1). If an observed effect is actually spurious (as described in hypothesis H0), we want the probability of inferring it is genuine to be low. That is given by the p-value in a properly designed test. So the severity with which H1 may be inferred is high, 1- p. However, data dredging, multiple testing and cherry picking can result in frequently inferring there is a genuine effect erroneously. The nominal p-value may be low, but the actual error probability associated with such an inference is high. Unsurprisingly, such data dredged effects often disappear in attempted replications with stricter protocols, and the random variation goes a different way. Methods where inferential assessments require knowing the relevant error probabilities of the method producing x may be called ‘error statistical methods’ (for example, significance levels, type 1 and type 2 errors, confidence levels).[4]

However, as occurred to me in 1991, there are many cases where severity is satisfied despite data-dredging, double counting For some examples:  Consider searching a full database for a DNA match with a criminal’s DNA, where we suppose the probabilities of false negatives and false positives are both extremely low. Since the probability is very high of a mismatch with person i, if i were not the criminal (and a nonmatch virtually excludes the person), a match with i warrants inferring that i is the criminal. Nor need the data dredged hypotheses be prespecified to be tested with severity by data—even where those data were “used” to arrive at or select the hypothesis inferred. Another favorite example is from the data analysis of the 1919 eclipse data. The same data were used to arrive at, as well as to test, the source of one set of blurred eclipse data during the tests of the Einstein deflection effect.[5]

“Metascientific research has supported the prediction that due to more transparent reporting, preregistration increases the ability of peers to evaluate the severity of tests compared to non-preregistered studies” (Lakens et al. 2024b, 10).  So the first interesting thesis I extract from their discussion is, roughly, that important reforms (prespecification, registered reports) get their rationale within a philosophy of science (falsification, severe testing) and a philosophy of statistics (error statistical).

From understanding that the rationale of predesignation is severity, we understand when data dredging, double counting, adaptive trials, etc. are kosher: namely, when severity is satisfied.[7]

It’s socio-epistemological and logical. The second thesis is that there is an important “social epistemological” basis for preregistration and registered reports in significance testing. It points to one of the ways that inquiry and auditing can be arranged in practice to satisfy the formal criteria of tests.  A severe testing account, in my view, was never intended to just give a definition. Methodological, social and institutional arrangements are part of making severity operative in practice. I have lately been describing it as a kind of anti-deceit epistemology. Methodological and even institutional arrangements are part of critically auditing claims in practice. Lakens et al. (2024b, 2) put it this way: “the current approach to science is to implement methodological procedures that allow peers to transparently evaluate whether decisions introduce bias. Preregistration is such a methodological procedure and, when implemented well, makes it possible to criticize decisions related to the statistical analyses perceived to reduce the severity of tests”.

But we should be careful with sliding into the view that there is “no logical basis to treat a predesignated and a post-designated  hypothesis that is created before or after looking at data any differently” unless one means only to convey the position of purely formal logics of confirmation. That would be so under a philosophy of statistics where inferential appraisal is insensitive to the error probability of the procedure, such as those that follow the (strong) likelihood principle.  By contrast, the logic of error statistical methods breaks down with biasing selection effects. Their error-statistical guarantees no longer hold.

When the error statistic logic breaks down. Consider, for example, optional stopping. Holders of the likelihood principle aver that “the import of the sequence of n data actually observed will be exactly the same as it would be had you planned to take exactly n observations in the first place” (Edwards et al.1963, 238–39). According to the Bayesian Eric-Jan Wagenmakers, (2007, 785), “if the sampling plan is ignored, the researcher is able to always reject the null hypothesis, even if it is true. This example is sometimes used to argue that any statistical framework should somehow take the sampling plan into account. Some people feel that ‘optional stopping’ amounts to cheating . . ..This feeling is, however, contradicted by a mathematical analysis.” The trouble is that the mathematical analysis presupposes a measure of evidence (the Bayes factor) that is insensitive to error probabilities. While the Bayesian assessment of the evidence remains the same, the probability of erroneously attaining such an assessment grows. “The likelihood principle implies. . . the irrelevance of predesignation, of whether an hypothesis was thought of beforehand or was introduced to explain known effects” (Rosenkrantz 1977,122). Since the import of the evidence is through the likelihood ratio, the evidence for H1 is the same whether H1 is constructed or predesignated. But it is not the same evidence for an error statistician. [6]

Thus, there appears to be a tension between popular calls for preregistration—arguably, one of the most promising ways to boost replication—and accounts where error probabilities do not enter in interpreting results: Bayes Factors, Bayesian posteriors, likelihood ratios. It may be argued, however, that even those who follow the likelihood principle care about the expected error probabilities (or operating characteristics) of a procedure at the planning stages, before the data are collected. As the likelihoodist Richard Royall (2000b, 776) remarks: “The probabilities of weak and of misleading evidence are certainly relevant to the planning of studies. . .. But after the experiment is completed, these probabilities are not relevant for interpreting the results.” It would still, of course, be relevant for them to point out if a given significance test had an invalid p-value. [6] But the implications for the post-data critique of the inference is unclear.

Please share your thoughts and queries in the comments to this post.

NOTES

[1] Lakens is one of the guest editors for the Topical Collection, along with Kent Staley, Wendy Parker and I.

[2] “We define preregistration as a complete description of all information related to planned analyses (including the experimental design, measures, data preprocessing, and when statistical tests will corroborate or falsify predictions) that is demonstrably created without access to the data that will be analysed” (Lakens et al 2024, 2).

[3] A registered report goes further in requiring a detailed study proposal to be reviewed, followed by a commitment to publish the results before data are collected. Lakens et al (2024, 1-2) explain this concept in more detail as well as trace its evolution.

[4] Because error probabilities are based on the sampling distribution of the test statistic, these methods are often called “sampling theory methods”, but “error statistics” seems more apt.

[5] I will be like Lakatos and put interesting material in the footnotes. The recognition around 1990 that data-dredging, double use of data, violations of “novel prediction,” need not violate severity, in contrast to the reigning vs Popper-Lakatos philosophy is what led to my putting forth my severity notion.

I might mention a (real) example that first convinced me that [use-novelty] is not necessary for a good test. Here evidence was used to construct as well as to test a hypothesis of the form

H(x): x dented my 1976 Camaro.

The procedure was to hunt for a car whose tail fin perfectly matched the dent in my Camaro’s fender to construct a hypothesis about the likely make of the car that dented it. … it is practically impossible for the dent to have the features it has unless it was created by a specific type of car tail fin. (Mayo 1996, 276)

For a more recent discussion, please see Excursion 1 tour II in this excerpt of my book Mayo (2018) and my paper Mayo (2025).

[6] It is often argued that the error probability report is only concerned with long-run performance, not the evidence at hand. For the severe tester, what bothers us about pejorative data dredging is not about long-runs–even though they do damage the reliability of performance. It is that a poor job has been done in the case at hand in distinguishing genuine from spurious effects. More generally, as a method’s error probability changes, so does its capability of mitigating and prevent known ways to arrive at systematically misleading results.

REFERENCES

Benjamini, Y. (2016). It’s not the P-values’ Fault: Supplement to Wasserstein and Lazar, “The ASA statement on p-values: Context, process, and purpose”. The American Statistician 70, doi.org/10.1080/00031305.2016.1154108.

Edwards, W., Lindman, H. and Savage, L. (1963). Bayesian statistical inference for psychological research. Psychological Review 70: 193–242.

Lakens, D. (2019). The value of preregistration for Psychological Science: A conceptual Analysis. Japanese Psychological Review 62: 221–30.

Lakens, D. (2024a). When and how to deviate from a preregistration. Collabra Psychology 10(1), Article 117094.

Lakens, D., Mesquida, C., Rasti, S., & Ditroilo, M. (2024b). The benefits of preregistration and registered reports. Evidence-Based Toxicology 2(1).

Mayo, D. G. (1996). Error and the Growth of Experimental Knowledge (EGEK). Chicago: University of Chicago Press.

Mayo, D. G. (2018). Statistical inference as severe testing: How to get beyond the statistics wars, Cambridge: Cambridge University Press, 2018.

Mayo, D. G. (2025). Severe testing: Error statistics versus Bayes factor tests. British Journal for the Philosophy of Science.

Nosek, B., Ebersole, C., DeHaven, A., & Mellor, D. (2018). The preregistration revolution. Proc. Natl. Acad. Sci. 115(11) 2600-2606.

Rosenkrantz, R. (1977). Inference, method, and decision: Towards a Bayesian philosophy of science. D. Reidel.

Royall, R. (2000a,b). On the probability of observing misleading statistical evidence, and Rejoinder. Journal of the American Statistical Association, 95, pp. 760–68, 773–80.

Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p values. Psychonomic Bulletin and Review 14: 779–804.

Categories: predesignation, preregistration, SEV26 | 3 Comments

Preregistration, promises and pitfalls, continued v2

..

In my last post, I sketched some first remarks I would have made had I been able to travel to London to fulfill my invitation to speak at a Royal Society conference, March 4 and 5, 2024, on “the promises and pitfalls of preregistration.” This is a continuation. It’s a welcome consequence of today’s statistical crisis of replication that some social sciences are taking a page from medical trials and calling for preregistration of sampling protocols and full reporting. In 2018, Brian Nosek and others wrote of the “Preregistration Revolution”, as part of open science initiatives. Continue reading

Categories: Bayesian/frequentist, Likelihood Principle, preregistration, Severity | 3 Comments

The F.D.A.’s controversial ruling on an Alzheimer’s drug (letter from a reader)(ii)

I was watching Biogen’s stock (BIIB) climb over 100 points yesterday because its Alzheimer’s drug, aducanumab [brand name: Aduhelm], received surprising FDA approval.  I hadn’t been following the drug at all (it’s enough to try and track some Covid treatments/vaccines). I knew only that the FDA panel had unanimously recommended not to approve it last year, and the general sentiment was that it was heading for FDA rejection yesterday. After I received an email from Geoff Stuart[i] asking what I thought, I found out a bit more. He wrote: Continue reading

Categories: PhilStat/Med, preregistration | 10 Comments

On the current state of play in the crisis of replication in psychology: some heresies

.

The replication crisis has created a “cold war between those who built up modern psychology and those” tearing it down with failed replications–or so I read today [i]. As an outsider (to psychology), the severe tester is free to throw some fuel on the fire on both sides. This is a short update on my post “Some ironies in the replication crisis in social psychology” from 2014.

Following the model from clinical trials, an idea gaining steam is to prespecify a “detailed protocol that includes the study rationale, procedure and a detailed analysis plan” (Nosek et.al. 2017). In this new paper, they’re called registered reports (RRs). An excellent start. I say it makes no sense to favor preregistration and deny the relevance to evidence of optional stopping and outcomes other than the one observed. That your appraisal of the evidence is altered when you actually see the history supplied by the RR is equivalent to worrying about biasing selection effects when they’re not written down; your statistical method should pick up on them (as do p-values, confidence levels and many other error probabilities). There’s a tension between the RR requirements and accounts following the Likelihood Principle (no need to name names [ii]). Continue reading

Categories: Error Statistics, preregistration, reforming the reformers, replication research | 9 Comments

For Statistical Transparency: Reveal Multiplicity and/or Just Falsify the Test (Remark on Gelman and Colleagues)

images-31

.

Gelman and Loken (2014) recognize that even without explicit cherry picking there is often enough leeway in the “forking paths” between data and inference so that by artful choices you may be led to one inference, even though it also could have gone another way. In good sciences, measurement procedures should interlink with well-corroborated theories and offer a triangulation of checks– often missing in the types of experiments Gelman and Loken are on about. Stating a hypothesis in advance, far from protecting from the verification biases, can be the engine that enables data to be “constructed”to reach the desired end [1].

[E]ven in settings where a single analysis has been carried out on the given data, the issue of multiple comparisons emerges because different choices about combining variables, inclusion and exclusion of cases…..and many other steps in the analysis could well have occurred with different data (Gelman and Loken 2014, p. 464).

An idea growing out of this recognition is to imagine the results of applying the same statistical procedure, but with different choices at key discretionary junctures–giving rise to a multiverse analysis, rather than a single data set (Steegen, Tuerlinckx, Gelman, and Vanpaemel 2016). One lists the different choices thought to be plausible at each stage of data processing. The multiverse displays “which constellation of choices corresponds to which statistical results” (p. 797). The result of this exercise can, at times, mimic the delineation of possibilities in multiple testing and multiple modeling strategies. Continue reading

Categories: Bayesian/frequentist, Error Statistics, Gelman, P-values, preregistration, reproducibility, Statistics | 9 Comments

Preregistration Challenge: My email exchange

images-2

.

David Mellor, from the Center for Open Science, emailed me asking if I’d announce his Preregistration Challenge on my blog, and I’m glad to do so. You win $1,000 if your properly preregistered paper is published. The recent replication effort in psychology showed, despite the common refrain – “it’s too easy to get low P-values” – that in preregistered replication attempts it’s actually very difficult to get small P-values. (I call this the “paradox of replication”[1].) Here’s our e-mail exchange from this morning:

          Dear Deborah Mayod,

I’m reaching out to individuals who I think may be interested in our recently launched competition, the Preregistration Challenge (https://cos.io/prereg). Based on your blogging, I thought it could be of interest to you and to your readers.

In case you are unfamiliar with it, preregistration specifies in advance the precise study protocols and analytical decisions before data collection, in order to separate the hypothesis-generating exploratory work from the hypothesis testing confirmatory work. 

Though required by law in clinical trials, it is virtually unknown within the basic sciences. We are trying to encourage this new behavior by offering 1,000 researchers $1000 prizes for publishing the results of their preregistered work. 

Please let me know if this is something you would consider blogging about or sharing in other ways. I am happy to discuss further. 

Best,

David
David Mellor, PhD

Project Manager, Preregistration Challenge, Center for Open Science

 

Deborah Mayo To David:                                                                          10:33 AM (1 hour ago)

David: Yes I’m familiar with it, and I hope that it encourages people to avoid data-dependent determinations that bias results. It shows the importance of statistical accounts that can pick up on such biasing selection effects. On the other hand, coupling prereg with some of the flexible inference accounts now in use won’t really help. Moreover, there may, in some fields, be a tendency to research a non-novel, fairly trivial result.

And if they’re going to preregister, why not go blind as well?  Will they?

Best,

Mayo Continue reading

Categories: Announcement, preregistration, Statistical fraudbusting, Statistics | 11 Comments

Blog at WordPress.com.