Data centers: how about an adversarial collaboration?

.

Data Centers: How About an Adversarial Collaboration?

I live in a state booming with data centers.  They’re also a booming political issue, weirdly scrambling some of the familiar political divides such as Steve Bannon and Bernie Sanders on the same side supporting a proposed temporary ban on building new AI data centers. Here in Virginia, a local delegate Josh Cole claims to be “90% sure that he is against data centers,” but feels he’s moving toward a moratorium. In NY, where I spend part of my time, there’s already a moratorium on (hyperscale) data centers. [i]

We’ve been talking about adversarial collaborations as of late, perhaps the controversy can be put to an adversarial collaboration. An adversarial collaboration (AC) allows rivals who hold opposing hypotheses or positions to work together under a shared framework to design tests of competing views. I blogged on the paper, “Teams of Rivals” by Ceci, Clark, Jussim and Williams (2025) here.“ The strong motivation each side’s members will feel to severely test the other side’s predictions should inspire greater confidence in the collaboration’s eventual conclusions” (Ceci et al., 2025), provided various strictures are held. AC’s, they argue, are more effective than mere open science or preregistration. It can happen that both sides of an issue continue to replicate their results!

I don’t think an adversarial collaboration on data centers is far-fetched. Even if the data science explosion is inevitable, as I think it is, understanding the disagreements and arriving at any warranted mitigation might thereby be advanced. An AC on the data center debate would at least move the conversation away from emotional meetings and secret lobbying toward a structured, data-driven negotiation. The huge infrastructure growth has been great for the stock market–at least for now. But sufficient public opposition could lead, and is already leading, to restrictions with serious consequences for a market where we know big tech companies have borrowed enormous sums to finance.

The first step in an adversarial collaboration of this sort–which admittedly is not the type typically envisioned in science–is translating vague public resistance and corporate talking points into testable hypotheses. What exactly are the disagreements? Or at least the testable elements? I can imagine the rival hypotheses could be something like:

  • Public advocate hypothesis (pro-data center moratorium): Electricity demands of data centers are putting serious additional stress on the grid, risking blackouts and raised electricity prices.
  • Tech Industry Hypothesis: Our efficiency and green power investments will stabilize the grid and subsidize clean energy; stopping data center growth would halt progress.

The Adversarial Experiment: Suppose an experiment ran for a year or two in an area with lots of large data centers. Experts might identify periods when electricity demand is especially high—during peak periods and summer heatwaves, for example—and randomly select some of those times for the data centers to completely disconnect from the grid and run on their own batteries for several hours. At other times they would operate normally. The opposing sides would agree in advance on what to measure—grid stress, electricity prices, risks of blackouts, or whatever—and make quantitative predictions to test the effect attributable to the data centers. How much additional stress do they put on the grid, and by how much do they affect prices or other risks during these high-demand periods? The results might also point to whether mitigation is needed and if so what kind. The point is to design a test with a good chance of showing each side flawed, just if it is flawed. Both sides would need to agree on a neutral third party to design the empirical tests. Of course, both sides would bring out much more in their defense, this is just an outsider’s approximation of how a testable element might go.

Members of the two sides won’t shift their general stance, you might say. True, but that wouldn’t preclude positive payoffs such as meaningful safeguards, if public fears are warranted, and a basis for limiting public backlash and restrictions, if they are not. At the very least, people will feel listened to. Similar criticisms about water use resulted in the adoption of water-saving closed-loop cooling at data centers. Ideally both sides in an AC move past all or nothing stances

What do people think?

What about the presumably less testable sources of the data center disagreement? What was the term Paul Slovic used in talking about risk perception? Yes, I remember: dread risk. Risks that are outside one’s control, involuntary exposure, unfamiliar new technology with potentially large consequences. Don’t forget the ugliness. Enormous, mostly windowless buildings covering hundreds or thousands of football fields (with entire campuses) springing up without the general community being involved. Another term used in a podcast appropriately called “Why Everyone Hates AI Data Centers,” is “pain sponge”. Data centers have become a “pain sponge” for lots of issues—electricity prices, water, noise, secrecy, land use, AI, distribution of benefits, distrust of powerful firms and billionaires. Only some of these are testable and fixable. Then there’s the hum–a constant low-frequency sound: a drone, a whistle, an airplane engine, a lawn mower that never stops. Fortunately, I read that the newer centers, using the latest AI chips require new data centers to switch from air to the much quieter direct chip liquid cooling.

I love the New Yorker cartoon above, which zeros in on what at least some of this fabulous computing and storage capacity actually does.[ii]

Use the comments to share your thoughts.

[i] A poll from University of Pennsylvania finds that about 60% oppose the construction of new data centers in their area, up (a statistically significant) 12 percentage points from a survey fielded in February and March. Interestings, the opposition was greatest among adults under 30 (70%) and declined to 57% among those 65 and older.  https://almanac.upenn.edu/articles/opposition-to-local-data-centers-rises-sharply

[ii] It turns out the cartoon is closer to the truth than I thought. The data centers are essentially digital hoarders—they’d rather build a giant new closet than clean out the old one. Apparently it would cost too much to filter junk. Maybe one day they’ll have a neat way to do it. While even aggressively deleting our digital junk is unlikely to stop the data center explosion (which is largely driven by massive computing and processing power rather than just storage space, and the desire to train on junk), it wouldn’t hurt. At least people should be aware, and I doubt most are. I read that something like 80% of the data stored is of this “dark” and useless sort. I welcome knowledgeable inputs on this!

RELATED BLOG POST:

November 1, 2025: Severity and Adversarial Collaborations i

REFERENCE

Ceci, S. J., Clark, C. J., Jussim, L., & Williams, W. M. (2024). Adversarial collaboration: An undervalued approach in behavioral science. American Psychologist. Advance online publication. https://dx.doi.org/10.1037/amp0001391

Synthese Topical Collection on Severity and learning from error (CFP here)

 

Categories: adversarial collaboration, AI data centers | Leave a comment

Preregistration has a socio-epistemological and a logical rationale

.

I will use this banner for posts that seem relevant for our Synthese Topical Collection on Severity and learning from error (CFP here). Many of the issues in today’s meta-methodology interconnect with philosophy of statistics and epistemology, and I am keen to highlight posts that touch on this. Consider preregistration. It’s a welcome consequence of today’s statistical crisis of replication that some social sciences are taking a page from medical trials and calling for preregistration of sampling protocols and full reporting. In 2018, Brian Nosek and others wrote of the “Preregistration Revolution”, as part of open science initiatives. The topic was the focus of a 2024 conference in London, which I was unable to attend, but for which I wrote these two posts here and here. Continue reading

Categories: predesignation, preregistration, SEV26 | 3 Comments

2026 David Cox Foundations of Statistics Award: Peter McCullagh

I am pleased to share that Professor Peter McCullagh has received the 2026 Sir David R. Cox Foundations of Statistics Award, given by the American Statistical Association (ASA).  Below is the announcement from the JUNE 1, 2026 issue of AMSTAT NEWS.

Peter McCullagh to Give David Cox Foundations of Statistics Lecture

.

For foundational contributions to statistical science that have shaped both the theoretical underpinnings and applied practice of the discipline across more than four decades, Peter McCullagh is the third recipient of the David R. Cox Foundations of Statistics Award, presented by the American Statistical Association. McCullagh will receive the award and deliver a lecture titled “What Is a Regression Model?” at the Joint Statistical Meetings in Boston Massachusetts at 10:30 a.m. on August 5. Continue reading

Categories: David R. Cox Foundations of Statistics Award | Leave a comment

Can You Make Me More Capable? Of Art, Astrophysics and AI

.

Can You Make Me More Capable? Of Art, Astrophysics and AI

Since the pandemic, I have returned to an old passion of mine–drawing, especially life drawing. I have always loved it. One year I even won my high school’s art award. Over the years, academic work had crowded it out, except for the occasional conference poster or sketching faculty during meetings. During one of the dark pandemic days I wondered if there were any “drop-in” life drawing classes nearby. It turned out there were two, one in a big old house within walking distance (this was NYC). What a great way to overcome some of the social isolation of those years—even when masked. These classes were pretty full, and post pandemic, the number of offerings has grown by leaps and bounds. This seemed somewhat paradoxical to me. At a time when AI can generate beautiful drawings and paintings in virtually any style within seconds, people were still spending hours struggling to sketch a live model. Why? Clearly because it is great fun, relaxing, and an enjoyable (and, in my case, unusual) social activity; getting better at it expands our ability to impart our own creative perspectives. It’s not so much the drawing we want, but the creative power to create them in unique ways. Continue reading

Categories: life-drawing and astrophysics | 1 Comment

Happy belated birthday Sir David Cox

15 July 1924-18 January 2022

Last week, July 15, was Sir David Cox’s birthday. [1]  It was 23 years ago that I first got to know Cox after I (boldly) invited him to be in a session I was organizing on philosophy of statistics for  the Second Erich L. Lehmann Symposium held in May 19–22, 2004; Rice University, Texas. I invited him by email, which seemed too informal back in 2023. To my surprise he said yes. Reasons for my surprise were, for one thing, the conference was in the United States while he was at Oxford. For another, Erich Lehmann had been a prominent student of Jerzy Neyman at Berkeley and had developed statistical significance testing in the Neymanian tradition that Cox wasn’t too fond of. Readers of this blog will recall how Fisher (1955) criticized Neyman for converting “his” significance tests into “acceptance procedures” more suitable for technology than science: Continue reading

Categories: Sir David Cox | 4 Comments

Announcement: CFP Synthese Topical Collection:  Severity and Learning from Error

.

I hope that many readers of this blog will consider contributing to this!

ANNOUNCEMENT SEV26

 

Synthese Topical Collection CFP:  Severity and Learning from Error

This Topical Collection examines how inquiry learns from error by focusing on a basic principle of evidence in science, statistics, medicine, law, epistemology, and day-to-day learning: a claim is not well-tested, known or epistemically warranted, if it is based on a method that makes it easy to accept, conclude or infer the claim, even if it is false. Such a claim may accord well with the data, but it has not passed a stringent or severe test. While this overarching intuition is widely shared, the problem of how to understand or satisfy it remains unsolved. C. S. Peirce emphasizes randomization and (what is now called) pre-designation to achieve self-correcting methods. Popper viewed severity in terms of satisfying novel predictive success and surviving stringent attempts at falsification. Deborah Mayo (1996, 2018) combines elements from Popper and Peirce with the use of error probabilities from statistical methods: proposed solutions to problems earn warrant by surviving probes that were capable of showing them wrong or inadequate. This Topical Collection takes “severity” to be a broad meta-level concept according to which a claim – whether a report of a perception, a prediction, a hypothesis, or part of a model – is assessed according to whether, and how readily, its errors and inadequacies would have been found, if present. Continue reading

Categories: Error Statistics, SEV 26, severity | Leave a comment

‘Low power’ and an all too standard error (continuation of “don’t turn power on its head”)

.

“In my opinion, a great deal of confusion about statistics can be traced to the fact that the point estimate is seen as being the be all and end all, the expression of uncertainty being forgotten….to provide a point estimate without also providing a standard error is, indeed, an all too standard error.”

Stephen Senn: “Error point: the importance of knowing how much you don’t know”

 

In my previous blogpost, (“How not to turn power on its head”), I argued, in relation to a one-sided test of mean μ (e.g., H0: µ  0 vs H1: µ > 0 with known SE):

If POW(μ′) is high (e.g., over .5), then a just significant result is poor evidence that μ > μ′; while if POW(μ′) is low (e.g., less than .2), it is good evidence that μ > μ′ where μ′ is a value greater than 0 (provided assumptions for these claims hold approximately).

Continue reading

Categories: power, reforming the reformers | 2 Comments

How not to turn power on its head

.

In giving some informal remarks about power at a seminar a couple of weeks ago, I proposed that the tendency to turn the notion of power on its head might be avoided by imagining we need to define a test’s error probabilities in terms of its power alone. We can refer to the power against the null hypothesis, rather than alluding to a type 1 error probability, for example. What do I mean by turning power on its head? I mean, at least here, supposing that a test provides poor evidence of discrepancies that the test has low power to detect.  Continue reading

Categories: power | 3 Comments

Error and the Growth of Experimental Knowledge cover: 30 years ago

30 years ago today, Chicago Press sent me a draft version of this cover for Error and the Growth of Experimental Knowledge for my approval (except the fuchsia and mustard in “ERROR” were switched). At first I thought it was so cartoony that it might be an April 1 joke! I had sent them a picture I drew (now in the preface), but they didn’t think that worked for a cover. They were right. It’s a fabulous cover!

To access EGEK.

Categories: Error and the Growth of Experimental Knowledge | 4 Comments

Comments on “The ASA p-value statement 10 years on” (ii)

.

Given how much I’ve blogged about the 2016 ASA p-value statement, the 2019 Executive Editor’s editorial in The American Statistician (TAS), the 2020 ASA (President’s) Task Force, and the various casualties of the related teeth pulling, I thought I should say something about the recent article by Robert Matthews in Significance (March 2026): “The ASA p-value statement 10 years on: An event of statistical significance?” He begins: “Ten years ago this month, the American Statistical Association (ASA) took the unprecedented step of issuing a statement on one of the most controversial issues in statistics: the use and abuse of p-values.” The Statement is here, 2016 ASA Statement on P-Values and Statistical Significance [1]. The Executive director of the ASA, Ronald Wasserstein, invited me to be a ”philosophical observer” at the meeting which gave rise to the 2016 statement. Although the 2016 ASA statement wasn’t radically controversial, at least as compared to the 2019 Executive Editor’s editorial, which I’ll get to in a minute, it was met with critical reactions on all sides. Stephen Senn provides a figure displaying relationships between reactions. Here’s how Matthews’ article begins: Continue reading

Categories: abandon statistical significance, ASA Task Force on Significance and Replicability, P-values, significance tests, stat wars and their casualties | 26 Comments

Power and Severity with nonsignificant results: more power puzzles? (ii)

The concept of a test’s power, originating in Neyman-Pearson’s early work, by and large, is a pre-data concept for purposes of specifying a test (notably, determining worthwhile sample size), and choosing between tests. In some papers, however, Neyman lists a third goal for power: to interpret test results post data much in the spirit of what is often called “power analysis”. This is to determine the discrepancy from a null hypothesis that may be ruled out, given nonsignificant results. One example is in a paper “The Problem of Inductive Inference” (Neyman 1955)–already a surprising title for behaviorist Neyman. The reason I’m bringing this up is that it has direct bearing on some of today’s most puzzling (and problematic) post-data uses of power. Interestingly, in that 1955 paper, Neyman is talking to none other than the logical positivist philosopher of confirmation, Rudof Carnap:

I am concerned with the term “degree of confirmation” introduced by Carnap.  …We have seen that the application of the locally best one-sided test to the data … failed to reject the hypothesis [that the n observations come from a source in which the null hypothesis is true].  The question is: does this result “confirm” the hypothesis that H0 is true of the particular data set? (Neyman, pp 40-41).

Neyman continues: Continue reading

Categories: Neyman's Nursery, power analysis | Tags: , , , | Leave a comment

Continuing the blizzard of 26 power puzzles

 

.The mayor of NYC offered $30 an hour to help shovel the ~ 30 inches of snow that fell last Sunday and Monday. From what I hear, it was a very effective program. Here’s a little power puzzle to very easily shovel through [1]

Suppose you are reading about a result x  that is just statistically significant at level α (i.e., P-value = α) in a one-sided test T+ of the mean of a Normal distribution with n iid samples, and (for simplicity) known σ:   H0: µ ≤  0 against H1: µ >  0. I have heard some people say:

A. If the test’s power to detect alternative µ’ is very low, then the just statistically significant x is poor evidence of a discrepancy (from the null) corresponding to µ’.  (i.e., there’s poor evidence that  µ > µ’ ). I am keeping symbols as simple as possible. *See point on language in notes.

They will generally also hold that if POW(µ’) is reasonably high (at least .5), then the inference to µ > µ’ is warranted, or at least not problematic.

I have heard other people say:

B. If the test’s power to detect alternative µ’ is very low, then the just statistically significant x is good evidence of a discrepancy (from the null) corresponding to µ’ (i.e., there’s good evidence that  µ > µ’).

They will generally also hold that if POW(µ’) is reasonably high (at least .5), then the inference to µ > µ’ is unwarranted.

Which is correct, from the perspective of the (error statistical) philosophy, within which power and associated tests are defined? Continue reading

Categories: blizzard of 26 power puzzles, power, reforming the reformers | 1 Comment

A Blizzard of Power Puzzles Replicate in Meta-Research

.

I often say that the most misunderstood concept in error statistics is power. One week ago, stuck in the blizzard of 2026 in NYC —exciting, if also a bit unnerving, with airports closed for two and a half days and no certainty of when I might fly out—I began collecting the many power howlers I’ve discussed in the past, because some of them are being replicated in todays meta-research about replication failure! Apparently, mistakes about statistical concepts replicate quite reliably—even when statistically significant effects do not. Others I find in medical reports of clinical trials of treatments I’m trying to evaluate in real life! Here’s one variant: A statistically significant result in a clinical trial with fairly high (e.g.,  .8) power to detect an impressive improvement δ’ is taken as good evidence of its impressive improvement δ’. Often the high power of .8 is even used as a (posterior) probability of the hypothesis of improvement being δ’. [0] If these do not immediately strike you as fallacious, compare:

  • If the house is fully ablaze, then very probably the fire alarm goes off.
  • If the fire alarm goes off, then very probably the house is fully ablaze.

The first bullet is saying the fire alarm has high power to detect the house being fully ablaze. It does not mean the converse in the second bullet. Continue reading

Categories: blizzard of 26, power, SIST, statistical significance tests | Tags: , , | 11 Comments

Leisurely Cruise February 2026: power, shpower, positive predictive value

2025-6 Leisurely Cruise

The following is the February stop of our leisurely cruise (meeting 6 from my 2020 Seminar at the LSE). There was a guest speaker, Professor David Hand. Slides and videos are below. Ship StatInfasSt may head back to port or continue for an additional stop or two, if there is interest. Although I often say on this blog that the classical notion of power, as defined by Neyman and Pearson, is one of the most misunderstood notions in stat foundations. I did not know, in writing SIST, just how ingrained those misconceptions would become. I’ll write more on this in my next post. (The following is from SIST pp. 354-356, the pages are provided below)

Shpower and Retrospective Power Analysis

It’s unusual to hear books condemn an approach in a hush-hush sort of way without explaining what’s so bad about it. This is the case with something called post hoc power analysis, practiced by some who live on the outskirts of Power Peninsula. Psst, don’t go there. We hear “there’s a sinister side to statistical power, … I’m referring to post hoc power” (Cumming 2012, pp. 340-1), also called observed power and retrospective (retro) power. I will be calling it shpower analysis. It distorts the logic of ordinary power analysis (from insignificant results). The “post hoc” part comes in because it’s based on the observed results. The trouble is that ordinary power analysis is also post-data. The criticisms are often wrongly taken to reject both. Continue reading

Categories: 2025-2026 Leisurely Cruise, power | Leave a comment

Severe testing of deep learning models of cognition (ii)

.

From time to time I hear of an application of the severe testing philosophy in intriguing ways in fields I know very little about. An example is a recent article by cognitive psychologist Jeffrey Bowers and colleagues (2023): “On the importance of severely testing deep learning models of cognition” (abstract below). Because deep neural networks (DNNs)–advanced machine learning models–seem to recognize images of objects at a similar or even better rate than humans, many researchers suppose DNNs learn to recognize objects in a way similar to humans. However, Bowers and colleagues argue that, on closer inspection, the evidence is remarkably weak, and “in order to address this problem, we argue that the philosophy of severe testing is needed”.

The problem is this. Deep learning models, after all, consist of millions of (largely uninterpretable) parameters. Without understanding how the black box model moves from inputs to outputs, it’s easy to see why observed correlations can easily occur even where the DNN output is due to a variety of factors other than using a similar mechanism as the human visual system. From the standpoint of severe testing, this is a familiar mistake. For data to provide evidence for a claim, it does not suffice that the claim agrees with data, the method must have been capable of revealing the claim to be false, (just) if it is. Here the type of claim of interest is that a given algorithmic model uses similar features or mechanisms as humans to categorize images.[1] The problem isn’t the engineering one of getting more accurate algorithmic models, the problem is inferring claim C: DNNs mimic human cognition in some sense (they focus on vision), even though C has not been well probed. Continue reading

Categories: severity and deep learning models | 5 Comments

(JAN #2) Leisurely cruise January 2026: Excursion 4 Tour II: 4.4 “Do P-Values Exaggerate the Evidence?”

2026-26 Cruise

Our second stop in 2026 on the leisurely tour of SIST is Excursion 4 Tour II which you can read here. This criticism of statistical significance tests takes a number of forms. Here I consider the best known.  The bottom line is that one should not suppose that quantities measuring different things ought to be equal. At the bottom you will see links to posts discussing this issue, each with a large number of comments. The comments from readers are of interest! We will have a zoom meeting Fri Jan 23 11AM ET on these last two posts.*If you want to join us, contact us.

getting beyond…

Excerpt from Excursion 4 Tour II*

4.4 Do P-Values Exaggerate the Evidence? Continue reading

Categories: 2026 Leisurely Cruise, frequentist/Bayesian, P-values | Leave a comment

(JAN #1) Leisurely Cruise January 2026: Excursion 4 Tour I: The Myth of “The Myth of Objectivity” (Mayo 2018, CUP)

2025-26 Cruise

Our first stop in 2026 on the leisurely tour of SIST is Excursion 4 Tour I which you can read here. I hope that this will give you the chutzpah to push back in 2026, if you hear that objectivity in science is just a myth. This leisurely tour may be a bit more leisurely than I intended, but this is philosophy, so slow blogging is best. (Plus, we’ve had some poor sailing weather). Please use the comments to share thoughts.

.

Tour I The Myth of “The Myth of Objectivity”*

Objectivity in statistics, as in science more generally, is a matter of both aims and methods. Objective science, in our view, aims to find out what is the case as regards aspects of the world [that hold] independently of our beliefs, biases and interests; thus objective methods aim for the critical control of inferences and hypotheses, constraining them by evidence and checks of error. (Cox and Mayo 2010, p. 276) [i]

Continue reading

Categories: 2026 Leisurely Cruise, objectivity, Statistical Inference as Severe Testing | Leave a comment

Midnight With Birnbaum: Happy New Year 2026!

.

Anyone here remember that old Woody Allen movie, “Midnight in Paris,” where the main character (I forget who plays it, I saw it on a plane), a writer finishing a novel, steps into a cab that mysteriously picks him up at midnight and transports him back in time where he gets to run his work by such famous authors as Hemingway and Virginia Wolf?  (It was a new movie when I began the blog in 2011.) He is wowed when his work earns their approval and he comes back each night in the same mysterious cab…Well, ever since I began this blog in 2011, I imagine being picked up in a mysterious taxi at midnight on New Year’s Eve, and lo and behold, find myself in the 1960s New York City, in the company of Allan Birnbaum who is is looking deeply contemplative, perhaps studying his 1962 paper…Birnbaum reveals some new and surprising twists this year! [i] 

(The pic on the left is the only blurry image I have of the club I’m taken to.) It has been a decade since  I published my article in Statistical Science (“On the Birnbaum Argument for the Strong Likelihood Principle”), which includes  commentaries by A. P. David, Michael Evans, Martin and Liu, D. A. S. Fraser, Jan Hannig, and Jan Bjornstad. David Cox, who very sadly did in January 2022, is the one who encouraged me to write and publish it. Not only does the (Strong) Likelihood Principle (LP or SLP) remain at the heart of many of the criticisms of Neyman-Pearson (N-P) statistics and of error statistics in general, but a decade after my 2014 paper, it is more central than ever–even if it is often unrecognized.

OUR EXCHANGE:

ERROR STATISTICIAN: It’s wonderful to meet you Professor Birnbaum; I’ve always been extremely impressed with the important impact your work has had on philosophical foundations of statistics.  I happen to have published on your famous argument about the likelihood principle (LP).  (whispers: I can’t believe this!) Continue reading

Categories: Birnbaum, CHAT GPT, Likelihood Principle, Sir David Cox | Leave a comment

For those who want to binge read the (Strong) Likelihood Principle in 2025

.

David Cox’s famous “weighing machine” example” from my last post is thought to have caused “a subtle earthquake” in foundations of statistics. It’s been 11 years since I published my Statistical Science article on this, Mayo (2014), which includes several commentators, but the issue is still mired in controversy. It’s generally dismissed as an annoying, mind-bending puzzle on which those in statistical foundations tend to hold absurdly strong opinions. Mostly it has been ignored. Yet I sense that 2026 is the year that people will return to it again. It’s at least touched upon in Roderick Little’s new book (pic below). This post gives some background, and collects the essential links that you would need if you want to delve into it. Many readers know that each year I return to the issue on New Year’s Eve…. But that’s tomorrow.

By the way, this is not part of our lesurely tour of SIST. In fact, the argument is not even in SIST, although the SLP (or LP) arises a lot. But if you want to go off the beaten track with me to the SLP conundrum, here’s your opportunity. Continue reading

Categories: 11 years ago, Likelihood Principle | Leave a comment

67 Years of Cox’s (1958) Chestnut: Excerpt from Excursion 3 Tour II

2025-26 Cruise

.

We’re stopping to consider one of the “chestnuts” in the exhibits of “chestnuts and howlers” in Excursion 3 (Tour II) of Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars (SIST 2018). It is now 67 years since Cox gave his famous weighing machine example in Sir David Cox (1958)[1]. It will play a vital role in our discussion of the (strong) Likelihood Principle later this week. The excerpt is from SIST (pp. 170-173).

Exhibit (vi): Two Measuring Instruments of Different Precisions. Did you hear about the frequentist who, knowing she used a scale that’s right only half the time, claimed her method of weighing is right 75% of the time? 

She says, “I flipped a coin to decide whether to use a scale that’s right 100% of the time, or one that’s right only half the time, so, overall, I’m right 75% of the time.” (She wants credit because she could have used a better scale, even knowing she used a lousy one.)

Basis for the joke: An N-P test bases error probability on all possible outcomes or measurements that could have occurred in repetitions, but did not. Continue reading

Categories: 2025 leisurely cruise, Birnbaum, Likelihood Principle | Leave a comment

Blog at WordPress.com.