Dr Max Grossmann

Multiple hypothesis testing misconceptions

Posted: 2026-09-27 · Last updated: 2026-09-28 ·

In my experience, experimental economists handle multiple hypothesis testing (MHT) in one of two ways: they adjust nothing, or they adjust everything that moves. Both are generally wrong.

Which tests to adjust for is not really a statistical question. It is an epistemological question: what will you claim, and how can we avoid errors in what we think we know? The most useful distinction I know is whether your hypotheses compete (Hochberg and Tamhane, 1987, ch. 1). Hypotheses that do not compete come in two kinds, so there are three cases:

  • Hypotheses compete if any one of them being significant would give you a finding. You must adjust for them.
  • Hypotheses are complementary if you claim a finding only when all of them are significant. You do not need to adjust.
  • Hypotheses are separate if each supports its own claim, which stands or falls regardless of the others. Each forms its own family, and you do not adjust across them.

Everything below follows from this distinction. Before you see the data, write down which of your hypotheses compete, which complement each other, and which are separate, and how that will change what you will write in a paper. After that, choosing the correction is typically simple.

Hochberg and Tamhane make the point with two examples that differ only in the claim. Several new drugs are tested against a common control, and any of them could be approved: the treatments compete, and a joint error rate is required (their Example 3.1). Several school boards evaluate their own programs against a common control: the treatments are “noncompeting,” and no joint error rate is required (their Example 3.5). In my experience, many economists are unaware of this contrast, although it settles most questions about what to adjust for.

Competing hypotheses form a family

A family is the set of hypotheses over which you want to control the risk of a false conclusion (Hochberg and Tamhane, 1987, ch. 1, §2.1). In practice, a family is a set of competing hypotheses.

Here is why. Take ten independent tests, each at 5%, and suppose all ten nulls are true. The chance that at least one of them rejects is $1-0.95^{10}\approx40\%$. Each individual test is perfectly valid. But if your rule is “I have a world-historical scientific finding if any of these rejects,” you are running a 40% test, not a 5% test. Crucially, preregistering all ten tests does not per se change this fact.

To find out whether two tests compete, ask yourself: had the other result been significant instead, would I have used it to advance the paper? If yes, they compete and belong in the same family. This is an old test: Hochberg and Tamhane (1987, ch. 1, §2.1) count every inference that “would have been from the start equally plausible” (Fisher) or that would have made the researcher “sit up and take notice” (Putter), not only the ones you report. This applies to outcomes, treatments, subgroups, specifications, and even separate experiments. Moving a test to a different table or to the appendix does not per se change its role.

Boulesteix and Hoffmann (2024) reach a similar conclusion: adjust “only where authors use the significance of statistical tests to weight the reporting, discussion, and interpretation of their findings,” whether the study is labeled exploratory or confirmatory. They go further and argue that reporting all tests “with equal emphasis, regardless of the significance of the findings” makes adjustment unnecessary. Equal emphasis alone does not make adjustment unnecessary. If any one significant result would still count as a finding, the tests compete, however evenly you report them.

One experiment, three stories

Suppose a theory predicts that an information treatment raises both participants’ beliefs about Group A’s cooperation and Group B’s cooperation. You preregister two one-sided tests, one per outcome. What you need to adjust for depends entirely on what you will claim.

Story 1: either result is a finding. A positive effect on beliefs for Group A would be a finding. So would a positive effect for Group B. The two hypotheses compete. Suppose neither effect exists, so both nulls are true. If you test each at 5%, your chance of wrongly rejecting at least one of them, and thus reporting a finding that is not there, can be close to 10%. So adjust: Bonferroni, for example, tests each at 2.5% (Holm is better; see below). Adjustment is not free: it trades fewer false positives for more false negatives, because each test loses power. This is why you should adjust only for hypotheses that actually compete.

Story 2: you need both. You will claim the predicted pattern only if both tests reject. The hypotheses are complementary. Test each at 5% and do not adjust. The p-value for the joint claim is simply the larger of the two:

\[ p_{\text{both}}=\max(p_{\text{Group A}},p_{\text{Group B}}). \]

If this joint claim is false, at least one of the two effects is zero or negative. To wrongly claim that both are positive, you must wrongly reject that outcome’s null, which here happens at most 5% of the time. This holds for any dependence between the tests and for any number of required results. It is called an intersection–union test (Berger, 1982; Hochberg and Tamhane, 1987, Example 3.2; Benjamini, 2010, §6.2).

You may read elsewhere that “must succeed on every measure” calls for familywise error control (e.g., Hochberg and Tamhane, 1987, Example 3.4). For the joint claim, it does not. Since you claim anything only when every test rejects, any false claim requires rejecting some true null, and each test has size 5%. So even the probability that any of your individual claims is false stays at or below 5%. The price is power: every result you require makes the joint claim harder to establish.

Story 3: you hope for both, but either will do. This is the most common case, and it is Story 1 in disguise. If one test fails and the other gets promoted to the abstract, the hypotheses were competing all along, whatever you called them in your preregistration. Psaradakis (2000) documents a close analogue from time-series econometrics. Researchers run a battery of tests for nonlinearity, and Barnett et al. (1997, p. 189), whom he quotes, argue that “some of the ‘competing’ tests could be viewed as complementary, rather than competing.” But the series is declared nonlinear if any one test rejects. The tests therefore compete, whatever they are called, and the chance of falsely rejecting linearity grows with every test added. (There, all tests share a single null, linearity, but the logic is the same: an “any one will do” rule needs adjustment.)

Story 2 only works if you fix the required pattern in advance and stick to it. If you check several patterns and report whichever one holds up, those patterns compete, and you are back in Story 1.

Robustness checks do not compete with your main analysis

Suppose you have four competing primary treatment comparisons, and you adjust for that family of four. Then you re-estimate each comparison under six alternative specifications. Do you now need to adjust for a family of 28? No.

Why? The robustness checks do not compete with your main tests. Your preregistered main specification decides which effects you claim, and the checks only show how sensitive those estimates are. They cannot make the main tests reject more often, and no finding is selected from them, so they need no correction (Benjamini, 2010, §4.4: unadjusted inference is appropriate where neither selection nor simultaneity is at play). If a check undermines a finding, say so and temper the claim.

This changes as soon as you let the checks do more. If any specification can establish the effect, the specifications compete, and you must adjust for all of them. If you require the effect under every prespecified specification, they are complementary, and Story 2 applies. To summarize, a robustness check that can rescue a failed main result is a competing hypothesis.

Mechanisms compete with each other, not with the main effect

Commonly, a paper has some main hypothesis and then asks why it happens. Suppose an information treatment raises contributions, and you test three candidate channels: the treatment shifts beliefs about others’ contributions, shifts perceived social norms, or improves understanding of the game. These tests provide evidence about candidate mechanisms; they do not by themselves identify how the treatment changed contributions.

The main effect and the mechanisms go in separate families. The main test decides whether there is an effect. The mechanism tests explain it, and none of them can substitute for the main result. So you test the main effect at 5% without adjusting for the mechanisms.

Test the mechanisms only if the main effect is significant. Clinical trials apply the same logic when they test secondary endpoints only if the primary endpoint is positive (Chowdhry et al., 2024). This gatekeeping costs you nothing if your claim about mechanisms is conditional on first establishing the main effect, and it keeps the chance of any false claim across the main test and the mechanisms at or below 5%. If there is no main effect, any false claim requires first wrongly rejecting the main null, which happens at most 5% of the time. If there is a main effect, errors can only come from the mechanism tests, which you control at 5% (see next paragraph). If you instead test the mechanisms regardless, the chance of at least one false claim across the two families can exceed 5%.

The mechanisms, however, compete with each other. If beliefs had not moved but norms had, you would have told a norms story instead. Any one of the three would have become “the mechanism™.” Adjust within the family of three, for example with Holm. If your theory instead requires all three channels to move, they are complementary, and Story 2 applies: no adjustment.

The same caveat as for robustness checks applies. If a significant mechanism would let you claim success when the main effect fails (“no effect on contributions, but the treatment moved beliefs”), then the mechanisms compete with the main effect too, and all four tests belong in one family.

Sharing data does not create competition

Two school boards each evaluate their own program against a shared control group to save money. Each board makes its own decision based on its own comparison, regardless of what happens to the other. The hypotheses do not compete, so each board tests at 5%. Hochberg and Tamhane (1987, Example 3.5) call these “noncompeting treatments”: each inference is autonomous, although the tests are statistically dependent through the shared control. If neither program works, the probability of at least one false finding across both boards exceeds 5%, but that event is simply not decision-relevant.

This is why the issue is epistemological, not statistical: nothing in the data or the tests distinguishes the boards' situation from the one below; only the claims do.

But suppose that now a researcher goes through the same results looking for a successful program to feature in a paper. Either program would do. The hypotheses now compete, and the two tests form a family. The data did not change; the claim did. Likewise, putting unrelated questions into the same dataset or the same paper does not make them compete. Justify each family by the claims you make.

What to do in practice

  1. Preregister your claims, not just your tests. State which hypotheses compete, which are complementary, and which are separate, and what you will conclude if the main prediction fails. Also fix outcomes and specifications (Olken, 2015). Twenty preregistered regressions without a stated claim structure leave plenty of room to pick a story after the fact.
  2. Do not count coefficients. Ten control variables do not add ten tests as long as you make no claims about them. But if you run ten regressions and report whichever gives a significant treatment effect, those ten compete. Adjust for all ten, including the ones that failed.
  3. Use Holm for competing hypotheses. The familywise error rate (FWER) is the probability of at least one false rejection in the family. Holm’s (1979) procedure controls it for any dependence between the tests and whether or not some effects are real. It rejects everything Bonferroni rejects, and sometimes more. In R: p.adjust(p_primary, method = "holm"), then compare the adjusted p-values with 0.05. If you screen many hypotheses and can live with a few false positives, control the false discovery rate (FDR) instead: the expected share of false rejections among all rejections, counted as zero when nothing is rejected (Benjamini, 2010). The standard procedure is Benjamini–Hochberg, p.adjust(p, method = "BH"), which is valid for independent or positively dependent tests; under arbitrary dependence, use method = "BY".
  4. Exploit correlation for power. Correlated outcomes still need adjustment. But methods that account for the correlation gain power over Holm. Use List, Shaikh, and Vayalinkal (2023), who provide Stata code. Alternatively, combine the outcomes into a single prespecified index (Anderson, 2008), which replaces the competing hypotheses with one. If you have several indices, they compete in turn, and Anderson adjusts across them. Note that an effect on the index does not establish an effect on each component. If your claim is only that the treatment moved at least one outcome, you can also combine the p-values directly with the Cauchy combination test (Liu and Xie, 2020): transform each p-value to \(\tan((0.5-p_i)\pi)\), average, and compare the average with a standard Cauchy distribution. The resulting p-value is easy to compute, remains approximately valid when the tests are correlated (most accurately for small p-values), and has good power when only a few outcomes move. Like the index, it does not tell you which outcome moved.
  5. Report the family. State which hypotheses you treated as competing and which error rate you controlled. Report effect sizes along with raw and adjusted p-values. Say whether your confidence intervals are individual or simultaneous (Hothorn, Bretz, and Westfall, 2008). In general, a significant joint test of “no treatment has any effect” does not license unadjusted follow-up tests of which treatments have an effect (Hochberg and Tamhane, 1987, ch. 1, §1).

Determining the right family requires thinking (remember that?). You need to know what you will claim. Once you do, the rule is simple: competing hypotheses need adjustment to control FWER for the claim; complementary hypotheses, and hypotheses supporting genuinely separate claims, do not require adjustment within the same family.

References