Equivalence testing
Also known as: Tests of equivalence
Formally testing that an effect is too small to matter, not just absent.
What it means
Equivalence testing is a procedure for providing positive evidence that an effect is negligibly small, rather than merely failing to find one. The researcher first defines a smallest effect size of interest, then tests whether the observed effect is significantly inside those equivalence bounds, commonly via two one-sided tests (TOST). This repairs a key flaw of conventional testing, where a non-significant result is wrongly read as proof of no effect, when it may just reflect low power. Widely used in pharmacology to show a generic drug performs like the original, it is increasingly adopted to interpret null results credibly across the behavioral sciences.
How the test works
The procedure runs two one-sided tests against the equivalence bounds rather than against zero. One test asks whether the effect is reliably above the lower bound; the other whether it is reliably below the upper bound. Only if both reject does the effect fall inside the equivalence region. An equivalent and often clearer formulation checks whether the whole confidence interval lies within the bounds, using a 90 percent interval to match a 5 percent test on each side. That 90 percent width is not a mistake: because each one-sided test spends its error rate in a single direction, the matching two-sided interval is narrower than the usual 95 percent one. Passing does not prove the effect is exactly zero, only that it is too small to reach the bound.
Setting the bounds is the hard part
The conclusion is only as meaningful as the equivalence bound chosen, and that choice is a judgment about what counts as trivial, not a statistical output. In regulated bioequivalence the bound is fixed by convention: a generic must keep the ratio of key pharmacokinetic measures within 80 to 125 percent of the reference drug. Behavioral research has no such standard, so the smallest effect size of interest must be argued from theory, from prior effects, from the resolution of the measure, or from the smallest change a decision-maker would act on. Bounds set after seeing the data, or set wide enough to guarantee a pass, drain the result of meaning. Because the whole inference rests on this number, it should be fixed before collection and reported alongside the estimate.
Where it shows up
Equivalence testing began in pharmacology, where regulators demand positive proof that a generic behaves like the branded original before approval, and it remains the backbone of bioequivalence and non-inferiority trials. In the behavioral sciences it has spread as a way to interpret null results honestly: large replication projects use it to declare an original effect smaller than anything worth caring about, rather than reporting an ambiguous non-significant result. Product teams apply the same logic to experiments, showing a redesign moved a metric by less than a pre-agreed threshold and can be shipped for other reasons. It also appears in assay validation, manufacturing, and instrument comparison, wherever the useful claim is sameness rather than difference.
Limits and what it cannot do
Equivalence testing is still null-hypothesis testing with the hypotheses swapped, so it inherits the same dependence on power: a small, noisy study can fail to establish equivalence for the same reason it fails to detect an effect, and a wide interval simply leaves the question open. It never proves the effect is zero, only that it is unlikely to exceed the bound. Pairing it with a conventional test yields four readable outcomes, since an effect can be significant, equivalent, both, or neither, which is more informative than either test alone. Bayesian alternatives such as a region of practical equivalence answer a related question in the language of posterior probability, and are usually reported alongside the frequentist test rather than in place of it.
Examples
Concluding a new pill is equivalent to the standard one by showing the difference falls reliably within a pre-set, clinically trivial margin.
Rather than reporting 'no significant difference' between two checkout buttons, a team shows the effect sits inside a half-point band they agreed in advance was too small to act on.
A large replication concludes the original effect, if it exists at all, is smaller than anything worth caring about — evidence of absence, not merely a failure to find it.
An education agency judges an online course equivalent to classroom teaching by showing the exam-score gap stays inside the small band it declared beforehand as educationally trivial, not merely non-significant.
A factory switching to a cheaper alloy supplier runs an equivalence test on tensile strength, confirming new batches sit within the tolerance engineers set in advance rather than just passing a difference test.
First described in Bioequivalence testing; Schuirmann's TOST (1987); Lakens (2017) in psychology.
Key references
- Lakens, D., Scheel, A. M., & Isager, P. M. (2018). Equivalence Testing for Psychological Research: A Tutorial. Advances in Methods and Practices in Psychological Science, 1(2), 259-269. doi.org/10.1177/2515245918770963
- Kruschke, J. K. (2018). Rejecting or Accepting Parameter Values in Bayesian Estimation. Advances in Methods and Practices in Psychological Science, 1(2), 270-280. doi.org/10.1177/2515245918771304
- Lakens, D. (2017). Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta-Analyses. Social Psychological and Personality Science, 8(4), 355-362. doi.org/10.1177/1948550617697177
- Schuirmann, D. J. (1987). A Comparison of the Two One-Sided Tests Procedure and the Power Approach for Assessing the Equivalence of Average Bioavailability. Journal of Pharmacokinetics and Biopharmaceutics, 15(6), 657-680. doi.org/10.1007/BF01068419