Behavioral Science Dictionary

Bayes factor

Methods & Evidence

A ratio measuring how strongly the data favor one model over another.

What it means

A Bayes factor is the ratio of the marginal likelihoods of the data under two competing models or hypotheses, summarizing how much the evidence shifts the balance between them. It is the multiplicative bridge from prior odds to posterior odds, and unlike a p-value it can quantify support for the null hypothesis, not merely fail to reject it. Conventional benchmarks describe factors of 3, 10, or 100 as moderate, strong, or decisive evidence, though these labels are heuristic. Bayes factors depend on the priors placed on parameters within each model, so they reward predictive accuracy while penalizing needless complexity.

What a marginal likelihood is

The number under each model is its marginal likelihood: the probability the model assigned to the data before seeing them, averaged over every parameter value the prior entertained. A sharp, well-placed prior concentrates its bets and scores high when the data land where predicted; a vague prior spreads the same probability thinly across many possible datasets, so even a good fit earns a smaller average. This averaging is the Occam mechanism at work, penalizing complexity automatically because flexibility is paid for in dispersed prior-predictive probability. The Bayes factor is one such average divided by another, which is why it rewards a model for predicting the data, not merely accommodating them.

Why the prior does real work

Because each marginal likelihood integrates over a prior, the Bayes factor inherits that prior's shape, and the dependence does not wash out as data accumulate. Widen the effect-size prior under the alternative and you dilute its predictions, pulling the factor back toward the null; this is the Jeffreys-Lindley paradox, in which a fixed p-value can coincide with ever-stronger Bayesian support for the null as the sample grows. Diffuse or improper priors make the factor arbitrary or undefined, because a constant with no natural scale fails to cancel. A reported Bayes factor is therefore a statement about a specific prior, and honest reporting pairs it with a sensitivity analysis across plausible widths.

Getting the number

Only in conjugate or low-dimensional problems does the marginal likelihood have a closed form; elsewhere it is a hard integral that must be estimated. For nested models the Savage-Dickey density ratio offers a shortcut: the Bayes factor equals the posterior height divided by the prior height at the null value, sidestepping the integral entirely. For general models, bridge sampling and its variants estimate each marginal likelihood from posterior draws and are the current workhorse in cognitive science, while naive harmonic-mean estimators are notoriously unstable and best avoided. Estimation noise is real, so a factor reported to three digits usually overstates the precision the sampler actually delivered.

Default tests and optional stopping

To spare users from specifying priors, psychology adopted default Bayes factors built on Jeffreys-Zellner-Siow priors: a Cauchy distribution on standardized effect size that is scale-free and yields a ready-made Bayesian t-test, now routine in software such as JASP. A second attraction is sequential use. Because the Bayes factor is a ratio of marginal likelihoods rather than a tail-area probability, you may inspect it after every observation and stop once the evidence is decisive, without the alpha inflation that dooms repeated significance testing. The evidence simply updates, with no penalty for looking. That freedom is genuine but conditional: it still assumes the two models being compared are the ones worth weighing, and de Heide and Grunwald caution that optional stopping can still bite a Bayesian who wants frequentist error control or whose priors are improper.

What it does not tell you

A Bayes factor compares two specified models and nothing else. It is silent on whether either is any good, so a decisive ratio can merely mean the data favor the less wrong of two poor candidates. It is also not a posterior probability: turning it into belief requires prior odds the analyst must supply, so the labels 3, 10, and 100 describe strength of evidence, not the chance a hypothesis is true. Critics press further: Tendeiro and Kiers argue that posterior model probabilities answer researchers' questions more directly, and Robert doubts that default priors carry any defensible meaning. Recent work even documents reversals, where two reasonable analyses of the same data point opposite ways.

Examples

A Bayes factor of 12 in favor of an effect indicates the data are twelve times more probable under the effect model than under the null.

An A/B test of a redesigned checkout button returns a Bayes factor of 8 for the null: not merely no significant difference, but positive evidence the button changes nothing.

A replication reports the same direction of effect but a Bayes factor near 1, meaning its data are about equally probable either way, so it settles nothing at all.

In forensic science, the likelihood ratio a court hears is a Bayes factor: it weighs how much more probable a DNA match is if the suspect, rather than an unrelated person, left the trace.

Cosmologists rank a model that includes dark energy against one that omits it by the Bayes factor between them, letting the data's averaged fit, not a significance cutoff, decide which universe to prefer.

First described in Harold Jeffreys (1935, 1961); Kass & Raftery (1995).

Key references

  1. Tendeiro, J. N., & Kiers, H. A. L. (2019). A review of issues about null hypothesis Bayesian testing. Psychological Methods, 24(6), 774-795. doi.org/10.1037/met0000221
  2. Gronau, Q. F., Sarafoglou, A., Matzke, D., Ly, A., Boehm, U., Marsman, M., Leslie, D. S., Forster, J. J., Wagenmakers, E.-J., & Steingroever, H. (2017). A tutorial on bridge sampling. Journal of Mathematical Psychology, 81, 80-97. doi.org/10.1016/j.jmp.2017.09.005
  3. Robert, C. P. (2016). The expected demise of the Bayes factor. Journal of Mathematical Psychology, 72, 33-37. doi.org/10.48550/arXiv.1506.08292
  4. Rouder, J. N., Speckman, P. L., Sun, D., Morey, R. D., & Iverson, G. (2009). Bayesian t tests for accepting and rejecting the null hypothesis. Psychonomic Bulletin & Review, 16(2), 225-237. doi.org/10.3758/PBR.16.2.225
  5. Kass, R. E., & Raftery, A. E. (1995). Bayes factors. Journal of the American Statistical Association, 90(430), 773-795. doi.org/10.1080/01621459.1995.10476572

← All 1001 terms