Behavioral Science Dictionary

Convergent validity

Methods & Evidence

Measures of the same construct should agree with one another.

What it means

Convergent validity is the extent to which different measures intended to capture the same construct correlate strongly with each other. High convergence is evidence that the measures tap a shared underlying concept rather than method-specific quirks. It is established alongside discriminant validity, typically within a multitrait-multimethod matrix that simultaneously checks that same-construct measures cohere and different-construct measures diverge. Convergence using different methods is especially persuasive, because agreement across formats argues that the construct, not the measurement procedure, is driving the result.

How it is quantified

Convergent validity is read off correlations, but the exact numbers depend on the framework. In a multitrait-multimethod matrix, the diagnostic entries are the monotrait-heteromethod cells, the same trait measured two ways, which should be substantial and clearly larger than the heterotrait correlations surrounding them. In confirmatory factor analysis, the evidence comes from standardized factor loadings, ideally above roughly 0.70, so each item shares about half its variance with the factor, and from the average variance extracted, where Fornell and Larcker's rule of thumb asks each construct's AVE to exceed 0.50. These cutoffs are conventions, not laws. A measure can clear every one of them and still be capturing something other than what it names.

The multimethod requirement

Agreement is only compelling when the methods genuinely differ. Two self-report scales given minutes apart will correlate partly because they share a format, a response style, and a passing mood, shared method variance that inflates the correlation without saying anything about the construct. Campbell and Fiske's central move was to demand convergence across unlike methods: self-report against observation, a questionnaire against a behavioural task, a survey against administrative records. When measures that have almost nothing in common except their target still line up, method artefacts become an implausible explanation and the shared construct becomes the parsimonious one. Convergence within a single method is weak evidence, easily mistaken for validity but consistent with everyone simply answering questionnaires the same way.

What it cannot establish alone

High convergence does not, by itself, certify a measure. If two scales correlate at 0.90, they may be tapping one construct, or they may be near-duplicates that both miss the intended target, or both carry the same nuisance such as acquiescence. This is why convergent validity is inert without its partner: a measure must also diverge from things it should not track. In practice the evidence is routinely misread. Reviews of applied work find that the Fornell-Larcker check is frequently misapplied, with researchers comparing average variance extracted against the wrong quantity or treating one passed threshold as proof. Newer diagnostics such as the heterotrait-monotrait ratio were proposed precisely because the older rules missed real discriminant-validity problems.

Related but distinct

Convergent validity is easy to confuse with reliability. Reliability asks whether a single measure agrees with itself, across items, occasions, or raters; convergent validity asks whether distinct measures of the same thing agree with each other. A perfectly reliable scale can still lack convergent validity if nothing it should resemble moves with it. Both sit under the older umbrella of construct validity, the question Cronbach and Meehl framed of whether a test measures the theoretical attribute it claims to. Convergent and discriminant validity are the two operational halves of that larger question: one shows the construct is really there, the other shows it is not just some other construct wearing a new name.

Examples

A new depression questionnaire shows convergent validity by correlating highly with an established depression scale and with clinician ratings.

A wrist tracker's sleep scores earn convergent validity by lining up with a sleep lab's readings and with the wearer's own diary, three very different methods pointing the same way.

A two-question burnout screener is trusted once its scores track the full inventory and match supervisors' independent ratings of the same staff; agreement across formats is what persuades.

A new reading-comprehension test gains convergent validity when children's scores track both their teachers' independent reading grades and their performance on an established standardized literacy exam.

In a customer-loyalty study, a survey's trust scale shows convergent validity when its scores line up not just with related self-report items but with an unlike measure of the same customers, such as observed repeat-purchase behaviour or complaint records.

First described in Campbell & Fiske (1959).

Key references

  1. Rönkkö, M., & Cho, E. (2022). An updated guideline for assessing discriminant validity. Organizational Research Methods, 25(1), 6-14. doi.org/10.1177/1094428120968614
  2. Henseler, J., Ringle, C. M., & Sarstedt, M. (2015). A new criterion for assessing discriminant validity in variance-based structural equation modeling. Journal of the Academy of Marketing Science, 43(1), 115-135. doi.org/10.1007/s11747-014-0403-8
  3. Fornell, C., & Larcker, D. F. (1981). Evaluating structural equation models with unobservable variables and measurement error. Journal of Marketing Research, 18(1), 39-50. doi.org/10.1177/002224378101800104
  4. Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81-105. doi.org/10.1037/h0046016
  5. Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. doi.org/10.1037/h0040957

← All 1001 terms