Factor analysis
A statistical method that distills many correlated measures into a few underlying dimensions.
What it means
Factor analysis is a statistical technique that explains the pattern of correlations among many observed variables by positing a smaller number of unobserved underlying variables, called factors, that the measured items are taken to reflect. The logic is that items which rise and fall together do so because they tap a common latent dimension, so the method identifies clusters of inter-correlated items and estimates how strongly each item 'loads' on each factor. Exploratory factor analysis lets the data suggest how many factors exist and what they comprise, useful for discovering structure, while confirmatory factor analysis tests whether a pre-specified structure fits, used to validate measurement models. It is the workhorse behind much of psychometrics, from deriving the 'Big Five' personality dimensions to establishing that diverse cognitive tests share a general factor. Its results, however, hinge on judgment—how many factors to retain, how to rotate them, and how to name them—so it can impose tidy structure that is more interpretive than real. It matters because it underlies how psychological constructs are defined, measured, and validated.
How the math finds the factors
Factor analysis begins with the matrix of correlations among the observed items and splits each item's variance into two parts: common variance, shared with other items, and unique variance, made of item-specific content plus measurement error. Only the common variance is modeled. The method searches for a small set of factors whose weighted combinations reproduce the observed correlations as closely as possible. The loadings are those weights, and an item's communality is the share of its variance the factors account for. Eigenvalues summarize how much variance each candidate factor captures. Because a factor is defined entirely by a pattern of covariation, it is a mathematical summary of how things move together, not an entity that has been located or observed anywhere.
The decisions that shape the answer
Several choices, none forced by the data, change what emerges. How many factors to keep is the most consequential. Kaiser's popular 'eigenvalue greater than one' rule, still a default in older software, reliably extracts too many factors; Horn's parallel analysis, which compares the real eigenvalues against those from random data of the same size, performs far better and is now the recommended standard. Rotation makes loadings easier to read without changing overall fit, and oblique rotations that let factors correlate are usually more realistic than orthogonal ones. A survey by Fabrigar and colleagues found researchers routinely defaulting to the weakest options. Naming the retained factors is interpretation, and a persuasive label can lend a cluster more reality than the numbers themselves warrant.
Not the same as principal component analysis
The two are routinely confused, and many published 'factor analyses' are in fact principal component analyses run by default. The distinction is real. Principal component analysis draws no line between common and unique variance; it simply repackages the total variance of the items into a smaller set of weighted composites, with no latent variable assumed to cause the correlations. Common factor analysis models only shared variance and treats the factor as an unobserved common cause of the items. In practice the two often yield similar-looking solutions, but components tend to absorb error into the loadings and can inflate them, and they answer a different question. When the goal is to understand underlying constructs rather than merely compress data, factor analysis is the appropriate tool.
What it cannot settle
Factor analysis describes how items covary; it cannot prove that a factor names anything real, and its tidy output invites over-reading. Two classic traps recur. Reification treats a statistical factor as a discovered object, as when a general intelligence factor is spoken of as a substance in the brain. The jingle-jangle fallacy assumes identically named scales measure the same construct and differently named ones measure different constructs, when the factor structure may say otherwise. Structure can also fail to travel: testing the Big Five inventory among Tsimane forager-farmers in Bolivia, Gurven and colleagues recovered only two broad dimensions (industriousness and prosociality), suggesting the familiar five-factor solution partly reflects the literate, industrialized samples it was built on rather than a human universal.
Examples
Administering dozens of personality questionnaire items and finding they collapse onto five correlated clusters is the empirical basis for the five-factor model of personality.
An engagement survey asks forty questions, and factor analysis shows the answers really move on three dimensions: pay, workload and trust in management. The board sees three charts, not forty.
Shoppers rate dozens of attributes across a supermarket's ranges and the answers collapse onto roughly two dimensions. Calling them value and quality is a judgment call, not a result the maths hands over.
A clinic scores patients on twenty depression and anxiety symptoms, and factor analysis shows the items load on two correlated dimensions, low mood and physical arousal, guiding which subscales the questionnaire should report.
Feeding dozens of national indicators — income, schooling, life expectancy, crime — into a factor analysis can reveal they move on just two underlying dimensions, letting an index rank countries on a few scores.
First described in Charles Spearman (1904).
Key references
- Gurven, M., von Rueden, C., Massenkoff, M., Kaplan, H., & Lero Vie, M. (2013). How universal is the Big Five? Testing the five-factor model of personality variation among forager-farmers in the Bolivian Amazon. Journal of Personality and Social Psychology, 104(2), 354-370. doi.org/10.1037/a0030841
- Costello, A. B., & Osborne, J. W. (2005). Best practices in exploratory factor analysis: Four recommendations for getting the most from your analysis. Practical Assessment, Research, and Evaluation, 10, Article 7. doi.org/10.7275/jyj1-4868
- Fabrigar, L. R., Wegener, D. T., MacCallum, R. C., & Strahan, E. J. (1999). Evaluating the use of exploratory factor analysis in psychological research. Psychological Methods, 4(3), 272-299. doi.org/10.1037/1082-989X.4.3.272
- Horn, J. L. (1965). A rationale and test for the number of factors in factor analysis. Psychometrika, 30(2), 179-185. doi.org/10.1007/BF02289447
- Spearman, C. (1904). "General intelligence," objectively determined and measured. The American Journal of Psychology, 15(2), 201-293. doi.org/10.2307/1412107
Where this comes up
- Response Bias in Surveys: A Behavioral Science PerspectiveSurvey responses aren't always accurate. Learn about the most common types of response bias, why they occur, and practi…