Feeling of knowing
The sense that you could recognize an answer you currently can't recall.
What it means
The feeling of knowing is a metacognitive judgment that unrecalled information is nonetheless stored and would be recognized, or retrieved later, if cued. On the dominant account it is inferred rather than read directly off the memory store, drawing on how familiar the question itself feels and on whatever partial or related material a failed retrieval attempt turns up. Such judgments predict later recognition above chance, a result reproduced across laboratories, materials and decades. But it is a modest within-person correlation rather than a dependable signal about any single item, and it differs across materials and criterion tests; a familiar cue alone can raise the feeling without teaching anything about the answer. It is closely related to, but distinguishable from, the tip-of-the-tongue state. The construct matters because people use it to decide what to restudy and whether to keep searching, though it forecasts recognition rather than unaided production, and cue priming has been shown to shift those choices without improving recall.
The original demonstration
Joseph Hart's 1965 experiments supplied both the name and the method. Participants answered general-information questions; for each one they failed, they predicted whether they could pick the correct answer from a set of alternatives; then they took that multiple-choice test. The sequence is known as the recall-judge-recognize procedure, and most later work uses a variant of it. Hart found that unrecalled items participants said they would recognize were indeed recognized far more often than items they said they would not, the latter falling close to guessing. What is striking is not the accuracy but that the ratings carried information at all, from people who had just failed to produce those answers. Hart read this as evidence of a monitor with some direct view of the memory store. The finding replicated widely; the interpretation did not. Direct access was still the leading reading in the mid-1980s; what displaced it was a run of experiments from the late 1980s into the early 1990s, roughly a quarter-century after Hart, favouring accounts in which the sensation is assembled after the fact from whatever the failed retrieval attempt threw up.
Why it happens
Two candidate mechanisms dominated the debate from the late 1980s onward, and neither is a description of the contents of memory. The first is cue familiarity: the judgment tracks how familiar the question feels, not anything about the answer. Reder and Ritter made the case with arithmetic problems: repeated exposure to a problem's terms raised participants' quick sense that they could retrieve the answer while teaching them nothing about it. Schwartz and Metcalfe found the complementary dissociation: priming the cue raised ratings; making the target more retrievable did not. The second is accessibility. Koriat's 1993 account treats the sensation as a by-product of the retrieval attempt. The more partial material surfaces — fragments, letters, associated facts, a sense of the answer's shape — the stronger the feeling, and the mechanism is blind to whether that material is correct. Ratings rose with the sheer quantity of accessible information, independently of later recognition. That evidence came mainly from nonsense letter strings rather than real knowledge, a limitation the follow-up addressed. Koriat and Levy-Sadot later argued that the two are stages rather than rivals: familiarity acts early and largely determines whether a search is launched, accessibility afterwards. The interaction they predicted appeared for delayed judgments but not immediate ones.
How it is measured
Almost all of the quantitative literature reports a Goodman-Kruskal gamma correlation, computed within each participant, between the ratings given to unrecalled items and whether they were later recognized. Nelson compared the candidate measures in 1984 and made the case for gamma. The literature also distinguishes resolution, ranking one's own items correctly, from calibration, whether the level of the feeling matches the actual probability of success. Gamma has since come under sustained attack. Masson and Rotello showed in 2009 that its measured value departs systematically from its true value under ordinary conditions, because response bias intrudes on a statistic adopted precisely to avoid such problems. Vuorre and Metcalfe added a second: because multiple-choice criterion tests permit correct guessing, resolution measures are biased downwards as first-order performance falls, even in simulations that hold the true first-order to second-order relationship fixed. Two groups can therefore differ in gamma without differing in monitoring at all — which is why signal-detection alternatives such as type-2 d-prime are now often reported alongside it, and why comparisons across populations differing in raw memory carry a confound. Two practical cautions follow. Gamma from a handful of items is very noisy, and usable items are capped by the design: only recall failures qualify for a rating. And a rating that predicts four-alternative recognition need not predict cued or free recall as well, so studies using different criterion tests are not comparable.
What the evidence shows
The core effect is solid. Above-chance prediction of later recognition for unrecalled items has been reproduced across laboratories, materials and decades. No pooled estimate of its size exists for the literature as a whole, and criterion tests vary too much for one figure to travel. What is not solid is the interpretation that people are reading their own memory. Koriat's 1995 follow-up, which moved the account from nonsense strings to general-knowledge questions, makes the point. Grading questions by how often the answers they precipitate are actually correct, he found the feeling-recognition correlation was positive for questions where accessible partial information tended to be right, and nil or negative where it tended to be wrong. The judgment is accurate only where the cues feeding it are valid; elsewhere it is systematically misleading rather than merely noisy. The best-powered synthesis available concerns aging. Devaluez, Mazancieux and Souchay pooled twenty studies comparing 922 younger and 966 older adults, yielding thirty effects, with gamma as the accuracy measure. Semantic material showed no age difference across eight effects (g = -0.10, 95% CI [-0.29, 0.10]). Episodic material did show a gap across twenty-two effects (g = 0.53, 95% CI [0.28, 0.78]) — but heterogeneity was severe (I-squared = 79%), the funnel plot was asymmetric in a way consistent with publication bias (z = 2.99, p = .003), and recognition performance moderated the effect decisively, the estimate no longer significant at zero age difference in recognition. In a separate pooled dataset of individual participants, matching older adults with the strongest recognition against younger adults with the weakest shrank the gap to g = 0.09, 95% CI [-0.24, 0.41], while recall itself differed enormously (g = 1.33). Their reading is that the gap reflects first-order memory differences and the known interaction between gamma, guessing and low performance, not a separate decline in monitoring. The pooled figure describes a memory difference wearing metacognitive clothing. The judgment also guides behaviour. Hanczakowski, Zawadzka and Cockcroft-McKay found that people choosing which unrecalled items to restudy preferred those carrying a strong feeling of knowing, and that restudy was more effective for them. But priming the cues raised ratings and shifted choices without improving recall: the control function inherits the vulnerability of the judgment behind it.
Where it breaks down
The most consequential failure mode follows from the familiarity findings. A question whose terms have been encountered recently will feel answerable whether or not the answer was ever learned, so the sensation is easiest to manufacture where wording is repeated: campaign copy, internal communications, revision material skimmed rather than tested. There is also a mismatch of criterion. The judgment forecasts recognition, and it is respectable at that. Most situations that matter demand production — stating the figure, drafting the clause, naming the person — and a strong sense of recognizing the answer says much less about generating it unaided. What the literature supports is a within-person rank correlation of moderate size, estimated from few observations, not a dependable signal about any one item. The graded laboratory rating should also be kept apart from its neighbours: a tip-of-the-tongue state, in which a blocked name arrives with its first letter and rhythm attached, has its own literature and its own still-live direct-access accounts, while judgments of learning and retrospective confidence are different measurements again. The mechanism, finally, remains only partly settled: familiarity and accessibility both contribute, the sequencing proposed for them holds under some conditions and not others, and their integration remains a working compromise.
Examples
Unable to recall a capital city, you feel confident you'd pick it correctly from a list — and usually you can.
A support agent cannot state an obscure clause in a refund policy from memory, but feels sure the wording would be obvious on sight, so searches the handbook instead of escalating the ticket. The feeling here is doing its useful job: directing search effort rather than supplying an answer.
A learner revising for an exam skips a topic on the grounds that it would come back if seen, then cannot produce any of it when the paper asks for a written answer. The judgment was calibrated to recognition; the test demanded recall.
A questionnaire repeats phrasing that respondents have met many times in recent advertising. What the familiarity work predicts, though it has not been established for survey instruments specifically, is that the terms alone should lift the sense of knowing the answer and depress use of the do-not-know option while leaving actual accuracy untouched.
Someone stuck on a colleague's surname retrieves the first letter and a sense of its rhythm, and concludes with confidence that the name is there. This is the tip-of-the-tongue form of the experience, sharper than the graded rating the laboratory procedure collects and studied under a heading of its own. Sometimes the name is indeed there; sometimes the fragments belong to a different name entirely, and the confidence tracks the fragments rather than the target.
First described in Hart (1965); cue-familiarity account by Reder & Ritter (1992).
Key references
- Hart, J. T. (1965). Memory and the feeling-of-knowing experience. Journal of Educational Psychology, 56(4), 208-216. doi.org/10.1037/h0022263
- Nelson, T. O. (1984). A comparison of current measures of the accuracy of feeling-of-knowing predictions. Psychological Bulletin, 95(1), 109-133. doi.org/10.1037/0033-2909.95.1.109
- Reder, L. M., & Ritter, F. E. (1992). What determines initial feeling of knowing? Familiarity with question terms, not with the answer. Journal of Experimental Psychology: Learning, Memory, and Cognition, 18(3), 435-451. doi.org/10.1037/0278-7393.18.3.435
- Schwartz, B. L., & Metcalfe, J. (1992). Cue familiarity but not target retrievability enhances feeling-of-knowing judgments. Journal of Experimental Psychology: Learning, Memory, and Cognition, 18(5), 1074-1083. doi.org/10.1037/0278-7393.18.5.1074
- Koriat, A. (1993). How do we know that we know? The accessibility model of the feeling of knowing. Psychological Review, 100(4), 609-639. doi.org/10.1037/0033-295X.100.4.609
- Koriat, A. (1995). Dissociating knowing and the feeling of knowing: Further evidence for the accessibility model. Journal of Experimental Psychology: General, 124(3), 311-333. doi.org/10.1037/0096-3445.124.3.311
- Koriat, A., & Levy-Sadot, R. (2001). The combined contributions of the cue-familiarity and accessibility heuristics to feelings of knowing. Journal of Experimental Psychology: Learning, Memory, and Cognition, 27(1), 34-53. doi.org/10.1037/0278-7393.27.1.34
- Masson, M. E. J., & Rotello, C. M. (2009). Sources of bias in the Goodman-Kruskal gamma coefficient measure of association: Implications for studies of metacognitive processes. Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(2), 509-527. doi.org/10.1037/a0014876
- Hanczakowski, M., Zawadzka, K., & Cockcroft-McKay, C. (2014). Feeling of knowing and restudy choices. Psychonomic Bulletin & Review, 21(6), 1617-1622. doi.org/10.3758/s13423-014-0619-0
- Vuorre, M., & Metcalfe, J. (2022). Measures of relative metacognitive accuracy are confounded with task performance in tasks that permit guessing. Metacognition and Learning, 17(2), 269-291. doi.org/10.1007/s11409-020-09257-1
- Devaluez, M., Mazancieux, A., & Souchay, C. (2023). Episodic and semantic feeling-of-knowing in aging: A systematic review and meta-analysis. Scientific Reports, 13, Article 16439. doi.org/10.1038/s41598-023-36251-9