Active Recall Statistics: 45 Findings on the Testing Effect
Students who read a passage once and then practiced recalling it remembered 61% of it one week later, compared with 40% for students who read it four times (Roediger and Karpicke, 2006). Across 222 classroom studies and 48,478 students, quizzing raised achievement by about half a standard deviation (Yang et al., 2021). Below are 45 statistics on active recall, retrieval practice and the testing effect, each checked against the original paper.
Key statistics
Students who studied a prose passage once and recalled it three times remembered 61% of it after one week, compared with 40% for students who read it four times (Roediger and Karpicke, 2006).
On a test given five minutes after studying, the same pattern reversed: repeated readers recalled 83% and repeated testers 71% (Roediger and Karpicke, 2006).
Retrieval practice beat concept mapping on a science test one week later, 67% vs 45% correct, an effect size of d = 1.50 (Karpicke and Blunt, 2011).
101 of 120 students (84%) learned more from retrieval practice than from building concept maps (Karpicke and Blunt, 2011).
A meta-analysis of 159 effect sizes from 61 studies found an average testing effect of g = 0.50 compared with restudying (Rowland, 2014).
A meta-analysis of 272 effect sizes from 118 articles and 15,427 participants found practice tests beat other learning conditions by g = 0.61 (Adesope, Trevisan and Sundararajan, 2017).
Across 222 classroom studies with 48,478 students, quizzing raised academic achievement by g = 0.499 (Yang et al., 2021).
Testing effects roughly doubled when learners got feedback: g = 0.73 with feedback vs g = 0.39 without (Rowland, 2014).
Sixth graders scored 94% on chapter exam questions they had been quizzed on in class, compared with 81% on questions they had not (Roediger, Agarwal, McDaniel and McDermott, 2011).
Only 11% of 177 college students reported practicing retrieval when studying, while 84% reported rereading (Karpicke, Butler and Roediger, 2009).
Answering questions before a lesson improved learning of the questioned content by g = 0.66 but had no measurable effect on other content (g = 0.01) (King-Shepard et al., 2025).
Landmark experiments
Two lab studies, one from 2006 and one from 2011, supply most of the numbers people quote about active recall, and both show that retrieval wins on delayed tests but not on immediate ones.
Roediger and Karpicke (2006), Psychological Science
Washington University undergraduates read short prose passages (about 256 to 275 words, on "The Sun" and "Sea Otters") and then either reread them or wrote down everything they could remember, with no feedback. Source: Roediger and Karpicke, 2006.
Experiment 1 (120 students), one week later: students who had taken one recall test remembered 56% of the passage, compared with 42% for students who had restudied it (d = 0.83).
Experiment 1, two days later:68% for tested students vs 54% for restudiers (d = 0.95).
Experiment 1, five minutes later: restudying won, 81% vs 75%.
Experiment 2 (180 students), one week later: read once and recall three times (STTT) produced 61% recall, read three times and recall once (SSST) 56%, and read four times (SSSS) 40%. This is the source of the widely quoted "61% vs 40%" figure.
Experiment 2, five minutes later: the order flipped, with SSSS at 83%, SSST at 78% and STTT at 71%.
Forgetting over the week: the four-reading group lost 52% of what they had recalled initially, the single-test group 28%, and the repeated-test group only 14%.
Confidence: the students who reread four times were the most confident they would remember the passage a week later, and they recalled the least.
Repeated reading gives the best score five minutes later and the worst score a week later. Repeated recall does the opposite.
Karpicke and Blunt (2011), Science
This study pitted retrieval practice against concept mapping, an elaborative technique widely used in science classes, with learning time matched. Source: Karpicke and Blunt, 2011.
Experiment 1 (80 students): on a short-answer test one week later, retrieval practice scored 0.67 vs 0.45 for concept mapping, which the authors describe as "about a 50% improvement" (d = 1.50).
Concept mapping was not significantly better than simply spending extra time rereading.
Experiment 2 (120 students): retrieval practice again won on the short-answer test (d = 1.07), and it also won when the final test was to draw a concept map from memory (d = 1.01).
101 of 120 students (84%) performed better after retrieval practice than after concept mapping.
The advantage held for inference questions that required connecting several ideas, not just for verbatim recall. Students predicted the reverse: they expected rereading and concept mapping to work better.
Meta-analyses: how big is the testing effect?
Pooled across hundreds of studies, retrieval practice improves test performance by roughly half a standard deviation compared with restudying, with larger effects over longer delays.
Major meta-analyses of retrieval practice and the testing effect
Meta-analysis
Scope
Size of evidence base
Overall effect
Rowland (2014), Psychological Bulletin
Testing vs restudying, mostly lab studies
159 effect sizes, 61 studies (1975 to 2013)
g = 0.50
Adesope, Trevisan and Sundararajan (2017), Review of Educational Research
Yang, Luo, Vadillo, Yu and Shanks (2021), Psychological Bulletin
Quizzing in real classrooms
573 effects, 222 studies, 48,478 students
g = 0.499 (0.427 after bias correction)
King-Shepard et al. (2025), Educational Psychology Review
Prequestions asked before learning
About 12,000 participants, multilevel model
g = 0.66 for prequestioned content; g = 0.01 for other content
Rowland (2014)
81% of the 159 effect sizes favored testing, 18% favored restudying and 1% were zero. Rowland, 2014
Final tests given one day or more after learning showed larger testing effects (g = 0.69) than tests given the same day (g = 0.41).
Published studies averaged g = 0.58 and unpublished ones g = 0.25. Both were reliably above zero, but the gap is a reminder that the headline figure may be inflated by publication bias.
Final tests that required recall showed bigger benefits than recognition tests. In the high-exposure subset, cued recall (g = 0.70) and free recall (g = 0.79) final tests beat recognition final tests (g = 0.32).
Adesope, Trevisan and Sundararajan (2017)
The random-effects estimate was higher than the fixed-effects one: g = 0.70 vs g = 0.61. Adesope et al., 2017
Classroom studies (g = 0.67) produced effects similar to laboratory studies (g = 0.62).
Secondary school students showed the largest effect (g = 0.83), compared with 0.64 for primary and 0.60 for postsecondary students.
A gap of 1 to 6 days between practice test and final test produced the largest effect (g = 0.82), compared with g = 0.56 for gaps under one day.
Multiple-choice practice tests worked well (g = 0.70), which undercuts the idea that only free recall counts as real retrieval.
Pan and Rickard (2018): transfer
Retrieval practice transfers to new questions and contexts, but less than it helps on the same questions. Across 192 transfer effect sizes, the average benefit was d = 0.40, largest for application and inference questions and weakest for untested material that had only been seen during initial study. After correcting for publication bias, the authors found transfer often disappeared when favorable conditions (matching responses, elaborated retrieval, good initial test performance) were absent. Pan and Rickard, 2018
Classroom and real-exam evidence
The testing effect holds up in real courses graded on real exams, where it is about the same size as in the lab.
Yang et al. (2021): quizzing beat no activity or filler by g = 0.610 and beat restudying by g = 0.330. Against other elaborative strategies such as note-taking and summarizing, the difference (g = 0.095) was not statistically significant. Yang et al., 2021
Quizzing helped conceptual learning (g = 0.644) at least as much as fact learning (g = 0.524), and it helped problem solving (g = 0.453).
Effects grew with treatment length, from g = 0.385 for a single class to g = 0.624 for interventions lasting longer than a semester.
By school level: elementary g = 0.328, middle school 0.597, high school 0.655, university 0.486.
Agarwal, Nunes and Blunt (2021) systematically reviewed applied school and classroom research: from nearly 2,000 abstracts they coded 50 experiments yielding 49 effect sizes (n = 5,374), and 57% showed medium or large benefits. Only 6% of the experiments were run outside WEIRD countries. Agarwal et al., 2021
Roediger et al. (2011): sixth grade social studies students (142 took part in the first experiment) scored 94% on quizzed items vs 81% on non-quizzed items in chapter exams, and 79% vs 67% on the end-of-semester exam. The researchers estimated that quizzing all items would have lifted grades from a B minus to an A. Roediger et al., 2011
In the same study, 97% of students said the in-class quizzes increased their learning and 65% said the quizzes reduced their test anxiety.
Larsen, Butler and Roediger (2009): in a randomized trial with 40 pediatric and emergency medicine residents, repeated testing with feedback produced final scores of 39% vs 26% for repeated study of a review sheet, measured about six months later (d = 0.91). Larsen et al., 2009
Dunlosky et al. (2013) rated 10 common study techniques and gave practice testing a "high utility" rating (along with distributed practice), while rereading and highlighting were rated "low utility." Dunlosky et al., 2013
Most students still reread
Karpicke, Butler and Roediger (2009) surveyed 177 college students about how they study. Source.
84% listed rereading as a strategy, and 55% said it was their number one strategy.
Only 11% (19 of 177) described practicing retrieval, and 1% (2 students) ranked it first.
40% reported using flashcards and 43% doing practice problems, though the authors note both can be used without real retrieval.
Asked what they would do after reading a textbook chapter, 57% chose to reread and only 18% chose to test themselves.
If making practice questions is the part that stops you, StudyCards AI turns a PDF of your notes or slides into flashcards you can quiz yourself on. For a rough estimate of what regular retrieval could do for your own study hours, try the flashcard effectiveness calculator.
How much, how often, and when
In classrooms, testing the same material more than once gives bigger gains, and quizzes after teaching beat quizzes before it.
Yang et al. (2021) found gains rose from g = 0.444 for items quizzed once to 0.601 for twice and 0.642 for three times or more. Quizzes allowing unlimited attempts reached g = 0.762.
Adesope et al. (2017) found the opposite pattern in their broader, mostly lab sample: a single practice test (g = 0.70) outperformed two or more (g = 0.51). The authors link this partly to differences in retention interval, so the question is not settled.
Post-class quizzes (g = 0.536) outperformed pre-class quizzes (g = 0.186), though pre-class quizzes still helped (Yang et al., 2021).
In real classrooms, quizzing the same material more often was linked to larger learning gains.
Pretesting effect
Trying to answer a question before you have learned the material improves memory for that specific material, as long as you see the correct answer afterward.
King-Shepard et al. (2025): a multilevel meta-analysis of more than 60 years of research found prequestions improved learning of the prequestioned content by g = 0.66, with no general benefit for other content in the same lesson (g = 0.01). Prequestions combined with feedback beat prequestions alone. King-Shepard et al., 2025
Richland, Kornell and Kao (2009): in all 5 experiments, students who tried to answer questions about a vision essay before reading it did better on a later test than students given extra reading time, even when only the pretest questions they had answered wrong were analyzed. Richland et al., 2009
Kornell, Hays and Bjork (2009): across 6 experiments using questions participants could not answer (including fictional trivia), unsuccessful retrieval attempts followed by the answer beat simply reading the question and answer together. Kornell et al., 2009
A 2023 review by Pan and Carpenter concluded that the benefit depends on having "an opportunity to study the correct answers afterwards" and that its size varies with procedure and how learning is measured. Pan and Carpenter, 2023
Feedback matters most after wrong answers, and without it, retrieval practice only helps if you get a fair share of answers right.
Rowland (2014): g = 0.73 with feedback vs g = 0.39 without. In no-feedback studies where learners got 50% or less right on the practice test, the testing effect vanished (g = 0.03); above 75% initial success it was g = 0.56.
Yang et al. (2021): classroom quizzes with corrective feedback produced g = 0.537 vs g = 0.374 without.
Adesope et al. (2017) found no significant difference: g = 0.63 with feedback and g = 0.60 without. Meta-analyses disagree on how much feedback adds, but none find it hurts.
Pashler et al. (2005): 258 people learned Luganda to English word pairs. Showing the correct answer after a wrong response raised one-week retention of those items by 494% relative to no feedback, while feedback after correct answers made little difference. Pashler et al., 2005
In the same study, confidence tracked accuracy closely: final-test answers given with very low to very high confidence were correct 15%, 37%, 61%, 84% and 90% of the time.
Our own data: difficulty ratings and self-graded answers
In our analysis of 18,217 flashcard reviews on StudyCards AI (September 2025 to July 2026), cards that students tagged "easy" were marked correct 95.1% of the time (n = 429), "medium" 72.3% (n = 16,456) and "hard" 47.4% (n = 365).
Read these numbers with three caveats.
Both the difficulty tag and "correct" are self-reports, entered at the same moment after the answer is revealed. This measures consistency between two self-judgments, not calibration against an objective score.
"Medium" is the default when a student does not pick a rating, so the medium group (95% of reviews) is mostly unrated cards. Only 794 reviews carry a deliberate easy or hard tag.
It is usage data from one app, so it shows correlation, not cause.
That fits the lab pattern. People judge individual answers reasonably well (Pashler's confidence data), but they misjudge which study strategy works: in Roediger and Karpicke (2006) and Karpicke and Blunt (2011), students expected rereading or concept mapping to beat retrieval, and they were wrong.
Common claims that are wrong or misleading
Several popular active recall numbers are misquoted versions of real findings.
"Active recall gives 80% retention after a week vs 40% for rereading." No condition in Roediger and Karpicke (2006) reached 80% after a week. The real figures are 61% (read once, recall three times) vs 40% (read four times). The only numbers near 80% are from the five-minute test, where rereading won.
"An effect size of 0.50 means you remember 50% more." It doesn't. Hedges g = 0.50 means the average tested student scores half a standard deviation above the average restudier. How that converts to percentage points depends on how spread out the scores are.
"Active recall improves retention by 50%." The closest source is Karpicke and Blunt (2011), where retrieval practice scored 0.67 vs 0.45 for concept mapping, "about a 50% improvement" in one experiment against one alternative. It is not a general rule. In Roediger and Karpicke's Experiment 1 the relative gain was 33% (56% vs 42%).
"Feedback improves retention by 494%." That is a relative gain for word pairs that learners had got wrong, compared with no feedback at all (Pashler et al., 2005). Relative gains look huge when the starting point is small, and the figure says little about feedback in general.
"Testing only helps rote memorization." Classroom quizzing helped conceptual learning (g = 0.644) and inference questions (Karpicke and Blunt, 2011). The real limit is narrower: benefits for untested material are smaller (g = 0.321 in Yang et al., 2021) and transfer depends on conditions (Pan and Rickard, 2018).
We included peer-reviewed experiments, meta-analyses and systematic reviews, favoring meta-analyses for effect sizes and landmark experiments for concrete percentages. Every number was checked against the original paper (full text where available, otherwise the publisher or PubMed abstract) in September 2026. We did not use secondary statistics roundups. Effect sizes are reported as the authors reported them (Hedges g or Cohen d), with the comparison condition named, because an effect against "no activity" is always larger than one against restudying. Our own StudyCards AI figures are labeled as app usage data.
You are welcome to use these statistics in your own writing. Please credit StudyCards AI with a link to this page.
APA
Groves, C. (2026). Active Recall Statistics: 45 Findings on the Testing Effect. StudyCards AI. https://studycardsai.com/blog/active-recall-statistics
HTML link
<a href="https://studycardsai.com/blog/active-recall-statistics">Active Recall Statistics: 45 Findings on the Testing Effect</a>
Frequently Asked Questions
Does active recall actually work?
Yes. A 2021 meta-analysis in Psychological Bulletin pooled 222 classroom studies with 48,478 students and found that quizzing raised achievement by g = 0.499, about half a standard deviation (Yang et al., 2021). Laboratory meta-analyses land in the same range: g = 0.50 across 159 effect sizes (Rowland, 2014) and g = 0.61 across 272 (Adesope et al., 2017).
What is the effect size of retrieval practice?
Most meta-analyses put it between g = 0.40 and g = 0.70. Rowland (2014) reported g = 0.50 against restudying, Adesope et al. (2017) g = 0.61, and Yang et al. (2021) g = 0.499 in real classrooms. Transfer to new questions or contexts is smaller, d = 0.40 (Pan and Rickard, 2018). An effect of 0.50 means half a standard deviation, not 50% better.
How much better is active recall than rereading?
In Roediger and Karpicke (2006), students who read a passage once and then recalled it three times remembered 61% of it a week later, compared with 40% for students who read it four times. Rereading actually won on a test five minutes later (83% vs 71%), which is why it feels effective while you are studying.
What percentage of students use active recall?
Few use it deliberately. In a survey of 177 college students, 84% listed rereading as a study strategy and 55% called it their number one strategy, while only 11% described practicing retrieval and just 1% ranked it first (Karpicke, Butler and Roediger, 2009). Asked what they would do after reading a chapter, only 18% chose to test themselves.
Does the pretesting effect work?
Yes, for the specific material that was pretested. A 2025 multilevel meta-analysis found that answering questions before a lesson improved learning of that prequestioned content by g = 0.66, but showed no general benefit for other content in the same lesson (g = 0.01). Adding feedback to the prequestions made the effect stronger (King-Shepard et al., 2025).
Do you need feedback for retrieval practice to work?
It helps, especially when you get answers wrong. Rowland (2014) found testing effects of g = 0.73 with feedback and g = 0.39 without. Without feedback, studies where learners recalled 50% or less on the practice test showed essentially no benefit (g = 0.03). In classrooms, feedback raised the effect from g = 0.374 to g = 0.537 (Yang et al., 2021).