Spaced Repetition Statistics (2026): 40+ Research Findings
Spacing out study sessions beat cramming the same material into one session in 259 of 271 comparisons in the largest review of the spacing effect (Cepeda et al., 2006, Psychological Bulletin). Newer meta-analyses put the advantage at a moderate to large effect size: d = 0.54 in real classrooms (Mawson and Kang, 2025) and SMD = 0.78 across 21,415 medical learners (Maye and Hurley, 2026).
Below are 40+ spaced repetition statistics, each checked against the original paper or dataset, with the sample size and the limits of each finding stated plainly. Where a popular version of a number is misleading, we say so.
Key statistics
Spaced study produced better recall than massed study in 259 of 271 comparisons with equal study time (Cepeda et al., 2006).
A 2026 meta-analysis of 13 medical education studies covering 21,415 learners found spaced repetition beat standard study techniques on objective tests, SMD = 0.78 (Maye and Hurley, 2026).
Across 22 classroom reports with more than 3,000 students, distributed practice beat massed practice with an effect size of d = 0.54 (Mawson and Kang, 2025).
A meta-analysis of 63 studies and 112 effect sizes found spaced practice beat massed practice with a mean weighted effect size of d = 0.46 (Donovan and Radosevich, 1999).
Spacing out retrieval practice beat massed retrieval practice with an effect size of g = 0.74 across 29 studies (Latimier, Peyre and Ramus, 2021).
Expanding review intervals were no better than evenly spaced intervals, g = 0.034 (Latimier, Peyre and Ramus, 2021).
With the same study time, the best review gap increased final recall by 64% compared with no gap, in a study of 1,354 learners (Cepeda et al., 2008).
The optimal review gap fell from about 43% of a 7-day test delay to about 8% of a 350-day test delay (Cepeda et al., 2008).
Spacing beat massing for 90% of flashcard learners, yet 72% believed massing had worked better (Kornell, 2009).
FSRS-6 predicted Anki recall with a log loss of 0.3456, versus 0.616 for Anki's default SM-2, across 9,999 user collections (open-spaced-repetition srs-benchmark, 2025).
68.3% of 560 US medical students surveyed in 2020 used Anki (Halperin et al., 2024).
Meta-analyses of the spacing effect
Every major meta-analysis finds that spaced practice beats massed practice, with effect sizes from about 0.46 to 0.78 depending on the task and setting.
839 assessments of distributed practice from 317 experiments in 184 articles went into the largest review of the spacing effect in verbal learning (Cepeda et al., 2006).
In the massed versus spaced subset of that review, only 12 of 271 comparisons showed no benefit or a negative effect from spacing. That is where the widely quoted "259 of 271" comes from (Cepeda et al., 2006).
Averaged across all retention intervals, final recall was 36.7% after massed study and 47.3% after spaced study. Study time was equal in every comparison (Cepeda et al., 2006).
For retention intervals of 8 to 30 days, recall was 32.8% massed versus 62.2% spaced, though this rests on only 6 comparisons (Cepeda et al., 2006).
63 studies and 112 effect sizes produced a mean weighted effect of d = 0.46 (95% CI 0.42 to 0.50) in favor of spaced practice (Donovan and Radosevich, 1999).
Studies with low methodological rigor reported an effect of d = 1.22, while moderate and high rigor studies both reported d = 0.40 (Donovan and Radosevich, 1999).
The effect varied sharply by task: d = 0.97 for simple, mostly physical (psychomotor) tasks, falling to d = 0.42, 0.11 and 0.07 for the three more complex task clusters (Donovan and Radosevich, 1999).
Screening more than 3,000 articles left 22 classroom reports with 31 effect sizes and more than 3,000 students. Distributed practice won with d = 0.54 (95% CI 0.31 to 0.77) (Mawson and Kang, 2025).
Of 542 records screened, 13 studies with 21,415 learners went into a medical education meta-analysis that found SMD = 0.78 (95% CI 0.56 to 0.99, p < 0.0001) for spaced repetition over standard studying (Maye and Hurley, 2026).
Across 29 studies and 39 effect sizes, spacing out retrieval practice beat massing it with g = 0.74 (Latimier, Peyre and Ramus, 2021).
Distributed practice was one of only 2 of 10 learning techniques rated "high utility" in a major review, alongside practice testing (Dunlosky et al., 2013).
Spacing effect meta-analyses compared
Meta-analysis
What was compared
Size
Effect size
Donovan and Radosevich (1999)
Spaced vs massed practice on task performance, motor to cognitive tasks
63 studies, 112 effect sizes
d = 0.46
Cepeda et al. (2006)
Spaced vs massed verbal recall, equal study time
271 comparisons, about 14,800 participants
47.3% vs 36.7% recall (effect sizes too sparse to pool)
Latimier et al. (2021)
Spaced vs massed retrieval practice
29 studies, 39 effect sizes
g = 0.74
Mawson and Kang (2025)
Distributed vs massed practice in real classrooms
22 reports, 31 effect sizes, 3,000+ students
d = 0.54
Maye and Hurley (2026)
Spaced repetition vs standard study, medical education
13 studies, 21,415 learners
SMD = 0.78
The 21,415-learner meta-analysis, checked
The "21,415 learners, SMD 0.78" figure is real, but it comes from a medical education meta-analysis by two authors in The Clinical Teacher, and it measures short-term test performance.
Who: J. A. Maye and F. Hurley of the Royal Devon and Exeter Hospital, UK, published in The Clinical Teacher, volume 23, issue 2, April 2026 (PubMed 41601436). PubMed is the database that indexes it, not the publisher.
Population: undergraduate and postgraduate medical learners only. The databases were searched in February 2025, and 14 studies made the systematic review, 13 of which were pooled.
What counted as spaced repetition: faculty-made or third-party flashcards, multiple-choice questions sent by email or through continuing medical education, and spaced classroom quizzes. It is not a study of Anki alone.
Result:SMD = 0.78, 95% CI 0.56 to 0.99, on objective tests. The authors say more work is needed on longer-term performance.
Optimal spacing gaps
The best gap between study sessions gets longer the longer you need to remember something, but it shrinks as a share of that time (Cepeda et al., 2008).
1,354 participants learned 32 obscure facts, reviewed them after a gap of up to 3.5 months, and took a final test up to 350 days later, across 26 gap and delay combinations (Cepeda et al., 2008).
For the same total study time, the optimal gap produced 64% more recall (d = 1.1) and 26% better recognition (d = 1.5) than reviewing with no gap.
For test delays of 7, 35, 70 and 350 days, the estimated best gaps for recall were about 3, 8, 12 and 27 days. That is 43%, 23%, 17% and 8% of the test delay.
Compared with no gap, the best gap improved recall by 10%, 59%, 111% and 77% at those four test delays.
The authors summarize the pattern as an optimal gap of about 20 to 40% of a 1-week test delay falling to about 5 to 10% of a 1-year delay.
In a 9-year study of 4 participants learning 300 foreign-language word pairs, 13 sessions spaced 56 days apart gave retention comparable to 26 sessions spaced 14 days apart (Bahrick et al., 1993). With only 4 learners, treat it as a proof of concept.
In 216 college students learning a math procedure, spreading 10 practice problems over two sessions made no difference on a test 1 week later but an "extremely large" difference 4 weeks later (Rohrer and Taylor, 2006).
The best review gap grows in days as the test gets further away, but falls from 43% to 8% of the delay.
Cramming can match spacing on a test the next day, but spacing wins clearly once the test is weeks away.
Of the 271 massed versus spaced comparisons in Cepeda et al. (2006), 229 used retention intervals under 10 minutes and only 7 used intervals longer than a week. The "259 of 271" figure is mostly short-term lab evidence (Cepeda et al., 2006).
Even at retention intervals under 1 minute, spaced presentations improved final-test performance by 9 percentage points over massed presentations (Cepeda et al., 2006).
Spacing was also more effective than cramming, defined as massing study on the last day before the test, in three flashcard experiments with GRE-type word pairs (Kornell, 2009).
Studying one large stack of flashcards beat studying four smaller stacks separately, because the large stack spaces each card further apart (Kornell, 2009).
Spaced practice improved retention (d = 0.51) about as much as it improved performance during acquisition (d = 0.45); the difference was not significant (Donovan and Radosevich, 1999).
Most learners think massed study works better, even right after spacing has produced better results for them.
Spacing was more effective than massing for 90% of participants, yet after the first session 72% believed massing had been more effective (Kornell, 2009).
When people learned painting styles from multiple artists, spacing the examples improved later identification of new paintings, but participants still rated massing as more effective, even after their own test scores showed the opposite (Kornell and Bjork, 2008).
Rereading and highlighting are what most students report using, and both were rated low utility; distributed practice and practice testing were the only techniques rated high (Dunlosky et al., 2013).
Algorithms: SM-2 vs FSRS benchmark numbers
On the open-spaced-repetition srs-benchmark, FSRS predicts whether an Anki user will remember a card far more accurately than SM-2, the algorithm behind Anki's long-standing default scheduler.
The benchmark dataset holds review logs from about 10,000 Anki users and about 727 million reviews (srs-benchmark).
Excluding same-day reviews, 349,923,850 reviews from 9,999 collections are used for evaluation, with older reviews used for training and newer ones for testing.
In the last published table that included SM-2 (August 2025), FSRS-6 scored an average log loss of 0.3456, against 0.616 for Anki's default SM-2 and 0.722 for original SM-2. Lower is better (srs-benchmark, August 2025 README).
On the same table, FSRS-6 with recency weighting had a lower log loss than Anki's default SM-2 for 99.5% of collections. The maintainers note that SM-2 was not designed to output probabilities, so extra formulas were added to test it.
The current table no longer lists SM-2. The newest version, FSRS-7 with recency weighting, scores a log loss of 0.3370, and the best entry, the neural network RWKV-Instant, scores 0.2773 using many more input features (srs-benchmark).
FSRS is built into Anki from version 23.10, and its default desired retention is 90%. The Anki manual warns that workload rises very quickly above 90% and recommends staying below 97% (Anki manual).
Duolingo's half-life regression model cut prediction error by 45% or more compared with baselines, using 12.9 million student-word practice records (Settles and Meeder, 2016).
An adaptive spaced education system that stopped repeating mastered questions let 62 medical students reach comparable test scores with fewer questions, a 38% gain in learning efficiency (Kerfoot, 2010).
srs-benchmark results, 9,999 Anki collections (unweighted, August 2025 table)
Algorithm
Log loss (lower is better)
RMSE bins (lower is better)
AUC (higher is better)
FSRS-6
0.3456
0.0653
0.7051
FSRS-6, default parameters
0.3661
0.0927
0.6956
Anki SM-2 (default)
0.616
0.1724
0.6133
SM-2 (original)
0.722
0.2031
0.6026
FSRS-6 has 44% lower log loss than Anki's default SM-2 and 52% lower than original SM-2 on real Anki review histories.
How to switch and what retention to pick is covered in our guide to the Anki FSRS algorithm.
Spaced repetition in medical education
Medical students are the heaviest users of spaced repetition software, and the evidence linking it to exam scores is mostly correlational.
68.3% of 560 US medical students from 102 schools reported using Anki in a 2020 survey (Halperin et al., 2024).
Among 72 students at one school, 31% used Anki, and each additional 1,700 unique Anki cards was associated with about one extra point on USMLE Step 1 after controlling for other academic and psychological factors (Deng, Gluckstein and Larsen, 2015). This is an association, not proof that Anki raised scores.
In a spaced education game with 1,470 urologists from 63 countries, the median physician answered 48% of questions correctly at first and retired 98% of them by the end (Kerfoot and Baker, 2012).
Widely repeated numbers that need a correction
Several spaced repetition statistics circulate online in a form the original source does not support.
"254 studies found spacing better in 259 of 271 comparisons." Close, but the paper never prints "259". Cepeda et al. (2006) report that "only 12 of 271 comparisons" showed no or a negative spacing effect. The 254 is the total of the "number of studies" column in their Table 1, and the 14,811 participants are summed across study and condition combinations. The review as a whole covered 317 experiments in 184 articles.
"A PubMed study of 21,415 learners proves Anki works." The meta-analysis is by Maye and Hurley in The Clinical Teacher. It covers medical learners only, pools 13 studies, and mixes flashcards, emailed quiz questions and classroom quizzes. It supports spaced repetition in medical education, not Anki specifically.
"You need expanding intervals." Expanding schedules showed no reliable advantage over evenly spaced ones (g = 0.034), although expanding schedules did better when items were tested more times (Latimier, Peyre and Ramus, 2021).
"Use a gap of 10 to 20% of the time until your exam." No single ratio holds. In Cepeda et al. (2008), the best ratio ranged from 43% for a 1-week test to 8% for a test nearly a year out.
"Duolingo's algorithm boosted engagement 12%." The 12% rise in daily return rate for any activity came from comparing two versions of Duolingo's own half-life regression model over two weeks with 3.3 million students, not from comparing spaced repetition with no spaced repetition (Settles and Meeder, 2016).
"FSRS cuts your reviews by X%." The srs-benchmark measures how well an algorithm predicts recall. It does not report review counts or learning gains, so it cannot be the source of a workload percentage.
Our own data: 18,217 flashcard reviews
Real-world usage data from our own app shows repeated retrieval working, but it is too noisy to isolate the spacing effect itself.
We analyzed 18,217 flashcard reviews from 667 students on StudyCards AI between September 2025 and July 2026 (We Analyzed 18,000 Flashcard Reviews).
Accuracy on the same card rose from 69.0% on the first attempt to 97.7% by the fourth (n = 15,104 down to n = 215).
Sessions on decks of 50 or more cards were completed 7.8% of the time, versus 25.0% for decks of 10 or fewer. Note the tension with Kornell (2009): bigger stacks space cards further apart, which helps memory, but they are also more likely to be abandoned.
We looked for a relationship between the length of the gap between reviews and recall and did not find a clean enough signal to publish. When students choose their own review times, the timing is confounded with everything else about how they study.
If you want spaced repetition without typing out cards, StudyCards AI turns a lecture PDF into a flashcard deck you can review on a schedule.
Methodology
We included peer-reviewed meta-analyses, controlled experiments, official software documentation and the public srs-benchmark repository. Every number was checked against the primary source (the journal abstract, full text, PubMed record or repository file) in September 2026; secondary statistics roundups were not used. Where a paper reports several variants of a result, we quote the one the authors headline and say which it is. Our own data is observational and labeled as such. For background on how spaced repetition works, see what is spaced repetition and our round-up of the latest spaced repetition research.
Sources
Bahrick, H. P., Bahrick, L. E., Bahrick, A. S., and Bahrick, P. E. (1993). Maintenance of foreign language vocabulary and the spacing effect. Psychological Science, 4(5), 316 to 321. doi:10.1111/j.1467-9280.1993.tb00571.x
Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., and Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354 to 380. PubMed 16719566
Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T., and Pashler, H. (2008). Spacing effects in learning: A temporal ridgeline of optimal retention. Psychological Science, 19(11), 1095 to 1102. PubMed 19076480
Deng, F., Gluckstein, J. A., and Larsen, D. P. (2015). Student-directed retrieval practice is a predictor of medical licensing examination performance. Perspectives on Medical Education, 4(6), 308 to 313. PubMed 26498443
Donovan, J. J., and Radosevich, D. J. (1999). A meta-analytic review of the distribution of practice effect: Now you see it, now you don't. Journal of Applied Psychology, 84(5), 795 to 805. doi:10.1037/0021-9010.84.5.795
Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., and Willingham, D. T. (2013). Improving students' learning with effective learning techniques. Psychological Science in the Public Interest, 14(1), 4 to 58. PubMed 26173288
Halperin, S. J., Zhu, J. R., Francis, J. S., and Grauer, J. N. (2024). Are medical school curricula adapting with their students? A survey on how medical students study and how it relates to USMLE Step 1 scores. Journal of Medical Education and Curricular Development, 11. PubMed 38268729
Kerfoot, B. P. (2010). Adaptive spaced education improves learning efficiency: A randomized controlled trial. Journal of Urology, 183(2), 678 to 681. PubMed 20022032
Kerfoot, B. P., and Baker, H. (2012). An online spaced-education game for global continuing medical education: A randomized trial. Annals of Surgery, 256(1), 33 to 38. PubMed 22664558
Kornell, N. (2009). Optimising learning using flashcards: Spacing is more effective than cramming. Applied Cognitive Psychology, 23(9), 1297 to 1317. doi:10.1002/acp.1537
Kornell, N., and Bjork, R. A. (2008). Learning concepts and categories: Is spacing the "enemy of induction"? Psychological Science, 19(6), 585 to 592. doi:10.1111/j.1467-9280.2008.02127.x
Latimier, A., Peyre, H., and Ramus, F. (2021). A meta-analytic review of the benefit of spacing out retrieval practice episodes on retention. Educational Psychology Review, 33, 959 to 987. doi:10.1007/s10648-020-09572-8
Mawson, R. D., and Kang, S. H. K. (2025). The distributed practice effect on classroom learning: A meta-analytic review of applied research. Behavioral Sciences, 15(6), 771. PubMed 40564553
Maye, J. A., and Hurley, F. (2026). The effectiveness of spaced repetition in medical education: A systematic review and meta-analysis. The Clinical Teacher, 23(2), e70353. PubMed 41601436
Rohrer, D., and Taylor, K. (2006). The effects of overlearning and distributed practise on the retention of mathematics knowledge. Applied Cognitive Psychology, 20(9), 1209 to 1224. doi:10.1002/acp.1266
Settles, B., and Meeder, B. (2016). A trainable spaced repetition model for language learning. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, 1848 to 1858. ACL Anthology P16-1174
You are welcome to use these statistics in your own writing. Please credit StudyCards AI with a link to this page.
APA
Groves, C. (2026). Spaced Repetition Statistics (2026): 40+ Research Findings. StudyCards AI. https://studycardsai.com/blog/spaced-repetition-statistics
HTML link
<a href="https://studycardsai.com/blog/spaced-repetition-statistics">Spaced Repetition Statistics (2026): 40+ Research Findings</a>
Frequently Asked Questions
What did the 21,415 learner spaced repetition meta-analysis find?
Maye and Hurley (The Clinical Teacher, 2026) pooled 13 medical education studies covering 21,415 learners and found spaced repetition beat standard study methods on objective tests, with a standardized mean difference of 0.78 (95% CI 0.56 to 0.99). The interventions included flashcards, emailed multiple-choice questions and spaced quizzes, so it is not an Anki-only result.
How much better is spaced repetition than cramming?
In Cepeda et al. (2006), spaced study beat massed study in 259 of 271 comparisons with equal study time, and average final recall rose from 36.7% to 47.3%. Kornell (2009) found spacing beat cramming-style massing for 90% of participants learning vocabulary with flashcards, even though 72% believed massing had worked better.
What is the effect size of the spacing effect?
It depends on the setting. Donovan and Radosevich (1999) found d = 0.46 across 63 studies. A 2025 meta-analysis of classroom studies by Mawson and Kang found d = 0.54. Latimier et al. (2021) found g = 0.74 for spaced versus massed retrieval practice, and a 2026 medical education meta-analysis reported SMD = 0.78.
What is the optimal spacing interval for studying?
The best gap grows with how long you need to remember, but shrinks as a share of that time. In Cepeda et al. (2008), with 1,354 learners, the best review gap for recall was about 43% of a 7-day test delay but only about 8% of a 350-day delay, roughly 27 days for a test a year away.
Is FSRS better than SM-2 at scheduling reviews?
At predicting recall, yes. In the open-spaced-repetition srs-benchmark (9,999 Anki collections, about 350 million reviews), FSRS-6 had an average log loss of 0.3456 versus 0.616 for Anki's default SM-2. The benchmark measures prediction accuracy, not how much you learn or how many reviews you do.
What does spaced repetition research from 2025 and 2026 show?
Two recent meta-analyses agree with decades of lab work. Mawson and Kang (2025) found d = 0.54 for distributed practice across 22 classroom reports with more than 3,000 students. Maye and Hurley (2026) found SMD = 0.78 across 21,415 medical learners. Both call for more research on long-term retention.