By ·

Does Flashcard Length Matter? We Tested It On 18,000 Reviews

"Keep your flashcards short" is one of the most repeated pieces of study advice on the internet. It shows up in every Anki guide, every med school study blog, every Reddit thread about active recall. Our own guide to making good flashcards says it too. It's usually stated as fact and almost never with data behind it.

We had 18,168 real flashcard reviews sitting in our database, each one linked to the exact text of the card being answered. So we checked. The short version: the advice looks strongly supported until you run one additional test, at which point most of the effect disappears. Both halves of that are worth showing.

Key Takeaways

Why anyone thinks length matters in the first place

The theoretical case for short cards comes from cognitive load theory. The distinction that matters here is between intrinsic load, the inherent difficulty of the material, and extraneous load, which a 2025 review in Health Education & Behavior defines as the part of the process that does not facilitate learning and is imposed by suboptimal instructional techniques or information presentation. The same authors put the consequence plainly: the more working memory a learner spends on extraneous load, the fewer resources remain for the material itself.

A wordy flashcard question is, in theory, pure extraneous load. The concept being tested hasn't changed, but the student has to parse more text before recall can start. That's the mechanism the advice assumes, and it's covered in more depth in our pieces on cognitive load theory and working memory limits. It's a reasonable prediction. The question is how much it shows up in practice.

Where this data comes from

StudyCards AI turns an uploaded PDF into a deck of flashcards. Every graded review during a study session logs whether the answer was correct, how long it took, and how the student rated the card's difficulty afterward. Because we store the card text alongside the review, we can measure the question a student actually saw against how they performed on it.

This analysis covers 18,168 reviews from 676 students across 15,083 unique flashcards and 727 decks, logged between September 21, 2025 and July 31, 2026. Following our earlier analysis of the same review log, we excluded reviews taking longer than 10 minutes, which are almost always someone leaving a tab open rather than genuinely deliberating. We checked that no single account dominates the totals.

Methodology note: this is observational usage data from one flashcard app, not a controlled experiment. Nobody was randomly assigned short or long cards. That distinction turns out to be the entire story of this post, so we've shown the analysis that undercuts our own headline rather than stopping at the result that looked cleanest.

The headline result: shorter questions get answered correctly more often

Bucketing every review by the character length of its question produces a clear downward slope. Questions of 20-39 characters are answered correctly 75.4% of the time (n=2,665). At 40-59 characters it's 75.1% (n=5,689). Then it starts falling: 70.6% at 60-79 characters (n=5,532), 65.9% at 80-99 (n=2,614), and 63.2% at 100-119 (n=772).

That's a 12-point spread across the range holding roughly 95% of all reviews, and it's monotonic through the buckets with meaningful sample sizes. Above 120 characters the numbers get erratic and the sample sizes collapse into the low hundreds and then into the dozens, so we've left them off the chart rather than pretend a bucket of 25 reviews tells us anything.

Accuracy declines as question length grows Questions of 2-19 characters: 71.9% correct, n=572. 20-39 characters: 75.4% correct, n=2,665. 40-59 characters: 75.1% correct, n=5,689. 60-79 characters: 70.6% correct, n=5,532. 80-99 characters: 65.9% correct, n=2,614. 100-119 characters: 63.2% correct, n=772. Data: StudyCards AI internal usage, 2025-2026. 0% 25% 50% 75% 100% 71.9% 2-19 n=572 75.4% 20-39 n=2,665 75.1% 40-59 n=5,689 70.6% 60-79 n=5,532 65.9% 80-99 n=2,614 63.2% 100-119 n=772 Question length (characters) Source: StudyCards AI internal usage data (2025-2026), 17,844 reviews shown
Accuracy peaks in the 20-59 character range and declines steadily as questions get longer.
Share on X

The obvious objection: maybe long cards are just harder cards

The first thing to rule out is that question length is standing in for topic difficulty. A long question might simply be attached to a more complicated concept, in which case length isn't doing anything and we'd just be measuring hard material.

We can test that, because students rate each card easy, medium, or hard after answering it. Splitting on the student's own difficulty rating, the length gap persists inside every band. On cards rated easy, short questions score 95.3% against 91.7% for long ones. On medium, 74.7% against 68.6%. On hard, 51.9% against 41.9%. The gap is there at every level of perceived difficulty, and it's actually widest on the hardest material.

At this point the finding looked solid. It's a clean monotonic trend across large samples, and it survives the most obvious confound. This is the point where it would be easy to publish.

The test that broke it

There's a second confound, and it's less obvious. Every number above pools all students together. That means the comparison isn't only "short cards versus long cards," it's also "the students who happened to encounter short cards versus the students who happened to encounter long cards." If weaker students tend to study decks with longer questions, perhaps because of the subject they're studying or the density of the PDF they uploaded, then the length effect and the student effect are tangled together.

This is a well-documented failure mode rather than an exotic one. Kievit and colleagues, writing in Frontiers in Psychology, describe how the direction of an association at the population level may be reversed within the subgroups comprising that population, a pattern known as Simpson's paradox. Their sharpest illustration is that two variables can correlate positively across a population of individuals but negatively within each individual over time. An aggregate flashcard statistic is exactly the kind of number that can behave this way.

The way to separate them is to compare each student against themselves. We restricted the data to students who reviewed a decent number of both short cards (under 60 characters) and long ones (80 or more), then measured each person's own accuracy gap between the two.

The effect largely disappeared. Among students with at least 20 reviews of each type (n=26), mean accuracy was 66.0% on short cards and 63.6% on long ones. That's a 2.4-point gap, not 12. Loosening the threshold to 10 reviews of each type (n=68) gives 72.3% against 70.7%, a 1.6-point gap. Tightening it to 30 reviews of each (n=13) gives 69.9% against 67.4%, a 2.5-point gap. The result is stable across all three cutoffs, so this isn't an artifact of where we drew the line.

The length gap shrinks when compared within students Pooled across all reviews, short questions score 75.4% and long questions 63.2%, a 12.2 point gap. Comparing each student against themselves at three sample thresholds, the gap is 1.6 points at 10 or more reviews each, 2.4 points at 20 or more, and 2.5 points at 30 or more. Data: StudyCards AI internal usage, 2025-2026. 0 3 6 9 12 15 Accuracy gap (points) 12.2 Pooled all reviews 1.6 Within n=68 2.4 Within n=26 2.5 Within n=13 Pooled comparison vs. same-student comparison at three sample thresholds Source: StudyCards AI internal usage data (2025-2026), 18,168 reviews
The 12-point pooled gap shrinks to under 3 points once each student is compared against their own performance on both card types.

One more number makes the point sharper. Of those 26 students, only 15 scored better on short cards. The other 11 did better on long ones. If card length were exerting a strong, consistent pull on recall, we'd expect that split to be lopsided. Fifteen out of 26 is close enough to a coin flip that we can't read much into the direction for any individual student.

So the honest summary is: there is probably a small real effect, somewhere in the range of 2 to 3 percentage points, and roughly four fifths of the dramatic version of this finding was students differing from each other rather than cards differing from each other.

Answer length does much less

We ran the same bucketing on the answer side and the pattern is far weaker. Accuracy moves from 73.1% on very short answers up to 75.4% at 60-89 characters, then down to 64.9% at 180-209 characters, and then back up to 75.7% on the longest bucket. That last reversal is the tell. A real effect shouldn't bounce like that.

Our read is that answer length isn't doing much on its own, and we're not going to build a recommendation on a pattern that doesn't hold its direction. It makes intuitive sense that the question matters more than the answer here, since the question is what the student has to parse and hold in mind before recall even begins, while the answer is mostly read after the attempt is over.

Response time tells a different story than accuracy

Median response time rises with question length exactly as you'd expect, from 4.6 seconds at 20-39 characters to 5.5 seconds at 60-79 characters. Then it falls back to 4.6 seconds at 100-119 characters and 3.0 seconds beyond that.

That reversal at the long end is interesting. Combined with the accuracy drop in the same range, the most plausible reading is that past a certain length some students stop genuinely attempting the card and start skimming and guessing. Answering faster while also answering worse is not what careful reading looks like. Skipping the retrieval attempt matters because retrieval is where the learning happens, which is the whole basis of the testing effect. We can't confirm that interpretation from this data, but it's the reading most consistent with both curves.

What this means for writing flashcards

Three things this data supports, stated at the strength the evidence actually justifies:

  1. Aim for questions in the 20-60 character range, roughly 4 to 10 words, where accuracy was highest.
  2. Treat this as a mild preference, not a rule. The within-student effect is 2 to 3 points, not 12.
  3. Don't bother agonizing over answer length. We found no reliable signal there.

It's worth putting the size of this effect in perspective against the things that clearly do matter. In our earlier analysis of the same review log, accuracy on a repeated card climbed from 69% to 97.7% across four attempts, and session completion fell from 25% to 7.8% as decks grew past 50 cards. Those are large, robust effects. Two or three points of card-length difference is not in the same category. If you're deciding where to spend effort, spacing your reviews properly and keeping decks short enough to finish will do far more than rewording your questions.

The broader takeaway is about how to read study statistics generally, including ours. A pooled comparison across thousands of data points can look authoritative and still be mostly measuring who the people are rather than what was done to them. The 12-point version of this finding would have made a better headline, and it would have been wrong by roughly a factor of five.

If you're turning lecture notes or a textbook chapter into flashcards, StudyCards AI generates them automatically from a PDF upload, with questions that land in the short range by default. For writing them by hand, our guides to making good flashcards and the psychology behind effective flashcards cover the parts of card design that carry more weight than length.

Frequently Asked Questions

Does flashcard question length affect how well you remember?

Across 18,168 flashcard reviews, accuracy fell from 75.4% on questions of 20-39 characters to 63.2% on questions of 100-119 characters. However, when we compared the same student on both short and long cards, the gap shrank to roughly 2 percentage points, suggesting most of the headline difference is about which students get long cards rather than about card length itself.

How long should a flashcard question be?

In this dataset the best-performing range was roughly 20 to 59 characters, about 4 to 10 words, where accuracy sat at 75%. Accuracy declined steadily above 60 characters. The effect is modest once you account for differences between students, so treat this as a mild preference for shorter phrasing rather than a rule.

Why did the effect shrink when you controlled for the student?

The aggregate comparison pools every review together, so it mixes two things: whether long cards are harder, and whether students who encounter long cards are weaker overall. Comparing each student against themselves removes the second factor. The gap dropped from about 12 points to between 1.6 and 2.5 points, meaning most of the aggregate effect was between students, not within them.

Does a longer flashcard answer also hurt accuracy?

Much less clearly. Answer length showed a far weaker and less consistent pattern than question length, moving from 75.4% correct at 60-89 characters down to 64.9% at 180-209 characters, but then rising again at the longest lengths. We do not consider the answer-length pattern reliable enough to draw a conclusion from.

Was any personal or identifiable data used in this analysis?

No. All figures are aggregate counts and percentages computed directly against the production database. No individual answers, user names, or flashcard contents were extracted or reviewed.

Generate Anki flashcards from PDFs