How confident should we be?

The intuition is tidy: space your repetitions close together at first, then push them gradually further apart as memory consolidates. Spaced-repetition software — Anki is the best-known example — builds this assumption into its algorithm. The model has a name, expanding retrieval practice, and it sounds like applied neuroscience.

An adult alone at a desk mid-recall with the book shut
FIG. 2The second session, days later, on material that has been allowed to fade a little.

The actual evidence is surprisingly modest. Studies comparing expanding schedules against equal-interval spacing — gaps that stay the same length — have repeatedly struggled to show a reliable advantage for expanding gaps. Careful reviews of the research have found that when expanding and equal schedules are matched for total study time and number of repetitions, the difference in final recall is small and inconsistent. Some experiments favour expanding; others favour equal spacing; many find no difference worth reporting.

Part of the problem is confounding. Early spaced-repetition research often compared expanding schedules against massed practice — cramming — where almost anything spaced wins. The spacing effect itself is robust; the expanding variant of it is not.

Disentangling “expanding gaps” from “more early practice” requires tight experimental control that many published studies lack.

There is also a measurement complication. Expanding schedules front-load practice, so items get more repetitions early. Equal schedules stretch the same repetitions more evenly.

A cluttered desk with a handwritten notebook, glasses, mouse and lamp illuminating the scene

One mechanism often cited for the advantage is retrieval difficulty: an expanding gap should hit each test at exactly the moment memory is fading, maximising the effort needed to reconstruct it and therefore maximising consolidation. The logic is coherent, but the empirical payoff is elusive. Difficulty at retrieval does matter — that finding is solid — yet it does not follow that expanding gaps are reliably better at calibrating that difficulty than equal gaps are. Individual material, individual forgetting rates and the length of the final retention interval all interact in ways no single algorithm captures cleanly.

What the software actually implements is often a hybrid: expanding gaps as a default, adjusted by how easily each card was answered. That responsiveness is probably more important than the expanding structure itself. A system that detects failure and shortens the next interval is doing something genuinely useful; a system that mechanically widens gaps on a fixed schedule is running ahead of the evidence.

The honest summary: distributed practice beats massing, consistently and with a large effect. Whether those distributed sessions should use growing or fixed gaps remains an open question that review software has quietly answered by implementation rather than proof.