What effect size measures — and why significance alone misleads
When a study finds that retrieval practice beats re-reading, it can report that result as statistically significant — meaning the difference is unlikely to be noise — without telling you how large that difference actually is. Effect size fills that gap. The most common metric, Cohen’s d, expresses the gap between two groups in units of standard deviation. A d of 0.2 is conventionally small, 0.5 medium, 0.8 large. But these thresholds are rough guides, not laws; what counts as meaningful depends entirely on what you are measuring and at what cost.
The testing effect, in which one retrieval attempt outperforms an equivalent period of re-reading, is among the more robust findings in memory research. Henry Roediger and Jeffrey Karpicke at Washington University in St. Louis have reported effect sizes in the medium-to-large range under lab conditions, and the result has replicated reasonably well.

The spacing effect — spreading study across time rather than massing it — has a similarly long pedigree, traceable to Hermann Ebbinghaus’s self-experiments in the 1880s, and carries effect sizes that tend to be moderate but consistent.
Far transfer, the kind that would justify grand claims about learning to learn, produces effect sizes that are typically small and frequently indistinguishable from zero.
Interleaving is a different story. The core finding — that mixing problem types during practice improves later test performance compared to blocking — is real, but the effect sizes reported in many studies are modest, the boundary conditions are contested, and some of the clearest demonstrations involve relatively simple material over short intervals.
Key numbers
- These thresholds come from Jacob Cohen’s 1988 work; they are widely used and widely misapplied.
Cognitive load theory, developed by John Sweller, rests on a well-supported account of working memory’s limits, but individual predictions from the theory vary considerably in how well they replicate and how large their effects are outside the lab.
Effect size also interacts with study design in ways that inflate optimism. A finding that sounds transformative in a press release may carry a d of 0.15 on the outcome anyone actually cares about.

None of this discredits the field. A small effect that is genuine, cheap to apply and cumulative over time can still matter. What it discredits is the habit of treating statistical significance as the end of the conversation. The number behind the asterisk — and what that number means in context — is where the real question starts.
