What effect size measures — and why significance alone misleads

When a study finds that retrieval practice beats re-reading, it can report that result as statistically significant — meaning the difference is unlikely to be noise — without telling you how large that difference actually is. Effect size fills that gap. The most common metric, Cohen’s d, expresses the gap between two groups in units of standard deviation. A d of 0.2 is conventionally small, 0.5 medium, 0.8 large. But these thresholds are rough guides, not laws; what counts as meaningful depends entirely on what you are measuring and at what cost.

Replicates broadly4 Robust, bounded8 Method note4 Mixed3 Shape replicates1 Contested1
How the 21 entries on this site are marked. The classification is this site’s own, applied from the strength of evidence each entry describes — not a published index.

The testing effect, in which one retrieval attempt outperforms an equivalent period of re-reading, is among the more robust findings in memory research. Henry Roediger and Jeffrey Karpicke at Washington University in St. Louis have reported effect sizes in the medium-to-large range under lab conditions, and the result has replicated reasonably well.

A quiet library carrel seen from behind
FIG. 2Most of the estimates on this site were produced in rooms like this, on adults working alone.

The spacing effect — spreading study across time rather than massing it — has a similarly long pedigree, traceable to Hermann Ebbinghaus’s self-experiments in the 1880s, and carries effect sizes that tend to be moderate but consistent.

Far transfer, the kind that would justify grand claims about learning to learn, produces effect sizes that are typically small and frequently indistinguishable from zero.

Interleaving is a different story. The core finding — that mixing problem types during practice improves later test performance compared to blocking — is real, but the effect sizes reported in many studies are modest, the boundary conditions are contested, and some of the clearest demonstrations involve relatively simple material over short intervals.

Key numbers

Cohen’s d = 0.2small effect (conventional threshold)
Cohen’s d = 0.5medium effect
Cohen’s d = 0.8large effect
  • These thresholds come from Jacob Cohen’s 1988 work; they are widely used and widely misapplied.

Cognitive load theory, developed by John Sweller, rests on a well-supported account of working memory’s limits, but individual predictions from the theory vary considerably in how well they replicate and how large their effects are outside the lab.

Effect size also interacts with study design in ways that inflate optimism. A finding that sounds transformative in a press release may carry a d of 0.15 on the outcome anyone actually cares about.

Cluttered writer's desk with open notebooks, magnifying glass, pens and lamps casting warm light

None of this discredits the field. A small effect that is genuine, cheap to apply and cumulative over time can still matter. What it discredits is the habit of treating statistical significance as the end of the conversation. The number behind the asterisk — and what that number means in context — is where the real question starts.