The landscape
In 2015, the Open Science Collaboration published its attempt to reproduce 100 psychology experiments in Science. Roughly 60 percent failed to replicate at the original effect size or significance level — a result that landed hard across the discipline. Cognitive psychology fared noticeably better than social psychology, which is relevant here: most of the core learning-science findings rest on cognitive rather than social mechanisms, and many were already supported by decades of independent replications before the crisis sharpened anyone’s attention.
The spacing effect is perhaps the most robust result in the entire field. Hermann Ebbinghaus documented it in the 1880s in Germany, working on himself with nonsense syllables, and independent laboratories have since confirmed distributed practice across languages, age groups, materials and retention intervals. Effect sizes are consistently large. This one held.

The testing effect — the finding that retrieving information strengthens memory more than an equivalent period of re-study — has also replicated well. Henry Roediger and Jeffrey Karpicke’s influential 2006 experiment at Washington University in St. Louis showed large advantages for retrieval practice over re-reading on a final test one week later. Multiple independent labs have reproduced the core finding. Some boundary conditions matter — the advantage shrinks when feedback is absent and material is very difficult — but the basic effect is real and substantial.
Laboratory studies, particularly those involving mathematics and motor learning, consistently show better retention and transfer from interleaved practice. But the effect size varies considerably depending on the domain and the difficulty of the material, and some classroom-based studies have struggled to show the same advantage. The effect is real; how far it travels beyond the laboratory remains an open question.
What shrank or collapsed
Cognitive load theory, developed by John Sweller, rests on a solid foundation in Alan Baddeley’s working memory research and has generated useful, replicable findings — notably the split-attention and expertise-reversal effects. But several specific predictions derived from the theory have proven harder to replicate cleanly, and the theory itself has been criticised for being difficult to falsify in a strict sense. The core intuitions survive; the precise mechanisms are contested.

The more serious casualties sit mostly outside the core memory-and-retrieval cluster. Learning styles — the idea that individuals have preferred modalities (visual, auditory, kinaesthetic) that should be matched to instruction — never had strong experimental support, and what evidence existed has not survived scrutiny. It is not a contested finding; it is effectively a failed one.
Key moment
- 2015Open Science Collaboration publishes results of 100 replication attempts in Science; ~60 percent fail to replicate at original effect size
Judgements of learning are reliably miscalibrated — people consistently overestimate how well they have learned, particularly after fluent re-reading — and this finding replicates robustly. The bias is real.

What this means for the field
The pattern that emerges is roughly this: findings built on controlled laboratory paradigms, replicated across multiple independent groups over decades, and tied to plausible cognitive mechanisms have generally survived. The effects documented by researchers at Washington University in St. Louis, Purdue University and the University of California, Los Angeles — retrieval practice, spacing, interleaving — constitute a more defensible evidence base than the broader pop-psychology layer that grew around them.
Effect size matters as much as replication status. A finding can replicate reliably and still be too small to be practically significant. Conversely, an effect that replicates across dozens of studies with a large effect size — as spacing does — represents something genuinely useful to know. The replication crisis did not destroy cognitive learning science; it clarified which parts of it deserve confidence and which parts were always softer than their advocates implied.
