NextPass
The evidence

Content is necessary. It just isn’t sufficient.

Content builds understanding. The bigger gains come when information is turned into repeated practice with clear, criterion-referenced feedback, spaced over time. That sentence is the whole pitch — and this page is the research it stands on, with the confidence levels showing.
01How to read this page

Confidence levels, showing

Every claim is tagged with what stands behind it. Replicated meta-analyses carry the weight. Findings from medicine are labeled medicine. Industry research is labeled industry. Early AI studies are labeled early, with sample sizes.

Plain language first, statistics second

Learning research reports gains as effect sizes — statistical units that mean little outside academia. So every claim here is stated in plain terms: how far a method moves an average learner. The studies' own published figures are preserved beneath each claim, footnote-style, for accuracy.

One honest gap, up front

No single study pits the full consume → practice → feedback → re-practice loop against content alone. The case is component-wise: every component has large, replicated, independent effects that stack in the same direction. The missing study is at the bottom of this page.

No zombie stats

You won't find the Learning Pyramid (“people remember 10% of what they read…”), “10,000 hours,” or Ebbinghaus percentages here. They're debunked or unverifiable — and a page built on them wouldn't deserve your trust.

02The failure mode

Programs don’t fail in the room. They fail in the weeks after it.

PEER-REVIEWED SURVEY · 150 ORGANIZATIONS

Application decays: ~62% right after training, ~44% at six months, ~34% a year out.

The most citable decay numbers available — a peer-reviewed survey of training professionals across 150 organizations. It isn't an RCT, so we phrase it the way it deserves: this is what the people who run training estimate happens to it.

Saks & Belcourt, 2006

INDUSTRY RESEARCH · LABELED, ALWAYS

Industry estimates put “scrap learning” — training delivered but never applied — at 45–85% of what's bought.

CEB (now Gartner) put it at 45%. Brinkerhoff's field research puts sustained non-application near 80–85%: roughly 20% never try, and another ~65% try and then revert. Against ATD's ~$1,254 average annual training spend per employee, even the conservative end is real money buying nothing.

CEB 2014 · Brinkerhoff · ATD 2025 State of the Industry

PEER-REVIEWED · LANDMARK STUDY

Passive learners feel trained — right up until the delayed test.

In the classic comparison, re-studiers actually beat practicers on the immediate five-minute check, then lost badly a week later: 40% retained versus 61%. An end-of-module quiz is measuring the five-minute version of a learner.

Roediger & Karpicke, 2006 · Psychological Science

IN THE LOOP  NextPass optimizes for the delayed test — the real conversation, weeks later — not the completion quiz.

03Component 1 — retrieval

Producing the skill beats reviewing it.

A voice pass is retrieval practice in its purest form: the learner must produce the skill live, out loud — not recognize it on a quiz.

THREE CONVERGING META-ANALYSES

Across hundreds of studies, retrieval practice beats re-study — and the advantage grows after a day.

The edge is enough to move an average learner from the middle of the pack into roughly the top 30% — and it gets bigger once more than a day has passed, the interval that matters for real conversations. Corroborated by a 272-study meta-analysis and a classroom-only one, and retrieval improves transfer to related tasks, not just recall.

Rowland, 2014 · Adesope et al., 2017 · Yang et al., 2021Published figures: g ≈ 0.50 overall; ≈ 0.69 beyond one day

PEER-REVIEWED · SCIENCE

Once material is “learned,” more re-exposure adds almost nothing. Continued retrieval holds ~80% vs ~36% a week later.

After the first successful recall, continued testing — not continued re-study — is what made memory durable. “We'll just have them re-watch the module” is the intervention this result retires.

Karpicke & Roediger, 2008 · Science

10-TECHNIQUE UTILITY REVIEW

The two most common study behaviors are the two lowest-utility techniques known.

Re-reading and highlighting rate LOW utility in the definitive ten-technique review — and they're exactly what content-only delivery produces. The only two techniques rated HIGH: practice testing and spaced practice. A practice loop adds both.

Dunlosky et al., 2013 · Psychological Science in the Public Interest

04Component 2 — feedback

Feedback lifts performance — when it's built right.

META-ANALYSES · 131 + 435 STUDIES

On average, feedback improves performance. And more than a third of feedback interventions backfire anyway.

Across 607 comparisons from 131 studies, the average lift is solid — an average performer climbing from the middle of the pack into roughly the top 35%. But over a third of the time, feedback made performance worse. What separates helping from harming: task-focused, actionable feedback measured against clear criteria, versus vague or person-focused feedback. A newer 435-study meta-analysis lands in the same place.

Kluger & DeNisi, 1996 · Wisniewski, Zierer & Hattie, 2020Published figures: d = 0.41 (1996); d = 0.48 (2020)

IN THE LOOP  This is the argument for rubric-scored debriefs. NextPass feedback is structurally the kind that helps: scored against your rubric, tied to the learner's own words, never about the person. Ad-hoc human feedback varies on exactly the dimension that decides the outcome.

FEEDBACK TYPOLOGY · QUALITATIVE

Task and process feedback carry the value. Praise carries roughly none.

The canonical typology of feedback finds information about the task and how to improve it high-value, and feedback about the person close to worthless for learning. Rubrics force feedback into the high-value lane.

Hattie & Timperley, 2007

05Component 3 — spacing

Spaced practice beats cramming.

META-ANALYSES

Spacing practice over weeks beats cramming it — enough to move an average learner from the middle of the pack into roughly the top 30%.

One of the oldest, most replicated results in learning science, in skill tasks and in real classrooms alike — and the effects grow at longer retention intervals, which is where working life happens.

Donovan & Radosevich, 1999 · 2025 classroom meta-analysisPublished figures: d = 0.46 (skill tasks); d = 0.54 (classrooms)

META-ANALYSIS

When the thing being spaced is retrieval itself, the advantage grows further still.

Practice sessions distributed over time beat the same sessions massed together — a jump to roughly the top 25% — and the gap widens as the retention interval grows.

Latimier, Peyre & Ramus, 2021Published figure: g ≈ 0.74

IN THE LOOP  The one-day workshop is massed practice by design. A practicum runs the other way: short scored passes, spaced across the weeks of a cohort.

06The assembled loop

The full loop has a name in the literature — and forty years of results.

Model the skill, rehearse it, get structured feedback: the literature calls this Behavior Modeling Training. NextPass is that loop, delivered as software — your curriculum models the skill, the role-play is the rehearsal, the debrief is the feedback.

META-ANALYSIS · 117 STUDIES

Modeling + rehearsal + feedback produces skill gains about as large as training research ever measures — and its behavior effects held or grew over time while knowledge faded.

117 studies, mostly in real workplaces. The skill gains put an average learner roughly in the top 15%. On-the-job behavior change was more modest — and it was the effect that lasted. That reframes the KPI: knowledge scores decay; rehearsed behavior is the part that sticks.

Taylor, Russ-Eft & Chan, 2005 · Journal of Applied PsychologyPublished figures: knowledge/skill d ≈ 1.0; behavior d ≈ 0.25–0.4, durable

META-ANALYSIS · 117 STUDIES

The same meta-analysis names what maximizes transfer — and it reads like a product spec.

Practicing learner-relevant scenarios. Seeing both good and bad models. Setting goals. Involving supervisors. Scenario variety across realistic characters, briefings that frame what good and bad sound like, and rubric goals per session are that list, shipped.

Taylor et al., 2005 · transfer moderators

META-ANALYSIS · 335 STUDIES

In leadership training — the same skill family as management conversations — practice, feedback, and spacing are what drive transfer.

Across 335 studies, Lacerenza et al. report effect sizes translating to approximately +25% learning, +28% on-the-job behavior change, and +20% job performance for well-designed programs. Information + demonstration + practice beats any subset; if you can only pick one method, practice beats information delivery; spaced multi-session scheduling improves transfer even when knowledge scores don't move.

Lacerenza et al., 2017 · Journal of Applied Psychology

07Mastery, not seat time

The strongest simulation evidence is from medicine. We label it that way.

Medicine stopped trusting lectures first, so it has the most rigorous simulation research. We use it as the existence proof for the mechanism — simulated practice changing real behavior — not as a claim about our domain.

META-ANALYSIS · 609 STUDIES · MEDICINE

Against no practice, simulation moved everything: knowledge, skills, on-the-job behavior — and real patient outcomes.

609 studies and 35,226 trainees across the health professions. The knowledge and skill gains are among the largest in training research — enough to put an average trainee roughly in the top 12% of the unpracticed; smaller but real effects carried all the way through to how trainees behaved on the job and how their patients fared. Distributed, multi-day practice was associated with larger effects.

Cook et al., 2011 · JAMAPublished figures: knowledge 1.20; skills ≈ 1.1; behavior ≈ 0.8; patient outcomes 0.50

META-ANALYSIS · 82 STUDIES · MEDICINE

Practicing to a rubric-defined standard beats practicing for a fixed amount of time — by one of the largest margins in the simulation literature.

Mastery learning — advance when the rubric says you're ready, not when the seat-time counter runs out — beat time-bound simulation across 82 studies, the average mastery-trained learner landing roughly in the top 12% of the time-trained group. The honest trade-off, reported in the same analysis: mastery takes more learner time. Completion ≠ competence, in one finding.

Cook et al., 2013 · Academic MedicinePublished figure: 1.17 vs non-mastery simulation

IN THE LOOP  It's why progression in a practicum is score-gated, not seat-timed — and why the readout reports mastery, not completion.

SINGLE STUDY · SMALL N · MEDICINE

Practice-to-criterion research suggests structured, spaced practice can roughly halve the practice attempts needed to reach a defined standard.

In a proficiency-based simulator study, a structured spaced schedule reached competence in 16 repetitions versus 32. Small sample, procedural domain — we cite it as directional, not as proof.

Proficiency-based surgical simulator study

08The AI-specific evidence

The AI role-play research is early. We say so, every time.

The mature findings above carry the weight; these early studies are consistent with the mechanism. Sample sizes are shown because that's what honest looks like.

RCT · N = 86 · EARLY

Adding rubric-style feedback to AI role-play improved skill mastery by 17% and self-efficacy by 27% over AI role-play alone — and mastery was the outcome that transferred to new situations.

The most on-point study in existence: AI conversation practice plus rubric-grounded feedback, for interpersonal skills. The domain-grounded feedback also measured 28% closer to expert feedback than a vanilla general model's — the case for encoding a curriculum rather than wrapping a chatbot.

Lin et al., 2024 (IMBUE) · ACL

PILOT RCT · N = 19 · EARLY

In a pilot RCT, AI role-play matched or beat peer role-play on several scored scenarios — and learners rated it higher on autonomy, repeatability, and structured feedback.

Peers won on one thing: authenticity. That clause is why NextPass is built to sit under live facilitation and real conversations, not replace them.

Lee et al., 2025 · OSCE preparation

QUASI-EXPERIMENT · N = 53 · EARLY

GenAI role-play produced speaking gains statistically similar to human-partner practice — with higher intrinsic motivation and self-efficacy.

A practice partner for every learner, at scale, with nothing measurable lost and motivation gained.

Cheng et al., 2024

MIXED-METHODS · ADJACENT DOMAIN

AI conversation practice measurably reduced speaking anxiety — and the most anxious learners improved the most.

Adjacent domain (language speaking), converging with counselor-trainee reports of practicing without the “human gaze.” The learners who avoid role-play in front of colleagues are the ones private practice finally reaches.

Ding et al., 2025 · converging studies

09What the evidence doesn't show — yet

The missing study, named. And who's positioned to run it.

No peer-reviewed RCT yet compares a full AI-voice-role-play-plus-rubric-feedback loop against content-only delivery for management conversation skills in organizations. The components are proven; this exact assembly, in this exact domain, is young. That's why every AI-specific claim on this site is phrased as early research — the mature findings above do the carrying.

One more thing the literature insists on: transfer isn’t a content problem — or a practice-only problem. For open skills like conversation, the work environment (manager support, transfer climate) moves outcomes far more than it does for technical skills (Blume et al., 2010 · 89-study meta-analysis). Practice loops improve the odds; they don’t replace a manager who reinforces the behavior. We’d rather tell you that now than have your L&D team tell us later.

The missing evidence is generatable — and the platform captures what researchers struggle to get: voluntary pass counts, per-dimension rubric-score trajectories across sessions, pre/post self-efficacy. As cohorts accumulate, we intend to publish what we find — including the parts that don’t flatter us.

Another pass, whenever they need one

The research describes a loop. Come watch one run.

A 30-minute walkthrough of a live practicum — real scenarios, real rubric scoring, the real readout — and what your program would look like on the platform.