Every source the science pages draw on, with its link and the label it has to carry. Effect sizes, sample sizes and study designs live here rather than in the page copy.Where a claim did not survive verification, the wording we are allowed to use is recorded next to the source.
Mar, Li, Nguyen & Ta (2021). Memory and comprehension of narrative versus expository texts: a meta-analysis. Psychonomic Bulletin & Review. Open the source Narrative advantage g = .55 across 150 effect sizes. The authors flag publication bias and warn against forcing every subject into narrative form.Green & Brock (2000). The role of transportation in the persuasiveness of public narratives. Journal of Personality and Social Psychology. Open the source Established and validated a scale for narrative transportation.
Spacing and retrieval
Kim & Webb (2022). The effects of spacing on second language learning: a meta-analysis. Language Learning. Open the source 48 experiments, N = 3,411. Medium-to-large spacing effect. Equal and expanding schedules came out statistically equivalent.Adesope, Trevisan & Sundararajan (2017). Rethinking the use of tests: a meta-analysis of practice testing. Review of Educational Research. Open the source 272 effect sizes. Practice testing beats restudying at g = 0.51 and beats no further practice at g = 0.93.Hwang (2024). Interleaving as an undesirable difficulty for low-achieving L2 learners. Language Learning. Open the source 107 low-achieving adolescent learners. Pure interleaving hurt this group, which is why new material is blocked before it is interleaved.Interleaved practice of L2 grammar tenses (2024). Four experiments on interleaved versus blocked practice of verb tenses. Learning and Instruction. Open the source d = .52 on conjugation and d = .36 on tense identification at a one-week delay.Pan et al. (2019). Systematic alternation for study trials with randomized practice in L2 grammar. Journal of Applied Research in Memory and Cognition. Open the source
Grammar instruction and feedback
Norris & Ortega (2000). Effectiveness of L2 instruction: a research synthesis and quantitative meta-analysis. Language Learning. Open the source Focused instruction produces large gains and explicit treatments outperform implicit ones. Outcome measures in this literature lean toward explicit knowledge. This is part of the evidence base we cite where Krashen's input hypothesis is described as contested.Li (2010). The effectiveness of corrective feedback in SLA: a meta-analysis. Language Learning. Open the source d = .61 to .64, maintained on delayed post-tests. The second half of the evidence base behind the Krashen wording.Lyster & Saito (2010). Oral feedback in classroom SLA: a meta-analysis. Studies in Second Language Acquisition. Open the source 15 classroom studies, N = 827. Significant and durable effects, and prompts outperformed recasts.Xu & Zeng (2023). The timing of corrective feedback in second language learning. Frontiers in Psychology. Open the source A systematic review, not primary research. Immediate feedback was better or equal in 85 percent of 20 comparisons, moderated by explicitness.Surface-level versus deep-level writing feedback (2024). A meta-analysis of written corrective feedback across 200 comparisons. Learning and Instruction. Open the source Form-only feedback improved surface accuracy at g = 0.58 and cost deep-level writing outcomes at g = -0.23.Shute (2008). Focus on formative feedback. Review of Educational Research. Open the source Formative feedback works when it is timely, specific and small enough to act on. A review of the literature, so it constrains how we write a correction rather than proving ours lands.
Reading and level control
Nakanishi (2015). A meta-analysis of extensive reading research. TESOL Quarterly. Open the source d = 0.46 against control conditions and d = 0.71 pre-post.Day & Bamford (2002). Top ten principles for teaching extensive reading. Reading in a Foreign Language. Open the source The easy-material principle: one or two unknown words per page for beginners.Lexical coverage replication (2025). Comprehension across 90, 95, 98 and 100 percent lexical coverage. Reading in a Foreign Language. Open the source The 98 percent threshold did not replicate. Only full coverage was clearly ahead, at d = 1.3 and d = 1.4 over 90 and 95 percent. Treat coverage as a matching heuristic.Malik, Mayhew, Piech & Bicknell (2024). From Tarzan to Tolkien: controlling the language proficiency level of LLMs for content generation. Findings of ACL 2024. Open the source Company-affiliated peer-reviewed research: two of the authors are Duolingo staff. Level-control error fell from 0.57 to 0.30 with per-level exemplars, and to 0.15 with best-of-3 plus a scorer.Imperial & Tayyar Madabushi (2023). Flesch or Fumble? Evaluating readability standard alignment of instruction-tuned language models. GEM Workshop at EMNLP 2023. Open the source Naive prompting drifted about one CEFR level high. A2 story completions hit the target 0 to 13 percent of the time. A dated baseline on 2023-era models.
Diagnosis and learner models
Alderson & Huhta (2005). The development of a suite of computer-based diagnostic tests based on the Common European Framework. Language Testing. Open the source DIALANG: the first major CEFR-based diagnostic system. Online, fourteen languages, built for diagnosis and feedback rather than certification.O'Keeffe & Mark (2017). The English Grammar Profile of learner competence. International Journal of Corpus Linguistics. Open the source Over 1,200 corpus-derived grammar can-do statements from the Cambridge Learner Corpus.Bayesian Knowledge Tracing, a 25-year PRISMA review (2024). Twenty-five years of Bayesian Knowledge Tracing. User Modeling and User-Adapted Interaction. Open the source Enhanced variants beat vanilla BKT at predicting answers. Prediction accuracy is what this field measures. Validation against learning outcomes is thin.Bodily, Kay, Aleven, Jivet, Davis, Xhakaj & Verbert (2018). Open learner models and learning analytics dashboards: a systematic review. LAK'18. Open the source 102 articles, 107 systems, and no pooled effect on learning outcomes. Open learner models are the better-evaluated dashboard family and outcome evidence remains thin. Nothing here supports a claim that reading a dashboard improves learning.Prescriptive learning analytics dashboards. From descriptive to prescriptive learning analytics dashboards. Open Research Online record, The Open University. Open the source A small-sample case study in which learners valued explicit recommendations to study material over charts alone. Design evidence. Efficacy is unevaluated here, and this entry's citation metadata was carried from our own research record rather than re-verified against the publisher page.Rodriguez (2005). Three options are optimal for multiple-choice items: a meta-analysis of 80 years of research. Educational Measurement: Issues and Practice. Open the source Three functional options performed at least as well as four once the fourth option was a weak distractor. Fewer options does not mean less guessing: with three choices a blind guess is still right about a third of the time.Butler & Roediger (2008). Feedback enhances the positive effects and reduces the negative effects of multiple-choice testing. Memory & Cognition. Open the source A wrong option a learner selects can be remembered as true. Feedback after the test reduced that effect in the reported experiments, which is why reviewed explanations arrive after a check-up rather than never.
Reminders, goals and habit
Nobbe, Breitwieser, Biedermann & Brod (2024). Smartphone-based study reminders can be a double-edged sword. npj Science of Learning. Open the source Reminders raised same-day study odds (OR 1.77) and lowered study on non-reminder days (OR 0.45). The authors recommend fading reminders out. One 36-day trial with 85 children.Kizilcec et al. (2020). Scaling up behavioral science interventions in online education. PNAS. Open the source 247 courses, 269,169 learners. Light-touch nudges lost roughly an order of magnitude of their effect at scale and suited one-time actions rather than sustained habits.Bälter et al. (2023). Effect of personalized email-based reminders on participants' timeliness in an online education program. JMIR Formative Research. Open the source A progress-aware email moved on-schedule rates from 56 to 70 percent against a generic email at the same cadence. A 39-person crossover pilot, so the evidence is thin.Beshears, Lee, Milkman, Mislavsky & Wisdom (2021). Creating exercise habits using incentives: the trade-off between flexibility and routinization. Management Science. Open the source N = 2,508. Rigid daily windows produced fewer sessions than flexible incentives, during the programme and after it ended. A field experiment on gym attendance, so transfer to study habits is an inference.Lally, van Jaarsveld, Potts & Wardle (2010). How are habits formed: modelling habit formation in the real world. European Journal of Social Psychology. Open the source Median 66 days to automaticity, and missing a single day did not materially hurt habit formation.Wang, Wang, Wang, Wind & Gill (2024). A systematic review and meta-analysis of self-determination-theory-based interventions in education. Learning and Motivation. Open the source Bias-corrected pooled effect on intrinsic motivation g = 0.58, with extreme heterogeneity. The paper's need-specific sub-finding did not survive our verification and is cited nowhere here.Silverman & Barasch (2022). On or off track: how (broken) streaks affect consumer decisions. Journal of Consumer Research. Open the source The effect is driven by how the log is displayed. A repair option attenuates the harm of a broken streak.
Tutoring and AI tutors
Bastani, Bastani, Sungu, Ge, Kabakcı & Mariman (2025). Generative AI without guardrails can harm learning: evidence from high school mathematics. PNAS. Open the source Unguarded chatbot access lifted assisted grades by 48 percent and cost 17 percent on unassisted exams once removed. A hint-only tutor neutralised the harm without adding a gain. Our AI Tutor implements the guarded condition: it opens with recall or a hint and reveals a worked answer only under the learner's answer policy.Wang et al. (2024). Tutor CoPilot: a human-AI approach for scaling real-time expertise. arXiv 2410.03017. Open the source A preprint on K-12 mathematics. LLM guidance for human tutors raised topic mastery by 4 points, and by 9 points for students of lower-rated tutors.Education Next (2024). Two-sigma tutoring: separating science fiction from science fact. Education Next 24(2). Open the source A secondary source. Bloom's two-sigma graph was hand-drawn and illustrative rather than fitted to data. Tutoring meta-analyses land around 0.33 to 0.37 standard deviations, and none of 96 studies reached two sigma.VanLehn (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist. Open the source Human tutoring around d = 0.79, well short of two sigma.Kulik & Fletcher (2016). Effectiveness of intelligent tutoring systems: a meta-analytic review. Review of Educational Research. Open the source Intelligent tutoring systems around d = 0.76, median 0.66.Chi, Siler, Jeong, Yamauchi & Hausmann (2001). Learning from human tutoring. Cognitive Science. Open the source Tutoring gains tracked what the learner constructed rather than what the tutor explained. This is why our Tutor asks you to recall or repair before it explains, and it says nothing about how well this product does it.Graesser, Lu, Jackson, Mitchell, Ventura, Olney & Louwerse (2004). AutoTutor: a tutor with dialogue in natural language. Behavior Research Methods, Instruments, & Computers. Open the source The staged prompt-hint-answer structure this product borrows. We took the sequence, not the results: AutoTutor's own effect sizes are not evidence about Infinite Story.Aleven and colleagues. Help seeking and help design in interactive learning environments. Author-hosted copy, Carnegie Mellon University. Open the source Documents help avoidance, rapid clicking through hints, and the need for a learner to make sense of a hint before it helps. Cited for those design observations only. This is an author-hosted copy and its published venue line was carried from our own research record rather than re-verified against a publisher page.AI tutoring in an undergraduate physics course (2025). AI tutoring outperforms in-class active learning. Scientific Reports. Open the source One structured AI tutor in one university course, using authoritative solutions and controlled sequencing. Promising contextual evidence for the shape, not evidence of general efficacy. Its citation metadata was carried from our own research record rather than re-verified against the publisher page.
Company evidence and the product landscape
Kessler (2023). An independent head-to-head comparison of major language-learning apps. Computer Assisted Language Learning. Open the source Found no outcome difference between the apps compared.Settles & Meeder (2016). A trainable spaced repetition model for language learning. ACL 2016. Open the source Company evidence, as reported by Duolingo: their half-life regression model beat Leitner and Pimsleur baselines and lifted any-activity retention by 12 percent in an A/B test. An independent replication put its ranking accuracy barely above chance.Yancey & Settles (2020). A sleeping, recovering bandit algorithm for optimizing recurring notifications. KDD 2020. Open the source Company evidence, as reported by Duolingo: optimising reminder content moved daily actives by about half a percent, with new-user retention around two percent. Repeating one message scored below random.Babbel efficacy study (2016). Babbel efficacy study. Company white paper. Open the source Company evidence, as reported by Babbel. Vendor-funded, and the reported rate collapses from 18 points per hour to roughly 5 past novice level.Gamification misuse in Duolingo (L@S 2022). A qualitative case study of streak, XP and leaderboard fixation. ACM Learning @ Scale 2022. Open the source Peer-reviewed and independent of the company: 30,618 forum comments plus interviews. Fixating on streaks and XP instead of learning degrades outcomes.
Getting the most out of it
Follow a link before believing a number. Several of these are single small trials.
Read the label on company evidence. Vendor-funded results are reported here as reported by the vendor.
Treat a secondary source as a pointer to the primary one.
Under the hood
Each claim used on these pages went through three independent refute-framed verification votes against its primary source. A claim survived at two non-refuted votes.Refuted and caveat-surviving claims are either dropped or restricted to a fixed public wording, which is what the labels above record. Claims that were extracted but never verified are used internally and cited nowhere.