Learning with AI: what the research really says
Studies on AI effectiveness in learning show positive results under specific conditions and null or negative effects in other configurations. The decisive variable is not the technology but the pedagogical design: AI that forces retrieval effort improves outcomes, while AI that does the work for the learner degrades them. This finding has been robust in the scientific literature since VanLehn's (2011) work on intelligent tutoring systems, confirmed by meta-analyses from the Institute for Education Sciences.
What positive studies show
Well-designed Intelligent Tutoring Systems (ITS) produce documented, significant learning gains. Kurt VanLehn's (2011) meta-analysis, covering several dozen comparative studies, found that the best ITS produce gains comparable to one-to-one human tutoring — around 0.76 standard deviations above standard classroom teaching. Researchers call this Bloom's "two sigma problem": individualised teaching produces results far superior to mass instruction, and ITS come close.
These gains appear consistently for structured, assessable skills: mathematics, physics, languages, programming, science. They appear among learners with stable learning conditions — learners who have the time, access to technology and sufficient motivation to engage with the system.
What studies qualify
Effects vary greatly by implementation. A system that simply distributes video content with an AI interface bolted on produces no measurable gain versus video without AI. The quality of the underlying pedagogical content matters more than technological sophistication: poor content well distributed by AI remains poor content.
Available meta-analyses of MOOCs show completion rates between 3% and 15% and highly variable learning effects depending on learners' level of active engagement. Platforms that integrate active retrieval mechanisms (questions, exercises, regular assessments) achieve significantly better outcomes than those offering essentially passive video.
Effects on complex behavioural skills — leadership, interpersonal communication, creativity, conflict management — are far less well documented.
What studies warn about
Using generative AI to produce answers without any retrieval effort from the learner degrades learning. This is the most important and most ignored finding in commercial discourse.
The mechanism has been documented since classic work on the testing effect or retrieval practice effect: forcing yourself to recall information from memory produces far superior long-term retention than re-reading or passively listening. Reading an AI-generated answer is a form of passive re-reading: the information is easily absorbed, but 30-day retention is very poor.
The sense of understanding that comes from reading a well-phrased model output is deceptive. Researchers call this the "fluency illusion": because the answer is fluent and coherent, the learner feels they have understood, yet shallow comprehension without retrieval effort does not produce durable memory.
How to read these studies rigorously
Four criteria help assess the quality of a study on AI effectiveness in training.
Funding. A study commissioned by an AI solution vendor on its own product does not carry the same weight as an independent study.
Control group. A study that compares "before AI" and "after AI" in the same group without a control group cannot isolate the AI effect. A randomised controlled trial is the gold standard.
Population scope. Findings from primary-school pupils in mathematics in the United States do not automatically transfer to senior managers in professional training in France.
Measurement duration. Short-term and six-month effects can diverge significantly. Studies that measure only immediate effects systematically overestimate effectiveness.
Frequently asked questions
Is AI more effective than traditional training?
It can be, for certain types of learning under certain conditions. The relevant question is not AI versus traditional, but which combination for which objective and with which pedagogical scenario.
Where can you find robust studies?
Peer-reviewed education journals (Journal of Educational Psychology, Learning and Instruction), OECD publications and meta-analyses from the Institute for Education Sciences. Vendor white papers without peer review are sales collateral, not scientific evidence.
Do gains last?
When active retrieval is present throughout the learning path, gains are sustained at six months. When content is consumed passively, retention collapses within weeks in line with the Ebbinghaus curve.