SAVRN
Search Contact SAVRN

Savrn Insights · The Superintelligence Transition · Chapter 3 of 11

Disclosure: prepared for Savrn, which has a commercial interest in AI infrastructure and publishes the seven trackers cited here. Not an independent institutional review. Not investment, medical, legal, or tax advice.

Chapter 3 · Schooling, K-12

Does AI Help Students Learn? What 10 K-12 Studies Actually Show

50 min read11,528 wordsOpen as its own page

A copper school desk and notebook with a ghosted blueprint robot hand reaching in to write the answers
Savrn Insights · Chapter 3 illustration

It usually happens at the kitchen table. Your kid has a laptop open, an AI tool in one tab and the homework in another, and the answers are coming fast. Too fast. And you ask yourself the question every parent I know is asking right now: is this tool helping my kid learn, or is it doing the homework for them? That's the question this chapter takes to the AI in education research, and I want to answer it the way I'd answer any question about a system I was about to bet money on: by reading what the studies actually did, not what the brochures say they did.

Here's why it matters. Schools are buying. Districts are signing contracts, teachers are being handed new tools, and families are being told their children will get "personalized learning." Some of that is backed by real evidence. Some of it is not. And the difference between the two is often invisible unless you know exactly where to look: who was in the study, what the comparison group got, whether the kid was tested with the tool or without it, and whether anyone independent checked the work.

I reviewed ten study programs spanning kindergarten through high school. The good news is real. Structured systems, tools that support the teacher rather than replace the teacher, and supervised programs have produced measurable benefits for real students. The uncomfortable part is also real, and it's the part most people skip. In the best-designed high-school experiment in this set, an unrestricted chatbot made students look much better on practice and then left them worse off than students who never had it, once the tool was taken away. Same subject. Same kind of students. The guardrails made the difference.

So this chapter will not tell you AI is good for kids or bad for kids. It will tell you which arrangements have been tested, what they showed, where they stop, and what questions to ask before your school spends money or your child spends hours. Capability is not benefit. A powerful model sitting in a browser tab is not the same thing as a child who can now do the work on her own.

3.1 How to Read AI in Education Research Without Getting Fooled#

Before we get to a single study, you need to know what kind of review this is and what rules I'm using to read the numbers. I'm spending time here on purpose. Most bad decisions about educational technology don't come from bad studies. They come from good studies read badly.

3.1.1 Scope: Three Grade Bands, Ten Study Programs#

This chapter reviews the evidence for assisted learning across the three school bands that organize most adoption decisions: kindergarten through grade 4, grades 5 through 8, and grades 9 through 12. It examines ten study programs. I prioritized randomized field evaluations, independently measured outcomes, original research reports, and, where one exists, an independent evidence review. Policy guidance shows up later in the chapter, but only as guidance on safeguards. It is never used as proof that a tool improves learning.

The evidence is international. The studies come from the United States, the United Kingdom, Nigeria, and Turkey-connected research settings, and I don't assume that a result transfers unchanged across curricula, languages, staffing arrangements, or school systems. A British "Year 7" is not automatically an American seventh grade. A Nigerian first-year senior secondary student is not automatically an American ninth grader. Where the source population matters to how you should read a result, I keep it attached to the result.

This is a targeted review of decision-relevant research. It is not a systematic review, and I'm not claiming to have located every relevant study in the world. I also don't calculate a pooled effect size, because the interventions, comparison conditions, ages, subjects, and outcomes differ too much for an average to mean anything. Averaging a kindergarten reading package with a high-school chatbot experiment would produce a number, and that number would be a rumor. Positive, null, uncertain, and adverse findings are all kept, with the same weight they carry in the original work.

3.1.2 The Four Interpretation Rules#

Four interpretation rules govern every number in this chapter. Learn them once and you'll read every education headline differently for the rest of your life.

  • Standard deviations are not percentages. A result of 0.23 standard deviations is a standardized score difference, not 23 percent more learning.
  • Percentage points are not percent. A pass rate rising from 62 percent to 66 percent is a four-percentage-point increase, not a 66 percent learning gain.
  • Relative percentages need their denominator. A reported 17 percent reduction in an exam score is not a 17-percentage-point change unless the study defines it that way.
  • Assignment is not use. The effect of offering a program differs from an estimate for people who actively use it. Selecting only the active users invites selection bias unless the analysis deals with it.

3.1.3 The Bundle Rule#

There's one structural rule on top of those four. Software, extra instruction, teacher training, devices, supervision, and curriculum changes usually arrive together as a bundle. A result for the package is not automatically a result for the model alone. Nearly every study in this chapter bundles something, and the bundle is part of what each result means.

3.2 What AI in Education Research Shows Across Three Grade Bands#

The evidence supports a conditional conclusion, and I'm asking you to hold the whole thing, not just the half you like: specific, well-designed learning systems can improve particular educational outcomes, but access to a powerful model does not itself establish educational value. The support for both halves of that sentence runs through the kindergarten reading trial, the ASSISTments technical report, the teacher preparation trial, and the high-school mathematics experiment published in PNAS.

So the practical question for a school, a district, or a parent is not whether to embrace or reject a technology category. It's which combinations of students, teachers, instructional methods, tools, and safeguards produce independently measured benefits that justify their costs. That's a harder question. It's also the only one worth asking.

Here are the ten study programs, sorted by band:

Grade band Study programs Direction of evidence
Kindergarten through grade 4 Kindergarten reading (A2i package); early mathematics applications; literacy engagement support Positive for teacher-directed and bounded packages; engagement without established literacy gains
Grades 5 through 8 ASSISTments; Tutor CoPilot; Khanmigo remediation package; teacher preparation time Small or qualified gains; session mastery without established annual attainment; bounded teacher-time savings
Grades 9 through 12 PNAS mathematics assistants; Nigerian English program; Cognitive Tutor Algebra I Adverse risk from unrestricted assistance; positive supervised program; historical adaptive-system evidence
Figure 3.1Data

Ten study programs by grade band and type of assistance

ADULT-FACINGSTRUCTURED SOFTWARESTUDENT-FACING CONVERSATIONALK-4A2i kindergarten readingEarly math appsLiteracy engagement supportGrades 5-8Tutor CoPilotTeacher prep timeASSISTmentsKhanmigo staffed packageGrades 9-12Cognitive Tutor Algebra IPNAS unrestricted assistantNigeria supervised EnglishFavorableSmall or qualifiedNull or uncertainAdverse
Where each program sits. Tutor CoPilot also includes grade 3 and 4 sessions. The PNAS experiment's constrained tutor largely avoided the unrestricted assistant's penalty. Direction summarizes the primary finding; see the text for certainty and limits.

The sections that follow take each band in turn, study by study, with the design, the finding, and the boundary attached to each. One framing note first. The schooling debate is usually staged as a contest between "innovation" and "caution." The evidence I reviewed supports neither banner. It supports a procurement discipline: name the learner, the intervention, the comparison, the outcome, and the cost, and then look at what the comparison actually showed. That's the whole game.

3.3 AI Tutoring Studies, Kindergarten Through Grade 4#

The early-grade evidence is most useful when it tells you three things: what role the adult plays, what skill is being taught, and how limited the child's direct interaction with the tool is. The studies in this band test assessment-informed teaching, bounded mathematics applications, and a literacy platform with and without human engagement support. They do not test what happens when a young child is handed an unrestricted conversational assistant, and I won't pretend they do.

That gap matters. If you're the parent of a six-year-old and someone tells you "the research shows AI helps young children learn," the right response is: which AI, doing what, with which adult, measured how? In this band, the answers are always some version of "a structured tool, working through or alongside a teacher, measured outside the tool."

3.3.1 Kindergarten Reading: The Tool Behind the Teacher#

The strongest early-grade result in the reviewed set comes from a 2011 cluster-randomized study of Individualized Student Instruction for Kindergarten. The evaluation involved 556 students, 44 teachers, and 14 schools. Students in intervention classrooms outperformed comparison students on a latent measure of reading skills, with a reported effect size of 0.52. (Al Otaiba and colleagues, original study record)

Here's the part most people skip: the intervention. It used assessment-informed recommendations associated with Assessment2Instruction (A2i), combined with professional development and classroom support, to help teachers individualize instruction. The system recommended instructional amounts and groupings. The teacher still organized instruction, interpreted student needs, and worked directly with children. Information flowed from student assessments, to recommendations for the teacher, and then to differentiated classroom instruction. This was not an open chatbot teaching a kindergartner on its own. (Connor, research synthesis and intervention description)

Four boundaries attach to the 0.52 estimate, and I'd want every one of them in any school board presentation that cites it.

First, it concerns a latent reading construct, a statistically estimated measure of reading skill. It is not a universal improvement across every reading test or every early-grade student.

Second, the publicly available study abstract does not report a confidence interval for the estimate. Without an interval, you can't see how much uncertainty surrounds the 0.52.

Third, the result supports the tested instructional package including adult implementation, not a software-only effect. A school considering a similar approach should test the complete delivery package, including the time and training required for teachers to act on the recommendations.

Fourth, a disclosure belongs in the record: the later A2i synthesis discloses the author's equity interest in Learning Ovations, which matters when assessing the broader presentation of the intervention's evidence. (Study abstract, Disclosure in the synthesis)

A disclosed interest does not make a study wrong. It tells you where to look harder. I'd say the same about anybody's claims, including ours.

The defensible practical reading is that intelligent support can operate behind the teacher rather than in place of the teacher. That's a useful finding. It's also narrower than the one usually advertised.

What a parent should take from it. If your child's kindergarten uses a system like this, the useful question isn't "does my child use the AI?" It's "how does the teacher use what the system tells her, and how do you check whether my child is reading better without it?"

Common misreading. "AI raised kindergarten reading scores by half a standard deviation." No. A teacher-led instructional package, informed by assessment-based recommendations and supported by training, was associated with that difference on a latent reading measure in this trial. Take the teacher out of that sentence and you've described a study nobody ran.

3.3.2 Early Mathematics: Structured Apps, Not Open Conversation#

A United Kingdom randomized trial published in 2019 evaluated interactive mathematics applications with children aged four and five in their first compulsory school year. The trial randomized 461 children across 12 schools, with 389 children from 11 schools available at posttest. This is age-adjacent evidence for the youngest part of the K-12 range, not an exact equivalent of a U.S. kindergarten. (Outhwaite and colleagues, full article)

Here's what the study actually did. Children used the applications for approximately 30 minutes per day over 12 weeks, supervised by classroom staff. Three groups were compared. One group received app-based mathematics in addition to its normal mathematics activities. A second group used the apps in place of a daily small-group mathematics activity. The control group continued normal teaching.

Comparison with normal teaching Reported result Interpretation
Additional app-based mathematics time Progress effect size 0.31; 95% confidence interval 0.06 to 0.55 A positive result for additional structured mathematics practice; additional mathematics exposure is part of the treatment
App use replacing a daily small-group activity Progress effect size 0.21; 95% confidence interval -0.03 to 0.46; significance from a one-tailed test More tentative: the two-sided interval includes zero
Supplementary versus time-equivalent groups No statistically significant difference Not proof that the two approaches are equivalent
Figure 3.7Data

Early mathematics apps with four and five year olds

-0.100.20.40.6Progress effect size versus normal teaching, 95% CIApps added to normal mathextra structured practice0.31 (0.06 to 0.55)Apps replacing a small-group activitytwo-sided interval includes zero0.21 (-0.03 to 0.46)
United Kingdom randomized trial, 461 children in 12 schools, 389 at posttest, about 30 minutes a day for 12 weeks, supervised. No significant difference between the two app groups, which is not proof they are equivalent.

Source: Outhwaite et al., 2019

The supplementary result is clearly positive, but those children got their normal math plus the apps, so some of that gain may simply be more math. The replacement result is the more interesting policy number and the less certain one: its interval crosses zero. And "no statistically significant difference" between the app groups means the study couldn't tell them apart, not that they are equivalent.

The applications supplied bounded tasks, visual and spoken instructions, immediate feedback, repeated practice, and progression requirements. Children were not relying on unrestricted generated explanations, and the assessment was administered outside the app. (Intervention and assessment methods)

This evidence supports a specific proposition: structured digital practice can contribute to early mathematics learning under tested classroom conditions. It does not establish that replacing play, teacher interaction, or broad early-childhood experiences with more screen time improves overall development. The distance between those two propositions is the distance between what the trial measured and what a vendor might hope it implies.

What a teacher should take from it. If you're deciding whether to use math apps in a reception or kindergarten room, the trial suggests two things. Bounded, feedback-rich practice under adult supervision can help. And the evidence is firmer when app time is added than when it replaces your own small-group work, so be careful about what you give up to make room for it.

3.3.3 Literacy Engagement: More Minutes, No Proven Reading Gain#

A June 2026 working paper examined two randomized trials involving 355 students: one in an after-school setting covering grades 1 through 5, and one in an in-school setting covering grades 1 through 3. Both treatment and comparison students had access to the same literacy platform. The treatment added human support intended to encourage engagement, not to provide reading instruction. (Robinson and colleagues, full working paper)

The support differed between settings. The in-school trial used middle-school peer support, which means the support people can't all be assumed to be adults. And the platform can't be assumed to be a general-purpose language model. (Study design)

Measure After-school, grades 1 through 5 In-school, grades 1 through 3
Randomized sample 174 students 181 students
Comparison-group mean platform minutes per week, including nonuse 2.18 minutes 5.23 minutes
Added weekly platform time from human support About 1.00 minute; marginal at the 10% level About 4.42 minutes; significant at the 5% level
End-of-year literacy result No statistically significant improvement No statistically significant improvement

Because both groups had the platform, this is not a randomized comparison of the platform against no platform. What it shows is more awkward and more valuable: the added human engagement component increased some usage measures without producing a statistically established reading gain in either trial. (Full working paper)

For decision-making, this is a warning against counting licenses, logins, or available tutoring hours as learning outcomes. A school should distinguish four different numbers that routinely get compressed into one dashboard metric: scheduled time, actual engaged time, completed practice, and independently measured reading growth.

Read that again. Scheduled time is what the district planned. Engaged time is what kids actually did. Completed practice is what the software logged. Reading growth is what you were paying for. A vendor report that shows the first three and not the fourth has shown you activity, not results.

3.3.4 Grades 3 and 4: Tutor Support Across the Band Line#

Tutor CoPilot, a generative assistant used by human mathematics tutors, included substantial grade 3 and grade 4 representation in its session-level analysis. Because the study crosses the band boundary, I treat it in full in Section 3.4.2. One rule applies here: its pooled effect must not be recast as a separate verified effect for third graders or fourth graders. Pooled results are pooled. I won't manufacture grade-specific precision the study did not produce. (Wang and colleagues, full study)

3.3.5 What Schools Can Responsibly Tell Families of Young Children#

For this age band, the most defensible explanation a school can give families is that tools may help a teacher identify what a child needs next, or provide a bounded opportunity to practice. The positive studies I reviewed involve specified instructional designs and adult implementation, not simply giving a young child a general-purpose assistant. (A2i synthesis, early mathematics trial)

Here's an illustration, clearly labeled as an illustration and not an additional measured result. A teacher reviews assessment information, selects a short activity, observes the child, and checks the same skill later without the tool. A parent receives an explanation of the learning goal and can ask whether the child is improving independently, rather than being asked to accept an opaque "personalization" claim.

Claims that these systems reduce parents' stress, save families money, or replace home reading support would require direct household evidence. Those outcomes were not the causal endpoints of any study in this section, and I won't borrow them from the household chapter, which reaches its own, more conditional conclusions in Chapter 6.

Checklist for parents of children in kindergarten through grade 4:

  • Does my child interact with the tool directly, or does the teacher use it to plan?
  • What exactly can my child ask it, and what can it answer?
  • Which adult is watching while my child uses it?
  • How does the teacher check whether my child can do the skill without the tool?
  • What is my child not doing (play, reading aloud, small-group time) during the minutes the tool is in use?

3.4 AI Homework Help and Tutoring Research, Grades 5 Through 8#

The middle-grade evidence covers three distinguishable arrangements: structured feedback on mathematics practice, assistance delivered through a human tutor, and a conversational tool embedded in a staffed remediation program. Add a fourth study about teacher preparation time, and you have the most varied band in this chapter. The findings are informative exactly as long as those arrangements stay separate. (ASSISTments report, Tutor CoPilot paper, Khanmigo school experiment)

3.4.1 ASSISTments: Why Independent Review Matters#

A North Carolina trial assigned 63 schools to ASSISTments or to business-as-usual mathematics instruction, involving 102 grade 7 mathematics classrooms across 41 districts. The platform provided feedback and hints to students and reports to teachers, with professional development and coaching supporting implementation. (Feng, Huang, and Collins, technical report)

Focal students used the intervention in grade 7 during the 2019 to 2020 school year. The longer-term outcome was the grade 8 state mathematics assessment in spring 2021. The grade 8 outcome analysis included 5,991 students and reported an effect size of 0.10, with a p-value of .011. (Technical report)

That's useful evidence of a modest later assessment difference. It is not a claim of a 10 percent increase in learning. (Rule one: standard deviations are not percentages.) The study also ran through pandemic disruption, and students taking high-school-level mathematics instead of the relevant grade 8 test were excluded from that test's analytic sample. That exclusion matters for reading the result, because it removed a group of students from the comparison. (Technical report, WWC sample description)

Now the independent review, which requires particular care. The April 2026 What Works Clearinghouse review rates the study "Meets WWC standards with reservations," citing high cluster-level attrition with baseline equivalence in the analytic groups. It identifies uncertain effects for the mathematics achievement domain, with a positive grade 8 state-test result and a nonsignificant mathematics-readiness result in a smaller sample. An earlier review displayed on the same page was more favorable, so a citation to the earlier rating alone would leave out material later evaluation. (Latest WWC review, Review history and outcomes)

Figure 3.3Comparison

One study, two records: the trial report and the independent review

TRIAL REPORTGrade 8 state math test:effect size 0.10, p = .01163 schools, 102 grade 7 classrooms5,991 students in the grade 8 analysisPandemic disruption; some studentsexcluded from the analytic sampleWWC REVIEW, APRIL 2026Meets standards withreservationsHigh cluster-level attrition withbaseline equivalenceUncertain effects in the math domainEarlier, more favorable rating supersededDEFENSIBLE READINGPositive state-test result,qualified overallHold both recordsNot a 10% learning gainNot evidence for unrestrictedconversational tutoring
ASSISTments in North Carolina. A district that reads only the trial report gets an inflated picture; one that reads only the review might discard a promising program.

Source: Technical report · What Works Clearinghouse

The defensible conclusion: the state-test result is positive, but the broader independent assessment is qualified, not uniformly confirmatory. Faster identification of errors and better-targeted follow-up are plausible mechanisms consistent with the intervention. They are not separately isolated causal effects. And the result belongs to the ASSISTments instructional package and its tested comparison conditions. It is not evidence for unrestricted conversational tutoring.

This study is my standing example of why independent review matters. The trial team's report and the federal reviewer's later assessment of the same study produce different confidence levels, and both belong in the record. A district that reads only the first gets an inflated estimate. A district that reads only the second might throw out a promising program. The discipline is to hold both, which means holding a smaller and more accurate conclusion than either source alone suggests.

3.4.2 Tutor CoPilot: Helping the Human Tutor#

Tutor CoPilot randomized 900 tutors to access or no access to an adult-facing assistant that suggested ways to respond to students during live online mathematics tutoring. The study identified 1,787 participating students in nine schools. The analyzed session sample contained 4,136 sessions taught by the remaining participating tutors. (Full study, January 2025 version)

Here's the design choice that makes this study worth your attention. The assistant could suggest a question or an explanation, and the tutor decided what to use. The child still talked to a human tutor. What changed was the support available to that tutor. That's a very different arrangement from putting a chatbot in front of a student.

Outcome or boundary Finding
Main randomized assignment result Exit-ticket passing rose from approximately 62% to 66%, a four-percentage-point increase; p < .01
Lower-rated tutor subgroup A nine-percentage-point improvement; a subgroup result, not the overall effect
End-of-year mathematics assessment No statistically significant improvement
Exposure period Approximately two months
Coverage limitation Grade 5 and grade 6 sessions support relevance to this band; the final session sample does not establish effects for grades 7 and 8
Publication status Preprint; not treated here as a peer-reviewed journal finding

The result supports near-term lesson mastery under this tutor-assisted arrangement. It does not establish durable improvement in a student's overall mathematics attainment. Developer and provider involvement disclosed in the paper stays visible in my assessment. For a district, the next questions are whether the lesson-level improvement persists on independent assessments, and whether it's still worth it after training, supervision, and full delivery costs are included. (Outcomes and disclosures)

The four-percentage-point headline and the null end-of-year result must be held together. They describe different links of the benefit chain from Chapter 2. The first is task performance within sessions. The second is the real-world outcome a parent actually cares about. An adoption decision made on the first number alone is exactly the kind of decision this report exists to prevent.

Notice also the subgroup finding. Lower-rated tutors improved by nine percentage points. That's intriguing, because it suggests the tool may do the most good where tutoring quality is weakest. But it's a subgroup result, not the overall effect, and subgroup results are where hopeful readers go to find the number they wanted. Treat it as a question for the next study, not an answer for this one.

3.4.3 Khanmigo Study Results in a Staffed Remediation Package#

An August 2026 working paper examined Khan Academy with Khanmigo in grades 6 through 8 across 18 middle schools in Hamilton County, Tennessee, over two school years. Randomization involved 53 grade-within-school clusters, and the analysis included 6,902 student-term observations. That's not 6,902 distinct children. The same student can show up more than once across terms. (Oreopoulos and Low, full working paper)

The treatment combined a mathematics practice platform, conversational support, assigned learning paths, and scheduled remediation periods staffed by adults. The comparison was existing remediation, which itself could include teacher-led work and other digital learning products. The experiment therefore does not isolate the incremental effect of Khanmigo over otherwise identical Khan Academy use. There's a bundle on both sides of the comparison, which limits what either side can claim. (Design and comparison conditions)

The paper's pooled MAP result was an increase of approximately 1.26 national percentile-rank points per term, with a conventional standard error of 0.60. A small-cluster robustness test gave p = .0511, which makes the certainty of the main pooled result sensitive to the inference method. (Primary results, Robustness analysis)

That robustness detail is easy to skim past, so let me slow down on it. With only 53 clusters, the way you calculate uncertainty matters. Under the conventional method, the pooled result looks solid. Under a method built for a small number of clusters, it sits right at the edge of the usual significance threshold. Neither calculation is a trick. Together they tell you the result is real enough to take seriously and fragile enough not to oversell.

Annual standardized estimate Reported result Appropriate reading
First year 0.020 standard deviations, SE 0.035; not statistically significant No demonstrated first-year annual gain
Second year 0.084 standard deviations, SE 0.041; significant at the conventional 5% level A modest positive second-year package result
Stacked annual estimate 0.062 standard deviations, SE 0.035; marginal at the 10% level Promising but less certain than a simple "proven annual gain" statement
Figure 3.4Data

Khanmigo staffed remediation package: annual estimates

-0.100.10.2Standard deviations (bars are approximate 95% intervals computed as estimate ± 1.96 SE)Year 1not significant0.020 (SE 0.035)Year 2significant at 5%0.084 (SE 0.041)Stackedmarginal at 10%0.062 (SE 0.035)
Grades 6 to 8, 18 Hamilton County middle schools, 53 randomized clusters, 6,902 student-term observations. The year-to-year difference is not itself significant, and the package does not isolate the chatbot from the rest of the program.

Source: Oreopoulos and Low, working paper

The difference between the first-year and second-year annual estimates was not itself statistically significant. So one significant year and one nonsignificant year must not be described as a proven improvement between years. The report also notes that the preregistered state-test co-primary outcome had not yet been incorporated into the analysis. And its usage findings show that having a conversational tool available did not automatically translate into sustained explanatory dialogue. (Annual results, Outcome availability and usage analysis)

The appropriate conclusion is that this staffed remediation package produced modest, qualified evidence of benefit. It is not yet a clean demonstration that adding a chatbot alone improves outcomes for middle-school students.

3.4.4 Teacher Preparation Time: A Different Kind of Outcome#

A randomized trial involving 259 teachers in 68 English secondary schools tested ChatGPT with a preparation guide for Year 7 and Year 8 science lessons. The population is relevant to the middle-grade discussion, but it keeps its original label: United Kingdom school-year labels are not automatically U.S. grades. (NFER/EEF report)

In the measured later weeks of the trial, intervention teachers reported approximately 56.2 minutes per week preparing the specified lessons, compared with 81.5 minutes in the comparison group. That's roughly 25.3 minutes, or 31 percent, less preparation time for the measured work, with a reported time-ratio confidence interval of 0.53 to 0.90. (Primary results)

Figure 3.6Data

Teacher lesson preparation time, measured weeks of the trial

0255075100Comparison teachers81.5 min/weekChatGPT plus preparation guide56.2 min/weekAbout 25.3 minutes (31%) less; time-ratio CI 0.53 to 0.90; teacher diaries; 211 of 259 teachers analyzed.
Year 7 and 8 science lessons in 68 English secondary schools. This is time preparing specified lessons, not total workload, and the trial did not measure student achievement.

Source: NFER / EEF evaluation

The outcome came from teacher diaries, and the primary analysis included 211 teachers rather than every randomized participant. A blind review of 30 submitted lessons detected no quality difference, but a 30-lesson sample does not prove universal quality equivalence. The study did not measure student achievement. (Methods, attrition, and quality assessment)

This establishes a bounded operational benefit worth testing locally: 31 percent less time preparing specified science lessons under a preparation guide. It is not a 31 percent reduction in teachers' total workload, and it's not a measured improvement in anything students learned. Whether the saved minutes went to individualized teaching is a question the trial did not ask, and I won't answer it on the trial's behalf.

What a teacher should take from it. If your school offers AI tools for lesson prep, this is the best evidence in the set that it can save time on a defined task with a guide. Ask for the guide, not just the login. And keep an eye on verification time: checking generated material is real work, and the trial's quality review was small.

3.4.5 What the Middle-Grade Evidence Supports#

The evidence supports trials of bounded mathematics feedback, adult-mediated tutoring support, and carefully evaluated remediation packages, with separate measurement of student attainment and teacher time. It does not support treating all conversational access as equivalent, or stretching a positive session-level result into a full-year learning guarantee. (ASSISTments review, Tutor CoPilot, Khanmigo evaluation)

Here's an illustrative workflow, proposed as a design principle for evaluation rather than claimed as a proven effect: students attempt a problem, receive a limited hint, explain their reasoning, and later solve a related problem without assistance. Every clause in that sequence is doing evaluative work. The last clause is the one most existing programs leave out.

For parents of middle schoolers dealing with AI homework help at home, the research doesn't hand you a rule, but the pattern across these studies points to a practical habit. The studies that showed gains used tools that gave hints and feedback, kept an adult involved, and measured the student without help. You can copy that structure at the kitchen table: let the tool give a hint, have your child explain the step back to you, and then have them try a similar problem with the laptop closed. That's not a tested intervention. It's the evaluation logic of the better studies, applied at home.

3.5 Does ChatGPT Help Students Learn? High-School Evidence#

High school is where the difference between producing an answer and acquiring a skill becomes impossible to avoid. A teenager with a capable chatbot can produce a lot of correct-looking work. Whether she can do that work herself next month is a separate question, and in this band the evidence answers it directly. The reviewed studies include an adverse effect from unrestricted assistance, a promising supervised program, and an older adaptive intervention whose effects differed by implementation year. (PNAS mathematics experiment, World Bank English evaluation, RAND algebra evaluation)

3.5.1 The PNAS Experiment: Better Practice, Weaker Learning#

A peer-reviewed 2025 study in PNAS ran a field experiment with nearly 1,000 high-school students in mathematics. It compared three conditions: a GPT-4-based general assistant, a more constrained tutoring version, and a control condition. The tutoring version incorporated teacher-provided instructional material and restrictions intended to guide students rather than simply hand them answers. (Bastani and colleagues, original paper)

Now here's what happened. The unrestricted assistant increased assisted practice performance by 48 percent relative to the control group. The tutoring version increased it by 127 percent. Then access was removed for subsequent exams, and the picture flipped. Students in the unrestricted-assistant condition performed 17 percent worse than the control group, while the constrained tutor largely mitigated that adverse effect. (Reported experimental results)

Read that again. The students with the open chatbot looked better on practice and then did worse on the exam than students who never had it.

Figure 3.2Data

Better practice, weaker learning: the PNAS high-school math inversion

control groupAssisted practice, unrestricted+48%Assisted practice, tutor version+127%Later unaided exam, unrestricted-17%Later unaided exam, tutor versionpenalty largely mitigated; no gain establishedACCESS REMOVED
Nearly 1,000 high-school math students. Percentages are relative to the control group on the study's own measures, not percentage points. The tutor version avoided the penalty; that is not the same as a demonstrated gain in unaided learning.

Source: Bastani et al., PNAS 2025

These percentages describe the study's measured performance. They must not be rewritten as percentage-point changes, and they must not be generalized to every subject. (Rule three: relative percentages need their denominator.) The tutoring result also must not be turned into a claim of a demonstrated positive independent-learning effect just because it avoided the unrestricted tool's penalty. Avoiding a penalty is not the same as producing a gain, and I decline to launder one into the other.

What the experiment demonstrates, with unusual clarity, is the third link of the benefit chain failing in the wild. Both assistants made practice look better. Only one of them left students able to perform after the assistance was gone. The design implication follows directly: completed assignments can't be read as learning, and any deployment that doesn't include an unaided measurement is flying blind on the outcome that matters most.

This is not an argument against educational assistance. It's evidence that the structure of the assistance is part of the intervention and has to be evaluated. The same kind of tool, used by the same kind of students, in the same subject, produced learning protection in one form and learning loss in the other, depending on the guardrails around it.

This is the study that answers the kitchen-table question most directly. If the tool is doing the work, your child's practice will look great and her independent performance may suffer. If the tool is built to guide instead of answer, that penalty can largely go away. The only way to know which one you've got is to test your child without it.

3.5.2 The Unaided-Assessment Rule for Any School Deployment#

The PNAS design generalizes into a simple rule, and it's the single most useful thing a school can take from this chapter. Any deployment should have three stages: assisted practice, removal of assistance, and an independent measurement of what the student can now do alone.

Figure 3.5Framework

The unaided-assessment rule for any school deployment

01 BaselineMeasure the skillbefore any assistance02 AssistedpracticeHints, feedback, ortutoring under setlimits03 AssistanceremovedThe step mostdeployments nevertake04 Unaided testSame skill, no tool,independent grading05 Delayed andtransfer testWeeks later, onunrehearsed problems
Generalized from the PNAS design into a proposed evaluation workflow. Proposed, not tested as a standard.

Here's a hypothetical to make it concrete. A high school wants to pilot an AI writing assistant in tenth-grade English. Under the unaided-assessment rule, the pilot plan would say, before anything starts: students will use the tool for drafting practice during the unit; at the end of the unit, they will write an in-class essay without the tool; that essay will be scored by teachers who don't know which students used the tool; and the comparison will be against classes that didn't use it. If the school can't commit to that last essay, it isn't running an evaluation. It's running a subscription.

Nothing in that illustration is a measured result. It's the evaluation structure the PNAS study makes impossible to ignore.

3.5.3 Nigerian Senior Secondary English: Promising Supervised Use#

A World Bank working paper evaluated a six-week program using Microsoft Copilot, powered by GPT-4, for first-year senior secondary students learning English in Nigeria. The reported effect was 0.23 standard deviations on English and 0.31 standard deviations on a composite assessment combining English, knowledge of the technology, and digital skills. (World Bank publication record and abstract)

The delivery details matter a lot here. The program ran in nine public schools in Benin City, with 1,328 randomized students and 759 in the final analytic sample presented by the authors. It comprised twelve 90-minute after-school sessions over six weeks, with students working in pairs, curriculum-oriented prompts, and teacher guidance and monitoring. So the intervention combined a model with additional structured learning time, access arrangements, prompts, peer interaction, and supervision. The result is evidence for that package, not a model-only effect holding all other instructional resources constant. (Authors' methodological presentation)

The drop from the randomized sample to the analyzed sample is a material limitation. The authors report attrition-related robustness analyses rather than ignoring the problem. Those analyses strengthen the interpretation, but they don't make missing outcome data irrelevant. (Sample flow and robustness presentation)

Four claims, four treatments:

Claim Evidence-based treatment
English learning improved in the evaluated program Supported by the reported 0.23-standard-deviation English result
The composite result was an English-only effect Incorrect: the 0.31 estimate combined English with technology knowledge and digital skills
Students completed years of schooling in six weeks Not established; schooling-equivalent language is a benchmark conversion, not observed completion of additional school years
The same effect holds for American seniors or every high-school subject Not established by this population and subject

The World Bank evaluation also reports heterogeneous effects, including larger gains for female students and for students with higher baseline academic performance. A positive average therefore shouldn't automatically be presented as evidence that the intervention closes achievement gaps. If anything, the reported pattern for baseline performance points the other way. (World Bank study abstract)

This is promising evidence for supervised use in a specified secondary-school setting, not a universal promise. The canonical World Bank abstract supplies the principal effect estimates. The authors' presentation supplies sample-flow and implementation detail, and I'm not presenting my reading of it as a complete audit of the full working paper.

3.5.4 Cognitive Tutor Algebra I: The Version Lesson#

The Cognitive Tutor Algebra I evaluation studied a blended curriculum combining classroom instruction and software across 147 schools, including 73 high schools and 74 middle schools. The high-school sample was predominantly ninth graders, and the intervention was an adaptive subject-specific system, not a modern general-purpose conversational model. (RAND study summary, journal-publication record)

The evaluation found no statistically significant first-year benefit and a positive second-year high-school result. RAND's later addendum reported a significant second-year high-school effect of 0.21 standard deviations, substantively consistent with the earlier analysis. (Study summary, 2014 analytical addendum)

Two cautions are structural. First, the two implementation years involved different student cohorts, not two years of exposure for the same students. So the findings shouldn't be described as proof that an individual learner must use the product for two years to benefit, or as proof that teacher experience caused the difference. Second, the system's era matters. A bounded subject system producing a measurable benefit at scale does not establish that a more general or newer model will outperform it.

The relevance here is historical and methodological. This study shows both that bounded systems can work at scale and that results can differ by implementation year.

3.5.5 Four High-School Years Are Not One Proven Population#

The high-school band includes four years, but the evidence base doesn't justify four grade-specific benefit estimates. The algebra evidence is concentrated in grade 9. The Nigerian study concerns first-year senior secondary students in a different school system. The mathematics chatbot experiment does not establish a separately verified effect for every U.S. high-school year. (RAND sample description, World Bank population description, PNAS study)

Senior-year readiness, independent research, college transition, career preparation, and sustained writing development remain open research questions. They are not assumed consequences of a mathematics or English trial. The absence of a verified estimate in this targeted review is not a claim that no relevant research exists anywhere. It's a refusal to fill the gap with the nearest available number. For what happens after graduation, Chapter 4 takes up college and workforce training on its own evidence.

3.6 How to Explain Effect Sizes to Parents#

Every number in this chapter will eventually get repeated at a PTA meeting, in a school newsletter, or in a parent group chat. Most of them will get mangled on the way. Here's how a teacher, principal, or board member can explain them without adding a single new number, using only the four interpretation rules from Section 3.1.

Start with what a standard deviation is not. When a study says 0.23 standard deviations, say: "This is a way of comparing score differences across different tests. It is not 23 percent more learning." If a parent asks whether 0.23 is big, the fair answer is that it depends on the students, the test, the length of the program, and the cost, which is why the study details matter more than the number.

Say "points" when you mean points. When Tutor CoPilot's exit-ticket passing rose from about 62 percent to 66 percent, that's four percentage points. Say "four more students out of every hundred passed the exit ticket," not "a 66 percent gain" and not "learning went up 6 percent."

Always say "compared with what." When the PNAS study reports that unrestricted-assistant students performed 17 percent worse on later exams, the comparison is the control group in that study. It's a relative change, not a 17-point drop. Say the comparison out loud every time.

Say who got the offer, not just who used it. If a result is reported for everyone assigned to the program, say so. If a vendor shows you results only for "active users," explain that the kids who use a tool the most may have been different from the start, and that the fair comparison is between the groups that were offered it and the groups that weren't.

Name the package. Never say "AI raised reading scores." Say "a teacher-led reading program that used assessment-based recommendations was associated with higher reading scores." It's longer. It's also true.

Here's a hypothetical script for a principal at a parent night, built only from numbers already in this chapter: "Our pilot is modeled on programs where the tool helps the teacher rather than replacing her. In one published study, a similar arrangement helped students pass short end-of-lesson checks more often, but it did not show a significant change on end-of-year tests. So we're going to measure your children at the end of the year, without the tool, and we'll report back either way."

That last sentence is the one parents remember. It tells them the school is testing for the outcome they care about.

3.7 Who Decides What: The Audience Framework#

The same educational outcome can matter to several audiences, but relevance must not be confused with proof of an additional downstream benefit. This table translates the findings into decision questions. It's an analytical framework, not a new causal study.

Audience Interaction to evaluate Demonstrated or directly studied value Additional evidence needed
Student, kindergarten through grade 4 Teacher-selected instruction or supervised, bounded practice Reading and mathematics outcomes in specified instructional packages Independent retention, developmental appropriateness, transfer, effects outside tested skills
Student, grades 5 through 8 Hints, feedback, human tutoring, staffed remediation Small or qualified assessment gains and improved session mastery in particular trials Which component causes improvement; durable learning; effects for the particular grade and learner
Student, grades 9 through 12 Curriculum-linked tutoring with independent assessment Positive supervised-program results and a demonstrated risk from unrestricted answer assistance Graduation, college readiness, employment, cross-subject transfer, year-specific outcomes
Parent or caregiver Understand the goal, supervision, data use, and the child's independent progress Child-level outcomes can inform decisions; household outcomes were not directly established Family time, stress, affordability, usability, access measured directly
Teacher or tutor Review suggestions, adapt instruction, prepare materials Bounded preparation-time savings and improved tutoring-session outcomes Total workload, verification burden, sustainability, actual redeployment of saved time
School board or district Select, fund, monitor, and discontinue programs Study-specific attainment and operational findings as starting benchmarks, not local guarantees Local comparative outcomes, full costs, access, incident reporting, procurement controls
Civic group or infrastructure builder Evaluate the relationship between educational services and physical infrastructure This chapter establishes no site-specific power, water, fiscal, or community finding Separate utility, siting, resource-use, fiscal, and community evidence

Take a stay-at-home parent raising three children. The immediate value of this chapter for that parent is a clearer way to ask whether a school's proposed tool actually helps each child learn. It would be premature to turn these studies into an assumed number of hours saved at home or a promised improvement in family quality of life, and Chapter 6 shows why the household evidence doesn't fill that gap either.

3.8 Conditions for Responsible AI in Schools#

UNICEF's December 2025 Guidance on AI and Children 3.0 recommends age-appropriate systems, protection of children's data, human oversight of consequential decisions, inclusion, and demonstrated educational efficacy before large-scale deployment. It specifically emphasizes keeping teachers central to education rather than replacing them with automated systems. This guidance is a policy and child-rights framework. It is not a randomized estimate of learning gains, and it is not a statement of every jurisdiction's law. (UNICEF Guidance on AI and Children 3.0)

Here's my proposed translation of that framework into an evaluation standard:

  • Adult accountability. Name the teacher, school administrator, or other responsible adult who can intervene, correct an error, or stop use.
  • Age-appropriate interaction. Specify what the student can ask, what the system can answer, and how the design differs between an early reader and an older adolescent.
  • Instructional boundaries. Identify when the system may give a hint, an explanation, a complete solution, or no response.
  • Data minimization. Document what information is collected, retained, shared, and used for training. Don't treat a child's full educational history as a default input.
  • Human recourse. Provide a practical way to challenge an incorrect recommendation or harmful interaction.
  • Accessibility and alternatives. Test the actual interface with intended users, and provide a meaningful alternative when a student cannot or should not use it.
  • Consequential-decision exclusion. Don't infer that a tool shown useful for practice is validated for placement, discipline, diagnosis, or high-stakes assessment.

These controls are consistent with UNICEF's emphasis on developmental appropriateness, privacy, inclusion, and human control. They still need to be translated into applicable school policies, contract terms, and legal requirements before deployment.

Look at "instructional boundaries" next to the PNAS result. The difference between the tool that hurt independent performance and the tool that largely avoided that harm was, in large part, a set of instructional boundaries. This condition isn't bureaucratic box-checking. It's the variable the best experiment in the set turned on.

One distribution caveat belongs here. This targeted set of studies does not supply a consistent set of independently verified effects for students with disabilities, every language group, every income group, or every grade. Local evaluation should state which populations are represented and should not fill missing subgroup evidence with assumptions. A positive average, as the World Bank heterogeneity results show, can coexist with widening gaps.

3.9 AI in Schools Evidence: Measuring Value Consistently#

The central evaluation question for any school deployment should be: what can the student do, what can the adult do, and what does the complete intervention cost compared with a credible alternative? Student learning and operational efficiency should be reported separately before anyone tries to combine them.

3.9.1 The Common Evidence Record#

Every consequential education claim should keep the same fields, so a favorable headline can't shed its comparison group, grade level, or uncertainty when it's repeated in a public document:

Field Required description
Population Grade, age if available, location, language, baseline achievement, eligibility, relevant access conditions
Intervention Product and version, student-facing or adult-facing role, instructional restrictions, curriculum, supervision, training
Comparator What the other group actually received, including teacher time and other technology
Exposure Scheduled time, realized use, duration, adherence, attrition
Assignment and analysis Randomization unit, numbers assigned, analytic sample, unique students versus repeated observations, clustering
Primary outcome Named assessment or operational measure, timing, whether assistance remained available
Effect and uncertainty Metric, estimate, interval or standard error when reported, significance, robustness qualifications
Durability and transfer Delayed tests, unassisted performance, unfamiliar problems, broader outcomes where measured
Cost and burden Software, devices, connectivity, implementation, training, supervision, verification, displaced activities
Independence and status Peer review, working-paper status, provider involvement, funding, independent review where available
Permitted conclusion The narrow statement actually supported, plus the most tempting unsupported extension

3.9.2 The Eight-Dimension School Evaluation Scorecard#

For future evaluations, I recommend measurement across eight dimensions:

  1. Unassisted learning, with delayed follow-up.
  2. Transfer to unrehearsed problems.
  3. Engagement, reported alongside nonuse.
  4. Adult work, including verification and error resolution.
  5. Access and outcomes by student group, with reasons for nonparticipation.
  6. Safety, privacy, and human recourse incidents.
  7. Total delivery cost per participating student, against the alternative.
  8. Displacement: what the intervention replaced.

A local benefit claim should depend on a comparison and outcome defined in advance, not on a retrospective pick of the most favorable dashboard metric. The reviewed studies show why: engagement increased without established literacy gains in one program, and assisted performance rose while later independent performance fell in another. (Literacy engagement evaluation, high-school mathematics experiment)

3.9.3 Why This Chapter Makes No Return-on-Investment Claim#

I don't convert teacher minutes, standardized test differences, or exit-ticket improvements into a single financial return. That conversion would require explicit assumptions about implementation costs, the value and actual use of saved time, the durability of effects, and the alternative use of school resources. And an inexpensive model interaction is not the full cost of educational delivery. Training, supervision, curriculum integration, devices, connectivity, verification, and program management all belong in the economic comparison before a district, or a company like ours, makes an affordability claim. Any document that quotes a per-student software price as the cost of the program has answered a question no one asked.

3.10 Questions to Ask Your School Board Before It Buys an AI Tutor#

Everything above boils down to a list you can bring to a public comment period or a board work session. Each question comes straight from the scorecard, the common evidence record, or the conditions in Section 3.8. None requires technical expertise to ask. All of them require a real answer.

About the evidence

  • Which study does the vendor cite, and did that study test the same product, grade, subject, and arrangement we're buying?
  • What did the comparison group in that study actually receive?
  • Is the cited result a peer-reviewed paper, a working paper, a preprint, or a vendor report, and has any independent reviewer rated it?
  • Were students tested with the tool available or without it?

About our own pilot

  • What outcome will we measure, and did we write it down before the pilot started?
  • Will students be assessed later without the tool, and will we compare them against students who didn't use it?
  • Will we report usage alongside nonuse, and report which students didn't participate and why?
  • Will we measure teacher time, including time spent checking what the tool produces?

About safeguards

  • Which adult is accountable for each classroom's use, and who can stop it?
  • What can students ask, what can the system answer, and when will it give a hint instead of a full solution?
  • What student data is collected, retained, shared, or used for training?
  • How does a student or parent challenge a wrong or harmful response?
  • Is this tool being used, or planned for use, in placement, discipline, diagnosis, or high-stakes testing?

About cost and trade-offs

  • What is the total delivery cost per participating student, including training, supervision, devices, connectivity, and verification, not just the license?
  • What are we giving up to make room for it: teacher small-group time, tutoring hours, something else?
  • What result would cause us to stop, and when will we decide?

3.11 Evidence Boundaries and Open Questions#

3.11.1 Grade-by-Grade Coverage Map#

The grade-band coverage deserves explicit mapping, because "covered" here means covered by the selected studies, not by all relevant research worldwide:

Segment Coverage in this review Boundary
Kindergarten Direct kindergarten reading trial; age-adjacent U.K. early mathematics trial Does not establish unrestricted conversational use for young children
Grades 1 and 2 Included in the literacy-platform support trials No statistically established incremental reading gain from added engagement support
Grades 3 and 4 Included in literacy support and Tutor CoPilot sessions Pooled results are not separate effects for each grade
Grade 5 Included in one literacy trial and CoPilot session data Less direct grade-specific coverage than the grouped label suggests
Grades 6 through 8 CoPilot grade 6 sessions; grades 6 through 8 remediation trial; grade 7 ASSISTments with grade 8 outcomes Different interventions, assessment timings, levels of certainty
Grade 9 or early secondary Algebra evidence concentrated in grade 9; Nigerian first-year senior secondary English evidence Foreign school stages are not automatically U.S. grade equivalents
Grades 10 through 12 High-school mathematics experiment and some broader high-school representation No separate verified sophomore, junior, or senior benefit estimate

3.11.2 Publication Status Stays Attached#

Publication status stays attached to each finding rather than being flattened into a single "research-backed" label. The kindergarten reading study is peer-reviewed. The early mathematics trial is a peer-reviewed 2019 article. The literacy engagement and Khanmigo studies are working papers. Tutor CoPilot is a preprint. ASSISTments carries both a research report and the April 2026 WWC review, with its reservations kept. The teacher preparation trial is an NFER/EEF randomized evaluation report. The PNAS mathematics study is peer-reviewed. The Nigerian evaluation is a World Bank working paper whose canonical abstract and authors' presentation were reviewed, not subjected to a claimed full-paper audit. And the algebra evaluation rests on RAND's publication record, summary, and 2014 addendum.

Peer review doesn't remove the need to examine design, sample loss, measurement, and generalization. A working paper isn't automatically uninformative. Status attaches, and it stays attached. For how this report grades evidence classes across every chapter, see Chapter 11.

3.11.3 What the Research Has Not Settled#

The evidence does not yet settle these questions:

  • Whether observed improvements persist across later grades and survive removal of the tool.
  • Whether routine assistance strengthens or weakens verification, reasoning, and recognizing uncertainty.
  • What sustained effects look like for writing, source evaluation, and original work, beyond this mathematics-heavy set of studies.
  • Which interventions improve college readiness, skilled work, and adult civic participation for older students.
  • Whether families experience measurable reductions in administrative burden or tutoring cost.
  • Who benefits, who disengages, and what design changes are needed for language, disability, and access differences.
  • Whether tested programs outperform a credible alternative use of the same money and adult time.
  • Whether effects persist after a model update, interface redesign, pricing change, or guardrail change.

These are limits on what this chapter can conclude. They are not reasons to disregard its positive findings.

3.12 The Boundary Between Learning Evidence and Infrastructure#

I'll be direct about who is writing this. This report is prepared by Savrn, which designs and builds AI factories and works on the infrastructure side of this industry, and our stated objective is a stronger relationship between that infrastructure and community well-being. I'm disclosing that commercial context rather than dressing it up as institutional independence.

Educational benefit and infrastructure acceptability are separate propositions. Even a strong learning result does not establish that a proposed facility has acceptable power demand, water use, public costs, land-use effects, or community commitments.

The seven Savrn trackers belong alongside this chapter as a separate evidence layer, not as supporting citations for learning gains. Their role is to inform infrastructure questions while education claims stay grounded in education research. A tracker entry starts a question and never ends one. The public-value chain I'd propose should be treated as a set of questions, not an assumed causal sequence:

  1. Does the educational intervention improve a defined outcome?
  2. Can schools deliver it reliably, safely, and affordably?
  3. What infrastructure is actually needed to provide that service?
  4. What costs, resource demands, and benefits arise from a particular facility?
  5. Can affected people verify commitments and obtain remedies?

That separation keeps this chapter useful to a parent, teacher, or civic committee that never buys a Savrn service. It also stops a legitimate education finding from being stretched into an argument for a specific development. Chapter 10 examines the infrastructure side on its own evidence.

3.13 Conclusion: What AI Can and Can't Do for Your Kid's Learning#

So, back to the kitchen table. Is the tool helping your kid learn, or doing the homework for them? The evidence says that depends almost entirely on how the tool is built and how it's used, and that you can't tell which from the homework itself. You can only tell by checking what your child can do without it.

The defensible position is neither blanket adoption nor blanket rejection. Educational decisions should rest on clearly identified learners, interventions, comparison conditions, outcomes, costs, and uncertainties, with the same standard of evidence applied to favorable and unfavorable results. The thesis that comes out of that is narrower and more useful than either blanket optimism or blanket rejection: educational value should be established by what learners can do independently and what educators can demonstrably deliver, under clearly described conditions.

The evidence here supports real optimism of a specific kind. Structured systems, adult-facing support, and supervised programs have produced measurable benefits for real students, and the PNAS experiment shows that the adverse outcome is not inevitable. It's a design failure with a known fix. What the evidence does not support is the idea that capability turns into learning by itself. Between the model and the child stands a chain of design decisions. This chapter has tried to show, study by study, exactly where that chain holds and exactly where, left unattended, it breaks.

3.14 Frequently Asked Questions#

Does AI help students learn, according to research?

Sometimes, under specific conditions. In this review of ten K-12 study programs, structured systems, tools that support teachers, and supervised programs produced measurable gains, such as 0.52 on a latent kindergarten reading measure in a teacher-led package. But in a 2025 PNAS experiment, students with an unrestricted GPT-4 assistant scored 17 percent worse than the control group once the tool was removed. Design and supervision decide the outcome.

Does ChatGPT help students learn or just do their homework?

The strongest evidence says it can do both, depending on design. In the PNAS study of nearly 1,000 high-school math students, an unrestricted GPT-4 assistant raised assisted practice performance 48 percent, then those students did 17 percent worse than controls on exams without it. A constrained tutor version largely avoided that penalty, though avoiding a penalty is not proof of a gain.

What did the Khanmigo study find?

An August 2026 working paper on Khan Academy with Khanmigo in grades 6 through 8 in Hamilton County, Tennessee found 0.020 standard deviations in year one (not significant), 0.084 in year two (significant at 5 percent), and 0.062 stacked (marginal at 10 percent). The treatment was a staffed remediation package, so it does not isolate the chatbot, and a small-cluster robustness test gave p = .0511.

Do AI tutoring studies show gains on end-of-year tests?

Not consistently. Tutor CoPilot, which gave human tutors AI suggestions, raised exit-ticket passing from about 62 percent to 66 percent over roughly two months, but found no statistically significant end-of-year math improvement. ASSISTments showed a 0.10 effect size on a grade 8 state test, yet the April 2026 What Works Clearinghouse review rated the study "with reservations" and found uncertain effects overall.

Does AI save teachers time?

For a defined task, one trial says yes. In an NFER/EEF randomized trial of 259 teachers in 68 English secondary schools, teachers using ChatGPT with a preparation guide reported about 56.2 minutes a week preparing specified Year 7 and Year 8 science lessons, versus 81.5 minutes for comparison teachers, 31 percent less. It did not measure total workload or student achievement.

Is AI safe for young children in school?

The research here does not test unrestricted chatbots with young children. The positive early-grade studies used teacher-directed or bounded tools, such as math apps used about 30 minutes a day for 12 weeks under staff supervision. UNICEF's December 2025 guidance calls for age-appropriate design, data protection, human oversight, and demonstrated efficacy before large-scale deployment, but it is policy guidance, not a learning study.

What does an effect size of 0.23 mean for my child?

It is a standardized score difference between groups, not 23 percent more learning. The 0.23 figure comes from a World Bank evaluation of a six-week supervised English program in Nigeria, with 759 students in the final analytic sample out of 1,328 randomized. It describes that program and population on average, not a guaranteed result for any individual child or for U.S. students.

What should parents ask a school before it adopts an AI tutor?

Ask what study supports the product and whether it tested the same grade and arrangement, what the comparison group received, and whether students will be tested later without the tool. Ask which adult is accountable, what data is collected, how errors are challenged, and the total cost beyond the license. In the studies reviewed, results depended on supervision, structure, and unaided measurement.

Keep listening

A copper studio microphone and headphones on a stack of books, with blueprint sound waves flowing into a sketched open book

The audio edition · Read by Bella

Listen to the whole report

13 episodes, 2 h 26 min. Each one walks a chapter's argument, its studies, and their limits, so you can listen instead of read. It plays straight through, one chapter into the next.

Up next · Chapter 1

AI Safety Promises and Public Trust: How to Test What AI Labs Say

0:00 / 11:57
  1. Read
  2. Read
  3. Read
  4. Read
  5. Read
  6. Read
  7. Read
  8. Read
  9. Read
  10. Read
  11. Read
  12. Read
  13. Read