# Chapter 3: Does AI Help Students Learn? What 10 K-12 Studies Actually Show

Part of The Superintelligence Transition by Chad Everett Harris, Founder and CEO, Savrn. Published September 24, 2026. Evidence cutoff September 23, 2026.

Canonical: https://savrn.com/blog/the-superintelligence-transition/ai-in-k12-education
Full edition: https://savrn.com/blog/the-superintelligence-transition

Disclosure: prepared for Savrn, which has a commercial interest in AI infrastructure and publishes the seven trackers cited here. Not an independent institutional review.

**Key takeaways**

- In a 2025 PNAS field experiment with nearly 1,000 high-school math students, an unrestricted GPT-4 assistant raised assisted practice performance 48 percent, but once access was removed those students scored 17 percent worse than the control group on later exams.
- A 2011 kindergarten reading trial of 556 students reported an effect size of 0.52 on a latent reading measure, but the tool worked behind the teacher, and the abstract reports no confidence interval.
- Tutor CoPilot lifted exit-ticket passing from about 62 percent to 66 percent over roughly two months, yet produced no statistically significant end-of-year math improvement, and the paper is still a preprint.
- The Khanmigo remediation package in Hamilton County, Tennessee showed 0.020 standard deviations in year one (not significant) and 0.084 in year two, and it does not isolate the chatbot from the rest of the staffed program.
- Teachers using ChatGPT with a preparation guide spent about 56.2 minutes a week preparing specified science lessons versus 81.5 minutes for comparison teachers, a 31 percent saving on that work only, with no student achievement measured.
- A World Bank evaluation in Nigeria found 0.23 standard deviations on English after a six-week supervised program, with 759 of 1,328 randomized students in the final analytic sample and larger gains for female and higher-achieving students.
It usually happens at the kitchen table. Your kid has a laptop open, an AI tool in one tab and the homework in another, and the answers are coming fast. Too fast. And you ask yourself the question every parent I know is asking right now: is this tool helping my kid learn, or is it doing the homework for them? That's the question this chapter takes to the AI in education research, and I want to answer it the way I'd answer any question about a system I was about to bet money on: by reading what the studies actually did, not what the brochures say they did.

Here's why it matters. Schools are buying. Districts are signing contracts, teachers are being handed new tools, and families are being told their children will get "personalized learning." Some of that is backed by real evidence. Some of it is not. And the difference between the two is often invisible unless you know exactly where to look: who was in the study, what the comparison group got, whether the kid was tested with the tool or without it, and whether anyone independent checked the work.

I reviewed ten study programs spanning kindergarten through high school. The good news is real. Structured systems, tools that support the teacher rather than replace the teacher, and supervised programs have produced measurable benefits for real students. The uncomfortable part is also real, and it's the part most people skip. In the best-designed high-school experiment in this set, an unrestricted chatbot made students look much better on practice and then left them worse off than students who never had it, once the tool was taken away. Same subject. Same kind of students. The guardrails made the difference.

So this chapter will not tell you AI is good for kids or bad for kids. It will tell you which arrangements have been tested, what they showed, where they stop, and what questions to ask before your school spends money or your child spends hours. Capability is not benefit. A powerful model sitting in a browser tab is not the same thing as a child who can now do the work on her own.

### 3.1 How to Read AI in Education Research Without Getting Fooled

Before we get to a single study, you need to know what kind of review this is and what rules I'm using to read the numbers. I'm spending time here on purpose. Most bad decisions about educational technology don't come from bad studies. They come from good studies read badly.

#### 3.1.1 Scope: Three Grade Bands, Ten Study Programs

This chapter reviews the evidence for assisted learning across the three school bands that organize most adoption decisions: kindergarten through grade 4, grades 5 through 8, and grades 9 through 12. It examines ten study programs. I prioritized randomized field evaluations, independently measured outcomes, original research reports, and, where one exists, an independent evidence review. Policy guidance shows up later in the chapter, but only as guidance on safeguards. It is never used as proof that a tool improves learning.

The evidence is international. The studies come from the United States, the United Kingdom, Nigeria, and Turkey-connected research settings, and I don't assume that a result transfers unchanged across curricula, languages, staffing arrangements, or school systems. A British "Year 7" is not automatically an American seventh grade. A Nigerian first-year senior secondary student is not automatically an American ninth grader. Where the source population matters to how you should read a result, I keep it attached to the result.

This is a targeted review of decision-relevant research. It is not a systematic review, and I'm not claiming to have located every relevant study in the world. I also don't calculate a pooled effect size, because the interventions, comparison conditions, ages, subjects, and outcomes differ too much for an average to mean anything. Averaging a kindergarten reading package with a high-school chatbot experiment would produce a number, and that number would be a rumor. Positive, null, uncertain, and adverse findings are all kept, with the same weight they carry in the original work.

#### 3.1.2 The Four Interpretation Rules

Four interpretation rules govern every number in this chapter. Learn them once and you'll read every education headline differently for the rest of your life.

- **Standard deviations are not percentages.** A result of 0.23 standard deviations is a standardized score difference, not 23 percent more learning.
- **Percentage points are not percent.** A pass rate rising from 62 percent to 66 percent is a four-percentage-point increase, not a 66 percent learning gain.
- **Relative percentages need their denominator.** A reported 17 percent reduction in an exam score is not a 17-percentage-point change unless the study defines it that way.
- **Assignment is not use.** The effect of offering a program differs from an estimate for people who actively use it. Selecting only the active users invites selection bias unless the analysis deals with it.

#### 3.1.3 The Bundle Rule

There's one structural rule on top of those four. Software, extra instruction, teacher training, devices, supervision, and curriculum changes usually arrive together as a bundle. A result for the package is not automatically a result for the model alone. Nearly every study in this chapter bundles something, and the bundle is part of what each result means.

> **Operator's note: Nobody buys a transformer by itself**
>
> When you build large power loads, you learn fast that nothing works as a single component. A transformer is useless without the switchgear, the protection scheme, the crew that maintains it, and the utility agreement behind it. If a site performs well, you don't credit the transformer. You credit the system, and you ask which parts of the system you'd have to copy to get the same result somewhere else.
> 
> School technology is the same. A study that shows gains from a tool delivered with training, coaching, extra time, and supervision is telling you the system worked. If your district buys the license and skips the rest, you have not bought what was tested. I'd tell a school board the same thing I'd tell an investor looking at a facility: price the whole system, not the part with the logo on it.
### 3.2 What AI in Education Research Shows Across Three Grade Bands

The evidence supports a conditional conclusion, and I'm asking you to hold the whole thing, not just the half you like: **specific, well-designed learning systems can improve particular educational outcomes, but access to a powerful model does not itself establish educational value.** The support for both halves of that sentence runs through the [kindergarten reading trial](https://eric.ed.gov/?id=EJ963694), the [ASSISTments technical report](https://files.eric.ed.gov/fulltext/ED630781.pdf), the [teacher preparation trial](https://files.eric.ed.gov/fulltext/ED680921.pdf), and the [high-school mathematics experiment published in PNAS](https://www.pnas.org/doi/10.1073/pnas.2422633122).

So the practical question for a school, a district, or a parent is not whether to embrace or reject a technology category. It's which combinations of students, teachers, instructional methods, tools, and safeguards produce independently measured benefits that justify their costs. That's a harder question. It's also the only one worth asking.

Here are the ten study programs, sorted by band:

| Grade band | Study programs | Direction of evidence |
|---|---|---|
| Kindergarten through grade 4 | Kindergarten reading (A2i package); early mathematics applications; literacy engagement support | Positive for teacher-directed and bounded packages; engagement without established literacy gains |
| Grades 5 through 8 | ASSISTments; Tutor CoPilot; Khanmigo remediation package; teacher preparation time | Small or qualified gains; session mastery without established annual attainment; bounded teacher-time savings |
| Grades 9 through 12 | PNAS mathematics assistants; Nigerian English program; Cognitive Tutor Algebra I | Adverse risk from unrestricted assistance; positive supervised program; historical adaptive-system evidence |

*Figure 3.1 (interactive on the page).*
The sections that follow take each band in turn, study by study, with the design, the finding, and the boundary attached to each. One framing note first. The schooling debate is usually staged as a contest between "innovation" and "caution." The evidence I reviewed supports neither banner. It supports a procurement discipline: name the learner, the intervention, the comparison, the outcome, and the cost, and then look at what the comparison actually showed. That's the whole game.

### 3.3 AI Tutoring Studies, Kindergarten Through Grade 4

The early-grade evidence is most useful when it tells you three things: what role the adult plays, what skill is being taught, and how limited the child's direct interaction with the tool is. The studies in this band test assessment-informed teaching, bounded mathematics applications, and a literacy platform with and without human engagement support. They do not test what happens when a young child is handed an unrestricted conversational assistant, and I won't pretend they do.

That gap matters. If you're the parent of a six-year-old and someone tells you "the research shows AI helps young children learn," the right response is: which AI, doing what, with which adult, measured how? In this band, the answers are always some version of "a structured tool, working through or alongside a teacher, measured outside the tool."

#### 3.3.1 Kindergarten Reading: The Tool Behind the Teacher

The strongest early-grade result in the reviewed set comes from a 2011 cluster-randomized study of Individualized Student Instruction for Kindergarten. The evaluation involved 556 students, 44 teachers, and 14 schools. Students in intervention classrooms outperformed comparison students on a latent measure of reading skills, with a reported effect size of 0.52. ([Al Otaiba and colleagues, original study record](https://eric.ed.gov/?id=EJ963694))

Here's the part most people skip: the intervention. It used assessment-informed recommendations associated with Assessment2Instruction (A2i), combined with professional development and classroom support, to help teachers individualize instruction. The system recommended instructional amounts and groupings. The teacher still organized instruction, interpreted student needs, and worked directly with children. Information flowed from student assessments, to recommendations for the teacher, and then to differentiated classroom instruction. This was not an open chatbot teaching a kindergartner on its own. ([Connor, research synthesis and intervention description](https://pmc.ncbi.nlm.nih.gov/articles/PMC5854500/))

Four boundaries attach to the 0.52 estimate, and I'd want every one of them in any school board presentation that cites it.

First, it concerns a latent reading construct, a statistically estimated measure of reading skill. It is not a universal improvement across every reading test or every early-grade student.

Second, the publicly available study abstract does not report a confidence interval for the estimate. Without an interval, you can't see how much uncertainty surrounds the 0.52.

Third, the result supports the tested instructional package including adult implementation, not a software-only effect. A school considering a similar approach should test the complete delivery package, including the time and training required for teachers to act on the recommendations.

Fourth, a disclosure belongs in the record: the later A2i synthesis discloses the author's equity interest in Learning Ovations, which matters when assessing the broader presentation of the intervention's evidence. ([Study abstract](https://eric.ed.gov/?id=EJ963694), [Disclosure in the synthesis](https://pmc.ncbi.nlm.nih.gov/articles/PMC5854500/))

A disclosed interest does not make a study wrong. It tells you where to look harder. I'd say the same about anybody's claims, including ours.

**Evidence card**

- Title: Individualized Student Instruction for Kindergarten (2011)
- Design: Cluster-randomized trial of assessment-informed instruction using A2i recommendations, with professional development and classroom support.
- Population: 556 kindergarten students, 44 teachers, 14 schools.
- Finding: Intervention students outperformed comparison students on a latent measure of reading skills, reported effect size 0.52.
- Limit: Latent reading construct only; no confidence interval in the public abstract; package includes adult implementation; later synthesis discloses the author's equity interest in Learning Ovations.
- Source: [Al Otaiba and colleagues, ERIC record](https://eric.ed.gov/?id=EJ963694)
The defensible practical reading is that intelligent support can operate behind the teacher rather than in place of the teacher. That's a useful finding. It's also narrower than the one usually advertised.

**What a parent should take from it.** If your child's kindergarten uses a system like this, the useful question isn't "does my child use the AI?" It's "how does the teacher use what the system tells her, and how do you check whether my child is reading better without it?"

**Common misreading.** "AI raised kindergarten reading scores by half a standard deviation." No. A teacher-led instructional package, informed by assessment-based recommendations and supported by training, was associated with that difference on a latent reading measure in this trial. Take the teacher out of that sentence and you've described a study nobody ran.

#### 3.3.2 Early Mathematics: Structured Apps, Not Open Conversation

A United Kingdom randomized trial published in 2019 evaluated interactive mathematics applications with children aged four and five in their first compulsory school year. The trial randomized 461 children across 12 schools, with 389 children from 11 schools available at posttest. This is age-adjacent evidence for the youngest part of the K-12 range, not an exact equivalent of a U.S. kindergarten. ([Outhwaite and colleagues, full article](https://www.imagineworldwide.org/wp-content/uploads/2019/08/3.8-Raising-early-achievement-in-math-with-interactive-apps-A-randomized-control-trial.pdf))

Here's what the study actually did. Children used the applications for approximately 30 minutes per day over 12 weeks, supervised by classroom staff. Three groups were compared. One group received app-based mathematics in addition to its normal mathematics activities. A second group used the apps in place of a daily small-group mathematics activity. The control group continued normal teaching.

| Comparison with normal teaching | Reported result | Interpretation |
|---|---|---|
| Additional app-based mathematics time | Progress effect size 0.31; 95% confidence interval 0.06 to 0.55 | A positive result for additional structured mathematics practice; additional mathematics exposure is part of the treatment |
| App use replacing a daily small-group activity | Progress effect size 0.21; 95% confidence interval -0.03 to 0.46; significance from a one-tailed test | More tentative: the two-sided interval includes zero |
| Supplementary versus time-equivalent groups | No statistically significant difference | Not proof that the two approaches are equivalent |

*Figure 3.7 (interactive on the page).*
The supplementary result is clearly positive, but those children got their normal math plus the apps, so some of that gain may simply be more math. The replacement result is the more interesting policy number and the less certain one: its interval crosses zero. And "no statistically significant difference" between the app groups means the study couldn't tell them apart, not that they are equivalent.

The applications supplied bounded tasks, visual and spoken instructions, immediate feedback, repeated practice, and progression requirements. Children were not relying on unrestricted generated explanations, and the assessment was administered outside the app. ([Intervention and assessment methods](https://www.imagineworldwide.org/wp-content/uploads/2019/08/3.8-Raising-early-achievement-in-math-with-interactive-apps-A-randomized-control-trial.pdf))

**Evidence card**

- Title: Interactive Mathematics Apps in the First School Year (2019)
- Design: Randomized trial with three arms: apps added to normal math, apps replacing a daily small-group activity, and normal teaching; about 30 minutes per day for 12 weeks, supervised by classroom staff.
- Population: 461 children aged four and five across 12 United Kingdom schools randomized; 389 children from 11 schools at posttest.
- Finding: Supplementary app time, progress effect size 0.31 (95% CI 0.06 to 0.55); replacement app time, 0.21 (95% CI -0.03 to 0.46, one-tailed significance); no significant difference between app groups.
- Limit: Age-adjacent to U.S. kindergarten, not equivalent; supplementary gain includes extra math exposure; replacement interval includes zero; bounded apps, not conversational AI.
- Source: [Outhwaite and colleagues, full article](https://www.imagineworldwide.org/wp-content/uploads/2019/08/3.8-Raising-early-achievement-in-math-with-interactive-apps-A-randomized-control-trial.pdf)
This evidence supports a specific proposition: structured digital practice can contribute to early mathematics learning under tested classroom conditions. It does not establish that replacing play, teacher interaction, or broad early-childhood experiences with more screen time improves overall development. The distance between those two propositions is the distance between what the trial measured and what a vendor might hope it implies.

**What a teacher should take from it.** If you're deciding whether to use math apps in a reception or kindergarten room, the trial suggests two things. Bounded, feedback-rich practice under adult supervision can help. And the evidence is firmer when app time is added than when it replaces your own small-group work, so be careful about what you give up to make room for it.

#### 3.3.3 Literacy Engagement: More Minutes, No Proven Reading Gain

A June 2026 working paper examined two randomized trials involving 355 students: one in an after-school setting covering grades 1 through 5, and one in an in-school setting covering grades 1 through 3. Both treatment and comparison students had access to the same literacy platform. The treatment added human support intended to encourage engagement, not to provide reading instruction. ([Robinson and colleagues, full working paper](https://edworkingpapers.com/sites/default/files/ai26-1451.pdf))

The support differed between settings. The in-school trial used middle-school peer support, which means the support people can't all be assumed to be adults. And the platform can't be assumed to be a general-purpose language model. ([Study design](https://edworkingpapers.com/sites/default/files/ai26-1451.pdf))

| Measure | After-school, grades 1 through 5 | In-school, grades 1 through 3 |
|---|---|---|
| Randomized sample | 174 students | 181 students |
| Comparison-group mean platform minutes per week, including nonuse | 2.18 minutes | 5.23 minutes |
| Added weekly platform time from human support | About 1.00 minute; marginal at the 10% level | About 4.42 minutes; significant at the 5% level |
| End-of-year literacy result | No statistically significant improvement | No statistically significant improvement |

Because both groups had the platform, this is not a randomized comparison of the platform against no platform. What it shows is more awkward and more valuable: the added human engagement component increased some usage measures without producing a statistically established reading gain in either trial. ([Full working paper](https://edworkingpapers.com/sites/default/files/ai26-1451.pdf))

**Evidence card**

- Title: Human Engagement Support for a Literacy Platform (2026)
- Design: Two randomized trials; both arms had the same literacy platform; treatment added human support to encourage engagement, not to teach reading.
- Population: 355 students total: 174 in an after-school trial (grades 1 through 5) and 181 in an in-school trial (grades 1 through 3).
- Finding: Added about 1.00 minute per week after school (marginal at 10%) and about 4.42 minutes per week in school (significant at 5%); no statistically significant end-of-year literacy improvement in either trial.
- Limit: Working paper; tests added engagement support, not the platform against no platform; in-school support came from middle-school peers; platform not assumed to be a general-purpose language model.
- Source: [Robinson and colleagues, EdWorkingPapers](https://edworkingpapers.com/sites/default/files/ai26-1451.pdf)
For decision-making, this is a warning against counting licenses, logins, or available tutoring hours as learning outcomes. A school should distinguish four different numbers that routinely get compressed into one dashboard metric: scheduled time, actual engaged time, completed practice, and independently measured reading growth.

Read that again. Scheduled time is what the district planned. Engaged time is what kids actually did. Completed practice is what the software logged. Reading growth is what you were paying for. A vendor report that shows the first three and not the fourth has shown you activity, not results.

#### 3.3.4 Grades 3 and 4: Tutor Support Across the Band Line

Tutor CoPilot, a generative assistant used by human mathematics tutors, included substantial grade 3 and grade 4 representation in its session-level analysis. Because the study crosses the band boundary, I treat it in full in Section 3.4.2. One rule applies here: its pooled effect must not be recast as a separate verified effect for third graders or fourth graders. Pooled results are pooled. I won't manufacture grade-specific precision the study did not produce. ([Wang and colleagues, full study](https://arxiv.org/pdf/2410.03017))

#### 3.3.5 What Schools Can Responsibly Tell Families of Young Children

For this age band, the most defensible explanation a school can give families is that tools may help a teacher identify what a child needs next, or provide a bounded opportunity to practice. The positive studies I reviewed involve specified instructional designs and adult implementation, not simply giving a young child a general-purpose assistant. ([A2i synthesis](https://pmc.ncbi.nlm.nih.gov/articles/PMC5854500/), [early mathematics trial](https://www.imagineworldwide.org/wp-content/uploads/2019/08/3.8-Raising-early-achievement-in-math-with-interactive-apps-A-randomized-control-trial.pdf))

Here's an illustration, clearly labeled as an illustration and not an additional measured result. A teacher reviews assessment information, selects a short activity, observes the child, and checks the same skill later without the tool. A parent receives an explanation of the learning goal and can ask whether the child is improving independently, rather than being asked to accept an opaque "personalization" claim.

Claims that these systems reduce parents' stress, save families money, or replace home reading support would require direct household evidence. Those outcomes were not the causal endpoints of any study in this section, and I won't borrow them from the household chapter, which reaches its own, more conditional conclusions in [Chapter 6](#ai-for-families-households).

**Checklist for parents of children in kindergarten through grade 4:**

- Does my child interact with the tool directly, or does the teacher use it to plan?
- What exactly can my child ask it, and what can it answer?
- Which adult is watching while my child uses it?
- How does the teacher check whether my child can do the skill without the tool?
- What is my child not doing (play, reading aloud, small-group time) during the minutes the tool is in use?

### 3.4 AI Homework Help and Tutoring Research, Grades 5 Through 8

The middle-grade evidence covers three distinguishable arrangements: structured feedback on mathematics practice, assistance delivered through a human tutor, and a conversational tool embedded in a staffed remediation program. Add a fourth study about teacher preparation time, and you have the most varied band in this chapter. The findings are informative exactly as long as those arrangements stay separate. ([ASSISTments report](https://files.eric.ed.gov/fulltext/ED630781.pdf), [Tutor CoPilot paper](https://arxiv.org/pdf/2410.03017), [Khanmigo school experiment](https://edworkingpapers.com/sites/default/files/ai26-1551.pdf))

#### 3.4.1 ASSISTments: Why Independent Review Matters

A North Carolina trial assigned 63 schools to ASSISTments or to business-as-usual mathematics instruction, involving 102 grade 7 mathematics classrooms across 41 districts. The platform provided feedback and hints to students and reports to teachers, with professional development and coaching supporting implementation. ([Feng, Huang, and Collins, technical report](https://files.eric.ed.gov/fulltext/ED630781.pdf))

Focal students used the intervention in grade 7 during the 2019 to 2020 school year. The longer-term outcome was the grade 8 state mathematics assessment in spring 2021. The grade 8 outcome analysis included 5,991 students and reported an effect size of 0.10, with a p-value of .011. ([Technical report](https://files.eric.ed.gov/fulltext/ED630781.pdf))

That's useful evidence of a modest later assessment difference. It is not a claim of a 10 percent increase in learning. (Rule one: standard deviations are not percentages.) The study also ran through pandemic disruption, and students taking high-school-level mathematics instead of the relevant grade 8 test were excluded from that test's analytic sample. That exclusion matters for reading the result, because it removed a group of students from the comparison. ([Technical report](https://files.eric.ed.gov/fulltext/ED630781.pdf), [WWC sample description](https://ies.ed.gov/ncee/WWC/Study/94267))

Now the independent review, which requires particular care. The April 2026 What Works Clearinghouse review rates the study **"Meets WWC standards with reservations,"** citing high cluster-level attrition with baseline equivalence in the analytic groups. It identifies uncertain effects for the mathematics achievement domain, with a positive grade 8 state-test result and a nonsignificant mathematics-readiness result in a smaller sample. An earlier review displayed on the same page was more favorable, so a citation to the earlier rating alone would leave out material later evaluation. ([Latest WWC review](https://ies.ed.gov/ncee/WWC/Study/94267), [Review history and outcomes](https://ies.ed.gov/ncee/WWC/Study/94267))

*Figure 3.3 (interactive on the page).*
**Evidence card**

- Title: ASSISTments in North Carolina (WWC review, April 2026)
- Design: School-level randomized trial of a feedback and hint platform with teacher reports, professional development, and coaching, against business-as-usual math instruction.
- Population: 63 schools, 102 grade 7 math classrooms, 41 districts; 5,991 students in the grade 8 outcome analysis.
- Finding: Grade 8 state math assessment effect size 0.10, p = .011; WWC rates the study "Meets WWC standards with reservations" with uncertain effects in the math achievement domain.
- Limit: High cluster-level attrition; pandemic disruption; students taking high-school-level math excluded from the grade 8 test sample; nonsignificant math-readiness result in a smaller sample; earlier, more favorable WWC rating superseded.
- Source: [Feng, Huang, and Collins, technical report](https://files.eric.ed.gov/fulltext/ED630781.pdf)
The defensible conclusion: the state-test result is positive, but the broader independent assessment is qualified, not uniformly confirmatory. Faster identification of errors and better-targeted follow-up are plausible mechanisms consistent with the intervention. They are not separately isolated causal effects. And the result belongs to the ASSISTments instructional package and its tested comparison conditions. It is not evidence for unrestricted conversational tutoring.

This study is my standing example of why independent review matters. The trial team's report and the federal reviewer's later assessment of the same study produce different confidence levels, and both belong in the record. A district that reads only the first gets an inflated estimate. A district that reads only the second might throw out a promising program. The discipline is to hold both, which means holding a smaller and more accurate conclusion than either source alone suggests.

#### 3.4.2 Tutor CoPilot: Helping the Human Tutor

Tutor CoPilot randomized 900 tutors to access or no access to an adult-facing assistant that suggested ways to respond to students during live online mathematics tutoring. The study identified 1,787 participating students in nine schools. The analyzed session sample contained 4,136 sessions taught by the remaining participating tutors. ([Full study, January 2025 version](https://arxiv.org/pdf/2410.03017))

Here's the design choice that makes this study worth your attention. The assistant could suggest a question or an explanation, and the tutor decided what to use. The child still talked to a human tutor. What changed was the support available to that tutor. That's a very different arrangement from putting a chatbot in front of a student.

| Outcome or boundary | Finding |
|---|---|
| Main randomized assignment result | Exit-ticket passing rose from approximately 62% to 66%, a four-percentage-point increase; p < .01 |
| Lower-rated tutor subgroup | A nine-percentage-point improvement; a subgroup result, not the overall effect |
| End-of-year mathematics assessment | No statistically significant improvement |
| Exposure period | Approximately two months |
| Coverage limitation | Grade 5 and grade 6 sessions support relevance to this band; the final session sample does not establish effects for grades 7 and 8 |
| Publication status | Preprint; not treated here as a peer-reviewed journal finding |

The result supports near-term lesson mastery under this tutor-assisted arrangement. It does not establish durable improvement in a student's overall mathematics attainment. Developer and provider involvement disclosed in the paper stays visible in my assessment. For a district, the next questions are whether the lesson-level improvement persists on independent assessments, and whether it's still worth it after training, supervision, and full delivery costs are included. ([Outcomes and disclosures](https://arxiv.org/pdf/2410.03017))

**Evidence card**

- Title: Tutor CoPilot (2025 preprint)
- Design: Randomized assignment of 900 tutors to access or no access to a generative assistant that suggested responses during live online math tutoring; tutors chose what to use.
- Population: 1,787 students in nine schools; 4,136 analyzed sessions; grade 5 and grade 6 sessions relevant to this band, with substantial grade 3 and grade 4 representation.
- Finding: Exit-ticket passing rose from about 62% to 66%, four percentage points, p < .01; lower-rated tutors improved nine percentage points (subgroup).
- Limit: No statistically significant end-of-year math improvement; about two months of exposure; no established effect for grades 7 and 8; preprint; developer and provider involvement disclosed.
- Source: [Wang and colleagues, arXiv](https://arxiv.org/pdf/2410.03017)
The four-percentage-point headline and the null end-of-year result must be held together. They describe different links of the benefit chain from [Chapter 2](#ai-productivity-evidence). The first is task performance within sessions. The second is the real-world outcome a parent actually cares about. An adoption decision made on the first number alone is exactly the kind of decision this report exists to prevent.

Notice also the subgroup finding. Lower-rated tutors improved by nine percentage points. That's intriguing, because it suggests the tool may do the most good where tutoring quality is weakest. But it's a subgroup result, not the overall effect, and subgroup results are where hopeful readers go to find the number they wanted. Treat it as a question for the next study, not an answer for this one.

#### 3.4.3 Khanmigo Study Results in a Staffed Remediation Package

An August 2026 working paper examined Khan Academy with Khanmigo in grades 6 through 8 across 18 middle schools in Hamilton County, Tennessee, over two school years. Randomization involved 53 grade-within-school clusters, and the analysis included 6,902 student-term observations. That's not 6,902 distinct children. The same student can show up more than once across terms. ([Oreopoulos and Low, full working paper](https://edworkingpapers.com/sites/default/files/ai26-1551.pdf))

The treatment combined a mathematics practice platform, conversational support, assigned learning paths, and scheduled remediation periods staffed by adults. The comparison was existing remediation, which itself could include teacher-led work and other digital learning products. The experiment therefore does not isolate the incremental effect of Khanmigo over otherwise identical Khan Academy use. There's a bundle on both sides of the comparison, which limits what either side can claim. ([Design and comparison conditions](https://edworkingpapers.com/sites/default/files/ai26-1551.pdf))

The paper's pooled MAP result was an increase of approximately 1.26 national percentile-rank points per term, with a conventional standard error of 0.60. A small-cluster robustness test gave p = .0511, which makes the certainty of the main pooled result sensitive to the inference method. ([Primary results](https://edworkingpapers.com/sites/default/files/ai26-1551.pdf), [Robustness analysis](https://edworkingpapers.com/sites/default/files/ai26-1551.pdf))

That robustness detail is easy to skim past, so let me slow down on it. With only 53 clusters, the way you calculate uncertainty matters. Under the conventional method, the pooled result looks solid. Under a method built for a small number of clusters, it sits right at the edge of the usual significance threshold. Neither calculation is a trick. Together they tell you the result is real enough to take seriously and fragile enough not to oversell.

| Annual standardized estimate | Reported result | Appropriate reading |
|---|---|---|
| First year | 0.020 standard deviations, SE 0.035; not statistically significant | No demonstrated first-year annual gain |
| Second year | 0.084 standard deviations, SE 0.041; significant at the conventional 5% level | A modest positive second-year package result |
| Stacked annual estimate | 0.062 standard deviations, SE 0.035; marginal at the 10% level | Promising but less certain than a simple "proven annual gain" statement |

*Figure 3.4 (interactive on the page).*
The difference between the first-year and second-year annual estimates was not itself statistically significant. So one significant year and one nonsignificant year must not be described as a proven improvement between years. The report also notes that the preregistered state-test co-primary outcome had not yet been incorporated into the analysis. And its usage findings show that having a conversational tool available did not automatically translate into sustained explanatory dialogue. ([Annual results](https://edworkingpapers.com/sites/default/files/ai26-1551.pdf), [Outcome availability and usage analysis](https://edworkingpapers.com/sites/default/files/ai26-1551.pdf))

**Evidence card**

- Title: Khan Academy With Khanmigo in Hamilton County Remediation (2026)
- Design: Cluster-randomized experiment over two school years; treatment combined math practice platform, conversational support, assigned learning paths, and staffed remediation periods; comparison was existing remediation.
- Population: Grades 6 through 8 in 18 middle schools, Hamilton County, Tennessee; 53 grade-within-school clusters; 6,902 student-term observations (not distinct children).
- Finding: Pooled MAP gain of about 1.26 national percentile-rank points per term (SE 0.60); annual estimates 0.020 SD in year one (not significant), 0.084 SD in year two (significant at 5%), 0.062 SD stacked (marginal at 10%).
- Limit: Working paper; small-cluster robustness test p = .0511; does not isolate Khanmigo from Khan Academy or the staffed program; year-to-year difference not significant; preregistered state-test co-primary outcome not yet incorporated.
- Source: [Oreopoulos and Low, EdWorkingPapers](https://edworkingpapers.com/sites/default/files/ai26-1551.pdf)
The appropriate conclusion is that this staffed remediation package produced modest, qualified evidence of benefit. It is not yet a clean demonstration that adding a chatbot alone improves outcomes for middle-school students.

#### 3.4.4 Teacher Preparation Time: A Different Kind of Outcome

A randomized trial involving 259 teachers in 68 English secondary schools tested ChatGPT with a preparation guide for Year 7 and Year 8 science lessons. The population is relevant to the middle-grade discussion, but it keeps its original label: United Kingdom school-year labels are not automatically U.S. grades. ([NFER/EEF report](https://files.eric.ed.gov/fulltext/ED680921.pdf))

In the measured later weeks of the trial, intervention teachers reported approximately 56.2 minutes per week preparing the specified lessons, compared with 81.5 minutes in the comparison group. That's roughly 25.3 minutes, or 31 percent, less preparation time for the measured work, with a reported time-ratio confidence interval of 0.53 to 0.90. ([Primary results](https://files.eric.ed.gov/fulltext/ED680921.pdf))

*Figure 3.6 (interactive on the page).*
The outcome came from teacher diaries, and the primary analysis included 211 teachers rather than every randomized participant. A blind review of 30 submitted lessons detected no quality difference, but a 30-lesson sample does not prove universal quality equivalence. The study did not measure student achievement. ([Methods, attrition, and quality assessment](https://files.eric.ed.gov/fulltext/ED680921.pdf))

**Evidence card**

- Title: ChatGPT for Science Lesson Preparation, NFER/EEF Trial
- Design: Randomized trial of ChatGPT with a preparation guide for specified science lessons; outcome from teacher diaries in the measured later weeks.
- Population: 259 teachers in 68 English secondary schools, Year 7 and Year 8 science; 211 teachers in the primary analysis.
- Finding: About 56.2 minutes per week preparing specified lessons versus 81.5 minutes for comparison teachers, roughly 25.3 minutes or 31 percent less; time-ratio confidence interval 0.53 to 0.90.
- Limit: Self-reported diaries; not all randomized teachers analyzed; blind quality review of only 30 lessons; student achievement not measured; specified lessons only, not total workload.
- Source: [NFER/EEF report via ERIC](https://files.eric.ed.gov/fulltext/ED680921.pdf)
This establishes a bounded operational benefit worth testing locally: 31 percent less time preparing *specified science lessons under a preparation guide*. It is not a 31 percent reduction in teachers' total workload, and it's not a measured improvement in anything students learned. Whether the saved minutes went to individualized teaching is a question the trial did not ask, and I won't answer it on the trial's behalf.

**What a teacher should take from it.** If your school offers AI tools for lesson prep, this is the best evidence in the set that it can save time on a defined task with a guide. Ask for the guide, not just the login. And keep an eye on verification time: checking generated material is real work, and the trial's quality review was small.

#### 3.4.5 What the Middle-Grade Evidence Supports

The evidence supports trials of bounded mathematics feedback, adult-mediated tutoring support, and carefully evaluated remediation packages, with separate measurement of student attainment and teacher time. It does not support treating all conversational access as equivalent, or stretching a positive session-level result into a full-year learning guarantee. ([ASSISTments review](https://ies.ed.gov/ncee/WWC/Study/94267), [Tutor CoPilot](https://arxiv.org/pdf/2410.03017), [Khanmigo evaluation](https://edworkingpapers.com/sites/default/files/ai26-1551.pdf))

Here's an illustrative workflow, proposed as a design principle for evaluation rather than claimed as a proven effect: students attempt a problem, receive a limited hint, explain their reasoning, and later solve a related problem without assistance. Every clause in that sequence is doing evaluative work. The last clause is the one most existing programs leave out.

**For parents of middle schoolers dealing with AI homework help at home**, the research doesn't hand you a rule, but the pattern across these studies points to a practical habit. The studies that showed gains used tools that gave hints and feedback, kept an adult involved, and measured the student without help. You can copy that structure at the kitchen table: let the tool give a hint, have your child explain the step back to you, and then have them try a similar problem with the laptop closed. That's not a tested intervention. It's the evaluation logic of the better studies, applied at home.

### 3.5 Does ChatGPT Help Students Learn? High-School Evidence

High school is where the difference between producing an answer and acquiring a skill becomes impossible to avoid. A teenager with a capable chatbot can produce a lot of correct-looking work. Whether she can do that work herself next month is a separate question, and in this band the evidence answers it directly. The reviewed studies include an adverse effect from unrestricted assistance, a promising supervised program, and an older adaptive intervention whose effects differed by implementation year. ([PNAS mathematics experiment](https://www.pnas.org/doi/10.1073/pnas.2422633122), [World Bank English evaluation](https://openknowledge.worldbank.org/entities/publication/15e1ff08-15ae-4f7a-b2a8-d146e6c113ee), [RAND algebra evaluation](https://www.rand.org/pubs/research_briefs/RB9746.html))

#### 3.5.1 The PNAS Experiment: Better Practice, Weaker Learning

A peer-reviewed 2025 study in *PNAS* ran a field experiment with nearly 1,000 high-school students in mathematics. It compared three conditions: a GPT-4-based general assistant, a more constrained tutoring version, and a control condition. The tutoring version incorporated teacher-provided instructional material and restrictions intended to guide students rather than simply hand them answers. ([Bastani and colleagues, original paper](https://www.pnas.org/doi/10.1073/pnas.2422633122))

Now here's what happened. The unrestricted assistant increased assisted practice performance by 48 percent relative to the control group. The tutoring version increased it by 127 percent. Then access was removed for subsequent exams, and the picture flipped. Students in the unrestricted-assistant condition performed 17 percent worse than the control group, while the constrained tutor largely mitigated that adverse effect. ([Reported experimental results](https://www.pnas.org/doi/10.1073/pnas.2422633122))

Read that again. The students with the open chatbot looked better on practice and then did worse on the exam than students who never had it.

*Figure 3.2 (interactive on the page).*
These percentages describe the study's measured performance. They must not be rewritten as percentage-point changes, and they must not be generalized to every subject. (Rule three: relative percentages need their denominator.) The tutoring result also must not be turned into a claim of a demonstrated positive independent-learning effect just because it avoided the unrestricted tool's penalty. Avoiding a penalty is not the same as producing a gain, and I decline to launder one into the other.

**Evidence card**

- Title: GPT-4 Assistants in High-School Mathematics, PNAS (2025)
- Design: Peer-reviewed field experiment with three conditions: a GPT-4-based general assistant, a constrained tutoring version using teacher-provided material and guidance restrictions, and a control; assistance removed for subsequent exams.
- Population: Nearly 1,000 high-school mathematics students.
- Finding: Assisted practice performance rose 48% (unrestricted) and 127% (tutor) relative to control; on subsequent unaided exams, the unrestricted group performed 17% worse than control, while the tutor version largely mitigated that adverse effect.
- Limit: Relative percentages, not percentage points; mathematics only; avoiding the penalty is not evidence of a positive independent-learning gain for the tutor version; no separately verified effect for each high-school year.
- Source: [Bastani and colleagues, PNAS](https://www.pnas.org/doi/10.1073/pnas.2422633122)
What the experiment demonstrates, with unusual clarity, is the third link of the benefit chain failing in the wild. Both assistants made practice look better. Only one of them left students able to perform after the assistance was gone. The design implication follows directly: completed assignments can't be read as learning, and any deployment that doesn't include an unaided measurement is flying blind on the outcome that matters most.

This is not an argument against educational assistance. It's evidence that the structure of the assistance is part of the intervention and has to be evaluated. The same kind of tool, used by the same kind of students, in the same subject, produced learning protection in one form and learning loss in the other, depending on the guardrails around it.

This is the study that answers the kitchen-table question most directly. If the tool is doing the work, your child's practice will look great and her independent performance may suffer. If the tool is built to guide instead of answer, that penalty can largely go away. The only way to know which one you've got is to test your child without it.

> **Operator's note: Pull the backup and see what holds**
>
> In a large power facility, you don't learn anything about resilience while every backup system is running. You learn it when you load-test with the backup removed. If the load holds, you've got a real system. If it drops, you've got a system that only looked healthy because something else was carrying it.
> 
> The unaided test in education works the same way. Assisted practice scores are the facility running with the backup online. The exam without the tool is the load test with the backup pulled. The PNAS students with the open chatbot looked great until the backup came out, and then they dropped below the students who'd never had one. I would never sign off on a facility without that test. I don't think a school should sign off on a learning tool without it either.
#### 3.5.2 The Unaided-Assessment Rule for Any School Deployment

The PNAS design generalizes into a simple rule, and it's the single most useful thing a school can take from this chapter. Any deployment should have three stages: assisted practice, removal of assistance, and an independent measurement of what the student can now do alone.

*Figure 3.5 (interactive on the page).*
Here's a hypothetical to make it concrete. A high school wants to pilot an AI writing assistant in tenth-grade English. Under the unaided-assessment rule, the pilot plan would say, before anything starts: students will use the tool for drafting practice during the unit; at the end of the unit, they will write an in-class essay without the tool; that essay will be scored by teachers who don't know which students used the tool; and the comparison will be against classes that didn't use it. If the school can't commit to that last essay, it isn't running an evaluation. It's running a subscription.

Nothing in that illustration is a measured result. It's the evaluation structure the PNAS study makes impossible to ignore.

#### 3.5.3 Nigerian Senior Secondary English: Promising Supervised Use

A World Bank working paper evaluated a six-week program using Microsoft Copilot, powered by GPT-4, for first-year senior secondary students learning English in Nigeria. The reported effect was 0.23 standard deviations on English and 0.31 standard deviations on a composite assessment combining English, knowledge of the technology, and digital skills. ([World Bank publication record and abstract](https://openknowledge.worldbank.org/entities/publication/15e1ff08-15ae-4f7a-b2a8-d146e6c113ee))

The delivery details matter a lot here. The program ran in nine public schools in Benin City, with 1,328 randomized students and 759 in the final analytic sample presented by the authors. It comprised twelve 90-minute after-school sessions over six weeks, with students working in pairs, curriculum-oriented prompts, and teacher guidance and monitoring. So the intervention combined a model with additional structured learning time, access arrangements, prompts, peer interaction, and supervision. The result is evidence for that package, not a model-only effect holding all other instructional resources constant. ([Authors' methodological presentation](https://voxdev.org/sites/default/files/2026-03/From%20Chalkboards%20to%20Chatbots.pdf))

The drop from the randomized sample to the analyzed sample is a material limitation. The authors report attrition-related robustness analyses rather than ignoring the problem. Those analyses strengthen the interpretation, but they don't make missing outcome data irrelevant. ([Sample flow and robustness presentation](https://voxdev.org/sites/default/files/2026-03/From%20Chalkboards%20to%20Chatbots.pdf))

Four claims, four treatments:

| Claim | Evidence-based treatment |
|---|---|
| English learning improved in the evaluated program | Supported by the reported 0.23-standard-deviation English result |
| The composite result was an English-only effect | Incorrect: the 0.31 estimate combined English with technology knowledge and digital skills |
| Students completed years of schooling in six weeks | Not established; schooling-equivalent language is a benchmark conversion, not observed completion of additional school years |
| The same effect holds for American seniors or every high-school subject | Not established by this population and subject |

The World Bank evaluation also reports heterogeneous effects, including larger gains for female students and for students with higher baseline academic performance. A positive average therefore shouldn't automatically be presented as evidence that the intervention closes achievement gaps. If anything, the reported pattern for baseline performance points the other way. ([World Bank study abstract](https://openknowledge.worldbank.org/entities/publication/15e1ff08-15ae-4f7a-b2a8-d146e6c113ee))

**Evidence card**

- Title: Microsoft Copilot for English in Benin City, Nigeria (World Bank working paper)
- Design: Randomized evaluation of a six-week after-school program: twelve 90-minute sessions, students in pairs, curriculum-oriented prompts, teacher guidance and monitoring, using Microsoft Copilot powered by GPT-4.
- Population: First-year senior secondary students in nine public schools in Benin City; 1,328 randomized, 759 in the final analytic sample.
- Finding: 0.23 standard deviations on English; 0.31 standard deviations on a composite of English, technology knowledge, and digital skills; larger gains for female students and higher baseline performers.
- Limit: Substantial sample loss between randomization and analysis; package includes extra time, prompts, peers, and supervision; composite is not English only; schooling-equivalent language is a benchmark conversion; not established for U.S. students or other subjects.
- Source: [World Bank publication record and abstract](https://openknowledge.worldbank.org/entities/publication/15e1ff08-15ae-4f7a-b2a8-d146e6c113ee)
This is promising evidence for supervised use in a specified secondary-school setting, not a universal promise. The canonical World Bank abstract supplies the principal effect estimates. The authors' presentation supplies sample-flow and implementation detail, and I'm not presenting my reading of it as a complete audit of the full working paper.

#### 3.5.4 Cognitive Tutor Algebra I: The Version Lesson

The Cognitive Tutor Algebra I evaluation studied a blended curriculum combining classroom instruction and software across 147 schools, including 73 high schools and 74 middle schools. The high-school sample was predominantly ninth graders, and the intervention was an adaptive subject-specific system, not a modern general-purpose conversational model. ([RAND study summary](https://www.rand.org/pubs/research_briefs/RB9746.html), [journal-publication record](https://www.rand.org/pubs/external_publications/EP50410.html))

The evaluation found no statistically significant first-year benefit and a positive second-year high-school result. RAND's later addendum reported a significant second-year high-school effect of 0.21 standard deviations, substantively consistent with the earlier analysis. ([Study summary](https://www.rand.org/pubs/research_briefs/RB9746.html), [2014 analytical addendum](https://www.rand.org/pubs/working_papers/WR1050.html))

Two cautions are structural. First, the two implementation years involved different student cohorts, not two years of exposure for the same students. So the findings shouldn't be described as proof that an individual learner must use the product for two years to benefit, or as proof that teacher experience caused the difference. Second, the system's era matters. A bounded subject system producing a measurable benefit at scale does not establish that a more general or newer model will outperform it.

**Evidence card**

- Title: Cognitive Tutor Algebra I (RAND, with 2014 addendum)
- Design: Large evaluation of a blended algebra curriculum combining classroom instruction with adaptive, subject-specific software, across two implementation years.
- Population: 147 schools, including 73 high schools and 74 middle schools; high-school sample predominantly ninth graders.
- Finding: No statistically significant first-year benefit; positive second-year high-school result; 2014 addendum reports a significant second-year high-school effect of 0.21 standard deviations.
- Limit: Different student cohorts in each year, not two years of exposure for the same students; older adaptive system, not a general-purpose conversational model; does not establish that newer models would do better.
- Source: [RAND study summary](https://www.rand.org/pubs/research_briefs/RB9746.html)
The relevance here is historical and methodological. This study shows both that bounded systems can work at scale and that results can differ by implementation year. 

#### 3.5.5 Four High-School Years Are Not One Proven Population

The high-school band includes four years, but the evidence base doesn't justify four grade-specific benefit estimates. The algebra evidence is concentrated in grade 9. The Nigerian study concerns first-year senior secondary students in a different school system. The mathematics chatbot experiment does not establish a separately verified effect for every U.S. high-school year. ([RAND sample description](https://www.rand.org/pubs/research_briefs/RB9746.html), [World Bank population description](https://openknowledge.worldbank.org/entities/publication/15e1ff08-15ae-4f7a-b2a8-d146e6c113ee), [PNAS study](https://www.pnas.org/doi/10.1073/pnas.2422633122))

Senior-year readiness, independent research, college transition, career preparation, and sustained writing development remain open research questions. They are not assumed consequences of a mathematics or English trial. The absence of a verified estimate in this targeted review is not a claim that no relevant research exists anywhere. It's a refusal to fill the gap with the nearest available number. For what happens after graduation, [Chapter 4](#ai-college-workforce-training) takes up college and workforce training on its own evidence.

### 3.6 How to Explain Effect Sizes to Parents

Every number in this chapter will eventually get repeated at a PTA meeting, in a school newsletter, or in a parent group chat. Most of them will get mangled on the way. Here's how a teacher, principal, or board member can explain them without adding a single new number, using only the four interpretation rules from Section 3.1.

**Start with what a standard deviation is not.** When a study says 0.23 standard deviations, say: "This is a way of comparing score differences across different tests. It is not 23 percent more learning." If a parent asks whether 0.23 is big, the fair answer is that it depends on the students, the test, the length of the program, and the cost, which is why the study details matter more than the number.

**Say "points" when you mean points.** When Tutor CoPilot's exit-ticket passing rose from about 62 percent to 66 percent, that's four percentage points. Say "four more students out of every hundred passed the exit ticket," not "a 66 percent gain" and not "learning went up 6 percent."

**Always say "compared with what."** When the PNAS study reports that unrestricted-assistant students performed 17 percent worse on later exams, the comparison is the control group in that study. It's a relative change, not a 17-point drop. Say the comparison out loud every time.

**Say who got the offer, not just who used it.** If a result is reported for everyone assigned to the program, say so. If a vendor shows you results only for "active users," explain that the kids who use a tool the most may have been different from the start, and that the fair comparison is between the groups that were offered it and the groups that weren't.

**Name the package.** Never say "AI raised reading scores." Say "a teacher-led reading program that used assessment-based recommendations was associated with higher reading scores." It's longer. It's also true.

Here's a hypothetical script for a principal at a parent night, built only from numbers already in this chapter: "Our pilot is modeled on programs where the tool helps the teacher rather than replacing her. In one published study, a similar arrangement helped students pass short end-of-lesson checks more often, but it did not show a significant change on end-of-year tests. So we're going to measure your children at the end of the year, without the tool, and we'll report back either way."

That last sentence is the one parents remember. It tells them the school is testing for the outcome they care about.

### 3.7 Who Decides What: The Audience Framework

The same educational outcome can matter to several audiences, but relevance must not be confused with proof of an additional downstream benefit. This table translates the findings into decision questions. It's an analytical framework, not a new causal study.

| Audience | Interaction to evaluate | Demonstrated or directly studied value | Additional evidence needed |
|---|---|---|---|
| Student, kindergarten through grade 4 | Teacher-selected instruction or supervised, bounded practice | Reading and mathematics outcomes in specified instructional packages | Independent retention, developmental appropriateness, transfer, effects outside tested skills |
| Student, grades 5 through 8 | Hints, feedback, human tutoring, staffed remediation | Small or qualified assessment gains and improved session mastery in particular trials | Which component causes improvement; durable learning; effects for the particular grade and learner |
| Student, grades 9 through 12 | Curriculum-linked tutoring with independent assessment | Positive supervised-program results and a demonstrated risk from unrestricted answer assistance | Graduation, college readiness, employment, cross-subject transfer, year-specific outcomes |
| Parent or caregiver | Understand the goal, supervision, data use, and the child's independent progress | Child-level outcomes can inform decisions; household outcomes were not directly established | Family time, stress, affordability, usability, access measured directly |
| Teacher or tutor | Review suggestions, adapt instruction, prepare materials | Bounded preparation-time savings and improved tutoring-session outcomes | Total workload, verification burden, sustainability, actual redeployment of saved time |
| School board or district | Select, fund, monitor, and discontinue programs | Study-specific attainment and operational findings as starting benchmarks, not local guarantees | Local comparative outcomes, full costs, access, incident reporting, procurement controls |
| Civic group or infrastructure builder | Evaluate the relationship between educational services and physical infrastructure | This chapter establishes no site-specific power, water, fiscal, or community finding | Separate utility, siting, resource-use, fiscal, and community evidence |

Take a stay-at-home parent raising three children. The immediate value of this chapter for that parent is a clearer way to ask whether a school's proposed tool actually helps each child learn. It would be premature to turn these studies into an assumed number of hours saved at home or a promised improvement in family quality of life, and [Chapter 6](#ai-for-families-households) shows why the household evidence doesn't fill that gap either.

### 3.8 Conditions for Responsible AI in Schools

UNICEF's December 2025 Guidance on AI and Children 3.0 recommends age-appropriate systems, protection of children's data, human oversight of consequential decisions, inclusion, and demonstrated educational efficacy before large-scale deployment. It specifically emphasizes keeping teachers central to education rather than replacing them with automated systems. This guidance is a policy and child-rights framework. It is not a randomized estimate of learning gains, and it is not a statement of every jurisdiction's law. ([UNICEF Guidance on AI and Children 3.0](https://www.unicef.org/innocenti/media/11991/file/UNICEF-Innocenti-Guidance-on-AI-and-Children-3-2025.pdf))

Here's my proposed translation of that framework into an evaluation standard:

- **Adult accountability.** Name the teacher, school administrator, or other responsible adult who can intervene, correct an error, or stop use.
- **Age-appropriate interaction.** Specify what the student can ask, what the system can answer, and how the design differs between an early reader and an older adolescent.
- **Instructional boundaries.** Identify when the system may give a hint, an explanation, a complete solution, or no response.
- **Data minimization.** Document what information is collected, retained, shared, and used for training. Don't treat a child's full educational history as a default input.
- **Human recourse.** Provide a practical way to challenge an incorrect recommendation or harmful interaction.
- **Accessibility and alternatives.** Test the actual interface with intended users, and provide a meaningful alternative when a student cannot or should not use it.
- **Consequential-decision exclusion.** Don't infer that a tool shown useful for practice is validated for placement, discipline, diagnosis, or high-stakes assessment.

These controls are consistent with UNICEF's emphasis on developmental appropriateness, privacy, inclusion, and human control. They still need to be translated into applicable school policies, contract terms, and legal requirements before deployment.

Look at "instructional boundaries" next to the PNAS result. The difference between the tool that hurt independent performance and the tool that largely avoided that harm was, in large part, a set of instructional boundaries. This condition isn't bureaucratic box-checking. It's the variable the best experiment in the set turned on.

One distribution caveat belongs here. This targeted set of studies does not supply a consistent set of independently verified effects for students with disabilities, every language group, every income group, or every grade. Local evaluation should state which populations are represented and should not fill missing subgroup evidence with assumptions. A positive average, as the World Bank heterogeneity results show, can coexist with widening gaps.

### 3.9 AI in Schools Evidence: Measuring Value Consistently

The central evaluation question for any school deployment should be: what can the student do, what can the adult do, and what does the complete intervention cost compared with a credible alternative? Student learning and operational efficiency should be reported separately before anyone tries to combine them.

#### 3.9.1 The Common Evidence Record

Every consequential education claim should keep the same fields, so a favorable headline can't shed its comparison group, grade level, or uncertainty when it's repeated in a public document:

| Field | Required description |
|---|---|
| Population | Grade, age if available, location, language, baseline achievement, eligibility, relevant access conditions |
| Intervention | Product and version, student-facing or adult-facing role, instructional restrictions, curriculum, supervision, training |
| Comparator | What the other group actually received, including teacher time and other technology |
| Exposure | Scheduled time, realized use, duration, adherence, attrition |
| Assignment and analysis | Randomization unit, numbers assigned, analytic sample, unique students versus repeated observations, clustering |
| Primary outcome | Named assessment or operational measure, timing, whether assistance remained available |
| Effect and uncertainty | Metric, estimate, interval or standard error when reported, significance, robustness qualifications |
| Durability and transfer | Delayed tests, unassisted performance, unfamiliar problems, broader outcomes where measured |
| Cost and burden | Software, devices, connectivity, implementation, training, supervision, verification, displaced activities |
| Independence and status | Peer review, working-paper status, provider involvement, funding, independent review where available |
| Permitted conclusion | The narrow statement actually supported, plus the most tempting unsupported extension |

#### 3.9.2 The Eight-Dimension School Evaluation Scorecard

For future evaluations, I recommend measurement across eight dimensions:

1. Unassisted learning, with delayed follow-up.
2. Transfer to unrehearsed problems.
3. Engagement, reported alongside nonuse.
4. Adult work, including verification and error resolution.
5. Access and outcomes by student group, with reasons for nonparticipation.
6. Safety, privacy, and human recourse incidents.
7. Total delivery cost per participating student, against the alternative.
8. Displacement: what the intervention replaced.

A local benefit claim should depend on a comparison and outcome defined in advance, not on a retrospective pick of the most favorable dashboard metric. The reviewed studies show why: engagement increased without established literacy gains in one program, and assisted performance rose while later independent performance fell in another. ([Literacy engagement evaluation](https://edworkingpapers.com/sites/default/files/ai26-1451.pdf), [high-school mathematics experiment](https://www.pnas.org/doi/10.1073/pnas.2422633122))

#### 3.9.3 Why This Chapter Makes No Return-on-Investment Claim

I don't convert teacher minutes, standardized test differences, or exit-ticket improvements into a single financial return. That conversion would require explicit assumptions about implementation costs, the value and actual use of saved time, the durability of effects, and the alternative use of school resources. And an inexpensive model interaction is not the full cost of educational delivery. Training, supervision, curriculum integration, devices, connectivity, verification, and program management all belong in the economic comparison before a district, or a company like ours, makes an affordability claim. Any document that quotes a per-student software price as the cost of the program has answered a question no one asked.

### 3.10 Questions to Ask Your School Board Before It Buys an AI Tutor

Everything above boils down to a list you can bring to a public comment period or a board work session. Each question comes straight from the scorecard, the common evidence record, or the conditions in Section 3.8. None requires technical expertise to ask. All of them require a real answer.

**About the evidence**

- Which study does the vendor cite, and did that study test the same product, grade, subject, and arrangement we're buying?
- What did the comparison group in that study actually receive?
- Is the cited result a peer-reviewed paper, a working paper, a preprint, or a vendor report, and has any independent reviewer rated it?
- Were students tested with the tool available or without it?

**About our own pilot**

- What outcome will we measure, and did we write it down before the pilot started?
- Will students be assessed later without the tool, and will we compare them against students who didn't use it?
- Will we report usage alongside nonuse, and report which students didn't participate and why?
- Will we measure teacher time, including time spent checking what the tool produces?

**About safeguards**

- Which adult is accountable for each classroom's use, and who can stop it?
- What can students ask, what can the system answer, and when will it give a hint instead of a full solution?
- What student data is collected, retained, shared, or used for training?
- How does a student or parent challenge a wrong or harmful response?
- Is this tool being used, or planned for use, in placement, discipline, diagnosis, or high-stakes testing?

**About cost and trade-offs**

- What is the total delivery cost per participating student, including training, supervision, devices, connectivity, and verification, not just the license?
- What are we giving up to make room for it: teacher small-group time, tutoring hours, something else?
- What result would cause us to stop, and when will we decide?

### 3.11 Evidence Boundaries and Open Questions

#### 3.11.1 Grade-by-Grade Coverage Map

The grade-band coverage deserves explicit mapping, because "covered" here means covered by the selected studies, not by all relevant research worldwide:

| Segment | Coverage in this review | Boundary |
|---|---|---|
| Kindergarten | Direct kindergarten reading trial; age-adjacent U.K. early mathematics trial | Does not establish unrestricted conversational use for young children |
| Grades 1 and 2 | Included in the literacy-platform support trials | No statistically established incremental reading gain from added engagement support |
| Grades 3 and 4 | Included in literacy support and Tutor CoPilot sessions | Pooled results are not separate effects for each grade |
| Grade 5 | Included in one literacy trial and CoPilot session data | Less direct grade-specific coverage than the grouped label suggests |
| Grades 6 through 8 | CoPilot grade 6 sessions; grades 6 through 8 remediation trial; grade 7 ASSISTments with grade 8 outcomes | Different interventions, assessment timings, levels of certainty |
| Grade 9 or early secondary | Algebra evidence concentrated in grade 9; Nigerian first-year senior secondary English evidence | Foreign school stages are not automatically U.S. grade equivalents |
| Grades 10 through 12 | High-school mathematics experiment and some broader high-school representation | No separate verified sophomore, junior, or senior benefit estimate |

#### 3.11.2 Publication Status Stays Attached

Publication status stays attached to each finding rather than being flattened into a single "research-backed" label. The kindergarten reading study is peer-reviewed. The early mathematics trial is a peer-reviewed 2019 article. The literacy engagement and Khanmigo studies are working papers. Tutor CoPilot is a preprint. ASSISTments carries both a research report and the April 2026 WWC review, with its reservations kept. The teacher preparation trial is an NFER/EEF randomized evaluation report. The PNAS mathematics study is peer-reviewed. The Nigerian evaluation is a World Bank working paper whose canonical abstract and authors' presentation were reviewed, not subjected to a claimed full-paper audit. And the algebra evaluation rests on RAND's publication record, summary, and 2014 addendum.

Peer review doesn't remove the need to examine design, sample loss, measurement, and generalization. A working paper isn't automatically uninformative. Status attaches, and it stays attached. For how this report grades evidence classes across every chapter, see [Chapter 11](#ai-accountability-framework).

#### 3.11.3 What the Research Has Not Settled

The evidence does not yet settle these questions:

- Whether observed improvements persist across later grades and survive removal of the tool.
- Whether routine assistance strengthens or weakens verification, reasoning, and recognizing uncertainty.
- What sustained effects look like for writing, source evaluation, and original work, beyond this mathematics-heavy set of studies.
- Which interventions improve college readiness, skilled work, and adult civic participation for older students.
- Whether families experience measurable reductions in administrative burden or tutoring cost.
- Who benefits, who disengages, and what design changes are needed for language, disability, and access differences.
- Whether tested programs outperform a credible alternative use of the same money and adult time.
- Whether effects persist after a model update, interface redesign, pricing change, or guardrail change.

These are limits on what this chapter can conclude. They are not reasons to disregard its positive findings.

### 3.12 The Boundary Between Learning Evidence and Infrastructure

I'll be direct about who is writing this. This report is prepared by Savrn, which designs and builds AI factories and works on the infrastructure side of this industry, and our stated objective is a stronger relationship between that infrastructure and community well-being. I'm disclosing that commercial context rather than dressing it up as institutional independence.

Educational benefit and infrastructure acceptability are separate propositions. Even a strong learning result does not establish that a proposed facility has acceptable power demand, water use, public costs, land-use effects, or community commitments.

The seven Savrn trackers belong alongside this chapter as a separate evidence layer, not as supporting citations for learning gains. Their role is to inform infrastructure questions while education claims stay grounded in education research. A tracker entry starts a question and never ends one. The public-value chain I'd propose should be treated as a set of questions, not an assumed causal sequence:

1. Does the educational intervention improve a defined outcome?
2. Can schools deliver it reliably, safely, and affordably?
3. What infrastructure is actually needed to provide that service?
4. What costs, resource demands, and benefits arise from a particular facility?
5. Can affected people verify commitments and obtain remedies?

That separation keeps this chapter useful to a parent, teacher, or civic committee that never buys a Savrn service. It also stops a legitimate education finding from being stretched into an argument for a specific development. [Chapter 10](#data-centers-community-impact) examines the infrastructure side on its own evidence.

> **Operator's note: The same test applies to us**
>
> Savrn is built around behind-the-meter power and closed-loop cooling with a zero-makeup-water design goal. Those are design goals and publisher statements. They are not results, and I'm not going to let them ride on the back of anybody's reading scores.
> 
> I've asked every study in this chapter to show what happens when the support is removed, to name its comparison, and to report its null results as loudly as its good ones. Hold us to that. When a Savrn facility makes a claim about water or power, the fair response is the same one I'd give a school board: show me the measured outcome, show me the comparison, and show me the denominator. If we can't, the claim doesn't count yet. Operator first means no infrastructure without customers, and it also means no claims without records.
### 3.13 Conclusion: What AI Can and Can't Do for Your Kid's Learning

So, back to the kitchen table. Is the tool helping your kid learn, or doing the homework for them? The evidence says that depends almost entirely on how the tool is built and how it's used, and that you can't tell which from the homework itself. You can only tell by checking what your child can do without it.

The defensible position is neither blanket adoption nor blanket rejection. Educational decisions should rest on clearly identified learners, interventions, comparison conditions, outcomes, costs, and uncertainties, with the same standard of evidence applied to favorable and unfavorable results. The thesis that comes out of that is narrower and more useful than either blanket optimism or blanket rejection: **educational value should be established by what learners can do independently and what educators can demonstrably deliver, under clearly described conditions.**

The evidence here supports real optimism of a specific kind. Structured systems, adult-facing support, and supervised programs have produced measurable benefits for real students, and the PNAS experiment shows that the adverse outcome is not inevitable. It's a design failure with a known fix. What the evidence does not support is the idea that capability turns into learning by itself. Between the model and the child stands a chain of design decisions. This chapter has tried to show, study by study, exactly where that chain holds and exactly where, left unattended, it breaks.

### 3.14 Frequently Asked Questions

#### Does AI help students learn, according to research?

Sometimes, under specific conditions. In this review of ten K-12 study programs, structured systems, tools that support teachers, and supervised programs produced measurable gains, such as 0.52 on a latent kindergarten reading measure in a teacher-led package. But in a 2025 PNAS experiment, students with an unrestricted GPT-4 assistant scored 17 percent worse than the control group once the tool was removed. Design and supervision decide the outcome.

#### Does ChatGPT help students learn or just do their homework?

The strongest evidence says it can do both, depending on design. In the PNAS study of nearly 1,000 high-school math students, an unrestricted GPT-4 assistant raised assisted practice performance 48 percent, then those students did 17 percent worse than controls on exams without it. A constrained tutor version largely avoided that penalty, though avoiding a penalty is not proof of a gain.

#### What did the Khanmigo study find?

An August 2026 working paper on Khan Academy with Khanmigo in grades 6 through 8 in Hamilton County, Tennessee found 0.020 standard deviations in year one (not significant), 0.084 in year two (significant at 5 percent), and 0.062 stacked (marginal at 10 percent). The treatment was a staffed remediation package, so it does not isolate the chatbot, and a small-cluster robustness test gave p = .0511.

#### Do AI tutoring studies show gains on end-of-year tests?

Not consistently. Tutor CoPilot, which gave human tutors AI suggestions, raised exit-ticket passing from about 62 percent to 66 percent over roughly two months, but found no statistically significant end-of-year math improvement. ASSISTments showed a 0.10 effect size on a grade 8 state test, yet the April 2026 What Works Clearinghouse review rated the study "with reservations" and found uncertain effects overall.

#### Does AI save teachers time?

For a defined task, one trial says yes. In an NFER/EEF randomized trial of 259 teachers in 68 English secondary schools, teachers using ChatGPT with a preparation guide reported about 56.2 minutes a week preparing specified Year 7 and Year 8 science lessons, versus 81.5 minutes for comparison teachers, 31 percent less. It did not measure total workload or student achievement.

#### Is AI safe for young children in school?

The research here does not test unrestricted chatbots with young children. The positive early-grade studies used teacher-directed or bounded tools, such as math apps used about 30 minutes a day for 12 weeks under staff supervision. UNICEF's December 2025 guidance calls for age-appropriate design, data protection, human oversight, and demonstrated efficacy before large-scale deployment, but it is policy guidance, not a learning study.

#### What does an effect size of 0.23 mean for my child?

It is a standardized score difference between groups, not 23 percent more learning. The 0.23 figure comes from a World Bank evaluation of a six-week supervised English program in Nigeria, with 759 students in the final analytic sample out of 1,328 randomized. It describes that program and population on average, not a guaranteed result for any individual child or for U.S. students.

#### What should parents ask a school before it adopts an AI tutor?

Ask what study supports the product and whether it tested the same grade and arrangement, what the comparison group received, and whether students will be tested later without the tool. Ask which adult is accountable, what data is collected, how errors are challenged, and the total cost beyond the license. In the studies reviewed, results depended on supervision, structure, and unaided measurement.
