Search public pages, research tools, and SAVRN solutions.
Savrn Insights · Research Report
The Superintelligence Transition: A Human Agenda From Kindergarten to Retirement
What the evidence actually says about AI in schools, jobs, homes, town halls, retirement accounts, clinics, and the data centers underneath all of it. Favorable, null, and adverse findings, every number with its source and its limit attached.
CH
Chad Everett HarrisFounder and CEO, Savrn
Published
Sep 24, 2026
Evidence cutoff
Sep 23, 2026
Length
116,808 words · 8h 28m
Sources
83 external, plus 7 Savrn trackers
Figures
63
Disclosure: prepared for Savrn, which has a commercial interest in AI infrastructure and publishes the seven trackers cited here. Not an independent institutional review. Not investment, medical, legal, or tax advice.
Every chapter stands on its own. Start with Chapter 1, or jump to data centers and communities. The argument that ties all eleven chapters together, with its seven findings, what would prove it wrong, and the method behind it, closes the page in The Research Thesis, after the evidence register.
The audio edition · Read by Bella
Listen to the whole report
13 episodes, 2 h 26 min. Each one walks a chapter's argument, its studies, and their limits, so you can listen instead of read. It plays straight through, one chapter into the next.
Up next · Chapter 1
AI Safety Promises and Public Trust: How to Test What AI Labs Say
I have spent my career around industries that make big promises about power and pace. Utilities promise capacity. Developers promise timelines. Equipment vendors promise uptime. Every one of those promises sounds solid in the press release, and every one of them either holds up against a number, a date, and a penalty, or it doesn't. So when the companies building the most capable AI systems in the world publish essays about AI safety commitments and then ship new models days later, I don't think your real question is whether they're heroes or hypocrites. Your real question is simpler and harder: should I believe them?
A parent deciding whether a tutoring app belongs in the house, a teacher weighing a classroom tool, a county commissioner looking at a power contract, a worker whose job is being redesigned around software: every one of them is being asked, in one form or another, to take a company's word about what a system does and how carefully it was released.
Here's what the evidence in this chapter will show, including the uncomfortable parts. The September 2026 essays do not forbid releases, so a release is not automatically a broken promise. But the longest public record we have of a lab making a specific, quantified safety pledge, OpenAI's 2023 superalignment commitment, was not carried out as stated, according to reporting by several outlets. The public knows something is off: Pew's surveys show concern climbing for five years. And the tools the public needs to judge any single release, a named baseline, an operational safety bar, an evaluator with publication rights, are mostly missing. I'm going to give you a test you can run yourself, a worksheet to run it on, and a standard for what a release explanation should contain. Then I'm going to run that same test on my own company, because a standard that exempts its author is not a standard.
1.1 AI Safety Commitments in September 2026: What Happened#
In September 2026, two of the most prominent laboratories building advanced AI systems published essays asking the industry to slow down. Days later, both companies released new products. The sequence produced the kind of headlines that write themselves: the warnings were theater, the restraint was marketing, and the releases were proof of hypocrisy. Or, from the opposite camp: the warnings were responsible leadership, the releases were routine progress, and the critics were scoring points without reading what was actually said.
This report takes neither side, because neither side can be supported by the documentary record alone. What the record does support is more interesting than either polemic. It supports a real question about the relationship between public warnings, private pacing, and product releases. It supports holding companies to the specific conditions they themselves proposed. And it supports a communication standard that would let a parent, teacher, civic member, worker, or resident judge releases on evidence rather than on trust or suspicion.
Everything that follows in this research depends on this chapter's method. The later chapters cover kindergarten classrooms, college physics courses, customer-support queues, household kitchens, civic committee rooms, adviser offices, clinics, and the substations and water lines of the physical infrastructure that powers all of it. Every one of those settings involves a person making a decision with incomplete information.
The September 2026 record is a useful place to learn what complete information would look like, because it is the moment when the industry's own leaders told the public what should count. They handed the rest of us a grading rubric. This chapter picks it up.
It is also a useful place to learn what institutions owe the public when they ask for trust. Savrn, the company I founded and for which this report was prepared, states a purpose of improving the connection between data center infrastructure and communities, including avoiding burdens on community power and water. I treat that purpose as an objective that requires measurement and enforceable commitments, not as a verified operating claim. The same procedural standard I apply here to the laboratories applies to every institution in these pages, Savrn included.
The chapters that follow are long because the questions are long. A teacher's decision about a classroom tool and a community's decision about a power contract are not questions that can be answered well in a paragraph, and I don't try. What I offer instead is a way of reading claims, any claims, from any source, that survives contact with the evidence. That skill transfers. It is the actual product of this entire report.
If you read nothing else, read Section 1.5 and the worksheet in Section 1.5.5.
1.2 Separating Four Claims: Statements, Releases, Consistency, Harm#
Public calls for restraint followed by product releases deserve scrutiny. But proximity on a calendar does not establish a broken commitment. The reviewed statements call for safety-conditioned pacing, not a blanket end to training or product releases. Dario Amodei's essay, "We Must Pace the Frontier", argues that the pace of capability improvement must slow; Jakub Pachocki's essay, "An Alien Mind", argues that scaling must be constrained by confidence in safety. Both are statements of position, and both explicitly preserve room for continued releases under stated conditions.
So the fair question is not "did a release follow a warning?" It is a narrower and harder question: does an organization's actual conduct satisfy the particular conditions it publicly proposed? Answering that requires separating four distinct claims.
Claim
Can it be answered from the record?
What it takes
1. What was said
Yes
Dated public records
2. What was released
Yes
Dated public records
3. Whether the two are consistent
In principle, yes
Comparing conduct against the specific conditions proposed, not against a caricature of them
4. Whether the communication caused harm
Not as of September 23, 2026
Causal evidence about comprehension, trust, and behavior that the documentary record does not contain
The first two claims are answerable. The third is answerable in principle, with a disciplined test. The fourth, as of the September 23, 2026 evidence cutoff of this report, is not established in either direction, and a report that pretends otherwise is selling something.
Public discussion of AI routinely collapses all four claims into one. A dramatic sequence of headlines feels like evidence. It is not. It is a sequence of headlines.
The distinction also protects critics from their own worst habit: grading companies on a commitment the company never made. If a laboratory proposes conditional pacing and ships a product that plausibly satisfies its stated conditions, the shipping is not a confession. If it proposes evaluator-confirmed safety bars and cannot name an evaluator, the proposal is a press strategy. Everything depends on the content of the commitment, which is why this chapter spends its effort on what was actually written rather than on what was felt.
1.3 A Bounded Chronology of Statements and Releases#
Four dated records anchor this chapter. Each is presented with what it verifies and what it cannot.
Record
Verified content
Evidentiary limit
OpenAI chief scientist's September 6, 2026 essay
Jakub Pachocki wrote that scaling must be constrained by confidence in safety, and expressed hope that voluntary slowdowns would become commonplace until shared safety bars were established. (An Alien Mind)
A statement of position, not an independently assessed risk threshold or a commitment to stop every update.
Anthropic chief executive's September 2026 essay
Dario Amodei wrote, "We must slow the pace at which we improve the capabilities of AI models," while explicitly distinguishing pacing from halting training or technical progress. (We Must Pace the Frontier)
The essay's rationale and proposed evaluation arrangements are not proof that those arrangements were implemented adequately.
The update date is not evidence that every model described on that page was first released on that date.
Figure 1.1Timeline
The public record, 2023 to 2026: commitments, departures, essays, releases
Documentary record only. Dates come from the cited announcements and from secondary reporting where the chapter says so. A sequence on a calendar does not by itself establish a broken commitment or a satisfied one.
The chronology is bounded on purpose. It includes the records this chapter can verify and excludes the rumor, inference, and retrospective interpretation that surround them. A wider chronology could be assembled, but it would add heat, not light, to the question that matters here.
1.3.1 Two features of the chronology worth noticing#
First, the pacing essays and the releases are separated by roughly two weeks, not by years. Whatever else the sequence shows, it shows that the industry's leading safety voices expected the public to evaluate releases against a standard those voices had just articulated. That makes the consistency test in Section 1.5 a live question rather than an academic one.
Second, the two September 22 records are announcements of record, not independent assessments. Anthropic describes evaluations and safeguards associated with Claude Opus 5.5; OpenAI records additions to the GPT-6 family. What those descriptions establish is what the companies say. Whether the described evaluations are sufficient is precisely the kind of question the announcements cannot answer about themselves.
1.4 What AI Pacing Means: Conditional Commitments, Not Halts#
The most quoted sentence from Amodei's essay is the demand to slow the pace of capability improvement. The most important sentence is the clarification that follows it. In his own words:
"To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this." (Amodei's original essay)
Pachocki's essay is constructed the same way. He describes pursuing technical solutions, building defensive systems, and withholding further scaling as needed, rather than announcing an unconditional stop. (Pachocki's original essay)
Two implications follow, and both cut against the easy storylines.
First, treating every later release as an automatic contradiction is ruled out by the text itself. A company can release a product while sincerely holding the pacing position, provided the release satisfies the conditions the essay describes: adequate alignment and safeguard time, third-party evaluator confirmation. Whether any particular release did satisfy those conditions is a factual question this record cannot settle, but the position itself does not forbid releases.
Second, the pacing position is not empty. It proposes something specific and testable: that deployment decisions should be conditioned on safety evidence confirmed by evaluators who are not the company itself. A company that claims to hold this position has accepted a standard it can fail. That is what makes the position worth taking seriously, and what makes the consistency test in the next section possible.
The records therefore support scrutiny of conditional commitments. They do not, by themselves, prove either hypocrisy or satisfactory compliance.
1.4.1 The longer arc: OpenAI superalignment, 2023 to 2024#
The September 2026 essays did not arrive from nowhere, and the record of what laboratories have publicly promised is longer than a single month.
In July 2023, OpenAI announced a dedicated superalignment effort, co-led by then-chief scientist Ilya Sutskever and alignment head Jan Leike, stating that it would dedicate 20 percent of its secured compute to the problem of steering and controlling systems "much smarter than us" within four years. The announcement described superintelligence as potentially "the most impactful technology humanity has ever invented," warned that its power could lead to "the disempowerment of humanity or even human extinction," and conceded that current alignment techniques would not scale to superintelligence because humans would not be able to reliably supervise such systems. (OpenAI, Introducing Superalignment)
The 2023 announcement is worth pausing on, because it establishes the pattern the 2026 essays repeat: a laboratory publicly defining a risk, publicly committing resources against it, and explicitly acknowledging that the commitment's success is uncertain. You can weigh the sincerity and the follow-through of such commitments for yourself. What I insist on is only that the weighing happen against the text of the commitments, with dates, rather than against the mood of the moment.
1.4.2 What happened to the superalignment commitment#
The follow-through is checkable, because the 2023 commitment named a number and a deadline. What happened next is on the public record, and I'm relying on secondary reporting for it, so I'll name the reporters.
OpenAI announces superalignment: 20 percent of secured compute, four years
May 14, 2024
Ilya Sutskever announces his departure from OpenAI
May 17, 2024
Jan Leike resigns publicly, citing safety culture and compute
Within days of those departures
OpenAI dissolves the Superalignment team and redistributes its remaining staff
October 2024
OpenAI's AGI Readiness team dissolves when its leader, Miles Brundage, resigns
February 2026
OpenAI's Mission Alignment team is disbanded after roughly sixteen months
Leike wrote on his way out that "over the past years, safety culture and processes have taken a backseat to shiny products" and that his team had been "sailing against the wind," struggling to obtain the compute its mandate required. The 20-percent-of-compute, four-year commitment announced in July 2023 thus ended, as a dedicated team, at roughly the one-year mark. Reporting at the time, including Fortune's, as relayed in these accounts, indicated that the team's requests for its pledged compute had been denied before the dissolution. The successor structure did not stabilize either, according to the same reporting.
Read that again: a four-year commitment, with a number attached, ended as a dedicated team at about one year.
1.4.3 What the superalignment record does and does not prove#
This record does not by itself prove that OpenAI's safety work was abandoned. The company stated that safety responsibilities were redistributed rather than ended, and some departures were unattributed. Those facts belong in the same paragraph as the departures, and I'm keeping them there.
What the record does is more precise. It shows that the 2023 commitment, which was specific, dated, and quantified, was not carried out as stated, and that the people most identified with it left while saying so in public. When the same organization's chief scientist writes in September 2026 that scaling must be constrained by confidence in safety, the earlier record is not a footnote to that essay. It is the first data point for the seven-part test in Section 1.5.
The record is checkable only because the commitment had a number and a deadline. Vague commitments can't be caught falling short. That's an argument for specificity, not against it.
1.4.4 A symmetry worth recording: the Jacob Coxon resignation#
One more symmetry deserves notice. In September 2026, according to reporting by Business Standard, an Anthropic researcher named Jacob Coxon resigned and warned publicly about the direction of the frontier race. That is the same move Leike made at OpenAI in 2024, now from inside the company whose chief executive had just published the pacing essay.
I draw no conclusion from the symmetry. I record it, because a reader evaluating either company's September conduct is entitled to the full pattern of who has left these organizations, and what they said on their way out. A departure is testimony, not a verdict. It tells you where to look; it doesn't tell you what you'll find.
1.5 The Missing Denominator: A Seven-Part Test for AI Safety Commitments#
A claim of "slowing" is incomplete unless two things are identified: the activity being slowed, and the baseline it is being compared with. This sounds pedantic until you notice how often the word is used without either.
Show me the denominator. Slowing what, compared with what, over what period?
The problem is structural. "Slower" could refer to training scale, capability improvement, external access, deployment scope, autonomy, or release frequency, and these can move in opposite directions simultaneously. A laboratory can slow its training schedule while expanding its product rollout. It can ship fewer models while each model gains more capability. It can restrict access tiers while widening deployment scope within each tier. Any of these patterns would be accurately described as "slowing" by someone who chose the favorable metric, and accurately described as "accelerating" by someone who chose another.
Figure 1.2Concept
The missing denominator: "slowing down" has to name an activity and a baseline
Conceptual diagram. Directions are illustrative and describe no company. A claim of slowing is incomplete until it names which activity slowed and against which baseline.
Fewer releases than last month does not necessarily mean slower capability growth. A calm product quarter and a jump in capability are not mutually exclusive.
This report proposes a seven-part consistency test. It is a proposed standard, not a finding that any company has passed or failed it.
Test element
Disclosure needed
Why the distinction matters
Activity
Training scale, capability improvement, external access, deployment scope, autonomy, or release frequency
A slower training schedule and a larger product rollout can coexist.
Baseline
Previously planned or expected activity over a stated period
Fewer releases than last month does not necessarily mean slower capability growth.
Safety bar
Stated criteria, test population, failure conditions, and uncertainty
"Safe enough" is not reproducible without an operational standard.
Evaluation
Evaluator access, independence, methods, exclusions, and publication rights
An evaluation's existence does not establish what it could test.
Decision
Whether evidence caused a delay, restriction, redesign, or cancellation
A process with no possible effect on deployment is not meaningful restraint.
Residual risk
Unresolved failures and limits on detection or monitoring
Passing selected tests does not settle every deployment context.
Change control
Version, access tier, tool permissions, and post-release updates
A result for one configuration need not cover a later one.
Each element answers a failure mode in public discussion:
Activity prevents metric shopping.
Baseline prevents phantom slowdowns.
Safety bar prevents "safe enough" from meaning whatever is convenient on a given day.
Evaluation distinguishes an evaluation from a press release about an evaluation.
Decision is the heart of the matter: a review process that has never delayed, restricted, redesigned, or cancelled anything is not restraint, whatever it is called.
Residual risk prevents a clean evaluation sheet from being read as a clean system.
Change control prevents a result for one model version from being silently extended to the next.
If you only remember one element, remember the decision element. That's the whole game. Any review process can produce paperwork. Only a real one can produce a "no."
1.5.2 Why the International AI Safety Report can't settle September#
It's tempting to reach for a big scientific synthesis to answer the September question. The International AI Safety Report 2026 synthesizes scientific evidence on capabilities, risks, and mitigation, but its evidence cutoff precedes December 2025. It therefore cannot validate the particular September 2026 releases in this chronology, and I don't use it that way.
For this report, the absence of a reviewed independent evaluation is recorded as an evidence gap, not as proof that no evaluation exists. A defensible adverse conclusion would need to identify an actual commitment, its relevant conditions, the conduct assessed, and the evidence of noncompliance. That standard protects everyone: it holds critics to the same standard, and it gives companies a clear path to demonstrate compliance rather than asserting it.
To see the test in motion, consider a purely hypothetical laboratory whose conduct will be scored against it. The illustration is invented to demonstrate the test; it describes no real company.
Suppose the laboratory announces that it is "slowing down for safety." The activity element asks: which of the six activities has slowed? Suppose its answer is release frequency: it shipped two models this year instead of four. The baseline element asks: compared with what? If last year's four releases were itself an unusual burst and the two-year average is two, the "slowdown" is a return to normal, and the claim dissolves into arithmetic. Suppose, though, that the schedule really did halve. The safety-bar element asks: what criteria determined that these two releases were safe to ship? If the answer is internal red-teaming with unpublished failure conditions, the bar is not operational; no third party could reproduce the judgment, and no critic could falsify it.
Now suppose the laboratory names an external evaluation firm. The evaluation element asks: what access did the evaluators have, what methods did they use, what could they not test, and may they publish? If the evaluators saw a frozen configuration that differs from the shipped product, the evaluation result may not cover the release. The decision element asks the question that separates process from theater: did the evaluation, or anything else, ever cause a delay, a restriction, a redesign, or a cancellation? If the answer across the company's history is no, if every evaluation has ended in a release, then either the evaluators have never found a problem worth acting on (possible, but worth scrutinizing) or the process cannot say no. A review that cannot say no is not a review.
The residual-risk element asks what the evaluation did not cover, and the change-control element asks whether next quarter's update inherits this quarter's result. A passing grade for version 1.0 says nothing about version 1.1 with new tool permissions.
None of these questions requires insider access. All of them can be asked by any journalist, regulator, or interested citizen, and all of them have answers that are either on the public record or conspicuously absent from it. That is what makes the test usable. It converts "trust us" into a checklist.
Run the hypothetical lab through the worksheet in Section 1.5.5 and nearly every row reads "undisclosed." That doesn't say the lab is reckless. It says the public can't verify the claim.
The test is designed to be run, not just admired. A working journalist, an analyst at a regulator, or a member of a state legislature's technology committee can apply it to any laboratory's public record in a focused afternoon, using four steps.
Step one: fix the commitment's text before fixing its meaning. Pull the primary statement (the essay, the system card, the testimony) and quote the operative sentences verbatim, with their dates. Paraphrase is where grading goes to die. "We are committed to safety" and "we will dedicate 20 percent of our compute to alignment within four years" are different kinds of sentences, and only the second kind can fail. If the commitment contains no number, date, criterion, or named evaluator, that is itself the finding: the commitment is not operational, and no conduct can be scored against it.
Step two: assemble the conduct record on the commitment's own terms. Collect the releases, the evaluations, and the organizational changes that postdate the commitment. The record assembled in Section 1.4 is an example of this step done properly: dates of the commitment, dates of the departures, the dissolution, the reported denial of pledged compute. None of it required a subpoena.
Step three: score each test element as supported, contradicted, or undisclosed. The three-way scoring is the point. A loud public argument pressures everyone toward supported or contradicted; the real output of most scoring sessions is a column of undisclosed entries, and that column is the story. An undisclosed safety bar is not proof of a failed safety bar. It is proof that the public cannot verify the claim.
Step four: publish the worksheet. The test's value compounds if the scoring is reproducible. When a regulator's worksheet and a journalist's worksheet disagree about the same element, the disagreement identifies exactly which record needs to be produced. That is how a procedural standard does its work: not by producing verdicts, but by producing answerable questions.
One caution belongs with the method. Running the test well requires resisting the two shortcuts that make public argument about AI so unproductive:
The sympathy shortcut grades a company on intentions and difficulty rather than on disclosed conduct.
The villainy shortcut treats every undisclosed element as proof of the worst case.
The worksheet format exists to make both shortcuts visible. A column of "undisclosed" entries is a demand for records, not a conviction.
1.5.5 Run the test yourself: a copyable worksheet#
Here's the worksheet. Copy it into a document or spreadsheet, fill in the commitment text at the top, and score each row. It uses only the seven elements above; nothing here requires inside information.
Header (fill in before scoring):
Organization:
Commitment text, quoted verbatim:
Date and source of the commitment:
Conduct being scored (release, evaluation, organizational change) and its date:
Who scored this, and when:
Element
What to look for
Supported
Contradicted
Undisclosed
Activity
Which activity is said to be slowing: training scale, capability improvement, external access, deployment scope, autonomy, or release frequency
[ ]
[ ]
[ ]
Baseline
The previously planned or expected activity, over a stated period, that the claim is compared with
[ ]
[ ]
[ ]
Safety bar
Stated criteria, the test population, failure conditions, and stated uncertainty
[ ]
[ ]
[ ]
Evaluation
Who evaluated, what access they had, how independent they are, their methods and exclusions, and whether they may publish
[ ]
[ ]
[ ]
Decision
Any documented delay, restriction, redesign, or cancellation caused by evidence
[ ]
[ ]
[ ]
Residual risk
Unresolved failures, and limits on detection or monitoring
[ ]
[ ]
[ ]
Change control
The version, access tier, and tool permissions tested, and how post-release updates are handled
[ ]
[ ]
[ ]
How to score each row:
Mark Supported only when a dated public record shows the element was disclosed and the conduct matches it. Write the source link in your notes.
Mark Contradicted only when a dated public record shows conduct that conflicts with the stated condition. Name the record. The superalignment sequence in Section 1.4.2 is an example of what contradicting evidence looks like: a stated number and deadline, and a reported outcome that did not match them.
Mark Undisclosed when the public record is silent. That is the default, not a failure of your research.
Before you publish, ask one question: could someone who disagrees with me reproduce my scores from my links? If not, you've produced an opinion with a table around it.
For a parent: when an app says a new version is "safer," ask which activity changed and compared with what. If there's no answer, the claim hasn't been made yet.
For a teacher or school leader: ask whether the version your students use is the version that was evaluated. That's the change-control row, and it's the one most easily missed.
For a board member or investor: ask management which rows of this worksheet the company could fill in today with public records. The undisclosed rows are the governance gap.
For a county commissioner or resident: use the same seven rows on any infrastructure proposal, including one from a company like mine. The questions translate directly.
1.6 What the Record Cannot Support: Consequences and Causal Inference#
This report consequently does not state that the September communication damaged the economy. It also does not assume the communication was harmless. The question is unresolved within this review, and saying so plainly is load-bearing: the chapters on infrastructure and household life quantify what is known about those domains, and their credibility depends on not smuggling unexamined assumptions in through the front door.
The temptation to draw causal conclusions from the September sequence is strongest in exactly the places where the evidence is weakest. Watch for three moves in public argument.
Timeline juxtaposition. Placing an announcement date beside a market movement or a project delay and inviting the reader to connect them.
Denominator switching. Citing the number of headlines, tweets, or alarmed op-eds as if attention were harm.
Attribution by mood. Arguing that because the communication felt chaotic, it must have caused concrete damage somewhere.
Each move produces the feeling of evidence. None of them is evidence. A number without a denominator is a rumor, and a count of alarmed headlines has no denominator at all.
1.6.1 Public Trust in AI: What the Pew Surveys Show, 2021 to 2026#
If the documentary record cannot tell us whether the September communication caused damage, survey research can at least describe the audience that received it. That description matters for communication design, and it comes with a hard caveat: attitudes are not harms, and a worried public is not an injured public.
The Pew AI survey trend: concern up, excitement down. The Pew Research Center has tracked American attitudes toward AI since 2021, and the trend is one-directional. In June 2021, 37 percent of U.S. adults said the increased use of AI in daily life made them more concerned than excited, while 18 percent said the reverse. By June 2026, the concerned share had climbed to 52 percent and the excited share had fallen to 9 percent, with the remainder reporting equal measures of both. Those 2026 figures come from survey data published by Pew Research Center in August 2026.
The shift was not a single post-ChatGPT shock. The concerned share jumped between 2022 and 2023 (from 38 to 52 percent) and has held near that level since, at 51, 50, and 52 percent in 2024, 2025, and 2026 respectively. Concern has continued to climb among adults under 30 even after leveling off among older groups, and worry about job loss is still rising.
Pew survey year
U.S. adults more concerned than excited
2021 (June)
37 percent (18 percent more excited)
2022
38 percent
2023
52 percent
2024
51 percent
2025
50 percent
2026 (June)
52 percent (9 percent more excited)
Figure 1.4Data
Public concern about AI rose and stayed high; experts and the public see it differently
Share of U.S. adults more concerned than excited about increased AI use in daily life, 2021 to 2026, and the 2025 expert versus public comparison. Attitudes are not harms; these are survey measures of opinion.
A second finding sharpens the picture. When Pew compared the general public with people who work as AI researchers and developers in 2025, the two groups were nearly mirror images: 47 percent of experts said they were more excited than concerned about AI in daily life, against 11 percent of the public. On the twenty-year outlook, 56 percent of experts expected a positive impact versus 17 percent of the public. (Pew Research Center, April 2025)
Read together, these results describe a public that has heard the warnings, absorbed them, and is waiting to be convinced. They do not describe a public that has been harmed by any particular communication event. They also describe a trust gap between the people building these systems and everyone else, and that gap is precisely what a communication standard has to bridge.
A company that reads the Pew trend as a marketing problem to be solved with reassurance has missed the finding. The public's concern has proven durable across four years of product improvements, which suggests it attaches to the pattern of conduct, not to any single headline. That is the strongest argument for the procedural standard in this chapter: the only communication strategy with any evidence behind it is a verifiable one.
Common misreadings to avoid:
"52 percent are concerned, so AI is harming people." Concern is an attitude. The Pew data do not measure injury.
"Concern jumped in 2023, so ChatGPT caused it." The data show when the jump happened, not why.
"Experts are more optimistic, so the public is misinformed." The comparison shows a gap. It does not say which side is right.
"The September essays drove the 2026 number." The 2026 survey was fielded in June 2026, before the September essays were published.
1.6.4 Two studies that would answer the harm question#
What would resolve the question of whether the September communication caused harm? Two studies, both feasible.
The first is a communication study. It would randomly assign adults to source-faithful versions of a release explanation: a headline-only version, a capability-and-limits version, and a version including independent evaluation and residual uncertainty. It would measure understanding of the actual commitment, recognition of uncertainty, ability to identify appropriate uses, and willingness to revise a decision after contrary evidence.
One subtlety deserves emphasis: a person can understand a statement more accurately while trusting the organization less. That result would not necessarily represent a communication failure, and a study that treated trust as the only outcome would miss the point. Exposure to alarming claims should not be tested on young children without appropriate specialist and ethical oversight.
The second is a business study. It would need documented changes in procurement, financing, construction, or demand, and a credible comparison separating communication effects from interest rates, utility constraints, customer commitments, and the other events that move markets and project schedules. Merely placing announcement dates beside market or project movements would not establish causation. Anyone can line up two timelines; the trick is that the events in them usually share causes rather than cause each other.
Neither study exists in the reviewed record. Proposing them is not a dodge. It is the difference between a report that grades the available evidence and one that invents evidence to fill a gap.
1.7 A Responsible AI Release Standard: Six Questions#
Whatever the September record shows, one conclusion needs no further study: the way capable systems are explained to the public is not working well enough. The practical question is what a working explanation would contain.
This report proposes that a public explanation of a consequential release or deployment should answer six questions in plain language:
What changed? Not what might change, not what the technology could eventually do. What, specifically, is different as of this date.
Who can use it? The permitted age range, access tier, professional context, and the adult or institutional role required.
What evidence supports the change? The strongest study, its population and design, and where to read it.
What remains unreliable? The failure modes, the populations the evidence does not cover, and the uncertainty.
Who is accountable? A named role, not a brand. Someone who can answer for the outcome.
How can the public challenge an outcome? A working correction and appeal path, with a stated response time.
This standard is proposed, not yet tested. Chapter 11 returns to it as part of the shared accountability framework.
1.7.1 The six questions and their failure patterns#
Each question in the standard exists because a specific, recurring failure pattern made it necessary. I name the patterns generically, because the point is the pattern rather than any one offender.
"What changed?" fails as a headline. Headlines report the most dramatic thing a system could conceivably do, which is rarely the thing that shipped. The reader who absorbs only the headline arrives at the product page expecting a capability that is not there, and the gap between expectation and product does more damage to trust than a candid limitation ever could. The question forces the communicator to lead with the shipped fact.
"Who can use it?" fails as a demographic afterthought. Age gates, professional restrictions, and institutional requirements are usually buried in policy documents that no one reads before forming an opinion. A parent who learns, after the fact, that a product was never intended for a nine-year-old has been failed by the communication, whatever the terms of service said.
"What evidence supports the change?" fails as a number without a study. "Our model scores 94 percent" is not evidence until the reader can learn what the other participants were, what the test population was, and whether anyone independent ran the test. The question requires the citation, not the statistic.
"What remains unreliable?" fails by omission most of all. This is the question that separates explanation from advocacy, and it is the one most often answered with silence. The cost of answering it is a shorter, less shareable announcement. The cost of omitting it is discovered later, by the user, in the worst possible context: a wrong answer in front of a student, a patient, a judge, a customer.
"Who is accountable?" fails as a brand. Brands do not answer questions; people with names and roles do. An accountability chain that terminates at a corporate identity terminates nowhere, because no individual's judgment is on the line.
"How can the public challenge an outcome?" fails as a black box. A correction path that requires expertise the user does not have, or persistence most users will not muster, is not a correction path. The test is procedural, not rhetorical: can an ordinary user, with an ordinary wrong answer, get it fixed within a stated time? If the answer is unknown, the sixth question has not been answered.
The six questions form a sequence with a deliberate structure. The first three establish the positive claim; the last three establish what happens when the claim fails. An explanation that answers only the first three is a sales document. One that answers all six is the minimum a consequential deployment owes its public.
The same six questions produce different sentences for different readers.
For a parent: the permitted age range and adult role, the learning objective, what child information is collected, and how to stop or report a problem.
For an educator: curriculum scope, independent learning outcomes, teacher workload, and whether students retain competence without assistance.
For a civic committee: the source documents, alternatives considered, omitted viewpoints, and the authority that makes the decision.
For a worker: productivity claims separated from monitoring, changes in job responsibility, training requirements, and pay.
For an infrastructure builder: product capability announcements separated from executed capacity commitments and local approval obligations.
For a resident: the project's actual power, water, fiscal, and operating terms, without distant software benefits presented as compensation already delivered locally.
That infrastructure-builder line is the one I live with. A product announcement from a model developer is not a signed capacity commitment, and neither one is a local approval. When those three get blurred together in a public meeting, residents end up evaluating a promise about software as if it were a promise about their substation.
Notice what the standard refuses to do. It does not promise brevity at the cost of uncertainty; the fourth question is the one most corporate communications omit, and it is the one this report considers non-negotiable. It does not treat the audience as a single public; a parent and a power-system engineer need different sentences, and a standard that serves both says so. And it does not confuse drafting with action. A model that can write the explanation does not thereby earn the right to publish it; Chapter 11 returns to that separation as a general control.
A concise explanation should not remove material uncertainty. The practical target is a decision the reader can inspect, not reassurance at any cost.
1.7.4 A worked example: reading a release note in a meeting#
Here's a hypothetical. A school district's technology committee receives a vendor's one-page note announcing that its classroom assistant now runs on a newer model. The note says the update is "more capable and safer." No statistics are given, so I won't invent any.
Walk the six questions. What changed? "More capable" doesn't say. Who can use it? The note doesn't mention age ranges or whether a teacher must supervise. What evidence supports the change? None is cited. What remains unreliable? Silent. Who is accountable? The vendor's brand name. How can a family challenge an outcome? Not stated.
That note answers zero of six. The committee doesn't need to reject the product; it needs to send the six questions back in writing and decide after the answers arrive, then run the seven-part test on the word "safer." Neither tool gives a verdict. Both give better questions.
1.8 Terminology and the Scope of the Report's Title#
The title of this report, "The Superintelligence Transition: A Human Agenda From Kindergarten to Retirement," is an organizing theme, not a finding. It does not claim that any system available as of September 23, 2026 has achieved a universal capability threshold beyond human intelligence. When the historical record discusses such thresholds, I quote it and attribute it.
OpenAI's 2023 explanation of superintelligence, published when it announced its superalignment effort, described systems "much smarter than us" and stated the organization's belief that such systems could arrive within the decade. That was a publisher's prediction, and I treat it as one. (OpenAI, Introducing Superalignment)
Throughout the chapters that follow, the text uses "current models," "advanced models," "structured software," or the name of the specific intervention when that is the accurate description. Original organization names, study titles, quotations, and regulatory language retain their wording, including "AI" where changing it would misrepresent the record. This report does not treat "no longer artificial" as a technical finding, and it is wary of any text, in either direction, that treats terminology as evidence.
The reason for this discipline is practical. The lifespan questions in this report (whether a child learns, whether a worker earns, whether a resident's water bill changes) do not depend on what the most capable system is called. They depend on what a specific tool does for a specific person under specific conditions. Precision about words is how the report keeps its attention on those conditions.
Terminology discipline also has a defensive purpose. In a heated public argument, the side that controls the definition of the threshold often appears to control the conclusion. By declining to treat any label as a finding, this report denies both the optimists and the catastrophists that shortcut. The chapters measure tasks, outcomes, and people. Whatever the most capable system of the moment is called, those measurements stand or fall on their own. Capability is not benefit, and a name is not a capability.
This report is a targeted, multi-source desk-research synthesis with an evidence cutoff of September 23, 2026. It is not a preregistered systematic review, and it makes no claim of exhaustive search or complete study census.
The review prioritizes original research, government evaluations, official records, and primary company statements when the claim concerns what a company said. Secondary and institutional accounts are labeled where used, as with the Business Insider, Newsweek, The Next Web, and Business Standard reporting in this chapter. Where a chapter relies on a publication record, abstract, presentation, or research summary rather than a full-paper audit, it says so.
Favorable, null, mixed, and adverse findings are all retained. No pooled estimate is calculated across the chapters of this report, because the populations, interventions, outcomes, comparison conditions, and study designs are materially different. Combining them into a single number would manufacture a precision that none of the underlying studies supports. The full evidence-class table lives in Chapter 11.
Eight interpretation rules govern every chapter:
Relevant human outcome. Evidence must bear on learning, task performance, access, welfare, work, financial decisions, health, participation, infrastructure effects, or accountability.
Traceable record. Claims require a retrievable source and an identifiable population, process, document, or observation period.
Technology fidelity. Non-generative programs may illuminate mechanisms or alternatives but are not presented as evidence that a current language model reproduces their effects.
Version fidelity. Material updates are considered without silently replacing an earlier study's sample or estimate.
Outcome fidelity. Time saved, output quality, learning, earnings, health, satisfaction, and community benefit remain different measures.
Transfer discipline. A finding is not generalized to a new age, country, occupation, product, or facility without acknowledging the gap.
Conflict visibility. Disputes and source limitations are described rather than resolved by choosing the most convenient number.
These rules mirror the seven-part test: version fidelity is change control applied to research. I hold the research to the same discipline I ask of the labs.
1.9.1 Two evidence chains that never substitute for each other#
Every chapter in this report runs along one of two evidence chains, and it helps to see them side by side before the rest of the report begins.
The first chain runs from a tool to a task to a human outcome. A tutoring program (tool) is used for practice problems (task), and the question is whether the student learns (human outcome). Most of the chapters on schooling, work, households, finance, and health live on this chain.
The second chain runs from a facility to its resources, contracts, and obligations to a community outcome. A data center (facility) draws power and water under specific agreements and approvals (resources, contracts, obligations), and the question is what happens to the community's bills, supply, tax base, and services (community outcome). Chapter 10 lives on this chain.
Figure 1.3Concept
Two evidence chains that interact but never substitute
A useful learning tool does not prove a facility benefits its neighbors, and a poorly run facility does not prove an application lacks value. Each chain needs its own evidence.
The chains interact, but they never substitute for one another. Evidence that a tool helps a student says nothing about whether a facility burdens a town's water, and evidence that a facility paid its taxes says nothing about whether the software it runs helps anyone learn. Borrowing credit from one chain to pay a debt on the other is the most common trick in debates about AI infrastructure.
Two instruments hold this report to its own standard, and both are unusual enough to describe. The first is a citation ledger: a mechanically generated index of every source-linked passage in the manuscript, each with its document, section, and line location. The ledger is an audit aid, not a second body of evidence; a source that appears in five chapters has not become five studies. The second is a source verification log: every one of the 83 distinct sources cited in this report was checked during preparation of this edition. Sources that blocked automated access were spot-checked through independent retrieval paths; sources that had died were either replaced with a canonical location of the same work or flagged. The verification record is published with the report so that readers can repeat any check, and the evidence register at the back of the report is where to start.
The reason for this apparatus is simple. This report's only currency is traceability. A claim a reader cannot check is, for this report's purposes, not a claim but a rumor with formatting. The apparatus is imperfect (a live link does not certify a true page, and a dead link does not falsify a study) but it converts the report's factual posture from an assertion into a procedure.
The same logic governs the seven Savrn trackers referenced throughout the report: a tracker entry starts a question and never ends one.
1.9.3 Falsifiability: what evidence would change these conclusions#
A research synthesis that cannot say what would prove it wrong is a position paper wearing a lab coat. This report's central claims are falsifiable in specific, nameable ways, and stating them here is part of the method.
The conditional thesis, that capability, adoption, and benefit are not interchangeable, and that defined applications can improve particular outcomes, would be weakened by credible evidence that the reviewed null and adverse results failed to replicate, or that the favorable results replicate readily across the populations and settings where this report marks transfer limits.
A well-run, adequately powered trial showing that unrestricted conversational assistance improved unaided later performance in high-school mathematics would directly contradict the interpretation this report places on the PNAS experiment; the interpretation would have to change, and the schooling chapter, Chapter 3, says so explicitly.
Conversely, if the field trials proposed in Chapter 11 repeatedly find that workflows built on the seven Savrn trackers produce no improvement in decision quality over ordinary research practice, the trackers' public-value claim should be withdrawn. Savrn's interest in that claim is exactly why the withdrawal test is worth stating in advance.
The September 2026 analysis in this chapter is falsifiable through the seven-part test itself. If either laboratory publishes the elements the test calls for (the operational safety bar, the evaluator arrangements with publication rights, evidence that the process has said no at least once), the "undisclosed" entries in any fair worksheet fill in, and the scrutiny this chapter argues for resolves in the companies' favor. That outcome is not a failure of this report's method. It is the method succeeding.
Two structural boundaries deserve restating at the outset.
First, this report is prepared for Savrn, a company with a commercial interest in the data center infrastructure industry, and it does not describe itself as an independent institutional review. I'm the founder and CEO. You should read every chapter knowing that. Sponsorship disclosure appears wherever tracker findings support an argument.
Second, "completed" describes the manuscript and its evidence package. It does not mean the proposed workflows have undergone field trials, that outside experts have peer-reviewed this report, or that any facility has passed a site-specific impact assessment.
1.10 Conclusion: A Procedural Standard for Public Trust in AI#
The public record of September 2026 supports a real question about the relationship between warnings, pacing proposals, and continued releases. The language of the reviewed proposals, in Amodei's essay and Pachocki's essay, rules out treating every later release as an automatic contradiction. Both things are true at once, and a serious analysis has to hold them together rather than choosing one.
My conclusion from that record is procedural:
Ask for a measurable commitment.
Evaluate conduct against the specific conditions of that commitment, using the seven-part test in this chapter.
Record what is said, what is released, what is evaluated, and by whom.
Let neither corporate assurances nor public suspicion substitute for evidence.
The rest of this report applies that procedural standard to the questions that matter across a human lifetime: what these systems do for a kindergartner, a high-school sophomore, a college student, a new worker, a midcareer parent, a civic committee, an investor, an older adult, and the community that hosts the physical infrastructure underneath all of it. The standard does not change from chapter to chapter. Only the decisions do.
Did OpenAI and Anthropic break their AI safety commitments in September 2026?
The record does not show that. Both September 2026 essays called for safety-conditioned pacing, not a halt, and Amodei wrote that pacing "does not mean halting model training or technical progress." Releases on September 22 are consistent with that wording if they met the stated conditions, such as third party evaluator confirmation. Whether they did is a factual question the public record, as of September 23, 2026, does not settle.
What happened to OpenAI's superalignment team?
OpenAI announced superalignment in July 2023, pledging 20 percent of secured compute over four years. According to reporting by Business Insider, Newsweek, and The Next Web, Ilya Sutskever announced his departure on May 14, 2024, Jan Leike resigned publicly on May 17, and the team was dissolved within days, at roughly the one-year mark. OpenAI stated that safety responsibilities were redistributed rather than ended.
What does AI pacing mean?
In Dario Amodei's September 2026 essay, pacing means slowing the rate of capability improvement while ensuring companies take adequate time to align and safeguard models, and for third party evaluators to confirm this. It explicitly does not mean halting training or technical progress. That makes pacing a conditional commitment: a release can be consistent with it, but only if the stated conditions are met and shown.
How can I tell if an AI company is really slowing down?
Ask two things first: slowing what, and compared with what? "Slower" could mean training scale, capability improvement, external access, deployment scope, autonomy, or release frequency, and these can move in opposite directions. Then check the safety bar, the evaluation, whether evidence ever caused a delay or cancellation, residual risk, and change control. This seven-part test is proposed, not yet applied by any regulator.
What did the 2026 Pew AI survey find about public trust in AI?
In survey data from June 2026, published by Pew Research Center in August 2026, 52 percent of U.S. adults said increased use of AI in daily life made them more concerned than excited, up from 37 percent in June 2021. Only 9 percent were more excited, down from 18 percent. The survey measures attitudes, not harm, and does not isolate any single event.
What should a responsible AI release announcement include?
This report proposes six plain-language questions: what changed as of this date, who can use it, what evidence supports the change, what remains unreliable, who is accountable by named role, and how the public can challenge an outcome with a stated response time. The first three make the claim; the last three cover failure. The standard is proposed and has not yet been tested.
Does Savrn hold itself to the same test?
Yes. Savrn's behind-the-meter power and zero-makeup-water cooling are design goals and publisher statements, not verified operating results. Run through the seven-part test today, most rows read undisclosed or not yet measurable. This report is prepared for Savrn, which has a commercial interest in data center infrastructure, and it is not an independent institutional review.
Chapter 2 · Capability vs. Benefit
Does AI Actually Improve Productivity? The Standard of Proof
Every vendor deck I have ever seen has a productivity number on it. Forty percent faster. Twice the output. Hours back every week. So the question most readers bring to AI productivity studies is simple: does AI make workers more productive, and if it does, is the number on the slide the number I will actually get? That is the right question, and it deserves a better answer than the slide gives it.
I have never bought serious equipment on a spec sheet alone. When you build large power loads, you learn early that a nameplate rating describes a machine under the manufacturer's test conditions, not the machine in your building, on your power, with your cooling, run by your crew, on the worst day of the summer. The spec sheet is real information. It is also the beginning of the diligence, not the end of it. Scaling a large site taught me to respect that gap between rated and delivered, because that gap is where an operator makes or loses money. The same logic applies here, with one difference: an AI productivity claim is usually a spec sheet for a person working with a tool, and people vary far more than machines do.
Here is what the evidence will show. The best experiments we have found real gains on some bounded tasks, real harm on at least one task that looked like it belonged in the same category, and a developer result that moved from a measured slowdown in 2025 to an unsettled picture in 2026. The uncomfortable part, for my industry and for anyone selling these tools, is that none of these studies supports a universal productivity multiplier. The useful part is that the same studies show exactly how to test a claim before you sign for it. That method is the subject of this chapter, and it is the standard of proof the rest of this report uses.
2.1 Capability vs. Benefit: What AI Productivity Studies Measure#
Every claim about what a capable system does for a person hides a chain of at least three questions.
Can the system produce good output? That is a question about the model.
Does a person, using the system for a specific task, perform that task better? That is a question about the person and the model together, in one task, on one day.
Is the person better off in learning, in earnings, in health, in time that actually becomes theirs, once every cost of using the system is counted? That is a question about a life.
Three different outcomes: model output, task performance, welfare
Where the four anchor studies sit. Each measured task performance. None measured long-run welfare such as earnings or retained skill, which is why no task result should be quoted as a life outcome.
The divergence is not an anomaly to explain away. It is the central finding of this chapter and the organizing fact for the rest of this report. A model's output quality, a user's task performance, and the user's long-term welfare are different outcomes, connected by links that can each break on their own. A system can produce fluent text that saves a professional forty minutes and still leave that professional no better at writing. A system can ace the tasks chosen to show its strengths and fail the task its users assumed fell within them. A system can speed up work in one year and slow it down in another, for reasons that have as much to do with the measurement as with the model.
Capability is not benefit. I will say that a few more times in this report, because almost every public argument about AI slides from the first question to the third without stopping at the second.
So the operating question is not whether a system is powerful. It is whether a specific person, using a specific configuration, for a specific task, achieves a better outcome after verification, correction, supervision, and every other cost is included.
That question is answerable. The experiments below answer it, in both directions, for four task environments. Everything else in this chapter (the benefit chain, the cost framework, the adoption process, the contract checklist) is machinery for keeping that question answerable when a vendor, a school board, an employer, or a reader's own enthusiasm tries to collapse it back into "is the system powerful?"
A note on scope. This chapter reviews task-level and workplace experiments because that is where the causal evidence is strongest. It does not review benchmark results, capability demonstrations, or model-card claims, because those answer the first question (can the system produce good output?) rather than the second and third. Benchmarks have their place. That place is not a claim about human benefit.
2.2 What the Best AI Productivity Studies Actually Show#
Four study programs anchor this chapter. Each one is presented with its design, its finding, and its boundary: the sentence that says what the result does not establish. The boundaries are not hedges. They are the load-bearing part of the evidence. A finding without its boundary is how a careful study turns into a marketing number.
2.2.1 The ChatGPT Writing Study: Noy and Zhang (2023)#
The cleanest favorable result in the reviewed set comes from a preregistered experiment published in Science in 2023. Shakked Noy and Whitney Zhang assigned 453 college-educated professionals to occupation-specific writing tasks (press releases, delicate emails, the kind of short document that fills a knowledge worker's week) and randomly gave half of them access to ChatGPT. (Noy and Zhang, Science)
Here's what the study actually found. The result was immediate and large. Participants with access completed the tasks in 40 percent less time on average, and independent evaluators rated their output 18 percent higher in quality.
The distribution matters as much as the average. The tool compressed the performance distribution: the least productive writers improved the most, so the gap between the bottom and the top of the pool narrowed. A follow-up survey found that participants in the assisted group were about twice as likely as the control group to report using the tool in their actual jobs two weeks later (34 percent versus 18 percent), a difference that persisted at two months.
Now the boundary, and it has four parts.
First, these are task outcomes. They are not measured changes in annual earnings, employment, or total organizational productivity. The experiment measured one bounded writing session.
Second, the study did not administer a later writing test without the tool, so it cannot establish that participants became more capable writers. What it established is that the person-plus-tool system, under the experiment's conditions, outperformed the person alone. Those are different findings, and the difference is the whole subject of Section 2.4.
Third, whether the saved time became output, leisure, or simply more tasks is a question the experiment deliberately left open. Forty percent less time on a press release is not forty percent more value for the organization. It is forty percent of that task's time released into something, and what that something is was not measured.
Fourth, the result describes GPT-3.5-era assistance on short, self-contained documents in 2023. It is a historical estimate for that configuration, not a permanent property of "AI."
Common misreadings. "ChatGPT makes workers 40 percent more productive" is the version that shows up in decks. It is wrong on three counts: the population was professionals doing short writing tasks, the outcome was time on that task, and the tool was a 2023 version.
What this means for you. If you are an employer looking at drafting tasks that resemble these (short, self-contained, low consequence, easy to review), this is the strongest evidence in the reviewed set that assistance can help, and it justifies a controlled test of your own. If you are a worker, the result supports using the tool for drafts you are going to review anyway. It does not support the idea that the tool is making you a better writer; nobody tested that.
2.2.2 The Jagged Frontier Study: Dell'Acqua and Colleagues#
The second study program is the one this report relies on most heavily, because it was designed to find where assistance fails. Fabrizio Dell'Acqua and colleagues ran a field experiment with 758 consultants at Boston Consulting Group, with tasks deliberately chosen on both sides of the tested model's capability frontier. Some tasks were ones the model demonstrably handled well. One complex analytical task was selected because the researchers judged it outside the model's reliable competence. (Dell'Acqua et al., working paper)
On the inside-frontier tasks, the pattern resembled the writing experiment. Assisted participants completed approximately 12.2 percent more tasks and worked roughly 25 percent faster, with higher assessed quality.
On the outside-frontier task, the result inverted. The working paper reports correct answers for:
approximately 84.4 percent of the control group, which had no tool;
70.6 percent of the group with model access;
60 percent of the group that received model access plus an overview of the model's capabilities.
Read that again. Participants who knew the most about what the tool could do did the worst. They did worse than the group that simply had access, and dramatically worse than the group with no tool at all.
Figure 2.2Data
Inside the frontier assistance helped; outside it, the most-briefed group did worst
On tasks inside the tested model's frontier, assisted consultants completed about 12.2% more tasks and worked about 25% faster with higher assessed quality. On the one task chosen to sit outside it, correctness fell from 84.4% to 70.6% to 60%. Figures are from the reviewed working-paper version.
The boundary: the model, the tasks, and the evaluation design are bounded. A successful creative assignment does not establish reliable quantitative or strategic judgment.
But the study's lasting contribution is not the favorable number. It is the demonstration that the sign of the effect depends on the fit between task and tool, and that the users least equipped to know where the frontier lies (the ones who trusted the tool on the hardest task) absorbed the harm. The researchers' phrase for this pattern, the "jagged frontier," has been widely adopted, sometimes as a slogan. In the study itself it is a measured result: assistance helped on the tasks the researchers selected as inside, and hurt on the task they selected as outside.
That distinction matters because the slogan version gets used to excuse failures after the fact ("well, that task was outside the frontier"). The study did something harder. It picked the tasks in advance, labeled which side of the line each one sat on, and then measured. If you want to use the jagged frontier idea responsibly, you have to do the same thing: decide before the pilot which tasks you believe are inside, and test the ones you are least sure about.
A version note. This belongs here because it shows a discipline this report applies everywhere. The consulting research was subsequently published in Organization Science, according to Harvard Business School's account of the publication, and the published version carries its own figures. This chapter labels the numerical estimates above as originating in the reviewed working-paper version and does not silently combine them with different numbers in the later institutional summary. Later publication may clarify a finding. It must not become an excuse to choose whichever estimate makes the stronger promotional statement.
Common misreadings. "AI makes consultants 25 percent faster" drops the outside-frontier task entirely. "Training people on AI makes them worse" overreaches in the other direction; Section 2.3 explains what the overview result does and does not show.
2.2.3 The METR Developer Study: 2025 and the 2026 Update#
The third study program is the adverse result, and this report treats it with the same care as the favorable ones.
In early 2025, the research organization METR ran a randomized study of 16 experienced open-source developers working on 246 tasks in repositories they knew well. The developers were paid for their time and randomly assigned to work with or without early-2025 AI tools. The result: assistance increased completion time by 19 percent, with a reported interval of 2 percent to 39 percent longer. (METR, early-2025 study)
This is a small study, and its authors said so. Sixteen developers, one task environment, one generation of tools, one point in a fast-moving field. It is a historical result for that sample and setup, not a universal or current verdict on coding assistance.
But it is a randomized result, on a real task environment, with experienced professionals, which is exactly the population enthusiasts assumed benefited most. It earned its place in the public record by contradicting that assumption, and it complicates every productivity number quoted without a date attached.
In February 2026, METR reported a follow-up with 57 developers and more than 800 tasks. The point estimates moved toward speedups: approximately 18 percent less time for returning participants and 4 percent less for newly recruited participants. Their confidence intervals, however, ran from -38% to +9% for returning participants and from -15% to +9% for new ones. Both intervals include no effect.
The authors also identified selection and measurement problems that limit a reliable current estimate. Some participants avoided contributing tasks they did not want to perform without assistance, and simultaneous or asynchronous work complicated the measurement of time. The later evidence neither erases the earlier randomized result nor supplies an uncomplicated replacement productivity number. (METR, February 2026 update)
Figure 2.3Data
The moving developer estimate: why productivity numbers need dates
The 2025 randomized result showed a slowdown. The 2026 update pointed toward speedups, but both intervals include no effect and the authors say measurement problems prevent a reliable current estimate.
The boundary runs in both directions. The 2025 slowdown cannot be quoted as the current state of coding assistance. The 2026 update cannot be quoted as a settled speedup, and it should not be presented as a confirmed reversal.
What the pair establishes is subtler and more useful. The measured effect of assistance on expert work moves with the tools, the task selection, and the denominator. That is why Section 2.5 treats every productivity estimate as dated, and why Section 2.6 requires retesting after material changes.
Look closely at the selection problem, because it is not a technicality. If participants could steer away from contributing the tasks they did not want to do without the tool, then the pool of tasks being compared is no longer the pool of work the developers actually do. The tasks where the tool was most wanted may be the ones missing from the comparison. That changes what the estimate means, and it changes it in a direction you cannot read off the headline.
What this means for you. If you manage engineers, neither METR number is your number. Both are strong arguments for measuring your own team, on your own repositories, with a comparison, and for writing down the tool version when you do.
2.2.4 The Four AI Productivity Studies Side by Side#
Set side by side, the four programs show a pattern no single one of them shows. (The two METR eras appear as separate rows because they are one program measured in two periods.)
Point estimates moved toward speedups; intervals include no effect; selection and measurement problems limit a reliable estimate
Neither erases 2025 nor settles a current number
Three lessons survive the differences.
First, the effect of assistance is task-specific, not system-general. The same population of capable professionals can be helped, hurt, or unaffected depending on the task in front of them. The consulting study shows this inside a single sample.
Second, the users most at risk are often the ones with the most trust in the tool. In the consulting study, the harm concentrated among participants who had been taught the tool's strengths.
Third, every estimate in the table is dated. None of them is a property of "AI." Each is a property of a configuration (model version, task set, population, measurement method) that existed at a specific time and may not exist now.
A fourth lesson is about the reader's own inference habits. The favorable studies are the ones most often quoted, and the adverse one is most often quoted with its sample size attached while the favorable ones go bare. This report applies the same boundary discipline in both directions. The writing experiment's 40 percent is a task-level finding with no earnings evidence behind it. The METR slowdown is one randomized study of sixteen developers. Both are exactly as strong as their designs and no stronger.
The consulting experiment's outside-frontier result deserves its own section, because it isolates the failure mode that matters most for every high-consequence setting in this report: the moment when a tool's output is fluent, confident, and wrong.
The numbers again: 84.4 percent correct in the control group, 70.6 percent with model access, 60 percent with model access plus a capability overview. The gap between the second and third groups is the finding that should keep institutional designers up at night. Training about the tool, the intervention almost every responsible-adoption program begins with, made performance worse on the task where the tool was unreliable. Participants armed with an overview of the model's strengths trusted it on a task those strengths did not cover.
2.3.1 What the Training Result Does and Does Not Show#
The result does not establish that training is useless, and this report does not draw that conclusion. It shows why the particular training, user behavior, task difficulty, and evaluation conditions must be tested rather than treated as generic safeguards.
Training that teaches what a tool can do, without teaching where it stops, manufactures the exact overconfidence the overview group displayed. A training program that wants to avoid this failure mode has to do something harder than listing capabilities. It has to show users the frontier from the wrong side, with tasks where the tool fails visibly, so that the user's calibrated distrust is built on experience rather than caveats.
In practice, that suggests a different kind of onboarding. Instead of a slide of use cases, staff work through examples from their own job where the tool produces a confident wrong answer, and they practice catching it. Whether that design works is itself a hypothesis; it has not been tested in the studies reviewed here. But it is the design the evidence points toward, and it is testable.
2.3.2 Plausible Output Is Cheap; Verified Output Is Not#
The same logic extends beyond training to output itself. Plausible output is cheap. The model that drafts a fluent paragraph, a confident diagnosis, a polished financial recommendation, or a clean block of code is producing exactly what it was built to produce: text that reads like competence.
Whether the competence is present is a separate question, answerable only by verification against something that is not the model: a source, a calculation, a measurement, a domain expert. That is why the benefit chain in the next section places task performance, not output quality, at its second link, and why the verification burden appears in the cost accounting in Section 2.5 as a first-class expense rather than an afterthought.
Here's the part most people skip. Verification is not free, and it does not scale with the tool. It scales with the person doing it and the consequence of getting it wrong. A one-paragraph internal email needs a glance. A benefits determination, a medication instruction, or a structural calculation needs someone qualified, with time, checking against an independent source. If the business case assumes the glance and the task needs the check, the savings are fictional.
There is a school-setting analogue that the education chapters develop fully, starting with Chapter 3. A student whose assisted practice looks excellent while unaided competence quietly erodes is producing the same kind of plausible output, and the erosion is invisible until the assistance is removed. The mechanism is identical. The cost of discovering it is paid later, by someone else, usually in a higher-stakes setting.
For parents and teachers: the signal to watch is not how good the homework looks. It is whether the student can explain it, or do a similar problem, with the tool closed. That is the independent-capability test in Section 2.4, and it is the one most often missing.
How, then, should an institution reason from "this system is capable" to "this system benefits our people"? This report proposes a chain of six links, each of which needs its own evidence rather than borrowing certainty from the previous one.
Link
Question
Example measure
Access
Can the intended person use the service?
Successful completion across language, disability, device, and connectivity conditions
Task performance
Does the workflow improve the immediate task?
Accuracy, completion time, omissions, correction burden
Independent capability
Can the person still understand or perform without it?
Unaided assessment, explanation, transfer to a new problem
Real-world outcome
Does the change matter beyond the task?
Learning retained, completed benefit application, resolved service request
Distribution
Who benefits and who bears costs?
Outcomes by relevant subgroup, including nonusers
Durability
Does the benefit persist as conditions change?
Follow-up outcome, version retest, staff turnover sensitivity
Figure 2.4Framework
The six-link benefit chain, and where the reviewed evidence reaches
An editorial framework, not a validated scale. Each link needs its own evidence. The highlighted link is the one pilots skip most often because skipping it makes results look better.
The chain is an editorial framework for this report, not a validated universal scale. Its purpose is diagnostic. A task can improve at the second link while failing at the third or fourth, and the framework's job is to make that failure visible rather than letting the second-link success stand in for the whole chain.
Access. Access failures are invisible in averaged results. A tool that works well for English-speaking users with new laptops and high-bandwidth connections may not exist at all for the population an institution actually serves. A pilot that never measures completion across language, disability, and connectivity conditions will never know.
Task performance. These are the failures the reviewed experiments measure best. All four study programs in this chapter live mainly at this link.
Independent capability. These are the failures the experiments miss most often. The writing experiment had no unaided retest, and the education chapters show how consequential that omission can be when the population is a student rather than a professional.
Real-world outcome. This is where task gains become, or fail to become, a completed application, a resolved ticket, a retained skill. A faster draft is a task gain. A benefit application that gets approved is a real-world outcome.
Distribution. This is where averages hide. The customer-support study in Chapter 5 shows large average gains coexisting with negligible effects for the most experienced workers. Any institution that adopts on the average is making a decision about real individuals while looking at a statistic.
Durability. This is where dated estimates live. The METR pair is, in effect, a measurement of the sixth link, and its lesson is that a benefit measured once is a benefit measured once.
2.4.2 Where the Reviewed Evidence Actually Reaches#
Place the four programs on the chain and the gaps become obvious. All of them live mainly at link two. The writing study reaches toward link four only through self-reported job use, and it does not test link three. The consulting study's outside-frontier task is a direct measurement of a link-two failure, with a distribution result attached: the overview group fared worst. The METR pair, taken together, says something about link six, because the measured effect did not hold still across tool generations and study designs. None of the four measures access across language, disability, or connectivity, and none follows real-world outcomes such as earnings or organizational productivity. That is not a criticism of the studies. It is a map of what is still open.
The chain also explains this report's structural rule against pooled estimates. A finding that stops at link two and a finding that reaches link four are not two votes for the same proposition. Combining them into one average effect would manufacture a precision that neither supports. Averaging the writing study's 40 percent time reduction with the METR slowdown would give you a number that describes no task, no population, and no tool.
Here's a hypothetical, to show the chain in use. It contains no real statistics.
A county benefits office is considering an assistant that drafts responses to residents' questions about program eligibility. The vendor presents a time-savings figure from its own pilot.
Access: Does the pilot include residents who write in languages other than English, use screen readers, or contact the office by phone? If the pilot only measured web users, link one is untested.
Task performance: Did caseworkers produce accurate responses faster, counting the time they spent checking the draft against current program rules? If review time was not logged, link two is only half measured.
Independent capability: Can a newer caseworker still explain an eligibility rule without the tool? If the office plans to rely on the assistant for training new staff, this link matters a great deal.
Real-world outcome: Did more residents complete applications, or get correct answers the first time? Faster drafts that lead to more resubmissions are not a benefit.
Distribution: Did residents with complex cases get worse service while simple cases got faster? An average can hide that.
Durability: What happens when the vendor updates the model or the program rules change? Is there a retest scheduled?
The vendor's time-savings figure, whatever it is, answers part of link two. The office's decision depends on all six.
2.5 AI Total Cost of Ownership, Not Subscription Price#
The cheapest number attached to a capable system is its subscription price, and it is almost never the relevant one. The proposed cost boundary for any adoption decision includes:
acquisition;
devices and connectivity;
integration;
staff training;
verification;
correction;
human escalation;
accessibility;
security;
exit.
These are analytical categories. This report does not assign a universal dollar amount to them, because the amounts are institution-specific in ways that matter.
Figure 2.5Framework
Total cost is a stack; the subscription is one layer
Proposed cost boundary for any adoption decision. Verification and correction are first-class costs. Saved time is capacity, not cash, unless staffing or spending actually changes.
Two quantities should remain separate, and their separation should be visible in every internal business case.
Net staff time saved equals baseline staff time minus assisted staff time, where assisted staff time includes review and correction.
Cost per additional successful outcome equals the assisted total cost minus the comparison total cost, divided by the assisted successful outcomes minus the comparison successful outcomes.
The second calculation is meaningful only when the outcome definitions, periods, and populations are comparable and the denominator is positive and credibly estimated. If the assisted approach does not produce more successful outcomes than the comparison, the denominator is zero or negative and the ratio tells you nothing useful. If a program costs less but produces fewer successful outcomes, the result should be reported as a trade-off, not compressed into an attractive savings statistic. A school that spends less per student but graduates fewer of them has not discovered efficiency.
Show me the denominator. That is the single most useful sentence you can say in a procurement meeting. A cost-per-outcome figure without a clearly defined, comparably measured count of successful outcomes is not a figure. A number without a denominator is a rumor.
2.5.2 The Two Accounting Errors Behind Most Inflated Claims#
Two accounting errors account for most inflated claims.
The first is treating saved time as cash. Unless staffing, purchased services, or another expenditure actually changes, saved time is capacity, not savings. Whether that capacity becomes productive output, better work, or absorbed slack is another outcome to measure.
The second is omitting verification and correction from the assisted-time column. The writing experiment's participants submitted lightly edited or unedited model output at high rates. In low-consequence settings, that behavior is the source of the time saving. In high-consequence settings, it is the source of the errors. An institution that counts the draft and not the review is running a cost account that would fail an audit in any other category of expenditure.
The denominator problem has a measurement dimension as well, and the METR follow-up illustrates it. When participants can choose which tasks to contribute, and when work happens simultaneously or asynchronously, the time denominator stops meaning what it appears to mean. A productivity evaluation can change meaning when the task pool or the time measurement changes, and the change is invisible unless the measurement method is reported alongside the estimate. (METR update)
The same error shows up in every sector, under different names:
For a school, it is comparing only students who voluntarily used a tutor with everyone else.
For a civic committee, it is comparing a polished machine summary against an unfinished manual draft rather than against a completed, quality-controlled alternative.
For an employer, it is measuring only the staff who kept using the tool after the pilot and calling their numbers the effect of the tool.
Every one of these is a denominator substitution, and every one of them inflates the apparent benefit in the same direction.
2.5.4 A Maintenance Rule for Every Productivity Number#
This report proposes a maintenance rule for every productivity claim: attach tool version, task, worker experience, observation period, and evaluation design to every estimate, and recheck after substantial product or workflow changes rather than treating a historical study as a permanent coefficient. A productivity number without a date is not information. It is a fossil.
Given all of the above, how should an institution actually adopt? This report proposes a process that begins with one consequential task rather than an organization-wide promise. The owner documents the baseline, chooses a non-model alternative, defines the outcome and stopping rules, and tests a limited configuration before scaling.
The process has seven requirements.
Specify the task. Identify what the tool may do and what it may not decide. A drafting assistant has a different risk profile from a decision system, and the specification must say which one is being adopted.
Preserve a comparison. Use random assignment where practical, or a transparent comparison that acknowledges confounding. The four study programs in this chapter earn their authority from their comparisons; an adoption process without one earns nothing.
Measure corrections. Count factual errors, missing exceptions, source failures, and the time spent repairing them. The correction log is the institution's version of the consulting study's outside-frontier task: it is where the true cost of fluent output becomes visible.
Test unaided competence. Include it where education, professional judgment, or safety depends on retained skill. This is the third link of the benefit chain, and it is the one most often skipped because skipping it makes the pilot look better.
Check distribution. Examine whether language, disability, experience, or access changes results. The average is a policy decision about individuals; the subgroup results are the actual individuals.
Include nonusers. Preserve a functional route for people who cannot or do not choose to use the system. An adoption that removes the non-digital route has coerced adoption, whatever the consent language says.
Retest material changes. A new model, prompt, retrieval source, or permission can alter the evaluated intervention. The METR pair is the standing example: the intervention changed, the estimate changed, and the only right response was to remeasure.
A stopping rule is a condition, written before the pilot starts, that ends or pauses the pilot. Examples of the kind of rule an owner might write (as categories, not thresholds): the correction log shows errors of a type the task cannot tolerate; a subgroup does measurably worse than under the comparison; the vendor changes the model and no retest is scheduled. The point of writing them in advance is the same as the point of the consulting study labeling tasks in advance. It stops you from deciding after the fact that the bad result did not count.
NIST's AI Risk Management Framework and its generative AI profile support an approach grounded in governance, context, measurement, and risk management rather than performance testing alone. They are voluntary guidance, not empirical proof that a particular deployment is effective or compliant, and this report uses them as structure, not as certification.
If a vendor tells you its product is "aligned with the NIST framework," that is a publisher statement about process. It does not tell you the product works for your task. The seven steps above are how you find out.
2.6.3 Why the Process Is Deliberately Unglamorous#
No part of this process generates a headline, and all of it generates records. That is the point. An institution that has run this process on one consequential task has something better than confidence: it has a baseline, a comparison, a correction log, and a stopping rule. Those four artifacts are what the word "evidence-based" means when it is meant literally.
2.7 Before You Sign the AI Contract: A Buyer's Checklist#
Everything in this chapter can be turned into questions a buyer asks before signing. This checklist is derived only from the requirements above. It is a proposed workflow, not a tested instrument, and it is meant to be printed and brought into the meeting.
The task and the claim
[ ] Which specific task will the tool perform, and what is it not allowed to decide? Is it a drafting assistant or a decision system?
[ ] Which productivity or quality figure is the vendor citing, and which of the three outcomes (output quality, task performance, welfare) does it actually measure?
[ ] What tool version, task, worker experience, observation period, and evaluation design produced that figure, and when?
[ ] If the figure comes from a published study, which version of the study? Does the vendor mix figures from different versions?
The comparison
[ ] What was the comparison group, and what did it actually receive (no tool, a different tool, a completed manual process)?
[ ] Was assignment random? If not, what confounding is acknowledged?
[ ] Could participants choose which tasks to include? Were only continuing users counted?
The frontier
[ ] Which of our tasks do we believe are inside the tool's reliable competence, and which are we unsure about?
[ ] Will our pilot include tasks where we expect the tool to fail, so staff learn to recognize failure?
The costs
[ ] Have we priced acquisition, devices and connectivity, integration, staff training, verification, correction, human escalation, accessibility, security, and exit?
[ ] Does our assisted-time figure include review and correction time?
[ ] Is any "saving" in the business case tied to a real change in staffing, purchased services, or other spending? If not, have we labeled it capacity?
[ ] Is the cost per additional successful outcome calculated with comparable outcome definitions, periods, and populations, and a positive, credible denominator?
The people
[ ] Will we test whether staff (or students) can still perform without the tool where skill, judgment, or safety depends on it?
[ ] Will results be broken out by language, disability, experience, and access?
[ ] Is there a working route for people who cannot or choose not to use the system?
The contract terms
[ ] Will the vendor notify us before changing the model, prompts, retrieval sources, or permissions?
[ ] Do we have the right to retest after a material change, and a plan to do it?
[ ] What are our stopping rules, written before the pilot starts?
[ ] What does exit cost, and can we get our data out?
[ ] Is any governance or framework claim (for example, alignment with NIST guidance) presented as a publisher statement about process rather than proof of effectiveness?
One more instrument belongs in the standard of proof, because the chapters that follow cite many kinds of evidence and the kinds are not interchangeable. This report uses an eight-class editorial labeling system throughout. It is an editorial convention, not a formal evidence-grading methodology such as GRADE, and its purpose is communicative: to keep each claim attached to what its underlying evidence can actually support.
The eight classes are randomized comparison, quasi-experimental analysis, observational or survey evidence, technical evaluation, simulation or forecast, administrative or official record, publisher statement, and proposed workflow. The full table, with what each class can and cannot establish on its own, lives in Chapter 11.
For this chapter, three points are enough.
The writing, consulting, and METR studies are randomized or experimental comparisons. They can establish an effect of the assigned intervention under their designs. They cannot establish universal transfer or long-term benefit on their own.
Publisher statements (model cards, announcements, vendor white papers, and Savrn's own design goals) are legitimate evidence about what an organization says. This report cites them for that purpose and never for more.
Proposed workflows, the class to which most of this report's own recommendations belong (including the six-link chain, the seven-step process, and the checklist above), carry a standing label: implementable and testable, not yet tested here. That label is what separates a research agenda from a marketing deck.
The studies in this chapter are about professionals, consultants, and developers, but the method applies to everyone who is being told AI will make something better. Here is how it translates.
Workers. The best evidence says assistance can speed up bounded tasks that resemble the ones tested, and that it can hurt on tasks outside the tool's reliable competence, especially if you trust it because you have been told what it is good at. Keep a habit of checking output against something that is not the tool. If your employer measures your output with the tool, ask how saved time will be used: more work, better work, or a change in your role.
Employers and managers. Do not adopt on the average. Pick one consequential task, run the seven steps, and write down the tool version. If the business case says "hours saved," ask where the hours go and whether any spending changes.
Board members and investors. Ask which of the three outcomes a productivity claim measures, and what the comparison received. Treat any figure without a date and a study design attached as a publisher statement. Ask whether verification and correction are in the cost model.
County commissioners and residents. When a public office proposes an AI tool, ask whether residents who cannot or will not use it keep a working route, and whether results will be reported by language, disability, and access. Those are links one and five of the chain, and they are the ones most likely to be skipped.
The reviewed evidence does not support a universal productivity multiplier. It supports testing bounded tasks and retaining both favorable and unfavorable results, including material changes in later evidence. That conclusion rests on the writing trial, the consulting trial, and the developer studies together, not on any one of them.
For the rest of this report, "value" means an identified outcome for an identified person or institution, with a defensible comparison and explicit costs. Faster or more convincing output is evidence of value only when it advances that outcome.
This definition is stricter than the one in most public discussion, and it is stricter on the enthusiasts than on the skeptics. The skeptic's position ("we don't know yet") is often the accurate state of the evidence. The enthusiast's ("this changes everything") bears the full burden of the six-link chain. I build infrastructure for this industry, and I am telling you the burden sits on our side of the table. That's the whole game.
The chapters that follow apply this standard across the lifespan: what the evidence supports for a kindergartner, a high-schooler, a college student, a new worker, a midcareer professional, a household, a civic committee, an investor, an older adult, and the community under the infrastructure. In each setting, the questions are the ones this chapter has built. What is the task? Who is the person? What is the comparison? What does it cost, all in? And does the benefit survive when the assistance is gone?
On some bounded tasks, yes. In a preregistered experiment, 453 professionals with ChatGPT finished writing tasks 40% faster with 18% higher rated quality. But a consulting experiment found assistance lowered correct answers on a task outside the model's competence, and a 2025 developer study found a 19% slowdown. The effect depends on the task, the tool version, and the person, so no single productivity multiplier is supported.
What did the Noy and Zhang ChatGPT writing study find?
In a preregistered experiment published in Science in 2023, 453 college-educated professionals given ChatGPT completed occupation-specific writing tasks in 40% less time, and evaluators rated their output 18% higher. The weakest writers improved most. The study measured one bounded session with GPT-3.5-era tools, with no unaided writing retest and no measured effect on earnings, employment, or organizational productivity.
What is the jagged frontier study?
It is a field experiment by Dell'Acqua and colleagues with 758 Boston Consulting Group consultants. On tasks inside the model's capability frontier, assisted consultants completed about 12.2% more tasks and worked about 25% faster. On a task outside it, correct answers fell from 84.4% without the tool to 70.6% with it and 60% with it plus a capability overview, per the working paper.
Did the METR developer study show AI slows programmers down?
METR's randomized 2025 study of 16 experienced open-source developers on 246 tasks found early-2025 AI tools increased completion time by 19%, with an interval of 2% to 39% longer. It is a small, historical result for one setup and one generation of tools. It is not a current verdict on coding assistance, and later evidence has not settled a replacement number.
What did METR's 2026 update show?
The February 2026 follow-up with 57 developers and more than 800 tasks moved toward speedups: about 18% less time for returning participants (interval -38% to +9%) and 4% less for new ones (interval -15% to +9%). Both intervals include no effect, and the authors flagged selection and measurement problems, so it is not a confirmed reversal.
How do I calculate the total cost of ownership for an AI tool?
Count acquisition, devices and connectivity, integration, training, verification, correction, human escalation, accessibility, security, and exit, not just the subscription. Keep net staff time saved (baseline time minus assisted time, including review) separate from cost per additional successful outcome. The second figure only works with comparable outcomes and a positive, credibly estimated denominator. This report assigns no universal dollar amounts.
Is time saved by AI the same as money saved?
No. Unless staffing, purchased services, or another expenditure actually changes, saved time is capacity, not cash savings. That capacity may become more output, better work, or absorbed slack, and which one happens has to be measured. Business cases that multiply hours saved by an hourly rate, while leaving out review and correction time, overstate the benefit.
Why can training people on AI tools make results worse?
In the consulting experiment, the group that got model access plus an overview of the model's capabilities had the lowest share of correct answers (60%) on the task outside the model's competence, below access alone (70.6%) and no tool (84.4%). This does not show training is useless. It shows training that teaches only strengths can build overconfidence, so training designs must be tested.
Chapter 3 · Schooling, K-12
Does AI Help Students Learn? What 10 K-12 Studies Actually Show
It usually happens at the kitchen table. Your kid has a laptop open, an AI tool in one tab and the homework in another, and the answers are coming fast. Too fast. And you ask yourself the question every parent I know is asking right now: is this tool helping my kid learn, or is it doing the homework for them? That's the question this chapter takes to the AI in education research, and I want to answer it the way I'd answer any question about a system I was about to bet money on: by reading what the studies actually did, not what the brochures say they did.
Here's why it matters. Schools are buying. Districts are signing contracts, teachers are being handed new tools, and families are being told their children will get "personalized learning." Some of that is backed by real evidence. Some of it is not. And the difference between the two is often invisible unless you know exactly where to look: who was in the study, what the comparison group got, whether the kid was tested with the tool or without it, and whether anyone independent checked the work.
I reviewed ten study programs spanning kindergarten through high school. The good news is real. Structured systems, tools that support the teacher rather than replace the teacher, and supervised programs have produced measurable benefits for real students. The uncomfortable part is also real, and it's the part most people skip. In the best-designed high-school experiment in this set, an unrestricted chatbot made students look much better on practice and then left them worse off than students who never had it, once the tool was taken away. Same subject. Same kind of students. The guardrails made the difference.
So this chapter will not tell you AI is good for kids or bad for kids. It will tell you which arrangements have been tested, what they showed, where they stop, and what questions to ask before your school spends money or your child spends hours. Capability is not benefit. A powerful model sitting in a browser tab is not the same thing as a child who can now do the work on her own.
3.1 How to Read AI in Education Research Without Getting Fooled#
Before we get to a single study, you need to know what kind of review this is and what rules I'm using to read the numbers. I'm spending time here on purpose. Most bad decisions about educational technology don't come from bad studies. They come from good studies read badly.
3.1.1 Scope: Three Grade Bands, Ten Study Programs#
This chapter reviews the evidence for assisted learning across the three school bands that organize most adoption decisions: kindergarten through grade 4, grades 5 through 8, and grades 9 through 12. It examines ten study programs. I prioritized randomized field evaluations, independently measured outcomes, original research reports, and, where one exists, an independent evidence review. Policy guidance shows up later in the chapter, but only as guidance on safeguards. It is never used as proof that a tool improves learning.
The evidence is international. The studies come from the United States, the United Kingdom, Nigeria, and Turkey-connected research settings, and I don't assume that a result transfers unchanged across curricula, languages, staffing arrangements, or school systems. A British "Year 7" is not automatically an American seventh grade. A Nigerian first-year senior secondary student is not automatically an American ninth grader. Where the source population matters to how you should read a result, I keep it attached to the result.
This is a targeted review of decision-relevant research. It is not a systematic review, and I'm not claiming to have located every relevant study in the world. I also don't calculate a pooled effect size, because the interventions, comparison conditions, ages, subjects, and outcomes differ too much for an average to mean anything. Averaging a kindergarten reading package with a high-school chatbot experiment would produce a number, and that number would be a rumor. Positive, null, uncertain, and adverse findings are all kept, with the same weight they carry in the original work.
Four interpretation rules govern every number in this chapter. Learn them once and you'll read every education headline differently for the rest of your life.
Standard deviations are not percentages. A result of 0.23 standard deviations is a standardized score difference, not 23 percent more learning.
Percentage points are not percent. A pass rate rising from 62 percent to 66 percent is a four-percentage-point increase, not a 66 percent learning gain.
Relative percentages need their denominator. A reported 17 percent reduction in an exam score is not a 17-percentage-point change unless the study defines it that way.
Assignment is not use. The effect of offering a program differs from an estimate for people who actively use it. Selecting only the active users invites selection bias unless the analysis deals with it.
There's one structural rule on top of those four. Software, extra instruction, teacher training, devices, supervision, and curriculum changes usually arrive together as a bundle. A result for the package is not automatically a result for the model alone. Nearly every study in this chapter bundles something, and the bundle is part of what each result means.
3.2 What AI in Education Research Shows Across Three Grade Bands#
So the practical question for a school, a district, or a parent is not whether to embrace or reject a technology category. It's which combinations of students, teachers, instructional methods, tools, and safeguards produce independently measured benefits that justify their costs. That's a harder question. It's also the only one worth asking.
Here are the ten study programs, sorted by band:
Grade band
Study programs
Direction of evidence
Kindergarten through grade 4
Kindergarten reading (A2i package); early mathematics applications; literacy engagement support
Positive for teacher-directed and bounded packages; engagement without established literacy gains
Grades 5 through 8
ASSISTments; Tutor CoPilot; Khanmigo remediation package; teacher preparation time
Small or qualified gains; session mastery without established annual attainment; bounded teacher-time savings
Grades 9 through 12
PNAS mathematics assistants; Nigerian English program; Cognitive Tutor Algebra I
Ten study programs by grade band and type of assistance
Where each program sits. Tutor CoPilot also includes grade 3 and 4 sessions. The PNAS experiment's constrained tutor largely avoided the unrestricted assistant's penalty. Direction summarizes the primary finding; see the text for certainty and limits.
The sections that follow take each band in turn, study by study, with the design, the finding, and the boundary attached to each. One framing note first. The schooling debate is usually staged as a contest between "innovation" and "caution." The evidence I reviewed supports neither banner. It supports a procurement discipline: name the learner, the intervention, the comparison, the outcome, and the cost, and then look at what the comparison actually showed. That's the whole game.
3.3 AI Tutoring Studies, Kindergarten Through Grade 4#
The early-grade evidence is most useful when it tells you three things: what role the adult plays, what skill is being taught, and how limited the child's direct interaction with the tool is. The studies in this band test assessment-informed teaching, bounded mathematics applications, and a literacy platform with and without human engagement support. They do not test what happens when a young child is handed an unrestricted conversational assistant, and I won't pretend they do.
That gap matters. If you're the parent of a six-year-old and someone tells you "the research shows AI helps young children learn," the right response is: which AI, doing what, with which adult, measured how? In this band, the answers are always some version of "a structured tool, working through or alongside a teacher, measured outside the tool."
3.3.1 Kindergarten Reading: The Tool Behind the Teacher#
The strongest early-grade result in the reviewed set comes from a 2011 cluster-randomized study of Individualized Student Instruction for Kindergarten. The evaluation involved 556 students, 44 teachers, and 14 schools. Students in intervention classrooms outperformed comparison students on a latent measure of reading skills, with a reported effect size of 0.52. (Al Otaiba and colleagues, original study record)
Here's the part most people skip: the intervention. It used assessment-informed recommendations associated with Assessment2Instruction (A2i), combined with professional development and classroom support, to help teachers individualize instruction. The system recommended instructional amounts and groupings. The teacher still organized instruction, interpreted student needs, and worked directly with children. Information flowed from student assessments, to recommendations for the teacher, and then to differentiated classroom instruction. This was not an open chatbot teaching a kindergartner on its own. (Connor, research synthesis and intervention description)
Four boundaries attach to the 0.52 estimate, and I'd want every one of them in any school board presentation that cites it.
First, it concerns a latent reading construct, a statistically estimated measure of reading skill. It is not a universal improvement across every reading test or every early-grade student.
Second, the publicly available study abstract does not report a confidence interval for the estimate. Without an interval, you can't see how much uncertainty surrounds the 0.52.
Third, the result supports the tested instructional package including adult implementation, not a software-only effect. A school considering a similar approach should test the complete delivery package, including the time and training required for teachers to act on the recommendations.
Fourth, a disclosure belongs in the record: the later A2i synthesis discloses the author's equity interest in Learning Ovations, which matters when assessing the broader presentation of the intervention's evidence. (Study abstract, Disclosure in the synthesis)
A disclosed interest does not make a study wrong. It tells you where to look harder. I'd say the same about anybody's claims, including ours.
The defensible practical reading is that intelligent support can operate behind the teacher rather than in place of the teacher. That's a useful finding. It's also narrower than the one usually advertised.
What a parent should take from it. If your child's kindergarten uses a system like this, the useful question isn't "does my child use the AI?" It's "how does the teacher use what the system tells her, and how do you check whether my child is reading better without it?"
Common misreading. "AI raised kindergarten reading scores by half a standard deviation." No. A teacher-led instructional package, informed by assessment-based recommendations and supported by training, was associated with that difference on a latent reading measure in this trial. Take the teacher out of that sentence and you've described a study nobody ran.
3.3.2 Early Mathematics: Structured Apps, Not Open Conversation#
A United Kingdom randomized trial published in 2019 evaluated interactive mathematics applications with children aged four and five in their first compulsory school year. The trial randomized 461 children across 12 schools, with 389 children from 11 schools available at posttest. This is age-adjacent evidence for the youngest part of the K-12 range, not an exact equivalent of a U.S. kindergarten. (Outhwaite and colleagues, full article)
Here's what the study actually did. Children used the applications for approximately 30 minutes per day over 12 weeks, supervised by classroom staff. Three groups were compared. One group received app-based mathematics in addition to its normal mathematics activities. A second group used the apps in place of a daily small-group mathematics activity. The control group continued normal teaching.
Comparison with normal teaching
Reported result
Interpretation
Additional app-based mathematics time
Progress effect size 0.31; 95% confidence interval 0.06 to 0.55
A positive result for additional structured mathematics practice; additional mathematics exposure is part of the treatment
App use replacing a daily small-group activity
Progress effect size 0.21; 95% confidence interval -0.03 to 0.46; significance from a one-tailed test
More tentative: the two-sided interval includes zero
Supplementary versus time-equivalent groups
No statistically significant difference
Not proof that the two approaches are equivalent
Figure 3.7Data
Early mathematics apps with four and five year olds
United Kingdom randomized trial, 461 children in 12 schools, 389 at posttest, about 30 minutes a day for 12 weeks, supervised. No significant difference between the two app groups, which is not proof they are equivalent.
The supplementary result is clearly positive, but those children got their normal math plus the apps, so some of that gain may simply be more math. The replacement result is the more interesting policy number and the less certain one: its interval crosses zero. And "no statistically significant difference" between the app groups means the study couldn't tell them apart, not that they are equivalent.
The applications supplied bounded tasks, visual and spoken instructions, immediate feedback, repeated practice, and progression requirements. Children were not relying on unrestricted generated explanations, and the assessment was administered outside the app. (Intervention and assessment methods)
This evidence supports a specific proposition: structured digital practice can contribute to early mathematics learning under tested classroom conditions. It does not establish that replacing play, teacher interaction, or broad early-childhood experiences with more screen time improves overall development. The distance between those two propositions is the distance between what the trial measured and what a vendor might hope it implies.
What a teacher should take from it. If you're deciding whether to use math apps in a reception or kindergarten room, the trial suggests two things. Bounded, feedback-rich practice under adult supervision can help. And the evidence is firmer when app time is added than when it replaces your own small-group work, so be careful about what you give up to make room for it.
3.3.3 Literacy Engagement: More Minutes, No Proven Reading Gain#
A June 2026 working paper examined two randomized trials involving 355 students: one in an after-school setting covering grades 1 through 5, and one in an in-school setting covering grades 1 through 3. Both treatment and comparison students had access to the same literacy platform. The treatment added human support intended to encourage engagement, not to provide reading instruction. (Robinson and colleagues, full working paper)
The support differed between settings. The in-school trial used middle-school peer support, which means the support people can't all be assumed to be adults. And the platform can't be assumed to be a general-purpose language model. (Study design)
Measure
After-school, grades 1 through 5
In-school, grades 1 through 3
Randomized sample
174 students
181 students
Comparison-group mean platform minutes per week, including nonuse
2.18 minutes
5.23 minutes
Added weekly platform time from human support
About 1.00 minute; marginal at the 10% level
About 4.42 minutes; significant at the 5% level
End-of-year literacy result
No statistically significant improvement
No statistically significant improvement
Because both groups had the platform, this is not a randomized comparison of the platform against no platform. What it shows is more awkward and more valuable: the added human engagement component increased some usage measures without producing a statistically established reading gain in either trial. (Full working paper)
For decision-making, this is a warning against counting licenses, logins, or available tutoring hours as learning outcomes. A school should distinguish four different numbers that routinely get compressed into one dashboard metric: scheduled time, actual engaged time, completed practice, and independently measured reading growth.
Read that again. Scheduled time is what the district planned. Engaged time is what kids actually did. Completed practice is what the software logged. Reading growth is what you were paying for. A vendor report that shows the first three and not the fourth has shown you activity, not results.
3.3.4 Grades 3 and 4: Tutor Support Across the Band Line#
Tutor CoPilot, a generative assistant used by human mathematics tutors, included substantial grade 3 and grade 4 representation in its session-level analysis. Because the study crosses the band boundary, I treat it in full in Section 3.4.2. One rule applies here: its pooled effect must not be recast as a separate verified effect for third graders or fourth graders. Pooled results are pooled. I won't manufacture grade-specific precision the study did not produce. (Wang and colleagues, full study)
3.3.5 What Schools Can Responsibly Tell Families of Young Children#
For this age band, the most defensible explanation a school can give families is that tools may help a teacher identify what a child needs next, or provide a bounded opportunity to practice. The positive studies I reviewed involve specified instructional designs and adult implementation, not simply giving a young child a general-purpose assistant. (A2i synthesis, early mathematics trial)
Here's an illustration, clearly labeled as an illustration and not an additional measured result. A teacher reviews assessment information, selects a short activity, observes the child, and checks the same skill later without the tool. A parent receives an explanation of the learning goal and can ask whether the child is improving independently, rather than being asked to accept an opaque "personalization" claim.
Claims that these systems reduce parents' stress, save families money, or replace home reading support would require direct household evidence. Those outcomes were not the causal endpoints of any study in this section, and I won't borrow them from the household chapter, which reaches its own, more conditional conclusions in Chapter 6.
Checklist for parents of children in kindergarten through grade 4:
Does my child interact with the tool directly, or does the teacher use it to plan?
What exactly can my child ask it, and what can it answer?
Which adult is watching while my child uses it?
How does the teacher check whether my child can do the skill without the tool?
What is my child not doing (play, reading aloud, small-group time) during the minutes the tool is in use?
3.4 AI Homework Help and Tutoring Research, Grades 5 Through 8#
The middle-grade evidence covers three distinguishable arrangements: structured feedback on mathematics practice, assistance delivered through a human tutor, and a conversational tool embedded in a staffed remediation program. Add a fourth study about teacher preparation time, and you have the most varied band in this chapter. The findings are informative exactly as long as those arrangements stay separate. (ASSISTments report, Tutor CoPilot paper, Khanmigo school experiment)
A North Carolina trial assigned 63 schools to ASSISTments or to business-as-usual mathematics instruction, involving 102 grade 7 mathematics classrooms across 41 districts. The platform provided feedback and hints to students and reports to teachers, with professional development and coaching supporting implementation. (Feng, Huang, and Collins, technical report)
Focal students used the intervention in grade 7 during the 2019 to 2020 school year. The longer-term outcome was the grade 8 state mathematics assessment in spring 2021. The grade 8 outcome analysis included 5,991 students and reported an effect size of 0.10, with a p-value of .011. (Technical report)
That's useful evidence of a modest later assessment difference. It is not a claim of a 10 percent increase in learning. (Rule one: standard deviations are not percentages.) The study also ran through pandemic disruption, and students taking high-school-level mathematics instead of the relevant grade 8 test were excluded from that test's analytic sample. That exclusion matters for reading the result, because it removed a group of students from the comparison. (Technical report, WWC sample description)
Now the independent review, which requires particular care. The April 2026 What Works Clearinghouse review rates the study "Meets WWC standards with reservations," citing high cluster-level attrition with baseline equivalence in the analytic groups. It identifies uncertain effects for the mathematics achievement domain, with a positive grade 8 state-test result and a nonsignificant mathematics-readiness result in a smaller sample. An earlier review displayed on the same page was more favorable, so a citation to the earlier rating alone would leave out material later evaluation. (Latest WWC review, Review history and outcomes)
Figure 3.3Comparison
One study, two records: the trial report and the independent review
ASSISTments in North Carolina. A district that reads only the trial report gets an inflated picture; one that reads only the review might discard a promising program.
The defensible conclusion: the state-test result is positive, but the broader independent assessment is qualified, not uniformly confirmatory. Faster identification of errors and better-targeted follow-up are plausible mechanisms consistent with the intervention. They are not separately isolated causal effects. And the result belongs to the ASSISTments instructional package and its tested comparison conditions. It is not evidence for unrestricted conversational tutoring.
This study is my standing example of why independent review matters. The trial team's report and the federal reviewer's later assessment of the same study produce different confidence levels, and both belong in the record. A district that reads only the first gets an inflated estimate. A district that reads only the second might throw out a promising program. The discipline is to hold both, which means holding a smaller and more accurate conclusion than either source alone suggests.
Tutor CoPilot randomized 900 tutors to access or no access to an adult-facing assistant that suggested ways to respond to students during live online mathematics tutoring. The study identified 1,787 participating students in nine schools. The analyzed session sample contained 4,136 sessions taught by the remaining participating tutors. (Full study, January 2025 version)
Here's the design choice that makes this study worth your attention. The assistant could suggest a question or an explanation, and the tutor decided what to use. The child still talked to a human tutor. What changed was the support available to that tutor. That's a very different arrangement from putting a chatbot in front of a student.
Outcome or boundary
Finding
Main randomized assignment result
Exit-ticket passing rose from approximately 62% to 66%, a four-percentage-point increase; p < .01
Lower-rated tutor subgroup
A nine-percentage-point improvement; a subgroup result, not the overall effect
End-of-year mathematics assessment
No statistically significant improvement
Exposure period
Approximately two months
Coverage limitation
Grade 5 and grade 6 sessions support relevance to this band; the final session sample does not establish effects for grades 7 and 8
Publication status
Preprint; not treated here as a peer-reviewed journal finding
The result supports near-term lesson mastery under this tutor-assisted arrangement. It does not establish durable improvement in a student's overall mathematics attainment. Developer and provider involvement disclosed in the paper stays visible in my assessment. For a district, the next questions are whether the lesson-level improvement persists on independent assessments, and whether it's still worth it after training, supervision, and full delivery costs are included. (Outcomes and disclosures)
The four-percentage-point headline and the null end-of-year result must be held together. They describe different links of the benefit chain from Chapter 2. The first is task performance within sessions. The second is the real-world outcome a parent actually cares about. An adoption decision made on the first number alone is exactly the kind of decision this report exists to prevent.
Notice also the subgroup finding. Lower-rated tutors improved by nine percentage points. That's intriguing, because it suggests the tool may do the most good where tutoring quality is weakest. But it's a subgroup result, not the overall effect, and subgroup results are where hopeful readers go to find the number they wanted. Treat it as a question for the next study, not an answer for this one.
3.4.3 Khanmigo Study Results in a Staffed Remediation Package#
An August 2026 working paper examined Khan Academy with Khanmigo in grades 6 through 8 across 18 middle schools in Hamilton County, Tennessee, over two school years. Randomization involved 53 grade-within-school clusters, and the analysis included 6,902 student-term observations. That's not 6,902 distinct children. The same student can show up more than once across terms. (Oreopoulos and Low, full working paper)
The treatment combined a mathematics practice platform, conversational support, assigned learning paths, and scheduled remediation periods staffed by adults. The comparison was existing remediation, which itself could include teacher-led work and other digital learning products. The experiment therefore does not isolate the incremental effect of Khanmigo over otherwise identical Khan Academy use. There's a bundle on both sides of the comparison, which limits what either side can claim. (Design and comparison conditions)
The paper's pooled MAP result was an increase of approximately 1.26 national percentile-rank points per term, with a conventional standard error of 0.60. A small-cluster robustness test gave p = .0511, which makes the certainty of the main pooled result sensitive to the inference method. (Primary results, Robustness analysis)
That robustness detail is easy to skim past, so let me slow down on it. With only 53 clusters, the way you calculate uncertainty matters. Under the conventional method, the pooled result looks solid. Under a method built for a small number of clusters, it sits right at the edge of the usual significance threshold. Neither calculation is a trick. Together they tell you the result is real enough to take seriously and fragile enough not to oversell.
Annual standardized estimate
Reported result
Appropriate reading
First year
0.020 standard deviations, SE 0.035; not statistically significant
No demonstrated first-year annual gain
Second year
0.084 standard deviations, SE 0.041; significant at the conventional 5% level
A modest positive second-year package result
Stacked annual estimate
0.062 standard deviations, SE 0.035; marginal at the 10% level
Promising but less certain than a simple "proven annual gain" statement
Grades 6 to 8, 18 Hamilton County middle schools, 53 randomized clusters, 6,902 student-term observations. The year-to-year difference is not itself significant, and the package does not isolate the chatbot from the rest of the program.
The difference between the first-year and second-year annual estimates was not itself statistically significant. So one significant year and one nonsignificant year must not be described as a proven improvement between years. The report also notes that the preregistered state-test co-primary outcome had not yet been incorporated into the analysis. And its usage findings show that having a conversational tool available did not automatically translate into sustained explanatory dialogue. (Annual results, Outcome availability and usage analysis)
The appropriate conclusion is that this staffed remediation package produced modest, qualified evidence of benefit. It is not yet a clean demonstration that adding a chatbot alone improves outcomes for middle-school students.
3.4.4 Teacher Preparation Time: A Different Kind of Outcome#
A randomized trial involving 259 teachers in 68 English secondary schools tested ChatGPT with a preparation guide for Year 7 and Year 8 science lessons. The population is relevant to the middle-grade discussion, but it keeps its original label: United Kingdom school-year labels are not automatically U.S. grades. (NFER/EEF report)
In the measured later weeks of the trial, intervention teachers reported approximately 56.2 minutes per week preparing the specified lessons, compared with 81.5 minutes in the comparison group. That's roughly 25.3 minutes, or 31 percent, less preparation time for the measured work, with a reported time-ratio confidence interval of 0.53 to 0.90. (Primary results)
Figure 3.6Data
Teacher lesson preparation time, measured weeks of the trial
Year 7 and 8 science lessons in 68 English secondary schools. This is time preparing specified lessons, not total workload, and the trial did not measure student achievement.
The outcome came from teacher diaries, and the primary analysis included 211 teachers rather than every randomized participant. A blind review of 30 submitted lessons detected no quality difference, but a 30-lesson sample does not prove universal quality equivalence. The study did not measure student achievement. (Methods, attrition, and quality assessment)
This establishes a bounded operational benefit worth testing locally: 31 percent less time preparing specified science lessons under a preparation guide. It is not a 31 percent reduction in teachers' total workload, and it's not a measured improvement in anything students learned. Whether the saved minutes went to individualized teaching is a question the trial did not ask, and I won't answer it on the trial's behalf.
What a teacher should take from it. If your school offers AI tools for lesson prep, this is the best evidence in the set that it can save time on a defined task with a guide. Ask for the guide, not just the login. And keep an eye on verification time: checking generated material is real work, and the trial's quality review was small.
The evidence supports trials of bounded mathematics feedback, adult-mediated tutoring support, and carefully evaluated remediation packages, with separate measurement of student attainment and teacher time. It does not support treating all conversational access as equivalent, or stretching a positive session-level result into a full-year learning guarantee. (ASSISTments review, Tutor CoPilot, Khanmigo evaluation)
Here's an illustrative workflow, proposed as a design principle for evaluation rather than claimed as a proven effect: students attempt a problem, receive a limited hint, explain their reasoning, and later solve a related problem without assistance. Every clause in that sequence is doing evaluative work. The last clause is the one most existing programs leave out.
For parents of middle schoolers dealing with AI homework help at home, the research doesn't hand you a rule, but the pattern across these studies points to a practical habit. The studies that showed gains used tools that gave hints and feedback, kept an adult involved, and measured the student without help. You can copy that structure at the kitchen table: let the tool give a hint, have your child explain the step back to you, and then have them try a similar problem with the laptop closed. That's not a tested intervention. It's the evaluation logic of the better studies, applied at home.
3.5 Does ChatGPT Help Students Learn? High-School Evidence#
High school is where the difference between producing an answer and acquiring a skill becomes impossible to avoid. A teenager with a capable chatbot can produce a lot of correct-looking work. Whether she can do that work herself next month is a separate question, and in this band the evidence answers it directly. The reviewed studies include an adverse effect from unrestricted assistance, a promising supervised program, and an older adaptive intervention whose effects differed by implementation year. (PNAS mathematics experiment, World Bank English evaluation, RAND algebra evaluation)
3.5.1 The PNAS Experiment: Better Practice, Weaker Learning#
A peer-reviewed 2025 study in PNAS ran a field experiment with nearly 1,000 high-school students in mathematics. It compared three conditions: a GPT-4-based general assistant, a more constrained tutoring version, and a control condition. The tutoring version incorporated teacher-provided instructional material and restrictions intended to guide students rather than simply hand them answers. (Bastani and colleagues, original paper)
Now here's what happened. The unrestricted assistant increased assisted practice performance by 48 percent relative to the control group. The tutoring version increased it by 127 percent. Then access was removed for subsequent exams, and the picture flipped. Students in the unrestricted-assistant condition performed 17 percent worse than the control group, while the constrained tutor largely mitigated that adverse effect. (Reported experimental results)
Read that again. The students with the open chatbot looked better on practice and then did worse on the exam than students who never had it.
Figure 3.2Data
Better practice, weaker learning: the PNAS high-school math inversion
Nearly 1,000 high-school math students. Percentages are relative to the control group on the study's own measures, not percentage points. The tutor version avoided the penalty; that is not the same as a demonstrated gain in unaided learning.
These percentages describe the study's measured performance. They must not be rewritten as percentage-point changes, and they must not be generalized to every subject. (Rule three: relative percentages need their denominator.) The tutoring result also must not be turned into a claim of a demonstrated positive independent-learning effect just because it avoided the unrestricted tool's penalty. Avoiding a penalty is not the same as producing a gain, and I decline to launder one into the other.
What the experiment demonstrates, with unusual clarity, is the third link of the benefit chain failing in the wild. Both assistants made practice look better. Only one of them left students able to perform after the assistance was gone. The design implication follows directly: completed assignments can't be read as learning, and any deployment that doesn't include an unaided measurement is flying blind on the outcome that matters most.
This is not an argument against educational assistance. It's evidence that the structure of the assistance is part of the intervention and has to be evaluated. The same kind of tool, used by the same kind of students, in the same subject, produced learning protection in one form and learning loss in the other, depending on the guardrails around it.
This is the study that answers the kitchen-table question most directly. If the tool is doing the work, your child's practice will look great and her independent performance may suffer. If the tool is built to guide instead of answer, that penalty can largely go away. The only way to know which one you've got is to test your child without it.
3.5.2 The Unaided-Assessment Rule for Any School Deployment#
The PNAS design generalizes into a simple rule, and it's the single most useful thing a school can take from this chapter. Any deployment should have three stages: assisted practice, removal of assistance, and an independent measurement of what the student can now do alone.
Figure 3.5Framework
The unaided-assessment rule for any school deployment
Generalized from the PNAS design into a proposed evaluation workflow. Proposed, not tested as a standard.
Here's a hypothetical to make it concrete. A high school wants to pilot an AI writing assistant in tenth-grade English. Under the unaided-assessment rule, the pilot plan would say, before anything starts: students will use the tool for drafting practice during the unit; at the end of the unit, they will write an in-class essay without the tool; that essay will be scored by teachers who don't know which students used the tool; and the comparison will be against classes that didn't use it. If the school can't commit to that last essay, it isn't running an evaluation. It's running a subscription.
Nothing in that illustration is a measured result. It's the evaluation structure the PNAS study makes impossible to ignore.
A World Bank working paper evaluated a six-week program using Microsoft Copilot, powered by GPT-4, for first-year senior secondary students learning English in Nigeria. The reported effect was 0.23 standard deviations on English and 0.31 standard deviations on a composite assessment combining English, knowledge of the technology, and digital skills. (World Bank publication record and abstract)
The delivery details matter a lot here. The program ran in nine public schools in Benin City, with 1,328 randomized students and 759 in the final analytic sample presented by the authors. It comprised twelve 90-minute after-school sessions over six weeks, with students working in pairs, curriculum-oriented prompts, and teacher guidance and monitoring. So the intervention combined a model with additional structured learning time, access arrangements, prompts, peer interaction, and supervision. The result is evidence for that package, not a model-only effect holding all other instructional resources constant. (Authors' methodological presentation)
The drop from the randomized sample to the analyzed sample is a material limitation. The authors report attrition-related robustness analyses rather than ignoring the problem. Those analyses strengthen the interpretation, but they don't make missing outcome data irrelevant. (Sample flow and robustness presentation)
Four claims, four treatments:
Claim
Evidence-based treatment
English learning improved in the evaluated program
Supported by the reported 0.23-standard-deviation English result
The composite result was an English-only effect
Incorrect: the 0.31 estimate combined English with technology knowledge and digital skills
Students completed years of schooling in six weeks
Not established; schooling-equivalent language is a benchmark conversion, not observed completion of additional school years
The same effect holds for American seniors or every high-school subject
Not established by this population and subject
The World Bank evaluation also reports heterogeneous effects, including larger gains for female students and for students with higher baseline academic performance. A positive average therefore shouldn't automatically be presented as evidence that the intervention closes achievement gaps. If anything, the reported pattern for baseline performance points the other way. (World Bank study abstract)
This is promising evidence for supervised use in a specified secondary-school setting, not a universal promise. The canonical World Bank abstract supplies the principal effect estimates. The authors' presentation supplies sample-flow and implementation detail, and I'm not presenting my reading of it as a complete audit of the full working paper.
3.5.4 Cognitive Tutor Algebra I: The Version Lesson#
The Cognitive Tutor Algebra I evaluation studied a blended curriculum combining classroom instruction and software across 147 schools, including 73 high schools and 74 middle schools. The high-school sample was predominantly ninth graders, and the intervention was an adaptive subject-specific system, not a modern general-purpose conversational model. (RAND study summary, journal-publication record)
The evaluation found no statistically significant first-year benefit and a positive second-year high-school result. RAND's later addendum reported a significant second-year high-school effect of 0.21 standard deviations, substantively consistent with the earlier analysis. (Study summary, 2014 analytical addendum)
Two cautions are structural. First, the two implementation years involved different student cohorts, not two years of exposure for the same students. So the findings shouldn't be described as proof that an individual learner must use the product for two years to benefit, or as proof that teacher experience caused the difference. Second, the system's era matters. A bounded subject system producing a measurable benefit at scale does not establish that a more general or newer model will outperform it.
The relevance here is historical and methodological. This study shows both that bounded systems can work at scale and that results can differ by implementation year.
3.5.5 Four High-School Years Are Not One Proven Population#
The high-school band includes four years, but the evidence base doesn't justify four grade-specific benefit estimates. The algebra evidence is concentrated in grade 9. The Nigerian study concerns first-year senior secondary students in a different school system. The mathematics chatbot experiment does not establish a separately verified effect for every U.S. high-school year. (RAND sample description, World Bank population description, PNAS study)
Senior-year readiness, independent research, college transition, career preparation, and sustained writing development remain open research questions. They are not assumed consequences of a mathematics or English trial. The absence of a verified estimate in this targeted review is not a claim that no relevant research exists anywhere. It's a refusal to fill the gap with the nearest available number. For what happens after graduation, Chapter 4 takes up college and workforce training on its own evidence.
Every number in this chapter will eventually get repeated at a PTA meeting, in a school newsletter, or in a parent group chat. Most of them will get mangled on the way. Here's how a teacher, principal, or board member can explain them without adding a single new number, using only the four interpretation rules from Section 3.1.
Start with what a standard deviation is not. When a study says 0.23 standard deviations, say: "This is a way of comparing score differences across different tests. It is not 23 percent more learning." If a parent asks whether 0.23 is big, the fair answer is that it depends on the students, the test, the length of the program, and the cost, which is why the study details matter more than the number.
Say "points" when you mean points. When Tutor CoPilot's exit-ticket passing rose from about 62 percent to 66 percent, that's four percentage points. Say "four more students out of every hundred passed the exit ticket," not "a 66 percent gain" and not "learning went up 6 percent."
Always say "compared with what." When the PNAS study reports that unrestricted-assistant students performed 17 percent worse on later exams, the comparison is the control group in that study. It's a relative change, not a 17-point drop. Say the comparison out loud every time.
Say who got the offer, not just who used it. If a result is reported for everyone assigned to the program, say so. If a vendor shows you results only for "active users," explain that the kids who use a tool the most may have been different from the start, and that the fair comparison is between the groups that were offered it and the groups that weren't.
Name the package. Never say "AI raised reading scores." Say "a teacher-led reading program that used assessment-based recommendations was associated with higher reading scores." It's longer. It's also true.
Here's a hypothetical script for a principal at a parent night, built only from numbers already in this chapter: "Our pilot is modeled on programs where the tool helps the teacher rather than replacing her. In one published study, a similar arrangement helped students pass short end-of-lesson checks more often, but it did not show a significant change on end-of-year tests. So we're going to measure your children at the end of the year, without the tool, and we'll report back either way."
That last sentence is the one parents remember. It tells them the school is testing for the outcome they care about.
The same educational outcome can matter to several audiences, but relevance must not be confused with proof of an additional downstream benefit. This table translates the findings into decision questions. It's an analytical framework, not a new causal study.
Audience
Interaction to evaluate
Demonstrated or directly studied value
Additional evidence needed
Student, kindergarten through grade 4
Teacher-selected instruction or supervised, bounded practice
Reading and mathematics outcomes in specified instructional packages
Bounded preparation-time savings and improved tutoring-session outcomes
Total workload, verification burden, sustainability, actual redeployment of saved time
School board or district
Select, fund, monitor, and discontinue programs
Study-specific attainment and operational findings as starting benchmarks, not local guarantees
Local comparative outcomes, full costs, access, incident reporting, procurement controls
Civic group or infrastructure builder
Evaluate the relationship between educational services and physical infrastructure
This chapter establishes no site-specific power, water, fiscal, or community finding
Separate utility, siting, resource-use, fiscal, and community evidence
Take a stay-at-home parent raising three children. The immediate value of this chapter for that parent is a clearer way to ask whether a school's proposed tool actually helps each child learn. It would be premature to turn these studies into an assumed number of hours saved at home or a promised improvement in family quality of life, and Chapter 6 shows why the household evidence doesn't fill that gap either.
UNICEF's December 2025 Guidance on AI and Children 3.0 recommends age-appropriate systems, protection of children's data, human oversight of consequential decisions, inclusion, and demonstrated educational efficacy before large-scale deployment. It specifically emphasizes keeping teachers central to education rather than replacing them with automated systems. This guidance is a policy and child-rights framework. It is not a randomized estimate of learning gains, and it is not a statement of every jurisdiction's law. (UNICEF Guidance on AI and Children 3.0)
Here's my proposed translation of that framework into an evaluation standard:
Adult accountability. Name the teacher, school administrator, or other responsible adult who can intervene, correct an error, or stop use.
Age-appropriate interaction. Specify what the student can ask, what the system can answer, and how the design differs between an early reader and an older adolescent.
Instructional boundaries. Identify when the system may give a hint, an explanation, a complete solution, or no response.
Data minimization. Document what information is collected, retained, shared, and used for training. Don't treat a child's full educational history as a default input.
Human recourse. Provide a practical way to challenge an incorrect recommendation or harmful interaction.
Accessibility and alternatives. Test the actual interface with intended users, and provide a meaningful alternative when a student cannot or should not use it.
Consequential-decision exclusion. Don't infer that a tool shown useful for practice is validated for placement, discipline, diagnosis, or high-stakes assessment.
These controls are consistent with UNICEF's emphasis on developmental appropriateness, privacy, inclusion, and human control. They still need to be translated into applicable school policies, contract terms, and legal requirements before deployment.
Look at "instructional boundaries" next to the PNAS result. The difference between the tool that hurt independent performance and the tool that largely avoided that harm was, in large part, a set of instructional boundaries. This condition isn't bureaucratic box-checking. It's the variable the best experiment in the set turned on.
One distribution caveat belongs here. This targeted set of studies does not supply a consistent set of independently verified effects for students with disabilities, every language group, every income group, or every grade. Local evaluation should state which populations are represented and should not fill missing subgroup evidence with assumptions. A positive average, as the World Bank heterogeneity results show, can coexist with widening gaps.
3.9 AI in Schools Evidence: Measuring Value Consistently#
The central evaluation question for any school deployment should be: what can the student do, what can the adult do, and what does the complete intervention cost compared with a credible alternative? Student learning and operational efficiency should be reported separately before anyone tries to combine them.
Every consequential education claim should keep the same fields, so a favorable headline can't shed its comparison group, grade level, or uncertainty when it's repeated in a public document:
Field
Required description
Population
Grade, age if available, location, language, baseline achievement, eligibility, relevant access conditions
Intervention
Product and version, student-facing or adult-facing role, instructional restrictions, curriculum, supervision, training
Comparator
What the other group actually received, including teacher time and other technology
Exposure
Scheduled time, realized use, duration, adherence, attrition
Assignment and analysis
Randomization unit, numbers assigned, analytic sample, unique students versus repeated observations, clustering
Primary outcome
Named assessment or operational measure, timing, whether assistance remained available
Effect and uncertainty
Metric, estimate, interval or standard error when reported, significance, robustness qualifications
Durability and transfer
Delayed tests, unassisted performance, unfamiliar problems, broader outcomes where measured
Peer review, working-paper status, provider involvement, funding, independent review where available
Permitted conclusion
The narrow statement actually supported, plus the most tempting unsupported extension
3.9.2 The Eight-Dimension School Evaluation Scorecard#
For future evaluations, I recommend measurement across eight dimensions:
Unassisted learning, with delayed follow-up.
Transfer to unrehearsed problems.
Engagement, reported alongside nonuse.
Adult work, including verification and error resolution.
Access and outcomes by student group, with reasons for nonparticipation.
Safety, privacy, and human recourse incidents.
Total delivery cost per participating student, against the alternative.
Displacement: what the intervention replaced.
A local benefit claim should depend on a comparison and outcome defined in advance, not on a retrospective pick of the most favorable dashboard metric. The reviewed studies show why: engagement increased without established literacy gains in one program, and assisted performance rose while later independent performance fell in another. (Literacy engagement evaluation, high-school mathematics experiment)
3.9.3 Why This Chapter Makes No Return-on-Investment Claim#
I don't convert teacher minutes, standardized test differences, or exit-ticket improvements into a single financial return. That conversion would require explicit assumptions about implementation costs, the value and actual use of saved time, the durability of effects, and the alternative use of school resources. And an inexpensive model interaction is not the full cost of educational delivery. Training, supervision, curriculum integration, devices, connectivity, verification, and program management all belong in the economic comparison before a district, or a company like ours, makes an affordability claim. Any document that quotes a per-student software price as the cost of the program has answered a question no one asked.
3.10 Questions to Ask Your School Board Before It Buys an AI Tutor#
Everything above boils down to a list you can bring to a public comment period or a board work session. Each question comes straight from the scorecard, the common evidence record, or the conditions in Section 3.8. None requires technical expertise to ask. All of them require a real answer.
About the evidence
Which study does the vendor cite, and did that study test the same product, grade, subject, and arrangement we're buying?
What did the comparison group in that study actually receive?
Is the cited result a peer-reviewed paper, a working paper, a preprint, or a vendor report, and has any independent reviewer rated it?
Were students tested with the tool available or without it?
About our own pilot
What outcome will we measure, and did we write it down before the pilot started?
Will students be assessed later without the tool, and will we compare them against students who didn't use it?
Will we report usage alongside nonuse, and report which students didn't participate and why?
Will we measure teacher time, including time spent checking what the tool produces?
About safeguards
Which adult is accountable for each classroom's use, and who can stop it?
What can students ask, what can the system answer, and when will it give a hint instead of a full solution?
What student data is collected, retained, shared, or used for training?
How does a student or parent challenge a wrong or harmful response?
Is this tool being used, or planned for use, in placement, discipline, diagnosis, or high-stakes testing?
About cost and trade-offs
What is the total delivery cost per participating student, including training, supervision, devices, connectivity, and verification, not just the license?
What are we giving up to make room for it: teacher small-group time, tutoring hours, something else?
What result would cause us to stop, and when will we decide?
The grade-band coverage deserves explicit mapping, because "covered" here means covered by the selected studies, not by all relevant research worldwide:
Segment
Coverage in this review
Boundary
Kindergarten
Direct kindergarten reading trial; age-adjacent U.K. early mathematics trial
Does not establish unrestricted conversational use for young children
Grades 1 and 2
Included in the literacy-platform support trials
No statistically established incremental reading gain from added engagement support
Grades 3 and 4
Included in literacy support and Tutor CoPilot sessions
Pooled results are not separate effects for each grade
Grade 5
Included in one literacy trial and CoPilot session data
Less direct grade-specific coverage than the grouped label suggests
Grades 6 through 8
CoPilot grade 6 sessions; grades 6 through 8 remediation trial; grade 7 ASSISTments with grade 8 outcomes
Different interventions, assessment timings, levels of certainty
Grade 9 or early secondary
Algebra evidence concentrated in grade 9; Nigerian first-year senior secondary English evidence
Foreign school stages are not automatically U.S. grade equivalents
Grades 10 through 12
High-school mathematics experiment and some broader high-school representation
No separate verified sophomore, junior, or senior benefit estimate
Publication status stays attached to each finding rather than being flattened into a single "research-backed" label. The kindergarten reading study is peer-reviewed. The early mathematics trial is a peer-reviewed 2019 article. The literacy engagement and Khanmigo studies are working papers. Tutor CoPilot is a preprint. ASSISTments carries both a research report and the April 2026 WWC review, with its reservations kept. The teacher preparation trial is an NFER/EEF randomized evaluation report. The PNAS mathematics study is peer-reviewed. The Nigerian evaluation is a World Bank working paper whose canonical abstract and authors' presentation were reviewed, not subjected to a claimed full-paper audit. And the algebra evaluation rests on RAND's publication record, summary, and 2014 addendum.
Peer review doesn't remove the need to examine design, sample loss, measurement, and generalization. A working paper isn't automatically uninformative. Status attaches, and it stays attached. For how this report grades evidence classes across every chapter, see Chapter 11.
Whether observed improvements persist across later grades and survive removal of the tool.
Whether routine assistance strengthens or weakens verification, reasoning, and recognizing uncertainty.
What sustained effects look like for writing, source evaluation, and original work, beyond this mathematics-heavy set of studies.
Which interventions improve college readiness, skilled work, and adult civic participation for older students.
Whether families experience measurable reductions in administrative burden or tutoring cost.
Who benefits, who disengages, and what design changes are needed for language, disability, and access differences.
Whether tested programs outperform a credible alternative use of the same money and adult time.
Whether effects persist after a model update, interface redesign, pricing change, or guardrail change.
These are limits on what this chapter can conclude. They are not reasons to disregard its positive findings.
3.12 The Boundary Between Learning Evidence and Infrastructure#
I'll be direct about who is writing this. This report is prepared by Savrn, which designs and builds AI factories and works on the infrastructure side of this industry, and our stated objective is a stronger relationship between that infrastructure and community well-being. I'm disclosing that commercial context rather than dressing it up as institutional independence.
Educational benefit and infrastructure acceptability are separate propositions. Even a strong learning result does not establish that a proposed facility has acceptable power demand, water use, public costs, land-use effects, or community commitments.
The seven Savrn trackers belong alongside this chapter as a separate evidence layer, not as supporting citations for learning gains. Their role is to inform infrastructure questions while education claims stay grounded in education research. A tracker entry starts a question and never ends one. The public-value chain I'd propose should be treated as a set of questions, not an assumed causal sequence:
Does the educational intervention improve a defined outcome?
Can schools deliver it reliably, safely, and affordably?
What infrastructure is actually needed to provide that service?
What costs, resource demands, and benefits arise from a particular facility?
Can affected people verify commitments and obtain remedies?
That separation keeps this chapter useful to a parent, teacher, or civic committee that never buys a Savrn service. It also stops a legitimate education finding from being stretched into an argument for a specific development. Chapter 10 examines the infrastructure side on its own evidence.
3.13 Conclusion: What AI Can and Can't Do for Your Kid's Learning#
So, back to the kitchen table. Is the tool helping your kid learn, or doing the homework for them? The evidence says that depends almost entirely on how the tool is built and how it's used, and that you can't tell which from the homework itself. You can only tell by checking what your child can do without it.
The defensible position is neither blanket adoption nor blanket rejection. Educational decisions should rest on clearly identified learners, interventions, comparison conditions, outcomes, costs, and uncertainties, with the same standard of evidence applied to favorable and unfavorable results. The thesis that comes out of that is narrower and more useful than either blanket optimism or blanket rejection: educational value should be established by what learners can do independently and what educators can demonstrably deliver, under clearly described conditions.
The evidence here supports real optimism of a specific kind. Structured systems, adult-facing support, and supervised programs have produced measurable benefits for real students, and the PNAS experiment shows that the adverse outcome is not inevitable. It's a design failure with a known fix. What the evidence does not support is the idea that capability turns into learning by itself. Between the model and the child stands a chain of design decisions. This chapter has tried to show, study by study, exactly where that chain holds and exactly where, left unattended, it breaks.
Does AI help students learn, according to research?
Sometimes, under specific conditions. In this review of ten K-12 study programs, structured systems, tools that support teachers, and supervised programs produced measurable gains, such as 0.52 on a latent kindergarten reading measure in a teacher-led package. But in a 2025 PNAS experiment, students with an unrestricted GPT-4 assistant scored 17 percent worse than the control group once the tool was removed. Design and supervision decide the outcome.
Does ChatGPT help students learn or just do their homework?
The strongest evidence says it can do both, depending on design. In the PNAS study of nearly 1,000 high-school math students, an unrestricted GPT-4 assistant raised assisted practice performance 48 percent, then those students did 17 percent worse than controls on exams without it. A constrained tutor version largely avoided that penalty, though avoiding a penalty is not proof of a gain.
What did the Khanmigo study find?
An August 2026 working paper on Khan Academy with Khanmigo in grades 6 through 8 in Hamilton County, Tennessee found 0.020 standard deviations in year one (not significant), 0.084 in year two (significant at 5 percent), and 0.062 stacked (marginal at 10 percent). The treatment was a staffed remediation package, so it does not isolate the chatbot, and a small-cluster robustness test gave p = .0511.
Do AI tutoring studies show gains on end-of-year tests?
Not consistently. Tutor CoPilot, which gave human tutors AI suggestions, raised exit-ticket passing from about 62 percent to 66 percent over roughly two months, but found no statistically significant end-of-year math improvement. ASSISTments showed a 0.10 effect size on a grade 8 state test, yet the April 2026 What Works Clearinghouse review rated the study "with reservations" and found uncertain effects overall.
Does AI save teachers time?
For a defined task, one trial says yes. In an NFER/EEF randomized trial of 259 teachers in 68 English secondary schools, teachers using ChatGPT with a preparation guide reported about 56.2 minutes a week preparing specified Year 7 and Year 8 science lessons, versus 81.5 minutes for comparison teachers, 31 percent less. It did not measure total workload or student achievement.
Is AI safe for young children in school?
The research here does not test unrestricted chatbots with young children. The positive early-grade studies used teacher-directed or bounded tools, such as math apps used about 30 minutes a day for 12 weeks under staff supervision. UNICEF's December 2025 guidance calls for age-appropriate design, data protection, human oversight, and demonstrated efficacy before large-scale deployment, but it is policy guidance, not a learning study.
What does an effect size of 0.23 mean for my child?
It is a standardized score difference between groups, not 23 percent more learning. The 0.23 figure comes from a World Bank evaluation of a six-week supervised English program in Nigeria, with 759 students in the final analytic sample out of 1,328 randomized. It describes that program and population on average, not a guaranteed result for any individual child or for U.S. students.
What should parents ask a school before it adopts an AI tutor?
Ask what study supports the product and whether it tested the same grade and arrangement, what the comparison group received, and whether students will be tested later without the tool. Ask which adult is accountable, what data is collected, how errors are challenged, and the total cost beyond the license. In the studies reviewed, results depended on supervision, structure, and unaided measurement.
Chapter 4 · College, Trades, First Jobs
AI in College, Skilled Trades Training, and the First Years of Work
Picture a 19-year-old at a kitchen table with three open tabs, and a pile of AI in higher education research headlines telling them what to think. One is a university acceptance letter. One is an apprenticeship application for an electrical program. One is a job posting that pays next Friday. Every adult in that kid's life has an opinion, and now every one of those paths comes with an AI pitch attached: the college has a chatbot, the training center has a simulator, and the job board offers to write the resume. The real question underneath all of it is simple. Which of these will actually make me able to do something someone will pay for? That is the question this chapter takes to the evidence, and it is the question most of those headlines get quoted around rather than answering.
I care about this one personally. I have spent my career building large power infrastructure. When you build at industrial scale, you hire electricians, technicians, and operators, and you learn fast that there is exactly one credential that matters on an energized site: can this person do the work, safely, when nobody is standing over them? A diploma tells me someone finished something. A certificate tells me someone passed something. Neither one tells me what happens when a breaker trips at 2 a.m. and the procedure on the screen doesn't match what's in front of them. "Can do the work" is the whole credential. Everything else is a proxy.
That operator's lens is how I read this chapter's eleven study programs. Some of them are good news. An enrollment chatbot helped committed students actually show up. A carefully designed physics tutor beat a strong classroom on short-term tests. A resume tool raised hiring on one platform. A workplace assistant raised output per hour. Two long-running workforce programs posted real earnings gains. Here's the part most people skip: the chapter also contains a randomized experiment where adults learning a new technical skill with an AI assistant came out understanding it less well, with no time saved. And the trades evidence, the part closest to my world, is thin, short, and about welding only.
So the finding is not "AI helps" or "AI hurts." It is that enrolling, finishing an assignment, demonstrating a skill, getting hired, and performing on the job are five different achievements. A tool can win one and leave the other four untouched. I'm going to walk through each study with its population, its comparison, and its limit attached, then show you what a real competence sequence looks like on a site where mistakes are measured in arc flash, and finish with practical advice for students and job seekers who want to use these tools without hollowing out the very skill they're trying to build.
4.1 AI in Higher Education Research: Five Different Wins#
This chapter picks up where the K-12 evidence in Chapter 3 leaves off, in the years where education turns into employment. The organizing idea is practical: helping a person enroll, finish an assignment, demonstrate a skill, obtain work, and perform at work are different achievements. A system can accomplish one while leaving the others untouched, and the reviewed evidence does exactly that. Improvements in college enrollment, immediate subject learning, hiring, and workplace output sit alongside uncertain vocational results and an experiment showing weaker independent technical understanding after assisted practice. (College enrollment experiment, college physics experiment, hiring experiment, workplace study, welding experiment, technical skill experiment)
The standard I hold this chapter to follows from that. Educational assistance should produce demonstrable capability, not merely a more polished submission. Where the goal is administration or workplace production rather than learning, the evaluation should say so plainly and measure that goal directly. Confusing the two is how institutions end up measuring the polish and calling it the capability.
I selected eleven decision-relevant study programs, prioritizing original studies, government evaluations, and identifiable comparison conditions. College, vocational training, and employment are treated as overlapping pathways, not a ladder where every learner must climb to a four-year degree. That matters to me. Some of the most capable people I've ever put on a site never set foot in a lecture hall, and some of the most credentialed people I've interviewed couldn't read a one-line diagram.
This is a targeted synthesis, not a systematic review. It offers no pooled effect and makes no claim to have found every study. Publication status is tracked separately from design: a randomized working paper and a peer-reviewed observational study each keep both labels, and I never merge them into a single "proven" bucket. A working paper can be well designed. A journal article can be observational. You need both facts to judge the finding.
The technology boundaries need restating. "Superintelligence" names the transition this report investigates. It is not a verified capability level of the tools reviewed here. The evidence in this chapter includes generative tutoring, non-generative messaging support, virtual and augmented reality, and training programs with no model-based intervention at all.
Why include old and non-generative systems in a report about the AI transition? Because they identify mechanisms and credible alternatives. If a text-message enrollment bot moves outcomes, that tells you something about what a newer assistant would have to beat. If an intensive program with no AI at all produces seven-year earnings gains, that is the benchmark any AI-enabled pathway should be measured against. What I won't do is relabel the effect of an older system as evidence that a newer general-purpose model produces the same outcome. Capability is not benefit, and a result earned by one tool doesn't transfer to another because they share a chat window.
Figure 4.1Framework
The education-to-work pathway and where each study's evidence stops
Proposed stage model. Each stage has its own outcome and decision gate. No reviewed study evaluates the whole sequence as one treatment. Gray stages have no anchor study in this review.
The funnel view is the most useful way to read everything that follows. Each of the eleven programs sits at one stage of the education-to-work sequence, from choosing a pathway through progressing in a job. Each one evidences its own stage and stops there. When you see a headline that jumps from one stage to another, from "students liked the tutor" to "graduates earn more," you're watching someone cross a gate the evidence never crossed.
4.2 Can a College Chatbot Get Admitted Students to Show Up?#
The first stage of the funnel is the least glamorous and, for a lot of families, the most consequential. A student can be admitted, intend to go, and still never arrive because a form, a deposit, a transcript, or a financial-aid step fell through the cracks over the summer. Researchers call that "summer melt." It is an administrative failure, and it turns out to be one of the places where a narrow, well-built tool can help.
Page and Gehlbach evaluated Pounce, a university enrollment-support chatbot, in a randomized study involving 7,489 admitted Georgia State University students, 1,948 of whom had already committed to attend. The intervention used text outreach, a knowledge base, student-specific enrollment information, and escalation to admissions staff. It was not an open-ended generative tutor. It answered procedural questions and prompted action on enrollment tasks, while staff handled anything it couldn't resolve. (Original AERA Open article)
Among already-committed students, the treatment increased enrollment at Georgia State by 3.3 percentage points, with a standard error of 1.6 percentage points. The authors describe this as a 21 percent relative reduction in summer melt, meaning admitted students who intend to enroll but fail to complete the steps. The full-sample university-enrollment estimate of 1.2 percentage points was not statistically significant, and neither was the committed-subgroup estimate for attendance at any postsecondary institution. (Study results, Full-sample and attendance estimates)
Figure 4.2Data
Pounce enrollment chatbot at Georgia State
The authors describe the committed-subgroup result as a 21% relative reduction in summer melt. It does not show higher college attendance overall, learning, completion, or earnings.
Read those two numbers side by side. The 3.3-point gain applies to students who had already said yes. Across everyone admitted, the 1.2-point estimate could not be distinguished from zero. And for the committed group, the study could not show more students attending college anywhere, only more attending this particular university. That's a useful result for Georgia State's enrollment office. It is a much weaker result for the national question of whether chatbots send more young people to college.
What the study demonstrates is narrow and useful: more already-committed students enrolled at the participating institution under tested conditions. What it does not demonstrate is a general increase in college attendance across all admitted students, stronger academic learning, degree completion, or lifetime earnings. The result belongs to the funnel's first stage. Treating it as evidence about the later stages is the exact pathway confusion this chapter exists to prevent.
Common misreading: "A chatbot cut summer melt by 21 percent." That's the authors' relative framing of a subgroup result. Show me the denominator: it's committed students at one university, and the absolute change was 3.3 points with a standard error of 1.6.
The implementation lesson follows from the design. Separate administrative assistance from academic authority. An enrollment assistant should explain deadlines and procedures from current institutional records, show where those records came from, and escalate eligibility questions, financial-aid disputes, and unusual circumstances to authorized staff. Every clause there is doing work. Current records, visible sourcing, and human escalation are what separated this intervention from an unsupported advice bot. Strip those out and you have a different tool that happens to use the same text thread.
For a parent: if your student's college offers a text assistant for enrollment steps, use it for what it's good at, which is reminders and procedural answers. For anything involving money, eligibility, or an exception, get a human name and a written answer.
4.2.2 Course Reminders at Georgia State: A More Mixed Result#
The 2026 version of "Let's Chat" reports randomized course-support outreach covering 2,483 students in introductory government and microeconomics courses at Georgia State. The intervention added personalized course messages and knowledge-base support to a population that already had access to the university's general retention chatbot. So this was a test of additional course-specific support, not a comparison against no support at all. That comparison asymmetry disciplines everything you can say about it. (Meyer and colleagues, working paper)
The pooled numeric course-grade estimate was +1.80 points, but it was not statistically significant after the study's multiple-comparison adjustment. The estimated four-percentage-point increase in earning a B or higher carried an adjusted p-value of .090. A stronger result for women in microeconomics was a subgroup finding, not the overall effect. The working paper did not establish broad improvements in credits from other courses, overall semester GPA, or subsequent course enrollment. (Pooled and subgroup results, Additional outcomes)
The practical implication cuts against a common deployment reflex. Reminder systems should be evaluated against the support students actually receive, and message response rates should stay what they are, which is implementation measures. They are not substitutes for course completion or independent learning. A chatbot that students answer is a communications channel. Whether it is an educational intervention is a separate question, and this study answers it only conditionally.
Common misreading: "Students who got AI course nudges did better." The pooled estimate was positive and not significant after adjustment. The subgroup result for women in microeconomics is worth following up, and it is a lead for the next study, not a finding for the brochure.
For a college administrator: before you buy a course-nudge product, write down what your students already receive. If they already have a general chatbot, advising, and early-alert emails, the question is not "does the new tool help versus nothing" but "does it help on top of all that." This study is the template for asking it correctly.
4.3 Does an AI Tutor Help College Students Learn?#
This is the section that generates the most headlines, and the one where I'd ask you to slow down the most. There is one strong college learning study in this chapter's set. It is good work. It is also the finding most likely to be stretched past what it measured.
4.3.1 The Harvard Physics AI Tutor: Promising, Bounded Evidence#
Kestin and colleagues studied a custom GPT-4-based tutor in an introductory Harvard physics course, using randomized crossover sequences across two lessons. The study reported 194 eligible participants, with 142 tutor-condition and 174 classroom-condition posttest observations. Those are observations across crossover conditions, not two independent samples of students, so you can't add them up and call it a head count. (Scientific Reports article)
The comparison is the first thing to understand. The tutor was tested against an in-person active-learning class, not a passive lecture and not the absence of instruction. That's a demanding comparison, and it's why the result earns attention. The tutor itself was a purpose-built instructional package: deliberately structured prompts, expert-written solutions, guided problem progression, and instructional materials. It was not a student opening a general chatbot and typing "explain this."
The reported adjusted effect on immediate posttest performance was 0.63 standard deviations in favor of the tutor condition. Median tutor use was approximately 49 minutes, against the authors' estimate of roughly 60 instructional minutes in the classroom condition after excluding assessment time. (Regression results, Time comparison)
Figure 4.3Comparison
The designed physics tutor: a bounded win
Harvard introductory physics. A real finding with a narrow scope: a carefully engineered tutoring workflow beat a strong comparison on immediate outcomes.
Four qualifications attach, and each one matters to a decision-maker.
Assessment timing. Outcomes followed individual lessons. The study did not measure what students remembered months later or whether they finished the course or the degree.
Comparison time. You'll sometimes see "49 versus 75 minutes." That figure mixes tutor use with a full class period that includes tests. The appropriate comparison uses the study's own instructional-time accounting, which is roughly 49 against roughly 60.
Design generality. The result does not establish that any generic chatbot equals the studied tutor, that a tutor replaces a university course, or that the short-run advantage persists into later coursework.
Cost. Building the tutor required substantial expert preparation. The learning result must not be converted into a claim that the course became cheaper to deliver. (Study design and limitations, Development process)
What the study supports is the proposition that a carefully engineered tutoring workflow can outperform a strong instructional comparison on selected immediate outcomes. That is a real finding. It is also the most commonly overstated finding in this report's corpus, because "GPT tutor beat Harvard physics class" travels much farther than "a purpose-built instructional package designed by the course's own experts outperformed that same course's active-learning sessions on immediate posttests across two lessons, at unknown total cost, with no long-term measure." Both sentences describe the same study. Only one of them is evidence.
Common misreading: "Students learn twice as much from AI tutors." The study reported an adjusted 0.63 standard deviation advantage on immediate posttests. It did not report a multiple of learning, and it did not test a generic tool.
For a student: the takeaway is not "use ChatGPT for physics." It is that tutoring built around structured problems, expert solutions, and step-by-step progression can work in the short run. If your school offers a tutor built that way by your own instructors, that's a different product from an open chat window, and the evidence here is about the first one.
For a professor or department chair: the design ingredients are the finding. If you want to replicate the result, budget for the expert preparation, keep the active-learning comparison, and add the delayed measure the original study didn't have.
4.3.2 The College Evaluation Standard I'd Hold Every Pilot To#
A college pilot should distinguish two legitimate goals: helping students produce work, and helping students become able to produce comparable work independently. Both are fine goals. They are not the same goal. The program should declare which one is primary before collecting results, because the measurement sequence differs and the temptation to promote the easier goal after the fact is constant.
Here is the assessment sequence I'd propose, in order:
Baseline task. What can the student do before assistance?
Assisted practice. The tool is available.
Unaided assessment. The tool is removed.
Delayed assessment. Weeks later, still without the tool.
Transfer task. A problem that differs from the practiced examples.
For high-stakes claims, graders should be blind to assignment where feasible, and the comparison should preserve instructional time or explicitly account for differences. Recommended measures:
course completion
independent reasoning
citation accuracy
delayed retention
academic-integrity incidents
accessibility
student verification time
instructor workload
total delivery cost
Satisfaction and confidence can supplement those measures. They should not replace them. A student can be delighted by a tool that is quietly eroding the skill they enrolled to acquire, and the high-school mathematics experiment in Chapter 3 is the standing demonstration: students with unrestricted assistance did better during assisted practice and then performed 17 percent worse than the control group once the assistance was removed.
Questions to bring to a faculty senate or board meeting about an AI pilot:
Which goal is primary, producing work or building independent capability, and when was that written down?
What happens on the unaided assessment, and is there a delayed one?
Who grades, and do they know which students used the tool?
What did the comparison group receive, and for how long?
What did the tool cost to build, license, and support, counting faculty time?
How many academic-integrity cases, accessibility complaints, and citation errors were logged?
4.4 AI in Skilled Trades Training: What the Welding Studies Show#
Now we get to the part of this chapter closest to where I've spent my working life. I want to be precise about how little the evidence covers, because the trades are where overclaiming costs the most. A sloppy essay costs a grade. A sloppy termination costs a lot more.
4.4.1 Virtual Welding: A Null Result Is Not Proof of Equivalence#
Wells and Miller randomly assigned 101 university volunteers to live welding, virtual-reality welding, or two mixed sequences. Participants received approximately 30 minutes of practice, and three certified welding inspectors evaluated their subsequent live welds using a structured rubric. (Journal of Agricultural Education article)
The four groups' average weld scores differed, but the overall comparison did not reach the conventional 5 percent threshold: F(3,97) = 2.235, p = .089. This finding establishes neither virtual training's superiority nor its equivalence. A nonsignificant difference is not a noninferiority finding. Treating it as one is among the most common statistical errors in technology procurement. (Statistical results)
Read that again, because it's the sentence that saves budgets. "No significant difference" means the study couldn't tell the groups apart with confidence. It does not mean the study showed they were the same. To show equivalence, you have to design for it up front, decide how close is close enough, and power the study to detect that. This study wasn't built to do that, and nobody should cite it as if it were.
The implementation conditions bound the result further. Approximately 69 percent of participants had prior welding experience. The virtual system's visual guidance cues were disabled, and its audio was not functioning. The short practice period and immediate assessment do not establish long-term retention, injury reduction, employment, or successful professional certification. The study is relevant to training design. It is not evidence that trades can be taught through simulation. (Participant and implementation details, Study scope)
Common misreading: "Research shows VR welding is as good as the real thing." It doesn't. It shows a short, partially impaired VR setup couldn't be statistically separated from live practice in a mostly experienced group. That's a reason to run a better study, not a reason to close a welding bay.
4.4.2 WeldAR: Promising Transfer on a Narrow Skill Measure#
The 2026 WeldAR manuscript describes a crossover study with 24 novice participants comparing augmented-reality guidance during actual welding with video instruction. Participants received both conditions in randomized order, followed by unassisted practice measurements. The authors identify the manuscript as accepted for CHI 2026. (WeldAR author manuscript)
The system displayed real-time guidance on torch movement and positioning. The study found improved unassisted motion performance: the first-period comparison favored augmented guidance with p = .032. Note what was measured. The endpoint was adherence to movement targets, not a certified weld-integrity test or an occupational qualification.
The limits are real. The sample was small. The comparison was video instruction, not individualized coaching from an expert. Crossover order and carryover complicate interpretation. Tracking problems, headset weight, visibility, and the short observation period further limit any deployment claim. (Study results and endpoint definitions, Limitations and implementation findings)
The right role for this kind of support is supplementary practice with qualified supervision and independent verification. It should not authorize unsupervised work, replace required safety instruction, or let a motion score stand in for inspection of the finished work. The distinction between a motion endpoint and a weld-integrity endpoint is the whole story. The study measured whether trainees moved the torch the way the system asked. It did not measure whether the weld holds.
For a training director: a device that improves how trainees move is worth piloting as extra practice reps. Put it under an instructor, keep the inspection standard exactly where it is, and track whether AR-trained cohorts pass the same destructive or visual tests at the same or better rates. That's the number you'd need before changing anything else.
4.4.3 Infrastructure Trades and the Competence Sequence#
Let me start with the boundary. The welding studies concern welding. They are not about data center electrical commissioning, high-voltage work, refrigeration qualification, or emergency response. Their findings cannot establish competency in those different occupations, and this report will not transfer them. (Virtual welding study, WeldAR scope)
That leaves a gap between what the research covers and what people in my industry are going to be asked to decide. So here is how I think about it, labeled as what it is: operator experience, not evidence.
On a large power site, the work that can hurt people is the work where the equipment is energized, the energy is stored, or the system is under pressure. The people doing that work need to recognize a hazard they have never seen in exactly that configuration, stop, and call it. That capability isn't a fact you retrieve. It's judgment built through supervised repetition, checked by someone qualified who is willing to say "not yet."
A model can help someone prepare for that. It can explain a concept three different ways, quiz someone on a procedure, walk through why a sequence matters. I've got no objection to any of that. What it cannot be is the controlling authority on what's safe to do next. A training pipeline whose safety-critical step is "the model said so" is not a training pipeline. It's a liability with a login.
Here's the reason, stated plainly. On an energized site, the procedure is not a suggestion, and the person following it has to own it. If the model is right nearly every time and wrong once, the one wrong answer is the one that matters, and a trainee who learned to defer to the screen won't catch it. The skill you're building is exactly the skill of catching it. So any pipeline that routes the final safety judgment through a tool is training the opposite of what the job requires.
For a proposed infrastructure training program, the relevant sequence looks like this:
Occupational analysis. Define the actual tasks, hazards, and decisions in the job, not a generic job title.
Qualified instruction. A person who holds the relevant qualification teaches it.
Controlled practice. Deenergized, simulated, or mockup equipment where mistakes are cheap.
Independent practical assessment. Someone other than the instructor, and certainly not the tool, verifies that the trainee can perform the task.
Supervised field experience. Real equipment, real conditions, a qualified person present and accountable.
Documented authorization. A written record that this person is cleared for this specific work.
Model-generated explanations can assist preparation at steps 2 and 3. They should not be the controlling safety procedure at any step. That's the line, and it doesn't move because a vendor demo looked good.
For a county commissioner or economic development official: when a facility proposal promises "local workforce training," ask which occupations, which of the six steps the program actually provides, who performs the independent assessment, and where the supervised field experience happens. A promise of AI-powered training that stops at step 3 is a promise of practice, not of qualified workers.
For a young person considering an electrical or mechanical trade: use AI tools to understand theory, check your math, and prepare for written exams. Don't use them to decide whether something is safe. The person who signs off on your work should be a qualified human, and one day that person should be you.
4.5 Apprenticeship Outcomes and Employer-Connected Pathways#
If the college and trades evidence is mostly about short-run tests, the apprenticeship and workforce evidence is where you finally get to see paychecks. Neither of the two programs here depends on a model. I include them because they are the benchmark. Any AI-enabled pathway that wants to claim workforce value should be measured against programs like these, on outcomes like these.
4.5.1 Apprenticeship Outcomes: Positive Estimates With Selection Limits#
A September 2025 U.S. Department of Labor evaluation examined registered and unregistered apprenticeship programs supported by Scaling Apprenticeship and Closing the Skills Gap grants. For registered apprenticeships, it compared participants with public workforce-service customers and with selected community-college students, using weighting and regression rather than random assignment. (Department of Labor evaluation)
At the ninth quarter after enrollment, registered-apprenticeship participants showed estimated employment advantages of 7.8 percentage points over the workforce-service comparison and 8.8 percentage points over the community-college comparison. Estimated quarterly earnings differences were $3,230 and $4,693 respectively, in the report's stated units. The two comparisons used different samples: 1,271 apprentices for the workforce-service comparison and 1,124 for the community-college comparison. (Employment estimates, Earnings estimates)
Figure 4.6Data
Registered apprenticeship: estimated employment and earnings advantages
Quarter-specific historical estimates from a quasi-experimental Department of Labor evaluation. Many apprentices were already employed; unobserved selection cannot be fully removed.
Now the limits, which are as important as the numbers. Many apprentices were already employed when they enrolled. Statistical adjustment cannot fully eliminate differences in employer selection, motivation, or other unobserved characteristics. Think about who gets into a registered apprenticeship: someone an employer chose, who chose the program back, and who showed up. Those same traits could raise earnings on their own. Weighting and regression can account for what the researchers could measure. They can't account for what they couldn't. (Samples and identification limitations)
One discrepancy belongs on the record. The report's executive summary, results chapter, and Table A.10 give the lower earnings estimate as $3,230, while the report's Chapter 6 gives $3,320. I retain the $3,230 estimate supported by the results table rather than quietly reconciling the conflicting prose. If you check the report yourself, you should find the same number I quote, and you should also find the one place where it doesn't match.
These figures should stay what they are: quarter-specific historical study estimates. They are not annualized promises to prospective students. Multiplying a ninth-quarter difference by four and printing it on a flyer turns a careful estimate into a sales claim. They also don't measure the contribution of any model, and they don't imply that every apprenticeship or every entering worker receives the same benefit.
Common misreading: "Apprenticeship adds a fixed amount to your yearly pay." Nobody in the report said that. The study reported a quarterly difference at one point in time against one comparison group, from a design that can't fully rule out selection.
For a young person weighing an apprenticeship: these results are encouraging, and they're one reason I'd never tell a 19-year-old that a four-year degree is the only serious path. Ask the specific program for its own completion and placement records, including people who didn't finish.
4.5.2 Year Up Earnings: The Long-Run Benchmark for the Whole Package#
The Year Up evaluation randomly assigned 2,544 eligible young adults aged 18 to 24 to program access or to a comparison group that could pursue other community services. The intervention combined six months of training with six months of internships, alongside stipends, advising, mentoring, professional-skills development, and job-placement support. (ACF long-term evaluation)
The wage-record analysis covered 2,495 participants, and the prespecified confirmatory outcome was average quarterly earnings in follow-up quarters 23 and 24. The estimated impact was $1,895 per quarter, approximately 28 percent above the comparison group's $6,901 quarterly average, with p < .001. The study followed outcomes for approximately seven years. It did not observe a ten-year or lifetime effect, and the report's longer projections are assumptions about future persistence, not additional measured follow-up. (Confirmatory result, Follow-up and projections)
Figure 4.4Data
Year Up: average quarterly earnings, follow-up quarters 23 and 24
2,544 young adults randomized; wage records for 2,495. Observed for about seven years; longer projections in the report are assumptions. The package was not a model intervention and its components were not isolated.
The benefits shouldn't hide the opportunity costs. Reported earnings were lower for the treatment group in the first year, while participants devoted time to the program, with stipends reported separately rather than counted as wage earnings. That dip is real money for a young adult with rent due. And the package's long-run earnings effect did not establish improvement in every measured education, employment, or well-being outcome. (Annual results and other outcome domains)
Year Up supplies the benchmark this report's thesis needs: a workforce pathway evaluated on sustained outcomes rather than course registrations. It does not show which component produced the gain, whether the training, the internship, the stipend, the mentoring, or the combination. And it does not show that adding a model to a less intensive program reproduces the result. The defensible lesson is about the shape of credible workforce evidence (randomization, prespecified long-run earnings outcomes, published opportunity costs) more than about any single number.
That's the whole game for anyone proposing an AI-enabled workforce program. Year Up's design is what "proven" looks like in this field. It took years of follow-up, a real comparison group, an outcome chosen in advance, and the willingness to publish the first-year dip. A program that reports enrollments, satisfaction scores, and "learners served" is reporting inputs. Year Up reported earnings seven years out.
4.5.3 Designing a Workforce Pathway That Can Be Evaluated#
Workforce research programs should compare an assisted program with a credible alternative pathway, including conventional instruction and employer-supported practice. A comparison with "no training" can be informative, but it doesn't establish that the proposed approach is the best use of the money.
The proposed delivery package should specify, in writing:
paid or unpaid participation
childcare and transportation support where relevant
prerequisites
employer commitments
occupational assessments
placement assistance
the consequences of noncompletion
These are design requirements for evaluation. They are not findings that any particular program already meets them. I list them because every one is a place where a program can look good on paper and fail the participant. An unpaid program loses the people who can't afford to go without wages. A program without transportation support loses the people who live farthest from the site. A program without real employer commitments trains people for jobs that aren't there.
Most young people's first direct encounter with AI in the job market is the resume. The evidence here is better than you might expect, and more specific than the headlines.
4.6.1 AI Resume Writing Study: Presentation Without Invention#
Wiles, Munyikwa, and Horton studied algorithmic writing assistance in a randomized experiment involving approximately 481,000 new registrants on an online labor platform. The intervention suggested improvements to resume writing. It did not generate a complete fictional professional history. The research was subsequently published in Management Science in 2025. (Research and publication record, detailed author manuscript)
In the detailed manuscript's main analysis of 194,701 approved profiles with nonempty resumes, the first-month hiring effect was approximately 0.247 percentage points against a control hiring rate of 3.093 percent, about 8 percent in relative terms. "Eight percent more likely to be hired" must not be rewritten as an eight-percentage-point increase. The first is a small bump on a small base. The second would be enormous. (Manuscript estimates and sample)
The analysis did not find evidence of lower employer satisfaction, but a nonsignificant satisfaction difference is not proof of exact equivalence. The study concerns hiring on a particular platform. It is not about first-ever employment for all participants, a national increase in available jobs, or long-run career stability. (Study outcomes and scope)
A number without a denominator is a rumor, and this study is the cleanest example in the chapter. Roughly 3 in 100 control-group profiles got a hire in the first month. With help, it was a little more than that. Useful, especially at platform scale. Not a job guarantee.
The acceptable use follows from the design: improve the clarity of accurate information supplied by the applicant. A workforce program should prohibit invented credentials, fabricated experience, false references, and unreviewed submissions. Assistance that polishes a truthful resume and assistance that fabricates one are different interventions with the same interface. The boundary between them has to be drawn in policy, not left to the model's discretion.
Speaking as someone who has hired for jobs where a fabricated qualification can get someone hurt: a resume that overstates what a person can do isn't a small ethical lapse on a power site. It puts the wrong person in front of the wrong equipment. Every experienced hiring manager I know checks the claims that matter. The resume that wins is the clear, true one.
4.6.2 Workplace Support: A Meaningful, Setting-Specific Result#
Brynjolfsson, Li, and Raymond's 2025 Quarterly Journal of Economics article studied a staggered introduction of a conversational assistant among 5,172 customer-support agents. The assistant suggested responses and relevant information, and agents could accept, edit, or ignore the suggestions. The preferred estimate was approximately 15 percent more customer issues resolved per hour, with larger benefits for less-skilled and less-experienced workers. This was a difference-in-differences analysis of deployment in one company, not a randomized national workforce experiment. (Published study, Productivity estimates and design)
The heterogeneity cuts both ways. The study examined performance during outages, when the assistant was unavailable, and found patterns consistent with learning. It also reported limited benefits and some quality deterioration among the most-skilled workers. Those findings caution against both extremes: assuming assistance always substitutes for learning, or assuming every worker benefits equally. (Heterogeneity and learning analyses)
The result supports testing supervised assistance during onboarding in comparable workflows. It does not establish a 15 percent wage gain, 15 percent fewer employees, or a transferable 15 percent benefit in electrical work, medicine, law, or other unrelated occupations. Output per hour is a real outcome. It is not a wage outcome, a hiring outcome, or an economy-wide outcome, and Chapter 5 follows where each of those conversions breaks down.
For an employer onboarding new hires: this is the most encouraging finding in the chapter for your purposes, with a condition attached. The gains were largest for newer workers, and the outage evidence suggests some of what they learned stuck. Track quality along with speed, watch your most experienced people for the deterioration the study reported, and schedule periodic work without the assistant so you can see what your new hires can do on their own.
4.7 Protecting Competence in the First Years of Work#
The customer-support study is the optimistic case: assistance that seems to leave some learning behind. The next study is the adverse case, and it's the one I'd most want a new graduate, a new hire, and their manager to read together.
4.7.1 The Trio Learning Experiment: An Adverse Randomized Result#
Shen and Tamkin's 2026 preprint reports a randomized experiment with 52 programmers learning the unfamiliar Python Trio library. The assisted group had access to a GPT-4o-based coding assistant. Both groups then took a quiz without model assistance. The assisted group scored lower on immediate comprehension, with a reported standardized difference of Cohen's d = 0.738, p = .010, while the average task-completion-time difference was not statistically significant. The assessment covered concepts, code reading, and debugging after a short task. It did not measure long-term professional performance. (Study methods, Main results and limitations)
Put those two results together. The assisted group understood the library less well, and it didn't finish meaningfully faster. That's the worst of both: the skill cost without the speed payoff.
The source presentation requires care. The research page reports mean scores of 50 percent and 67 percent. The manuscript reports a 4.15-point difference on a 27-point quiz and describes it as 17 percent. Those statements don't produce one arithmetically interchangeable percentage measure, so I use the reported standardized effect rather than harmonizing the conflicting descriptions. The authors also identified interaction patterns associated with stronger or weaker quiz performance, but participants were not randomized to those patterns. That observational analysis does not prove that a particular prompting technique causes better learning. (Anthropic research summary, manuscript wording, Exploratory analysis)
This is vendor-affiliated research, and the authors' affiliations should stay visible. Its adverse result is informative precisely because the incentives ran the other way: a lab that sells assistance published a study showing assistance eroding comprehension. I respect that, and I'd like to see more of it from every company in this industry, including mine.
The study does not establish that all assistance, all technical work, or all tutoring produces the same effect. What it establishes, for this specific configuration, is the pattern Chapter 3 documented in high-school mathematics showing up in adult technical learning: successful assisted execution alongside measurably weaker unaided understanding, with no speed benefit to offset it. (Affiliations and limitations)
Figure 4.5Comparison
Competence erosion appears in two different populations
Two designs, two populations, the same pattern: assisted execution alongside weaker unaided understanding. Neither shows that all assistance erodes skill.
Two populations, two designs, one warning. In the high-school mathematics experiment covered in Chapter 3, students given unrestricted assistance performed 17 percent worse than the control group on later exams taken without it. In the Trio experiment, adult programmers who learned with an assistant scored lower on an unaided quiz by d = 0.738. The populations differ, the tools differ, and the measures differ, so the two numbers can't be combined or compared directly. What they share is the direction: help during practice, less capability after.
Common misreading: "AI makes programmers worse." The study says that in this setup, learning a new library with this assistant for a short task led to lower immediate comprehension. It says nothing about experienced programmers using assistance on familiar work, and nothing about whether the gap persists.
For a new hire: the first year of a job is when you build the understanding you'll draw on for the next ten. If your employer hands you an assistant, use it, and also make yourself do some of the hard parts without it. You want to be the person who can debug the thing when the assistant is wrong.
For a manager: the Trio study and the customer-support study point the same practical direction. Measure what new people can do without the tool at regular intervals. If output is climbing and unaided capability is flat or falling, you're borrowing against your future bench.
4.7.2 The Maintenance Rule for Productivity Claims#
METR's February 2026 update revisited its earlier finding that experienced open-source developers took 19 percent longer with early-2025 tools. Later estimates pointed toward faster completion, approximately 18 percent less time for returning participants and 4 percent less for newly recruited participants. But their confidence intervals (-38% to +9% and -15% to +9%) included no effect, and the researchers described selection and measurement problems that prevented a reliable updated estimate. These results should not be presented as a confirmed reversal or a definitive current productivity uplift. (METR update)
The implication is a maintenance rule I propose for every productivity claim: attach the tool version, task, worker experience, observation period, and evaluation design to every estimate, and recheck after substantial product or workflow changes rather than treating a historical study as a permanent coefficient. A productivity number without a date is not information. It's a fossil.
I run infrastructure, and in infrastructure we recommission. A system that tested fine at startup gets tested again after a major change, because the old test no longer describes the new system. Productivity claims about AI tools deserve the same discipline. The tools change, and the workflows around them change with them. A study from last year describes last year's tool.
The studies above anchor individual stages of a longer pathway. The complete process below is proposed, not proven. Each stage has a separate outcome and a decision gate, so success at one stage can't hide failure at the next.
Stage
Proposed interaction and human responsibility
Outcome to measure
Decision gate
Pathway choice
Learner and advisor compare actual prerequisites, costs, completion requirements, and employment pathways; models may summarize linked records
Accuracy of understanding, accessible options, ability to explain trade-offs
No steering based only on a market forecast or tracker score
Follow-up includes noncompleters and adverse outcomes
None of the reviewed studies evaluates this entire process as one connected treatment. The research anchors differ by stage: enrollment (Pounce), learning (the physics tutor and Let's Chat), practice (the welding studies), pathways (apprenticeship and Year Up), hiring (the resume study), workplace performance (customer support), and competence (Trio and METR). The gates exist precisely because the evidence is stage-specific. An institution that treats any single stage's evidence as evidence for the whole pathway has recreated, at organizational scale, the confusion between assisted performance and acquired skill that runs through this entire report.
Here's a hypothetical to show how the table works. A community partnership announces an "AI-powered pathway into technical careers." It cites the physics tutor for learning, the resume study for hiring, and the customer-support study for job performance. Run it through the gates. Did anyone measure unaided learning in this program's courses, with a delayed test? Did technical practice end in an independent assessment, or in a simulator score? Did the employer provide supervised work-based experience, or a logo on the brochure? Does the follow-up include people who dropped out? Each "no" is a stage where the borrowed evidence stops applying. The program might still be good. It just hasn't shown it yet.
4.9 For Students and Job Seekers: How to Use AI Without Losing the Skill#
Everything above was written for decision-makers. This section is for the 19-year-old at the kitchen table, and for anyone at the start of a working life. It isn't a new set of findings. It's what the chapter's findings imply when you're the one using the tool.
Before you open an assistant, ask yourself the same question I'd ask a college pilot: am I trying to produce this piece of work, or am I trying to become someone who can produce it? Both are legitimate. A cover letter due tonight is a production task. A physics problem set in the course you need for your major is a learning task. The Trio experiment and the high-school mathematics experiment both show what happens when a learning task gets treated like a production task: the work gets done, and the understanding doesn't arrive.
Try it first, then ask. Attempt the problem on your own before you get help. The baseline step in the college evaluation standard exists for a reason; you need to know what you can do alone.
Use structured help over open-ended answers. The physics tutor that worked was built around guided problems and expert solutions, not a model handing over results. If your course offers a tutor designed by your instructors, prefer it to a general chatbot for coursework.
Test yourself without the tool. Close the window and redo a problem from scratch. Then do it again a week later. If you can't, you practiced with a crutch, and now you know.
Try a problem that looks different. Transfer is the real test. If you can only solve the version you practiced, you learned the example, not the idea.
Keep your resume true. The resume study found that clearer writing helped hiring on one platform. It tested polish, not invention. Use assistance to say accurately what you did, and never to add what you didn't.
Get the important answers from a person. For enrollment deadlines, the Pounce design is the model: procedural questions to the assistant, financial aid, eligibility, and exceptions to staff who can put it in writing.
Never let a tool make a safety call. If you go into a trade, use AI to study, and learn to make the safety judgment yourself under a qualified supervisor. That's the skill employers in my world are paying for.
When you compare a degree, a trade, and a job, the evidence in this chapter suggests some questions worth asking each option:
A college: What does the institution measure about its AI tools? Does anyone check whether students can do the work without them?
A training program or apprenticeship: What happens after the simulator? Who performs the independent practical assessment, and is there supervised field experience? What are the completion and placement records, including people who didn't finish?
An intensive workforce program: Is it paid? The Year Up evaluation showed a real long-run earnings gain along with a first-year earnings dip, so plan for the dip.
A first job: If they hand you an assistant, will anyone help you build skills you can use without it?
4.10 What AI in Higher Education Research Permits You to Claim#
The selected postsecondary evidence is concentrated in a small number of subjects and work settings. It should not be described as a comprehensive evaluation of higher education, every skilled trade, or every form of early employment. Unanswered questions include persistence through graduation, independent writing and research across semesters, accommodations and outcomes for learners with disabilities, differences by language and access, effects in community-college technical programs, transfer across occupations, and full delivery costs. These define research needs, not presumed benefits or harms.
The defensible public claim is this: specific forms of intelligent support can improve particular educational and employment-related outcomes under documented conditions, while other configurations produce uncertain results or weaken independent learning.
The stronger claim, that a new facility produces a better local education-to-employment pathway, remains unestablished here. It would require a specified local program, actual access, employer participation, independent skill evidence, placement and retention records, and a separate accounting of the facility's community effects. Chapter 10 takes up those community effects directly.
Claim
Status
Describe the studied outcome, population, comparison, date, uncertainty, and implementation conditions
Permitted
Convert short-term learning into lifetime earnings
Not permitted
Convert applicant hiring on one platform into aggregate job creation
Not permitted
Convert announced infrastructure investment into verified local placements
Not permitted
Use the seven Savrn trackers to investigate the physical and institutional context of potential training pathways
Permitted
Describe the trackers themselves as proven educational interventions
Not permitted
Propose an accountable research partnership
Permitted
Describe such a partnership as already operating without documentation
Not permitted
That table applies to Savrn as much as to anyone. We build around behind-the-meter power and a closed-loop cooling design with a zero-makeup-water goal, and those are publisher statements until someone can check them. The same goes for any workforce benefit we might someday describe. A tracker entry starts a question. It never ends one. If we propose a training partnership, you should expect a named program, named employers, independent assessment, and placement records before we call it a result.
4.11 Conclusion: Specific Help, Measured Capability#
The research favors specificity over either enthusiasm or rejection. The useful questions are what the assistance does, for whom, against which alternative, over what period, and with what effect on the person's ability to perform without it.
The transition from education into work carries two distinct obligations that no tool can merge: help people complete real tasks, and maintain the competence needed to judge the result. The enrollment chatbot, the physics tutor, the resume tool, and the customer-support assistant all show that the first obligation can be met under the right conditions. The Trio experiment and the high-school mathematics result show that the second can quietly fail at the same time. The welding studies show how little we know yet about the trades. The apprenticeship and Year Up evaluations show what credible, long-run workforce evidence looks like, and none of it depends on a model.
The seven Savrn trackers can support source literacy and investigation of infrastructure-related opportunities, but the proof of educational and workforce value has to come from measured human outcomes. The pathway from a facility announcement to a young adult's paycheck passes through too many unproven stages to be walked in one claim. I'd rather walk it one gate at a time, and show you each one.
What does AI in higher education research actually show?
It shows specific wins under specific conditions. A Georgia State enrollment chatbot raised enrollment among committed students by 3.3 points, and a purpose-built Harvard physics tutor beat an active-learning class by 0.63 standard deviations on immediate tests. But the chatbot's full-sample effect was not significant, the tutor study had no long-term measure, and a separate experiment found weaker unaided understanding after AI-assisted learning.
Did an AI tutor really beat a Harvard physics class?
A custom GPT-4 tutor built with expert-written solutions and structured prompts outperformed in-person active-learning sessions by 0.63 standard deviations on immediate posttests across two lessons, with 194 eligible participants. The study did not measure long-term retention or course completion, did not test a generic chatbot, and required substantial expert preparation, so it does not show lower teaching costs.
Does using AI to write a resume help you get hired?
In a randomized experiment on one online labor platform, resume writing suggestions raised first-month hiring by about 0.247 percentage points on a 3.093 percent base, roughly 8 percent relative, in an analysis of 194,701 profiles. That is not an 8-point jump. The tool improved presentation of real experience; it says nothing about fabricated credentials, which should never be used.
Can VR or AR replace hands-on skilled trades training?
The evidence does not show that. A virtual welding study of 101 volunteers found no significant difference among groups (p = .089), which is not proof of equivalence, and cues and audio were impaired. A 24-person AR study improved unassisted torch motion (p = .032) but did not test weld integrity. Neither addresses electrical, high-voltage, or other infrastructure trades.
Do apprenticeships increase earnings?
A September 2025 Department of Labor evaluation estimated that registered apprentices had employment 7.8 and 8.8 points higher, and quarterly earnings $3,230 and $4,693 higher, than two comparison groups at the ninth quarter. It was not randomized, many apprentices were already employed, and selection could not be fully ruled out. One report chapter lists $3,320 instead of $3,230.
How much does Year Up increase earnings?
A randomized evaluation of 2,544 young adults aged 18 to 24 found quarterly earnings $1,895 higher in quarters 23 and 24, about 28 percent above the $6,901 comparison average (p < .001). Earnings were lower in the first year while participants trained, the study observed about seven years, and it does not identify which program component caused the gain.
Can using AI while learning make you worse at a skill?
In one randomized experiment, 52 programmers learning a new Python library with a GPT-4o-based assistant scored lower on an unaided quiz than those without it (d = 0.738, p = .010), with no significant time savings. It was a small, short study from vendor-affiliated authors, and it does not show the same effect for all tasks or experienced workers.
Does AI make new workers more productive?
In one company, a conversational assistant raised customer issues resolved per hour by about 15 percent among 5,172 support agents, with larger gains for newer and less-skilled workers. It was not randomized, the most-skilled agents saw limited benefit and some quality decline, and the result does not translate into wages, headcount, or other occupations.
Chapter 5 · Work and Career Change
Will AI Take My Job? Exposure, Displacement, and Career Change
It's late. You can't sleep. You open your phone and type the question: will AI take my job? You're not asking about productivity statistics. You're asking whether the paycheck that covers the mortgage, the kid's braces, and the parent you help take care of will still be there in three years. That's the real question, and most of what you'll find in the search results answers a different one.
I've spent my career building and operating large infrastructure, most of it in Texas, and I've read a lot of spreadsheets. Here's something every operator learns early: output per hour and a worker's paycheck are different line items. They sit on different rows. One can go up while the other stays flat, or goes down, and nothing about the first number decides the second. What decides the second is a person, usually in a meeting you're not invited to, choosing how the gains get split. That's the part of the AI and jobs debate almost nobody talks about, and it's the part that matters most to you.
This chapter looks at the best evidence I could find on AI job displacement research, working life, and career change. Four study programs anchor it: a large workplace deployment in customer support, a global index of which occupations are exposed to generative AI, a Danish study that tracked earnings and hours after ChatGPT arrived, and a randomized evaluation that followed a job training program for ten years. The uncomfortable part cuts both ways. The evidence doesn't support the prophets of mass unemployment, and it doesn't support the people promising that everyone gets a raise. Both guarantees fail, and I'll show you exactly where. Then I'll give you something more useful than a forecast: what to do if your occupation shows up on an "exposed" list, how employers can split gains in good faith, what workforce boards should demand, and why an announced facility, including the kind my industry builds, is not the same thing as a job in your town.
5.1 Will AI Take My Job? The Question Behind the Question#
The productivity question (can a person produce more in an hour with assistance) is the question the technology press asks. It's also the smallest of the questions a working adult actually lives with.
Think about what you want from work. Stable earnings, for one. But also safer work, useful training, some control over your schedule, a way to move up or move over, and maybe the ability to stay employed while you care for a parent or a child with needs. Any serious account of AI at work has to hold all of those outcomes in view at once, because a tool can raise output per hour while making several of them worse. The research this chapter reviews contains exactly that combination, and I'm not going to hide it.
So when you ask "will AI take my job," I'd suggest breaking it into four smaller questions that evidence can actually address:
Will the tasks in my job change? This is an exposure question. The ILO index speaks to it.
Has that change shown up in people's earnings and hours yet? This is a measured labor-market question. The Danish study speaks to it.
When a workplace adopts a tool, who gains and who doesn't? This is a deployment question. The customer support study speaks to it.
If I need to change careers, does training work? This is a transition question. WorkAdvance speaks to it, with an important caveat: it didn't test AI training at all.
This chapter covers the middle of the working lifespan: established employment, career change, and economic participation. Chapter 4 covers college and the start of a career; Chapter 2 covers the task-level productivity experiments in more detail. Where those chapters matter here, I'll point you to them.
5.2 The AI Productivity Customer Service Study, Read Carefully#
The strongest workplace evidence in this research remains Brynjolfsson, Li, and Raymond's 2025 study in the Quarterly Journal of Economics. If you've seen a headline that AI made customer service workers 15 percent more productive, this is almost certainly where it came from. It deserves a careful read, because the headline and the study are not the same thing. (Brynjolfsson, Li, and Raymond)
Here's what the study actually did. The researchers studied 5,172 customer support agents at a business-software company. The company rolled out an AI assistant that suggested responses and surfaced relevant information while agents handled customer issues. The agents kept control: they could accept a suggestion, edit it, or ignore it entirely.
The design was quasi-experimental. The assistant was adopted in a staggered way across the workforce, and the researchers used difference-in-differences estimation, which compares how outcomes changed for agents who got access against how they changed for agents who hadn't yet. That is a respectable design. It is not a company-wide randomized trial, and the distinction matters for how much weight the estimate can carry. In a randomized trial, a coin flip decides who gets the tool, so nothing about the agents or their managers can explain the difference. In a staggered rollout, the timing of access was not a coin flip, and the analysis has to assume that the groups would have moved in parallel without the tool.
Access increased issues resolved per hour by approximately 15 percent on average. Read that again, and focus on the words "on average." The gains were larger among less-experienced and lower-skilled workers. The most skilled workers saw little productivity improvement. And the study reports evidence of small quality declines for some higher-skilled workers. (Published study)
Figure 5.2Data
The 15 percent that wasn't one number
The published average hides the structure: gains concentrated among newer and lower-skilled agents, with the most skilled gaining little. Bar lengths show direction for the subgroups, not measured magnitudes.
Here's the part most people skip. The internal structure of this result matters more than the headline. The gains concentrated among newer and lower-skilled workers. That pattern is consistent with the assistant passing along pieces of the organization's accumulated know-how to the people who hadn't absorbed it yet. It's the same compression pattern Noy and Zhang observed in their professional writing experiment, which Chapter 2 covers in detail: the tool narrows the gap between the least and most experienced.
At the other end, the most skilled workers gained little and showed some quality deterioration. That's consistent with a tool whose suggestions are tuned to the typical case becoming a distraction, or an anchor, for people whose judgment already exceeds it. If you've ever had a well-meaning helper hand you the textbook answer when you already knew the exception, you know the feeling.
Put those together and something odd emerges. A deployment that quotes the 15 percent average is describing a workforce that does not exist. No individual worker experienced the average. The workers nearest the tool's design center (the experienced, skilled agents whose best practices the tool arguably drew on) benefited least.
5.2.3 What the study doesn't tell you about wages#
This is evidence of a particular workflow's effect. It is not a universal labor-productivity multiplier you can apply to any job. It does not establish that the average worker received a corresponding wage increase. It does not establish that staffing could be cut by the same percentage without other consequences. (Study scope and outcomes)
The 15 percent was the company's throughput. Whether any of it reached the agents' paychecks is a question the study could not answer. That's not a criticism of the researchers; it's a statement about what the study measured. It is, however, a sharp warning about how the result gets used. If you hear someone say "AI raised productivity 15 percent, so workers will earn more," they've converted a throughput number into a wage claim the evidence never made. That's the AI impact on wages question in one sentence: we don't know from this study, and anyone telling you otherwise is forecasting.
One bookkeeping note. This study also appears in Chapter 4, where it anchors the discussion of onboarding new workers. Its repeated use here connects career stages. This report counts it once, not twice, as evidence.
5.2.4 Common misreadings of the customer support result#
Misreading 1: "AI makes workers 15 percent more productive." It raised issues resolved per hour by about 15 percent on average for one workflow at one company. That's a narrower sentence, and it's the true one.
Misreading 2: "So the company can cut 15 percent of its staff." The study explicitly doesn't establish that. Customer demand, quality, escalation rates, and the knowledge that experienced agents carry all sit outside the throughput number.
Misreading 3: "AI helps everyone." The gains were uneven. The most skilled agents gained little, and some showed small quality declines.
Misreading 4: "AI is bad for experts." The study shows little productivity gain and some small quality declines among some higher-skilled workers in this workflow. It does not show experts were harmed across the board, or that the same would happen in another job.
If you're newer in your job: the pattern suggests assistance may help you get up to speed faster. That's real value. But speed on assisted tasks isn't the same as knowing the job. Chapter 4 describes a technical-learning experiment in which assisted work coexisted with weaker unaided understanding. Keep checking whether you could do the task without the tool.
If you're experienced: watch for the anchor effect. If the tool's suggestion pulls you toward the typical answer when your judgment says the case is unusual, trust the judgment and document why. Your quality is part of your value, and the study suggests it's the thing most at risk for people like you.
If you're a manager: don't report the average without the spread. Ask how gains and quality changed by experience level, and ask what happened to pay.
5.3 Exposure Is Not Displacement: ILO Generative AI Exposure and the Danish Evidence#
The customer support study tells you what happened in one workplace. It can't tell you what's happening across an economy. Two evidence programs measure the labor market at that larger scale, and together they tell a two-part story that resists every sweeping narrative.
5.3.1 The ILO exposure index: one in four, and 3.3 percent#
The International Labour Organization's refined 2025 index estimates that approximately one in four workers globally is in an occupation with some exposure to generative AI. It estimates that 3.3 percent of global employment falls in its highest exposure category. (ILO working paper)
The index classifies task exposure. It is not an estimate that one in four jobs will disappear. The distance between those two sentences is where most public commentary lives.
What does "exposure" actually mean? It measures which tasks a system could, in principle, affect. Whether it actually affects them depends on things the index doesn't contain: adoption, demand, regulation, costs, complementary skills, and employer decisions. Exposure can legitimately motivate investigation. It's a good reason to look at how tasks are changing, what training people need, and what choices an organization is making. It cannot determine an individual worker's future.
So here's a translation I'd like you to carry around. When you hear "exposed occupation," hear "occupation whose task list overlaps the tool's capabilities." Don't hear "occupation on its way out." Those are different claims, and only the first one is what the index measures.
5.3.2 The Danish study: still waters, rapid currents#
The second program is a study of Denmark's labor market. Its March 2026 revision is now titled Still Waters, Rapid Currents. It links survey evidence with administrative records and finds precise null effects on earnings and recorded hours over its early observation period. It rules out effects larger than approximately 2 percent for those measured outcomes over roughly two years after ChatGPT's launch, while documenting changes in work and occupational movement. (Revised NBER working paper)
"Precise null" is a phrase worth understanding. A lot of studies fail to find an effect because they're too small or too noisy to detect one; that's an imprecise null, and it tells you very little. A precise null is different. It means the study had enough statistical power to say not just "we didn't find an effect" but "if there were an effect on these outcomes, it would have to be smaller than about 2 percent." That's a real finding.
This is a working paper about a particular country, period, and set of outcomes. It is not evidence that future displacement cannot occur. It is not evidence that no individual was affected. It is not evidence that task changes have no welfare consequences.
What it does establish is economically meaningful. Two years after the most widely adopted general-purpose AI tool in history entered one of the world's most digital labor markets, the measured earnings and hours effects were statistically indistinguishable from zero, and tightly enough to exclude anything above roughly 2 percent. The water was still.
The study's own title records the other half. The currents beneath the surface (task content, occupational movement) were moving. The measured outcomes simply hadn't caught them yet. That's why I'd never cite this study as proof that "AI won't affect jobs." It says the earnings and hours line hadn't moved in that window. It also says the work itself was changing.
Figure 5.1Comparison
Exposure is not displacement
Left: ILO refined 2025 index. Right: the Danish linked survey and administrative study (working paper, March 2026 revision) shown schematically; the gray band marks the approximately 2% bound the study rules out, not a plotted interval.
Held together, the two programs discipline prediction in both directions.
The exposure index says the task overlap is widespread and real. The Danish study says that in the first two years, in one country, that task overlap did not convert into measured earnings or hours displacement. Both can be true at the same time, because the conversion from exposure to displacement depends on decisions (adoption, pricing, reorganization, demand) that neither study measures.
Anyone who tells you which way those decisions will fall is forecasting, not reporting. That includes people selling fear and people selling reassurance. It includes me, which is why I'm not going to give you a forecast.
Here's a simple table that separates what each study measured from what people often claim it says.
Study
What it measured
What it found
What it does not show
ILO refined index (2025)
Task overlap between occupations and generative AI
About one in four workers globally in an occupation with some exposure; 3.3% of global employment in the highest category
That any share of jobs will disappear
Danish study (March 2026 revision)
Earnings and recorded hours, linked survey and administrative data
Precise nulls; effects above roughly 2% ruled out over about two years; changes in work and occupational movement documented
That future displacement can't happen, or that no individual was affected
Customer support study (2025)
Issues resolved per hour at one company
About 15% average increase, concentrated among newer and lower-skilled agents
A wage increase, or that staff could be cut by 15%
5.3.4 Common misreadings of exposure and displacement#
"One in four jobs will be replaced by AI." The ILO index says about one in four workers is in an occupation with some exposure. Exposure is task overlap. The index doesn't estimate replacement.
"The Danish study proves AI doesn't cost jobs." It found precise nulls on earnings and recorded hours over roughly two years, in one country. It documented changes in work and occupational movement. It explicitly doesn't rule out future displacement or individual harm.
"If my occupation is in the top exposure category, I'm in trouble." The top category covers 3.3 percent of global employment. Being in it tells you the task overlap is high. It tells you nothing about what your employer will decide, what customers will demand, or what new tasks will appear.
"No effect on earnings means no effect on workers." Earnings and hours are two outcomes. Task content, stress, autonomy, and the path to the next job are others. The Danish study's own title tells you the currents were moving.
Let's say you've seen a list, or an online calculator, or a news graphic, and your occupation is on it. Maybe it's in the highest category. Here's practical guidance built only on what the evidence in this chapter supports. It's not a forecast. It's a way to keep your options open while the actual decisions are still being made.
5.4.1 Translate "exposed" into your own task list#
The ILO index works at the level of occupations and tasks. You work at the level of your actual week. Sit down and list what you do: the tasks, roughly how much time each takes, and which ones require your judgment, your physical presence, your license, or your relationships.
Then sort them into three rough piles:
Tasks a tool could plausibly draft or retrieve. First drafts, summaries, standard responses, looking things up.
Tasks where the tool might help but you must verify. Anything where a fluent wrong answer would cause harm.
Tasks that remain human by requirement or by nature. Safety-critical physical work, licensed judgments, relationships with customers or patients, supervising others.
This is exactly the first stage of the working-life process in Section 5.6. You're doing for yourself what a responsible employer should do for the whole team.
5.4.2 Find out what your employer is actually deciding#
Exposure becomes displacement (or doesn't) through decisions. So find out what decisions are on the table. Good questions to ask your manager, your union representative, or HR:
Which tools are we adopting, for which tasks, and on what schedule?
How will we measure whether the tool helps? Will quality and rework be measured, or only speed?
If time is saved, what happens to it: pay, workload, new responsibilities, or reduced staffing?
Will workers be able to see what's measured about them and challenge incorrect records?
What training is offered, and will it be tied to actual job requirements?
You may not get complete answers. But the answers you get, and the ones you don't, tell you a lot.
The customer support study found the biggest gains among newer workers. Chapter 4 describes an experiment in which programmers learning an unfamiliar library with an assistant scored lower on an unaided comprehension quiz afterward (Shen and Tamkin preprint). Put those two findings side by side and a practical rule falls out: use the tool, but regularly check that you can still do the core of your job without it. Your unaided competence is portable. It goes with you if the tool changes, the vendor changes, or you change employers.
Career change AI tools can produce a polished resume and cover letter in minutes. A polished application should not be mistaken for readiness to perform the role. Before you invest in a pivot:
Get verified job requirements from actual employers, not from a generated summary.
Look for independent skill assessment, not just course completion.
Ask any training program for its placement, retention, and earnings records, and whether it has a comparison group. Section 5.5 explains why.
Here's a hypothetical to make this concrete. No real person, and no statistics beyond what the chapter already cites.
Imagine a claims processor at a regional insurer. She reads a news story saying her occupation is in the "highest exposure" category and spends a sleepless night assuming her job is gone.
In the morning, she lists her week. A good part of it is reading documents and drafting standard letters, the kind of work a tool could plausibly draft. Another part is spotting inconsistencies in claims and deciding when something needs a closer look, work where a fluent wrong answer would be costly. And some of it is talking with upset customers and training the newest hire, work that stays human.
She asks her supervisor the questions in Section 5.4.2. She learns that the company is piloting a drafting tool, but hasn't decided what to do with any time saved, and hasn't planned to measure rework. That tells her two things. First, the decision about her job hasn't been made yet, so exposure hasn't become anything. Second, there's a gap she can help fill: she volunteers to help track errors during the pilot, which puts her closer to the judgment work and gives her a record of quality she can point to later.
She also keeps doing a share of her drafting without the tool, so her unaided skill stays sharp. Nothing in this example guarantees her job. It replaces a forecast she couldn't verify with information she could.
5.5 Does Job Training Work? What WorkAdvance Shows After Ten Years#
If the answer to "will AI take my job" is "maybe some tasks, and the decisions aren't made yet," the natural next question is whether retraining works. WorkAdvance is this chapter's non-AI benchmark, and it earns its place by testing what real workforce development looks like when evaluated rigorously. (MDRC ten-year report)
The program tested employer-connected sectoral training. That means training built around the needs of specific industries and connected to employers in them, rather than generic job-readiness classes. It's the kind of intervention that AI enthusiasts sometimes assume becomes obsolete the moment a tool arrives.
The randomized evaluation followed 2,564 participants across four providers. Recruitment began in 2011 to 2013. The evaluation used long-term administrative earnings records, which means outcomes came from official records rather than from what participants remembered or reported. And because it was randomized, the control group was set by lottery, which makes the comparison clean.
At St. Nicks Alliance, year-ten earnings averaged $35,218 for the program group and $26,638 for the control group. That's a difference of $8,580, with a reported p-value of .005. The other three providers did not show statistically significant effects on the report's year-ten confirmatory earnings outcome. (Ten-year results)
Figure 5.3Data
WorkAdvance at year ten: provider variation is the finding
One provider produced a large, durable earnings gain; three structurally similar providers did not reach significance on the confirmatory outcome. Providers two to four are unnamed here to avoid implying a ranking.
The provider variation is the finding. One provider's program produced a large, durable earnings gain a decade out. Three structurally similar providers produced none that reached statistical significance.
That's worth sitting with. Same program model. Same evaluation. Same randomized design. Wildly different results depending on who ran it and where. The result is not a claim that St. Nicks Alliance's approach will transfer to your town. It demonstrates two things at once: that durable gains from training are possible, and that implementation and context decide whether you get them.
5.5.3 What WorkAdvance does and doesn't tell us about AI#
WorkAdvance reminds us of a distinction this report enforces everywhere: the trial did not test a generative AI training curriculum. It tested sectoral training.
So whatever AI can add to training, the claim "the tool replaces the program" has no support here. The program's own results depend on details no tool configures: the provider, the employer connections, the local labor market. If three of four well-designed providers didn't produce a significant ten-year earnings gain, a chatbot that promises to "reskill" you by itself is making a claim nobody has tested.
The reverse misreading is also wrong. WorkAdvance doesn't show that training is futile. One provider produced a gain of $8,580 per year, ten years out, in a randomized design. That's the kind of result any workforce leader should want to understand and try to reproduce, carefully, with measurement.
5.5.4 What this means for workers weighing a training program#
Ask any program three questions. Does it connect to documented employer demand? Does it track completion, placement, retention, and earnings? Does it have, or will it allow, a comparison group? A program that can answer all three is behaving like the evidence says training should behave. A program that answers with testimonials is asking you to trust it.
For a workforce discussion involving an infrastructure company like mine, the implications are concrete: connect training to documented employer demand, and track completion, placement, retention, and earnings. That is not a basis for advertising a guaranteed return to a particular course or facility. Chapter 4's permitted-claims discipline applies unchanged here.
So far, I've described what studies found. Now here's what I propose. The following process is for employers, workforce boards, training institutions, and workers. It is evidence-informed, but it has not been evaluated as an integrated program. I'm presenting it as a proposal, not a proven intervention.
Stage
Decision
Required evidence
1. Map work
Which tasks are changing, and which responsibilities remain human?
Observed task inventory, safety and professional requirements, worker input
2. Establish baseline
What constitutes successful work today?
Quality, throughput, rework, customer outcomes, workload, and compensation
3. Choose support
Is the need information, practice, workflow redesign, or formal qualification?
Map work. Exposure lists work at the occupation level. Real decisions happen at the task level, in a specific workplace. Observed task inventories (what people actually do, not what the job description says) plus worker input are how you find out which tasks are changing and which responsibilities have to stay human for safety, professional, or legal reasons.
Establish baseline. You can't tell whether something improved if you didn't measure it before. The baseline includes compensation and workload on purpose. If you measure only throughput before and after, you'll only be able to report on throughput.
Choose support. Not every problem needs an AI tool. Sometimes the need is better information, more practice, a redesigned workflow, or a formal qualification. Comparing with simpler tools and with human support keeps you clear about whether the new tool is the right answer or just the newest one.
Test bounded use. Try it on defined tasks, compare performance, analyze errors, and check unaided skill. That last piece matters because Chapter 2 documents a consulting experiment where assistance helped on some tasks and hurt on a task outside the tool's reliable competence, and Chapter 4 documents weaker unaided comprehension after assisted learning.
Train and qualify. The test is whether the person can perform the actual job, demonstrated independently under relevant conditions. Course completion is not the test.
Distribute gains. This is a decision, not an outcome. More on this in Section 5.7.
Follow the transition. Watch retention, earnings, hours, injury or error records, and what workers say about the experience over time, not just in the launch month.
First, the worker should be able to understand what is measured and challenge incorrect records. Monitoring intensity is not itself a productivity benefit, and an employer's output gain should not be silently represented as an improvement in employee welfare. The 15 percent in Section 5.2 was the company's throughput. Whether any of it reached the agents is a question the study could not answer, and a process that never asks it has taken the employer's side by default.
Second, the gains-distribution stage is a decision, not an outcome. Time saved can become higher pay, lighter workload, new responsibilities, or layoffs. Which one happens is an organizational choice that evidence should document rather than assume.
Here's a hypothetical walk-through. Picture a mid-sized property management company adopting an AI assistant for tenant communications.
Map work: The team lists the tasks. Routine maintenance acknowledgments and lease reminders could be drafted. Disputes, fair housing questions, and anything involving safety stay with trained staff, by requirement.
Establish baseline: Before launch, the company records response times, complaint rates, how often letters need rework, staff workload, and current pay bands.
Choose support: The team asks whether better templates would solve most of the problem. Some of it, yes. The pilot proceeds only for the parts templates don't handle.
Test bounded use: For a defined period, some staff use the tool on routine messages. The company tracks errors, especially confident wrong statements about lease terms, and has staff periodically handle messages without the tool.
Train and qualify: New staff must show they can handle a dispute call and a lease question without assistance before working on their own.
Distribute gains: Management writes down what happens to saved time. In this hypothetical, they choose to move staff toward in-person inspections and tenant relationships rather than cut positions, and they say so in writing. A different company might choose otherwise; the point is that the choice is recorded.
Follow the transition: Over the following year, they track retention, error records, and what staff and tenants say.
Notice what this example doesn't include: a promised productivity percentage. The process doesn't need one. It needs the records.
5.7 For Employers: How to Split the Gains in Good Faith#
If you run a business, this section is for you. I run one too, and I'll tell you what I believe: the question of how productivity gains get split is the most important decision in this whole debate, and it's yours.
When a tool saves time, that time goes somewhere. Broadly, it can become:
Higher pay for the people whose work got more productive.
Lighter workload, meaning less overtime, less burnout, more time on hard cases.
New responsibilities, meaning people take on work that used to be out of reach.
Reduced staffing, meaning layoffs or not replacing people who leave.
All four are legal. All four are choices. The evidence in this chapter doesn't tell you which to choose. What it tells you is that none of them happens automatically. The customer support study measured throughput; it did not show a wage pass-through. If your plan assumes one will happen by itself, you don't have a plan.
Write down the decision. Before rollout, document what you intend to do with saved time. After rollout, document what you actually did.
Report both columns. Put business outcomes (throughput, quality, customer outcomes) beside worker outcomes (earnings, hours, workload, retention). A report with only the business column is an investor update, not an evaluation.
Measure by experience level. The customer support study found gains concentrated among newer workers and small quality declines among some of the most skilled. Your averages will hide the same kind of spread unless you break them out.
Protect your experts' judgment. Give experienced staff explicit permission to override the tool and a way to flag when suggestions pull them in the wrong direction.
Let workers see and challenge their records. More monitoring is not more productivity.
Count the people who leave. If workers who struggle with the tool quit or stop using it, your remaining-user numbers will look better than reality. Section 5.9 explains why.
Don't advertise a wage effect you haven't measured. If pay went up, show the records. If it didn't, don't imply it did.
I'm an operator, and I'll give you an operator's reason, not a sermon. The skilled workers in the customer support study gained little and some showed small quality declines. Those are the people who carry your institutional knowledge, the knowledge the tool arguably draws on for everyone else. If you treat the tool's average as a reason to squeeze them, you're drawing down the very asset that made the tool useful. Splitting gains in a way you can defend in writing isn't charity. It's how you keep the people who know how the place actually works.
Different working contexts require different boundaries. The research reviewed here does not justify filling the gaps between them with invented effect estimates, so I won't. The following are proposed implementation boundaries, not study findings for each named population.
Established professionals. Use assistance for preparation, retrieval, and drafting while testing whether review remains substantive. Faster output should not remove responsibility for judgments the professional is qualified to make. Chapter 4's technical-learning experiment shows that assisted execution can coexist with eroding comprehension in exactly this kind of population.
Trades and technical workers. Use bounded practice and documentation without replacing physical demonstration, supervised experience, or required qualification. A generated explanation of a procedure is not authorization to perform hazardous work. I've spent enough time around high-voltage and heavy mechanical systems to say this plainly: the written procedure is where safety starts, and the supervised hands-on sign-off is where it's proven.
Small businesses. Compare total operating costs and corrected outcomes, not only the subscription price. A service that saves drafting time but creates compliance or customer-service rework may not be beneficial. The correction column of the ledger decides, not the drafting column.
Workers changing careers. Require verified job requirements and independent skill assessment. A polished application should not be mistaken for readiness to perform the role.
Workers with care responsibilities or disabilities. Measure accessibility and usable flexibility directly. Do not assume that remote software access removes scheduling, assistive-technology, or transportation barriers.
5.8.1 A small-business worked example (hypothetical)#
Imagine a two-person bookkeeping practice that subscribes to a drafting assistant for client emails and monthly summaries. The subscription is cheap, and the drafts come fast. But a few summaries mislabel expense categories, and catching those errors takes the owner's evening. A client notices one before she does.
On the drafting column, the tool looks like a win. On the correction column, it's closer to a wash, maybe worse once you count the client's trust. This is the small-business version of the whole chapter: you have to count the rework, not just the speed, before you know whether the tool is helping. No figure in this example is a statistic; it's a way to set up your own ledger.
The proposed evaluation reports business and worker outcomes side by side: task quality, total review time, customer outcomes, worker earnings, hours, workload, retention, and independent competence where relevant. The pairing is the safeguard. A program that reports only the business column is an investor update, not an evaluation.
Attrition needs special treatment, because it's the subtlest failure mode in this research. If people who struggle with a system leave the workplace or stop using the tool, then measuring only the remaining users overstates success by construction.
Think about what that means. Suppose a company rolls out a tool, and the workers it doesn't suit either quit or quietly stop using it. Six months later, the company measures the users who remain and reports great numbers. Those numbers can be accurate about the people measured and still be misleading about the program. The surviving-user average is the measurement equivalent of the consulting study's outside-frontier task, described in Chapter 2: it looks like evidence, and it's an artifact of who got counted.
Training evaluation should retain a comparison group where feasible and report the offer of participation separately from completion. WorkAdvance's provider variation is a reminder that successful participants and successful programs are not interchangeable analytical categories. (MDRC evaluation)
Here's why the distinction matters. "People who completed our program earn more" may just mean that the people most likely to succeed anyway were the ones who finished. "People offered our program earn more than a comparable group who weren't offered it" is a claim about the program. The first is about participants. The second is about the program. WorkAdvance's randomized design is what let it make the second kind of claim, and it's why the provider-by-provider results mean something.
5.9.3 Questions to ask any AI workplace or training report#
Show me the denominator. How many people started, and how many are in the numbers?
Is there a comparison group? How was it formed?
Are worker outcomes (earnings, hours, workload, retention) reported next to business outcomes?
Are results broken out by experience or skill level?
Is quality reported, or only speed?
Are the results from people offered the tool or program, or only from those who stuck with it?
5.10 For Workforce Boards: What to Demand Before You Fund or Endorse#
Workforce boards sit at a pressure point. You're asked to endorse training programs, partner with employers, and respond to announcements of new facilities. Here's what the evidence in this chapter supports.
WorkAdvance is the clearest lesson: one of four providers produced a significant ten-year earnings gain, and three did not. The program label was the same. What differed was implementation and context. So when a provider proposes an "AI-ready" or "future of work" curriculum, ask about the things that made the difference in a real trial: connection to documented employer demand, the quality of the provider, and whether outcomes will be tracked.
A training seat is a commitment. A completed credential is an outcome. A placement in a job is a different outcome. A job still held a year later is another. Keep them in separate columns in every report you publish. Chapter 4 lays out the claims that are and aren't permitted when training meets infrastructure investment, and it applies here unchanged.
The ILO index can help you decide which occupations in your region deserve a closer look. It can't tell you which workers will lose jobs. Use it to prioritize conversations with employers about task changes, not to announce which jobs are disappearing.
Imagine a county workforce board hears that a large technology facility has been announced nearby. A local official proposes a new training program marketed as preparing residents for "hundreds of jobs."
A board applying this chapter would ask: Which employer has documented demand for which roles? Are those construction roles or ongoing operating roles? (Section 5.11 explains why that matters.) How many operating roles are direct employees versus contractors? What qualifications do they require? Will the program track placement, retention, and earnings, and report offer separately from completion?
If the answers come back as a press release, the board has learned that the demand isn't documented yet. It can still fund exploration and employer conversations. What it shouldn't do is market a guaranteed job pipeline that no employer has committed to.
5.11 Announced Capital Is Not a Local Job: The Tracker Boundary#
This section is about my own industry, so I'll be especially careful.
Capital, permits, and project schedules can identify questions about future labor demand. They do not establish vacancies, hiring dates, qualifications, wage levels, or durable employment. Those must come from employers, executed contracts, training partners, and observed hiring outcomes, not from project announcements alone. (Capital Atlas scope, Permits Tracker scope, Delay Watchlist scope)
The forbidden inference deserves exact statement, because it's the one Savrn has the strongest commercial temptation to make: announced capital is not a local job.
A billion-dollar facility announcement may produce hundreds of construction jobs for a defined period and dozens of permanent operating roles. Or the schedule may slip. Or the phase may be cancelled. Or the work may be contracted to specialists brought in from elsewhere. The seven Savrn trackers are instruments for discovering which of those worlds a community is in. They are not evidence of employment outcomes, and the chapters of this report that quote them never use them as such. A tracker entry starts a question. It never ends one.
If you're a resident, reporter, or board member, here's how I'd use them. Use the Capital Atlas to see what's been announced and ask who the employer is and what roles are documented. Use the Permits Tracker to see what's been permitted and ask what construction schedule the permit implies and who's doing the work. Use the Delay Watchlist to see whether a project's timeline has moved, and ask what that does to any training pipeline built around it. In every case, the answer you're looking for comes from an employer, a contract, or a hiring record, not from the tracker.
5.12 Conclusion: What the Evidence Supports About AI and Your Job#
The defensible working-life goal is not to predict a single labor-market future. It's to improve a worker's practical ability to adapt, verify competence, and share in measured gains, while preserving evidence of adverse outcomes.
The research supports optimism of a disciplined kind. AI assistance measurably raised throughput in one well-studied customer support workflow, with the gains concentrated among newer workers and small quality declines among some of the most skilled. Sectoral training produced a large, durable earnings gain at one of four providers, and none that reached significance at the other three. And two years of data from Denmark after ChatGPT's arrival showed earnings and hours effects near zero, precise enough to rule out anything above roughly 2 percent, with task content and occupational movement shifting underneath.
So, will AI take my job? The future isn't written in any of these numbers. It will be written in the adoption, pricing, and organizational decisions that the numbers can't see. That's why the process in this chapter puts the worker's outcomes beside the company's. And it's why the burden of proof stays where this report has kept it from the first page: on the claim, not the skeptic. Capability is not benefit. Who gets the benefit is a decision, and you have every right to ask who's making it.
The best evidence reviewed here doesn't support a single answer. The ILO estimates about one in four workers globally is in an occupation with some exposure to generative AI, but exposure means task overlap, not job loss. A Danish study found precise null effects on earnings and hours, ruling out effects above roughly 2 percent over about two years after ChatGPT's launch, while noting changes in work. What happens next depends on employer decisions.
What does it mean if my occupation is exposed to AI?
In the ILO's refined 2025 index, exposure means an occupation's tasks overlap with what generative AI could, in principle, affect. About one in four workers globally is in an occupation with some exposure, and 3.3 percent of global employment is in the highest category. The index doesn't include adoption, demand, costs, regulation, or employer decisions, so it can't tell you whether your job will disappear.
Did AI make customer service workers more productive?
In a study of 5,172 agents at one business-software company, AI assistance raised issues resolved per hour by about 15 percent on average. The design was quasi-experimental with staggered adoption, not a randomized trial. Gains were larger for newer and lower-skilled agents; the most skilled saw little improvement and some showed small quality declines. The study does not show that wages rose.
Does AI raise wages?
The evidence reviewed here doesn't show it. The customer support study measured throughput, about 15 percent more issues resolved per hour on average, not pay, and does not establish a corresponding wage increase. The Danish study found precise null effects on earnings, ruling out effects above roughly 2 percent over about two years. Whether productivity gains become higher pay is an employer decision.
Has AI already reduced earnings or hours?
Not in the Danish evidence. The March 2026 revision of the study, titled Still Waters, Rapid Currents, links survey and administrative data and finds precise null effects on earnings and recorded hours over roughly two years after ChatGPT's launch, ruling out effects above about 2 percent. It covers one country and period and does not prove future displacement can't happen or that no individual was affected.
Does job training work if I need to change careers?
Sometimes, and it depends heavily on who runs it. In the randomized WorkAdvance evaluation of 2,564 participants across four providers, St. Nicks Alliance participants earned $35,218 in year ten versus $26,638 for controls, a gain of $8,580 (p = .005). The other three providers showed no statistically significant year-ten earnings effect. The trial tested sectoral training, not AI training.
Will a new data center in my area create local jobs?
An announcement alone can't tell you. Virginia's JLARC describes a typical 250,000-square-foot data center as employing about 50 full-time workers in ongoing operations, roughly half contractors, while construction can peak around 1,500 workers over roughly 12 to 18 months. Those are different jobs over different periods. Real answers come from employers, executed contracts, and observed hiring records, not capital announcements.
What should I ask my employer about AI at work?
Ask which tools are being adopted and for which tasks, how success will be measured (including quality and rework, not only speed), and what happens to saved time: higher pay, lighter workload, new responsibilities, or reduced staffing. Ask whether you can see and challenge what's measured about you. The customer support study showed throughput gains without establishing wage gains, so the split is a decision worth asking about.
Chapter 6 · Households and Families
AI for Families: What Actually Helps Households, and What Doesn't
It's 10 p.m. The kids are finally asleep, and you notice the permission slip. It's due tomorrow, it needs a signature, it mentions a fee, and the pickup time on page two doesn't match the one in the email. That is the moment most people are thinking about when they ask whether AI for parents is worth anything. Not a demo. Not a benchmark. The question is simple: will this thing take work off my plate, or will it hand me a new job checking its work at the worst hour of the day?
I want to answer that question the way I'd answer it about a power contract or a cooling system: by looking at what was measured, on whom, against what. A household is not a task list. It's a small institution with no HR department, no legal team, no on-call clinician, and no separation between its chief financial officer and the person who spots that permission slip at 10 p.m.
Here's what the evidence shows, and I'll give you the uncomfortable part up front. The burden is real, large, and unevenly shared. The interventions with measured benefits for families are almost all structured programs with a human, an institution, or an expert built into them. The direct evidence that a general-purpose chatbot saves families time is thin: a survey of comfortable respondents and a pilot with five mothers. And in the two places where the stakes are highest, medical decisions and loneliness, randomized trials found the tools pointed in the wrong direction. None of that means you should keep AI out of your house. It means you should give it a job description.
6.1 What Should AI for Parents Actually Be Asked to Do?#
Household assistance is not one outcome, because household life is not one outcome. A system can help a parent remember a deadline, complete an application, support a child's learning, reduce energy use, or help care for an aging relative. The same system can fail at diagnosis, at relationship support, at privacy, or at the fair division of responsibility between the adults in the house. Success at one of those jobs tells you nothing about the others.
The studies I reviewed for this chapter split into positive, null, mixed, and adverse findings in roughly equal measure. So there's no defensible blanket claim in either direction, for or against. The answer depends on the task, the design, and who stays in charge.
The practical scenario Savrn asked me to test was a familiar one: a parent raising three children while a spouse works outside the home. I use that scenario later as a worked illustration. But I deliberately widened the lens to single-parent households, dual-earner households, multigenerational homes, grandparent-led families, households shaped by disability, and people doing unpaid caregiving for a relative.
The reason is simple. A tool should not be judged only in the household most able to configure it, supervise it, and correct it. A parent with time, a fast connection, fluent English, and a background in reading fine print can catch a chatbot's mistakes. An exhausted caregiver with a spotty connection may not. If a tool only works when the user is already well resourced, that's a finding about the user, not about the tool.
The standard of evidence follows the stakes. I gave the most weight to randomized field studies, administrative outcomes (did the person actually enroll, did the bill actually change), government time-use data, and peer-reviewed synthesis. I included direct evidence of general-purpose model use wherever it existed, and I say plainly when it is thin.
I also included older interventions that don't use generative AI at all: text-message programs, energy reports, human-assisted benefits outreach. Those studies show which mechanisms have measured value inside real households, and they give you a credible alternative to hold any new AI product against. If a cheap text program already moved an outcome, an expensive chatbot has to beat it, not just exist.
Here is the conditional claim this chapter defends, and I'd ask you to hold it whole rather than quoting half of it: assistance should reduce verified burden or improve a defined outcome while leaving consequential judgment, relationship responsibility, and appeal rights with people.
Every word in that sentence is doing work. "Verified" means somebody measured it. "Defined outcome" means you said in advance what better looks like. "Leaving judgment with people" means the tool can prepare the decision, not make it. Convenience, usage, confidence, and an attractive answer are how a product feels, not household outcomes. Capability is not benefit, and the gap between the two is where families get hurt.
6.2 How Much Time Do Parents Spend on Childcare? The Burden Baseline#
Before you can say a tool saves time, you need to know how much time is on the table. The best national source in the United States is the American Time Use Survey, run by the Bureau of Labor Statistics.
6.2.1 What the 2025 American Time Use Survey Shows#
According to the 2025 American Time Use Survey release from the Bureau of Labor Statistics, adults living with a child under age six averaged 2.3 hours of primary childcare per day. Women in that group averaged 2.8 hours; men averaged 1.7 hours. Adults living with a child under 13 averaged 5.1 hours of secondary childcare per day, which BLS defines as being responsible for a child while doing something else.
Figure 6.1Data
The household burden baseline
Published BLS aggregates. These establish that care takes substantial time, not that any tool can recover it.
Those numbers are the scale that any household tool would have to move. They are also easy to misuse. Three disciplines apply.
First, don't add them together. The two categories use different populations (a child under six versus a child under 13) and measure different things. Primary care is attention devoted to the child: feeding, bathing, reading, driving to practice. Secondary care is supervision folded into something else, like cooking dinner while a toddler plays on the floor. Adding 2.3 and 5.1 manufactures a number that no one actually lived. Show me the denominator. When two figures have different denominators, the sum is a rumor.
Second, a baseline is not a remedy. These figures establish that household care occupies a large share of the day. They do not establish that any particular system can recover any of it. A parent reading to a four-year-old is not a process in need of automation.
Third, the 2025 survey has a hole in it. BLS notes that approximately 6,100 people were interviewed and that a federal shutdown created a period of missing data whose effect on the estimates BLS says cannot be quantified. A baseline with a known hole is still the best baseline available. It is not a precise one, and anyone using it to promise specific minutes back should say so.
6.2.2 The Mental Load Nobody Puts on the Calendar#
The size of the burden matters. So does who carries it. A peer-reviewed study of U.S. parents in the Journal of Marriage and Family separates two kinds of cognitive labor: the daily mental tracking that never appears on a calendar (who's out of clean socks, which form is due, what the pediatrician said last time) and episodic household planning (the summer schedule, the move, the holiday). It finds that mothers carry more of the core daily tasks.
This is the research most people have in mind when they search for AI mental load research. Here's the part most people skip: this study describes who reports responsibility. It does not test a technological remedy. A product that cites it as proof an app can move the load is borrowing credibility it hasn't earned.
6.2.3 The Closest Direct Test: Croatian Mothers and ChatGPT#
The closest direct evidence on a general-purpose model in household use is a Croatian mixed-methods study, and it's exploratory in exactly the places that matter.
Read that sample description again. Urban, highly educated, partnered, financially stable. Those are the households best equipped to get value out of a general tool and to catch its errors. Even there, the study measured perceived value, not minutes saved. And the most important finding for anyone hoping AI would rebalance the mental load is the quiet one: responsibility stayed where it was. If the mother still owns the list, and the chatbot just helps her work the list faster, the load hasn't been shared. It's been sped up.
A 2026 survey from Lurie Children's reported that 81 percent of 1,004 U.S. parents had used intelligent systems for parenting tasks, with reported average savings of 58 minutes per week.
That's a striking number, and it's worth knowing. It's also self-report. The published survey methodology does not specify a probability sampling frame, does not describe weighting, and does not give a separate respondent count for the time-savings question. So we don't know how representative the sample is, or how many people answered the 58-minute question.
The right way to read this: it's adoption evidence. It tells you parents are already using these tools and believe they help. It does not tell you what the tools cause.
Here's my summary of the baseline. The burden is real, large, and unevenly distributed. The evidence that general-purpose assistance reduces it is, as of this report's evidence cutoff, perceived value among comfortable survey respondents and five mothers in a two-week pilot.
That gap between the size of the problem and the strength of the remedy evidence is the central fact of this chapter. Everything that follows is about how to act sensibly inside that gap: use what's been shown to work, test what hasn't, and keep a hand on the wheel.
What this means for you as a parent: if an app claims it will give you hours back each week, ask what that number is based on. If the answer is a survey of its own users, you're looking at a satisfaction score, not a time study.
6.3 Can AI Help Families Apply for Benefits? Information vs. Completion#
If you want to know what actually helps households with paperwork, the strongest evidence I found doesn't involve a chatbot at all. It involves food assistance, older adults, and a very clean experiment.
Here's what the study actually did. One group got nothing new, one got information about SNAP, and one got that information plus trained staff who did the heavy lifting on the application. The enrollment results in NBER Working Paper 24652 form a ladder. Over nine months, 5.8 percent of the control group enrolled. Information alone raised that to 10.5 percent. Information plus human application help raised it to 17.6 percent. The treatment differences were statistically significant.
Figure 6.2Data
Information is not completion: SNAP enrollment after nine months
31,888 Pennsylvania households with a Medicaid-enrolled person 60 or older not on SNAP. The effect came from helping people enter the process, not from changing approval, and the helpers were people, not software.
Now look at the number most people skip. Among the people who actually applied, roughly 75 percent were approved in every arm. That tells you what the intervention did and didn't do. It didn't change who qualified or make the agency more lenient. It helped more people get into the process and through it. The barrier wasn't eligibility. It was completion.
The cost figures deserve the same care. The estimated intervention cost per additional enrollee was approximately $20 for the information arm and approximately $60 for information plus assistance. Those calculations left out applicant time, the government's processing costs, and the benefit payments themselves. They are the cost of the outreach, not the full cost or value of an enrollment. Anyone quoting "$60 per enrollee" as the whole economic story is quoting half a ledger.
6.3.2 What the SNAP Result Means for AI Benefits Help#
This study supports a process insight. It does not support automating eligibility.
The biggest jump on that ladder came from human application help layered on top of information. Real people assessed eligibility, handled documents, and followed up. That's a standing caution for any product that proposes to replace the helper with a chat window. Maybe a well-designed tool could reproduce part of that effect. Nobody has shown it yet, and the burden of proof sits with the tool.
What a household assistant can reasonably do is the preparatory work around the application: organize records, explain the official instructions in plain language, and help you write down the questions to ask the agency. What it must not do is decide whether you're entitled to something. An authorized institution decides that. Good assistance keeps four things intact: the original form, the agency's current rule, the exact version you submitted, and your right to a human appeal.
Common misreading #1: "Information worked, so an AI that explains benefits is enough." Information roughly moved enrollment from 5.8 to 10.5 percent. Adding hands-on human help moved it to 17.6 percent. Explanation is the smaller half of the effect.
Common misreading #2: "The approval rate stayed the same, so the help didn't matter." The approval rate staying near 75 percent is the evidence that the help worked without lowering standards. More people applied, and they were approved at the same rate.
What this means for a county benefits office or a nonprofit: if you're deciding between funding a chatbot and funding trained application assisters, this is the best evidence in the file, and it favors the assisters. If you pilot a chatbot, run it against the human option, not against nothing.
6.3.3 Lower-Risk and Higher-Risk Administrative Tasks#
Family administration is a long list: school forms, insurance claims, utility accounts, medical portals, benefits renewals. The tasks aren't equally risky, and the design should reflect that.
Lower-risk candidate tasks include:
Extracting dates from a school notice.
Creating a draft checklist from a set of instructions.
Comparing a draft form you've filled in with the official instructions.
Translating a message the family wrote, so a family member can review it before it's sent.
Higher-risk tasks include:
Asserting that someone is legally eligible for a benefit or program.
Submitting anything under penalty of perjury.
Accepting contract terms.
Spending money.
Disclosing a child's or a dependent adult's data.
The line between the two lists is consequence. On the first list, a mistake is annoying and catchable. On the second, a mistake can cost money, create legal exposure, or put a dependent's information somewhere it can't be pulled back.
No reviewed experiment establishes the reliability of a general-purpose model across these administrative tasks. That's worth saying twice, because product marketing tends to imply otherwise. If you're running a pilot, measure the things that matter: missed obligations, invented requirements, incorrect field entries, review time, completion, successful resolution, and correction effort. Do not measure the number of summaries generated. That's a throughput metric for a task whose value is accuracy. A tool that produces a thousand fast summaries with one invented deadline has not saved you anything on the day that deadline was the one you trusted.
6.4 AI Parenting Apps Study Results: What Structured Programs Show#
When people search for an AI parenting apps study, they usually hope to find a trial of the app on their phone. What the evidence actually offers is more useful and less flattering to open-ended chat: two randomized programs in which structure, humans, and small actionable steps did the work.
6.4.1 The Blended Digital and Human Program in China#
The trial methods describe a program adapted from ParentText, a rule-based chatbot. Caregivers received daily modules through WeChat and joined weekly or twice-weekly group discussions led by headteachers and social workers. It was not an unrestricted conversation with a generative model. It was not a software-only intervention either. That distinction is the meaning of the finding.
On the primary outcomes, caregivers in the program reported more early learning and stimulation at home (beta 1.79, 95% confidence interval 0.24 to 3.34) and less total caregiver-perpetrated violence (incidence-rate ratio 0.87, 95% confidence interval 0.80 to 0.96). Emotional violence, measured separately, did not differ significantly between groups. That null belongs in the headline alongside the positives.
The follow-up and limitations set hard boundaries. Once the immediate assessment was done, the waitlist classes got the program too, so there's no untreated comparison at twelve months and no causal long-term effect to claim. Twelve-month attrition in the original intervention group was 41.9 percent. Every outcome came from caregivers describing their own behavior.
So the study supports the blended program's immediate effects, in its setting. It doesn't show a long-term causal effect, it doesn't show the result transfers to other countries or schools, and it doesn't show a general model benefit.
The READY4K methods are almost aggressively simple. Three texts a week. One shares a fact. One suggests a specific activity. One follows up with encouragement. Control families weren't left with nothing; they got occasional administrative messages, which is a fairer comparison than silence.
Among the 821 children who had spring assessments, the pooled literacy effect was 0.109 standard deviations. The results varied: the first cohort's overall effect was not statistically significant, and the larger estimates showed up among children who started with lower skills. Parent and teacher survey response was incomplete. The analysis also has a 2019 Journal of Human Resources publication record.
A 0.109 standard deviation effect is small in absolute terms. For a program that costs a few texts a week and asks parents for minutes, not hours, that's a meaningful signal, especially for the kids who started behind.
6.4.3 Why Tiny, Structured Prompts Beat Open Chat (So Far)#
Put the two studies side by side and a pattern shows up. Both programs were structured. Both used predefined content. The Chinese program put trained humans in the loop every week. READY4K sent tiny, well-timed, actionable prompts and nothing else.
The READY4K mechanism is nearly the opposite of an open-ended conversational agent. It doesn't wait for the parent to think of a question. It doesn't produce long answers. It asks for one small action that fits into a routine. That difference should guide design, not be waved past.
What these studies do not establish:
That more messages are better. READY4K tested three a week. It didn't test thirty.
That an open chatbot should talk directly with a child. Neither program did that.
That a generated school communication is accurate without review by the parent and the school.
What this means for parents: the best-evidenced digital help for young kids' learning looks like a short nudge you can act on in two minutes, not a conversation you have to manage. If an app is built that way, it's closer to the evidence.
What this means for teachers and school leaders: if you're choosing a family engagement tool, look for predefined, reviewed content and a fixed, light cadence. Ask the vendor what the comparison group received in any study they cite.
6.5 AI Caregiver Support for Dementia: A Small Average, a Wide Future#
Many households aren't just raising kids. They're caring for a parent or grandparent, sometimes both at once. Dementia caregiving is one of the heaviest forms of unpaid work a family can take on, and it's an area where digital support has been studied for years.
In the JMIR review, the pooled burden effect was SMD -0.26 (95 percent confidence interval -0.42 to -0.10). Negative means less burden. The interval doesn't cross zero, so on average, the evidence favors intervention.
Then comes the number that should change how you read the first one. Heterogeneity was 73.6 percent, which means the trials disagreed with each other far more than chance alone would explain. And the 95 percent prediction interval ran from -1.10 to +0.58.
Figure 6.6Data
Average effect versus the next implementation: dementia caregiver eHealth
The confidence interval describes the average. The prediction interval describes what a new program might do, from substantial benefit to harm. Only three studies were low risk of bias; certainty moderate.
Read carefully, this is two findings, and they answer two different questions.
The confidence interval answers: what's the average effect across programs like these? Answer: a modest reduction in burden, and we're fairly sure it's real.
The prediction interval answers: what might the next specific program do? Answer: anywhere from a substantial benefit (-1.10) through little or no effect to an adverse result (+0.58), where caregivers end up more burdened.
If you're a caregiver deciding whether to try one particular app, the second question is the one you're actually asking. You aren't buying the average. You're buying one program, and the evidence says that program's effect is hard to predict from the category.
The review's risk of bias and certainty assessment adds more caution. Only three of the 35 studies were rated low risk of bias. The overall evidence was judged moderate certainty, and it relied heavily on subjective outcomes, meaning caregivers rating their own burden. The subgroup analysis found that human-supported interventions carried a larger point estimate than self-guided ones, but the test for a difference between those subgroups was not significant. And because the review covered web, mobile, video, and hybrid programs, it does not isolate what a general-purpose generative model does.
My summary: caregiver technology support is promising, modest on average, and unpredictable for any specific deployment. There's a hint, not a demonstration, that the versions with a human involved work better.
Notice the design feature of the Age and Ageing trial before anything else: every plan the system generated passed two senior dementia-care experts before a caregiver saw it. This was supervised by construction.
On the main outcome, the burden contrasts were null between groups. Among the 201 participants who completed follow-up, burden went down within the intervention arm, but the differences between the intervention and control groups at weeks six and twelve were not statistically significant. A within-arm decline without a between-group difference is exactly the pattern you'd expect if some of the improvement would have happened anyway. That's why control groups exist.
The reported safety results point the other way. The authors report 16 intervention participants and 29 controls with at least one safety event, p = .037. No individual category of safety event differed significantly on its own.
6.5.3 Two Inconsistencies I'm Recording, Not Fixing#
I found two inconsistencies in this paper that I can't resolve from the published text, and I'm recording them rather than smoothing them over.
First, the safety table reports the control group's percentage as 28.7 percent. But the paper states 103 control completers, and 29 affected participants out of 103 comes to approximately 28.2 percent. Those published counts and the percentage don't reconcile.
Second, the paper's analysis statements say both that all 201 completers were analyzed and that the analysis was intention-to-treat. Those can't both describe the same analysis in the usual sense, because intention-to-treat analyzes everyone randomized, which was 250.
Neither issue proves the result wrong. Both mean I keep the safety finding as reported and do not convert it into a precise promise that this kind of system reduces household risk by some specific amount.
The study methods and limitations add attrition, complete-case presentation, the heavy expert review, and the short 12-week window. Put it together, and the trial supports further testing of a supervised knowledge-graph system. It does not support autonomous caregiving, and it does not support replacing clinical evaluation.
6.5.4 What This Means for Caregivers and Older Adults#
If you're caring for someone with dementia: digital support has a real, modest average benefit in trials, and the ones with human support built in look at least as promising as the ones without. Before you commit, ask whether a person reviews what the tool tells you and whether you can reach a human expert. The best-designed trial in this section had both.
If you're an older adult whose family is setting up tools on your behalf: you should be told what the tool records about you, who sees it, and how to turn it off. That applies to the dependent adult as much as to the child.
If you're a health system or agency buying caregiver support: remember the prediction interval. The category's average won't tell you what your specific program will do. Plan to measure burden against a comparison group from day one.
6.6 Home Energy Reports: What Household Feedback Can Measure#
Energy is where household decisions and infrastructure meet, and it's where I have the most operating experience. So I want to be especially careful here.
The mechanism in Allcott's 2011 analysis is social comparison: you see how your use stacks up against your neighbors, plus a few tips. Because high-use households cut more than low-use households, the 2 percent average isn't the result for every family.
The boundaries here are multiple, and each one is load-bearing.
The outcome definition was kilowatt-hours. Not a guaranteed lower bill, not better reliability, not a model-based agent, and not a community's acceptance of a data center.
The welfare effect is ambiguous. The author notes that the household's costs of conserving weren't observed. Lower use bought with a colder house, extra effort, or forgone convenience isn't automatically a net gain for the family.
A useful household energy interface has to show its work. For any household infrastructure question, a good tool should show the consumption period, the rate structure, the units, the comparison group, and the actions actually available to that household. And it should never imply that resident conservation offsets an unmeasured industrial load, or shift responsibility for system planning onto families.
What this means for residents: if a utility or a company shows you an energy comparison, check the period, the units, and who you're being compared with. A lower number of kilowatt-hours is a measured result. A lower bill, a more reliable grid, or a "fair share" of a new facility's impact are separate claims that need separate evidence.
6.7 Where AI for Parents Must Stop: Medical Advice and Loneliness#
Two randomized experiments mark the firmest boundaries in this chapter. Both are about things families already use chatbots for: figuring out whether a symptom is serious, and having someone to talk to.
6.7.1 The ChatGPT Medical Advice Study: Strong Models, Weaker Decisions#
The study methods test the thing that actually matters. Researchers didn't just grade a model's answer. They measured what the people using it concluded.
When researchers prompted the models directly, each one performed strongly. The people using those models did not reproduce that performance. According to the human decision results, control participants had 1.76 times the odds of identifying a relevant condition compared with the pooled model users. On disposition, meaning what level of care to seek, accuracy did not differ significantly between each model group and control. Every group, with or without a model, tended to underestimate how urgent the situation was.
Read that again. The model knew. The person with the model did worse at naming a relevant condition than the person without it.
The limits are real. These were simulated scenarios, not people who were actually sick and scared. And the models have changed since the 2024 data collection. But the study shows why families shouldn't infer safe decision support from a model's medical test scores. The gap between what a model can do when an expert prompts it and what a worried parent decides after using it is the whole story, and in this trial it ran in the direction of harm.
What that means in your house: a household system can help you organize symptoms and list questions for a professional. It is not the final authority on emergency care.
6.7.2 The AI Loneliness Study: Daily Chats, More Loneliness#
The design and sample are large for this kind of question. Because the daily conversations were encouraged, not required, the study also measured whether people followed through. Per the intervention and first-stage result, assignment increased the number of days with personal conversations by about 5.8.
Now the outcome. The loneliness estimate went up, not down: a single-item loneliness measure increased by 0.168 points on a 0 to 10 scale, which is 0.057 standard deviations, with a multiplicity-adjusted p-value of .0002. The other primary results showed no significant effect on happiness, and several other well-being estimates were small and only marginal after adjusting for multiple comparisons.
The effect is small in size. It is also precisely estimated and in the opposite direction from the usual pitch.
The limits cut both ways. It was an encouragement design. It lasted one month. There was no alternative reflection control, such as asking a comparison group to keep a journal, so we can't tell whether the effect comes from talking to an AI specifically or from spending daily time on personal reflection. It's a working paper, and it didn't test specialized therapy tools. So it does not show that every structured support tool harms well-being.
What it does show: you cannot presume that replacing or supplementing personal interaction with unstructured model conversations will reduce loneliness. In this trial, the measured effect went the other way.
Taken together, these two studies set the chapter's firmest household boundaries. A household system organizes and escalates. It does not diagnose. It does not substitute for relationship.
These aren't stylistic cautions or my personal preference. Each boundary rests on a preregistered randomized experiment with a statistically significant adverse result. That's a higher evidentiary bar than most of the benefits claimed for these tools.
For older adults and the families who care about them: a companion app is not a plan for isolation. People are.
6.8 The Household Evidence Register: Strong vs. Perceived Evidence#
Here is every entry that informs this chapter, with what each contributes and where each stops.
ID
Evidence and design
Contribution
Principal boundary
H01
2025 ATUS; weighted national time-use survey
Household-care time baseline
Not an intervention effect; missing 2025 diary period
H02
Croatian survey n=369 plus uncontrolled five-person pilot
Direct exploratory evidence on generative household organization
No objective time outcome, randomization, control, or redistribution result
H03
Chinese cluster RCT, n=541 caregivers
Immediate parenting and child-protection outcomes from blended rule-based and human support
One preschool, self-report, no long-term untreated comparison
Adverse evidence on unstructured personal conversation and loneliness
One-month encouragement design; working paper; not specialized therapy
H11
Lurie Children's 2026 survey, n=1,004 parents
Reported uses, perceived burden, time savings
Adoption survey without causal or observed-time measurement
Figure 6.5Map
Where household evidence is measured and where it is only perceived
The eleven household register entries by evidence class. The only direct evidence on general-purpose household use (H02, H11) is exploratory or self-reported.
Two reading rules. First, the register counts evidence entries, not eleven independent randomized trials. H07 synthesizes many studies that may overlap with individual trial entries, so its 3,388 caregivers must not be added to the others as if they were a separate pool. Second, publication status is not the same as design. H10 is a working paper from CESifo. H11 is a public survey report from Lurie Children's. H01 is official statistics from BLS. H02 is exploratory mixed-methods research. Each deserves the weight of what it is, no more. The full evidence-class scale is defined in Chapter 11.
Scan the register by what was measured, and the split is plain. The measured outcomes (enrollment, literacy scores, kilowatt-hours, condition identification, loneliness scores, caregiver burden in trials) come almost entirely from structured programs or from adverse tests of general models. The evidence that general-purpose AI for parents saves time comes from self-report and a five-person pilot. That's the difference between strong evidence and perceived evidence, and it's the difference that matters when someone asks you to pay for a subscription.
6.9 A Household Operating Process for AI for Parents#
So what should a family actually do? Below is a proposed control process. I want to be clear about its status: it's a design derived from the evidence above, not a tested end-to-end household intervention. Each stage leaves an evidence trail and has a gate.
That looks heavy for a permission slip. In practice, most of it takes seconds: you know who needs what by when, you know it's low stakes, you don't paste in your kid's medical history, and you check the draft against the paper. The table matters most when the task gets serious, because that's when people skip steps.
Qualified professional or emergency service remains authoritative
Relationship judgment, discipline, emotional support
Offer bounded reflection, not substitution
People retain responsibility; serious risk escalates to human support
Figure 6.3Framework
The household authority ladder
Proposed default roles for a household assistant, from the ten-stage operating process. The higher the consequence, the narrower the system's role, and authority never leaves the person.
The architecture of that table is the chapter's thesis in one picture: the higher the consequence, the narrower the system's default role, and the authority never leaves the person.
6.9.3 A Family's AI Rules: What to Let It Do, What to Never Let It Do#
If you want something you can print and stick on the fridge, here's the authority matrix translated into house rules.
Let it do these, and check its work:
Pull dates, times, fees, and required items out of school notices and put them in a draft list.
Draft checklists, grocery lists, and meal ideas, with an adult checking allergies, prices, and what's actually in the pantry.
Summarize a school or agency document, as long as you can see which line of the original each point came from.
Translate a message the family wrote, so someone can read the translation before it goes out.
Help prepare a benefits application by organizing documents and explaining the official instructions.
Compare utility plans or bills, showing the period, the units, and the rate structure.
Make it ask first, every time:
Any purchase or payment.
Any application filing or account change.
Any sharing of a family member's data with anyone.
Any message sent to a teacher, agency, doctor, or other third party on your behalf.
Silence is not a yes. A shared family account is not a yes. The adult who owns the account approves the exact action, the exact amount, and the exact recipient.
Never let it do these:
Decide whether a symptom, a fall, a fever, or a medication question is an emergency. It can organize the facts and help you call. The professional decides.
Make a decision about a child's safety or protection.
Decide discipline, or settle an argument between family members.
Serve as a child's or a lonely relative's main source of company or emotional support.
Assert that your family is legally eligible for something, or sign anything under penalty of perjury.
Watch the household continuously and infer what you need without being asked.
And one rule over all the others: any family member with authority must be able to stop it, see what it did, and delete what it kept. A tool that can't be stopped isn't assistance. It's a new dependent.
6.9.4 Five Questions Before You Let an App Near Your Kids' Data#
The "minimize data" stage is the one families skip most, because it's invisible until something goes wrong. Before any app sees a child's name, school, health information, photos, or location, ask these five questions. If the app can't answer them in plain language, that's your answer.
What exactly will it collect, and does the task need it? A permission slip checklist needs a date and a fee. It doesn't need your child's diagnosis or home address. Redact what the task doesn't require.
Who else will see it? Look for the list of recipients, not just a promise. Unintended recipients are one of the privacy outcomes this chapter says a household pilot should measure.
How long does it keep it, and can I set that? The retention setting is part of the record in the ten-stage process for a reason.
Can I delete it, and will it tell me when deletion is done? Stopping use is not the same as deleting data. A household must be able to stop and recover.
Who in this family has authority to say yes? A child can't consent for themselves in the way an adult can, and a dependent adult's data needs the same care. If the person clicking "agree" isn't the person with authority and a justified need, stop there.
6.10 How to Test Whether AI Actually Helps a Household#
6.10.1 An Illustrative Workflow: A Parent, Three Children, One School Notice#
Here's a hypothetical. It's an illustrative workflow, not a real family, an observed deployment, or a promise of time savings.
A parent is caring for three children while a spouse is at work. A school activity notice comes home. The parent gives the assistant a redacted copy and asks for a checklist. The assistant identifies the date, the pickup arrangement, the required items, the fees, and the consent requirements, and it attaches each field to the passage in the notice it came from. Where the notice doesn't answer something, the assistant says so instead of guessing.
The parent compares the draft with the notice. The pickup time conflicts between two pages, so the parent checks with the school. Then the parent decides who owns each task. A calendar entry, a payment, a message to the school, or a signed form each requires a separate approval. The system does not infer agreement just because two adults share a household.
The outcome record captures total preparation and verification time, which fields were corrected, the final deadline, whether the event was missed, and who completed the work. And here's the test for the mental load question: if the other parent performs an assigned task, that counts as redistribution only if the action and the responsibility actually moved. A shared list on a screen is not redistribution. The spouse picking up the kid, paying the fee, and remembering next time is.
The same sequence works for a grocery list or a meal plan, but the parent stays responsible for allergies, suitability, actual prices, and what ingredients are on hand. No reviewed study establishes that an unrestricted generated meal plan lowers food spending or meets nutritional needs. For a medical concern, the workflow stops short of diagnosis: it organizes questions and any existing professional instructions, but it can't be the authority on urgency or treatment. For a benefits form, it keeps the official requirements in view and drafts entries, while the applicant and the authorized agency control submission, eligibility, and appeal.
6.10.2 What a Real Household Pilot Should Measure#
Measuring the whole task means recruiting varied family structures and comparing against a credible existing method. Not an artificial "do nothing" group. Families already use calendars, search, school portals, relatives, and professionals, and the new tool has to beat that.
Primary outcomes: net time after verification, missed obligations, substantive error rate, successful completion, caregiver burden, dependent safety, and human control.
Distribution: results by household structure, income, language, disability, digital access, caregiving intensity, and prior skill.
Relationship outcomes: who owns the task before and after, partner visibility, conflict, and whether responsibility concentrates on one person.
Proposed net time saved equals the baseline time for the same-quality task minus all the time spent in the assisted workflow: setup allocation, data preparation, prompting, verification, correction, execution, and follow-up.
That's the whole game. Faster drafting is not a net saving if supervision and correction take longer than the old way. A pilot should preregister one or two primary outcomes per task category, and it should not search across dozens of measures and publish only the favorable ones.
6.11 Evidence Gaps and Permitted Claims About AI for Families#
The household-specific evidence I located does not establish durable reductions in total administrative time, grocery spending, relationship conflict, or unequal responsibility from general-purpose assistants.
Effects for single parents, low-connectivity households, non-English users, people with disabilities, and families in acute crisis can't be inferred from digitally comfortable survey respondents. The positive interventions in this chapter usually include an institution, predefined content, expert review, or human assistance. Research should test whether those components are necessary before anyone presents a cheaper automated substitute as equivalent. Housing search, insurance disputes, debt advice, legal forms, food safety, and emergency caregiving all need domain-specific evidence before strong efficacy claims are made.
Structured digital and human programs have improved certain parenting, application, energy, and caregiver outcomes in defined settings
Permitted
General-purpose assistance has been proven to save every family time, redistribute mental load, or improve relationships
Not permitted
Families can use original-source infrastructure records to investigate possible effects on bills, water, services, jobs, and public finances
Permitted
A tracker entry, announced investment, model capability, or educational benefit proves that a specific facility improves household welfare
Not permitted
A household pilot can test whether source-linked assistance improves comprehension and reduces burden
Permitted
Usage, confidence, a generated answer, or a product demonstration substitutes for measured household outcomes
Not permitted
On the infrastructure rows: the seven Savrn trackers exist so families and communities can find the original records behind a facility. A tracker entry starts a question. It never ends one.
6.12 Conclusion: Bounded Assistance, Not Delegation#
The evidence supports bounded assistance, not household delegation as a default. The most credible benefits in this chapter came from interventions built around a particular barrier, an accountable institution, a comparison condition, and a measurable outcome. You've now seen that pattern in schooling, in workforce programs, and here again.
So the household standard is practical: reduce verified work without hiding who remains responsible. Families should be able to identify the source, understand the action, withhold approval, correct an error, protect a dependent, and stop the system. Use AI for the permission slip. Check its work. Keep the judgment, the relationships, and the off switch in your own hands.
Not in any controlled study yet. A 2026 Lurie Children's survey found 81% of 1,004 U.S. parents used AI for parenting tasks and reported saving 58 minutes a week, but that is self-report without a disclosed probability sample. A Croatian pilot with five mothers measured no objective time savings. Real savings must count the time spent checking and correcting the AI's work, which no reviewed study has measured.
Can AI reduce the mental load for moms?
The evidence so far says no redistribution has been shown. A study of 369 employed Croatian mothers and a five-mother ChatGPT pilot found participants valued help with planning and scheduling, but responsibility largely stayed with mothers. Helping one person work the list faster is not the same as moving the list to someone else. That requires a partner who actually takes on the task.
Is it safe to use ChatGPT for medical advice about my child?
Use it to organize symptoms and questions, not to decide urgency. In a randomized study of 1,298 U.K. adults using ten clinical scenarios, people without a chatbot had 1.76 times the odds of identifying a relevant condition compared with chatbot users, and every group tended to underestimate urgency. The scenarios were simulated and the models have since changed, but a professional or emergency service should make the call.
Do AI chatbots make people less lonely?
A large trial suggests not. In a July 2026 working paper covering 12,356 French adults, people encouraged to have daily personal conversations with a generative AI for 28 days reported loneliness 0.168 points higher on a 0 to 10 scale (adjusted p = .0002). The effect was small, the study lasted one month, and it had no journaling comparison, but the direction was the opposite of the usual promise.
Are AI parenting apps backed by research?
Structured programs are, open chat apps are not yet. A blended program in China combining rule-based WeChat modules with teacher and social worker groups improved caregiver-reported learning activities (beta 1.79) and reduced reported violence (IRR 0.87) at the immediate follow-up, in one preschool of 541 caregivers. The READY4K text program raised literacy by 0.109 standard deviations. Neither tested an open-ended generative chatbot.
Does technology help dementia caregivers?
On average, a little. A 2026 meta-analysis of 35 trials with 3,388 caregivers found interactive eHealth reduced burden by SMD -0.26. But the prediction interval ran from -1.10 to +0.58, so a specific new program could help a lot, do nothing, or add burden. Only three studies were low risk of bias, and the review did not isolate general-purpose generative AI.
Can AI help my family apply for benefits like SNAP?
It can help you prepare, but the strongest evidence favors human help. In a randomized trial of 31,888 Pennsylvania households with an older adult, enrollment was 5.8% with no outreach, 10.5% with information, and 17.6% with information plus staff who filled out and submitted applications. The trial involved no AI. Let a tool organize documents, and let the agency decide eligibility.
What should I never let an AI do for my family?
Never let it decide whether something is a medical emergency, make a child safety or discipline decision, act as a child's or lonely relative's main companion, assert legal eligibility, or spend money and share data without your specific approval. Those limits come from randomized trials with adverse findings on medical decisions and loneliness, and from the principle that authority stays with the responsible adult.
Chapter 7 · Civic Life and Local Government
AI in Local Government: Public Comments, Chatbots, and Civic Decisions
I've stood at the microphone in enough public hearings to know how a local decision really gets made. A project can live or die in the space of a few minutes, and what decides it usually isn't the engineering. It's what the people in the folding chairs understand, or think they understand, about what's being proposed, what it costs, and who carries the risk. When that understanding is shaky, the loudest version of the story wins. So when I hear people pitch AI in local government, the question I ask is the one I think most residents are quietly asking: can AI make local decisions more understandable without making them for us?
That question matters because the stakes are local and concrete. A park gets built or it doesn't. A drainage project gets funded or the next storm floods the same street. A city chatbot tells a small business owner the right rule or the wrong one. These aren't abstract policy debates. They're the decisions that shape whether people trust their own government.
Here's what the evidence in this chapter shows, including the uncomfortable part. The research is real and some of it is encouraging. People in controlled experiments preferred group statements drafted by a machine. A government-commissioned tool found most of the themes human analysts found in thousands of public comments. But that same tool, in its blinded test, also produced a large share of themes that reviewers didn't accept, and it missed about a quarter of the real ones. And the most public test of a city chatbot ended in an official audit that the city disputed, with the two sides counting success in different ways. None of this supports letting a model decide. All of it supports a narrower, more useful goal: a public record that more people can read, check, and argue about on equal terms.
The practical goal of better civic research is modest, and that's exactly why it's achievable. Make budgets, proposals, rules, and alternatives understandable without requiring every resident to become a spreadsheet specialist. That goal does not mean a generated answer should decide which park to build, which drainage plan to fund, or which development to approve. That distinction is the whole chapter.
Most of the confusion about AI in local government comes from collapsing two very different jobs into one. The first job is explanation: taking a 300-page budget or a stack of engineering documents and helping a resident find the part that answers their question. The second job is judgment: deciding which competing goods a community should prioritize. The first job is about information. The second is about values and legal authority. A tool can help with the first. The evidence I reviewed gives no support for handing it the second.
7.1.1 Three Forms of Civic Evidence That Must Not Be Merged#
The evidence for this domain comes in three distinct forms, and they must not be blended into one comforting claim that automated civic decisions work.
First, there is a large experimental program on group statements, published in Science, that tests assisted deliberation under controlled conditions (Science deliberation study). Second, there is a government-commissioned evaluation, published by the United Kingdom Department for Transport with the Alan Turing Institute in December 2025, that tests a consultation-analysis tool in a blinded coding exercise (Department for Transport evaluation). Third, there is an official audit, issued by the New York City comptroller on December 30, 2025, that identifies shortcomings in a public-service system and records the agency's dispute of those findings (New York City audit).
Each form answers a different question:
Evidence form
The question it asks
What it can support
What it cannot support
Controlled deliberation experiments
Do people prefer assisted group statements under study conditions?
Further work on assisted synthesis
Claims that a generated consensus is correct, legitimate, or durable
Blinded consultation evaluation
How accurately does a tool reproduce human-coded themes when graders cannot see the source?
Assisted coding inside an expert-reviewed process
Claims that a summary is complete, or that savings were realized
Adversarial audit
Was a deployed system managed and measured well enough to defend?
Demands for published denominators and evaluation plans
Claims that all public-service assistants fail, or that the dispute is resolved
Figure 7.1Comparison
Three forms of civic evidence answer three different questions
Each form supports a different claim. None supports delegating a community's decision to a model.
A reader who wants all three questions answered before trusting a civic tool is not being timid. The evidence itself is organized that way. A preference result doesn't tell you about accuracy. An accuracy test doesn't tell you whether a real deployment was managed well. And an audit of one deployment doesn't tell you what a well-designed system could do. You need each piece for what it is.
7.1.2 The Thesis: Legible Records, Accountable Decisions#
Here's the position this chapter defends, conditional on the record above. Assistance should widen access to a common factual record and make the official process more inspectable. The decision itself, meaning the weighing of competing priorities that residents reasonably hold, stays with the authorized public body.
A tool that makes politics disappear is not a promise this evidence supports. A tool that makes the record legible is. I'll come back to that line at the end, because every study below points toward it from a different direction.
7.2 The Habermas Machine Study: What 5,734 Participants Showed#
The most prominent research program in assisted civic deliberation is the Habermas Machine work published in Science. It involved 5,734 United Kingdom participants across a series of experiments (Science study).
That number needs its qualifier immediately. This was not a single field trial of 5,734 residents making a binding municipal decision. It was a structured program of comparative experiments, and the findings describe what happened inside them. When you see this study cited as proof that "AI can find common ground for a whole country," the person citing it has dropped the qualifier.
The system generated and revised group statements from participants' own views and critiques. In practice, that meant iterative rounds. A model synthesized positions, participants reacted, and the synthesis was revised in light of those reactions. The comparison group was not "nothing." It was statements produced by human mediators working on the same task.
In the studied comparisons, participants preferred the machine's statements to those produced by human mediators. That's a real finding about stated preference under the experiment's conditions, and I don't want to wave it away. Drafting a statement that a disagreeing group can live with is hard work. A tool that helps people see a shared draft faster has value.
It is also a narrower finding than the headline sometimes suggests. Preference for a statement is not evidence that the statement is factually correct. It isn't evidence that the statement is legally legitimate. It isn't evidence that the agreement lasts once a policy is implemented. And it isn't evidence that the statement benefits people who weren't in the room.
Think about how a city council actually works. A resolution that everyone in the chamber likes can still rest on a wrong cost estimate. It can still exceed the council's legal authority. It can still collapse the first time the bill comes due. And it can still ignore the renters, the future residents, and the neighbors downstream who never showed up. Preference inside a study tells you nothing about any of those.
7.2.3 The Virtual Citizens' Assembly and What Convergence Means#
The research program also examined a virtual citizens' assembly and reported convergence in expressed positions across rounds. The authors are direct about what this supports: further investigation of assisted synthesis.
Here's the part most people skip. Convergence inside a moderated exchange is not consent to a policy. The study does not establish that a generated consensus reflects what the community would choose after costs land, after implementation stalls, or after the minority that lost the synthesis has to live with the result. People can move toward each other's language in a structured session and still disagree sharply once the tax bill or the construction noise arrives.
7.2.4 Common Misreadings of the Habermas Machine Study#
Misreading
Why it's wrong
"AI mediators are better than human mediators."
Participants preferred the machine's statements in the studied comparisons. Preference is not a measure of mediation quality in real disputes.
"5,734 people reached consensus with AI."
The 5,734 were spread across a series of experiments, not one assembly making one decision.
"Convergence means the community agreed."
Convergence in expressed positions across rounds is not consent to a policy, and the authors frame it as grounds for more research.
"We can use this to settle local disputes."
Nothing in the study tested binding local decisions, implementation, or effects on people outside the experiment.
7.2.5 What a Defensible Civic Application Looks Like#
This is where the chapter's proposed civic application takes shape. The defensible product of assisted deliberation is an inspectable summary of agreement and disagreement. It keeps minority objections and the underlying submissions visible, rather than presenting one polished paragraph as the community's final voice. A resident who disagrees with the majority conclusion should be able to find their concern, stated in their terms, with a path back to the original submissions.
Consensus and accuracy also have to stay separate even when everyone is working from the same record. Residents can reasonably accept identical cost, flood-risk, or service data and still disagree, because they attach different weights to recreation, taxes, access, safety, and distribution. A synthesis that papers over those weight differences does not resolve them. It hides them.
The Habermas Machine results justify building better tools for seeing one another's positions. They do not justify letting the tool decide which positions count.
What this means for residents: if a city or civic group presents a machine-drafted "common ground" statement, ask to see the submissions it was built from and the objections it did not absorb. What this means for elected officials: a statement people prefer is a starting draft for deliberation, not a substitute for your vote or your legal duty to consider the record.
7.3 AI Public Consultation Analysis: The Department for Transport Test#
The strongest domain-specific evaluation in this chapter is the United Kingdom Department for Transport's December 2025 technical evaluation of a consultation-analysis tool, conducted with the Alan Turing Institute (Technical evaluation).
This is the kind of study that matters most to people who run public comment processes. Any city, county, or agency that opens a consultation can receive far more written responses than staff can read carefully by the deadline. The promise of AI public consultation analysis is that a tool reads everything, groups it into themes, and hands staff a map. The question is how good that map is.
The centerpiece was a blinded theme-generation test: 11 consultation questions, approximately 9,100 responses, and 165 human-validated reference themes. "Blinded" matters here. It means the graders judging whether a theme was right could not see which output came from where. That removes one of the easiest ways for an evaluation to flatter a tool, which is a reviewer who knows the machine wrote something and grades it more generously or more harshly because of that.
Two tasks were tested, and they are different. Theme generation asks the tool to produce the list of themes present in the responses. Response-to-theme mapping asks the tool to sort individual responses under themes. Generating the list is not the same as sorting individual responses under it, and this chapter won't substitute one number for the other.
7.3.2 The Results: Recall, Precision, and F1 in Plain English#
The blinded results were mixed, and the mixture is informative (Evaluation findings). Theme generation achieved approximately 0.75 recall, 0.50 precision, and 0.59 F1. The separate blinded response-to-theme mapping task achieved approximately 0.75 F1.
If those terms are new to you, here's what they mean for a committee reading a summary:
Recall answers: of the themes the human reference found, how many did the tool also find? A recall of 0.75 means the tool found most of the human-validated themes but left about a quarter of them out.
Precision answers: of the themes the tool produced, how many were judged relevant? A precision of 0.50 means that, under the evaluation's matching procedure, roughly half of the generated themes were judged relevant, and roughly half were not.
F1 combines the two into one score. It's useful for comparison, but it hides which kind of error is driving the result. For a public process, you want to know both.
Figure 7.2Data
Blinded consultation-analysis results
Two different tasks; do not substitute one number for the other. The stronger results in the report came from nonblinded live use and are not shown.
Put those together and the failure mode is subtle. A confident, readable summary can simultaneously add themes nobody argued for and drop themes that real constituents submitted. The output looks more settled than the underlying record is.
7.3.3 A Worked Example: 100 Themes, Translated Into Arithmetic#
Here's a hypothetical to make those rates concrete. This is an arithmetic illustration only. It uses the stated blinded rates (about 0.75 recall and about 0.50 precision) and assumes, for simplicity, that each matched generated theme corresponds to exactly one reference theme. The real evaluation's matching procedure is described in the report; this example does not reproduce it, and no real consultation had these counts.
Suppose a human team, reading every comment carefully, validated 100 distinct themes in a hypothetical consultation.
Recall of 0.75. The tool finds 75 of those 100 themes. It misses 25. Those 25 themes were submitted by real people and are absent from the summary.
Precision of 0.50. Half of what the tool produces is judged relevant. If the 75 matched themes are the relevant half, the tool produced about 150 themes in total. That means about 75 generated themes did not match anything reviewers accepted.
What the committee sees. A summary listing roughly 150 themes, which looks thorough. Inside it, about 75 are well supported, about 75 are not, and 25 real themes are nowhere to be found.
The check. With precision of 0.50 and recall of 0.75, the F1 arithmetic comes out near 0.60, consistent with the approximately 0.59 reported.
Hypothetical count (illustration only)
Number
Human-validated reference themes
100
Reference themes the tool found (recall about 0.75)
75
Reference themes the tool missed
25
Total themes the tool produced (precision about 0.50)
about 150
Generated themes reviewers did not accept
about 75
Read that again. A longer list is not a better list. In this illustration, the summary is half again as long as the human reference, and it is still missing a quarter of what residents actually said. A committee member skimming it would come away thinking the process was more complete than it was.
Now ask which 25 themes are most likely to be the missing ones. The evaluation doesn't give that breakdown, so I won't pretend it does. But common sense about how summarization works says the risk is highest for concerns that are rare, oddly phrased, or raised by only a few people. Those are often exactly the concerns a public process exists to hear: the one family whose driveway floods, the business owner whose access road gets cut, the neighborhood that's always outnumbered.
That failure mode has a specific remedy, and the evaluation itself points toward it. The appropriate inference from the blinded results is conditional: assisted coding can be useful inside an expert-reviewed process, but omission and invention both need explicit testing.
In practice, a committee should inspect a sample of original comments rather than judging the output by readability alone. The sample should deliberately include dissenting and low-frequency concerns, which are exactly the themes a 0.75-recall tool is most likely to drop. And someone should check a sample of the tool's themes against the comments they claim to summarize, because a 0.50-precision tool will produce themes that don't hold up.
7.3.5 Blinded Versus Nonblinded: A Reporting Rule#
Two reporting disciplines attach to this evaluation, and both matter for anyone citing it.
First, the report's stronger results came from nonblinded live use. Those results should never be presented as though they came from the blinded test. The blinded numbers are the approximately 0.75 recall, 0.50 precision, and 0.59 F1 for theme generation, and the approximately 0.75 F1 for mapping. If a vendor slide or a staff memo quotes a higher number from this report, ask which part of the report it came from.
Second, the report's time and cost savings are estimates against an assumed manual process. That's a modeled comparison, not a randomized measurement of realized savings across consultations. An estimate against an assumed baseline is a planning input, not a demonstrated saving. It can help a budget office decide whether a pilot is worth running. It can't tell a council that the money was saved.
7.4 How to Read an AI Summary of Public Comments: A Resident's Checklist#
If your city, county, school board, or transportation agency starts using AI to summarize public comments, you don't need to be a data scientist to hold it accountable. You need a few questions and the willingness to ask them out loud. Everything on this list follows from the evidence above; none of it requires new facts about any particular tool.
Before you trust the summary, find out:
Was AI used, and for which step? Theme generation and sorting individual comments are different tasks with different error rates, as the Department for Transport evaluation showed. Ask which one the tool did.
How many comments were received, and how many were analyzed? Show me the denominator. A theme described as "common" means nothing unless you know common out of what.
Can I read the original submissions? The summary is a derivative. The comments are the record. If the originals aren't available, the summary can't be checked.
Did a person review a sample of the originals against the summary? Ask who, how many, and whether the sample included minority and low-frequency concerns.
Is my concern in there, in my terms? Search for your own comment. If you can't find your point, or it has been merged into a broader theme that changes its meaning, say so on the record.
Are there themes in the summary that you don't recognize from the hearing or the comment file? A 0.50-precision tool in a blinded test produced many themes reviewers didn't accept. Ask staff to point to the submissions behind any theme that seems to come from nowhere.
Are counts presented with their denominators? "Many residents raised traffic" is not a count. "A stated number of the total comments raised traffic" is.
Is there a correction route? If you find an error, who do you tell, and will the correction be visible?
Was there a way to comment without having your submission processed by an AI system? A fair process offers a non-model route.
Who makes the decision, and when? The summary is an input. Make sure you know which public body decides and when you can speak to it.
Red flags that should make you slow down:
A summary that reads smoothly but gives no counts, no denominators, and no links to originals.
A claim of time or cost savings with no note on whether it was modeled or measured.
A single "community consensus" paragraph with no section on disagreement.
Themes attributed to "residents" that nobody at the hearing remembers hearing.
Any statement that the tool's recommendation is the preferred option.
None of these red flags proves the summary is wrong. They tell you it hasn't yet earned your trust. That's a reasonable position for a resident to take, and it's one staff should welcome, because the checks that catch errors also protect the people who publish the summary.
7.5 For City Staff: A Minimum Standard Before Publishing an AI-Generated Summary#
If you work inside a city, county, or agency, you're the person who has to make this work under a deadline. Here's a minimum standard I'd want in place before any AI-generated summary of public input goes to a council, commission, or the public. It's built directly from the three evidence streams in this chapter: the omission and invention risk shown in the blinded consultation test, the denominator disputes in the MyCity audit, and the gap between preference and accuracy in the deliberation experiments.
Separate personal information from public comments. Distinguish public comments from personal identifying information before either touches a model. Do not upload private constituent records into unapproved systems.
Decide which task the tool is doing. Theme generation, response mapping, and drafting prose summaries are different jobs. Name the job in your methods note.
Keep a non-model route. Residents who don't want their submission processed by a system should have another way to participate.
Publish the denominator. State the total number of submissions received, the number analyzed, and any exclusions with reasons.
Run an omission check. Have a qualified person read a sample of original comments, deliberately including dissenting and low-frequency concerns, and confirm each appears in the summary in a form its author would recognize.
Run an invention check. For a sample of generated themes, trace each back to specific submissions. Remove or flag any theme that can't be traced.
Preserve disagreement. Include a section on minority objections and unresolved disagreement. Do not collapse the record into one consensus paragraph.
Link to the record. Make the original submissions available alongside the summary, subject to privacy law.
Label estimates as estimates. If you cite time or cost savings, say whether they were measured locally or modeled against an assumed baseline.
State the decision authority. Name the body that decides, the date, and how residents can still participate.
Keep a visible correction log. When a summary is found wrong, append the correction visibly rather than silently replacing the mistaken version.
Track correction time. Record how much staff time went into reviewing and correcting the tool's output. That's the real measure of what the tool saved, and it will be far more useful to your budget office than a projection.
Collect satisfaction, but don't stop there. Satisfaction should never substitute for comprehension or factual accuracy.
Minimum standard item
Evidence it responds to
Omission check with minority concerns in the sample
Blinded recall of about 0.75 in theme generation
Invention check tracing themes to submissions
Blinded precision of about 0.50 in theme generation
Published denominators and exclusions
Denominator dispute in the MyCity audit
Estimates labeled as modeled
Modeled savings in the Department for Transport report
Disagreement section, not just consensus
Preference versus accuracy in the Habermas Machine experiments
Visible correction log
Accountability principle in section 7.9
This standard won't make a tool accurate. It makes the tool's errors findable. For a public body, that's the difference that matters.
7.6 The NYC MyCity Chatbot Audit and the Disputed Denominator#
The third evidence stream is adversarial by design. On December 30, 2025, the New York City comptroller released an audit of the Office of Technology and Innovation's MyCity system (Official audit). The audit assessed a broader digital-service program, not a chatbot in isolation. It identified shortcomings including inconsistent chatbot responses and weaknesses in management and evaluation.
7.6.1 Why Program Spending Is Not Chatbot Spending#
The audit's spending figures illustrate a denominator discipline this whole report keeps returning to. The program-level spending cited in the audit is not a chatbot-only cost, and it should not be represented as one.
Think about how any public digital service is built. A system that includes case management, portal infrastructure, and human staff has a total program cost. A chatbot inside that system has a share of the cost, and the audit record does not isolate that share in any way this chapter can defend. Any chart in this report that attached the full program figure to the chatbot alone would be manufacturing a fact.
Figure 7.5Concept
Denominator discipline in the MyCity audit
Schematic, not to scale; no dollar values are plotted. The comptroller and the agency used different denominators, and both appear in the published report.
You'll see this mistake in headlines and on social media. "The city spent this much on a chatbot that gave wrong answers." If the figure is the program total, that sentence is false even when every individual word in it sounds plausible. A number without a denominator is a rumor.
7.6.2 The City's Dispute and the Competing Denominators#
The city did not accept the audit. The Office of Technology and Innovation disputed the findings and recommendations, including the auditors' interpretation of accuracy and the use of voluntary user feedback (Audit and agency response). The published report contains both the agency response and the auditors' rejoinder.
The two sides use different denominators and assessment choices. They differ on which questions were tested, how refusals and non-answers were treated, and what counts as a correct answer. Each of those choices can move an accuracy figure substantially, and none of them is purely technical. Deciding that a refusal counts as "not wrong" is a policy judgment. So is deciding that a partially correct answer counts as correct.
7.6.3 Two Measurement Principles Every Government Chatbot Needs#
Two principles fall out of that dispute, and they apply far beyond New York.
Voluntary feedback is not a representative satisfaction survey. The residents who bother to rate a chatbot are not a random sample of the residents who needed it. People who got a quick answer may click a thumbs-up. People who gave up in frustration may never rate anything. People who got a wrong answer and didn't know it was wrong may rate it highly.
An accuracy percentage is uninterpretable standing alone. It requires the questions tested, the exclusions, the grading rules, the treatment of refusals, and the definition of a correct answer. Here's a hypothetical to make that concrete. An 85 percent accuracy figure with unknown question sampling tells a policymaker almost nothing. An 85 percent accuracy figure with the test set, grading protocol, and refusal policy published tells a policymaker something worth acting on. Same number, completely different value.
Who decided what was correct, and against what source
The treatment of refusals and non-answers
Counting a refusal as correct or incorrect can swing the figure
The definition of a correct answer
Partial answers and outdated answers need a rule
7.6.4 What the MyCity Case Does and Does Not Establish#
This case should be stated precisely. It does not establish that all public-service assistants fail, and this chapter won't claim that. It establishes that a deployed system was audited, that real disagreements existed about how success was counted, and that neither side's number can be responsibly quoted without its denominator.
Distinguishing disagreement over evidence from a verified resolution is not a technicality. In public administration, it's the difference between accountability and argument by press release. If you're a resident, a reporter, or a council member, the defensible sentence is: an official audit found shortcomings, the agency disputed them, and the two sides counted differently. Anything stronger in either direction goes beyond the record.
7.7 Participatory Budgeting: Participation Is an Institution, Not an Interface#
A scoping review of participatory budgeting, the practice of giving residents direct authority over a defined slice of public money, located 37 studies reported across 39 articles (Campbell and colleagues). The body of evidence is concentrated in Brazil and dominated by nonrandomized designs. It found suggestive benefits in some settings, and it did not provide a general causal guarantee of improved health or public services.
7.7.1 Why a Review Without AI Belongs in an AI Chapter#
That review predates generative models and did not test any. It belongs here for a structural reason. A better information interface does not remove the need for legitimate participation rules, actual budget authority, inclusion, and implementation.
Participatory budgeting works, where it works, because an institution ceded real money and real rules to a defined public. It doesn't work because residents gained a better window into a budget they still don't control. An assistant that explains the process but leaves the authority untouched has improved the interface of a process whose value was never in the interface. Capability is not benefit.
Notice also what the review's limits teach. Even for an established civic practice that predates generative AI, the evidence is mostly nonrandomized and geographically concentrated, and the benefits are suggestive rather than guaranteed. If that's the actual state of the evidence for a well-established civic institution, anyone claiming that a new AI citizen engagement tool will reliably improve outcomes is claiming far more than the research base allows.
The standard that follows is direct: assistance should make the official process more accessible and more inspectable. It should not create an unofficial parallel process in which only technically confident participants can influence what gets summarized.
Here's why that risk is real. If a generated summary is the version of public input that staff actually read, then whoever controls the summary controls the input. The residents without the time, language fluency, or confidence to check it have been quietly disenfranchised by a convenience. Nobody voted to shut them out. The workflow did it.
What this means for board members and commissioners: before approving an AI citizen engagement tool, ask whether it changes who has authority, or only how information is displayed. If only the display changes, judge it as a display tool. Don't credit it with the benefits of real participation.
7.8 A Civic Decision Workflow for AI in Local Government#
What follows is a proposed operating process, not a field-tested municipal intervention. It's designed to support a park, drainage, service, or facility decision without inventing a local case or assuming the preferred answer. Each stage names a required work product and the verification or authority that attaches to it.
Stage
Required work product
Verification and authority
1. Define the decision
Exact question, responsible body, deadline, legal scope, affected population
Confirm against official notice and governing documents
2. Establish the baseline
Existing budget, service condition, asset condition, and relevant risks
Reconcile periods, definitions, and responsible agencies
3. Build alternatives
Status quo, minimum-change option, and substantive alternatives
Do not compare only the preferred option with an unrealistic alternative
4. Normalize costs
Initial, recurring, maintenance, financing, contingency, and end-of-life costs
Qualified staff validate arithmetic and assumptions
5. Analyze distribution
Who receives benefits, pays, loses access, or bears risk
Include renters, nonusers, adjacent residents, and future users where relevant
6. Summarize participation
Themes, counts with clear denominators, original submissions, minority concerns
Actual costs, service outcomes, distribution, and complaints
Compare with the adopted commitments, not just the announcement
Figure 7.3Framework
The eight-stage civic decision workflow
Proposed, not field tested. A model may extract, explain, organize, and draft at any stage. It may not invent engineering estimates, certify legal authority, or complete an incomplete dataset.
7.8.1 Where the Model Helps and Where It Must Stop#
Model assistance has a real but bounded role in this workflow. It may extract line items, explain terminology, organize questions, or draft alternative summaries. It should not invent missing engineering estimates, certify legal authority, or convert an incomplete dataset into a finding that an option is optimal.
The boundary is consistent with everything earlier in the chapter: synthesis and navigation, yes; fabrication of the record, no.
Stage by stage, that boundary looks like this:
Stage 1. A model can help draft a plain-language version of the decision question. It cannot confirm the body's legal scope. That comes from the official notice and governing documents.
Stage 2. A model can pull line items out of a budget document. A person reconciles periods and definitions.
Stage 3. A model can draft descriptions of alternatives staff have defined. It should not invent an alternative's cost or feasibility.
Stage 4. A model can lay out cost categories. Qualified staff validate the arithmetic and assumptions.
Stage 5. A model can organize who is affected, based on data staff supply. It cannot fill in missing data about who lacks service.
Stage 6. A model can help group comments. The omission and invention checks from section 7.5 apply in full.
Stage 7. The model has no role in the choice. The authorized public body decides and records its reasons.
Stage 8. A model can help compare actual results with adopted commitments. The data must come from the delivery record, not from the announcement.
Here's a hypothetical to show how the stages fit together. Imagine a city deciding whether to renovate an existing recreation facility, build a new one elsewhere, or do minimal repairs. I'm not describing any real city, and there are no real numbers here.
At stage 1, staff write the exact question and name the body that decides, confirmed against the official notice. At stage 2, they document the current facility's condition and budget. At stage 3, they build three options: status quo with minimal repairs, renovation, and new construction. That rule about avoiding unrealistic alternatives matters here. If the only comparison is between a well-developed new facility plan and a deliberately weak "do nothing" option, the process has been rigged before anyone speaks.
At stage 4, every option gets the same cost categories, including maintenance and end-of-life costs, not just construction. At stage 5, staff ask who gains, who pays, and who loses access, including people who don't use the facility today. At stage 6, public comments are summarized with denominators, originals are linked, and a human sample review checks for dropped and invented themes. At stage 7, the council chooses and writes down why, including what it still doesn't know. At stage 8, a year later, someone compares what was delivered with what was promised.
At no point does the model rank the options or recommend one. At several points, it saves staff time on reading and organizing. That's the right division of labor.
Apply the workflow to a neighborhood park decision. The proposed evidence packet would identify construction and maintenance costs, accessibility, travel distance, utilization assumptions, drainage implications, and who currently lacks service. Each field carries its source and uncertainty. Values that can't be obtained are marked unavailable, not filled.
That last rule is the one most likely to be broken, and it's the one that matters most. A language model asked to complete a table will tend to complete it. A blank cell marked "unavailable" tells the truth. A plausible-looking number that nobody measured is a liability.
The packet's most important property is what it refuses to do. A generated ranking should not be the final product. The useful product is a comparison in which residents can change the explicit weights, such as more weight on access for carless households or less on programmed field hours, and see whether the conclusion depends on an unverified assumption. A ranking that survives weight changes is robust. One that flips the moment a resident adjusts a slider has told the community something valuable about how little the analysis supports it.
Here's a hypothetical illustration of that sensitivity test, with no real values. Two park sites are compared. When travel distance for residents without cars is weighted heavily, one site comes out ahead. When programmed field hours are weighted heavily, the other does. If the utilization assumption behind field hours turns out to be unverified, the community has learned that the case for the second site rests on a number nobody checked. That's exactly the kind of finding a public hearing should surface before a vote, not after.
A drainage decision raises the stakes, because the underlying record is engineering. The packet would identify the engineering study, the design storm, the model assumptions, downstream effects, maintenance responsibilities, and the consequences if performance falls short.
A language model may help residents work through those records by translating terminology and mapping which document answers which question. It should not replace the relevant engineering assessment, and no summary should outrun the study it summarizes. If the engineering study didn't model a particular downstream effect, the summary shouldn't imply that it did. If the study's assumptions are uncertain, the summary should say so in the same place it states the conclusion.
This boundary protects both sides. The community is protected from a confident explanation of a model that was never run. The decision-maker is protected from the liability that attaches the moment an unofficial summary is treated as the technical record. An accessible explanation is valuable only if it preserves the limits of the underlying technical work.
Figure 7.4Framework
Two evidence packets: a park and a drainage project
Illustrative packet specifications. A ranking that flips when a resident changes a weight has revealed how little the analysis supports it.
Packet feature
Parks comparison
Drainage comparison
Required fields
Construction and maintenance costs, accessibility, travel distance, utilization assumptions, drainage implications, who lacks service
Engineering study, design storm, model assumptions, downstream effects, maintenance responsibilities, consequences of underperformance
Source and uncertainty
Tagged on every field
Tagged on every field
Missing values
Marked unavailable, not filled
Marked unavailable, not filled
Key test
Weight sensitivity: does the conclusion flip when residents adjust priorities?
Engineering boundary: does any summary claim more than the study supports?
Model's role
Explain, organize, let residents change weights
Translate terms, map documents to questions
Model must not
Produce the final ranking
Replace the engineering assessment
What this means for residents at a drainage hearing: ask which engineering study the plan relies on, what design storm it assumed, and what happens if the system performs below that design. If a plain-language summary answers those questions, check that its answers match the study itself.
7.9 Evaluating AI Citizen Engagement: Outcomes and Safeguards#
The proposed primary outcome for evaluating civic assistance is not the number of pages summarized, the response time, or the satisfaction score. It's the proportion of residents who can correctly identify the alternatives, the major costs, the unresolved assumptions, the decision authority, and the opportunity to participate.
Comprehension is the outcome, because a process that produces fluent summaries and confused residents has failed at the only thing a democracy requires of its public. That's the whole game. If a city adopts an AI tool and residents understand their options no better than before, the tool hasn't delivered a civic benefit, whatever else it has done.
Omission rates against human-validated themes. The 0.50 precision and 0.75 recall from the Department for Transport evaluation are the right template: measure both what was dropped and what was added.
Unsupported statements in summaries.
Staff correction time.
Accessibility for residents with disabilities and limited language fluency.
Participation breadth. Who took part, not just how many.
Survival of minority views in the final record.
Satisfaction should be collected but never substituted for comprehension or factual accuracy. A resident can be satisfied with a summary that dropped their neighbor's objection.
The safeguards are structural rather than technological:
A non-model route for residents who don't want their submission processed by a system.
Privacy separation. Avoid uploading private constituent records into unapproved systems, and distinguish public comments from personal identifying information before either touches a model.
An explicit correction log. When a summary is found wrong, the correction is appended visibly rather than silently replacing the mistaken version. Silent correction is how a summary acquires an authority it never earned.
What this means for advisers and consultants who sell or implement these tools: propose comprehension as the success metric in your contract. If your tool can't show that residents understood more, you haven't shown what the public is paying for.
The seven Savrn trackers can help locate capital, delay, policy, water, grid, permit, and scarcity records relevant to a civic question: a data center's proposed water use, a moratorium under consideration, a permit record that anchors who committed to what. They are the Capital Atlas, the Delay Watchlist, the Moratorium Tracker, the Water Tracker, the Grid Watchlist, the Permits Tracker, and the Scarcity Tracker.
Their publisher-described scopes make them a starting index, not a complete municipal evidence file, and a tracker entry does not establish the effects of a local decision. The companion cross-audience map in the research package specifies the records that must be added before a civic conclusion is defensible: local budgets, engineering studies, service baselines, and the official notice itself.
A tracker entry starts a question. It never ends one. An infrastructure record can inform a parks-and-drainage packet, but it cannot substitute for the park's own budget or the drainage study's own assumptions. Chapter 11 develops this boundary in full, and Chapter 10 applies it to data center projects in particular.
7.11 Conclusion: Make the Record Legible, Keep the Decision Public#
The evidence supports cautious use of assistance for deliberation and document analysis. It comes with material limitations in theme detection and unresolved disagreement in public-service evaluation. It does not support delegating a community's priorities or statutory decisions to a model.
So back to the question I started with: can AI make local decisions more understandable without making them for us? The evidence says it can help with the first part, under supervision, with checks that catch what it drops and what it invents. It says nothing that would justify the second part.
The goal is wider access to a common factual record, with visible disagreement and accountable decisions. That's a more defensible public benefit than promising a tool will make politics disappear. Politics, meaning the weighing of competing goods by a community that has to live with the result, is not a defect in the process. It is the process.
The evidence in this report covers three uses: assisted deliberation, where a model drafts group statements; consultation analysis, where a tool groups public comments into themes; and public-service chatbots. Each was tested differently, through controlled experiments, a blinded evaluation, and an official audit. None of the evidence supports letting a model make the decision itself, which should stay with the authorized public body.
Can AI accurately summarize public comments?
Partly. In a blinded UK Department for Transport evaluation covering 11 questions, about 9,100 responses, and 165 reference themes, a consultation tool reached about 0.75 recall and 0.50 precision in theme generation. It found most real themes but missed about a quarter, and roughly half its generated themes were not judged relevant. Human review of original comments remains necessary.
What did the Habermas Machine study find?
Published in Science, the Habermas Machine program involved 5,734 UK participants across a series of experiments. Participants preferred machine-drafted group statements to those from human mediators, and a virtual citizens' assembly showed convergence across rounds. Preference is not evidence of accuracy, legitimacy, or lasting agreement, and the authors frame the results as grounds for further research.
What did the NYC MyCity chatbot audit find?
The New York City comptroller's December 30, 2025 audit of the MyCity digital-service program identified inconsistent chatbot responses and weaknesses in management and evaluation. The Office of Technology and Innovation disputed the findings, including how accuracy was interpreted and the use of voluntary feedback. The two sides used different denominators, so the dispute is recorded, not resolved.
Did the MyCity chatbot cost what the audit's spending figure says?
No such conclusion is supported. The spending figures in the audit are program-level costs for a broader digital-service system, not chatbot-only costs. The audit record does not isolate the chatbot's share in a way this report can defend, so attaching the full program figure to the chatbot alone would misstate the record.
Does AI save cities time on public consultations?
The Department for Transport report estimated time and cost savings, but those were modeled against an assumed manual process, not measured in randomized comparisons of real consultations. Its stronger performance results came from nonblinded live use. A city considering a tool should treat the savings as a planning estimate and measure its own staff time, including review and correction.
Does participatory budgeting improve public services?
A scoping review by Campbell and colleagues found 37 studies across 39 articles, concentrated in Brazil and mostly nonrandomized. It found suggestive benefits in some settings but no general causal guarantee of improved health or public services. The review predates generative AI, and its lesson is that participation works through real authority over money and rules, not a better interface.
What should residents ask when a city uses an AI summary of public comments?
Ask how many comments were received and analyzed, whether originals are available, whether a person checked a sample including minority concerns, and whether any themes cannot be traced to real submissions. Ask for a visible correction log and a non-AI route to comment. Finally, confirm which public body makes the decision and when you can still speak.
Chapter 8 · Investing and Retirement Money
AI Financial Advice: Can You Trust a Chatbot With Your Retirement?
I have raised capital, and I have sat on both sides of the diligence table. Here's the thing you learn fast in that room: a confident answer and a correct answer are different products. They can come in the same package, and they often sound identical, but only one of them survives a second question. So when people ask me whether AI financial advice is any good, whether they can hand a chatbot their 401(k) question or their emergency fund math and trust what comes back, I hear the question underneath. It isn't "does the tool sound smart?" It obviously does. The real question is whether it helps you make a better-supported decision about your money, one you could defend if someone you trust asked you to explain it.
Before we go one paragraph further, the disclaimer, and I mean it plainly. I am not a financial adviser. Nothing in this chapter is investment, tax, or legal advice. I am not going to pick securities, forecast returns, estimate what anything is worth today, or tell you what withdrawal rate your household should use. What I am going to do is walk you through what the research actually found about AI and money, what the regulators have actually said, and a process you can use to check any answer, whether it came from a chatbot, from this report, or from a person with a nice office.
This matters because the evidence cuts in an uncomfortable direction. Three recent studies point the same way. Different AI tools give materially different answers to the same financial question. A simulation of lifetime saving found the advice follows the textbook in broad strokes but stumbles in ways that matter. And a controlled experiment found that AI recommendations measurably move where people put their pension money, without showing that the moved money does any better. Read that again. The tool changes behavior. Nobody has yet shown it improves outcomes. In finance, that gap can stay invisible for years and then show up at the exact moment it's most expensive to fix.
The relevant question for any financial assistant, human or machine, is whether it helps a person or a professional adviser make a better-supported decision. It is not whether the tool produces a confident, well-organized portfolio explanation. Those are two different accomplishments. The evidence reviewed in this chapter separates them cleanly: research documents material differences across tools given identical inputs (Journal of Financial Planning study), modeled weaknesses in life-cycle advice (life-cycle advice working paper), and experimental evidence that recommendations can shift pension allocations without demonstrating that the shifted allocations perform better (pension-choice preprint).
Think about what "better-supported" means in practice. A better-supported decision is one where you know which facts it rests on, which of those facts were verified, which assumptions filled the gaps, what the realistic alternatives were, and who is answerable if it goes wrong. A persuasive explanation can have none of those things and still read beautifully. In fact, fluent writing is exactly what these tools are built to produce, which is why fluency tells you so little.
This chapter evaluates research and professional process. It does not recommend products, rank tools, or tell you what to buy. Any generated financial text, from this report or from any tool, deserves the same treatment: it is material to be checked against your verified household facts and, where the stakes warrant it, reviewed by a qualified professional who is accountable for the recommendation.
That last phrase, "accountable for the recommendation," does a lot of work in the pages that follow. I'll come back to it in the regulatory section and again in the section for advisers. For now, hold onto the idea that accountability is not a feeling of being advised. It's a named person who owes you a duty and can be asked to explain.
Financial decisions sit at the intersection of everything this report has covered so far. A household's financial life touches schooling costs, career changes, housing, caregiving, health, and retirement. Those are the same life stages the earlier chapters walked through, from Chapter 3 on schooling to Chapter 6 on households. Money is the thread that ties them together, and a mistake in one stage tends to show up in the next.
The accountability problem is also sharpest here. When an AI tool gives a student a wrong answer on homework, the error usually surfaces on the next quiz. When it gives a household a wrong assumption about retirement, the error can sit quietly for a decade. The gap between a persuasive answer and a correct one may remain invisible for years and surface only when it is most expensive to fix. That is the core reason I treat AI financial advice with more caution than almost any other consumer use in this report.
If you are a worker weighing a retirement plan election, a parent thinking about education costs, a midcareer adult changing jobs, or an older adult managing drawdown, this chapter gives you a way to test what a chatbot tells you. If you are a financial adviser, broker, or compliance officer, the regulatory and process sections speak to you directly, and there is a section written for you in 8.6. If you are an investor looking at infrastructure, including the data center buildout that my own industry is part of, section 8.10 explains how to use public trackers, including ours, without mistaking a list for diligence.
8.2 Does ChatGPT Investment Advice Change From Tool to Tool?#
Short answer: yes, and by a lot. Here's what the study actually did.
The design is simple, and simple is a strength here. The researchers wrote standardized financial scenarios, held the core facts of each scenario constant, and gave them to seven free generative tools. The responses were collected in August 2025, and the study was published in June 2026. Because the inputs were identical, any difference in the outputs comes from the tools, not from the person asking. That is the whole value of a standardized prompt design: it takes the user out of the equation so you can see what the machine does on its own.
What came back was not one answer. It was a spread of answers. The study found differences in recommendations across tools, which means that a person who asks the same money question of two different chatbots can walk away with two different plans.
8.2.2 The emergency fund spread: $19,500 to $37,500#
The cleanest illustration is one emergency fund scenario. Across the seven tools, recommended fund sizes ranged from $19,500 to $37,500. That is an $18,000 spread for the same household facts.
Figure 8.1Data
Same scenario, seven tools: recommended emergency fund
One standardized emergency-fund scenario. The study shows variation, not which answer was right; it followed no real investors.
I want to be precise about what that number is and is not. It is an observed experimental output: what the tools did in August 2025 when given one standardized scenario. It is not this report's recommendation for an emergency fund. It is not a claim about what any household should hold. If you remember one thing from this section, remember that the range describes the tools, not you.
Here's the part most people skip. An $18,000 difference in an emergency fund recommendation is not a rounding error for most families. It's the difference between money sitting in cash and money going toward debt, a down payment, or retirement contributions. If two tools disagree by that much on the same facts, at least one of them is making assumptions you haven't seen. Probably both are.
The design disciplines its conclusions, and the authors' approach deserves respect for that. It used a small number of standardized scenarios. It did not follow actual investors. It did not measure realized returns, realized losses, or financial well-being. Differences among answers therefore establish variation, not error.
That distinction matters more than it sounds. The study cannot say which tool was closest to right for a real client, because no real client's outcome was tracked. Maybe $19,500 was the better answer for that scenario. Maybe $37,500 was. Maybe neither. The study was not built to tell you, and it doesn't pretend to.
What it does establish is narrower and still useful. A user who shops the same question across tools will receive materially different guidance. And an authoritative tone, or partial agreement among a few tools, is not a substitute for verified household facts and a defensible analytical method.
8.2.4 Two common misreadings of the variation finding#
The first misreading is "the tools are wrong." The study does not show that. It shows they disagree. Disagreement is a warning sign, but it isn't proof that any particular answer is bad.
The second misreading is the opposite: "just ask three chatbots and go with the majority." Partial agreement among a few tools is still not a substitute for your actual facts. Tools can share the same blind spot. If none of them asked about your income stability or your dependents, their agreement tells you they made similar guesses, not that the guesses were right.
8.2.5 The habit that matters: check completeness, consistency, and assumptions#
The operational implication of this study is a habit rather than a product. When you get a financial answer from any AI tool, check three things.
Completeness: did the response state the assumptions it made about income stability, existing savings, and dependents? Consistency: are those assumptions your actual circumstances? Sensitivity: does the recommendation change if you correct an assumption?
A tool that never states its assumptions makes you do all the verification work without telling you what needs verifying. That's the trap. The answer arrives clean and confident, and the guesses underneath it are invisible.
Here's a hypothetical to make it concrete. Imagine a household asks two different chatbots how large an emergency fund should be. One tool silently assumes a single, stable salary. The other silently assumes variable income. They return very different numbers, and each explains its number with calm authority. Neither answer is useless, but neither can be judged until the household knows which assumption each tool made. The fix is to ask the tool, directly, "What did you assume about my income, my savings, and who depends on me?" and then correct anything that's wrong and ask again. If the number moves a lot when one assumption changes, that tells you where your real decision lives.
8.3 AI Retirement Planning Study: Simulation Is Not a Retirement Outcome#
The second study goes deeper into time. Instead of asking one question once, it asks what happens to AI advice across a simulated lifetime of saving, working, losing a job, and retiring.
The March 2026 working paper uses human-written prompts and simulated economic lives to assess model-generated recommendations. That means the researchers built modeled individuals who move through a working life inside a simulation, and they asked AI systems what those individuals should do at each point. The modeled individuals are not people followed through decades of actual saving and retirement. They are constructs inside the authors' model.
That design has a real advantage. You can't ethically or practically run a forty-year experiment on real people's retirement money to see which chatbot does best. A simulation lets you ask "what would this advice do over a lifetime?" in a controlled way. The cost is that every answer comes back wrapped in the model's assumptions.
Within its model, the authors report findings worth taking seriously. The advice could broadly reflect life-cycle principles: save more when income is high, draw down in retirement. That's the textbook shape, and the tools got the shape.
But inside the same model, the advice still produced three weaknesses. First, excessive accumulation, meaning the simulated individuals piled up more than the model suggested was sensible. Second, weak responses to modeled job loss, meaning the advice didn't adjust well when the simulated person's income was interrupted. Third, differences associated with the characteristics of the prompt itself, meaning the way a question was worded changed what came back.
That third finding connects directly to the variation study in section 8.2. If the wording of a prompt changes the advice, then two people with the same situation who describe it differently may get different plans. The tool is responding to the description, not to the life.
Each of those is a finding within the paper's model, its assumptions, and its tested systems. None is an estimate of actual wealth lost by real clients, and none should be quoted as one.
Because this distinction is so central to responsible communication, I'll state it as a rule. A simulated wealth difference can identify a concern worth testing. It should never be described as an observed retirement loss or a guaranteed consequence of using a particular tool.
The verb matters. "The simulation produced excessive accumulation under these assumptions" is a reportable finding. "This tool cost retirees X dollars" is not, unless real retirees were followed and the dollars were measured. If you read a headline that turns a simulation into a dollar loss for real people, you're reading a misquote, whether the writer meant it or not.
This isn't pedantry. Once a modeled number gets described as a real loss, it travels. It shows up in sales pitches, in op-eds, in policy arguments. And it can cut both ways: a simulated weakness gets inflated into a scandal, or a simulated strength gets inflated into a promise. Either way, people make real decisions on a number that was never measured in the real world.
8.3.4 When scenario analysis helps, and when it misleads#
Scenario analysis is useful when its assumptions are visible. A base case, an adverse case, and an alternative case with consistent assumptions let a household see how sensitive a conclusion is to things nobody can know in advance. That's good practice, and it's exactly what a careful adviser does.
Scenario analysis becomes misleading the moment a model-generated path is presented as a forecast that has already been validated. The line between those two uses is not technical. It's a labeling choice, and it sits entirely within the control of whoever communicates the output.
Here's a hypothetical. A midcareer worker asks an AI tool to project how long her savings will last in retirement. The tool returns a smooth line showing money lasting well into old age. That line is one scenario under one set of assumptions. If the tool labels it as such, shows an adverse case where markets do worse or a job ends early, and tells her which assumptions drive the difference, she has something useful. If the tool presents the smooth line as "your retirement," she has a forecast dressed as a fact. Same math, different label, completely different risk.
If you use AI for retirement planning, ask it to show you at least one adverse scenario, and ask what it assumed about job loss. The working paper's finding about weak responses to modeled job loss is exactly the kind of blind spot you want to surface yourself. And if you rephrase your question and get a meaningfully different answer, take that as information: the tool is sensitive to wording, which means you should anchor on your verified facts rather than on whichever phrasing produced the answer you liked.
8.4 Does AI Advice Move Pension Choices? The South Korean Experiment#
The third study is the one I consider the hinge of this chapter, because it's the only one of the three that measures cause and effect on human decisions.
An August 2026 preprint reports an experiment with 400 employed South Korean adults aged 35 to 55 who had defined-contribution pension plans. Participants made incentivized hypothetical allocation decisions. The researchers varied the recommendations and the explanatory rationales experimentally.
That design is powerful for one specific reason. When you randomly vary what advice people receive, you can isolate the causal effect of receiving a recommendation from everything else about the participant: their education, their confidence, their prior beliefs. You're not guessing whether the advice changed behavior. You're measuring it.
"Incentivized" means the choices carried some stake within the experiment, which tends to make people take them more seriously than a pure survey would. "Hypothetical" means the allocations were not changes to their actual pension portfolios. Both words matter.
The effect was real and measurable. The reported pass-through estimate was 0.368, with a 95 percent confidence interval of 0.303 to 0.433. In plain terms, roughly 37 cents of every dollar of recommended allocation shift showed up in the participant's own choice, on the experiment's scale. Recommendations moved behavior.
The confidence interval tells you the estimate is not a fluke of noise. Even at the low end of the interval, recommendations moved choices meaningfully. At the high end, they moved them more. Whatever the precise figure, the direction is clear: people in this experiment followed the advice partway.
Figure 8.2Data
Influence is not benefit
Recommendations measurably moved allocations. The experiment did not show better measured portfolio performance and says nothing about actual retirement returns.
What the experiment did not find matters just as much, and I'm going to give it the same weight. It did not establish an improvement in its measured portfolio-performance outcome. And it cannot speak to actual long-term retirement returns at all: the allocations were hypothetical, and the horizon was an experimental session, not a working lifetime.
Persuasion is an outcome. It is not automatically a benefit. A recommendation that moves an allocation and a recommendation that improves a retirement are different claims, and this study supports only the first.
Capability is not benefit. Read the pair together: influence demonstrated, benefit not. That's the finding.
The study's status also disciplines its use. It is a preprint, which means it has been made public before completing the formal peer-review process. Its allocations were not changes to participants' actual pension portfolios. It should be treated as evidence about how people respond to advice under the experiment's conditions (educated, employed, midcareer adults making incentivized hypothetical choices) rather than as a retirement-policy evaluation of any live system.
That population detail is worth pausing on. The participants were employed, aged 35 to 55, and already had defined-contribution plans. Someone younger, older, less familiar with pensions, or under financial stress might respond differently, more or less. The experiment doesn't tell us. Generalizing from it to "everyone follows AI advice 37 percent of the way" would be a misreading.
8.4.5 Why this finding is the hinge of the chapter#
Put the three studies side by side. The variation study shows tools disagree. The simulation shows the advice has modeled blind spots. The pension experiment shows people act on the advice anyway, partway, whether or not it's good.
That combination is what makes accountability urgent. If recommendations measurably move people's money, then whoever or whatever generates them operates inside a system where someone must be answerable for their quality. A tool that nobody listens to can be wrong without much harm. A tool that moves roughly a third of the recommended shift into real choices, if that pattern held outside the lab, needs a named human standing behind it.
For a worker making a pension election: notice when you're being moved. If an AI explanation makes you want to change your allocation, that's the moment to slow down and check the assumptions, not the moment to act. The experiment suggests the pull is real.
For an employer or plan sponsor thinking about offering AI guidance to employees: the evidence here shows the guidance will likely influence choices. It does not show those choices will be better. Those are two separate questions, and only one has been answered.
For an adviser: your clients may arrive already partway persuaded by a chatbot. Part of your job is now to surface what the tool assumed and whether those assumptions fit the client in front of you.
8.5 Who Is Accountable for AI Financial Advice? FINRA and the SEC#
The regulatory record on this point is short, clear, and narrower than it's sometimes quoted. Two documents do most of the work.
8.5.1 FINRA Regulatory Notice 24-09 on generative AI#
FINRA's June 2024 Regulatory Notice 24-09 states that its existing rules remain applicable when member firms use generative systems, including supervision and communications requirements. In other words, a firm does not step outside its obligations because a third-party tool supplied the content. If a member firm sends a client something an AI drafted, the firm's supervision and communications rules apply to it the same way they would if an employee had written it by hand.
That's the core of FINRA's position on generative AI as the notice states it: not a new rulebook, but a reminder that the old one still applies.
8.5.2 The SEC's 2019 investment adviser interpretation#
Look at those two phrases and think about the studies in this chapter. "A reasonable understanding of the client's objectives" is exactly what a tool lacks when it silently assumes your income, savings, and dependents, as section 8.2 warned. "A reasonable basis for investment advice" is exactly what's missing when a simulated path is presented as a validated forecast, as section 8.3 warned. The standard was written long before these tools, and it fits them well.
A caution on scope. How the interpretation applies depends on the adviser-client relationship and the relevant legal context. This chapter is not a complete statement of current law. Securities advice, tax advice, insurance recommendations, and estate planning can each involve different responsibilities and different qualified professionals. If your question crosses those lines, so does the set of people you may need.
Figure 8.4Framework
Where accountability attaches in financial advice
FINRA Regulatory Notice 24-09 governs member firms, not every source of financial information. Not legal advice.
First misreading: FINRA governs everyone who talks about money. It doesn't. Notice 24-09 governs member firms. It should not be described as governing every person who offers financial information, and it says nothing one way or the other about a neighbor, a content creator, or a general-purpose chatbot outside the broker-dealer structure. If you get advice from a general chatbot on your phone, don't assume FINRA's notice covers that conversation.
Second misreading: a human signature fixes everything. A human signature at the bottom of machine-drafted content does not transfer accountability to the human unless the human actually reviewed it. A signature without substantive review satisfies a form and fails the standard this report proposes. The point of having a named professional is that the professional did the thinking, or at least checked it line by line. A rubber stamp is not review.
8.5.4 The operating principle: responsibility has to land somewhere#
Here is the principle I take from the regulatory record and the evidence together. AI assistance may expand research capacity: organizing disclosures, comparing fee structures, drafting scenario analyses for review. It should not obscure who is responsible for the recommendation.
Responsibility attaches to a named, qualified, accountable professional, or it attaches nowhere. And in financial advice, "nowhere" is the same as attaching to the client, who is usually the person least equipped to carry it. That's the whole game. When you use a chatbot on your own, you are the accountable party, whether you realize it or not.
8.6 For Advisers: Using AI Without Losing Your Fiduciary Footing#
This section is for financial advisers, registered representatives, and the compliance people who supervise them. Everything in it follows from the two regulatory documents and three studies above. None of it is legal advice, and your firm's counsel and compliance function remain the authority on how the rules apply to your practice.
The useful role for AI in an advisory practice is research capacity, not judgment. Organizing a client's disclosures into one place. Comparing fee structures side by side. Drafting base, adverse, and alternative scenarios for you to review. Flagging inconsistent definitions across documents. Surfacing questions you haven't asked the client yet. All of that can make your process more complete and more traceable, which is the direction the evidence points.
The variation study shows that different tools return materially different recommendations on identical inputs. If your workflow uses one tool and a colleague uses another, your firm may be producing inconsistent advice without anyone deciding to.
The life-cycle simulation shows that the wording of a prompt can change the output. If you describe a client's situation loosely, the tool responds to your description, not to the client. Your prompt becomes part of your advice, and it needs the same care.
The pension experiment shows recommendations move behavior. If a client sees an AI-drafted recommendation, they may lean toward it before you've reviewed it. Draft material that reaches a client before review is not a draft anymore in any practical sense.
FINRA's notice means member firms' supervision and communications rules still apply to AI-generated content. The SEC interpretation means the duties of care and loyalty, including a reasonable understanding of the client's objectives and a reasonable basis for the advice, don't pause because a tool did the first pass.
Here's a set of habits that follows from all of that.
Keep the client's verified facts separate from the tool's assumptions, in writing, so anyone reviewing the file can see which is which.
Require the tool's output to state its assumptions, and check each one against the client record.
Treat every AI-generated scenario as a scenario, labeled with its assumptions, never as a forecast.
Review before anything reaches the client. A signature on unreviewed content fails the standard, as section 8.5 explained.
Record who reviewed what, and when, so accountability is traceable to a named person.
Revisit prior recommendations when circumstances change rather than letting an old answer stand by default.
Limit the tool's permissions to what the task requires. Drafting is not trading authority.
None of this is exotic. It's what a careful practice already does. The difference is that AI makes it easier to skip steps without noticing, because the output looks finished before it is.
8.7 An Eight-Stage Process for Lifetime Financial Decisions#
The following process is proposed. It is not an empirically validated advice product, and I'm not claiming any study tested it. Its purpose is to separate four things that generated content tends to blur together: verified client facts, analytical assumptions, professional judgment, and client authorization. Each stage names its work product and its boundary.
Stage
Work product
Boundary
1. Establish the household position
Verified assets, liabilities, income, obligations, dependents, coverage, and liquidity needs
Do not infer missing facts from a demographic label
2. Clarify the objective
Priority, time horizon, risk capacity, preferences, and constraints
Willingness to take risk is not the same as capacity to absorb loss
3. Identify alternatives
Reasonable options, including delay or no change
Avoid presenting a single generated recommendation as the only solution
4. Assemble evidence
Dated primary disclosures, verified costs, applicable rules, and documented assumptions
A search summary is not a substitute for the underlying record
5. Analyze scenarios
Base, adverse, and alternative cases with consistent assumptions
Scenarios are not promises or realized outcomes
6. Review the recommendation
Qualified review of fit, conflicts, uncertainty, and implementation
The tool does not become the accountable adviser
7. Authorize and execute
Explicit client decision through approved systems
Drafting authority is not trading or payment authority
8. Monitor and revisit
Agreed review scope and triggers for changed circumstances
A prior answer should not remain current by default
Figure 8.3Framework
The eight-stage lifetime decision process
Proposed, not a validated advice product. The two highlighted stages are where generated advice most often fails without anyone noticing.
Stage 1 is about facts, not profiles. A tool that knows your age and job title may fill in the rest with averages for people like you. The boundary says don't. Your household is not a demographic label, and a plan built on inferred facts is a plan for someone else.
Stage 2 is where the objective gets defined, and it holds the first of two silent failure points, which section 8.8 covers in depth.
Stage 3 insists on alternatives, including doing nothing or waiting. A generated answer tends to arrive as the answer. The boundary keeps it as one option among several.
Stage 4 is the evidence file: dated disclosures, verified costs, applicable rules, and assumptions written down. A search summary, including an AI summary, points you toward the record. It isn't the record.
Stage 5 is scenario analysis in the sense section 8.3 described: base, adverse, and alternative cases with consistent assumptions. Scenarios are not promises.
Stage 6 is review by a qualified person. This is the stage FINRA's notice and the SEC's interpretation speak to most directly, and it's the stage the boundary protects: the tool does not become the accountable adviser.
Stage 7 separates deciding from doing. The client makes an explicit decision through approved systems. A tool that can draft a recommendation should not, by that fact alone, be able to execute a trade or move money.
Stage 8 holds the second silent failure point: a prior answer should not remain current by default.
8.7.2 The two stages where AI advice fails silently#
Two stages deserve emphasis because they are where generated advice most often fails without anyone noticing.
The first is stage 2's boundary: willingness to take risk is not the same as capacity to absorb loss. A retiree may cheerfully accept volatility on paper while lacking the years or income to recover from it, and a model that asks only about attitude will miss the distinction entirely.
The second is stage 8's boundary: a prior answer should not remain current by default. An allocation that fit a household at 35 with two incomes does not remain suitable at 58 after a layoff, and no answer should carry an unexamined shelf life. Chatbot conversations are especially prone to this. The answer sits in your history looking just as authoritative a year later as the day you got it.
The same structure serves a young worker choosing benefits, a family planning education costs, a midcareer adult changing jobs, or an older adult managing retirement. The calculations and the professional involvement differ across those situations. The framework does not imply identical recommendations, and using one framework does not make the outputs interchangeable. What stays constant is the separation of facts, assumptions, judgment, and authorization.
8.8 Why Willingness vs Capacity Matters More Than Your Risk Quiz#
Most online risk questionnaires, and most chatbot conversations about investing, ask some version of "how would you feel if your portfolio dropped?" That's a question about willingness. It matters. But it's only half the picture, and the evidence in this chapter suggests it's the half AI tools are most likely to rely on.
Willingness is psychological: how comfortable you are with ups and downs. Capacity is financial: how much loss your household can absorb without derailing a goal, given your time horizon, your income, your obligations, and your other resources.
They can point in opposite directions. Someone can be calm about volatility and have very little room to recover from it. Someone else can be nervous about every dip and have plenty of room. A plan that fits your willingness but exceeds your capacity can feel fine right up until it isn't.
Here's a hypothetical, with no real numbers, to show how the distinction plays out.
Imagine a couple nearing retirement. One spouse has always enjoyed following markets and describes themselves as comfortable with risk. Asked by a chatbot how they would react to a sharp drop, they answer sincerely: they'd hold steady. The tool, working from that answer, suggests an allocation that leans toward growth.
Now look at what the questionnaire didn't ask. The couple has a short time until they plan to start drawing on savings. One of them has recently lost a job, so household income has dropped. They are helping support an adult child. Their emergency savings are modest. Every one of those facts reduces capacity, meaning the household's ability to wait out a loss or replace it with new income.
Their willingness is high. Their capacity is limited. A tool that asks only about attitude sees the first and misses the second entirely. That is the exact failure the stage 2 boundary is written to catch.
Notice that the job loss detail also connects to the life-cycle working paper. Inside its model, that paper found weak responses to modeled job loss. I'm not saying the paper's simulation predicts what would happen to this hypothetical couple. It doesn't; its findings live inside its model. I'm saying it points to the kind of change a careful process should force you to re-examine.
If you're using AI to think about risk, tell it your capacity facts, not just your feelings: your time horizon, how stable your income is, who depends on you, what else you could draw on in a bad year. Then ask it directly whether its recommendation changes when those facts are included. If it doesn't ask for them in the first place, that's a sign it was working from willingness alone.
For advisers, the distinction is part of what a reasonable understanding of the client's objectives, in the SEC interpretation's language, looks like in practice. A risk quiz is an input. It isn't the objective.
8.9 How to Pressure-Test an AI Answer About Your Money: 10 Questions#
Everything in this chapter boils down to a set of questions you can ask of any financial answer, whether it came from ChatGPT, another AI tool, a website, or a person. Each question comes from a finding or a boundary above. Take this list into the conversation.
What did you assume about me? Ask the tool to list its assumptions about income stability, existing savings, and dependents. The variation study shows different tools give materially different answers on identical inputs, so hidden assumptions matter.
Are those assumptions actually true for me? Compare each one to your verified facts, not to what's typical for someone your age.
What changes if I correct one assumption? If the recommendation moves a lot, you've found where your real decision lives.
Would another tool say the same thing? If you check, expect differences. Treat disagreement as a reason to look at assumptions, and treat agreement as no substitute for your own facts.
Does the answer change if I word the question differently? The life-cycle working paper found differences associated with prompt characteristics. If rewording moves the answer, anchor on facts.
What's the adverse scenario? Ask for a base case, an adverse case, and an alternative with consistent assumptions. A single smooth path is one scenario, not a forecast.
Is this a scenario or a promise? Check the label. A simulated or projected result is not a realized outcome.
Does this reflect my capacity to absorb loss, or only my willingness? See section 8.8. Tell the tool your time horizon, income, and obligations.
Is this an educational example or advice for me? A general example is not personalized advice. If you can't tell which you received, assume it's general.
Who is accountable for this recommendation? If the answer is "no one," then it's you. For decisions where the stakes warrant it, bring the answer to a qualified professional who will review it and stand behind it.
One more habit that isn't a question: notice when an answer is moving you. The pension experiment found recommendations shifted hypothetical allocations with a pass-through of 0.368 among 400 employed South Korean adults, without showing better measured performance. Feeling persuaded is not evidence that the advice is good.
8.10 Evidence Use by an Investment Adviser: Trackers and Diligence#
For company and infrastructure research, the useful role of AI assistance is organizational: sorting verified disclosures, identifying inconsistent definitions across documents, surfacing unanswered questions, and preparing a comparison that a professional can review line by line.
The prohibited roles follow from the same evidence disciplines as the rest of this report. Do not invent figures. Do not fill missing contract terms with industry averages. Do not equate infrastructure scarcity with a security's expected return. That last one deserves emphasis because it's tempting in my industry. A shortage of something, whether power, sites, or capacity, can be real without telling you anything about whether a particular investment will pay off.
Capital disclosures, grid constraints, and permits concern different units and different stages of a project. Combining them into a synthetic investment fact without reconciliation manufactures precision that no single record contains. A capital announcement, a grid operator's constraint, and a permit filing may all concern related projects, but they are not three measurements of the same thing. Stitching them into one confident number is exactly the kind of move this chapter has warned against.
A tracker entry starts a question. It never ends one.
The proposed adviser checklist for any infrastructure-linked security or project requires confirming seven things:
The legal entity
The asset ownership
The contract layer
The project stage
The customer obligations
The costs
The financing terms
A tracker record should direct the analyst toward those questions. It flags that a project exists, that a permit was filed, that a grid operator raised a constraint. It should not answer the seven questions by implication. Chapter 11 sets out the full evidence standard this implies. The point here is that an index entry is the beginning of diligence, never the end of it.
Figure 8.5Framework
From tracker entry to adviser diligence
A tracker entry flags that a project, permit, or constraint exists. It starts diligence and never ends it. Savrn publishes these trackers and has a commercial interest in the sector.
8.11 How Should AI Financial Advice Be Evaluated?#
If someone wants to show that a financial assistant helps people, here's what a responsible trial would measure: factual accuracy, suitability of the process, assumption completeness, source freshness, error correction, and user comprehension.
It would also test whether users can distinguish an educational example from personalized advice. The standardized-prompt advice study shows that different tools answer the same question differently, so users can't take that distinction for granted. And it would test whether users understand adverse scenarios rather than only favorable ones.
One measurement trap deserves its own section. Short-term market performance alone is an inadequate test of advice quality. A risky recommendation can perform well by chance over any short window. A suitable, diversified decision can experience a loss in the same period.
An evaluation that scores recommendations by hindsight returns is measuring luck with a lag. The evaluation needs a defined objective, such as a suitable allocation for this household's horizon and capacity, and an appropriate horizon, not retrospective selection of favorable returns.
This matters for anyone reading marketing claims. If a tool's pitch rests on "our recommendations would have beaten the market last year," that is a claim about a window, not about advice quality. Show me the denominator: how many recommendations, over what period, measured against what objective.
8.11.2 Privacy and execution permissions are a separate test#
Privacy and execution permissions should be assessed separately from advice quality, because the two have nothing to do with each other. A tool that can explain a brokerage statement does not need unrestricted access to transact, move money, or expose a family's financial records.
The capability that justifies access is the capability that should receive it, and no more. Before you connect an AI tool to a financial account, ask what it can do with that connection, not just what it can tell you. Reading is one permission. Moving money is another. They should never arrive as a bundle by default.
8.12 Conclusion: Feeling Advised vs Being Well Advised#
The evidence reviewed in this chapter shows three things. AI financial advice can vary materially across tools given identical inputs (standardized advice study). Simulated life-cycle advice shows modeled weaknesses inside its model (life-cycle simulation). And recommendations can measurably influence allocation decisions (pension experiment).
It does not establish that unreviewed advice improves actual lifetime wealth or retirement security. The strongest causal finding, the 0.368 pass-through among 400 South Korean adults making hypothetical pension choices, is a finding about influence. Influence is not benefit.
The proposed path to value is a more complete, more traceable, and more understandable professional process, with the client keeping authority and the adviser keeping responsibility. Use the tools for legwork. Use the eight stages to keep facts, assumptions, judgment, and authorization apart. Use the ten questions every time an answer sounds sure of itself.
Financial confidence is not itself evidence of financial improvement. In this domain, more than any other in this report, the difference between feeling advised and being well advised is measured in years. The next chapter, Chapter 9, carries the same question into health and aging, where the stakes run just as long.
The research does not show that it reliably improves outcomes. A June 2026 Journal of Financial Planning study found seven free AI tools gave materially different recommendations on identical scenarios, including emergency fund sizes from $19,500 to $37,500 in one case. The study measured variation, not which answer was right, and it did not track real investors, so treat any AI answer as a starting point to check against your own verified facts.
Can I use ChatGPT for investment advice?
You can use it to organize information and generate questions, but it is not an accountable adviser. The studies reviewed here show AI tools vary across identical inputs and can move pension choices without shown performance gains. Ask the tool to state its assumptions, check them against your facts, and bring significant decisions to a qualified professional who will review the recommendation and stand behind it.
How much should my emergency fund be according to AI?
That depends on which tool you ask, which is the point. In one standardized scenario tested in August 2025, seven free AI tools recommended emergency funds ranging from $19,500 to $37,500 for the same facts. That range is an observed output from the study, not a recommendation for your household. Ask any tool what it assumed about your income, savings, and dependents.
Does AI advice actually change how people invest their retirement money?
In one experiment, yes. An August 2026 preprint studied 400 employed South Korean adults aged 35 to 55 making incentivized hypothetical pension allocations. AI recommendations passed through at an estimated 0.368 (95% CI 0.303 to 0.433). The study did not show improved measured portfolio performance, and the allocations were not changes to real pension accounts.
Did a study show AI costs retirees money?
No study reviewed here measured real retirement losses. A March 2026 working paper simulated economic lives and found, inside its model, excessive accumulation and weak responses to modeled job loss. Those are findings within the paper's model and assumptions, not measured losses to actual clients, and they should not be quoted as dollars lost by real retirees.
What does FINRA say about generative AI?
FINRA Regulatory Notice 24-09, issued in June 2024, states that existing FINRA rules, including supervision and communications requirements, remain applicable when member firms use generative AI. It applies to member firms. It does not govern every person or general chatbot that offers financial information outside the broker-dealer structure.
Is a financial adviser still responsible if they use AI?
The SEC's 2019 interpretation describes an investment adviser's duties of care and loyalty, including a reasonable understanding of client objectives and a reasonable basis for advice, and application depends on the relationship and legal context. A signature on machine-drafted content does not transfer accountability unless the human actually reviewed it. The tool never becomes the accountable adviser.
What is the difference between risk willingness and risk capacity?
Willingness is how comfortable you feel with ups and downs. Capacity is how much loss your household can absorb given its time horizon, income, obligations, and resources. They can point in opposite directions, and a tool that asks only how you'd feel about a drop can miss capacity entirely. Give any tool your capacity facts and ask whether its answer changes.
Chapter 9 · Health, Aging, Retirement
AI in Healthcare and Aging: What the Evidence Says for Patients and Older Adults
Maybe you're the one who drives to the appointments now. You keep a folder, or a notes app, with your mother's medication list, and you're never quite sure it matches what the pharmacy has. Your father's phone buzzes all day with texts about a package he never ordered, a toll he supposedly didn't pay, a bank account that has been "locked." And somewhere in the middle of all that, someone tells you that AI in healthcare research has shown these tools can read scans better than doctors, keep lonely seniors company, and catch scammers before they strike. The question you actually have is simpler and harder: which of this is real, and what should I let near my parent?
That's the question this chapter answers. Before anything else, a clear statement: this chapter is evidence synthesis, not medical advice. I'm a builder who reads studies, not a clinician. Nothing here tells you what to do about a specific symptom, drug, dose, or diagnosis. For that, you talk to your parent's doctor, pharmacist, or care team, and if something is urgent you call for emergency help. What I can do is show you what four pieces of research actually found, where each one stops, and how to think about the tools landing on your family's kitchen table.
Here's why it matters. The second half of life is where a lot of the stakes concentrate: more appointments, more prescriptions, more paperwork, more isolation for some people, and more people trying to take their money. Products are arriving to help with every one of those. Some of the underlying research is excellent. Some of it is thin. And the marketing rarely tells you which is which.
The evidence has a good part and an uncomfortable part. The good part: one of the largest randomized evaluations in this entire report, a Swedish mammography screening trial with more than one hundred thousand women, found that a specific AI-supported workflow did as well as the standard approach on its main safety outcome and found more cancers without more false alarms. The uncomfortable part: a physician trial showed that a model that performed well by itself did not measurably improve the doctors who used it, and the best controlled review on loneliness in older adults cannot tell you whether these tools help, do nothing, or hurt. Capability is not benefit. This chapter is about the distance between the two.
Before we look at a single study, I want to set the scoreboard, because the wrong scoreboard makes every study look better or worse than it is.
The value of assistance in the second half of life is broader than any clinical score. It includes independence, access to services, social connection, protection from exploitation, and support for caregivers. A diagnostic number or a satisfying conversation is not an adequate substitute for measuring those outcomes. If I only told you how accurate a model was, I'd be answering a question no family actually lives with.
Think about what your parent would say if you asked what a good year looks like. Probably something like: I got to stay in my own home. I could still drive to church, or the lake, or my sister's. I didn't miss my appointments. Nobody cleaned out my savings. I saw the grandkids. I made my own decisions. Almost none of that shows up in a benchmark. Some of it shows up in clinical trials. Most of it has to be measured deliberately, or it doesn't get measured at all.
The reviewed research spans a large screening trial, a small physician-reasoning trial, and two systematic reviews of interventions for older adults. These studies concern different technologies and different populations. They cannot be stacked into one claim that conversational systems improve health or solve loneliness. So I'm presenting them as what they are: one of the largest randomized evaluations in this report (the MASAI trial) sitting next to one of its most instructive null results (the JAMA Network Open physician trial), with the older-adult evidence (the BMC Geriatrics randomized-trial review) described plainly as thin.
Alongside those, I'll bring in two items that Chapter 6 covers in more depth: a caregiver meta-analysis and a medical self-assessment trial with ordinary adults. And I'll use a Federal Trade Commission report on fraud against older adults and a W3C accessibility guidance note, because a health chapter that ignores money and usability is not describing the life your parent is living.
Here's the one rule I'll keep coming back to. I call it the household boundary: assistance prepares, organizes, and explains; professionals diagnose, treat, and respond to urgency. It's a recommended safeguard. It is not a claim that every health-navigation product on the market has been evaluated against it, and it is not a verdict on any particular product. It's the line I'd draw for my own family based on what the evidence can and can't support today.
What does that look like in practice? A tool that helps your mother write down her three questions before a cardiology appointment is on the right side of the line. A tool that helps her understand the discharge instructions her nurse already gave her, so she can ask better follow-up questions, is on the right side. A tool that tells her whether her chest tightness is "probably nothing" is on the wrong side. The difference is not how smart the tool is. The difference is who holds responsibility for the decision.
9.2 AI in Healthcare Research: How the MASAI Mammography Trial Worked#
If someone tells you AI has been "proven" in medicine, there's a decent chance the study they're half-remembering is this one or something like it. So let's read it properly.
The strongest clinical evidence in the reviewed corpus is the Swedish MASAI trial, published in The Lancet. It randomized 105,934 women either to AI-supported mammography screening or to standard double reading, the standard practice in which two readers review each exam. After exclusions, 105,915 women were included in the reported analysis. (Lancet trial record)
Now the most important sentence in this section. The intervention was a specific screening system embedded in a radiologist workflow. It was not a general conversational model acting as an autonomous physician. Radiologists were still in the loop. The thresholds, the reading process, and the follow-up pathway all belonged to an organized screening program. That scope condition is the first thing you should carry out of this chapter, because nearly every misuse of this trial begins by dropping it.
The primary outcome was interval cancer. Interval cancers are cancers diagnosed between screening rounds: after a woman had a screening exam that did not flag a cancer, and before her next scheduled screen. They're the standard signal for whether a screening program is catching what it should. If a new approach to reading mammograms misses more cancers, you'd expect interval cancers to go up. So the primary question was essentially a safety question: does the supported workflow let more cancers slip through than the standard one?
The reported interval-cancer rates were 1.55 per 1,000 participants in the supported workflow and 1.76 per 1,000 in standard double reading. That's a ratio of 0.88, with a 95 percent confidence interval of 0.65 to 1.18. The result met the trial's noninferiority criterion. It did not demonstrate a statistically significant reduction in interval cancer. And the upper bound of that interval, 1.18, means the data are consistent with the supported workflow being somewhat worse, as well as somewhat better. (Primary outcome, Lancet)
Read that again. The point estimate leans in a good direction. The interval says the true answer could sit on either side of "no difference." The trial's claim is that the new workflow was not unacceptably worse, and that is a real, valuable claim. It is not a claim that the new workflow prevented cancers.
The secondary performance figures are more positive and equally scoped. Sensitivity was 80.5 percent with the supported workflow versus 73.8 percent with standard reading. Specificity was approximately 98.5 percent in both groups. In plain terms, the supported workflow caught a larger share of the cancers that were there, and the improvement in detection did not come with a compensating rise in false alarms. (MASAI trial)
Those are screening-performance results inside a defined radiologist workflow. They are not evidence of reduced mortality, because the trial was not designed to measure mortality. And they are not evidence about consumer self-diagnosis, where there is no radiologist, no calibrated threshold, and no defined workflow at all.
Figure 9.1Data
The MASAI mammography result, stated precisely
The interval includes values on both sides of 1.0, so the data are consistent with the supported workflow being somewhat better or somewhat worse on interval cancer. Detection rose without more false alarms.
This trial is the report's best example of a benefit claim stated precisely without exaggeration. Here are two sentences you might see written about it.
The first: "The AI-supported workflow achieved a noninferior interval-cancer outcome in this screening program, with higher sensitivity and similar specificity." The data support that.
The second: "The technology prevented 12 percent of cancers." The data do not support that. The 0.88 ratio is where that 12 percent comes from, but the direction of the point estimate does not survive its confidence interval, and mortality was never measured. The distance between those two sentences is the distance between evidence and marketing.
Misreading 1: "AI reads mammograms better than doctors." The trial compared two workflows, both involving radiologists. It did not test a machine against a human in isolation, and it did not remove the radiologist.
Misreading 2: "So my mother can upload her scan to an app." Nothing in this trial concerns consumer apps. The system operated inside a screening program with trained readers and defined follow-up. Take the workflow away and you've taken away the thing that was tested.
Misreading 3: "Fewer women will die of breast cancer." Maybe, someday, if detection gains translate into outcomes. This trial did not measure that, so no one can claim it from this trial.
Misreading 4: "Noninferior means it didn't work." No. Noninferior means it met a predefined standard of not being unacceptably worse on the primary safety outcome. Combined with the sensitivity result, that's a meaningful finding for a screening program deciding how to organize its reading work.
If you're an adult child, the practical takeaway is modest and useful. If your parent's screening program uses an AI-supported reading workflow, the best large trial in this area gives no reason for alarm about its primary safety outcome, within the limits above. If you have questions about how your parent's screening is read, ask the screening provider. That's a question for them, not for a chatbot.
If you're a health system board member or administrator, the MASAI result supports evaluating a defined system inside a defined workflow, with your own monitoring of interval cancers, sensitivity, and specificity. It does not support deploying a different system and assuming the same result.
9.3 How to Read a Screening Study: Sensitivity, Specificity, and Noninferiority in Plain English#
A lot of families tune out when they hit these words, and I understand why. But these three terms are how screening research talks, and once you have them, you can read a headline and know within a minute whether it's telling you the truth. I'll explain each using only MASAI's own numbers. No new data, just the concepts.
9.3.1 Sensitivity: of the cancers that were there, how many were found?#
Sensitivity answers one question: among people who actually have the condition, what share does the test correctly flag? In MASAI, sensitivity was 80.5 percent in the supported workflow and 73.8 percent in standard reading.
Here's a way to picture it. Imagine a room holding every woman in the trial who truly had breast cancer at the time of screening. Sensitivity is the share of that room the screening process correctly identified. A higher number means fewer cancers were missed at that screen. The supported workflow's higher sensitivity means it correctly flagged a larger share of that room.
What sensitivity does not tell you: how many healthy women were flagged by mistake. For that, you need the next term.
9.3.2 Specificity: of the people who were fine, how many were correctly left alone?#
Specificity asks the mirror question: among people who do not have the condition, what share does the test correctly clear? In MASAI, specificity was approximately 98.5 percent in both groups.
Picture a second, much larger room holding every woman who did not have cancer. Specificity is the share of that room the screening process correctly left alone. The rest got a false alarm: a callback, more imaging, maybe a biopsy, and a lot of worry.
Here's the part most people skip. You can raise sensitivity by flagging more people. Flag everyone and you'll catch every cancer, with sensitivity at its maximum and specificity collapsing, because you've also called back every healthy woman. That's why sensitivity alone is never enough. The MASAI result matters because sensitivity went up while specificity stayed at about the same level in both arms. The detection gain wasn't bought with more false alarms.
MASAI's primary result was reported as interval cancers per 1,000 participants: 1.55 in the supported workflow and 1.76 in standard reading. "Per 1,000" is a denominator, and a number without a denominator is a rumor. Interval cancers are rare events in a screening population, which is exactly why the trial needed more than one hundred thousand women to say anything about them.
The ratio of 0.88 compares the two rates. A ratio below 1.0 means the supported arm's rate was lower. A ratio of 1.0 would mean no difference. A ratio above 1.0 would mean the supported arm's rate was higher.
9.3.4 Confidence intervals: the range the data can't rule out#
The 95 percent confidence interval of 0.65 to 1.18 is the range of true ratios that are reasonably consistent with the data. Think of it as the study telling you, "Given what we saw, the true answer is probably somewhere in here."
Notice that 1.0 sits inside that range. The low end (0.65) would be a meaningful improvement. The high end (1.18) would be somewhat worse. When an interval contains 1.0 for a ratio, the study has not shown a statistically significant difference in either direction. That's why the correct statement is "no demonstrated reduction," not "a 12 percent reduction."
This is the concept people most often get backward. Most studies you hear about ask, "Is the new thing better?" A noninferiority trial asks a different question: "Is the new thing not unacceptably worse than the standard?" That's the right question when the new approach might offer other advantages and the main worry is that it could quietly let more harm through.
Before the trial starts, the researchers set a boundary for how much worse would count as unacceptable. If the confidence interval stays on the acceptable side of that boundary, the new approach is declared noninferior. MASAI met its noninferiority criterion on interval cancer. That tells you the supported workflow passed the safety test the trial set for it. It does not tell you the supported workflow was superior on that outcome, because the trial did not show that.
Noninferior is not the same as equivalent, and it's not the same as better. It's a pass on a specific safety bar.
9.3.6 A hypothetical: reading a headline in two minutes#
Here's a hypothetical to practice on. Suppose you see a post that says, "Huge study proves AI catches breast cancer that doctors miss and cuts cancer rates." You now have four questions to ask, and MASAI answers them:
What was compared? Two radiologist workflows, one supported by a specific system. Not AI versus doctors.
What was the primary outcome, and did it show superiority or noninferiority? Interval cancer, noninferior, no significant reduction.
Did sensitivity rise without specificity falling? Yes: 80.5 versus 73.8 percent sensitivity, specificity near 98.5 percent in both.
Was mortality measured? No.
So the headline gets one piece right (more cancers were detected in this workflow) and two pieces wrong (it wasn't AI versus doctors, and it didn't show fewer cancers or deaths). That took two minutes, and you didn't need a medical degree to do it.
Of the women without cancer, what share were correctly cleared?
About 98.5% in both groups
Interval cancer rate
How many cancers showed up between screens, per 1,000 screened?
1.55 vs 1.76 per 1,000
Ratio with 95% CI
How do the two rates compare, and how sure are we?
0.88, CI 0.65 to 1.18
Noninferiority
Was the new workflow not unacceptably worse?
Criterion met
Mortality
Did fewer women die?
Not measured
9.4 AI Diagnosis Accuracy vs Doctors: What the Physician Trial Found#
The second clinical study is small, and its lesson is out of proportion to its size. If you only remember one study from this chapter besides MASAI, make it this one.
A randomized trial published in JAMA Network Open involved 50 physicians working through structured diagnostic vignettes: written clinical cases designed to test diagnostic reasoning. Physicians were randomized either to conventional resources or to access to a language model. The outcome was a diagnostic-reasoning score graded on those cases. (JAMA Network Open trial)
Median diagnostic-reasoning scores were 76 percent and 74 percent, with the model-access group ahead by an adjusted difference of two percentage points. The 95 percent confidence interval for that difference ran from -4 to 8. The difference was not statistically significant. (Trial results)
Walk through that interval the same way we did for MASAI. A difference of zero sits inside it. The data are consistent with the tool making physicians a bit worse, making no difference, or making them moderately better. The study can't tell those apart. Under the tested conditions, access to the model did not produce a demonstrated improvement.
The vignette structure matters for interpretation. These were structured diagnostic cases, not observed patient outcomes. Fifty physicians grading hypothetical patients is a laboratory for reasoning, not a field trial of care.
But the finding that matters most is in an exploratory analysis. The model alone, without the physician, performed strongly on the same vignettes. And that strong standalone performance did not translate into a demonstrated improvement for the physicians who used it under the tested conditions. The tool was better than its effect. (Exploratory analysis, JAMA Network Open)
"Exploratory" is a flag you should respect. It means the comparison wasn't the trial's primary question, and it should be treated as a signal to investigate, not a settled result. Even so, the direction of the gap is the lesson.
Figure 9.2Data
Capability is not workflow effect: the physician trial
Structured vignettes, not patient outcomes. In an exploratory analysis the model alone performed strongly on the same cases; that strength did not show up as a demonstrated gain for the physicians using it.
That gap, standalone capability versus workflow effect, is the single most repeated pattern in this report. It showed up in consulting tasks in Chapter 2, in education in Chapter 3, and in the household studies in Chapter 6. Here it appears in its clearest clinical form, and it explains why I don't infer benefit from capability demonstrations. The performance of the model and the performance of the professional workflow require separate evaluation. A vendor who reports only the first has told you half of what you need.
Chapter 6 covers a companion result on the consumer side. In a randomized study published in Nature Medicine, Bean and colleagues recruited 1,298 U.K. adults to work through ten physician-authored clinical vignettes using GPT-4o, Llama 3, Command R+, or their usual resources. Each model performed strongly when directly prompted, and the people using those models did not reproduce that performance: control participants had 1.76 times the odds of identifying a relevant condition compared with pooled model users, while disposition accuracy did not differ significantly between each model group and control. All groups tended to underestimate acuity. That study used simulated scenarios, not people experiencing symptoms, and the models have changed since its 2024 data collection. (Nature Medicine study)
Put the two together and the pattern is hard to miss. A strong model in the hands of physicians did not measurably help on vignettes. Strong models in the hands of ordinary adults were associated with worse condition identification than usual resources on vignettes. Neither result says these tools can never help. Both say help cannot be assumed from a model's score.
9.4.5 What this result does and does not establish#
This result does not establish that clinical assistance cannot help. It establishes that help cannot be assumed, and that the evaluation has to happen at the level where the help is supposed to arrive: the clinician's actual decision, with the actual time pressure, liability, and patient context of real practice.
Common misreading: "The AI was smarter than the doctors, so the doctors were the problem." That reads a small exploratory comparison as a verdict on physicians. What the trial shows is that putting a capable tool next to a professional did not, under these conditions, produce a measurable gain. Why that happened is a question for further research, not a conclusion.
Common misreading: "The AI didn't help, so it's useless." A null result in 50 physicians on vignettes is not proof of no effect anywhere. The interval runs up to 8 points. The accurate reading is "not demonstrated," not "disproven."
For households, the boundary follows directly. Use assistance to prepare questions, organize records, and understand clinician-provided instructions. Seek professional care for diagnosis, treatment, and urgent symptoms. That boundary is a safeguard, not a verdict on any particular product.
Here's a hypothetical to make it concrete. Your father comes home from a visit with a printed after-visit summary and a new prescription, and he's confused about when to take it relative to his other pills. A reasonable use of an assistant: help him turn his confusion into a short list of specific questions, then call the pharmacist or the clinic with that list. An unreasonable use: asking the assistant to decide the timing and acting on its answer. The first keeps the professional in charge of the decision. The second quietly moves the decision to a tool nobody evaluated for that job.
If you're a clinician or a practice manager, the physician trial is a reason to evaluate any assistant at the level of your team's real decisions, not at the level of the vendor's benchmark. Ask for evidence that the workflow improved, not just that the model scored well.
9.5 AI for Seniors: Loneliness, Depression, and Social Isolation Are Not the Same Thing#
This is the section where I expect the most disagreement, because companionship products for older adults are being sold with a lot of warmth and not much evidence. I'm going to be careful here, because the people on the other end of these products are someone's parents.
A May 2026 review in BMC Geriatrics gathered eight randomized trials that included 611 older adults across heterogeneous interventions: social robots, voice systems, and cognitive activities. It is the most current controlled synthesis of its kind in the reviewed corpus. (BMC Geriatrics review)
"Heterogeneous" matters. These weren't eight tests of the same product. A social robot and a voice system and a cognitive activity program are different things. Pooling them tells you something about the category, and less about any one product.
For depression, the pooled estimate across seven trials was small, with Hedges' g of -0.25 and a 95 percent confidence interval of -0.48 to -0.02. The review rated that evidence moderate certainty.
Some plain English. Hedges' g is a standardized effect size: it expresses the difference between groups in units of how spread out the scores were, so that studies using different depression scales can be combined. Here, a negative value means lower depression scores in the intervention groups. A value of -0.25 is small by the usual conventions. The interval runs from -0.48 to -0.02, which means it just barely excludes zero.
So: a small effect, with an interval that barely excludes zero, rated moderate certainty. That's a finding of real but limited promise. Not nothing, and not enough.
For loneliness, the same review pooled three trials with 190 participants and reported g of -0.67 with a very wide interval of -2.57 to 1.23 and high heterogeneity. (BMC Geriatrics review)
Look at the point estimate alone and you'd think this was a big win, bigger than the depression effect. That's the trap. The interval runs from a very large benefit to a substantial harm. That's the statistical signature of "we do not know." Three trials, 190 people, high heterogeneity: the result does not establish a reliable average reduction in loneliness, and no responsible summary of this review can present it as one.
If you see a product page citing this review as evidence that its companion device reduces loneliness, you now know that's not what the review found.
The interventions lasted roughly two to twelve weeks and were not a uniform test of current general-purpose conversational models. They justify neither a broad claim about durable companionship nor any claim that a system can substitute for human contact. Two to twelve weeks is a demonstration window. Loneliness and social isolation are measured in years.
And those three words, depression, loneliness, social isolation, are not interchangeable. Depression is a clinical condition. Loneliness is the felt experience of lacking connection. Social isolation is an objective lack of contact. A person can be isolated without feeling lonely, or lonely in a full house. A product that improves one does not automatically improve the others, and a study that measured one tells you little about the rest.
9.5.5 The January 2026 Psychological Medicine review#
A separate January 2026 review in Psychological Medicine reported favorable before-and-after changes in similar outcomes. But its pooled estimates were within-group comparisons, because controlled evidence was limited. (Psychological Medicine review)
Within-group means comparing people to themselves: their scores before the program versus after. The problem is that a lot of things change between "before" and "after" besides the intervention. Time passes. People who enroll during a hard stretch often improve anyway, which statisticians call regression to the mean. Someone showed up repeatedly over the study period and paid attention to them. Within-group change can't separate the intervention from any of that. Such estimates are observations, not effects.
The two reviews should not be averaged together as independent confirmations. Their comparison methods differ (one pooled controlled contrasts, the other pooled within-group change), and their underlying studies may overlap, so an apparently stronger result in one review does not resolve the uncertainty in the other. Averaging them would manufacture a precision neither review claims.
Here's a hypothetical of how this goes wrong. A vendor slide reads: "Two 2026 meta-analyses confirm companion technology reduces loneliness in seniors." One of those reviews found a loneliness interval spanning benefit to harm. The other measured change within the same people, without a control group, and may share studies with the first. Two weak signals, one of which may be partly the same signal counted twice, do not add up to a strong one.
The plain state of the evidence:
Depression: a small, moderate-certainty signal.
Loneliness: unresolved uncertainty.
Companionship replacement: no basis for claims of any kind.
Figure 9.3Data
Depression and loneliness are different outcomes
BMC Geriatrics review of eight randomized trials, 611 older adults, interventions of about 2 to 12 weeks. The January 2026 Psychological Medicine review pooled within-group change and is not comparable, so it is not plotted or averaged.
9.5.7 The caregiver evidence belongs in the same frame#
If you're the adult child in this story, you're also a caregiver, and there's evidence about you too. Chapter 6 reviewed digital caregiver interventions. A registered 2026 systematic review in JMIR included 35 randomized trials with 3,388 informal dementia caregivers. Interactive eHealth interventions produced a pooled burden effect of SMD -0.26 (95% CI -0.42 to -0.10), but heterogeneity was 73.6 percent and the 95 percent prediction interval ran from -1.10 to +0.58. (JMIR caregiver intervention review)
The confidence interval describes the average, which favors the interventions. The prediction interval describes what the next specific program might do, which could be substantial benefit, little effect, or an adverse result. Only three studies were rated low risk of bias, the evidence was judged moderate certainty, and it relied heavily on subjective outcomes. The review covered web, mobile, video, and hybrid programs and does not isolate a general-purpose generative model effect.
So that evidence concerns caregiver support, with small average burden effects and substantial uncertainty about transfer to new settings and to general-purpose models. It is not a comprehensive solution for older adults' health or independence. For more, see Chapter 6.
For adult children: If you're considering a companion device or voice assistant for a parent, go in with a small, specific goal (for example, a daily check-in routine your parent wants), and watch whether human contact increases, stays the same, or quietly shrinks. The evidence doesn't tell you it will help with loneliness. Your own observation, over months, is the only evaluation you'll get.
For older adults: You get to decide whether you want one of these. If someone buys you a device "so you won't be lonely," it's fair to say that's not what you need, or to try it on your own terms and put it away if it doesn't earn its place.
For senior living operators and program funders: A pilot of two to twelve weeks is a demonstration. If you're deciding whether to fund a program at scale, ask for controlled comparisons, the outcome that matters to residents, and follow-up that outlasts the novelty.
9.6 A Later-Life Workflow That Starts With the Person's Own Goals#
Now for the practical part. If the evidence can't tell you which product to buy, what can you actually do? My answer is a workflow, not a product. And the first step in that workflow isn't a feature. It's a conversation.
Every proposed use in this section requires local testing. I'm claiming no demonstrated effect for any item merely because it sounds plausible. The workflow's first step is the person's own goals and consent, and that means separating two lists: the tasks the person wants help with, and the tasks other people would prefer to automate for their own convenience.
Those two lists diverge more often than any product demo suggests, and the divergence is where autonomy gets lost. Your mother may want help keeping her appointments straight. She may not want you getting a notification every time she opens the refrigerator. Both might be "safety features." Only one is hers.
Here's a hypothetical. A family sets up a shared calendar, medication reminders, and a location-sharing app for a father who has recently been widowed. The adult children feel reassured. The father feels watched, stops carrying his phone on walks, and misses calls. Nothing in that setup was malicious, and nothing in it started with his goals. The fix isn't a better app. It's going back and asking him which of those he actually wants.
Appointment preparation: Organize the person's questions and existing records without diagnosing from incomplete information.
Medication communication: Help prepare a verified list for a clinician or pharmacist; do not autonomously change doses or treatment.
Service navigation: Explain documented eligibility and application steps, while preserving a human route for exceptions.
Care coordination: Share only information the person has authorized, with clear roles and a current contact plan.
Daily independence: Test whether reminders or interfaces actually help the person complete a task rather than simply increasing caregiver monitoring.
Emergency continuity: Keep an accessible plan that works when the digital service, connectivity, or power is unavailable.
Figure 9.4Framework
A later-life workflow that starts with the person's goals
Proposed uses, each requiring local testing. The test at the center: does the person do more and decide more, or are they simply monitored more?
Each item has a verb that keeps a person in charge: organize, prepare, explain, share, test, keep. None of them says decide, prescribe, or diagnose. That's deliberate. It's the household boundary from section 9.1 turned into a daily routine.
9.6.3 Daily independence inverts the usual measurement#
Two items carry more weight than their bullet size suggests. The first is daily independence. It flips the usual measurement on its head. A reminder system that lets a caregiver watch more closely while the older adult completes fewer tasks independently has moved in the wrong direction, even if every dashboard trend line points up.
That's worth sitting with. Most eldercare technology gets judged by the people buying it, and the people buying it are often the adult children. So the metrics that get reported tend to be the ones that reassure the buyer: alerts sent, check-ins logged, dashboards viewed. The question that matters to the person living with the technology is different: can I do more of my own life than I could before? If the answer is no, the product is serving the family's anxiety, not the person's independence.
9.6.4 Emergency continuity is the test most systems fail silently#
The second is emergency continuity. It's the test most systems fail without anyone noticing, because nobody runs the test until the day it matters. A plan that exists only inside an app is not a plan during the outage when it's most needed.
Here's how it might look for one family, as an illustration only.
An older woman who lives alone has a cardiology appointment next week. She's comfortable with her tablet and wants help getting organized. Step zero: she says she wants help with preparing for appointments and keeping her medication list straight, and she does not want her children reading her messages.
For appointment preparation, she uses an assistant to turn her scattered notes into a short list of questions and to pull together the dates of her recent tests from the paperwork she already has. She doesn't ask it what her symptoms mean. For medication communication, she and her daughter build a list from the actual pill bottles, and she brings it to the pharmacist to verify. For care coordination, she decides her daughter can see the appointment calendar, and nothing else. For daily independence, she tries medication reminders for a month and checks whether she's taking her pills more reliably on her own, not whether her daughter is getting alerts. For emergency continuity, a printed copy of the verified list and her contacts sits on the refrigerator, and her daughter has one too.
Nothing in that example proves any tool works. That's the point. The workflow is how a family tests, locally, whether assistance is helping the person it's supposed to help.
9.7 A Family Checklist for Using AI Around a Parent's Health#
This is the list I'd want in my pocket. Take it to a family meeting, a sibling group text, or a conversation with your parent. It doesn't require any technical skill, and it doesn't replace anything a clinician tells you.
Did your parent ask for this? Write down what they want help with, in their words. If the list is mostly your worries, stop and talk first.
Who decides? Agree that your parent (or their legally authorized representative, where one exists) makes decisions about their care, and that the assistant never does.
What's the boundary? Assistance prepares, organizes, and explains. Professionals diagnose, treat, and respond to urgency. Say it out loud so everyone hears the same rule.
What information goes in? Share only what's needed for the task, and only what your parent has agreed to share.
If something feels urgent, call the clinician, a nurse line, or emergency services. Don't ask a chatbot first.
Remember the evidence from the Nature Medicine vignette study: people using models did not match the models' standalone performance, and all groups tended to underestimate acuity. A tool's confidence is not a safety signal.
Keep paper backups for medications, contacts, and key instructions.
Test the plan: phone off, internet down, power out. Does your parent still know what to do?
Set a date to review what's working, with your parent in the room.
Ask one question every time: is my parent doing more of their own life, or are we just watching more closely?
9.8 Elder Fraud and AI: Fraud Protection Without Treating Older Adults as Helpless#
Back to that phone full of scam texts. This section is about money, and it's also about dignity, because the fastest way to fail an older adult on fraud is to treat them as if they can't think.
The Federal Trade Commission reported that adults aged 60 and older reported $2.4 billion in fraud losses in 2024. Three qualifiers attach to that number, and they belong in the same breath. (FTC older-adult report announcement)
These are reported losses, not a complete population estimate. Many fraud losses are never reported. The true total is unknown.
The report does not attribute the total to generative systems or any particular technology. If you see "$2.4 billion lost to AI scams," that's not what the FTC said.
The same reporting distinguishes the likelihood of reporting a loss from the amount lost in particular scam categories. How often a group reports losing money and how much they lose when they do are different measurements.
9.8.2 The equity point: aggregate losses are not a verdict on capability#
That third qualifier carries an equity point I insist on. Aggregate loss data do not support treating every older adult as less capable of recognizing deception. Many older adults report fraud at higher rates than younger adults precisely because they check their statements and know how to complain. The stereotype of universal vulnerability is not in the FTC's numbers, and it shouldn't be smuggled into products or policies built on them.
Common misreading: "Older people fall for scams more." The data described here don't say that. They say a certain amount of loss was reported by people in a certain age group. Reporting behavior, the kinds of scams involved, and the amounts at stake all shape that number. Turning it into a statement about the mental capacity of an entire generation is a stereotype, not a finding.
Why does this matter practically? Because a protection design built on the stereotype takes control away. It locks accounts, routes every transaction through a child, and treats every decision as suspect. A protection design built on the evidence keeps the person in control and gives them better tools and a pause button.
Independently verified contact information. When a message says it's from the bank, the pharmacy, a government office, or a grandchild, the person contacts that party through a number or address they already trust, not the one in the message.
A pause before unusual transfers. Urgency is the scammer's main tool. A built-in waiting period, set by the person, takes it away.
Trusted-person involvement when the person has authorized it. A named person the older adult chose, involved in the ways the older adult agreed to. Not automatic oversight.
A clear escalation route. Everyone knows who to call and what to do if something looks wrong, including how to report it.
Here's the limit, stated plainly. This review contains no trial establishing that a particular automated scam detector prevents real-world losses. That study has not been done in the reviewed corpus, and any product claiming otherwise is ahead of its evidence.
Automated warnings need evaluation, and the evaluation is double-sided. This is the same sensitivity and specificity logic from section 9.3, applied to your father's phone.
A system that produces frequent false alarms interferes with legitimate activity and trains the person to dismiss warnings. That habit collapses exactly when a real alarm arrives. That's the specificity problem. A reassuring false negative is worse: the tool says the message looks fine, the person relaxes, and the scam goes through with a stamp of approval. That's the sensitivity problem.
Fraud protection, like everything else in this chapter, is a measured-outcome question, not a feature-list question. "Our tool flags scams" is a feature. "Our tool reduced losses among users compared with a similar group, without so many false alarms that people stopped listening" would be an outcome. I haven't seen that outcome in the reviewed evidence.
Figure 9.5Comparison
Fraud protection with its qualifiers attached
Proposed protective workflow: independently verified contacts, a pause before unusual transfers, trusted-person involvement when authorized, and a clear escalation route.
As an illustration: an older man gets a frantic call from someone claiming to be his grandson, saying he's in trouble and needs money wired today. Under a stereotype-based design, his account might be frozen until one of his children approves any transfer, which also blocks him from paying his own contractor next week. Under the workflow above, he has already agreed with his family on a simple rule: any urgent money request gets a pause and a call back to a number he already has. He hangs up, calls his grandson's real number, and finds out the truth. He stays in charge. The pause did the work.
Nothing in that example depends on artificial intelligence. That's worth noticing. The strongest protections here are process, not software.
For older adults: You're allowed to hang up. You're allowed to call back on a number you trust. Real institutions don't punish you for verifying.
For adult children: Offer tools, not supervision. Ask your parent what help they want with fraud, and agree on a pause rule and a call-back rule together.
For banks, carriers, and product teams: Evaluate your warnings on both sides. Count false alarms and missed scams. Don't design protections that treat every customer over 60 as incapable.
W3C's cognitive-accessibility guidance recommends usable content and interfaces beyond basic technical conformance. It is published as a Working Group Note, which makes it design guidance, not a new normative conformance standard and not an effect estimate for any particular product. (W3C cognitive accessibility guidance)
That distinction matters if you're a buyer. A vendor saying "we follow the W3C cognitive accessibility guidance" is saying something about design intent. It is not evidence that their product improves outcomes for anyone.
The implementation standard I'm proposing draws on that guidance:
Clear language. Plain words, short instructions, no jargon.
Recognizable controls. Buttons and actions that look like what they do and stay where people expect them.
Recoverable errors. A mistake shouldn't be a trap. People should be able to undo and try again.
Understandable consent. A person should know what they're agreeing to, especially about what information is shared and with whom.
Testing with the intended users. This is the item most often skipped: testing with the actual people the product is for, rather than a convenience sample of developers and their relatives.
Age alone should not be treated as a proxy for disability, digital skill, or desired level of assistance. The 70-year-old retired systems engineer and the 70-year-old with early cognitive decline are different users who happen to share a birth year. Any design or policy that treats them as one demographic will fail at least one of them.
This is the same equity point as the fraud section, from a different angle. Designing for "seniors" as a single group tends to produce either condescension or neglect. Designing for real people with real needs, and testing with them, is harder and works better.
The final requirement is structural: a public service should remain available when a person declines automated interaction. That's a proposed equity and continuity requirement, not an empirical finding that existing services meet it. And if you've tried to reach a bank, an airline, or a benefits office by telephone lately, you already know how often the requirement goes unmet.
The non-digital route is the emergency continuity principle from section 9.6, applied to public services. If the only way to reach help runs through an app or a chatbot, then everyone who can't or won't use that channel has been quietly shut out.
Retirement decisions connect finance, housing, relationships, health, transport, and community participation. That's why this report gives the topic two chapters rather than a section. Chapter 8 addressed research on advice quality and declined to convert simulations into prescribed plans. This chapter doesn't convert them either. I'm not going to tell you how to allocate a retirement account or when to claim benefits.
The proposed later-life assessment asks five questions of any assistance an older adult uses:
Does it reduce missed obligations (appointments, bills, renewals)?
Does it improve access to services and care?
Does it support meaningful contact with other people?
Does it preserve autonomy?
Does it avoid harm?
The outcome set is deliberately two-sided. Adverse dependence, unwanted surveillance, incorrect advice, caregiver workload, and the availability of human support must be measured alongside user satisfaction. A companionship product that scores well on satisfaction while quietly displacing human contact has failed the assessment I'm proposing, whatever its survey numbers say.
Here's the part most people skip. Satisfaction is easy to measure and pleasant to report. Displacement is hard to measure and uncomfortable to report. So if nobody deliberately measures the hard side, the easy side wins by default, and products get judged by how much people like them rather than whether their lives got better. That's the whole game.
9.10.3 Follow-up separates a demonstration from a life#
Follow-up is the discipline that separates a demonstration from a life. The short durations (roughly two to twelve weeks) and heterogeneous interventions in the older-adult trial review limit every claim about durable effects, and a short successful demonstration can differ materially from months of routine use. (BMC Geriatrics randomized-trial review)
Any evaluation that ends when the novelty ends has measured novelty. A new device in the house gets attention for a while. Grandkids ask about it. The person tries it out. Then it becomes furniture, or it becomes part of life. Only follow-up tells you which.
For older adults: You're the one who decides whether a tool is worth keeping. Judge it by whether you can do more, decide more, and connect more.
For adult children and caregivers: Measure your own workload too, and be candid about it. The caregiver evidence shows small average effects and wide uncertainty for any specific program, so watch whether a tool actually lightens your load or just changes its shape.
For advisers and service providers: Keep a human route open, keep your claims inside your evidence, and don't use aggregate statistics to stereotype your clients.
For board members and funders: Fund evaluations with control groups, two-sided outcomes, and follow-up beyond the pilot window. Chapter 11 sets out the full accountability framework and evidence classes I use to judge these claims.
9.11 Conclusion: What AI in Healthcare Research Supports Today#
The strongest clinical example in this chapter is a narrowly defined, professionally supervised screening workflow in which a model-support system achieved noninferior interval-cancer outcomes and improved sensitivity without losing specificity (MASAI).
The conversational and older-adult evidence contains important null results and real uncertainty: a two-point physician-workflow difference that is statistically indistinguishable from zero (physician trial), a loneliness estimate whose interval spans benefit to harm, and no controlled basis for companionship-replacement claims (older-adult review). Those findings support differentiated evaluation, task by task, workflow by workflow, outcome by outcome. They support neither blanket endorsement nor blanket rejection.
The standard I'm proposing is assistance that preserves independence and access to people, not a system that replaces human support because replacement is easier to count. Retirement should remain a stage of agency and participation. The measure of success is an older adult who can do more, decide more, and connect more, not one who has been efficiently managed.
For the adult child at the kitchen table with the folder, the pill bottles, and the phone full of scam texts: you don't need to become an expert. You need a boundary (tools prepare, professionals decide), a workflow that starts with your parent's goals, a paper backup, a pause rule for money, and the habit of asking what the evidence actually measured. That's enough to use these tools without being used by them.
The Swedish MASAI trial analyzed 105,915 women and found 1.55 interval cancers per 1,000 with AI-supported screening versus 1.76 with standard double reading, a ratio of 0.88 (95% CI 0.65 to 1.18). That met the noninferiority criterion but was not a significant reduction. Sensitivity rose from 73.8% to 80.5% with specificity near 98.5% in both groups. Mortality was not measured, and the result applies to a radiologist workflow.
Is AI more accurate than doctors at diagnosis?
The evidence reviewed here does not show that. In a randomized trial of 50 physicians on structured vignettes, access to a language model changed median reasoning scores by an adjusted 2 points (76% vs 74%, 95% CI -4 to 8), not statistically significant. The model alone performed strongly in an exploratory analysis, but that capability did not translate into a demonstrated improvement for the physicians using it.
Do AI companions reduce loneliness in older adults?
The best controlled evidence reviewed cannot say. A May 2026 BMC Geriatrics review pooled only three trials with 190 participants for loneliness and found g -0.67 with a 95% CI of -2.57 to 1.23, which spans large benefit to substantial harm. Interventions lasted roughly 2 to 12 weeks, and there is no controlled basis for claims that a device can replace human contact.
Can AI help with depression in seniors?
A May 2026 review of randomized trials found a small effect on depression across seven trials (Hedges' g -0.25, 95% CI -0.48 to -0.02), rated moderate certainty. The interventions included social robots, voice systems, and cognitive activities lasting roughly 2 to 12 weeks, not a uniform test of current chatbots. It is a signal of limited promise, not a treatment, and not medical advice.
Is it safe to use a chatbot to check my parent's symptoms?
This chapter is not medical advice, and the evidence supports a clear boundary: assistance can prepare questions and organize records, while professionals diagnose, treat, and handle urgency. In a Nature Medicine vignette study of 1,298 U.K. adults, people using models identified relevant conditions less often than those using usual resources, and all groups tended to underestimate acuity. For urgent symptoms, contact a clinician or emergency services.
How much money do older adults lose to scams, and is AI to blame?
The FTC reported that adults aged 60 and older reported $2.4 billion in fraud losses in 2024. Those are reported losses only, so the real total is unknown, and the FTC did not attribute the total to generative AI or any particular technology. The data also do not support treating every older adult as less able to spot deception.
Do AI scam detectors protect seniors from fraud?
No trial in the reviewed evidence shows that a particular automated scam detector prevents real-world losses. Warnings also need two-sided evaluation: frequent false alarms train people to ignore alerts, and a reassuring miss can be worse. The protective workflow proposed here relies on verified contact information, a pause before unusual transfers, a trusted person the older adult chooses, and a clear escalation route.
How should families set up AI tools for an aging parent?
Start with the parent's own goals and consent, not the family's convenience. Use tools for appointment preparation, verified medication lists for a pharmacist, service navigation, and authorized care coordination. Check whether the parent completes more tasks independently rather than simply being monitored more, and keep a paper plan that works if the phone, internet, or power fails. Every use needs local testing.
Chapter 10 · Data Centers and Communities
Do Data Centers Raise Electric Bills and Drain Water? Power, Water, Jobs, and Taxes
I've spent my career building large power infrastructure. So when people ask me about data center impact on communities, I don't answer from a think tank. I answer from the builder's side of the table at the public meeting, the side where you sit with your drawings and your load letters while the people in the folding chairs decide whether they trust you. Here's what that side of the table teaches you fast: the community does not care what you call the building. Data center, mine, campus, refinery. None of that matters. They care about four things. The meter on the side of their house. The water line that runs to their kitchen. The noise at two in the morning. And the tax bill that shows up every year whether they wanted the project or not.
Those are the right four things to care about, and this chapter is organized around them. The questions people type into search engines are the same ones they ask at the microphone. Do data centers raise electric bills? How much water does a data center use? How many jobs does a data center actually create? Does the tax revenue make up for the incentives? Each one has a real answer that depends on records, and each one gets muddied by numbers that are true somewhere and misapplied everywhere else.
Here's what the evidence will show, including the uncomfortable part. The national numbers on electricity and water are large, well-constructed, and model-based, and they tell you to plan without telling you where. The best state-level audit I've seen found that one state's rates were fairly allocating costs at the time of the study, and in the same breath modeled a residential bill increase by 2040 under growth scenarios. One utility tariff shows a serious attempt to put the risk of underused capacity on the big customer instead of on households, and it's an attempt, not a proven outcome. The job numbers are real and routinely misread. The fiscal numbers vary so much by locality that the word "typical" barely applies. And no one has produced a complete national inventory of data center water use.
Now the part I owe you before anything else. Savrn is built around behind-the-meter power and a zero-makeup-water design goal. Those are Savrn's claims. They are exactly the kind of claims this chapter teaches you to test, so the tests apply to us first. I'm not writing this to win an argument with residents. I'm writing it so a resident, a county commissioner, or a utility regulator can walk into the room with better questions than the developer expects, including when the developer is me.
10.1 Data Center Impact on Communities: Two Claims to Keep Apart#
I open with a boundary, because nearly every bad argument about computing infrastructure, on both sides, comes from violating it. A useful learning tool does not prove that a particular data center benefits its neighbors. Conversely, a poorly designed facility does not establish that every application running on computing infrastructure lacks value. The application's benefit and the facility's local effects are separate claims, supported by separate evidence, and I keep them apart on every page of this chapter.
That separation isn't evasive. It's how the evidence is actually organized. The chapters on schooling and working life established that specific, bounded applications of these systems can produce measured benefit in defined settings. Nothing in those findings says anything about what a given facility does to its host community's power bills, water supply, noise environment, or tax base. And nothing a community experiences locally says anything about whether a tutoring program three states away works. Both directions of the inference are unsupported, and I won't make either.
Figure 10.1Concept
Two claims that must never be merged
A useful application does not prove a facility benefits its host community, and the reverse inference fails too.
The merge happens because each side wants the other side's evidence to carry its conclusion. A promoter who can point at a promising study of students or workers wants that benefit to travel down the transmission line and land in the county that hosts the servers. It doesn't. The students who benefit may live nowhere near the facility, and the county's bill, water, and noise questions are untouched by their test scores.
The opponent's version runs the other way. A facility that was sited badly, loud at night, or approved on thin fiscal analysis becomes proof that the whole technology is worthless. It isn't. A badly run gas station does not prove cars have no value. It proves that particular gas station needed better conditions.
Here's the part most people skip. Keeping the claims apart helps residents more than it helps developers. When the benefit argument is allowed to stand in for the local argument, the local questions never get answered, because the room has already been told the project is "good for the future." Separate the claims and the local questions come back to the center of the table, where they belong.
10.1.2 What Savrn's stated objective is, and is not#
Savrn's stated objective of improving the relationship between data centers and communities is treated here as a commitment to test, not as evidence that any particular facility has already met it. The community standard I propose is outcome-based: identify the actual power and water dependencies, allocate costs transparently, measure local operating effects, and make commitments enforceable. Communities should be able to approve, modify, defer, or reject a project on that record.
One more boundary, and it matters. This chapter is national research and documented regulatory cases. It is not a site-specific engineering, fiscal, or environmental assessment, and no sentence in it should be read as one. If you are deciding on a specific parcel, a specific feeder, or a specific tax agreement, the records for that site govern. This chapter tells you which records to ask for and how to read them.
Resident: You don't need to form an opinion on artificial intelligence to form an opinion on a facility. Ask about your meter, your water, your nights, and your taxes, and ask for the records behind each answer.
County commissioner: Separate the economic development pitch from the local obligations. A company's national story is not a condition in your agreement.
Utility regulator: National load growth is context. The docket in front of you is about a specific tariff, a specific service territory, and a specific allocation of costs.
Investor: A community that can see the record is a community that can say yes with confidence. Projects approved on vague claims carry risk that surfaces later.
10.2 How Much Electricity Do Data Centers Use Nationally?#
Lawrence Berkeley National Laboratory's December 2024 report estimated that United States data centers used approximately 176 TWh of electricity in 2023, about 4.4 percent of national electricity consumption. The qualifiers attach immediately: these are model-based estimates, not a complete facility-level meter census, and they describe a national total assembled from modeled server categories, utilization rates, cooling assumptions, and infrastructure characteristics. (LBNL report)
The same report presented 2028 scenarios of approximately 325 to 580 TWh, equivalent to 6.7 to 12 percent of projected national electricity consumption under its assumptions. Those numbers need three labels that are often dropped. They are scenarios from a 2024 report, not actual 2026 consumption. They are ranges across modeled futures, not a statistical confidence interval. And the underlying total covers data centers broadly; it should not be described as electricity used exclusively by generative models. The report's separate discussion of cryptocurrency mining belongs to a different series and should not be silently added to or confused with the principal data center figures. (LBNL scenarios and methods)
Figure 10.2Data
National data-center electricity: 2023 estimate and 2028 scenarios
Lawrence Berkeley National Laboratory, December 2024. Model-based estimates, not a meter census. The 2028 band is a range across modeled futures, not a statistical interval, and says nothing about any specific grid.
10.2.1 What a model-based national estimate actually is#
Here's what the report actually did, in plain English. Nobody walked around the country reading every data center meter and adding them up. That census doesn't exist. Instead, the researchers built the total from parts: how many servers of each kind are out there, how hard they typically run, how much cooling and power conversion each kind of facility needs on top of the computing itself, and how those assumptions add up across the country. That's a respectable way to estimate a national quantity when the direct measurement isn't available. It also means the number inherits every assumption that went into it.
So when you see "176 TWh" or "4.4 percent," read it as "the best available modeled estimate of the national total for 2023." Don't read it as "the sum of every bill paid by every facility." The difference matters when someone uses the number to argue about precision it doesn't have.
10.2.2 Why a scenario range is not a confidence interval#
This is one of the most common misreadings I see, and it matters. A confidence interval comes from statistics: you measured a sample, and the interval describes uncertainty about the true value given that sample. A scenario range is something else. The researchers asked, in effect, "if the future goes this way, what's the total? If it goes that way, what's the total?" The low end and the high end are different stories about the future, each with its own assumptions about equipment, efficiency, and growth.
That means you can't say "the true 2028 value is probably somewhere in the middle." The range doesn't carry that meaning. And you certainly can't say "data centers will use 580 TWh in 2028," because that's picking one scenario and dropping the word scenario. Read that again: the high end is a scenario, not a forecast, and the low end is a scenario too.
"Generative models use 176 TWh." No. The total covers data centers broadly. It includes the whole range of computing those facilities do.
"Data centers use 12 percent of U.S. electricity." No. That's the top of a 2028 scenario range expressed as a share of projected consumption, under the report's assumptions. The 2023 estimate is about 4.4 percent.
"Add crypto and it's even bigger." Maybe, but the report treats cryptocurrency mining as a separate series. Adding the two together yourself, without the report's own method, creates a number the report never produced. I have no interest in hiding crypto load. I have an interest in not double-counting it or smuggling it into a different series.
10.2.4 What national scale can and cannot tell a county#
The right use of national evidence is modest: it establishes the importance of planning. A national total does not determine whether a specific feeder, transmission zone, utility, or community can accommodate a particular load. The question "can this grid serve this facility" is answered by interconnection studies, feeder data, and utility resource plans for that service territory. Those records exist at a different level of the system than the LBNL national model. Scale tells you to plan; it cannot tell you where.
If a speaker at a hearing cites the national total to argue that your county's grid is fine, or that your county's grid is doomed, the number is being asked to do something it can't do. The response is simple and polite: "That's a national model estimate. What do the interconnection study and the utility's resource plan for our service territory say?"
Before we look at bills, you need a working picture of how a large load gets connected to the grid and who pays for what. I'm going to explain this conceptually. No numbers here, because the numbers are always specific to a utility, a tariff, and a site, and a chapter that invented them would be doing exactly what I'm warning you about. What I can give you is the map of the machinery, from someone who has spent a lot of time inside it.
10.3.1 Interconnection: asking permission to plug in#
A large facility can't just run a wire to the nearest pole. The utility, and often the regional grid operator, has to study whether the system can carry the new load without overloading lines, transformers, or substations, and without degrading service to everyone else. That study is the interconnection process. It asks where the load connects, how much it draws, how fast it ramps, and what upgrades the system needs to serve it safely.
The output is a set of findings and, usually, a list of required upgrades. Some of those upgrades serve only the new customer, like a dedicated substation. Others strengthen shared parts of the system that other customers also use. That distinction, dedicated versus shared, is where most of the cost-allocation argument starts.
I have sat in enough utility interconnection meetings to know that the study is where reality shows up. The press release says a facility "will" be a certain size. The interconnection study says what the system can actually serve, when, and with what upgrades. When a community wants to know whether a project is real, the interconnection record is often more informative than the announcement.
10.3.2 Tariffs: the rulebook for what a customer pays#
A tariff is the published set of rates, terms, and conditions a utility uses to charge a class of customers, approved by the state utility commission. Residential customers have tariffs. Commercial and industrial customers have tariffs. Some states now have tariffs written specifically for very large new loads like data centers, as the Ohio example later in this chapter shows.
The tariff matters because it decides which costs the large customer pays directly and which costs get spread across the customer base through general rates. When a utility builds something to serve growth, the money has to come from somewhere. If the large customer's tariff and contract cover it, households don't. If they don't, some share can flow into the rates everyone pays. That's the whole game in rate cases: who carries which cost, and what happens if the load doesn't show up as promised.
Here's a risk most residents never hear about. A utility plans and sometimes builds for a load that a customer said it would bring. If that customer then uses much less than planned, delays for years, or cancels, the utility may be left with investments that were sized for demand that never arrived. Without protections, those stranded costs can end up in general rates.
Protections take a few forms: minimum payments tied to contracted capacity whether or not the customer uses it, collateral posted up front, exit fees for leaving early, and reimbursement if the customer cancels before the service is energized. None of these are exotic. They're how you make the party who created the risk carry it.
10.3.4 Curtailment: can the facility turn down when the grid is stressed#
Curtailment means reducing a load on request, usually when the grid is strained. Some facilities can shift or pause part of their computing; others run workloads that can't be interrupted. A claim that a facility "can curtail" is only meaningful if it says what the contractual right is, what loads are excluded as critical, whether the response has been tested, and what compensation, if any, applies.
This is a place where operating experience is directly relevant. A flexible load that actually drops when asked is a grid asset. A load that says it's flexible but has never been tested, or that excludes most of its consumption as critical, isn't the same thing. Ask for the contract and the test record.
10.3.5 Backup power: what runs when the main supply doesn't#
Every serious facility has a plan for when its primary supply fails or is down for maintenance. Sometimes that's on-site generators. Sometimes it's the grid itself. The backup plan affects the community in three ways: it may rely on public infrastructure at the worst possible moments, it may involve fuel deliveries and storage, and backup equipment gets tested on a schedule, which has noise and emissions consequences. A power story that skips backup is half a story.
10.3.6 Cost allocation: the question behind every bill question#
Put the pieces together and "do data centers raise electricity bills" turns into a sharper question: under the governing tariff and contracts, which costs of serving this load are paid by the large customer, which are spread across other customers, and who carries the risk if the load changes? That question can be answered with records. The general question can't.
10.4 Do Data Centers Raise Electricity Bills? What Virginia's Audit Found#
Virginia's legislative audit, the Joint Legislative Audit and Review Commission's 2024 study of data centers, is the most thorough state-level fiscal examination in the reviewed research. It found that existing rates appropriately allocated current costs at the time of its study, while also warning that future growth could increase system costs. These findings address different periods and different questions. They are not contradictory, and quoting one without the other misrepresents both. (JLARC data center report)
The report modeled a potential increase of roughly $14 to $37 per month in a typical Dominion residential bill by 2040 under its examined scenarios. Three limits apply. That is a modeled 2040 outcome under stated assumptions, not an observed 2026 bill increase. It is a Virginia-specific result from a Virginia-specific load forecast, not a nationwide figure. And it is a range across scenarios, not a prediction that any particular project produces any particular effect. A campaign flyer that converts the top of that range into a present-tense bill increase is manufacturing a fact the audit never produced. (JLARC scenario analysis)
Figure 10.7Data
Virginia's modeled 2040 bill range
Virginia-specific scenarios. JLARC also found rates appropriately allocated current costs at the time of its study. Both findings belong in the same sentence.
10.4.1 Why "fair now" and "risk later" are both true#
People on both sides of a hearing like to quote half of this audit. The supporter says, "Virginia's own auditors found the rates were fair." The opponent says, "Virginia's own auditors found bills could go up." Both are quoting accurately, and both are misrepresenting the report.
Here's the plain reading. At the time of the study, the cost of serving existing data centers was being allocated appropriately through existing rates. That's a finding about the present, about costs that had already been incurred and a customer base that already existed. Separately, the auditors looked ahead and said growth could raise system costs. That's a finding about a future that depends on how much load arrives, what has to be built to serve it, and how the rules allocate those new costs. One statement is a snapshot. The other is a warning about a trajectory. You need both to understand the report.
Take the three limits one at a time, because each one gets dropped in public argument.
Modeled, 2040, under stated assumptions. It isn't a bill anyone received. It's what the model produced for a future year given assumptions about load growth and the investments needed to serve it.
Virginia-specific. It comes from Virginia's load forecast and a typical Dominion residential bill. Your state has a different utility, different tariffs, different growth, and different generation. The number doesn't travel.
A range across scenarios. The low and high ends are different scenarios. Nobody should present the high end as the expected outcome, or the low end as the guaranteed outcome.
A number without a denominator is a rumor. In this case the denominator is Virginia's system, Dominion's typical residential customer, the year 2040, and the auditors' scenarios. Strip those away and what's left is a rumor with a dollar sign.
10.4.3 What a resident should take from the Virginia audit#
The useful lesson for a resident anywhere isn't the dollar range. It's the structure. A state with serious data center growth looked carefully and found that the present allocation was fair while the future carried real risk. That tells you where to look in your own state: not at whether today's rates are fair in the abstract, but at whether the rules for new large loads protect households from the cost of future growth. That's a tariff and contract question, which is exactly what the Ohio case addresses.
10.5 The AEP Ohio Data Center Tariff: Shifting Risk to the Large Customer#
AEP Ohio provides the documented counterexample on the allocation side. The Public Utilities Commission of Ohio approved its data center tariff settlement in July 2025, with the utility listing an effective date of July 23, 2025. The published process includes minimum contract-capacity ramps, a term extending beyond the ramp period, potential collateral and exit obligations, and reimbursement requirements for specified cancellations or delays before energization. (Commission announcement, utility tariff process)
The accurate label for that design is this: a documented attempt to shift underuse and cancellation risk toward the large customer rather than leaving it with residential ratepayers by default. It is not a measured finding that all ratepayer risk has been eliminated. The ramp, collateral, and exit terms address specific risks under specific contracts, and residual risk depends on enforcement, counterparties, and facts that only time and dockets will produce.
Figure 10.8Framework
The AEP Ohio data-center tariff: how risk is allocated
PUCO approved the settlement in July 2025; the utility lists July 23, 2025 as effective. A documented attempt to shift underuse and cancellation risk toward the large customer, not proof that all ratepayer risk is gone. Check the governing tariff and orders before relying on this summary.
10.5.1 What each tariff element does, in plain English#
Each piece of the Ohio structure answers one of the risks I described in Section 10.3.
Minimum contract-capacity ramps. The customer commits to paying for a minimum share of the capacity it asked for, stepping up over a ramp period. If it uses less, it still pays the minimum. That protects other customers from paying for capacity reserved for a load that doesn't arrive.
A term extending beyond the ramp. The commitment doesn't end the moment the ramp finishes. It runs long enough to matter against the life of the investments made to serve the customer.
Collateral. Where required, the customer posts security up front, so if it fails financially, there's something to draw on besides the rate base.
Exit obligations. Leaving early has a cost, which discourages speculative requests and compensates the system when someone walks away.
Reimbursement before energization. If the customer cancels or delays in specified ways before it's even connected, it reimburses costs the utility incurred on its behalf.
Put together, the structure says: the party that creates the planning risk carries it. That's the right principle. Whether it works in practice depends on the contracts, the enforcement, and the financial strength of the counterparties.
It doesn't prove that no Ohio household will ever pay any cost associated with data center growth. It doesn't prove the terms will be enforced as written in every case. It doesn't prove that the same design would work, or be approved, in another state. And it doesn't tell you anything about water, noise, jobs, or taxes. It's one well-documented tool for one category of risk.
10.5.3 A summary page is a map, not the territory#
One more discipline applies to every utility webpage cited in this report: a project decision requires review of the governing tariff, applicable commission orders, executed agreements, and current litigation status. A summary page is a map, not the territory. I cite the utility page and the commission announcement because they document that the structure exists. If you're a commissioner or a regulator relying on it, pull the actual tariff sheets and orders, and check whether anything has changed since.
10.5.4 What this means for a resident in another state#
If your state doesn't have a large-load tariff like this, that's not proof that households are exposed. Existing tariffs and contracts may already carry some of these protections. But it's a question worth asking out loud: "If this customer uses less than it reserved, or cancels, who pays for what was built to serve it?" If nobody in the room can answer with a document, that's your answer for now.
10.6 Behind the Meter Power and "Not Taking Community Power"#
Community meetings about data centers generate a phrase I propose to retire: "the facility will not take community power." It is not evaluable as stated, and unevaluable claims resolve in favor of whoever makes them. Replace it with bounded questions and the evidence each requires.
Question
Evidence required
What serves the facility in normal operation?
Actual supply arrangement, interconnection configuration, load profile, and delivery obligations
What happens during outages or maintenance?
Backup fuel, grid dependence, restoration arrangements, and tested operating procedures
Who pays for upgrades?
Effective tariff, executed agreements, cost-allocation analysis, and residual-cost treatment
Who bears underuse or cancellation risk?
Minimum payments, collateral, termination terms, credit support, and enforceability
Can the facility curtail?
Contractual rights, tested response, excluded critical loads, and compensation
What is verified after operation begins?
Metered load, outages, compliance, and public reporting against commitments
Figure 10.3Framework
Retiring "not taking community power": six questions that replace it
A behind-the-meter design must still answer all six. Savrn builds around behind-the-meter power; that is a design claim this test applies to first.
Every row in that table maps to a way the slogan can be true on paper and false in practice.
Normal operation. A facility can be served mostly by one source and still depend on another. The supply arrangement and interconnection configuration tell you what's physically connected and what's contractually promised. The load profile tells you how much and when.
Outages and maintenance. This is where "not taking community power" most often breaks. If the backup supply is the grid, then at exactly the moments when the primary source is down, the facility leans on the public system. That may be perfectly acceptable, but it has to be stated, planned for, and paid for.
Upgrades. Even a facility with its own generation may require utility work: a connection for backup, a substation modification, a transmission change. Who pays for that work is a tariff and contract question, not a slogan question.
Underuse and cancellation. If the utility builds anything in reliance on this customer, what happens if the customer shrinks or leaves? Section 10.5 shows one way to answer that.
Curtailment. If the facility draws from the grid at all, can it reduce that draw when the grid is stressed? Is that right written down, has it been tested, and which loads are exempt?
Verification. Promises made before construction need to be checked after operation. Metered load, recorded outages, compliance reports, and public reporting against the original commitments are how anyone finds out whether the power story was true.
One configuration deserves its own caution. A behind-the-meter design, generation on site feeding the facility directly, is not, by its label alone, proof of independence from public infrastructure. The claim must include backup operation, fuel supply, interconnection arrangements, emissions, and the other dependencies relevant to the actual configuration. A facility whose backup is the grid, whose fuel arrives by public road, and whose interconnection requires utility coordination is not independent of public infrastructure because its primary supply is on site. Labels are not evidence.
Let me say what "behind the meter" means in plain terms. Normally, power flows from the utility system, through a meter, into a customer. Behind-the-meter generation sits on the customer's side of that meter, producing power used on site. It can reduce what the customer draws from the grid, sometimes to very little in normal operation. What it doesn't automatically do is remove the facility from every public system. Fuel may travel by public road or pipeline. Emissions go into the shared air. Backup may come from the grid. And the connection, if there is one, still needs coordination with the utility.
That's why a behind-the-meter power data center has to answer all six questions just like any other facility. The answers may be different, and may be better for the community on some rows. But "behind the meter" is a configuration, not a verdict.
Here's a hypothetical, with no real site or numbers. A developer tells a county board its facility will run on on-site generation and "won't take community power." A commissioner walks the six rows.
Normal operation: the developer shows the on-site generation design and a connection to the utility. The commissioner asks what that connection is for and how much it can carry.
Outages: the developer says the grid provides backup. Now the commissioner knows the facility does depend on community power at specific times, and asks how often, how much, and under what terms.
Upgrades: the connection requires some utility work. Who pays is in the tariff and agreement, which the commissioner asks to see.
Underuse: if the utility builds anything for the backup connection, the commissioner asks what protects other customers if the facility leaves.
Curtailment: if the facility draws from the grid during outages, can it limit that draw during system emergencies?
Verification: what gets reported publicly after operation, and how often?
None of those answers is a trap. They're the difference between a slogan and a record. By the end, the commissioner can say something accurate like "this facility is designed to meet its normal load on site, relies on the grid for backup under these terms, and pays for these upgrades." That's a sentence a community can decide on.
10.7 Data Center Water Use: Define the Numerator First#
LBNL estimated approximately 66 billion liters of direct United States data center water consumption in 2023, and nearly 800 billion liters of indirect consumption associated with electricity supply. Those two categories describe different locations and different system boundaries. Direct use is water the facility withdraws for cooling and operations, while indirect use is water consumed upstream in the generation of the electricity the facility consumes. Neither figure is a site-specific permit reading or a meter total, and the two should never be added, swapped, or cherry-picked to suit a conclusion. (LBNL water analysis)
The Congressional Research Service's July 2026 review states that the United States Geological Survey has not conducted a systematic national assessment isolating data center water use. That is an important admission about the state of the evidence: available datasets and estimates should not be presented as a complete, reconciled inventory, because no complete, reconciled inventory exists at the national level. (CRS water report)
Figure 10.4Data
Define the numerator: direct versus indirect water, U.S. data centers, 2023
LBNL national estimates. Neither is a permit reading or meter total. Project comparisons must separate withdrawal from consumption, potable from nonpotable, actual from permitted, and normal from drought operation.
Direct water is the water the building itself uses: mostly cooling, plus other operations. If a facility uses evaporative cooling, some of that water leaves as vapor and doesn't come back to the local source. That's the 66 billion liter category at the national level.
Indirect water is upstream. Many power plants use water for their own cooling, and some of that water is consumed. When a facility uses electricity, part of that upstream water consumption is attributable to it. That's the nearly 800 billion liter category. It happens wherever the electricity is generated, which may be far from the facility.
Here's why the boundary matters. A resident worried about their town's wells is asking mostly about direct use, and specifically about the local source. A policymaker worried about a river basin that supplies several power plants may care a great deal about indirect use. Both questions are legitimate. They're different questions, and they have different numerators.
Adding them. Direct plus indirect gives you a number that mixes two locations and two system boundaries. It might be useful in a carefully defined total-footprint analysis. Dropped into a local debate, it implies all that water comes from the host community, which it doesn't.
Swapping them. Using the indirect figure when talking about a town's water supply, or using only the direct figure to argue the whole footprint is small.
Cherry-picking them. Quoting whichever number suits the conclusion and ignoring the other.
Show me the denominator. And while you're at it, show me the numerator's boundary.
10.7.3 The vocabulary you need at a water hearing#
The appropriate project comparison has to be built record by record, distinguishing:
Withdrawal versus consumption. Withdrawal is water taken from a source. Consumption is water not returned. A facility can withdraw a lot and return most of it, or withdraw less and consume most of it.
Discharge quantity and quality. What goes back, how much, and in what condition.
Potable versus nonpotable. Drinking-quality water competes directly with household supply. Nonpotable water may not.
Direct versus upstream. On site versus at the power plant.
Actual versus permitted. A permit sets what's allowed. Actual use is what's measured. They're rarely the same.
Normal versus drought operation. A facility's water needs on a mild day and on the hottest week of a dry year may be very different, and the second is the one a community worries about.
A lower onsite cooling requirement does not itself answer questions about electricity-related water, construction water, domestic use, emergency operations, or local supply constraints. It answers one question about one design element.
10.7.4 Five common water claims and their minimum boundaries#
I propose minimum boundaries for the claims most often heard in public debate:
Proposed claim
Minimum acceptable boundary
"No water cooling"
Specify the cooling process and whether other facility water uses remain
"Closed loop"
Identify initial fill, makeup, maintenance losses, blowdown if applicable, and heat-rejection method
"Uses reclaimed water"
Identify source, competing uses, treatment requirements, supply reliability, and discharge
"Low water intensity"
State denominator, period, load, weather, and whether the figure is measured or designed
"No impact on residents"
Requires local availability, quality, rate, and reliability evidence; a cooling specification is insufficient
A few notes on reading that table.
"No water cooling" can be true of the cooling process and still leave restrooms, landscaping, cleaning, fire systems, and construction water. The claim has to say which.
"Closed loop" sounds like zero water, and it isn't automatically. A loop has to be filled once. It may need makeup water to replace losses. Maintenance drains and refills parts of it. Some systems bleed off water to control mineral buildup, which is called blowdown. And the heat has to go somewhere, so the heat-rejection method matters: if the loop's heat is rejected through an evaporative tower, the loop is closed but the facility still consumes water.
"Uses reclaimed water" is often good news, but reclaimed water may already have other users, may need treatment, and may not be reliable in every season.
"Low water intensity" is a ratio, and a ratio without a stated denominator is meaningless. Per what? Over what period? At what load? In what weather? Measured, or designed?
"No impact on residents" is the biggest claim, and no cooling specification can support it on its own. It needs local evidence on supply, quality, rates, and reliability.
LBNL's own analysis makes the final point: there are trade-offs among cooling choices, including the relationship between onsite water use and energy requirements. A design assessed across only one resource can improve its headline number while worsening its total footprint. An air-cooled design that raises electricity demand raises indirect water consumption too. Success declared by minimizing one reported number is not success. (LBNL cooling analysis)
This one lands close to home for me, and I'll come back to it in the Savrn section. A zero-makeup-water goal addresses direct water. It doesn't erase the electricity question, and the electricity question carries its own water footprint depending on how the power is generated. Any credible water claim names both.
10.7.6 What this means for a resident worried about their well#
Your concern is local and specific, so your questions should be too. Where does the facility's water come from? Is it the same aquifer or utility that serves your home? Is it potable or nonpotable? How much is permitted, and how much will be reported as actually used? What happens in a drought? What gets discharged, and where? If the answers are a cooling brochure, you haven't been answered yet. National numbers, including the ones in this section, can't tell you about your well.
Virginia's JLARC report describes a typical 250,000-square-foot data center as employing about 50 full-time workers in ongoing operations, approximately half of them contractors, while construction can involve a peak workforce around 1,500 over a roughly 12 to 18 month build period. Those are different categories across different time periods, and treating them as interchangeable is the most common job-count error in this debate. A 1,500-person construction peak and a 50-person operating staff are both real, and neither is the other. (JLARC employment findings)
The same report uses broader economic modeling that includes indirect and induced activity: supported jobs in supply chains and local spending. Such estimates answer a legitimate economic question about regional activity, but they should never be represented as the number of people permanently employed inside facilities. A multiplier is a model output, not a headcount, and a community deciding whether a project serves it needs to know which number it is looking at.
Figure 10.5Data
Jobs by phase and denominator
Virginia JLARC figures for a typical facility. Whether host-neighborhood residents get any of these jobs is a separate distribution question.
10.8.1 Three different job numbers, three different questions#
Construction peak. How many people are on site at the busiest point of the build. JLARC's figure for a typical facility is around 1,500 over roughly 12 to 18 months. It's temporary by definition, and the peak is a peak, not an average.
Ongoing operations. How many people work the facility once it's running. JLARC's figure is about 50 full-time workers, about half of them contractors. That's the permanent number, and it includes people who aren't the operator's own employees.
Indirect and induced. Jobs supported elsewhere in the economy through suppliers and local spending. That's a modeled estimate of regional activity, not people inside the building.
When you build at industrial scale, the construction phase and the operating phase are different worlds. Construction brings trades and trucks. Operations are a much smaller, steadier crew. Both matter to a town. Neither should be sold as the other. That's operator experience, and it matches what the Virginia audit describes; it's not a substitute for the audit.
10.8.2 What the local employment record should contain#
The proposed local employment record states construction person-hours or job-years, peak headcount, ongoing employees, ongoing contractors, wage and qualification information, local-resident participation, and the period observed. Training seats and job announcements stay separate from completed training and actual employment. A commitment to train is a commitment, not a workforce.
This is also a distribution question, and I insist on it. A project can create regional employment without employing many residents of the host neighborhood. That possibility should be evaluated and disclosed, not hidden by an aggregate number that averages the region over the host community. The people who hear the construction traffic and live near the substation are entitled to know whether the jobs are theirs or the next county's.
"This project will create 1,500 jobs." Maybe at the construction peak, for a limited period, if the project resembles JLARC's typical facility. Not permanently.
"Data centers only create 50 jobs, so they're worthless." The ongoing number is small relative to construction and to the size of the building. It isn't zero, and the fiscal and other effects still have to be weighed on their own records.
"The study says thousands of jobs are supported." Supported jobs from economic modeling are real estimates of regional activity. They aren't headcount inside the facility, and they aren't necessarily local.
10.9 Data Center Tax Revenue: Gross Revenue Is Not Net Benefit#
JLARC documents substantial variation in the share of local revenue associated with data centers among mature host localities. There is no typical answer to "what does a data center contribute to a local budget," because the answer depends on the locality's tax structure, land values, and incentive arrangements. The same audit documents a significant sales-and-use-tax exemption, which illustrates why tax receipts and incentives belong in the same fiscal analysis: a jurisdiction that exempts equipment from sales tax and then counts the property-tax revenue as pure benefit has counted one side of a ledger with two sides. (JLARC fiscal findings)
The proposed fiscal model separates recurring revenues, one-time receipts, abatements, exemptions, public infrastructure costs, service costs, financing obligations, asset depreciation, and closure or redevelopment costs. It also identifies which jurisdiction receives each benefit and bears each obligation. That last item is where fiscal analysis most often quietly fails: the revenue lands in one government's books, the road and water upgrades land in another's, and no single ledger shows the project whole.
Here's what each line means in practice:
Recurring revenues: taxes paid year after year, such as property taxes on land, buildings, and equipment where they apply.
One-time receipts: fees and payments tied to permitting or construction that don't repeat.
Abatements and exemptions: revenue the jurisdiction or state chose not to collect as an incentive. These belong on the same page as the revenue.
Public infrastructure costs: roads, water and sewer extensions, and other public works needed to serve the site.
Service costs: fire, emergency response, inspection, and other services the facility will use.
Financing obligations: any public borrowing tied to the project.
Asset depreciation: equipment taxed on value loses value over time, and so does the revenue it produces.
Closure or redevelopment: what happens to the site, and who pays, if the facility shuts down.
First, an increase in the local tax base does not automatically reduce a household's tax bill. That conclusion requires the actual budget, assessment rules, tax rates, spending choices, and incidence analysis, and it sometimes comes out the other way. A county might use new revenue to lower rates, expand services, pay down debt, or some combination. Only the budget decisions determine what a household sees.
Second, scenario analysis should include delayed construction, lower realized load, changes in equipment value, tenant failure, and early closure. Those scenarios are risk tests, not predictions that a project will fail. A fiscal model that only runs the developer's case is a brochure with decimal points.
Here's a hypothetical, with no real numbers. A county is presented with a revenue projection for a proposed facility. The finance director asks five questions. What incentives or exemptions apply, at the state and local level, and what is their value against the projected revenue? Which jurisdiction receives the revenue, and which pays for the road and water work? How does projected revenue change as equipment depreciates? What does the projection look like if construction slips, if the load comes in lower, or if the facility closes early? And what does the county intend to do with any net revenue? A presentation that can answer all five is a fiscal analysis. One that can't is a sales document.
The Virginia audit identifies noise and proximity to residential areas as material local issues and discusses limitations in existing approaches to regulating them. Regional emissions totals, however well constructed, do not by themselves determine exposure at a particular property, which depends on stack location, terrain, operating hours, and what stands between the source and the receiver. (JLARC community and environmental findings)
The proposed operating record includes measurement location, frequency characteristics, duration, operating mode, background conditions, and applicable limits. Generator testing, emergency operation, maintenance, and nighttime conditions belong in the record when relevant. Evaluating only a favorable daytime demonstration is a measurement choice, and it should be labeled as one.
In plain terms: where was the meter, what kind of sound was measured (a low hum travels and annoys differently from a higher-pitched whine), for how long, with the facility doing what, compared to what background, and against what limit? Noise that's fine at noon on a weekday can be a different story at night when the background drops. Anyone who has lived near large mechanical equipment knows that. I've built around it. The record should show the hard conditions, not only the easy ones.
A regional emissions total is useful for regional planning. It doesn't tell a family whether the air at their property is affected. Exposure depends on where the stacks are, the terrain, when equipment runs, and what's between the source and the house. That's why backup generation and its testing schedule belong in the community's record, and why on-site generation, including behind-the-meter designs, needs a local emissions answer and not only a regional one.
Land-use review should cover construction traffic, drainage, setbacks, emergency access, and the compatibility of actual operations with nearby uses. And a permit establishes a regulatory status within its scope; it does not demonstrate all future operating conditions, and it does not replace continuing compliance measurement. The document that authorizes operation is the beginning of the accountability record, not the end of it.
10.11 A Community Benefit Compact That Can Change a Project#
The following compact is proposed as an implementation mechanism. It is not presented as an existing Savrn agreement or a proven universal template. It's what a credible accountability mechanism would have to contain, derived from the failure modes documented throughout this chapter.
Baseline: Publish the initial resource, fiscal, and community conditions against which commitments will be measured.
Obligations: Define each power, water, noise, employment, tax, and reporting commitment with a responsible entity and deadline.
Verification: Identify measurement methods, access to records, reviewer qualifications, and treatment of confidential information.
Remedy: Specify notice, correction periods, escalation, and enforceable consequences for material failures.
Participation: Provide accessible reporting and a process for residents to challenge a measurement or interpretation.
Change control: Require reassessment when load, cooling, generation, ownership, or operating conditions materially change.
End of life: Allocate decommissioning, restoration, and unresolved public obligations before operation begins.
Figure 10.6Framework
The proposed community benefit compact
Proposed, not an existing Savrn agreement. Measurement without consequence does not meet the standard.
One requirement does the real work: the compact must allow an adverse result to change the project. A reporting process that cannot alter behavior, such as a dashboard that accumulates violations no one acts on, is insufficient for the accountability standard this report proposes. Measurement without consequence is surveillance of the community, by the community, for no one.
Baseline comes first because you can't measure a change without knowing the starting point. Water levels, noise at night, traffic, local revenue, and service costs all need a before picture.
Obligations turn promises into commitments with names and dates. "We'll be a good neighbor" isn't an obligation. "The operator will report measured nighttime sound at these locations each quarter" is.
Verification answers who checks, how, and with what access. Confidential business information is real, and the compact should say how it's handled rather than using it as a reason to report nothing.
Remedy is where most agreements go soft. Notice, a correction period, escalation, and consequences that actually apply. Without it, the other elements are paperwork.
Participation gives residents a way to challenge a number or an interpretation. Communities often notice problems before any report does.
Change control matters because projects change after approval. More load, a different cooling design, new generation, or a new owner can each change the community's exposure. The compact should require a fresh look when that happens.
End of life is the element everyone skips, because it feels far away. It isn't. Deciding before operation who pays to decommission and restore a site is far easier than deciding after a company has left.
10.12 The Savrn Test: Holding Our Own Claims to This Standard#
Now I turn the chapter on my own company. Savrn designs, builds, and delivers AI factories, purpose-built data centers. I've told you already that the community doesn't care what you call the building. So I won't hide behind the name. Everything in this chapter applies to Savrn at full strength.
Savrn is built around two design claims that matter here: behind-the-meter power, and closed-loop cooling with a zero-makeup-water design goal. Those are publisher statements. They are commitments to be verified, not achievements. Savrn has not achieved any metric in this chapter, and I won't write a sentence implying otherwise.
10.12.1 Behind-the-meter power must answer all six questions#
Section 10.6 says a behind-the-meter label is not proof of independence from public infrastructure. That sentence was written for everyone, and it's written for us first. For any Savrn facility, a community should be able to get documented answers to each of the six:
Normal operation: the actual supply arrangement, the interconnection configuration if any, the load profile, and delivery obligations.
Outages and maintenance: what serves the load when on-site generation is down, what fuel backs it up, whether the grid is involved, and whether those procedures have been tested.
Upgrades: any utility work required, and who pays under the governing tariff and agreements.
Underuse and cancellation: what protects other customers if anything is built in reliance on us and we use less or leave.
Curtailment: if we draw on the grid at any point, what the contractual right to curtail is, how it's been tested, and which loads are excluded.
Verification: what we report publicly after operation begins about metered load, outages, and compliance against commitments.
The accurate version of our claim is not "Savrn won't take community power." It's "Savrn is designed to meet its load behind the meter, and here are the documented answers to all six questions." If we can't produce those answers for a specific site, the community should treat our claim as unverified.
The behind-the-meter caution names other dependencies too: fuel supply and emissions. On-site generation needs fuel, and fuel moves on public infrastructure. On-site generation produces emissions into shared air. Those belong in our record, measured at the locations that matter to neighbors, as Section 10.10 describes.
10.12.2 Zero makeup water must pass the table's boundaries#
Our water goal touches two rows of the five-claim table directly.
"Closed loop." The minimum boundary is to identify initial fill, makeup, maintenance losses, blowdown if applicable, and heat-rejection method. A zero-makeup-water goal is a claim about one of those items. It still has to state the initial fill, how maintenance losses are handled, whether blowdown applies, and how the heat leaves the system. If the heat-rejection method consumed water, "closed loop" would be true of the loop and misleading about the facility. We have to show the whole path.
"No water cooling." The minimum boundary is to specify the cooling process and whether other facility water uses remain. Even a facility that uses no water for cooling has people, restrooms, cleaning, fire protection, and construction. Those uses have to be named, not assumed away.
And then LBNL's trade-off applies to us with full force. A design that reduces onsite water can raise electricity requirements, and electricity has its own water footprint depending on how it's generated. A zero-makeup-water goal is a direct-water commitment. It does not by itself answer the indirect-water question, and I won't let it be used to imply that it does. If we ever reported only the favorable number, you should hold us to the sentence I've already written: success declared by minimizing one reported number is not success.
Finally, "no impact on residents" is a claim I won't make from a cooling design. That row of the table requires local availability, quality, rate, and reliability evidence. A cooling specification is insufficient, including ours.
10.12.3 The other rows: jobs, taxes, noise, compact#
Our design claims are about power and water, but the chapter's other tests apply too. Job statements should separate construction peak, ongoing employees, ongoing contractors, and modeled supported jobs, and should disclose local-resident participation. Fiscal statements should put incentives on the same page as revenue and name which jurisdiction bears which cost. Noise records should include nighttime and testing conditions. And any community agreement we propose should contain all seven compact elements, including a remedy that lets an adverse result change what we do.
10.13 A Resident's Field Guide for the Public Hearing#
You don't need an engineering degree to ask good questions. You need the right questions and the right documents. Here's a field guide you can print and bring.
Power
- What serves this facility in normal operation, and is there any connection to the utility system?
- What serves it during outages and maintenance? Is the grid the backup?
- Who pays for any utility upgrades, and under which tariff?
- If the facility uses less power than it reserved, or cancels, who pays for what was built?
- Can the facility reduce its grid draw during emergencies, and has that been tested?
Water
- Where does the water come from, and is it the same source that serves homes?
- Is it potable or nonpotable?
- How much is withdrawn, how much is consumed, and how much is discharged, and where?
- What's the permitted amount, and will actual use be reported?
- What happens in a drought?
- If the claim is "closed loop," what are the initial fill, makeup, maintenance losses, blowdown, and heat-rejection method?
Jobs
- What's the construction peak, and for how long?
- How many ongoing employees and ongoing contractors?
- How many are expected to be local residents, and how will that be reported?
Taxes
- What incentives, abatements, or exemptions apply?
- Which government gets the revenue, and which pays for roads, water, and services?
- What happens to revenue as equipment depreciates, and if the facility closes early?
Noise, air, land
- Where will sound be measured, and does that include nighttime and generator testing?
- What backup generation is on site, how often is it tested, and what are the emissions at nearby homes?
- What are the setbacks, drainage plans, and construction traffic routes?
Accountability
- Is there a written agreement with a baseline, verification, and remedies?
- Can residents challenge a measurement?
- What happens if the project changes after approval, or closes?
When you hear "we won't take community power," "we use no water," or "this project brings 1,500 jobs," don't argue the slogan. Ask for the row. "Which of the six power questions does that answer, and where's the document?" "Is that direct or indirect water, and withdrawal or consumption?" "Is that construction peak or ongoing?" A calm request for the record is more powerful than any counterclaim, because it can't be dismissed as opinion.
Don't bring the national totals as proof about your county. Don't present the Virginia bill range as your future bill. Don't treat a missing document as proof of wrongdoing; treat it as a gap to be filled before a decision. And don't let anyone, developer or opponent, merge the application argument into the facility argument. Your meter, your water, your nights, and your taxes are enough of an agenda.
10.14 For County Commissioners and Utility Regulators#
You're the people who turn questions into conditions. That's a heavier job than asking.
Put obligations in writing. Every power, water, noise, employment, tax, and reporting commitment you rely on in approval should appear in an enforceable document with a responsible entity and a deadline.
Require a baseline before construction. Without it, you can't evaluate any later claim.
Build the whole fiscal ledger. Revenue, incentives, infrastructure, services, depreciation, and closure costs, with the jurisdiction that bears each one. Run the delay, low-load, and early-closure scenarios.
Separate job categories in the record. Construction peak, ongoing employees, ongoing contractors, and modeled supported jobs, plus local-resident participation.
Condition change control. Material changes in load, cooling, generation, or ownership should trigger reassessment.
Allocate end of life up front. Decommissioning and restoration obligations should be settled before operation.
Keep your options open. On the record in front of you, you can approve, modify, defer, or reject. Deferral to obtain missing records is a legitimate decision.
Separate present allocation from future risk. The Virginia audit shows both can be true at once. Examine whether the rules for new large loads protect other customers from growth-driven costs.
Look at the risk-allocation tools. The AEP Ohio tariff documents minimum capacity ramps, term, collateral, exit obligations, and reimbursement before energization. Whether a similar structure fits your jurisdiction is your call, based on your own record.
Test behind-the-meter claims. Ask what the grid provides during outages, what interconnection is required, and who pays for it.
Make curtailment real. If flexibility is claimed, ask for contractual terms, tests, and excluded loads.
Require post-operation verification. Metered load and outage reporting against commitments is how a commission learns whether its assumptions held.
One boundary from Chapter 6 belongs here. Home energy report programs can change household electricity use through social-comparison feedback, but no household conservation program converts an unmeasured industrial obligation into a community benefit. Families should not be asked to offset an industrial load through conservation, and the two evidence chains do not substitute for each other in either direction.
10.15 How to Read a Data Center Announcement in Five Minutes#
Announcements come with big numbers. Here's a five-minute read that tells you what the announcement actually establishes.
Minute one: What stage is this? Is it an idea, a land purchase, a filed application, an approved permit, a signed utility agreement, construction, or operation? An announcement is not a project. A permit is not construction.
Minute two: What's the power number, and what kind is it? Is it a requested load, a contracted capacity, or a measured draw? Is it all behind the meter, all grid, or a mix? Does it mention backup?
Minute three: What's the water claim, and what's its boundary? "No water cooling," "closed loop," "reclaimed," "low intensity"? Run it against the five-row table. Is it direct only?
Minute four: What's the jobs and money claim? Is it construction peak, ongoing, or modeled supported jobs? Is the investment figure a dollar spent locally, or a headline capital commitment? Are incentives mentioned?
Minute five: What's enforceable? Is there any commitment with a deadline, a verification method, and a remedy? Or is it all "will" and "expects"?
At the end of five minutes, write one sentence: "This announcement establishes X and does not yet establish Y." That's the most useful sentence you can bring to a neighbor or a meeting.
The seven Savrn trackers organize relevant categories of public evidence about this infrastructure: capital commitments, documented delays, moratorium actions, water-related disclosures, grid-operator constraints, permits, and scarcity indicators. They do not replace local engineering, utility, environmental, employment, or fiscal records. Their own published scopes distinguish disclosures, documented status, permits, modeled indicators, and measurement boundaries, categories with different evidentiary weight that the trackers themselves do not collapse. (Capital Atlas, Delay Watchlist, Moratorium Tracker, Water Tracker, Grid Watchlist, Permits Tracker, Scarcity Tracker)
The rule is the same everywhere in this report: a tracker entry starts a question and never ends one. Chapter 11 states a forbidden inference for each tracker. The ones most relevant here:
Capital Atlas: a dollar announced is not a dollar spent locally, a tax benefit to residents, a permanent job, or an investment return until the corresponding record shows it.
Delay Watchlist: an entry does not prove failure, establish its cause, or predict permanent cancellation.
Moratorium Tracker: a proposal is not law, a pause is not permanent, and a measure in one jurisdiction does not govern another.
Water Tracker: a cooling figure does not prove zero total use, a permit does not equal actual consumption, and a regional total does not establish local household harm.
Grid Operator Watchlist: queue position does not prove energization, statewide capacity does not prove local deliverability, and one tariff does not apply to every project.
Permits Tracker: a permit does not equal construction, a load request does not equal contracted demand, and planned generation is not available capacity.
Scarcity Tracker: a modeled index is not realized revenue, a liquid market price, a guaranteed financing basis, or proof that residents benefit.
A responsible public claim about a facility should be reproducible from original records and should identify its observation date, because these records change. And a missing record is an evidence gap, not proof of harmlessness and not proof of wrongdoing. The asymmetry matters in both directions: absence of a permit dispute does not certify a clean project, and absence of a clean audit does not prove malfeasance. Gaps stay gaps until someone fills them with records.
10.17 Conclusion: Real Numbers on Data Center Impact on Communities#
National energy and water research establishes the scale of the planning questions: 176 TWh and 4.4 percent of national electricity in 2023 as a model-based estimate, with 2028 scenarios running roughly 325 to 580 TWh; about 66 billion liters of direct water consumption with a much larger indirect footprint, and no completed national inventory. The Virginia audit and the Ohio tariff illustrate local effects and possible allocation mechanisms: a $14 to $37 modeled monthly bill range by 2040 under Dominion scenarios, a typical facility's 50 ongoing jobs beside its 1,500-person construction peak, material variation in local fiscal effects, and a contractual architecture for putting underuse risk on the large customer. None of these substitutes for the actual conditions of a proposed facility. (LBNL, JLARC, AEP Ohio)
Savrn's stated community objective becomes credible through measured, bounded, enforceable commitments: the compact of Section 10.11, tested against baselines, with remedies that bite. That includes our own behind-the-meter and zero-makeup-water goals, which stand as commitments until records verify them. The standard I propose is not to win an argument against residents. It's to make a project's benefits, burdens, alternatives, and obligations visible enough for a legitimate decision. That's the same standard Chapter 7 proposed for parks and drainage, applied here to the facilities this report's subject depends on. The application benefits and the facility effects stay separate. Both deserve real numbers.
It depends on the tariff and contracts. Virginia's JLARC found existing rates appropriately allocated current costs at the time of its 2024 study, but warned future growth could raise system costs and modeled a possible $14 to $37 monthly increase in a typical Dominion residential bill by 2040. That is a Virginia scenario range, not an observed bill increase and not a national figure.
How much electricity do data centers use in the US?
Lawrence Berkeley National Laboratory estimated about 176 TWh in 2023, roughly 4.4% of national electricity, using a model rather than a meter census. Its 2028 scenarios run about 325 to 580 TWh, or 6.7 to 12% of projected consumption. Those are scenarios, not a confidence interval, and cover data centers broadly, not only generative models.
How much water does a data center use?
No single answer exists for every facility. LBNL estimated about 66 billion liters of direct U.S. data center water consumption in 2023 and nearly 800 billion liters of indirect consumption tied to electricity generation. These are national modeled figures at different boundaries, and CRS reports that USGS has not systematically assessed data center water use nationally.
How many jobs does a data center create?
Virginia's JLARC describes a typical 250,000-square-foot data center employing about 50 full-time workers in ongoing operations, roughly half of them contractors. Construction can peak around 1,500 workers over roughly 12 to 18 months. Modeled indirect and induced jobs are regional estimates, not people working inside the facility, and none of these show how many jobs go to local residents.
What is behind the meter power for a data center?
Behind-the-meter power is generation on site that feeds the facility directly, on the customer's side of the utility meter. It can reduce grid draw, but the label alone does not prove independence from public infrastructure. Backup supply, fuel delivery, emissions, interconnection, and who pays for upgrades must still be documented for the actual configuration.
What is the AEP Ohio data center tariff?
It is a data center tariff settlement approved by the Public Utilities Commission of Ohio in July 2025, effective July 23, 2025 according to the utility. It includes minimum contract-capacity ramps, a term beyond the ramp, potential collateral and exit obligations, and reimbursement for specified cancellations or delays before energization. It shifts risk toward large customers but is not proof that ratepayer risk is eliminated.
Do data centers pay a lot in local taxes?
JLARC found substantial variation in the share of local revenue from data centers among mature host localities, so there is no typical answer. The same audit documents a significant sales-and-use-tax exemption. Revenue must be weighed against incentives, infrastructure and service costs, depreciation, and closure risk, and a larger tax base does not automatically lower household taxes.
What should residents ask at a data center public hearing?
Ask what powers the facility normally and during outages, who pays for upgrades and bears cancellation risk, where the water comes from and how much is consumed versus withdrawn, how many jobs are ongoing versus construction, what incentives apply, where noise is measured at night, and whether a written agreement includes remedies that can change the project.
Chapter 11 · The Accountability Standard
An AI Accountability Framework: Evidence Classes, Claim Records, and Pilot Protocols
Every chapter in this report kept arriving at the same place. Schools, households, workplaces, city halls, clinics, retirement accounts, power grids. Different evidence, different people, same requirement. So this chapter is my attempt at an AI accountability framework that anyone can pick up and use, whether you run a district, sit on a utility commission, or are trying to decide whether to trust a chatbot with your mother's medication list. The real question underneath all of it is simple: when someone tells you an AI system works, how do you know what that claim is actually worth, and who is on the hook if it's wrong?
Here's why I care about that question the way I do. In infrastructure, you do not energize a site on a promise. You energize it on a commissioning record. Somebody tested the breaker, somebody signed the sheet, somebody knows which relay trips first, and every one of those facts has a name and a date attached. I learned early, building large power loads, that a load that big doesn't care how confident the brochure sounded. It cares whether the record matches the physics. This chapter is the commissioning record for claims.
What the evidence will show is uncomfortable in both directions. The tools that institutions already lean on, including NIST's voluntary framework, give you structure but certify nothing. The strongest trials in this report hold up only when their scope conditions travel with them. And the standard I propose here, including every template, is itself a proposed workflow: the lowest rung on the evidence ladder it defines. I'll hold it to that label, and I'll hold Savrn's own claims to it too.
11.1 The Common Requirement Behind Every AI Accountability Framework#
Across education, household life, work, civic decisions, finance, health, and infrastructure, the preceding ten chapters kept arriving at the same practical requirement: a claim must remain connected to its evidence, its limits, and the person authorized to act. The requirement wore a different uniform in each chapter.
In schooling (Chapter 3), it showed up as the difference between assisted practice and unaided learning.
In the household (Chapter 6), it showed up as the difference between faster drafting and redistributed responsibility.
In civic life (Chapter 7), it was the difference between a readable summary and an accurate one. A clean summary of public comments can read beautifully and still leave out the objection your neighbor filed.
In finance and health (Chapter 8 and Chapter 9), it was the difference between a persuasive output and a measured outcome.
In infrastructure (Chapter 10), it was the difference between an announced commitment and a metered, enforceable one. A press release is a sentence. A meter is a record.
Figure 11.1Framework
One requirement across every domain
The recurring requirement from Chapters 1 to 10, stated once.
Line those up and you'll see one pattern. In every domain, the failure mode is a claim that drifts loose from its anchor. The number keeps circulating after its population, its comparison group, and its limit have fallen away. The recommendation keeps moving after the person who was supposed to approve it has been skipped. That's the whole game: keep the claim tied to the record, and keep the record tied to a person with the authority to act.
This chapter turns that recurring requirement into a complete proposed process. Nothing in it is an independently validated intervention or a legal-compliance certification. It is an evidence-informed operating standard, offered for testing and adoption by institutions that want the benefits described in the earlier chapters without inheriting the failure modes.
11.1.1 Where the NIST AI Risk Management Framework Fits#
If you work inside an institution, you've probably heard someone mention NIST. The governance backdrop for this chapter is NIST's voluntary risk-management framework, which organizes work around four functions: governance, mapping, measurement, and management. NIST also publishes a generative-model profile that addresses risks including confabulation, privacy, information integrity, and human interaction. (NIST AI Risk Management Framework, NIST generative AI profile)
Those documents give a research-informed governance basis, and I lean on their structure throughout. But be clear about what they are. They do not certify a product. They do not prove that any deployment produces benefits. A vendor who says "we follow the NIST framework" has told you about their process, not about your outcome.
The framework page also describes continuing work and revisions. That matters, because a draft or concept note should never be presented as a finalized binding standard. The same discipline applies to this chapter's own proposals as much as to NIST's. (NIST framework status page)
11.2 The Eight Evidence Classes: What Each Kind of Study Can Say#
This whole report relies on a classification of evidence types. I state it here in full because everything downstream depends on it, and other chapters link back to this table rather than repeating it. It is an editorial labeling system, not a formal evidence-grading methodology such as GRADE. The labels organize judgment. They do not replace it.
Class
What it can establish
What it cannot establish alone
Randomized comparison
An effect of the assigned intervention under the study's design and assumptions
Universal transfer, long-term benefit, or attribution to one component of a bundle
Quasi-experimental analysis
An estimated effect under an explicit identification strategy
Freedom from unmeasured confounding without supporting assumptions
Observational or survey evidence
Associations, reported experience, measured prevalence within scope
Causal benefit from adoption
Technical evaluation
Performance on specified inputs and scoring rules
Real-world welfare or reliable operation outside the evaluated task
Simulation or forecast
Conditional results under stated assumptions
Observed outcomes or guaranteed future results
Administrative or official record
A documented policy, status, expenditure, or measured quantity
A causal effect beyond the record's scope
Publisher statement
What an organization says about its product, method, or commitment
Independent verification of the claim
Proposed workflow
A process that can be implemented and tested
A benefit already demonstrated
Figure 11.2Framework
The eight evidence classes, with one example from this report
Every public summary must keep the class of its evidence. A simulation stays a simulation beside a trial; a publisher statement stays a publisher statement.
Every public-facing summary derived from this report should preserve the class of its underlying evidence. Three rules follow, and I'd tape them to the wall of any communications office:
A simulation must not become a real-world outcome merely because it appears beside an empirical study.
A publisher statement must not become verification because the publisher is reputable.
A proposed workflow must not become a finding because it is well constructed.
You've watched this report apply the classes repeatedly. The METR developer studies were treated as evaluations whose meaning shifted with changing tools. The JLARC bill figures were scenarios, not observations. The LBNL national totals were modeled estimates, not meter data. The classification is not bookkeeping. It is the mechanism that keeps sound claims sound.
11.2.1 Each Class, With One Example From This Report#
A table is easy to nod at and hard to use. So here is each rung, what it looks like in practice, and the most common way people misread it.
Randomized comparison. The Swedish MASAI trial randomized 105,934 women to model-supported mammography screening or standard double reading (MASAI trial record). Randomization is what lets you say the difference came from the assigned workflow rather than from who chose it. The common misreading is to treat a randomized result as universal: the effect belongs to that screening system, inside that radiologist workflow, under that design. It is not evidence about consumer self-diagnosis, and it did not measure mortality.
Quasi-experimental analysis. The apprenticeship evaluation cited in the evidence register found that registered participation was associated with positive estimated employment and earnings effects, and the register flags its quasi-experimental design and possible unobserved selection (DOL apprenticeship evaluation). The misreading is to treat the estimate as if randomization had removed selection. It didn't. The identification strategy carries assumptions, and the claim is only as good as those assumptions.
Observational or survey evidence. A survey can tell you how many people report using a tool, or how they feel about it. It can measure prevalence within its scope. The misreading is causal: "people who use it are doing better, so it made them better." Surveys can't carry that sentence alone.
Technical evaluation. The UK Department for Transport's blinded test of a consultation-analysis tool measured performance on specified inputs under a scoring procedure: theme generation reached approximately 0.75 recall, 0.50 precision, and 0.59 F1 (Department for Transport evaluation). The misreading is to treat benchmark performance as real-world welfare. A tool that scores well on the test set has not yet shown that any committee made a better decision.
Simulation or forecast. Virginia's JLARC study modeled a potential increase of roughly $14 to $37 per month in a typical Dominion residential bill by 2040 under its examined scenarios (JLARC data-center report). That is a conditional result under stated assumptions. The misreading is turning it into a present-tense bill increase on a flyer. Chapter 10 calls that manufacturing a fact the audit never produced.
Administrative or official record. The AEP Ohio data-center tariff is a documented, commission-approved policy with minimum contract-capacity ramps and exit obligations (AEP Ohio data-center tariff). A record like that establishes what the policy says. The misreading is to treat the record as proof of an effect beyond its scope: a tariff on paper is not evidence of what happened to anyone's bill.
Publisher statement. Savrn's own design goals, including behind-the-meter power and a zero-makeup-water cooling design goal, are publisher statements. So are the publisher-described methods of the seven Savrn trackers. They tell you what we say. They are not independent verification of it. I'll come back to this in Section 11.6, because it's the part most people would expect me to skip.
Proposed workflow. This chapter. The claim record, the five-stage process, the pilot protocol, the four-touch sequence, the eight-step audience workflow. All of it can be implemented and tested. None of it is a benefit already demonstrated.
11.3 The Claim Record: Ten Fields for AI Claim Verification#
The proposed record for any consequential claim contains enough information for a different reviewer to understand the claim and its limits without calling the author. Not every record needs a numerical estimate. Every consequential claim needs a clear basis.
Identity. A stable claim identifier, version, responsible editor, and chapter or use case. Without an ID and a version, nobody can tell whether two people quoting "the study" mean the same sentence, and corrections can't find their target.
Statement. The narrow proposition being asserted, without promotional extensions. Write the sentence the evidence supports, not the sentence marketing would prefer.
Population and setting. Who, where, when, and under what institutional conditions. A finding from radiologists in a national screening program is not a finding about a consumer app.
Intervention and comparator. What changed and what it was compared with. "Better" is meaningless until you know "better than what."
Outcome. The exact measure, denominator, time horizon, and whether it was primary, secondary, exploratory, or self-reported. Show me the denominator.
Estimate and uncertainty. The effect, the interval where available, and the study's treatment of statistical significance. A point estimate without its interval is half a fact.
Technology boundary. Model and version when stated, or explicit identification of a non-generative intervention. A result from a specific screening system is not a result about general conversational models.
Source. Original URL, publication date, status, and relevant update.
Limitations. Selection, attrition, measurement, transfer, sponsorship, and conflicting evidence.
Action boundary. What the evidence permits, what it does not permit, and who must approve consequential use.
The last field is the one I'd fight hardest to keep. Most evidence summaries stop at "here's what the study found." The action boundary forces you to say what a decision-maker may do with it, what they may not do, and whose signature is required. That is where a finding turns into accountability.
11.3.2 A Filled-In Claim Record: The MASAI Mammography Trial#
Here's the ten-field record applied to the strongest clinical trial in this report, using only what Chapter 9 reports from the trial record. Where the report doesn't state a value, the record says so rather than guessing. That is a feature of the template, not a gap in it.
Field
Entry
Identity
Example ID: CH09-MASAI-01; version 1; responsible editor named by the using institution; use case: clinical screening evidence (Chapter 9)
Statement
In this screening workflow, model-supported reading achieved a noninferior interval-cancer outcome compared with standard double reading, with higher sensitivity and similar specificity
Population and setting
Women in the Swedish MASAI trial: 105,934 randomized, 105,915 included in the reported analysis after exclusions; screening inside a radiologist workflow
Intervention and comparator
A specific screening-support system embedded in a radiologist workflow, compared with standard double reading
Outcome
Primary: interval cancer (cancers diagnosed between screening rounds), per 1,000 participants. Secondary: sensitivity and specificity. Mortality was not measured
Estimate and uncertainty
Interval cancer 1.55 per 1,000 (supported) vs 1.76 per 1,000 (standard); ratio 0.88, 95% CI 0.65 to 1.18. Met the noninferiority criterion; did not show a statistically significant reduction. Sensitivity 80.5% vs 73.8%; specificity approximately 98.5% in both groups
Technology boundary
A specific screening system within a supervised radiologist workflow; not a general conversational model acting as an autonomous physician. Model version: not stated in this report's summary
Source
Lancet trial record via PubMed; publication date and later updates to be recorded from the original record at time of use
Limitations
The upper bound of 1.18 means the data are consistent with the supported workflow being somewhat worse as well as somewhat better on interval cancer; no mortality outcome; no evidence about consumer self-diagnosis; sponsorship and conflicts to be recorded from the original paper
Action boundary
Permits: citing a noninferior interval-cancer result and improved sensitivity without a rise in false alarms, within a defined radiologist workflow. Does not permit: "prevented 12 percent of cancers," mortality claims, or consumer-diagnosis claims. Consequential use requires approval by the clinical authority responsible for the screening program
Read the action-boundary row again. Chapter 9 said it plainly: "noninferior interval-cancer outcome in this screening workflow" is a finding the data support, and "the technology prevented 12 percent of cancers" is not, because the direction of the point estimate doesn't survive its confidence interval and mortality was never measured. The claim record makes that distance visible on one page. Anybody who later writes the second sentence has to delete a row to do it.
It's worth running the same record on the null result that sat beside MASAI in Chapter 9, because a standard that only works on good news isn't a standard. A randomized trial of 50 physicians working structured diagnostic vignettes found median diagnostic-reasoning scores of 76 percent with conventional resources and 74 percent with access to a language model, an adjusted difference of two percentage points with a confidence interval spanning -4 to 8, not statistically significant (JAMA Network Open trial). The model alone performed strongly on the same vignettes in an exploratory analysis, and that did not translate into a demonstrated gain for the physicians who used it. Its action boundary would read: permits a statement that no workflow gain was demonstrated under the tested conditions; does not permit citing standalone model performance as evidence of improved physician decisions or patient outcomes. Capability is not benefit.
11.3.3 A Blank Claim-Record Template You Can Copy#
Copy this into a document, a spreadsheet, or a shared form. One record per consequential claim.
CLAIM RECORD
1. Identity
Claim ID:
Version:
Responsible editor:
Chapter or use case:
2. Statement (the narrow proposition, no promotional extensions):
3. Population and setting (who, where, when, institutional conditions):
4. Intervention and comparator (what changed; compared with what):
5. Outcome
Exact measure:
Denominator:
Time horizon:
Primary / secondary / exploratory / self-reported:
6. Estimate and uncertainty
Effect:
Interval (if available):
Significance as treated by the study:
7. Technology boundary (model and version if stated, or "non-generative"):
8. Source
Original URL:
Publication date:
Status (final, preprint, draft, disputed):
Relevant update or correction:
9. Limitations
Selection:
Attrition:
Measurement:
Transfer:
Sponsorship:
Conflicting evidence:
10. Action boundary
Evidence permits:
Evidence does not permit:
Approver for consequential use:
Evidence class (from the eight-class table):
Not stated in source (list fields left blank on purpose):
I added two lines at the bottom that aren't among the ten fields. The evidence-class line keeps the rung attached to the claim. The "not stated in source" line is where you write down what the original didn't tell you, so a blank field reads as a known gap rather than an oversight.
The consolidated evidence register in the back matter of this report is organized on these fields, covering the major study programs and principal permitted claims. A separate citation ledger indexes source-linked passages across every one of the 83 distinct sources cited in this report. It is an audit aid, not a claim that every cited publication is independent or equally strong. Several entries describe the same programs from different angles, and repeated appearances of one study do not turn it into multiple independent studies.
11.4 From Question to Action: A Five-Stage AI Governance Framework#
A claim record describes evidence. A process decides what to do with it. The proposed process has five stages, and each stage has an owner.
Figure 11.3Framework
From question to action: the five-stage process
Proposed operating standard. The highlighted control is non-negotiable: a model's ability to draft an email, filing, or instruction confers no permission to send or execute it.
The process begins with the person's question, intended use, and authority to act. Researching options, recommending an action, approving it, and executing it are distinct permissions, and a system that can do the first has no claim on the fourth.
The owner identifies affected people, potential harms, private data, legal or professional obligations, and available non-model alternatives. If a task can be completed adequately with a simpler, lower-risk process, the evaluation should include that alternative. "The model did it" is not a result if a checklist would have done it as well.
Here's a hypothetical. A county clerk's office wants to use an assistant to answer residents' questions about permit deadlines. Intake asks: Who is asking, and what will they do with the answer? Who is authorized to state a deadline officially? What happens to a resident who relies on a wrong date? Is there a simpler alternative, such as a single maintained deadline table on the county website? If the table would work as well, the pilot has to compare against it, not against nothing.
The researcher retrieves original documents, records dates and versions, and distinguishes quotations from interpretation. Conflicting sources remain visible until resolved. A newer source is not automatically better if it concerns a different population or denominator.
Two cases from Chapter 7 show why this stage can't be skipped. The Department for Transport consultation-analysis evaluation illustrates why summaries require omission and invention checks: a theme generator with 0.50 precision needs sampling against the original comments, not a readability judgment. Under the evaluation's matching procedure, roughly half of generated themes were judged relevant, and a recall of 0.75 meant a quarter of human-validated themes were left out (Department for Transport evaluation). The New York City MyCity audit illustrates why denominators and grading rules must be inspectable before an accuracy figure means anything: the comptroller and the agency used different denominators and assessment choices, and the agency disputed the findings (MyCity audit and agency response).
The analyst separates five categories that generated output blurs by default: observed facts, derived calculations, assumptions, forecasts, and value judgments. Another reviewer should be able to reproduce material calculations and identify which assumptions change the conclusion.
The challenge review asks what evidence would reverse the recommendation. It tests the strongest reasonable alternative explanation rather than comparing only against a weak opposing argument. A recommendation that survives only against a strawman has not been tested.
Here's the part most people skip. The five-way split is not academic. Take one sentence from a hypothetical staff memo: "The facility will raise household bills by $37 a month, which is unacceptable." Split it. Is "$37" observed or forecast? If it came from the JLARC scenario range, it's the top of a modeled 2040 range for Dominion's Virginia territory, not an observation, and not a figure about any other utility. Is "will" a calculation or an assumption? Is "unacceptable" a value judgment? It is, and that's legitimate, but it belongs in its own column. Once the sentence is split, the challenge review has something to grab.
The decision record states the options, relevant evidence, uncertainty, distributional effects, and accountable decision-maker. The user or authorized institution approves consequential action through its own process.
A model's ability to generate an email, application, financial instruction, or public statement does not confer permission to send or execute it. Drafting and action remain separate controls, and the separation is a feature, not friction to be optimized away.
The owner compares actual outcomes with the original commitments, records incidents and corrections, and reopens the decision when material assumptions fail. A successful launch is not the final outcome.
The METR developer-study update is the standing example of why. Tools changed, task selection changed, and the measured effect changed with them, which is why a process must accommodate changing tools rather than freezing a historical estimate as permanent truth (METR update). The later evidence neither erases the earlier randomized result nor supplies a clean replacement number.
The following protocol is proposed for a school, employer, community organization, service provider, or household-support program that wants to test assistance fairly. It is not a claim that such a pilot has already been conducted.
Element
Required specification
Objective
One primary outcome that matters to the intended beneficiary
Population
Eligibility, exclusions, consent, accessibility, and expected setting
Comparison
Current practice and, where relevant, a simpler alternative
Assignment
Randomization where feasible, or an explicit nonrandom design with limitations
Intervention
Tool version, instructions, retrieval sources, human support, and permissions
Cost boundary
Acquisition, implementation, use, review, correction, escalation, and exit
Quality boundary
Error definition, missing-information rules, independent grading, and serious-harm criteria
Follow-up
Immediate and later outcomes appropriate to the claim
Analysis
Intention-to-treat where applicable, missing-data handling, subgroup plan, and uncertainty
Decision rule
Conditions to continue, revise, pause, or stop
Publication
Favorable, null, adverse, and inconclusive findings retained
Most of the table explains itself once you've read the chapters before it. A few lines deserve a sentence of their own.
Objective. One primary outcome, chosen for the beneficiary, not five outcomes with the best one promoted after the fact. For a tutoring pilot, that probably means unaided learning, because Chapter 3 showed how easily assisted practice gets mistaken for it.
Population and comparison. A pilot that silently excludes people with disabilities, limited connectivity, or a different first language will report an average nobody in those groups will experience. And the comparison should include the simpler alternative from Stage One, not just "no tool."
Intervention. Tool version, instructions, retrieval sources, human support, and permissions. If any of these change mid-pilot, you're now testing a different intervention. The METR pair is the reason this line exists.
Cost and quality boundaries. Acquisition is the smallest cost line; review, correction, escalation, and exit are where real costs hide. Define an error and a serious harm before you see output, and use graders who don't know which arm produced what, so nobody gets to redefine failure after an incident.
Analysis and decision rule. Intention-to-treat means analyzing people in the group they were assigned to, including those who stopped using the tool. The decision rule (continue, revise, pause, stop) is written before results arrive. Without it, every result becomes a reason to continue.
Publication. Favorable, null, adverse, and inconclusive findings retained. This is the element sponsors most want to soften, which is why it's on the list.
11.5.2 Sample Size, Feasibility, and the Mislabel Problem#
Two elements carry the most weight.
First, sample size should be calculated from the intended effect, outcome variability, design, and acceptable uncertainty, rather than selected for convenience and then described as definitive. No universal sample size is prescribed here, because the right number depends on the question. The physician trial in Chapter 9 is a useful reminder: with 50 physicians, a confidence interval running from -4 to 8 points could not rule out a meaningful gain or a meaningful loss.
Second, a short pilot can establish feasibility or expose failures without establishing long-term effectiveness. Its public description should state which of those purposes it actually served. A feasibility study marketed as an effectiveness study is not a gray area. It is a mislabel.
Here's a hypothetical on the mislabel. A workforce program runs a six-week pilot of an assistant that helps job seekers draft cover letters. Participants finish more applications, and staff like it. That's useful feasibility evidence: the tool worked in the setting, people used it, nothing broke badly. It is not evidence that anyone got hired faster or earned more. If the program's annual report says "AI pilot improves employment outcomes," it has promoted a feasibility study to an effectiveness claim. The correct sentence is shorter and less exciting: "A six-week feasibility pilot found the tool usable; employment effects were not measured."
Follow-up
Immediate outcome timing:
Later outcome timing:
Analysis
Intention-to-treat (yes / not applicable, reason):
Missing-data handling:
Subgroups pre-specified:
How uncertainty will be reported:
Sample size and how it was calculated:
Every stage above needs an owner, and some responsibilities can't be handed to software no matter how capable it gets. Here is the proposed matrix.
Role
Proposed responsibility
Cannot be delegated merely by using a model
Sponsor
Funding disclosure, scope, publication rights, and access to adverse findings
Candor about commercial interests
Research editor
Claim accuracy, version control, and correction process
Source fidelity
Domain specialist
Technical, educational, clinical, financial, or engineering review
Professional judgment within their remit
Deployment owner
Permissions, access, monitoring, incident handling, and exit
Operational responsibility
Affected person
Informed participation and a route to challenge
The burden of detecting every hidden error
Public authority
Lawful procedure and legitimate decisions
Statutory authority and public accountability
Figure 11.5Framework
Accountability by role: what cannot be delegated to a model
Proposed six-role matrix. The sponsor row applies to this report: prepared for Savrn, not an independent institutional review.
Read the right-hand column slowly. It's the most important column in this chapter. Each entry names something a model can help with and can never own. A model can check a citation; the research editor still owns source fidelity. A model can flag an anomaly in a scan; the clinician still owns the judgment. A model can draft a zoning summary; the public authority still owns the decision and answers for it.
The affected-person row runs the other direction. It says what must not be pushed onto the person at the end of the chain. A resident, a patient, a parent, or a job seeker should get informed participation and a route to challenge. They should not carry the burden of detecting every hidden error. If your deployment only works when the least-resourced person in the system catches the machine's mistakes, the deployment doesn't work.
11.6.1 The Sponsorship Row Applies to This Report#
The sponsorship row is this report's own, and it deserves direct treatment. This research was prepared for Savrn and concerns an industry in which Savrn has a commercial interest. It does not describe itself as an independent institutional review, and no outside peer-review panel is represented as having approved it.
The proposed publication policy, which this report applies to itself, is to preserve material unfavorable findings even when they weaken a commercial narrative. A sponsor should not be able to relabel an adverse result as missing data simply because it is inconvenient. You've seen that policy operating throughout: the null physician-workflow trial is reported as fully as the positive mammography trial, and the loneliness evidence is presented with its uncertainty intact. For loneliness, Chapter 9 reported a review that pooled three trials with 190 participants at g of -0.67 with an interval of -2.57 to 1.23 and high heterogeneity, which spans a very large benefit to a substantial harm and does not establish a reliable reduction (older-adult randomized-trial review). That result would have been easy to round up into a hopeful sentence. It wasn't.
The brief should be understandable without the chapter, but it must not contradict or materially overstate it. Every layer closer to the public compresses. The rule is that compression may shorten but never strengthen.
Here's what that rule looks like in practice. Take the MASAI record from Section 11.3.2 and compress it layer by layer:
Layer
Acceptable compression
Unacceptable strengthening
Original source
Full trial record
Not applicable
Claim register
Noninferior interval cancer; ratio 0.88 (95% CI 0.65 to 1.18); sensitivity 80.5% vs 73.8%; specificity about 98.5% both; radiologist workflow
Dropping the interval or the workflow scope
Research chapter
Noninferior on interval cancer, higher sensitivity without more false alarms, in a defined screening workflow
"Reduced interval cancers"
Public brief
In a large Swedish screening trial, model-supported reading caught cancers at least as well as standard double reading, inside a radiologist workflow
"AI prevents breast cancer"
Each row is shorter than the one above it. None is stronger. When you check a brief, walk it backward through the layers: can every sentence be traced to the register and then to the source without gaining strength along the way?
For outreach built around the seven Savrn trackers, the research in this report can support a non-promotional four-touch sequence:
First touch, the human question: Explain one practical decision faced by a teacher, household, worker, or civic member.
Second touch, the evidence: Present one favorable finding and its main limitation, with the original source.
Third touch, the infrastructure connection: Show which tracker can locate a relevant record and which local facts are still needed.
Fourth touch, the decision standard: Provide a short method for comparing alternatives, asking questions, and checking an answer.
A hypothetical sequence for a county commissioner audience might run: first, the question of how to read a data-center proposal before a hearing; second, the JLARC finding that existing Virginia rates appropriately allocated current costs at the time of its study, alongside its warning that future growth could increase system costs, both linked to the JLARC data-center report; third, how the Grid Operator Watchlist can locate a relevant tariff record, and which local interconnection studies are still needed; fourth, the eight-step workflow in Section 11.9.3. No touch asks anyone to support anything.
This is a proposed content sequence, not a claim that the campaign improves participation or decision quality. Any later campaign evaluation should measure comprehension and useful action, not just opens and clicks. The evaluation itself belongs to the pilot protocol of Section 11.5, with its favorable, null, and adverse findings all retained.
The proposed refresh policy uses triggers rather than pretending every source becomes stale on the same schedule.
Live records get rechecked before consequential use. Product configurations, tariffs, permits, and live project status can change without notice. The AEP Ohio tariff of Chapter 10 is exactly the kind of record that can change by commission order; the Public Utilities Commission of Ohio approved its data-center tariff settlement in July 2025 (PUCO announcement, AEP Ohio tariff process). Before anyone relies on its terms, someone should open the current version.
Historical trials stay historical, but get checked. A trial's result doesn't expire, but it can be corrected, retracted, or followed up. Before an estimate is quoted as current, check for those.
Each correction has five parts. It should identify the affected claim, the original wording, the corrected wording, the reason, and the version. Here's a blank you can copy:
CORRECTION NOTICE
Claim ID affected:
Original wording:
Corrected wording:
Reason for correction:
New version number and date:
Layers updated (brief / chapter / register / outreach):
Material corrections propagate. A material correction should reach summaries and outreach content, not remain hidden in the longest document. If the brief said it wrong, the brief gets fixed. A correction buried on page 400 while the wrong sentence keeps circulating on the one-pager isn't a correction.
The cutoff is a boundary, not a warranty. The September 23, 2026 evidence cutoff of this edition is not a promise that every linked webpage will remain unchanged or that all subsequent evidence has been incorporated.
11.9 The Seven Savrn Trackers and the Four-Layer Chain#
The seven Savrn trackers organize evidence about data-center capital, timing, public policy, water, grid governance, permits and power development, and modeled compute scarcity:
Their public value depends on the decision being made and the additional records brought to that decision. The publisher-described methods do not establish a family's bill, a student's learning, a resident's tax burden, a worker's job, an investor's return, or the net effect of a facility. Under the table in Section 11.2, the trackers' method descriptions are publisher statements, and the records they index keep whatever class they had at the source. A tracker entry starts a question. It never ends one.
The common structure connecting the trackers to any audience in this report is a four-layer chain:
Tracker record: What the indexed source says, with scope, status, and observation date.
Local or institutional record: The budget, tariff, permit, contract, employer record, school policy, or household baseline relevant to the decision.
Causal or allocation analysis: The method connecting the project or policy to the outcome.
Decision: The responsible person or authority weighs the evidence and alternatives.
Figure 11.4Framework
The four-layer tracker chain and its forbidden inferences
A tracker is an evidence discovery instrument, not an answer engine. No study reviewed here tests the seven trackers as an intervention.
Skipping from the first layer to the fourth produces confident but unsupported conclusions. The tracker is an evidence-discovery instrument, not a universal answer engine.
Here's a hypothetical walk through the chain. A resident sees a large capital announcement for a nearby project in the Capital Atlas and wonders whether it will lower her property taxes. Layer one: the tracker record shows the announcement, its source, its status, and when it was observed. Layer two: she needs the county budget, the assessment rules, the tax rates, and any incentive agreement. Layer three: someone has to connect the project to the tax bill, which Chapter 10 warned can come out either way, because an increase in the local tax base does not automatically reduce a household's bill and incentives belong on the same ledger as receipts. Layer four: the county's elected body makes budget decisions, and she can bring her question to them with the records in hand. The tracker got her to the right question. It didn't answer it.
Each tracker carries a specific forbidden inference. I state them here once, in consolidated form, because the pattern matters more than any single entry.
Capital Atlas: A dollar announced is a dollar spent locally, a tax benefit to residents, a permanent job, or an investment return. It is none of these until the corresponding record shows it.
Delay Watchlist: An entry proves failure, establishes its cause, or predicts permanent cancellation.
Moratorium Tracker: A proposal is law, a pause is permanent, or a measure in one jurisdiction governs another.
Water Tracker: A cooling figure proves zero total use, a permit equals actual consumption, or a regional total establishes local household harm.
Grid Operator Watchlist: Queue position proves energization, state-wide capacity proves local deliverability, or one tariff applies to every project.
Permits Tracker: A permit equals construction, a load request equals contracted demand, or planned generation is available capacity.
Scarcity Tracker: A modeled index is realized revenue, a liquid market price, a guaranteed financing basis, or proof that residents benefit.
Look at the shape they share. Every forbidden inference takes a record of one kind and treats it as a record of a more consequential kind. An announcement becomes spending. A proposal becomes law. A queue position becomes power. A model becomes money. That's the same drift the evidence classes are built to stop, applied to infrastructure records.
The proposed interface for all of this begins with a human question rather than a tracker name. A resident asking about a bill should not need to know that the relevant evidence may appear in a grid docket, a tariff, a capital disclosure, and a public budget.
The eight-step audience workflow is proposed for evaluation:
State the question and decision date.
Select relevant records with dates and statuses preserved.
Open the original sources.
Identify the local records still missing.
Separate observed facts, calculated values, scenarios, and preferences.
Compare reasonable alternatives, including no change.
Route consequential interpretation to the qualified person.
Record the decision, assumptions, commitments, and later results.
No study reviewed in this series tests the seven trackers as a combined intervention, and this report says so wherever they appear.
11.9.4 Evaluate for Comprehension, Not Persuasion#
Evaluation, when it comes, should measure comprehension rather than persuasion. The correct initial study is not whether exposure makes residents more supportive of data centers. It is whether a source-linked workflow improves factual comprehension and decision quality, measured by correct identification of:
project stage,
source scope,
relevant authority,
unresolved evidence, and
prohibited inference.
Participants who conclude that the evidence is insufficient should be retained in the analysis. "Unable to determine" can be the correct answer and must not be scored as a failure to adopt.
Here's a hypothetical test item to show what that means. A participant is shown a Moratorium Tracker entry describing a proposed pause in a neighboring county and asked: "Does this pause apply to a project in your county?" The correct answer is no, because a measure in one jurisdiction does not govern another. Asked "Is the pause now law?", the correct answer depends on the status field, and if the entry says "proposed," the answer is no. Asked "Will the pause be permanent?", the correct answer is that the record can't tell you. A participant who writes "unable to determine" on that last item got it right.
A short comprehension study should never be promoted as evidence of improved public budgets, family welfare, portfolio returns, or community trust. If we ever run one, the publication commitment of Section 11.5 applies to it in full.
11.10 How a School, a Utility, and a Family Would Each Use This Standard#
The instruments above can look like they were built for large institutions with research staff. They weren't. Here are three short illustrations, each labeled as a hypothetical, each using only the tools defined in this chapter.
A district is offered a reading assistant for grades 5 through 8. The curriculum director starts at intake: the question is whether students read better on their own, the decision date is the spring budget vote, and the superintendent is the approver. A simpler alternative, more small-group time with a reading specialist, goes into the comparison.
At evidence assembly, the vendor's materials are classified with the eight classes. A benchmark score is a technical evaluation. A teacher survey is survey evidence. A study "showing gains" gets a claim record, and the population, comparator, and outcome fields reveal whether the gains were measured with the tool open or after it was removed.
The district then writes a one-page pilot protocol. The objective is unaided reading performance at the end of the semester. Assignment is by classroom, randomized where the schedule allows, with the limitation stated where it doesn't. The permissions line says the assistant may suggest and explain but may not submit graded work. The decision rule says the district stops if unaided performance doesn't improve relative to the comparison, whatever the satisfaction scores say. The publication line commits to posting results for parents whether they're favorable or not.
The accountability matrix names the curriculum director as deployment owner, a reading specialist as domain specialist, and parents and students as affected persons with a route to challenge. The board is the public authority. None of this required a research department. It required a page of writing and six names.
11.10.2 A Utility or Utility Commission (Hypothetical)#
A utility is asked by a large-load customer to connect a new facility, and a commissioner wants to know what it means for residential customers. The four-layer chain frames the work.
Layer one: tracker records show the project's announcement in the Capital Atlas and a related entry in the Grid Operator Watchlist. Both are observed dates and statuses, nothing more. Layer two: the utility's staff pull the actual tariff, the interconnection study, and the resource plan. If a tariff with minimum contract-capacity ramps and exit obligations exists, like the one AEP Ohio operates under, it's an administrative record of the allocation mechanism. Layer three: someone runs the allocation analysis, and the analyst separates observed load from forecast load and names which assumptions move the residential result. A scenario range like Virginia's is labeled as a scenario. Layer four: the commission decides, on the record, through its own procedure.
The claim record earns its keep at the hearing. When an advocate says the project will raise bills by a specific amount, staff can ask which class that number belongs to. When a developer says the project will pay its own way, staff can ask for the record that makes the commitment enforceable, not the press release that announces it. Both sides get held to the same table.
An adult daughter is using a chat assistant to help organize her father's medications and appointments after a hospital stay. She doesn't need a protocol document. She needs three pieces of this chapter.
First, intake and authority: the assistant can help her research and draft, but her father's physician and pharmacist are the ones authorized to decide dosing, and her father is the one whose consent matters. Second, the drafting-and-action separation: the assistant can draft a question list for the next appointment, and it doesn't get to change a dose or cancel a prescription. Third, evidence classes as a filter: if she reads that an AI system caught cancers in a screening trial, she can recognize that result as belonging to a radiologist workflow, not to a chatbot reading her father's discharge papers. Chapter 9 said it directly: the trial is not evidence about consumer self-diagnosis.
She might also keep a simplified record of what she learns, a family version of the claim record: what the source was, what it said, what it didn't say, and who she needs to ask. That's four lines instead of ten. It's still a record.
11.11 Publication Gates and Actual Completion Status#
This report completes a desk-research manuscript covering the requested life stages, infrastructure questions, tracker mappings, and proposed operating process. It does not claim completion of original field trials, independent peer review, legal review, clinical review, engineering certification, or a site-specific impact assessment. A reader or institution should treat any summary that implies otherwise as inaccurate.
Before using this research as an institutional policy or project justification, the proposed external gates are:
Gate
What it checks
Status in this edition
Domain review
Technical, educational, clinical, financial, or engineering accuracy
Not completed
Sponsor-conflict review
Whether Savrn's commercial interest shaped any claim
Not completed
Local-record verification
Whether the national or state evidence matches the local records
Not completed; depends on each use
Accessibility testing
Whether the materials work for the people they're meant for
Not completed
Approval by the responsible authority
Legitimate decision under the institution's own procedure
Belongs to each using institution
Those are implementation requirements, not unfinished chapters. The distinction matters. A proposal that labels its remaining gates is ready for the next step. A proposal that hides them is not ready for any step.
11.12 Conclusion: Evidence That Stays Inspectable#
The common standard is evidence that remains inspectable from source to decision and from decision to measured result. The proposal running through every chapter of this report is to use increasingly capable systems to strengthen that chain rather than conceal its weak links: to make summaries more checkable, records more accessible, alternatives more visible, and commitments more enforceable, while keeping the decision with the person or institution authorized to make it.
That's what a commissioning record does for a power facility. It doesn't make the equipment better. It makes the equipment's condition visible to the people who have to decide whether to close the breaker. The claim record, the pilot protocol, the accountability matrix, and the four-layer chain do the same for claims about AI.
This report can support discussion and practical evaluation now. Claims of proven program effectiveness, independent validation, or community benefit should wait for the corresponding evidence. The next evidentiary step is implementation testing under the shared protocol of Section 11.5, not stronger adjectives.
In this report, an AI accountability framework is a proposed operating standard that keeps every consequential claim connected to its evidence, its limits, and the person authorized to act. It combines eight evidence classes, a ten-field claim record, a five-stage decision process, an eleven-element pilot protocol, and a six-role accountability matrix. It is a proposed workflow offered for testing, not an independently validated intervention or a legal-compliance certification.
Does following the NIST AI Risk Management Framework mean an AI tool is safe or effective?
No. NIST's voluntary framework organizes risk work around governance, mapping, measurement, and management, and its generative-model profile addresses risks such as confabulation, privacy, and information integrity. It provides a governance basis, but it does not certify a product or prove that any deployment produces benefits. The framework page also describes continuing work and revisions, so drafts should not be presented as final binding standards.
What are the eight evidence classes for evaluating AI claims?
The classes are randomized comparison, quasi-experimental analysis, observational or survey evidence, technical evaluation, simulation or forecast, administrative or official record, publisher statement, and proposed workflow. Each can establish something specific and nothing more on its own. A benchmark score shows performance on specified inputs, not real-world welfare. A publisher statement shows what an organization says, not independent verification. This is an editorial labeling system, not a formal grading method such as GRADE.
What should a claim record include?
Ten fields: identity, statement, population and setting, intervention and comparator, outcome, estimate and uncertainty, technology boundary, source, limitations, and action boundary. The goal is that a different reviewer can understand the claim and its limits without calling the author. The action boundary matters most, because it states what the evidence permits, what it does not permit, and who must approve consequential use.
How should a school or employer run an AI pilot evaluation?
Use the eleven-element protocol: objective, population, comparison, assignment, intervention, cost boundary, quality boundary, follow-up, analysis, decision rule, and publication. Pick one primary outcome that matters to the beneficiary, compare against current practice and a simpler alternative, randomize where feasible, and retain favorable, null, adverse, and inconclusive findings. A short pilot can show feasibility, but it should not be described as proof of long-term effectiveness.
What did the MASAI mammography trial actually show?
MASAI randomized 105,934 Swedish women to model-supported screening or standard double reading. Interval cancer was 1.55 versus 1.76 per 1,000 (ratio 0.88, 95% CI 0.65 to 1.18), which met noninferiority but was not a statistically significant reduction. Sensitivity rose from 73.8% to 80.5% with specificity near 98.5% in both groups. Mortality was not measured, and the result applies to a radiologist workflow, not consumer self-diagnosis.
Can the Savrn trackers tell me whether a data center will raise my electric bill?
Not by themselves. The seven Savrn trackers are evidence-discovery instruments: a tracker entry starts a question and never ends one. To reach a bill conclusion you need the four-layer chain: the tracker record, the local tariff and budget records, an allocation analysis connecting the project to rates, and a decision by the responsible authority. No study reviewed in this report tests the trackers as a combined intervention.
Has this AI accountability framework been independently reviewed?
No. It was prepared for Savrn, which has a commercial interest in the data-center industry, and it does not describe itself as an independent institutional review. No outside peer-review panel has approved it, and domain, sponsor-conflict, legal, clinical, and engineering reviews have not been completed. Its publication policy is to keep unfavorable findings even when they weaken a commercial narrative, such as the null 50-physician trial reported beside the positive MASAI result.
Back matter · Evidence Register and Glossary
Evidence Register, Glossary, Sources, and Disclosures
I've asked you to take a lot on trust over eleven chapters. You shouldn't. Every claim I made traces to a row below, and every row traces to a source you can open yourself. This is the ai research evidence register for the whole report: the study, the narrowest claim it can carry, and the limit that claim runs into. If I stretched a finding somewhere, this is where you catch me.
When you build large power loads, you learn to trust the meter over the drawing. A summary is the drawing; the study is the meter. Here's the part most people skip: the limit column matters as much as the claim column. Check my work.
12.1 How to Use the AI Research Evidence Register#
The register indexes the major empirical and documentary evidence used in this report, organized by study program or record rather than by sentence. Exact estimates, denominators, methods, and citations remain in the chapter where each study is interpreted.
It does not make heterogeneous outcomes directly comparable, and one study cited in several chapters is still one study.
Here's how I'd read each row:
Start with the ID. The letter gives the theme: E education, P pathways into work, H household, C capability, W work, V civic life, F finance, L later life, I infrastructure, G governance, T trackers.
Open the source. Every row carries at least one link.
Read the permitted claim as a ceiling, not a floor.
Read the limit before you repeat the claim.
Find the evidence class. The eight classes are defined in Chapter 11. A working paper is not a published trial, and a publisher statement is not verification.
Many rows are nulls, mixed results, or adverse findings, kept on purpose.
Figure 12.1Data
What the register contains, by domain
Entries are evidence programs or records, not independent studies; one program can appear in several chapters and is counted once here.
12.2 The Consolidated Evidence and Claim Register#
Changing tools, task selection, and time measurement complicate a current uplift estimate.
The update explicitly does not provide one definitive productivity number.
How to read this table. Packages beat products. The positive results (E01, E02, E06, E09, P07) mostly come from structured programs with people around the technology, and attribution to the model alone is not established. E08 and P10 cut the other way with equal weight: unassisted performance or comprehension fell. P04 is a null, not proof of equivalence. Ask what surrounds the tool, not just which tool it is.
Parents reported particular uses and perceived time effects.
Self-report adoption survey without causal measurement.
How to read this table. READY4K (H04), SNAP outreach (H05), and Home Energy Reports (H06) are not model studies; they show what structured information and human help can do, a benchmark newer tools must meet. The rows that test model access are sobering: H09 found no reliable gain in condition identification, and H10 found adverse loneliness evidence. H01, H02, and H11 describe behavior and belief, not causation.
12.2.3 Capability, Work, Civic Life, Finance, and Later Life#
Favorable within-group changes warrant further study.
Within-group comparisons do not isolate causal effects.
How to read this table. The capability rows (C01 through C03) show assistance speeding one task, hurting accuracy on another, and slowing experienced developers in a third. W01 measures exposure, not displacement. V02 shows measurable performance alongside omissions and false themes. The finance rows show no realized investor outcome. L01 is a strong workflow trial that still does not show reduced mortality, and the nulls in L02 and L03 sit at full size.
12.2.4 Infrastructure, Governance, and Public Record#
Product updates occurred after public calls for safety-conditioned pacing.
Sequence alone does not establish contradiction or sufficient safety.
How to read this table. This is my industry's table. I01 is model-based national scale, not site meter data. I02 says no complete national water inventory exists. I03 is Virginia only, and I04 is an allocation mechanism, not proof every risk is gone. G03 through G05 are publisher statements: what an organization says, not whether it did it. G05's sequence of events is neither proven contradiction nor proven safety.
Spot price, liquid settlement, or guaranteed revenue
How to read this table. The trackers are not studies, so the columns change. Savrn publishes them and has a commercial interest in what they cover. A tracker entry starts a question and never ends one: an announced dollar is not a local job, and a queue record is not energization. The four-layer chain in Chapter 11 runs from tracker record to local record to causal or allocation analysis to decision; skipping steps produces confident, unsupported conclusions.
The evidence supports conditional, bounded conclusions and keeps favorable, null, mixed, and adverse findings side by side. It does not establish that every proposed workflow works, or that a useful application automatically justifies a facility. Capability is not benefit. The next step is implementation testing under the shared protocol of Chapter 11, not stronger adjectives.
Assistance, delegation, substitution. Defined in the Terminology and Conventions section of The Research Thesis. Assistance: a system prepares, drafts, or explains under human review. Delegation: a human authorizes a specific action. Substitution: the system performs a function previously done by a person without per-case authorization.
Behind-the-meter. Power supplied on the customer's side of the utility meter rather than drawn entirely through the public grid.
Confidence interval. The range of values consistent with the data at a stated confidence level. A 95 percent interval spanning benefit and harm, such as the loneliness estimate's -2.57 to 1.23, means the study cannot tell them apart.
Denominator discipline. Reporting every figure with its population, comparator, and limit in the same breath.
Difference-in-differences. A quasi-experimental method comparing how an outcome changed for a group that got something against a group that did not, over the same period. P09 used it. It works only if both groups would otherwise have moved in parallel.
Effect size (Hedges' g, Cohen's d). Differences between groups in standard-deviation units, so studies on different scales can be compared. Values near 0.2 are typically called small.
Evidence class. One of eight editorial labels, from randomized comparison through proposed workflow, defined in Chapter 11, stating what evidence can and cannot establish.
F1 score. The harmonic mean of precision and recall. The 0.59 F1 in the DfT theme test (V02) reflects weakness in both directions at once.
Forbidden inference. A cross-layer conclusion the evidence does not support, such as announced capital equaling local jobs.
Heterogeneity. Variation in results across people, settings, or studies. High heterogeneity (as in H07) means one pooled average hides very different effects.
Intention-to-treat. Comparing groups as randomized, regardless of what participants actually used. Analyzing only engaged users can manufacture effects.
Interconnection. The process and agreement for connecting a large generator or load to the grid, usually through a queue and studies. A queue position is not energization.
Interval cancer. A cancer diagnosed between scheduled screening rounds; the MASAI trial's primary outcome (L01).
Noninferiority. A design testing whether a new approach is no worse than a standard by more than a preset margin. Meeting it is not demonstrating improvement.
Pass-through estimate. In the pension experiment (F03), how much a recommended allocation change showed up in participants' choices: 0.368, 95% CI 0.303 to 0.433, in hypothetical incentivized choices.
Percentage point. The plain difference between two percentages. Moving from 10 to 15 percent is 5 percentage points, or a 50 percent relative increase.
Precision and recall. Precision is the share of flagged items that are correct; recall is the share of correct items that were flagged. Recall of 0.75 means a quarter of validated themes were missed.
Prediction interval. In a meta-analysis, the range where a new, similar program's effect would be expected to fall. A wide one (as in H07) means a small average benefit promises nothing in your setting.
Preprint. A paper posted publicly before, or without, journal peer review. E05, P10, and F03 are preprints.
Scenario, forecast, simulation. A conditional result under stated assumptions; a planning input, not an observation. The LBNL 2028 electricity range and the JLARC 2040 bill range are scenarios.
Sensitivity and specificity. Sensitivity is the share of true cases detected (MASAI: 80.5 percent supported versus 73.8 percent standard); specificity is the share of non-cases correctly cleared (about 98.5 percent in both arms).
Standard deviation. How spread out values are around their average; the unit effect sizes are measured in.
Tariff. A utility's regulator-approved schedule of rates and terms for a customer class, such as the AEP Ohio data center tariff (I04). Tariffs can change by commission order.
The seven Savrn trackers. The Capital Atlas, Delay Watchlist, Moratorium Tracker, Water Tracker, Grid Operator Watchlist, Permits and Power Development Tracker, and Scarcity Tracker: an evidence-discovery instrument.
TWh (terawatt-hour). One billion kilowatt-hours. LBNL estimated about 176 TWh of data center use in 2023, with 2028 scenarios of roughly 325 to 580 TWh.
Withdrawal and consumption (water). Withdrawal is water taken from a source; consumption is water not returned. Both must be distinguished from discharge, permits, and modeled estimates.
Within-group comparison. Change in one group over time with no control group. It cannot separate the intervention from time, attention, or regression to the mean (see L04).
Working paper. A paper circulated by an author or institution such as NBER or CESifo before or instead of journal publication. Estimates can change.
12.4 Source Verification: How I Checked Every AI Research Source#
On September 23, 2026, every distinct URL in the report's citation ledger received an automated browser-like request, with redirects followed and a 25-second timeout, plus manual spot-checks of representative blocked sources and one replacement search. Three rules governed drafting:
Blocked sources are cited with an access note, and no claim may depend on content that could not be inspected directly or through an independent fetch path.
Pages that change after the cutoff are quoted as of the cutoff.
A source that fails verification is logged and downgraded to "as cited in the September 23, 2026 edition," never silently dropped or upgraded.
Live pages that block automated access (major publishers, government sites, OpenAI). Spot-checks confirmed content matches; OpenAI "Introducing Superalignment" was verified verbatim. Cited with an access note.
HTTP 203, nonstandard success
2
PubMed record for MASAI (PMID 41620232) and JMIR e78568; both live.
Recovered on retry
1
ERIC EJ963694 (kindergarten reading trial) returned HTTP 200 on retry.
HTTP 404, dead link
1
SNAP outreach manuscript on economics.mit.edu (truncated filename). Replaced.
12.4.3 The One Correction: SNAP Outreach Experiment#
The SNAP outreach experiment (H05) originally linked to a manuscript on the MIT economics site that returned a 404. I replaced it with the canonical location of the same study: Finkelstein and Notowidigdo, "Take-up and Targeting: Experimental Evidence from SNAP," NBER Working Paper 24652, later published in the QJE.
The full text was retrieved and figures confirmed. The study covered about 30,000 elderly individuals likely eligible for SNAP. Over nine months, enrollment was 5.8 percent in the control group, 10.5 percent with information only, and 17.6 percent with information plus application assistance. The verification log rounded those to 6, 11, and 18 percent; this report uses the figures as published. Chapter 6 cites the NBER record as a working paper with later QJE publication.
In my own words: this research was prepared for Savrn, the company I run. Savrn operates in an industry this report examines and publishes the seven trackers used throughout. Read every page with that in mind. What I did about it was keep the unfavorable findings at full size and label Savrn's own design goals as publisher statements. What I can't do is certify my own independence. Nobody can.
Formal statement. This report was prepared for Savrn and concerns an industry in which Savrn has a commercial interest, including through the seven Savrn trackers described throughout. Savrn's publisher role and commercial interest are stated wherever tracker findings support an argument, consistent with the public-claim controls set out in Chapter 11. This report does not describe itself as an independent institutional review. No outside peer-review panel is represented as having approved it. The proposed publication policy, preservation of material unfavorable findings even when they weaken a commercial narrative, has been applied throughout the drafting. The sponsor-conflict review listed among the publication gates in Chapter 11 is an external requirement for any institutional use of this report, not a step this report can perform on itself.
The evidence cutoff for this edition is September 23, 2026. All findings, product references, regulatory records, tracker statuses, and URLs reflect that date. The cutoff is a boundary, not a warranty: linked pages may change, records may be corrected, and later evidence is not incorporated. Product configurations, tariffs, permits, and live project status should be rechecked against original records before consequential use.
The refresh policy uses triggers rather than pretending every source goes stale on the same schedule. Product configurations, tariffs, permits, and live project status should be rechecked before consequential use; the AEP Ohio tariff in Chapter 10 is exactly the kind of record that can change by commission order. Historical trials remain historical records, but they should be checked for corrections, retractions, and material follow-up before their estimates are quoted as current.
Each correction should identify the affected claim, original wording, corrected wording, reason, and version. A material correction should propagate to summaries and outreach content, not stay hidden in the longest document.
Chad Everett Harris is the Founder and CEO of Savrn. He is a serial infrastructure entrepreneur based in Dallas. He has spent his career building large power infrastructure. He now builds Savrn, and he works from an "operator first" philosophy: no infrastructure without customers, and vertical integration as a defense.
Savrn designs, builds, and delivers AI factories, purpose-built data centers. Its design is built around behind-the-meter power and closed-loop cooling with a zero-makeup-water design goal. Those are publisher statements, held to the same evidence test as every other organization's claims. Savrn also publishes the seven Savrn trackers listed in the register.
It is an index of every major study and record behind this report. Each row gives the source link, the narrowest claim the evidence permits, and its primary limit. Exact estimates and denominators stay in the chapter that interprets each study, and one study cited in several chapters is still one study.
How were the sources in this AI report verified?
Every distinct URL in the citation ledger got an automated browser-like check as of September 23, 2026, plus manual spot-checks of blocked sources. Sixty-four loaded directly, 17 were blocked but confirmed another way, 2 returned a nonstandard success code, 1 recovered on retry, and 1 was dead and replaced.
What was the SNAP study correction?
The dead MIT economics link now cites NBER Working Paper 24652 by Finkelstein and Notowidigdo. Among about 30,000 elderly people likely eligible, nine-month enrollment was 5.8 percent for controls, 10.5 percent with information, and 17.6 percent with information plus application help. It tested human assistance, not an AI model.
Who sponsored this AI research report, and why does it matter?
No. It was prepared for Savrn, which has a commercial interest in the data center industry the report examines and publishes the seven Savrn trackers. It is not an independent institutional review, no outside peer-review panel approved it, and an external sponsor-conflict review is required before institutional use.
Why are the Savrn trackers in the evidence register?
Because the report uses them, and anything the report uses belongs where you can check it. The tracker table lists what each record can support and what it cannot support alone. An announced dollar is not a local job. A tracker entry starts a question and never ends one.
What does the evidence cutoff date mean?
The cutoff is September 23, 2026, and findings, records, tracker statuses, and URLs reflect that date. It is a boundary, not a warranty: pages can change and later evidence is not included. Recheck tariffs, permits, product configurations, and live project status before any consequential decision.
Research Thesis · The Research Thesis
The Research Thesis: Why the AI Transition Is an Institutional Question
This is the last section of the report, and it's the one I'd hand you first if you asked me what the whole thing argues. Eleven chapters and more than a hundred thousand words come down to one claim, and it isn't the claim you'd expect from someone who builds AI infrastructure for a living. I'm not arguing that the transition will go well. I'm not arguing that it will go badly. I'm arguing that, right now, it's an institutional question. It gets decided by whether the people and organizations that adopt these systems keep every claim attached to its evidence, its limits, and the person accountable for acting on it.
That's the research thesis behind this AI research report 2026. Everything above this section is the evidence for it. What follows is the argument laid out in order: why this research exists, the premises the thesis rests on, the seven findings that carry it, what would prove it wrong, how it was tested, and the terms and disclosures you need to judge it yourself.
The move toward increasingly capable AI systems usually gets told as a story about the systems: what they can do, how fast they're improving, and when they might cross some threshold. This report tells it as a story about people. It follows the human life stages those systems will actually pass through, from a child's first years of school to an older adult's retirement and care, and asks one question at each stage: do people end up better off, and how do we know? The answer the evidence gives is consistent. The gains are real, measurable, and smaller and more conditional than the public conversation suggests. Some well-designed uses help. Some well-performing tools did nothing measurable once real people used them. At least one made later learning worse. And the loudest numbers in the infrastructure debate are often forecasts dressed as facts. Across all of it, the difference between a gain and a null was rarely the model. It was the institution around the model. That's the thesis.
Figure T.1Data
Headline findings, each with its limit attached
Eight findings from across the report. Color marks direction, not importance. Every number is shown with the population and the limit that travel with it.
I've spent my career building large power infrastructure, and now I build Savrn out of Dallas. Savrn designs and builds AI factories, data centers built around behind-the-meter power and closed-loop cooling. You might have expected a report with my name on it to tell you the machines are coming and everything will be fine. It doesn't. It puts favorable, null, and adverse findings side by side, and several of them cut against the easiest talking points my own industry uses. I asked for it that way on purpose.
When you build at industrial scale, you learn one rule early: the spreadsheet does not get a vote. The transformer either holds the load or it doesn't. The water either comes back or it doesn't. A county either trusts you after the first public meeting or it doesn't. Operators who believe their own pitch decks don't stay operators very long.
"Operator first" is how we run Savrn. It means no infrastructure without customers, and vertical integration as a defense against depending on things you can't control. Applied to evidence, operator first means something just as plain. Don't build a conclusion before you have the demand for it, meaning the actual question a real person is trying to answer. Don't borrow certainty from someone else's benchmark. And own every link in the chain between a claim and the person who has to act on it, because the weak link is the one that fails under load.
So the question this research was built to answer is simple to state and hard to answer well: when AI shows up in a classroom, a job, a kitchen table budget, a doctor's office, a county commission agenda, or a retirement plan, what do we actually know about whether people end up better off?
The evidence base is a fixed corpus assembled as of September 23, 2026. It includes randomized trials, government evaluations, official audits, national statistics, and documented regulatory cases. Eleven chapters walk the lifespan:
A thesis is only as strong as the steps under it. Here are the five premises this one stands on, in the order the chapters establish them. Each premise is carried by specific studies, and each study keeps its limit.
Capability is real, and it's measurable. In defined tasks, assistance has produced measured gains: faster professional writing, improved tutor behavior, higher screening sensitivity in a supervised radiology workflow.
Capability does not travel on its own. A model's standalone performance is not the performance of the professional, student, or household that uses it. Between the two sits a chain of human and institutional links, and some of them are weak.
Benefit concentrates where institutions are strong. Measured gains show up where a defined barrier, an accountable institution, a real comparison condition, and a measurable outcome appear together.
Claims lose their denominators in transit. Between the study and the headline, the population, the comparator, and the limit fall away, and a bounded finding becomes a general promise or a general threat.
Trust follows the evidence state. Public trust attaches to inspectable records, published adverse findings, and enforceable commitments, not to reassurance.
Conclusion. If capability is real but doesn't travel on its own, if benefit depends on the institution, if claims shed their limits in transit, and if trust follows the record, then the outcome of the transition is decided by institutional discipline: whether claims stay tied to their evidence, their limits, and the people authorized to act on them. That's not a hedge. It's a practical claim about where the leverage is. The leverage is in the process around the tool, and that process is something a school board, an employer, a hospital, a county, and a household can control.
The seven findings below are how the chapters support those five premises.
The synthesis compresses eleven chapters into seven findings. Each one carries its numbers exactly as the underlying studies report them, with the population, the comparison, and the limit kept in the same place. If you jumped straight here from the table of contents, this is the fastest way into this AI evidence review.
T.4.1 The transition is measurable, and the measurements are smaller than the discourse#
The strongest causal evidence in this report is real but bounded. Three studies anchor that claim, and each one comes with a limit you have to carry along with the headline.
These are real findings with real limits, and this report treats every one of them as both. The writing result tells you what happened on a set of tasks. It does not tell you that a company's output rose or that anybody got a raise. The mammography result tells you what a particular system did inside a particular clinical process run by radiologists. It does not tell you a chatbot can read your scan. The SNAP result tells you that a human who helps fill out the form triples enrollment compared with doing nothing. It tells you nothing about software, because no software was tested.
Here's the part most people skip. The biggest effect in this set, the SNAP jump, came from people helping people. That doesn't make AI irrelevant to benefits access. It tells you what the bar is: a new tool has to beat an arrangement that already works, not a straw man.
What this means for you. If you're an employer, the writing study justifies a task-level pilot, not a headcount plan. If you're on a hospital board, MASAI justifies looking at supervised screening workflows, not replacing clinical judgment. If you run a benefits program, the SNAP result tells you the comparison group your AI pilot must beat is a trained human helper, not an empty inbox. Deeper treatment sits in Chapter 2, Chapter 6, and Chapter 9.
T.4.2 Capability and benefit are separated by a chain with weak links#
A model's standalone performance is not the performance of the professional, student, or household that uses it. That pattern repeats across the whole corpus.
Physicians. In a physician-reasoning trial in JAMA Network Open, the model alone scored well, but physicians with access gained a statistically indistinguishable two points. The trial used 50 physicians and diagnostic vignettes, not live patients.
High school students. In a high school mathematics experiment published in PNAS, unrestricted assistance improved assisted practice while reducing later unaided performance. The practice looked better. The learning got worse.
Software developers. A METR study from early 2025 found assistance slowed experienced developers (n = 16). A later METR update (n = 57) moved toward speedups, while the measurement problems stayed unresolved.
Read that again: a strong model, a small and uncertain gain for doctors. A helpful tutor, weaker learning for students. A slowdown that turned into a possible speedup when the tools and tasks changed. Capability is not benefit. Any institution that buys capability on the basis of standalone benchmarks is buying the wrong outcome.
Common misreading. "The METR update proves the slowdown was wrong." It doesn't. It proves the number moved when the tools and the task selection moved, and that an estimate needs a date on it. That's why the report treats METR as technical evaluation, whose meaning shifts with changing tools. Chapter 5 walks through it.
What this means for you. A parent or teacher should ask whether a homework tool is being judged by the work it helps produce or by what the student can do alone a week later. Those are different outcomes, and the PNAS study shows they can move in opposite directions. Chapter 3 covers it grade band by grade band.
T.4.3 The biggest institutional failure is denominator discipline#
The lesson this report repeats most often is technical and unglamorous: a number without its population, comparator, and limit attached is not information. A number without a denominator is a rumor. Four examples from the corpus show how it goes wrong.
Modeled 2040 bill ranges are not present-tense rate increases; the Virginia JLARC scenarios give roughly $14 to $37 per month under examined assumptions
"The project will create 1,500 jobs."
A construction peak of about 1,500 workers over 12 to 18 months is not an operating staff of about 50
"Scores went up after we adopted the tool."
A within-group improvement is not an effect
This report proposes denominator discipline as an operating standard, not a writing style. Every consequential claim should carry a claim record with its evidence class attached, so a different reviewer can see what the number counts, what it's compared against, and where it stops being true. The claim record is laid out in Chapter 11, and the MyCity dispute gets full treatment in Chapter 7.
T.4.4 Assistance helps most where the institution is strongest#
Across every chapter, measured benefit concentrates where four conditions show up together:
A defined barrier
An accountable institution
A real comparison condition
A measurable outcome
Tutor CoPilot worked in a structured tutoring nonprofit. The MASAI system worked inside a radiologist workflow. Year Up worked as a comprehensive program package with earnings gains sustained to year ten, and it was not a model intervention at all.
The implication cuts against both the hype and the dismissal. The same tool that transforms one setting can do nothing measurable in another, and the difference is usually institutional, not technical. That's the whole game for anyone deciding about a deployment: the question isn't "is this model good?" but "is our institution the kind where assistance has been shown to work, and can we measure whether it did?"
What this means for you. A school board, workforce board, or employer should spend as much time on the program around the tool as on the tool. Year Up is in this report precisely because it shows what a sustained, measured pathway gain looks like, and none of it came from a model. See Chapter 4 for the pathway evidence and Chapter 3 for Tutor CoPilot.
T.4.5 Public trust is an evidence state, not a messaging problem#
In 2026 survey data published in August 2026, Pew Research Center found 52 percent of Americans more concerned than excited about AI in daily life, and 9 percent more excited than concerned. The 2023 to 2024 public record shows why trust has to be earned structurally. The superalignment team was dissolved within days of its leaders' departures in May 2024, and a resignation letter saying "safety culture and processes have taken a backseat to shiny products" entered the public record.
Trust attaches to inspectable denominators, published adverse findings, and enforceable commitments. It doesn't attach to reassurance. I'll say that plainly as someone whose company lives or dies on community trust: you can't message your way out of a record. Chapter 1 lays out the chronology and a proposed public communication standard.
Common misreading. "People are worried because they don't understand the technology." The survey measures attitudes; it doesn't measure understanding, and it doesn't show that any single communication event caused the concern. Read with the Chapter 1 record, it describes a public waiting to be shown evidence, not one waiting for a better slogan.
T.4.6 The infrastructure chapter is this report's test of its own method#
This is the chapter where Savrn has the most to gain or lose, so it's where the method has to hold hardest. The application benefits documented in Chapters 2 through 6 say nothing about whether a specific data center benefits its host community. The report refuses that inference in both directions: a useful chatbot does not justify a facility, and a bad chatbot does not condemn one.
What the evidence does support:
National electricity estimates with stated scenario ranges.Lawrence Berkeley National Laboratory estimated 176 TWh, or 4.4 percent of national consumption, in 2023, and 325 to 580 TWh, or 6.7 to 12 percent, by 2028 under stated assumptions. The first is an estimate of the past. The second is a scenario range, not a forecast of your county.
Documented allocation mechanisms. The AEP Ohio data center tariff uses a minimum-commitment architecture that shows how cost risk can be assigned to large new loads.
A proposed community benefit compact. Its enforceability test is simple: an adverse result must be able to change the project. A compact that can't change anything isn't a compact.
A thesis about institutions has to hand institutions something they can use. Chapter 11 consolidates this research into an operating process:
an eight-class evidence labeling system
a ten-field claim record
a five-stage question-to-action process that keeps drafting and action separate
an eleven-element pilot protocol
a six-role accountability matrix
the seven-tracker boundary: a tracker entry starts a question; it never ends one
None of it is validated. All of it is testable. The report says so on every page where it appears, and I'm saying it here too. If you take nothing else from this report, take the process, test it in your own institution, and publish what happens, including the parts that don't work.
A research thesis that can't lose isn't a thesis. It's a slogan. So here's what the evidence would have to show, in future studies, for this one to fail. Each test is stated against a finding already in this report, so you can check any new study against it.
If the thesis is wrong, you would expect to see
What this report's evidence currently shows
Standalone model scores reliably predicting what users achieve with the model
In the 50-physician vignette trial, a strong standalone model and a statistically indistinguishable two-point gain for physicians with access
Assisted practice reliably turning into unaided learning
In the PNAS high school mathematics experiment, better assisted practice and weaker later unaided performance
Effects that hold across settings regardless of the institution around the tool
Gains concentrated where a defined barrier, an accountable institution, a real comparison, and a measurable outcome appear together
A model intervention beating a trained human helper in a randomized comparison
The largest effect in the set, SNAP enrollment rising from 5.8 percent to 17.6 percent, came from human application help, and no model was tested
Modeled infrastructure burdens matching later observed records without adjustment
The 325 to 580 TWh range for 2028 is a scenario under stated assumptions, and the JLARC bill ranges are modeled 2040 scenarios, not observed rate changes
Public trust rising on messaging alone, with no change in the record
No evidence in this corpus shows a communication event causing a change in trust in either direction
If well-designed studies start filling the left column, the thesis weakens, and this report should say so in its next edition. That's the same standard it asks of AI labs in Chapter 1 and of vendors in Chapter 11: publish the adverse result, and let it change the conclusion.
Common misreading. "An institutional thesis means the technology doesn't matter." It doesn't mean that. Better systems widen what's possible. The thesis says that what's possible becomes what's delivered only through institutions that measure it, and that's where the evidence shows the gap.
If I had to boil every chapter of this report down to one sentence, it would be this: keep the denominator attached, and never mistake capability for benefit.
Those two habits do most of the work in every chapter. Denominator discipline means any figure (cost, accuracy, time, jobs) travels with the population it measures, what it's compared with, and the limit of its interpretation, in the same sentence. "Capability is not benefit" means a system's performance on its own tells you what it can do on specified inputs, not what happens when a tired teacher, a busy physician, or a worried retiree actually uses it inside a real institution.
Here's a hypothetical to show how the two habits work together. Say a vendor tells a school district that its math assistant "scores in the top tier on benchmark problems" and "students using it completed far more practice." Denominator discipline asks: more than whom, measured on what population, over how long, and compared with which alternative? Capability-is-not-benefit asks: does more assisted practice mean more learning when the student works alone? The PNAS study in Chapter 3 shows that it can mean less. Two questions, and the pitch has to come back with evidence or it doesn't go forward.
Try the same two questions on an infrastructure claim. A developer says a project brings "1,500 jobs." Show me the denominator: is that the construction peak over 12 to 18 months, or the operating staff of about 50? A utility forecast says bills could rise. Show me the denominator: is that an observed rate increase or a modeled 2040 scenario like the JLARC range? The habit is the same whether you're a parent or a county commissioner.
This report is a desk-research synthesis of a fixed corpus assembled for the Savrn research series of September 2026. It includes peer-reviewed randomized trials and experiments, government-commissioned evaluations and technical reports, official audits with agency responses, national statistics, systematic reviews, working papers and preprints (labeled as such), regulatory records, and publisher-described tracker methodologies. The citation ledger indexes each source. A verification pass confirmed the status of every cited URL as of the cutoff, and the results are recorded in the verification log that accompanies the manuscript. One example of what that pass caught: a dead link for the SNAP study was replaced with its canonical version, NBER Working Paper 24652.
Studies were chosen for their relevance to decisions in the eleven chapter domains and for evidentiary strength within each domain. Older interventions that don't involve generative models were kept where they show which mechanisms have measured value. That's why Year Up and the SNAP outreach experiment are here: they set the bar a new tool has to clear. The corpus is not a systematic review of all published research. It's a curated evidence base, and this report claims no more than that.
No claim exceeds its evidence class. A simulation is never written as an observation, and a publisher statement is never written as verification.
Every number carries its population, comparator, and limit in the same breath.
Favorable, null, adverse, and mixed findings get equal structural prominence. No chapter leads with its strongest study and buries its nulls.
You can check the third rule yourself. The null physician trial sits right next to the positive mammography trial in the Seven Findings. The PNAS learning loss sits next to the Tutor CoPilot gain. The METR slowdown sits next to its update.
The prose is written for the person making a decision, not the specialist: plain constructions, contractions where they read naturally, technical terms defined at first use, and no promotional register. Where the underlying finding is uncertain, the sentence structure shows the uncertainty up front rather than asserting confidence and hedging afterward. This edition is written in my voice as founder of Savrn. The evidence descriptions are held to the same standard regardless of whose voice carries them, and my experience as an operator explains why a question matters. It never stands in for evidence.
It is not personalized financial, medical, legal, or tax advice.
It is not a site-specific engineering, fiscal, or environmental assessment.
It is not an independent institutional review (see the disclosure below).
It is not a claim that any proposed workflow in it has been field-tested.
The publication gates in Section 11.10 of Chapter 11 apply to any institutional use: domain review, sponsor-conflict review, local-record verification, accessibility testing, and approval by the responsible authority.
Assistance means a system prepares, organizes, drafts, or explains under human review.
Delegation means a human authorizes a specific action.
Substitution means the system performs a function a person used to perform, without per-case authorization.
The report's standard generally supports the first, conditions the second on explicit approval, and requires evidence the third almost never has. Most of the positive findings above are assistance findings. Almost none of the evidence supports substitution.
Every claim is labeled with one of eight editorial evidence classes: randomized comparison, quasi-experimental analysis, observational or survey evidence, technical evaluation, simulation or forecast, administrative or official record, publisher statement, and proposed workflow. The full table, with what each class can and cannot establish, is defined once in Chapter 11. It's a labeling system, not a formal grading methodology such as GRADE. The labels organize judgment; they don't replace it.
The practical rule: a claim can't be stronger than its class. A simulation is never written up as an observation. A publisher statement, including every statement Savrn makes about its own design, is never written up as verification.
Any figure, whether cost, accuracy, time, or jobs, is reported with the population it measures, what it's compared with, and the limit of its interpretation, all in the same sentence. You've already seen this applied to the JLARC bill scenarios and the construction job peak.
Where the evidence supports a bounded positive statement, the report makes it and states its limit. Where it doesn't, the report names the claim as not permitted rather than softening it with qualifiers. This convention comes from the research package and is kept because it's clearer than graded adjectives like "promising" or "emerging."
A forbidden inference is a specific cross-layer conclusion the evidence does not support, named explicitly where it comes up. Two examples: that announced capital equals local jobs, or that a cooling figure proves zero total water use. Each of the seven Savrn trackers carries its own forbidden inference, consolidated in Section 11.9.
Spot price, liquid settlement, or guaranteed revenue
Throughout, a tracker entry is treated as an evidence-discovery instrument. It starts a question and never ends one. These are Savrn products, and Savrn has a commercial interest in them, which is exactly why the boundary is written down.
All findings, product references, and regulatory records reflect the September 23, 2026 evidence cutoff. That date is a boundary on this edition. It is not a warranty that linked pages are unchanged or that later evidence has been folded in. Tariffs, permits, and project status should be rechecked against original records before anyone relies on them for a consequential decision.
This report is a single long page with a sidebar table of contents, so nobody has to read it front to back. Each chapter stands on its own and links to the others. If you're sharing it with someone, point them to the path that matches the decision in front of them.
Figure T.2Map
How the report is organized: one lifespan, one infrastructure layer, one standard
The report walks the lifespan in order, from kindergarten to retirement, then turns to the physical infrastructure and the shared standard. The seven Savrn trackers run across the lifespan as evidence discovery tools, never as proof of outcomes.
The lifespan runs from kindergarten to retirement, and the seven Savrn trackers run across it as a separate axis, because infrastructure questions show up at every stage of life, from a school budget to a retiree's electric bill. Keeping those two axes separate is deliberate: the tool-to-outcome chain and the facility-to-community chain interact, but neither substitutes for the other.
Read this section's thesis, argument, and seven findings, then Sections 11.2 through 11.6 of Chapter 11: the evidence classes, the claim record, the question-to-action process, the pilot protocol, and the accountability matrix. That's the operating core.
That means a school system, employer, health provider, financial firm, or public body. Read your domain's chapter first, then Chapter 11. Each domain chapter ends with a proposed evaluation suited to its setting, and the pilot protocol in Section 11.5 is the common skeleton.
If you are a...
Start here
Then read
School leader, teacher, or parent of a K-12 student
T.9.3 If you're a researcher or journalist auditing a claim#
Go to the chapter's figures, then to the Consolidated Evidence and Claim Register just above this section. Every consequential claim in this report traces to a register entry, and every register entry states its primary limit. The citation ledger in the research package indexes the source-linked passages for audit, covering every one of the 83 distinct sources cited in this report. Repeated citations of one study are not independent confirmations.
T.9.4 If you're a resident facing a local infrastructure decision#
Chapter 10's six-question power test and five-claim water boundary table are designed to be used in a public meeting. Chapter 7's eight-stage civic workflow applies to any proposal on the agenda. Bring them. Ask the questions out loud. That includes asking them of Savrn.
Numbers in this report travel with their populations, comparators, and limits in the same breath. The text distinguishes observed facts, modeled scenarios, and proposed workflows throughout. Where the evidence is thin, it says so instead of reaching for an adjective. That isn't caution for its own sake. It's the mechanism that keeps the strong claims strong.
This report was prepared for Savrn, and it concerns an industry in which Savrn has a commercial interest, including through the seven Savrn trackers. I'm the founder and CEO. You should read it knowing that.
Here's what that disclosure means in practice. This report does not describe itself as an independent institutional review. No outside peer-review panel is represented as having approved it. The publication policy it proposes, preserving material unfavorable findings even when they weaken a commercial narrative, has been applied throughout. Savrn's publisher role and commercial interest are stated wherever tracker findings support an argument. The sponsor-conflict review listed among the publication gates in Section 11.10 is an external requirement for any institution that wants to use this report. It's not a step the report can perform on itself.
What is the central thesis of this AI research report?
That whether the AI transition goes well is, right now, an institutional question. Across eleven chapters, measured gains showed up where a defined barrier, an accountable institution, a real comparison, and a measurable outcome appeared together, and nulls showed up where they didn't. The outcome depends on whether claims stay attached to their evidence, their limits, and the people authorized to act on them.
What is this AI research report 2026 about?
It's a desk-research synthesis of a fixed evidence corpus, assembled as of September 23, 2026, that follows AI through the human lifespan: schooling, training, work, households, civic life, finance, health, and the infrastructure behind it. It reports favorable, null, and adverse findings with equal prominence and closes with a proposed, untested standard for evidence and accountability.
What evidence would prove the thesis wrong?
Well-designed studies showing that standalone model scores reliably predict what users achieve, that assisted practice reliably becomes unaided learning, or that a model beats a trained human helper in a randomized comparison. Right now the evidence runs the other way: a 50-physician trial found a statistically indistinguishable two-point gain, and the largest benefits effect came from human application help.
Does AI improve productivity according to research?
In defined tasks, sometimes. A preregistered experiment with 453 professionals found writing tasks done about 40 percent faster with about 18 percent higher assessed quality, but it measured task outcomes, not earnings or firm output. METR found assistance slowed 16 experienced developers in early 2025, and a later update with 57 moved toward speedups while measurement problems stayed unresolved.
Is this report independent?
No. It was prepared for Savrn, which has a commercial interest in the data center industry and publishes the seven Savrn trackers. It does not describe itself as an independent institutional review, and no outside peer-review panel approved it. It does apply a policy of keeping unfavorable findings, and it lists a sponsor-conflict review as an external requirement for institutional use.
Does AI help students learn?
It depends on the design. Tutor CoPilot improved tutor behavior in a structured tutoring nonprofit. But a high school mathematics experiment published in PNAS found that unrestricted assistance improved assisted practice while reducing later unaided performance. Better homework output is not the same as better learning, and the evidence shows those two can move in opposite directions.
Can AI diagnose better than doctors?
The evidence here doesn't support that claim. In a 50-physician trial using diagnostic vignettes, the model alone scored well, but physicians with access gained a statistically indistinguishable two points. The strongest clinical result, the MASAI trial of 105,934 women, involved a specific screening system inside a radiologist workflow, with sensitivity of 80.5 percent versus 73.8 percent.
Will data centers raise my electric bill?
This report can't answer that for your household. Virginia JLARC scenarios show roughly $14 to $37 per month by 2040 under examined assumptions, which are modeled ranges, not present-tense rate increases. Nationally, LBNL estimated 176 TWh in 2023 and a scenario range of 325 to 580 TWh by 2028. Your answer depends on your utility's tariff and local records.
What does "denominator discipline" mean?
It means every figure, whether cost, accuracy, time, or jobs, is reported with the population it measures, what it's compared with, and the limit of its interpretation, in the same sentence. A construction peak of about 1,500 workers over 12 to 18 months, for example, is not an operating staff of about 50, and program spending is not chatbot-only cost.
What is the evidence cutoff for this report?
September 23, 2026. All findings, product references, and regulatory records reflect that date. It's a boundary on this edition, not a guarantee that linked pages haven't changed or that later evidence is included. Tariffs, permits, and project statuses should be rechecked against the original records before any consequential decision.