# The Research Thesis: Why the AI Transition Is an Institutional Question

Part of The Superintelligence Transition by Chad Everett Harris, Founder and CEO, Savrn. Published September 24, 2026. Evidence cutoff September 23, 2026.

Canonical: https://savrn.com/blog/the-superintelligence-transition/research-thesis
Full edition: https://savrn.com/blog/the-superintelligence-transition

Disclosure: prepared for Savrn, which has a commercial interest in AI infrastructure and publishes the seven trackers cited here. Not an independent institutional review.

**Key takeaways**

- The thesis: whether the AI transition goes well is, right now, an institutional question, decided by whether claims stay attached to their evidence, their limits, and the people authorized to act on them.
- Capability is real but bounded: 453 professionals finished writing tasks about 40 percent faster with about 18 percent higher assessed quality, a task outcome, not earnings or firm productivity.
- Capability is not benefit: in a 50-physician vignette trial, access to a language model moved diagnostic scores by a statistically indistinguishable two points, even though the model alone scored well.
- The bar is set by people helping people: information plus human application help raised SNAP enrollment from 5.8 percent to 17.6 percent among 31,888 eligible Pennsylvania households, with no model involved.
- Forecasts are not facts: national data center electricity was 176 TWh (4.4 percent of U.S. consumption) in 2023, while 325 to 580 TWh for 2028 is a scenario under stated assumptions.
- Trust follows the evidence: Pew's 2026 data, published in August 2026, found 52 percent of Americans more concerned than excited about AI in daily life and 9 percent more excited than concerned.
This is the last section of the report, and it's the one I'd hand you first if you asked me what the whole thing argues. Eleven chapters and more than a hundred thousand words come down to one claim, and it isn't the claim you'd expect from someone who builds AI infrastructure for a living. I'm not arguing that the transition will go well. I'm not arguing that it will go badly. I'm arguing that, right now, it's an institutional question. It gets decided by whether the people and organizations that adopt these systems keep every claim attached to its evidence, its limits, and the person accountable for acting on it.

That's the research thesis behind this AI research report 2026. Everything above this section is the evidence for it. What follows is the argument laid out in order: why this research exists, the premises the thesis rests on, the seven findings that carry it, what would prove it wrong, how it was tested, and the terms and disclosures you need to judge it yourself.

### T.1 The Thesis in One Paragraph

The move toward increasingly capable AI systems usually gets told as a story about the systems: what they can do, how fast they're improving, and when they might cross some threshold. This report tells it as a story about people. It follows the human life stages those systems will actually pass through, from a child's first years of school to an older adult's retirement and care, and asks one question at each stage: do people end up better off, and how do we know? The answer the evidence gives is consistent. The gains are real, measurable, and smaller and more conditional than the public conversation suggests. Some well-designed uses help. Some well-performing tools did nothing measurable once real people used them. At least one made later learning worse. And the loudest numbers in the infrastructure debate are often forecasts dressed as facts. Across all of it, the difference between a gain and a null was rarely the model. It was the institution around the model. That's the thesis.

*Figure T.1 (interactive on the page).*
### T.2 Why This Research Exists

I've spent my career building large power infrastructure, and now I build Savrn out of Dallas. Savrn designs and builds AI factories, data centers built around behind-the-meter power and closed-loop cooling. You might have expected a report with my name on it to tell you the machines are coming and everything will be fine. It doesn't. It puts favorable, null, and adverse findings side by side, and several of them cut against the easiest talking points my own industry uses. I asked for it that way on purpose.

When you build at industrial scale, you learn one rule early: the spreadsheet does not get a vote. The transformer either holds the load or it doesn't. The water either comes back or it doesn't. A county either trusts you after the first public meeting or it doesn't. Operators who believe their own pitch decks don't stay operators very long.

"Operator first" is how we run Savrn. It means no infrastructure without customers, and vertical integration as a defense against depending on things you can't control. Applied to evidence, operator first means something just as plain. Don't build a conclusion before you have the demand for it, meaning the actual question a real person is trying to answer. Don't borrow certainty from someone else's benchmark. And own every link in the chain between a claim and the person who has to act on it, because the weak link is the one that fails under load.

So the question this research was built to answer is simple to state and hard to answer well: when AI shows up in a classroom, a job, a kitchen table budget, a doctor's office, a county commission agenda, or a retirement plan, what do we actually know about whether people end up better off?

The evidence base is a fixed corpus assembled as of September 23, 2026. It includes randomized trials, government evaluations, official audits, national statistics, and documented regulatory cases. Eleven chapters walk the lifespan:

1. [Public trust and the framing of the transition](#ai-safety-public-trust)
2. [How capability translates, or fails to translate, into benefit](#ai-productivity-evidence)
3. [Kindergarten through high school](#ai-in-k12-education)
4. [College, training, and the first years of work](#ai-college-workforce-training)
5. [Working life and career change](#ai-jobs-displacement)
6. [Households and family life](#ai-for-families-households)
7. [Neighborhoods and civic decisions](#ai-local-government-civic)
8. [Investment advice and lifetime finance](#ai-financial-advice)
9. [Health, aging, and retirement](#ai-healthcare-aging)
10. [The physical infrastructure and its host communities](#data-centers-community-impact)
11. [A shared standard for decisions and accountability](#ai-accountability-framework)

### T.3 The Argument, Step by Step

A thesis is only as strong as the steps under it. Here are the five premises this one stands on, in the order the chapters establish them. Each premise is carried by specific studies, and each study keeps its limit.

1. **Capability is real, and it's measurable.** In defined tasks, assistance has produced measured gains: faster professional writing, improved tutor behavior, higher screening sensitivity in a supervised radiology workflow.
2. **Capability does not travel on its own.** A model's standalone performance is not the performance of the professional, student, or household that uses it. Between the two sits a chain of human and institutional links, and some of them are weak.
3. **Benefit concentrates where institutions are strong.** Measured gains show up where a defined barrier, an accountable institution, a real comparison condition, and a measurable outcome appear together.
4. **Claims lose their denominators in transit.** Between the study and the headline, the population, the comparator, and the limit fall away, and a bounded finding becomes a general promise or a general threat.
5. **Trust follows the evidence state.** Public trust attaches to inspectable records, published adverse findings, and enforceable commitments, not to reassurance.

**Conclusion.** If capability is real but doesn't travel on its own, if benefit depends on the institution, if claims shed their limits in transit, and if trust follows the record, then the outcome of the transition is decided by institutional discipline: whether claims stay tied to their evidence, their limits, and the people authorized to act on them. That's not a hedge. It's a practical claim about where the leverage is. The leverage is in the process around the tool, and that process is something a school board, an employer, a hospital, a county, and a household can control.

The seven findings below are how the chapters support those five premises.

### T.4 The Seven Findings That Carry the Thesis

The synthesis compresses eleven chapters into seven findings. Each one carries its numbers exactly as the underlying studies report them, with the population, the comparison, and the limit kept in the same place. If you jumped straight here from the table of contents, this is the fastest way into this AI evidence review.

#### T.4.1 The transition is measurable, and the measurements are smaller than the discourse

The strongest causal evidence in this report is real but bounded. Three studies anchor that claim, and each one comes with a limit you have to carry along with the headline.

**Evidence card**

- Title: Noy and Zhang, professional writing experiment (2023)
- Design: Preregistered randomized experiment published in Science
- Population: 453 college-educated professionals doing occupation-specific writing tasks
- Finding: Tasks completed about 40 percent faster, with about 18 percent higher assessed quality on graded tasks
- Limit: Task outcomes, not earnings, employment, or organizational productivity; no unaided retest
- Source: [Noy and Zhang in Science](https://www.science.org/doi/10.1126/science.adh2586)
**Evidence card**

- Title: MASAI mammography screening trial
- Design: Randomized trial of a model-supported screening workflow versus standard double reading
- Population: 105,934 women randomized in Sweden
- Finding: Met noninferiority for interval cancer, with sensitivity of 80.5 percent versus 73.8 percent and specificity near 98.5 percent in both arms
- Limit: A specific screening system inside a radiologist workflow, not a general conversational model acting as an autonomous physician
- Source: [MASAI trial record on PubMed](https://pubmed.ncbi.nlm.nih.gov/41620232/)
**Evidence card**

- Title: SNAP outreach experiment, Finkelstein and Notowidigdo
- Design: Randomized outreach experiment with three arms: status quo, information, and information plus human application assistance
- Population: 31,888 eligible Pennsylvania households
- Finding: Enrollment of 5.8 percent in control, 10.5 percent with information alone, and 17.6 percent with information plus assistance
- Limit: An older likely-eligible population in one state, and not a model intervention at all
- Source: [NBER Working Paper 24652](https://www.nber.org/papers/w24652)
These are real findings with real limits, and this report treats every one of them as both. The writing result tells you what happened on a set of tasks. It does not tell you that a company's output rose or that anybody got a raise. The mammography result tells you what a particular system did inside a particular clinical process run by radiologists. It does not tell you a chatbot can read your scan. The SNAP result tells you that a human who helps fill out the form triples enrollment compared with doing nothing. It tells you nothing about software, because no software was tested.

Here's the part most people skip. The biggest effect in this set, the SNAP jump, came from people helping people. That doesn't make AI irrelevant to benefits access. It tells you what the bar is: a new tool has to beat an arrangement that already works, not a straw man.

**What this means for you.** If you're an employer, the writing study justifies a task-level pilot, not a headcount plan. If you're on a hospital board, MASAI justifies looking at supervised screening workflows, not replacing clinical judgment. If you run a benefits program, the SNAP result tells you the comparison group your AI pilot must beat is a trained human helper, not an empty inbox. Deeper treatment sits in [Chapter 2](#ai-productivity-evidence), [Chapter 6](#ai-for-families-households), and [Chapter 9](#ai-healthcare-aging).

#### T.4.2 Capability and benefit are separated by a chain with weak links

A model's standalone performance is not the performance of the professional, student, or household that uses it. That pattern repeats across the whole corpus.

- **Physicians.** In a [physician-reasoning trial in JAMA Network Open](https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2825395), the model alone scored well, but physicians with access gained a statistically indistinguishable two points. The trial used 50 physicians and diagnostic vignettes, not live patients.
- **High school students.** In a [high school mathematics experiment published in PNAS](https://www.pnas.org/doi/10.1073/pnas.2422633122), unrestricted assistance improved assisted practice while reducing later unaided performance. The practice looked better. The learning got worse.
- **Software developers.** A [METR study from early 2025](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/) found assistance slowed experienced developers (n = 16). A [later METR update](https://metr.org/blog/2026-02-24-uplift-update/) (n = 57) moved toward speedups, while the measurement problems stayed unresolved.

Read that again: a strong model, a small and uncertain gain for doctors. A helpful tutor, weaker learning for students. A slowdown that turned into a possible speedup when the tools and tasks changed. Capability is not benefit. Any institution that buys capability on the basis of standalone benchmarks is buying the wrong outcome.

**Common misreading.** "The METR update proves the slowdown was wrong." It doesn't. It proves the number moved when the tools and the task selection moved, and that an estimate needs a date on it. That's why the report treats METR as technical evaluation, whose meaning shifts with changing tools. [Chapter 5](#ai-jobs-displacement) walks through it.

**What this means for you.** A parent or teacher should ask whether a homework tool is being judged by the work it helps produce or by what the student can do alone a week later. Those are different outcomes, and the PNAS study shows they can move in opposite directions. [Chapter 3](#ai-in-k12-education) covers it grade band by grade band.

#### T.4.3 The biggest institutional failure is denominator discipline

The lesson this report repeats most often is technical and unglamorous: a number without its population, comparator, and limit attached is not information. A number without a denominator is a rumor. Four examples from the corpus show how it goes wrong.

| Claim as it often circulates | What the record actually supports |
|---|---|
| "The city's chatbot cost this much." | Program-level spending is not chatbot-only cost ([New York City Comptroller MyCity audit](https://comptroller.nyc.gov/reports/audit-report-on-the-new-york-city-office-of-technology-and-innovations-mycity-system/)) |
| "Data centers will raise your bill by this much." | Modeled 2040 bill ranges are not present-tense rate increases; the [Virginia JLARC scenarios](https://jlarc.virginia.gov/landing-2024-data-centers-in-virginia.asp) give roughly $14 to $37 per month under examined assumptions |
| "The project will create 1,500 jobs." | A construction peak of about 1,500 workers over 12 to 18 months is not an operating staff of about 50 |
| "Scores went up after we adopted the tool." | A within-group improvement is not an effect |

This report proposes denominator discipline as an operating standard, not a writing style. Every consequential claim should carry a claim record with its evidence class attached, so a different reviewer can see what the number counts, what it's compared against, and where it stops being true. The claim record is laid out in [Chapter 11](#ai-accountability-framework), and the MyCity dispute gets full treatment in [Chapter 7](#ai-local-government-civic).

> **Operator's note: Why a power guy cares about denominators**
>
> Anyone who has sat through utility interconnection meetings learns fast that a megawatt figure means nothing until you know whether it's nameplate, contracted, or actually drawn, and over what hours. Same with water: withdrawn, consumed, or discharged are three different numbers. The AI debate is full of figures that skipped that question. Jobs quoted at construction peak. Bills quoted from a 2040 scenario as if they landed on this month's statement. I'd rather lose an argument with the right denominator than win one with the wrong one, because the wrong one always comes back to find you, usually at a public hearing.
#### T.4.4 Assistance helps most where the institution is strongest

Across every chapter, measured benefit concentrates where four conditions show up together:

1. A defined barrier
2. An accountable institution
3. A real comparison condition
4. A measurable outcome

[Tutor CoPilot](https://arxiv.org/pdf/2410.03017) worked in a structured tutoring nonprofit. The MASAI system worked inside a radiologist workflow. [Year Up](https://acf.gov/sites/default/files/documents/opre/year%20up%20long-term%20impact%20report_apr2022.pdf) worked as a comprehensive program package with earnings gains sustained to year ten, and it was not a model intervention at all.

The implication cuts against both the hype and the dismissal. The same tool that transforms one setting can do nothing measurable in another, and the difference is usually institutional, not technical. That's the whole game for anyone deciding about a deployment: the question isn't "is this model good?" but "is our institution the kind where assistance has been shown to work, and can we measure whether it did?"

**What this means for you.** A school board, workforce board, or employer should spend as much time on the program around the tool as on the tool. Year Up is in this report precisely because it shows what a sustained, measured pathway gain looks like, and none of it came from a model. See [Chapter 4](#ai-college-workforce-training) for the pathway evidence and Chapter 3 for Tutor CoPilot.

#### T.4.5 Public trust is an evidence state, not a messaging problem

In 2026 survey data published in August 2026, [Pew Research Center](https://www.pewresearch.org/short-reads/2026/08/18/young-adults-in-the-us-are-increasingly-wary-of-ai-concerned-it-will-take-jobs/) found 52 percent of Americans more concerned than excited about AI in daily life, and 9 percent more excited than concerned. The 2023 to 2024 public record shows why trust has to be earned structurally. The superalignment team was dissolved within days of its leaders' departures in May 2024, and a resignation letter saying "safety culture and processes have taken a backseat to shiny products" entered the public record.

Trust attaches to inspectable denominators, published adverse findings, and enforceable commitments. It doesn't attach to reassurance. I'll say that plainly as someone whose company lives or dies on community trust: you can't message your way out of a record. [Chapter 1](#ai-safety-public-trust) lays out the chronology and a proposed public communication standard.

**Common misreading.** "People are worried because they don't understand the technology." The survey measures attitudes; it doesn't measure understanding, and it doesn't show that any single communication event caused the concern. Read with the Chapter 1 record, it describes a public waiting to be shown evidence, not one waiting for a better slogan.

#### T.4.6 The infrastructure chapter is this report's test of its own method

This is the chapter where Savrn has the most to gain or lose, so it's where the method has to hold hardest. The application benefits documented in Chapters 2 through 6 say nothing about whether a specific data center benefits its host community. The report refuses that inference in both directions: a useful chatbot does not justify a facility, and a bad chatbot does not condemn one.

What the evidence does support:

- **National electricity estimates with stated scenario ranges.** [Lawrence Berkeley National Laboratory](https://escholarship.org/content/qt32d6m0d1/qt32d6m0d1.pdf) estimated 176 TWh, or 4.4 percent of national consumption, in 2023, and 325 to 580 TWh, or 6.7 to 12 percent, by 2028 under stated assumptions. The first is an estimate of the past. The second is a scenario range, not a forecast of your county.
- **Documented allocation mechanisms.** The [AEP Ohio data center tariff](https://www.aepohio.com/company/about/rates/data-center-tariff/) uses a minimum-commitment architecture that shows how cost risk can be assigned to large new loads.
- **A proposed community benefit compact.** Its enforceability test is simple: an adverse result must be able to change the project. A compact that can't change anything isn't a compact.

> **Operator's note: Holding Savrn to the same test**
>
> Savrn's design goals include behind-the-meter power and closed-loop cooling with a zero-makeup-water design goal. Under this report's own rules, those are publisher statements. They are what we say about our design, not verified outcomes, and nothing here claims we've achieved any metric. A cooling figure doesn't prove zero total water use. A power arrangement doesn't prove we aren't affecting a community's grid until the records show it. I want residents to apply the six-question power test and the water boundary table in [Chapter 10](#data-centers-community-impact) to us first. If our claims can't survive that test, they shouldn't survive at all.
#### T.4.7 The shared standard is the deliverable

A thesis about institutions has to hand institutions something they can use. Chapter 11 consolidates this research into an operating process:

- an eight-class evidence labeling system
- a ten-field claim record
- a five-stage question-to-action process that keeps drafting and action separate
- an eleven-element pilot protocol
- a six-role accountability matrix
- the seven-tracker boundary: a tracker entry starts a question; it never ends one

None of it is validated. All of it is testable. The report says so on every page where it appears, and I'm saying it here too. If you take nothing else from this report, take the process, test it in your own institution, and publish what happens, including the parts that don't work.

### T.5 What Would Prove This Thesis Wrong

A research thesis that can't lose isn't a thesis. It's a slogan. So here's what the evidence would have to show, in future studies, for this one to fail. Each test is stated against a finding already in this report, so you can check any new study against it.

| If the thesis is wrong, you would expect to see | What this report's evidence currently shows |
|---|---|
| Standalone model scores reliably predicting what users achieve with the model | In the 50-physician vignette trial, a strong standalone model and a statistically indistinguishable two-point gain for physicians with access |
| Assisted practice reliably turning into unaided learning | In the PNAS high school mathematics experiment, better assisted practice and weaker later unaided performance |
| Effects that hold across settings regardless of the institution around the tool | Gains concentrated where a defined barrier, an accountable institution, a real comparison, and a measurable outcome appear together |
| A model intervention beating a trained human helper in a randomized comparison | The largest effect in the set, SNAP enrollment rising from 5.8 percent to 17.6 percent, came from human application help, and no model was tested |
| Modeled infrastructure burdens matching later observed records without adjustment | The 325 to 580 TWh range for 2028 is a scenario under stated assumptions, and the JLARC bill ranges are modeled 2040 scenarios, not observed rate changes |
| Public trust rising on messaging alone, with no change in the record | No evidence in this corpus shows a communication event causing a change in trust in either direction |

If well-designed studies start filling the left column, the thesis weakens, and this report should say so in its next edition. That's the same standard it asks of AI labs in Chapter 1 and of vendors in Chapter 11: publish the adverse result, and let it change the conclusion.

**Common misreading.** "An institutional thesis means the technology doesn't matter." It doesn't mean that. Better systems widen what's possible. The thesis says that what's possible becomes what's delivered only through institutions that measure it, and that's where the evidence shows the gap.

### T.6 The One Idea to Take With You

If I had to boil every chapter of this report down to one sentence, it would be this: **keep the denominator attached, and never mistake capability for benefit.**

Those two habits do most of the work in every chapter. Denominator discipline means any figure (cost, accuracy, time, jobs) travels with the population it measures, what it's compared with, and the limit of its interpretation, in the same sentence. "Capability is not benefit" means a system's performance on its own tells you what it can do on specified inputs, not what happens when a tired teacher, a busy physician, or a worried retiree actually uses it inside a real institution.

Here's a hypothetical to show how the two habits work together. Say a vendor tells a school district that its math assistant "scores in the top tier on benchmark problems" and "students using it completed far more practice." Denominator discipline asks: more than whom, measured on what population, over how long, and compared with which alternative? Capability-is-not-benefit asks: does more assisted practice mean more learning when the student works alone? The PNAS study in Chapter 3 shows that it can mean less. Two questions, and the pitch has to come back with evidence or it doesn't go forward.

Try the same two questions on an infrastructure claim. A developer says a project brings "1,500 jobs." Show me the denominator: is that the construction peak over 12 to 18 months, or the operating staff of about 50? A utility forecast says bills could rise. Show me the denominator: is that an observed rate increase or a modeled 2040 scenario like the JLARC range? The habit is the same whether you're a parent or a county commissioner.

### T.7 How the Thesis Was Tested: Method

#### T.7.1 Corpus

This report is a desk-research synthesis of a fixed corpus assembled for the Savrn research series of September 2026. It includes peer-reviewed randomized trials and experiments, government-commissioned evaluations and technical reports, official audits with agency responses, national statistics, systematic reviews, working papers and preprints (labeled as such), regulatory records, and publisher-described tracker methodologies. The citation ledger indexes each source. A verification pass confirmed the status of every cited URL as of the cutoff, and the results are recorded in the verification log that accompanies the manuscript. One example of what that pass caught: a dead link for the SNAP study was replaced with its canonical version, [NBER Working Paper 24652](https://www.nber.org/papers/w24652).

#### T.7.2 Selection

Studies were chosen for their relevance to decisions in the eleven chapter domains and for evidentiary strength within each domain. Older interventions that don't involve generative models were kept where they show which mechanisms have measured value. That's why Year Up and the SNAP outreach experiment are here: they set the bar a new tool has to clear. The corpus is not a systematic review of all published research. It's a curated evidence base, and this report claims no more than that.

#### T.7.3 Synthesis rules

Three rules governed every chapter:

1. **No claim exceeds its evidence class.** A simulation is never written as an observation, and a publisher statement is never written as verification.
2. **Every number carries its population, comparator, and limit** in the same breath.
3. **Favorable, null, adverse, and mixed findings get equal structural prominence.** No chapter leads with its strongest study and buries its nulls.

You can check the third rule yourself. The null physician trial sits right next to the positive mammography trial in the Seven Findings. The PNAS learning loss sits next to the Tutor CoPilot gain. The METR slowdown sits next to its update.

#### T.7.4 Writing

The prose is written for the person making a decision, not the specialist: plain constructions, contractions where they read naturally, technical terms defined at first use, and no promotional register. Where the underlying finding is uncertain, the sentence structure shows the uncertainty up front rather than asserting confidence and hedging afterward. This edition is written in my voice as founder of Savrn. The evidence descriptions are held to the same standard regardless of whose voice carries them, and my experience as an operator explains why a question matters. It never stands in for evidence.

#### T.7.5 What this report is not

- It is not personalized financial, medical, legal, or tax advice.
- It is not a site-specific engineering, fiscal, or environmental assessment.
- It is not an independent institutional review (see the disclosure below).
- It is not a claim that any proposed workflow in it has been field-tested.

The publication gates in Section 11.10 of [Chapter 11](#ai-accountability-framework) apply to any institutional use: domain review, sponsor-conflict review, local-record verification, accessibility testing, and approval by the responsible authority.

### T.8 Terminology and Conventions

A handful of terms do a lot of work in this report. Here they are in plain language.

#### T.8.1 Assistance, delegation, and substitution

- **Assistance** means a system prepares, organizes, drafts, or explains under human review.
- **Delegation** means a human authorizes a specific action.
- **Substitution** means the system performs a function a person used to perform, without per-case authorization.

The report's standard generally supports the first, conditions the second on explicit approval, and requires evidence the third almost never has. Most of the positive findings above are assistance findings. Almost none of the evidence supports substitution.

#### T.8.2 Evidence classes

Every claim is labeled with one of eight editorial evidence classes: randomized comparison, quasi-experimental analysis, observational or survey evidence, technical evaluation, simulation or forecast, administrative or official record, publisher statement, and proposed workflow. The full table, with what each class can and cannot establish, is defined once in [Chapter 11](#ai-accountability-framework). It's a labeling system, not a formal grading methodology such as GRADE. The labels organize judgment; they don't replace it.

The practical rule: a claim can't be stronger than its class. A simulation is never written up as an observation. A publisher statement, including every statement Savrn makes about its own design, is never written up as verification.

#### T.8.3 Denominator discipline

Any figure, whether cost, accuracy, time, or jobs, is reported with the population it measures, what it's compared with, and the limit of its interpretation, all in the same sentence. You've already seen this applied to the JLARC bill scenarios and the construction job peak.

#### T.8.4 Permitted and prohibited claims

Where the evidence supports a bounded positive statement, the report makes it and states its limit. Where it doesn't, the report names the claim as not permitted rather than softening it with qualifiers. This convention comes from the research package and is kept because it's clearer than graded adjectives like "promising" or "emerging."

#### T.8.5 Forbidden inference

A forbidden inference is a specific cross-layer conclusion the evidence does not support, named explicitly where it comes up. Two examples: that announced capital equals local jobs, or that a cooling figure proves zero total water use. Each of the seven Savrn trackers carries its own forbidden inference, consolidated in Section 11.9.

#### T.8.6 The seven Savrn trackers

Savrn publishes seven trackers, and they appear throughout this report:

| Tracker | What its record can support | What it cannot support alone |
|---|---|---|
| [Capital Atlas](https://savrn.com/data-center-capital-atlas) | Identification and classification of financing and capital disclosures | Local spending, jobs, tax benefit, or return |
| [Delay Watchlist](https://savrn.com/data-center-delay-tracker) | Dated evidence of identified delay and related statuses | Cause, permanent failure, or a lost household benefit |
| [Moratorium Tracker](https://savrn.com/data-center-moratorium-tracker) | Status and scope of indexed state and local actions | Complete census or cross-jurisdiction applicability |
| [Water Tracker](https://savrn.com/data-center-water-tracker) | Categorized water claims, records, and measurement boundaries | Local impact without utility and facility evidence |
| [Grid Operator Watchlist](https://savrn.com/data-center-grid-operator-watchlist) | Authorities, tariffs, dockets, queue, and curtailment records | Energization, household rate effect, or site reliability |
| [Permits and Power Development Tracker](https://savrn.com/data-center-permits-tracker) | Selected permits and development signals by stage | Construction, operation, revenue, or net benefit |
| [Scarcity Tracker](https://savrn.com/ai-index/scarcity-tracker) | A modeled signal under published assumptions | Spot price, liquid settlement, or guaranteed revenue |

Throughout, a tracker entry is treated as an evidence-discovery instrument. It starts a question and never ends one. These are Savrn products, and Savrn has a commercial interest in them, which is exactly why the boundary is written down.

#### T.8.7 Evidence cutoff

All findings, product references, and regulatory records reflect the September 23, 2026 evidence cutoff. That date is a boundary on this edition. It is not a warranty that linked pages are unchanged or that later evidence has been folded in. Tariffs, permits, and project status should be rechecked against original records before anyone relies on them for a consequential decision.

### T.9 How to Use This Report From Here

This report is a single long page with a sidebar table of contents, so nobody has to read it front to back. Each chapter stands on its own and links to the others. If you're sharing it with someone, point them to the path that matches the decision in front of them.

*Figure T.2 (interactive on the page).*
The lifespan runs from kindergarten to retirement, and the seven Savrn trackers run across it as a separate axis, because infrastructure questions show up at every stage of life, from a school budget to a retiree's electric bill. Keeping those two axes separate is deliberate: the tool-to-outcome chain and the facility-to-community chain interact, but neither substitutes for the other.

#### T.9.1 If you have twenty minutes

Read this section's thesis, argument, and seven findings, then Sections 11.2 through 11.6 of [Chapter 11](#ai-accountability-framework): the evidence classes, the claim record, the question-to-action process, the pilot protocol, and the accountability matrix. That's the operating core.

#### T.9.2 If you're deciding about a deployment

That means a school system, employer, health provider, financial firm, or public body. Read your domain's chapter first, then Chapter 11. Each domain chapter ends with a proposed evaluation suited to its setting, and the pilot protocol in Section 11.5 is the common skeleton.

| If you are a... | Start here | Then read |
|---|---|---|
| School leader, teacher, or parent of a K-12 student | [Chapter 3: Schooling, K-12](#ai-in-k12-education) | [Chapter 11](#ai-accountability-framework), Section 11.5 |
| College leader, trainer, or early career worker | [Chapter 4: College and training](#ai-college-workforce-training) | [Chapter 5](#ai-jobs-displacement) |
| Employer or worker weighing productivity claims | [Chapter 2: Capability to benefit](#ai-productivity-evidence) | [Chapter 5: Jobs and careers](#ai-jobs-displacement) |
| Household or family caregiver | [Chapter 6: Households](#ai-for-families-households) | [Chapter 9](#ai-healthcare-aging) |
| Local official or civic volunteer | [Chapter 7: Civic decisions](#ai-local-government-civic) | [Chapter 10](#data-centers-community-impact) |
| Financial adviser or saver | [Chapter 8: Financial advice](#ai-financial-advice) | [Chapter 11](#ai-accountability-framework) |
| Clinician, older adult, or health system leader | [Chapter 9: Health and aging](#ai-healthcare-aging) | [Chapter 2](#ai-productivity-evidence) |
| Investor, board member, or anyone judging AI claims in public | [Chapter 1: Public trust](#ai-safety-public-trust) | [Chapter 11](#ai-accountability-framework) |

#### T.9.3 If you're a researcher or journalist auditing a claim

Go to the chapter's figures, then to the [Consolidated Evidence and Claim Register](#evidence-register) just above this section. Every consequential claim in this report traces to a register entry, and every register entry states its primary limit. The citation ledger in the research package indexes the source-linked passages for audit, covering every one of the 83 distinct sources cited in this report. Repeated citations of one study are not independent confirmations.

#### T.9.4 If you're a resident facing a local infrastructure decision

Chapter 10's six-question power test and five-claim water boundary table are designed to be used in a public meeting. Chapter 7's eight-stage civic workflow applies to any proposal on the agenda. Bring them. Ask the questions out loud. That includes asking them of Savrn.

#### T.9.5 A note on style

Numbers in this report travel with their populations, comparators, and limits in the same breath. The text distinguishes observed facts, modeled scenarios, and proposed workflows throughout. Where the evidence is thin, it says so instead of reaching for an adjective. That isn't caution for its own sake. It's the mechanism that keeps the strong claims strong.

### T.10 Sponsorship Disclosure

This report was prepared for Savrn, and it concerns an industry in which Savrn has a commercial interest, including through the seven Savrn trackers. I'm the founder and CEO. You should read it knowing that.

Here's what that disclosure means in practice. This report does not describe itself as an independent institutional review. No outside peer-review panel is represented as having approved it. The publication policy it proposes, preserving material unfavorable findings even when they weaken a commercial narrative, has been applied throughout. Savrn's publisher role and commercial interest are stated wherever tracker findings support an argument. The sponsor-conflict review listed among the publication gates in Section 11.10 is an external requirement for any institution that wants to use this report. It's not a step the report can perform on itself.

The full disclosure sits with the [Consolidated Evidence and Claim Register](#evidence-register) directly above this section.

> **Operator's note: Why I'd publish evidence against my own industry**
>
> Operator first means no infrastructure without customers. The customer for this report is the reader who has to make a real decision: a teacher choosing a tool, a worker weighing a job change, a county commissioner voting on a rezoning. If I hand that reader a report with the nulls sanded off, I've built capacity nobody can safely use. The findings that cut against the industry, the modeled bills treated as facts, the benchmark scores that didn't translate, are the load tests. A structure you won't load test isn't finished.
### T.11 Frequently Asked Questions

#### What is the central thesis of this AI research report?

That whether the AI transition goes well is, right now, an institutional question. Across eleven chapters, measured gains showed up where a defined barrier, an accountable institution, a real comparison, and a measurable outcome appeared together, and nulls showed up where they didn't. The outcome depends on whether claims stay attached to their evidence, their limits, and the people authorized to act on them.

#### What is this AI research report 2026 about?

It's a desk-research synthesis of a fixed evidence corpus, assembled as of September 23, 2026, that follows AI through the human lifespan: schooling, training, work, households, civic life, finance, health, and the infrastructure behind it. It reports favorable, null, and adverse findings with equal prominence and closes with a proposed, untested standard for evidence and accountability.

#### What evidence would prove the thesis wrong?

Well-designed studies showing that standalone model scores reliably predict what users achieve, that assisted practice reliably becomes unaided learning, or that a model beats a trained human helper in a randomized comparison. Right now the evidence runs the other way: a 50-physician trial found a statistically indistinguishable two-point gain, and the largest benefits effect came from human application help.

#### Does AI improve productivity according to research?

In defined tasks, sometimes. A preregistered experiment with 453 professionals found writing tasks done about 40 percent faster with about 18 percent higher assessed quality, but it measured task outcomes, not earnings or firm output. METR found assistance slowed 16 experienced developers in early 2025, and a later update with 57 moved toward speedups while measurement problems stayed unresolved.

#### Is this report independent?

No. It was prepared for Savrn, which has a commercial interest in the data center industry and publishes the seven Savrn trackers. It does not describe itself as an independent institutional review, and no outside peer-review panel approved it. It does apply a policy of keeping unfavorable findings, and it lists a sponsor-conflict review as an external requirement for institutional use.

#### Does AI help students learn?

It depends on the design. Tutor CoPilot improved tutor behavior in a structured tutoring nonprofit. But a high school mathematics experiment published in PNAS found that unrestricted assistance improved assisted practice while reducing later unaided performance. Better homework output is not the same as better learning, and the evidence shows those two can move in opposite directions.

#### Can AI diagnose better than doctors?

The evidence here doesn't support that claim. In a 50-physician trial using diagnostic vignettes, the model alone scored well, but physicians with access gained a statistically indistinguishable two points. The strongest clinical result, the MASAI trial of 105,934 women, involved a specific screening system inside a radiologist workflow, with sensitivity of 80.5 percent versus 73.8 percent.

#### Will data centers raise my electric bill?

This report can't answer that for your household. Virginia JLARC scenarios show roughly $14 to $37 per month by 2040 under examined assumptions, which are modeled ranges, not present-tense rate increases. Nationally, LBNL estimated 176 TWh in 2023 and a scenario range of 325 to 580 TWh by 2028. Your answer depends on your utility's tariff and local records.

#### What does "denominator discipline" mean?

It means every figure, whether cost, accuracy, time, or jobs, is reported with the population it measures, what it's compared with, and the limit of its interpretation, in the same sentence. A construction peak of about 1,500 workers over 12 to 18 months, for example, is not an operating staff of about 50, and program spending is not chatbot-only cost.

#### What is the evidence cutoff for this report?

September 23, 2026. All findings, product references, and regulatory records reflect that date. It's a boundary on this edition, not a guarantee that linked pages haven't changed or that later evidence is included. Tariffs, permits, and project statuses should be rechecked against the original records before any consequential decision.
