SAVRN
Search Contact SAVRN

Savrn Insights · The Superintelligence Transition · The Research Thesis

Disclosure: prepared for Savrn, which has a commercial interest in AI infrastructure and publishes the seven trackers cited here. Not an independent institutional review. Not investment, medical, legal, or tax advice.

Research Thesis · The Research Thesis

The Research Thesis: Why the AI Transition Is an Institutional Question

28 min read6,476 wordsOpen as its own page

A copper fraction sculpture, a cube resting on a bar over a wide block, beside the same cube floating alone in blueprint lines
Savrn Insights · Research Thesis illustration

This is the last section of the report, and it's the one I'd hand you first if you asked me what the whole thing argues. Eleven chapters and more than a hundred thousand words come down to one claim, and it isn't the claim you'd expect from someone who builds AI infrastructure for a living. I'm not arguing that the transition will go well. I'm not arguing that it will go badly. I'm arguing that, right now, it's an institutional question. It gets decided by whether the people and organizations that adopt these systems keep every claim attached to its evidence, its limits, and the person accountable for acting on it.

That's the research thesis behind this AI research report 2026. Everything above this section is the evidence for it. What follows is the argument laid out in order: why this research exists, the premises the thesis rests on, the seven findings that carry it, what would prove it wrong, how it was tested, and the terms and disclosures you need to judge it yourself.

T.1 The Thesis in One Paragraph#

The move toward increasingly capable AI systems usually gets told as a story about the systems: what they can do, how fast they're improving, and when they might cross some threshold. This report tells it as a story about people. It follows the human life stages those systems will actually pass through, from a child's first years of school to an older adult's retirement and care, and asks one question at each stage: do people end up better off, and how do we know? The answer the evidence gives is consistent. The gains are real, measurable, and smaller and more conditional than the public conversation suggests. Some well-designed uses help. Some well-performing tools did nothing measurable once real people used them. At least one made later learning worse. And the loudest numbers in the infrastructure debate are often forecasts dressed as facts. Across all of it, the difference between a gain and a null was rarely the model. It was the institution around the model. That's the thesis.

Figure T.1Data

Headline findings, each with its limit attached

FAVORABLE40%less time, 453 professionalsWriting tasks (Noy andZhang)Limit: task outcome, no unaidedretest, no earningsFAVORABLE80.5%vs 73.8% sensitivityMASAI mammographyLimit: noninferior, not provensuperior; no mortalityFAVORABLE17.6%vs 5.8% SNAP enrollmentInformation plus humanhelpLimit: older eligiblePennsylvania adults; not aSMALL OR QUALIFIED15%more issues resolved per hourCustomer-support assistantLimit: one firm,quasi-experimental, not a wageNULL OR UNCERTAIN+2 pts95% CI -4 to 8, 50 physiciansPhysician diagnosticreasoningLimit: vignettes; notsignificantNULL OR UNCERTAIN~0earnings and hours, DenmarkTwo years after ChatGPTLimit: rules out effects aboveabout 2%; working paperADVERSE-17%later unaided exam performanceUnrestricted high-schoolmath helpLimit: one subject andconfigurationADVERSE+0.168loneliness on a 0 to 10 scaleDaily personalconversationsLimit: one-month encouragementdesignFavorableSmall or qualifiedNull or uncertainAdverse
Eight findings from across the report. Color marks direction, not importance. Every number is shown with the population and the limit that travel with it.

Source: Noy and Zhang · MASAI · NBER w24652 · QJE · JAMA Netw Open · NBER w33777 · PNAS · CESifo

T.2 Why This Research Exists#

I've spent my career building large power infrastructure, and now I build Savrn out of Dallas. Savrn designs and builds AI factories, data centers built around behind-the-meter power and closed-loop cooling. You might have expected a report with my name on it to tell you the machines are coming and everything will be fine. It doesn't. It puts favorable, null, and adverse findings side by side, and several of them cut against the easiest talking points my own industry uses. I asked for it that way on purpose.

When you build at industrial scale, you learn one rule early: the spreadsheet does not get a vote. The transformer either holds the load or it doesn't. The water either comes back or it doesn't. A county either trusts you after the first public meeting or it doesn't. Operators who believe their own pitch decks don't stay operators very long.

"Operator first" is how we run Savrn. It means no infrastructure without customers, and vertical integration as a defense against depending on things you can't control. Applied to evidence, operator first means something just as plain. Don't build a conclusion before you have the demand for it, meaning the actual question a real person is trying to answer. Don't borrow certainty from someone else's benchmark. And own every link in the chain between a claim and the person who has to act on it, because the weak link is the one that fails under load.

So the question this research was built to answer is simple to state and hard to answer well: when AI shows up in a classroom, a job, a kitchen table budget, a doctor's office, a county commission agenda, or a retirement plan, what do we actually know about whether people end up better off?

The evidence base is a fixed corpus assembled as of September 23, 2026. It includes randomized trials, government evaluations, official audits, national statistics, and documented regulatory cases. Eleven chapters walk the lifespan:

  1. Public trust and the framing of the transition
  2. How capability translates, or fails to translate, into benefit
  3. Kindergarten through high school
  4. College, training, and the first years of work
  5. Working life and career change
  6. Households and family life
  7. Neighborhoods and civic decisions
  8. Investment advice and lifetime finance
  9. Health, aging, and retirement
  10. The physical infrastructure and its host communities
  11. A shared standard for decisions and accountability

T.3 The Argument, Step by Step#

A thesis is only as strong as the steps under it. Here are the five premises this one stands on, in the order the chapters establish them. Each premise is carried by specific studies, and each study keeps its limit.

  1. Capability is real, and it's measurable. In defined tasks, assistance has produced measured gains: faster professional writing, improved tutor behavior, higher screening sensitivity in a supervised radiology workflow.
  2. Capability does not travel on its own. A model's standalone performance is not the performance of the professional, student, or household that uses it. Between the two sits a chain of human and institutional links, and some of them are weak.
  3. Benefit concentrates where institutions are strong. Measured gains show up where a defined barrier, an accountable institution, a real comparison condition, and a measurable outcome appear together.
  4. Claims lose their denominators in transit. Between the study and the headline, the population, the comparator, and the limit fall away, and a bounded finding becomes a general promise or a general threat.
  5. Trust follows the evidence state. Public trust attaches to inspectable records, published adverse findings, and enforceable commitments, not to reassurance.

Conclusion. If capability is real but doesn't travel on its own, if benefit depends on the institution, if claims shed their limits in transit, and if trust follows the record, then the outcome of the transition is decided by institutional discipline: whether claims stay tied to their evidence, their limits, and the people authorized to act on them. That's not a hedge. It's a practical claim about where the leverage is. The leverage is in the process around the tool, and that process is something a school board, an employer, a hospital, a county, and a household can control.

The seven findings below are how the chapters support those five premises.

T.4 The Seven Findings That Carry the Thesis#

The synthesis compresses eleven chapters into seven findings. Each one carries its numbers exactly as the underlying studies report them, with the population, the comparison, and the limit kept in the same place. If you jumped straight here from the table of contents, this is the fastest way into this AI evidence review.

T.4.1 The transition is measurable, and the measurements are smaller than the discourse#

The strongest causal evidence in this report is real but bounded. Three studies anchor that claim, and each one comes with a limit you have to carry along with the headline.

These are real findings with real limits, and this report treats every one of them as both. The writing result tells you what happened on a set of tasks. It does not tell you that a company's output rose or that anybody got a raise. The mammography result tells you what a particular system did inside a particular clinical process run by radiologists. It does not tell you a chatbot can read your scan. The SNAP result tells you that a human who helps fill out the form triples enrollment compared with doing nothing. It tells you nothing about software, because no software was tested.

Here's the part most people skip. The biggest effect in this set, the SNAP jump, came from people helping people. That doesn't make AI irrelevant to benefits access. It tells you what the bar is: a new tool has to beat an arrangement that already works, not a straw man.

What this means for you. If you're an employer, the writing study justifies a task-level pilot, not a headcount plan. If you're on a hospital board, MASAI justifies looking at supervised screening workflows, not replacing clinical judgment. If you run a benefits program, the SNAP result tells you the comparison group your AI pilot must beat is a trained human helper, not an empty inbox. Deeper treatment sits in Chapter 2, Chapter 6, and Chapter 9.

T.4.2 Capability and benefit are separated by a chain with weak links#

A model's standalone performance is not the performance of the professional, student, or household that uses it. That pattern repeats across the whole corpus.

Read that again: a strong model, a small and uncertain gain for doctors. A helpful tutor, weaker learning for students. A slowdown that turned into a possible speedup when the tools and tasks changed. Capability is not benefit. Any institution that buys capability on the basis of standalone benchmarks is buying the wrong outcome.

Common misreading. "The METR update proves the slowdown was wrong." It doesn't. It proves the number moved when the tools and the task selection moved, and that an estimate needs a date on it. That's why the report treats METR as technical evaluation, whose meaning shifts with changing tools. Chapter 5 walks through it.

What this means for you. A parent or teacher should ask whether a homework tool is being judged by the work it helps produce or by what the student can do alone a week later. Those are different outcomes, and the PNAS study shows they can move in opposite directions. Chapter 3 covers it grade band by grade band.

T.4.3 The biggest institutional failure is denominator discipline#

The lesson this report repeats most often is technical and unglamorous: a number without its population, comparator, and limit attached is not information. A number without a denominator is a rumor. Four examples from the corpus show how it goes wrong.

Claim as it often circulates What the record actually supports
"The city's chatbot cost this much." Program-level spending is not chatbot-only cost (New York City Comptroller MyCity audit)
"Data centers will raise your bill by this much." Modeled 2040 bill ranges are not present-tense rate increases; the Virginia JLARC scenarios give roughly $14 to $37 per month under examined assumptions
"The project will create 1,500 jobs." A construction peak of about 1,500 workers over 12 to 18 months is not an operating staff of about 50
"Scores went up after we adopted the tool." A within-group improvement is not an effect

This report proposes denominator discipline as an operating standard, not a writing style. Every consequential claim should carry a claim record with its evidence class attached, so a different reviewer can see what the number counts, what it's compared against, and where it stops being true. The claim record is laid out in Chapter 11, and the MyCity dispute gets full treatment in Chapter 7.

T.4.4 Assistance helps most where the institution is strongest#

Across every chapter, measured benefit concentrates where four conditions show up together:

  1. A defined barrier
  2. An accountable institution
  3. A real comparison condition
  4. A measurable outcome

Tutor CoPilot worked in a structured tutoring nonprofit. The MASAI system worked inside a radiologist workflow. Year Up worked as a comprehensive program package with earnings gains sustained to year ten, and it was not a model intervention at all.

The implication cuts against both the hype and the dismissal. The same tool that transforms one setting can do nothing measurable in another, and the difference is usually institutional, not technical. That's the whole game for anyone deciding about a deployment: the question isn't "is this model good?" but "is our institution the kind where assistance has been shown to work, and can we measure whether it did?"

What this means for you. A school board, workforce board, or employer should spend as much time on the program around the tool as on the tool. Year Up is in this report precisely because it shows what a sustained, measured pathway gain looks like, and none of it came from a model. See Chapter 4 for the pathway evidence and Chapter 3 for Tutor CoPilot.

T.4.5 Public trust is an evidence state, not a messaging problem#

In 2026 survey data published in August 2026, Pew Research Center found 52 percent of Americans more concerned than excited about AI in daily life, and 9 percent more excited than concerned. The 2023 to 2024 public record shows why trust has to be earned structurally. The superalignment team was dissolved within days of its leaders' departures in May 2024, and a resignation letter saying "safety culture and processes have taken a backseat to shiny products" entered the public record.

Trust attaches to inspectable denominators, published adverse findings, and enforceable commitments. It doesn't attach to reassurance. I'll say that plainly as someone whose company lives or dies on community trust: you can't message your way out of a record. Chapter 1 lays out the chronology and a proposed public communication standard.

Common misreading. "People are worried because they don't understand the technology." The survey measures attitudes; it doesn't measure understanding, and it doesn't show that any single communication event caused the concern. Read with the Chapter 1 record, it describes a public waiting to be shown evidence, not one waiting for a better slogan.

T.4.6 The infrastructure chapter is this report's test of its own method#

This is the chapter where Savrn has the most to gain or lose, so it's where the method has to hold hardest. The application benefits documented in Chapters 2 through 6 say nothing about whether a specific data center benefits its host community. The report refuses that inference in both directions: a useful chatbot does not justify a facility, and a bad chatbot does not condemn one.

What the evidence does support:

  • National electricity estimates with stated scenario ranges. Lawrence Berkeley National Laboratory estimated 176 TWh, or 4.4 percent of national consumption, in 2023, and 325 to 580 TWh, or 6.7 to 12 percent, by 2028 under stated assumptions. The first is an estimate of the past. The second is a scenario range, not a forecast of your county.
  • Documented allocation mechanisms. The AEP Ohio data center tariff uses a minimum-commitment architecture that shows how cost risk can be assigned to large new loads.
  • A proposed community benefit compact. Its enforceability test is simple: an adverse result must be able to change the project. A compact that can't change anything isn't a compact.

T.4.7 The shared standard is the deliverable#

A thesis about institutions has to hand institutions something they can use. Chapter 11 consolidates this research into an operating process:

  • an eight-class evidence labeling system
  • a ten-field claim record
  • a five-stage question-to-action process that keeps drafting and action separate
  • an eleven-element pilot protocol
  • a six-role accountability matrix
  • the seven-tracker boundary: a tracker entry starts a question; it never ends one

None of it is validated. All of it is testable. The report says so on every page where it appears, and I'm saying it here too. If you take nothing else from this report, take the process, test it in your own institution, and publish what happens, including the parts that don't work.

T.5 What Would Prove This Thesis Wrong#

A research thesis that can't lose isn't a thesis. It's a slogan. So here's what the evidence would have to show, in future studies, for this one to fail. Each test is stated against a finding already in this report, so you can check any new study against it.

If the thesis is wrong, you would expect to see What this report's evidence currently shows
Standalone model scores reliably predicting what users achieve with the model In the 50-physician vignette trial, a strong standalone model and a statistically indistinguishable two-point gain for physicians with access
Assisted practice reliably turning into unaided learning In the PNAS high school mathematics experiment, better assisted practice and weaker later unaided performance
Effects that hold across settings regardless of the institution around the tool Gains concentrated where a defined barrier, an accountable institution, a real comparison, and a measurable outcome appear together
A model intervention beating a trained human helper in a randomized comparison The largest effect in the set, SNAP enrollment rising from 5.8 percent to 17.6 percent, came from human application help, and no model was tested
Modeled infrastructure burdens matching later observed records without adjustment The 325 to 580 TWh range for 2028 is a scenario under stated assumptions, and the JLARC bill ranges are modeled 2040 scenarios, not observed rate changes
Public trust rising on messaging alone, with no change in the record No evidence in this corpus shows a communication event causing a change in trust in either direction

If well-designed studies start filling the left column, the thesis weakens, and this report should say so in its next edition. That's the same standard it asks of AI labs in Chapter 1 and of vendors in Chapter 11: publish the adverse result, and let it change the conclusion.

Common misreading. "An institutional thesis means the technology doesn't matter." It doesn't mean that. Better systems widen what's possible. The thesis says that what's possible becomes what's delivered only through institutions that measure it, and that's where the evidence shows the gap.

T.6 The One Idea to Take With You#

If I had to boil every chapter of this report down to one sentence, it would be this: keep the denominator attached, and never mistake capability for benefit.

Those two habits do most of the work in every chapter. Denominator discipline means any figure (cost, accuracy, time, jobs) travels with the population it measures, what it's compared with, and the limit of its interpretation, in the same sentence. "Capability is not benefit" means a system's performance on its own tells you what it can do on specified inputs, not what happens when a tired teacher, a busy physician, or a worried retiree actually uses it inside a real institution.

Here's a hypothetical to show how the two habits work together. Say a vendor tells a school district that its math assistant "scores in the top tier on benchmark problems" and "students using it completed far more practice." Denominator discipline asks: more than whom, measured on what population, over how long, and compared with which alternative? Capability-is-not-benefit asks: does more assisted practice mean more learning when the student works alone? The PNAS study in Chapter 3 shows that it can mean less. Two questions, and the pitch has to come back with evidence or it doesn't go forward.

Try the same two questions on an infrastructure claim. A developer says a project brings "1,500 jobs." Show me the denominator: is that the construction peak over 12 to 18 months, or the operating staff of about 50? A utility forecast says bills could rise. Show me the denominator: is that an observed rate increase or a modeled 2040 scenario like the JLARC range? The habit is the same whether you're a parent or a county commissioner.

T.7 How the Thesis Was Tested: Method#

T.7.1 Corpus#

This report is a desk-research synthesis of a fixed corpus assembled for the Savrn research series of September 2026. It includes peer-reviewed randomized trials and experiments, government-commissioned evaluations and technical reports, official audits with agency responses, national statistics, systematic reviews, working papers and preprints (labeled as such), regulatory records, and publisher-described tracker methodologies. The citation ledger indexes each source. A verification pass confirmed the status of every cited URL as of the cutoff, and the results are recorded in the verification log that accompanies the manuscript. One example of what that pass caught: a dead link for the SNAP study was replaced with its canonical version, NBER Working Paper 24652.

T.7.2 Selection#

Studies were chosen for their relevance to decisions in the eleven chapter domains and for evidentiary strength within each domain. Older interventions that don't involve generative models were kept where they show which mechanisms have measured value. That's why Year Up and the SNAP outreach experiment are here: they set the bar a new tool has to clear. The corpus is not a systematic review of all published research. It's a curated evidence base, and this report claims no more than that.

T.7.3 Synthesis rules#

Three rules governed every chapter:

  1. No claim exceeds its evidence class. A simulation is never written as an observation, and a publisher statement is never written as verification.
  2. Every number carries its population, comparator, and limit in the same breath.
  3. Favorable, null, adverse, and mixed findings get equal structural prominence. No chapter leads with its strongest study and buries its nulls.

You can check the third rule yourself. The null physician trial sits right next to the positive mammography trial in the Seven Findings. The PNAS learning loss sits next to the Tutor CoPilot gain. The METR slowdown sits next to its update.

T.7.4 Writing#

The prose is written for the person making a decision, not the specialist: plain constructions, contractions where they read naturally, technical terms defined at first use, and no promotional register. Where the underlying finding is uncertain, the sentence structure shows the uncertainty up front rather than asserting confidence and hedging afterward. This edition is written in my voice as founder of Savrn. The evidence descriptions are held to the same standard regardless of whose voice carries them, and my experience as an operator explains why a question matters. It never stands in for evidence.

T.7.5 What this report is not#

  • It is not personalized financial, medical, legal, or tax advice.
  • It is not a site-specific engineering, fiscal, or environmental assessment.
  • It is not an independent institutional review (see the disclosure below).
  • It is not a claim that any proposed workflow in it has been field-tested.

The publication gates in Section 11.10 of Chapter 11 apply to any institutional use: domain review, sponsor-conflict review, local-record verification, accessibility testing, and approval by the responsible authority.

T.8 Terminology and Conventions#

A handful of terms do a lot of work in this report. Here they are in plain language.

T.8.1 Assistance, delegation, and substitution#

  • Assistance means a system prepares, organizes, drafts, or explains under human review.
  • Delegation means a human authorizes a specific action.
  • Substitution means the system performs a function a person used to perform, without per-case authorization.

The report's standard generally supports the first, conditions the second on explicit approval, and requires evidence the third almost never has. Most of the positive findings above are assistance findings. Almost none of the evidence supports substitution.

T.8.2 Evidence classes#

Every claim is labeled with one of eight editorial evidence classes: randomized comparison, quasi-experimental analysis, observational or survey evidence, technical evaluation, simulation or forecast, administrative or official record, publisher statement, and proposed workflow. The full table, with what each class can and cannot establish, is defined once in Chapter 11. It's a labeling system, not a formal grading methodology such as GRADE. The labels organize judgment; they don't replace it.

The practical rule: a claim can't be stronger than its class. A simulation is never written up as an observation. A publisher statement, including every statement Savrn makes about its own design, is never written up as verification.

T.8.3 Denominator discipline#

Any figure, whether cost, accuracy, time, or jobs, is reported with the population it measures, what it's compared with, and the limit of its interpretation, all in the same sentence. You've already seen this applied to the JLARC bill scenarios and the construction job peak.

T.8.4 Permitted and prohibited claims#

Where the evidence supports a bounded positive statement, the report makes it and states its limit. Where it doesn't, the report names the claim as not permitted rather than softening it with qualifiers. This convention comes from the research package and is kept because it's clearer than graded adjectives like "promising" or "emerging."

T.8.5 Forbidden inference#

A forbidden inference is a specific cross-layer conclusion the evidence does not support, named explicitly where it comes up. Two examples: that announced capital equals local jobs, or that a cooling figure proves zero total water use. Each of the seven Savrn trackers carries its own forbidden inference, consolidated in Section 11.9.

T.8.6 The seven Savrn trackers#

Savrn publishes seven trackers, and they appear throughout this report:

Tracker What its record can support What it cannot support alone
Capital Atlas Identification and classification of financing and capital disclosures Local spending, jobs, tax benefit, or return
Delay Watchlist Dated evidence of identified delay and related statuses Cause, permanent failure, or a lost household benefit
Moratorium Tracker Status and scope of indexed state and local actions Complete census or cross-jurisdiction applicability
Water Tracker Categorized water claims, records, and measurement boundaries Local impact without utility and facility evidence
Grid Operator Watchlist Authorities, tariffs, dockets, queue, and curtailment records Energization, household rate effect, or site reliability
Permits and Power Development Tracker Selected permits and development signals by stage Construction, operation, revenue, or net benefit
Scarcity Tracker A modeled signal under published assumptions Spot price, liquid settlement, or guaranteed revenue

Throughout, a tracker entry is treated as an evidence-discovery instrument. It starts a question and never ends one. These are Savrn products, and Savrn has a commercial interest in them, which is exactly why the boundary is written down.

T.8.7 Evidence cutoff#

All findings, product references, and regulatory records reflect the September 23, 2026 evidence cutoff. That date is a boundary on this edition. It is not a warranty that linked pages are unchanged or that later evidence has been folded in. Tariffs, permits, and project status should be rechecked against original records before anyone relies on them for a consequential decision.

T.9 How to Use This Report From Here#

This report is a single long page with a sidebar table of contents, so nobody has to read it front to back. Each chapter stands on its own and links to the others. If you're sharing it with someone, point them to the path that matches the decision in front of them.

Figure T.2Map

How the report is organized: one lifespan, one infrastructure layer, one standard

THE LIFESPANK-4Ch 3Grades 5-8Ch 3Grades 9-12Ch 3College andtradesCh 4WorkinglifeCh 5HouseholdCh 6Civic lifeCh 7MoneyCh 8Health andagingCh 9CROSS-CUTTINGChapter 1 public trust · Chapter 2 standard of proof · Chapter 11 shared accountability standardTHE SEVEN SAVRN TRACKERSCapital Atlas · Delay · Moratorium · Water · Grid Operator · Permits and Power · Scarcity (a tracker entry starts a question)CHAPTER 10 · THE PHYSICAL INFRASTRUCTURE UNDER ALL OF ITPower, water, jobs, taxes, noise. A separate evidence chain that never substitutes for the application evidence above.
The report walks the lifespan in order, from kindergarten to retirement, then turns to the physical infrastructure and the shared standard. The seven Savrn trackers run across the lifespan as evidence discovery tools, never as proof of outcomes.

The lifespan runs from kindergarten to retirement, and the seven Savrn trackers run across it as a separate axis, because infrastructure questions show up at every stage of life, from a school budget to a retiree's electric bill. Keeping those two axes separate is deliberate: the tool-to-outcome chain and the facility-to-community chain interact, but neither substitutes for the other.

T.9.1 If you have twenty minutes#

Read this section's thesis, argument, and seven findings, then Sections 11.2 through 11.6 of Chapter 11: the evidence classes, the claim record, the question-to-action process, the pilot protocol, and the accountability matrix. That's the operating core.

T.9.2 If you're deciding about a deployment#

That means a school system, employer, health provider, financial firm, or public body. Read your domain's chapter first, then Chapter 11. Each domain chapter ends with a proposed evaluation suited to its setting, and the pilot protocol in Section 11.5 is the common skeleton.

If you are a... Start here Then read
School leader, teacher, or parent of a K-12 student Chapter 3: Schooling, K-12 Chapter 11, Section 11.5
College leader, trainer, or early career worker Chapter 4: College and training Chapter 5
Employer or worker weighing productivity claims Chapter 2: Capability to benefit Chapter 5: Jobs and careers
Household or family caregiver Chapter 6: Households Chapter 9
Local official or civic volunteer Chapter 7: Civic decisions Chapter 10
Financial adviser or saver Chapter 8: Financial advice Chapter 11
Clinician, older adult, or health system leader Chapter 9: Health and aging Chapter 2
Investor, board member, or anyone judging AI claims in public Chapter 1: Public trust Chapter 11

T.9.3 If you're a researcher or journalist auditing a claim#

Go to the chapter's figures, then to the Consolidated Evidence and Claim Register just above this section. Every consequential claim in this report traces to a register entry, and every register entry states its primary limit. The citation ledger in the research package indexes the source-linked passages for audit, covering every one of the 83 distinct sources cited in this report. Repeated citations of one study are not independent confirmations.

T.9.4 If you're a resident facing a local infrastructure decision#

Chapter 10's six-question power test and five-claim water boundary table are designed to be used in a public meeting. Chapter 7's eight-stage civic workflow applies to any proposal on the agenda. Bring them. Ask the questions out loud. That includes asking them of Savrn.

T.9.5 A note on style#

Numbers in this report travel with their populations, comparators, and limits in the same breath. The text distinguishes observed facts, modeled scenarios, and proposed workflows throughout. Where the evidence is thin, it says so instead of reaching for an adjective. That isn't caution for its own sake. It's the mechanism that keeps the strong claims strong.

T.10 Sponsorship Disclosure#

This report was prepared for Savrn, and it concerns an industry in which Savrn has a commercial interest, including through the seven Savrn trackers. I'm the founder and CEO. You should read it knowing that.

Here's what that disclosure means in practice. This report does not describe itself as an independent institutional review. No outside peer-review panel is represented as having approved it. The publication policy it proposes, preserving material unfavorable findings even when they weaken a commercial narrative, has been applied throughout. Savrn's publisher role and commercial interest are stated wherever tracker findings support an argument. The sponsor-conflict review listed among the publication gates in Section 11.10 is an external requirement for any institution that wants to use this report. It's not a step the report can perform on itself.

The full disclosure sits with the Consolidated Evidence and Claim Register directly above this section.

T.11 Frequently Asked Questions#

What is the central thesis of this AI research report?

That whether the AI transition goes well is, right now, an institutional question. Across eleven chapters, measured gains showed up where a defined barrier, an accountable institution, a real comparison, and a measurable outcome appeared together, and nulls showed up where they didn't. The outcome depends on whether claims stay attached to their evidence, their limits, and the people authorized to act on them.

What is this AI research report 2026 about?

It's a desk-research synthesis of a fixed evidence corpus, assembled as of September 23, 2026, that follows AI through the human lifespan: schooling, training, work, households, civic life, finance, health, and the infrastructure behind it. It reports favorable, null, and adverse findings with equal prominence and closes with a proposed, untested standard for evidence and accountability.

What evidence would prove the thesis wrong?

Well-designed studies showing that standalone model scores reliably predict what users achieve, that assisted practice reliably becomes unaided learning, or that a model beats a trained human helper in a randomized comparison. Right now the evidence runs the other way: a 50-physician trial found a statistically indistinguishable two-point gain, and the largest benefits effect came from human application help.

Does AI improve productivity according to research?

In defined tasks, sometimes. A preregistered experiment with 453 professionals found writing tasks done about 40 percent faster with about 18 percent higher assessed quality, but it measured task outcomes, not earnings or firm output. METR found assistance slowed 16 experienced developers in early 2025, and a later update with 57 moved toward speedups while measurement problems stayed unresolved.

Is this report independent?

No. It was prepared for Savrn, which has a commercial interest in the data center industry and publishes the seven Savrn trackers. It does not describe itself as an independent institutional review, and no outside peer-review panel approved it. It does apply a policy of keeping unfavorable findings, and it lists a sponsor-conflict review as an external requirement for institutional use.

Does AI help students learn?

It depends on the design. Tutor CoPilot improved tutor behavior in a structured tutoring nonprofit. But a high school mathematics experiment published in PNAS found that unrestricted assistance improved assisted practice while reducing later unaided performance. Better homework output is not the same as better learning, and the evidence shows those two can move in opposite directions.

Can AI diagnose better than doctors?

The evidence here doesn't support that claim. In a 50-physician trial using diagnostic vignettes, the model alone scored well, but physicians with access gained a statistically indistinguishable two points. The strongest clinical result, the MASAI trial of 105,934 women, involved a specific screening system inside a radiologist workflow, with sensitivity of 80.5 percent versus 73.8 percent.

Will data centers raise my electric bill?

This report can't answer that for your household. Virginia JLARC scenarios show roughly $14 to $37 per month by 2040 under examined assumptions, which are modeled ranges, not present-tense rate increases. Nationally, LBNL estimated 176 TWh in 2023 and a scenario range of 325 to 580 TWh by 2028. Your answer depends on your utility's tariff and local records.

What does "denominator discipline" mean?

It means every figure, whether cost, accuracy, time, or jobs, is reported with the population it measures, what it's compared with, and the limit of its interpretation, in the same sentence. A construction peak of about 1,500 workers over 12 to 18 months, for example, is not an operating staff of about 50, and program spending is not chatbot-only cost.

What is the evidence cutoff for this report?

September 23, 2026. All findings, product references, and regulatory records reflect that date. It's a boundary on this edition, not a guarantee that linked pages haven't changed or that later evidence is included. Tariffs, permits, and project statuses should be rechecked against the original records before any consequential decision.

Keep listening

A copper studio microphone and headphones on a stack of books, with blueprint sound waves flowing into a sketched open book

The audio edition · Read by Bella

Listen to the whole report

13 episodes, 2 h 26 min. Each one walks a chapter's argument, its studies, and their limits, so you can listen instead of read. It plays straight through, one chapter into the next.

Up next · Chapter 1

AI Safety Promises and Public Trust: How to Test What AI Labs Say

0:00 / 11:57
  1. Read
  2. Read
  3. Read
  4. Read
  5. Read
  6. Read
  7. Read
  8. Read
  9. Read
  10. Read
  11. Read
  12. Read
  13. Read