Chapter 1 · Public Trust and AI Safety
AI Safety Promises and Public Trust: How to Test What AI Labs Say

I have spent my career around industries that make big promises about power and pace. Utilities promise capacity. Developers promise timelines. Equipment vendors promise uptime. Every one of those promises sounds solid in the press release, and every one of them either holds up against a number, a date, and a penalty, or it doesn't. So when the companies building the most capable AI systems in the world publish essays about AI safety commitments and then ship new models days later, I don't think your real question is whether they're heroes or hypocrites. Your real question is simpler and harder: should I believe them?
A parent deciding whether a tutoring app belongs in the house, a teacher weighing a classroom tool, a county commissioner looking at a power contract, a worker whose job is being redesigned around software: every one of them is being asked, in one form or another, to take a company's word about what a system does and how carefully it was released.
Here's what the evidence in this chapter will show, including the uncomfortable parts. The September 2026 essays do not forbid releases, so a release is not automatically a broken promise. But the longest public record we have of a lab making a specific, quantified safety pledge, OpenAI's 2023 superalignment commitment, was not carried out as stated, according to reporting by several outlets. The public knows something is off: Pew's surveys show concern climbing for five years. And the tools the public needs to judge any single release, a named baseline, an operational safety bar, an evaluator with publication rights, are mostly missing. I'm going to give you a test you can run yourself, a worksheet to run it on, and a standard for what a release explanation should contain. Then I'm going to run that same test on my own company, because a standard that exempts its author is not a standard.
1.1 AI Safety Commitments in September 2026: What Happened#
In September 2026, two of the most prominent laboratories building advanced AI systems published essays asking the industry to slow down. Days later, both companies released new products. The sequence produced the kind of headlines that write themselves: the warnings were theater, the restraint was marketing, and the releases were proof of hypocrisy. Or, from the opposite camp: the warnings were responsible leadership, the releases were routine progress, and the critics were scoring points without reading what was actually said.
This report takes neither side, because neither side can be supported by the documentary record alone. What the record does support is more interesting than either polemic. It supports a real question about the relationship between public warnings, private pacing, and product releases. It supports holding companies to the specific conditions they themselves proposed. And it supports a communication standard that would let a parent, teacher, civic member, worker, or resident judge releases on evidence rather than on trust or suspicion.
1.1.1 Why this chapter comes first#
Everything that follows in this research depends on this chapter's method. The later chapters cover kindergarten classrooms, college physics courses, customer-support queues, household kitchens, civic committee rooms, adviser offices, clinics, and the substations and water lines of the physical infrastructure that powers all of it. Every one of those settings involves a person making a decision with incomplete information.
The September 2026 record is a useful place to learn what complete information would look like, because it is the moment when the industry's own leaders told the public what should count. They handed the rest of us a grading rubric. This chapter picks it up.
It is also a useful place to learn what institutions owe the public when they ask for trust. Savrn, the company I founded and for which this report was prepared, states a purpose of improving the connection between data center infrastructure and communities, including avoiding burdens on community power and water. I treat that purpose as an objective that requires measurement and enforceable commitments, not as a verified operating claim. The same procedural standard I apply here to the laboratories applies to every institution in these pages, Savrn included.
1.1.2 What this report asks of you#
The chapters that follow are long because the questions are long. A teacher's decision about a classroom tool and a community's decision about a power contract are not questions that can be answered well in a paragraph, and I don't try. What I offer instead is a way of reading claims, any claims, from any source, that survives contact with the evidence. That skill transfers. It is the actual product of this entire report.
If you read nothing else, read Section 1.5 and the worksheet in Section 1.5.5.
1.2 Separating Four Claims: Statements, Releases, Consistency, Harm#
Public calls for restraint followed by product releases deserve scrutiny. But proximity on a calendar does not establish a broken commitment. The reviewed statements call for safety-conditioned pacing, not a blanket end to training or product releases. Dario Amodei's essay, "We Must Pace the Frontier", argues that the pace of capability improvement must slow; Jakub Pachocki's essay, "An Alien Mind", argues that scaling must be constrained by confidence in safety. Both are statements of position, and both explicitly preserve room for continued releases under stated conditions.
So the fair question is not "did a release follow a warning?" It is a narrower and harder question: does an organization's actual conduct satisfy the particular conditions it publicly proposed? Answering that requires separating four distinct claims.
| Claim | Can it be answered from the record? | What it takes |
|---|---|---|
| 1. What was said | Yes | Dated public records |
| 2. What was released | Yes | Dated public records |
| 3. Whether the two are consistent | In principle, yes | Comparing conduct against the specific conditions proposed, not against a caricature of them |
| 4. Whether the communication caused harm | Not as of September 23, 2026 | Causal evidence about comprehension, trust, and behavior that the documentary record does not contain |
The first two claims are answerable. The third is answerable in principle, with a disciplined test. The fourth, as of the September 23, 2026 evidence cutoff of this report, is not established in either direction, and a report that pretends otherwise is selling something.
1.2.1 Why collapsing the four claims misleads#
Public discussion of AI routinely collapses all four claims into one. A dramatic sequence of headlines feels like evidence. It is not. It is a sequence of headlines.
The distinction also protects critics from their own worst habit: grading companies on a commitment the company never made. If a laboratory proposes conditional pacing and ships a product that plausibly satisfies its stated conditions, the shipping is not a confession. If it proposes evaluator-confirmed safety bars and cannot name an evaluator, the proposal is a press strategy. Everything depends on the content of the commitment, which is why this chapter spends its effort on what was actually written rather than on what was felt.
1.3 A Bounded Chronology of Statements and Releases#
Four dated records anchor this chapter. Each is presented with what it verifies and what it cannot.
| Record | Verified content | Evidentiary limit |
|---|---|---|
| OpenAI chief scientist's September 6, 2026 essay | Jakub Pachocki wrote that scaling must be constrained by confidence in safety, and expressed hope that voluntary slowdowns would become commonplace until shared safety bars were established. (An Alien Mind) | A statement of position, not an independently assessed risk threshold or a commitment to stop every update. |
| Anthropic chief executive's September 2026 essay | Dario Amodei wrote, "We must slow the pace at which we improve the capabilities of AI models," while explicitly distinguishing pacing from halting training or technical progress. (We Must Pace the Frontier) | The essay's rationale and proposed evaluation arrangements are not proof that those arrangements were implemented adequately. |
| Anthropic's September 22 announcement | Anthropic announced Claude Opus 5.5 and described evaluations and safeguards associated with the release. (Anthropic's Claude Opus 5.5 announcement) | A publisher's account does not independently establish the sufficiency of its safety work. |
| OpenAI's September 22 product-page update | The GPT-6 page records the addition of GPT-6 Sol and GPT-6 Luna to the family. (OpenAI's dated GPT-6 product update) | The update date is not evidence that every model described on that page was first released on that date. |
The public record, 2023 to 2026: commitments, departures, essays, releases
Source: OpenAI, Introducing Superalignment · Business Insider · The Next Web · Pachocki · Amodei · Anthropic · OpenAI GPT-6
The chronology is bounded on purpose. It includes the records this chapter can verify and excludes the rumor, inference, and retrospective interpretation that surround them. A wider chronology could be assembled, but it would add heat, not light, to the question that matters here.
1.3.1 Two features of the chronology worth noticing#
First, the pacing essays and the releases are separated by roughly two weeks, not by years. Whatever else the sequence shows, it shows that the industry's leading safety voices expected the public to evaluate releases against a standard those voices had just articulated. That makes the consistency test in Section 1.5 a live question rather than an academic one.
Second, the two September 22 records are announcements of record, not independent assessments. Anthropic describes evaluations and safeguards associated with Claude Opus 5.5; OpenAI records additions to the GPT-6 family. What those descriptions establish is what the companies say. Whether the described evaluations are sufficient is precisely the kind of question the announcements cannot answer about themselves.
1.4 What AI Pacing Means: Conditional Commitments, Not Halts#
The most quoted sentence from Amodei's essay is the demand to slow the pace of capability improvement. The most important sentence is the clarification that follows it. In his own words:
"To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this." (Amodei's original essay)
Pachocki's essay is constructed the same way. He describes pursuing technical solutions, building defensive systems, and withholding further scaling as needed, rather than announcing an unconditional stop. (Pachocki's original essay)
Two implications follow, and both cut against the easy storylines.
First, treating every later release as an automatic contradiction is ruled out by the text itself. A company can release a product while sincerely holding the pacing position, provided the release satisfies the conditions the essay describes: adequate alignment and safeguard time, third-party evaluator confirmation. Whether any particular release did satisfy those conditions is a factual question this record cannot settle, but the position itself does not forbid releases.
Second, the pacing position is not empty. It proposes something specific and testable: that deployment decisions should be conditioned on safety evidence confirmed by evaluators who are not the company itself. A company that claims to hold this position has accepted a standard it can fail. That is what makes the position worth taking seriously, and what makes the consistency test in the next section possible.
The records therefore support scrutiny of conditional commitments. They do not, by themselves, prove either hypocrisy or satisfactory compliance.
1.4.1 The longer arc: OpenAI superalignment, 2023 to 2024#
The September 2026 essays did not arrive from nowhere, and the record of what laboratories have publicly promised is longer than a single month.
In July 2023, OpenAI announced a dedicated superalignment effort, co-led by then-chief scientist Ilya Sutskever and alignment head Jan Leike, stating that it would dedicate 20 percent of its secured compute to the problem of steering and controlling systems "much smarter than us" within four years. The announcement described superintelligence as potentially "the most impactful technology humanity has ever invented," warned that its power could lead to "the disempowerment of humanity or even human extinction," and conceded that current alignment techniques would not scale to superintelligence because humans would not be able to reliably supervise such systems. (OpenAI, Introducing Superalignment)
The 2023 announcement is worth pausing on, because it establishes the pattern the 2026 essays repeat: a laboratory publicly defining a risk, publicly committing resources against it, and explicitly acknowledging that the commitment's success is uncertain. You can weigh the sincerity and the follow-through of such commitments for yourself. What I insist on is only that the weighing happen against the text of the commitments, with dates, rather than against the mood of the moment.
1.4.2 What happened to the superalignment commitment#
The follow-through is checkable, because the 2023 commitment named a number and a deadline. What happened next is on the public record, and I'm relying on secondary reporting for it, so I'll name the reporters.
According to reporting by Business Insider, in its running account of safety-leadership departures, Newsweek, in its account of the 2026 Coxon resignation, which recaps the 2024 departures, and The Next Web, on the successor teams, the sequence ran like this:
| Date | Event (as reported) |
|---|---|
| July 2023 | OpenAI announces superalignment: 20 percent of secured compute, four years |
| May 14, 2024 | Ilya Sutskever announces his departure from OpenAI |
| May 17, 2024 | Jan Leike resigns publicly, citing safety culture and compute |
| Within days of those departures | OpenAI dissolves the Superalignment team and redistributes its remaining staff |
| October 2024 | OpenAI's AGI Readiness team dissolves when its leader, Miles Brundage, resigns |
| February 2026 | OpenAI's Mission Alignment team is disbanded after roughly sixteen months |
Leike wrote on his way out that "over the past years, safety culture and processes have taken a backseat to shiny products" and that his team had been "sailing against the wind," struggling to obtain the compute its mandate required. The 20-percent-of-compute, four-year commitment announced in July 2023 thus ended, as a dedicated team, at roughly the one-year mark. Reporting at the time, including Fortune's, as relayed in these accounts, indicated that the team's requests for its pledged compute had been denied before the dissolution. The successor structure did not stabilize either, according to the same reporting.
Read that again: a four-year commitment, with a number attached, ended as a dedicated team at about one year.
1.4.3 What the superalignment record does and does not prove#
This record does not by itself prove that OpenAI's safety work was abandoned. The company stated that safety responsibilities were redistributed rather than ended, and some departures were unattributed. Those facts belong in the same paragraph as the departures, and I'm keeping them there.
What the record does is more precise. It shows that the 2023 commitment, which was specific, dated, and quantified, was not carried out as stated, and that the people most identified with it left while saying so in public. When the same organization's chief scientist writes in September 2026 that scaling must be constrained by confidence in safety, the earlier record is not a footnote to that essay. It is the first data point for the seven-part test in Section 1.5.
The record is checkable only because the commitment had a number and a deadline. Vague commitments can't be caught falling short. That's an argument for specificity, not against it.
1.4.4 A symmetry worth recording: the Jacob Coxon resignation#
One more symmetry deserves notice. In September 2026, according to reporting by Business Standard, an Anthropic researcher named Jacob Coxon resigned and warned publicly about the direction of the frontier race. That is the same move Leike made at OpenAI in 2024, now from inside the company whose chief executive had just published the pacing essay.
I draw no conclusion from the symmetry. I record it, because a reader evaluating either company's September conduct is entitled to the full pattern of who has left these organizations, and what they said on their way out. A departure is testimony, not a verdict. It tells you where to look; it doesn't tell you what you'll find.
1.5 The Missing Denominator: A Seven-Part Test for AI Safety Commitments#
A claim of "slowing" is incomplete unless two things are identified: the activity being slowed, and the baseline it is being compared with. This sounds pedantic until you notice how often the word is used without either.
Show me the denominator. Slowing what, compared with what, over what period?
The problem is structural. "Slower" could refer to training scale, capability improvement, external access, deployment scope, autonomy, or release frequency, and these can move in opposite directions simultaneously. A laboratory can slow its training schedule while expanding its product rollout. It can ship fewer models while each model gains more capability. It can restrict access tiers while widening deployment scope within each tier. Any of these patterns would be accurately described as "slowing" by someone who chose the favorable metric, and accurately described as "accelerating" by someone who chose another.
The missing denominator: "slowing down" has to name an activity and a baseline
Fewer releases than last month does not necessarily mean slower capability growth. A calm product quarter and a jump in capability are not mutually exclusive.
1.5.1 The seven elements#
This report proposes a seven-part consistency test. It is a proposed standard, not a finding that any company has passed or failed it.
| Test element | Disclosure needed | Why the distinction matters |
|---|---|---|
| Activity | Training scale, capability improvement, external access, deployment scope, autonomy, or release frequency | A slower training schedule and a larger product rollout can coexist. |
| Baseline | Previously planned or expected activity over a stated period | Fewer releases than last month does not necessarily mean slower capability growth. |
| Safety bar | Stated criteria, test population, failure conditions, and uncertainty | "Safe enough" is not reproducible without an operational standard. |
| Evaluation | Evaluator access, independence, methods, exclusions, and publication rights | An evaluation's existence does not establish what it could test. |
| Decision | Whether evidence caused a delay, restriction, redesign, or cancellation | A process with no possible effect on deployment is not meaningful restraint. |
| Residual risk | Unresolved failures and limits on detection or monitoring | Passing selected tests does not settle every deployment context. |
| Change control | Version, access tier, tool permissions, and post-release updates | A result for one configuration need not cover a later one. |
Each element answers a failure mode in public discussion:
- Activity prevents metric shopping.
- Baseline prevents phantom slowdowns.
- Safety bar prevents "safe enough" from meaning whatever is convenient on a given day.
- Evaluation distinguishes an evaluation from a press release about an evaluation.
- Decision is the heart of the matter: a review process that has never delayed, restricted, redesigned, or cancelled anything is not restraint, whatever it is called.
- Residual risk prevents a clean evaluation sheet from being read as a clean system.
- Change control prevents a result for one model version from being silently extended to the next.
If you only remember one element, remember the decision element. That's the whole game. Any review process can produce paperwork. Only a real one can produce a "no."
1.5.2 Why the International AI Safety Report can't settle September#
It's tempting to reach for a big scientific synthesis to answer the September question. The International AI Safety Report 2026 synthesizes scientific evidence on capabilities, risks, and mitigation, but its evidence cutoff precedes December 2025. It therefore cannot validate the particular September 2026 releases in this chronology, and I don't use it that way.
For this report, the absence of a reviewed independent evaluation is recorded as an evidence gap, not as proof that no evaluation exists. A defensible adverse conclusion would need to identify an actual commitment, its relevant conditions, the conduct assessed, and the evidence of noncompliance. That standard protects everyone: it holds critics to the same standard, and it gives companies a clear path to demonstrate compliance rather than asserting it.
1.5.3 An illustrative scoring exercise#
To see the test in motion, consider a purely hypothetical laboratory whose conduct will be scored against it. The illustration is invented to demonstrate the test; it describes no real company.
Suppose the laboratory announces that it is "slowing down for safety." The activity element asks: which of the six activities has slowed? Suppose its answer is release frequency: it shipped two models this year instead of four. The baseline element asks: compared with what? If last year's four releases were itself an unusual burst and the two-year average is two, the "slowdown" is a return to normal, and the claim dissolves into arithmetic. Suppose, though, that the schedule really did halve. The safety-bar element asks: what criteria determined that these two releases were safe to ship? If the answer is internal red-teaming with unpublished failure conditions, the bar is not operational; no third party could reproduce the judgment, and no critic could falsify it.
Now suppose the laboratory names an external evaluation firm. The evaluation element asks: what access did the evaluators have, what methods did they use, what could they not test, and may they publish? If the evaluators saw a frozen configuration that differs from the shipped product, the evaluation result may not cover the release. The decision element asks the question that separates process from theater: did the evaluation, or anything else, ever cause a delay, a restriction, a redesign, or a cancellation? If the answer across the company's history is no, if every evaluation has ended in a release, then either the evaluators have never found a problem worth acting on (possible, but worth scrutinizing) or the process cannot say no. A review that cannot say no is not a review.
The residual-risk element asks what the evaluation did not cover, and the change-control element asks whether next quarter's update inherits this quarter's result. A passing grade for version 1.0 says nothing about version 1.1 with new tool permissions.
None of these questions requires insider access. All of them can be asked by any journalist, regulator, or interested citizen, and all of them have answers that are either on the public record or conspicuously absent from it. That is what makes the test usable. It converts "trust us" into a checklist.
Run the hypothetical lab through the worksheet in Section 1.5.5 and nearly every row reads "undisclosed." That doesn't say the lab is reckless. It says the public can't verify the claim.
1.5.4 Applying the test: a four-step method#
The test is designed to be run, not just admired. A working journalist, an analyst at a regulator, or a member of a state legislature's technology committee can apply it to any laboratory's public record in a focused afternoon, using four steps.
Step one: fix the commitment's text before fixing its meaning. Pull the primary statement (the essay, the system card, the testimony) and quote the operative sentences verbatim, with their dates. Paraphrase is where grading goes to die. "We are committed to safety" and "we will dedicate 20 percent of our compute to alignment within four years" are different kinds of sentences, and only the second kind can fail. If the commitment contains no number, date, criterion, or named evaluator, that is itself the finding: the commitment is not operational, and no conduct can be scored against it.
Step two: assemble the conduct record on the commitment's own terms. Collect the releases, the evaluations, and the organizational changes that postdate the commitment. The record assembled in Section 1.4 is an example of this step done properly: dates of the commitment, dates of the departures, the dissolution, the reported denial of pledged compute. None of it required a subpoena.
Step three: score each test element as supported, contradicted, or undisclosed. The three-way scoring is the point. A loud public argument pressures everyone toward supported or contradicted; the real output of most scoring sessions is a column of undisclosed entries, and that column is the story. An undisclosed safety bar is not proof of a failed safety bar. It is proof that the public cannot verify the claim.
Step four: publish the worksheet. The test's value compounds if the scoring is reproducible. When a regulator's worksheet and a journalist's worksheet disagree about the same element, the disagreement identifies exactly which record needs to be produced. That is how a procedural standard does its work: not by producing verdicts, but by producing answerable questions.
One caution belongs with the method. Running the test well requires resisting the two shortcuts that make public argument about AI so unproductive:
- The sympathy shortcut grades a company on intentions and difficulty rather than on disclosed conduct.
- The villainy shortcut treats every undisclosed element as proof of the worst case.
The worksheet format exists to make both shortcuts visible. A column of "undisclosed" entries is a demand for records, not a conviction.
1.5.5 Run the test yourself: a copyable worksheet#
Here's the worksheet. Copy it into a document or spreadsheet, fill in the commitment text at the top, and score each row. It uses only the seven elements above; nothing here requires inside information.
Header (fill in before scoring):
- Organization:
- Commitment text, quoted verbatim:
- Date and source of the commitment:
- Conduct being scored (release, evaluation, organizational change) and its date:
- Who scored this, and when:
| Element | What to look for | Supported | Contradicted | Undisclosed |
|---|---|---|---|---|
| Activity | Which activity is said to be slowing: training scale, capability improvement, external access, deployment scope, autonomy, or release frequency | [ ] | [ ] | [ ] |
| Baseline | The previously planned or expected activity, over a stated period, that the claim is compared with | [ ] | [ ] | [ ] |
| Safety bar | Stated criteria, the test population, failure conditions, and stated uncertainty | [ ] | [ ] | [ ] |
| Evaluation | Who evaluated, what access they had, how independent they are, their methods and exclusions, and whether they may publish | [ ] | [ ] | [ ] |
| Decision | Any documented delay, restriction, redesign, or cancellation caused by evidence | [ ] | [ ] | [ ] |
| Residual risk | Unresolved failures, and limits on detection or monitoring | [ ] | [ ] | [ ] |
| Change control | The version, access tier, and tool permissions tested, and how post-release updates are handled | [ ] | [ ] | [ ] |
How to score each row:
- Mark Supported only when a dated public record shows the element was disclosed and the conduct matches it. Write the source link in your notes.
- Mark Contradicted only when a dated public record shows conduct that conflicts with the stated condition. Name the record. The superalignment sequence in Section 1.4.2 is an example of what contradicting evidence looks like: a stated number and deadline, and a reported outcome that did not match them.
- Mark Undisclosed when the public record is silent. That is the default, not a failure of your research.
Before you publish, ask one question: could someone who disagrees with me reproduce my scores from my links? If not, you've produced an opinion with a table around it.
1.5.6 What this means for you#
- For a parent: when an app says a new version is "safer," ask which activity changed and compared with what. If there's no answer, the claim hasn't been made yet.
- For a teacher or school leader: ask whether the version your students use is the version that was evaluated. That's the change-control row, and it's the one most easily missed.
- For a board member or investor: ask management which rows of this worksheet the company could fill in today with public records. The undisclosed rows are the governance gap.
- For a county commissioner or resident: use the same seven rows on any infrastructure proposal, including one from a company like mine. The questions translate directly.
1.6 What the Record Cannot Support: Consequences and Causal Inference#
The documentary records reviewed in this chapter do not estimate the causal effect of these statements on investment, facility construction, children's well-being, or parental understanding. They describe positions and products, not controlled measurements of downstream outcomes. (Amodei's pacing statement, Pachocki's OpenAI statement, Anthropic's announcement, OpenAI's product update)
This report consequently does not state that the September communication damaged the economy. It also does not assume the communication was harmless. The question is unresolved within this review, and saying so plainly is load-bearing: the chapters on infrastructure and household life quantify what is known about those domains, and their credibility depends on not smuggling unexamined assumptions in through the front door.
The temptation to draw causal conclusions from the September sequence is strongest in exactly the places where the evidence is weakest. Watch for three moves in public argument.
- Timeline juxtaposition. Placing an announcement date beside a market movement or a project delay and inviting the reader to connect them.
- Denominator switching. Citing the number of headlines, tweets, or alarmed op-eds as if attention were harm.
- Attribution by mood. Arguing that because the communication felt chaotic, it must have caused concrete damage somewhere.
Each move produces the feeling of evidence. None of them is evidence. A number without a denominator is a rumor, and a count of alarmed headlines has no denominator at all.
1.6.1 Public Trust in AI: What the Pew Surveys Show, 2021 to 2026#
If the documentary record cannot tell us whether the September communication caused damage, survey research can at least describe the audience that received it. That description matters for communication design, and it comes with a hard caveat: attitudes are not harms, and a worried public is not an injured public.
The Pew AI survey trend: concern up, excitement down. The Pew Research Center has tracked American attitudes toward AI since 2021, and the trend is one-directional. In June 2021, 37 percent of U.S. adults said the increased use of AI in daily life made them more concerned than excited, while 18 percent said the reverse. By June 2026, the concerned share had climbed to 52 percent and the excited share had fallen to 9 percent, with the remainder reporting equal measures of both. Those 2026 figures come from survey data published by Pew Research Center in August 2026.
The shift was not a single post-ChatGPT shock. The concerned share jumped between 2022 and 2023 (from 38 to 52 percent) and has held near that level since, at 51, 50, and 52 percent in 2024, 2025, and 2026 respectively. Concern has continued to climb among adults under 30 even after leveling off among older groups, and worry about job loss is still rising.
| Pew survey year | U.S. adults more concerned than excited |
|---|---|
| 2021 (June) | 37 percent (18 percent more excited) |
| 2022 | 38 percent |
| 2023 | 52 percent |
| 2024 | 51 percent |
| 2025 | 50 percent |
| 2026 (June) | 52 percent (9 percent more excited) |
Public concern about AI rose and stayed high; experts and the public see it differently
Source: Pew Research Center, Aug 2026 · Pew Research Center, Apr 2025
1.6.2 The expert versus public gap#
A second finding sharpens the picture. When Pew compared the general public with people who work as AI researchers and developers in 2025, the two groups were nearly mirror images: 47 percent of experts said they were more excited than concerned about AI in daily life, against 11 percent of the public. On the twenty-year outlook, 56 percent of experts expected a positive impact versus 17 percent of the public. (Pew Research Center, April 2025)
1.6.3 How to read the Pew trend (and how not to)#
Read together, these results describe a public that has heard the warnings, absorbed them, and is waiting to be convinced. They do not describe a public that has been harmed by any particular communication event. They also describe a trust gap between the people building these systems and everyone else, and that gap is precisely what a communication standard has to bridge.
A company that reads the Pew trend as a marketing problem to be solved with reassurance has missed the finding. The public's concern has proven durable across four years of product improvements, which suggests it attaches to the pattern of conduct, not to any single headline. That is the strongest argument for the procedural standard in this chapter: the only communication strategy with any evidence behind it is a verifiable one.
Common misreadings to avoid:
- "52 percent are concerned, so AI is harming people." Concern is an attitude. The Pew data do not measure injury.
- "Concern jumped in 2023, so ChatGPT caused it." The data show when the jump happened, not why.
- "Experts are more optimistic, so the public is misinformed." The comparison shows a gap. It does not say which side is right.
- "The September essays drove the 2026 number." The 2026 survey was fielded in June 2026, before the September essays were published.
1.6.4 Two studies that would answer the harm question#
What would resolve the question of whether the September communication caused harm? Two studies, both feasible.
The first is a communication study. It would randomly assign adults to source-faithful versions of a release explanation: a headline-only version, a capability-and-limits version, and a version including independent evaluation and residual uncertainty. It would measure understanding of the actual commitment, recognition of uncertainty, ability to identify appropriate uses, and willingness to revise a decision after contrary evidence.
One subtlety deserves emphasis: a person can understand a statement more accurately while trusting the organization less. That result would not necessarily represent a communication failure, and a study that treated trust as the only outcome would miss the point. Exposure to alarming claims should not be tested on young children without appropriate specialist and ethical oversight.
The second is a business study. It would need documented changes in procurement, financing, construction, or demand, and a credible comparison separating communication effects from interest rates, utility constraints, customer commitments, and the other events that move markets and project schedules. Merely placing announcement dates beside market or project movements would not establish causation. Anyone can line up two timelines; the trick is that the events in them usually share causes rather than cause each other.
Neither study exists in the reviewed record. Proposing them is not a dodge. It is the difference between a report that grades the available evidence and one that invents evidence to fill a gap.
1.7 A Responsible AI Release Standard: Six Questions#
Whatever the September record shows, one conclusion needs no further study: the way capable systems are explained to the public is not working well enough. The practical question is what a working explanation would contain.
This report proposes that a public explanation of a consequential release or deployment should answer six questions in plain language:
- What changed? Not what might change, not what the technology could eventually do. What, specifically, is different as of this date.
- Who can use it? The permitted age range, access tier, professional context, and the adult or institutional role required.
- What evidence supports the change? The strongest study, its population and design, and where to read it.
- What remains unreliable? The failure modes, the populations the evidence does not cover, and the uncertainty.
- Who is accountable? A named role, not a brand. Someone who can answer for the outcome.
- How can the public challenge an outcome? A working correction and appeal path, with a stated response time.
This standard is proposed, not yet tested. Chapter 11 returns to it as part of the shared accountability framework.
1.7.1 The six questions and their failure patterns#
Each question in the standard exists because a specific, recurring failure pattern made it necessary. I name the patterns generically, because the point is the pattern rather than any one offender.
"What changed?" fails as a headline. Headlines report the most dramatic thing a system could conceivably do, which is rarely the thing that shipped. The reader who absorbs only the headline arrives at the product page expecting a capability that is not there, and the gap between expectation and product does more damage to trust than a candid limitation ever could. The question forces the communicator to lead with the shipped fact.
"Who can use it?" fails as a demographic afterthought. Age gates, professional restrictions, and institutional requirements are usually buried in policy documents that no one reads before forming an opinion. A parent who learns, after the fact, that a product was never intended for a nine-year-old has been failed by the communication, whatever the terms of service said.
"What evidence supports the change?" fails as a number without a study. "Our model scores 94 percent" is not evidence until the reader can learn what the other participants were, what the test population was, and whether anyone independent ran the test. The question requires the citation, not the statistic.
"What remains unreliable?" fails by omission most of all. This is the question that separates explanation from advocacy, and it is the one most often answered with silence. The cost of answering it is a shorter, less shareable announcement. The cost of omitting it is discovered later, by the user, in the worst possible context: a wrong answer in front of a student, a patient, a judge, a customer.
"Who is accountable?" fails as a brand. Brands do not answer questions; people with names and roles do. An accountability chain that terminates at a corporate identity terminates nowhere, because no individual's judgment is on the line.
"How can the public challenge an outcome?" fails as a black box. A correction path that requires expertise the user does not have, or persistence most users will not muster, is not a correction path. The test is procedural, not rhetorical: can an ordinary user, with an ordinary wrong answer, get it fixed within a stated time? If the answer is unknown, the sixth question has not been answered.
The six questions form a sequence with a deliberate structure. The first three establish the positive claim; the last three establish what happens when the claim fails. An explanation that answers only the first three is a sales document. One that answers all six is the minimum a consequential deployment owes its public.
1.7.2 How the standard adapts to each audience#
The same six questions produce different sentences for different readers.
- For a parent: the permitted age range and adult role, the learning objective, what child information is collected, and how to stop or report a problem.
- For an educator: curriculum scope, independent learning outcomes, teacher workload, and whether students retain competence without assistance.
- For a civic committee: the source documents, alternatives considered, omitted viewpoints, and the authority that makes the decision.
- For a worker: productivity claims separated from monitoring, changes in job responsibility, training requirements, and pay.
- For an infrastructure builder: product capability announcements separated from executed capacity commitments and local approval obligations.
- For a resident: the project's actual power, water, fiscal, and operating terms, without distant software benefits presented as compensation already delivered locally.
That infrastructure-builder line is the one I live with. A product announcement from a model developer is not a signed capacity commitment, and neither one is a local approval. When those three get blurred together in a public meeting, residents end up evaluating a promise about software as if it were a promise about their substation.
1.7.3 What the standard refuses to do#
Notice what the standard refuses to do. It does not promise brevity at the cost of uncertainty; the fourth question is the one most corporate communications omit, and it is the one this report considers non-negotiable. It does not treat the audience as a single public; a parent and a power-system engineer need different sentences, and a standard that serves both says so. And it does not confuse drafting with action. A model that can write the explanation does not thereby earn the right to publish it; Chapter 11 returns to that separation as a general control.
A concise explanation should not remove material uncertainty. The practical target is a decision the reader can inspect, not reassurance at any cost.
1.7.4 A worked example: reading a release note in a meeting#
Here's a hypothetical. A school district's technology committee receives a vendor's one-page note announcing that its classroom assistant now runs on a newer model. The note says the update is "more capable and safer." No statistics are given, so I won't invent any.
Walk the six questions. What changed? "More capable" doesn't say. Who can use it? The note doesn't mention age ranges or whether a teacher must supervise. What evidence supports the change? None is cited. What remains unreliable? Silent. Who is accountable? The vendor's brand name. How can a family challenge an outcome? Not stated.
That note answers zero of six. The committee doesn't need to reject the product; it needs to send the six questions back in writing and decide after the answers arrive, then run the seven-part test on the word "safer." Neither tool gives a verdict. Both give better questions.
1.8 Terminology and the Scope of the Report's Title#
The title of this report, "The Superintelligence Transition: A Human Agenda From Kindergarten to Retirement," is an organizing theme, not a finding. It does not claim that any system available as of September 23, 2026 has achieved a universal capability threshold beyond human intelligence. When the historical record discusses such thresholds, I quote it and attribute it.
OpenAI's 2023 explanation of superintelligence, published when it announced its superalignment effort, described systems "much smarter than us" and stated the organization's belief that such systems could arrive within the decade. That was a publisher's prediction, and I treat it as one. (OpenAI, Introducing Superalignment)
Throughout the chapters that follow, the text uses "current models," "advanced models," "structured software," or the name of the specific intervention when that is the accurate description. Original organization names, study titles, quotations, and regulatory language retain their wording, including "AI" where changing it would misrepresent the record. This report does not treat "no longer artificial" as a technical finding, and it is wary of any text, in either direction, that treats terminology as evidence.
The reason for this discipline is practical. The lifespan questions in this report (whether a child learns, whether a worker earns, whether a resident's water bill changes) do not depend on what the most capable system is called. They depend on what a specific tool does for a specific person under specific conditions. Precision about words is how the report keeps its attention on those conditions.
Terminology discipline also has a defensive purpose. In a heated public argument, the side that controls the definition of the threshold often appears to control the conclusion. By declining to treat any label as a finding, this report denies both the optimists and the catastrophists that shortcut. The chapters measure tasks, outcomes, and people. Whatever the most capable system of the moment is called, those measurements stand or fall on their own. Capability is not benefit, and a name is not a capability.
1.9 Research Method and Evidentiary Boundaries#
This report is a targeted, multi-source desk-research synthesis with an evidence cutoff of September 23, 2026. It is not a preregistered systematic review, and it makes no claim of exhaustive search or complete study census.
The review prioritizes original research, government evaluations, official records, and primary company statements when the claim concerns what a company said. Secondary and institutional accounts are labeled where used, as with the Business Insider, Newsweek, The Next Web, and Business Standard reporting in this chapter. Where a chapter relies on a publication record, abstract, presentation, or research summary rather than a full-paper audit, it says so.
Favorable, null, mixed, and adverse findings are all retained. No pooled estimate is calculated across the chapters of this report, because the populations, interventions, outcomes, comparison conditions, and study designs are materially different. Combining them into a single number would manufacture a precision that none of the underlying studies supports. The full evidence-class table lives in Chapter 11.
Eight interpretation rules govern every chapter:
- Relevant human outcome. Evidence must bear on learning, task performance, access, welfare, work, financial decisions, health, participation, infrastructure effects, or accountability.
- Traceable record. Claims require a retrievable source and an identifiable population, process, document, or observation period.
- Design fidelity. Randomized, quasi-experimental, observational, simulated, technical, documentary, and publisher claims remain distinct categories.
- Technology fidelity. Non-generative programs may illuminate mechanisms or alternatives but are not presented as evidence that a current language model reproduces their effects.
- Version fidelity. Material updates are considered without silently replacing an earlier study's sample or estimate.
- Outcome fidelity. Time saved, output quality, learning, earnings, health, satisfaction, and community benefit remain different measures.
- Transfer discipline. A finding is not generalized to a new age, country, occupation, product, or facility without acknowledging the gap.
- Conflict visibility. Disputes and source limitations are described rather than resolved by choosing the most convenient number.
These rules mirror the seven-part test: version fidelity is change control applied to research. I hold the research to the same discipline I ask of the labs.
1.9.1 Two evidence chains that never substitute for each other#
Every chapter in this report runs along one of two evidence chains, and it helps to see them side by side before the rest of the report begins.
The first chain runs from a tool to a task to a human outcome. A tutoring program (tool) is used for practice problems (task), and the question is whether the student learns (human outcome). Most of the chapters on schooling, work, households, finance, and health live on this chain.
The second chain runs from a facility to its resources, contracts, and obligations to a community outcome. A data center (facility) draws power and water under specific agreements and approvals (resources, contracts, obligations), and the question is what happens to the community's bills, supply, tax base, and services (community outcome). Chapter 10 lives on this chain.
Two evidence chains that interact but never substitute
The chains interact, but they never substitute for one another. Evidence that a tool helps a student says nothing about whether a facility burdens a town's water, and evidence that a facility paid its taxes says nothing about whether the software it runs helps anyone learn. Borrowing credit from one chain to pay a debt on the other is the most common trick in debates about AI infrastructure.
1.9.2 Citation tracking and source verification#
Two instruments hold this report to its own standard, and both are unusual enough to describe. The first is a citation ledger: a mechanically generated index of every source-linked passage in the manuscript, each with its document, section, and line location. The ledger is an audit aid, not a second body of evidence; a source that appears in five chapters has not become five studies. The second is a source verification log: every one of the 83 distinct sources cited in this report was checked during preparation of this edition. Sources that blocked automated access were spot-checked through independent retrieval paths; sources that had died were either replaced with a canonical location of the same work or flagged. The verification record is published with the report so that readers can repeat any check, and the evidence register at the back of the report is where to start.
The reason for this apparatus is simple. This report's only currency is traceability. A claim a reader cannot check is, for this report's purposes, not a claim but a rumor with formatting. The apparatus is imperfect (a live link does not certify a true page, and a dead link does not falsify a study) but it converts the report's factual posture from an assertion into a procedure.
The same logic governs the seven Savrn trackers referenced throughout the report: a tracker entry starts a question and never ends one.
1.9.3 Falsifiability: what evidence would change these conclusions#
A research synthesis that cannot say what would prove it wrong is a position paper wearing a lab coat. This report's central claims are falsifiable in specific, nameable ways, and stating them here is part of the method.
The conditional thesis, that capability, adoption, and benefit are not interchangeable, and that defined applications can improve particular outcomes, would be weakened by credible evidence that the reviewed null and adverse results failed to replicate, or that the favorable results replicate readily across the populations and settings where this report marks transfer limits.
A well-run, adequately powered trial showing that unrestricted conversational assistance improved unaided later performance in high-school mathematics would directly contradict the interpretation this report places on the PNAS experiment; the interpretation would have to change, and the schooling chapter, Chapter 3, says so explicitly.
Conversely, if the field trials proposed in Chapter 11 repeatedly find that workflows built on the seven Savrn trackers produce no improvement in decision quality over ordinary research practice, the trackers' public-value claim should be withdrawn. Savrn's interest in that claim is exactly why the withdrawal test is worth stating in advance.
The September 2026 analysis in this chapter is falsifiable through the seven-part test itself. If either laboratory publishes the elements the test calls for (the operational safety bar, the evaluator arrangements with publication rights, evidence that the process has said no at least once), the "undisclosed" entries in any fair worksheet fill in, and the scrutiny this chapter argues for resolves in the companies' favor. That outcome is not a failure of this report's method. It is the method succeeding.
1.9.4 Sponsorship and completion boundaries#
Two structural boundaries deserve restating at the outset.
First, this report is prepared for Savrn, a company with a commercial interest in the data center infrastructure industry, and it does not describe itself as an independent institutional review. I'm the founder and CEO. You should read every chapter knowing that. Sponsorship disclosure appears wherever tracker findings support an argument.
Second, "completed" describes the manuscript and its evidence package. It does not mean the proposed workflows have undergone field trials, that outside experts have peer-reviewed this report, or that any facility has passed a site-specific impact assessment.
1.10 Conclusion: A Procedural Standard for Public Trust in AI#
The public record of September 2026 supports a real question about the relationship between warnings, pacing proposals, and continued releases. The language of the reviewed proposals, in Amodei's essay and Pachocki's essay, rules out treating every later release as an automatic contradiction. Both things are true at once, and a serious analysis has to hold them together rather than choosing one.
My conclusion from that record is procedural:
- Ask for a measurable commitment.
- Evaluate conduct against the specific conditions of that commitment, using the seven-part test in this chapter.
- Record what is said, what is released, what is evaluated, and by whom.
- Let neither corporate assurances nor public suspicion substitute for evidence.
The rest of this report applies that procedural standard to the questions that matter across a human lifetime: what these systems do for a kindergartner, a high-school sophomore, a college student, a new worker, a midcareer parent, a civic committee, an investor, an older adult, and the community that hosts the physical infrastructure underneath all of it. The standard does not change from chapter to chapter. Only the decisions do.
1.11 Frequently Asked Questions#
Did OpenAI and Anthropic break their AI safety commitments in September 2026?
The record does not show that. Both September 2026 essays called for safety-conditioned pacing, not a halt, and Amodei wrote that pacing "does not mean halting model training or technical progress." Releases on September 22 are consistent with that wording if they met the stated conditions, such as third party evaluator confirmation. Whether they did is a factual question the public record, as of September 23, 2026, does not settle.
What happened to OpenAI's superalignment team?
OpenAI announced superalignment in July 2023, pledging 20 percent of secured compute over four years. According to reporting by Business Insider, Newsweek, and The Next Web, Ilya Sutskever announced his departure on May 14, 2024, Jan Leike resigned publicly on May 17, and the team was dissolved within days, at roughly the one-year mark. OpenAI stated that safety responsibilities were redistributed rather than ended.
What does AI pacing mean?
In Dario Amodei's September 2026 essay, pacing means slowing the rate of capability improvement while ensuring companies take adequate time to align and safeguard models, and for third party evaluators to confirm this. It explicitly does not mean halting training or technical progress. That makes pacing a conditional commitment: a release can be consistent with it, but only if the stated conditions are met and shown.
How can I tell if an AI company is really slowing down?
Ask two things first: slowing what, and compared with what? "Slower" could mean training scale, capability improvement, external access, deployment scope, autonomy, or release frequency, and these can move in opposite directions. Then check the safety bar, the evaluation, whether evidence ever caused a delay or cancellation, residual risk, and change control. This seven-part test is proposed, not yet applied by any regulator.
What did the 2026 Pew AI survey find about public trust in AI?
In survey data from June 2026, published by Pew Research Center in August 2026, 52 percent of U.S. adults said increased use of AI in daily life made them more concerned than excited, up from 37 percent in June 2021. Only 9 percent were more excited, down from 18 percent. The survey measures attitudes, not harm, and does not isolate any single event.
What should a responsible AI release announcement include?
This report proposes six plain-language questions: what changed as of this date, who can use it, what evidence supports the change, what remains unreliable, who is accountable by named role, and how the public can challenge an outcome with a stated response time. The first three make the claim; the last three cover failure. The standard is proposed and has not yet been tested.
Does Savrn hold itself to the same test?
Yes. Savrn's behind-the-meter power and zero-makeup-water cooling are design goals and publisher statements, not verified operating results. Run through the seven-part test today, most rows read undisclosed or not yet measurable. This report is prepared for Savrn, which has a commercial interest in data center infrastructure, and it is not an independent institutional review.
