Chapter 11 · The Accountability Standard
An AI Accountability Framework: Evidence Classes, Claim Records, and Pilot Protocols

Every chapter in this report kept arriving at the same place. Schools, households, workplaces, city halls, clinics, retirement accounts, power grids. Different evidence, different people, same requirement. So this chapter is my attempt at an AI accountability framework that anyone can pick up and use, whether you run a district, sit on a utility commission, or are trying to decide whether to trust a chatbot with your mother's medication list. The real question underneath all of it is simple: when someone tells you an AI system works, how do you know what that claim is actually worth, and who is on the hook if it's wrong?
Here's why I care about that question the way I do. In infrastructure, you do not energize a site on a promise. You energize it on a commissioning record. Somebody tested the breaker, somebody signed the sheet, somebody knows which relay trips first, and every one of those facts has a name and a date attached. I learned early, building large power loads, that a load that big doesn't care how confident the brochure sounded. It cares whether the record matches the physics. This chapter is the commissioning record for claims.
What the evidence will show is uncomfortable in both directions. The tools that institutions already lean on, including NIST's voluntary framework, give you structure but certify nothing. The strongest trials in this report hold up only when their scope conditions travel with them. And the standard I propose here, including every template, is itself a proposed workflow: the lowest rung on the evidence ladder it defines. I'll hold it to that label, and I'll hold Savrn's own claims to it too.
11.1 The Common Requirement Behind Every AI Accountability Framework#
Across education, household life, work, civic decisions, finance, health, and infrastructure, the preceding ten chapters kept arriving at the same practical requirement: a claim must remain connected to its evidence, its limits, and the person authorized to act. The requirement wore a different uniform in each chapter.
In schooling (Chapter 3), it showed up as the difference between assisted practice and unaided learning.
In the household (Chapter 6), it showed up as the difference between faster drafting and redistributed responsibility.
In civic life (Chapter 7), it was the difference between a readable summary and an accurate one. A clean summary of public comments can read beautifully and still leave out the objection your neighbor filed.
In finance and health (Chapter 8 and Chapter 9), it was the difference between a persuasive output and a measured outcome.
In infrastructure (Chapter 10), it was the difference between an announced commitment and a metered, enforceable one. A press release is a sentence. A meter is a record.
One requirement across every domain
Line those up and you'll see one pattern. In every domain, the failure mode is a claim that drifts loose from its anchor. The number keeps circulating after its population, its comparison group, and its limit have fallen away. The recommendation keeps moving after the person who was supposed to approve it has been skipped. That's the whole game: keep the claim tied to the record, and keep the record tied to a person with the authority to act.
This chapter turns that recurring requirement into a complete proposed process. Nothing in it is an independently validated intervention or a legal-compliance certification. It is an evidence-informed operating standard, offered for testing and adoption by institutions that want the benefits described in the earlier chapters without inheriting the failure modes.
11.1.1 Where the NIST AI Risk Management Framework Fits#
If you work inside an institution, you've probably heard someone mention NIST. The governance backdrop for this chapter is NIST's voluntary risk-management framework, which organizes work around four functions: governance, mapping, measurement, and management. NIST also publishes a generative-model profile that addresses risks including confabulation, privacy, information integrity, and human interaction. (NIST AI Risk Management Framework, NIST generative AI profile)
Those documents give a research-informed governance basis, and I lean on their structure throughout. But be clear about what they are. They do not certify a product. They do not prove that any deployment produces benefits. A vendor who says "we follow the NIST framework" has told you about their process, not about your outcome.
The framework page also describes continuing work and revisions. That matters, because a draft or concept note should never be presented as a finalized binding standard. The same discipline applies to this chapter's own proposals as much as to NIST's. (NIST framework status page)
11.2 The Eight Evidence Classes: What Each Kind of Study Can Say#
This whole report relies on a classification of evidence types. I state it here in full because everything downstream depends on it, and other chapters link back to this table rather than repeating it. It is an editorial labeling system, not a formal evidence-grading methodology such as GRADE. The labels organize judgment. They do not replace it.
| Class | What it can establish | What it cannot establish alone |
|---|---|---|
| Randomized comparison | An effect of the assigned intervention under the study's design and assumptions | Universal transfer, long-term benefit, or attribution to one component of a bundle |
| Quasi-experimental analysis | An estimated effect under an explicit identification strategy | Freedom from unmeasured confounding without supporting assumptions |
| Observational or survey evidence | Associations, reported experience, measured prevalence within scope | Causal benefit from adoption |
| Technical evaluation | Performance on specified inputs and scoring rules | Real-world welfare or reliable operation outside the evaluated task |
| Simulation or forecast | Conditional results under stated assumptions | Observed outcomes or guaranteed future results |
| Administrative or official record | A documented policy, status, expenditure, or measured quantity | A causal effect beyond the record's scope |
| Publisher statement | What an organization says about its product, method, or commitment | Independent verification of the claim |
| Proposed workflow | A process that can be implemented and tested | A benefit already demonstrated |
The eight evidence classes, with one example from this report
Every public-facing summary derived from this report should preserve the class of its underlying evidence. Three rules follow, and I'd tape them to the wall of any communications office:
- A simulation must not become a real-world outcome merely because it appears beside an empirical study.
- A publisher statement must not become verification because the publisher is reputable.
- A proposed workflow must not become a finding because it is well constructed.
You've watched this report apply the classes repeatedly. The METR developer studies were treated as evaluations whose meaning shifted with changing tools. The JLARC bill figures were scenarios, not observations. The LBNL national totals were modeled estimates, not meter data. The classification is not bookkeeping. It is the mechanism that keeps sound claims sound.
11.2.1 Each Class, With One Example From This Report#
A table is easy to nod at and hard to use. So here is each rung, what it looks like in practice, and the most common way people misread it.
Randomized comparison. The Swedish MASAI trial randomized 105,934 women to model-supported mammography screening or standard double reading (MASAI trial record). Randomization is what lets you say the difference came from the assigned workflow rather than from who chose it. The common misreading is to treat a randomized result as universal: the effect belongs to that screening system, inside that radiologist workflow, under that design. It is not evidence about consumer self-diagnosis, and it did not measure mortality.
Quasi-experimental analysis. The apprenticeship evaluation cited in the evidence register found that registered participation was associated with positive estimated employment and earnings effects, and the register flags its quasi-experimental design and possible unobserved selection (DOL apprenticeship evaluation). The misreading is to treat the estimate as if randomization had removed selection. It didn't. The identification strategy carries assumptions, and the claim is only as good as those assumptions.
Observational or survey evidence. A survey can tell you how many people report using a tool, or how they feel about it. It can measure prevalence within its scope. The misreading is causal: "people who use it are doing better, so it made them better." Surveys can't carry that sentence alone.
Technical evaluation. The UK Department for Transport's blinded test of a consultation-analysis tool measured performance on specified inputs under a scoring procedure: theme generation reached approximately 0.75 recall, 0.50 precision, and 0.59 F1 (Department for Transport evaluation). The misreading is to treat benchmark performance as real-world welfare. A tool that scores well on the test set has not yet shown that any committee made a better decision.
Simulation or forecast. Virginia's JLARC study modeled a potential increase of roughly $14 to $37 per month in a typical Dominion residential bill by 2040 under its examined scenarios (JLARC data-center report). That is a conditional result under stated assumptions. The misreading is turning it into a present-tense bill increase on a flyer. Chapter 10 calls that manufacturing a fact the audit never produced.
Administrative or official record. The AEP Ohio data-center tariff is a documented, commission-approved policy with minimum contract-capacity ramps and exit obligations (AEP Ohio data-center tariff). A record like that establishes what the policy says. The misreading is to treat the record as proof of an effect beyond its scope: a tariff on paper is not evidence of what happened to anyone's bill.
Publisher statement. Savrn's own design goals, including behind-the-meter power and a zero-makeup-water cooling design goal, are publisher statements. So are the publisher-described methods of the seven Savrn trackers. They tell you what we say. They are not independent verification of it. I'll come back to this in Section 11.6, because it's the part most people would expect me to skip.
Proposed workflow. This chapter. The claim record, the five-stage process, the pilot protocol, the four-touch sequence, the eight-step audience workflow. All of it can be implemented and tested. None of it is a benefit already demonstrated.
11.3 The Claim Record: Ten Fields for AI Claim Verification#
The proposed record for any consequential claim contains enough information for a different reviewer to understand the claim and its limits without calling the author. Not every record needs a numerical estimate. Every consequential claim needs a clear basis.
11.3.1 The Ten Fields, Explained#
- Identity. A stable claim identifier, version, responsible editor, and chapter or use case. Without an ID and a version, nobody can tell whether two people quoting "the study" mean the same sentence, and corrections can't find their target.
- Statement. The narrow proposition being asserted, without promotional extensions. Write the sentence the evidence supports, not the sentence marketing would prefer.
- Population and setting. Who, where, when, and under what institutional conditions. A finding from radiologists in a national screening program is not a finding about a consumer app.
- Intervention and comparator. What changed and what it was compared with. "Better" is meaningless until you know "better than what."
- Outcome. The exact measure, denominator, time horizon, and whether it was primary, secondary, exploratory, or self-reported. Show me the denominator.
- Estimate and uncertainty. The effect, the interval where available, and the study's treatment of statistical significance. A point estimate without its interval is half a fact.
- Technology boundary. Model and version when stated, or explicit identification of a non-generative intervention. A result from a specific screening system is not a result about general conversational models.
- Source. Original URL, publication date, status, and relevant update.
- Limitations. Selection, attrition, measurement, transfer, sponsorship, and conflicting evidence.
- Action boundary. What the evidence permits, what it does not permit, and who must approve consequential use.
The last field is the one I'd fight hardest to keep. Most evidence summaries stop at "here's what the study found." The action boundary forces you to say what a decision-maker may do with it, what they may not do, and whose signature is required. That is where a finding turns into accountability.
11.3.2 A Filled-In Claim Record: The MASAI Mammography Trial#
Here's the ten-field record applied to the strongest clinical trial in this report, using only what Chapter 9 reports from the trial record. Where the report doesn't state a value, the record says so rather than guessing. That is a feature of the template, not a gap in it.
| Field | Entry |
|---|---|
| Identity | Example ID: CH09-MASAI-01; version 1; responsible editor named by the using institution; use case: clinical screening evidence (Chapter 9) |
| Statement | In this screening workflow, model-supported reading achieved a noninferior interval-cancer outcome compared with standard double reading, with higher sensitivity and similar specificity |
| Population and setting | Women in the Swedish MASAI trial: 105,934 randomized, 105,915 included in the reported analysis after exclusions; screening inside a radiologist workflow |
| Intervention and comparator | A specific screening-support system embedded in a radiologist workflow, compared with standard double reading |
| Outcome | Primary: interval cancer (cancers diagnosed between screening rounds), per 1,000 participants. Secondary: sensitivity and specificity. Mortality was not measured |
| Estimate and uncertainty | Interval cancer 1.55 per 1,000 (supported) vs 1.76 per 1,000 (standard); ratio 0.88, 95% CI 0.65 to 1.18. Met the noninferiority criterion; did not show a statistically significant reduction. Sensitivity 80.5% vs 73.8%; specificity approximately 98.5% in both groups |
| Technology boundary | A specific screening system within a supervised radiologist workflow; not a general conversational model acting as an autonomous physician. Model version: not stated in this report's summary |
| Source | Lancet trial record via PubMed; publication date and later updates to be recorded from the original record at time of use |
| Limitations | The upper bound of 1.18 means the data are consistent with the supported workflow being somewhat worse as well as somewhat better on interval cancer; no mortality outcome; no evidence about consumer self-diagnosis; sponsorship and conflicts to be recorded from the original paper |
| Action boundary | Permits: citing a noninferior interval-cancer result and improved sensitivity without a rise in false alarms, within a defined radiologist workflow. Does not permit: "prevented 12 percent of cancers," mortality claims, or consumer-diagnosis claims. Consequential use requires approval by the clinical authority responsible for the screening program |
Read the action-boundary row again. Chapter 9 said it plainly: "noninferior interval-cancer outcome in this screening workflow" is a finding the data support, and "the technology prevented 12 percent of cancers" is not, because the direction of the point estimate doesn't survive its confidence interval and mortality was never measured. The claim record makes that distance visible on one page. Anybody who later writes the second sentence has to delete a row to do it.
It's worth running the same record on the null result that sat beside MASAI in Chapter 9, because a standard that only works on good news isn't a standard. A randomized trial of 50 physicians working structured diagnostic vignettes found median diagnostic-reasoning scores of 76 percent with conventional resources and 74 percent with access to a language model, an adjusted difference of two percentage points with a confidence interval spanning -4 to 8, not statistically significant (JAMA Network Open trial). The model alone performed strongly on the same vignettes in an exploratory analysis, and that did not translate into a demonstrated gain for the physicians who used it. Its action boundary would read: permits a statement that no workflow gain was demonstrated under the tested conditions; does not permit citing standalone model performance as evidence of improved physician decisions or patient outcomes. Capability is not benefit.
11.3.3 A Blank Claim-Record Template You Can Copy#
Copy this into a document, a spreadsheet, or a shared form. One record per consequential claim.
CLAIM RECORD
1. Identity
Claim ID:
Version:
Responsible editor:
Chapter or use case:
2. Statement (the narrow proposition, no promotional extensions):
3. Population and setting (who, where, when, institutional conditions):
4. Intervention and comparator (what changed; compared with what):
5. Outcome
Exact measure:
Denominator:
Time horizon:
Primary / secondary / exploratory / self-reported:
6. Estimate and uncertainty
Effect:
Interval (if available):
Significance as treated by the study:
7. Technology boundary (model and version if stated, or "non-generative"):
8. Source
Original URL:
Publication date:
Status (final, preprint, draft, disputed):
Relevant update or correction:
9. Limitations
Selection:
Attrition:
Measurement:
Transfer:
Sponsorship:
Conflicting evidence:
10. Action boundary
Evidence permits:
Evidence does not permit:
Approver for consequential use:
Evidence class (from the eight-class table):
Not stated in source (list fields left blank on purpose):
I added two lines at the bottom that aren't among the ten fields. The evidence-class line keeps the rung attached to the claim. The "not stated in source" line is where you write down what the original didn't tell you, so a blank field reads as a known gap rather than an oversight.
11.3.4 The Register and the Citation Ledger#
The consolidated evidence register in the back matter of this report is organized on these fields, covering the major study programs and principal permitted claims. A separate citation ledger indexes source-linked passages across every one of the 83 distinct sources cited in this report. It is an audit aid, not a claim that every cited publication is independent or equally strong. Several entries describe the same programs from different angles, and repeated appearances of one study do not turn it into multiple independent studies.
11.4 From Question to Action: A Five-Stage AI Governance Framework#
A claim record describes evidence. A process decides what to do with it. The proposed process has five stages, and each stage has an owner.
From question to action: the five-stage process
11.4.1 Stage One: Intake and Authority#
The process begins with the person's question, intended use, and authority to act. Researching options, recommending an action, approving it, and executing it are distinct permissions, and a system that can do the first has no claim on the fourth.
The owner identifies affected people, potential harms, private data, legal or professional obligations, and available non-model alternatives. If a task can be completed adequately with a simpler, lower-risk process, the evaluation should include that alternative. "The model did it" is not a result if a checklist would have done it as well.
Here's a hypothetical. A county clerk's office wants to use an assistant to answer residents' questions about permit deadlines. Intake asks: Who is asking, and what will they do with the answer? Who is authorized to state a deadline officially? What happens to a resident who relies on a wrong date? Is there a simpler alternative, such as a single maintained deadline table on the county website? If the table would work as well, the pilot has to compare against it, not against nothing.
11.4.2 Stage Two: Evidence Assembly#
The researcher retrieves original documents, records dates and versions, and distinguishes quotations from interpretation. Conflicting sources remain visible until resolved. A newer source is not automatically better if it concerns a different population or denominator.
Two cases from Chapter 7 show why this stage can't be skipped. The Department for Transport consultation-analysis evaluation illustrates why summaries require omission and invention checks: a theme generator with 0.50 precision needs sampling against the original comments, not a readability judgment. Under the evaluation's matching procedure, roughly half of generated themes were judged relevant, and a recall of 0.75 meant a quarter of human-validated themes were left out (Department for Transport evaluation). The New York City MyCity audit illustrates why denominators and grading rules must be inspectable before an accuracy figure means anything: the comptroller and the agency used different denominators and assessment choices, and the agency disputed the findings (MyCity audit and agency response).
11.4.3 Stage Three: Analysis and Challenge#
The analyst separates five categories that generated output blurs by default: observed facts, derived calculations, assumptions, forecasts, and value judgments. Another reviewer should be able to reproduce material calculations and identify which assumptions change the conclusion.
The challenge review asks what evidence would reverse the recommendation. It tests the strongest reasonable alternative explanation rather than comparing only against a weak opposing argument. A recommendation that survives only against a strawman has not been tested.
Here's the part most people skip. The five-way split is not academic. Take one sentence from a hypothetical staff memo: "The facility will raise household bills by $37 a month, which is unacceptable." Split it. Is "$37" observed or forecast? If it came from the JLARC scenario range, it's the top of a modeled 2040 range for Dominion's Virginia territory, not an observation, and not a figure about any other utility. Is "will" a calculation or an assumption? Is "unacceptable" a value judgment? It is, and that's legitimate, but it belongs in its own column. Once the sentence is split, the challenge review has something to grab.
11.4.4 Stage Four: Decision and Authorization#
The decision record states the options, relevant evidence, uncertainty, distributional effects, and accountable decision-maker. The user or authorized institution approves consequential action through its own process.
A model's ability to generate an email, application, financial instruction, or public statement does not confer permission to send or execute it. Drafting and action remain separate controls, and the separation is a feature, not friction to be optimized away.
11.4.5 Stage Five: Measurement and Correction#
The owner compares actual outcomes with the original commitments, records incidents and corrections, and reopens the decision when material assumptions fail. A successful launch is not the final outcome.
The METR developer-study update is the standing example of why. Tools changed, task selection changed, and the measured effect changed with them, which is why a process must accommodate changing tools rather than freezing a historical estimate as permanent truth (METR update). The later evidence neither erases the earlier randomized result nor supplies a clean replacement number.
11.5 A Complete AI Pilot Evaluation Protocol#
The following protocol is proposed for a school, employer, community organization, service provider, or household-support program that wants to test assistance fairly. It is not a claim that such a pilot has already been conducted.
| Element | Required specification |
|---|---|
| Objective | One primary outcome that matters to the intended beneficiary |
| Population | Eligibility, exclusions, consent, accessibility, and expected setting |
| Comparison | Current practice and, where relevant, a simpler alternative |
| Assignment | Randomization where feasible, or an explicit nonrandom design with limitations |
| Intervention | Tool version, instructions, retrieval sources, human support, and permissions |
| Cost boundary | Acquisition, implementation, use, review, correction, escalation, and exit |
| Quality boundary | Error definition, missing-information rules, independent grading, and serious-harm criteria |
| Follow-up | Immediate and later outcomes appropriate to the claim |
| Analysis | Intention-to-treat where applicable, missing-data handling, subgroup plan, and uncertainty |
| Decision rule | Conditions to continue, revise, pause, or stop |
| Publication | Favorable, null, adverse, and inconclusive findings retained |
11.5.1 Why Each Element Is There#
Most of the table explains itself once you've read the chapters before it. A few lines deserve a sentence of their own.
Objective. One primary outcome, chosen for the beneficiary, not five outcomes with the best one promoted after the fact. For a tutoring pilot, that probably means unaided learning, because Chapter 3 showed how easily assisted practice gets mistaken for it.
Population and comparison. A pilot that silently excludes people with disabilities, limited connectivity, or a different first language will report an average nobody in those groups will experience. And the comparison should include the simpler alternative from Stage One, not just "no tool."
Intervention. Tool version, instructions, retrieval sources, human support, and permissions. If any of these change mid-pilot, you're now testing a different intervention. The METR pair is the reason this line exists.
Cost and quality boundaries. Acquisition is the smallest cost line; review, correction, escalation, and exit are where real costs hide. Define an error and a serious harm before you see output, and use graders who don't know which arm produced what, so nobody gets to redefine failure after an incident.
Analysis and decision rule. Intention-to-treat means analyzing people in the group they were assigned to, including those who stopped using the tool. The decision rule (continue, revise, pause, stop) is written before results arrive. Without it, every result becomes a reason to continue.
Publication. Favorable, null, adverse, and inconclusive findings retained. This is the element sponsors most want to soften, which is why it's on the list.
11.5.2 Sample Size, Feasibility, and the Mislabel Problem#
Two elements carry the most weight.
First, sample size should be calculated from the intended effect, outcome variability, design, and acceptable uncertainty, rather than selected for convenience and then described as definitive. No universal sample size is prescribed here, because the right number depends on the question. The physician trial in Chapter 9 is a useful reminder: with 50 physicians, a confidence interval running from -4 to 8 points could not rule out a meaningful gain or a meaningful loss.
Second, a short pilot can establish feasibility or expose failures without establishing long-term effectiveness. Its public description should state which of those purposes it actually served. A feasibility study marketed as an effectiveness study is not a gray area. It is a mislabel.
Here's a hypothetical on the mislabel. A workforce program runs a six-week pilot of an assistant that helps job seekers draft cover letters. Participants finish more applications, and staff like it. That's useful feasibility evidence: the tool worked in the setting, people used it, nothing broke badly. It is not evidence that anyone got hired faster or earned more. If the program's annual report says "AI pilot improves employment outcomes," it has promoted a feasibility study to an effectiveness claim. The correct sentence is shorter and less exciting: "A six-week feasibility pilot found the tool usable; employment effects were not measured."
11.5.3 A One-Page Pilot Protocol Template#
Copy this. Fill every line before the pilot starts. Lines left blank should be marked "not applicable" with a reason.
``` PILOT PROTOCOL (one page) Program / institution: Protocol version and date: Pilot purpose: [ ] Feasibility [ ] Effectiveness (choose one)
- Objective Primary outcome (one), and why it matters to the beneficiary:
- Population Eligibility: Exclusions: Consent process: Accessibility provisions (language, disability, connectivity): Setting:
- Comparison Current practice: Simpler alternative tested (or reason none):
- Assignment Method (randomized / nonrandom design named): Known limitations of the design:
- Intervention Tool and version: Instructions / prompts: Retrieval sources: Human support available: Permissions (what the tool may draft vs execute):
- Cost boundary Acquisition / implementation / use / review / correction / escalation / exit:
- Quality boundary Error definition: Missing-information rule: Independent grading method: Serious-harm criteria and reporting route:
- Follow-up Immediate outcome timing: Later outcome timing:
- Analysis Intention-to-treat (yes / not applicable, reason): Missing-data handling: Subgroups pre-specified: How uncertainty will be reported: Sample size and how it was calculated:
- Decision rule Continue if: Revise if: Pause if: Stop if:
- Publication Where results will be posted: Commitment: favorable, null, adverse, and inconclusive findings retained Sponsor and funding disclosure:
Signed (deployment owner): Date: Signed (domain specialist): Date: ```
The two signature lines at the bottom aren't decoration. They connect the protocol to the accountability matrix in the next section.
11.6 Accountability by Role: Who Owns Which Part#
Every stage above needs an owner, and some responsibilities can't be handed to software no matter how capable it gets. Here is the proposed matrix.
| Role | Proposed responsibility | Cannot be delegated merely by using a model |
|---|---|---|
| Sponsor | Funding disclosure, scope, publication rights, and access to adverse findings | Candor about commercial interests |
| Research editor | Claim accuracy, version control, and correction process | Source fidelity |
| Domain specialist | Technical, educational, clinical, financial, or engineering review | Professional judgment within their remit |
| Deployment owner | Permissions, access, monitoring, incident handling, and exit | Operational responsibility |
| Affected person | Informed participation and a route to challenge | The burden of detecting every hidden error |
| Public authority | Lawful procedure and legitimate decisions | Statutory authority and public accountability |
Accountability by role: what cannot be delegated to a model
Read the right-hand column slowly. It's the most important column in this chapter. Each entry names something a model can help with and can never own. A model can check a citation; the research editor still owns source fidelity. A model can flag an anomaly in a scan; the clinician still owns the judgment. A model can draft a zoning summary; the public authority still owns the decision and answers for it.
The affected-person row runs the other direction. It says what must not be pushed onto the person at the end of the chain. A resident, a patient, a parent, or a job seeker should get informed participation and a route to challenge. They should not carry the burden of detecting every hidden error. If your deployment only works when the least-resourced person in the system catches the machine's mistakes, the deployment doesn't work.
11.6.1 The Sponsorship Row Applies to This Report#
The sponsorship row is this report's own, and it deserves direct treatment. This research was prepared for Savrn and concerns an industry in which Savrn has a commercial interest. It does not describe itself as an independent institutional review, and no outside peer-review panel is represented as having approved it.
The proposed publication policy, which this report applies to itself, is to preserve material unfavorable findings even when they weaken a commercial narrative. A sponsor should not be able to relabel an adverse result as missing data simply because it is inconvenient. You've seen that policy operating throughout: the null physician-workflow trial is reported as fully as the positive mammography trial, and the loneliness evidence is presented with its uncertainty intact. For loneliness, Chapter 9 reported a review that pooled three trials with 190 participants at g of -0.67 with an interval of -2.57 to 1.23 and high heterogeneity, which spans a very large benefit to a substantial harm and does not establish a reliable reduction (older-adult randomized-trial review). That result would have been easy to round up into a hopeful sentence. It wasn't.
11.7 Public Communication in Four Layers#
The proposed publication system has four layers:
- A concise public brief.
- A full research chapter.
- The claim register.
- Original-source access.
The brief should be understandable without the chapter, but it must not contradict or materially overstate it. Every layer closer to the public compresses. The rule is that compression may shorten but never strengthen.
Here's what that rule looks like in practice. Take the MASAI record from Section 11.3.2 and compress it layer by layer:
| Layer | Acceptable compression | Unacceptable strengthening |
|---|---|---|
| Original source | Full trial record | Not applicable |
| Claim register | Noninferior interval cancer; ratio 0.88 (95% CI 0.65 to 1.18); sensitivity 80.5% vs 73.8%; specificity about 98.5% both; radiologist workflow | Dropping the interval or the workflow scope |
| Research chapter | Noninferior on interval cancer, higher sensitivity without more false alarms, in a defined screening workflow | "Reduced interval cancers" |
| Public brief | In a large Swedish screening trial, model-supported reading caught cancers at least as well as standard double reading, inside a radiologist workflow | "AI prevents breast cancer" |
Each row is shorter than the one above it. None is stronger. When you check a brief, walk it backward through the layers: can every sentence be traced to the register and then to the source without gaining strength along the way?
11.7.1 The Four-Touch Outreach Sequence#
For outreach built around the seven Savrn trackers, the research in this report can support a non-promotional four-touch sequence:
- First touch, the human question: Explain one practical decision faced by a teacher, household, worker, or civic member.
- Second touch, the evidence: Present one favorable finding and its main limitation, with the original source.
- Third touch, the infrastructure connection: Show which tracker can locate a relevant record and which local facts are still needed.
- Fourth touch, the decision standard: Provide a short method for comparing alternatives, asking questions, and checking an answer.
A hypothetical sequence for a county commissioner audience might run: first, the question of how to read a data-center proposal before a hearing; second, the JLARC finding that existing Virginia rates appropriately allocated current costs at the time of its study, alongside its warning that future growth could increase system costs, both linked to the JLARC data-center report; third, how the Grid Operator Watchlist can locate a relevant tariff record, and which local interconnection studies are still needed; fourth, the eight-step workflow in Section 11.9.3. No touch asks anyone to support anything.
This is a proposed content sequence, not a claim that the campaign improves participation or decision quality. Any later campaign evaluation should measure comprehension and useful action, not just opens and clicks. The evaluation itself belongs to the pilot protocol of Section 11.5, with its favorable, null, and adverse findings all retained.
11.8 Corrections and Refresh Rules#
The proposed refresh policy uses triggers rather than pretending every source becomes stale on the same schedule.
Live records get rechecked before consequential use. Product configurations, tariffs, permits, and live project status can change without notice. The AEP Ohio tariff of Chapter 10 is exactly the kind of record that can change by commission order; the Public Utilities Commission of Ohio approved its data-center tariff settlement in July 2025 (PUCO announcement, AEP Ohio tariff process). Before anyone relies on its terms, someone should open the current version.
Historical trials stay historical, but get checked. A trial's result doesn't expire, but it can be corrected, retracted, or followed up. Before an estimate is quoted as current, check for those.
Each correction has five parts. It should identify the affected claim, the original wording, the corrected wording, the reason, and the version. Here's a blank you can copy:
CORRECTION NOTICE
Claim ID affected:
Original wording:
Corrected wording:
Reason for correction:
New version number and date:
Layers updated (brief / chapter / register / outreach):
Material corrections propagate. A material correction should reach summaries and outreach content, not remain hidden in the longest document. If the brief said it wrong, the brief gets fixed. A correction buried on page 400 while the wrong sentence keeps circulating on the one-pager isn't a correction.
The cutoff is a boundary, not a warranty. The September 23, 2026 evidence cutoff of this edition is not a promise that every linked webpage will remain unchanged or that all subsequent evidence has been incorporated.
11.9 The Seven Savrn Trackers and the Four-Layer Chain#
The seven Savrn trackers organize evidence about data-center capital, timing, public policy, water, grid governance, permits and power development, and modeled compute scarcity:
- Data Center Capital Atlas
- Data Center Delay Watchlist
- Data Center Moratorium Tracker
- Data Center Water Tracker
- Grid Operator Watchlist
- Permits and Power Development Tracker
- Scarcity Tracker
Their public value depends on the decision being made and the additional records brought to that decision. The publisher-described methods do not establish a family's bill, a student's learning, a resident's tax burden, a worker's job, an investor's return, or the net effect of a facility. Under the table in Section 11.2, the trackers' method descriptions are publisher statements, and the records they index keep whatever class they had at the source. A tracker entry starts a question. It never ends one.
11.9.1 The Four-Layer Chain#
The common structure connecting the trackers to any audience in this report is a four-layer chain:
- Tracker record: What the indexed source says, with scope, status, and observation date.
- Local or institutional record: The budget, tariff, permit, contract, employer record, school policy, or household baseline relevant to the decision.
- Causal or allocation analysis: The method connecting the project or policy to the outcome.
- Decision: The responsible person or authority weighs the evidence and alternatives.
The four-layer tracker chain and its forbidden inferences
Skipping from the first layer to the fourth produces confident but unsupported conclusions. The tracker is an evidence-discovery instrument, not a universal answer engine.
Here's a hypothetical walk through the chain. A resident sees a large capital announcement for a nearby project in the Capital Atlas and wonders whether it will lower her property taxes. Layer one: the tracker record shows the announcement, its source, its status, and when it was observed. Layer two: she needs the county budget, the assessment rules, the tax rates, and any incentive agreement. Layer three: someone has to connect the project to the tax bill, which Chapter 10 warned can come out either way, because an increase in the local tax base does not automatically reduce a household's bill and incentives belong on the same ledger as receipts. Layer four: the county's elected body makes budget decisions, and she can bring her question to them with the records in hand. The tracker got her to the right question. It didn't answer it.
11.9.2 Seven Forbidden Inferences#
Each tracker carries a specific forbidden inference. I state them here once, in consolidated form, because the pattern matters more than any single entry.
- Capital Atlas: A dollar announced is a dollar spent locally, a tax benefit to residents, a permanent job, or an investment return. It is none of these until the corresponding record shows it.
- Delay Watchlist: An entry proves failure, establishes its cause, or predicts permanent cancellation.
- Moratorium Tracker: A proposal is law, a pause is permanent, or a measure in one jurisdiction governs another.
- Water Tracker: A cooling figure proves zero total use, a permit equals actual consumption, or a regional total establishes local household harm.
- Grid Operator Watchlist: Queue position proves energization, state-wide capacity proves local deliverability, or one tariff applies to every project.
- Permits Tracker: A permit equals construction, a load request equals contracted demand, or planned generation is available capacity.
- Scarcity Tracker: A modeled index is realized revenue, a liquid market price, a guaranteed financing basis, or proof that residents benefit.
Look at the shape they share. Every forbidden inference takes a record of one kind and treats it as a record of a more consequential kind. An announcement becomes spending. A proposal becomes law. A queue position becomes power. A model becomes money. That's the same drift the evidence classes are built to stop, applied to infrastructure records.
11.9.3 The Eight-Step Audience Workflow#
The proposed interface for all of this begins with a human question rather than a tracker name. A resident asking about a bill should not need to know that the relevant evidence may appear in a grid docket, a tariff, a capital disclosure, and a public budget.
The eight-step audience workflow is proposed for evaluation:
- State the question and decision date.
- Select relevant records with dates and statuses preserved.
- Open the original sources.
- Identify the local records still missing.
- Separate observed facts, calculated values, scenarios, and preferences.
- Compare reasonable alternatives, including no change.
- Route consequential interpretation to the qualified person.
- Record the decision, assumptions, commitments, and later results.
No study reviewed in this series tests the seven trackers as a combined intervention, and this report says so wherever they appear.
11.9.4 Evaluate for Comprehension, Not Persuasion#
Evaluation, when it comes, should measure comprehension rather than persuasion. The correct initial study is not whether exposure makes residents more supportive of data centers. It is whether a source-linked workflow improves factual comprehension and decision quality, measured by correct identification of:
- project stage,
- source scope,
- relevant authority,
- unresolved evidence, and
- prohibited inference.
Participants who conclude that the evidence is insufficient should be retained in the analysis. "Unable to determine" can be the correct answer and must not be scored as a failure to adopt.
Here's a hypothetical test item to show what that means. A participant is shown a Moratorium Tracker entry describing a proposed pause in a neighboring county and asked: "Does this pause apply to a project in your county?" The correct answer is no, because a measure in one jurisdiction does not govern another. Asked "Is the pause now law?", the correct answer depends on the status field, and if the entry says "proposed," the answer is no. Asked "Will the pause be permanent?", the correct answer is that the record can't tell you. A participant who writes "unable to determine" on that last item got it right.
A short comprehension study should never be promoted as evidence of improved public budgets, family welfare, portfolio returns, or community trust. If we ever run one, the publication commitment of Section 11.5 applies to it in full.
11.10 How a School, a Utility, and a Family Would Each Use This Standard#
The instruments above can look like they were built for large institutions with research staff. They weren't. Here are three short illustrations, each labeled as a hypothetical, each using only the tools defined in this chapter.
11.10.1 A School District (Hypothetical)#
A district is offered a reading assistant for grades 5 through 8. The curriculum director starts at intake: the question is whether students read better on their own, the decision date is the spring budget vote, and the superintendent is the approver. A simpler alternative, more small-group time with a reading specialist, goes into the comparison.
At evidence assembly, the vendor's materials are classified with the eight classes. A benchmark score is a technical evaluation. A teacher survey is survey evidence. A study "showing gains" gets a claim record, and the population, comparator, and outcome fields reveal whether the gains were measured with the tool open or after it was removed.
The district then writes a one-page pilot protocol. The objective is unaided reading performance at the end of the semester. Assignment is by classroom, randomized where the schedule allows, with the limitation stated where it doesn't. The permissions line says the assistant may suggest and explain but may not submit graded work. The decision rule says the district stops if unaided performance doesn't improve relative to the comparison, whatever the satisfaction scores say. The publication line commits to posting results for parents whether they're favorable or not.
The accountability matrix names the curriculum director as deployment owner, a reading specialist as domain specialist, and parents and students as affected persons with a route to challenge. The board is the public authority. None of this required a research department. It required a page of writing and six names.
11.10.2 A Utility or Utility Commission (Hypothetical)#
A utility is asked by a large-load customer to connect a new facility, and a commissioner wants to know what it means for residential customers. The four-layer chain frames the work.
Layer one: tracker records show the project's announcement in the Capital Atlas and a related entry in the Grid Operator Watchlist. Both are observed dates and statuses, nothing more. Layer two: the utility's staff pull the actual tariff, the interconnection study, and the resource plan. If a tariff with minimum contract-capacity ramps and exit obligations exists, like the one AEP Ohio operates under, it's an administrative record of the allocation mechanism. Layer three: someone runs the allocation analysis, and the analyst separates observed load from forecast load and names which assumptions move the residential result. A scenario range like Virginia's is labeled as a scenario. Layer four: the commission decides, on the record, through its own procedure.
The claim record earns its keep at the hearing. When an advocate says the project will raise bills by a specific amount, staff can ask which class that number belongs to. When a developer says the project will pay its own way, staff can ask for the record that makes the commitment enforceable, not the press release that announces it. Both sides get held to the same table.
11.10.3 A Family (Hypothetical)#
An adult daughter is using a chat assistant to help organize her father's medications and appointments after a hospital stay. She doesn't need a protocol document. She needs three pieces of this chapter.
First, intake and authority: the assistant can help her research and draft, but her father's physician and pharmacist are the ones authorized to decide dosing, and her father is the one whose consent matters. Second, the drafting-and-action separation: the assistant can draft a question list for the next appointment, and it doesn't get to change a dose or cancel a prescription. Third, evidence classes as a filter: if she reads that an AI system caught cancers in a screening trial, she can recognize that result as belonging to a radiologist workflow, not to a chatbot reading her father's discharge papers. Chapter 9 said it directly: the trial is not evidence about consumer self-diagnosis.
She might also keep a simplified record of what she learns, a family version of the claim record: what the source was, what it said, what it didn't say, and who she needs to ask. That's four lines instead of ten. It's still a record.
11.11 Publication Gates and Actual Completion Status#
This report completes a desk-research manuscript covering the requested life stages, infrastructure questions, tracker mappings, and proposed operating process. It does not claim completion of original field trials, independent peer review, legal review, clinical review, engineering certification, or a site-specific impact assessment. A reader or institution should treat any summary that implies otherwise as inaccurate.
Before using this research as an institutional policy or project justification, the proposed external gates are:
| Gate | What it checks | Status in this edition |
|---|---|---|
| Domain review | Technical, educational, clinical, financial, or engineering accuracy | Not completed |
| Sponsor-conflict review | Whether Savrn's commercial interest shaped any claim | Not completed |
| Local-record verification | Whether the national or state evidence matches the local records | Not completed; depends on each use |
| Accessibility testing | Whether the materials work for the people they're meant for | Not completed |
| Approval by the responsible authority | Legitimate decision under the institution's own procedure | Belongs to each using institution |
Those are implementation requirements, not unfinished chapters. The distinction matters. A proposal that labels its remaining gates is ready for the next step. A proposal that hides them is not ready for any step.
11.12 Conclusion: Evidence That Stays Inspectable#
The common standard is evidence that remains inspectable from source to decision and from decision to measured result. The proposal running through every chapter of this report is to use increasingly capable systems to strengthen that chain rather than conceal its weak links: to make summaries more checkable, records more accessible, alternatives more visible, and commitments more enforceable, while keeping the decision with the person or institution authorized to make it.
That's what a commissioning record does for a power facility. It doesn't make the equipment better. It makes the equipment's condition visible to the people who have to decide whether to close the breaker. The claim record, the pilot protocol, the accountability matrix, and the four-layer chain do the same for claims about AI.
This report can support discussion and practical evaluation now. Claims of proven program effectiveness, independent validation, or community benefit should wait for the corresponding evidence. The next evidentiary step is implementation testing under the shared protocol of Section 11.5, not stronger adjectives.
11.13 Frequently Asked Questions#
What is an AI accountability framework?
In this report, an AI accountability framework is a proposed operating standard that keeps every consequential claim connected to its evidence, its limits, and the person authorized to act. It combines eight evidence classes, a ten-field claim record, a five-stage decision process, an eleven-element pilot protocol, and a six-role accountability matrix. It is a proposed workflow offered for testing, not an independently validated intervention or a legal-compliance certification.
Does following the NIST AI Risk Management Framework mean an AI tool is safe or effective?
No. NIST's voluntary framework organizes risk work around governance, mapping, measurement, and management, and its generative-model profile addresses risks such as confabulation, privacy, and information integrity. It provides a governance basis, but it does not certify a product or prove that any deployment produces benefits. The framework page also describes continuing work and revisions, so drafts should not be presented as final binding standards.
What are the eight evidence classes for evaluating AI claims?
The classes are randomized comparison, quasi-experimental analysis, observational or survey evidence, technical evaluation, simulation or forecast, administrative or official record, publisher statement, and proposed workflow. Each can establish something specific and nothing more on its own. A benchmark score shows performance on specified inputs, not real-world welfare. A publisher statement shows what an organization says, not independent verification. This is an editorial labeling system, not a formal grading method such as GRADE.
What should a claim record include?
Ten fields: identity, statement, population and setting, intervention and comparator, outcome, estimate and uncertainty, technology boundary, source, limitations, and action boundary. The goal is that a different reviewer can understand the claim and its limits without calling the author. The action boundary matters most, because it states what the evidence permits, what it does not permit, and who must approve consequential use.
How should a school or employer run an AI pilot evaluation?
Use the eleven-element protocol: objective, population, comparison, assignment, intervention, cost boundary, quality boundary, follow-up, analysis, decision rule, and publication. Pick one primary outcome that matters to the beneficiary, compare against current practice and a simpler alternative, randomize where feasible, and retain favorable, null, adverse, and inconclusive findings. A short pilot can show feasibility, but it should not be described as proof of long-term effectiveness.
What did the MASAI mammography trial actually show?
MASAI randomized 105,934 Swedish women to model-supported screening or standard double reading. Interval cancer was 1.55 versus 1.76 per 1,000 (ratio 0.88, 95% CI 0.65 to 1.18), which met noninferiority but was not a statistically significant reduction. Sensitivity rose from 73.8% to 80.5% with specificity near 98.5% in both groups. Mortality was not measured, and the result applies to a radiologist workflow, not consumer self-diagnosis.
Can the Savrn trackers tell me whether a data center will raise my electric bill?
Not by themselves. The seven Savrn trackers are evidence-discovery instruments: a tracker entry starts a question and never ends one. To reach a bill conclusion you need the four-layer chain: the tracker record, the local tariff and budget records, an allocation analysis connecting the project to rates, and a decision by the responsible authority. No study reviewed in this report tests the trackers as a combined intervention.
Has this AI accountability framework been independently reviewed?
No. It was prepared for Savrn, which has a commercial interest in the data-center industry, and it does not describe itself as an independent institutional review. No outside peer-review panel has approved it, and domain, sponsor-conflict, legal, clinical, and engineering reviews have not been completed. Its publication policy is to keep unfavorable findings even when they weaken a commercial narrative, such as the null 50-physician trial reported beside the positive MASAI result.
