# Chapter 7: AI in Local Government: Public Comments, Chatbots, and Civic Decisions

Part of The Superintelligence Transition by Chad Everett Harris, Founder and CEO, Savrn. Published September 24, 2026. Evidence cutoff September 23, 2026.

Canonical: https://savrn.com/blog/the-superintelligence-transition/ai-local-government-civic
Full edition: https://savrn.com/blog/the-superintelligence-transition

Disclosure: prepared for Savrn, which has a commercial interest in AI infrastructure and publishes the seven trackers cited here. Not an independent institutional review.

**Key takeaways**

- The Habermas Machine experiments in *Science* involved 5,734 UK participants across a series of controlled experiments and found people preferred machine-drafted group statements to human mediators' statements, which is evidence about preference, not about accuracy or legitimacy.
- In a blinded test run by the UK Department for Transport with the Alan Turing Institute (11 questions, about 9,100 responses, 165 reference themes), an AI consultation tool reached roughly 0.75 recall, 0.50 precision, and 0.59 F1 on theme generation, so it both missed real themes and added ones reviewers did not accept.
- The same evaluation's stronger live-use results were nonblinded, and its time and cost savings were modeled against an assumed manual process, not measured in randomized comparisons.
- The New York City comptroller's December 30, 2025 audit of the MyCity system found inconsistent chatbot responses and weak evaluation, and the city's technology office disputed the findings, so neither side's accuracy number can be quoted without its denominator.
- A scoping review of participatory budgeting found 37 studies across 39 articles, mostly nonrandomized and concentrated in Brazil, with suggestive benefits in some settings and no general causal guarantee.
- The defensible role for AI in local government is making the public record easier to read and check, while the authorized public body keeps the decision.
I've stood at the microphone in enough public hearings to know how a local decision really gets made. A project can live or die in the space of a few minutes, and what decides it usually isn't the engineering. It's what the people in the folding chairs understand, or think they understand, about what's being proposed, what it costs, and who carries the risk. When that understanding is shaky, the loudest version of the story wins. So when I hear people pitch AI in local government, the question I ask is the one I think most residents are quietly asking: can AI make local decisions more understandable without making them for us?

That question matters because the stakes are local and concrete. A park gets built or it doesn't. A drainage project gets funded or the next storm floods the same street. A city chatbot tells a small business owner the right rule or the wrong one. These aren't abstract policy debates. They're the decisions that shape whether people trust their own government.

Here's what the evidence in this chapter shows, including the uncomfortable part. The research is real and some of it is encouraging. People in controlled experiments preferred group statements drafted by a machine. A government-commissioned tool found most of the themes human analysts found in thousands of public comments. But that same tool, in its blinded test, also produced a large share of themes that reviewers didn't accept, and it missed about a quarter of the real ones. And the most public test of a city chatbot ended in an official audit that the city disputed, with the two sides counting success in different ways. None of this supports letting a model decide. All of it supports a narrower, more useful goal: a public record that more people can read, check, and argue about on equal terms.

### 7.1 What Can AI in Local Government Actually Do?

The practical goal of better civic research is modest, and that's exactly why it's achievable. Make budgets, proposals, rules, and alternatives understandable without requiring every resident to become a spreadsheet specialist. That goal does not mean a generated answer should decide which park to build, which drainage plan to fund, or which development to approve. That distinction is the whole chapter.

Most of the confusion about AI in local government comes from collapsing two very different jobs into one. The first job is explanation: taking a 300-page budget or a stack of engineering documents and helping a resident find the part that answers their question. The second job is judgment: deciding which competing goods a community should prioritize. The first job is about information. The second is about values and legal authority. A tool can help with the first. The evidence I reviewed gives no support for handing it the second.

#### 7.1.1 Three Forms of Civic Evidence That Must Not Be Merged

The evidence for this domain comes in three distinct forms, and they must not be blended into one comforting claim that automated civic decisions work.

First, there is a large experimental program on group statements, published in *Science*, that tests assisted deliberation under controlled conditions ([Science deliberation study](https://www.science.org/doi/10.1126/science.adq2852)). Second, there is a government-commissioned evaluation, published by the United Kingdom Department for Transport with the Alan Turing Institute in December 2025, that tests a consultation-analysis tool in a blinded coding exercise ([Department for Transport evaluation](https://assets.publishing.service.gov.uk/media/696f6641011505255b2d4203/ai-consultation-analysis-tool-CAT-evaluation.pdf)). Third, there is an official audit, issued by the New York City comptroller on December 30, 2025, that identifies shortcomings in a public-service system and records the agency's dispute of those findings ([New York City audit](https://comptroller.nyc.gov/reports/audit-report-on-the-new-york-city-office-of-technology-and-innovations-mycity-system/)).

Each form answers a different question:

| Evidence form | The question it asks | What it can support | What it cannot support |
|---|---|---|---|
| Controlled deliberation experiments | Do people prefer assisted group statements under study conditions? | Further work on assisted synthesis | Claims that a generated consensus is correct, legitimate, or durable |
| Blinded consultation evaluation | How accurately does a tool reproduce human-coded themes when graders cannot see the source? | Assisted coding inside an expert-reviewed process | Claims that a summary is complete, or that savings were realized |
| Adversarial audit | Was a deployed system managed and measured well enough to defend? | Demands for published denominators and evaluation plans | Claims that all public-service assistants fail, or that the dispute is resolved |

*Figure 7.1 (interactive on the page).*
A reader who wants all three questions answered before trusting a civic tool is not being timid. The evidence itself is organized that way. A preference result doesn't tell you about accuracy. An accuracy test doesn't tell you whether a real deployment was managed well. And an audit of one deployment doesn't tell you what a well-designed system could do. You need each piece for what it is.

#### 7.1.2 The Thesis: Legible Records, Accountable Decisions

Here's the position this chapter defends, conditional on the record above. Assistance should widen access to a common factual record and make the official process more inspectable. The decision itself, meaning the weighing of competing priorities that residents reasonably hold, stays with the authorized public body.

A tool that makes politics disappear is not a promise this evidence supports. A tool that makes the record legible is. I'll come back to that line at the end, because every study below points toward it from a different direction.

### 7.2 The Habermas Machine Study: What 5,734 Participants Showed

The most prominent research program in assisted civic deliberation is the Habermas Machine work published in *Science*. It involved 5,734 United Kingdom participants across a series of experiments ([Science study](https://www.science.org/doi/10.1126/science.adq2852)).

That number needs its qualifier immediately. This was not a single field trial of 5,734 residents making a binding municipal decision. It was a structured program of comparative experiments, and the findings describe what happened inside them. When you see this study cited as proof that "AI can find common ground for a whole country," the person citing it has dropped the qualifier.

**Evidence card**

- Title: The Habermas Machine experiments (*Science*)
- Design: A series of comparative experiments in which a model generated and revised group statements from participants' own views and critiques, compared against statements produced by human mediators; the program also examined a virtual citizens' assembly.
- Population: 5,734 United Kingdom participants across the experiments, not a single binding municipal decision.
- Finding: In the studied comparisons, participants preferred the machine's statements to those of human mediators; the virtual assembly showed convergence in expressed positions across rounds.
- Limit: Preference is not evidence of factual correctness, legal legitimacy, durability after implementation, or benefit to people outside the room; the authors frame the results as support for further investigation of assisted synthesis.
- Source: [Science deliberation study](https://www.science.org/doi/10.1126/science.adq2852)
#### 7.2.1 How the Habermas Machine Worked

The system generated and revised group statements from participants' own views and critiques. In practice, that meant iterative rounds. A model synthesized positions, participants reacted, and the synthesis was revised in light of those reactions. The comparison group was not "nothing." It was statements produced by human mediators working on the same task.

In the studied comparisons, participants preferred the machine's statements to those produced by human mediators. That's a real finding about stated preference under the experiment's conditions, and I don't want to wave it away. Drafting a statement that a disagreeing group can live with is hard work. A tool that helps people see a shared draft faster has value.

#### 7.2.2 What Preference Does and Does Not Prove

It is also a narrower finding than the headline sometimes suggests. Preference for a statement is not evidence that the statement is factually correct. It isn't evidence that the statement is legally legitimate. It isn't evidence that the agreement lasts once a policy is implemented. And it isn't evidence that the statement benefits people who weren't in the room.

Think about how a city council actually works. A resolution that everyone in the chamber likes can still rest on a wrong cost estimate. It can still exceed the council's legal authority. It can still collapse the first time the bill comes due. And it can still ignore the renters, the future residents, and the neighbors downstream who never showed up. Preference inside a study tells you nothing about any of those.

#### 7.2.3 The Virtual Citizens' Assembly and What Convergence Means

The research program also examined a virtual citizens' assembly and reported convergence in expressed positions across rounds. The authors are direct about what this supports: further investigation of assisted synthesis.

Here's the part most people skip. Convergence inside a moderated exchange is not consent to a policy. The study does not establish that a generated consensus reflects what the community would choose after costs land, after implementation stalls, or after the minority that lost the synthesis has to live with the result. People can move toward each other's language in a structured session and still disagree sharply once the tax bill or the construction noise arrives.

#### 7.2.4 Common Misreadings of the Habermas Machine Study

| Misreading | Why it's wrong |
|---|---|
| "AI mediators are better than human mediators." | Participants preferred the machine's statements in the studied comparisons. Preference is not a measure of mediation quality in real disputes. |
| "5,734 people reached consensus with AI." | The 5,734 were spread across a series of experiments, not one assembly making one decision. |
| "Convergence means the community agreed." | Convergence in expressed positions across rounds is not consent to a policy, and the authors frame it as grounds for more research. |
| "We can use this to settle local disputes." | Nothing in the study tested binding local decisions, implementation, or effects on people outside the experiment. |

#### 7.2.5 What a Defensible Civic Application Looks Like

This is where the chapter's proposed civic application takes shape. The defensible product of assisted deliberation is an inspectable summary of agreement and disagreement. It keeps minority objections and the underlying submissions visible, rather than presenting one polished paragraph as the community's final voice. A resident who disagrees with the majority conclusion should be able to find their concern, stated in their terms, with a path back to the original submissions.

Consensus and accuracy also have to stay separate even when everyone is working from the same record. Residents can reasonably accept identical cost, flood-risk, or service data and still disagree, because they attach different weights to recreation, taxes, access, safety, and distribution. A synthesis that papers over those weight differences does not resolve them. It hides them.

The Habermas Machine results justify building better tools for seeing one another's positions. They do not justify letting the tool decide which positions count.

**What this means for residents:** if a city or civic group presents a machine-drafted "common ground" statement, ask to see the submissions it was built from and the objections it did not absorb. **What this means for elected officials:** a statement people prefer is a starting draft for deliberation, not a substitute for your vote or your legal duty to consider the record.

### 7.3 AI Public Consultation Analysis: The Department for Transport Test

The strongest domain-specific evaluation in this chapter is the United Kingdom Department for Transport's December 2025 technical evaluation of a consultation-analysis tool, conducted with the Alan Turing Institute ([Technical evaluation](https://assets.publishing.service.gov.uk/media/696f6641011505255b2d4203/ai-consultation-analysis-tool-CAT-evaluation.pdf)).

This is the kind of study that matters most to people who run public comment processes. Any city, county, or agency that opens a consultation can receive far more written responses than staff can read carefully by the deadline. The promise of AI public consultation analysis is that a tool reads everything, groups it into themes, and hands staff a map. The question is how good that map is.

**Evidence card**

- Title: Consultation-analysis tool technical evaluation (UK Department for Transport with the Alan Turing Institute, December 2025)
- Design: Blinded theme-generation test comparing tool output with human-validated reference themes, plus a separate blinded response-to-theme mapping task; the report also includes nonblinded live use and modeled time and cost savings.
- Population: 11 consultation questions, approximately 9,100 responses, and 165 human-validated reference themes.
- Finding: Theme generation reached approximately 0.75 recall, 0.50 precision, and 0.59 F1; response-to-theme mapping reached approximately 0.75 F1.
- Limit: Stronger results came from nonblinded live use and must not be presented as blinded results; savings are estimates against an assumed manual process, not randomized measurements of realized savings.
- Source: [Department for Transport evaluation](https://assets.publishing.service.gov.uk/media/696f6641011505255b2d4203/ai-consultation-analysis-tool-CAT-evaluation.pdf)
#### 7.3.1 What the Blinded Test Actually Measured

The centerpiece was a blinded theme-generation test: 11 consultation questions, approximately 9,100 responses, and 165 human-validated reference themes. "Blinded" matters here. It means the graders judging whether a theme was right could not see which output came from where. That removes one of the easiest ways for an evaluation to flatter a tool, which is a reviewer who knows the machine wrote something and grades it more generously or more harshly because of that.

Two tasks were tested, and they are different. Theme generation asks the tool to produce the list of themes present in the responses. Response-to-theme mapping asks the tool to sort individual responses under themes. Generating the list is not the same as sorting individual responses under it, and this chapter won't substitute one number for the other.

#### 7.3.2 The Results: Recall, Precision, and F1 in Plain English

The blinded results were mixed, and the mixture is informative ([Evaluation findings](https://assets.publishing.service.gov.uk/media/696f6641011505255b2d4203/ai-consultation-analysis-tool-CAT-evaluation.pdf)). Theme generation achieved approximately 0.75 recall, 0.50 precision, and 0.59 F1. The separate blinded response-to-theme mapping task achieved approximately 0.75 F1.

If those terms are new to you, here's what they mean for a committee reading a summary:

- **Recall** answers: of the themes the human reference found, how many did the tool also find? A recall of 0.75 means the tool found most of the human-validated themes but left about a quarter of them out.
- **Precision** answers: of the themes the tool produced, how many were judged relevant? A precision of 0.50 means that, under the evaluation's matching procedure, roughly half of the generated themes were judged relevant, and roughly half were not.
- **F1** combines the two into one score. It's useful for comparison, but it hides which kind of error is driving the result. For a public process, you want to know both.

*Figure 7.2 (interactive on the page).*
Put those together and the failure mode is subtle. A confident, readable summary can simultaneously add themes nobody argued for and drop themes that real constituents submitted. The output looks more settled than the underlying record is.

#### 7.3.3 A Worked Example: 100 Themes, Translated Into Arithmetic

Here's a hypothetical to make those rates concrete. This is an arithmetic illustration only. It uses the stated blinded rates (about 0.75 recall and about 0.50 precision) and assumes, for simplicity, that each matched generated theme corresponds to exactly one reference theme. The real evaluation's matching procedure is described in the report; this example does not reproduce it, and no real consultation had these counts.

Suppose a human team, reading every comment carefully, validated 100 distinct themes in a hypothetical consultation.

1. **Recall of 0.75.** The tool finds 75 of those 100 themes. It misses 25. Those 25 themes were submitted by real people and are absent from the summary.
2. **Precision of 0.50.** Half of what the tool produces is judged relevant. If the 75 matched themes are the relevant half, the tool produced about 150 themes in total. That means about 75 generated themes did not match anything reviewers accepted.
3. **What the committee sees.** A summary listing roughly 150 themes, which looks thorough. Inside it, about 75 are well supported, about 75 are not, and 25 real themes are nowhere to be found.
4. **The check.** With precision of 0.50 and recall of 0.75, the F1 arithmetic comes out near 0.60, consistent with the approximately 0.59 reported.

| Hypothetical count (illustration only) | Number |
|---|---|
| Human-validated reference themes | 100 |
| Reference themes the tool found (recall about 0.75) | 75 |
| Reference themes the tool missed | 25 |
| Total themes the tool produced (precision about 0.50) | about 150 |
| Generated themes reviewers did not accept | about 75 |

Read that again. A longer list is not a better list. In this illustration, the summary is half again as long as the human reference, and it is still missing a quarter of what residents actually said. A committee member skimming it would come away thinking the process was more complete than it was.

Now ask which 25 themes are most likely to be the missing ones. The evaluation doesn't give that breakdown, so I won't pretend it does. But common sense about how summarization works says the risk is highest for concerns that are rare, oddly phrased, or raised by only a few people. Those are often exactly the concerns a public process exists to hear: the one family whose driveway floods, the business owner whose access road gets cut, the neighborhood that's always outnumbered.

#### 7.3.4 The Remedy the Evaluation Points Toward

That failure mode has a specific remedy, and the evaluation itself points toward it. The appropriate inference from the blinded results is conditional: assisted coding can be useful inside an expert-reviewed process, but omission and invention both need explicit testing.

In practice, a committee should inspect a sample of original comments rather than judging the output by readability alone. The sample should deliberately include dissenting and low-frequency concerns, which are exactly the themes a 0.75-recall tool is most likely to drop. And someone should check a sample of the tool's themes against the comments they claim to summarize, because a 0.50-precision tool will produce themes that don't hold up.

#### 7.3.5 Blinded Versus Nonblinded: A Reporting Rule

Two reporting disciplines attach to this evaluation, and both matter for anyone citing it.

First, the report's stronger results came from nonblinded live use. Those results should never be presented as though they came from the blinded test. The blinded numbers are the approximately 0.75 recall, 0.50 precision, and 0.59 F1 for theme generation, and the approximately 0.75 F1 for mapping. If a vendor slide or a staff memo quotes a higher number from this report, ask which part of the report it came from.

Second, the report's time and cost savings are estimates against an assumed manual process. That's a modeled comparison, not a randomized measurement of realized savings across consultations. An estimate against an assumed baseline is a planning input, not a demonstrated saving. It can help a budget office decide whether a pilot is worth running. It can't tell a council that the money was saved.

> **Operator's note: I've seen what happens when the modeled number becomes the promised number**
>
> When you build large power loads, you learn to keep two columns separate: what the model projected and what the meter actually read. Projections are necessary. You can't plan without them. But the moment a projected figure gets repeated in a meeting as if it were measured, it becomes a commitment that nobody actually made. Consultation tools are no different. If your city adopts one because a report estimated time savings, write down that it was an estimate, and then measure what your own staff actually spent. That's the only number you'll be able to defend a year later.
### 7.4 How to Read an AI Summary of Public Comments: A Resident's Checklist

If your city, county, school board, or transportation agency starts using AI to summarize public comments, you don't need to be a data scientist to hold it accountable. You need a few questions and the willingness to ask them out loud. Everything on this list follows from the evidence above; none of it requires new facts about any particular tool.

**Before you trust the summary, find out:**

1. **Was AI used, and for which step?** Theme generation and sorting individual comments are different tasks with different error rates, as the Department for Transport evaluation showed. Ask which one the tool did.
2. **How many comments were received, and how many were analyzed?** Show me the denominator. A theme described as "common" means nothing unless you know common out of what.
3. **Can I read the original submissions?** The summary is a derivative. The comments are the record. If the originals aren't available, the summary can't be checked.
4. **Did a person review a sample of the originals against the summary?** Ask who, how many, and whether the sample included minority and low-frequency concerns.
5. **Is my concern in there, in my terms?** Search for your own comment. If you can't find your point, or it has been merged into a broader theme that changes its meaning, say so on the record.
6. **Are there themes in the summary that you don't recognize from the hearing or the comment file?** A 0.50-precision tool in a blinded test produced many themes reviewers didn't accept. Ask staff to point to the submissions behind any theme that seems to come from nowhere.
7. **Are counts presented with their denominators?** "Many residents raised traffic" is not a count. "A stated number of the total comments raised traffic" is.
8. **Is there a correction route?** If you find an error, who do you tell, and will the correction be visible?
9. **Was there a way to comment without having your submission processed by an AI system?** A fair process offers a non-model route.
10. **Who makes the decision, and when?** The summary is an input. Make sure you know which public body decides and when you can speak to it.

**Red flags that should make you slow down:**

- A summary that reads smoothly but gives no counts, no denominators, and no links to originals.
- A claim of time or cost savings with no note on whether it was modeled or measured.
- A single "community consensus" paragraph with no section on disagreement.
- Themes attributed to "residents" that nobody at the hearing remembers hearing.
- Any statement that the tool's recommendation is the preferred option.

None of these red flags proves the summary is wrong. They tell you it hasn't yet earned your trust. That's a reasonable position for a resident to take, and it's one staff should welcome, because the checks that catch errors also protect the people who publish the summary.

### 7.5 For City Staff: A Minimum Standard Before Publishing an AI-Generated Summary

If you work inside a city, county, or agency, you're the person who has to make this work under a deadline. Here's a minimum standard I'd want in place before any AI-generated summary of public input goes to a council, commission, or the public. It's built directly from the three evidence streams in this chapter: the omission and invention risk shown in the blinded consultation test, the denominator disputes in the MyCity audit, and the gap between preference and accuracy in the deliberation experiments.

#### 7.5.1 Before You Run the Tool

- **Separate personal information from public comments.** Distinguish public comments from personal identifying information before either touches a model. Do not upload private constituent records into unapproved systems.
- **Decide which task the tool is doing.** Theme generation, response mapping, and drafting prose summaries are different jobs. Name the job in your methods note.
- **Keep a non-model route.** Residents who don't want their submission processed by a system should have another way to participate.

#### 7.5.2 Before You Publish

- **Publish the denominator.** State the total number of submissions received, the number analyzed, and any exclusions with reasons.
- **Run an omission check.** Have a qualified person read a sample of original comments, deliberately including dissenting and low-frequency concerns, and confirm each appears in the summary in a form its author would recognize.
- **Run an invention check.** For a sample of generated themes, trace each back to specific submissions. Remove or flag any theme that can't be traced.
- **Preserve disagreement.** Include a section on minority objections and unresolved disagreement. Do not collapse the record into one consensus paragraph.
- **Link to the record.** Make the original submissions available alongside the summary, subject to privacy law.
- **Label estimates as estimates.** If you cite time or cost savings, say whether they were measured locally or modeled against an assumed baseline.
- **State the decision authority.** Name the body that decides, the date, and how residents can still participate.

#### 7.5.3 After You Publish

- **Keep a visible correction log.** When a summary is found wrong, append the correction visibly rather than silently replacing the mistaken version.
- **Track correction time.** Record how much staff time went into reviewing and correcting the tool's output. That's the real measure of what the tool saved, and it will be far more useful to your budget office than a projection.
- **Collect satisfaction, but don't stop there.** Satisfaction should never substitute for comprehension or factual accuracy.

| Minimum standard item | Evidence it responds to |
|---|---|
| Omission check with minority concerns in the sample | Blinded recall of about 0.75 in theme generation |
| Invention check tracing themes to submissions | Blinded precision of about 0.50 in theme generation |
| Published denominators and exclusions | Denominator dispute in the MyCity audit |
| Estimates labeled as modeled | Modeled savings in the Department for Transport report |
| Disagreement section, not just consensus | Preference versus accuracy in the Habermas Machine experiments |
| Visible correction log | Accountability principle in section 7.9 |

This standard won't make a tool accurate. It makes the tool's errors findable. For a public body, that's the difference that matters.

### 7.6 The NYC MyCity Chatbot Audit and the Disputed Denominator

The third evidence stream is adversarial by design. On December 30, 2025, the New York City comptroller released an audit of the Office of Technology and Innovation's MyCity system ([Official audit](https://comptroller.nyc.gov/reports/audit-report-on-the-new-york-city-office-of-technology-and-innovations-mycity-system/)). The audit assessed a broader digital-service program, not a chatbot in isolation. It identified shortcomings including inconsistent chatbot responses and weaknesses in management and evaluation.

**Evidence card**

- Title: Audit of the MyCity system (New York City Comptroller, December 30, 2025)
- Design: Official audit of the Office of Technology and Innovation's MyCity digital-service program, published with the agency's response and the auditors' rejoinder.
- Population: A broader digital-service program, not a chatbot in isolation.
- Finding: The audit identified inconsistent chatbot responses and weaknesses in management and evaluation.
- Limit: The agency disputed the findings and recommendations, including the auditors' interpretation of accuracy and the use of voluntary user feedback; the two sides used different denominators and assessment choices, and the dispute is recorded, not resolved.
- Source: [New York City Comptroller audit report](https://comptroller.nyc.gov/reports/audit-report-on-the-new-york-city-office-of-technology-and-innovations-mycity-system/)
#### 7.6.1 Why Program Spending Is Not Chatbot Spending

The audit's spending figures illustrate a denominator discipline this whole report keeps returning to. The program-level spending cited in the audit is not a chatbot-only cost, and it should not be represented as one.

Think about how any public digital service is built. A system that includes case management, portal infrastructure, and human staff has a total program cost. A chatbot inside that system has a share of the cost, and the audit record does not isolate that share in any way this chapter can defend. Any chart in this report that attached the full program figure to the chatbot alone would be manufacturing a fact.

*Figure 7.5 (interactive on the page).*
You'll see this mistake in headlines and on social media. "The city spent this much on a chatbot that gave wrong answers." If the figure is the program total, that sentence is false even when every individual word in it sounds plausible. A number without a denominator is a rumor.

#### 7.6.2 The City's Dispute and the Competing Denominators

The city did not accept the audit. The Office of Technology and Innovation disputed the findings and recommendations, including the auditors' interpretation of accuracy and the use of voluntary user feedback ([Audit and agency response](https://comptroller.nyc.gov/reports/audit-report-on-the-new-york-city-office-of-technology-and-innovations-mycity-system/)). The published report contains both the agency response and the auditors' rejoinder.

The two sides use different denominators and assessment choices. They differ on which questions were tested, how refusals and non-answers were treated, and what counts as a correct answer. Each of those choices can move an accuracy figure substantially, and none of them is purely technical. Deciding that a refusal counts as "not wrong" is a policy judgment. So is deciding that a partially correct answer counts as correct.

#### 7.6.3 Two Measurement Principles Every Government Chatbot Needs

Two principles fall out of that dispute, and they apply far beyond New York.

**Voluntary feedback is not a representative satisfaction survey.** The residents who bother to rate a chatbot are not a random sample of the residents who needed it. People who got a quick answer may click a thumbs-up. People who gave up in frustration may never rate anything. People who got a wrong answer and didn't know it was wrong may rate it highly.

**An accuracy percentage is uninterpretable standing alone.** It requires the questions tested, the exclusions, the grading rules, the treatment of refusals, and the definition of a correct answer. Here's a hypothetical to make that concrete. An 85 percent accuracy figure with unknown question sampling tells a policymaker almost nothing. An 85 percent accuracy figure with the test set, grading protocol, and refusal policy published tells a policymaker something worth acting on. Same number, completely different value.

| To interpret a chatbot accuracy claim, you need | Why |
|---|---|
| The questions tested | Easy questions inflate accuracy; rare questions reveal failures |
| The exclusions | Removing hard cases changes the result |
| The grading rules | Who decided what was correct, and against what source |
| The treatment of refusals and non-answers | Counting a refusal as correct or incorrect can swing the figure |
| The definition of a correct answer | Partial answers and outdated answers need a rule |

#### 7.6.4 What the MyCity Case Does and Does Not Establish

This case should be stated precisely. It does not establish that all public-service assistants fail, and this chapter won't claim that. It establishes that a deployed system was audited, that real disagreements existed about how success was counted, and that neither side's number can be responsibly quoted without its denominator.

Distinguishing disagreement over evidence from a verified resolution is not a technicality. In public administration, it's the difference between accountability and argument by press release. If you're a resident, a reporter, or a council member, the defensible sentence is: an official audit found shortcomings, the agency disputed them, and the two sides counted differently. Anything stronger in either direction goes beyond the record.

> **Operator's note: Show me the denominator, including mine**
>
> Every operator I respect asks the same question when someone quotes a performance number: out of what? I asked it of equipment vendors. I ask it of my own team. And I expect residents to ask it of Savrn. When we describe our design goals, such as behind-the-meter power and closed-loop cooling with a zero-makeup-water design goal, those are publisher statements about what we're building toward. They're not measured results, and they deserve the same test this chapter applies to a city chatbot: what was measured, over what period, with what exclusions, and who checked. If I can't answer that, you shouldn't repeat the claim. That rule doesn't get suspended because the claim is mine.
### 7.7 Participatory Budgeting: Participation Is an Institution, Not an Interface

A scoping review of participatory budgeting, the practice of giving residents direct authority over a defined slice of public money, located 37 studies reported across 39 articles ([Campbell and colleagues](https://pmc.ncbi.nlm.nih.gov/articles/PMC6029380/)). The body of evidence is concentrated in Brazil and dominated by nonrandomized designs. It found suggestive benefits in some settings, and it did not provide a general causal guarantee of improved health or public services.

**Evidence card**

- Title: Scoping review of participatory budgeting (Campbell and colleagues)
- Design: Scoping review of published research on participatory budgeting.
- Population: 37 studies reported across 39 articles, concentrated in Brazil and dominated by nonrandomized designs.
- Finding: Suggestive benefits in some settings.
- Limit: No general causal guarantee of improved health or public services; the review predates generative models and did not test any.
- Source: [Campbell and colleagues, scoping review](https://pmc.ncbi.nlm.nih.gov/articles/PMC6029380/)
#### 7.7.1 Why a Review Without AI Belongs in an AI Chapter

That review predates generative models and did not test any. It belongs here for a structural reason. A better information interface does not remove the need for legitimate participation rules, actual budget authority, inclusion, and implementation.

Participatory budgeting works, where it works, because an institution ceded real money and real rules to a defined public. It doesn't work because residents gained a better window into a budget they still don't control. An assistant that explains the process but leaves the authority untouched has improved the interface of a process whose value was never in the interface. Capability is not benefit.

Notice also what the review's limits teach. Even for an established civic practice that predates generative AI, the evidence is mostly nonrandomized and geographically concentrated, and the benefits are suggestive rather than guaranteed. If that's the actual state of the evidence for a well-established civic institution, anyone claiming that a new AI citizen engagement tool will reliably improve outcomes is claiming far more than the research base allows.

#### 7.7.2 The Risk of a Parallel Process

The standard that follows is direct: assistance should make the official process more accessible and more inspectable. It should not create an unofficial parallel process in which only technically confident participants can influence what gets summarized.

Here's why that risk is real. If a generated summary is the version of public input that staff actually read, then whoever controls the summary controls the input. The residents without the time, language fluency, or confidence to check it have been quietly disenfranchised by a convenience. Nobody voted to shut them out. The workflow did it.

**What this means for board members and commissioners:** before approving an AI citizen engagement tool, ask whether it changes who has authority, or only how information is displayed. If only the display changes, judge it as a display tool. Don't credit it with the benefits of real participation.

### 7.8 A Civic Decision Workflow for AI in Local Government

What follows is a proposed operating process, not a field-tested municipal intervention. It's designed to support a park, drainage, service, or facility decision without inventing a local case or assuming the preferred answer. Each stage names a required work product and the verification or authority that attaches to it.

| Stage | Required work product | Verification and authority |
|---|---|---|
| 1. Define the decision | Exact question, responsible body, deadline, legal scope, affected population | Confirm against official notice and governing documents |
| 2. Establish the baseline | Existing budget, service condition, asset condition, and relevant risks | Reconcile periods, definitions, and responsible agencies |
| 3. Build alternatives | Status quo, minimum-change option, and substantive alternatives | Do not compare only the preferred option with an unrealistic alternative |
| 4. Normalize costs | Initial, recurring, maintenance, financing, contingency, and end-of-life costs | Qualified staff validate arithmetic and assumptions |
| 5. Analyze distribution | Who receives benefits, pays, loses access, or bears risk | Include renters, nonusers, adjacent residents, and future users where relevant |
| 6. Summarize participation | Themes, counts with clear denominators, original submissions, minority concerns | Human sample review and a correction route |
| 7. Record the choice | Adopted alternative, reasons, unresolved uncertainty, commitments | Authorized public body retains the decision |
| 8. Measure delivery | Actual costs, service outcomes, distribution, and complaints | Compare with the adopted commitments, not just the announcement |

*Figure 7.3 (interactive on the page).*
#### 7.8.1 Where the Model Helps and Where It Must Stop

Model assistance has a real but bounded role in this workflow. It may extract line items, explain terminology, organize questions, or draft alternative summaries. It should not invent missing engineering estimates, certify legal authority, or convert an incomplete dataset into a finding that an option is optimal.

The boundary is consistent with everything earlier in the chapter: synthesis and navigation, yes; fabrication of the record, no.

Stage by stage, that boundary looks like this:

- **Stage 1.** A model can help draft a plain-language version of the decision question. It cannot confirm the body's legal scope. That comes from the official notice and governing documents.
- **Stage 2.** A model can pull line items out of a budget document. A person reconciles periods and definitions.
- **Stage 3.** A model can draft descriptions of alternatives staff have defined. It should not invent an alternative's cost or feasibility.
- **Stage 4.** A model can lay out cost categories. Qualified staff validate the arithmetic and assumptions.
- **Stage 5.** A model can organize who is affected, based on data staff supply. It cannot fill in missing data about who lacks service.
- **Stage 6.** A model can help group comments. The omission and invention checks from section 7.5 apply in full.
- **Stage 7.** The model has no role in the choice. The authorized public body decides and records its reasons.
- **Stage 8.** A model can help compare actual results with adopted commitments. The data must come from the delivery record, not from the announcement.

#### 7.8.2 A Hypothetical Walk-Through

Here's a hypothetical to show how the stages fit together. Imagine a city deciding whether to renovate an existing recreation facility, build a new one elsewhere, or do minimal repairs. I'm not describing any real city, and there are no real numbers here.

At stage 1, staff write the exact question and name the body that decides, confirmed against the official notice. At stage 2, they document the current facility's condition and budget. At stage 3, they build three options: status quo with minimal repairs, renovation, and new construction. That rule about avoiding unrealistic alternatives matters here. If the only comparison is between a well-developed new facility plan and a deliberately weak "do nothing" option, the process has been rigged before anyone speaks.

At stage 4, every option gets the same cost categories, including maintenance and end-of-life costs, not just construction. At stage 5, staff ask who gains, who pays, and who loses access, including people who don't use the facility today. At stage 6, public comments are summarized with denominators, originals are linked, and a human sample review checks for dropped and invented themes. At stage 7, the council chooses and writes down why, including what it still doesn't know. At stage 8, a year later, someone compares what was delivered with what was promised.

At no point does the model rank the options or recommend one. At several points, it saves staff time on reading and organizing. That's the right division of labor.

#### 7.8.3 A Parks Comparison

Apply the workflow to a neighborhood park decision. The proposed evidence packet would identify construction and maintenance costs, accessibility, travel distance, utilization assumptions, drainage implications, and who currently lacks service. Each field carries its source and uncertainty. Values that can't be obtained are marked unavailable, not filled.

That last rule is the one most likely to be broken, and it's the one that matters most. A language model asked to complete a table will tend to complete it. A blank cell marked "unavailable" tells the truth. A plausible-looking number that nobody measured is a liability.

The packet's most important property is what it refuses to do. A generated ranking should not be the final product. The useful product is a comparison in which residents can change the explicit weights, such as more weight on access for carless households or less on programmed field hours, and see whether the conclusion depends on an unverified assumption. A ranking that survives weight changes is robust. One that flips the moment a resident adjusts a slider has told the community something valuable about how little the analysis supports it.

Here's a hypothetical illustration of that sensitivity test, with no real values. Two park sites are compared. When travel distance for residents without cars is weighted heavily, one site comes out ahead. When programmed field hours are weighted heavily, the other does. If the utilization assumption behind field hours turns out to be unverified, the community has learned that the case for the second site rests on a number nobody checked. That's exactly the kind of finding a public hearing should surface before a vote, not after.

#### 7.8.4 A Drainage Comparison

A drainage decision raises the stakes, because the underlying record is engineering. The packet would identify the engineering study, the design storm, the model assumptions, downstream effects, maintenance responsibilities, and the consequences if performance falls short.

A language model may help residents work through those records by translating terminology and mapping which document answers which question. It should not replace the relevant engineering assessment, and no summary should outrun the study it summarizes. If the engineering study didn't model a particular downstream effect, the summary shouldn't imply that it did. If the study's assumptions are uncertain, the summary should say so in the same place it states the conclusion.

This boundary protects both sides. The community is protected from a confident explanation of a model that was never run. The decision-maker is protected from the liability that attaches the moment an unofficial summary is treated as the technical record. An accessible explanation is valuable only if it preserves the limits of the underlying technical work.

*Figure 7.4 (interactive on the page).*
| Packet feature | Parks comparison | Drainage comparison |
|---|---|---|
| Required fields | Construction and maintenance costs, accessibility, travel distance, utilization assumptions, drainage implications, who lacks service | Engineering study, design storm, model assumptions, downstream effects, maintenance responsibilities, consequences of underperformance |
| Source and uncertainty | Tagged on every field | Tagged on every field |
| Missing values | Marked unavailable, not filled | Marked unavailable, not filled |
| Key test | Weight sensitivity: does the conclusion flip when residents adjust priorities? | Engineering boundary: does any summary claim more than the study supports? |
| Model's role | Explain, organize, let residents change weights | Translate terms, map documents to questions |
| Model must not | Produce the final ranking | Replace the engineering assessment |

**What this means for residents at a drainage hearing:** ask which engineering study the plan relies on, what design storm it assumed, and what happens if the system performs below that design. If a plain-language summary answers those questions, check that its answers match the study itself.

### 7.9 Evaluating AI Citizen Engagement: Outcomes and Safeguards

#### 7.9.1 The Primary Outcome Is Comprehension

The proposed primary outcome for evaluating civic assistance is not the number of pages summarized, the response time, or the satisfaction score. It's the proportion of residents who can correctly identify the alternatives, the major costs, the unresolved assumptions, the decision authority, and the opportunity to participate.

Comprehension is the outcome, because a process that produces fluent summaries and confused residents has failed at the only thing a democracy requires of its public. That's the whole game. If a city adopts an AI tool and residents understand their options no better than before, the tool hasn't delivered a civic benefit, whatever else it has done.

#### 7.9.2 Secondary Outcomes

Secondary outcomes should include:

- **Omission rates against human-validated themes.** The 0.50 precision and 0.75 recall from the Department for Transport evaluation are the right template: measure both what was dropped and what was added.
- **Unsupported statements** in summaries.
- **Staff correction time.**
- **Accessibility** for residents with disabilities and limited language fluency.
- **Participation breadth.** Who took part, not just how many.
- **Survival of minority views** in the final record.

Satisfaction should be collected but never substituted for comprehension or factual accuracy. A resident can be satisfied with a summary that dropped their neighbor's objection.

#### 7.9.3 Structural Safeguards

The safeguards are structural rather than technological:

- **A non-model route** for residents who don't want their submission processed by a system.
- **Privacy separation.** Avoid uploading private constituent records into unapproved systems, and distinguish public comments from personal identifying information before either touches a model.
- **An explicit correction log.** When a summary is found wrong, the correction is appended visibly rather than silently replacing the mistaken version. Silent correction is how a summary acquires an authority it never earned.

**What this means for advisers and consultants** who sell or implement these tools: propose comprehension as the success metric in your contract. If your tool can't show that residents understood more, you haven't shown what the public is paying for.

### 7.10 The Tracker Boundary in Civic Decisions

The seven Savrn trackers can help locate capital, delay, policy, water, grid, permit, and scarcity records relevant to a civic question: a data center's proposed water use, a moratorium under consideration, a permit record that anchors who committed to what. They are the [Capital Atlas](https://savrn.com/data-center-capital-atlas), the [Delay Watchlist](https://savrn.com/data-center-delay-tracker), the [Moratorium Tracker](https://savrn.com/data-center-moratorium-tracker), the [Water Tracker](https://savrn.com/data-center-water-tracker), the [Grid Watchlist](https://savrn.com/data-center-grid-operator-watchlist), the [Permits Tracker](https://savrn.com/data-center-permits-tracker), and the [Scarcity Tracker](https://savrn.com/ai-index/scarcity-tracker).

Their publisher-described scopes make them a starting index, not a complete municipal evidence file, and a tracker entry does not establish the effects of a local decision. The companion cross-audience map in the research package specifies the records that must be added before a civic conclusion is defensible: local budgets, engineering studies, service baselines, and the official notice itself.

A tracker entry starts a question. It never ends one. An infrastructure record can inform a parks-and-drainage packet, but it cannot substitute for the park's own budget or the drainage study's own assumptions. [Chapter 11](#ai-accountability-framework) develops this boundary in full, and [Chapter 10](#data-centers-community-impact) applies it to data center projects in particular.

> **Operator's note: Our own trackers get the same rule**
>
> I'd be doing exactly what this chapter warns against if I let Savrn's trackers become the answer rather than the index. They're a place to start looking. If you're a county commissioner weighing a data center proposal, a tracker entry can tell you a permit or a moratorium exists and where to look for it. It can't tell you what that project will do to your water system or your grid. For that you need the official record, the engineering, and the local baseline, and you need your own public body making the call. That's how I'd want my own projects judged, and it's how I think every project should be.
### 7.11 Conclusion: Make the Record Legible, Keep the Decision Public

The evidence supports cautious use of assistance for deliberation and document analysis. It comes with material limitations in theme detection and unresolved disagreement in public-service evaluation. It does not support delegating a community's priorities or statutory decisions to a model.

Each study earns a specific, limited conclusion. The [Habermas Machine experiments](https://www.science.org/doi/10.1126/science.adq2852) justify better synthesis tools under human inspection. The [Department for Transport consultation evaluation](https://assets.publishing.service.gov.uk/media/696f6641011505255b2d4203/ai-consultation-analysis-tool-CAT-evaluation.pdf) justifies assisted coding inside an expert-reviewed process with explicit omission and invention testing. The [MyCity public-service audit](https://comptroller.nyc.gov/reports/audit-report-on-the-new-york-city-office-of-technology-and-innovations-mycity-system/) justifies demanding published denominators before believing any accuracy claim.

So back to the question I started with: can AI make local decisions more understandable without making them for us? The evidence says it can help with the first part, under supervision, with checks that catch what it drops and what it invents. It says nothing that would justify the second part.

The goal is wider access to a common factual record, with visible disagreement and accountable decisions. That's a more defensible public benefit than promising a tool will make politics disappear. Politics, meaning the weighing of competing goods by a community that has to live with the result, is not a defect in the process. It is the process.

### 7.12 Frequently Asked Questions

#### How is AI being used in local government?

The evidence in this report covers three uses: assisted deliberation, where a model drafts group statements; consultation analysis, where a tool groups public comments into themes; and public-service chatbots. Each was tested differently, through controlled experiments, a blinded evaluation, and an official audit. None of the evidence supports letting a model make the decision itself, which should stay with the authorized public body.

#### Can AI accurately summarize public comments?

Partly. In a blinded UK Department for Transport evaluation covering 11 questions, about 9,100 responses, and 165 reference themes, a consultation tool reached about 0.75 recall and 0.50 precision in theme generation. It found most real themes but missed about a quarter, and roughly half its generated themes were not judged relevant. Human review of original comments remains necessary.

#### What did the Habermas Machine study find?

Published in Science, the Habermas Machine program involved 5,734 UK participants across a series of experiments. Participants preferred machine-drafted group statements to those from human mediators, and a virtual citizens' assembly showed convergence across rounds. Preference is not evidence of accuracy, legitimacy, or lasting agreement, and the authors frame the results as grounds for further research.

#### What did the NYC MyCity chatbot audit find?

The New York City comptroller's December 30, 2025 audit of the MyCity digital-service program identified inconsistent chatbot responses and weaknesses in management and evaluation. The Office of Technology and Innovation disputed the findings, including how accuracy was interpreted and the use of voluntary feedback. The two sides used different denominators, so the dispute is recorded, not resolved.

#### Did the MyCity chatbot cost what the audit's spending figure says?

No such conclusion is supported. The spending figures in the audit are program-level costs for a broader digital-service system, not chatbot-only costs. The audit record does not isolate the chatbot's share in a way this report can defend, so attaching the full program figure to the chatbot alone would misstate the record.

#### Does AI save cities time on public consultations?

The Department for Transport report estimated time and cost savings, but those were modeled against an assumed manual process, not measured in randomized comparisons of real consultations. Its stronger performance results came from nonblinded live use. A city considering a tool should treat the savings as a planning estimate and measure its own staff time, including review and correction.

#### Does participatory budgeting improve public services?

A scoping review by Campbell and colleagues found 37 studies across 39 articles, concentrated in Brazil and mostly nonrandomized. It found suggestive benefits in some settings but no general causal guarantee of improved health or public services. The review predates generative AI, and its lesson is that participation works through real authority over money and rules, not a better interface.

#### What should residents ask when a city uses an AI summary of public comments?

Ask how many comments were received and analyzed, whether originals are available, whether a person checked a sample including minority concerns, and whether any themes cannot be traced to real submissions. Ask for a visible correction log and a non-AI route to comment. Finally, confirm which public body makes the decision and when you can still speak.
