Chapter 2 · Capability vs. Benefit
Does AI Actually Improve Productivity? The Standard of Proof

Every vendor deck I have ever seen has a productivity number on it. Forty percent faster. Twice the output. Hours back every week. So the question most readers bring to AI productivity studies is simple: does AI make workers more productive, and if it does, is the number on the slide the number I will actually get? That is the right question, and it deserves a better answer than the slide gives it.
I have never bought serious equipment on a spec sheet alone. When you build large power loads, you learn early that a nameplate rating describes a machine under the manufacturer's test conditions, not the machine in your building, on your power, with your cooling, run by your crew, on the worst day of the summer. The spec sheet is real information. It is also the beginning of the diligence, not the end of it. Scaling a large site taught me to respect that gap between rated and delivered, because that gap is where an operator makes or loses money. The same logic applies here, with one difference: an AI productivity claim is usually a spec sheet for a person working with a tool, and people vary far more than machines do.
Here is what the evidence will show. The best experiments we have found real gains on some bounded tasks, real harm on at least one task that looked like it belonged in the same category, and a developer result that moved from a measured slowdown in 2025 to an unsettled picture in 2026. The uncomfortable part, for my industry and for anyone selling these tools, is that none of these studies supports a universal productivity multiplier. The useful part is that the same studies show exactly how to test a claim before you sign for it. That method is the subject of this chapter, and it is the standard of proof the rest of this report uses.
2.1 Capability vs. Benefit: What AI Productivity Studies Measure#
Every claim about what a capable system does for a person hides a chain of at least three questions.
- Can the system produce good output? That is a question about the model.
- Does a person, using the system for a specific task, perform that task better? That is a question about the person and the model together, in one task, on one day.
- Is the person better off in learning, in earnings, in health, in time that actually becomes theirs, once every cost of using the system is counted? That is a question about a life.
These questions have different answers, and the reviewed experiments prove it in the most direct way available: by measuring them separately and watching them diverge. Assistance improved selected writing tasks in the Noy and Zhang writing experiment in Science. It helped with some consulting tasks but harmed performance on another in the Dell'Acqua consulting working paper. And it produced different results across developer settings and study periods in the METR 2025 developer trial and the METR 2026 developer follow-up.
Three different outcomes: model output, task performance, welfare
The divergence is not an anomaly to explain away. It is the central finding of this chapter and the organizing fact for the rest of this report. A model's output quality, a user's task performance, and the user's long-term welfare are different outcomes, connected by links that can each break on their own. A system can produce fluent text that saves a professional forty minutes and still leave that professional no better at writing. A system can ace the tasks chosen to show its strengths and fail the task its users assumed fell within them. A system can speed up work in one year and slow it down in another, for reasons that have as much to do with the measurement as with the model.
Capability is not benefit. I will say that a few more times in this report, because almost every public argument about AI slides from the first question to the third without stopping at the second.
2.1.1 The Operating Question#
So the operating question is not whether a system is powerful. It is whether a specific person, using a specific configuration, for a specific task, achieves a better outcome after verification, correction, supervision, and every other cost is included.
That question is answerable. The experiments below answer it, in both directions, for four task environments. Everything else in this chapter (the benefit chain, the cost framework, the adoption process, the contract checklist) is machinery for keeping that question answerable when a vendor, a school board, an employer, or a reader's own enthusiasm tries to collapse it back into "is the system powerful?"
2.1.2 Why This Chapter Skips Benchmarks#
A note on scope. This chapter reviews task-level and workplace experiments because that is where the causal evidence is strongest. It does not review benchmark results, capability demonstrations, or model-card claims, because those answer the first question (can the system produce good output?) rather than the second and third. Benchmarks have their place. That place is not a claim about human benefit.
2.2 What the Best AI Productivity Studies Actually Show#
Four study programs anchor this chapter. Each one is presented with its design, its finding, and its boundary: the sentence that says what the result does not establish. The boundaries are not hedges. They are the load-bearing part of the evidence. A finding without its boundary is how a careful study turns into a marketing number.
2.2.1 The ChatGPT Writing Study: Noy and Zhang (2023)#
The cleanest favorable result in the reviewed set comes from a preregistered experiment published in Science in 2023. Shakked Noy and Whitney Zhang assigned 453 college-educated professionals to occupation-specific writing tasks (press releases, delicate emails, the kind of short document that fills a knowledge worker's week) and randomly gave half of them access to ChatGPT. (Noy and Zhang, Science)
Here's what the study actually found. The result was immediate and large. Participants with access completed the tasks in 40 percent less time on average, and independent evaluators rated their output 18 percent higher in quality.
The distribution matters as much as the average. The tool compressed the performance distribution: the least productive writers improved the most, so the gap between the bottom and the top of the pool narrowed. A follow-up survey found that participants in the assisted group were about twice as likely as the control group to report using the tool in their actual jobs two weeks later (34 percent versus 18 percent), a difference that persisted at two months.
Now the boundary, and it has four parts.
First, these are task outcomes. They are not measured changes in annual earnings, employment, or total organizational productivity. The experiment measured one bounded writing session.
Second, the study did not administer a later writing test without the tool, so it cannot establish that participants became more capable writers. What it established is that the person-plus-tool system, under the experiment's conditions, outperformed the person alone. Those are different findings, and the difference is the whole subject of Section 2.4.
Third, whether the saved time became output, leisure, or simply more tasks is a question the experiment deliberately left open. Forty percent less time on a press release is not forty percent more value for the organization. It is forty percent of that task's time released into something, and what that something is was not measured.
Fourth, the result describes GPT-3.5-era assistance on short, self-contained documents in 2023. It is a historical estimate for that configuration, not a permanent property of "AI."
Common misreadings. "ChatGPT makes workers 40 percent more productive" is the version that shows up in decks. It is wrong on three counts: the population was professionals doing short writing tasks, the outcome was time on that task, and the tool was a 2023 version.
What this means for you. If you are an employer looking at drafting tasks that resemble these (short, self-contained, low consequence, easy to review), this is the strongest evidence in the reviewed set that assistance can help, and it justifies a controlled test of your own. If you are a worker, the result supports using the tool for drafts you are going to review anyway. It does not support the idea that the tool is making you a better writer; nobody tested that.
2.2.2 The Jagged Frontier Study: Dell'Acqua and Colleagues#
The second study program is the one this report relies on most heavily, because it was designed to find where assistance fails. Fabrizio Dell'Acqua and colleagues ran a field experiment with 758 consultants at Boston Consulting Group, with tasks deliberately chosen on both sides of the tested model's capability frontier. Some tasks were ones the model demonstrably handled well. One complex analytical task was selected because the researchers judged it outside the model's reliable competence. (Dell'Acqua et al., working paper)
On the inside-frontier tasks, the pattern resembled the writing experiment. Assisted participants completed approximately 12.2 percent more tasks and worked roughly 25 percent faster, with higher assessed quality.
On the outside-frontier task, the result inverted. The working paper reports correct answers for:
- approximately 84.4 percent of the control group, which had no tool;
- 70.6 percent of the group with model access;
- 60 percent of the group that received model access plus an overview of the model's capabilities.
Read that again. Participants who knew the most about what the tool could do did the worst. They did worse than the group that simply had access, and dramatically worse than the group with no tool at all.
Inside the frontier assistance helped; outside it, the most-briefed group did worst
Source: Dell'Acqua et al., working paper
The boundary: the model, the tasks, and the evaluation design are bounded. A successful creative assignment does not establish reliable quantitative or strategic judgment.
But the study's lasting contribution is not the favorable number. It is the demonstration that the sign of the effect depends on the fit between task and tool, and that the users least equipped to know where the frontier lies (the ones who trusted the tool on the hardest task) absorbed the harm. The researchers' phrase for this pattern, the "jagged frontier," has been widely adopted, sometimes as a slogan. In the study itself it is a measured result: assistance helped on the tasks the researchers selected as inside, and hurt on the task they selected as outside.
That distinction matters because the slogan version gets used to excuse failures after the fact ("well, that task was outside the frontier"). The study did something harder. It picked the tasks in advance, labeled which side of the line each one sat on, and then measured. If you want to use the jagged frontier idea responsibly, you have to do the same thing: decide before the pilot which tasks you believe are inside, and test the ones you are least sure about.
A version note. This belongs here because it shows a discipline this report applies everywhere. The consulting research was subsequently published in Organization Science, according to Harvard Business School's account of the publication, and the published version carries its own figures. This chapter labels the numerical estimates above as originating in the reviewed working-paper version and does not silently combine them with different numbers in the later institutional summary. Later publication may clarify a finding. It must not become an excuse to choose whichever estimate makes the stronger promotional statement.
Common misreadings. "AI makes consultants 25 percent faster" drops the outside-frontier task entirely. "Training people on AI makes them worse" overreaches in the other direction; Section 2.3 explains what the overview result does and does not show.
2.2.3 The METR Developer Study: 2025 and the 2026 Update#
The third study program is the adverse result, and this report treats it with the same care as the favorable ones.
In early 2025, the research organization METR ran a randomized study of 16 experienced open-source developers working on 246 tasks in repositories they knew well. The developers were paid for their time and randomly assigned to work with or without early-2025 AI tools. The result: assistance increased completion time by 19 percent, with a reported interval of 2 percent to 39 percent longer. (METR, early-2025 study)
This is a small study, and its authors said so. Sixteen developers, one task environment, one generation of tools, one point in a fast-moving field. It is a historical result for that sample and setup, not a universal or current verdict on coding assistance.
But it is a randomized result, on a real task environment, with experienced professionals, which is exactly the population enthusiasts assumed benefited most. It earned its place in the public record by contradicting that assumption, and it complicates every productivity number quoted without a date attached.
In February 2026, METR reported a follow-up with 57 developers and more than 800 tasks. The point estimates moved toward speedups: approximately 18 percent less time for returning participants and 4 percent less for newly recruited participants. Their confidence intervals, however, ran from -38% to +9% for returning participants and from -15% to +9% for new ones. Both intervals include no effect.
The authors also identified selection and measurement problems that limit a reliable current estimate. Some participants avoided contributing tasks they did not want to perform without assistance, and simultaneous or asynchronous work complicated the measurement of time. The later evidence neither erases the earlier randomized result nor supplies an uncomplicated replacement productivity number. (METR, February 2026 update)
The moving developer estimate: why productivity numbers need dates
Source: METR 2025 · METR 2026 update
The boundary runs in both directions. The 2025 slowdown cannot be quoted as the current state of coding assistance. The 2026 update cannot be quoted as a settled speedup, and it should not be presented as a confirmed reversal.
What the pair establishes is subtler and more useful. The measured effect of assistance on expert work moves with the tools, the task selection, and the denominator. That is why Section 2.5 treats every productivity estimate as dated, and why Section 2.6 requires retesting after material changes.
Look closely at the selection problem, because it is not a technicality. If participants could steer away from contributing the tasks they did not want to do without the tool, then the pool of tasks being compared is no longer the pool of work the developers actually do. The tasks where the tool was most wanted may be the ones missing from the comparison. That changes what the estimate means, and it changes it in a direction you cannot read off the headline.
What this means for you. If you manage engineers, neither METR number is your number. Both are strong arguments for measuring your own team, on your own repositories, with a comparison, and for writing down the tool version when you do.
2.2.4 The Four AI Productivity Studies Side by Side#
Set side by side, the four programs show a pattern no single one of them shows. (The two METR eras appear as separate rows because they are one program measured in two periods.)
| Study program | Population and design | Finding | Boundary |
|---|---|---|---|
| Professional writing (Noy and Zhang) | 453 professionals, preregistered RCT, occupation-specific writing tasks | Time fell 40%, assessed quality rose 18% | Task outcomes, not earnings, employment, or organizational productivity; no unaided retest |
| Consulting (Dell'Acqua et al.) | 758 consultants, field experiment, tasks on both sides of the capability frontier | Inside frontier: about 12.2% more tasks, about 25% faster; outside frontier: correctness fell from 84.4% to 70.6% to 60% | Bounded tasks and model; sign of effect depends on task and tool fit |
| Developers, 2025 (METR) | 16 experienced developers, 246 tasks, randomized | Completion time rose 19% (interval 2% to 39%) | Small, historical, one setup; not a current verdict |
| Developers, 2026 (METR update) | 57 developers, 800+ tasks | Point estimates moved toward speedups; intervals include no effect; selection and measurement problems limit a reliable estimate | Neither erases 2025 nor settles a current number |
Three lessons survive the differences.
First, the effect of assistance is task-specific, not system-general. The same population of capable professionals can be helped, hurt, or unaffected depending on the task in front of them. The consulting study shows this inside a single sample.
Second, the users most at risk are often the ones with the most trust in the tool. In the consulting study, the harm concentrated among participants who had been taught the tool's strengths.
Third, every estimate in the table is dated. None of them is a property of "AI." Each is a property of a configuration (model version, task set, population, measurement method) that existed at a specific time and may not exist now.
A fourth lesson is about the reader's own inference habits. The favorable studies are the ones most often quoted, and the adverse one is most often quoted with its sample size attached while the favorable ones go bare. This report applies the same boundary discipline in both directions. The writing experiment's 40 percent is a task-level finding with no earnings evidence behind it. The METR slowdown is one randomized study of sixteen developers. Both are exactly as strong as their designs and no stronger.
2.3 Why Plausible Output Is Not Enough#
The consulting experiment's outside-frontier result deserves its own section, because it isolates the failure mode that matters most for every high-consequence setting in this report: the moment when a tool's output is fluent, confident, and wrong.
The numbers again: 84.4 percent correct in the control group, 70.6 percent with model access, 60 percent with model access plus a capability overview. The gap between the second and third groups is the finding that should keep institutional designers up at night. Training about the tool, the intervention almost every responsible-adoption program begins with, made performance worse on the task where the tool was unreliable. Participants armed with an overview of the model's strengths trusted it on a task those strengths did not cover.
2.3.1 What the Training Result Does and Does Not Show#
The result does not establish that training is useless, and this report does not draw that conclusion. It shows why the particular training, user behavior, task difficulty, and evaluation conditions must be tested rather than treated as generic safeguards.
Training that teaches what a tool can do, without teaching where it stops, manufactures the exact overconfidence the overview group displayed. A training program that wants to avoid this failure mode has to do something harder than listing capabilities. It has to show users the frontier from the wrong side, with tasks where the tool fails visibly, so that the user's calibrated distrust is built on experience rather than caveats.
In practice, that suggests a different kind of onboarding. Instead of a slide of use cases, staff work through examples from their own job where the tool produces a confident wrong answer, and they practice catching it. Whether that design works is itself a hypothesis; it has not been tested in the studies reviewed here. But it is the design the evidence points toward, and it is testable.
2.3.2 Plausible Output Is Cheap; Verified Output Is Not#
The same logic extends beyond training to output itself. Plausible output is cheap. The model that drafts a fluent paragraph, a confident diagnosis, a polished financial recommendation, or a clean block of code is producing exactly what it was built to produce: text that reads like competence.
Whether the competence is present is a separate question, answerable only by verification against something that is not the model: a source, a calculation, a measurement, a domain expert. That is why the benefit chain in the next section places task performance, not output quality, at its second link, and why the verification burden appears in the cost accounting in Section 2.5 as a first-class expense rather than an afterthought.
Here's the part most people skip. Verification is not free, and it does not scale with the tool. It scales with the person doing it and the consequence of getting it wrong. A one-paragraph internal email needs a glance. A benefits determination, a medication instruction, or a structural calculation needs someone qualified, with time, checking against an independent source. If the business case assumes the glance and the task needs the check, the savings are fictional.
2.3.3 The Classroom Version of the Same Failure#
There is a school-setting analogue that the education chapters develop fully, starting with Chapter 3. A student whose assisted practice looks excellent while unaided competence quietly erodes is producing the same kind of plausible output, and the erosion is invisible until the assistance is removed. The mechanism is identical. The cost of discovering it is paid later, by someone else, usually in a higher-stakes setting.
For parents and teachers: the signal to watch is not how good the homework looks. It is whether the student can explain it, or do a similar problem, with the tool closed. That is the independent-capability test in Section 2.4, and it is the one most often missing.
2.4 The Six-Link Benefit Chain#
How, then, should an institution reason from "this system is capable" to "this system benefits our people"? This report proposes a chain of six links, each of which needs its own evidence rather than borrowing certainty from the previous one.
| Link | Question | Example measure |
|---|---|---|
| Access | Can the intended person use the service? | Successful completion across language, disability, device, and connectivity conditions |
| Task performance | Does the workflow improve the immediate task? | Accuracy, completion time, omissions, correction burden |
| Independent capability | Can the person still understand or perform without it? | Unaided assessment, explanation, transfer to a new problem |
| Real-world outcome | Does the change matter beyond the task? | Learning retained, completed benefit application, resolved service request |
| Distribution | Who benefits and who bears costs? | Outcomes by relevant subgroup, including nonusers |
| Durability | Does the benefit persist as conditions change? | Follow-up outcome, version retest, staff turnover sensitivity |
The six-link benefit chain, and where the reviewed evidence reaches
The chain is an editorial framework for this report, not a validated universal scale. Its purpose is diagnostic. A task can improve at the second link while failing at the third or fourth, and the framework's job is to make that failure visible rather than letting the second-link success stand in for the whole chain.
2.4.1 What Each Link Catches#
Access. Access failures are invisible in averaged results. A tool that works well for English-speaking users with new laptops and high-bandwidth connections may not exist at all for the population an institution actually serves. A pilot that never measures completion across language, disability, and connectivity conditions will never know.
Task performance. These are the failures the reviewed experiments measure best. All four study programs in this chapter live mainly at this link.
Independent capability. These are the failures the experiments miss most often. The writing experiment had no unaided retest, and the education chapters show how consequential that omission can be when the population is a student rather than a professional.
Real-world outcome. This is where task gains become, or fail to become, a completed application, a resolved ticket, a retained skill. A faster draft is a task gain. A benefit application that gets approved is a real-world outcome.
Distribution. This is where averages hide. The customer-support study in Chapter 5 shows large average gains coexisting with negligible effects for the most experienced workers. Any institution that adopts on the average is making a decision about real individuals while looking at a statistic.
Durability. This is where dated estimates live. The METR pair is, in effect, a measurement of the sixth link, and its lesson is that a benefit measured once is a benefit measured once.
2.4.2 Where the Reviewed Evidence Actually Reaches#
Place the four programs on the chain and the gaps become obvious. All of them live mainly at link two. The writing study reaches toward link four only through self-reported job use, and it does not test link three. The consulting study's outside-frontier task is a direct measurement of a link-two failure, with a distribution result attached: the overview group fared worst. The METR pair, taken together, says something about link six, because the measured effect did not hold still across tool generations and study designs. None of the four measures access across language, disability, or connectivity, and none follows real-world outcomes such as earnings or organizational productivity. That is not a criticism of the studies. It is a map of what is still open.
2.4.3 Why This Report Does Not Pool Estimates#
The chain also explains this report's structural rule against pooled estimates. A finding that stops at link two and a finding that reaches link four are not two votes for the same proposition. Combining them into one average effect would manufacture a precision that neither supports. Averaging the writing study's 40 percent time reduction with the METR slowdown would give you a number that describes no task, no population, and no tool.
2.4.4 A Worked Example: A County Benefits Office#
Here's a hypothetical, to show the chain in use. It contains no real statistics.
A county benefits office is considering an assistant that drafts responses to residents' questions about program eligibility. The vendor presents a time-savings figure from its own pilot.
- Access: Does the pilot include residents who write in languages other than English, use screen readers, or contact the office by phone? If the pilot only measured web users, link one is untested.
- Task performance: Did caseworkers produce accurate responses faster, counting the time they spent checking the draft against current program rules? If review time was not logged, link two is only half measured.
- Independent capability: Can a newer caseworker still explain an eligibility rule without the tool? If the office plans to rely on the assistant for training new staff, this link matters a great deal.
- Real-world outcome: Did more residents complete applications, or get correct answers the first time? Faster drafts that lead to more resubmissions are not a benefit.
- Distribution: Did residents with complex cases get worse service while simple cases got faster? An average can hide that.
- Durability: What happens when the vendor updates the model or the program rules change? Is there a retest scheduled?
The vendor's time-savings figure, whatever it is, answers part of link two. The office's decision depends on all six.
2.5 AI Total Cost of Ownership, Not Subscription Price#
The cheapest number attached to a capable system is its subscription price, and it is almost never the relevant one. The proposed cost boundary for any adoption decision includes:
- acquisition;
- devices and connectivity;
- integration;
- staff training;
- verification;
- correction;
- human escalation;
- accessibility;
- security;
- exit.
These are analytical categories. This report does not assign a universal dollar amount to them, because the amounts are institution-specific in ways that matter.
Total cost is a stack; the subscription is one layer
2.5.1 Two Numbers That Must Stay Separate#
Two quantities should remain separate, and their separation should be visible in every internal business case.
Net staff time saved equals baseline staff time minus assisted staff time, where assisted staff time includes review and correction.
Cost per additional successful outcome equals the assisted total cost minus the comparison total cost, divided by the assisted successful outcomes minus the comparison successful outcomes.
The second calculation is meaningful only when the outcome definitions, periods, and populations are comparable and the denominator is positive and credibly estimated. If the assisted approach does not produce more successful outcomes than the comparison, the denominator is zero or negative and the ratio tells you nothing useful. If a program costs less but produces fewer successful outcomes, the result should be reported as a trade-off, not compressed into an attractive savings statistic. A school that spends less per student but graduates fewer of them has not discovered efficiency.
Show me the denominator. That is the single most useful sentence you can say in a procurement meeting. A cost-per-outcome figure without a clearly defined, comparably measured count of successful outcomes is not a figure. A number without a denominator is a rumor.
2.5.2 The Two Accounting Errors Behind Most Inflated Claims#
Two accounting errors account for most inflated claims.
The first is treating saved time as cash. Unless staffing, purchased services, or another expenditure actually changes, saved time is capacity, not savings. Whether that capacity becomes productive output, better work, or absorbed slack is another outcome to measure.
The second is omitting verification and correction from the assisted-time column. The writing experiment's participants submitted lightly edited or unedited model output at high rates. In low-consequence settings, that behavior is the source of the time saving. In high-consequence settings, it is the source of the errors. An institution that counts the draft and not the review is running a cost account that would fail an audit in any other category of expenditure.
2.5.3 The Denominator Problem in Measurement#
The denominator problem has a measurement dimension as well, and the METR follow-up illustrates it. When participants can choose which tasks to contribute, and when work happens simultaneously or asynchronously, the time denominator stops meaning what it appears to mean. A productivity evaluation can change meaning when the task pool or the time measurement changes, and the change is invisible unless the measurement method is reported alongside the estimate. (METR update)
The same error shows up in every sector, under different names:
- For a school, it is comparing only students who voluntarily used a tutor with everyone else.
- For a civic committee, it is comparing a polished machine summary against an unfinished manual draft rather than against a completed, quality-controlled alternative.
- For an employer, it is measuring only the staff who kept using the tool after the pilot and calling their numbers the effect of the tool.
Every one of these is a denominator substitution, and every one of them inflates the apparent benefit in the same direction.
2.5.4 A Maintenance Rule for Every Productivity Number#
This report proposes a maintenance rule for every productivity claim: attach tool version, task, worker experience, observation period, and evaluation design to every estimate, and recheck after substantial product or workflow changes rather than treating a historical study as a permanent coefficient. A productivity number without a date is not information. It is a fossil.
2.6 A Controlled Adoption Process in Seven Steps#
Given all of the above, how should an institution actually adopt? This report proposes a process that begins with one consequential task rather than an organization-wide promise. The owner documents the baseline, chooses a non-model alternative, defines the outcome and stopping rules, and tests a limited configuration before scaling.
The process has seven requirements.
- Specify the task. Identify what the tool may do and what it may not decide. A drafting assistant has a different risk profile from a decision system, and the specification must say which one is being adopted.
- Preserve a comparison. Use random assignment where practical, or a transparent comparison that acknowledges confounding. The four study programs in this chapter earn their authority from their comparisons; an adoption process without one earns nothing.
- Measure corrections. Count factual errors, missing exceptions, source failures, and the time spent repairing them. The correction log is the institution's version of the consulting study's outside-frontier task: it is where the true cost of fluent output becomes visible.
- Test unaided competence. Include it where education, professional judgment, or safety depends on retained skill. This is the third link of the benefit chain, and it is the one most often skipped because skipping it makes the pilot look better.
- Check distribution. Examine whether language, disability, experience, or access changes results. The average is a policy decision about individuals; the subgroup results are the actual individuals.
- Include nonusers. Preserve a functional route for people who cannot or do not choose to use the system. An adoption that removes the non-digital route has coerced adoption, whatever the consent language says.
- Retest material changes. A new model, prompt, retrieval source, or permission can alter the evaluated intervention. The METR pair is the standing example: the intervention changed, the estimate changed, and the only right response was to remeasure.
2.6.1 Stopping Rules#
A stopping rule is a condition, written before the pilot starts, that ends or pauses the pilot. Examples of the kind of rule an owner might write (as categories, not thresholds): the correction log shows errors of a type the task cannot tolerate; a subgroup does measurably worse than under the comparison; the vendor changes the model and no retest is scheduled. The point of writing them in advance is the same as the point of the consulting study labeling tasks in advance. It stops you from deciding after the fact that the bad result did not count.
2.6.2 Where NIST Fits#
NIST's AI Risk Management Framework and its generative AI profile support an approach grounded in governance, context, measurement, and risk management rather than performance testing alone. They are voluntary guidance, not empirical proof that a particular deployment is effective or compliant, and this report uses them as structure, not as certification.
If a vendor tells you its product is "aligned with the NIST framework," that is a publisher statement about process. It does not tell you the product works for your task. The seven steps above are how you find out.
2.6.3 Why the Process Is Deliberately Unglamorous#
No part of this process generates a headline, and all of it generates records. That is the point. An institution that has run this process on one consequential task has something better than confidence: it has a baseline, a comparison, a correction log, and a stopping rule. Those four artifacts are what the word "evidence-based" means when it is meant literally.
2.7 Before You Sign the AI Contract: A Buyer's Checklist#
Everything in this chapter can be turned into questions a buyer asks before signing. This checklist is derived only from the requirements above. It is a proposed workflow, not a tested instrument, and it is meant to be printed and brought into the meeting.
The task and the claim
- [ ] Which specific task will the tool perform, and what is it not allowed to decide? Is it a drafting assistant or a decision system?
- [ ] Which productivity or quality figure is the vendor citing, and which of the three outcomes (output quality, task performance, welfare) does it actually measure?
- [ ] What tool version, task, worker experience, observation period, and evaluation design produced that figure, and when?
- [ ] If the figure comes from a published study, which version of the study? Does the vendor mix figures from different versions?
The comparison
- [ ] What was the comparison group, and what did it actually receive (no tool, a different tool, a completed manual process)?
- [ ] Was assignment random? If not, what confounding is acknowledged?
- [ ] Could participants choose which tasks to include? Were only continuing users counted?
The frontier
- [ ] Which of our tasks do we believe are inside the tool's reliable competence, and which are we unsure about?
- [ ] Will our pilot include tasks where we expect the tool to fail, so staff learn to recognize failure?
The costs
- [ ] Have we priced acquisition, devices and connectivity, integration, staff training, verification, correction, human escalation, accessibility, security, and exit?
- [ ] Does our assisted-time figure include review and correction time?
- [ ] Is any "saving" in the business case tied to a real change in staffing, purchased services, or other spending? If not, have we labeled it capacity?
- [ ] Is the cost per additional successful outcome calculated with comparable outcome definitions, periods, and populations, and a positive, credible denominator?
The people
- [ ] Will we test whether staff (or students) can still perform without the tool where skill, judgment, or safety depends on it?
- [ ] Will results be broken out by language, disability, experience, and access?
- [ ] Is there a working route for people who cannot or choose not to use the system?
The contract terms
- [ ] Will the vendor notify us before changing the model, prompts, retrieval sources, or permissions?
- [ ] Do we have the right to retest after a material change, and a plan to do it?
- [ ] What are our stopping rules, written before the pilot starts?
- [ ] What does exit cost, and can we get our data out?
- [ ] Is any governance or framework claim (for example, alignment with NIST guidance) presented as a publisher statement about process rather than proof of effectiveness?
2.8 Evidence Classes: A Short Guide#
One more instrument belongs in the standard of proof, because the chapters that follow cite many kinds of evidence and the kinds are not interchangeable. This report uses an eight-class editorial labeling system throughout. It is an editorial convention, not a formal evidence-grading methodology such as GRADE, and its purpose is communicative: to keep each claim attached to what its underlying evidence can actually support.
The eight classes are randomized comparison, quasi-experimental analysis, observational or survey evidence, technical evaluation, simulation or forecast, administrative or official record, publisher statement, and proposed workflow. The full table, with what each class can and cannot establish on its own, lives in Chapter 11.
For this chapter, three points are enough.
- The writing, consulting, and METR studies are randomized or experimental comparisons. They can establish an effect of the assigned intervention under their designs. They cannot establish universal transfer or long-term benefit on their own.
- Publisher statements (model cards, announcements, vendor white papers, and Savrn's own design goals) are legitimate evidence about what an organization says. This report cites them for that purpose and never for more.
- Proposed workflows, the class to which most of this report's own recommendations belong (including the six-link chain, the seven-step process, and the checklist above), carry a standing label: implementable and testable, not yet tested here. That label is what separates a research agenda from a marketing deck.
2.9 What AI Productivity Studies Mean for You#
The studies in this chapter are about professionals, consultants, and developers, but the method applies to everyone who is being told AI will make something better. Here is how it translates.
Workers. The best evidence says assistance can speed up bounded tasks that resemble the ones tested, and that it can hurt on tasks outside the tool's reliable competence, especially if you trust it because you have been told what it is good at. Keep a habit of checking output against something that is not the tool. If your employer measures your output with the tool, ask how saved time will be used: more work, better work, or a change in your role.
Employers and managers. Do not adopt on the average. Pick one consequential task, run the seven steps, and write down the tool version. If the business case says "hours saved," ask where the hours go and whether any spending changes.
Board members and investors. Ask which of the three outcomes a productivity claim measures, and what the comparison received. Treat any figure without a date and a study design attached as a publisher statement. Ask whether verification and correction are in the cost model.
County commissioners and residents. When a public office proposes an AI tool, ask whether residents who cannot or will not use it keep a working route, and whether results will be reported by language, disability, and access. Those are links one and five of the chain, and they are the ones most likely to be skipped.
2.10 Conclusion: Value as a Defined Claim#
The reviewed evidence does not support a universal productivity multiplier. It supports testing bounded tasks and retaining both favorable and unfavorable results, including material changes in later evidence. That conclusion rests on the writing trial, the consulting trial, and the developer studies together, not on any one of them.
For the rest of this report, "value" means an identified outcome for an identified person or institution, with a defensible comparison and explicit costs. Faster or more convincing output is evidence of value only when it advances that outcome.
This definition is stricter than the one in most public discussion, and it is stricter on the enthusiasts than on the skeptics. The skeptic's position ("we don't know yet") is often the accurate state of the evidence. The enthusiast's ("this changes everything") bears the full burden of the six-link chain. I build infrastructure for this industry, and I am telling you the burden sits on our side of the table. That's the whole game.
The chapters that follow apply this standard across the lifespan: what the evidence supports for a kindergartner, a high-schooler, a college student, a new worker, a midcareer professional, a household, a civic committee, an investor, an older adult, and the community under the infrastructure. In each setting, the questions are the ones this chapter has built. What is the task? Who is the person? What is the comparison? What does it cost, all in? And does the benefit survive when the assistance is gone?
2.11 Frequently Asked Questions#
Does AI make workers more productive?
On some bounded tasks, yes. In a preregistered experiment, 453 professionals with ChatGPT finished writing tasks 40% faster with 18% higher rated quality. But a consulting experiment found assistance lowered correct answers on a task outside the model's competence, and a 2025 developer study found a 19% slowdown. The effect depends on the task, the tool version, and the person, so no single productivity multiplier is supported.
What did the Noy and Zhang ChatGPT writing study find?
In a preregistered experiment published in Science in 2023, 453 college-educated professionals given ChatGPT completed occupation-specific writing tasks in 40% less time, and evaluators rated their output 18% higher. The weakest writers improved most. The study measured one bounded session with GPT-3.5-era tools, with no unaided writing retest and no measured effect on earnings, employment, or organizational productivity.
What is the jagged frontier study?
It is a field experiment by Dell'Acqua and colleagues with 758 Boston Consulting Group consultants. On tasks inside the model's capability frontier, assisted consultants completed about 12.2% more tasks and worked about 25% faster. On a task outside it, correct answers fell from 84.4% without the tool to 70.6% with it and 60% with it plus a capability overview, per the working paper.
Did the METR developer study show AI slows programmers down?
METR's randomized 2025 study of 16 experienced open-source developers on 246 tasks found early-2025 AI tools increased completion time by 19%, with an interval of 2% to 39% longer. It is a small, historical result for one setup and one generation of tools. It is not a current verdict on coding assistance, and later evidence has not settled a replacement number.
What did METR's 2026 update show?
The February 2026 follow-up with 57 developers and more than 800 tasks moved toward speedups: about 18% less time for returning participants (interval -38% to +9%) and 4% less for new ones (interval -15% to +9%). Both intervals include no effect, and the authors flagged selection and measurement problems, so it is not a confirmed reversal.
How do I calculate the total cost of ownership for an AI tool?
Count acquisition, devices and connectivity, integration, training, verification, correction, human escalation, accessibility, security, and exit, not just the subscription. Keep net staff time saved (baseline time minus assisted time, including review) separate from cost per additional successful outcome. The second figure only works with comparable outcomes and a positive, credibly estimated denominator. This report assigns no universal dollar amounts.
Is time saved by AI the same as money saved?
No. Unless staffing, purchased services, or another expenditure actually changes, saved time is capacity, not cash savings. That capacity may become more output, better work, or absorbed slack, and which one happens has to be measured. Business cases that multiply hours saved by an hourly rate, while leaving out review and correction time, overstate the benefit.
Why can training people on AI tools make results worse?
In the consulting experiment, the group that got model access plus an overview of the model's capabilities had the lowest share of correct answers (60%) on the task outside the model's competence, below access alone (70.6%) and no tool (84.4%). This does not show training is useless. It shows training that teaches only strengths can build overconfidence, so training designs must be tested.
