SAVRN
Search Contact SAVRN

Savrn Insights · The Superintelligence Transition · Chapter 6 of 11

Disclosure: prepared for Savrn, which has a commercial interest in AI infrastructure and publishes the seven trackers cited here. Not an independent institutional review. Not investment, medical, legal, or tax advice.

Chapter 6 · Households and Families

AI for Families: What Actually Helps Households, and What Doesn't

42 min read9,551 wordsOpen as its own page

A copper kitchen table with bills, lunchbox and keys under a sketched blueprint chat bubble and phone
Savrn Insights · Chapter 6 illustration

It's 10 p.m. The kids are finally asleep, and you notice the permission slip. It's due tomorrow, it needs a signature, it mentions a fee, and the pickup time on page two doesn't match the one in the email. That is the moment most people are thinking about when they ask whether AI for parents is worth anything. Not a demo. Not a benchmark. The question is simple: will this thing take work off my plate, or will it hand me a new job checking its work at the worst hour of the day?

I want to answer that question the way I'd answer it about a power contract or a cooling system: by looking at what was measured, on whom, against what. A household is not a task list. It's a small institution with no HR department, no legal team, no on-call clinician, and no separation between its chief financial officer and the person who spots that permission slip at 10 p.m.

Here's what the evidence shows, and I'll give you the uncomfortable part up front. The burden is real, large, and unevenly shared. The interventions with measured benefits for families are almost all structured programs with a human, an institution, or an expert built into them. The direct evidence that a general-purpose chatbot saves families time is thin: a survey of comfortable respondents and a pilot with five mothers. And in the two places where the stakes are highest, medical decisions and loneliness, randomized trials found the tools pointed in the wrong direction. None of that means you should keep AI out of your house. It means you should give it a job description.

6.1 What Should AI for Parents Actually Be Asked to Do?#

Household assistance is not one outcome, because household life is not one outcome. A system can help a parent remember a deadline, complete an application, support a child's learning, reduce energy use, or help care for an aging relative. The same system can fail at diagnosis, at relationship support, at privacy, or at the fair division of responsibility between the adults in the house. Success at one of those jobs tells you nothing about the others.

The studies I reviewed for this chapter split into positive, null, mixed, and adverse findings in roughly equal measure. So there's no defensible blanket claim in either direction, for or against. The answer depends on the task, the design, and who stays in charge.

6.1.1 Which Families This Chapter Is Written For#

The practical scenario Savrn asked me to test was a familiar one: a parent raising three children while a spouse works outside the home. I use that scenario later as a worked illustration. But I deliberately widened the lens to single-parent households, dual-earner households, multigenerational homes, grandparent-led families, households shaped by disability, and people doing unpaid caregiving for a relative.

The reason is simple. A tool should not be judged only in the household most able to configure it, supervise it, and correct it. A parent with time, a fast connection, fluent English, and a background in reading fine print can catch a chatbot's mistakes. An exhausted caregiver with a spotty connection may not. If a tool only works when the user is already well resourced, that's a finding about the user, not about the tool.

6.1.2 How I Weighed the Evidence#

The standard of evidence follows the stakes. I gave the most weight to randomized field studies, administrative outcomes (did the person actually enroll, did the bill actually change), government time-use data, and peer-reviewed synthesis. I included direct evidence of general-purpose model use wherever it existed, and I say plainly when it is thin.

I also included older interventions that don't use generative AI at all: text-message programs, energy reports, human-assisted benefits outreach. Those studies show which mechanisms have measured value inside real households, and they give you a credible alternative to hold any new AI product against. If a cheap text program already moved an outcome, an expensive chatbot has to beat it, not just exist.

6.1.3 The Thesis, in One Sentence#

Here is the conditional claim this chapter defends, and I'd ask you to hold it whole rather than quoting half of it: assistance should reduce verified burden or improve a defined outcome while leaving consequential judgment, relationship responsibility, and appeal rights with people.

Every word in that sentence is doing work. "Verified" means somebody measured it. "Defined outcome" means you said in advance what better looks like. "Leaving judgment with people" means the tool can prepare the decision, not make it. Convenience, usage, confidence, and an attractive answer are how a product feels, not household outcomes. Capability is not benefit, and the gap between the two is where families get hurt.

6.2 How Much Time Do Parents Spend on Childcare? The Burden Baseline#

Before you can say a tool saves time, you need to know how much time is on the table. The best national source in the United States is the American Time Use Survey, run by the Bureau of Labor Statistics.

6.2.1 What the 2025 American Time Use Survey Shows#

According to the 2025 American Time Use Survey release from the Bureau of Labor Statistics, adults living with a child under age six averaged 2.3 hours of primary childcare per day. Women in that group averaged 2.8 hours; men averaged 1.7 hours. Adults living with a child under 13 averaged 5.1 hours of secondary childcare per day, which BLS defines as being responsible for a child while doing something else.

Figure 6.1Data

The household burden baseline

PRIMARY CHILDCARE, ADULTS WITH A CHILD UNDER 6 (HOURS/DAY)All adults2.3 hWomen2.8 hMen1.7 hSECONDARY CHILDCARE, CHILD UNDER 135.1 hSupervision folded into other activity.Different denominator: never add to primary care.2025 American Time Use Survey, about 6,100 respondents; a federal shutdown created a missing-data period BLS says cannot be quantified.
Published BLS aggregates. These establish that care takes substantial time, not that any tool can recover it.

Source: BLS ATUS 2025

Those numbers are the scale that any household tool would have to move. They are also easy to misuse. Three disciplines apply.

First, don't add them together. The two categories use different populations (a child under six versus a child under 13) and measure different things. Primary care is attention devoted to the child: feeding, bathing, reading, driving to practice. Secondary care is supervision folded into something else, like cooking dinner while a toddler plays on the floor. Adding 2.3 and 5.1 manufactures a number that no one actually lived. Show me the denominator. When two figures have different denominators, the sum is a rumor.

Second, a baseline is not a remedy. These figures establish that household care occupies a large share of the day. They do not establish that any particular system can recover any of it. A parent reading to a four-year-old is not a process in need of automation.

Third, the 2025 survey has a hole in it. BLS notes that approximately 6,100 people were interviewed and that a federal shutdown created a period of missing data whose effect on the estimates BLS says cannot be quantified. A baseline with a known hole is still the best baseline available. It is not a precise one, and anyone using it to promise specific minutes back should say so.

6.2.2 The Mental Load Nobody Puts on the Calendar#

The size of the burden matters. So does who carries it. A peer-reviewed study of U.S. parents in the Journal of Marriage and Family separates two kinds of cognitive labor: the daily mental tracking that never appears on a calendar (who's out of clean socks, which form is due, what the pediatrician said last time) and episodic household planning (the summer schedule, the move, the holiday). It finds that mothers carry more of the core daily tasks.

This is the research most people have in mind when they search for AI mental load research. Here's the part most people skip: this study describes who reports responsibility. It does not test a technological remedy. A product that cites it as proof an app can move the load is borrowing credibility it hasn't earned.

6.2.3 The Closest Direct Test: Croatian Mothers and ChatGPT#

The closest direct evidence on a general-purpose model in household use is a Croatian mixed-methods study, and it's exploratory in exactly the places that matter.

Read that sample description again. Urban, highly educated, partnered, financially stable. Those are the households best equipped to get value out of a general tool and to catch its errors. Even there, the study measured perceived value, not minutes saved. And the most important finding for anyone hoping AI would rebalance the mental load is the quiet one: responsibility stayed where it was. If the mother still owns the list, and the chatbot just helps her work the list faster, the load hasn't been shared. It's been sped up.

6.2.4 What Parents Say in Surveys#

A 2026 survey from Lurie Children's reported that 81 percent of 1,004 U.S. parents had used intelligent systems for parenting tasks, with reported average savings of 58 minutes per week.

That's a striking number, and it's worth knowing. It's also self-report. The published survey methodology does not specify a probability sampling frame, does not describe weighting, and does not give a separate respondent count for the time-savings question. So we don't know how representative the sample is, or how many people answered the 58-minute question.

The right way to read this: it's adoption evidence. It tells you parents are already using these tools and believe they help. It does not tell you what the tools cause.

6.2.5 The Gap at the Center of This Chapter#

Here's my summary of the baseline. The burden is real, large, and unevenly distributed. The evidence that general-purpose assistance reduces it is, as of this report's evidence cutoff, perceived value among comfortable survey respondents and five mothers in a two-week pilot.

That gap between the size of the problem and the strength of the remedy evidence is the central fact of this chapter. Everything that follows is about how to act sensibly inside that gap: use what's been shown to work, test what hasn't, and keep a hand on the wheel.

What this means for you as a parent: if an app claims it will give you hours back each week, ask what that number is based on. If the answer is a survey of its own users, you're looking at a satisfaction score, not a time study.

6.3 Can AI Help Families Apply for Benefits? Information vs. Completion#

If you want to know what actually helps households with paperwork, the strongest evidence I found doesn't involve a chatbot at all. It involves food assistance, older adults, and a very clean experiment.

6.3.1 The SNAP Outreach Experiment#

Here's what the study actually did. One group got nothing new, one got information about SNAP, and one got that information plus trained staff who did the heavy lifting on the application. The enrollment results in NBER Working Paper 24652 form a ladder. Over nine months, 5.8 percent of the control group enrolled. Information alone raised that to 10.5 percent. Information plus human application help raised it to 17.6 percent. The treatment differences were statistically significant.

Figure 6.2Data

Information is not completion: SNAP enrollment after nine months

0%5%10%15%20%Status quo5.8%Information only10.5%Information plus human applicationhelp17.6%About 75% of applicants were approved in every arm. Cost per added enrollee about $20 and $60, excluding applicant time,processing, and benefits.
31,888 Pennsylvania households with a Medicaid-enrolled person 60 or older not on SNAP. The effect came from helping people enter the process, not from changing approval, and the helpers were people, not software.

Source: Finkelstein and Notowidigdo, NBER w24652

Now look at the number most people skip. Among the people who actually applied, roughly 75 percent were approved in every arm. That tells you what the intervention did and didn't do. It didn't change who qualified or make the agency more lenient. It helped more people get into the process and through it. The barrier wasn't eligibility. It was completion.

The cost figures deserve the same care. The estimated intervention cost per additional enrollee was approximately $20 for the information arm and approximately $60 for information plus assistance. Those calculations left out applicant time, the government's processing costs, and the benefit payments themselves. They are the cost of the outreach, not the full cost or value of an enrollment. Anyone quoting "$60 per enrollee" as the whole economic story is quoting half a ledger.

6.3.2 What the SNAP Result Means for AI Benefits Help#

This study supports a process insight. It does not support automating eligibility.

The biggest jump on that ladder came from human application help layered on top of information. Real people assessed eligibility, handled documents, and followed up. That's a standing caution for any product that proposes to replace the helper with a chat window. Maybe a well-designed tool could reproduce part of that effect. Nobody has shown it yet, and the burden of proof sits with the tool.

What a household assistant can reasonably do is the preparatory work around the application: organize records, explain the official instructions in plain language, and help you write down the questions to ask the agency. What it must not do is decide whether you're entitled to something. An authorized institution decides that. Good assistance keeps four things intact: the original form, the agency's current rule, the exact version you submitted, and your right to a human appeal.

Common misreading #1: "Information worked, so an AI that explains benefits is enough." Information roughly moved enrollment from 5.8 to 10.5 percent. Adding hands-on human help moved it to 17.6 percent. Explanation is the smaller half of the effect.

Common misreading #2: "The approval rate stayed the same, so the help didn't matter." The approval rate staying near 75 percent is the evidence that the help worked without lowering standards. More people applied, and they were approved at the same rate.

What this means for a county benefits office or a nonprofit: if you're deciding between funding a chatbot and funding trained application assisters, this is the best evidence in the file, and it favors the assisters. If you pilot a chatbot, run it against the human option, not against nothing.

6.3.3 Lower-Risk and Higher-Risk Administrative Tasks#

Family administration is a long list: school forms, insurance claims, utility accounts, medical portals, benefits renewals. The tasks aren't equally risky, and the design should reflect that.

Lower-risk candidate tasks include:

  • Extracting dates from a school notice.
  • Creating a draft checklist from a set of instructions.
  • Comparing a draft form you've filled in with the official instructions.
  • Translating a message the family wrote, so a family member can review it before it's sent.

Higher-risk tasks include:

  • Asserting that someone is legally eligible for a benefit or program.
  • Submitting anything under penalty of perjury.
  • Accepting contract terms.
  • Spending money.
  • Disclosing a child's or a dependent adult's data.

The line between the two lists is consequence. On the first list, a mistake is annoying and catchable. On the second, a mistake can cost money, create legal exposure, or put a dependent's information somewhere it can't be pulled back.

No reviewed experiment establishes the reliability of a general-purpose model across these administrative tasks. That's worth saying twice, because product marketing tends to imply otherwise. If you're running a pilot, measure the things that matter: missed obligations, invented requirements, incorrect field entries, review time, completion, successful resolution, and correction effort. Do not measure the number of summaries generated. That's a throughput metric for a task whose value is accuracy. A tool that produces a thousand fast summaries with one invented deadline has not saved you anything on the day that deadline was the one you trusted.

6.4 AI Parenting Apps Study Results: What Structured Programs Show#

When people search for an AI parenting apps study, they usually hope to find a trial of the app on their phone. What the evidence actually offers is more useful and less flattering to open-ended chat: two randomized programs in which structure, humans, and small actionable steps did the work.

6.4.1 The Blended Digital and Human Program in China#

The trial methods describe a program adapted from ParentText, a rule-based chatbot. Caregivers received daily modules through WeChat and joined weekly or twice-weekly group discussions led by headteachers and social workers. It was not an unrestricted conversation with a generative model. It was not a software-only intervention either. That distinction is the meaning of the finding.

On the primary outcomes, caregivers in the program reported more early learning and stimulation at home (beta 1.79, 95% confidence interval 0.24 to 3.34) and less total caregiver-perpetrated violence (incidence-rate ratio 0.87, 95% confidence interval 0.80 to 0.96). Emotional violence, measured separately, did not differ significantly between groups. That null belongs in the headline alongside the positives.

The follow-up and limitations set hard boundaries. Once the immediate assessment was done, the waitlist classes got the program too, so there's no untreated comparison at twelve months and no causal long-term effect to claim. Twelve-month attrition in the original intervention group was 41.9 percent. Every outcome came from caregivers describing their own behavior.

So the study supports the blended program's immediate effects, in its setting. It doesn't show a long-term causal effect, it doesn't show the result transfers to other countries or schools, and it doesn't show a general model benefit.

6.4.2 READY4K: Three Small Messages a Week#

The READY4K methods are almost aggressively simple. Three texts a week. One shares a fact. One suggests a specific activity. One follows up with encouragement. Control families weren't left with nothing; they got occasional administrative messages, which is a fairer comparison than silence.

Among the 821 children who had spring assessments, the pooled literacy effect was 0.109 standard deviations. The results varied: the first cohort's overall effect was not statistically significant, and the larger estimates showed up among children who started with lower skills. Parent and teacher survey response was incomplete. The analysis also has a 2019 Journal of Human Resources publication record.

A 0.109 standard deviation effect is small in absolute terms. For a program that costs a few texts a week and asks parents for minutes, not hours, that's a meaningful signal, especially for the kids who started behind.

6.4.3 Why Tiny, Structured Prompts Beat Open Chat (So Far)#

Put the two studies side by side and a pattern shows up. Both programs were structured. Both used predefined content. The Chinese program put trained humans in the loop every week. READY4K sent tiny, well-timed, actionable prompts and nothing else.

The READY4K mechanism is nearly the opposite of an open-ended conversational agent. It doesn't wait for the parent to think of a question. It doesn't produce long answers. It asks for one small action that fits into a routine. That difference should guide design, not be waved past.

What these studies do not establish:

  • That more messages are better. READY4K tested three a week. It didn't test thirty.
  • That an open chatbot should talk directly with a child. Neither program did that.
  • That a generated school communication is accurate without review by the parent and the school.

What this means for parents: the best-evidenced digital help for young kids' learning looks like a short nudge you can act on in two minutes, not a conversation you have to manage. If an app is built that way, it's closer to the evidence.

What this means for teachers and school leaders: if you're choosing a family engagement tool, look for predefined, reviewed content and a fixed, light cadence. Ask the vendor what the comparison group received in any study they cite.

6.5 AI Caregiver Support for Dementia: A Small Average, a Wide Future#

Many households aren't just raising kids. They're caring for a parent or grandparent, sometimes both at once. Dementia caregiving is one of the heaviest forms of unpaid work a family can take on, and it's an area where digital support has been studied for years.

6.5.1 The Meta-Analysis: What 35 Trials Show#

In the JMIR review, the pooled burden effect was SMD -0.26 (95 percent confidence interval -0.42 to -0.10). Negative means less burden. The interval doesn't cross zero, so on average, the evidence favors intervention.

Then comes the number that should change how you read the first one. Heterogeneity was 73.6 percent, which means the trials disagreed with each other far more than chance alone would explain. And the 95 percent prediction interval ran from -1.10 to +0.58.

Figure 6.6Data

Average effect versus the next implementation: dementia caregiver eHealth

-1-0.500.5Standardized mean difference in caregiver burden (negative favors intervention)Pooled burden effect (95% CI)35 RCTs, 3,388 caregiversSMD -0.26 (-0.42 to -0.10)Prediction interval (95%)heterogeneity 73.6%-1.10 to +0.58
The confidence interval describes the average. The prediction interval describes what a new program might do, from substantial benefit to harm. Only three studies were low risk of bias; certainty moderate.

Source: JMIR 2026 systematic review

Read carefully, this is two findings, and they answer two different questions.

  • The confidence interval answers: what's the average effect across programs like these? Answer: a modest reduction in burden, and we're fairly sure it's real.
  • The prediction interval answers: what might the next specific program do? Answer: anywhere from a substantial benefit (-1.10) through little or no effect to an adverse result (+0.58), where caregivers end up more burdened.

If you're a caregiver deciding whether to try one particular app, the second question is the one you're actually asking. You aren't buying the average. You're buying one program, and the evidence says that program's effect is hard to predict from the category.

The review's risk of bias and certainty assessment adds more caution. Only three of the 35 studies were rated low risk of bias. The overall evidence was judged moderate certainty, and it relied heavily on subjective outcomes, meaning caregivers rating their own burden. The subgroup analysis found that human-supported interventions carried a larger point estimate than self-guided ones, but the test for a difference between those subgroups was not significant. And because the review covered web, mobile, video, and hybrid programs, it does not isolate what a general-purpose generative model does.

My summary: caregiver technology support is promising, modest on average, and unpredictable for any specific deployment. There's a hint, not a demonstration, that the versions with a human involved work better.

6.5.2 The Expert-Reviewed Recommender Trial#

Notice the design feature of the Age and Ageing trial before anything else: every plan the system generated passed two senior dementia-care experts before a caregiver saw it. This was supervised by construction.

On the main outcome, the burden contrasts were null between groups. Among the 201 participants who completed follow-up, burden went down within the intervention arm, but the differences between the intervention and control groups at weeks six and twelve were not statistically significant. A within-arm decline without a between-group difference is exactly the pattern you'd expect if some of the improvement would have happened anyway. That's why control groups exist.

The reported safety results point the other way. The authors report 16 intervention participants and 29 controls with at least one safety event, p = .037. No individual category of safety event differed significantly on its own.

6.5.3 Two Inconsistencies I'm Recording, Not Fixing#

I found two inconsistencies in this paper that I can't resolve from the published text, and I'm recording them rather than smoothing them over.

First, the safety table reports the control group's percentage as 28.7 percent. But the paper states 103 control completers, and 29 affected participants out of 103 comes to approximately 28.2 percent. Those published counts and the percentage don't reconcile.

Second, the paper's analysis statements say both that all 201 completers were analyzed and that the analysis was intention-to-treat. Those can't both describe the same analysis in the usual sense, because intention-to-treat analyzes everyone randomized, which was 250.

Neither issue proves the result wrong. Both mean I keep the safety finding as reported and do not convert it into a precise promise that this kind of system reduces household risk by some specific amount.

The study methods and limitations add attrition, complete-case presentation, the heavy expert review, and the short 12-week window. Put it together, and the trial supports further testing of a supervised knowledge-graph system. It does not support autonomous caregiving, and it does not support replacing clinical evaluation.

6.5.4 What This Means for Caregivers and Older Adults#

If you're caring for someone with dementia: digital support has a real, modest average benefit in trials, and the ones with human support built in look at least as promising as the ones without. Before you commit, ask whether a person reviews what the tool tells you and whether you can reach a human expert. The best-designed trial in this section had both.

If you're an older adult whose family is setting up tools on your behalf: you should be told what the tool records about you, who sees it, and how to turn it off. That applies to the dependent adult as much as to the child.

If you're a health system or agency buying caregiver support: remember the prediction interval. The category's average won't tell you what your specific program will do. Plan to measure burden against a comparison group from day one.

6.6 Home Energy Reports: What Household Feedback Can Measure#

Energy is where household decisions and infrastructure meet, and it's where I have the most operating experience. So I want to be especially careful here.

The mechanism in Allcott's 2011 analysis is social comparison: you see how your use stacks up against your neighbors, plus a few tips. Because high-use households cut more than low-use households, the 2 percent average isn't the result for every family.

The boundaries here are multiple, and each one is load-bearing.

The outcome definition was kilowatt-hours. Not a guaranteed lower bill, not better reliability, not a model-based agent, and not a community's acceptance of a data center.

The welfare effect is ambiguous. The author notes that the household's costs of conserving weren't observed. Lower use bought with a colder house, extra effort, or forgone convenience isn't automatically a net gain for the family.

A useful household energy interface has to show its work. For any household infrastructure question, a good tool should show the consumption period, the rate structure, the units, the comparison group, and the actions actually available to that household. And it should never imply that resident conservation offsets an unmeasured industrial load, or shift responsibility for system planning onto families.

What this means for residents: if a utility or a company shows you an energy comparison, check the period, the units, and who you're being compared with. A lower number of kilowatt-hours is a measured result. A lower bill, a more reliable grid, or a "fair share" of a new facility's impact are separate claims that need separate evidence.

6.7 Where AI for Parents Must Stop: Medical Advice and Loneliness#

Two randomized experiments mark the firmest boundaries in this chapter. Both are about things families already use chatbots for: figuring out whether a symptom is serious, and having someone to talk to.

6.7.1 The ChatGPT Medical Advice Study: Strong Models, Weaker Decisions#

The study methods test the thing that actually matters. Researchers didn't just grade a model's answer. They measured what the people using it concluded.

When researchers prompted the models directly, each one performed strongly. The people using those models did not reproduce that performance. According to the human decision results, control participants had 1.76 times the odds of identifying a relevant condition compared with the pooled model users. On disposition, meaning what level of care to seek, accuracy did not differ significantly between each model group and control. Every group, with or without a model, tended to underestimate how urgent the situation was.

Read that again. The model knew. The person with the model did worse at naming a relevant condition than the person without it.

The limits are real. These were simulated scenarios, not people who were actually sick and scared. And the models have changed since the 2024 data collection. But the study shows why families shouldn't infer safe decision support from a model's medical test scores. The gap between what a model can do when an expert prompts it and what a worried parent decides after using it is the whole story, and in this trial it ran in the direction of harm.

What that means in your house: a household system can help you organize symptoms and list questions for a professional. It is not the final authority on emergency care.

6.7.2 The AI Loneliness Study: Daily Chats, More Loneliness#

The design and sample are large for this kind of question. Because the daily conversations were encouraged, not required, the study also measured whether people followed through. Per the intervention and first-stage result, assignment increased the number of days with personal conversations by about 5.8.

Now the outcome. The loneliness estimate went up, not down: a single-item loneliness measure increased by 0.168 points on a 0 to 10 scale, which is 0.057 standard deviations, with a multiplicity-adjusted p-value of .0002. The other primary results showed no significant effect on happiness, and several other well-being estimates were small and only marginal after adjusting for multiple comparisons.

The effect is small in size. It is also precisely estimated and in the opposite direction from the usual pitch.

The limits cut both ways. It was an encouragement design. It lasted one month. There was no alternative reflection control, such as asking a comparison group to keep a journal, so we can't tell whether the effect comes from talking to an AI specifically or from spending daily time on personal reflection. It's a working paper, and it didn't test specialized therapy tools. So it does not show that every structured support tool harms well-being.

What it does show: you cannot presume that replacing or supplementing personal interaction with unstructured model conversations will reduce loneliness. In this trial, the measured effect went the other way.

6.7.3 Two Boundaries Backed by Adverse Trials#

Figure 6.4Comparison

Two firm household boundaries, each backed by a randomized trial

MEDICAL SELF-ASSESSMENTControls had 1.76 times the odds of naming arelevant condition1,298 U.K. adults, ten physician-written vignettesGPT-4o, Llama 3, Command R+ vs usual resourcesModels scored well alone; users did notSimulated cases, 2024-era modelsPERSONAL CONVERSATIONLoneliness rose 0.168 on a 0 to 10 scaleAnalytic n = 12,356 French adults0.057 SD; adjusted p = .000228-day encouragement designWorking paper; no journaling control
A household system organizes and escalates. It does not diagnose, and it does not substitute for relationships.

Source: Bean et al., Nature Medicine · Fréget, Reshef, and Senik, CESifo

Taken together, these two studies set the chapter's firmest household boundaries. A household system organizes and escalates. It does not diagnose. It does not substitute for relationship.

These aren't stylistic cautions or my personal preference. Each boundary rests on a preregistered randomized experiment with a statistically significant adverse result. That's a higher evidentiary bar than most of the benefits claimed for these tools.

For older adults and the families who care about them: a companion app is not a plan for isolation. People are.

6.8 The Household Evidence Register: Strong vs. Perceived Evidence#

Here is every entry that informs this chapter, with what each contributes and where each stops.

ID Evidence and design Contribution Principal boundary
H01 2025 ATUS; weighted national time-use survey Household-care time baseline Not an intervention effect; missing 2025 diary period
H02 Croatian survey n=369 plus uncontrolled five-person pilot Direct exploratory evidence on generative household organization No objective time outcome, randomization, control, or redistribution result
H03 Chinese cluster RCT, n=541 caregivers Immediate parenting and child-protection outcomes from blended rule-based and human support One preschool, self-report, no long-term untreated comparison
H04 READY4K randomized program, n=1,031 families Low-burden prompts and child literacy outcome Non-generative program; incomplete surveys, cohort variation
H05 SNAP outreach RCT, initial n=31,888 households Information and human application-assistance effects Older likely-eligible Pennsylvania population; not a model intervention
H06 Home Energy Reports, about 600,000 households Measured reduction in household electricity use Feedback program, not a facility-impact or generative-system study
H07 Dementia eHealth meta-analysis, 35 RCTs, n=3,388 Small average caregiver-burden effect with explicit heterogeneity Not specific to general-purpose generative models
H08 Knowledge-graph dementia-care RCT, n=250 randomized Personalized, expert-reviewed recommendations and safety-event result Short, selected sample; burden between-group estimates null
H09 Medical-vignette RCT, n=1,298 adults Adverse evidence on human use of language models for condition identification Simulated scenarios, three 2024-era models
H10 Personal-conversation preregistered RCT, analytic n=12,356 Adverse evidence on unstructured personal conversation and loneliness One-month encouragement design; working paper; not specialized therapy
H11 Lurie Children's 2026 survey, n=1,004 parents Reported uses, perceived burden, time savings Adoption survey without causal or observed-time measurement
Figure 6.5Map

Where household evidence is measured and where it is only perceived

RANDOMIZEDH03 parentingH04 READY4KH05 SNAPH06 energyH08 dementia rec.H09 medicalH10 lonelinessMETA-ANALYSISH07 dementia eHealthOFFICIAL STATISTICSH01 ATUSSURVEY / SELF-REPORTH11 Lurie surveyEXPLORATORY PILOTH02 Croatia
The eleven household register entries by evidence class. The only direct evidence on general-purpose household use (H02, H11) is exploratory or self-reported.

Two reading rules. First, the register counts evidence entries, not eleven independent randomized trials. H07 synthesizes many studies that may overlap with individual trial entries, so its 3,388 caregivers must not be added to the others as if they were a separate pool. Second, publication status is not the same as design. H10 is a working paper from CESifo. H11 is a public survey report from Lurie Children's. H01 is official statistics from BLS. H02 is exploratory mixed-methods research. Each deserves the weight of what it is, no more. The full evidence-class scale is defined in Chapter 11.

Scan the register by what was measured, and the split is plain. The measured outcomes (enrollment, literacy scores, kilowatt-hours, condition identification, loneliness scores, caregiver burden in trials) come almost entirely from structured programs or from adverse tests of general models. The evidence that general-purpose AI for parents saves time comes from self-report and a five-person pilot. That's the difference between strong evidence and perceived evidence, and it's the difference that matters when someone asks you to pay for a subscription.

6.9 A Household Operating Process for AI for Parents#

So what should a family actually do? Below is a proposed control process. I want to be clear about its status: it's a design derived from the evidence above, not a tested end-to-end household intervention. Each stage leaves an evidence trail and has a gate.

6.9.1 The Ten Stages#

Stage Household action and system limit Evidence retained Gate
Define the task A family member states the concrete need and desired outcome Request, owner, deadline, affected people Do not begin from continuous surveillance or hidden inference
Classify consequence Sort as informational, administrative, financial, educational, health, safety, legal, or relational Risk class and reason Health, safety, legal, financial execution, and dependent-care decisions require heightened review
Minimize data Use only necessary information; redact child, health, financial, location, identity data where possible Data fields, consent, retention setting No dependent's sensitive data without authority and justified need
Acquire sources Retrieve current school, agency, utility, provider, contract, or original public records Original URL or document, issuer, date, version Model output is never its own source
Draft assistance Generate a summary, checklist, comparison, question list, or draft Tool and version, prompt, output No autonomous submission or purchase by default
Verify Compare every consequential field and claim with the original record or qualified professional Reviewer, discrepancies, corrected version Unresolved conflicts remain visible
Approve The authorized adult chooses, edits, refuses, or escalates Final approval and scope Silence and nonresponse are not consent
Execute narrowly A person submits or explicitly authorizes a constrained action Exact payload, destination, timestamp, receipt Separate approval for spending, disclosure, commitment, or third-party communication
Monitor outcome Record completion, error, benefit, burden, unintended consequences Outcome and comparison with baseline Usage is not success
Correct and delete Appeal, retract, notify affected people, apply retention policy Correction record and deletion status A household must be able to stop use and recover

That looks heavy for a permission slip. In practice, most of it takes seconds: you know who needs what by when, you know it's low stakes, you don't paste in your kid's medical history, and you check the draft against the paper. The table matters most when the task gets serious, because that's when people skip steps.

6.9.2 The Task Authority Matrix#

The companion matrix assigns each kind of task a default role for the system.

Task type Default role Human authority required
Calendar extraction, draft checklist, meal ideas, household inventory Suggest Adult reviews allergies, dates, cost, feasibility
School summary, translation, benefits preparation, utility comparison Prepare Adult checks original records, approves communication or submission
Purchase, application filing, account change, data sharing Execute only with specific approval Account holder reviews exact terms, destination, amount, data
Symptom triage, medication, child protection, emergency decision Organize and escalate only Qualified professional or emergency service remains authoritative
Relationship judgment, discipline, emotional support Offer bounded reflection, not substitution People retain responsibility; serious risk escalates to human support
Figure 6.3Framework

The household authority ladder

SuggestCalendar extraction, checklists, meal ideas, inventoryAdult reviews dates, allergies, costPrepareSchool summaries, translation, benefits prep, utility comparisonAdult checks originals, approves sendingExecute only with specific approvalPurchases, filings, account changes, data sharingAccount holder reviews exact termsOrganize and escalate onlySymptoms, medication, child protection, emergenciesProfessional or emergency service decidesBounded reflection, not substitutionRelationship judgment, discipline, emotional supportPeople keep responsibilityconsequence rises, system role narrows
Proposed default roles for a household assistant, from the ten-stage operating process. The higher the consequence, the narrower the system's role, and authority never leaves the person.

The architecture of that table is the chapter's thesis in one picture: the higher the consequence, the narrower the system's default role, and the authority never leaves the person.

6.9.3 A Family's AI Rules: What to Let It Do, What to Never Let It Do#

If you want something you can print and stick on the fridge, here's the authority matrix translated into house rules.

Let it do these, and check its work:

  1. Pull dates, times, fees, and required items out of school notices and put them in a draft list.
  2. Draft checklists, grocery lists, and meal ideas, with an adult checking allergies, prices, and what's actually in the pantry.
  3. Summarize a school or agency document, as long as you can see which line of the original each point came from.
  4. Translate a message the family wrote, so someone can read the translation before it goes out.
  5. Help prepare a benefits application by organizing documents and explaining the official instructions.
  6. Compare utility plans or bills, showing the period, the units, and the rate structure.

Make it ask first, every time:

  1. Any purchase or payment.
  2. Any application filing or account change.
  3. Any sharing of a family member's data with anyone.
  4. Any message sent to a teacher, agency, doctor, or other third party on your behalf.

Silence is not a yes. A shared family account is not a yes. The adult who owns the account approves the exact action, the exact amount, and the exact recipient.

Never let it do these:

  1. Decide whether a symptom, a fall, a fever, or a medication question is an emergency. It can organize the facts and help you call. The professional decides.
  2. Make a decision about a child's safety or protection.
  3. Decide discipline, or settle an argument between family members.
  4. Serve as a child's or a lonely relative's main source of company or emotional support.
  5. Assert that your family is legally eligible for something, or sign anything under penalty of perjury.
  6. Watch the household continuously and infer what you need without being asked.

And one rule over all the others: any family member with authority must be able to stop it, see what it did, and delete what it kept. A tool that can't be stopped isn't assistance. It's a new dependent.

6.9.4 Five Questions Before You Let an App Near Your Kids' Data#

The "minimize data" stage is the one families skip most, because it's invisible until something goes wrong. Before any app sees a child's name, school, health information, photos, or location, ask these five questions. If the app can't answer them in plain language, that's your answer.

  1. What exactly will it collect, and does the task need it? A permission slip checklist needs a date and a fee. It doesn't need your child's diagnosis or home address. Redact what the task doesn't require.
  2. Who else will see it? Look for the list of recipients, not just a promise. Unintended recipients are one of the privacy outcomes this chapter says a household pilot should measure.
  3. How long does it keep it, and can I set that? The retention setting is part of the record in the ten-stage process for a reason.
  4. Can I delete it, and will it tell me when deletion is done? Stopping use is not the same as deleting data. A household must be able to stop and recover.
  5. Who in this family has authority to say yes? A child can't consent for themselves in the way an adult can, and a dependent adult's data needs the same care. If the person clicking "agree" isn't the person with authority and a justified need, stop there.

6.10 How to Test Whether AI Actually Helps a Household#

6.10.1 An Illustrative Workflow: A Parent, Three Children, One School Notice#

Here's a hypothetical. It's an illustrative workflow, not a real family, an observed deployment, or a promise of time savings.

A parent is caring for three children while a spouse is at work. A school activity notice comes home. The parent gives the assistant a redacted copy and asks for a checklist. The assistant identifies the date, the pickup arrangement, the required items, the fees, and the consent requirements, and it attaches each field to the passage in the notice it came from. Where the notice doesn't answer something, the assistant says so instead of guessing.

The parent compares the draft with the notice. The pickup time conflicts between two pages, so the parent checks with the school. Then the parent decides who owns each task. A calendar entry, a payment, a message to the school, or a signed form each requires a separate approval. The system does not infer agreement just because two adults share a household.

The outcome record captures total preparation and verification time, which fields were corrected, the final deadline, whether the event was missed, and who completed the work. And here's the test for the mental load question: if the other parent performs an assigned task, that counts as redistribution only if the action and the responsibility actually moved. A shared list on a screen is not redistribution. The spouse picking up the kid, paying the fee, and remembering next time is.

The same sequence works for a grocery list or a meal plan, but the parent stays responsible for allergies, suitability, actual prices, and what ingredients are on hand. No reviewed study establishes that an unrestricted generated meal plan lowers food spending or meets nutritional needs. For a medical concern, the workflow stops short of diagnosis: it organizes questions and any existing professional instructions, but it can't be the authority on urgency or treatment. For a benefits form, it keeps the official requirements in view and drafts entries, while the applicant and the authorized agency control submission, eligibility, and appeal.

6.10.2 What a Real Household Pilot Should Measure#

Measuring the whole task means recruiting varied family structures and comparing against a credible existing method. Not an artificial "do nothing" group. Families already use calendars, search, school portals, relatives, and professionals, and the new tool has to beat that.

  • Primary outcomes: net time after verification, missed obligations, substantive error rate, successful completion, caregiver burden, dependent safety, and human control.
  • Distribution: results by household structure, income, language, disability, digital access, caregiving intensity, and prior skill.
  • Relationship outcomes: who owns the task before and after, partner visibility, conflict, and whether responsibility concentrates on one person.
  • Privacy outcomes: sensitive fields disclosed, unintended recipients, retention, permission errors, and successful deletion.
  • Economic outcomes: subscription and device cost, professional time, corrections, applicant time, and avoided or added costs.
  • Follow-up beyond the immediate task.
  • Adverse outcomes, named explicitly: medical delay, incorrect filing, missed deadline, unauthorized spending, manipulation, overreliance, isolation, and child exposure.

6.10.3 The Net Time Formula#

Proposed net time saved equals the baseline time for the same-quality task minus all the time spent in the assisted workflow: setup allocation, data preparation, prompting, verification, correction, execution, and follow-up.

That's the whole game. Faster drafting is not a net saving if supervision and correction take longer than the old way. A pilot should preregister one or two primary outcomes per task category, and it should not search across dozens of measures and publish only the favorable ones.

6.11 Evidence Gaps and Permitted Claims About AI for Families#

6.11.1 What We Still Don't Know#

The household-specific evidence I located does not establish durable reductions in total administrative time, grocery spending, relationship conflict, or unequal responsibility from general-purpose assistants.

Effects for single parents, low-connectivity households, non-English users, people with disabilities, and families in acute crisis can't be inferred from digitally comfortable survey respondents. The positive interventions in this chapter usually include an institution, predefined content, expert review, or human assistance. Research should test whether those components are necessary before anyone presents a cheaper automated substitute as equivalent. Housing search, insurance disputes, debt advice, legal forms, food safety, and emergency caregiving all need domain-specific evidence before strong efficacy claims are made.

6.11.2 What You Can and Can't Claim#

Claim Status
Structured digital and human programs have improved certain parenting, application, energy, and caregiver outcomes in defined settings Permitted
General-purpose assistance has been proven to save every family time, redistribute mental load, or improve relationships Not permitted
Families can use original-source infrastructure records to investigate possible effects on bills, water, services, jobs, and public finances Permitted
A tracker entry, announced investment, model capability, or educational benefit proves that a specific facility improves household welfare Not permitted
A household pilot can test whether source-linked assistance improves comprehension and reduces burden Permitted
Usage, confidence, a generated answer, or a product demonstration substitutes for measured household outcomes Not permitted

On the infrastructure rows: the seven Savrn trackers exist so families and communities can find the original records behind a facility. A tracker entry starts a question. It never ends one.

6.12 Conclusion: Bounded Assistance, Not Delegation#

The evidence supports bounded assistance, not household delegation as a default. The most credible benefits in this chapter came from interventions built around a particular barrier, an accountable institution, a comparison condition, and a measurable outcome. You've now seen that pattern in schooling, in workforce programs, and here again.

So the household standard is practical: reduce verified work without hiding who remains responsible. Families should be able to identify the source, understand the action, withhold approval, correct an error, protect a dependent, and stop the system. Use AI for the permission slip. Check its work. Keep the judgment, the relationships, and the off switch in your own hands.

6.13 Frequently Asked Questions#

Does AI actually save parents time?

Not in any controlled study yet. A 2026 Lurie Children's survey found 81% of 1,004 U.S. parents used AI for parenting tasks and reported saving 58 minutes a week, but that is self-report without a disclosed probability sample. A Croatian pilot with five mothers measured no objective time savings. Real savings must count the time spent checking and correcting the AI's work, which no reviewed study has measured.

Can AI reduce the mental load for moms?

The evidence so far says no redistribution has been shown. A study of 369 employed Croatian mothers and a five-mother ChatGPT pilot found participants valued help with planning and scheduling, but responsibility largely stayed with mothers. Helping one person work the list faster is not the same as moving the list to someone else. That requires a partner who actually takes on the task.

Is it safe to use ChatGPT for medical advice about my child?

Use it to organize symptoms and questions, not to decide urgency. In a randomized study of 1,298 U.K. adults using ten clinical scenarios, people without a chatbot had 1.76 times the odds of identifying a relevant condition compared with chatbot users, and every group tended to underestimate urgency. The scenarios were simulated and the models have since changed, but a professional or emergency service should make the call.

Do AI chatbots make people less lonely?

A large trial suggests not. In a July 2026 working paper covering 12,356 French adults, people encouraged to have daily personal conversations with a generative AI for 28 days reported loneliness 0.168 points higher on a 0 to 10 scale (adjusted p = .0002). The effect was small, the study lasted one month, and it had no journaling comparison, but the direction was the opposite of the usual promise.

Are AI parenting apps backed by research?

Structured programs are, open chat apps are not yet. A blended program in China combining rule-based WeChat modules with teacher and social worker groups improved caregiver-reported learning activities (beta 1.79) and reduced reported violence (IRR 0.87) at the immediate follow-up, in one preschool of 541 caregivers. The READY4K text program raised literacy by 0.109 standard deviations. Neither tested an open-ended generative chatbot.

Does technology help dementia caregivers?

On average, a little. A 2026 meta-analysis of 35 trials with 3,388 caregivers found interactive eHealth reduced burden by SMD -0.26. But the prediction interval ran from -1.10 to +0.58, so a specific new program could help a lot, do nothing, or add burden. Only three studies were low risk of bias, and the review did not isolate general-purpose generative AI.

Can AI help my family apply for benefits like SNAP?

It can help you prepare, but the strongest evidence favors human help. In a randomized trial of 31,888 Pennsylvania households with an older adult, enrollment was 5.8% with no outreach, 10.5% with information, and 17.6% with information plus staff who filled out and submitted applications. The trial involved no AI. Let a tool organize documents, and let the agency decide eligibility.

What should I never let an AI do for my family?

Never let it decide whether something is a medical emergency, make a child safety or discipline decision, act as a child's or lonely relative's main companion, assert legal eligibility, or spend money and share data without your specific approval. Those limits come from randomized trials with adverse findings on medical decisions and loneliness, and from the principle that authority stays with the responsible adult.

Keep listening

A copper studio microphone and headphones on a stack of books, with blueprint sound waves flowing into a sketched open book

The audio edition · Read by Bella

Listen to the whole report

13 episodes, 2 h 26 min. Each one walks a chapter's argument, its studies, and their limits, so you can listen instead of read. It plays straight through, one chapter into the next.

Up next · Chapter 1

AI Safety Promises and Public Trust: How to Test What AI Labs Say

0:00 / 11:57
  1. Read
  2. Read
  3. Read
  4. Read
  5. Read
  6. Read
  7. Read
  8. Read
  9. Read
  10. Read
  11. Read
  12. Read
  13. Read