# Chapter 9: AI in Healthcare and Aging: What the Evidence Says for Patients and Older Adults

Part of The Superintelligence Transition by Chad Everett Harris, Founder and CEO, Savrn. Published September 24, 2026. Evidence cutoff September 23, 2026.

Canonical: https://savrn.com/blog/the-superintelligence-transition/ai-healthcare-aging
Full edition: https://savrn.com/blog/the-superintelligence-transition

Disclosure: prepared for Savrn, which has a commercial interest in AI infrastructure and publishes the seven trackers cited here. Not an independent institutional review.

**Key takeaways**

- In the Swedish MASAI trial (105,915 women analyzed), AI-supported mammography screening had 1.55 interval cancers per 1,000 versus 1.76 with standard double reading, a ratio of 0.88 (95% CI 0.65 to 1.18): noninferior, not a proven reduction.
- MASAI's supported workflow found more cancers (sensitivity 80.5% vs 73.8%) with specificity near 98.5% in both groups, but the trial did not measure mortality.
- In a randomized trial of 50 physicians working structured diagnostic vignettes, access to a language model changed median reasoning scores by an adjusted 2 points (76% vs 74%, 95% CI -4 to 8), not statistically significant, even though the model alone performed strongly.
- A May 2026 review of 8 randomized trials with 611 older adults found a small, moderate-certainty effect on depression (g -0.25, 95% CI -0.48 to -0.02) from interventions lasting roughly 2 to 12 weeks.
- For loneliness, the same review pooled only 3 trials with 190 participants (g -0.67, 95% CI -2.57 to 1.23), an interval running from large benefit to substantial harm, so the answer is unknown.
- The FTC reported that adults aged 60 and older reported $2.4 billion in fraud losses in 2024; those are reported losses, not a full count, and the total is not attributed to AI or any particular technology.
Maybe you're the one who drives to the appointments now. You keep a folder, or a notes app, with your mother's medication list, and you're never quite sure it matches what the pharmacy has. Your father's phone buzzes all day with texts about a package he never ordered, a toll he supposedly didn't pay, a bank account that has been "locked." And somewhere in the middle of all that, someone tells you that AI in healthcare research has shown these tools can read scans better than doctors, keep lonely seniors company, and catch scammers before they strike. The question you actually have is simpler and harder: which of this is real, and what should I let near my parent?

That's the question this chapter answers. Before anything else, a clear statement: **this chapter is evidence synthesis, not medical advice.** I'm a builder who reads studies, not a clinician. Nothing here tells you what to do about a specific symptom, drug, dose, or diagnosis. For that, you talk to your parent's doctor, pharmacist, or care team, and if something is urgent you call for emergency help. What I can do is show you what four pieces of research actually found, where each one stops, and how to think about the tools landing on your family's kitchen table.

Here's why it matters. The second half of life is where a lot of the stakes concentrate: more appointments, more prescriptions, more paperwork, more isolation for some people, and more people trying to take their money. Products are arriving to help with every one of those. Some of the underlying research is excellent. Some of it is thin. And the marketing rarely tells you which is which.

The evidence has a good part and an uncomfortable part. The good part: one of the largest randomized evaluations in this entire report, a Swedish mammography screening trial with more than one hundred thousand women, found that a specific AI-supported workflow did as well as the standard approach on its main safety outcome and found more cancers without more false alarms. The uncomfortable part: a physician trial showed that a model that performed well by itself did not measurably improve the doctors who used it, and the best controlled review on loneliness in older adults cannot tell you whether these tools help, do nothing, or hurt. Capability is not benefit. This chapter is about the distance between the two.

### 9.1 What Counts as Value in Later Life

Before we look at a single study, I want to set the scoreboard, because the wrong scoreboard makes every study look better or worse than it is.

The value of assistance in the second half of life is broader than any clinical score. It includes independence, access to services, social connection, protection from exploitation, and support for caregivers. A diagnostic number or a satisfying conversation is not an adequate substitute for measuring those outcomes. If I only told you how accurate a model was, I'd be answering a question no family actually lives with.

Think about what your parent would say if you asked what a good year looks like. Probably something like: I got to stay in my own home. I could still drive to church, or the lake, or my sister's. I didn't miss my appointments. Nobody cleaned out my savings. I saw the grandkids. I made my own decisions. Almost none of that shows up in a benchmark. Some of it shows up in clinical trials. Most of it has to be measured deliberately, or it doesn't get measured at all.

#### 9.1.1 What the research in this chapter covers

The reviewed research spans a large screening trial, a small physician-reasoning trial, and two systematic reviews of interventions for older adults. These studies concern different technologies and different populations. They cannot be stacked into one claim that conversational systems improve health or solve loneliness. So I'm presenting them as what they are: one of the largest randomized evaluations in this report ([the MASAI trial](https://pubmed.ncbi.nlm.nih.gov/41620232/)) sitting next to one of its most instructive null results ([the JAMA Network Open physician trial](https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2825395)), with the older-adult evidence ([the BMC Geriatrics randomized-trial review](https://pmc.ncbi.nlm.nih.gov/articles/PMC13321516/)) described plainly as thin.

Alongside those, I'll bring in two items that Chapter 6 covers in more depth: a caregiver meta-analysis and a medical self-assessment trial with ordinary adults. And I'll use a Federal Trade Commission report on fraud against older adults and a W3C accessibility guidance note, because a health chapter that ignores money and usability is not describing the life your parent is living.

#### 9.1.2 The household boundary

Here's the one rule I'll keep coming back to. I call it the household boundary: **assistance prepares, organizes, and explains; professionals diagnose, treat, and respond to urgency.** It's a recommended safeguard. It is not a claim that every health-navigation product on the market has been evaluated against it, and it is not a verdict on any particular product. It's the line I'd draw for my own family based on what the evidence can and can't support today.

What does that look like in practice? A tool that helps your mother write down her three questions before a cardiology appointment is on the right side of the line. A tool that helps her understand the discharge instructions her nurse already gave her, so she can ask better follow-up questions, is on the right side. A tool that tells her whether her chest tightness is "probably nothing" is on the wrong side. The difference is not how smart the tool is. The difference is who holds responsibility for the decision.

### 9.2 AI in Healthcare Research: How the MASAI Mammography Trial Worked

If someone tells you AI has been "proven" in medicine, there's a decent chance the study they're half-remembering is this one or something like it. So let's read it properly.

#### 9.2.1 What the trial actually did

The strongest clinical evidence in the reviewed corpus is the Swedish MASAI trial, published in The Lancet. It randomized 105,934 women either to AI-supported mammography screening or to standard double reading, the standard practice in which two readers review each exam. After exclusions, 105,915 women were included in the reported analysis. ([Lancet trial record](https://pubmed.ncbi.nlm.nih.gov/41620232/))

Now the most important sentence in this section. The intervention was a specific screening system embedded in a radiologist workflow. It was not a general conversational model acting as an autonomous physician. Radiologists were still in the loop. The thresholds, the reading process, and the follow-up pathway all belonged to an organized screening program. That scope condition is the first thing you should carry out of this chapter, because nearly every misuse of this trial begins by dropping it.

#### 9.2.2 What the trial measured

The primary outcome was interval cancer. Interval cancers are cancers diagnosed between screening rounds: after a woman had a screening exam that did not flag a cancer, and before her next scheduled screen. They're the standard signal for whether a screening program is catching what it should. If a new approach to reading mammograms misses more cancers, you'd expect interval cancers to go up. So the primary question was essentially a safety question: does the supported workflow let more cancers slip through than the standard one?

The reported interval-cancer rates were 1.55 per 1,000 participants in the supported workflow and 1.76 per 1,000 in standard double reading. That's a ratio of 0.88, with a 95 percent confidence interval of 0.65 to 1.18. The result met the trial's noninferiority criterion. It did not demonstrate a statistically significant reduction in interval cancer. And the upper bound of that interval, 1.18, means the data are consistent with the supported workflow being somewhat worse, as well as somewhat better. ([Primary outcome, Lancet](https://pubmed.ncbi.nlm.nih.gov/41620232/))

Read that again. The point estimate leans in a good direction. The interval says the true answer could sit on either side of "no difference." The trial's claim is that the new workflow was not unacceptably worse, and that is a real, valuable claim. It is not a claim that the new workflow prevented cancers.

#### 9.2.3 The secondary results

The secondary performance figures are more positive and equally scoped. Sensitivity was 80.5 percent with the supported workflow versus 73.8 percent with standard reading. Specificity was approximately 98.5 percent in both groups. In plain terms, the supported workflow caught a larger share of the cancers that were there, and the improvement in detection did not come with a compensating rise in false alarms. ([MASAI trial](https://pubmed.ncbi.nlm.nih.gov/41620232/))

Those are screening-performance results inside a defined radiologist workflow. They are not evidence of reduced mortality, because the trial was not designed to measure mortality. And they are not evidence about consumer self-diagnosis, where there is no radiologist, no calibrated threshold, and no defined workflow at all.

**Evidence card**

- Title: MASAI mammography screening trial (The Lancet)
- Design: Randomized trial comparing an AI-supported screening workflow with standard double reading by radiologists in Sweden
- Population: 105,934 women randomized; 105,915 included in the reported analysis after exclusions
- Finding: Interval cancers were 1.55 per 1,000 (supported) versus 1.76 per 1,000 (standard), ratio 0.88, 95% CI 0.65 to 1.18, meeting the noninferiority criterion; sensitivity was 80.5% versus 73.8%, with specificity about 98.5% in both groups
- Limit: Not a statistically significant reduction in interval cancer; the interval includes a somewhat worse result; mortality was not measured; the result applies to a specific system inside a radiologist workflow, not to consumer tools or general chatbots
- Source: [The Lancet via PubMed](https://pubmed.ncbi.nlm.nih.gov/41620232/)
*Figure 9.1 (interactive on the page).*
#### 9.2.4 Two sentences, one supported and one not

This trial is the report's best example of a benefit claim stated precisely without exaggeration. Here are two sentences you might see written about it.

The first: "The AI-supported workflow achieved a noninferior interval-cancer outcome in this screening program, with higher sensitivity and similar specificity." The data support that.

The second: "The technology prevented 12 percent of cancers." The data do not support that. The 0.88 ratio is where that 12 percent comes from, but the direction of the point estimate does not survive its confidence interval, and mortality was never measured. The distance between those two sentences is the distance between evidence and marketing.

#### 9.2.5 Common misreadings of MASAI

**Misreading 1: "AI reads mammograms better than doctors."** The trial compared two workflows, both involving radiologists. It did not test a machine against a human in isolation, and it did not remove the radiologist.

**Misreading 2: "So my mother can upload her scan to an app."** Nothing in this trial concerns consumer apps. The system operated inside a screening program with trained readers and defined follow-up. Take the workflow away and you've taken away the thing that was tested.

**Misreading 3: "Fewer women will die of breast cancer."** Maybe, someday, if detection gains translate into outcomes. This trial did not measure that, so no one can claim it from this trial.

**Misreading 4: "Noninferior means it didn't work."** No. Noninferior means it met a predefined standard of not being unacceptably worse on the primary safety outcome. Combined with the sensitivity result, that's a meaningful finding for a screening program deciding how to organize its reading work.

#### 9.2.6 What this means for you

If you're an adult child, the practical takeaway is modest and useful. If your parent's screening program uses an AI-supported reading workflow, the best large trial in this area gives no reason for alarm about its primary safety outcome, within the limits above. If you have questions about how your parent's screening is read, ask the screening provider. That's a question for them, not for a chatbot.

If you're a health system board member or administrator, the MASAI result supports evaluating a defined system inside a defined workflow, with your own monitoring of interval cancers, sensitivity, and specificity. It does not support deploying a different system and assuming the same result.

### 9.3 How to Read a Screening Study: Sensitivity, Specificity, and Noninferiority in Plain English

A lot of families tune out when they hit these words, and I understand why. But these three terms are how screening research talks, and once you have them, you can read a headline and know within a minute whether it's telling you the truth. I'll explain each using only MASAI's own numbers. No new data, just the concepts.

#### 9.3.1 Sensitivity: of the cancers that were there, how many were found?

Sensitivity answers one question: among people who actually have the condition, what share does the test correctly flag? In MASAI, sensitivity was 80.5 percent in the supported workflow and 73.8 percent in standard reading.

Here's a way to picture it. Imagine a room holding every woman in the trial who truly had breast cancer at the time of screening. Sensitivity is the share of that room the screening process correctly identified. A higher number means fewer cancers were missed at that screen. The supported workflow's higher sensitivity means it correctly flagged a larger share of that room.

What sensitivity does not tell you: how many healthy women were flagged by mistake. For that, you need the next term.

#### 9.3.2 Specificity: of the people who were fine, how many were correctly left alone?

Specificity asks the mirror question: among people who do not have the condition, what share does the test correctly clear? In MASAI, specificity was approximately 98.5 percent in both groups.

Picture a second, much larger room holding every woman who did not have cancer. Specificity is the share of that room the screening process correctly left alone. The rest got a false alarm: a callback, more imaging, maybe a biopsy, and a lot of worry.

Here's the part most people skip. You can raise sensitivity by flagging more people. Flag everyone and you'll catch every cancer, with sensitivity at its maximum and specificity collapsing, because you've also called back every healthy woman. That's why sensitivity alone is never enough. The MASAI result matters because sensitivity went up while specificity stayed at about the same level in both arms. The detection gain wasn't bought with more false alarms.

#### 9.3.3 Rates per 1,000 and ratios

MASAI's primary result was reported as interval cancers per 1,000 participants: 1.55 in the supported workflow and 1.76 in standard reading. "Per 1,000" is a denominator, and a number without a denominator is a rumor. Interval cancers are rare events in a screening population, which is exactly why the trial needed more than one hundred thousand women to say anything about them.

The ratio of 0.88 compares the two rates. A ratio below 1.0 means the supported arm's rate was lower. A ratio of 1.0 would mean no difference. A ratio above 1.0 would mean the supported arm's rate was higher.

#### 9.3.4 Confidence intervals: the range the data can't rule out

The 95 percent confidence interval of 0.65 to 1.18 is the range of true ratios that are reasonably consistent with the data. Think of it as the study telling you, "Given what we saw, the true answer is probably somewhere in here."

Notice that 1.0 sits inside that range. The low end (0.65) would be a meaningful improvement. The high end (1.18) would be somewhat worse. When an interval contains 1.0 for a ratio, the study has not shown a statistically significant difference in either direction. That's why the correct statement is "no demonstrated reduction," not "a 12 percent reduction."

#### 9.3.5 Noninferiority: "not unacceptably worse"

This is the concept people most often get backward. Most studies you hear about ask, "Is the new thing better?" A noninferiority trial asks a different question: "Is the new thing not unacceptably worse than the standard?" That's the right question when the new approach might offer other advantages and the main worry is that it could quietly let more harm through.

Before the trial starts, the researchers set a boundary for how much worse would count as unacceptable. If the confidence interval stays on the acceptable side of that boundary, the new approach is declared noninferior. MASAI met its noninferiority criterion on interval cancer. That tells you the supported workflow passed the safety test the trial set for it. It does not tell you the supported workflow was superior on that outcome, because the trial did not show that.

Noninferior is not the same as equivalent, and it's not the same as better. It's a pass on a specific safety bar.

#### 9.3.6 A hypothetical: reading a headline in two minutes

Here's a hypothetical to practice on. Suppose you see a post that says, "Huge study proves AI catches breast cancer that doctors miss and cuts cancer rates." You now have four questions to ask, and MASAI answers them:

1. **What was compared?** Two radiologist workflows, one supported by a specific system. Not AI versus doctors.
2. **What was the primary outcome, and did it show superiority or noninferiority?** Interval cancer, noninferior, no significant reduction.
3. **Did sensitivity rise without specificity falling?** Yes: 80.5 versus 73.8 percent sensitivity, specificity near 98.5 percent in both.
4. **Was mortality measured?** No.

So the headline gets one piece right (more cancers were detected in this workflow) and two pieces wrong (it wasn't AI versus doctors, and it didn't show fewer cancers or deaths). That took two minutes, and you didn't need a medical degree to do it.

#### 9.3.7 Quick reference table

| Term | Plain question it answers | MASAI value |
|---|---|---|
| Sensitivity | Of the cancers present, what share were flagged? | 80.5% supported vs 73.8% standard |
| Specificity | Of the women without cancer, what share were correctly cleared? | About 98.5% in both groups |
| Interval cancer rate | How many cancers showed up between screens, per 1,000 screened? | 1.55 vs 1.76 per 1,000 |
| Ratio with 95% CI | How do the two rates compare, and how sure are we? | 0.88, CI 0.65 to 1.18 |
| Noninferiority | Was the new workflow not unacceptably worse? | Criterion met |
| Mortality | Did fewer women die? | Not measured |

### 9.4 AI Diagnosis Accuracy vs Doctors: What the Physician Trial Found

The second clinical study is small, and its lesson is out of proportion to its size. If you only remember one study from this chapter besides MASAI, make it this one.

#### 9.4.1 What the trial did

A randomized trial published in JAMA Network Open involved 50 physicians working through structured diagnostic vignettes: written clinical cases designed to test diagnostic reasoning. Physicians were randomized either to conventional resources or to access to a language model. The outcome was a diagnostic-reasoning score graded on those cases. ([JAMA Network Open trial](https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2825395))

#### 9.4.2 What it found

Median diagnostic-reasoning scores were 76 percent and 74 percent, with the model-access group ahead by an adjusted difference of two percentage points. The 95 percent confidence interval for that difference ran from -4 to 8. The difference was not statistically significant. ([Trial results](https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2825395))

Walk through that interval the same way we did for MASAI. A difference of zero sits inside it. The data are consistent with the tool making physicians a bit worse, making no difference, or making them moderately better. The study can't tell those apart. Under the tested conditions, access to the model did not produce a demonstrated improvement.

#### 9.4.3 The finding that should stay with you

The vignette structure matters for interpretation. These were structured diagnostic cases, not observed patient outcomes. Fifty physicians grading hypothetical patients is a laboratory for reasoning, not a field trial of care.

But the finding that matters most is in an exploratory analysis. The model alone, without the physician, performed strongly on the same vignettes. And that strong standalone performance did not translate into a demonstrated improvement for the physicians who used it under the tested conditions. The tool was better than its effect. ([Exploratory analysis, JAMA Network Open](https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2825395))

"Exploratory" is a flag you should respect. It means the comparison wasn't the trial's primary question, and it should be treated as a signal to investigate, not a settled result. Even so, the direction of the gap is the lesson.

**Evidence card**

- Title: Physician diagnostic reasoning trial (JAMA Network Open)
- Design: Randomized trial comparing conventional resources with access to a language model on structured diagnostic vignettes
- Population: 50 physicians
- Finding: Median diagnostic-reasoning scores of 76% and 74%, with the model-access group ahead by an adjusted 2 percentage points (95% CI -4 to 8), not statistically significant; in an exploratory analysis, the model alone performed strongly
- Limit: Small sample; hypothetical vignettes rather than real patients or outcomes; the standalone comparison was exploratory; tested conditions may not match real practice
- Source: [JAMA Network Open](https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2825395)
*Figure 9.2 (interactive on the page).*
#### 9.4.4 Capability versus workflow effect

That gap, standalone capability versus workflow effect, is the single most repeated pattern in this report. It showed up in consulting tasks in [Chapter 2](#ai-productivity-evidence), in education in [Chapter 3](#ai-in-k12-education), and in the household studies in [Chapter 6](#ai-for-families-households). Here it appears in its clearest clinical form, and it explains why I don't infer benefit from capability demonstrations. The performance of the model and the performance of the professional workflow require separate evaluation. A vendor who reports only the first has told you half of what you need.

Chapter 6 covers a companion result on the consumer side. In a randomized study published in Nature Medicine, Bean and colleagues recruited 1,298 U.K. adults to work through ten physician-authored clinical vignettes using GPT-4o, Llama 3, Command R+, or their usual resources. Each model performed strongly when directly prompted, and the people using those models did not reproduce that performance: control participants had 1.76 times the odds of identifying a relevant condition compared with pooled model users, while disposition accuracy did not differ significantly between each model group and control. All groups tended to underestimate acuity. That study used simulated scenarios, not people experiencing symptoms, and the models have changed since its 2024 data collection. ([Nature Medicine study](https://www.nature.com/articles/s41591-025-04074-y))

Put the two together and the pattern is hard to miss. A strong model in the hands of physicians did not measurably help on vignettes. Strong models in the hands of ordinary adults were associated with worse condition identification than usual resources on vignettes. Neither result says these tools can never help. Both say help cannot be assumed from a model's score.

#### 9.4.5 What this result does and does not establish

This result does not establish that clinical assistance cannot help. It establishes that help cannot be assumed, and that the evaluation has to happen at the level where the help is supposed to arrive: the clinician's actual decision, with the actual time pressure, liability, and patient context of real practice.

**Common misreading: "The AI was smarter than the doctors, so the doctors were the problem."** That reads a small exploratory comparison as a verdict on physicians. What the trial shows is that putting a capable tool next to a professional did not, under these conditions, produce a measurable gain. Why that happened is a question for further research, not a conclusion.

**Common misreading: "The AI didn't help, so it's useless."** A null result in 50 physicians on vignettes is not proof of no effect anywhere. The interval runs up to 8 points. The accurate reading is "not demonstrated," not "disproven."

#### 9.4.6 What this means for your family

For households, the boundary follows directly. Use assistance to prepare questions, organize records, and understand clinician-provided instructions. Seek professional care for diagnosis, treatment, and urgent symptoms. That boundary is a safeguard, not a verdict on any particular product.

Here's a hypothetical to make it concrete. Your father comes home from a visit with a printed after-visit summary and a new prescription, and he's confused about when to take it relative to his other pills. A reasonable use of an assistant: help him turn his confusion into a short list of specific questions, then call the pharmacist or the clinic with that list. An unreasonable use: asking the assistant to decide the timing and acting on its answer. The first keeps the professional in charge of the decision. The second quietly moves the decision to a tool nobody evaluated for that job.

If you're a clinician or a practice manager, the physician trial is a reason to evaluate any assistant at the level of your team's real decisions, not at the level of the vendor's benchmark. Ask for evidence that the workflow improved, not just that the model scored well.

### 9.5 AI for Seniors: Loneliness, Depression, and Social Isolation Are Not the Same Thing

This is the section where I expect the most disagreement, because companionship products for older adults are being sold with a lot of warmth and not much evidence. I'm going to be careful here, because the people on the other end of these products are someone's parents.

#### 9.5.1 The May 2026 BMC Geriatrics review

A May 2026 review in BMC Geriatrics gathered eight randomized trials that included 611 older adults across heterogeneous interventions: social robots, voice systems, and cognitive activities. It is the most current controlled synthesis of its kind in the reviewed corpus. ([BMC Geriatrics review](https://pmc.ncbi.nlm.nih.gov/articles/PMC13321516/))

"Heterogeneous" matters. These weren't eight tests of the same product. A social robot and a voice system and a cognitive activity program are different things. Pooling them tells you something about the category, and less about any one product.

#### 9.5.2 Depression: small, real, limited

For depression, the pooled estimate across seven trials was small, with Hedges' g of -0.25 and a 95 percent confidence interval of -0.48 to -0.02. The review rated that evidence moderate certainty.

Some plain English. Hedges' g is a standardized effect size: it expresses the difference between groups in units of how spread out the scores were, so that studies using different depression scales can be combined. Here, a negative value means lower depression scores in the intervention groups. A value of -0.25 is small by the usual conventions. The interval runs from -0.48 to -0.02, which means it just barely excludes zero.

So: a small effect, with an interval that barely excludes zero, rated moderate certainty. That's a finding of real but limited promise. Not nothing, and not enough.

#### 9.5.3 Loneliness: we do not know

For loneliness, the same review pooled three trials with 190 participants and reported g of -0.67 with a very wide interval of -2.57 to 1.23 and high heterogeneity. ([BMC Geriatrics review](https://pmc.ncbi.nlm.nih.gov/articles/PMC13321516/))

Look at the point estimate alone and you'd think this was a big win, bigger than the depression effect. That's the trap. The interval runs from a very large benefit to a substantial harm. That's the statistical signature of "we do not know." Three trials, 190 people, high heterogeneity: the result does not establish a reliable average reduction in loneliness, and no responsible summary of this review can present it as one.

If you see a product page citing this review as evidence that its companion device reduces loneliness, you now know that's not what the review found.

#### 9.5.4 Duration: weeks, not years

The interventions lasted roughly two to twelve weeks and were not a uniform test of current general-purpose conversational models. They justify neither a broad claim about durable companionship nor any claim that a system can substitute for human contact. Two to twelve weeks is a demonstration window. Loneliness and social isolation are measured in years.

And those three words, depression, loneliness, social isolation, are not interchangeable. Depression is a clinical condition. Loneliness is the felt experience of lacking connection. Social isolation is an objective lack of contact. A person can be isolated without feeling lonely, or lonely in a full house. A product that improves one does not automatically improve the others, and a study that measured one tells you little about the rest.

**Evidence card**

- Title: Randomized-trial review of interventions for older adults (BMC Geriatrics, May 2026)
- Design: Systematic review and meta-analysis of randomized trials of social robots, voice systems, and cognitive activities
- Population: 8 randomized trials, 611 older adults
- Finding: Depression (7 trials): Hedges' g -0.25, 95% CI -0.48 to -0.02, moderate certainty. Loneliness (3 trials, 190 participants): g -0.67, 95% CI -2.57 to 1.23, high heterogeneity
- Limit: Heterogeneous interventions lasting roughly 2 to 12 weeks; not a uniform test of current general-purpose conversational models; loneliness result spans large benefit to substantial harm and establishes no reliable average effect
- Source: [BMC Geriatrics via PMC](https://pmc.ncbi.nlm.nih.gov/articles/PMC13321516/)
#### 9.5.5 The January 2026 Psychological Medicine review

A separate January 2026 review in Psychological Medicine reported favorable before-and-after changes in similar outcomes. But its pooled estimates were within-group comparisons, because controlled evidence was limited. ([Psychological Medicine review](https://pmc.ncbi.nlm.nih.gov/articles/PMC12885342/))

Within-group means comparing people to themselves: their scores before the program versus after. The problem is that a lot of things change between "before" and "after" besides the intervention. Time passes. People who enroll during a hard stretch often improve anyway, which statisticians call regression to the mean. Someone showed up repeatedly over the study period and paid attention to them. Within-group change can't separate the intervention from any of that. Such estimates are observations, not effects.

#### 9.5.6 The two-review discipline

The two reviews should not be averaged together as independent confirmations. Their comparison methods differ (one pooled controlled contrasts, the other pooled within-group change), and their underlying studies may overlap, so an apparently stronger result in one review does not resolve the uncertainty in the other. Averaging them would manufacture a precision neither review claims.

Here's a hypothetical of how this goes wrong. A vendor slide reads: "Two 2026 meta-analyses confirm companion technology reduces loneliness in seniors." One of those reviews found a loneliness interval spanning benefit to harm. The other measured change within the same people, without a control group, and may share studies with the first. Two weak signals, one of which may be partly the same signal counted twice, do not add up to a strong one.

The plain state of the evidence:

- **Depression:** a small, moderate-certainty signal.
- **Loneliness:** unresolved uncertainty.
- **Companionship replacement:** no basis for claims of any kind.

*Figure 9.3 (interactive on the page).*
#### 9.5.7 The caregiver evidence belongs in the same frame

If you're the adult child in this story, you're also a caregiver, and there's evidence about you too. Chapter 6 reviewed digital caregiver interventions. A registered 2026 systematic review in JMIR included 35 randomized trials with 3,388 informal dementia caregivers. Interactive eHealth interventions produced a pooled burden effect of SMD -0.26 (95% CI -0.42 to -0.10), but heterogeneity was 73.6 percent and the 95 percent prediction interval ran from -1.10 to +0.58. ([JMIR caregiver intervention review](https://www.jmir.org/2026/1/e78568))

The confidence interval describes the average, which favors the interventions. The prediction interval describes what the next specific program might do, which could be substantial benefit, little effect, or an adverse result. Only three studies were rated low risk of bias, the evidence was judged moderate certainty, and it relied heavily on subjective outcomes. The review covered web, mobile, video, and hybrid programs and does not isolate a general-purpose generative model effect.

So that evidence concerns caregiver support, with small average burden effects and substantial uncertainty about transfer to new settings and to general-purpose models. It is not a comprehensive solution for older adults' health or independence. For more, see [Chapter 6](#ai-for-families-households).

#### 9.5.8 What this means for you

**For adult children:** If you're considering a companion device or voice assistant for a parent, go in with a small, specific goal (for example, a daily check-in routine your parent wants), and watch whether human contact increases, stays the same, or quietly shrinks. The evidence doesn't tell you it will help with loneliness. Your own observation, over months, is the only evaluation you'll get.

**For older adults:** You get to decide whether you want one of these. If someone buys you a device "so you won't be lonely," it's fair to say that's not what you need, or to try it on your own terms and put it away if it doesn't earn its place.

**For senior living operators and program funders:** A pilot of two to twelve weeks is a demonstration. If you're deciding whether to fund a program at scale, ask for controlled comparisons, the outcome that matters to residents, and follow-up that outlasts the novelty.

### 9.6 A Later-Life Workflow That Starts With the Person's Own Goals

Now for the practical part. If the evidence can't tell you which product to buy, what can you actually do? My answer is a workflow, not a product. And the first step in that workflow isn't a feature. It's a conversation.

#### 9.6.1 Step zero: the person's goals and consent

Every proposed use in this section requires local testing. I'm claiming no demonstrated effect for any item merely because it sounds plausible. The workflow's first step is the person's own goals and consent, and that means separating two lists: the tasks the person wants help with, and the tasks other people would prefer to automate for their own convenience.

Those two lists diverge more often than any product demo suggests, and the divergence is where autonomy gets lost. Your mother may want help keeping her appointments straight. She may not want you getting a notification every time she opens the refrigerator. Both might be "safety features." Only one is hers.

Here's a hypothetical. A family sets up a shared calendar, medication reminders, and a location-sharing app for a father who has recently been widowed. The adult children feel reassured. The father feels watched, stops carrying his phone on walks, and misses calls. Nothing in that setup was malicious, and nothing in it started with his goals. The fix isn't a better app. It's going back and asking him which of those he actually wants.

#### 9.6.2 The six-part workflow

- **Appointment preparation:** Organize the person's questions and existing records without diagnosing from incomplete information.
- **Medication communication:** Help prepare a verified list for a clinician or pharmacist; do not autonomously change doses or treatment.
- **Service navigation:** Explain documented eligibility and application steps, while preserving a human route for exceptions.
- **Care coordination:** Share only information the person has authorized, with clear roles and a current contact plan.
- **Daily independence:** Test whether reminders or interfaces actually help the person complete a task rather than simply increasing caregiver monitoring.
- **Emergency continuity:** Keep an accessible plan that works when the digital service, connectivity, or power is unavailable.

*Figure 9.4 (interactive on the page).*
Each item has a verb that keeps a person in charge: organize, prepare, explain, share, test, keep. None of them says decide, prescribe, or diagnose. That's deliberate. It's the household boundary from section 9.1 turned into a daily routine.

#### 9.6.3 Daily independence inverts the usual measurement

Two items carry more weight than their bullet size suggests. The first is daily independence. It flips the usual measurement on its head. A reminder system that lets a caregiver watch more closely while the older adult completes fewer tasks independently has moved in the wrong direction, even if every dashboard trend line points up.

That's worth sitting with. Most eldercare technology gets judged by the people buying it, and the people buying it are often the adult children. So the metrics that get reported tend to be the ones that reassure the buyer: alerts sent, check-ins logged, dashboards viewed. The question that matters to the person living with the technology is different: can I do more of my own life than I could before? If the answer is no, the product is serving the family's anxiety, not the person's independence.

#### 9.6.4 Emergency continuity is the test most systems fail silently

The second is emergency continuity. It's the test most systems fail without anyone noticing, because nobody runs the test until the day it matters. A plan that exists only inside an app is not a plan during the outage when it's most needed.

> **Operator's note: If it only works when the power is on, it isn't a plan**
>
> I've spent my career building power systems. When you build large power loads, you learn to design for the day things fail, because that day always comes. A facility that runs perfectly until the grid hiccups was never a reliable facility. It was a lucky one. We design redundancy, backup paths, and manual procedures because the whole point of infrastructure is to keep working when conditions are bad.
> 
> A family's health plan deserves the same discipline. If your mother's medication list lives only in an app, it's gone when her phone dies, when the cell network is down, or when the power is out after a storm. If the emergency contacts live only in a cloud account nobody else can open, they're gone too. The fix is not high tech: a current printed medication list, printed emergency contacts, a copy with a trusted person, and a plan for how she reaches help if the phone doesn't work. Test it the way we test backup systems. Turn the phone off and ask: what now?
> 
> I'll hold my own industry to this. Savrn is designed around behind-the-meter power, and I believe in that design. But a design goal is a publisher statement, not a result, and I expect the same scrutiny I'm asking you to apply here: show me how it performs when conditions go wrong, measured, not promised. A system that only works when the power and network are up is not a plan. That goes for data centers, and it goes for your parent's medication list.
#### 9.6.5 A hypothetical walk through the workflow

Here's how it might look for one family, as an illustration only.

An older woman who lives alone has a cardiology appointment next week. She's comfortable with her tablet and wants help getting organized. Step zero: she says she wants help with preparing for appointments and keeping her medication list straight, and she does not want her children reading her messages.

For appointment preparation, she uses an assistant to turn her scattered notes into a short list of questions and to pull together the dates of her recent tests from the paperwork she already has. She doesn't ask it what her symptoms mean. For medication communication, she and her daughter build a list from the actual pill bottles, and she brings it to the pharmacist to verify. For care coordination, she decides her daughter can see the appointment calendar, and nothing else. For daily independence, she tries medication reminders for a month and checks whether she's taking her pills more reliably on her own, not whether her daughter is getting alerts. For emergency continuity, a printed copy of the verified list and her contacts sits on the refrigerator, and her daughter has one too.

Nothing in that example proves any tool works. That's the point. The workflow is how a family tests, locally, whether assistance is helping the person it's supposed to help.

### 9.7 A Family Checklist for Using AI Around a Parent's Health

This is the list I'd want in my pocket. Take it to a family meeting, a sibling group text, or a conversation with your parent. It doesn't require any technical skill, and it doesn't replace anything a clinician tells you.

#### 9.7.1 Before you start

- **Did your parent ask for this?** Write down what they want help with, in their words. If the list is mostly your worries, stop and talk first.
- **Who decides?** Agree that your parent (or their legally authorized representative, where one exists) makes decisions about their care, and that the assistant never does.
- **What's the boundary?** Assistance prepares, organizes, and explains. Professionals diagnose, treat, and respond to urgency. Say it out loud so everyone hears the same rule.
- **What information goes in?** Share only what's needed for the task, and only what your parent has agreed to share.

#### 9.7.2 For appointments and records

- Use assistance to draft questions, not answers.
- Bring the question list to the appointment and let the clinician answer.
- Use it to understand instructions a clinician already gave, then confirm anything unclear with the clinic.
- Never treat a generated explanation of a symptom as a reason to skip or delay care.

#### 9.7.3 For medications

- Build the medication list from the actual bottles and packages, then have a pharmacist or clinician verify it.
- Never let a tool change a dose, stop a medicine, or add one. Those are clinical decisions.
- Keep a printed copy of the verified list where it can be found in an emergency.

#### 9.7.4 For urgent symptoms

- If something feels urgent, call the clinician, a nurse line, or emergency services. Don't ask a chatbot first.
- Remember the evidence from the Nature Medicine vignette study: people using models did not match the models' standalone performance, and all groups tended to underestimate acuity. A tool's confidence is not a safety signal.

#### 9.7.5 For companionship and daily life

- If a companion device or voice assistant is introduced, name the goal (a routine, a reminder, a way to call family) and check it after a few months.
- Watch whether human contact goes up or down. If it goes down, that's a warning sign, whatever the satisfaction survey says.
- Your parent can say no, and can change their mind.

#### 9.7.6 For continuity and review

- Keep paper backups for medications, contacts, and key instructions.
- Test the plan: phone off, internet down, power out. Does your parent still know what to do?
- Set a date to review what's working, with your parent in the room.
- Ask one question every time: is my parent doing more of their own life, or are we just watching more closely?

### 9.8 Elder Fraud and AI: Fraud Protection Without Treating Older Adults as Helpless

Back to that phone full of scam texts. This section is about money, and it's also about dignity, because the fastest way to fail an older adult on fraud is to treat them as if they can't think.

#### 9.8.1 What the FTC reported

The Federal Trade Commission reported that adults aged 60 and older reported $2.4 billion in fraud losses in 2024. Three qualifiers attach to that number, and they belong in the same breath. ([FTC older-adult report announcement](https://www.ftc.gov/news-events/news/press-releases/2025/12/ftc-issues-annual-report-congress-agencys-actions-protect-older-adults))

1. **These are reported losses, not a complete population estimate.** Many fraud losses are never reported. The true total is unknown.
2. **The report does not attribute the total to generative systems** or any particular technology. If you see "$2.4 billion lost to AI scams," that's not what the FTC said.
3. **The same reporting distinguishes the likelihood of reporting a loss from the amount lost** in particular scam categories. How often a group reports losing money and how much they lose when they do are different measurements.

**Evidence card**

- Title: FTC annual report to Congress on protecting older adults
- Design: Federal agency report based on consumer fraud reports
- Population: Adults aged 60 and older who reported fraud
- Finding: $2.4 billion in reported fraud losses in 2024
- Limit: Reported losses only, not a complete population estimate, because many losses go unreported; the total is not attributed to generative systems or any particular technology; reporting likelihood and amount lost are distinct measures
- Source: [Federal Trade Commission](https://www.ftc.gov/news-events/news/press-releases/2025/12/ftc-issues-annual-report-congress-agencys-actions-protect-older-adults)
#### 9.8.2 The equity point: aggregate losses are not a verdict on capability

That third qualifier carries an equity point I insist on. Aggregate loss data do not support treating every older adult as less capable of recognizing deception. Many older adults report fraud at higher rates than younger adults precisely because they check their statements and know how to complain. The stereotype of universal vulnerability is not in the FTC's numbers, and it shouldn't be smuggled into products or policies built on them.

**Common misreading: "Older people fall for scams more."** The data described here don't say that. They say a certain amount of loss was reported by people in a certain age group. Reporting behavior, the kinds of scams involved, and the amounts at stake all shape that number. Turning it into a statement about the mental capacity of an entire generation is a stereotype, not a finding.

Why does this matter practically? Because a protection design built on the stereotype takes control away. It locks accounts, routes every transaction through a child, and treats every decision as suspect. A protection design built on the evidence keeps the person in control and gives them better tools and a pause button.

#### 9.8.3 The proposed protective workflow

The workflow I'm proposing has four parts:

- **Independently verified contact information.** When a message says it's from the bank, the pharmacy, a government office, or a grandchild, the person contacts that party through a number or address they already trust, not the one in the message.
- **A pause before unusual transfers.** Urgency is the scammer's main tool. A built-in waiting period, set by the person, takes it away.
- **Trusted-person involvement when the person has authorized it.** A named person the older adult chose, involved in the ways the older adult agreed to. Not automatic oversight.
- **A clear escalation route.** Everyone knows who to call and what to do if something looks wrong, including how to report it.

Here's the limit, stated plainly. This review contains no trial establishing that a particular automated scam detector prevents real-world losses. That study has not been done in the reviewed corpus, and any product claiming otherwise is ahead of its evidence.

#### 9.8.4 Automated warnings cut both ways

Automated warnings need evaluation, and the evaluation is double-sided. This is the same sensitivity and specificity logic from section 9.3, applied to your father's phone.

A system that produces frequent false alarms interferes with legitimate activity and trains the person to dismiss warnings. That habit collapses exactly when a real alarm arrives. That's the specificity problem. A reassuring false negative is worse: the tool says the message looks fine, the person relaxes, and the scam goes through with a stamp of approval. That's the sensitivity problem.

Fraud protection, like everything else in this chapter, is a measured-outcome question, not a feature-list question. "Our tool flags scams" is a feature. "Our tool reduced losses among users compared with a similar group, without so many false alarms that people stopped listening" would be an outcome. I haven't seen that outcome in the reviewed evidence.

*Figure 9.5 (interactive on the page).*
#### 9.8.5 A hypothetical: the grandchild call

As an illustration: an older man gets a frantic call from someone claiming to be his grandson, saying he's in trouble and needs money wired today. Under a stereotype-based design, his account might be frozen until one of his children approves any transfer, which also blocks him from paying his own contractor next week. Under the workflow above, he has already agreed with his family on a simple rule: any urgent money request gets a pause and a call back to a number he already has. He hangs up, calls his grandson's real number, and finds out the truth. He stays in charge. The pause did the work.

Nothing in that example depends on artificial intelligence. That's worth noticing. The strongest protections here are process, not software.

#### 9.8.6 What this means for you

**For older adults:** You're allowed to hang up. You're allowed to call back on a number you trust. Real institutions don't punish you for verifying.

**For adult children:** Offer tools, not supervision. Ask your parent what help they want with fraud, and agree on a pause rule and a call-back rule together.

**For banks, carriers, and product teams:** Evaluate your warnings on both sides. Count false alarms and missed scams. Don't design protections that treat every customer over 60 as incapable.

### 9.9 Accessibility and the Non-Digital Route

A tool that isn't usable by the person it's meant for doesn't help, however good the underlying model is.

#### 9.9.1 What the W3C guidance is and isn't

W3C's cognitive-accessibility guidance recommends usable content and interfaces beyond basic technical conformance. It is published as a Working Group Note, which makes it design guidance, not a new normative conformance standard and not an effect estimate for any particular product. ([W3C cognitive accessibility guidance](https://www.w3.org/TR/coga-usable/))

That distinction matters if you're a buyer. A vendor saying "we follow the W3C cognitive accessibility guidance" is saying something about design intent. It is not evidence that their product improves outcomes for anyone.

#### 9.9.2 The proposed implementation standard

The implementation standard I'm proposing draws on that guidance:

- **Clear language.** Plain words, short instructions, no jargon.
- **Recognizable controls.** Buttons and actions that look like what they do and stay where people expect them.
- **Recoverable errors.** A mistake shouldn't be a trap. People should be able to undo and try again.
- **Understandable consent.** A person should know what they're agreeing to, especially about what information is shared and with whom.
- **Testing with the intended users.** This is the item most often skipped: testing with the actual people the product is for, rather than a convenience sample of developers and their relatives.

#### 9.9.3 Age is not a proxy

Age alone should not be treated as a proxy for disability, digital skill, or desired level of assistance. The 70-year-old retired systems engineer and the 70-year-old with early cognitive decline are different users who happen to share a birth year. Any design or policy that treats them as one demographic will fail at least one of them.

This is the same equity point as the fraud section, from a different angle. Designing for "seniors" as a single group tends to produce either condescension or neglect. Designing for real people with real needs, and testing with them, is harder and works better.

#### 9.9.4 The non-digital route

The final requirement is structural: a public service should remain available when a person declines automated interaction. That's a proposed equity and continuity requirement, not an empirical finding that existing services meet it. And if you've tried to reach a bank, an airline, or a benefits office by telephone lately, you already know how often the requirement goes unmet.

The non-digital route is the emergency continuity principle from section 9.6, applied to public services. If the only way to reach help runs through an app or a chatbot, then everyone who can't or won't use that channel has been quietly shut out.

> **Operator's note: Keep the phone line**
>
> When you run infrastructure that people depend on, you never take away the manual path. Automation is great until it isn't, and when it isn't, someone needs a way to reach a human who can override the system. I think public services owe residents the same thing. Offer the automated channel. Make it good. Measure whether it helps. But keep the phone line and the counter open for the person who needs them, because the day they need them is usually the day everything else has already failed.
### 9.10 Retirement as an Ongoing Life Stage

Retirement isn't a single decision made at 65. It's a long stretch of life, and the decisions in it keep coming.

#### 9.10.1 Why this topic gets two chapters

Retirement decisions connect finance, housing, relationships, health, transport, and community participation. That's why this report gives the topic two chapters rather than a section. [Chapter 8](#ai-financial-advice) addressed research on advice quality and declined to convert simulations into prescribed plans. This chapter doesn't convert them either. I'm not going to tell you how to allocate a retirement account or when to claim benefits.

#### 9.10.2 The proposed later-life assessment

The proposed later-life assessment asks five questions of any assistance an older adult uses:

1. Does it reduce missed obligations (appointments, bills, renewals)?
2. Does it improve access to services and care?
3. Does it support meaningful contact with other people?
4. Does it preserve autonomy?
5. Does it avoid harm?

The outcome set is deliberately two-sided. Adverse dependence, unwanted surveillance, incorrect advice, caregiver workload, and the availability of human support must be measured alongside user satisfaction. A companionship product that scores well on satisfaction while quietly displacing human contact has failed the assessment I'm proposing, whatever its survey numbers say.

Here's the part most people skip. Satisfaction is easy to measure and pleasant to report. Displacement is hard to measure and uncomfortable to report. So if nobody deliberately measures the hard side, the easy side wins by default, and products get judged by how much people like them rather than whether their lives got better. That's the whole game.

#### 9.10.3 Follow-up separates a demonstration from a life

Follow-up is the discipline that separates a demonstration from a life. The short durations (roughly two to twelve weeks) and heterogeneous interventions in the older-adult trial review limit every claim about durable effects, and a short successful demonstration can differ materially from months of routine use. ([BMC Geriatrics randomized-trial review](https://pmc.ncbi.nlm.nih.gov/articles/PMC13321516/))

Any evaluation that ends when the novelty ends has measured novelty. A new device in the house gets attention for a while. Grandkids ask about it. The person tries it out. Then it becomes furniture, or it becomes part of life. Only follow-up tells you which.

#### 9.10.4 What this means for each reader

**For older adults:** You're the one who decides whether a tool is worth keeping. Judge it by whether you can do more, decide more, and connect more.

**For adult children and caregivers:** Measure your own workload too, and be candid about it. The caregiver evidence shows small average effects and wide uncertainty for any specific program, so watch whether a tool actually lightens your load or just changes its shape.

**For advisers and service providers:** Keep a human route open, keep your claims inside your evidence, and don't use aggregate statistics to stereotype your clients.

**For board members and funders:** Fund evaluations with control groups, two-sided outcomes, and follow-up beyond the pilot window. [Chapter 11](#ai-accountability-framework) sets out the full accountability framework and evidence classes I use to judge these claims.

### 9.11 Conclusion: What AI in Healthcare Research Supports Today

The strongest clinical example in this chapter is a narrowly defined, professionally supervised screening workflow in which a model-support system achieved noninferior interval-cancer outcomes and improved sensitivity without losing specificity ([MASAI](https://pubmed.ncbi.nlm.nih.gov/41620232/)).

The conversational and older-adult evidence contains important null results and real uncertainty: a two-point physician-workflow difference that is statistically indistinguishable from zero ([physician trial](https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2825395)), a loneliness estimate whose interval spans benefit to harm, and no controlled basis for companionship-replacement claims ([older-adult review](https://pmc.ncbi.nlm.nih.gov/articles/PMC13321516/)). Those findings support differentiated evaluation, task by task, workflow by workflow, outcome by outcome. They support neither blanket endorsement nor blanket rejection.

The standard I'm proposing is assistance that preserves independence and access to people, not a system that replaces human support because replacement is easier to count. Retirement should remain a stage of agency and participation. The measure of success is an older adult who can do more, decide more, and connect more, not one who has been efficiently managed.

For the adult child at the kitchen table with the folder, the pill bottles, and the phone full of scam texts: you don't need to become an expert. You need a boundary (tools prepare, professionals decide), a workflow that starts with your parent's goals, a paper backup, a pause rule for money, and the habit of asking what the evidence actually measured. That's enough to use these tools without being used by them.

### 9.12 Frequently Asked Questions

#### What did the MASAI AI mammography study find?

The Swedish MASAI trial analyzed 105,915 women and found 1.55 interval cancers per 1,000 with AI-supported screening versus 1.76 with standard double reading, a ratio of 0.88 (95% CI 0.65 to 1.18). That met the noninferiority criterion but was not a significant reduction. Sensitivity rose from 73.8% to 80.5% with specificity near 98.5% in both groups. Mortality was not measured, and the result applies to a radiologist workflow.

#### Is AI more accurate than doctors at diagnosis?

The evidence reviewed here does not show that. In a randomized trial of 50 physicians on structured vignettes, access to a language model changed median reasoning scores by an adjusted 2 points (76% vs 74%, 95% CI -4 to 8), not statistically significant. The model alone performed strongly in an exploratory analysis, but that capability did not translate into a demonstrated improvement for the physicians using it.

#### Do AI companions reduce loneliness in older adults?

The best controlled evidence reviewed cannot say. A May 2026 BMC Geriatrics review pooled only three trials with 190 participants for loneliness and found g -0.67 with a 95% CI of -2.57 to 1.23, which spans large benefit to substantial harm. Interventions lasted roughly 2 to 12 weeks, and there is no controlled basis for claims that a device can replace human contact.

#### Can AI help with depression in seniors?

A May 2026 review of randomized trials found a small effect on depression across seven trials (Hedges' g -0.25, 95% CI -0.48 to -0.02), rated moderate certainty. The interventions included social robots, voice systems, and cognitive activities lasting roughly 2 to 12 weeks, not a uniform test of current chatbots. It is a signal of limited promise, not a treatment, and not medical advice.

#### Is it safe to use a chatbot to check my parent's symptoms?

This chapter is not medical advice, and the evidence supports a clear boundary: assistance can prepare questions and organize records, while professionals diagnose, treat, and handle urgency. In a Nature Medicine vignette study of 1,298 U.K. adults, people using models identified relevant conditions less often than those using usual resources, and all groups tended to underestimate acuity. For urgent symptoms, contact a clinician or emergency services.

#### How much money do older adults lose to scams, and is AI to blame?

The FTC reported that adults aged 60 and older reported $2.4 billion in fraud losses in 2024. Those are reported losses only, so the real total is unknown, and the FTC did not attribute the total to generative AI or any particular technology. The data also do not support treating every older adult as less able to spot deception.

#### Do AI scam detectors protect seniors from fraud?

No trial in the reviewed evidence shows that a particular automated scam detector prevents real-world losses. Warnings also need two-sided evaluation: frequent false alarms train people to ignore alerts, and a reassuring miss can be worse. The protective workflow proposed here relies on verified contact information, a pause before unusual transfers, a trusted person the older adult chooses, and a clear escalation route.

#### How should families set up AI tools for an aging parent?

Start with the parent's own goals and consent, not the family's convenience. Use tools for appointment preparation, verified medication lists for a pharmacist, service navigation, and authorized care coordination. Check whether the parent completes more tasks independently rather than simply being monitored more, and keep a paper plan that works if the phone, internet, or power fails. Every use needs local testing.
