Chapter 4 · College, Trades, First Jobs
AI in College, Skilled Trades Training, and the First Years of Work

Picture a 19-year-old at a kitchen table with three open tabs, and a pile of AI in higher education research headlines telling them what to think. One is a university acceptance letter. One is an apprenticeship application for an electrical program. One is a job posting that pays next Friday. Every adult in that kid's life has an opinion, and now every one of those paths comes with an AI pitch attached: the college has a chatbot, the training center has a simulator, and the job board offers to write the resume. The real question underneath all of it is simple. Which of these will actually make me able to do something someone will pay for? That is the question this chapter takes to the evidence, and it is the question most of those headlines get quoted around rather than answering.
I care about this one personally. I have spent my career building large power infrastructure. When you build at industrial scale, you hire electricians, technicians, and operators, and you learn fast that there is exactly one credential that matters on an energized site: can this person do the work, safely, when nobody is standing over them? A diploma tells me someone finished something. A certificate tells me someone passed something. Neither one tells me what happens when a breaker trips at 2 a.m. and the procedure on the screen doesn't match what's in front of them. "Can do the work" is the whole credential. Everything else is a proxy.
That operator's lens is how I read this chapter's eleven study programs. Some of them are good news. An enrollment chatbot helped committed students actually show up. A carefully designed physics tutor beat a strong classroom on short-term tests. A resume tool raised hiring on one platform. A workplace assistant raised output per hour. Two long-running workforce programs posted real earnings gains. Here's the part most people skip: the chapter also contains a randomized experiment where adults learning a new technical skill with an AI assistant came out understanding it less well, with no time saved. And the trades evidence, the part closest to my world, is thin, short, and about welding only.
So the finding is not "AI helps" or "AI hurts." It is that enrolling, finishing an assignment, demonstrating a skill, getting hired, and performing on the job are five different achievements. A tool can win one and leave the other four untouched. I'm going to walk through each study with its population, its comparison, and its limit attached, then show you what a real competence sequence looks like on a site where mistakes are measured in arc flash, and finish with practical advice for students and job seekers who want to use these tools without hollowing out the very skill they're trying to build.
4.1 AI in Higher Education Research: Five Different Wins#
This chapter picks up where the K-12 evidence in Chapter 3 leaves off, in the years where education turns into employment. The organizing idea is practical: helping a person enroll, finish an assignment, demonstrate a skill, obtain work, and perform at work are different achievements. A system can accomplish one while leaving the others untouched, and the reviewed evidence does exactly that. Improvements in college enrollment, immediate subject learning, hiring, and workplace output sit alongside uncertain vocational results and an experiment showing weaker independent technical understanding after assisted practice. (College enrollment experiment, college physics experiment, hiring experiment, workplace study, welding experiment, technical skill experiment)
The standard I hold this chapter to follows from that. Educational assistance should produce demonstrable capability, not merely a more polished submission. Where the goal is administration or workplace production rather than learning, the evaluation should say so plainly and measure that goal directly. Confusing the two is how institutions end up measuring the polish and calling it the capability.
4.1.1 How the Eleven Studies Were Chosen#
I selected eleven decision-relevant study programs, prioritizing original studies, government evaluations, and identifiable comparison conditions. College, vocational training, and employment are treated as overlapping pathways, not a ladder where every learner must climb to a four-year degree. That matters to me. Some of the most capable people I've ever put on a site never set foot in a lecture hall, and some of the most credentialed people I've interviewed couldn't read a one-line diagram.
This is a targeted synthesis, not a systematic review. It offers no pooled effect and makes no claim to have found every study. Publication status is tracked separately from design: a randomized working paper and a peer-reviewed observational study each keep both labels, and I never merge them into a single "proven" bucket. A working paper can be well designed. A journal article can be observational. You need both facts to judge the finding.
4.1.2 What "AI" Means in This Chapter#
The technology boundaries need restating. "Superintelligence" names the transition this report investigates. It is not a verified capability level of the tools reviewed here. The evidence in this chapter includes generative tutoring, non-generative messaging support, virtual and augmented reality, and training programs with no model-based intervention at all.
Why include old and non-generative systems in a report about the AI transition? Because they identify mechanisms and credible alternatives. If a text-message enrollment bot moves outcomes, that tells you something about what a newer assistant would have to beat. If an intensive program with no AI at all produces seven-year earnings gains, that is the benchmark any AI-enabled pathway should be measured against. What I won't do is relabel the effect of an older system as evidence that a newer general-purpose model produces the same outcome. Capability is not benefit, and a result earned by one tool doesn't transfer to another because they share a chat window.
The education-to-work pathway and where each study's evidence stops
The funnel view is the most useful way to read everything that follows. Each of the eleven programs sits at one stage of the education-to-work sequence, from choosing a pathway through progressing in a job. Each one evidences its own stage and stops there. When you see a headline that jumps from one stage to another, from "students liked the tutor" to "graduates earn more," you're watching someone cross a gate the evidence never crossed.
4.2 Can a College Chatbot Get Admitted Students to Show Up?#
The first stage of the funnel is the least glamorous and, for a lot of families, the most consequential. A student can be admitted, intend to go, and still never arrive because a form, a deposit, a transcript, or a financial-aid step fell through the cracks over the summer. Researchers call that "summer melt." It is an administrative failure, and it turns out to be one of the places where a narrow, well-built tool can help.
4.2.1 Pounce and the Reduction of Summer Melt#
Page and Gehlbach evaluated Pounce, a university enrollment-support chatbot, in a randomized study involving 7,489 admitted Georgia State University students, 1,948 of whom had already committed to attend. The intervention used text outreach, a knowledge base, student-specific enrollment information, and escalation to admissions staff. It was not an open-ended generative tutor. It answered procedural questions and prompted action on enrollment tasks, while staff handled anything it couldn't resolve. (Original AERA Open article)
Among already-committed students, the treatment increased enrollment at Georgia State by 3.3 percentage points, with a standard error of 1.6 percentage points. The authors describe this as a 21 percent relative reduction in summer melt, meaning admitted students who intend to enroll but fail to complete the steps. The full-sample university-enrollment estimate of 1.2 percentage points was not statistically significant, and neither was the committed-subgroup estimate for attendance at any postsecondary institution. (Study results, Full-sample and attendance estimates)
Pounce enrollment chatbot at Georgia State
Source: Page and Gehlbach, AERA Open
Read those two numbers side by side. The 3.3-point gain applies to students who had already said yes. Across everyone admitted, the 1.2-point estimate could not be distinguished from zero. And for the committed group, the study could not show more students attending college anywhere, only more attending this particular university. That's a useful result for Georgia State's enrollment office. It is a much weaker result for the national question of whether chatbots send more young people to college.
What the study demonstrates is narrow and useful: more already-committed students enrolled at the participating institution under tested conditions. What it does not demonstrate is a general increase in college attendance across all admitted students, stronger academic learning, degree completion, or lifetime earnings. The result belongs to the funnel's first stage. Treating it as evidence about the later stages is the exact pathway confusion this chapter exists to prevent.
Common misreading: "A chatbot cut summer melt by 21 percent." That's the authors' relative framing of a subgroup result. Show me the denominator: it's committed students at one university, and the absolute change was 3.3 points with a standard error of 1.6.
The implementation lesson follows from the design. Separate administrative assistance from academic authority. An enrollment assistant should explain deadlines and procedures from current institutional records, show where those records came from, and escalate eligibility questions, financial-aid disputes, and unusual circumstances to authorized staff. Every clause there is doing work. Current records, visible sourcing, and human escalation are what separated this intervention from an unsupported advice bot. Strip those out and you have a different tool that happens to use the same text thread.
For a parent: if your student's college offers a text assistant for enrollment steps, use it for what it's good at, which is reminders and procedural answers. For anything involving money, eligibility, or an exception, get a human name and a written answer.
4.2.2 Course Reminders at Georgia State: A More Mixed Result#
The 2026 version of "Let's Chat" reports randomized course-support outreach covering 2,483 students in introductory government and microeconomics courses at Georgia State. The intervention added personalized course messages and knowledge-base support to a population that already had access to the university's general retention chatbot. So this was a test of additional course-specific support, not a comparison against no support at all. That comparison asymmetry disciplines everything you can say about it. (Meyer and colleagues, working paper)
The pooled numeric course-grade estimate was +1.80 points, but it was not statistically significant after the study's multiple-comparison adjustment. The estimated four-percentage-point increase in earning a B or higher carried an adjusted p-value of .090. A stronger result for women in microeconomics was a subgroup finding, not the overall effect. The working paper did not establish broad improvements in credits from other courses, overall semester GPA, or subsequent course enrollment. (Pooled and subgroup results, Additional outcomes)
The practical implication cuts against a common deployment reflex. Reminder systems should be evaluated against the support students actually receive, and message response rates should stay what they are, which is implementation measures. They are not substitutes for course completion or independent learning. A chatbot that students answer is a communications channel. Whether it is an educational intervention is a separate question, and this study answers it only conditionally.
Common misreading: "Students who got AI course nudges did better." The pooled estimate was positive and not significant after adjustment. The subgroup result for women in microeconomics is worth following up, and it is a lead for the next study, not a finding for the brochure.
For a college administrator: before you buy a course-nudge product, write down what your students already receive. If they already have a general chatbot, advising, and early-alert emails, the question is not "does the new tool help versus nothing" but "does it help on top of all that." This study is the template for asking it correctly.
4.3 Does an AI Tutor Help College Students Learn?#
This is the section that generates the most headlines, and the one where I'd ask you to slow down the most. There is one strong college learning study in this chapter's set. It is good work. It is also the finding most likely to be stretched past what it measured.
4.3.1 The Harvard Physics AI Tutor: Promising, Bounded Evidence#
Kestin and colleagues studied a custom GPT-4-based tutor in an introductory Harvard physics course, using randomized crossover sequences across two lessons. The study reported 194 eligible participants, with 142 tutor-condition and 174 classroom-condition posttest observations. Those are observations across crossover conditions, not two independent samples of students, so you can't add them up and call it a head count. (Scientific Reports article)
The comparison is the first thing to understand. The tutor was tested against an in-person active-learning class, not a passive lecture and not the absence of instruction. That's a demanding comparison, and it's why the result earns attention. The tutor itself was a purpose-built instructional package: deliberately structured prompts, expert-written solutions, guided problem progression, and instructional materials. It was not a student opening a general chatbot and typing "explain this."
The reported adjusted effect on immediate posttest performance was 0.63 standard deviations in favor of the tutor condition. Median tutor use was approximately 49 minutes, against the authors' estimate of roughly 60 instructional minutes in the classroom condition after excluding assessment time. (Regression results, Time comparison)
The designed physics tutor: a bounded win
Four qualifications attach, and each one matters to a decision-maker.
- Assessment timing. Outcomes followed individual lessons. The study did not measure what students remembered months later or whether they finished the course or the degree.
- Comparison time. You'll sometimes see "49 versus 75 minutes." That figure mixes tutor use with a full class period that includes tests. The appropriate comparison uses the study's own instructional-time accounting, which is roughly 49 against roughly 60.
- Design generality. The result does not establish that any generic chatbot equals the studied tutor, that a tutor replaces a university course, or that the short-run advantage persists into later coursework.
- Cost. Building the tutor required substantial expert preparation. The learning result must not be converted into a claim that the course became cheaper to deliver. (Study design and limitations, Development process)
What the study supports is the proposition that a carefully engineered tutoring workflow can outperform a strong instructional comparison on selected immediate outcomes. That is a real finding. It is also the most commonly overstated finding in this report's corpus, because "GPT tutor beat Harvard physics class" travels much farther than "a purpose-built instructional package designed by the course's own experts outperformed that same course's active-learning sessions on immediate posttests across two lessons, at unknown total cost, with no long-term measure." Both sentences describe the same study. Only one of them is evidence.
Common misreading: "Students learn twice as much from AI tutors." The study reported an adjusted 0.63 standard deviation advantage on immediate posttests. It did not report a multiple of learning, and it did not test a generic tool.
For a student: the takeaway is not "use ChatGPT for physics." It is that tutoring built around structured problems, expert solutions, and step-by-step progression can work in the short run. If your school offers a tutor built that way by your own instructors, that's a different product from an open chat window, and the evidence here is about the first one.
For a professor or department chair: the design ingredients are the finding. If you want to replicate the result, budget for the expert preparation, keep the active-learning comparison, and add the delayed measure the original study didn't have.
4.3.2 The College Evaluation Standard I'd Hold Every Pilot To#
A college pilot should distinguish two legitimate goals: helping students produce work, and helping students become able to produce comparable work independently. Both are fine goals. They are not the same goal. The program should declare which one is primary before collecting results, because the measurement sequence differs and the temptation to promote the easier goal after the fact is constant.
Here is the assessment sequence I'd propose, in order:
- Baseline task. What can the student do before assistance?
- Assisted practice. The tool is available.
- Unaided assessment. The tool is removed.
- Delayed assessment. Weeks later, still without the tool.
- Transfer task. A problem that differs from the practiced examples.
For high-stakes claims, graders should be blind to assignment where feasible, and the comparison should preserve instructional time or explicitly account for differences. Recommended measures:
- course completion
- independent reasoning
- citation accuracy
- delayed retention
- academic-integrity incidents
- accessibility
- student verification time
- instructor workload
- total delivery cost
Satisfaction and confidence can supplement those measures. They should not replace them. A student can be delighted by a tool that is quietly eroding the skill they enrolled to acquire, and the high-school mathematics experiment in Chapter 3 is the standing demonstration: students with unrestricted assistance did better during assisted practice and then performed 17 percent worse than the control group once the assistance was removed.
Questions to bring to a faculty senate or board meeting about an AI pilot:
- Which goal is primary, producing work or building independent capability, and when was that written down?
- What happens on the unaided assessment, and is there a delayed one?
- Who grades, and do they know which students used the tool?
- What did the comparison group receive, and for how long?
- What did the tool cost to build, license, and support, counting faculty time?
- How many academic-integrity cases, accessibility complaints, and citation errors were logged?
4.4 AI in Skilled Trades Training: What the Welding Studies Show#
Now we get to the part of this chapter closest to where I've spent my working life. I want to be precise about how little the evidence covers, because the trades are where overclaiming costs the most. A sloppy essay costs a grade. A sloppy termination costs a lot more.
4.4.1 Virtual Welding: A Null Result Is Not Proof of Equivalence#
Wells and Miller randomly assigned 101 university volunteers to live welding, virtual-reality welding, or two mixed sequences. Participants received approximately 30 minutes of practice, and three certified welding inspectors evaluated their subsequent live welds using a structured rubric. (Journal of Agricultural Education article)
The four groups' average weld scores differed, but the overall comparison did not reach the conventional 5 percent threshold: F(3,97) = 2.235, p = .089. This finding establishes neither virtual training's superiority nor its equivalence. A nonsignificant difference is not a noninferiority finding. Treating it as one is among the most common statistical errors in technology procurement. (Statistical results)
Read that again, because it's the sentence that saves budgets. "No significant difference" means the study couldn't tell the groups apart with confidence. It does not mean the study showed they were the same. To show equivalence, you have to design for it up front, decide how close is close enough, and power the study to detect that. This study wasn't built to do that, and nobody should cite it as if it were.
The implementation conditions bound the result further. Approximately 69 percent of participants had prior welding experience. The virtual system's visual guidance cues were disabled, and its audio was not functioning. The short practice period and immediate assessment do not establish long-term retention, injury reduction, employment, or successful professional certification. The study is relevant to training design. It is not evidence that trades can be taught through simulation. (Participant and implementation details, Study scope)
Common misreading: "Research shows VR welding is as good as the real thing." It doesn't. It shows a short, partially impaired VR setup couldn't be statistically separated from live practice in a mostly experienced group. That's a reason to run a better study, not a reason to close a welding bay.
4.4.2 WeldAR: Promising Transfer on a Narrow Skill Measure#
The 2026 WeldAR manuscript describes a crossover study with 24 novice participants comparing augmented-reality guidance during actual welding with video instruction. Participants received both conditions in randomized order, followed by unassisted practice measurements. The authors identify the manuscript as accepted for CHI 2026. (WeldAR author manuscript)
The system displayed real-time guidance on torch movement and positioning. The study found improved unassisted motion performance: the first-period comparison favored augmented guidance with p = .032. Note what was measured. The endpoint was adherence to movement targets, not a certified weld-integrity test or an occupational qualification.
The limits are real. The sample was small. The comparison was video instruction, not individualized coaching from an expert. Crossover order and carryover complicate interpretation. Tracking problems, headset weight, visibility, and the short observation period further limit any deployment claim. (Study results and endpoint definitions, Limitations and implementation findings)
The right role for this kind of support is supplementary practice with qualified supervision and independent verification. It should not authorize unsupervised work, replace required safety instruction, or let a motion score stand in for inspection of the finished work. The distinction between a motion endpoint and a weld-integrity endpoint is the whole story. The study measured whether trainees moved the torch the way the system asked. It did not measure whether the weld holds.
For a training director: a device that improves how trainees move is worth piloting as extra practice reps. Put it under an instructor, keep the inspection standard exactly where it is, and track whether AR-trained cohorts pass the same destructive or visual tests at the same or better rates. That's the number you'd need before changing anything else.
4.4.3 Infrastructure Trades and the Competence Sequence#
Let me start with the boundary. The welding studies concern welding. They are not about data center electrical commissioning, high-voltage work, refrigeration qualification, or emergency response. Their findings cannot establish competency in those different occupations, and this report will not transfer them. (Virtual welding study, WeldAR scope)
That leaves a gap between what the research covers and what people in my industry are going to be asked to decide. So here is how I think about it, labeled as what it is: operator experience, not evidence.
On a large power site, the work that can hurt people is the work where the equipment is energized, the energy is stored, or the system is under pressure. The people doing that work need to recognize a hazard they have never seen in exactly that configuration, stop, and call it. That capability isn't a fact you retrieve. It's judgment built through supervised repetition, checked by someone qualified who is willing to say "not yet."
A model can help someone prepare for that. It can explain a concept three different ways, quiz someone on a procedure, walk through why a sequence matters. I've got no objection to any of that. What it cannot be is the controlling authority on what's safe to do next. A training pipeline whose safety-critical step is "the model said so" is not a training pipeline. It's a liability with a login.
Here's the reason, stated plainly. On an energized site, the procedure is not a suggestion, and the person following it has to own it. If the model is right nearly every time and wrong once, the one wrong answer is the one that matters, and a trainee who learned to defer to the screen won't catch it. The skill you're building is exactly the skill of catching it. So any pipeline that routes the final safety judgment through a tool is training the opposite of what the job requires.
For a proposed infrastructure training program, the relevant sequence looks like this:
- Occupational analysis. Define the actual tasks, hazards, and decisions in the job, not a generic job title.
- Qualified instruction. A person who holds the relevant qualification teaches it.
- Controlled practice. Deenergized, simulated, or mockup equipment where mistakes are cheap.
- Independent practical assessment. Someone other than the instructor, and certainly not the tool, verifies that the trainee can perform the task.
- Supervised field experience. Real equipment, real conditions, a qualified person present and accountable.
- Documented authorization. A written record that this person is cleared for this specific work.
Model-generated explanations can assist preparation at steps 2 and 3. They should not be the controlling safety procedure at any step. That's the line, and it doesn't move because a vendor demo looked good.
For a county commissioner or economic development official: when a facility proposal promises "local workforce training," ask which occupations, which of the six steps the program actually provides, who performs the independent assessment, and where the supervised field experience happens. A promise of AI-powered training that stops at step 3 is a promise of practice, not of qualified workers.
For a young person considering an electrical or mechanical trade: use AI tools to understand theory, check your math, and prepare for written exams. Don't use them to decide whether something is safe. The person who signs off on your work should be a qualified human, and one day that person should be you.
4.5 Apprenticeship Outcomes and Employer-Connected Pathways#
If the college and trades evidence is mostly about short-run tests, the apprenticeship and workforce evidence is where you finally get to see paychecks. Neither of the two programs here depends on a model. I include them because they are the benchmark. Any AI-enabled pathway that wants to claim workforce value should be measured against programs like these, on outcomes like these.
4.5.1 Apprenticeship Outcomes: Positive Estimates With Selection Limits#
A September 2025 U.S. Department of Labor evaluation examined registered and unregistered apprenticeship programs supported by Scaling Apprenticeship and Closing the Skills Gap grants. For registered apprenticeships, it compared participants with public workforce-service customers and with selected community-college students, using weighting and regression rather than random assignment. (Department of Labor evaluation)
At the ninth quarter after enrollment, registered-apprenticeship participants showed estimated employment advantages of 7.8 percentage points over the workforce-service comparison and 8.8 percentage points over the community-college comparison. Estimated quarterly earnings differences were $3,230 and $4,693 respectively, in the report's stated units. The two comparisons used different samples: 1,271 apprentices for the workforce-service comparison and 1,124 for the community-college comparison. (Employment estimates, Earnings estimates)
Registered apprenticeship: estimated employment and earnings advantages
Source: DOL evaluation, Sept 2025
Now the limits, which are as important as the numbers. Many apprentices were already employed when they enrolled. Statistical adjustment cannot fully eliminate differences in employer selection, motivation, or other unobserved characteristics. Think about who gets into a registered apprenticeship: someone an employer chose, who chose the program back, and who showed up. Those same traits could raise earnings on their own. Weighting and regression can account for what the researchers could measure. They can't account for what they couldn't. (Samples and identification limitations)
One discrepancy belongs on the record. The report's executive summary, results chapter, and Table A.10 give the lower earnings estimate as $3,230, while the report's Chapter 6 gives $3,320. I retain the $3,230 estimate supported by the results table rather than quietly reconciling the conflicting prose. If you check the report yourself, you should find the same number I quote, and you should also find the one place where it doesn't match.
These figures should stay what they are: quarter-specific historical study estimates. They are not annualized promises to prospective students. Multiplying a ninth-quarter difference by four and printing it on a flyer turns a careful estimate into a sales claim. They also don't measure the contribution of any model, and they don't imply that every apprenticeship or every entering worker receives the same benefit.
Common misreading: "Apprenticeship adds a fixed amount to your yearly pay." Nobody in the report said that. The study reported a quarterly difference at one point in time against one comparison group, from a design that can't fully rule out selection.
For a young person weighing an apprenticeship: these results are encouraging, and they're one reason I'd never tell a 19-year-old that a four-year degree is the only serious path. Ask the specific program for its own completion and placement records, including people who didn't finish.
4.5.2 Year Up Earnings: The Long-Run Benchmark for the Whole Package#
The Year Up evaluation randomly assigned 2,544 eligible young adults aged 18 to 24 to program access or to a comparison group that could pursue other community services. The intervention combined six months of training with six months of internships, alongside stipends, advising, mentoring, professional-skills development, and job-placement support. (ACF long-term evaluation)
The wage-record analysis covered 2,495 participants, and the prespecified confirmatory outcome was average quarterly earnings in follow-up quarters 23 and 24. The estimated impact was $1,895 per quarter, approximately 28 percent above the comparison group's $6,901 quarterly average, with p < .001. The study followed outcomes for approximately seven years. It did not observe a ten-year or lifetime effect, and the report's longer projections are assumptions about future persistence, not additional measured follow-up. (Confirmatory result, Follow-up and projections)
Year Up: average quarterly earnings, follow-up quarters 23 and 24
Source: ACF long-term impact report
The benefits shouldn't hide the opportunity costs. Reported earnings were lower for the treatment group in the first year, while participants devoted time to the program, with stipends reported separately rather than counted as wage earnings. That dip is real money for a young adult with rent due. And the package's long-run earnings effect did not establish improvement in every measured education, employment, or well-being outcome. (Annual results and other outcome domains)
Year Up supplies the benchmark this report's thesis needs: a workforce pathway evaluated on sustained outcomes rather than course registrations. It does not show which component produced the gain, whether the training, the internship, the stipend, the mentoring, or the combination. And it does not show that adding a model to a less intensive program reproduces the result. The defensible lesson is about the shape of credible workforce evidence (randomization, prespecified long-run earnings outcomes, published opportunity costs) more than about any single number.
That's the whole game for anyone proposing an AI-enabled workforce program. Year Up's design is what "proven" looks like in this field. It took years of follow-up, a real comparison group, an outcome chosen in advance, and the willingness to publish the first-year dip. A program that reports enrollments, satisfaction scores, and "learners served" is reporting inputs. Year Up reported earnings seven years out.
4.5.3 Designing a Workforce Pathway That Can Be Evaluated#
Workforce research programs should compare an assisted program with a credible alternative pathway, including conventional instruction and employer-supported practice. A comparison with "no training" can be informative, but it doesn't establish that the proposed approach is the best use of the money.
The proposed delivery package should specify, in writing:
- paid or unpaid participation
- childcare and transportation support where relevant
- prerequisites
- employer commitments
- occupational assessments
- placement assistance
- the consequences of noncompletion
These are design requirements for evaluation. They are not findings that any particular program already meets them. I list them because every one is a place where a program can look good on paper and fail the participant. An unpaid program loses the people who can't afford to go without wages. A program without transportation support loses the people who live farthest from the site. A program without real employer commitments trains people for jobs that aren't there.
4.6 AI Resume Writing and the First Job#
Most young people's first direct encounter with AI in the job market is the resume. The evidence here is better than you might expect, and more specific than the headlines.
4.6.1 AI Resume Writing Study: Presentation Without Invention#
Wiles, Munyikwa, and Horton studied algorithmic writing assistance in a randomized experiment involving approximately 481,000 new registrants on an online labor platform. The intervention suggested improvements to resume writing. It did not generate a complete fictional professional history. The research was subsequently published in Management Science in 2025. (Research and publication record, detailed author manuscript)
In the detailed manuscript's main analysis of 194,701 approved profiles with nonempty resumes, the first-month hiring effect was approximately 0.247 percentage points against a control hiring rate of 3.093 percent, about 8 percent in relative terms. "Eight percent more likely to be hired" must not be rewritten as an eight-percentage-point increase. The first is a small bump on a small base. The second would be enormous. (Manuscript estimates and sample)
The analysis did not find evidence of lower employer satisfaction, but a nonsignificant satisfaction difference is not proof of exact equivalence. The study concerns hiring on a particular platform. It is not about first-ever employment for all participants, a national increase in available jobs, or long-run career stability. (Study outcomes and scope)
A number without a denominator is a rumor, and this study is the cleanest example in the chapter. Roughly 3 in 100 control-group profiles got a hire in the first month. With help, it was a little more than that. Useful, especially at platform scale. Not a job guarantee.
The acceptable use follows from the design: improve the clarity of accurate information supplied by the applicant. A workforce program should prohibit invented credentials, fabricated experience, false references, and unreviewed submissions. Assistance that polishes a truthful resume and assistance that fabricates one are different interventions with the same interface. The boundary between them has to be drawn in policy, not left to the model's discretion.
Speaking as someone who has hired for jobs where a fabricated qualification can get someone hurt: a resume that overstates what a person can do isn't a small ethical lapse on a power site. It puts the wrong person in front of the wrong equipment. Every experienced hiring manager I know checks the claims that matter. The resume that wins is the clear, true one.
4.6.2 Workplace Support: A Meaningful, Setting-Specific Result#
Brynjolfsson, Li, and Raymond's 2025 Quarterly Journal of Economics article studied a staggered introduction of a conversational assistant among 5,172 customer-support agents. The assistant suggested responses and relevant information, and agents could accept, edit, or ignore the suggestions. The preferred estimate was approximately 15 percent more customer issues resolved per hour, with larger benefits for less-skilled and less-experienced workers. This was a difference-in-differences analysis of deployment in one company, not a randomized national workforce experiment. (Published study, Productivity estimates and design)
The heterogeneity cuts both ways. The study examined performance during outages, when the assistant was unavailable, and found patterns consistent with learning. It also reported limited benefits and some quality deterioration among the most-skilled workers. Those findings caution against both extremes: assuming assistance always substitutes for learning, or assuming every worker benefits equally. (Heterogeneity and learning analyses)
The result supports testing supervised assistance during onboarding in comparable workflows. It does not establish a 15 percent wage gain, 15 percent fewer employees, or a transferable 15 percent benefit in electrical work, medicine, law, or other unrelated occupations. Output per hour is a real outcome. It is not a wage outcome, a hiring outcome, or an economy-wide outcome, and Chapter 5 follows where each of those conversions breaks down.
For an employer onboarding new hires: this is the most encouraging finding in the chapter for your purposes, with a condition attached. The gains were largest for newer workers, and the outage evidence suggests some of what they learned stuck. Track quality along with speed, watch your most experienced people for the deterioration the study reported, and schedule periodic work without the assistant so you can see what your new hires can do on their own.
4.7 Protecting Competence in the First Years of Work#
The customer-support study is the optimistic case: assistance that seems to leave some learning behind. The next study is the adverse case, and it's the one I'd most want a new graduate, a new hire, and their manager to read together.
4.7.1 The Trio Learning Experiment: An Adverse Randomized Result#
Shen and Tamkin's 2026 preprint reports a randomized experiment with 52 programmers learning the unfamiliar Python Trio library. The assisted group had access to a GPT-4o-based coding assistant. Both groups then took a quiz without model assistance. The assisted group scored lower on immediate comprehension, with a reported standardized difference of Cohen's d = 0.738, p = .010, while the average task-completion-time difference was not statistically significant. The assessment covered concepts, code reading, and debugging after a short task. It did not measure long-term professional performance. (Study methods, Main results and limitations)
Put those two results together. The assisted group understood the library less well, and it didn't finish meaningfully faster. That's the worst of both: the skill cost without the speed payoff.
The source presentation requires care. The research page reports mean scores of 50 percent and 67 percent. The manuscript reports a 4.15-point difference on a 27-point quiz and describes it as 17 percent. Those statements don't produce one arithmetically interchangeable percentage measure, so I use the reported standardized effect rather than harmonizing the conflicting descriptions. The authors also identified interaction patterns associated with stronger or weaker quiz performance, but participants were not randomized to those patterns. That observational analysis does not prove that a particular prompting technique causes better learning. (Anthropic research summary, manuscript wording, Exploratory analysis)
This is vendor-affiliated research, and the authors' affiliations should stay visible. Its adverse result is informative precisely because the incentives ran the other way: a lab that sells assistance published a study showing assistance eroding comprehension. I respect that, and I'd like to see more of it from every company in this industry, including mine.
The study does not establish that all assistance, all technical work, or all tutoring produces the same effect. What it establishes, for this specific configuration, is the pattern Chapter 3 documented in high-school mathematics showing up in adult technical learning: successful assisted execution alongside measurably weaker unaided understanding, with no speed benefit to offset it. (Affiliations and limitations)
Competence erosion appears in two different populations
Source: PNAS · Shen and Tamkin preprint
Two populations, two designs, one warning. In the high-school mathematics experiment covered in Chapter 3, students given unrestricted assistance performed 17 percent worse than the control group on later exams taken without it. In the Trio experiment, adult programmers who learned with an assistant scored lower on an unaided quiz by d = 0.738. The populations differ, the tools differ, and the measures differ, so the two numbers can't be combined or compared directly. What they share is the direction: help during practice, less capability after.
Common misreading: "AI makes programmers worse." The study says that in this setup, learning a new library with this assistant for a short task led to lower immediate comprehension. It says nothing about experienced programmers using assistance on familiar work, and nothing about whether the gap persists.
For a new hire: the first year of a job is when you build the understanding you'll draw on for the next ten. If your employer hands you an assistant, use it, and also make yourself do some of the hard parts without it. You want to be the person who can debug the thing when the assistant is wrong.
For a manager: the Trio study and the customer-support study point the same practical direction. Measure what new people can do without the tool at regular intervals. If output is climbing and unaided capability is flat or falling, you're borrowing against your future bench.
4.7.2 The Maintenance Rule for Productivity Claims#
METR's February 2026 update revisited its earlier finding that experienced open-source developers took 19 percent longer with early-2025 tools. Later estimates pointed toward faster completion, approximately 18 percent less time for returning participants and 4 percent less for newly recruited participants. But their confidence intervals (-38% to +9% and -15% to +9%) included no effect, and the researchers described selection and measurement problems that prevented a reliable updated estimate. These results should not be presented as a confirmed reversal or a definitive current productivity uplift. (METR update)
The implication is a maintenance rule I propose for every productivity claim: attach the tool version, task, worker experience, observation period, and evaluation design to every estimate, and recheck after substantial product or workflow changes rather than treating a historical study as a permanent coefficient. A productivity number without a date is not information. It's a fossil.
I run infrastructure, and in infrastructure we recommission. A system that tested fine at startup gets tested again after a major change, because the old test no longer describes the new system. Productivity claims about AI tools deserve the same discipline. The tools change, and the workflows around them change with them. A study from last year describes last year's tool.
4.8 The Complete Education-to-Work Process#
The studies above anchor individual stages of a longer pathway. The complete process below is proposed, not proven. Each stage has a separate outcome and a decision gate, so success at one stage can't hide failure at the next.
| Stage | Proposed interaction and human responsibility | Outcome to measure | Decision gate |
|---|---|---|---|
| Pathway choice | Learner and advisor compare actual prerequisites, costs, completion requirements, and employment pathways; models may summarize linked records | Accuracy of understanding, accessible options, ability to explain trade-offs | No steering based only on a market forecast or tracker score |
| Access and enrollment | Administrative assistant explains institution-approved procedures; authorized staff resolve exceptions | Completed required steps, enrollment, unresolved cases, errors | Current source records and reliable human escalation |
| Foundational learning | Instructor defines curriculum; tutor provides bounded explanation and practice | Unassisted subject understanding and delayed retention | Gains must survive removal of assistance on the chosen assessment |
| Technical practice | Qualified instructor supervises simulation and hands-on work | Independently assessed practical quality, errors, transfer | No substitution of a simulation score for required field competence |
| Credential completion | Authorized institution evaluates defined requirements | Completion, assessment validity, time, total learner cost | Credential claim matches the issuer's actual standard |
| Work-based experience | Employer supplies supervised tasks and accountable mentoring | Competence, attendance, safe task performance, participant experience | Real placement and supervision, not a promotional partnership alone |
| Application and matching | Applicant improves truthful presentation and approves every submission | Interviews, offers, starts, fit, application time, misrepresentation | No fabricated qualifications or automatic consequential submissions |
| Onboarding | Worker receives assistance while supervisors review consequential outputs | Quality-adjusted output, escalation, independent troubleshooting | Productivity must not conceal deterioration in quality or competence |
| Progression | Employer and worker review training and job outcomes | Retention, earnings, hours, advancement, transferable skills | Follow-up includes noncompleters and adverse outcomes |
None of the reviewed studies evaluates this entire process as one connected treatment. The research anchors differ by stage: enrollment (Pounce), learning (the physics tutor and Let's Chat), practice (the welding studies), pathways (apprenticeship and Year Up), hiring (the resume study), workplace performance (customer support), and competence (Trio and METR). The gates exist precisely because the evidence is stage-specific. An institution that treats any single stage's evidence as evidence for the whole pathway has recreated, at organizational scale, the confusion between assisted performance and acquired skill that runs through this entire report.
Here's a hypothetical to show how the table works. A community partnership announces an "AI-powered pathway into technical careers." It cites the physics tutor for learning, the resume study for hiring, and the customer-support study for job performance. Run it through the gates. Did anyone measure unaided learning in this program's courses, with a delayed test? Did technical practice end in an independent assessment, or in a simulator score? Did the employer provide supervised work-based experience, or a logo on the brochure? Does the follow-up include people who dropped out? Each "no" is a stage where the borrowed evidence stops applying. The program might still be good. It just hasn't shown it yet.
4.9 For Students and Job Seekers: How to Use AI Without Losing the Skill#
Everything above was written for decision-makers. This section is for the 19-year-old at the kitchen table, and for anyone at the start of a working life. It isn't a new set of findings. It's what the chapter's findings imply when you're the one using the tool.
4.9.1 Decide Which Goal You're Serving#
Before you open an assistant, ask yourself the same question I'd ask a college pilot: am I trying to produce this piece of work, or am I trying to become someone who can produce it? Both are legitimate. A cover letter due tonight is a production task. A physics problem set in the course you need for your major is a learning task. The Trio experiment and the high-school mathematics experiment both show what happens when a learning task gets treated like a production task: the work gets done, and the understanding doesn't arrive.
4.9.2 Seven Habits the Evidence Supports#
- Try it first, then ask. Attempt the problem on your own before you get help. The baseline step in the college evaluation standard exists for a reason; you need to know what you can do alone.
- Use structured help over open-ended answers. The physics tutor that worked was built around guided problems and expert solutions, not a model handing over results. If your course offers a tutor designed by your instructors, prefer it to a general chatbot for coursework.
- Test yourself without the tool. Close the window and redo a problem from scratch. Then do it again a week later. If you can't, you practiced with a crutch, and now you know.
- Try a problem that looks different. Transfer is the real test. If you can only solve the version you practiced, you learned the example, not the idea.
- Keep your resume true. The resume study found that clearer writing helped hiring on one platform. It tested polish, not invention. Use assistance to say accurately what you did, and never to add what you didn't.
- Get the important answers from a person. For enrollment deadlines, the Pounce design is the model: procedural questions to the assistant, financial aid, eligibility, and exceptions to staff who can put it in writing.
- Never let a tool make a safety call. If you go into a trade, use AI to study, and learn to make the safety judgment yourself under a qualified supervisor. That's the skill employers in my world are paying for.
4.9.3 Choosing a Path With Clear Eyes#
When you compare a degree, a trade, and a job, the evidence in this chapter suggests some questions worth asking each option:
- A college: What does the institution measure about its AI tools? Does anyone check whether students can do the work without them?
- A training program or apprenticeship: What happens after the simulator? Who performs the independent practical assessment, and is there supervised field experience? What are the completion and placement records, including people who didn't finish?
- An intensive workforce program: Is it paid? The Year Up evaluation showed a real long-run earnings gain along with a first-year earnings dip, so plan for the dip.
- A first job: If they hand you an assistant, will anyone help you build skills you can use without it?
4.10 What AI in Higher Education Research Permits You to Claim#
The selected postsecondary evidence is concentrated in a small number of subjects and work settings. It should not be described as a comprehensive evaluation of higher education, every skilled trade, or every form of early employment. Unanswered questions include persistence through graduation, independent writing and research across semesters, accommodations and outcomes for learners with disabilities, differences by language and access, effects in community-college technical programs, transfer across occupations, and full delivery costs. These define research needs, not presumed benefits or harms.
The defensible public claim is this: specific forms of intelligent support can improve particular educational and employment-related outcomes under documented conditions, while other configurations produce uncertain results or weaken independent learning.
The stronger claim, that a new facility produces a better local education-to-employment pathway, remains unestablished here. It would require a specified local program, actual access, employer participation, independent skill evidence, placement and retention records, and a separate accounting of the facility's community effects. Chapter 10 takes up those community effects directly.
| Claim | Status |
|---|---|
| Describe the studied outcome, population, comparison, date, uncertainty, and implementation conditions | Permitted |
| Convert short-term learning into lifetime earnings | Not permitted |
| Convert applicant hiring on one platform into aggregate job creation | Not permitted |
| Convert announced infrastructure investment into verified local placements | Not permitted |
| Use the seven Savrn trackers to investigate the physical and institutional context of potential training pathways | Permitted |
| Describe the trackers themselves as proven educational interventions | Not permitted |
| Propose an accountable research partnership | Permitted |
| Describe such a partnership as already operating without documentation | Not permitted |
That table applies to Savrn as much as to anyone. We build around behind-the-meter power and a closed-loop cooling design with a zero-makeup-water goal, and those are publisher statements until someone can check them. The same goes for any workforce benefit we might someday describe. A tracker entry starts a question. It never ends one. If we propose a training partnership, you should expect a named program, named employers, independent assessment, and placement records before we call it a result.
4.11 Conclusion: Specific Help, Measured Capability#
The research favors specificity over either enthusiasm or rejection. The useful questions are what the assistance does, for whom, against which alternative, over what period, and with what effect on the person's ability to perform without it.
The transition from education into work carries two distinct obligations that no tool can merge: help people complete real tasks, and maintain the competence needed to judge the result. The enrollment chatbot, the physics tutor, the resume tool, and the customer-support assistant all show that the first obligation can be met under the right conditions. The Trio experiment and the high-school mathematics result show that the second can quietly fail at the same time. The welding studies show how little we know yet about the trades. The apprenticeship and Year Up evaluations show what credible, long-run workforce evidence looks like, and none of it depends on a model.
The seven Savrn trackers can support source literacy and investigation of infrastructure-related opportunities, but the proof of educational and workforce value has to come from measured human outcomes. The pathway from a facility announcement to a young adult's paycheck passes through too many unproven stages to be walked in one claim. I'd rather walk it one gate at a time, and show you each one.
4.12 Frequently Asked Questions#
What does AI in higher education research actually show?
It shows specific wins under specific conditions. A Georgia State enrollment chatbot raised enrollment among committed students by 3.3 points, and a purpose-built Harvard physics tutor beat an active-learning class by 0.63 standard deviations on immediate tests. But the chatbot's full-sample effect was not significant, the tutor study had no long-term measure, and a separate experiment found weaker unaided understanding after AI-assisted learning.
Did an AI tutor really beat a Harvard physics class?
A custom GPT-4 tutor built with expert-written solutions and structured prompts outperformed in-person active-learning sessions by 0.63 standard deviations on immediate posttests across two lessons, with 194 eligible participants. The study did not measure long-term retention or course completion, did not test a generic chatbot, and required substantial expert preparation, so it does not show lower teaching costs.
Does using AI to write a resume help you get hired?
In a randomized experiment on one online labor platform, resume writing suggestions raised first-month hiring by about 0.247 percentage points on a 3.093 percent base, roughly 8 percent relative, in an analysis of 194,701 profiles. That is not an 8-point jump. The tool improved presentation of real experience; it says nothing about fabricated credentials, which should never be used.
Can VR or AR replace hands-on skilled trades training?
The evidence does not show that. A virtual welding study of 101 volunteers found no significant difference among groups (p = .089), which is not proof of equivalence, and cues and audio were impaired. A 24-person AR study improved unassisted torch motion (p = .032) but did not test weld integrity. Neither addresses electrical, high-voltage, or other infrastructure trades.
Do apprenticeships increase earnings?
A September 2025 Department of Labor evaluation estimated that registered apprentices had employment 7.8 and 8.8 points higher, and quarterly earnings $3,230 and $4,693 higher, than two comparison groups at the ninth quarter. It was not randomized, many apprentices were already employed, and selection could not be fully ruled out. One report chapter lists $3,320 instead of $3,230.
How much does Year Up increase earnings?
A randomized evaluation of 2,544 young adults aged 18 to 24 found quarterly earnings $1,895 higher in quarters 23 and 24, about 28 percent above the $6,901 comparison average (p < .001). Earnings were lower in the first year while participants trained, the study observed about seven years, and it does not identify which program component caused the gain.
Can using AI while learning make you worse at a skill?
In one randomized experiment, 52 programmers learning a new Python library with a GPT-4o-based assistant scored lower on an unaided quiz than those without it (d = 0.738, p = .010), with no significant time savings. It was a small, short study from vendor-affiliated authors, and it does not show the same effect for all tasks or experienced workers.
Does AI make new workers more productive?
In one company, a conversational assistant raised customer issues resolved per hour by about 15 percent among 5,172 support agents, with larger gains for newer and less-skilled workers. It was not randomized, the most-skilled agents saw limited benefit and some quality decline, and the result does not translate into wages, headcount, or other occupations.
