Microdata, can resolve the young cohort
Stanford “Canaries” (ADP payroll): −16% relative employment for ages 22–25 in the most AI-exposed, automate-type occupations; older workers stable. Authors call it “consistent with,” not proven.4
AI covers the codifiable majority of work, while the residual it handles worst, tacit judgment, is built by doing the entry-level tasks AI absorbs first. A firm that automates its bottom rung saves cost this quarter and may be dismantling the process that produces its senior bench.
A compiled, source-verified research digest — every claim cites a downloaded source, every figure is drawn from the data behind it. Not a personal essay.
Generative AI covers the part of work that can be written down, which McKinsey estimates absorbs 60 to 70 percent of employees’ time today.9 The residual is tacit: judgment, common sense, accountability, and the know-how Michael Polanyi compressed into “we know more than we can tell.”7 That residual has never been taught directly. Workers build it by grinding through the codifiable majority, which is why Arrow could show in 1962 that “learning is the product of experience,“8 and why a 2026 general-equilibrium model treats entry-level tasks as the curriculum of a career and warns that automating them can “narrow the pipeline through which workers acquire skill,” tipping an economy into a human-capital trap.3
The evidence on whether the trap is running is early and split. When AI helps, it helps novices most. In a study of 5,172 customer-support agents, access to an AI assistant raised output 15% on average and 30% for the least experienced, who reached in two months a competence that had taken six.1 On the labor-market side, a Stanford/ADP study finds a 16% relative employment decline for workers aged 22–25 in the most AI-exposed occupations, concentrated where AI automates rather than augments.4 Two counterweights cut the other way. A precise-null Danish study replicates the early-career decline but finds it is not driven by firms adopting AI chatbots; the Yale Budget Lab detects no significant effect in survey data it concedes is underpowered for exactly this cohort.612 The strongest displacement prediction, for its part, comes from an interested party.16
The technically automatable share of work is large, and it grew when generative AI arrived. McKinsey’s 2023 analysis decomposed some 850 occupations into detailed work activities and concluded that generative AI and other technologies could automate work activities absorbing 60 to 70 percent of employees’ time today, up from a prior estimate of about half, with the jump driven by generative AI’s command of natural language.9 A 2025 update restates the figure as roughly 57% of U.S. work hours and repeats the boundary that matters: the figures measure technical potential in tasks, “not the inevitable loss of jobs.”15 This is the 80 of the paper’s colloquial 80/20. Most of what most jobs consist of, measured by time, is in principle reachable.
The 20 is what remains when the codifiable is stripped out. The corpus does not quantify this residual, and the 80/20 label is colloquial rather than measured. What it names is the split the rest of the argument turns on. AI reaches the codifiable majority; human value concentrates in the tacit remainder, in judgment, accountability for outcomes, and the handling of situations nobody has scripted. David Autor explains why that remainder resists automation. Many tasks are ones “people understand tacitly and accomplish effortlessly but for which neither computer programmers nor anyone else can enunciate the explicit ‘rules’ or procedures.” This is Polanyi’s paradox, named for the philosopher who observed in 1966 that “we know more than we can tell.”7 The tasks “that have proved most vexing to automate,” Autor writes, “are those demanding flexibility, judgment, and common sense.”7
Behavioral usage data runs in the same direction. The Anthropic Economic Index mapped roughly a million Claude.ai conversations onto the U.S. Department of Labor’s O*NET task taxonomy and found use leaning toward augmentation (57%) over automation (43%).10 Only about 4% of occupations use AI across at least three-quarters of their tasks, and roughly 36% for at least a quarter.10 It is a first-party report on a non-representative sample. The data covers Claude’s consumer Free and Pro tiers only, coding is overrepresented at 37.2% of queries, and the authors concede that some of what looks like automation in their data may in fact be augmentation.10 Within those limits, the picture is of AI diffused across the many tasks in the economy, with stronger impact on some groups of tasks than others, and the authors read their own index as pointing to a future in which most current jobs evolve rather than disappear.10 Autor’s warning covers both this and the McKinsey figures. Commentators, he argues, tend to overstate machine substitution for human labor and ignore the strong complementarities, and a substitution-only reading misses a central economic mechanism, the way automation raises the value of the tasks that workers uniquely supply.7
The residual has a structure, and Polanyi named it. Much of expert performance is knowable in the doing without being articulable as rules. “When we break an egg over the edge of a mixing bowl, identify a distinct species of birds based on a fleeting glimpse, write a persuasive paragraph, or develop a hypothesis to explain a poorly understood phenomenon, we are engaging in tasks that we only tacitly understand how to perform.”7 Because the rules cannot be written down, classical rule-based automation could not capture them.
Machine learning loosens the constraint by inferring the mapping from examples instead of explicit rules. That is how the support-agent system studied by Brynjolfsson, Li and Raymond could “distinguish successful behaviors of the top performers, including those they tacitly apply” and pass them down the experience curve; the authors describe these methods as enabling computers “to perform nonroutine tasks that rely on tacit knowledge and experience.”1 The loosening is partial and conditional. It needs abundant examples and clear feedback, and Autor argues that neither machine learning nor environmental control fully resolves Polanyi’s paradox for genuinely judgment-intensive work.7 It also carries a dependency. A model that imitates tacit behavior learns it from people who already possess the skill, so the model is downstream of the expertise it distills. If the formation of that expertise stops, the well the next model draws from stops refilling.
How the expertise gets formed has one of the most durable answers in economics. Kenneth Arrow’s 1962 paper states the generalization as one “so clear that all schools of thought must accept it”: “Learning is the product of experience. Learning can only take place through the attempt to solve a problem and therefore only takes place during activity.”8 He anchored it in production data. Labor-hours per airframe fell as a power function of cumulative output, and the Horndal iron works raised output per man-hour about 2% a year for fifteen years with no new investment, a gain “which can only be imputed to learning from experience.”8
Arrow added a second regularity, and for the AI question it is the load-bearing one. Learning from “repetition of essentially the same problem is subject to sharply diminishing returns… the stimulus situations must themselves be steadily evolving rather than merely repeating.”8 Expertise grows on a frontier of progressively harder problems. Apprenticeship, on this reading, is an economic arrangement for supplying stimulus situations at the right difficulty, and entry-level work is where the supply has always come from. A 2026 model by Afrouzi, Blanco, Drenik and Hurst, a Federal Reserve and academic team, makes the application to careers explicit. Entry-level tasks, they write, “are not merely low-value work — they are the curriculum through which workers accumulate the human capital that makes them productive later in their careers.”3
In the frame this paper opened with, the 20 is manufactured by doing the 80.
The trap follows from setting Arrow’s regularity next to what AI absorbs first. In the Afrouzi model, a fall in the price of AI does two things at once inside a single occupation. It substitutes for workers on entry-level tasks, and it complements the senior workers who create the new tasks the occupation performs.3 The substitution side can narrow the pipeline through which workers acquire skill; the complement side can help managers expand the task frontier and give juniors more to learn.3 The technology does not decide which side wins. Economies with high learning capacity admit paired stationary equilibria, the economy cannot smoothly move between them, and if coordination falls on the zero-learning branch, entry-level work is fully automated. That branch is the model’s human-capital trap.3
The trap needs no villain. The value of an entry-level task to future human-capital formation is not internalized by atomistic workers or firms, and when the skill is general rather than firm-specific, no single firm captures the return on training it.3 Each firm’s automation is individually rational, and the thinning of the senior pipeline lands on everyone. The model’s first-best correction operates at the level of a whole economy, a tax on automation profits paired with a subsidy on frontier-maintenance spending.3 The plain-language version of the risk, from the companion Fortune piece, is that firms automating entry roles risk hollowing out the pipeline of competent senior workers they may need in future, taking the short-term cost saving at the price of long-term stability.14
If the trap is running anywhere, the entry rung is where the machinery should be visible, and the best-identified evidence in the corpus sits on that rung. Brynjolfsson, Li and Raymond studied the staggered rollout of a generative-AI assistant across 5,172 customer-support agents and found productivity gains of 15% on average and 30% for the least skilled and least experienced, while the most experienced gained little and showed a small decline in conversation quality.1 Treated agents with two months of tenure performed as well as untreated agents with more than six months, and low-skill agents began communicating more like high-skill agents.1 The authors bound the finding’s scope themselves. These are medium-run effects in a single firm and a single occupation; the article “is not designed to shed light on the aggregate employment or wage effects of generative AI tools,” and in the longer run, the authors note, firms might respond by hiring more novices or by building systems “that replace labor altogether.”1 The equilibrium response, in the authors’ own telling, could land on either branch of the duality above.
Read plainly, the finding cuts against the trap thesis. The tool moved novices down the experience curve faster, and the authors report durable learning, with gains persisting through software outages when the AI was unavailable.1 That looks like AI accelerating skill formation at the entry level. The strongest direct evidence on the other side is thinner. Shen and Tamkin (2026), cited inside the Afrouzi model as preliminary support for its mechanism, randomly assigned workers to use AI and found it “reduced their accumulation of skills needed for career progression.”3 What separates the two findings is what got measured. The support-agent study measures output, issues resolved per hour and the communication patterns of high performers, and the outage evidence shows that some of the gain lodges in the worker rather than in the tool. Whether the lodged gain is the veteran’s judgment or the veteran’s phrasing is what output metrics cannot distinguish, and the skill-erosion result measures progression-relevant skill directly. The two findings are not yet reconciled.
A field experiment at BCG shows where the leveling stops. Dell’Acqua, Mollick and colleagues randomized 758 consultants to GPT-4 access and found that on 18 realistic consulting tasks inside the AI’s “jagged frontier,” the assisted group finished more tasks, faster, at markedly higher quality, and the skill gap compressed, with below-average performers gaining 43% against their own baseline versus 17% for the above-average.5 On a task chosen to sit outside the frontier, assisted consultants were 19 percentage points less likely to reach a correct answer, because AI tends to “produce incorrect, but plausible” output that people over-trust.5 Outside the frontier, what saves the human is tacit judgment about when a confident answer is wrong. That is the residual capability, and it takes experience to build.
The Stanford Digital Economy Lab’s Canaries in the Coal Mine? is the closest thing to a direct look at the rung, built on high-frequency administrative payroll data from ADP. Early-career workers aged 22–25 in the most AI-exposed occupations experienced a 16% relative employment decline after conditioning on firm-level shocks, while employment for experienced workers in the same occupations stayed stable or grew.4 The decline concentrates in applications of AI that automate work rather than those that most augment it, which is the split the trap model predicts should matter. Adjustment runs through employment rather than wages, suggesting possible wage stickiness, and the pattern survives excluding tech firms and remotable occupations and does not appear in pre-2022 data, including the COVID shock.4 One detail complicates the tidy young-versus-old divergence. In occupations with a low share of college graduates, the authors find “experience may serve as less of a buffer,” with divergent employment outcomes by AI exposure up to age 40.4
Broader observational data points the same way without isolating AI as the cause. SignalFire’s scan of more than 650 million professional profiles has new graduates at just 7% of big-tech hires in 2024, with new-grad hiring “down 25% from 2023 and over 50% from pre-pandemic levels in 2019,“11 and CNBC reports unemployment for labor-market new entrants at a “nine-year peak.”13
The Stanford team states the causal status plainly: “While we explore a variety of alternative explanations, we caution that the facts we document may in part be influenced by factors other than generative AI… our results are consistent with the hypothesis that generative AI has begun to affect entry-level employment.”4 The authors’ phrase is “consistent with,” which stops short of a causal claim. The SignalFire and CNBC data are observational and do not isolate AI from the post-2022 tech contraction, the rate environment, or an oversupply of graduates; CNBC’s own sources attribute part of the effect to “the rising share of young Americans obtaining four-year degrees.”1311
Three different negative figures circulate for the same phenomenon. The −16% is the firm-time-conditioned estimate (a 15 log-point decline); the −6% is the unconditioned change from late 2022 to September 2025; the −13% is the figure CNBC cites from the earlier (August 2025) preprint. They differ by controls and paper version, and this paper cites each to the source that states it.413
The strongest counterweight comes from Denmark, where Humlum and Vestergaard can do something the American studies cannot. They link large-scale adoption surveys to administrative records for the whole Danish labor market, which lets them split employment trends by whether firms actually adopted AI. Two years after ChatGPT’s launch, they estimate precise null effects on earnings and recorded hours at both the worker and workplace levels, ruling out effects larger than 2%, across 11 highly exposed occupations with no significant effects in any of them.6 On the entry rung specifically, they replicate the pattern of declining early-career employment in exposed occupations in Denmark and find “the declines are not driven by firms adopting AI chatbots.”6
The Yale Budget Lab reaches a similarly cautious null on U.S. survey data, and it is candid about what its data cannot see. The Current Population Survey is somewhat underpowered for subgroup analysis of the 22–27-year-old recent college graduates at the center of the dispute, and the lab notes that occupation-level exposure metrics could hide effects if half of the exposed occupations expand while the other half contract.12 Its most recent quarter shows a roughly half-percentage-point rise in unemployment for exposed occupations, larger for younger workers, both statistically insignificant as of the first quarter of 2026, and its summary allows that effects “may yet become evident in 2026 or 2027.”12
↑ Finds an entry-level effect
Stanford “Canaries” (ADP payroll): −16% relative employment for ages 22–25 in the most AI-exposed, automate-type occupations; older workers stable. Authors call it “consistent with,” not proven.4
SignalFire / CNBC: new-grad hiring down >50% at big tech since 2019; new-entrant unemployment at a nine-year peak. Confounded by the tech downturn and degree oversupply.1113
Humlum–Vestergaard (Denmark): precise null on earnings/hours (≤2% bound). Replicates the early-career decline but finds it is not driven by firms adopting chatbots.6
Yale Budget Lab (CPS / SDID): no significant employment or wage effect yet; explicitly notes its data is too underpowered to resolve the 22–27 subgroup. “May yet become evident in 2026 or 2027.”12
Finer microdata ←——— data granularity ———→ Broad survey aggregates
Read together, the studies locate the disagreement rather than cancelling out. The ones that find an entry-level effect use the finest data, ADP payroll records that can resolve the 22–25 cohort; the nulls use aggregates that by their own admission may be too coarse to catch a narrow early-career effect, or measure a different country and a different outcome. Both of this paper’s empirical pillars, the novice lift and the entry-rung decline, share Brynjolfsson as lead author, and the strongest counterweight directly contests the second. The most useful reconciliation is the Danish one: the early-career decline is real, but it may not be cleanly attributable to firm-level AI adoption, which would mean the labor market is reorganizing for reasons that AI is entangled with without being the sole driver. That reading is more defensible than either “AI is gutting entry-level jobs” or “nothing is happening.”
The same Danish data also holds the corpus’s only field sighting of the good equilibrium. Humlum and Vestergaard find employers absorbing AI through task reorganization, including “new tasks in content generation, AI oversight, and AI integration,” and adopters moving into higher-paying occupations where AI chatbots are more relevant, with switchers’ earnings growing 12 percentage points more, though still too few to move average earnings.6 That is the frontier-expansion branch of the trap model showing up in the field. The technology creates new tasks near the frontier, and the workers who engage with it climb. The dataset that undercuts the displacement story is the same one showing the high-learning mechanism running.
The “white-collar bloodbath” prediction. Dario Amodei’s May-2025 claim that AI “could wipe out roughly 50 percent of all entry-level white-collar jobs within five years,” pushing unemployment to 10–20%, is cited here as a named prediction from an interested party, Anthropic’s CEO. This corpus does not treat it as evidence. The original Axios interview could not be retrieved directly (HTTP 403); the wording is quoted verbatim inside the Stanford “Canaries” paper and corroborated across secondary reports.416 Amodei reportedly softened this framing in 2026, and it should not be read as a forecast this corpus endorses.
”Captures 80% of my style but not my soul.” This evocative line is sometimes attached to discussions of generative AI and creative work. No source in this topic’s corpus contains it, and no verifiable origin was found. It is recorded here as an unverified anecdote / aphorism and is not used as evidence anywhere in this paper. The substantive, sourced version of the same idea is Autor’s Polanyi point, the tacit residual that resists codification,7 and the jagged-frontier finding that AI produces “incorrect, but plausible” outputs outside its frontier.5
For a single firm the argument lands as a choice about deployment. AI is strong on the codifiable majority and most immediately useful at the bottom of the experience curve, where it lifts a novice to near-veteran output on in-frontier tasks.15 The tasks it most readily absorbs are the tasks through which the tacit residual has always been built.38 Automating them is individually rational, and the cost, a thinner senior bench years out, is a learning externality the firm does not fully bear, which is how the equilibrium can settle on the bad branch without anyone choosing it.3
The deployment choice maps onto the model’s two branches. The Stanford data shows entry-level declines concentrated where AI automates and growth where it augments,4 so deploying AI to expand what juniors can attempt, rather than to remove the attempt, is the nearest firm-level analogue of the model’s frontier maintenance. That translation is this paper’s inference, since the model’s own first-best is an economy-level tax-and-subsidy pair, and it says nothing about a single firm’s deployment policy.3 What a redesigned junior role looks like is less settled than the principle, and the corpus offers two anchors. Arrow’s condition is that the “stimulus situations must themselves be steadily evolving,“8 which for a junior role built around AI means the role must keep supplying harder problems, with supervising, correcting, and judging AI output replacing drafting from scratch as the evolving stimulus. And the BCG experiment observed two workable patterns of human-AI integration, “Centaurs” who divide and delegate subtasks between themselves and the AI, and “Cyborgs” who integrate the tool into their whole task flow.5 Those patterns are the closest thing in the evidence to concrete content for an apprenticeship built on AI supervision.
Two cautions bound all of this. The strongest displacement evidence is correlational and contested by precise nulls, so betting a workforce strategy on either extreme is unjustified.4612 And the trap is a stationary-equilibrium result; nothing in this corpus estimates how long a pipeline takes to visibly thin, which is the number a planning horizon needs and does not yet have. The defensible posture is to assume the tacit residual is real and durable, that it is built by doing, and that a firm automating away its own apprenticeship is optimizing a number it can see against one it cannot — until the senior bench it needed fails to exist.
The captured evidence does not settle several things. (1) Reinstatement vs starvation: whether AI complementing senior workers expands the task frontier fast enough to refresh the junior curriculum is, in the model, an equilibrium that could go either way. The Danish reorganization evidence is one field sighting of the good branch, and there is no empirical estimate yet of which branch whole economies are on.36 (2) Durability of the novice lift: the body argues an output-versus-judgment distinction, but the underlying findings remain unreconciled: novices reach veteran output with AI, and whether they internalize veteran judgment is what the cited Shen–Tamkin skill-erosion finding puts in doubt.13 (3) Will Polanyi’s paradox hold? Autor argues machine learning only partially circumvents the tacit-knowledge constraint; how far frontier models push that boundary is unresolved and is the subject of this corpus’s frontier-capability paper.7