Explicit knowledge · High cost of error
Augment with verification. AI can draft, but errors are expensive — keep human review and accountability. E.g. radiology scan-read with physician final sign-off.15
A job is a bundle of tasks. AI takes them up one at a time, and what happens to the job depends on the tasks it leaves behind, which is why occupation-level predictions of mass displacement keep missing.
A compiled, source-verified research digest — every claim cites a downloaded source, every figure is drawn from the data behind it. Not a personal essay.
Labor economics treats an occupation as a bundle of distinct tasks. A technology substitutes for the tasks it can do cheaply and complements the ones it cannot, and the fate of the job is the weighted result of both forces.1 Shifting the unit of analysis from the job to the task moves the headline numbers by a factor of five. A model that classified whole occupations put 47% of US employment at high risk of computerization;9 a task-based re-estimate of the same economies put the highly automatable share at roughly 9%.10 The reason is that automatability varies more within an occupation than across occupations. The Suitability-for-Machine-Learning research found few if any jobs in which every task suits the machine, so capturing the value means redesigning the job around the tasks that do.45 This paper opens with the sharpest field test of the job-level frame on record, Geoffrey Hinton’s 2016 call to stop training radiologists,15 and what has happened to that profession since.1416 It then works through the mechanics that produced the outcome — complementarity, displacement against reinstatement, and induced demand — and closes with a task-level decision rule, cost of error against knowledge type, for where to deploy AI and where to keep a human in the loop.
In 2016, onstage at a machine-learning conference in Toronto, Geoffrey Hinton told the audience it was “completely obvious” that within five years, ten at the most, AI would outperform humans at radiology tasks, and that “people should stop training radiologists now.”15 He told radiologists they were “like the coyote that’s already over the edge of the cliff but hasn’t yet looked down.”15 The prediction became the most-cited forecast of AI-driven job loss, and part of it held. The technology arrived, and hospitals adopted it at scale. What failed was the unit of analysis.
A decade on, the profession Hinton wrote off has grown, and the clearest case is the institution that adopted the technology hardest. Mayo Clinic built a 40-person AI team, licensed or developed more than 250 AI models, and over the same stretch expanded its radiology workforce by about 55%.14 The pattern holds beyond Mayo. The number of active US radiologists rose roughly 10% over ten years, into a documented shortage whose own workforce projections run to 2038, and pay has moved with the scarcity; Figure 1 carries the indicators.1516 Hinton himself has since narrowed the claim, saying he was speaking about image analysis and that radiologists would work alongside AI.15 The record of this reversal is press reporting (a Fortune feature, a TechCrunch report, an ACR bulletin), and its load-bearing figures were corroborated against independent outlets during verification.141516
The explanation is in the task list. Reading images is one of many things a radiologist does; the same person consults with physicians, monitors patients, and performs procedures. Automate the reading and the freed time flows to the rest. As the economist Christoph Herpfer put it in the Fortune reporting, “Even if you can automate one or two of those, you just expand the time you spend on the other tasks… the job itself won’t go away.”15 Two structural facts hold the arrangement in place. Medicare and Medicaid reimburse a study only when a licensed physician performs the final read, and accountability for a missed diagnosis cannot rest with a model.15 Both raise the effective cost of an error, and cost of error, as the decision frame at the end of this paper argues, is one of the two dimensions that decide where a human stays in the loop.
The captured corpus traces Hinton’s prediction to a 2016 Toronto machine-learning event, with the wording reproduced in Fortune (May 2026) and corroborated by TechCrunch (May 2025).1514 The original talk recording was not independently retrieved; the wording is cited to these reporting sources, not a primary transcript. The best-attested verbatim line is “completely obvious that within five years deep learning is going to do better than radiologists… we’ve got plenty of radiologists already”; “stop training radiologists now” is the widely reproduced reporting paraphrase, and the “coyote” sentence is sometimes rendered as a reporter’s characterization rather than a verbatim quote.
The radiology miss has a general form, and it was measured before Hinton spoke. In 2013 Carl Frey and Michael Osborne scored US occupations for their susceptibility to computerization and concluded that 47% of total US employment sat in the high-risk category.9 Their method was more careful than the headline suggests. They assessed occupations against three engineering bottlenecks — perception and manipulation, creative intelligence, and social intelligence — and assigned each occupation a single probability of being computerized.9 They also stated the limits themselves. “We make no attempt to estimate how many jobs will actually be automated,” they wrote, and because each probability describes an occupation being fully automated, their estimates “do not capture any within-occupation variation” from automating tasks that simply free a worker’s time for other tasks.9 So the authors of the 47% figure conceded the within-job point in the same paper. The figure traveled without the concession.
The OECD re-estimate made the correction quantitative. Arntz, Gregory and Zierahn re-ran the exercise accounting for the mix of tasks inside each occupation and found that roughly 9% of jobs across 21 OECD countries are highly automatable, with the US figure also near 9%.10 The same economies, measured at two units, five times apart. Both numbers date from 2013–2016 and concern computerization as it stood before large language models, so the dispute itself is a decade old; the exposure and usage measures later in this paper rerun it on current technology.
The reason a job outlasts its automated tasks starts with complementarity. David Autor’s 2015 essay observed that journalists and even expert commentators “tend to overstate the extent of machine substitution for human labor and ignore the strong complementarities between automation and labor that increase productivity, raise earnings, and augment demand for labor.”1 The mechanism behind those complementarities is Kremer’s O-ring production function. When production is a chain of tasks and any weak link can sink the output, making one link cheaper and more reliable raises the value of the human links that remain.1
The canonical case is the ATM. Bank-teller numbers were widely expected to collapse when the machines spread. Instead, US teller employment rose modestly from roughly 500,000 to 550,000 between 1980 and 2010, even as ATMs quadrupled from about 100,000 to 400,000 over 1995–2010.1 The machine cut the cost of running a branch, so banks opened more branches — urban branch counts rose more than 40% — and the teller’s work shifted from handling cash toward relationship banking.1 The task went to the machine. The job was redefined and persisted.
Tasks that cannot be substituted by automation are generally complemented by it… productivity improvements in one set of tasks almost necessarily increase the economic value of the remaining tasks. David H. Autor, “Why Are There Still So Many Jobs?”, Journal of Economic Perspectives, 2015
Complementarity comes with two caveats Autor states plainly. Higher value in the remaining tasks does not guarantee the worker captures it; if those tasks are ones almost anyone can supply, elastic labor supply keeps wages flat.1 And the pattern of who is touched shifts with the technology. Autor documents the polarization of the computer era, with employment growing at the top and bottom of the skill distribution while routine middle-skill work hollowed out, and judged in 2015 that polarization was “unlikely to continue very far into the foreseeable future.”1 The exposure measures covered below bear that judgment out: the AI wave reaches highest-paid analytic work first, the opposite end of the distribution from the routine middle.
Acemoglu and Restrepo turned complementarity into a formal engine with three named forces acting on what they call the task content of production. Automation lets capital take over tasks labor used to perform, a displacement effect that on its own lowers labor demand and the labor share.3 The same automation raises productivity, which makes output cheaper, expands demand, and can offset the displacement.3 And the economy creates new tasks in which labor holds comparative advantage, a reinstatement effect that pulls workers back into a widening range of work.3 In “The Race Between Machine and Man” the long-run condition is explicit: employment and wages grow sustainably only when reinstatement keeps pace with automation.2
The reinstatement effect has now been measured over a very long baseline. Autor, Chin, Salomons and Seegmiller traced US job titles from 1940 to 2018 and found that roughly 60% of 2018 employment sits in job titles that did not exist in 1940; among professionals, the share is 74%.8 Most of the work Americans do today had no name in 1940. The same paper carries the caution against extrapolating that comfort forward. The authors’ results suggest that “the demand-eroding effects of automation innovations have intensified in the last four decades while the demand-increasing effects of augmentation innovations have not.”8 The historical offset is real, and on this evidence its margin appears to have been narrowing for forty years. The ATM story is solid history; whether it extends forward depends on a balance that has been tilting toward displacement.
How much productivity the current wave delivers is Acemoglu’s later question, and his warning concerns automation adopted purely for task-level cost savings. That “so-so” automation removes the worker while adding little output.11 His macro arithmetic bounds the upside. Applying a version of Hulten’s theorem, where the GDP gain equals the share of tasks affected times the average cost saving per task, and taking average task-level labor cost savings of about 27%, he puts AI’s total-factor-productivity gain at no more than about 0.66% over ten years.11 McKinsey ran task accounting of the same family and reached a headline in the trillions. Decomposing occupations into detailed work activities, its 2023 study put generative AI’s potential at $2.6–4.4 trillion in annual value and judged that current technology could automate activities absorbing 60 to 70 percent of employees’ time.23 The two results measure different things, and the assumptions name the gap. Acemoglu counts what AI can profitably do at current capability within a decade. McKinsey counts technical potential at full implementation, and its own adoption scenarios stretch the realization out, with half of today’s work activities automated somewhere between 2030 and 2060 and a midpoint of 2045.23 They can both be right; they answer different questions on different clocks.
If the task is the unit, the working question is which tasks the current technology suits. Brynjolfsson and Mitchell answered it for machine learning in Science (2017) with a Suitability-for-Machine-Learning (SML) rubric, eight key criteria backed by a 21-item scoring instrument.4 A task suits ML when it maps well-defined inputs to well-defined outputs, when large labeled datasets and clear feedback exist, when some error is tolerable, when the underlying function is stable over time, and when it demands no long chains of common-sense reasoning, no detailed explanation of the decision, and no fine physical dexterity.4 Behind the rubric sits Polanyi’s paradox, “we know more than we can tell.” Tasks built on tacit knowledge humans cannot articulate (recognizing a face, comforting a patient) resisted earlier automation because nobody could write the rules; ML partially circumvents the paradox by inferring the input–output mapping from data, but only where the data and the feedback exist.4
The empirical companion paper scored 18,156 O*NET tasks across 964 occupations against the rubric, and the finding that matters most here concerns variance. The within-occupation standard deviation of task suitability came to about 17% of the mean, meaning the automatable and the resistant tasks sit side by side inside the same job.5 Most occupations in most industries contain at least some suitable tasks, and few if any occupations are suitable throughout.5 The authors’ recommendation followed: shift the debate “away from the common focus on full automation of many jobs… toward the redesign of jobs and reengineering of business processes.”5 They also found suitability nearly uncorrelated with wages (about −0.14 against the wage percentile), a first hint that this wave would land on a different slice of the workforce than the routine-biased automation before it.5
A good ML task maps clear inputs to clear outputs, has abundant labeled data, tolerates some error, stays stable over time, and needs no long reasoning chains, no explanation of its decisions, and no fine motor skill. A bad ML task runs on tacit judgment, sparse feedback, high stakes, shifting conditions, and accountability requirements. Most jobs are a mix of both.4
Two exposure measures then mapped which slice. Felten, Raj and Seamans built the AI Occupational Exposure (AIOE) index bottom-up, linking AI application areas to 52 human abilities in O*NET and on through tasks to occupations, and kept the measure deliberately neutral on whether AI substitutes for labor or complements it.6 The most-exposed occupations turn out to be higher-paid, higher-education analytic roles — genetic counselors, financial examiners, actuaries, accountants — while the least-exposed are physical and interpersonal trades like roofers, fitness trainers, and dancers.6 AIOE is itself an occupation-level score, a league table of the very unit this paper argues against. The measure survives the objection because exposure and displacement are different questions. The index ranks where the technology can reach, and it is silent by construction on what happens to the people it reaches.6
Eloundou, Manning, Mishkin and Rock (“GPTs are GPTs,” 2023) applied the same logic to large language models and found the exposure broad. Around 80% of the US workforce could have at least 10% of their tasks affected by LLMs, and roughly 19% could see half or more affected.7 The leverage sits in the tooling. LLMs alone could meaningfully speed up about 15% of tasks; software built on top of them raises the share to between 47% and 56%.7 Exposure again rose with income, and the authors read the pattern as the signature of a general-purpose technology.7 Their headline number aggregates task exposure back up to workers, so the same unit caveat applies, and “exposed” means technical potential; the study is explicit that it does not distinguish labor-augmenting from labor-displacing effects.7
Even a task fully handed to the machine need not shrink the total work around it, because cheaper output draws more demand. The principle is old. In The Coal Question (1865), William Stanley Jevons observed that “it is wholly a confusion of ideas to suppose that the economical use of fuel is equivalent to a diminished consumption. The very contrary is the truth.”18 Efficient engines lowered coal’s effective price, and Britain burned more of it. Jevons drew the labor analogy himself: “The economy of labour effected by the introduction of new machinery throws labourers out of employment for the moment. But such is the increased demand for the cheapened products, that eventually the sphere of employment is greatly widened.”18 His example was the seamstress, who he argued had “perhaps in no case been injured, but have often gained” from the sewing machine — the 1865 version of the ATM teller.18
The modern name for the mechanism is induced demand, and its condition is elasticity. When a task gets cheaper, whether total spending and employment rise or fall depends on the price elasticity of demand for the output. Brynjolfsson and Mitchell count that elasticity among six economic factors standing between “this task is automatable” and any employment outcome, which is why automatability alone settles nothing.4 Radiology met the condition. US imaging volumes rose about 25% from 2018 to early 2025, so the automated share of the reading arrived into a caseload that was itself growing.15
The paradox is now invoked for AI directly. Satya Nadella argued in January 2025 that “as AI gets more efficient and accessible, we will see its use skyrocket, turning it into a commodity we just can’t get enough of.”20 The executives doing the hiring hold both halves of the question at once. In a McKinsey survey reported by Stanford’s AI Index, software-engineer headcount is “expected to increase, consistent with the Jevons Paradox,” while 31% of the same respondents expect AI to reduce overall workforce size and 19% expect an increase.21
Jevons describes a tendency, and the tendency has conditions. It holds when demand for the cheapened output is elastic and unsatisfied; it fails when demand is saturated or the externalities (energy, hardware toxicity) are not priced in — a caution economists raised when applying the 160-year-old framework to AI compute.20 The induced-demand effect on radiology is observed (imaging volumes rose about 25% from 2018 to early 2025),15 but whether it generalizes across white-collar tasks at AI’s current capability is not yet settled by the captured evidence.
The economics of the split comes from Agrawal, Gans and Goldfarb. Machine intelligence, in their account, is a drop in the cost of prediction, and cheaper prediction raises the value of the complementary human input they call judgment (determining what the payoffs of a decision actually are).13 Prediction and judgment stay complements as long as the judgment is not too hard to supply, so the human’s residual role concentrates in exactly the tasks that are high-stakes and hard to specify.13 That result gives the task-level decision its two natural axes.
The first axis is knowledge type. An explicit task can be codified — clear input–output pairs, feedback, rules a rubric can score — and lands in the SML sweet spot. A tacit task runs on Polanyi’s unarticulated knowledge, the judgment and context nobody can write down, where ML is weakest.4 The second axis is cost of error. ML systems are statistical and rarely reach 100% accuracy, so tolerance for error is an adoption constraint in its own right, and where a wrong answer is expensive, requirements for accountability and explanation, both of which ML handles poorly, keep a human in the loop.4
↑ HIGH cost of error
Augment with verification. AI can draft, but errors are expensive — keep human review and accountability. E.g. radiology scan-read with physician final sign-off.15
Keep human-led. Judgment, no clear feedback signal — ML’s weakest zone. AI assists at the margins only. E.g. comforting a patient, high-stakes negotiation, novel strategy.4
Automate. Clear inputs and outputs, abundant data, error-tolerant — the sweet spot for ML. Redesign the job around the freed-up time. E.g. classification, transcription, routine drafting.5
Augment and experiment. Hard for AI unaided, but cheap mistakes make it a low-risk place to pilot human+AI workflows. E.g. brainstorming, first-draft creative work.25
EXPLICIT (codifiable) ←— knowledge type —→ TACIT (Polanyi)
Score the radiologist’s bundle on these axes and the outcome of the opening section follows. Scan-reading is explicit and data-rich, so the machine drafts it; the final read carries legal accountability and a high error cost, so it sits in the augment-with-verification cell, where the reimbursement rule had already placed it.15 Consulting with physicians and comforting patients are tacit work with high stakes, and they stay human-led. The bundle rearranged, and the job persisted.
The augmentation cells also have measured payoffs. In field studies gathered by Stanford’s AI Index, AI assistance raised the productivity of the least-skilled customer-support agents by 34% while the gain for the most skilled was indistinguishable from zero, and in consulting the split was roughly 43% against 16.5%.21 The Index describes this as AI’s equalizing effect on workplace performance, with gains varying by the worker’s initial skill level and running largest for the least skilled.21
The frame is a synthesis, and both axes come from evidence already on the table: the knowledge-type axis from the SML rubric, the cost-of-error axis from the error-tolerance and accountability constraints the same rubric records, and the theoretical bridge from the prediction–judgment split.413 Its hard cases are real. A task can look explicit until conditions shift, which is the rubric’s stability criterion doing its work, and an error cost that is rare but severe is easy to underweight until one lands.4 Scoring a bundle honestly is itself a judgment call, which is to say it belongs to the humans the frame keeps in the loop.
The live evidence sorts along the frame’s own lines. The Anthropic Economic Index mapped about a million Claude conversations to O*NET tasks and found usage leaning 57% to 43% toward augmentation over automation, with roughly 4% of jobs using AI for at least three-quarters of their tasks and about 36% using it for at least a quarter.25 Adoption, in other words, is running task by task and mostly as a complement, which is what the within-occupation variance predicted.5 The OECD reads the employment data the same direction so far. High-skilled workers are the most exposed yet have seen employment gains, which it attributes to a reinstatement effect at this early stage, and it finds little evidence to date of significant negative employment effects from AI.22
One dataset cuts the other way, and it cuts precisely where the frame says it should. The Stanford “Canaries in the Coal Mine?” working paper — high-frequency payroll data, November 2025, not yet peer-reviewed — finds a 16% relative employment decline for workers aged 22 to 25 in the most AI-exposed occupations, while employment for experienced workers in the same occupations stayed broadly stable.24 Where and how the decline lands matters as much as its size. The declines concentrate in occupations where AI automates rather than augments; where it augments, the effect is muted.24 The adjustment runs through employment rather than pay, and the authors’ reading is that firms may be shrinking junior inflows rather than displacing incumbents.24 The sharpest instance is the most AI-native profession of all. By September 2025, employment of software developers aged 22 to 25 had fallen nearly 20% from its late-2022 peak.24
A 16% relative decline is a job-level outcome for the people inside it, and the task frame owes them an account. The account is in the paper’s own interpretation: AI “may be automating the codifiable, checkable tasks that historically justified entry-level headcount, while complementing the judgment-, client-, and process-intensive tasks performed by experienced workers.”24 In the language of the frame, an entry-level role is the closest a real job comes to a pure bundle of explicit, lower-stakes tasks, because the tacit, high-stakes residue that anchors an experienced worker’s bundle has not been accumulated yet. Where a bundle collapses toward one quadrant, task displacement and job displacement stop being different things. The authors call their findings early large-scale evidence rather than a settled verdict, and that hedge is theirs to set. The direction, though, is what the frame predicts. The job survives where its bundle is mixed, and the entry rung is where the bundle is least mixed.
The instruction for leaders follows the unit. Map the job into its tasks and score each one for knowledge type and cost of error. Automate the explicit, low-stakes tasks, put verification around the explicit, high-stakes ones, and keep humans leading the tacit, high-stakes core. Then redesign the role around what remains, which is where Brynjolfsson, Mitchell and Rock located the value in the first place.5 The maxim popularized by Karim Lakhani’s HBR headline — humans with AI will replace humans without AI — holds at this level too, read as an instruction about tasks rather than a prophecy about jobs.17 The returns come from re-bundling the tasks, and the re-bundled job is what carries them.
”AI won’t replace you, but someone using AI will.” Attribute as “popularized by Lakhani” (HBR, Aug 2023),17 not “coined by” — the maxim is a folk saying with no single verifiable first author. The corpus does not assert a definitive originator.