Disuse · Algorithm aversion
Capable AI rejected. Triggered by seeing the system err; confidence in algorithms collapses faster than in humans after the same mistake. Dominant among experts and when the comparison is to one’s own judgment.35
Organizations bought the technology; most have yet to capture the value. The distance between the two is human and organizational. Trust, identity, workflow, and incentive decide whether a tool in hand becomes a tool in use.
A compiled, source-verified research digest — every claim cites a downloaded source, every figure is drawn from the data behind it. Not a personal essay.
By 2024–2025, enterprise AI adoption looked close to solved on paper, with 78% of organizations reporting AI use, up from 55% a year earlier.9 Over the same period more than 80% of companies reported no material earnings contribution from generative AI,8 and the Census Bureau’s firm-level survey put real US adoption in the low single digits by firm count.18 The gap between having AI and profiting from it is organizational, and this paper traces the mechanisms that produce it. The economics predicts the lag: general-purpose technologies require waves of intangible complementary investment — new processes, skills, business models — that depress measured productivity before they lift it.12 Rogers’ diffusion theory explains why reach runs ahead of depth.7 The behavioral literature supplies the conversion mechanism, a trust calibrated to what the system can actually do,6 and documents its failures in both directions, from abandonment after a single visible error3 to over-weighting of algorithmic advice.5 Field experiments find the measured gains concentrated in novices,13 and recent identity evidence finds workers hiding their AI use because using it is socially penalized.1220 The playbook that follows keeps people in control of outputs, rebuilds workflows around the tool, and treats managers and training as the unit of adoption.422
The headline surveys say adoption is close to solved. Stanford’s AI Index reports organizational AI use jumping to 78% in 2024 from 55% in 2023, with generative-AI use in at least one business function more than doubling from 33% to 71%.9 The other practitioner surveys confirm the pattern at every level of the stack. McKinsey’s November-2025 edition put organizational AI use at 88%,8 and Microsoft and LinkedIn found three-quarters of global knowledge workers already using generative AI at work, most of them bringing their own tools.11 The speed holds up under rigorous measurement too. A nationally representative US survey found work adoption of generative AI “as fast as the personal computer” and overall adoption faster than either PCs or the internet.17
The government’s measurement tells a much smaller story. The Census Bureau’s Business Trends and Outlook Survey, the highest-frequency official measure of firm behavior, found US firm-level AI use rising from 3.7% to 5.4% between September 2023 and February 2024, which is low single digits by firm count, concentrated in large firms and the Information sector.18 Both readings can be true, because they measure different events; a worker who has used a chatbot and a firm that deploys AI for business purposes are different units of adoption. The census data also carries the paper’s thesis in miniature. Firms that do adopt, BTOS finds, undergo organizational changes to accommodate the technology, “training staff, developing new workflows, and purchasing cloud services/storage” — the complementary-investment story, observed in official micro-data.18
The value numbers say even the wide reach is shallow. In McKinsey’s March-2025 edition, more than 80% of companies reported no material contribution to earnings from their generative-AI initiatives.8 The Stanford AI Index finds that among firms reporting any financial impact at all, most put it “at low levels,” with cost savings under 10% and revenue gains under 5% as the modal outcomes.9 The same survey that found record adoption speed also found that generative AI assists only 1–5% of total work hours, with reported time savings equivalent to about 1.4% of hours.17 Reach is wide; depth is thin. MIT’s Project NANDA compressed the gap into a contested but widely-cited figure, roughly 95% of enterprise generative-AI pilots delivering no measurable P&L impact, and located the cause in an organizational “learning gap,” the failure to integrate AI into workflows and structures.15
BTOS carries one more finding, and it cuts against the paper’s title. The most common reason firms give for non-adoption is “the inapplicability of AI to the business.”18 Part of the adoption gap, in other words, is a sober judgment that the tool has no task to do in a given firm, and no amount of trust-building or workflow redesign would change that. The corpus does not say how large that part is. The constraint this paper examines is the one that remains where AI demonstrably applies — licenses bought, tasks the field experiments show AI measurably helps — and the value still fails to arrive.
Two practitioner-credible syntheses name the cause directly. McKinsey’s own finding is that “the redesign of workflows has the biggest effect on an organization’s ability to see EBIT impact from its use of gen AI.”8 Harvard Business Review’s 2025 article opens from the same diagnosis: most firms struggle to capture value from AI “not because the technology fails—but because their people, processes, and politics do.”22 A diagnosis, though, names the failure without explaining it, and the explanation starts with the economics of why the lag was expected all along.
The gap between visible capability and invisible payoff has precedent, and it has a name. Brynjolfsson, Rock and Syverson opened their 2017 paper on “the modern productivity paradox” with the observation that AI systems “match or surpass human level performance in more and more domains” while “measured productivity growth has declined by half over the past decade” — a redux, they noted, of the Solow paradox of transformative technologies “everywhere but in the productivity statistics.”1 They weighed four explanations, false hopes, mismeasurement, redistribution, and implementation lags, and concluded that lags are “the biggest contributor”: the most capable AI “has not yet diffused widely,” and like other general-purpose technologies its full effect “won’t be realized until waves of complementary innovations are developed and implemented.”1
The decisive move is treating those complements, “the required adjustment costs, organizational changes, and new skills,” as a kind of intangible capital.1 Building that capital consumes resources now and pays out later, so during the build “total factor productivity growth will initially be underestimated because capital and labor are used to accumulate unmeasured intangible capital stocks,” and during the harvest measured productivity overshoots the truth.2 The measurement error traces a J, down and then up. Their calibration puts an intangibles-adjusted TFP measure 15.9% above the official figure by the end of 2017,2 and the historical record supplies the timescale, with electrification taking “a generation for the nature of factory layouts to be re-invented” before its benefits appeared.2 Their conclusion generalizes the pattern for any leader impatient for ROI: “the more transformative the new technology, the more likely its productivity effects will initially be underestimated.”2
The size of the eventual payoff is the contested part, and the leading skeptic works from the same underlying data. Daron Acemoglu, who drew on AI-exposure data shared by J-curve co-author Daniel Rock, estimates the macroeconomic upside of AI as “nontrivial but modest — no more than a 0.71% increase in total factor productivity over 10 years,” and below 0.55% once hard-to-learn tasks, which lack the objective feedback AI needs, are accounted for.16 Brynjolfsson sees a large payoff delayed by the complement-building lag; Acemoglu sees a payoff that is modest even at full diffusion. Both can be partly right. The J-curve says the gap we observe today is expected, and Acemoglu warns that the eventual top of the J may sit lower than the optimists assume. What the disagreement does settle for an operator is the ordering of work. On either account, the determinant of realized value is the complementary organizational investment, and that investment is where the human constraint lives.
The J-curve is a theory of measurement and intangible capital; the captured papers find it empirically for software and “small but growing” effects for AI-related intangibles as of 2017–20182 — they do not prove a large AI payoff is coming, only that an early productivity dip is consistent with one. Acemoglu’s modest ceiling and the J-curve’s optimistic delay are not yet reconciled by the evidence in this corpus; whether the top of the AI J-curve resembles 15.9% or sub-1% TFP is unsettled.
If the economics explains why the payoff lags, Everett Rogers’ diffusion theory explains the shape of adoption and why it stratifies. Rogers defines diffusion as “the process in which an innovation is communicated through certain channels over time among the members of a social system,” and identifies five perceived attributes that predict how fast an innovation spreads: relative advantage, compatibility, complexity, trialability, and observability.7 Relative advantage is “the strongest predictor,” and complexity, “the degree to which an innovation is perceived as relatively difficult to understand and use,” is the only attribute that works against adoption.7
Run generative AI through this rubric and the adoption-vs-value gap stops being a surprise. Relative advantage is large and visible for some tasks, drafting and summarizing and coding assistance among them, and weak or unproven for others; the field experiments below show it near zero for experienced experts on familiar work.1314 Trialability is unusually high, since anyone can open a chatbot, which helps explain the record top-line adoption speed.17 Observability is where the rubric turns against AI. Rogers notes software already has “a low level of observability, [so] its rate of adoption is quite slow,“7 and as the identity evidence later in this paper shows, with AI the observability of use triggers a social penalty, so people hide it.
The adopter categories supply the second reframe. Rogers splits any adopting population into innovators (2.5%), early adopters (13.5%), early majority (34%), late majority (34%), and laggards (16%), with adoption stalling at the early/late-majority crossover unless opinion leaders mediate and uncertainty falls.7 An organization, on this map, adopts as a distribution of these populations, and the stall is visible in the 2024–2025 data. BCG found overall regular use at 72% while frontline adoption sat at 51%.10 Slack’s Fall-2024 wave caught the curve flattening in time, with US adoption growing a single percentage point over three months, and excitement cooling six points among the overall population.12 Geography tells the same story at a larger scale: BCG’s country splits run from 92% regular use in India to 64% in the US and 51% in Japan.10 The model behind the chat window is the same in Mumbai and Tokyo, so where adoption varies that much across social systems while the technology is held constant, the technology is the wrong place to look for the cause.
Diffusion theory says observability and relative advantage matter; the trust literature says why they convert to use. John Lee and Katrina See’s foundational review establishes that “because people respond to technology socially, trust influences reliance on automation,” and that trust matters most “when complexity and unanticipated situations make a complete understanding of the automation impractical.”6 Their central construct is appropriate reliance, trust calibrated to the system’s true capability, and they name two symmetric failure modes: disuse, rejecting capable automation, and misuse, over-relying on imperfect automation.6 Both are miscalibration. The implication for an adoption program is that people should trust the system exactly where it is reliable and override it where it is unreliable; pushing trust upward across the board manufactures the second failure mode while curing the first.
Disuse has a well-documented behavioral engine. In the canonical study of algorithm aversion, Dietvorst, Simmons and Massey ran five incentivized experiments and found that people “are especially averse to algorithmic forecasters after seeing them perform, even when they see them outperform a human forecaster.”3 The engine is asymmetric confidence, since “people more quickly lose confidence in algorithmic than human forecasters after seeing them make the same mistake.”3 The aversion holds despite the algorithms’ real superiority; a meta-analysis of 136 studies found algorithms outperforming human forecasters by about 10% on average.3 Seeing the algorithm err is the trigger, which is why a single visible AI mistake can collapse adoption even when the system is winning on average.
The enterprise deployment data connects here, and it needs more care than the aversion label invites. Among workers who tried Microsoft Copilot and stopped, 44.2% cited distrust of answers as the reason.21 The instinct is to file that under aversion, as a barrier to train away. The paper’s own framework blocks the shortcut, because nothing in this corpus establishes that Copilot’s answers were reliable for those workers’ tasks, and on Lee and See’s account, distrust of an unreliable system is calibration doing its job. The same survey sharpens the point. ChatGPT converts 83.1% of US paid subscribers with workplace access against Copilot’s 35.8%, and workers given both choose ChatGPT 76% to Copilot’s 18%.21 A workforce that supposedly resists AI is walking past one tool to use another heavily. That behavior reads as rough calibration across tools, and it means tool quality and fit remain live variables on the technology side of the ledger.
The literature also runs the other direction. Logg, Minson and Moore document algorithm appreciation across six experiments, finding that lay people “adhere more to advice when they think it comes from an algorithm than from a person.”5 The appreciation “waned when people chose between an algorithm’s estimate and their own,” and it waned among experts — “experienced professionals, who make forecasts on a regular basis, relied less on algorithmic advice than lay people did, which hurt their accuracy.”5 Read together, aversion and appreciation map when each response appears, and the map has a clear axis: expertise, and whether the comparison is to one’s own judgment.
Two failure modes of reliance (Lee & See) × two behavioral defaults
Capable AI rejected. Triggered by seeing the system err; confidence in algorithms collapses faster than in humans after the same mistake. Dominant among experts and when the comparison is to one’s own judgment.35
Imperfect AI over-trusted. The mirror risk Lee & See warn of; the “algorithm appreciation” default in lay users can become automation bias where outputs are accepted without verification.65
Trust tracks the system’s true capability, task by task, with resolution and specificity — rely where it is reliable, override where it is unreliable.6
Letting people even slightly modify an imperfect algorithm sharply raises willingness to use it and improves their performance — the bridge from trust theory to the playbook.4
UNDER-RELIANCE ←——— miscalibrated trust ———→ OVER-RELIANCE
The trust map predicts that AI’s value, and the resistance to it, should vary sharply by skill. The field-experiment evidence confirms it. In “Generative AI at Work,” Brynjolfsson, Li and Raymond studied the staggered rollout of a conversational AI assistant across 5,179 customer-support agents and found productivity, measured as issues resolved per hour, rising 14% on average, with “a 34% improvement for novice and low-skilled workers” and “minimal impact on experienced and highly skilled workers.”13 The mechanism is dissemination: AI “disseminates the best practices of more able workers and helps newer workers move down the experience curve.”13 The tool lifts the bottom of the skill distribution by handing it the top’s instincts.
Now the counterpoint. METR ran a randomized controlled trial with 16 experienced open-source developers on mature repositories they had each worked on for an average of five years, randomizing 246 real tasks to allow or disallow early-2025 AI tools. The developers forecast a 24% speedup beforehand and estimated a 20% speedup afterward, and the measured result was a 19% slowdown.14 Outside experts had predicted 38–39% speedups.14 All three vantage points erred in the same direction. Experts confidently believed they were faster while measurably being slower. (The authors are careful about scope — small N, expert contributors on familiar complex code, explicitly no claim about novices or greenfield work.)
The asymmetry resolves the apparent contradiction in the algorithm literature, since experts under-rely exactly where lay novices appreciate and gain. It also reframes who resists and why. If AI’s value is largest for novices and smallest, sometimes negative, for experts, then the most skilled, highest-status people in an organization are precisely those for whom the relative advantage is weakest and the identity threat strongest. Resistance, for a senior expert, can be a locally accurate read of both the productivity math and the status math.
The same asymmetry reaches back to the value gap the paper opened with. If the measured gains concentrate at the bottom of the skill distribution and fade toward zero at the top, a firm that deploys AI to everyone should expect modest aggregate financials, because it is paying for the whole distribution and collecting from one tail. That is roughly what the value data records — the under-10% cost savings and under-5% revenue gains Stanford finds as the modal outcome.9 No source in this corpus draws that connection, so it stands here as this paper’s inference; every term in it is measured, and it converts the opening gap from an anomaly into something the field experiments would have predicted in advance.
The barrier most adoption plans underweight is social evaluation, the question of how using AI makes a worker look. Slack’s Workforce Index found just 7% of desk workers consider AI outputs “completely trustworthy,“12 and its Fall-2024 wave found something sharper than distrust: 48% of desk workers said they would be uncomfortable telling their manager they had used AI for a common task, with the leading reasons being that using AI feels like “cheating” (47%) and fear of being seen as less competent or lazy (46% each).12 Adoption is, for nearly half the workforce, something privately rational to hide. Beneath the discomfort sits a harder fear, with BCG finding 41% of workers who think their job will certainly or probably disappear entirely within ten years.10
Two studies move this from correlation to cause. In four preregistered experiments with 4,439 participants, Reif, Larrick and Soll found that “people who use AI at work anticipate and receive negative evaluations regarding their competence and motivation,” evaluations that “affect assessments of job candidates.”19 The belief that you will be judged for using AI is both held and justified. David Almog’s field experiment with 450 remote workers then priced the behavior it produces: “workers adopt AI recommendations at lower rates when their reliance on AI is visible to the evaluator, resulting in a measurable decline in task performance.”20 Making use observable cut reliance by 14% in relative terms and lowered accuracy, and reassuring the evaluator about the workers’ track record did not remove the effect; workers reported fearing that heavy AI reliance “signals a lack of confidence in their own judgment.”20 The penalty suppresses use and degrades measured performance at the same time.
Rogers’ framework gives this finding a second edge. Observability is the attribute that spreads innovations through social proof; people adopt what they watch respected peers adopt. For AI at work, the evidence above says observability administers a penalty instead, so the one diffusion attribute that should be driving uptake is the one suppressing it. Hidden use also corrupts the measurement layer, because the workers concealing their chat windows are invisible to the adoption dashboards their leaders rely on.
The evidence converges on a small number of levers, all of them organizational. The first comes out of the algorithm-aversion lab. In “Overcoming Algorithm Aversion,” Dietvorst, Simmons and Massey found that “participants were considerably more likely to choose to use an imperfect algorithm when they could modify its forecasts, and they performed better as a result,” and that “the preference for modifiable algorithms held even when participants were severely restricted in the modifications they could make.”4 The effect was “indicative of a desire for some control… not for a desire for greater control”; a slight, capped ability to adjust the output raised adoption, satisfaction, and willingness to use the algorithm again.4 The lever the experiments support is keeping a human in control of the output. This paper extends it one step, to aligning accountability with that control, and the extension should be read as an inference from the pattern, since Dietvorst tested control and never tested accountability. On that inference, people accept tools they can steer and are answerable for; they reject tools that displace their judgment while leaving them responsible for the result.
The second lever is workflow redesign ahead of tool deployment. McKinsey’s data points here, with workflow redesign having “the biggest effect on an organization’s ability to see EBIT impact,“8 and BCG locates “the next frontier: from adoption to value with end-to-end redesign,” noting that the roughly half of companies starting to reshape processes “invest more in their people—and it pays off.”10 MIT NANDA’s “learning gap” states the same point as a failure mode, since generic tools “stall in enterprise use since they don’t learn from or adapt to workflows.”15 Buying licenses is the cheap, visible move, and the Recon Analytics data shows where it ends, with a 35.8% Copilot conversion rate meaning roughly two-thirds of provisioned licenses sit idle.21 A license is access; adoption is what happens after the workflow is rebuilt around it.
The third lever addresses the identity barrier, and it runs through middle managers and training. Microsoft and LinkedIn found only 39% of people who use AI at work had received any company AI training, and the rest of their survey shows the same neglect from every angle, with few companies planning generative-AI training that year and leaders professing urgency while doubting their own organizations have a plan.11 BCG’s numbers agree, with 36% of employees satisfied with their AI training and “proper training, leadership support, and access to the right tools” named as what breaks the frontline ceiling.10 Managers carry the load for two reasons grounded in the evidence. In Rogers’ terms they are the local opinion leaders who “put their stamp of approval on a new idea by adopting it” and reduce uncertainty for the early and late majority.7 And because the identity penalty is administered by the manager’s perceived judgment, only the manager can lift it, by making AI use expected and sanctioned, which directly counters the 47% “feels like cheating” reflex.1219 HBR’s synthesis names the corollary failure. Where incentives are misaligned, workers may decline to surface AI-driven time savings out of self-interest, so the value never shows up in the numbers.22 (summarized from HBR’s argument — its body beyond the verbatim dek is paywalled; not a verbatim HBR quote, see verification note)
Most firms struggle to capture real value from AI not because the technology fails—but because their people, processes, and politics do. Harvard Business Review, “Overcoming the Organizational Barriers to AI Adoption,” 2025 (verbatim dek)
The playbook owes the reader one qualification before it can be priced. If Acemoglu’s ceiling is the right one — a TFP gain under 1% over ten years — the expected value of closing the human constraint shrinks proportionally, and so does the case for a heavy adoption program.16 The corpus cannot settle which ceiling is real. It does settle the ordering. On the optimistic account and the skeptical one alike, whatever value exists is captured only by organizations that do the complementary work, so the levers above are the path to the payoff at either size; what changes with the ceiling is how much a firm should spend pulling them.
For the operator, the program compresses. Invest through the lag instead of judging the tool by its first-year P&L; give people control of the output, with accountability aligned to that control; rebuild the workflow around the tool; and put managers in front of the identity penalty, aiming the deployment at the novices and lower-skill workers who measurably gain. On the technology side, one caution survives from the trust data. The same survey that shows two-thirds of Copilot licenses idle shows ChatGPT converting 83.1% of US paid subscribers with workplace access,21 so tool choice still moves adoption, and a flat claim that the technology is solved overstates what this corpus licenses. What it does license is the ordering: the models are further along than the organizations trying to use them, and the distance between the two is the human constraint.
The “Copilot usage fell from 25% to 5%” anecdote. This teaching-case figure is widely repeated but could not be traced to any primary source in this corpus, and is therefore not stated as fact anywhere above. The attributable, real adjacent data come from Recon Analytics (survey of 150,000+ respondents, Jul 2025–Jan 2026): a Microsoft Copilot workplace conversion rate of 35.8% — from which “~64% of licenses idle” is derived as the arithmetic inverse (100 − 35.8), recorded as derived, not a verbatim Recon claim — and 44.2% of lapsed Copilot users citing distrust of answers.21 Treat the “25%→5%” line as an unverifiable anecdote; cite the Recon figures instead.
MIT NANDA’s “95% of pilots fail.” The ~95%/~5% split is captured via Fortune’s editorially-vetted coverage, not the primary PDF (which would not render), and the headline has been publicly contested on methodology.15 It is used above only as a directional claim that organizational integration, not model capability, is the binding constraint, and is paired with the better-instrumented value-gap evidence (McKinsey #8, Stanford HAI #9, BCG #10) rather than leaned on alone.
Acemoglu TFP ceiling. The figures used (≤0.71% over 10 years; <0.55% conservative) are the verbatim numbers in the captured 2024 working paper; an upstream “0.66%” was checked against the source and rejected as not the figure in this version.16
Partial-capture sources. Lee & See (#6) is cited from its verbatim PubMed abstract plus standard summaries of its framework (full text paywalled); McKinsey (#8) figures are McKinsey’s own published wording surfaced via search (site blocked automated retrieval); the HBR (#22) body beyond the verbatim dek is summarized, not quoted; the METR (#14) and Almog (#20) studies are preprints. Each is flagged in the source manifest.