Paper 02

Where the Value Pools

Through 2024, AI revenue pooled in infrastructure, thinning by an order of magnitude at each step toward applications. In 2025, enterprise spending split almost evenly between the two. This paper maps where the value pools, what defends it, and the two dissents that contest the consensus.

18 verified sources A — The nature of the shift

A compiled, source-verified research digest — every claim cites a downloaded source, every figure is drawn from the data behind it. Not a personal essay.

Abstract

Historical technology revolutions split into an installation period that over-funds infrastructure and a deployment period in which value is realised through broad application and usage.9 Through 2024, AI ran the first half of that pattern: revenue pooled in semiconductors, hyperscaler capex and cloud, thinning by roughly an order of magnitude at each step up toward applications.6 In 2025 enterprise spending began to split, with $19 billion reaching the application layer against $18 billion for infrastructure.5 Two forces are redistributing the pool. The cost of a fixed level of machine intelligence is falling roughly 10× per year,1 and the foundation-model layer is converging on good enough for standard business tasks while its leaders churn.514 The durable moats sit at the two ends of the chain — scarce physical inputs below; proprietary data, workflow, distribution and the interface above71316 — though a sharp dissent argues the model and inference layer itself captured operating leverage in 2025–26.19

Prior waves funded the infrastructure first and paid out at the application layer

Carlota Perez’s history of technological revolutions gives the canonical map. Each 50–60-year cycle splits into an installation period, when financial capital floods into a new technology, builds its infrastructure and inflates an asset bubble, and a deployment period, when production capital takes over and the technology’s value is realised broadly across the economy through applications and usage.9 Railways ran this arc; so did steel and electricity, oil and mass production, and information and telecommunications from 1971.9 The turning point between the two phases is typically a financial crash.9 That last point sets a limit on how the framework can be used here, because no crash figures in this paper’s account. By Perez’s own criterion, AI cannot yet be placed in deployment, however deployment-like the 2025 spending data below looks. This paper leaves the phase question open and argues from the spend, the costs and the share data instead.

Two further frameworks say where in a value chain the money settles once a wave matures. Ben Thompson’s smiling curve, originated by Acer’s Stan Shih to explain the PC market, holds that value concentrates at the two ends of a chain — differentiated components on one side, brand, aggregation and distribution on the other — while the undifferentiated middle of assembly and integration gets margin-compressed.8 Aggregation Theory sharpens the right-hand end. Once the internet makes distribution of digital goods free, “suppliers can be commoditized leaving consumers/users as a first order priority,” and the firm that owns the customer relationship and user experience commoditises the suppliers beneath it.7 The theory awards that prize to whoever owns the end user, and it is silent on which layer the owner comes from. A model lab that owns the end-user relationship qualifies as an aggregator just as well as an application company does, and that possibility returns with the dissents at the end of this paper.

Why does the value lag the infrastructure rather than arriving with it? Brynjolfsson, Rock and Syverson’s productivity J-curve gives the peer-reviewed mechanism. General-purpose technologies “enable and require significant complementary investments, including co-invention of new processes, products, business models and human capital,” and those investments are largely intangible, mismeasured and slow to pay off.10 Productivity dips before it accelerates, and the returns flow to whoever builds the complementary assets around the technology.10 Read carefully, the J-curve names where value is created (inside adopting firms, in redesigned work) and stops short of naming which vendor layer captures it. McKinsey’s estimate belongs to the same family. Generative AI could add $2.6–4.4 trillion in annual economic value across 63 use cases, with about 75% concentrated in four functions (customer operations, marketing and sales, software engineering, and R&D), and with the potential to automate work activities that absorb 60 to 70 percent of employees’ time today.11 Those are figures for value to the economy and to adopting firms. They become application-vendor revenue only where adopting firms buy the complementary assets as products, a step the investor literature asserts and the economics alone does not supply. The 2025 purchasing data, taken up below, bears on exactly that step.

One caution applies to the whole corpus before the data arrives. Nearly every source arguing the up-the-stack case is an investor with a position. Summit, EQT, Sequoia and Menlo are investment firms, and Chamath Palihapitiya writes as an investor. This paper takes their data and their mechanisms, and treats their conclusions as claims to be tested.

The real opportunity isn’t in bigger models but in how AI models are applied.13 — Summit Partners, “Beyond Foundation Models: The Real Value of AI Lies in Applications” (2025)

The 2024 baseline: revenue pooled at the bottom of the stack

Map the AI stack as a vertical chain — semiconductors, hyperscaler infrastructure, cloud AI services, then foundation models and applications — and 2024 revenue fell by roughly an order of magnitude at each step up.6 Semiconductors earned an estimated $100–200B in AI revenue, led by Nvidia at a ~$105B run rate, while models and applications together sold only $5–10B, which Flaningam calls “a STEEP drop off in revenue from semiconductors, to data centers, to the cloud.”6 The chart below carries the layer-by-layer figures; the shape they make is a pyramid with its mass at the base.

2024 AI revenue / capex by stack layer (log scale, USD)$1B$10B$100BSemiconductors$100–200BHyperscaler capex$177BCloud AI services$20–25BModels + apps$5–10BBars run from the $1B axis origin to each layer’s reported upper figure on a log axis; ranges are in the labels.The hyperscaler figure is capital expenditure; the others are revenue.
Figure 1.Through 2024, AI dollars pooled in the bottom of the stack; revenue dropped roughly an order of magnitude at each step toward applications.Source: Flaningam, “The Current State of AI Markets,” Generative Value (2024).

That distribution is the setup for the whole debate, because it is not obviously sustainable. Sequoia’s David Cahn priced the implied bill. Take Nvidia’s run-rate data-centre revenue, double it for total cost of ownership, double it again for a 50% end-user gross margin, and you get the AI end-revenue the build-out implicitly requires.12 By mid-2024 that figure had risen from a “$200B question” to a “$600B question,” with the annual gap between implied and actual revenue widening from a $125B “hole” in September 2023 to roughly $500B in June 2024.12 The infrastructure spend is only justified if the value eventually shows up at the application and end-customer layer.12

A second caution attaches to the chart itself. Figure 1 maps where the dollars flow through, and where the profit sticks is a question the flow cannot answer. The hyperscaler bar is capital expenditure, as the caption notes, and none of the bars carries a margin. Value capture is in the end a margin question, and the open question at the close of this paper concedes the corpus lacks the margin data to settle it.

2025: enterprise spend reached the application layer

By 2025 the value had begun to arrive. Menlo Ventures’ December survey puts enterprise generative-AI spending at $37 billion for the year, up 3.2× on 2024, and finds it splitting almost evenly between the layers: $19 billion (51%) to applications against $18 billion (49%) to infrastructure.5 Set against the 2024 baseline, the movement is visible in two data points. In 2024, models and applications together sold $5–10 billion inside a flow dominated by chips and capex;6 a year later, applications alone took $19 billion, the larger half of enterprise spend.5 The two snapshots come from different analysts measuring different flows, so the comparison is loose; what it shows is spend reaching the layer the 2024 chart showed starved.

The buying pattern behind the split points the same way. Menlo found 76% of AI use cases purchased rather than built internally, up from 53% a year earlier, and AI deals converting at 47% against 25% for traditional SaaS.5 That is what a functioning application market looks like from the buy side. The infrastructure half kept pace. Model API spending reached $8.4 billion by mid-2025, more than doubling from $3.5 billion six months earlier, and that mid-year figure alone is already comparable to everything the models-plus-apps layer had sold across the whole of 2024.206

3.2×
YoY growth in enterprise gen-AI spend, $11.5B → $37B (2024→2025)
Menlo Ventures (2025)
>280×
Fall in inference cost for GPT-3.5-equivalent output in ~1.5 years
Stanford AI Index (2025)
$741B
Projected 2025 software-market spend — ~3× infrastructure
Summit Partners (2025)
1 in 8
Workers globally using AI each month
Innovation Endeavors (2025)

The first force: the cost of fixed-level intelligence is collapsing

The first force redistributing value is the falling price of machine intelligence. a16z’s Guido Appenzeller coined “LLMflation” for it: “for an LLM of equivalent performance, the cost is decreasing by 10× every year.”1 The cleanest documented case is GPT-3-class output, which fell from $60 per million tokens (GPT-3, November 2021) to $0.06 (Llama 3.2 3B), a 1,000× reduction over three years.1 Stanford’s AI Index documents the same collapse one capability tier up, and Figure 2 carries all three sourced series.3 The drivers are structural and stacked — a16z counts six, spanning hardware cost-performance, model efficiency and open-source competition — so the collapse does not depend on any single engine.1

Price to reach a fixed capability level (USD per million tokens, log scale)$60$6$0.6$0.06202120222024late’24$60$0.06 · 1,000× / 3 yr$20$0.07 · >280× / 1.5 yr$15$0.12 · 125× / 7 moGPT-3 quality (MMLU 42)GPT-3.5 equiv (MMLU 64.8)to Phi-4
Figure 2.Three documented price collapses to reach fixed capability levels. Prices are plotted on the log scale; the time-axis spacing is approximate; each endpoint is a literal data point.Source: a16z LLMflation (2024) and Stanford AI Index 2025, Ch.1.

The sourced figures describe a range. Epoch AI’s task-by-task study found prices falling “between 9× per year and 900× per year, with a median of 50× per year,” with the fastest trends starting after January 2024, after which the median rate rose from 50× to 200× per year.2 Different tasks commoditise at different speeds, and the frontier itself holds its price. OpenAI’s o1 launched at the same $60 per million output tokens as GPT-3 had in 2021.1 Chamath Palihapitiya, who draws the opposite value-capture conclusion from the same trend, supplies a corroborating figure: “the price of running a model has dropped 1,500× in six years, and intelligence is becoming free.”16

Verification note — the “10,000×” claim

A frequently-repeated figure holds that the cost of GPT-3.5-level intelligence fell ~10,000× over 2022–2026. No source in this corpus states that. The strongest primary (a16z) supports ~1,000× over three years for GPT-3-quality output;1 Stanford documents >280× in ~1.5 years for GPT-3.5-equivalent output;3 Epoch’s range tops out at 900×/year on specific tasks;2 Chamath cites 1,500× over six years.16 Treat 10,000× as an aggressive extrapolation. The defensible claim is “roughly 10× per year, two-to-three orders of magnitude over the era — unevenly, by task.”

The second force: models converge on good enough while the leaders churn

The second force is convergence at the model layer. EQT’s investors expect that “over time the hype around the foundational models will probably subside and they’re likely to become commoditized … a convergence point where all the models are good enough for most standard business applications.”14 Their portfolio companies already “use either open source or just very cheap models because the volume … tends to be very high.”14 Summit Partners describes the same dynamic from the supply side, where foundation-model builders “push for better performance at lower costs, often with shrinking margins.”13

The clearest evidence is enterprise share volatility. Menlo Ventures’ usage data shows OpenAI’s enterprise model share falling from 50% at the end of 2023 to 25% by mid-2025, recovering only to 27% by year-end.520 Anthropic and Google absorbed the share between them; Figure 3 carries the three trajectories. A layer where the leader can lose half its share in eighteen months is a contested input. Fortresses hold their share.

Enterprise LLM usage share, 2023 → year-end 2025 (%)502502023mid-2025YE 2025OpenAI 5027Anthropic 1240Google 721
Figure 3.Enterprise usage share at the model layer reshuffled hard: the 2023 leader more than halved by mid-2025. Mid-2025 and year-end columns come from two Menlo reports with slightly different provider groupings, and the time axis is not linear — the first interval spans 18–24 months, the second about six.Source: Menlo Ventures, State of GenAI in the Enterprise (Dec 2025) and Mid-Year LLM Market Update (2025).

The shape of the churn complicates the commoditisation story it seems to confirm. Menlo’s data shows the movement happening within providers far more than between them. Of enterprises surveyed, 66% upgraded models inside their existing provider and only 11% changed vendors, with the rest making no switch at all; yet within a month of Claude 4’s release it had taken 45% of Anthropic’s users, while Sonnet 3.5 fell from 83% to 16%.20 Switching vendors is rare, but repricing is constant, because every release resets what a dollar of inference buys. The buyers are also chasing capability ahead of price. Over the same period, open-source workloads fell from 19% to 13% of enterprise usage, with open models trailing the closed frontier by nine to twelve months.20 That decline qualifies the open-source driver in a16z’s list: open competition keeps pressing prices from below, and enterprises are nonetheless consolidating on closed frontier models, so the pressure on price and the pull of capability run at the same time.120

The same study records the compute centre of gravity shifting from training to inference, with 74% of startups now running majority-inference workloads, up from 48% a year earlier.20 Independent of any one report, Innovation Endeavors describes models becoming “10× cheaper, faster, and more capable year over year,” with capability churn rapid enough that models are characterised as obsoleting within weeks.18

On-device models: baseline capability now runs on a phone

The clearest single proof that baseline intelligence is commoditising is that it now runs on a phone. Microsoft’s Phi-3 technical report introduces phi-3-mini, a 3.8-billion-parameter model trained on 3.3 trillion tokens, “whose overall performance … rivals that of models such as Mixtral 8×7B and GPT-3.5 (e.g., phi-3-mini achieves 69% on MMLU and 8.38 on MT-bench), despite being small enough to be deployed on a phone.”4 Quantised to 4 bits it occupies roughly 1.8GB; the authors deployed it on an iPhone 14 with the A16 Bionic chip “running natively on-device and fully offline,” at more than 12 tokens per second.4 The “runs on a phone” claim, unlike the 10,000× claim, is fully verified.

The unit economics underneath reinforce it. Stanford’s AI Index records ML hardware costs for a fixed performance level dropping 30% per year while energy efficiency improves 40% annually; the B100 is 33.8× more energy-efficient than the 2016 P100.3 When frontier-2022 capability fits in 1.8GB on commodity silicon and the silicon itself gets 30% cheaper a year, the pricing power of the model layer erodes from below as well as from competition.43

Defensibility moved from model access to data, workflow, distribution and the interface

If models are a converging, cheapening input, defensibility has to come from somewhere else, and the corpus is consistent on where. EQT concludes that “the value is likely to accrue to the application layer and the product companies,” with the best-positioned firms being “those with pre-existing contracts, proprietary data or physical infrastructure.”14 Summit Partners locates durable advantage in application design — narrow domain focus, measurable ROI, integration with existing workflows — and treats model access as an input every competitor shares.13 Underneath both sits the J-curve’s mechanism, the returns flowing to the complementary assets built around the technology.10

The same sources draw a hard line through the middle of the application layer itself. M Accelerator’s case against thin “wrappers” is blunt: “AI wrappers don’t have moats because anyone can call the same APIs you’re using — your entire business model is one OpenAI update away from irrelevance.”15 Better prompts are discoverable, superior UI does not prevent switching when costs drop, and first-mover advantage evaporates when switching takes minutes; real defensibility appears only where proprietary data accumulates and network effects compound.15 The application-layer thesis is therefore conditional, and the condition is a hard problem, proprietary data and deep workflow lock-in.

Where defensible value sits in the AI stack

Differentiated component (down)

Scarce inputs — leading-edge chips, fabrication equipment, energy, HBM. Chamath’s “fulcrum assets.”16 Captures value through scarcity and chokepoint control.

Owned interface / distribution (up)

The customer relationship and workflow surface. Aggregation Theory’s prize: commoditise suppliers, own the user.7 Proprietary data and lock-in.14

Foundation models (middle)

Converging on “good enough,” share contested, margins compressing.145 The classic margin-compressed middle of the smiling curve; the dissents below contest this cell.

Thin wrappers (exposed)

No data, no lock-in, no distribution. “One OpenAI update away from irrelevance.”15 Squeezed from above and below.19

Value concentrates at the ends of the chain; the undifferentiated middle compresses — the smiling curve applied to AI8

Adobe shows the interface capturing the revenue, with one quote that does not check out

Adobe is the most-cited example of an incumbent monetising AI through an existing product surface. On its Q1 FY’25 call, CEO Shantanu Narayen described a three-stage approach — innovate, track usage, then “ensuring value and monetization” — and argued that “when somebody buys Creative Cloud or when somebody buys Document Cloud, in effect, they are actually monetizing AI.”17 AI-first products generated “more than $125 million” exiting Q1, expected to double by the end of fiscal ‘25.17 The revenue arrives where the workflow lives, at the application and interface layer, which is the thesis in miniature.

Verification note — the “interface layer” quote

A widely-circulated line attributes to Narayen the words “we make our money at the interface layer.” That verbatim phrasing was not found in the cited Diginomica reporting or any source located for this corpus, and is treated as unverified.17 The same article also contains no “$5B AI-influenced ARR” figure (Adobe’s $5.71B is total Q1 revenue; AI-first products were ”>$125M”). Rely on the verbatim quotes above; the paraphrase stays out of this corpus.

Two dissents: value pools down at the chokepoints, or the model layer keeps the leverage

The application-layer thesis is the corpus’s centre of gravity, and it is contested from two directions. The first dissent says value pools down the stack. Chamath argues computing eras are won by those controlling fulcrum assets — chokepoints where value concentrates — invoking the pattern of “Rockefeller had 90% of refining by 1880, Cisco had 85% of routing by 2000,” and locating today’s chokepoints in chip-manufacturing equipment, specialty materials and energy, far below the application layer.16

The second dissent is sharper, because it attacks the premise that the middle commoditised at all. Clarice Qiu’s “The Year the Commodity Layer Ate the Stack” argues that under specific 2025–26 conditions — high-willingness-to-pay workloads, persistent quality gaps favouring frontier models, and serving-stack innovation cutting cost-per-token faster than retail prices fell — the supposed commodity middle became “the operating-leverage layer,” like oil refining when demand outran capacity.19 Her headline exhibit is Claude Code reaching $2.5 billion in annual recurring revenue within nine months, and the corpus’s own market data supports the scale of it: Menlo records Claude holding 42% of code generation, double OpenAI’s 21%, and calls coding the “first breakout use case,” catalyst of a $1.9 billion ecosystem.1920 Her conclusion spares nothing for the thin middle. The most-exposed position, she writes, is “applications with no proprietary data, workflow lock-in, distribution advantage, or control over failure modes,” squeezed by labs capturing margin from above and cheap inference from below.19

The squeeze from above is Aggregation Theory pointed the other way. A model lab that owns the end-user relationship is an aggregator too, and Thompson’s framework awards it the same prize it awards any other interface owner; nothing in the up-the-stack reading rules that out.7 Claude Code is what the possibility looks like in practice.19

Weighed together, the evidence supports a layered position. Value pools at the ends of the chain, scarce physical inputs at the bottom and owned data, workflow and interface at the top, which is the smiling curve applied to AI.8 The model layer between them is genuinely contested, converging toward good enough for standard tasks14 yet able to capture operating leverage where quality gaps and willingness to pay hold.19 Every source in the corpus, bull and bear, denies the thin undifferentiated middle a durable claim on the value. And the 2025 spend split, $19 billion to applications against $18 billion to infrastructure, says the redistribution has begun, whether or not it yet counts as Perez’s deployment phase — a question her framework ties, typically, to a crash that has not arrived in this account.59

Open question

Whether 2025–26 model-layer operating leverage19 is a durable structural shift or a temporary spread that closes as capacity catches up is unresolved in the corpus. Qiu herself flags that semiconductor process nodes create stickier constraints than the temporary refining spreads in her analogy. Resolving it would require margin and pricing data the saved sources do not contain.

References

  1. Appenzeller, G. (2024). Welcome to LLMflation — LLM inference cost is going down fast. Andreessen Horowitz. Accessed 2026-06-16.
  2. Cottier, B., Snodin, B., Owen, D., & Adamczewski, T. (2025). LLM inference prices have fallen rapidly but unequally across tasks. Epoch AI. Accessed 2026-06-16.
  3. Maslej, N., et al. (2025). The 2025 AI Index Report (Chapter 1). Stanford HAI. Accessed 2026-06-16.
  4. Abdin, M., et al. (2024). Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone. arXiv:2404.14219 (Microsoft). Accessed 2026-06-16.
  5. Menlo Ventures (2025). 2025: The State of Generative AI in the Enterprise. Menlo Ventures. Accessed 2026-06-16.
  6. Flaningam, E. (2024). The Current State of AI Markets. Generative Value. Accessed 2026-06-16.
  7. Thompson, B. (2015). Aggregation Theory. Stratechery. Accessed 2026-06-16.
  8. Thompson, B. (2014). The Smiling Curve. Stratechery. Accessed 2026-06-16.
  9. Perez, C. (2002). Technological Revolutions and Financial Capital. Edward Elgar (Wikipedia overview). Accessed 2026-06-16.
  10. Brynjolfsson, E., Rock, D., & Syverson, C. (2018). The Productivity J-Curve: How Intangibles Complement General Purpose Technologies. NBER WP 25148. Accessed 2026-06-16.
  11. McKinsey & Company / MGI (2023). The economic potential of generative AI. McKinsey. Accessed 2026-06-16.
  12. Cahn, D. (2024). AI’s $600B Question. Sequoia Capital. Accessed 2026-06-16.
  13. Summit Partners (2025). Beyond Foundation Models: The Real Value of AI Lies in Applications. Summit Partners. Accessed 2026-06-16.
  14. EQT (ThinQ) (2025). Why AI Value Won’t Just Accrue to Base Models. EQT Group. Accessed 2026-06-16.
  15. M Accelerator (2025). Why AI Wrappers Don’t Have Moats. M Accelerator. Accessed 2026-06-16.
  16. Palihapitiya, C. (2026). Deep Dive: Where Value Accrues in the AI Stack. Chamath (Substack). Accessed 2026-06-16.
  17. du Preez, D. (2025). Adobe talks up AI monetization strategy as Wall Street’s disappointed at outlook provided. Diginomica. Accessed 2026-06-16.
  18. Innovation Endeavors (2025). State of Foundation Models 2025. Innovation Endeavors. Accessed 2026-06-16.
  19. Qiu, C. (2026). The Year the Commodity Layer Ate the Stack. Clarice Qiu (Substack). Accessed 2026-06-16.
  20. Menlo Ventures (2025). 2025 Mid-Year LLM Market Update. Menlo Ventures. Accessed 2026-06-16.