The Cost of Intelligence
It falls to zero, and strategy doesn't follow it down. The price of running a model is collapsing; building an advantage on top of it is where the work moves.
What this is. A strategy note for senior operators in financial services. The price of running a model is collapsing and will keep collapsing, and most organizations are building budgets, controls, and governance around that cost — apparatus for a problem that dissolves on its own, while the problem that stays gets almost none. Every external figure in this note is cited to a public source, and figures derived by this note rather than taken from a source are labelled as such in the text.
The cost of a fixed capability collapses; the price of the edge resets
Hold capability constant and the cost of producing it falls by around an order of magnitude a year. The model that first cleared a standard knowledge benchmark cost $60 per million tokens in November 2021; by November 2024 the cheapest model clearing the same bar cost six cents — a thousandfold fall in three years, and the origin of the working rule that fixed-capability inference gets ten times cheaper every year 1. The finding replicates. Stanford’s AI Index measured the same collapse one tier up, and a 2026 MIT-affiliated study of the same fixed-capability measure finds it still running through late 2025, at five to ten times a year on the cost-performance frontier 23.
There are two lines on the chart, though, and conflating them is where the strategic error begins. The floor, the cheapest way to buy a given capability, collapses toward zero. The frontier, the price of the newest capability at launch, behaves differently, and the record is messier than a slogan. For three generations the anchor genuinely held, with the price of OpenAI’s most powerful model landing at $60 per million tokens three times running: davinci’s blended rate in 2021, GPT-4’s output rate at launch in 2023, o1’s output rate at launch in 2024 456. The industry noticed the coincidence at the time 1. Through 2025 the flagship price broke downward, to $40 with o3 and then $10 with GPT-5 78. And in 2026 the premium slot re-formed above it. OpenAI’s pro tier now runs $180 per million output tokens, three times the old anchor, and the other frontier vendors price the same slot the same way, with Anthropic’s frontier at $50 and each new Google generation priced above its predecessor 91011. The reason the slot keeps re-forming is that “frontier” means the most compute anyone can usefully spend on a problem, and reasoning models made it possible to spend more. The cost of running frontier models is rising three to eighteen times a year even as fixed-capability cost collapses 3. The particular numbers wander. The structure holds: the edge stays priced like a scarce good.
How fast the floor falls turns out to depend on how you measure, and the disagreement is itself informative. Methods that track the cheapest route to a fixed benchmark score give multiples per year — ten times, and far more at some thresholds 1123 — because new models keep arriving underneath the old ones. Methods that build price indexes over market tiers give much slower numbers. Economy-tier token prices halve in about 1.1 years on that measure, and flagship prices barely decay at all, held up by a reasoning premium averaging 31.5 times non-reasoning rates 13. Gartner’s forward forecast for providers’ own inference cost (better than 90% cheaper by 2030, with the explicit caveat that loss-making providers may keep the gains rather than pass them through) sits in the same tens-of-percent-a-year territory 14. Both families of measurement are right about different things, and the study behind the tiered numbers compresses the reconciliation to “commoditization at the lower end and premium extraction at the frontier” 13. Ride the floor and your costs melt. Insist on the frontier and you will pay frontier prices indefinitely.
One caveat survives on the record, and it does not change the direction. The fastest measured declines are recent, and the analysts who measured them decline to promise they persist 12. What appears in no measured source is the tidy version, a one-time price war followed by a settling to ordinary rates. The nearest thing to a decomposition attributes essentially all of the cost reduction to software and architectural innovation and roughly none to hardware 13, though other sources here credit hardware a real share of the improvement 12.
The management conclusion follows from the shape of the curve. Building budget apparatus around token cost is calibrating to a number that is disappearing. Token cost is a melting asset. What deserves the apparatus is the set of costs and assets that do not melt.
Automating most of a call centre halves the cost, and tokens are almost none of what remains
Take ten customer-service agents handling roughly ten thousand contacts a month. Onshore, fully loaded — wages, supervision, facilities, QA, the compliance overhead a regulated firm carries — the industry’s 2026 pricing guides put that operation between roughly $42,000 and $65,000 a month 151617, with component-cost breakdowns suggesting a floor a couple of thousand lower 18, on one guide’s stated US average of about $26 per loaded agent-hour 15. Now move 80% of the contacts to a competent mid-tier model, and ask the question everyone asks: what does the AI cost to run?
Almost nothing, and the sources let you price it. Anthropic’s own worked example puts ten thousand support tickets through its cheapest current model for about $37 10; scale that through generous assumptions about context, retries, and a stronger model, and the token line lands in the low hundreds of dollars a month against a payroll of tens of thousands, under one percent of the bill. The arithmetic is this note’s own; the inputs are sourced. On a text channel, the tokens are a rounding error. A voice channel costs more. Assembled per-minute prices for AI voice agents run about $0.07 to $0.31, with the platform layer alone starting near $0.05 before model pass-through 192021, and inside the bundle the model is the small line, roughly a tenth of one platform’s worked per-minute total, with synthesis, transcription, telephony, and session infrastructure carrying the rest 2223. Even where the AI bill is real, intelligence is the smallest part of it.
But the operation’s cost does not fall 80%. Modelled through — and the after-figure is this note’s illustration, built on the sourced inputs above — it falls by about half. Three reasons. First, the residual 20% is the hardest 20%: escalations, genuine judgment, the angry customer, the regulated edge case, slower per contact and staffed by your better people. The research says the residual is structural. A pre-LLM study of customer-support automation found intent coverage asymptoting at 40%, with the remaining queries too messy and too rare to ever collect, and however much higher a modern system’s ceiling sits, the tail it never covers has the same shape 24. Second, you add an oversight function to run the machine and grade it, and that function is new cost. OpenAI’s own benchmark team found the same shape at the task level, where a frontier model’s naive advantage of 90 times faster and 474 times cheaper collapses to 1.4x and 1.6x once human review and rework are priced in 25. Third, in voice, the speech stack survives even when the intelligence is nearly free. The money stayed in human judgment and moved into the system that runs the AI; almost none of it moved to tokens.
The live case ran this experiment at full scale. Klarna’s assistant took two-thirds of chat volume in its first month and did the work, by the company’s own estimate, of seven hundred agents 26. Fifteen months later the CEO was telling Bloomberg that cost had been “a too predominant evaluation factor” and the result was lower quality, and Klarna was recruiting human agents again at pay starting at $41 an hour, for exactly the complex and sensitive cases, while the assistant kept the routine two-thirds 27. The company bought the human tier back.
Three things move the answer, and all of them belong in front of a board. If the deployment assists agents rather than replacing them, the measured gain is 14% on average, 34% for novices and little for the experienced 28. That is the ceiling on what “just add AI” does for an operation that keeps its people. Against an offshore baseline, where quotes run $6 to $14 an hour before a management overhead of 15 to 25 percent 16, the before-number for ten seats runs roughly $14,000 to $22,000 a month 15, and the saving shrinks to a fraction of the onshore case. And the more regulated and complex the work, the larger the residual that never leaves the human tier. A simple intake line automates most of its cost away. A compliance-heavy advisory line does not.
Handed the same model, everyone gets the same baseline; quality lives in what surrounds it
The strategy literature named this pattern before the technology existed. Carr’s argument about corporate IT, written in 2003, is that an input everyone can buy becomes a cost of doing business “that must be paid by all but provide distinction to none,” because what makes a resource strategic “is not ubiquity but scarcity” 29. The analysts covering the model market have reached the same verdict about foundation models — capability trends toward commodity, and advantage moves to the layer above 3031. So if the model is a commodity input, a deployment built on it is mostly the non-commodity assets expressed through that input, and an AI capability that is both high-quality and durable rests on five things a competitor with the same model cannot simply buy.
Proprietary data. The conditions here are the ones the venture literature spent a decade sharpening. Data defends when the sources are scarce or structurally locked to you and when it feeds a loop; raw volume defends little, because the marginal value of the next record falls while the cost of collecting it rises, and a static corpus erodes as competitors accumulate their own 24. In a regulated firm the structural lock is real. Your interactions, outcomes, and exceptions are generated by your licence and your book, and nobody else can collect them 3124.
Measurement. You cannot improve what you cannot grade. The practitioner literature is blunt here. Failed AI products almost always share one root cause, the absence of a working evaluation system, and a working one doubles as the data-curation, debugging, and fine-tuning engine, so the eval system and the improvement loop are substantially the same asset 32.
Encoded judgment. Your policies, your edge cases, the reasoning your best people apply, captured into the system. The first major causal field study of generative AI at work points to exactly this mechanism: its productivity gains came with “suggestive evidence” that the tool disseminated the best practices of the most able workers to everyone else 28.
The loop. The world is non-stationary and deployed models age. A systematic study of classical machine-learning deployments across four industries found quality degradation in 91% of model-dataset pairs as time passed since training — sometimes gradually, sometimes abruptly, and sometimes with no warning visible in the data 33. A static model decays; the mechanism that turns live outcomes back into improvement is the asset, and the model is only its current state.
Accountability. A human tier that owns the consequence. In financial services this is the licence to operate rather than a design choice, and it does not deflate when the model gets cheap. Klarna’s re-hiring is this asset being repurchased at market price. The CEO’s pledge that “there will be always a human if you want” is a brand promise, and the company judged it worth $41 an hour 27.
Put together, these five make a flywheel. The system deploys, captures what happens, grades it, improves, and deploys again, spinning on data only your operation generates. The venture version of the same conclusion, written as the model layer commoditized, runs through workflow, integration, trust, and distribution, and lands on the same point: “The new moats are the old moats” 31.
A saving every competitor gets is competed away; the durable move is redefining the service
The capability line rises through the org chart, and the pace is measured. The length of task an agent can complete autonomously has been doubling roughly every seven months for six years 34. The clearest human evidence sits at the bottom rung, where payroll data shows employment for 22-to-25-year-olds in the most AI-exposed occupations down 16% relative to peers while their seniors hold steady 35. The benchmark and usage records show the same rise from below — near-parity with experienced professionals on real multi-hour deliverables at the top of the measured range, task-shaped rather than job-shaped penetration in the middle of it 253637. The waterline rises from the bottom, and each year it covers work that last year needed a person.
The P&L from the call-centre section carries a trap. That roughly-half saving is available to every competitor running the same calculation, with the same models, the same voice vendors, the same playbook. Buffett shut Berkshire’s textile operation after years of watching cost-reducing investments that each looked like a winner, because once every mill made them, “their reduced costs became the baseline for reduced prices industrywide” 38. Every person at the parade stands on tiptoe and nobody sees better. Nordhaus measured the same mechanism economy-wide, finding that across half a century of US data producers kept about 2.2% of the surplus from technological advance, with the rest passing to customers 39. And the pattern is already running inside the AI stack itself, where commodity GPU compute is being competed toward marginal price 40 and model prices flipped from technology-driven to competition-driven decline in mid-2024 13. A surplus everyone can capture becomes the industry’s new baseline, and in a competitive, fee-sensitive market you will spend it to keep accounts you already hold.
You keep the surplus only through the things that do not commoditize, and the falling model price sharpens this, because when the input is handed to everyone, all the competitive weight lands on the part that is not. Which leaves the move that is hard to copy. Running today’s operation at half cost gets competed away; the durable move is to use execution-at-zero to deliver what rivals structurally won’t. The market is already repricing the unit of service: pay-per-resolution support is expected to run $1 to $7 per resolution in 2026, and one AI-first vendor’s live price card sells at $1.25 per resolution for AI-only service and $2.25 for hybrid 41. The financial-firm version is service that was never economical when a human had to deliver it: proactive outreach on every account rather than inbound only, continuous monitoring instead of periodic review, private-client depth at mass-market scale. The firms that win will be the ones that redefine what the service is now that delivering it costs almost nothing. Cutting the call-centre budget is the version of the move that every rival can copy.
The question that survives the falling price
What intelligence costs is settled. It goes to zero for any fixed capability, and the organization should plan accordingly — which means it should stop building apparatus to manage a melting asset. The question the falling price sharpens is what you build on top of free intelligence that a competitor with the same free intelligence cannot copy, and whether you use it to defend an old operation or to build a service they can’t match. The cost is melting. The moat is your data, your customers, your licence, and the speed of your loop. Manage those, and ignore the tokens.
On method. Figures are illustrative and move materially with workload mix, channel, and baseline; provider prices are quoted as of July 2026, and the 2026 frontier prices are listed prices at that date, not verified launch prices. Two figures are this note's own model rather than a single source — the scaled token line and the roughly-half after-figure in the call-centre section — and are labelled where they appear. The Klarna executive quotes originate in a paywalled Bloomberg interview and are cited via the Fortune and Customer Experience Dive reports that carry them.
Sources
- Appenzeller, "Welcome to LLMflation," a16z, 2024-11-12. https://a16z.com/llmflation-llm-inference-cost/
- Stanford HAI, 2025 AI Index Report, Research & Development chapter. https://hai.stanford.edu/ai-index/2025-ai-index-report/research-and-development
- Gundlach, Lynch, Mertens & Thompson, "The Price of Progress," arXiv 2511.23455, 2026. https://arxiv.org/abs/2511.23455
- OpenAI API pricing, December 2021 (Wayback). https://web.archive.org/web/20211223073823/https://openai.com/api/pricing/
- OpenAI pricing, March 16 2023, two days after GPT-4 launch (Wayback). https://web.archive.org/web/20230316024934/https://openai.com/pricing
- OpenAI API pricing, December 2024 / February 2025, o1 era (Wayback). https://web.archive.org/web/20250201145921/https://openai.com/api/pricing/
- OpenAI API pricing, April 19 2025, three days after o3 launch (Wayback). https://web.archive.org/web/20250419052710/https://openai.com/api/pricing/
- OpenAI, "Introducing GPT-5 for developers," 2025-08-07 (Wayback). https://web.archive.org/web/20250809142549/https://openai.com/index/introducing-gpt-5-for-developers/
- OpenAI API pricing, current, accessed 2026-07-02. https://developers.openai.com/api/docs/pricing
- Anthropic Claude platform pricing, accessed 2026-07-02. https://platform.claude.com/docs/en/about-claude/pricing
- Google Gemini API pricing, accessed 2026-07-02. https://ai.google.dev/gemini-api/docs/pricing
- Cottier et al., "LLM inference prices have fallen rapidly but unequally across tasks," Epoch AI, 2025-03-12. https://epoch.ai/data-insights/llm-inference-price-trends
- Du, "Tiered Super-Moore's Law," arXiv 2603.28576, 2026-03-30 (v1 preprint). https://arxiv.org/abs/2603.28576
- CIO Dive coverage of Gartner inference-cost forecast, 2026-03-25 (provider-side cost, fixed model size). https://www.ciodive.com/news/ai-inference-costs-drop-2030-gartner/815725/
- Contact Center USA, "Call Center Outsourcing Cost Per Hour in 2026". https://contactcenterusa.com/blog/call-center-outsourcing-cost-per-hour-2026
- Call Force Global, "Outsourced Call Center Pricing: 2026 Cost per Seat Guide". https://callforce.global/blog/call-center-outsourcing-cost/
- Hit Rate Solutions, "Call Center Pricing Guide: 2026 Rates". https://hitratesolutions.com/call-center-pricing-guide
- Nextiva, "How Much Does a Call Center Cost in 2026". https://www.nextiva.com/blog/call-center-cost.html
- Retell AI pricing, accessed 2026-07-02. https://www.retellai.com/pricing
- Deepgram pricing, accessed 2026-07-02. https://deepgram.com/pricing
- Vapi pricing, accessed 2026-07-02. https://vapi.ai/pricing
- LiveKit pricing (voice-agent per-minute component rates), accessed 2026-07-02. https://livekit.com/pricing
- Klariqo, "AI Voice Agent Cost Per Minute (2026)". https://klariqo.com/blog/voice-ai-cost-per-minute/
- Casado & Lauten, "The Empty Promise of Data Moats," a16z, 2019-05-09. https://a16z.com/the-empty-promise-of-data-moats/
- Patwardhan et al., "GDPval," OpenAI / arXiv 2510.04374, 2025-10-05. https://arxiv.org/abs/2510.04374
- Klarna press release, "AI assistant handles two-thirds of customer service chats in its first month," 2024-02-27. https://www.klarna.com/international/press/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month/
- Fortune (Ivanova) and CX Dive (Doerer), Klarna human-rehire coverage, 2025-05-09, carrying the Bloomberg interview of 2025-05-08. https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/
- Brynjolfsson, Li & Raymond, "Generative AI at Work," NBER w31161. https://www.nber.org/papers/w31161
- Carr, "IT Doesn't Matter," Harvard Business Review, May 2003 (captured from the author's repost). https://hbr.org/2003/05/it-doesnt-matter
- Evans, "How will OpenAI compete?," ben-evans.com, 2026-02-19. https://www.ben-evans.com/benedictevans/2026/2/19/how-will-openai-compete-nkg2x
- Chen, "The New Moats," Greylock, 2017 (2023 update). https://greylock.com/greymatter/the-new-moats/
- Husain, "Your AI Product Needs Evals," hamel.dev, 2024-03-29. https://hamel.dev/blog/posts/evals/
- Vela et al., "Temporal quality degradation in AI models," Scientific Reports, 2022. https://www.nature.com/articles/s41598-022-15245-z
- METR, "Measuring AI Ability to Complete Long Tasks," 2025-03-19 (data through Nov 2025). https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
- Brynjolfsson, Chandar & Chen, "Canaries in the Coal Mine?," Stanford Digital Economy Lab, 2025-11-13. https://digitaleconomy.stanford.edu/publications/canaries-in-the-coal-mine/
- Anthropic, "Introducing the Anthropic Economic Index," 2025-02-10. https://www.anthropic.com/news/the-anthropic-economic-index
- Anthropic, "Economic Index: New building blocks for understanding AI use," 2026-01-15. https://www.anthropic.com/research/economic-index-primitives
- Buffett, Chairman's Letter, Berkshire Hathaway 1985 Annual Report. https://www.berkshirehathaway.com/letters/1985.html
- Nordhaus, "Schumpeterian Profits in the American Economy," NBER w10433, 2004. https://www.nber.org/system/files/working_papers/w10433/w10433.pdf
- Cahn, "AI's $600B Question," Sequoia, 2024-06-20. https://www.sequoiacap.com/article/ais-600b-question/
- Crescendo, "Outsourced Customer Service Call Center Pricing Guide for 2026". https://www.crescendo.ai/blog/outsourced-call-center-pricing-guide