Paper 09

Business Model, Pricing & Governance

Agents change what a software company sells, what the selling costs, and who answers when the work goes wrong. All three move with how much control the buyer cedes to the machine.

27 verified sources C — Business model & the demand side

A compiled, source-verified research digest — every claim cites a downloaded source, every figure is drawn from the data behind it. Not a personal essay.

Abstract

AI agents decouple software’s value from the number of people using it, and three commercial facts move in response. The unit of sale shifts from the seat toward the work itself, so pricing climbs from subscriptions toward consumption and outcome meters.31 The delegated work carries a real cost of goods in compute, inference, and human oversight, which pulls gross margins well below the classic SaaS range.2 And the delegation draws legal exposure: voluntary frameworks, binding law, and live litigation in Mobley v. Workday are all working out who answers when an agent’s decision harms someone.1622

Outcome pricing, the far end of the ladder, is running in production at named vendors.69 Outcomes are also hard to attribute and easy to game, and roughly two-thirds of the incumbents Bain analysed have settled on hybrid platform-plus-consumption models instead.14 Underneath all three shifts sits one lever — how much control the buyer cedes to the agent. Ceded control is what the pricing meters, what the margin pays for, and what the liability follows.15

The unit of value moves from the seat to the work

The per-seat subscription was built on a proxy. A human operates the tool, so the number of humans logged in tracks the value delivered, and for two decades the proxy held. Agents break it, because an agent does the work itself and value attaches to the work completed. Andreessen Horowitz puts the proposition flatly: “Per-seat is no longer the atomic unit of software.”3 Bain & Company, analysing more than thirty SaaS vendors that added generative-AI capabilities, reaches a more measured version of the same conclusion — “Seats may not be dead, but they are no longer the only game in town.”1

Bain identifies two mechanisms doing the breaking. The first is decoupling. AI delivers value through automated background processes that need minimal human engagement, so value stops tracking headcount. The second is cost. Model inference, fine-tuning, and AI-specific R&D are substantial new expenses the vendor must recover somehow.1 The first mechanism gets the attention; the second is the one that reprices the business.

Underneath both mechanisms sits a design choice. Feng, McDonald, and Zhang, writing on agent autonomy, argue that “an agent’s level of autonomy can be treated as a deliberate design decision, separate from its capability and operational environment.”19 In other words, somebody chooses how much control the agent gets. Bessemer Venture Partners’ taxonomy of AI business models is the same choice expressed commercially. Copilots assist humans and price per seat or consumption, much like SaaS; agents execute entire workflows autonomously and price on outcome or workflow tied to tangible ROI; AI-enabled services blend automation with human oversight and price from consumption toward outcome, benchmarked against the cost of a full-time employee.2 Read as a dial running from assistance to delegation, the taxonomy carries the claim the rest of the evidence keeps confirming: how much control the buyer cedes sets what the vendor can charge for, what the service costs to run, and who answers when the work goes wrong.

Ceded work has a cost of goods, and the margin pays it

Bessemer states the cost mechanism bluntly in its 2026 pricing playbook: “AI economics ≠ SaaS economics. Every AI query costs money (compute, inference, human-in-the-loop). You’ll see 50–60% gross margins vs. 80–90% for traditional SaaS.”2 Microsoft, Bessemer notes, lost roughly $20 per user per month on GitHub Copilot at launch.2 A model in which marginal cost is real changes pricing from a packaging decision into a margin-survival decision. And the cost stack includes the governance posture itself. Human-in-the-loop is one of the inputs Bessemer names in the margin sentence, so keeping people supervising the agent is also a decision about gross margin.

Gross margin: classic SaaS vs AI-enabled software (%)100755025080–90%Traditional SaaS50–60%AI-enabled softwareMicrosoft initially lost ≈ $20 / user / month on GitHub Copilot. Lighter caps = top of each range.
Figure 1.The margin gap that forces the pricing question. Both ranges are Bessemer’s own estimate from a single playbook sentence; no second source in this corpus corroborates the figures.Source: Bessemer Venture Partners, “The AI Pricing Playbook for Founders,” 2026.

There is a measurement problem layered on top of the margin problem. McKinsey finds that only about 30% of software companies have published quantifiable ROI in dollar terms from real customer deployments, and warns that AI-enabling a full customer-service stack could imply a 60–80% increase in list prices.44b So the ask, as it stands, is that buyers pay materially more for a benefit most vendors cannot yet quantify. The buyers have noticed. One Fortune 100 HR executive told McKinsey’s researchers, “All of these copilots are supposed to make work more efficient with fewer people, but my business leaders are also saying they can’t reduce head count yet.”4 The buyer’s own costs compound the doubt, because McKinsey expects roughly $3 of change-management spending, on training and performance monitoring, for every $1 spent on model development.44b

The charge metric climbs from consumption to outcome, and cost risk climbs with it

Between the seat and the outcome, Bessemer maps three charge metrics, and they trade the same two variables against each other. Consumption pricing — tokens, API calls — gives clean margins and predictable costs, but customers do not think in tokens, so it works mainly for technical buyers. Workflow pricing, per task completed, sits closer to how work actually happens. Outcome pricing, per result delivered, offers maximum value alignment and puts maximum cost risk on the vendor.2 Bessemer compresses the trade into a sentence: “As you move from consumption → workflow → outcome-based pricing, you accept more cost risk for tighter value alignment.”2 And it insists the choice runs deeper than billing. “Your charge metric isn’t a billing decision — it’s a statement about what you believe your AI is worth and what you’re willing to stake your margins on to prove it.”2

The pricing ladder — value alignment and cost risk rise togetherVALUE ALIGNMENT → (and vendor cost risk)PRICE PER UNIT OF VALUE →Consumptiontokens / API calls · clean margins, predictable costWorkflowper task completed · closer to how work happensOutcomeper result delivered · max alignment, max cost risk
Figure 2.Bessemer’s charge-metric ladder. Moving up the rungs buys tighter alignment between price and customer value, and more variance lands on the vendor’s margin.Source: Bessemer Venture Partners, “The AI Pricing Playbook for Founders,” 2026.

The venture side reads the same ladder with more conviction. a16z names three pricing shifts AI forces — software becoming labour, outcome-based pricing, and variable costs, since foundation-model API calls scale with usage and break flat-fee logic.3 Its worked example is Zendesk, where customers historically paid roughly $115 per support-agent seat per month; as AI takes over the resolution work the seat was priced for, the natural pricing metric shifts toward “successful outcomes.”3 That single number — $115 per seat — is the thing the whole outcome-pricing movement is trying to replace.

Outcome pricing is live in production, and it is already drifting back toward a meter

Named vendors are charging on the ladder’s top rung. Bessemer calls Intercom’s $0.99 per resolved ticket “the gold standard” of outcome pricing,2 and Intercom’s official pricing for its Fin agent confirms the number: $0.99 per outcome, where a resolution means “no further help is requested after Fin’s last answer,” charged at most once per conversation however many actions the agent takes.6 The same page prices prospect qualification at $9.99, which shows “outcome” covering a menu of differently valued results; a resolved ticket and a qualified prospect are both outcomes, priced 10× apart.6

Salesforce shows the model evolving under commercial pressure. Agentforce launched with “an initial $2 fee per agent conversation”5 — a conversation, as Salesforce defined the metric, running from the first agent response until the issue is resolved, closed, or inactive for 24 hours. Within months Salesforce added Flex Credits at $0.10 per action, a Flex Agreement letting organisations convert user licences into credits and back, and per-user licensing, alongside pay-as-you-go and pre-commit payment options.7 The pivot answered buyer anxiety about the meter; Salesforce’s own research found that “90% of CIOs report that managing AI costs limits their ability to drive value.”7 Read together, the two announcements trace a trajectory that runs from an outcome meter back toward a consumption meter the buyer can predict.

Zendesk has built the most machinery around the billed unit. It announced in August 2024 that it was “first in CX industry to offer outcome-based pricing for AI agents,” with customers incurring costs “only for issues that are resolved autonomously by AI” and a starter usage level free.8 Its 2026 Relate launch added a verification layer aimed at the obvious dispute. Every charged resolution is “verified — both by the AI agent resolving the interaction end-to-end and independently confirmed by a dedicated AI evaluation model,” with spam and routine exchanges excluded.9 What Zendesk’s own pages do not carry is a dollar figure. The per-resolution price, clustering around $1.50, is documented only in third-party contract teardowns.9b

Harvey, the flagship legal vertical-AI vendor, has not climbed. It stays seat-anchored, on opaque enterprise pricing estimated at roughly $12,000–$16,800 per seat per year, and positions itself as “a labor cost substitute” priced at about 5–7% of associate labour cost. That is outcome logic in the pitch, seat logic in the invoice.25

Vendor / productUnitPriceSource character
Intercom Finper resolution (outcome)$0.99Official (fin.ai)
Intercom Finper qualification outcome$9.99Official (fin.ai)
Salesforce Agentforce (original)per conversation$2.00Press (Salesforce Ben)
Salesforce Agentforce Flexper action$0.10Official via MarTech
Zendesk AI agent (committed)per verified resolution≈ $1.50Third-party teardown
Zendesk legacy seatper agent / month$115a16z (the model being disrupted)
Harvey (legal, base)per seat / year$12k–$14.4kThird-party estimate
Per-unit AI-agent prices (USD, log scale)$0.10$1.00$10.00Salesforce Flex · per action — $0.10Intercom · per resolution — $0.99Zendesk · per verified resolution — ≈$1.50Salesforce · per conversation — $2.00Intercom · per qualification — $9.99
Figure 3.The going rate for an autonomous result spans two orders of magnitude depending on what is being charged for. A resolved ticket and a qualified prospect are both “outcomes,” priced 10× apart.Source: fin.ai/pricing; Salesforce Ben; MarTech (Salesforce); eesel AI Zendesk teardown.

Bessemer’s own bottom line hedges none of this. It is the bull case at full strength, stated as a claim about who wins.

AI doesn’t monetize access. It monetizes outcomes. The winners will charge for what their AI earns, not what it costs or what customers access.2 — Bessemer Venture Partners, “The AI Pricing Playbook for Founders” (2026)

Outcomes are hard to measure, and a better agent raises the bill

The claim has a measurement problem to survive, and the problem grows with the quality of the AI. Zendesk’s billing turns on a “Verified Resolution,” confirmed by a second language model that asks whether the reply actually solved the problem, and the teardown of those contracts lands on a perverse result: “a better-tuned AI agent costs more, not less,” because improved performance raises the count of verified resolutions on the bill.9b The buyer’s incentive to improve the product and the vendor’s incentive to bill for it point in opposite directions. Seat pricing never created that misalignment.

Attribution is the deeper problem. Lago, a billing-infrastructure vendor, argues that outcomes are deceptively hard to measure even with full data access. “If a user rolls their eyes, thinks ‘what a useless AI chatbot’ and slams their laptop shut in anger, they didn’t submit a ticket. Does the system count that as a resolution?“24 There is a boundedness problem too — “a proactive, outcome-based AI SDR might get 500 meetings in a week, but if you only have 2 salespeople, that’s counter-productive” — and Lago concludes that support tickets are a rare exception of clear, countable, uniform outcomes, while most professional work resists outcome billing because the results vary qualitatively or cannot be metered at all.24

Jason Lemkin of SaaStr raises the durability question from the buyer’s revealed preference. Proven, low-friction models win — payment processors charge per transaction, CRM vendors charge per seat — and outcome pricing may be “the cart driving the horse.” He cites a SaaStr Fund portfolio company whose outcome-based deal crossed $1m a year before the customer “quickly moved to a fixed contract.” He adds that as AI costs in B2B SaaS fall toward zero, the economic rationale for outcome pricing weakens; major vendors charge “$1–$3 per outcome-based resolution,” yet HubSpot and Box have largely abandoned the approach.18 “A pricing model is not a product. And a pricing model doesn’t make a mediocre product great.”18

The bull side has started to say the same thing. Bessemer’s 2026 playbook carries its own renewal warning. Soft ROI worked in 2025’s “AI adoption at all costs” environment, and “as pilots hit renewal, pricing must reflect actual value delivered, not promise.”2 The deals priced on promised outcomes meet that test at their first renewal cycle, with the measurement problem still open.

For a buyer, the same evidence reads as a term sheet. The meter needs a predictable cap, which is the anxiety that forced Salesforce’s pivot.7 The outcome needs a bound, or the vendor’s incentive runs toward Lago’s five hundred unusable meetings.24 And the definition of the billed outcome, and who verifies it, belongs in the contract, because as the market stands the verifier is the vendor’s own model.9

Verification note — the Zendesk “$1.50 per resolution” figure

Zendesk’s own newsroom pages confirm the mechanism — outcome-based pricing and dual AI verification of each charged resolution89 — but do not state a per-resolution dollar amount. The ”≈$1.50 (committed) / ≈$2.00 (pay-as-you-go)” figures come from third-party contract teardowns and search summaries of 2026 pricing pages, with sources clustering at ~$1.50 but ranging $1.20–$2.00.9b Lemkin independently brackets the category at “$1–$3 per outcome-based resolution.”18 Treat the dollar figure as third-party-sourced, not Zendesk-official. Intercom’s $0.99 and Salesforce’s $2.00 / $0.10, by contrast, are confirmed on primary/official material.657

Incumbents settle on hybrid while AI-natives run the outcome experiments

Among incumbents, the destination is already visible, and it is a hybrid. Bain’s analysis of 30-plus AI-enabled SaaS vendors found roughly 35% raised per-seat prices while bundling AI into existing tiers (Zoom, for one), and roughly 65% adopted hybrid models layering an AI usage or outcome meter on top of seat-based pricing (Adobe, Salesforce). 0% monetised AI purely as a standalone add-on, and 0% had fully transitioned to usage- or outcome-only pricing.1 The pure-outcome endpoint the venture commentary celebrates had, at the time of Bain’s study, zero incumbent adopters.

How 30+ AI-enabled SaaS vendors actually priced AI (% of analysed vendors)0%50%100%Hybrid: usage/outcome meter on seats~65%Raised per-seat price, AI bundled~35%AI as standalone add-on only0%Fully usage/outcome-only0%
Figure 4.The actual distribution of incumbent behaviour. Hybrid dominates, pure outcome pricing has no adopters among analysed incumbents, and the new meters layer on top of seats.Source: Bain & Company, “Per-Seat Software Pricing Isn’t Dead…,” 2025 (30+ vendors analysed).

Bain also explains why incumbents stall at the hybrid rung. Moving off per-seat runs into three walls: internal hurdles, because vendors lack the telemetry infrastructure and shared pricing metrics across silos that a usage or outcome meter requires; commercial obstacles, because sales teams accustomed to seat-based conversations have to be retrained; and customer resistance, because procurement teams are unfamiliar with value-based purchasing that demands upfront investment before measurable savings materialise.1 The hybrid equilibrium is what buyer demand for predictability plus those switching costs jointly produce. Seats stay because the machinery for anything else is still being built, on both sides of the contract.

The venture enthusiasm and the incumbent data stop conflicting once the population is named. a16z observes that AI-native companies — Decagon, Cursor, ElevenLabs — employ usage-based, outcome-based, or hybrid pricing, while “established software companies adding AI features tend to maintain per-seat or bundled approaches.”3 Deloitte’s finding that 83% of AI-native SaaS companies already offer usage-based pricing is a finding about the natives;21 Bain’s 0% is a finding about the incumbents.1 The ladder is being climbed by the companies born on it, and leaned against by the companies that would have to rebuild billing, sales motion, and procurement story to climb it. Both the bull case and the bear case are right about their own population.

The advisory prescriptions converge on the same hybrid structure. McKinsey’s is the fullest: blend per-user subscription fees with consumption-based metrics, similar to Microsoft’s Copilot structure, and meter consumption beyond a capacity cap — tokens processed per day, week, or month.4 The others sit inside the same envelope. Gartner’s projection, cited by Deloitte, has at least 40% of enterprise SaaS spend shifting toward usage-, agent-, or outcome-based pricing by 2030, which leaves a majority still anchored on seats or subscriptions,21 and Bessemer’s founder formula is a base subscription plus usage or outcome tiers that “provide predictability while capturing upside,” illustrated as a platform fee at roughly twice delivery cost plus outcome credits.2

50–60%
AI-enabled software gross margin vs 80–90% for classic SaaS
Bessemer (2026)
60–80%
Implied list-price increase to AI-enable a full customer-service stack
McKinsey (2025)
$1 : $3
Model-development spend vs change-management spend
McKinsey (2025)
≥ 40%
Enterprise SaaS spend on usage/agent/outcome pricing by 2030
Gartner via Deloitte (2026)
0%
Analysed incumbents fully transitioned to usage/outcome-only pricing
Bain (2025)

There is a forward-looking economic argument that outcome meters get cheaper to run as agents proliferate. An NBER chapter on the “Coasean Singularity” argues that AI agents “dramatically reduce transaction costs” by lowering “the costs of preference elicitation, contract enforcement, and identity verification,” which “expand[s] the feasible set of market designs.”17 Pricing tied to verifiable outcomes is exactly the kind of market mechanism that gets cheaper when verification is cheap, and Zendesk’s AI-verifier-on-AI-resolution architecture is that argument in microcosm. The same chapter warns that the transition “also introduce[s] frictions such as congestion and price obfuscation” and “novel regulatory challenges.”17

Voluntary frameworks and binding law converge on the same obligations

As agents take on consequential decisions, four reference frameworks have become the de facto governance stack. They differ on binding force and barely differ on substance. Accountability, traceability, risk management, and human oversight recur in each, and ISO/IEC 42001 is positioned as a complement to the NIST framework and the EU AI Act.26 Table 2 carries the inventory of publishers, dates, structures, and penalties.

The binding instrument is the EU AI Act (Regulation (EU) 2024/1689, in force 1 August 2024). It sorts systems into four risk tiers, bans the unacceptable tier outright, and subjects high-risk systems listed in Annex III to risk-management, documentation, data-governance, and human-oversight obligations, phased in over 6 to 36 months; penalties reach €35 million or 7% of global annual turnover for prohibited-practice violations.11 Annex III explicitly includes recruitment AI. The function at issue in Mobley v. Workday, algorithmic screening of job applicants, is work European law has already named high-risk.

The voluntary layer asks for the same things without the fine. NIST’s AI RMF, organised around four functions with governance cross-cutting the other three, names “Accountable and Transparent” among its seven characteristics of trustworthy AI.10 ISO/IEC 42001 turns the same substance into the first certifiable AI management-system standard, requiring AI impact assessment, lifecycle controls, and third-party supplier oversight.26 And beneath both sit the OECD AI Principles, the first intergovernmental AI standard, which make accountability an obligation of every AI actor “based on their roles, the context, and consistent with the state of the art.”23

FrameworkPublisher / yearCore structureBinding?Key figure
AI RMF 1.0 (AI 100-1)NIST · 2023Govern · Map · Measure · ManageVoluntary7 trustworthiness traits; GenAI Profile 600-1 (2024)
EU AI Act (2024/1689)European Union · 20244 risk tiers: unacceptable / high / limited / minimalBinding lawUp to €35m or 7% of global turnover
ISO/IEC 42001:2023ISO/IEC · 2023Plan-Do-Check-Act AI management systemVoluntary (certifiable)First international certifiable AIMS
OECD AI PrinciplesOECD · 2019 (am. 2024)5 values-based principles + 5 recommendationsSoft lawFirst intergovernmental AI standard

Frameworks are multiplying faster than the practice underneath them. Stanford’s 2025 AI Index records that in 2024 the OECD, EU, UN, and African Union all published responsible-AI frameworks, while the number of reported AI incidents hit a record 233 — a 56.4% jump over 2023.27 The same Index flags the gap the frameworks have not closed. Leaders in a McKinsey survey cite inaccuracy, regulatory compliance, and cybersecurity as their top responsible-AI risks, yet “not all are taking active steps to address them,” and standardised responsible-AI benchmarks remain sparse among major model developers.27

Accountability stays with the buyer, and Mobley tests whether it also reaches the vendor

Every layer of the stack lands on one principle, and IBM states it in procurement terms. When IBM asks who in an organisation should be accountable for responsible AI outcomes, the three most common answers it hears are “no one,” “we don’t use AI,” and “everyone” — answers it describes as “none of which are correct and all are concerning.” It concludes: “The principle that accountability cannot be outsourced remains critical; while AI can enhance efficiency and enable innovations, human oversight, judgment, and accountability must remain central.”16 The contracts do not rescue the buyer either. IBM finds that “most software and cloud vendor contracts lack explicit commitments that make them accountable for providing responsible AI,” and some include disclaimers removing liability outright.16 The OECD Principles encode the same duty by role, supported by mandatory traceability of datasets, processes, and decisions across the lifecycle.23

Outcome pricing gives that principle a commercial edge, because outcome billing puts the vendor in charge of measuring the thing it bills for. Zendesk’s charged resolutions are verified by Zendesk’s own evaluation model,9 the teardown of those contracts shows the bill rising as the agent improves,9b and IBM finds the standard vendor contract makes no responsible-AI commitments at all.16 Who audits the verifier is one question wearing two labels — a pricing dispute waiting to happen, and an accountability gap already documented.

Whether exposure also reaches the vendor is being decided in Mobley v. Workday (N.D. Cal., No. 3:23-cv-00770-RFL). Lead plaintiff Derek Mobley submitted over 100 applications through employers using Workday’s screening system and was rejected each time, alleging the AI tools “unfairly penalize older candidates.”14 The suit names the vendor; the employers who used the screen are not the defendants. The EEOC’s April 2024 amicus brief argues that Workday can be directly liable under Title VII, the ADA, and the ADEA on three theories — as an employment agency, as an indirect employer controlling access to opportunity, and as an agent of employers to whom hiring authority was delegated.22 Each theory tracks the delegation. The more of the hiring function the employers handed to Workday’s tools, the more directly the statutes reach the company that ran them; ceded control, here, is the theory of liability. The EEOC also rejects the size defence outright: “there is no authority for the novel proposition that an entity can be too big to qualify as an employment agency… whether Workday performs those tasks for one employer or thousands of employers.”22

The court has let the case advance and grow. In July 2024 it allowed the agent theory to proceed; in May 2025 it granted conditional nationwide certification of the ADEA collective, covering applicants aged 40 and older denied recommendations through Workday since 24 September 2020.1314 Per Workday’s own filings, roughly 1.1 billion applications were rejected during the period, with the collective potentially numbering in the “hundreds of millions” — one of the largest ever certified in employment litigation.13 A later order, filed 6 March 2026, preserved the ADEA disparate-impact claim while dismissing the related state-law FEHA claims as pled, noting that “attending a historically Black college can be used by AI screening tools as a proxy for race.”12 Workday maintains the suit “lacks merit” and stresses the rulings are preliminary.14

The lesson has two halves resting on different evidence. The buyer-side half — deploying a vendor’s AI does not move accountability off the deployer — rests on IBM’s contract finding and the OECD’s role-based duty.1623 What Mobley tests is the vendor-side half: whether the exposure created by a delegated, regulated function reaches the company that built and operated the screen. Whatever the outcome, the case does not relieve the deployer; the accountability that IBM and the OECD locate with the buying organisation sits where it was.

Verification note — what Mobley has and has not decided

These are rulings on motions to dismiss and a conditional collective certification — procedural milestones letting claims proceed, not a final finding that Workday discriminated.1314 The “agent of the employer” holding traces to the July 2024 order (documented via the EEOC amicus and law-firm analyses);2213 the saved court order (Doc. 267, 2026-03-06) is a later ruling addressing the ADEA and FEHA claims, not the original agent-theory order.12 The ~1.1 billion-applications figure is from Workday’s own filings as reported by counsel,13 not an independently audited number. Stated here as the litigation status, not as a determination of liability.

Autonomy is set by design, and price, margin, and liability move with it

The variable running underneath all of this has a formal literature, and the literature treats it as governable. Feng, McDonald, and Zhang define a five-level spectrum of agent autonomy by the user’s role — operator, collaborator, consultant, approver, observer — ranging from direct control to light supervision, and propose “AI autonomy certificates” to govern agent behaviour, calling autonomy “a double-edged sword.”19 The spectrum maps onto the familiar human-in-the-loop / human-on-the-loop / human-out-of-the-loop taxonomy. Operator and approver keep a human in or on the loop; observer approaches out-of-the-loop.

The risk claim attached to the spectrum is blunt. Mitchell, Ghosh, Luccioni, and Pistilli of Hugging Face argue, from the title down, that fully autonomous agents should not be developed, on the principle that “risks to people increase with the autonomy of a system: The more control a user cedes to an AI agent, the more risks to people arise.”15 Their preferred posture is constrained, semi-autonomous systems with human oversight.15 Two deployable postures follow. High autonomy on a narrow scope lets a tightly bounded task run end-to-end; broad scope at low autonomy permits wide-ranging action behind human verification gates. The danger zone pairs broad scope with high autonomy and no gate.

Two safe postures, one danger zone — agency × scope

High agency · narrow scope

A tightly-bounded task run end-to-end (resolve this ticket; reconcile this account). Safe because the blast radius is small. This is where outcome pricing works — discrete, measurable, contained.24

High agency · broad scope

The danger zone. Wide-ranging autonomous action with no verification gate. Risk rises with ceded control;15 autonomous pricing agents can even sustain collusion “without agreement, communication, or intent.”20

Low agency · narrow scope

A copilot that suggests and waits. Human-in-the-loop by construction. Low risk, but also soft ROI and weak pricing power — priced per seat.2

Low agency · broad scope

Wide reach behind human-on-the-loop verification gates (the approver role).19 Broad usefulness made safe by review — the posture most enterprise deployments should default to.

Autonomy is a governable lever set independently of capability;19 verification gates convert broad scope from danger zone to deployable

The same lever shows up in the pricing data. The cell where outcome pricing works in practice — Intercom’s resolutions, Zendesk’s verified resolutions — is high agency on narrow scope: discrete, contained, measurable tasks, exactly Lago’s “rare exception” of clear, countable outcomes.24 The cell that frightens regulators is high agency on broad scope. NBER work on AI-powered trading shows reinforcement-learning agents “autonomously sustain collusive supra-competitive profits without agreement, communication, or intent” — price-fixing emerging with no instruction to collude, degrading market efficiency.20 That is the empirical case for verification gates on any agent holding both broad scope and transactional authority. Set the dial deliberately, and the rest follows from the setting — the charge metric that fits, the margin the oversight costs, and the liability the delegation carries. Pair high agency with narrow scope, or broad scope with low agency behind gates, and price, govern, and accept liability accordingly.

Open question

Whether outcome-based pricing durably “sticks” or reverts to hybrid consumption is unresolved in the corpus. The bull case (a16z, Bessemer) and the bear case (Lemkin, Lago) both have evidence — confirmed live deployments on one side,65 a documented reversion to fixed contract and the “better AI costs more” perverse incentive on the other.189b Bain’s finding that 0% of incumbents had fully transitioned1 suggests hybrid is the stable equilibrium for now, but the sources do not contain the multi-year renewal-and-retention data needed to call the endpoint. The related governance question — whether vendor liability under Mobley survives to final judgment and generalises beyond hiring — is likewise open.

References

  1. Maltiel, E., & Sandberg, J. (2025). Per-Seat Software Pricing Isn’t Dead, but New Models Are Gaining Steam. Bain & Company. Accessed 2026-06-16.
  2. Bessemer Venture Partners (2026). The AI Pricing Playbook for Founders. Bessemer Venture Partners. Accessed 2026-06-16.
  3. a16z Enterprise team (2024). AI Is Driving a Shift Towards Outcome-Based Pricing (December 2024 Enterprise Newsletter). Andreessen Horowitz. Accessed 2026-06-16.
  4. McKinsey & Company (2025). Upgrading Software Business Models to Thrive in the AI Era. McKinsey & Company. Accessed 2026-06-16 (figures verified by corroboration).
  5. The Register (2025). McKinsey wonders how to sell AI apps with no measurable benefits. The Register. Accessed 2026-06-16 (independent corroboration of the McKinsey figures).
  6. Salesforce Ben editorial (2024). Salesforce’s Bold New Pricing Strategy: What You Need to Know. Salesforce Ben. Accessed 2026-06-16.
  7. Intercom / Fin (2026). Fin AI Agent Pricing. Intercom (fin.ai). Accessed 2026-06-16.
  8. MarTech (2025), reporting Salesforce’s May 2025 announcement. Salesforce Introduces New Agentforce Pricing Models (Flex Credits). MarTech.org. Accessed 2026-06-16.
  9. Zendesk (2024). Zendesk First in CX to Offer Outcome-Based Pricing for AI Agents. Zendesk newsroom. Accessed 2026-06-16.
  10. Zendesk (2026). Zendesk Introduces the Autonomous Service Workforce (Relate 2026). Zendesk newsroom. Accessed 2026-06-16 (verification mechanism; no dollar figure on page).
  11. eesel AI (2026). Understanding Zendesk AI Pricing: A Complete Pay-per-Resolution Guide. eesel AI. Accessed 2026-06-16 (third-party teardown; sources the ~$1.50 figure).
  12. NIST — Tabassi, E., et al. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. U.S. Dept. of Commerce / NIST. Accessed 2026-06-16.
  13. European Union (2024). EU AI Act — Regulation (EU) 2024/1689 (high-level summary). artificialintelligenceact.eu (Future of Life Institute). Accessed 2026-06-16.
  14. U.S. District Court, N.D. Cal. (2026). Mobley v. Workday, Inc., No. 3:23-cv-00770-RFL — Court Order (Doc. 267, filed 2026-03-06). govinfo.gov. Accessed 2026-06-16.
  15. Proskauer Rose LLP (2025). AI Bias Lawsuit Against Workday Reaches Next Stage (Conditional ADEA Certification). Proskauer Rose LLP. Accessed 2026-06-16.
  16. Holland & Knight LLP (2025). Federal Court Allows Collective Action Over Alleged AI Hiring Bias. Holland & Knight. Accessed 2026-06-16.
  17. Mitchell, M., Ghosh, A., Luccioni, A. S., & Pistilli, G. (2025). Fully Autonomous AI Agents Should Not Be Developed. arXiv:2502.02649 (Hugging Face). Accessed 2026-06-16.
  18. IBM (2025). Who Is Accountable for Responsible AI? IBM Think Insights. Accessed 2026-06-16.
  19. Shahidi, P., Rusak, G., Manning, B. S., Fradkin, A., & Horton, J. J. (2025). The Coasean Singularity? Demand, Supply, and Market Design with AI Agents. NBER. Accessed 2026-06-16.
  20. Lemkin, J. (2025). The Real Question for Outcome-Based Pricing Is If It Will Stick. SaaStr. Accessed 2026-06-16.
  21. Feng, K. J. K., McDonald, D. W., & Zhang, A. X. (2025). Levels of Autonomy for AI Agents. arXiv:2506.12469. Accessed 2026-06-16.
  22. Dou, W. W., Goldstein, I., & Ji, Y. (2025). AI-Powered Trading, Algorithmic Collusion, and Price Efficiency. NBER Working Paper 34054. Accessed 2026-06-16.
  23. Deloitte (2026). SaaS Meets AI Agents (TMT Predictions 2026). Deloitte Insights. Accessed 2026-06-16.
  24. U.S. Equal Employment Opportunity Commission (2024). Brief of EEOC as Amicus Curiae in Support of Plaintiff, Mobley v. Workday, Inc. EEOC. Accessed 2026-06-16.
  25. OECD (2019, amended 2024). OECD AI Principles (incl. Accountability). OECD. Accessed 2026-06-16.
  26. Lago team (2025). Outcome-Based Pricing Is NOT the Future. Lago (Substack). Accessed 2026-06-16.
  27. Metronome (2025). Harvey AI Pricing Index. Metronome. Accessed 2026-06-16 (pricing = third-party estimate; Harvey publishes none).
  28. ISO/IEC JTC 1/SC 42 (2023). ISO/IEC 42001:2023 — Artificial Intelligence Management System. ISO. Accessed 2026-06-16.
  29. Stanford HAI (2025). 2025 AI Index Report, Chapter 3: Responsible AI. Stanford Institute for Human-Centered AI. Accessed 2026-06-16.