High agency · narrow scope
A tightly-bounded task run end-to-end (resolve this ticket; reconcile this account). Safe because the blast radius is small. This is where outcome pricing works — discrete, measurable, contained.24
Agents change what a software company sells, what the selling costs, and who answers when the work goes wrong. All three move with how much control the buyer cedes to the machine.
A compiled, source-verified research digest — every claim cites a downloaded source, every figure is drawn from the data behind it. Not a personal essay.
AI agents decouple software’s value from the number of people using it, and three commercial facts move in response. The unit of sale shifts from the seat toward the work itself, so pricing climbs from subscriptions toward consumption and outcome meters.31 The delegated work carries a real cost of goods in compute, inference, and human oversight, which pulls gross margins well below the classic SaaS range.2 And the delegation draws legal exposure: voluntary frameworks, binding law, and live litigation in Mobley v. Workday are all working out who answers when an agent’s decision harms someone.1622
Outcome pricing, the far end of the ladder, is running in production at named vendors.69 Outcomes are also hard to attribute and easy to game, and roughly two-thirds of the incumbents Bain analysed have settled on hybrid platform-plus-consumption models instead.14 Underneath all three shifts sits one lever — how much control the buyer cedes to the agent. Ceded control is what the pricing meters, what the margin pays for, and what the liability follows.15
The per-seat subscription was built on a proxy. A human operates the tool, so the number of humans logged in tracks the value delivered, and for two decades the proxy held. Agents break it, because an agent does the work itself and value attaches to the work completed. Andreessen Horowitz puts the proposition flatly: “Per-seat is no longer the atomic unit of software.”3 Bain & Company, analysing more than thirty SaaS vendors that added generative-AI capabilities, reaches a more measured version of the same conclusion — “Seats may not be dead, but they are no longer the only game in town.”1
Bain identifies two mechanisms doing the breaking. The first is decoupling. AI delivers value through automated background processes that need minimal human engagement, so value stops tracking headcount. The second is cost. Model inference, fine-tuning, and AI-specific R&D are substantial new expenses the vendor must recover somehow.1 The first mechanism gets the attention; the second is the one that reprices the business.
Underneath both mechanisms sits a design choice. Feng, McDonald, and Zhang, writing on agent autonomy, argue that “an agent’s level of autonomy can be treated as a deliberate design decision, separate from its capability and operational environment.”19 In other words, somebody chooses how much control the agent gets. Bessemer Venture Partners’ taxonomy of AI business models is the same choice expressed commercially. Copilots assist humans and price per seat or consumption, much like SaaS; agents execute entire workflows autonomously and price on outcome or workflow tied to tangible ROI; AI-enabled services blend automation with human oversight and price from consumption toward outcome, benchmarked against the cost of a full-time employee.2 Read as a dial running from assistance to delegation, the taxonomy carries the claim the rest of the evidence keeps confirming: how much control the buyer cedes sets what the vendor can charge for, what the service costs to run, and who answers when the work goes wrong.
Bessemer states the cost mechanism bluntly in its 2026 pricing playbook: “AI economics ≠ SaaS economics. Every AI query costs money (compute, inference, human-in-the-loop). You’ll see 50–60% gross margins vs. 80–90% for traditional SaaS.”2 Microsoft, Bessemer notes, lost roughly $20 per user per month on GitHub Copilot at launch.2 A model in which marginal cost is real changes pricing from a packaging decision into a margin-survival decision. And the cost stack includes the governance posture itself. Human-in-the-loop is one of the inputs Bessemer names in the margin sentence, so keeping people supervising the agent is also a decision about gross margin.
There is a measurement problem layered on top of the margin problem. McKinsey finds that only about 30% of software companies have published quantifiable ROI in dollar terms from real customer deployments, and warns that AI-enabling a full customer-service stack could imply a 60–80% increase in list prices.44b So the ask, as it stands, is that buyers pay materially more for a benefit most vendors cannot yet quantify. The buyers have noticed. One Fortune 100 HR executive told McKinsey’s researchers, “All of these copilots are supposed to make work more efficient with fewer people, but my business leaders are also saying they can’t reduce head count yet.”4 The buyer’s own costs compound the doubt, because McKinsey expects roughly $3 of change-management spending, on training and performance monitoring, for every $1 spent on model development.44b
Between the seat and the outcome, Bessemer maps three charge metrics, and they trade the same two variables against each other. Consumption pricing — tokens, API calls — gives clean margins and predictable costs, but customers do not think in tokens, so it works mainly for technical buyers. Workflow pricing, per task completed, sits closer to how work actually happens. Outcome pricing, per result delivered, offers maximum value alignment and puts maximum cost risk on the vendor.2 Bessemer compresses the trade into a sentence: “As you move from consumption → workflow → outcome-based pricing, you accept more cost risk for tighter value alignment.”2 And it insists the choice runs deeper than billing. “Your charge metric isn’t a billing decision — it’s a statement about what you believe your AI is worth and what you’re willing to stake your margins on to prove it.”2
The venture side reads the same ladder with more conviction. a16z names three pricing shifts AI forces — software becoming labour, outcome-based pricing, and variable costs, since foundation-model API calls scale with usage and break flat-fee logic.3 Its worked example is Zendesk, where customers historically paid roughly $115 per support-agent seat per month; as AI takes over the resolution work the seat was priced for, the natural pricing metric shifts toward “successful outcomes.”3 That single number — $115 per seat — is the thing the whole outcome-pricing movement is trying to replace.
Named vendors are charging on the ladder’s top rung. Bessemer calls Intercom’s $0.99 per resolved ticket “the gold standard” of outcome pricing,2 and Intercom’s official pricing for its Fin agent confirms the number: $0.99 per outcome, where a resolution means “no further help is requested after Fin’s last answer,” charged at most once per conversation however many actions the agent takes.6 The same page prices prospect qualification at $9.99, which shows “outcome” covering a menu of differently valued results; a resolved ticket and a qualified prospect are both outcomes, priced 10× apart.6
Salesforce shows the model evolving under commercial pressure. Agentforce launched with “an initial $2 fee per agent conversation”5 — a conversation, as Salesforce defined the metric, running from the first agent response until the issue is resolved, closed, or inactive for 24 hours. Within months Salesforce added Flex Credits at $0.10 per action, a Flex Agreement letting organisations convert user licences into credits and back, and per-user licensing, alongside pay-as-you-go and pre-commit payment options.7 The pivot answered buyer anxiety about the meter; Salesforce’s own research found that “90% of CIOs report that managing AI costs limits their ability to drive value.”7 Read together, the two announcements trace a trajectory that runs from an outcome meter back toward a consumption meter the buyer can predict.
Zendesk has built the most machinery around the billed unit. It announced in August 2024 that it was “first in CX industry to offer outcome-based pricing for AI agents,” with customers incurring costs “only for issues that are resolved autonomously by AI” and a starter usage level free.8 Its 2026 Relate launch added a verification layer aimed at the obvious dispute. Every charged resolution is “verified — both by the AI agent resolving the interaction end-to-end and independently confirmed by a dedicated AI evaluation model,” with spam and routine exchanges excluded.9 What Zendesk’s own pages do not carry is a dollar figure. The per-resolution price, clustering around $1.50, is documented only in third-party contract teardowns.9b
Harvey, the flagship legal vertical-AI vendor, has not climbed. It stays seat-anchored, on opaque enterprise pricing estimated at roughly $12,000–$16,800 per seat per year, and positions itself as “a labor cost substitute” priced at about 5–7% of associate labour cost. That is outcome logic in the pitch, seat logic in the invoice.25
| Vendor / product | Unit | Price | Source character |
|---|---|---|---|
| Intercom Fin | per resolution (outcome) | $0.99 | Official (fin.ai) |
| Intercom Fin | per qualification outcome | $9.99 | Official (fin.ai) |
| Salesforce Agentforce (original) | per conversation | $2.00 | Press (Salesforce Ben) |
| Salesforce Agentforce Flex | per action | $0.10 | Official via MarTech |
| Zendesk AI agent (committed) | per verified resolution | ≈ $1.50 | Third-party teardown |
| Zendesk legacy seat | per agent / month | $115 | a16z (the model being disrupted) |
| Harvey (legal, base) | per seat / year | $12k–$14.4k | Third-party estimate |
Bessemer’s own bottom line hedges none of this. It is the bull case at full strength, stated as a claim about who wins.
AI doesn’t monetize access. It monetizes outcomes. The winners will charge for what their AI earns, not what it costs or what customers access.2 — Bessemer Venture Partners, “The AI Pricing Playbook for Founders” (2026)
The claim has a measurement problem to survive, and the problem grows with the quality of the AI. Zendesk’s billing turns on a “Verified Resolution,” confirmed by a second language model that asks whether the reply actually solved the problem, and the teardown of those contracts lands on a perverse result: “a better-tuned AI agent costs more, not less,” because improved performance raises the count of verified resolutions on the bill.9b The buyer’s incentive to improve the product and the vendor’s incentive to bill for it point in opposite directions. Seat pricing never created that misalignment.
Attribution is the deeper problem. Lago, a billing-infrastructure vendor, argues that outcomes are deceptively hard to measure even with full data access. “If a user rolls their eyes, thinks ‘what a useless AI chatbot’ and slams their laptop shut in anger, they didn’t submit a ticket. Does the system count that as a resolution?“24 There is a boundedness problem too — “a proactive, outcome-based AI SDR might get 500 meetings in a week, but if you only have 2 salespeople, that’s counter-productive” — and Lago concludes that support tickets are a rare exception of clear, countable, uniform outcomes, while most professional work resists outcome billing because the results vary qualitatively or cannot be metered at all.24
Jason Lemkin of SaaStr raises the durability question from the buyer’s revealed preference. Proven, low-friction models win — payment processors charge per transaction, CRM vendors charge per seat — and outcome pricing may be “the cart driving the horse.” He cites a SaaStr Fund portfolio company whose outcome-based deal crossed $1m a year before the customer “quickly moved to a fixed contract.” He adds that as AI costs in B2B SaaS fall toward zero, the economic rationale for outcome pricing weakens; major vendors charge “$1–$3 per outcome-based resolution,” yet HubSpot and Box have largely abandoned the approach.18 “A pricing model is not a product. And a pricing model doesn’t make a mediocre product great.”18
The bull side has started to say the same thing. Bessemer’s 2026 playbook carries its own renewal warning. Soft ROI worked in 2025’s “AI adoption at all costs” environment, and “as pilots hit renewal, pricing must reflect actual value delivered, not promise.”2 The deals priced on promised outcomes meet that test at their first renewal cycle, with the measurement problem still open.
For a buyer, the same evidence reads as a term sheet. The meter needs a predictable cap, which is the anxiety that forced Salesforce’s pivot.7 The outcome needs a bound, or the vendor’s incentive runs toward Lago’s five hundred unusable meetings.24 And the definition of the billed outcome, and who verifies it, belongs in the contract, because as the market stands the verifier is the vendor’s own model.9
Zendesk’s own newsroom pages confirm the mechanism — outcome-based pricing and dual AI verification of each charged resolution89 — but do not state a per-resolution dollar amount. The ”≈$1.50 (committed) / ≈$2.00 (pay-as-you-go)” figures come from third-party contract teardowns and search summaries of 2026 pricing pages, with sources clustering at ~$1.50 but ranging $1.20–$2.00.9b Lemkin independently brackets the category at “$1–$3 per outcome-based resolution.”18 Treat the dollar figure as third-party-sourced, not Zendesk-official. Intercom’s $0.99 and Salesforce’s $2.00 / $0.10, by contrast, are confirmed on primary/official material.657
Among incumbents, the destination is already visible, and it is a hybrid. Bain’s analysis of 30-plus AI-enabled SaaS vendors found roughly 35% raised per-seat prices while bundling AI into existing tiers (Zoom, for one), and roughly 65% adopted hybrid models layering an AI usage or outcome meter on top of seat-based pricing (Adobe, Salesforce). 0% monetised AI purely as a standalone add-on, and 0% had fully transitioned to usage- or outcome-only pricing.1 The pure-outcome endpoint the venture commentary celebrates had, at the time of Bain’s study, zero incumbent adopters.
Bain also explains why incumbents stall at the hybrid rung. Moving off per-seat runs into three walls: internal hurdles, because vendors lack the telemetry infrastructure and shared pricing metrics across silos that a usage or outcome meter requires; commercial obstacles, because sales teams accustomed to seat-based conversations have to be retrained; and customer resistance, because procurement teams are unfamiliar with value-based purchasing that demands upfront investment before measurable savings materialise.1 The hybrid equilibrium is what buyer demand for predictability plus those switching costs jointly produce. Seats stay because the machinery for anything else is still being built, on both sides of the contract.
The venture enthusiasm and the incumbent data stop conflicting once the population is named. a16z observes that AI-native companies — Decagon, Cursor, ElevenLabs — employ usage-based, outcome-based, or hybrid pricing, while “established software companies adding AI features tend to maintain per-seat or bundled approaches.”3 Deloitte’s finding that 83% of AI-native SaaS companies already offer usage-based pricing is a finding about the natives;21 Bain’s 0% is a finding about the incumbents.1 The ladder is being climbed by the companies born on it, and leaned against by the companies that would have to rebuild billing, sales motion, and procurement story to climb it. Both the bull case and the bear case are right about their own population.
The advisory prescriptions converge on the same hybrid structure. McKinsey’s is the fullest: blend per-user subscription fees with consumption-based metrics, similar to Microsoft’s Copilot structure, and meter consumption beyond a capacity cap — tokens processed per day, week, or month.4 The others sit inside the same envelope. Gartner’s projection, cited by Deloitte, has at least 40% of enterprise SaaS spend shifting toward usage-, agent-, or outcome-based pricing by 2030, which leaves a majority still anchored on seats or subscriptions,21 and Bessemer’s founder formula is a base subscription plus usage or outcome tiers that “provide predictability while capturing upside,” illustrated as a platform fee at roughly twice delivery cost plus outcome credits.2
There is a forward-looking economic argument that outcome meters get cheaper to run as agents proliferate. An NBER chapter on the “Coasean Singularity” argues that AI agents “dramatically reduce transaction costs” by lowering “the costs of preference elicitation, contract enforcement, and identity verification,” which “expand[s] the feasible set of market designs.”17 Pricing tied to verifiable outcomes is exactly the kind of market mechanism that gets cheaper when verification is cheap, and Zendesk’s AI-verifier-on-AI-resolution architecture is that argument in microcosm. The same chapter warns that the transition “also introduce[s] frictions such as congestion and price obfuscation” and “novel regulatory challenges.”17
As agents take on consequential decisions, four reference frameworks have become the de facto governance stack. They differ on binding force and barely differ on substance. Accountability, traceability, risk management, and human oversight recur in each, and ISO/IEC 42001 is positioned as a complement to the NIST framework and the EU AI Act.26 Table 2 carries the inventory of publishers, dates, structures, and penalties.
The binding instrument is the EU AI Act (Regulation (EU) 2024/1689, in force 1 August 2024). It sorts systems into four risk tiers, bans the unacceptable tier outright, and subjects high-risk systems listed in Annex III to risk-management, documentation, data-governance, and human-oversight obligations, phased in over 6 to 36 months; penalties reach €35 million or 7% of global annual turnover for prohibited-practice violations.11 Annex III explicitly includes recruitment AI. The function at issue in Mobley v. Workday, algorithmic screening of job applicants, is work European law has already named high-risk.
The voluntary layer asks for the same things without the fine. NIST’s AI RMF, organised around four functions with governance cross-cutting the other three, names “Accountable and Transparent” among its seven characteristics of trustworthy AI.10 ISO/IEC 42001 turns the same substance into the first certifiable AI management-system standard, requiring AI impact assessment, lifecycle controls, and third-party supplier oversight.26 And beneath both sit the OECD AI Principles, the first intergovernmental AI standard, which make accountability an obligation of every AI actor “based on their roles, the context, and consistent with the state of the art.”23
| Framework | Publisher / year | Core structure | Binding? | Key figure |
|---|---|---|---|---|
| AI RMF 1.0 (AI 100-1) | NIST · 2023 | Govern · Map · Measure · Manage | Voluntary | 7 trustworthiness traits; GenAI Profile 600-1 (2024) |
| EU AI Act (2024/1689) | European Union · 2024 | 4 risk tiers: unacceptable / high / limited / minimal | Binding law | Up to €35m or 7% of global turnover |
| ISO/IEC 42001:2023 | ISO/IEC · 2023 | Plan-Do-Check-Act AI management system | Voluntary (certifiable) | First international certifiable AIMS |
| OECD AI Principles | OECD · 2019 (am. 2024) | 5 values-based principles + 5 recommendations | Soft law | First intergovernmental AI standard |
Frameworks are multiplying faster than the practice underneath them. Stanford’s 2025 AI Index records that in 2024 the OECD, EU, UN, and African Union all published responsible-AI frameworks, while the number of reported AI incidents hit a record 233 — a 56.4% jump over 2023.27 The same Index flags the gap the frameworks have not closed. Leaders in a McKinsey survey cite inaccuracy, regulatory compliance, and cybersecurity as their top responsible-AI risks, yet “not all are taking active steps to address them,” and standardised responsible-AI benchmarks remain sparse among major model developers.27
Every layer of the stack lands on one principle, and IBM states it in procurement terms. When IBM asks who in an organisation should be accountable for responsible AI outcomes, the three most common answers it hears are “no one,” “we don’t use AI,” and “everyone” — answers it describes as “none of which are correct and all are concerning.” It concludes: “The principle that accountability cannot be outsourced remains critical; while AI can enhance efficiency and enable innovations, human oversight, judgment, and accountability must remain central.”16 The contracts do not rescue the buyer either. IBM finds that “most software and cloud vendor contracts lack explicit commitments that make them accountable for providing responsible AI,” and some include disclaimers removing liability outright.16 The OECD Principles encode the same duty by role, supported by mandatory traceability of datasets, processes, and decisions across the lifecycle.23
Outcome pricing gives that principle a commercial edge, because outcome billing puts the vendor in charge of measuring the thing it bills for. Zendesk’s charged resolutions are verified by Zendesk’s own evaluation model,9 the teardown of those contracts shows the bill rising as the agent improves,9b and IBM finds the standard vendor contract makes no responsible-AI commitments at all.16 Who audits the verifier is one question wearing two labels — a pricing dispute waiting to happen, and an accountability gap already documented.
Whether exposure also reaches the vendor is being decided in Mobley v. Workday (N.D. Cal., No. 3:23-cv-00770-RFL). Lead plaintiff Derek Mobley submitted over 100 applications through employers using Workday’s screening system and was rejected each time, alleging the AI tools “unfairly penalize older candidates.”14 The suit names the vendor; the employers who used the screen are not the defendants. The EEOC’s April 2024 amicus brief argues that Workday can be directly liable under Title VII, the ADA, and the ADEA on three theories — as an employment agency, as an indirect employer controlling access to opportunity, and as an agent of employers to whom hiring authority was delegated.22 Each theory tracks the delegation. The more of the hiring function the employers handed to Workday’s tools, the more directly the statutes reach the company that ran them; ceded control, here, is the theory of liability. The EEOC also rejects the size defence outright: “there is no authority for the novel proposition that an entity can be too big to qualify as an employment agency… whether Workday performs those tasks for one employer or thousands of employers.”22
The court has let the case advance and grow. In July 2024 it allowed the agent theory to proceed; in May 2025 it granted conditional nationwide certification of the ADEA collective, covering applicants aged 40 and older denied recommendations through Workday since 24 September 2020.1314 Per Workday’s own filings, roughly 1.1 billion applications were rejected during the period, with the collective potentially numbering in the “hundreds of millions” — one of the largest ever certified in employment litigation.13 A later order, filed 6 March 2026, preserved the ADEA disparate-impact claim while dismissing the related state-law FEHA claims as pled, noting that “attending a historically Black college can be used by AI screening tools as a proxy for race.”12 Workday maintains the suit “lacks merit” and stresses the rulings are preliminary.14
The lesson has two halves resting on different evidence. The buyer-side half — deploying a vendor’s AI does not move accountability off the deployer — rests on IBM’s contract finding and the OECD’s role-based duty.1623 What Mobley tests is the vendor-side half: whether the exposure created by a delegated, regulated function reaches the company that built and operated the screen. Whatever the outcome, the case does not relieve the deployer; the accountability that IBM and the OECD locate with the buying organisation sits where it was.
These are rulings on motions to dismiss and a conditional collective certification — procedural milestones letting claims proceed, not a final finding that Workday discriminated.1314 The “agent of the employer” holding traces to the July 2024 order (documented via the EEOC amicus and law-firm analyses);2213 the saved court order (Doc. 267, 2026-03-06) is a later ruling addressing the ADEA and FEHA claims, not the original agent-theory order.12 The ~1.1 billion-applications figure is from Workday’s own filings as reported by counsel,13 not an independently audited number. Stated here as the litigation status, not as a determination of liability.
The variable running underneath all of this has a formal literature, and the literature treats it as governable. Feng, McDonald, and Zhang define a five-level spectrum of agent autonomy by the user’s role — operator, collaborator, consultant, approver, observer — ranging from direct control to light supervision, and propose “AI autonomy certificates” to govern agent behaviour, calling autonomy “a double-edged sword.”19 The spectrum maps onto the familiar human-in-the-loop / human-on-the-loop / human-out-of-the-loop taxonomy. Operator and approver keep a human in or on the loop; observer approaches out-of-the-loop.
The risk claim attached to the spectrum is blunt. Mitchell, Ghosh, Luccioni, and Pistilli of Hugging Face argue, from the title down, that fully autonomous agents should not be developed, on the principle that “risks to people increase with the autonomy of a system: The more control a user cedes to an AI agent, the more risks to people arise.”15 Their preferred posture is constrained, semi-autonomous systems with human oversight.15 Two deployable postures follow. High autonomy on a narrow scope lets a tightly bounded task run end-to-end; broad scope at low autonomy permits wide-ranging action behind human verification gates. The danger zone pairs broad scope with high autonomy and no gate.
Two safe postures, one danger zone — agency × scope
A tightly-bounded task run end-to-end (resolve this ticket; reconcile this account). Safe because the blast radius is small. This is where outcome pricing works — discrete, measurable, contained.24
The danger zone. Wide-ranging autonomous action with no verification gate. Risk rises with ceded control;15 autonomous pricing agents can even sustain collusion “without agreement, communication, or intent.”20
A copilot that suggests and waits. Human-in-the-loop by construction. Low risk, but also soft ROI and weak pricing power — priced per seat.2
Wide reach behind human-on-the-loop verification gates (the approver role).19 Broad usefulness made safe by review — the posture most enterprise deployments should default to.
Autonomy is a governable lever set independently of capability;19 verification gates convert broad scope from danger zone to deployable
The same lever shows up in the pricing data. The cell where outcome pricing works in practice — Intercom’s resolutions, Zendesk’s verified resolutions — is high agency on narrow scope: discrete, contained, measurable tasks, exactly Lago’s “rare exception” of clear, countable outcomes.24 The cell that frightens regulators is high agency on broad scope. NBER work on AI-powered trading shows reinforcement-learning agents “autonomously sustain collusive supra-competitive profits without agreement, communication, or intent” — price-fixing emerging with no instruction to collude, degrading market efficiency.20 That is the empirical case for verification gates on any agent holding both broad scope and transactional authority. Set the dial deliberately, and the rest follows from the setting — the charge metric that fits, the margin the oversight costs, and the liability the delegation carries. Pair high agency with narrow scope, or broad scope with low agency behind gates, and price, govern, and accept liability accordingly.
Whether outcome-based pricing durably “sticks” or reverts to hybrid consumption is unresolved in the corpus. The bull case (a16z, Bessemer) and the bear case (Lemkin, Lago) both have evidence — confirmed live deployments on one side,65 a documented reversion to fixed contract and the “better AI costs more” perverse incentive on the other.189b Bain’s finding that 0% of incumbents had fully transitioned1 suggests hybrid is the stable equilibrium for now, but the sources do not contain the multi-year renewal-and-retention data needed to call the endpoint. The related governance question — whether vendor liability under Mobley survives to final judgment and generalises beyond hiring — is likewise open.