Essay 04

The Agentic Operating Model

When an agent is the default actor for incoming work and an orchestrator decides which one, the operating model is the whole game — and a wrong route in a regulated domain is a compliance event, not a bad answer.

Enterprise spending on generative AI grew from $1.7 billion in 2023 to $37 billion in 2025.9 Most of that money is going to the wrong layer. Firms buy models, wire them into a function, and call it transformation. The consultancies’ own guidance puts the weight elsewhere. BCG tells pacesetters to follow a 10-20-70 rule, with ten percent of the effort on algorithms, twenty on technology and data, and seventy on people and process.1 McKinsey’s published summaries carry the same proportion, framing roughly seventy percent of a successful AI transformation as people and process.2 Two caveats belong next to those numbers. Both rest on secondary captures, as the endnotes disclose, and BCG’s rule is a prescription about where effort should go rather than a measurement of where it paid off; two consultancy framings are weaker evidence than two independent measurements would be. Read at that weight, they still point one way. The part everyone is buying is the cheap part, and the part that decides whether you win is the part few firms are designing.

So design it. I run an operating model built on that premise, in a regulated domain, and what follows is the pattern I would defend: an engine in which the default actor for incoming work is an agent, an orchestrator decides which one, and the accountable humans stand where the regulator will come looking. The case for structure over technology is argued in a companion essay; here I describe what you build.

Agents are being bolted onto an operating model built for human workers

The dominant mistake is organisational. Deloitte found that eighty-four percent of companies have not redesigned a single job to fit AI, even while expecting heavy automation.3 They are bolting autonomous agents onto an operating model built for human workers and waiting for the gains. The gains do not come, because the old model was engineered, in Deloitte’s words, “for control, with stable processes, clear handoffs, and predictable outcomes.”3 An agent does not respect stable processes. It works in loops and fails in unfamiliar ways, and it moves at a volume no human handoff was designed to absorb. You cannot run new physics on old plumbing.

The same logic says efficiency is the wrong target. Aim at it directly and you optimise the shape you already have, hard-wiring the assumption that humans do the cognitive work into a faster version of the old org. Efficiency arrives as a by-product of AI-native design. You design for quality and for the new division of labour the capability makes possible, and the saving falls out the back. Chase it head-on and you will miss it.

The default actor is an agent, and an orchestrator decides which one

The organising claim of the design is that the default actor for any incoming unit of work is an agent, and an orchestrator decides which one. Three domains that would once have run as three departments with three stacks run instead on one engine.

An incoming question today arrives, a human reads it, decides whose problem it is, and routes it. In an AI-native model the agent is the default handler, and the routing itself becomes a system: a layer that reads the work, decomposes it if it spans domains, and sends each piece to the agent built for it. The literature has caught up to this. MIT Sloan’s working definition of agentic AI is already plural, “systems that incorporate multiple, different agents… orchestrating a task together,“4 and the trade press now describes the orchestration layer as a control plane sitting between the raw models and the applications, handling routing, task decomposition, memory, and policy. On that account the fundamental enterprise challenge is coordination, making many agents “operate reliably, securely, and coherently within enterprise constraints,” rather than model access or talent.5

Why separate agents under one orchestrator, rather than one general model that does everything? A pricing question and a compliance question are the same kind of work, cognition over context, but they carry different knowledge bases, different risk profiles, and different regulatory exposure. You want to tune and contain them independently, with a tight leash on the regulated one and a longer one elsewhere. That is the sense in which the engine is one and the agents are many. A single orchestrated system holds contained specialists, in place of three departmental stacks. California Management Review, reasoning separately, lands on the same lean. Agentic enterprises “increasingly deploy multiple specialized models,” because specialisation reduces hallucination and sharpens the alignment between an agent and what it has been asked to own.6

Routing is a judgment, and the exposure concentrates at the orchestrator

The hard problem in this engine is the orchestrator’s decision logic; the agents are the easier half. Routing is itself a judgment. In an ordinary domain a wrong route produces a bad answer, annoying and recoverable. In a regulated domain a wrong route is a compliance event. Send a question that needed a licensed human to an agent that closed it alone and the failure changes category, from quality problem to the kind of event that gets reported and examined. This is what separates a clever demo from an operating model a large regulated firm can actually run, and it is why the decision logic has to be designed as a first-class artefact, too consequential to leave inside a prompt.

The first piece of that artefact is decision rights. Deloitte states the rule plainly. For each kind of work, the design says “what agents can decide on their own, what requires human sign-off, and what needs to be escalated.”3 Get that triad right and the system runs. Get it wrong and you have either a bottleneck, where everything escalates, or a liability, where nothing does. In the drawn model below, the triad lives at the boundary between the agent layer and the human nodes; it is the rule the orchestrator consults before a route is committed, and the escalation layer enforces afterwards. Two refinements matter in practice. Decision rights are a firm-level question, what an agent may commit the business to without a human signature, and they sit apart from any sub-agent permission scheme; the answer is partly regulatory and partly risk appetite. And the rule is set per workflow. Deloitte’s own advice runs the same way, with leadership deciding which systems “can be made faster with less human oversight” and which should be deliberately slowed so that the oversight is sufficient.3

Deloitte also names the strangeness underneath the whole problem. Agents are “neither capital nor labor.”3 A firm knows how to fund and depreciate capital, and it knows how to hire, manage, and hold labour accountable. An agent fits neither category, which is exactly why decision rights, liability, and performance ownership go fuzzy around it. The operating model exists to resolve that ambiguity with structure, because the accounting categories will not do it for you.

The second piece is what I call the operating spine: the case state, context, history, handoff packaging, and audit trail that travel with a unit of work as it moves through the engine. The spine is what lets a tiny human layer serve a large base, because it does the context-assembly that would otherwise be human labour. When an escalation arrives with its full history attached, the human can apply judgment in minutes. When it arrives cold, the human re-does the work the system was supposed to absorb.

The third piece is the misroute itself. A misroute you cannot detect is the dangerous kind, so routing decisions get logged and audited like any other judgment the firm makes. The audit trail in the spine is what lets you find a wrong route after the fact, and the escalation thresholds are what catch the ones that surface in flight. An orchestrator without both is a router; with both it is an operating model.

The vertical columns become tiers around a shared core

Picture the conventional org as a set of vertical columns. Each function, whether sales, operations, or risk, stands as its own pillar with its own people, its own systems, its own slice of the data. The columns sit side by side and coordinate across the gaps between them. That shape is so familiar it reads as natural, and it is an artefact of the constraint that a function’s knowledge lived in its people’s heads and its filing cabinets.

AI inverts it. When data and infrastructure are shared, the columns fall over and become tiers. A single foundation runs underneath every function, and the functions grow out of it instead of sitting beside each other. The classic operating model decomposes a firm into process, organisation, location, information, suppliers, and management systems, six vertical concerns.7 The AI operating model collapses the data and information concerns into one shared floor and stands everything else on top of it. A function no longer owns its stack; it draws on a common one.

I draw the result as a core, read from the centre outward.

An engraved chandelier: concentric rings read from a central data core outward through a foundation ring of infrastructure, governance, evaluation and systems, then an agent layer, accountable human nodes outside it, a dashed compliance crossing, and an outer audience ring. Each spoke carries work outward and a return current of learning back to the core.
Fig. 1 — The operating model as a layered core. Data sits at dead centre, wrapped in a foundation of infrastructure, governance, evaluation and systems drawn with weight, because it is the part nobody looks at and the part everything stands on. The agent layer is talent grown into each function; the accountable humans sit outside it, one per spoke. A dashed compliance crossing separates the system from the audiences it serves, and every spoke runs two lines — work outward, learning back.

Read it from the middle. Data sits at dead centre, the asset everything else depends on. Around it sits a foundation ring of infrastructure, governance, evaluation, and systems. I draw this with depth, as a well rather than a flat disc, because in most operating models it is the layer nobody looks at and the layer everything secretly rests on.

Then the agent layer. Agents here are a layer of talent grown into the system, present inside each function rather than bolted to its side, and I mean talent as a category claim. Humans and agents are both categories of contributor, distinguished by their suitability for different jobs, and the work of the design is to distribute cognitive load across them. If calling agents talent sounds sentimental, look at the assumption underneath the reaction: that words like talent and team belong only to humans.

Then the humans, and this is the structural claim that matters most. The accountable humans sit outside the agent layer. One human owns each spoke, and one human may own several. Outside does the work that above used to do. The humans direct and supervise the agents beneath them and they carry the accountability, and no hierarchy of worth is drawn into the picture. What the position changes is when accountability attaches. In a traditional team accountability is continuous, owned start to finish by whoever holds the work. Here it crystallises immediately before delivery, at every audience crossing, internal or external, so it recurs at many points and is distributed across the spokes. That is a new drawing of an old logic; good organisations already run on juniors handing work up and seniors delivering it out, and this model makes the logic structural.

Past the humans sits the compliance crossing, dashed because everything that reaches an audience must pass through it. The reason it exists is a fact about regulators: they march up the chain until they find a human to hold responsible. As long as a single human is anywhere in the system, the legal environment will hold that person accountable for the system’s outputs, and the limit case makes the point cleanly, because a one-human, many-agent company is still a company with a fully accountable human. The crossing is a ring because every spoke passes through it, and I draw it in the same gold as the return current because the crossing is part of what makes the output worth anything to the audience receiving it. Compliance sits in the drawing as part of the value. Beyond it lies the audience ring, internal and external, the people each function serves.

And running through all of it is the element most diagrams omit. Every spoke is two lines. Work flows outward; data, signal, and learning flow back to the core. Even that pairing understates the traffic, because the flow is a loop. A unit of work iterates, system to agent, agent to human, human back to agent, each loop refining the last, and the cognition is jointly produced inside that loop. The human’s job in the loop is orchestration, plus the interventions only a human can make: judgment calls, audience crossings, accountability at delivery. Many of the exchanges are genuine thought partnership, which is why the framing in which the human owns the cognitive layer fails as a description of what actually happens.

Drawing the return line forces a question at every node: how does feedback actually get served here? Answer it node by node and the operating model becomes a connected system in which every node is a point of data capture, and the learning compounds in the institution instead of walking out in people’s heads. The trade-press version of the same point runs in the negative. Without a shared backbone, agents “cannot share state, learn from one another, or operate as part of a larger system.”5 This is a stronger claim than the usual data flywheel, which describes a product improving as more people use it. Here the org itself becomes an instrument that records what it learns.

A second drawing reaches the same layers and disagrees about the hub

Whether this structure is just my drawing is a fair question, and there is a check. California Management Review published a four-layer Agentic Operating Model that I had not seen when I drew mine.6 Its cognitive layer of specialised domain models corresponds to my agent layer. Its control layer carries guardrails, confidence thresholds, and escalation, the tier above the agents that I take up below. Its governance layer requires that “each agent is associated with a clear business owner, a defined risk profile, and documented decision boundaries,“6 which is one accountable human per spoke restated in academic register. Two people drawing the same structure from different starting points is the closest thing this young field has to corroboration.

The corroboration is worth having only with the disagreement attached. On the coordination layer, the routing tier, CMR argues for the opposite topology. It describes organisations shifting away from centralised hub-and-spoke orchestration toward decentralised swarms in which “agents operate via decentralized local rules and shared goals without a single point of failure.”6 My engine is a hub, one orchestrator that decides, and I hold the hub deliberately for the compliance reason already on the table. A swarm removes the single point of failure and removes the single point of answerability with it. In a regulated firm the routing decision is the judgment an examiner will eventually ask about, so you want one place where that judgment is made and owned, with the audit trail attached; a longer march up the chain is the last thing to hand an examiner. Outside regulated domains the swarm case is stronger, and CMR may well be right about where the general population of enterprises goes. Inside one, the hub is a design choice made for the same reason the compliance crossing is drawn as a ring.

Oversight migrates from approving each action to watching for the andon

If the human-on-the-loop pattern feels new, it is seventy years old. Toyota called it jidoka, “automation with human intelligence,” sometimes rendered automation with a human touch. A machine runs on its own, detects an abnormal condition, stops itself, and signals for a human.8 The loom that stopped when a thread broke let one operator run many machines instead of watching one, and that is the whole economic argument for an agent operating model, made in textiles in the early 1900s. Autonomous operation, with an automatic stop-and-escalate on the exception, lets a small human layer supervise a large base of work.8 The andon cord, which halts the line and summons attention when pulled, is the direct ancestor of the escalation layer in an agent system.

The modern version of the move is the migration from human-in-the-loop to human-on-the-loop. Early on, a human approves each agent action. That does not survive contact with volume. As the work scales, oversight shifts: humans “define objectives, constraints, and escalation thresholds, while agents operate independently within those boundaries,” intervening on the exception rather than the instance.6 Same move as the loom. The human stops watching every action and starts watching for the andon.

Which makes the escalation layer the hinge of the whole engine, and the tier most likely to be under-built, because it hides inside “the chatbot” and looks like a feature. Its job is to recognise when a question exceeds what the agent should close alone and to hand it up warm, with the operating spine’s full context attached, to the human who owns that spoke. Build it as a distinct, observable, independently tunable layer above the cognition. It is simultaneously your error-and-recovery architecture and the front door for your decision rights, and if you treat it as plumbing it will fail you at precisely the moments that matter.

The operating layer is mostly unbuilt, and the failure data cuts both ways

Almost none of this is built yet. Menlo Ventures, surveying enterprise deployments, found that only sixteen percent qualify as true agents; the rest are fixed-sequence or routing-based workflows wearing the word.9 Every other published measure points the same direction, with few pilots reaching production at scale,10 governance maturity trailing adoption badly,11 and the eighty-four percent redesign number from the opening sitting underneath all of it.3 Figure 2 carries the four measures with their sources.

The reality gap in agent operating models Four bars. Only 16 percent of enterprise deployments are true agents. Only 11 to 14 percent of pilots reach production at scale. Only 21 percent of organisations have mature agentic-AI governance. And 84 percent have not redesigned jobs to fit AI. 100% 0% Deployments that are true agents 16% Pilots reaching production at scale 11–14% With mature agentic-AI governance 21% Have NOT redesigned jobs for AI 84%
Fig. 2 — The operating layer is mostly unbuilt. Three measures of how far adoption has actually progressed sit low; the one number that runs high is the share of firms that have not redesigned the work. Sources: Menlo Ventures 2025 (true agents); FifthRow / 2026 aggregation (production at scale); Deloitte 2026 (mature governance); Deloitte 2026 (jobs not redesigned). The production figure rests on secondary aggregation — see the sources note.

I read that chart as the location of the advantage, and the reading has to be argued, because the same four numbers support a darker one. The hostile version says operating-model redesign is hard and mostly fails, and that the prescription in this essay is the thing the data shows not working. FifthRow’s aggregation of 2026 research puts pilots reaching production at scale at eleven to fourteen percent, which means roughly six in seven fail to deliver durable value.10 The figure rests on secondary aggregation, as the endnote discloses, but take it at face value and it needs answering.

The answer runs through the redesign number. With eighty-four percent of firms having redesigned no jobs at all, the pilots that failed were overwhelmingly pilots inside unredesigned operating models, agents bolted on, in Deloitte’s framing, to structures built for human workers.3 The failure data therefore reads more naturally as evidence about bolting-on than as evidence about redesign, because the redesign has mostly never been attempted. What the data cannot do is settle the question. No one has published a controlled result showing that the redesign itself, as distinct from AI adoption in general, produces the outcomes, a limit I return to at the close, so the opportunity claim stands on this argument rather than on proof. And the opportunity is priced. The same aggregation puts a number on the constraint this design treats as central: regulatory compliance adds twenty to fifty percent to orchestration budgets.10

The economics pull in this direction regardless of who wins the design race. Foundation Capital’s “service as software” thesis names the prize. When software stops being a tool you operate and becomes the worker that delivers the outcome, the addressable market shifts from software budgets to the roughly $4.6 trillion services-and-labour market, and pricing moves from the seat to the outcome.12 A firm that sells completed outcomes has to be organised as a delivery engine with supervision, which is to say it has to have built the operating model first. The pricing model and the org shape are the same decision seen twice.

The models underneath stay swappable, and staying swappable costs something

One design decision inside the engine deserves separate treatment, because most operating-model writing skips it: what the orchestration layer assumes about the models underneath. The market-share data argues against marrying a vendor. In Menlo’s survey of enterprise LLM spend, Anthropic holds forty percent, up from twelve percent in 2023, while OpenAI holds twenty-seven, down from fifty percent in 2023.9 An orchestration layer hard-wired to the 2023 leader would now be three years into an expensive divorce. Enterprises appear to see this coming; in FifthRow’s aggregation, seventy-six to eighty-one percent express concern over proprietary dependencies in agent memory, model integration, and orchestration tooling.10

The design response is a model-abstraction layer. Treat the models as commodity components behind it, and measure the cost of swapping one out as a first-class metric of the engine’s health. The honest complication is that this pulls against where the moat lives. The defensible asset in this design is the proprietary orchestration, the spine, the routing logic, the escalation thresholds, and the accumulated learning in the core, built around rented and swappable models. The more of that you build, the more specific it becomes to how your current models behave, and the more the agnosticism costs you. I hold this as an unresolved tension, priced by the swap-cost metric, because a design that pretends the tension away has usually just chosen lock-in without noticing.

What I am not claiming

Two limits stand, and both are load-bearing.

The pattern is scale-invariant in intent. The same structure describes a small team where each node is a person and a multi-thousand-person organisation where each node is a function, with only the granularity of a node changing. I believe this, and I have built the small version. I have not built the large one, and no one I can cite has published a controlled result showing the redesign itself, as distinct from AI adoption in general, produces the outcomes. There is no peer-reviewed quantification of operating-model redesign yet. The numbers in this essay measure adoption and governance maturity; they do not prove that drawing the core the way I draw it pays off. That is why this is an essay rather than a research finding, and why the convergence with CMR corroborates the shape without proving the payoff.

The second open problem is apprenticeship. If agents absorb the entry-level work that used to be how juniors learned the craft, the on-ramp disappears, and the structure does not solve this. My current thinking is that juniors enter somewhere new, as apprentice curators of the knowledge the agents run on and as apprentice handlers of escalations under supervision, but I hold that loosely and take it up properly in the essay on the AI-native contributor. I would rather leave the question standing than paper it over; the redesign has real costs, and this is one of them.

What endures in the design is the structure — the radial shape, the return current, agents as talent, the orchestration role, the data foundation. The division of labour drawn inside it is a snapshot, and the line between human and agent moves toward the agent over time. “Fewer, higher-leverage humans” shows up structurally as one human owning more spokes, with the shape itself holding. The discipline this demands of whoever runs it is bilevel and has to be held at once: keep questioning whether the foundational design still applies as the technology changes, and keep iterating the current instance so it does not go stale. The two pull against each other. Loyalty to a fixed model fails, and so does endless iteration that never asks whether the whole frame still fits. Holding both is most of the job, and the deliverable is the repeated act of designing against a moving capability, with the drawing as one act of that design.

Sources

  1. Boston Consulting Group, The Leader's Guide to Transforming with AI (the 10-20-70 rule), 2025–2026. https://www.bcg.com/featured-insights/the-leaders-guide-to-transforming-with-ai. Secondary capture: bcg.com returned HTTP 403 to automated retrieval; the 10-20-70 figures are quoted identically across multiple independent sources, which raises confidence, but the primary page was not directly retrieved.
  2. McKinsey & Company, Rewired: A Manual for Digital and AI Transformations (Lamarre, Smaje, Zemmel) and companion gen-AI insights, 2023–2025. https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/rewiring-for-the-era-of-gen-ai. Secondary capture: every mckinsey.com endpoint timed out on automated retrieval; the ~70% framing is drawn from McKinsey's own published summaries plus corroborating coverage and is not quoted verbatim.
  3. David Mallon, Brad Kreit and Natasha Buckley, "Rethinking operating models for humans with agents," Deloitte Insights, 2 April 2026. https://www.deloitte.com/us/en/insights/topics/talent/operating-models-for-humans-ai-agents.html.
  4. MIT Sloan (Ideas Made to Matter), "Agentic AI, explained," 2025, citing Sinan Aral and work by Horton et al. and Kellogg et al. https://mitsloan.mit.edu/ideas-made-to-matter/agentic-ai-explained.
  5. Rajashree Goswami, "The Agentic Orchestration Layer: The Missing Piece in Enterprise AI Stacks," CTO Magazine, 30 January 2026. https://ctomagazine.com/agentic-orchestration-layer-enterprise-ai-stacks/.
  6. Sandeep Saini, "Governing the Agentic Enterprise: A New Operating Model for Autonomous AI at Scale," California Management Review (Berkeley Haas), 20 March 2026. https://cmr.berkeley.edu/2026/03/governing-the-agentic-enterprise-a-new-operating-model-for-autonomous-ai-at-scale/.
  7. "Target operating model" (POLISM framework; Campbell/Ashridge operating-model lineage), Wikipedia, accessed 19 June 2026. https://en.wikipedia.org/wiki/Target_operating_model.
  8. Lean Enterprise Institute, "Jidoka" (Lean Lexicon), accessed 19 June 2026. https://www.lean.org/lexicon-terms/jidoka/.
  9. Tim Tully, Joff Redfern, Deedy Das and Derek Xiao, "2025: The State of Generative AI in the Enterprise," Menlo Ventures, 9 December 2025. https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/.
  10. FifthRow, "AI Agent Orchestration Goes Enterprise: The April 2026 Playbook," April 2026. https://www.fifthrow.com/blog/ai-agent-orchestration-goes-enterprise-the-april-2026-playbook-for-systematic-innovation-risk-and-value-at-scale. Vendor aggregation: the 11–14% production figure is attributed by FifthRow to third-party research (digitalapplied.com) not directly retrieved; treat as an industry estimate.
  11. Andy Bayiates, "Agentic AI is scaling faster than guardrails," Deloitte Insights, 24 April 2026, drawing on Deloitte's State of AI in the Enterprise (survey of 3,235 leaders across 24 countries). https://www.deloitte.com/us/en/insights/topics/emerging-technologies/ai-agents-scaling-faster.html.
  12. Foundation Capital, "AI Leads a Service-as-Software Paradigm Shift," 2024–2025. https://foundationcapital.com/ai-service-as-software/.