AI & Software
The Agentic Tax
Why flat subscriptions are breaking under agentic software, why per-seat, pure usage, and outcome pricing each fail alone, and the three-layer model that actually holds.

In June, two of the most widely used AI coding tools quietly changed how they charge. GitHub moved Copilot's agent features onto usage-based credits. Anthropic announced its command-line tool would follow two weeks later. Neither company framed it as a retreat, but that is what it was: a retreat from the flat monthly subscription, the pricing structure that built the entire SaaS industry, because the economics underneath it had stopped working.
What changed is what the software does when you are not looking. A conversation with a language model is one inference call, with a cost you can predict to the penny. An agent given the same goal makes dozens of calls, sometimes hundreds. It reasons, retrieves, calls a tool, checks its work, and reasons again, and because each step carries the accumulated context of every step before it, the calls get more expensive as the chain gets longer. Researchers at Stanford's Digital Economy Lab measured the difference and found agentic tasks consuming on the order of a thousand times the tokens of an ordinary chat. A thousand times is not a margin problem you can absorb. It is a different cost structure wearing the same product.
Most of the companies selling agentic software have not absorbed this, and the reason is mundane. They priced from the test environment. In testing, the data is clean and the task is bounded, and a workflow resolves in five steps. In production, the customer's data is a mess, the task has edge cases nobody anticipated, and the same workflow takes thirty steps, or fifty. The cost per task that the founders modeled before launch turns out to describe the easiest tenth of their traffic. A simple support ticket costs three cents of inference to resolve. A genuinely tangled one, with retrieval and tool calls and a long reasoning chain, can cost a dollar fifty. If the price is 99 cents per resolution, which is what Intercom charges, the easy tickets are funding the hard ones, and the vendor has taken on the cost variance of every customer's worst problems without ever deciding to.
Why each familiar model fails
It is worth being precise about why the familiar pricing models each fail here, because they fail for different reasons. Per-seat pricing assumes the amount of work scales with the number of people, and agents sever exactly that link; one employee can now set off thousands of compute events in an afternoon. Pure usage pricing has the opposite defect: it makes the customer's bill grow in proportion to how well the product is working, so the better the agent automates, the more painful the invoice, and the customer's finance team becomes the product's most motivated adversary. And outcome pricing, the model everyone in AI claims to be moving toward, works only while the cost of producing an outcome stays inside a narrow band. Agentic workflows are precisely what blows the band open.
Stop choosing one model
The conclusion I keep arriving at with the companies I advise is that no single model carries the load, and the answer is to stop choosing. A workable price for an agentic product has three parts. A flat platform fee at the bottom, because buyers need a number they can budget. We call this predictability of cost in pricing speak. A usage component in the middle, per workflow or per completed task, so that revenue grows with consumption rather than with seats that no longer mean anything. And an outcome component on top, a fee tied to a result both sides can count, which captures the upside but only once the relationship can bear it. The sequencing matters more than the arithmetic. The platform fee is there from the first day. Usage charges become meaningful as adoption deepens. The outcome layer comes last, because it depends on measurement infrastructure and on trust, and neither exists at the start of a contract.
The ACE test
Whatever metric sits in those upper layers has to survive three questions. Can the buyer verify what they are paying for, or is the agent's work a black box? Do the margins hold as usage grows, or does cost scale with context while price stays flat? And can the company actually meter and bill the thing without a year of engineering? I shorthand this as the ACE test, for auditability, cost-margin fit, and ease, and in my experience the first question is the one that kills. The moment a buyer cannot audit what the agent did, every conversation about outcome pricing turns into a conversation about discounts.
Pressure from the buyer's side
There is one more pressure building, this time from the buyer's side of the table. Gartner projects that by the end of this year, most large IT services contracts will carry performance clauses for AI: floors on resolution rates, accuracy thresholds, and price adjustments that trigger automatically when the system underperforms. Procurement teams have learned to write these clauses faster than vendors have learned to price for them. A vendor that has not built the performance floor into its own pricing will find it added in legal review, on terms it did not draft.
None of this means the agentic economy is unprofitable. It means the period in which nobody had to know what their software actually cost to run is over. For two decades, the marginal cost of serving a SaaS customer was close enough to zero that pricing could be pure psychology. Agents have ended that. The number on the pricing page now has a real number underneath it, and the companies that thrive will be the ones that found out what theirs was before their customers did.
See how we apply empirical pricing research in practice: Explore the C.O.R.E. roadmap & 78-artifact catalog