Economics · August 2026 · 4 min read
Your AI bill is becoming a payroll
Agentic AI does not consume like software. It consumes like labour: variable, workload-driven, and capable of tripling while you are in a budget meeting. Managing it takes payroll instincts, not licensing instincts.
Enterprise software used to be bought like office space: so many seats, so much a year, predictable to the pound. AI began that way too, and then agents arrived. An agent does not sit in a seat. It consumes intelligence by the token, in quantities that depend on the work it is given, the way it was built, and how enthusiastically your organisation discovers new uses for it. In other words, it behaves less like software and more like labour, and the finance function has noticed.
The economics underneath are moving in two directions at once. The price of the unit has collapsed: Stanford's AI Index recorded the cost of querying a GPT-3.5-class model falling more than 280-fold in under two years, and Epoch AI finds prices falling between ninefold and 900-fold per year depending on the task. Yet many AI bills keep growing, because agentic workflows multiply the number of model calls behind every piece of work faster than the unit price falls. Budgets written in the per-seat era deserve some sympathy: planners were forecasting a category that had not yet arrived.
The cap reflex
The instinctive response is the one finance always reaches for first: caps. Per-user token limits, frozen model tiers, approval gates on the expensive models. Caps stop the bleeding, which is why they happen, but as a policy they have a defect: they throttle consumption without knowing what the consumption was buying. The agent that was quietly saving a team two days a week hits the same ceiling as the experiment nobody remembers starting. A cap is a budget instrument pretending to be a value judgement.
Payroll instincts
Organisations already know how to manage a large, variable, value-producing cost, because they run one: payroll. Nobody manages payroll by capping salaries at random and hoping. They manage it by knowing who is doing what, what it produces, and whether the role is worth the cost. AI spend deserves the same instincts, which means the first investment is not a better model but visibility: which workflows, teams and agents consume what, at what cost, producing which outcomes. The attribution is genuinely harder than it sounds, because every provider structures its billing data differently, and the trail from a bill back to a workflow is often lost below the account level. It is also exactly why the discipline pays: once cost is attributed to outcomes, provisioning stops being a blunt limit and becomes an allocation decision. The expensive model goes where it earns its keep. The cheap model does the routine work. The agent nobody can justify is retired, on evidence.
The cheap model is a hypothesis
That allocation depends on knowing which model is the cheap one, and a price list will not tell you. Agentic work changes the arithmetic underneath it: the model is called repeatedly, carries a growing transcript back into every turn, and retries what it gets wrong. A study of eight frontier models on SWE-bench Verified, How do AI agents spend your money?, found runs of the same task differing by as much as thirty times in total tokens, and some models drawing more than 1.5 million tokens more than others across the same benchmark. A lower headline rate does not absorb gaps of that size.
The same work found that spending more did not buy accuracy. It tended to peak at an intermediate cost and flatten above it. The agents were also poor witnesses to their own consumption, underestimating what a task would take, so asking the system to forecast its own bill is not a control. Cheap and expensive are findings rather than labels a pricing page can assign, and they belong to your workload rather than to a benchmark. Public indices have started reporting average API cost per task alongside capability, which is the right unit. The number that governs your allocation is that one, measured on work your own people would accept as finished.
The part that outlasts the models
There is a temptation to defer all this until the technology settles. It will not settle soon, and that is the argument for acting now rather than against it: models, harnesses and pricing schemes will keep churning, but the need to see what intelligence you consume, what it costs and what it returns is permanent. Observability built today transfers intact across every model swap that follows. It may be the only part of a 2026 AI stack you can confidently expect still to be pulling its weight in 2030, and it is the difference between an AI bill you discover like a leak and one you manage like a payroll: understood, allocated, and working for you.
Written by Anthony Smith, Chief Technology Officer, Ballista.
Talking beats reading.
If your AI spend is growing faster than your understanding of it, describe what you are running. We will tell you what visibility into it would involve, and what it would let you decide.
Your message goes to the people who would do the work.
