Platforms and tools
A cap protects the number. It does not tell you what the number was buying
Agentic AI consumes tokens the way an operation consumes labour: variably, and in proportion to the work it is given rather than the seats it occupies. Governing that spend well means building the instrumentation to see what each workflow, team and agent actually consumes and what it produced, before deciding what it should be allowed to cost.
A cap is not a budget
When AI spend starts moving faster than the finance function expected, the reflex response is a limit: a frozen model tier, an approval gate, a hard ceiling per user or per team. In June 2026, reports described Walmart replacing unlimited use of its internal AI tools with fixed token budgets, and Uber setting a monthly spending cap in the region of $1,500 per person. Moves of that kind are a rational first response to a bill that has stopped behaving predictably, and they share a defect: a cap throttles consumption without knowing what the consumption was buying. The agent quietly saving a team two days a week hits the same ceiling as the experiment nobody remembers starting, because a cap set before the measurement exists cannot tell them apart.
The alternative is not a bigger cap. It is building the instrumentation that turns “we are spending too much” into “this workflow is spending this much, producing this, and here is what happens if we change it”. That is a capability we build and hand over as part of an AI estate, not a single dashboard bought once and left to drift out of date as the estate changes underneath it.
What we instrument
The starting requirement is unglamorous and easy to skip: every call to a model, wherever it originates, tagged back to the workflow, team and agent responsible for it, not merely to an account or an API key. Billing data alone will not give you this. Providers structure it differently, and the trail from an invoice line back to the piece of work that generated it is usually lost well below the account level, so the tagging has to happen at the point the call is made, inside the harness or the routing layer that sits between your systems and the model.
- ConsumptionTokens, calls and retries, attributed at the source
- Every model call tagged to the workflow, team and agent that made it, captured where the call happens rather than reconstructed later from a bill.
- OutcomeWhat the consumption produced
- The work the call was part of and whether it was accepted, corrected or discarded, so a token figure always sits next to what it bought.
- SignalRefusals, reroutes and degraded responses
- When a request is answered by a smaller or more heavily filtered model than the one it was sent to, or fails and is retried, recorded as an operational event rather than absorbed silently into the next call's cost.
None of this is a one-off audit. Consumption and outcome have to be captured continuously, because a system evaluated once at launch tells you nothing about the workflow that changed shape three months later, or the agent that started being used for a task nobody designed it for.
Setting a number without guessing
Once consumption and outcome are visible together, setting a budget stops being a guess about next quarter and becomes an allocation decision made on evidence. The expensive model goes where it is earning its keep. The cheap model does the routine work, once “cheap” has been established for that workload rather than assumed from a price list, because agentic work retries, carries a growing transcript into every turn, and can vary in total tokens by a large multiple between runs of what looks like the same task. The agent nobody can justify is retired, on evidence rather than suspicion.
The number itself should be a range tied to a workflow's measured baseline, not a round figure chosen to look responsible in a board pack, and it needs a standing review point: consumption per workflow checked against outcome on a fixed cadence, with the budget revised when either side of that relationship moves. A budget set once and left alone decays into exactly the blunt cap it was meant to replace.
Who needs this first
This is not equally urgent for every organisation. Most are still consuming too little AI for optimisation to matter more than adoption, and instrumentation built before there is meaningful volume to attribute is effort spent ahead of its value. The pressure shows up first, and most sharply, among organisations already running agentic workloads at scale: the ones exploring lower-cost or open-weight models, running mixed estates across several providers, or watching a bill triple between one budget meeting and the next. Judging that threshold, not treating every AI estate as equally exposed, is part of the engagement, not something we skip past to sell the instrumentation sooner.
Two different scarcities
Token discipline is usually framed as a demand problem: agentic workloads call models more often and more expensively than the queries most budgets were sized for. That is real, but it is not the only driver. Reports through mid-2026 pointed to constraints on the physical infrastructure behind model provision, including memory-chip supply, as a second and independent source of scarcity, one that shows up as price and availability changes an organisation did not cause and cannot instrument away.
The same measurement discipline has to surface both. A workflow whose costs are rising because it is doing more valuable work needs a different response than one whose costs are rising because its model tier became scarcer or more expensive industry-wide. Telling the two apart is why consumption is tracked against outcome rather than against spend alone: a demand-side rise shows up alongside more or better work done, and a supply-side rise does not.
This is the machinery behind observability and cost governance as an ongoing service, and the argument for why per-seat budget instincts do not survive agentic consumption is set out at greater length in Your AI bill is becoming a payroll.
Bring us the bill you cannot explain.
Describe the AI spend that is growing faster than your understanding of it. We will show you what instrumenting it would reveal, and what decisions that visibility would put in your hands.
Your message goes to the people who would do the work.
