Most model launches are treated as another round in the AI arms race. Companies should interpret this differently: work is in motion. The Gartner numbers are the warning sign. They show what happens when AI goes from answering questions to doing work, and when software goes from something that companies license based on location to doing work, they consume one action at a time.
Everyone keeps telling me that AI costs are uncontrollable because they are moving too fast. I don’t really buy it. Cloud also moved quickly. A runaway cluster could eat up a weekend budget years before anyone calls a workflow an agent. Speed makes the bill arrive faster, but doesn’t explain what the bill means.
1 View gallery


Roy Ravhon.
(Yarin Taranos)
What’s actually different now is where the work ends up. An agent completing a task can spend time in two places at the same time. It calls a model and therefore appears on the AI invoice. It also touches cloud systems to do the real work: computation, storage, data, logs, and calls to other tools. The engineering team sees a job. The model provider sees tokens. The cloud provider sees infrastructure. Nothing connects these parts by default. In practice, surprises often arise from unchecked agent runs, oversized context windows, and repeat loops that look small until they are added up.
That’s why the first surprise often comes in the form of an invoice, but the more difficult problem is not the invoice itself. It’s the missing link between work, personal responsibility and results. A coding agent can feel fast within a sprint. A support workflow may look successful as usage increases. A product feature can increase engagement. None of this tells the company which team incurred the costs, which product or customer they served, or whether the outcome was worth it.
This is where systems designed for the last decade of enterprise software begin to answer the wrong question. License trackers, procurement calendars and quarterly true-ups have been developed for licenses, contracts and predictable renewals. Agents don’t behave like that. They run on demand. They call tools. They try again. They carry context. They go through multiple systems before a single task is completed. It is not enough to extend the old software licensing logic to agent work, because the old logic was designed to answer a different question.
The cloud has already taught companies that fast-moving infrastructure can still be managed. At first, companies stared at AWS invoices as if the total amount was the solution. Over time, they discovered that the useful questions were more confusing. Shared services had to be split up. Resources without tags had to find an owner. For one product the cost might be justified, for another it could be pure waste. The bill was not the story. The work behind the bill was.
AI adds a second cost surface to this problem that is more opaque than the first. The market still pays too much attention to the model price. The model price is important, but it is not the unit that the company buys that matters. The company purchases a resolved ticket, a merged code change, a clearer forecast, a fraud check, or an internal workflow that used to take hours.
A cheap model can become expensive if it requires five retries and carries half the code base in context. An expensive model can be cheaper overall if it is cleanly processed, requires fewer steps and avoids rework. Focusing only on token price is like judging cloud efficiency based on the cost of an instance without asking what workload it serves.
Anthropic’s Claude Apps gateway is a sign of the same shift. It transforms Claude Code from a tool that developers run individually into an infrastructure that an organization can manage. This is useful, but it also highlights the limitations of tool-level controls: a company still needs to link AI usage to the underlying workflow, team, product and outcome.
The same pattern repeats throughout the AI stack. Each provider provides a different dashboard, billing language and user interface. The company still needs an answer: what work incurred the cost, who owns it, and whether the result justifies the expense.
The next management question is not how many people use AI. Acceptance was a useful signal during the pilot phase. Companies were told that employees were curious and the tools were spreading. In production the transfer is incomplete. It doesn’t tell you whether the work improved performance, reduced support burden, increased quality, or quietly turned a good product margin into a bad one.
This is already visible in software products that add AI capabilities. Usage can look good for two quarters, then someone calculates gross margin and the room goes quiet. The average user may be fine, while the most intensive users will consume enough inference, memory, or agent time to negate the economics of the feature. Recognizing the costs is just the starting point. The more difficult questions are pricing, packaging and ownership. If the AI work is bundled into the basic subscription, at some point someone will have to answer whether the price reflects what it actually costs to provide the feature.
The same problem occurs with internal productivity. A development team might feel faster with coding agents, but the finance department can’t “feel faster”. It needs to know which teams are consuming the budget, which workflows have changed, which projects have benefited, and which usage patterns are just expensive loops. While a support team can automate more interactions, the useful unit is not the number of AI conversations. It’s about the cost per case solved, the time to resolution and the quality of the response. The work unit is more important than the tool name.
This is where the discipline that has emerged around cloud costs comes in next. It started with turning cloud billing into accountability: allocation, budgets, forecasting, anomaly detection, and unit economics. AI does not replace this muscle. This makes muscle more important as costs move faster, traverse more systems and get closer to the product itself. Tokens are part of the story, but tokens alone do not answer the business question. The answer must connect model usage, cloud usage, team ownership, product context and outcome.
The companies that manage AI well will be the ones that look beyond the best model or lowest token rate. You will know which agent runs deserve more investment, which need a different model, which require policies, and which should be stopped because the economics don’t work. They will make it easier to use AI where it creates leverage and easier to stop where it creates costs without changing the outcome.
The first wave of enterprise AI was about access. The second wave is about accountability. The winners will be the companies that can look beyond the sample bill and say with confidence what each AI task cost and what it returned.
Roi Ravhon is co-founder and CEO of Finout and a member of the board of directors of the FinOps Foundation.
https://www.calcalistech.com/ctechnews/article/h1ymav84gg
