Business & Money

Budgeting for AI Spend That Keeps Growing

By Jim Vernon, Editor, AI Intelligence International · Published 15 February 2026 · Reviewed against our editorial standards · About the author

AI costs behave differently from traditional software. Seats are predictable; tokens are not, and usage grows precisely when the tool is succeeding.

Finance teams that treat AI as a normal subscription get surprised twice: first by growth, then by the shadow spend accumulating on departmental cards.

Key takeaways

  • Understand the unit economics: Usage pricing generally scales with input and output volume.
  • Forecast with volume bands, not a point: Produce low, expected and high volume scenarios and price all three.
  • Cap and alert before you need to: Set hard spending limits per environment and per key, plus alerts at defined thresholds.
  • Find the shadow spend: Most organisations have more AI subscriptions than the finance team knows about, purchased on cards by teams solving real problems.

Understand the unit economics

Usage pricing generally scales with input and output volume. Long context, retrieved documents and verbose responses all raise the cost of a single interaction, sometimes by an order of magnitude relative to a short exchange.

That means cost per task is a design choice, not a fixed property. Trimming retrieved context and constraining output length are ordinary engineering decisions with a direct line to the invoice.

Model the cost of one representative task before scaling anything. Multiply by realistic monthly volume, then double it for growth in usage as people find new uses.

Forecast with volume bands, not a point

Produce low, expected and high volume scenarios and price all three. The high scenario is the one to plan liquidity around, because it is what success looks like.

Include retries, evaluation runs and development usage, which are frequently omitted and often add a substantial share on top of production traffic.

Re-forecast monthly for the first six months. Early usage patterns are unstable and the first forecast is always wrong.

Cap and alert before you need to

Set hard spending limits per environment and per key, plus alerts at defined thresholds. The classic incident is a loop in a development script consuming a month's budget overnight.

Separate keys by team and workload so that cost attribution is possible without a forensic exercise. Shared keys make cost control impossible and blame inevitable.

Rate limits also serve as a safety mechanism: runaway usage is often the first symptom of a bug rather than of demand.

Find the shadow spend

Most organisations have more AI subscriptions than the finance team knows about, purchased on cards by teams solving real problems. Treat discovery as a data exercise, not a disciplinary one.

Audit card statements for the common categories, ask each team to list tools they use, and consolidate. Consolidation typically reduces cost and improves data governance simultaneously.

Replace the tools you cancel with an approved alternative on the same day. Removing capability without replacement pushes the spend further underground.

Reducing cost without reducing value

Three levers dominate: route simple tasks to cheaper models, cache repeated work, and shorten inputs. Together they routinely cut spend substantially with no user-visible change.

A fourth is scope: much of the spend in a mature deployment comes from a small number of high-volume, low-value use cases that nobody has reviewed since launch.

Review the top five cost drivers quarterly and ask whether each is still worth its price. Usually one is not.

Presenting it in the budget

Split the line into fixed seats, variable usage and implementation. Fixed and variable behave differently and blending them hides the risk.

Bring the cost calculator output and the subscription audit to the budget conversation. Showing the mechanism by which the number could grow, along with the caps you have set, is what converts an uncomfortable line into an approved one.

A worked example: a 40-person company's first full year

A forty-person agency started with two writing seats on a monthly plan. By month four, three departments had bought their own tools, an engineering team had opened an API account, and a transcription vendor was billing per hour of audio. Nobody had approved an increase; the line item simply grew.

When they finally added it up, the annual run rate was roughly four times the original budget, and about a third of it sat in accounts with fewer than five active users in the previous month. Usage-based spend was the fastest-growing part, because nothing capped it.

The fix was structural rather than punitive: one owner per vendor, hard spend caps on every usage-based account, and a monthly line in the management report. Spend stopped growing on its own without anyone being told to stop using the tools.

Separate the three kinds of AI cost

Seat licences are predictable and easy to right-size, but they drift upward because nobody removes leavers. Audit seats against the last thirty days of activity, not against the org chart.

Usage-based spend is the one that produces surprise invoices. It scales with traffic, with a retry loop, or with one enthusiastic experiment. Every usage account needs a hard cap, an alert at half the cap, and a named owner who gets the alert.

Build-and-maintain cost — the engineering time spent wiring tools together and keeping them working — is invisible on any invoice and frequently exceeds both other categories. Estimate it in days per quarter and put it in the same budget line, or the ROI case will be wrong.

Frequently asked questions

How much should a mid-sized company expect to spend?

It varies enormously by use case, which is why per-task modelling matters more than benchmarks. Model your own top three workflows and extrapolate.

Are annual commitments worth the discount?

Only where usage is proven and stable. In year one, optionality is usually worth more than the discount.

Should each team hold its own budget?

Yes, with central visibility. Local budgets create local discipline; central visibility prevents duplicate purchasing.

What is the fastest saving available?

Cancelling duplicate subscriptions found in an audit, followed by routing low-complexity traffic to cheaper models.

How often should the AI budget be reviewed?

Monthly for the first year, because usage patterns change quickly. Quarterly is enough once spend has been flat for two consecutive quarters.

Should each team hold its own budget?

Yes for seats, no for usage-based API accounts. Shared accounts with per-team tagging give you one place to enforce caps and still attribute cost.

Tools mentioned in this article

More in Business & Money

← All articles