20 playbooks for managing AI costs

Plus, key takeaways to help you level up fast.

Welcome executives and professionals. Some firms have burned through a year’s AI budget in a single quarter. Many manage tokens the same as cloud compute: reactively, when the bill arrives.

Those winning don't necessarily have better models or more compute, they manage token consumption with the same rigor as capital allocation.

I analyzed 117 playbooks to uncover the leading enterprise strategies for managing AI costs.

Here are the top 20 for executives:

BOSTON CONSULTING GROUP

Image source: Boston Consulting Group

Brief: BCG published the second in a series of insights on the cost of AI tokens and how companies can manage them. The first article examined the true costs of AI; this one details the challenges of measuring them in practice.

Breakdown:

  • AI works differently than traditional software as a service, changing the unit of management. Traditional FinOps isn't built for it.

  • Instead, a ratio BCG calls return on AI (RoAI) captures the full cost of the AI being applied across the business and its outcomes.

  • To account for costs and assess RoAI, companies need a workflow-level operating model that enables management to do three things well.

  • See what's happening, shape the cost, and either prove the value or stop (or minimize) the activity, as detailed in the image above.

Why it’s important: As AI moves into production, CFOs, CIOs, and CTOs inherit a substantial new cost: tokens consumed to produce outcomes. Rising agent use pushes the meter into overdrive. FinOps cannot keep pace, and CEOs will expect the C-suite to rise to the challenge.

DELOITTE

Image source: Deloitte simulation

Brief: Deloitte released a 28-page report on AI economics, exploring what tokens are, how different agentic models shape pricing, and strategies enterprises can use to optimize token use for maximum competitiveness.

Breakdown:

  • Tokens are the currency of AI economics, as vital as kilowatt hours are to electricity, yet harder to predict, making AI spend volatile.

  • AI costs appear as SaaS line items, metered API usage, or owned infra AI factories balancing GPUs, storage, networks, and energy.

  • Smaller, less predictable workloads may remain API-based, while scaled, high-value workloads shift to AI factories as economics stabilize.

  • The report provides TCO modeling and scenario analysis to show how AI costs scale and where cost inflection points emerge.

  • At 84B+ tokens annually, AI factories offer the lowest TCO. API costs scale linearly, while Neocloud hinges on GPU use (image above).

Why it’s important: AI has become the fastest-growing line item in enterprise IT budgets, consuming up to half of total spend at some firms. At the same time, rising sovereignty pressures and infrastructure control elevates token economics from an IT concern to a board-level issue for CFOs and investors.

INFOSYS

Image source: Infosys

Brief: Infosys published a paper giving executives a plain-language framework for understanding enterprise AI token economics: what drives costs, how to forecast them, how to govern them, and what the ROI of doing so looks like.

Breakdown:

  • AI token costs are poorly governed not because they're uncontrollable, but because their governance frameworks are less understood.

  • Build observability first, you cannot govern what you cannot see. Then add AI model routing: lowest effort, highest return.

  • Deploy quality guardrails in audit mode first, then enforce; add workflow budget controls as automated pipelines mature.

  • Executive ownership of governance decisions (image above) is consistently the single biggest predictor of a program's cost outcomes.

Why it’s important: AI token costs are architecture-driven, not user-driven: two decisions, how you route AI model calls and how you govern information retrieval, drive 78% of achievable savings. Enterprises that build governance before they scale spend 40-50% less, with no loss of capability.

DATABRICKS

Image source: Databricks

Brief: Databricks outlined a set of proven techniques for managing AI coding costs at scale, drawing on its own in-house experience and on conversations with digital-native companies including Stripe, Coinbase, Uber and Ramp.

Breakdown:

  • Move to open source and lower-cost models. Public benchmarks often mislead on real-world coding performance, so build internal evaluations.

  • Route requests and tasks automatically. Routing approaches fall into three categories: request-level, task-level and escalation/delegation.

  • Replace hard spending caps with visibility, tripwires and progressive friction: spend gates, then downshifting, then suspension.

  • Cut token overhead. Techniques include compressing active context more often and leveraging coding harnesses that are more token efficient.

Why it’s important: Exponential AI coding costs are not inevitable, they are a solvable engineering and governance problem. Chase the efficiency frontier (set of models that have the best price point for a given level of intelligence), preserve model flexibility, route intelligently, and cut token overhead.

OPENAI

Image source: OpenAI

Brief: OpenAI outlined practical steps for enterprise leaders to understand how AI is used across their organizations, control spend through targeted policies, and direct investment towards work that creates the most value.

Breakdown:

  • Leaders need a plain view of AI usage: who uses it, which models, how much capacity, and what kind of work it supports.

  • A more capable model may cost more per token, yet it can reach an acceptable result faster, with fewer attempts and less review.

  • Spend controls like workspace defaults, group limits, and overrides let leaders support high-value work without raising limits broadly.

  • OpenAI also addresses managing AI investments as a portfolio and matching the product, capacity, and support model to workflow demand.

Why it’s important: Token price alone does not show whether AI creates value. Leaders should look at useful work per dollar: tasks completed, time saved, and decisions improved. As teams move from chat to longer-running workflows, enterprises need clearer visibility into demand, spend, and risk.

ACCENTURE

Image source: Accenture

Brief: Accenture detailed why tokenomics is becoming the next C-suite discipline, following the launch of Accenture Tokenomics, an offering designed to help enterprises manage token spend by tying it to outcomes.

Breakdown:

  • Have the CFO and CIO jointly own a single enterprise view of AI cost, usage, and return, then ask where that spend earns its return.

  • Build controls into deployment from the very start, defining who can use which models, for what work, and under what safeguards. 

  • Route work to the right intelligence, separating what needs frontier models from what lighter models can handle, with limits by role.

  • Review consumption, outcomes, and routing decisions on a set cadence, because prices, model options, and workloads keep shifting.

Why it’s important: Some companies burn through a year's AI budget in a single quarter, managing tokens the way they manage cloud compute: reactively, when the bill arrives. Those that pull ahead do so by managing token consumption with the same rigor as capital allocation.

BAIN & COMPANY

Image source: Bain & Company

Brief: Bain released part 2 of its token economics series exploring why effective cost per task stays stubbornly flat despite falling model prices, and the practical moves to navigate the nonlinear shift from headcount to tokens.

Breakdown:

  • In part I, Bain laid out how agents, tokens, and data could replace 20% to 30% of headcount opex by 2028, with no clear transition.

  • Despite falling model prices, spend stays high as usage scales, agent work grows more complex, and firms default to frontier models.

  • Bain outlines variables that could shift this, from token-cost trends and on-premises inference to rightsizing models to tasks.

  • It then offers five moves for enterprises to manage costs, from a dedicated AI compute line item to metering the token cost of tasks.

Why it’s important: The opex shift from headcount to tokens isn't a budget problem; it's a structural transformation. The economics are unsettled and the path nonlinear. That's not a reason to wait, but to navigate the shift now, so when the curve breaks you act on hard data instead of intuition.

MCKINSEY & COMPANY

Image source: Reuters

Brief: McKinsey CFO Yuval Atsmon told Reuters how the firm is managing surging AI costs, why productivity gains haven't fundamentally changed project timelines, and how AI is reshaping pricing, talent, and the consulting model.

Breakdown:

  • AI spend is still small but growing 20-30% month over month; one lever is directing people to the right model for the right task. 

  • Teams now deliver more, but timelines are largely unchanged: client complexity has risen as AI forces firms to rethink their business.

  • A third of McKinsey's work is already priced on measurable client impact; Atsmon expects AI to speed the shift from team-based fees. 

  • McKinsey's hiring class is up about 25%, and Atsmon predicts generalist demand will return as AI lets one person span more domains. 

Why it’s important: Atsmon gives a view of AI economics at scale: costs are compounding fast, much of the productivity gain is reinvested in rising expectations, and pricing and talent models must adapt. It offers insights for executives weighing their own AI investments.

BAIN & COMPANY

Image source: Bain & Company

Brief: Bain & Company explored the opex shift from headcount to tokens, raising questions many leadership teams are not yet asking. These are not implementation issues. They’re structural.

Breakdown:

  • Many technology leaders now picture a future opex mix of 70-80% headcount and 20-30% token costs by 2028-2029.

  • Executives need to find unbudgeted millions. Token spend doesn't fit a conventional line item and requires an approval chain that doesn't exist.

  • Early data suggests the top 5% of users (those you can’t afford to throttle) often consume more tokens than the other 95% combined.

  • What does your talent pipeline look like when teams are 3 people, not 15? How do you run a dual operating model without tearing the org apart?

Why it’s important: The question is not whether the opex shift is directionally correct, but whether organizations can survive the transition. Companies will face overlapping costs, organizational whiplash, and periods where headcount and token spend rise together before efficiencies appear.

MCKINSEY & COMPANY

Image source: McKinsey & Company

Brief: McKinsey examined a core CIO challenge: balancing “run” spend that maintains systems with “change” spend that drives innovation and growth, as rising AI investment increases pressure on capital allocation.

Breakdown:

  • McKinsey’s framework assesses tech spend using two metrics: run intensity, revenue share spent on run activities, and change investment.

  • Firms are mapped across four archetypes: deliberate modernizers, strained transformers, lean operators, and heavy IT sustainers.

  • McKinsey finds deliberate modernizers are best positioned to drive value, typically allocating at least one third of spend to change investment.

  • These firms adopt lean architectures, standardise platforms, and reduce tech debt, lowering run costs while freeing capital for agentic AI.

Why it’s important: The companies that outperform will make deliberate decisions about what to simplify or retire. For CIOs, the task now is to reset the run–change balance, ensuring AI unlocks lasting returns with modern capabilities rather than reinforcing today’s complexity.

Gartner - How tech CEOs can improve AI visibility through token efficiency

McKinsey - The cost of intelligence

Cursor - Agent swarms and the new model economics

OpenAI - A scorecard for the AI age

Cognizant - From cost control to cost intelligence

McKinsey - Agentic economics

Forrester - AI cost management

EY - Unlocking agentic value: a new investment discipline

Microsoft - Tokenomics is the new headcount

Google - Guide to AI tokenomics

BCG - Return on AI: What CEOs need to know

McKinsey - Cost versus value

MIT - AI token costs must drop 90% to scale enterprise adoption

Deloitte - AI token cost accounting

MORE MUST-READ BREAKDOWNS

ENTERPRISE AI EXECUTIVE

Agentic and generative AI are evolving rapidly in the enterprise, driving a new era of AI transformation.

Twice a week, we review hundreds of the latest agentic and generative AI best practices, case studies, market dynamics and innovations to bring you what is driving material value — and why it’s important.

Example editions:

  • What Google Cloud CEO told Enterprise AI Executive.

  • Claude Mythos uncovers 10,000+ vulnerabilities.

  • Claude Mythos attacks: Executives’ 11-point defense plan.

  • Google’s enterprise multi-agent playbook.

  • Deloitte's agentic enterprise 2028 blueprint.

  • OpenAI's best practices from 300 implementations.

Found this valuable? Share with a colleague.
Received this from someone else? Sign up here.
Connect on LinkedIn.

Lewis Walker, Editor