Agentic AI may remove Finance from the decision chain. Explore how FP&A can shift from insight...

It’s interesting to see how fast the enterprise AI narrative has shifted from “if” to “how much.” Companies are routinely signing off on AI pilots. They believe they have clear vendor invoices and a limited downside. But then this happens: CFO approves a $50,000 pilot on Day 1; by Day 100, the company has spent multiples of that figure across line items no one can clearly explain.
Here is what makes AI spending difficult to track, where the hidden costs sit, and what must be done to master the economics of implementation.
Why AI Looks Cheap on Day 1
At the start, the numbers usually look manageable. A typical pilot budget basically resembles a standard SaaS contract:
Model API access: A fixed monthly commitment (e.g., $20K to OpenAI, Anthropic, or Google).
One vendor platform fee: A single line item for a copilot, RAG tool, or agent framework.
A small project team: Two engineers and a product manager (fully loaded).
A defined use case: For example, “summarise support tickets” or “draft sales emails.”
ROI math is simple: cost saved per ticket × ticket volume − license fee. The pilot gets approved.
This is the peak of financial clarity. Then comes the downhill.
What Shows Up by Day 100
By Day 100, the CFO opens the cloud bill and sees an unfamiliar pattern. In enterprise rollouts, I have observed that total AI-adjacent spend has grown 4 to 8 times the original pilot budget, and yet nobody can fully explain where the money went. Each team defensibly points to a “small” expense. The overspend is high, and attribution is impossible.
This fog has several sources, each practically invisible on Day 1.
The Token Meter Kept Running
SaaS seats are predictable, but large language models are billed by tokens consumed. Token usage grows faster than people expect. A pilot averaging 10,000 calls a day at 2,000 tokens each looks cheap. Once embedded into a customer-facing product, traffic can easily jump to 1 million calls a day. On top of that, retrieval-augmented prompts can be 10 to 50 times larger than the original prompts. The bill compounds across three axes simultaneously: more users, more tokens per user, and more expensive models being silently selected by autoscaling logic.
Shadow AI
Employees adopt AI faster than procurement. KPMG’s research on Shadow AI [1] points to a similar pattern: employees often turn to external AI tools because they are fast, accessible, and less restrictive than approved internal options. By Day 100, in many enterprise environments, a typical mid-sized enterprise may have 20 to 60 unsanctioned AI subscriptions either purchased on personal cards or buried inside other SaaS bundles, including Notion AI, GitHub Copilot, Grammarly, and Perplexity Pro. These rarely show up as “AI” in the general ledger. They are coded as “productivity,” “software,” or something more generic.
The Infrastructure Tail
Production AI requires a stack the pilot never paid for: vector databases, observability tools, guardrail layers, and evaluation harnesses. Each tool seems manageable on its own. But together they can exceed the model API bill itself.
Egress, Storage, and the Cloud Multiplier
RAG systems chunk, embed, and store documents in vector stores duplicated across regions. Embeddings are recomputed every time a document changes. Logs of prompts and completions are retained for compliance. The hyper-scaler bill gladly absorbs all of this under “storage” and “data transfer.”
The Unbudgeted Human Cost
Pilots assume existing engineers will “just integrate the API.” Things get complicated once in production: organisations suddenly need prompt engineers, ML platform engineers, red-teamers, and MLOps functions. Vendor-contractors at premium rates will fill some of these roles. Internal time spent on AI governance and training will never appear on any invoice.
And…The Costs That Arrive After Go-Live
Even the Day 100 picture captures only the directly attributable spend. The largest cost categories often haven’t yet landed on the books. They are queued up behind every successful deployment, waiting to be triggered after the system goes live for real users.
My observation is that CFOs who benchmark AI investment purely against the build effort consistently underestimate the total program cost by a factor of 2 to 3. A useful rule of thumb, drawn from large-scale enterprise rollouts, is the 1: 0.25: 1–2 ratio:
For every 100 person-days of implementation effort, expect roughly 25 person-days of training effort and a further 100 to 200 person-days of change management and business disruption costs.
The cost scope the CFO has considered so far is at best one-third of the true economic effort required to deliver the value promised to the board on Day 1.
The 25% Training Burden
Most firms underestimate the true cost of training. It requires role-based curricula, hands-on labs, and refresher cycles because models and UIs change every few months. The 25-person-day figure is paid in the salaries of people not doing their day jobs while they learn.
The 100–200% Change Management Tax
This sinks more AI programs than any model bill. It includes process redesign, dual-running costs (where old and new processes operate in parallel), human-in-the-loop quality assurance overhead, and the temporary productivity dip as users learn, distrust, verify, and re-verify model output. A 100-person-day implementation frequently carries a seven-figure organisational liability extending 6 to 18 months beyond go-live. Now, who could have possibly captured that on Day 1?
Why Finance Systems Miss the Real Bill
Let’s be fair to our finance leaders. CFO toolkits were designed for capex, headcount, and SaaS, all categories with predictable unit economics. That framework starts breaking down with AI:
- Traditional assumption: Costs are mostly fixed or step-fixed.
AI reality: AI costs are usage-metered and highly elastic. Traditional assumption: One contract per vendor.
AI reality: AI spend often comes through dozens of micro-contracts and pay-as-you-go meters.Traditional assumption: Spend lives in IT or one business unit.
AI reality: AI spend is distributed across every function that touches it.Traditional assumption: Outputs are deterministic.
AI reality: AI outputs are probabilistic, and quality assurance is itself a cost.Traditional assumption: Procurement gates new spend.
AI reality: Engineers can spin up GPU clusters with a credit card.Traditional assumption: Costs are vendor-invoiced.
AI reality: The largest AI costs are payroll-absorbed and never invoiced.
As a result of this, conventional FP&A reports show “cloud cost” or “software cost” rising, but the AI driver inside those buckets is invisible until a few quarters of overspend have accumulated.
The Pre-Approval Mandate: Rewriting the FP&A Gate
To prevent the Day 100 shock, FP&A teams must radically alter their gatekeeping behaviour before signing off on the AI pilot. Traditional procurement relies on signing a vendor contract and closing the folder. For AI, the business case format itself must be rewritten. Finance must mandate that any pilot proposal explicitly details its structural scaling boundaries:
The Volumetric Ceiling: At what exact traffic or user count will this pilot shut off or flag an alert, preventing silent API or token autoscaling from running amok?
The "Shadow" Audit: What existing software capabilities are being displaced, and how will procurement audit corporate credit cards for micro-subscriptions over the next 90 days?
The Downstream Burden Capitalisation: The proposal must include a pre-committed budget line for the internal payroll allocation required for training and process dual-running post-go-live.
The goal is simple: force the complexity of Day 100 into the financial modelling of Day 1.

Figure 1. How AI Pilot Costs Expand from Day 1 to Day 100
Finally… What AI FinOps Should Look Like
To actually make money from AI, we need to start treating cost observability as a first-class engineering and finance discipline. This direction is also reflected in the FinOps Foundation’s State of FinOps 2025 report [2], which shows FinOps expanding beyond public cloud into a broader “Cloud+” scope, including SaaS and AI-related cost visibility. A few things distinguish the firms that capture AI ROI from those that continue funding pilots without clear economic accountability:
Tag every token. Every API call should be tagged at the application layer with team, use case, environment, and customer. Without tag enforcement, no allocation is possible. Cloud-style FinOps practices must be extended to model APIs too.
Adopt a unit-economics view. Replace the question “How much are we spending on AI?” with “What is the cost per resolved ticket / generated lead / approved claim?” Unit economics expose whether usage growth is a feature (more revenue per dollar) or a leak (more cost per dollar).
Use different models for different workloads. Most production traffic does not need the most recent frontier model. A disciplined router, with a small model used by default and a premium model used on escalation, typically reduces inference spend by 60–80% in many enterprise settings while protecting output quality. This is consistent with the cost-performance logic described in FrugalGPT [3], which shows how LLM cascade strategies can reduce inference costs by routing different queries to different models.
Budget for training and disruption upfront. Apply the 1: 0.25: 1–2 ratio at business-case approval, not at go-live. Force every AI investment proposal to name the training owner, the change-management owner, and the productivity-dip assumption.
The Question Every CFO Should Still Be Able to Answer
The issue isn’t that AI can’t create value. It’s that AI economics are hard to track with current financial instruments. Both the Day 1 budget and the Day 100 budget are real, but the former is irrelevant, while the latter is unallocated.
Profit does not come from negotiating a better price with Anthropic or OpenAI. It comes from building the meters, tags, and unit economics that turn a fog of distributed spend into a portfolio of investments.
The CFOs who win the AI cycle are not the ones who approved the boldest pilots. They are the ones who can still answer this straight question on Day 100: how much did we spend, on what, and what did we get back? Until that question has an answer, your AI program is, financially speaking, just a pilot.
Note: Except where explicit source references are provided, all statistics and quantitative benchmarks in this piece reflect the author's proprietary observations and field-tested frameworks for enterprise AI economics.
Sources:
1. KPMG. (2025). Shadow AI is already here: Take control, reduce risk, unleash innovation. KPMG. https://kpmg.com/us/en/articles/2025/shadow-ai-already-here.html
2. FinOps Foundation. (2025). The State of FinOps 2025. https://data.finops.org/2025-report/
3. Chen, L., Zaharia, M., & Zou, J. (2023). FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. chrome-extension://efaidnbmnnnibpcajpcglclefindmkaj/https://arxiv.org/pdf/2305.05176
Subscribe to
FP&A Trends Digest

We will regularly update you on the latest trends and developments in FP&A. Take the opportunity to have articles written by finance thought leaders delivered directly to your inbox; watch compelling webinars; connect with like-minded professionals; and become a part of our global community.