Zero-Based Budgeting for Cloud and AI: Why 2026 Ended the Last-Year-Plus-Growth Era
80–85% of enterprises miss AI cost forecasts by 25% or more. Zero-based budgeting is the discipline that ends the last-year-plus-growth model for cloud and AI spend.
The CFO approved the AI budget in January. By April it was 40% over, and no one in the room could say which workload caused it.
That scene is playing out across the enterprise in 2026, and it exposes a budgeting model that quietly expired. For decades, technology budgets were built the way most budgets are built: take last year’s run rate, add a growth allowance, defend the delta. The base was assumed. Only the increment was scrutinised. That model worked when infrastructure cost behaved like rent — stable, predictable, contracted. It does not work when a single prompt costs anywhere from a fraction of a cent to fifty cents, when a GPU cluster sits idle 95% of the time while the invoice looks steady, and when 98% of FinOps teams are now managing a cost category that did not meaningfully exist three budget cycles ago.
Gartner has a name for the alternative, and it is not new. Zero-based budgeting — justify every expense from a blank slate, each period, against the business outcome it delivers — was conceived in the 1970s. What is new is that cloud and AI have made it the only budgeting discipline that survives contact with the 2026 technology estate.
Executive Summary
- 80–85% of organisations miss AI cost forecasts by more than 25% (Mavvrik and BenchmarkIT, 2026). Roll-forward budgeting cannot model consumption-based, token-metered economics.
- AI is now 19% of total cloud spending, up from 8% in 2023 (Gartner Q1 2026) — the fastest-growing cost category is the one the last-year-plus-growth model handles worst.
- As much as 95% of enterprise AI infrastructure spend is wasted at low GPU utilisation (VentureBeat Q1 2026), and the underutilisation penalty on identical H100 hardware runs 24× to 36× per token (arXiv, 2026).
- Only ~29% of executives can confidently measure AI ROI, against 79% who perceive productivity gains (IBM, 2026). The value is real; the budget defence is broken.
- The token invoice is one of nine AI cost buckets (FinOps X 2026). A forecast built on it alone is wrong by construction.
- Budgeting to business value rather than cost alone yields over 30% more profitability or market-share growth (Gartner). Zero is the method. Value is the destination.
What Zero-Based Budgeting Means for a Technology Estate
Zero-based budgeting rejects the single assumption that underpins most corporate budgets: that last year’s spend is a legitimate starting point. Gartner defines it as the process of justifying all expense items from scratch by providing a business case for each, distinct from traditional budgeting where historical spend patterns become the basis for future budgets. Every line begins at zero. Every line earns its funding by demonstrating alignment with a strategic outcome.
The principle is unglamorous and, in its original form, expensive to run. Executive inquiries to Gartner about it have surged, and yet only half of organisations implementing it realise the savings they expected — often because they treat it as a procedural change to the budgeting process rather than a principle for making better spending decisions. The failure is rarely the method. It is applying the method to the wrong things, or applying it once a year to a cost structure that changes daily.
Technology spend is exactly that cost structure. When the estate was a fixed set of servers and a handful of SaaS contracts, an annual zero-base was feasible if painful. In 2026 the estate is multi-cloud infrastructure, Kubernetes clusters, token-metered inference, Databricks and Snowflake workloads, agentic pipelines, and data centre — and the FinOps Foundation reports that this scope expanded faster than any organisation can staff for. Zero-basing it by hand is not slow. It is impossible. The discipline has to be automated into the platform that governs the spend, or it does not happen at all.
Why the Roll-Forward Budget Breaks in 2026
Roll-forward budgeting assumes a stable base. AI spend has no stable base. A single request can vary in cost by two orders of magnitude depending on model, context length, and output size, and that variability multiplies across thousands of daily requests and multiple vendors adopted independently by different teams. The result is a forecast that is wrong on arrival: 80–85% of organisations miss AI forecasts by more than 25% (Mavvrik and BenchmarkIT, 2026), and IDC projects G1000 organisations will face a 30% rise in underestimated AI infrastructure costs by 2027, driven not by reckless spending but by under-forecasting and expenses unique to AI workloads.
The scale of the exposure is what makes this a board problem rather than a FinOps inconvenience. Global AI spending is forecast at $2.5 trillion in 2026, a roughly 44% year-over-year jump (Gartner), and AI now represents 19% of total cloud spending against 8% in 2023 (Gartner Q1 2026). The fastest-growing category in the enterprise cost base is the one the incumbent budgeting model governs least well.
There is a deeper structural reason the model fails, and it is the one most forecasts miss. The token invoice — the number finance sees — is only one of nine AI cost buckets identified at FinOps X 2026. Inference infrastructure, idle GPUs, the KV cache the invoice never shows, evaluation and monitoring, model-churn integration costs, and more sit beneath the visible line. Organisations that cannot measure their AI ROI underestimate AI costs by 40–60% on average (McKinsey and IBM, 2026), because they count the licence and miss the iceberg. A budget built by rolling forward last year’s visible number therefore inherits last year’s blind spots and compounds them.
The GPU Cluster That Costs 36× More Than the Invoice Suggests
Consider the most expensive thing in a modern AI estate that no roll-forward budget will ever flag: an idle GPU.
At batch size one, a GPU spends most of its time waiting to load model weights from VRAM for a single active request, and utilisation falls below 5%. The invoice for that GPU-hour looks identical whether the hardware is saturated or nearly asleep. What changes — invisibly — is the cost per useful token. On identical H100 hardware, effective cost spans $0.21 to $15.25 per million output tokens depending on offered load, an underutilisation penalty of 24 to 36 times (arXiv concurrency-aware cost study, 2026). A GPU at 10% load turns a $13-per-million-token workload into a $130 one. Estimates place total wasted enterprise AI infrastructure spend as high as 95% at low utilisation (VentureBeat Q1 2026).
The Idle GPU the Invoice Hides
──────────────────────────────────────────────────────────
What finance sees What the workload actually costs
────────────────────── ──────────────────────────────────
GPU-hour: stable Batch size 1 → utilisation < 5%
Monthly invoice: flat Cost/M tokens at 10% load: $130
Line item: "AI compute" Cost/M tokens at target load: $13
──────────────────────────────────────────────────────────
Same hardware. Same invoice. 10× the cost per useful token.
Roll-forward budgeting cannot see this. A zero base can.
──────────────────────────────────────────────────────────
This is where tokenomics enters the budget conversation. Tokenomics is the unit economics of token-metered workloads — cost per million tokens, input/output ratio, cached-token ratio, GPU utilisation, and their drift over time. FinOps X 2026 laid out its maturity as a three-stage progression: spend visibility first, economics second, value as the destination. The optimisation levers are real and compounding — routing requests to right-sized models, quantisation, caching, continuous batching — but they only pay off if every layer is instrumented. Miss one layer and the savings collapse at the others.
The point for a budget owner is that none of this is legible from the top line. The token bill does not reveal the cached-token ratio. The cloud invoice does not reveal the idle cluster. You cannot zero-base what you cannot attribute, and you cannot attribute token, GPU, and data-platform cost to a named workload with a tool that was built to reconcile monthly cloud invoices. That gap between what the invoice shows and what the workload costs is precisely why so many 2026 AI budgets miss by 25% or more.
Zero Is the Method. Value Is the Target.
The most common way zero-based budgeting fails is subtle: teams treat it as a hunt to spend less. Gartner’s own research is blunt about the cost of that mistake. Unfocused zero-based exercises become opportunistic, seeking any opportunity to reduce spend without concern for business impact — and organisations that budget by business value instead of cost achieve over 30% more profitability or market-share growth. The name misleads. The goal was never zero. The goal is to fund what returns value and defund what does not.
For technology, that reframing is the entire game in 2026. Zero-basing a GPU cluster tells you the spend is unjustified if it sits idle. Value-basing it tells you whether the workload it runs earns its keep — whether cost per inference against the revenue that inference generates, or cost per verified outcome, clears the bar. Boards have stopped accepting efficiency projections as proof of value; in the 2026 budget cycle the directive to finance is to show measurable return, not promises. The uncomfortable truth is that most AI does generate value and most organisations cannot prove it, because they never set up the measurement.
This is why zero-basing and value-basing are a sequence, not a choice. Pressure-test each workload for strategic alignment first — does it map to an outcome anyone can name? Then pressure-test its cost. A workload with a named owner, a measurable outcome, and a defensible cost per unit survives the zero base. A workload that is none of those things is exactly what the discipline exists to surface. The organisations that keep their AI budgets through the 2026 cycle are not the ones with better models. They are the ones that baselined before they deployed, attributed every dollar to an owner, and can express the result in the P&L vocabulary a board accepts — cost per transaction, cost per customer, margin impact — rather than tokens and GPU-hours no director thinks in.
How to Zero-Base Cloud and AI Spend
Zero-basing a technology estate is five moves, and each one is a prerequisite for the next.
Normalise the full estate. A zero base is impossible if every data source is a separate conversation. Cloud, Kubernetes, AI workloads, Databricks, Snowflake, SaaS, and data centre have to resolve into one specification before any of them can be justified against a common outcome. The FinOps Open Cost and Usage Specification (FOCUS 1.4) exists for exactly this reason, and its adoption is what lets practitioners apply consistent principles across an increasingly complex landscape.
Attribute every dollar to an owner. FinOps X 2026 named direct-allocation percentage — the share of spend with a named owner — as a core maturity marker, because unattributed spend cannot be justified from zero. For AI specifically this means per-chain, per-model, per-environment attribution, not a shared account where GPU and egress costs disappear.
Build decision packages with kill/continue criteria. Each workload gets explicit thresholds — token budget caps, GPU idle limits, commitment coverage targets — enforced by policy rather than a quarterly spreadsheet review. Fund releases with kill/continue gates, not full transformations upfront.
Rebuild from outcomes, not run rate. The budget is assembled from cost per inference, cost per customer, and cost per verified outcome — the business results the spend is meant to produce — rather than last year’s number plus a percentage.
Govern continuously. Consumption-based spend drifts within the period. Real-time anomaly detection and automated enforcement keep the zero base valid after planning ends, so the discipline is a live control system rather than an annual event that is obsolete by February.
DigiUsher FinOps OS — What Makes a Zero Base Possible
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
┌─────────────────────────────────────────────────────┐
│ DigiUsher FinOps OS │
│ (FOCUS 1.4 Native Architecture) │
├──────────┬──────────┬──────────┬────────────────────┤
│ Cloud │ Kuberne- │ AI │ Data Platforms │
│ AWS/Az/ │ tes │ Bedrock/ │ Databricks / │
│ GCP/OCI │ EKS/AKS/ │ OpenAI/ │ Snowflake / │
│ │ GKE │ Vertex │ Redshift │
├──────────┴──────────┴──────────┴────────────────────┤
│ SaaS · Marketplace · On-Premises DC │
└─────────────────────────────────────────────────────┘
Normalise → Attribute → Package → Rebuild → Govern
│
▼
BYOC Secure Relay Proxy
(collector inside customer perimeter)
Where DigiUsher Fits
Every step above assumes two capabilities most cost tools do not have: normalisation across the full technology surface, and attribution down to the individual AI workload. DigiUsher was built as a FinOps Operating System to provide both, which is what turns zero-based technology budgeting from a slide into an operating practice.
Because DigiUsher is native to FOCUS 1.4 rather than adapted to it, it normalises AWS, Azure, GCP, OCI, and Alibaba Cloud alongside EKS, AKS, GKE, and OKE, the major AI providers, Databricks, Snowflake, SaaS, and data centre into a single attributable view. The zero base covers the full estate — not the 60% a single-category tool reaches before Kubernetes and data platforms fall out of scope. Its AI governance is native, not retrofitted from an infrastructure tool that predates the token economy: token budget caps, agentic kill-switches, GPU idle detection, and per-chain attribution give finance the workload-level justification a zero base demands. Real-time anomaly detection surfaces cost drift before the invoice lands, so the base holds through the period. AI-Specific P&L and unit-economics reporting express the outcome in board vocabulary, where zero-basing becomes value-basing.
The commercial model matters here more than it might appear. A platform priced as a percentage of the spend it governs has a structural incentive misalignment: it profits as your bill grows. DigiUsher’s flat enterprise licensing scales with the estate, not with the spend — vendor incentives aligned with the cost reduction a zero base is meant to produce. For regulated industries where billing data cannot leave the perimeter, BYOC deployment via the Secure Relay Proxy keeps the collector inside the customer’s environment, validated at institutional scale with ICICI Bank. A European energy utility used the same capabilities to remove €1M from its Databricks estate in 45 days — a zero base applied to a data platform that a cloud-only tool would never have reached.
DigiUsher is SOC 2 Type II certified and GDPR compliant, an AWS ISV Accelerate Partner listed on AWS Marketplace, Azure ISV Co-Sell Ready and MACC-eligible on Azure Marketplace and AppSource, and a Google Cloud Partner on GCP Marketplace. It is delivered globally by Infosys, Wipro, and Hexaware.
Frequently Asked Questions
What is zero-based budgeting for cloud and AI? Zero-based budgeting for cloud and AI is the practice of justifying every unit of technology spend from a zero base each budget period against a measurable business outcome, rather than carrying forward the prior period’s spend plus a growth allowance. Applied to technology the principle is unchanged from Gartner’s definition, but the cost structure is harder: cloud and AI spend is consumption-based and non-linear, so a token-metered workload or an idle GPU cannot be modelled as a fixed line. It requires per-workload attribution — every dollar tied to a named owner — because unattributed spend cannot be justified from zero. A FinOps Operating System such as DigiUsher makes this operational by normalising the full estate into one FOCUS 1.4 view.
Why does roll-forward budgeting break for AI spend in 2026? Because AI cost is consumption-based, non-linear, and largely invisible in the systems finance uses to plan. A single prompt can vary in cost by two orders of magnitude, and that variability across thousands of requests and multiple vendors makes a static forecast structurally unreliable — 80–85% of organisations miss AI forecasts by more than 25% (Mavvrik and BenchmarkIT, 2026). The token invoice is one of nine AI cost buckets (FinOps X 2026), so a forecast built on it omits eight of nine costs. The fix is to zero-base from workload-level attribution each period.
What is the difference between zero-based budgeting and value-based budgeting? Zero-based budgeting rebuilds the budget from a blank slate, requiring a business case for every expense; value-based budgeting ranks every justified expense against the value it returns. Gartner finds that budgeting by business value rather than cost alone yields over 30% more profitability or market-share growth. For AI, zero-basing tells you a cluster is unjustified if it sits idle; value-basing tells you whether the workload it runs earns enough business outcome to deserve funding. The two operate as a sequence — test strategic alignment first, then cost.
How much cloud and AI spend do enterprises waste on idle GPUs? Estimates place wasted enterprise AI infrastructure spend as high as 95% at low GPU utilisation (VentureBeat Q1 2026). At batch size one, GPU utilisation falls below 5%, and on identical H100 hardware effective cost spans $0.21 to $15.25 per million output tokens depending on load — a 24× to 36× penalty (arXiv, 2026). This is invisible to roll-forward budgeting because the invoice looks stable while cost per useful token multiplies. Governing it requires GPU idle detection and per-workload attribution, which DigiUsher provides natively.
How should enterprises apply zero-based budgeting to their technology estate? In five steps: normalise every cost source into one specification, attribute every dollar to a named owner, build decision packages with kill/continue criteria, rebuild the budget from business outcomes rather than run rate, and govern continuously with real-time anomaly detection. The FinOps Foundation reports scope has expanded faster than teams can staff, so this cannot be a manual annual exercise — a FinOps Operating System automates the normalisation, attribution, and enforcement.
How does DigiUsher enable zero-based budgeting across cloud and AI? By making full-estate normalisation and per-workload attribution operational. Built natively on FOCUS 1.4, DigiUsher normalises multi-cloud, Kubernetes, AI, Databricks, Snowflake, SaaS, and data centre into one attributable view. Its native AI governance — token budget caps, agentic kill-switches, GPU idle detection, per-chain attribution — gives finance the justification a zero base demands, and AI-Specific P&L expresses the result in board vocabulary. Flat licensing keeps incentives aligned; BYOC keeps regulated data in the perimeter, validated with ICICI Bank.
What are AI tokenomics and why do they matter for budgeting? AI tokenomics is the unit economics of token-metered workloads — cost per million tokens, input/output ratio, cached-token ratio, GPU utilisation, and their drift. It matters because the same workload on the same hardware can cost 24× to 36× more per token depending on stack configuration (arXiv, 2026), and the token invoice is one of nine cost buckets (FinOps X 2026). A budget built on the token bill alone is wrong by construction. DigiUsher provides per-token, per-chain attribution natively through its Meter module.
What happens when organisations fail to govern AI and cloud spend from a zero base? Cost drift compounds silently until it surfaces as a P&L problem after the fact. Organisations that cannot measure ROI underestimate AI costs by 40–60% (McKinsey/IBM, 2026), and only ~29% of executives can confidently measure AI ROI (IBM, 2026). Boards now demand measurable return, and projects without articulated outcomes are the first cut when budgets tighten. Without a zero base there is no mechanism to deprioritise low-impact spend — the real failure is overspending without a prioritisation mechanism. A FinOps Operating System restores that mechanism.
References
- FinOps Foundation — State of FinOps 2026
- Gartner — Use Zero-Based Budgeting to Rightsize Your Budget
- Gartner — Value-Based Budgeting, Because Zero Isn’t Valuable
- Gartner — Definition of Zero-Based Budgeting, Finance Glossary
- FinOps X 2026 Recap — AI Token Economics Explained (Mavvrik)
- AI Cost Statistics 2026: Forecasting, ROI, and Budget Risk (Mavvrik / BenchmarkIT)
- VentureBeat — Enterprise GPU Utilization: The AI Infrastructure Problem (Q1 2026 Market Tracker)
- Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation (arXiv, 2026)
- Spheron — GPU Cost Per Token Benchmark, LLM Inference 2026
- ClarityArc — FinOps and AI Cloud Cost Management 2026
- Computerworld — AI Budgets Soar, ROI Still Elusive (Forrester analysis)
- CIO Dive — FinOps Teams Gain Clout as AI Costs Climb
- Cloud Computing Industry Statistics 2026 (FinOps Foundation / Gartner data)
Roll-forward budgeting assumes last year was right. Cloud and AI made that assumption expensive: the discipline of 2026 is to fund what returns value and defund what cannot prove it — every workload, every period, from zero.
See what a zero base looks like across your full estate.
Related reading
- The CFO’s Guide to AI ROI: Answering the Board’s Hardest Question — May 20, 2026 — the value layer that turns an attributed AI budget into a defensible ROI story.
- GPU Cost Governance: Taming the Fastest-Growing Line in Your Cloud Bill — April 15, 2026 — idle-GPU detection and utilisation benchmarks in depth.
- Cloud Cost Optimization Is Dead. Long Live Technology Value Management. — February 10, 2026 — the strategic shift value-based budgeting formalises.
- Why FinOps Is Now a Board-Level Discipline — March 5, 2026 — the organisational pressure making zero-based technology budgeting urgent.
- AI Unit Economics: Building an AI-Specific P&L — June 1, 2026 — the unit-economics reporting behind value-based AI budgeting.
DigiUsher in 15 min
Give your board a cost-to-outcome view they can actually act on.
DigiUsher connects cloud and AI spend to the outcomes it funds, in language finance already speaks.
Book my 15-min discovery callNo hard pitch · specific to your stack
Continue Reading
More from the DigiUsher editorial team.
Board-Level AI ROI: Why $600B in Investment Is Delivering Single-Digit Returns — and What Fixes It
Only 7% of CFOs see high ROI from AI despite $270B in enterprise spend forecast for 2026. Here's the governance framework boards are demanding — and why cost attribution is the missing layer.
AI Unit Economics: Why Your AI Invoice Can No Longer Answer the Board’s Question
98% of FinOps teams now manage AI spend, but almost none can price a merged pull request. Here's the framework for AI unit economics in 2026 — and why the invoice alone can't answer the board's question.
Why Cloud and AI Budgets Fail Annual Planning
Annual planning was built for fixed assets. Cloud and AI spend is consumption-driven, non-linear, and breaks the 12-month budget cycle. Here is what replaces it.