
AI Infrastructure Costs Are Out of Control — Here's How to Fix Them
A healthcare AI company was running eight A100 GPU clusters around the clock. Monthly bill: $156,000. Nobody had measured actual GPU utilization. When someone finally did, the number was ugly — most of those GPUs were idle most of the time. After implementing proper Kubernetes scheduling and auto-scaling, the bill dropped to $34,000. Same workloads. Same velocity. The only thing that changed was how the GPUs were shared.
That pattern is not unusual. It's the default.
There's a name for this: the Warm Air Tax — the gap between what you pay for GPU infrastructure and the useful compute those GPUs actually deliver. Enterprise AI infrastructure spending hit $82 billion in Q2 2025 alone, up 166% year-over-year. The average enterprise monthly AI spend crossed $85,000. And GPU utilization across most organizations sits between 10% and 30%.
Run the math on a 100-GPU H100 cluster at 15% utilization and you get roughly $1.4 million per year in wasted capacity. The problem isn't that AI infrastructure is expensive. It's that most teams are paying the Warm Air Tax without knowing it exists.
The GPU Utilization Problem Nobody Measures
When a team provisions a GPU for a workload, Kubernetes assigns the entire device. If your inference service needs 6 GB of VRAM on an 80 GB H100, the remaining 74 GB sits idle. No other pod can touch it.
This is how default Kubernetes scheduling works. The device plugin model treats GPUs as indivisible units — request one, get the whole thing. For teams running a mix of small inference jobs, fine-tuning tasks, and batch processing, the result is predictable: most GPUs spend most of their time doing nothing.
The CNCF documented this across member organizations. Baseline utilization: 13%. With advanced scheduling, it climbed to 37%, with some implementations exceeding 80%.
The difference between 13% and 80% on a 100-GPU cluster is millions of dollars annually.
That's not an optimization opportunity. It's a broken default.
Three patterns that drive the waste
No fractional GPU sharing. Until recently, Kubernetes had no native mechanism for sharing a single GPU across pods. MPS (Multi-Process Service) existed but was unstable under mixed workloads. MIG (Multi-Instance GPU) partitions the device in hardware but requires pre-configuration and limits flexibility. Most teams defaulted to one pod per GPU, regardless of actual demand.
Fear of OOM. GPU out-of-memory errors kill workloads instantly — no swap, no graceful degradation. Teams pad their GPU memory requests by 2-3x to avoid surprises. A model that fits in 20 GB gets allocated 40 GB. The delta sits empty.
No visibility. Without DCGM Exporter or equivalent monitoring, teams can't see GPU utilization metrics. They don't know their fleet runs at 15% because nobody measured it. The GPU shows up as "allocated" in Kubernetes, and allocated looks the same as busy.
Fixing GPU Scheduling on Kubernetes
The scheduling layer is where you kill the Warm Air Tax.
Kubernetes 1.34 graduated Dynamic Resource Allocation (DRA) to general availability, replacing the rigid device plugin model with declarative hardware requirements. This is the infrastructure shift that makes fractional GPU sharing possible at the platform level. Two schedulers matter:
KAI Scheduler
NVIDIA open-sourced KAI Scheduler under Apache 2.0 in 2025. It's the production version of the Run:ai scheduler — battle-tested in multi-tenant GPU clusters before the acquisition and open-source release.
The key capability: fractional GPU allocation. A pod requests 0.3 GPUs, and three pods share one physical device without stepping on each other's memory. KAI also understands NVLink and NVSwitch topology, so a distributed training job needing 8 GPUs gets placed on nodes with the highest interconnect bandwidth — not scattered randomly across the cluster. For multi-GPU training, gang scheduling ensures jobs either get all their GPUs simultaneously or wait in queue. No partial allocation blocking resources.
CNCF member organizations that adopted this style of scheduling saw utilization move from 13% to 70-80% in advanced implementations — a 50-70% reduction in infrastructure spend.
A caveat on the 80% ceiling: it comes from CNCF community reports, and the conditions that produce it are specific — homogeneous workloads, single-tenant scheduling, purpose-built quota policies tuned over months. Most production clusters run a messy mix of inference, training, and batch jobs. Expect 40-60% as a realistic target for a well-tuned multi-tenant environment. That's still a 3-4x improvement over baseline, and it changes the economics completely.
Kueue for non-NVIDIA stacks
For teams not on the NVIDIA ecosystem, Kueue is the Kubernetes-native job queuing system from SIG Scheduling. It handles cluster-wide queues, tenant quotas with cohort borrowing, and admission control for gang scheduling. Less GPU-topology-aware than KAI, but it works with any hardware and integrates with the standard Kubernetes scheduler.
MIG partitioning for predictable workloads
NVIDIA's Multi-Instance GPU splits a single A100 or H100 into up to seven isolated instances, each with dedicated memory and compute. The trade-off: MIG profiles must be configured before workloads run, and the partitions are fixed until reconfigured.
MIG shines for predictable inference mixes — seven 10 GB endpoints sharing a single 80 GB H100, each with hardware-level memory isolation. For dynamic workloads where demand shifts throughout the day, fractional scheduling via KAI is the better fit.
Inference Optimization: Where the Real Money Is
Scheduling fixes how GPUs are shared. Inference optimization fixes what they do.
LLM inference overtook training as the primary cost center for most production AI systems. The a16z "LLMflation" analysis documents a 10x annual price drop for equivalent model capability — GPT-3 cost $60 per million tokens in 2021; a Llama 3.2 3B delivering comparable quality costs $0.06 per million tokens today. That's a 1,000x collapse in three years.
Yet enterprise AI spending keeps climbing. This is the Jevons paradox applied to GPUs: as inference gets cheaper, teams deploy more models into more workflows, and total spend increases even as unit cost falls. Optimization doesn't reduce your bill by making AI cheaper. It reduces your bill by eliminating the waste that inflates it. Different problem, different solutions:
Model routing
Not every request needs a frontier model. A customer service bot answering "what are your business hours?" doesn't need GPT-4o. Route simple tasks to smaller, cheaper models and reserve expensive capacity for complex reasoning.
The aggregate savings: 2-5x for mixed workloads, up to 12x for classification and extraction tasks where a fine-tuned 7B model matches frontier quality. LiteLLM provides a unified OpenAI-compatible proxy across 100+ providers with per-user cost tracking — routing becomes a configuration decision, not an engineering project.
A MILL5 case study puts numbers on this: a financial services firm paying $47,000/month for document classification migrated to a fine-tuned open-source model. Cost dropped 89% to roughly $5,200/month. Accuracy held.
vLLM and PagedAttention
vLLM dominates production inference for a reason. Its core innovation, PagedAttention, manages the KV cache the way an operating system manages virtual memory — reducing memory waste from 60-80% to under 4%.
The math: a Llama 3 8B model at 8K context consumes roughly 1 GB of KV cache per concurrent request. Forty concurrent users means 40 GB of cache on top of model weights. Without PagedAttention, most of that memory fragments. With it, the same GPU handles 2-3x more concurrent requests.
The benchmarks tell the rest: 793 tokens per second versus Ollama's 41 TPS. A 3.67x throughput advantage over HuggingFace TGI under normal load that widens to 24x under extreme concurrency.
Quantization
INT8 and INT4 quantization cuts model memory 2-4x and inference cost roughly 50%, with 95-99% accuracy retention. AWQ protects only the 1% of weights that matter most — achieving better generalization than GPTQ with no backpropagation required. FP8 is the default on Blackwell GPUs. FP4 is emerging for throughput-heavy workloads.
For most inference workloads — summarization, classification, chat, RAG — full-precision serving is money left on the table. The accuracy trade-off is measurable, and it's almost always worth it.
The rest of the inference stack
Four more levers, roughly ordered by implementation effort:
Speculative decoding. A smaller draft model generates candidate tokens; the larger model verifies them in parallel. Result: 2-3x latency reduction with zero quality loss. Both vLLM and SGLang support this natively.
Semantic caching. Cache responses for semantically similar queries. Customer service and FAQ endpoints hit the same questions repeatedly — caching these at the embedding level eliminates redundant inference calls entirely.
Prompt compression. LLMLingua compresses prompts by removing redundant tokens that don't affect output quality. An 800-token customer service prompt can drop to 200 tokens or less. The savings compound at scale.
Prefill/decode disaggregation. Run compute-heavy prefill on high-end GPUs, memory-bandwidth-bound decode on cheaper hardware. SGLang's disaggregated architecture does this natively. This is the most architecturally complex optimization — save it for after the fundamentals are working.
Stack these techniques and total inference cost drops significantly. The theoretical ceiling is 5-10x. The practical range is lower — not every lever applies to every workload, and the maximums assume perfect implementation of all seven simultaneously. Even applying just routing and quantization, the two easiest wins, cuts costs 2-4x on most inference workloads.
The Self-Hosting Decision
Self-hosting doesn't eliminate the Warm Air Tax. It just moves the bill from your cloud provider to your balance sheet.
At some point, the API spend gets large enough that owning the infrastructure makes financial sense. The question is where the line sits.
The break-even: at 150 million tokens per month — roughly 4,000 active enterprise users or 120,000 customer service sessions — self-hosting on an 8x H100 cluster costs approximately $9,800/month. The equivalent through OpenAI's GPT-4o API: $657,000/month. Annual savings: $7.7 million.
Below that threshold, cloud APIs win. The overhead of running your own GPU cluster doesn't justify the savings at lower volumes.
Three hidden costs that most break-even analyses skip:
Engineering labor eats the first year of savings. Two senior ML infrastructure engineers for 6 months: $150,000-$250,000, before a single request is served. Then one FTE ongoing.
Utilization risk is circular. The 10-30% utilization problem from the cloud section? It applies to hardware you own, too. Buy 8 H100s and run them at 15%, and your theoretical $9,800/month becomes $65,000/month in effective cost per useful compute.
Depreciation moves fast. AWS cut H100 instance prices 44% in June 2025 alone. The $300,000 you invest in H100 hardware today competes with H200 and B200 availability tomorrow.
The architecture that works: hybrid. Cloud APIs for experimentation, burst capacity, and low-volume endpoints where the operational overhead of ownership can't be justified. Private infrastructure for steady production workloads — predictable demand, high volume, where the workload profile is already proven on managed compute and the cluster design follows from data, not from guessing.
Trying to self-host before proving the workload profile is how teams end up with $300,000 in hardware running at 15%. The Warm Air Tax doesn't care whether the GPU is rented or owned.
What a Well-Run AI Infrastructure Stack Looks Like
The healthcare company from the opening — $156,000/month down to $34,000 — didn't buy new hardware or switch clouds. They changed how their existing cluster was managed. Here's what that stack looks like:
- NVIDIA GPU Operator — one Helm install handles drivers, container toolkit, DCGM monitoring, and MIG manager
- KAI Scheduler or Kueue — fractional GPU allocation and job queue management
- vLLM — inference serving with PagedAttention and continuous batching
- LiteLLM — model routing, per-request cost tracking, unified API gateway
- DCGM Exporter + Prometheus + Grafana — GPU utilization monitoring that catches cost drift before it compounds
This runs on any Kubernetes cluster with NVIDIA GPUs — cloud or on-premises. Every component is open source.
The hard part isn't installing the software. It's tuning gpu_memory_utilization thresholds, setting MIG profiles for your specific workload mix, building autoscaling that responds to inference queue depth instead of CPU load, and configuring the monitoring that catches cost drift before it compounds.
Need help? Talk to an engineer.
Every tool in this stack is free. The waste is in the defaults.
Frequently Asked Questions
What is AI infrastructure and why is it expensive?
AI infrastructure is the compute, storage, networking, and software stack required to train and serve machine learning models. Costs are high because GPU hardware runs $27,000-$40,000 per unit, cloud GPU instances cost $2-$15 per hour, and most organizations achieve only 10-30% utilization — meaning 70-90% of spending produces no useful work.
How do I reduce GPU costs on Kubernetes?
Deploy fractional GPU scheduling using NVIDIA's open-source KAI Scheduler or Kubernetes-native Kueue for job queuing. Enable MIG partitioning for mixed inference workloads. Add DCGM Exporter for real-time utilization monitoring. These changes alone move utilization from 10-30% to 37-80%, cutting effective per-workload costs by 50-70% without adding hardware.
When does self-hosting AI models make financial sense?
Self-hosting breaks even at roughly 150 million tokens per month — about 4,000 active enterprise users. Below that, cloud APIs are cheaper after accounting for engineering labor ($150K-$250K setup) and ongoing operations (one full-time engineer). Above that threshold, self-hosted infrastructure on 8x H100 GPUs saves $1.5-$7.7 million annually versus API pricing.
What is vLLM and why does it matter for inference costs?
vLLM is the dominant open-source LLM inference serving framework. Its PagedAttention mechanism reduces KV cache memory waste from 60-80% to under 4%, and it delivers 793 tokens per second versus Ollama's 41 TPS. vLLM supports continuous batching, tensor parallelism, speculative decoding, and quantization natively.
How much can inference optimization reduce AI infrastructure costs?
Combining model routing (2-5x savings), quantization (50% cost cut at 95%+ accuracy), PagedAttention (2-3x more concurrent requests), speculative decoding (2-3x latency reduction), and semantic caching for repetitive workloads delivers 3-10x total inference cost reduction depending on workload mix — significantly more than any hardware upgrade alone.
