The Math Nobody Wants to Look At

Six months ago I sat in a room with three different teams who’d all made the same discovery independently. They’d built impressive AI products. They’d shipped to production. And then the AWS bill showed up.

The pattern was identical across all three: inference costs weren’t the tail risk they’d assumed during planning. They were the whole dog. When Andreessen Horowitz published their 2025 State of AI report, they quantified what I’d been hearing anecdotally for months. At scaled startups, inference spending now consumes between 60% and 80% of total AI infrastructure budgets. Training? That’s a rounding error by comparison. This wasn’t a surprise to anyone who’s actually run these systems at scale, but seeing it in writing changes how leadership thinks about the problem.

The issue is structural. Training happens once. You pay that cost, you move on. Inference happens every single time a user interacts with your product. If you’re processing thousands or millions of requests daily, those token costs compound into something genuinely alarming. Most teams don’t feel the pain until they’re already committed to a specific architecture.

The Tiered Routing Architecture Is No Longer Optional

The smarter teams have stopped asking whether they should implement model routing. They’re asking how quickly they can get it running.

Here’s the reality that took time to sink in for most of us: not every request needs GPT-4. Actually, most requests don’t. A subset of your queries demand the full power of the largest models. The rest can get routed to something smaller and cheaper without degrading user experience. OpenAI’s gpt-4o-mini, released in July 2024, costs $0.15 per million input tokens. Compare that to gpt-4o at $5.00. That’s over a 30x price difference. You can’t ignore that math.

The Databricks State of Data and AI 2025 report found that 43% of enterprise AI teams have already implemented model routing or cascading strategies. This isn’t theoretical anymore. It’s table stakes in certain organizations. The idea is straightforward: classify incoming requests by complexity, route simple ones to cheaper models, reserve expensive models for queries that actually need them, measure quality carefully, iterate.

The infrastructure to do this has gotten substantially better. AWS introduced Bedrock Intelligent Prompt Routing in late 2024. The feature examines incoming prompts and automatically routes them to the most cost-effective model that can still meet your quality thresholds. AWS is targeting up to 30% cost reduction for teams that implement it. That’s not trivial when you’re running millions of inferences per day. For more details, check the AWS Bedrock Intelligent Prompt Routing overview.

Latency Isn’t Free Either, But There Are Tricks

Cost optimization creates a new problem: speed. Route everything to smaller models and you save money. But if responses get slower, users notice. That tension is real and worth acknowledging directly.

This is where speculative decoding enters the conversation. Google Cloud added support for this technique in Vertex AI Model Garden in late 2024. The approach uses a smaller, faster draft model to pre-fill token sequences, then a larger model validates and refines the output. The result is significant: inference latency can drop by 2x to 3x on large models. You’re maintaining quality while cutting response time. It’s not magic, but it’s close.

The tradeoff is complexity. You need to orchestrate two models instead of one, with monitoring and fallback logic on top of that. But if you’re already running at scale and price matters, the engineering investment is worth it. Teams that have implemented it report that users don’t perceive any quality degradation, just faster responses.

The Organizational Shift Required

Here’s what surprised me most while talking to teams about this: the technical solution is actually easier than the organizational one.

Most teams started their AI projects with a simple question: can we build this? They answered yes. They shipped it. Then the infrastructure team got involved and asked a different question: can we afford this? The answer was often no. But by then, product requirements had been locked in, customer expectations set, and the architecture hardened into place.

The teams executing well now handle this differently. They ask the cost question during architecture design, not after. They build observability into model routing decisions from day one. They measure not just latency and quality, but cost per transaction. They treat inference spend as a first-class constraint, like memory usage or database queries. It’s a mindset shift more than a technical one.

The vendors have noticed this too. AWS, Azure, and Google are all shipping tools specifically designed to make routing easier. This wasn’t on the roadmap two years ago. But when 70% of your infrastructure spend is going somewhere, you build tools to help people control it.

What Actually Matters Now

If you’re building or operating AI products at any meaningful scale, start measuring your inference costs today. Break them down by model, by request type, by user segment if possible. Find where the money is actually going. Most teams discover that a small percentage of their requests consume a large percentage of their budget. That concentration is where the leverage is.

Then get stubborn about cost per transaction. Track it like your business depends on it, because it does. Implement tiered routing. Experiment with cheaper models for specific use cases. Monitor quality carefully. The goal isn’t to cut corners. It’s to match model capacity to actual requirements.

The teams winning this game right now aren’t the ones who picked the most impressive model. They’re the ones who got disciplined about the economics early. They’ll outrun everyone else because they’ll still be profitable when others are still trying to figure out where all the money went.

What’s your inference budget looking like? Have you started measuring costs by model yet, or is that still on the backlog? I’m genuinely curious what’s working and what’s not in your stack.