# The Inference Paradox

> Token prices are falling by more than 90 percent. The cost of a finished agent workflow is going up fivefold. Both are true at once.

Source: https://savrn.com/blog/the-inference-paradox
Author: Chad Everett Harris
Published: 2026-08-23

---

On August 17, 2026, Gartner forecast that the cost of running an agentic AI workflow will rise more than fivefold through 2028 ([Gartner](https://www.gartner.com/en/newsroom/press-releases/2026-08-17-gartner-predicts-ai-inference-costs-per-agentic-workflow-will-increase-more-than-fivefold-through-2028)). It also expects the per-token price of frontier inference to fall by more than 90 percent by 2030. Gartner calls the combination the inference paradox. The visible unit price goes down. The cost of the thing you actually buy, a completed workflow, goes up.

## The equation

Cost per workflow = agents × steps × tokens per step × tier multiplier × price per token. Four terms trend up, set by architecture choices inside the enterprise. One trends down, set by suppliers. Tuned to Gartner's headline: 2.35 × 2.60 × 1.95 × 2.47 = 29.4 of complexity growth, times 0.17 of token deflation, equals 5.0.

## The evidence

- NVIDIA: reasoning can demand 100 times more compute; the vast majority of NVIDIA's compute today is inference.
- Anthropic: agents use about 4 times the tokens of chat, multi-agent systems about 15 times; token usage explains 80 percent of performance variance on BrowseComp ([Anthropic](https://www.anthropic.com/engineering/multi-agent-research-system)).
- Microsoft Research: agentic coding consumes 1,000 times the tokens of code chat, with 30 times variance on identical tasks ([Microsoft Research](https://www.microsoft.com/en-us/research/publication/how-do-ai-agents-spend-your-money-analyzing-and-predicting-token-consumption-in-agentic-coding-tasks/)).
- a16z: cost for equivalent performance falls 10 times a year; o1 launched at the same $60 per million tokens GPT-3 charged in 2021 ([a16z](https://a16z.com/llmflation-llm-inference-cost/)).
- Gartner: more than 40 percent of agentic AI projects canceled by end of 2027. MIT NANDA: about 5 percent of pilots reach rapid revenue impact. S&P Global: enterprises abandoning most AI initiatives jumped from 17 to 42 percent in a year.

## The trap

Five gates, each closed by default: renting the model layer, letting complexity grow team by team, no tier discipline, ungoverned swarms, no decision provenance.

## The way through

Four disciplines: ownership, governance, orchestration, provenance. Same 83 percent token deflation, three futures: a floor near 0.5×, Gartner's base of 5×, a ceiling near 14×. The difference is not the technology. It is the operating discipline.
