Agentic AI Workflows Drive 320% Spend Amidst 280x Token Plunge

Agentic AI Workflows Drive 320% Spend Amidst 280x Token Plunge
Key Takeaways

  • LLM token costs have fallen roughly 280-fold over two years, with equivalent model performance now available at around $0.40 per million tokens, down from $20 $60 per million in late 2021 to early 2023, according to Andreessen Horowitz’s November 2024 research.
  • Private generative AI investment reached $33.9 billion in 2024, even as token prices collapsed, a textbook Jevons Paradox: cheaper compute drives more consumption, not less.
  • A January 2026 arXiv paper found agentic workflows can consume up to 1,000 times more tokens than simple chat tasks, shifting the real cost management problem from token pricing to system architecture design.

Token prices have collapsed 280-fold in two years, yet enterprise AI budgets are growing faster than ever. The reason is architectural: as inference gets cheaper, enterprises are building far more complex systems to run on it, and those systems are expensive to operate at scale.

The Collapsing Cost of AI Inference

Equivalent model performance that cost $20 to $60 per million tokens in late 2021 to early 2023 now runs at around $0.40 per million. According to Andreessen Horowitz‘s November 2024 research, the cost of LLM inference for equivalent performance dropped by a factor of 10 every year, adding up to a roughly 1,000-fold reduction over three years since OpenAI introduced GPT-3 in November 2021. That pace is not uniform. Epoch AI‘s March 2025 analysis found that the rate of price decline varies by performance milestone, ranging from 9x to 900x per year, with the steepest drops arriving after January 2024. The drivers are familiar: GPU efficiency gains, model quantisation and sustained price competition among providers. For enterprise planning, the key point is that this is not a temporary discount, it is a structural repricing of AI compute.

Enterprise Spending Soars: The Jevons Paradox in Action

Cheaper compute was supposed to reduce AI costs. Instead, enterprise AI spending rose an estimated 320% over the same two-year period in which token prices fell 280-fold. This is the Jevons Paradox at work: efficiency gains get absorbed by expanded deployment, not returned as savings.

The mechanism is straightforward. Lower token prices unlocked use cases that were previously too expensive to run in production. Enterprises that once ran narrow, single-turn query tools are now deploying multi-step pipelines, retrieval-augmented systems and early agentic workflows, all of which consume substantially more tokens per task. Private generative AI investment alone reached $33.9 billion in 2024, a figure that reflects expansion in scope and scale, not cost optimisation on existing workloads. The savings from cheaper inference are being reinvested into more ambitious architecture, which drives consumption back up. Understanding how inference compute reshapes AI scaling is increasingly central to enterprise budget planning.

Agentic Workflows: Token Consumption Multiplied

The structural cost pressure is coming from architecture, not pricing. Agentic AI systems, which autonomously plan and execute tasks through multiple iterative LLM calls, consume tokens at a fundamentally different scale than single-turn interactions. A January 2026 arXiv paper, “Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering,” found that agentic tasks can consume up to 1,000 times more tokens than simpler code reasoning or chat tasks. The cost is driven primarily by input tokens, the context and instructions passed to the model on each call, rather than generated output. In agentic software development tasks specifically, the iterative code review stage alone accounts for an average of 59.4% of all tokens consumed, per that paper.

A single agentic operation can require the equivalent of 10 to 20 LLM calls, or five to 30 times more tokens per task than a direct query. At 280-times-cheaper token prices this is still manageable, but as agentic workflows scale and chain together, cumulative costs compound quickly. The optimisation problem shifts from negotiating a better token rate to redesigning how the system manages state, memory and task routing.

Rethinking AI Economics for Strategic Advantage

Enterprises moving from proofs of concept to platform-scale deployment are finding that AI’s cost profile is determined less by what they pay per token and more by how many tokens their workflows consume per task. That reframe has real consequences for procurement, architecture reviews and vendor selection.

Token costs will likely keep falling, which will keep unlocking new use cases and pushing total consumption higher. Agentic workflows, meanwhile, will keep growing in complexity and token appetite. Enterprises that track cost per workflow outcome rather than cost per token are better placed to manage that tension, and to make the case internally when AI budgets keep rising even as unit prices fall. For more analysis on enterprise AI strategy, visit our Enterprise AI section.

Morgan Blake
Morgan Blake

Morgan is a technology analyst covering enterprise AI strategy, automation, and business transformation. Morgan tracks how organisations are deploying AI at scale.

📰 Journalists welcome — cite Auton AI News with attribution. Press & Media → | press@autonainews.com