- Frontier model performance gains per unit of compute are visibly flattening, shifting competition from raw capability to cost efficiency, and OpenAI’s $2.3 billion inference bill in 2024 is the clearest evidence of where the pressure lands.
- The stock of human-generated public text data could be fully exhausted as early as 2028; synthetic data is the leading alternative, but research warns it risks “model collapse” through accumulated bias and increased hallucinations.
- Meta’s LLM-JEPA and Google’s Gemini Diffusion represent active investment in architectures beyond the Transformer, targeting the abstract reasoning and inference-cost limits that scaling alone has not solved.
OpenAI spent roughly $2.3 billion on inference in 2024, about 15 times what it cost to train GPT-4. That ratio is the clearest signal yet that the AI industry’s central problem has changed: less about building more capable models, more about affording to run them, as performance gains per dollar of compute flatten at the frontier.
Flattening Returns
The competitive pressure at the top of the model range has quietly shifted from raw capability to cost efficiency. Performance across large language models continues to advance, but the rate of improvement for a given compute input is visibly slowing. Adding more parameters or more training compute no longer reliably produces proportional gains. What that means practically is that the cost-to-performance ratio has replaced raw scale as the operative metric, and that creates very different competitive dynamics than the ones that defined the first wave of frontier model development. Llama 3 8B (2024) outperforming Falcon 180B (2023) despite a fraction of the parameters, by training on substantially more and higher-quality data, illustrates the point concisely. Brute compute is not the lever it was.
Running Out of Human-Written Text
The finite stock of human-generated public text data faces full utilisation in the coming years, a timeline accelerated by intensive data reuse across training runs, with some projections pointing to as early as 2028.
Synthetic data is the obvious response, but research warns that over-reliance on AI-generated training material can trigger what researchers call “model collapse”, a degradation cycle where models trained on their own outputs gradually accumulate bias and produce more hallucinations. How severe that effect is in practice, and at what ratio of synthetic-to-human data it becomes significant, remains an open question.
Where Abstract Reasoning Stalls
Scale helps with fluency. It helps less with thinking. A June 2026 report from MindStudio finds that on analogical reasoning tasks, those requiring pattern recognition across structural relationships rather than recall of trained patterns, larger models do not reliably outperform smaller ones. The distinction maps roughly onto what cognitive scientists call System 1 and System 2 processing: fast, fluent retrieval versus deliberate, abstract inference. Current scaling improves the former; the latter resists it. For applications that depend on novel problem-solving or generalisation beyond the training distribution, that gap is a real constraint.
The Inference Cost Problem
OpenAI’s $2.3 billion inference spend in 2024, roughly 15 times the cost of training GPT-4, captures something important about where the economics of AI deployment have landed. Inference costs this large signal that test-time compute is not a cheap operational afterthought to pre-training; it is increasingly where capability gains are being purchased, because additional pre-training is no longer reliably delivering them. The cost pressure is also one reason investment in alternative architectures has accelerated: if inference on Transformer-based models is this expensive, the incentive to find more efficient paths is substantial. Gartner has warned that organisations may be miscalculating AI costs by as much as 1,000% at scale a gap that becomes harder to ignore as inference bills grow.
Beyond the Transformer
Meta chief AI scientist Yann LeCun has argued that autoregressive LLMs are fundamentally limited and called for world models over word predictors. The alternatives attracting serious investment reflect that scepticism. State space models like Mamba scale linearly with context length, making million-token inputs tractable where attention-based models struggle. Meta’s Joint Embedding Predictive Architecture for language (LLM-JEPA, September 2025) targets genuine planning rather than next-token prediction. Google‘s Gemini Diffusion claims text generation roughly 10 times faster than autoregressive models, according to the company, though independent benchmarking remains limited. None of these has displaced the Transformer at scale. The investment in alternatives, however, is no longer speculative, it reflects a specific set of limitations that scaling has not resolved.
Power as the Hard Ceiling
Physical infrastructure is now the binding constraint for many AI build-outs. The International Energy Agency projects global data centre electricity use will roughly double by 2030, with AI driving an estimated 30% annual increase in consumption from accelerated compute servers, around four times the growth rate across other sectors combined. Worldwide AI infrastructure spending is projected to reach roughly $487 billion in 2026 and surpass $1 trillion by 2029, with a growing share going to land, power and network connectivity rather than semiconductors. A McKinsey survey from August 2026 reports that one in five organisations is already limiting AI use due to operating costs.
Efficiency as the New Metric
DeepSeek V4, released in early 2026is the clearest recent example of what efficiency-first development looks like at the frontier. The model’s approach, prioritising performance per dollar over raw capability, is the direction competitive pressure is now pushing across the industry. As Palo Alto’s CEO has argued, AI pricing itself may need to drop 90% by 2028 for enterprise deployment to remain viable at current usage trajectories. The metric that matters going forward is what a model costs to run against what it can actually do.



