Llama 3, DeepSeek Narrow Gap

Llama 3, DeepSeek Narrow Gap
Key Takeaways

  • Open-source models from DeepSeek, Moonshot and Zhipu now post coding scores within a few points of leading closed models, while costing 10x to 30x less per token, a gap that is making proprietary-only strategies harder to justify on cost grounds alone.
  • Meta’s Llama 3.1 scored 96.82% on the GSM8K mathematical reasoning benchmark, beating GPT-4o’s 94.24%, showing that open models are not just cheaper but can lead on specific high-value tasks.
  • Nebius Group reported one customer achieving up to a 26x inference cost reduction by switching to open models; at scale, that gap, roughly $2.50 versus $34 $60 per month per 10 million tokens, makes open-source the rational default for most enterprise workloads, with proprietary models reserved for tasks where the performance premium is worth paying.

Open-source AI has stopped playing catch-up. On several benchmark categories, models that anyone can download and self-host are now beating the closed systems that cost orders of magnitude more to run. The performance gap between the best open and proprietary models has narrowed to single digits on most real-world task benchmarks, and for many enterprise workloads, it has effectively closed.

Open-Source Models Close Performance Gaps with Proprietary Rivals

A June 2026 report found that open-weight systems from DeepSeekMoonshot and Zhipu now post coding scores within a few points of the best closed models. The gap is not uniform across task types, though, and that unevenness is worth paying attention to.

Take mathematical reasoning. Meta‘s Llama 3.1 scored 96.82% on the GSM8K benchmark, above GPT-4o’s 94.24%. On coding, GPT-4o retakes the lead at 92.07% accuracy. On graduate-level reasoning tasks, Llama 3 scored 35.7% against GPT-4’s 39.5%, a gap, but a smaller one than many expected. On the MMLU 5-shot test, Llama 3 reached 82% against GPT-4-turbo’s 86.4%. The picture that emerges is not one model winning outright but a patchwork: open models leading in some domains, trailing narrowly in others, and closing fast across the board.

Dramatic Cost Reductions Drive Open-Source Adoption

The economics are stark. A paper co-authored by Frank Nagle at the MIT Initiative on the Digital Economy found that open models cost, on average, six times less than closed models, $0.83 per million tokens versus $6.03 for proprietary alternatives, an 86% saving.

At real-world scale, those fractions compound fast. Nebius Group reported a customer cutting inference costs by up to 26 times by switching to open-source models. For an organisation processing 10 million tokens per month, roughly 7.5 million words, running Qwen3-235B works out to around $2.50 per month. Compare that to Claude 4.5 Sonnet at $60 per month or GPT-5 at $34.40 per month for the same volume. Document analysis, customer triage, code review: for most of these bread-and-butter enterprise tasks, the cost argument for open-source is difficult to dismiss. Self-hosting removes token-based pricing entirely, though it trades that cost for hardware and infrastructure overhead.

Enterprise Integration and Strategic Advantages

Wells Fargo is piloting Llama-based workflows. Goldman Sachs is using open-source models for research and automation tasks. IBM is integrating open-source large language models into internal productivity tools. The common thread is not just cost: it is control. Fine-tuning on proprietary data without vendor lock-in matters especially in regulated industries where the data itself cannot leave the building.

Intuit’s approach is instructive. The company built its Intuit Assist platform on a mix of open-source models, which lets its teams adopt new model improvements quickly and adapt them to tax, accounting and marketing workflows without rebuilding from scratch each time. That flexibility is harder to replicate when every capability upgrade requires renegotiating a vendor contract.

Gartner forecast that more than 60% of businesses would adopt open-source LLMs for at least one AI application by 2025, up from 25% in 2023. Whether that projection has held is harder to verify from public material, but the direction of enterprise adoption visible in public disclosures is consistent with it.

The Evolving Competitive Landscape and Future Outlook

Alibaba’s Qwen family crossed one billion Hugging Face downloads in January 2026, accounting for more than half of all open-model downloads globally, according to reports. That number matters because it captures not just adoption but the community flywheel that drives open-source improvement: more users, more fine-tunes, more pressure on the base models to improve.

Meta and Google releasing Llama and Gemma respectively is not altruism, it is a deliberate strategy to commoditise the model layer while competing on cloud infrastructure, hardware and the application stack on top. For enterprises, the practical consequence is that frontier-quality models are becoming a cost input rather than a competitive moat.

Proprietary models still lead on elite tasks: complex multi-step reasoning, the harder multimodal problems, the long-context edge cases, but the performance lead is narrowing. The lag between frontier closed models and the best open alternatives, once estimated at 18 months, has compressed to roughly 6 to 10 months. The decision for most enterprises is becoming less “open versus closed” and more a question of which workloads justify the premium. For more coverage of AI research and breakthroughs, visit our AI Research section.

Taylor Voss
Taylor Voss

Taylor Voss is a research correspondent covering AI breakthroughs, model releases, and the science shaping the future of artificial intelligence. Taylor translates complex research into clear, compelling stories for curious readers who want more than the headlines.

📰 Journalists welcome — cite Auton AI News with attribution. Press & Media → | press@autonainews.com