AI Agents
How Claude Opus 4.5 Broke the 80% SWE-bench Barrier, and What’s Overtaken It Since
When Anthropic released Claude Opus 4.5 in November 2025, it became the first model to clear 80% on…
Auton AI News covers SWE Bench Verified results, rankings, and breakthroughs as AI coding agents compete to solve real software engineering tasks.
When Anthropic released Claude Opus 4.5 in November 2025, it became the first model to clear 80% on…
Anthropic's Claude Opus 4.5 didn't just move the benchmark needle, it broke through the 80% mark on SWE-bench…