Anthropic's Fable 5.1 Tops Global Coding Benchmark at 38.8%, Ahead of GPT-6 Astra
Anthropic's Fable 5.1 now leads the SWE-bench coding benchmark at 38.8% — the first time Anthropic has held the top coding position against OpenAI and Google.
Top Stories
All Sections
Finance
About
12 results for "benchmarks"
Anthropic's Fable 5.1 now leads the SWE-bench coding benchmark at 38.8% — the first time Anthropic has held the top coding position against OpenAI and Google.
Cognition's SWE-2 matches the frontier coding leaderboard at 64% lower cost — and leads all benchmarked models on Terminal-Bench 2.1 at 92.8%.
Microsoft launched MDASH, a multi-model agentic security system that coordinates AI agents to detect and respond to threats at machine speed — and claims top scores on industry cybersecurity benchmarks.
A new study finds patients consistently withhold symptoms from AI health chatbots, creating a diagnostic gap that undermines real-world accuracy even as AI outperforms doctors on benchmarks.
China installs 295,000 industrial robots annually versus 34,200 in the US — an 8.6-to-1 gap that signals the US-China AI competition is moving from model benchmarks to factory floors.
An OpenAI reasoning model outperformed two experienced ER physicians in a real-world test at Beth Israel Deaconess Medical Center — one of the first head-to-head comparisons conducted with live clinical data rather than curated research benchmarks.
A study published April 29 finds AI systems produce correct answers while fundamentally failing to understand the underlying concepts — challenging benchmarks used to certify AI for high-stakes deployment.
DeepSeek released V4 Flash and V4 Pro on April 24 with top-tier coding benchmarks and a 1-million-token context window—open-source, again, at a moment when Silicon Valley assumed the gap was widening.
Frontier AI models now solve real software engineering tasks with near-perfect accuracy — but the same report finds leading AI systems are disclosing less about how they work than ever before.
The Stanford AI Index 2026 finds top AI agents complete complex scientific research tasks at half the rate of human PhD experts — a significant check on agentic AI hype, even as frontier models exceed 50% on Humanity's Last Exam.
The 2026 Stanford AI Index found China closed the US AI performance lead from 9.26% to 1.70% in one year — while major labs slashed transparency scores by 31%.
Claude Opus 4.6 and Gemini 3.1 Pro have passed 50% on Humanity's Last Exam — a benchmark designed to be unsolvable by current AI systems.