AI benchmarks
32 stories tagged AI benchmarks.
Ox Alpha: Anonymous AI Model Stumps the Industry
An anonymous AI model called Ox Alpha appeared on OpenRouter, beat top coding benchmarks, and triggered a forensic manhunt. Nobody's claimed it yet.
Harvey Tenet Legal AI Model: Cool Tech, Thin Proof
Harvey Tenet Legal AI Model: Cool Tech, Thin Proof
Harvey's Tenet model applies async RL to long-horizon legal tasks—genuinely interesting tech. But verified benchmarks are scarce. Here's what we know and what we don't.
Data Is Now the Hard Part of Building AI
Data Is Now the Hard Part of Building AI
At a recent YC Paper Club session, three AI researchers made the case that training data—not models or chips—is where the real work of building AI happens now.
AI Coding Agents Still Need a Human in the Loop
AI Coding Agents Still Need a Human in the Loop
Dexter Horthy built a fully automated software factory—then watched it corrupt his codebase. His case for keeping humans in the loop is harder to dismiss than most.
Claude Opus 5 Beats Fable 5 on Benchmarks at Half the Price
Claude Opus 5 Beats Fable 5 on Benchmarks at Half the Price
Anthropic's Claude Opus 5 outperforms Claude Fable 5 on most benchmarks at half the price. Here's what the numbers actually mean for developers.
Kimi K3 Benchmarks vs. Real-World Performance
Kimi K3 Benchmarks vs. Real-World Performance
Moonshot's Kimi K3 posts frontier-class benchmarks, but early testing reveals real gaps in reliability, speed, and cost. Here's what the numbers actually show.
How AI Is Actually Tested for Human-Level Intelligence
How AI Is Actually Tested for Human-Level Intelligence
ARC AGI 3 tasks AI with figuring out video games from scratch — and current models can barely start them. Here's what that reveals about the gap between AI and human reasoning.
Tencent HY3 Reviewed: Free, Open Source, and Uneven
Tencent HY3 Reviewed: Free, Open Source, and Uneven
Tencent's HY3 is a free, 295B open-source model with real agentic strengths—but benchmark scores and real-world output quality tell different stories.
GPT 5.6 Sol vs Fable 5: Early Numbers, Real Tradeoffs
GPT 5.6 Sol vs Fable 5: Early Numbers, Real Tradeoffs
GPT 5.6 Sol is half the price of Fable 5 — but is it half as good? Early benchmark comparisons, alignment regressions, and the politics reshaping who gets access.
AI Leaderboards Are Lying to You About State-of-the-Art
AI Leaderboards Are Lying to You About State-of-the-Art
Bertrand Charpentier of Pruna AI makes the case that 'state-of-the-art' is a broken concept—and that efficiency belongs in the same sentence as quality.
AI Benchmark Scores Are Broken. Here's Who's Fixing Them.
AI Benchmark Scores Are Broken. Here's Who's Fixing Them.
AI benchmark scores are less trustworthy than they look. Google DeepMind's Kaggle team is building open infrastructure to fix that—here's what you need to know.
GPT-5.5 Is Great, But You Might Not Notice—Here's Why
GPT-5.5 Is Great, But You Might Not Notice—Here's Why
OpenAI's GPT-5.5 dominates benchmarks and handles complex coding tasks, but many users won't feel the upgrade. We dig into the paradox.
GPT 5.5 vs DeepSeek V4: The Benchmarks Tell a Jagged Story
GPT 5.5 vs DeepSeek V4: The Benchmarks Tell a Jagged Story
OpenAI and DeepSeek released flagship models within 20 hours. The benchmark results reveal something more interesting than who's winning.
The Benchmark Paradox: What Qwen 3.6's Numbers Actually Mean
The Benchmark Paradox: What Qwen 3.6's Numbers Actually Mean
Qwen's new 27B model is beating models 10x its size—on paper. Here's what those benchmarks aren't telling you about AI performance.
OpenAI's GPT-5.5 Claims Speed Crown—But Costs 20% More
OpenAI's GPT-5.5 Claims Speed Crown—But Costs 20% More
GPT-5.5 promises faster AI coding with fewer tokens, but WorldofAI's tests reveal where it excels—and where it disappoints at premium pricing.
OpenAI's GPT-5.5: When the Benchmarks Don't Tell the Whole Story
OpenAI's GPT-5.5: When the Benchmarks Don't Tell the Whole Story
GPT-5.5 arrives with impressive real-world benchmarks and doubled pricing. But the coding results reveal tensions in how we measure AI capability.
Anthropic's Opus 4.7: The Enterprise Model You Can't Afford
Anthropic's Opus 4.7: The Enterprise Model You Can't Afford
Anthropic's Opus 4.7 excels at enterprise tasks but costs 35% more due to tokenizer changes. The upgrade everyone's complaining about, explained.
Three AI Models Just Dropped—Here's What Actually Matters
Three AI Models Just Dropped—Here's What Actually Matters
Meta's Muse Spark, Z.ai's GLM 5.1, and Anthropic's Managed Agents all launched this week. Here's what they're good at—and what they're not.
Anthropic's Mythos Launch: Security Theater or IPO Theater?
Anthropic's Mythos Launch: Security Theater or IPO Theater?
Anthropic's Project Glasswing positions Mythos as too dangerous to release. The timing before a $380B IPO raises questions about the narrative's purpose.
AI Benchmarks Are Breaking. Here's Why That Matters.
AI Benchmarks Are Breaking. Here's Why That Matters.
New ARC-AGI-3 benchmark exposes how AI models memorize rather than learn. Humans score 100%, frontier AI models score less than 1%. The gap reveals everything.