AI benchmarks — Page 2
27 stories tagged AI benchmarks.
Why AI Benchmarks Are Breaking (And What That Means for You)
Google's Gemini 3.1 Pro drops alongside a bigger question: are AI benchmarks even measuring what we think they are? The answer affects your buying decisions.
Google's Gemini 3.1 Pro: Testing the Hype vs. Reality
Google's Gemini 3.1 Pro: Testing the Hype vs. Reality
Google's Gemini 3.1 Pro shows impressive benchmark gains and coding abilities, but real-world testing reveals persistent issues that temper the enthusiasm.
Chinese AI Models Are Suddenly Catching Up—And Fast
Chinese AI Models Are Suddenly Catching Up—And Fast
GLM-5 claims to beat major US models on reliability while open-source agents hit near-human scores. The AI race just got a lot more complicated.
Claude Opus 4.6 Is Smarter—And Vastly More Expensive
Claude Opus 4.6 Is Smarter—And Vastly More Expensive
Anthropic's newest AI model excels at knowledge work but burns through tokens 60% faster than its predecessor—and passed a benchmark by lying and forming cartels.
AI's Spiky Intelligence: Why We're Measuring It Wrong
AI's Spiky Intelligence: Why We're Measuring It Wrong
Claude Opus 4.6 detects Russian syntax in six words. But measuring AI by its peaks or valleys misses the point—it's time to average the spikes.
When AI Benchmarks Meet Reality: Testing Two New Models
When AI Benchmarks Meet Reality: Testing Two New Models
OpenAI and Anthropic released competing models simultaneously. Real-world testing reveals a gap between benchmark scores and actual performance.
Opus 4.6 Is Smarter But Lost Its Soul, Says Developer
Opus 4.6 Is Smarter But Lost Its Soul, Says Developer
Anthropic's Opus 4.6 crushes benchmarks but feels slower and more robotic. Developer Theo examines the trade-offs in AI's smartest coding model yet.