AI model benchmarks
5 stories tagged AI model benchmarks.
GPT-6.1 Sol and Claude Sonnet 5.5 Test the Cost of AI
GPT-6.1 Sol used fewer tokens than Claude Sonnet 5.5 in one trial but took longer. Prices, benchmark settings and failed attempts complicate the efficiency claim.
Claude Opus 4.8: Honest Upgrade or Playing Catch-Up?
Claude Opus 4.8: Honest Upgrade or Playing Catch-Up?
Anthropic's Claude Opus 4.8 drops with better honesty, dynamic multi-agent workflows, and a $965B valuation. But is it enough to reclaim momentum from OpenAI?
Kimi K2.6 Nails Agent Tasks But Burns More Tokens Than Its Predecessor
Kimi K2.6 Nails Agent Tasks But Burns More Tokens Than Its Predecessor
Moonshot's Kimi K2.6 ranks #2 on OpenClaw with perfect usable fit, but costs more than K2.5 on basic coding. The efficiency tradeoff explained.
MiniMax M2.5 Claims to Match Top AI Models at 5% the Cost
MiniMax M2.5 Claims to Match Top AI Models at 5% the Cost
Chinese AI firm MiniMax releases M2.5, an open-source coding model claiming performance comparable to Claude and GPT-4 at dramatically lower prices.
This Tiny Open-Source OCR Model Just Beat Gemini Pro
This Tiny Open-Source OCR Model Just Beat Gemini Pro
GLM OCR is a 0.9B parameter model that outperforms Gemini Pro at reading handwriting, tables, and formulas—and it runs on your laptop for free.