Anthropic's Internal Model 2 and a Week of AI Shifts
Anthropic's internal Model 2 outpaces Mythos 5 on R&D benchmarks. Plus: DeepSeek pricing climbs, Gemini 3.7 Flash drops, and compute costs keep rising.
Written by AI. Marcus Chen-Ramirez

Photo: AI. Soraya Hadid
The most revealing thing about this particular week in AI isn't any single model release. It's the cluster of small signals all pointing in the same direction: the systems are getting more capable faster than the infrastructure—economic, regulatory, physical—can comfortably absorb them.
Start with Anthropic, because the story there is genuinely interesting in a way that the usual benchmark theater usually isn't.
The Model That Isn't (Yet)
Buried in Anthropic's latest risk report is a reference to something internally called "Model 2"—described as a Mythos-class model that's currently internal only and "overall slightly more capable than Mythos 5." The WorldofAI video flagged this disclosure, and it's worth sitting with what Anthropic actually revealed versus what the speculation around it has inflated.
The concrete data point: Anthropic tested Model 2 on CoBench v2, an internal benchmark built around historical AI R&D tasks that Anthropic's own researchers have actually performed. Mythos 5 scored 50.3% on this benchmark. Model 2 scored 62.8%—a 12.5 percentage point jump. For context on why that matters: Anthropic has stated that a model scoring around 85% on CoBench v2 could potentially substitute for its research staff on the tasks the benchmark represents.
That's the number that makes people sit up. Model 2 is at 62.8%. The "automate significant research work" threshold is at roughly 85%. That's a meaningful gap—but the trajectory from Mythos 5 to Model 2 closed about half that distance in a single internal iteration.
The WorldofAI presenter is appropriately careful here: "I wouldn't jump right away and say that this is RSI or that AI researchers are about to get completely replaced. This is one internal benchmark covering a particular set of historical tasks. So, let's keep the hype down."
The other thing worth noting: the risk report disclosing these figures was dated July 15th. Whatever Model 2 looks like now, those numbers are already weeks old. Anthropic has also acknowledged it hasn't run its full pre-deployment evaluation suite on the model yet—so "slightly more capable than Mythos 5" is a preliminary read, not a ceiling measurement.
Why isn't it released? The video suggests a practical constraint: compute. A model requiring substantially more inference resources than Mythos 5, deployed at consumer scale to millions of users, is an expensive proposition. Whether that's the whole story or just the convenient story is harder to know from the outside. For those tracking Anthropic's Mythos rollout and the cautious, staged way the company has approached its most powerful models, a deliberately slow internal deployment fits the pattern. And given how Claude Mythos has already broken benchmark ceilings in other evaluations, the restraint makes a kind of institutional sense.
The Watermark Controversy
Also from Anthropic this week: the ongoing fallout over Claude's text watermarking system, implemented to comply with EU AI Act transparency requirements. The initial reaction from users involved significant alarm—hidden Unicode characters, secret metadata, identifiers embedded in every response.
The reality is more mundane and more nuanced at the same time. Anthropic states the watermark doesn't add hidden characters, metadata, or extra tokens. Instead, it subtly influences word choice to create a statistical pattern that detection tools can recognize. The watermark can't identify you, your account, or the specific conversation—only that Claude was likely involved.
But the edge case generating the most friction: text you wrote yourself and later had Claude edit heavily could potentially read as watermarked Claude output. That's not a conspiracy; it's a technical consequence of how statistical watermarking works across collaborative human-AI text. The problem is real even if the implementation is less sinister than feared.
Anthropic is applying the watermark globally—not just in the EU—because it currently can't reliably restrict it by region. That's a practical explanation. It also means every Claude user worldwide is operating under a system designed primarily for European regulatory compliance, with limited ability to opt out.
Paying to Get Unstuck: The Codex Reset Model
OpenAI's Codex is rolling out paid usage resets. According to the WorldofAI breakdown, Pro subscribers—who pay $200 per month—are being offered a full weekly reset for $80. The catch generating the most complaints: redeeming the reset apparently pushes your next weekly reset date forward, rather than adding usage on top of your existing cycle. So it's not a top-up; it's borrowing from next week at a premium.
The video's framing here is pretty sharp: "The OpenAI team has definitely got us all addicted to the rate limit refreshes that they have essentially enabled over the past couple of weeks. And now that they have stopped it, they made it so that we have to rely on these refreshes and that way we would need to actually pay for that refresh."
Whether you read that as savvy product management or manufactured dependency is largely a matter of disposition toward the company. The pricing model itself will stress-test how much developer workflow lock-in Codex has actually achieved.
Separately, a significant performance upgrade to the Codex application itself is reportedly coming—focused on handling very long, complex conversations more efficiently. This is an infrastructure improvement, not a model capability jump, but one that should meaningfully improve the experience for developers running extended agentic sessions.
Gemini 3.7 Flash: Substance Behind the Flash Branding
Google's Gemini 3.7 Flash launched this week—three weeks after Gemini 3.6 Flash, which is either impressively fast iteration or a sign that the 3.6 Flash wasn't quite ready. The WorldofAI video suggests the latter: a quick follow-up partly meant to rehabilitate after the disappointing Gemini 3.5 Pro.
What the benchmark picture actually shows: improvements in coding performance, web development rankings, and document reasoning tasks, with the model launching at half the price per token of its predecessor. On the Arena web development leaderboard, it reportedly climbed from 19th to 8th place. It also outputs around 340 tokens per second, making it notably fast for inference-heavy agentic workflows—though like most Flash-tier models, it shows weaknesses in hallucination rates and certain generation tasks.
For developers building against Google's API, cheaper and faster with genuine capability improvements is a straightforward win. Whether it actually recaptures ground from Anthropic or OpenAI at the frontier is a different question.
DeepSeek V4 Pro: Solid, Not Stunning
DeepSeek's V4 Pro hit general availability this week with flexible reasoning modes—low, high, and max—designed for different workflow demands, plus native OpenAI API compatibility. The WorldofAI presenter tested it and landed on a verdict of "decent but underwhelming given the wait." Notably, some users are reporting that DeepSeek V4 Flash, the model released before Pro, actually outperforms Pro on certain front-end and visual coding tasks. Pro is more capable overall, particularly for agentic workloads—but "pro" and "newer" don't automatically mean better for every use case.
The bigger DeepSeek story, though, is pricing. The company is moving to peak and off-peak API pricing, with V4 Pro running at $1.32 per million input tokens and $3.96 per million output tokens at peak hours, with a 50% discount off-peak.
Those numbers are still cheaper than frontier Western labs. But DeepSeek's entire value proposition was radical cost efficiency matched to reasonable quality. As the presenter put it: "Deep Seek's ridiculously low pricing was always one of the biggest advantages cuz the model quality... wasn't always the best. Most openweight models obviously outperformed the DeepSeek models, but the pricing is what got people using this model."
Price increases from DeepSeek read as a market signal more than a DeepSeek-specific story. Serving increasingly capable models at scale costs money. Lots of it. The race to the bottom on inference pricing was never going to run forever.
The Hardware Floor Is Rising
GPU prices are following the same trajectory. Reports have circulated of the Nvidia RTX Pro 6000 climbing sharply at retail outlets in a short window, though the exact current price will vary. The directional trend is consistent across the hardware market: memory constraints, surging demand from inference workloads, and the insatiable appetite of model training are pushing costs up simultaneously.
NVIDIA's own Nemotron 3.5 Lightning—a 30-billion parameter mixture-of-experts model with only 3 billion active parameters—represents one answer to this pressure: pack more capability into smaller active footprints, making local and edge deployment viable for agents that need to run continuously without massive compute overhead.
The economic picture across every layer of the stack—API pricing, GPU hardware, subscription tiers—is converging on the same conclusion. AI is getting more expensive to deliver at the pace capability is expanding. The introductory pricing era is closing.
How users, developers, and enterprises recalibrate around that reality is the story that will matter most in the months ahead—more than any single benchmark number, however impressive.
Marcus Chen-Ramirez covers AI, software development, and the intersection of technology and society for Buzzrag.
AI Moves Fast. We Keep You Current.
Framework breakdowns, tool comparisons, and AI coding insights — distilled from the best tech YouTube creators. Free, weekly.
More Like This
Alibaba's Qwen 3.6 Max Tests Better Than Opus 4.5—At Half the Price
Alibaba's Qwen 3.6 Max Preview outperforms Claude Opus 4.5 in coding and agent workflows at $1.30 per million tokens. Here's what the tests actually show.
Google's Gemma 4 Turns Claude Code Into a Free Local Tool
Google's new Gemma 4 models let developers run Claude Code locally for free. Here's what works, what doesn't, and who this actually serves.
Claude Opus 4.8: Impressive Demos, Marginal Gains
Anthropic's Claude Opus 4.8 lands with better honesty, effort control, and stunning demos—but is the token cost worth marginal gains over Opus 4.7?
Anthropic's Claude Opus 4.6: The New AI Coding Benchmark
Anthropic's Claude Opus 4.6 brings a 1 million token context window and agentic capabilities. What does this mean for developers and knowledge workers?
Google Gemini 3.7 Flash: Coding Power at Low Cost
Google's Gemini 3.7 Flash arrives with serious coding benchmarks, a 1M-token context window, and pricing designed to scale. Here's what it actually means.
Gemini 3.7 Flash and the Agent Economics Race
Google's Gemini 3.7 Flash arrives three weeks after 3.6 with sharp gains in coding and agents—and a pricing strategy designed to buy market share fast.
AI Agents Break Zero Trust at the Last Mile
AI agents reason brilliantly but authenticate badly. Grant Miller explains why agentic systems shatter zero trust at the legacy integration point—and what fixes it.
Claude's /goal Command Can Manage Your AI Workspace
Mark Kashef demos /goal for Claude Code beyond code tasks—using it to clean, sharpen, and auto-maintain your agentic OS while you sleep.
RAG·vector embedding
2026-08-17This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.