AI Models and the Cybersecurity Threat Landscape
New frontier AI models are arriving as agencies warn of AI-assisted attacks on critical infrastructure. What the threat actually looks like—and what it doesn't.
Written by AI. Rachel "Rach" Kovacs

Photo: AI. Saskia Aaltonen
There's a version of the AI-and-cybersecurity story that gets told constantly, and it goes like this: scary new model drops, vague warnings about unprecedented capabilities, someone says "it's not a matter of if but when," and readers are left feeling vaguely menaced without knowing what to actually do. I've written about threats long enough to recognize when fear is the product being sold.
So when AI commentator Wes Roth recently put out a video connecting the wave of incoming frontier AI models to an active joint advisory from U.S. agencies about AI-assisted attacks on critical infrastructure, I wanted to look at what's actually being claimed—and what the evidence does and doesn't support.
The Advisory Is Real, and It Says Something Specific
Start with what's verifiable. CISA, the FBI, the NSA, the Department of Energy, and the EPA issued a joint advisory warning that threat actors are using AI to generate exploitation scripts targeting internet-exposed industrial control systems—specifically Siemens S7 programmable logic controllers used in power plants, water treatment facilities, and manufacturing operations.
The agencies' own language is worth sitting with: "This is not a theoretical risk. It's an active threat." And: "This represents an evolution in threat actor capabilities, dramatically reducing the technical expertise and time required to conduct cyber attacks."
That second quote is the one that matters most to me. The barrier to entry for sophisticated attacks has been high precisely because competent offensive security talent is expensive and, as Roth puts it in his video, "very well-paid and usually won't engage in these sort of shenanigans." State-backed actors could always clear that bar. Now the bar is coming down—and that changes the geometry of who can cause serious damage.
What the advisory doesn't say is equally important. It does not name a specific attack. It describes a pattern of persistent AI-assisted reconnaissance: agents systematically crawling internet-exposed systems, probing codebases for vulnerabilities, not necessarily exploiting them yet. The threat, at this moment, is the intelligence-gathering phase of what could become attacks. Roth is honest about this distinction, which I appreciate.
The Ox Alpha Distraction
Roth's video spends time on "Ox Alpha," an anonymously posted model that appeared on AI platforms and generated significant speculation about its origins—SSI (Ilya Sutskever's Safe Superintelligence company), a Chinese lab, or something else entirely. Roth's own assessment is that technical markers point to this being a Chinese model, likely from the GLM family. Benchmarks place it in mid-tier territory, not a frontier breakthrough.
The SSI thread is more interesting. Investor Gavin Baker, speaking on the Invest Like the Best podcast, stated that SSI planned to release a model in August. SSI itself has confirmed nothing—no announcement, no paper, no API access. The company was founded on a premise of silence followed by a sudden leap: no intermediate releases, just a "straight shot to superintelligence," as Roth describes their original framing.
That framing has apparently softened. Sutskever has more recently indicated that gradual releases would be part of "any plausible plan," partly so governments and the public can track what's being deployed. Whether that represents a genuine philosophical shift or just an acknowledgment that total opacity isn't commercially viable is genuinely unclear. Around the same time, SSI announced a partnership with Nvidia that significantly expands their access to compute infrastructure—which suggests they're scaling up research that they believe is worth scaling.
Gemini 4 and the Leadership Shift
The more grounded model-release story is Google's. On July 21st, Google confirmed it had begun what it called its "most ambitious pre-training run yet" for Gemini 4. Sundar Pichai, on Google's Q2 earnings call on July 23rd, described the model as "significantly larger" than its predecessors, with coding and autonomous agents as explicit priorities.
That's a meaningful directional shift. Demis Hassabis, who recently stepped down as DeepMind CEO to focus on research roles, was reportedly skeptical of the recursive self-improvement approach that most other major labs are pursuing—the idea of training models specifically to advance AI research itself. Sergey Brin is reportedly back at Google working on these coding models directly. Whatever internal dynamics drove those changes, the result is that Google is now publicly aligned with the same RSI-focused strategy as Anthropic, OpenAI, and xAI.
Unverified leaked evaluations suggest Gemini 4 may be competitive with current frontier models on coding benchmarks—but Roth flags these as unverified, and they should be treated as such. What isn't speculation: Google has publicly committed to this training run, and prediction markets currently put a high probability on release before the end of the year.
The Part Worth Taking Seriously
Here's where I want to zoom out from the model release calendar, because there's a structural point that gets lost in the hype cycle around benchmark numbers and anonymous model drops.
The cybersecurity threat from AI isn't primarily about one catastrophic attack enabled by one brilliant model. It's about scale and persistence. Roth frames it well: "Never before in the history of the world could we get agents that can work 24 hours a day that can be cloned infinitely to continuously go through and find these exploits."
Human offensive security talent is finite. It operates in business hours, takes sick days, has to sleep. AI agents doing reconnaissance don't. The NSA advisory describes exactly this: persistent, systematic probing—not dramatic infiltration, but patient cataloguing of every exposed surface. The threat isn't one genius hacker. It's an exhaustive search conducted at machine scale across every internet-exposed system, indefinitely.
The defensive implication of this is not "update your passwords." It's that organizations running critical infrastructure need to treat internet-exposed industrial control systems as the attack surface they actually are. Air-gapping what can be air-gapped. Patching what can be patched. Monitoring for reconnaissance patterns, not just active intrusions. These are the recommendations that come out of the joint advisory, and they're not glamorous, but they're real.
What I'm Watching
The cybersecurity industry has a long and profitable history of selling fear. The AI frontier labs have their own incentives to make their models sound both powerful and safe-but-dangerous-in-the-wrong-hands, simultaneously. Both of those distortions are live in this conversation.
But the joint advisory from multiple federal agencies describing active AI-assisted reconnaissance against critical infrastructure is not marketing copy. It's bureaucratic understatement if anything. Federal agencies don't typically lead with "this is not a theoretical risk" unless they want people to stop treating it as one.
The question worth sitting with isn't whether these new models are "superintelligence"—they almost certainly aren't, and that framing obscures more than it reveals. The question is whether the institutions running the infrastructure that keeps the lights on and water flowing are updating their threat models fast enough to match the pace at which the offense side is scaling up.
Given what the agencies are describing, that's not a hypothetical worth deferring.
Rachel "Rach" Kovacs is Buzzrag's cybersecurity and privacy correspondent.
More Like This
Seven Open-Source AI Tools Changing Development in 2026
From prompt testing to guardrail removal, these seven open-source AI tools represent a significant shift in how developers build—and what that means for security.
AI's Speed Problem: Hacks, Lawsuits, and Your Attack Surface
Google's zero-day warning, the OpenAI lawsuit pressure cooker, and why AI's speed makes old security hygiene dangerously obsolete.
AI Labs Call for a Global Pause Mechanism on AI
Top AI leaders signed a letter urging synthetic biology screening, while Anthropic published a stark assessment of recursive self-improvement and why a pause mechanism matters.
Google DeepMind Maps the Road From AGI to ASI
Google DeepMind's new paper treats AGI as a starting point, not a finish line. Here's what it actually argues—and what it leaves unresolved.
AI Benchmark Scores Are Broken. Here's Who's Fixing Them.
AI benchmark scores are less trustworthy than they look. Google DeepMind's Kaggle team is building open infrastructure to fix that—here's what you need to know.
9-Arm Skills: AI Agents Need Brakes, Not More Gas
A tiny GitHub repo called 9-arm-skills argues AI coding agents need behavioral constraints, not more power. The accountability implications go deeper than the code.
RAG·vector embedding
2026-08-24This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.