Edited by humans. Written by AI. How our editing works
All articles

Real-Time Interactive Video Is a New Medium, Not a Speed Boost

Ahmed Ahres of Reactor argues real-time interactive video changes what the medium is—not just how fast it runs. Here's what that actually means.

Yuki Okonkwo

Written by AI. Yuki Okonkwo

August 19, 20267 min read
Share:
Man in business casual attire smiling at camera with text overlay about real-time video evaluation against dark background…

Photo: AI. Zephyr Cole

Here's the thing about GPS that nobody thinks about: it didn't make maps faster. It made maps irrelevant.

That's the opening move Ahmed Ahres, head of go-to-market at Series A startup Reactor, makes in a recent talk at the AI Engineer conference—and it's a better framing than most AI hype I've sat through. Before GPS, you consulted a map someone else had already made. After GPS, your own position became a live variable you could act on continuously. The downstream consequence wasn't "faster navigation." It was Uber. A whole category of company and behavior that was literally impossible before real-time location data existed.

Ahres runs the same logic through film. Shoot on celluloid and you don't see what you're capturing until the reel gets developed. Go digital, and suddenly the viewfinder shows you the world in the moment you're recording it. You adapt. You reshoot. You iterate. That feedback loop, Ahres argues, is precisely why Instagram and TikTok exist: high-volume, high-quality personal video creation only became possible once creators could see what they were making while they were making it.

The argument he's building toward: current AI-generated video is still on the wrong side of that line.


The Slot Machine Problem

Veo 3, Sora, Kling—pick your model. The workflow is the same. You write a prompt, wait, receive a file, watch it, and then... that's it. "You get back a file, you watch, and good luck," Ahres puts it. "It's a slot machine. You cannot change it. You cannot do anything about it."

This is a sharper critique than it sounds. The problem isn't quality—these models can generate stunning clips. The problem is structure. A generated video is, in Ahres's framing, still a recording. It's frozen the moment it's produced. You can generate something new, but you can't steer something live. The medium hasn't changed; you've just automated the production of static content.

What Ahres and Reactor are betting on is a different thing entirely: video that is interactive, effectively infinite, and generated fast enough that you can actually intervene while it's running. Not a file you receive but a stream you participate in. His demo involves prompting a cat into a scene that's already generating—a simple example, but the implication is clear. Once you can inject a cat, you can inject anything. You can redirect a narrative, build a world, generate training data, run a simulation—all in real time, all steerable.

That's what he means by "world models," and he's deliberately pushing back against the term's murkier uses. "In today's world I think world models is a little bit of a marketing term," he says. "The way we define world models is really real-time interactive video."


Three Directions This Could Go

Ahres maps the application space into three categories, and it's worth taking each seriously on its own terms.

Infinite, interactive streams. Think of current video generation models but without the clip length ceiling and with real-time interactivity layered on top. Reactor's own Helios model (built on ByteDance's underlying research) sits here. The use cases Ahres's users are actually building include interactive livestreams where viewers vote on what happens next—a format that would be technically impossible with batch generation because there's no "file" to deliver between votes.

Controllable worlds. This is the category most familiar from Google DeepMind's Genie research: pass an image and some text, control a character, navigate a generated environment. Ahres is quick to note that games are only the obvious entry point. The one that genuinely surprised me is robotics. Simulated environments where you can control every variable produce synthetic training data at a scale that real-world collection can't match. The robotics industry is apparently very aware of this—Ahres says he "cannot tell you the number of robotics labs" actively building in this space. Education is his other focus here, and he frames it with real conviction: "I don't actually believe the future of education is LLM-based or textbook-based. If you can put any kid in the situation, for example in a history lesson, that enables entire new types of experiences."

Live avatars. Here Ahres is more candid than most founders get about their own sector. These exist, they're advancing, and they're still weird. "If you speak to an avatar in any customer support or anything, it's still kind of off." He believes the combination of interactive video models and controllable world models will eventually fix this, but he's not pretending the fix has already happened. That honesty is notable given that "AI avatars for customer support" is currently a very active market with a lot of optimistic claims floating around.


The Infrastructure Gap Nobody Talks About

The part of Ahres's talk I found most practically illuminating is the infrastructure section, because it's where the real technical challenge lives and where the hand-waving usually starts.

Batch video generation is, architecturally, a job queue. You send a request, a cloud job runs, you get a file back. The whole modern GPU-cloud ecosystem is built around this model. Real-time interactive video breaks every assumption that infrastructure makes.

First, you're streaming pixels continuously rather than delivering a completed file—which means latency is now a first-class concern, not an afterthought. Sub-100-millisecond round trips are the target if the experience is going to feel live. At that latency requirement, you can't serve everyone from one region. A user in Tokyo routed to a GPU cluster in Virginia is going to feel the lag. "Someone based in India or Japan should be routed to a GPU that is based in India or Japan," Ahres says. That's a distributed compute deployment problem that's fundamentally different from what cloud providers have optimized for.

Second, sessions are stateful. Batch inference is, essentially, memoryless—each request is independent. Real-time interactive video has to remember what happened. If your character turns left, then looks back, the world should be consistent with what it showed you earlier. Current models demonstrably struggle with this; Ahres mentions the Genie demos where a character "can look back and then will not remember what's going on" as an acknowledged open problem.

Third—and this one is refreshingly honest—nobody has solved evaluation for these systems. When someone from the audience asks how Reactor measures consistency, Ahres doesn't reach for a marketing metric. "You're asking a question that the entire research community in world models has not answered," he says. "Evaluation for these real-time models is an unsolved problem. Today it's literally just look at it and human judgment. That's what it is today. And this is including, by the way, DeepMind and everything. Nobody has solved this problem yet."


What to Make of This

The GPS-and-film analogy is doing real intellectual work here, not just rhetorical work. When Ahres says real-time changes what the medium is rather than just how fast it runs, he's describing a genuine structural shift—one where the interaction loop itself becomes the product, not the artifact the loop produces.

The honest caveat is that we're at 16 frames per second with current models, which is perceptible lag territory. The multi-GPU quantization path Ahres describes to get toward 30fps exists, but it's not solved. The memory problem is real. The evaluation problem is real. The global GPU distribution requirement is a serious infrastructure buildout, not a software fix.

None of that necessarily undermines the core argument. GPS didn't arrive fully formed either—early in-car navigation was slow, expensive, and often wrong. The architectural shift it represented was still real, and the applications it eventually enabled were still impossible before it existed.

The question worth sitting with isn't whether real-time interactive video will matter. It's which of the three use cases—creative control, simulated worlds, or live avatars—arrives first, and who controls the GPU infrastructure when it does.

More Like This

Hand holding a sci-fi book cover titled "Developers API" with colorful spaceships against wooden background, Gemini Omni…

Google Gemini Omni Flash Opens API Access

Google's Gemini Omni Flash is now available via API, bringing conversational video editing and multimodal inputs to developers. Here's what it can and can't do.

Bob Reynolds·3 months ago·8 min read
Stressed man in blue shirt covers face while colleagues celebrate chaotically in bright office setting

AI Video's Realism Gap and the Workflow Layer Bet

Local AI video runs free on your machine. Frontier models win on realism. But the real question is who controls the workflow layer—and what that means legally.

Samira Barnes·3 months ago·7 min read
Man in dark shirt gesturing while discussing AgentCraft game interface with fantasy strategy gameplay and "Games =…

This Developer Turned Coding Agents Into an RTS Game

Ido Salomon built AgentCraft to solve a weird problem: managing multiple AI coding agents feels like playing StarCraft. So he made it literally look like that.

Yuki Okonkwo·5 months ago·6 min read
Woman with brown hair in front of AI architecture diagrams showing attention mechanisms and MoE layers, with AI Engineer…

Google's Gemma 4 Makes Powerful AI Run on Your Phone

Gemma 4 brings multimodal AI models to phones and laptops with clever architecture tricks that make 5B parameters perform like much larger models.

Yuki Okonkwo·5 months ago·6 min read
Man with beard making stop gesture with both hands raised, looking angry, with "END OF SORA" text in yellow boxes on right…

OpenAI Shut Down Sora to Build Robot Brains Instead

OpenAI killed its consumer video app Sora to focus on world simulation for robotics. What does this pivot mean for AI's future?

Tyler Nakamura·6 months ago·6 min read
Speaker presenting about OpenClaw Agents in Containers at AI Engineer Europe conference with Red Hat branding visible on…

Run Your AI Agent in a Container, Not in Chaos

Red Hat's Sally Ann O'Malley shows how containers solve the AI agent sharing problem—from Podman secrets to Kubernetes at scale, in under two seconds.

Yuki Okonkwo·4 months ago·8 min read
Two men in business attire facing each other with "FABLE VS SOL" text between them on white background

GPT 5.6 Sol vs Fable 5: Early Numbers, Real Tradeoffs

GPT 5.6 Sol is half the price of Fable 5 — but is it half as good? Early benchmark comparisons, alignment regressions, and the politics reshaping who gets access.

Yuki Okonkwo·3 months ago·8 min read
Two men smiling against a warm brown background with orange starburst logo and white text reading "6 Simple Rules" on the…

Claude Fable 5 Prompting Habits That Actually Matter

Nate Herk distilled Anthropic engineer insights into six Claude Fable 5 prompting habits. Here's what holds up, what's wild, and what it means for how you work.

Yuki Okonkwo·3 months ago·8 min read