AI Built a Complete YouTube Video. Here's What It Got Wrong.
A creator typed five prompts. AI researched, scripted, animated, and edited a full video. The logistics worked. The storytelling didn't. Here's what that division means.
Written by AI. Bob Reynolds

Photo: AI. Liora Goldstein
The creator behind The Stack runs a faceless YouTube channel — no on-camera presence, no talking head, just narrated visuals. He noticed that several history channels using the same format were pulling substantial audiences with what looked like a repeatable production system. So he ran an experiment: could a single conversation with an AI assistant handle the entire pipeline, from picking a topic to delivering a finished, edited video?
He typed five prompts. The AI did the rest. The result is instructive, and not entirely in the way the framing suggests.
What the Pipeline Actually Did
Higgsfield is an AI video platform — you describe what you want, it generates the visuals, records narration, and assembles the edit. Think of it as a video production house that accepts text instructions instead of a film crew. It also offers something called an MCP server — a standardized interface that lets an AI assistant like Claude operate Higgsfield's tools directly, the same way a human user would click through the website, except the AI is doing the clicking programmatically.
The distinction matters more than it sounds. Without that connection, the creator was copying prompts from Claude and pasting them into Higgsfield's interface by hand — useful, but still a relay race with a human in the middle. With the MCP link in place, Claude could operate Higgsfield's tools directly, check its own output, fix what failed, and keep moving. The video editing pipeline that Claude has been developing across various integrations is maturing fast; this experiment sits squarely in that lineage.
Here is what Claude actually did once connected. It researched which history content formats were finding audiences and recommended first-person, day-in-the-life episodes about medieval daily life — identifying what it described as a gap in an otherwise crowded niche. It then chose a specific story: a grave digger in London during the winter of 1349, at the height of the Black Death. It wrote a full production brief, locked a complete script rather than leaving anything open to interpretation, asked two preference questions (visual style and narrator voice), and then — without being told to — generated 26 individual image assets covering characters, props, and settings before it touched a single frame of video. It built the video in segments, inspected each one as it came back, re-generated any that failed, then assembled the whole thing, added burned-in captions (text overlaid directly on the video rather than stored as a separate subtitle track), and upscaled the finished file to 1080p.
The credit cost for all of that: just over 1,000 Higgsfield credits, which according to Higgsfield's published pricing represents a meaningful but not prohibitive outlay depending on which plan you're on. One 4.5-minute video, narrated and edited, in exchange for five typed prompts.
The result, which the creator plays in full and unedited at the end of his video, looks better than most people would expect from a fully automated pipeline. One consistent hand-drawn visual world. A persistent central character. Historically accurate enough that the script correctly notes the iconic beaked plague doctor mask belongs to a later epidemic entirely — not the Black Death of 1349. That's the kind of detail a rushed human researcher might miss.
Where It Falls Apart
The glitches are cosmetic and fixable. A straw bed grows visibly on screen. A shovel briefly acquires a second head. The narrator's accent drifts between American and English across the runtime. These are the kinds of rendering errors you'd address by re-generating specific segments — the AI video quality problem that plagues most automated pipelines, and one that gets better with targeted iteration rather than wholesale regeneration.
The structural failure is different in kind. The creator puts it plainly: the story wanders instead of builds. It moves through events without escalating toward anything. The grave digger's day is presented as a series of accurate observations rather than a narrative with momentum and stakes. An audience that came for history gets an encyclopedia entry read aloud in the second person.
This is not a glitch. You cannot fix it by re-rendering a segment. The AI did not misfire on an instruction — it executed the brief competently and produced something that sounds like what it is: a well-researched summary organized by chronology rather than by human judgment about what the audience needs to feel next. The creator's own assessment: "if you want something publishable, you have to bring your own structure and your own script."
That's the line that separates what this pipeline can automate from what it cannot.
The Division of Labor
The research ran itself. The asset generation ran itself. The editing, assembly, and captioning ran itself. The thing that did not run itself — could not run itself — was the decision about what kind of experience the finished video should create for a viewer.
That is not a temporary limitation waiting on the next model release. It is a different category of task. Research and production logistics are optimization problems: given a goal, find the most efficient path. Narrative craft is a judgment problem: given infinite possible structures, choose the one that makes a stranger care. Those are not the same kind of work, and pretending the gap between them is closing faster than it is will cost you a channel.
The creator-as-operator model — where a human sets strategic parameters and AI executes production — works when the human has thought carefully about what they're automating. It breaks when the human hands over the judgment calls along with the logistics.
What The Stack's experiment demonstrates is that the logistics are now genuinely automated. The research phase, the asset preparation, the segment-by-segment build — these are solved problems, or close enough that the remaining errors are correctable. The craft questions — what story to tell, how to structure it, what emotional arc to build — remain human work, and they remain the most important work.
The grave digger knew London was burying its dead in rows, children placed carefully between adults, order maintained even at the worst of it. The AI found that detail, preserved it, delivered it accurately. What it could not do was make you feel the weight of that orderliness — the specifically human insistence on dignity when everything else was falling apart. That gap between accurate and resonant is where the creator still lives.
The grave digger wandered through his story without quite arriving anywhere. That's still the AI's problem to solve. Everything that got him on screen in the first place is no longer yours.
Bob Reynolds is Senior Technology Correspondent at Buzzrag.
AI Moves Fast. We Keep You Current.
Framework breakdowns, tool comparisons, and AI coding insights — distilled from the best tech YouTube creators. Free, weekly.
More Like This
Multica Wants to Turn AI Agents Into Project Managers
An open-source tool promises kanban boards for Claude and other coding agents. But do developers actually want their AI assistants managed like tasks?
Google Flow: Understanding the Credit Economics
Google Flow combines three AI models under one interface. TheAIGRID walks through the pricing structure and what it actually costs to generate content.
Matt Wolfe's YouTube Playbook: Money, AI & Workflow
Matt Wolfe opens the books on his YouTube AdSense, AI video workflow, and why he thinks faceless AI channels are mostly a losing bet.
NotebookLM Now Generates Short Videos Automatically
Google's NotebookLM can now turn your research notes into short educational videos. Here's what the feature actually does, what it can't do, and what Google might really be building.
Claude Opus 5 Is Verbose by Default. Here Is the Fix
Claude Opus 5 defaults to jargon-heavy, verbose output. Here's how to configure Claude Code's output style and custom skills to fix it.
Claude Code for Marketing: What the Course Gets Right
Nick Saraev's six-hour Claude Code marketing course has real ideas worth understanding—and one framework that every marketer should think hard about.
Build a Claude Code + Obsidian Command Center
Chase AI shows how to turn Obsidian into a Claude Code command center. Here's what the setup actually does—and what you should know before you build it.
TikTok Is Now a Serious App Marketing Tool
Julia Pintar of Playkit says TikTok is now the most effective free channel for app launches. Here's the playbook—and the questions it leaves open.
RAG·vector embedding
2026-08-17This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.