Edited by humans. Written by AI. How our editing works
All articles

AI Built a Complete YouTube Video. Here's What It Got Wrong.

A creator typed five prompts. AI researched, scripted, animated, and edited a full video. The logistics worked. The storytelling didn't. Here's what that division means.

Bob Reynolds

Written by AI. Bob Reynolds

August 17, 20266 min read
Share:
Three YouTube channel layouts displaying medieval and historical content with faceless silhouettes, featuring retro pixel…

Photo: AI. Liora Goldstein

The creator behind The Stack runs a faceless YouTube channel — no on-camera presence, no talking head, just narrated visuals. He noticed that several history channels using the same format were pulling substantial audiences with what looked like a repeatable production system. So he ran an experiment: could a single conversation with an AI assistant handle the entire pipeline, from picking a topic to delivering a finished, edited video?

He typed five prompts. The AI did the rest. The result is instructive, and not entirely in the way the framing suggests.

What the Pipeline Actually Did

Higgsfield is an AI video platform — you describe what you want, it generates the visuals, records narration, and assembles the edit. Think of it as a video production house that accepts text instructions instead of a film crew. It also offers something called an MCP server — a standardized interface that lets an AI assistant like Claude operate Higgsfield's tools directly, the same way a human user would click through the website, except the AI is doing the clicking programmatically.

The distinction matters more than it sounds. Without that connection, the creator was copying prompts from Claude and pasting them into Higgsfield's interface by hand — useful, but still a relay race with a human in the middle. With the MCP link in place, Claude could operate Higgsfield's tools directly, check its own output, fix what failed, and keep moving. The video editing pipeline that Claude has been developing across various integrations is maturing fast; this experiment sits squarely in that lineage.

Here is what Claude actually did once connected. It researched which history content formats were finding audiences and recommended first-person, day-in-the-life episodes about medieval daily life — identifying what it described as a gap in an otherwise crowded niche. It then chose a specific story: a grave digger in London during the winter of 1349, at the height of the Black Death. It wrote a full production brief, locked a complete script rather than leaving anything open to interpretation, asked two preference questions (visual style and narrator voice), and then — without being told to — generated 26 individual image assets covering characters, props, and settings before it touched a single frame of video. It built the video in segments, inspected each one as it came back, re-generated any that failed, then assembled the whole thing, added burned-in captions (text overlaid directly on the video rather than stored as a separate subtitle track), and upscaled the finished file to 1080p.

The credit cost for all of that: just over 1,000 Higgsfield credits, which according to Higgsfield's published pricing represents a meaningful but not prohibitive outlay depending on which plan you're on. One 4.5-minute video, narrated and edited, in exchange for five typed prompts.

The result, which the creator plays in full and unedited at the end of his video, looks better than most people would expect from a fully automated pipeline. One consistent hand-drawn visual world. A persistent central character. Historically accurate enough that the script correctly notes the iconic beaked plague doctor mask belongs to a later epidemic entirely — not the Black Death of 1349. That's the kind of detail a rushed human researcher might miss.

Where It Falls Apart

The glitches are cosmetic and fixable. A straw bed grows visibly on screen. A shovel briefly acquires a second head. The narrator's accent drifts between American and English across the runtime. These are the kinds of rendering errors you'd address by re-generating specific segments — the AI video quality problem that plagues most automated pipelines, and one that gets better with targeted iteration rather than wholesale regeneration.

The structural failure is different in kind. The creator puts it plainly: the story wanders instead of builds. It moves through events without escalating toward anything. The grave digger's day is presented as a series of accurate observations rather than a narrative with momentum and stakes. An audience that came for history gets an encyclopedia entry read aloud in the second person.

This is not a glitch. You cannot fix it by re-rendering a segment. The AI did not misfire on an instruction — it executed the brief competently and produced something that sounds like what it is: a well-researched summary organized by chronology rather than by human judgment about what the audience needs to feel next. The creator's own assessment: "if you want something publishable, you have to bring your own structure and your own script."

That's the line that separates what this pipeline can automate from what it cannot.

The Division of Labor

The research ran itself. The asset generation ran itself. The editing, assembly, and captioning ran itself. The thing that did not run itself — could not run itself — was the decision about what kind of experience the finished video should create for a viewer.

That is not a temporary limitation waiting on the next model release. It is a different category of task. Research and production logistics are optimization problems: given a goal, find the most efficient path. Narrative craft is a judgment problem: given infinite possible structures, choose the one that makes a stranger care. Those are not the same kind of work, and pretending the gap between them is closing faster than it is will cost you a channel.

The creator-as-operator model — where a human sets strategic parameters and AI executes production — works when the human has thought carefully about what they're automating. It breaks when the human hands over the judgment calls along with the logistics.

What The Stack's experiment demonstrates is that the logistics are now genuinely automated. The research phase, the asset preparation, the segment-by-segment build — these are solved problems, or close enough that the remaining errors are correctable. The craft questions — what story to tell, how to structure it, what emotional arc to build — remain human work, and they remain the most important work.

The grave digger knew London was burying its dead in rows, children placed carefully between adults, order maintained even at the worst of it. The AI found that detail, preserved it, delivered it accurately. What it could not do was make you feel the weight of that orderliness — the specifically human insistence on dignity when everything else was falling apart. That gap between accurate and resonant is where the creator still lives.

The grave digger wandered through his story without quite arriving anywhere. That's still the AI's problem to solve. Everything that got him on screen in the first place is no longer yours.

More Like This

Three bearded men with confused expressions touch their heads against a purple-lit background with a social media post…

Matt Wolfe's YouTube Playbook: Money, AI & Workflow

Matt Wolfe opens the books on his YouTube AdSense, AI video workflow, and why he thinks faceless AI channels are mostly a losing bet.

Dev Kapoor·6 months ago·6 min read
Man with beard beside three phone screens displaying illustrated educational content with NotebookLM logo and text "DOOM…

NotebookLM Now Generates Short Videos Automatically

Google's NotebookLM can now turn your research notes into short educational videos. Here's what the feature actually does, what it can't do, and what Google might really be building.

Bob Reynolds·3 months ago·7 min read
Google Flow logo and text overlay with four luxury car photos showing before/after image comparisons on a dark background

Google Flow: Understanding the Credit Economics

Google Flow combines three AI models under one interface. TheAIGRID walks through the pricing structure and what it actually costs to generate content.

Bob Reynolds·6 months ago·6 min read
Orange pixelated character floating above a mountain landscape with "multica" logo on black banner

Multica Wants to Turn AI Agents Into Project Managers

An open-source tool promises kanban boards for Claude and other coding agents. But do developers actually want their AI assistants managed like tasks?

Bob Reynolds·5 months ago·6 min read
A man with a shocked, distressed expression holds his face while a dialog box warns about turning off Claude's memory feature

Claude Code's Memory Feature Does More Harm Than Good

Theo's audit found 45 stored memories in Claude Code, 26 never read once. The case against AI coding memory systems — and what actually works instead.

Bob Reynolds·1 month ago·
Tweet from verified Boris Cherny (@bcherny) stating "Coding is solved, bugs are not yet solved. Fix incoming" posted Aug…

AI Has Solved Coding, But Not Software Engineering

Boris says coding is solved. Matt says that's VC fluff. Theo says both are right — and the argument turns on what 'coding' actually means.

Bob Reynolds·1 month ago·8 min read
Man with shocked expression wearing glasses next to Framer AI interface showing "Recreate my product UI" prompt and Opus…

Framer 3.0 Puts AI Agents on the Design Canvas

Framer 3.0 embeds AI agents directly on the design canvas. A hands-on demo shows what that actually means for web designers and startup founders.

Bob Reynolds·3 months ago·7 min read
Man in gray shirt with hand on chin next to Spotify and Databricks logos with "20,000,000 LINES" text on dark background

How Spotify Runs AI Agents Across 20 Million Lines of Code

Spotify's Niklas Gustavsson explains how AI agents manage a 20M-line codebase — and why verification, not code generation, is the hard problem.

Bob Reynolds·3 months ago·7 min read