Problem
Producing a long-form documentary video is a dozen separate jobs: research, scripting, storyboarding, visuals, narration, music, captions, editing, quality control, thumbnails, and publishing metadata. Doing that by hand for a 25-minute video takes days, and naive AI pipelines produce the wrong runtime, unverified claims, or silent audio gaps.
The project asks a harder question: can one command turn a topic into a finished, watchable documentary package with citations you can actually check?
Approach and architecture
The system is organized as a staged Python pipeline, each stage with a clear contract:
- Research gathers sources for the topic and verifies every citation URL actually resolves, failing loudly on placeholders.
- Script generates a multi-beat documentary script with a cold-open hook, then enforces the target duration (extending underweight beats at ~150 wpm) so a 25-minute target really is 25 minutes.
- Visuals routes each scene: stock footage where real footage exists, AI-generated clips for impossible shots, cinematic stills, and procedural fallbacks. Real footage is preferred over synthetic.
- Voiceover renders full narration via local TTS.
- Music lays a topic-matched ambient bed under the narration, ducked and loudness-normalized.
- Captions burn in karaoke-style word-highlight captions, with SRT/VTT exports for YouTube.
- Assembly joins scenes with FFmpeg crossfades and mixes the final audio to -16 LUFS.
- QC gate runs automated checks that fail the run on dark frames, digital silence, low audio, wrong resolution, or runtime drift.
- Packaging picks a thumbnail with title text and generates YouTube metadata: titles, hook-first description, chapters, tags, and sources.
Key decisions
Verify citations mechanically, not hopefully
The research stage checks that every citation URL returns HTTP 200. A documentary that cites dead links is worse than one with fewer sources, so verification is a hard gate rather than a suggestion.
Enforce runtime as a constraint, not an estimate
Script beats are extended until the word count matches the target duration at a realistic speaking pace. Runtime drift is then re-checked by the QC gate after assembly, closing the loop.
Prefer real footage, fall back deliberately
The visual router tries real stock footage first and only synthesizes what cannot be filmed. That ordering keeps the output grounded instead of drifting into generic AI-video sameness.
Challenges
The main engineering challenges were:
- keeping a nine-stage pipeline debuggable when any stage can fail a run
- matching narration pacing to visuals so scenes never feel rushed or padded
- building a QC gate strict enough to catch real defects without false-failing good runs
- generating YouTube packaging (chapters, hooks, tags) that reads as written for humans, not SEO filler
Outcome and learning
The pipeline produces complete long-form documentary packages from a single command, with demo outputs published to YouTube. The lasting takeaway is that long AI pipelines need the same discipline as production systems: explicit stage contracts, mechanical verification gates, and loud failures instead of silent degradation.