pdoom-video: The Music Video That Renders Like Code

A deterministic Three.js engine turns song time into pixels, using audio analysis, adaptive sampling, and a browser-to-FFmpeg export pipeline to make a music video that behaves more like software than a file.

10 to 12 min read • View on GitHub • More from mexicat

A wide editorial scene shows a drafting table where a score, a code editor, and a film strip are stitched together by small mechanical arms. At the center, a clockwork timeline wheel sends synchronized pulses into a Three.js scene frame and a render box, explaining that the video is computed from song time rather than assembled as a static file.
The project collapses editing, animation, and rendering into one deterministic timeline.
Key Takeaways

A music video that is really a function

The first surprise is not the visuals. It is the premise: the video is not a finished artifact so much as an evaluation of time. Give the engine a timestamp and it can reconstruct the same frame, the same blur, the same lyric state, and the same scene logic every time.

Every frame is a deterministic function of song time, so the live preview in the browser and the offline 1080p60 (or 4K60) export are identical.

mexicat (Giacomo Magnanini), Project Creator · mexicat/pdoom-video GitHub README

That matters because it changes the mental model. You are not rendering a clip that later gets edited. You are running a program where playback, seeking, and export all share the same source of truth.

Why this is not standard generative video

DimensionDiffusion video toolspdoom-video
Output modelSample pixels from noiseCompute pixels from code and time
DeterminismUsually unstable across runsExact for a given timestamp
Sync precisionGood enough for motion, weak for beat-lockFrame-level alignment to audio and lyrics
EditabilityPrompt and rerollChange code, timeline, or audio features
Rendering loopGenerate a clip, then inspect itRe-evaluate the scene at any moment
Best use casePlausible moving imageryMusic videos, typography, synchronized motion
Failure modeTemporal drift and mushy continuityNeeds careful engineering, but stays inspectable

This is not a dunk on diffusion systems. They solve a different problem. Their goal is convincing motion. `pdoom-video` is chasing repeatability, exact sync, and direct control over every beat.

The engine treats time as the source of truth

A single timestamp fans out into scene state, audio features, motion blur, and the final frame.

The core trick is simple to say and hard to implement well: the engine treats absolute song time as input. That means the renderer is not walking forward one frame at a time like a game loop. It is asking, "What should exist at 01:42.300?" and then rebuilding that state from the timeline, scene logic, and preroll.

// Conceptual shape of the engine
const t = songTimeSeconds;
const scene = timeline.sceneAt(t);
const state = scene.preroll ? scene.preroll(t) : scene.stateAt(t);
const audio = audioData.sampleAt(t);
const frame = engine.render(t, state, audio);

That design is what makes seeking sane. If the viewer jumps into the middle of a scene, the engine can replay the missing setup internally before it draws the visible frame. Without that, stateful motion would snap, pop, or drift.

The hidden trick: adaptive sampling for motion blur

A close-up engraving-style illustration shows layered sub-frame slices fanning out beneath one movie frame. Some slices are evenly spaced, while others split recursively as a small error gauge shrinks, explaining that motion blur is earned through convergence rather than smeared in post.
Motion blur is built by sampling until the frame converges.

This is the section where the project stops looking like a clever demo. The blur is not a post effect. It is computed by rendering subframes, measuring the difference between them, and sampling again until the error falls below tolerance.

The important detail is the recursive spacing of those subframes. A simple even spread can hide convergence problems and create temporal bias. The project instead keeps the sample set balanced as it grows, so the engine can compare a cheaper approximation against a denser one without shifting the result.

Why the error target matters

The `errRT` target is basically a quality gate. If the frame contains a whip pan, particle noise, or dense motion, the error rises and the renderer spends more effort. If the frame is calm, it stops early. That is how the system gets filmic blur without wasting time on static shots.

The timeline is the edit decision list

In a traditional edit system, the project file decides when a cut happens. Here, `timeline.ts` plays that role. Functions like `cut()` and `after()` anchor scene changes to beats and lyric events, so the story is stored as code instead of a mouse-driven timeline.

That choice also keeps the project modular. Scenes can be imported lazily, which matters when the renderer needs to cover a long song with a lot of heavy visual material. The engine only pulls in what it needs for the current time window.

Audio is not background. It is control data

The audio pipeline is doing real editorial work. Stem separation, envelopes, onset detection, and lyric alignment turn the song into a structured data set. Once that exists, typography can hit on words, drums can pulse motion, and scene accents can follow the song instead of merely reacting to it.

SignalWhat it gives the rendererWhy it matters
RMS and frequency bandsEnergy over timeDrives broad visual intensity
Drum hitsShort-lived pulsesCreates impact without hard binary jumps
Whisper and CTC alignmentWord-level lyric timingLocks typography to spoken or sung text
Stem separationVocal, drum, bass layersLets the visuals respond to specific instruments

The nice part is that none of this is decorative metadata. The analysis output becomes production input, which is why the typography can feel glued to the track instead of loosely inspired by it.

How the video leaves the browser

Export is where the project turns from creative coding into a production system. The browser renders RGBA buffers, streams them over WebSockets, and Bun hands them off to FFmpeg. Backpressure matters here because 4K renders can overwhelm memory fast if the pipeline is too eager.

That makes the final upload feel less like a terminal output and more like a mastered render from a studio pipeline. The browser is only one execution surface. The underlying machine is what matters.

What makes the authorship story unusual

A WSJ-style hedcut portrait of Giacomo Magnanini, generated from the verified GitHub avatar. The portrait identifies the creator behind the project and grounds the story in a real author rather than an abstract AI brand.

The authorship story is interesting because it is not "AI made art." It is a human creator using Claude Code as a co-designer for the style bible, scene transitions, and engine scaffolding. That is a different claim, and a much more precise one.

Claude Opus 5.5 does not generate pixels. It writes code that paints pixels. The output is a program, not a video file.

Trendshift Analysis, Tech Curator · Claude's Plan: A video by Opus 5.5

Why pdoom-video matters beyond one song

The broader lesson is that creative output can be deterministic without becoming dull. When the system is built well, code can produce media that is inspectable, reusable, and exact, while still feeling alive on screen.

That puts `pdoom-video` in a different category from both diffusion video and hand-authored motion graphics. It is a proof that agent-authored creative engines can be studio-grade systems, not just clever one-offs.

ApproachStrengthTrade-off
Diffusion videoFast, suggestive motionLess control over exact timing
Traditional motion graphicsPrecise, familiar workflowMore manual labor
pdoom-videoDeterministic, beat-locked, inspectableRequires careful engineering and data prep

For a single song, that may sound niche. It is not. It is a signal that the next generation of media tools may be built less like prompts and more like programs.