pdoom-video: The Music Video That Renders Like Code
A deterministic Three.js engine turns song time into pixels, using audio analysis, adaptive sampling, and a browser-to-FFmpeg export pipeline to make a music video that behaves more like software than a file.
- pdoom-video treats a music video as a function of song time, which makes the result reproducible, seekable, and inspectable instead of fragile.
- Its standout trick is adaptive motion blur, where the engine keeps sampling subframes until the frame converges instead of smearing motion after the fact.
- The timeline behaves like an edit decision list, while audio analysis supplies the timing data for lyric sync and reactive animation.
- The export path turns the browser into a renderer that streams raw frames to Bun and FFmpeg, so the finished video is really a software pipeline.
A music video that is really a function
The first surprise is not the visuals. It is the premise: the video is not a finished artifact so much as an evaluation of time. Give the engine a timestamp and it can reconstruct the same frame, the same blur, the same lyric state, and the same scene logic every time.
Every frame is a deterministic function of song time, so the live preview in the browser and the offline 1080p60 (or 4K60) export are identical.
That matters because it changes the mental model. You are not rendering a clip that later gets edited. You are running a program where playback, seeking, and export all share the same source of truth.
Why this is not standard generative video
| Dimension | Diffusion video tools | pdoom-video |
|---|---|---|
| Output model | Sample pixels from noise | Compute pixels from code and time |
| Determinism | Usually unstable across runs | Exact for a given timestamp |
| Sync precision | Good enough for motion, weak for beat-lock | Frame-level alignment to audio and lyrics |
| Editability | Prompt and reroll | Change code, timeline, or audio features |
| Rendering loop | Generate a clip, then inspect it | Re-evaluate the scene at any moment |
| Best use case | Plausible moving imagery | Music videos, typography, synchronized motion |
| Failure mode | Temporal drift and mushy continuity | Needs careful engineering, but stays inspectable |
This is not a dunk on diffusion systems. They solve a different problem. Their goal is convincing motion. `pdoom-video` is chasing repeatability, exact sync, and direct control over every beat.
The engine treats time as the source of truth
The core trick is simple to say and hard to implement well: the engine treats absolute song time as input. That means the renderer is not walking forward one frame at a time like a game loop. It is asking, "What should exist at 01:42.300?" and then rebuilding that state from the timeline, scene logic, and preroll.
// Conceptual shape of the engine
const t = songTimeSeconds;
const scene = timeline.sceneAt(t);
const state = scene.preroll ? scene.preroll(t) : scene.stateAt(t);
const audio = audioData.sampleAt(t);
const frame = engine.render(t, state, audio);
That design is what makes seeking sane. If the viewer jumps into the middle of a scene, the engine can replay the missing setup internally before it draws the visible frame. Without that, stateful motion would snap, pop, or drift.
The hidden trick: adaptive sampling for motion blur
This is the section where the project stops looking like a clever demo. The blur is not a post effect. It is computed by rendering subframes, measuring the difference between them, and sampling again until the error falls below tolerance.
The important detail is the recursive spacing of those subframes. A simple even spread can hide convergence problems and create temporal bias. The project instead keeps the sample set balanced as it grows, so the engine can compare a cheaper approximation against a denser one without shifting the result.
Why the error target matters
The `errRT` target is basically a quality gate. If the frame contains a whip pan, particle noise, or dense motion, the error rises and the renderer spends more effort. If the frame is calm, it stops early. That is how the system gets filmic blur without wasting time on static shots.
The timeline is the edit decision list
In a traditional edit system, the project file decides when a cut happens. Here, `timeline.ts` plays that role. Functions like `cut()` and `after()` anchor scene changes to beats and lyric events, so the story is stored as code instead of a mouse-driven timeline.
That choice also keeps the project modular. Scenes can be imported lazily, which matters when the renderer needs to cover a long song with a lot of heavy visual material. The engine only pulls in what it needs for the current time window.
Audio is not background. It is control data
The audio pipeline is doing real editorial work. Stem separation, envelopes, onset detection, and lyric alignment turn the song into a structured data set. Once that exists, typography can hit on words, drums can pulse motion, and scene accents can follow the song instead of merely reacting to it.
| Signal | What it gives the renderer | Why it matters |
|---|---|---|
| RMS and frequency bands | Energy over time | Drives broad visual intensity |
| Drum hits | Short-lived pulses | Creates impact without hard binary jumps |
| Whisper and CTC alignment | Word-level lyric timing | Locks typography to spoken or sung text |
| Stem separation | Vocal, drum, bass layers | Lets the visuals respond to specific instruments |
The nice part is that none of this is decorative metadata. The analysis output becomes production input, which is why the typography can feel glued to the track instead of loosely inspired by it.
How the video leaves the browser
Export is where the project turns from creative coding into a production system. The browser renders RGBA buffers, streams them over WebSockets, and Bun hands them off to FFmpeg. Backpressure matters here because 4K renders can overwhelm memory fast if the pipeline is too eager.
That makes the final upload feel less like a terminal output and more like a mastered render from a studio pipeline. The browser is only one execution surface. The underlying machine is what matters.
What makes the authorship story unusual
The authorship story is interesting because it is not "AI made art." It is a human creator using Claude Code as a co-designer for the style bible, scene transitions, and engine scaffolding. That is a different claim, and a much more precise one.
Claude Opus 5.5 does not generate pixels. It writes code that paints pixels. The output is a program, not a video file.
Why pdoom-video matters beyond one song
The broader lesson is that creative output can be deterministic without becoming dull. When the system is built well, code can produce media that is inspectable, reusable, and exact, while still feeling alive on screen.
That puts `pdoom-video` in a different category from both diffusion video and hand-authored motion graphics. It is a proof that agent-authored creative engines can be studio-grade systems, not just clever one-offs.
| Approach | Strength | Trade-off |
|---|---|---|
| Diffusion video | Fast, suggestive motion | Less control over exact timing |
| Traditional motion graphics | Precise, familiar workflow | More manual labor |
| pdoom-video | Deterministic, beat-locked, inspectable | Requires careful engineering and data prep |
For a single song, that may sound niche. It is not. It is a signal that the next generation of media tools may be built less like prompts and more like programs.