claude-real-video: The Video Pipeline That Teaches LLMs What to Ignore

A local-first tool that turns raw video into scene-aware keyframes, deduplicated grids, and a manifest an LLM can actually use.

6 to 8 min read • View on GitHub • More from HUANGCHIHHUNGLeo

A long film strip feeds into a mechanical sorter, where many near-identical frames are discarded into a waste pile. The surviving keyframes, transcript pages, and a manifest card are arranged neatly beside a chat window, showing how raw video becomes a compact package for an LLM to read.
`crv` is not trying to watch everything. It is trying to keep only the frames that still carry meaning.
Key Takeaways

The hard part is not watching video. It is deciding what counts

Most video tools start with a simple assumption: more frames means more understanding. `claude-real-video` rejects that idea. It assumes the real task is to compress visual evidence without throwing away the signal.

That matters because raw video is messy for LLMs. Fixed-interval sampling wastes budget on static slides, while naive transcript-only workflows miss the visual moments that actually change meaning. `crv` sits between those failures and asks a sharper question: which frames deserve to survive?

Point it at a URL or a file, and it pulls the frames that actually matter (every scene change, not a fixed quota), throws away the near-duplicates, transcribes the audio, and hands you a clean folder any LLM can read.

WSJ-style hedcut portrait of the project creator, rendered as black ink on a white background. The portrait provides a human anchor for the origin of the tool and helps connect the technical pipeline to its maintainer.

What `crv` actually produces

The output is plain, which is part of the appeal. A run gives you extracted keyframes, a transcript, a manifest, and optional grids. Nothing about that is exotic. The point is that the bundle is already shaped for a model to consume.

The pipeline is not just extraction. It is packaging the parts of video that a multimodal model can use without drowning in redundancy.

OutputWhat it containsWhy it matters
KeyframesScene-change frames plus deduplicated survivorsKeeps visual evidence compact without flattening the story
TranscriptLocal audio transcriptionPreserves spoken context for text-first reasoning
MANIFEST.txtA readable index of the packageGives an LLM a map instead of a pile of files
Grids3x3 contact sheets of chronological framesLets one image carry several moments at once
video_input → ffmpeg scene detection → dedup window → keyframes
          ↘ whisper transcript → MANIFEST.txt → LLM-ready bundle

The trick: scene detection plus a floor

The core move is simple to describe and hard to get right. `crv` does not trust only scene detection, and it does not trust only periodic sampling. It combines both. That means abrupt cuts are caught, but long static stretches still get coverage.

That is a small design choice with large consequences. A lecture with frozen slides should not turn into a desert of useless frames, but a screen recording with sudden UI changes should not glide past the important moment either. The floor keeps the system honest.

A close-up timeline shows an A-B-A pattern, with a speaker scene returning after a slide sequence. A naive adjacent comparison is fooled on the right side of the panel, while a sliding-window dedup lens correctly marks the repeated A as redundant and suppresses it. This explains why the tool compares against more than just the previous frame.
The A-B-A problem is where simple dedup logic breaks. A short memory of earlier frames fixes it.

Why deduplication matters more than it sounds

This is the subtle part of the repo. Adjacent-frame comparison only answers one question: did the image change from the last instant? That misses repeated scenes that return after an interruption. The result is wasted tokens and a model that keeps seeing the same visual idea as if it were new.

Sliding-window deduplication solves that A-B-A loop. It gives the pipeline a short memory, which is enough to suppress repeated interview cutaways, recurring speaker shots, and other returns that look novel only if you forgot the last few seconds.

In the code, that idea is backed by small RGB thumbnails and pixel-difference checks rather than grayscale shortcuts. The point is not academic purity. The point is to avoid collapsing visually distinct content into the same bucket.

# Conceptual sketch
for frame in candidate_frames:
    if scene_changed(frame) or periodic_floor_hit(frame):
        if not matches_any(frame, recent_retained_frames):
            keep(frame)
        else:
            skip(frame)
    recent_retained_frames = recent_retained_frames[-window_size:] + [frame]

Why this is local-first, not cloud-first

The privacy argument is real, but it is not the main story. Local-first here is mostly a product decision. It keeps raw video on your machine, and it lets you decide exactly what reaches an LLM provider.

That also changes the economics. If the pipeline can remove most of the redundancy before upload, then the model pays for meaning instead of entropy. For screen recordings, talks, and demos, that can be the difference between a usable workflow and a noisy one.

ApproachWhere processing happensWhat gets uploadedBest fitWeakness
`crv` local-firstOn your machineOnly selected frames, transcript, and manifestSensitive or long-form video analysisRequires local setup
Native video APICloud providerEntire video fileSimple questions and quick turnaroundsLess control over sampling and privacy
Fixed-interval frame grabberUsually local or cloudMany redundant framesSimple scripts and rough previewsWastes context on static scenes

Claude Code is the second product

The repo is more than a CLI. The `skills/claude-real-video/SKILL.md` file turns the pipeline into a tool an agent can reach for on its own. That matters because it shifts the project from preprocessing into workflow design.

In that mode, `crv` becomes a sensing layer for Claude Code. The model is not trying to reason about video directly. It is asking the right local utility to condense the file into something it can inspect.

What it beats, and what it does not

`crv` is not trying to beat every native video model on every task. If you just need a quick yes or no on a simple clip, a native multimodal model can be convenient. But convenience is not the same as control.

Where `crv` wins is portability, transparency, and token efficiency on messy real-world video. It is especially strong when the material is long, repetitive, private, or worth revisiting later as a local knowledge asset.

ToolSampling strategyPrivacy modelToken efficiencyBest use caseWeakness
`claude-real-video`Scene-aware plus sliding-window dedupLocal-firstHigh on repetitive videoLectures, demos, screen recordingsNeeds preprocessing
Gemini native videoProvider-controlled internal samplingCloud-firstStrong for simple queriesFast Q&A over a clipLess control, less transparency
Fixed frame extractorPeriodic frames onlyDepends on deploymentLowQuick baseline scriptsMisses cuts, wastes context
Broader video VLM frameworkVaries by model and configUsually mixedDepends on setupResearch and experimentationMore complex than a preprocessing layer

The best comparison is not really about who is smarter. It is about where the intelligence lives. `crv` moves intelligence into the preprocessing step, so the model receives a tighter bundle and can spend its context on reasoning instead of cleanup.

Why this project feels durable

That is the lasting idea here. `claude-real-video` sits at a useful intersection: local AI, token efficiency, and workflow design. It does not replace video models. It makes them easier to use well.

The strongest projects in this space are often the unglamorous ones. They do one job upstream, they do it carefully, and they remove friction everywhere else. `crv` looks like a video tool, but its real contribution is a better answer to a deeper question: what should a model not have to see?