void-model: VOID by Netflix: The Video Editor That Removes Objects and Their Physics

How a two-pass diffusion system, a vision-language reasoner, and counterfactual training data turn video inpainting into causal scene rewriting.

8 min read · Netflix/void-model

A stage-like scene shows an object removed from a table, but its effects still hang in the air: a tipping chair, a falling glass, and a displaced shadow. A second, cleaner version of the same scene is assembling beneath it, showing that the edit repairs the scene’s consequences rather than only masking the missing object.
VOID is not just erasing pixels. It is trying to reconstruct the world that should have followed after the object disappeared.
Key Takeaways

Most video erasers cheat. They hide the object, patch the hole, and hope you do not notice the broken shadow, the floating prop, or the motion that no longer makes sense. VOID starts from a sharper idea: if you remove the cause, you also have to repair the consequences.

That is what makes Netflix’s open-source model interesting. It is not framed as a generic video generator, but as a system for physically plausible inpainting, where the edit region grows to include the things the removed object affected. The result is closer to causal scene rewriting than to traditional object removal.

Why removing the object is the easy part

Classic video inpainting is good at continuity and bad at physics. It can make a person vanish, but it often leaves behind evidence that the scene has been tampered with: a shadow that no longer matches, a hand still gripping air, water that should have splashed, or a chair that should have moved.

VOID’s premise is that these are not separate bugs. They are symptoms of the same mistake: editing only the pixels the user selected, instead of the scene state the object was maintaining.

A close-up film strip passes through two gates. The first gate restores the missing object, while the second bends a ribbon of noise along a motion path to steady the frames. The image explains how warped-noise refinement reduces flicker and keeps texture aligned across time.
The second pass is not a cosmetic polish. It is the part that keeps the scene from crawling or flickering after the main reconstruction.

The causal pipeline: mask, reason, then generate

The repository is organized around that causal idea. User input starts with a selected object, then segmentation tracks the primary target, then a vision-language model expands the region to include affected areas, and only then does the diffusion model try to synthesize the edited video.

VOID’s pipeline is a loop, not a delete button. The system first understands what the object touched, then reconstructs the scene as if the object had never been there.

That middle stage is the project’s real differentiator. The VLM is not there for captions or vibes. It is there to reason about side effects, so the model can treat a shadow, splash, collision, or displaced prop as part of the edit, not as background noise.

What makes VOID different from ordinary inpainting

SystemWhat it removesHow it handles interactionsTemporal stabilityPositioning
VOIDThe object and its affected regionsExplicitly expands the edit with vision-language reasoning and quadmask conditioningUses warped-noise refinement in a second passOpen research model for causal scene repair
RunwayThe visible object and surrounding regionStrong visual cleanup, but not centered on causal repairProduct-level smoothing, depending on workflowCommercial editing tool
ProPainterThe masked objectPrimarily fills the hole left behindDesigned for inpainting qualityResearch baseline
DiffuEraserThe object regionFocuses on removal quality rather than causal effectsResearch-oriented stabilizationResearch baseline
MiniMax-RemoverThe object regionCompetitive object removal, less explicit about secondary effectsProduct or research dependentHybrid tool

The technical novelty is easier to see in the mask design. VOID does not rely on one binary mask. The model card describes a quadmask, which separates the primary object, overlap, affected regions, and background. That is a subtle change with a big consequence: the model knows what it is supposed to erase, what it should repair, and what it should preserve.

In other words, the system is not asked to guess everything from a single hole. It is given a richer contract about the scene.

The two-pass trick that steadies the video

The two-pass structure is the part engineers will remember. Pass one does the heavy lift of reconstruction. Pass two focuses on temporal consistency by reusing warped noise, instead of starting from fresh random noise on every frame.

That matters because random frame-by-frame sampling is one of the fastest ways to get video that looks fine in isolation and unstable in motion. Warped noise keeps the latent uncertainty aligned with optical flow, which helps suppress texture crawling and flicker.

# Simplified mental model of VOID's refinement loop
pass1 = diffusion_inpaint(masked_video, quadmask)
flow = estimate_optical_flow(pass1)
warped_noise = warp_latent_noise(seed_noise, flow)
pass2 = diffusion_refine(pass1, warped_noise)
output = pass2

This is a nice example of a research project making one clean engineering decision instead of ten flashy ones. The model does not just regenerate. It stabilizes its own regeneration.

How Netflix taught the model physics

The training story is the other half of the insight. VOID does not learn scene repair from generic edits alone. It trains on counterfactual examples generated with simulation tools such as Kubric and HUMOTO, which means the model is shown alternate worlds where the object never existed and the scene had to evolve differently.

That is the right kind of supervision for this problem. If you want a model to understand that a cup should fall when the hand disappears, you do not just show it a hole to fill. You show it what happens when the cause is removed.

The repo’s data-generation layer makes this explicit. Blender-based scene creation and physics-aware rendering give the model examples of the scene state before and after intervention, so the network can learn interactions rather than memorize masks.

Where VOID fits in the market

VOID is not trying to win the same race as every video editing product. Runway, ProPainter, DiffuEraser, and MiniMax-Remover all operate in the broad space of object removal, but VOID is aiming at a narrower and more ambitious target: repair the scene after causal intervention.

That difference matters in practice. A tool that removes a boom mic is useful. A tool that removes the boom mic, restores the occluded background, and fixes the shadow logic is a different class of instrument.

What this means for VFX workflows

The practical implication is workflow, not novelty. If the system can remove the object and automatically repair its side effects, then the editor’s job shifts from manual clean-up to review, correction, and taste.

That is a meaningful change for post-production. It does not eliminate judgment. It reduces the amount of labor spent undoing obvious physical inconsistencies after a shot has already been edited.

VOID is best understood as a proof that video editing can become causally aware. Once a model can reason about what a removed object was doing to the scene, object removal stops being a masking problem and becomes a scene reconstruction problem.