void-model: VOID by Netflix: The Video Editor That Removes Objects and Their Physics
How a two-pass diffusion system, a vision-language reasoner, and counterfactual training data turn video inpainting into causal scene rewriting.
- VOID treats video editing as a causality problem, so it expands the edit beyond the object itself to the effects that object was holding in place.
- Its core trick is a two-pass pipeline: reason about affected regions first, then refine the result with warped noise so the video stays temporally coherent.
- The model’s quadmask conditioning makes physical interactions explicit, which is why it feels like scene repair rather than ordinary inpainting.
- Netflix’s real move is workflow-level: it pushes cleanup from manual retouching toward controlled, physically plausible reconstruction.
Most video erasers cheat. They hide the object, patch the hole, and hope you do not notice the broken shadow, the floating prop, or the motion that no longer makes sense. VOID starts from a sharper idea: if you remove the cause, you also have to repair the consequences.
That is what makes Netflix’s open-source model interesting. It is not framed as a generic video generator, but as a system for physically plausible inpainting, where the edit region grows to include the things the removed object affected. The result is closer to causal scene rewriting than to traditional object removal.
Why removing the object is the easy part
Classic video inpainting is good at continuity and bad at physics. It can make a person vanish, but it often leaves behind evidence that the scene has been tampered with: a shadow that no longer matches, a hand still gripping air, water that should have splashed, or a chair that should have moved.
VOID’s premise is that these are not separate bugs. They are symptoms of the same mistake: editing only the pixels the user selected, instead of the scene state the object was maintaining.
The causal pipeline: mask, reason, then generate
The repository is organized around that causal idea. User input starts with a selected object, then segmentation tracks the primary target, then a vision-language model expands the region to include affected areas, and only then does the diffusion model try to synthesize the edited video.
That middle stage is the project’s real differentiator. The VLM is not there for captions or vibes. It is there to reason about side effects, so the model can treat a shadow, splash, collision, or displaced prop as part of the edit, not as background noise.
What makes VOID different from ordinary inpainting
| System | What it removes | How it handles interactions | Temporal stability | Positioning |
|---|---|---|---|---|
| VOID | The object and its affected regions | Explicitly expands the edit with vision-language reasoning and quadmask conditioning | Uses warped-noise refinement in a second pass | Open research model for causal scene repair |
| Runway | The visible object and surrounding region | Strong visual cleanup, but not centered on causal repair | Product-level smoothing, depending on workflow | Commercial editing tool |
| ProPainter | The masked object | Primarily fills the hole left behind | Designed for inpainting quality | Research baseline |
| DiffuEraser | The object region | Focuses on removal quality rather than causal effects | Research-oriented stabilization | Research baseline |
| MiniMax-Remover | The object region | Competitive object removal, less explicit about secondary effects | Product or research dependent | Hybrid tool |
The technical novelty is easier to see in the mask design. VOID does not rely on one binary mask. The model card describes a quadmask, which separates the primary object, overlap, affected regions, and background. That is a subtle change with a big consequence: the model knows what it is supposed to erase, what it should repair, and what it should preserve.
In other words, the system is not asked to guess everything from a single hole. It is given a richer contract about the scene.
The two-pass trick that steadies the video
The two-pass structure is the part engineers will remember. Pass one does the heavy lift of reconstruction. Pass two focuses on temporal consistency by reusing warped noise, instead of starting from fresh random noise on every frame.
That matters because random frame-by-frame sampling is one of the fastest ways to get video that looks fine in isolation and unstable in motion. Warped noise keeps the latent uncertainty aligned with optical flow, which helps suppress texture crawling and flicker.
# Simplified mental model of VOID's refinement loop
pass1 = diffusion_inpaint(masked_video, quadmask)
flow = estimate_optical_flow(pass1)
warped_noise = warp_latent_noise(seed_noise, flow)
pass2 = diffusion_refine(pass1, warped_noise)
output = pass2
This is a nice example of a research project making one clean engineering decision instead of ten flashy ones. The model does not just regenerate. It stabilizes its own regeneration.
How Netflix taught the model physics
The training story is the other half of the insight. VOID does not learn scene repair from generic edits alone. It trains on counterfactual examples generated with simulation tools such as Kubric and HUMOTO, which means the model is shown alternate worlds where the object never existed and the scene had to evolve differently.
That is the right kind of supervision for this problem. If you want a model to understand that a cup should fall when the hand disappears, you do not just show it a hole to fill. You show it what happens when the cause is removed.
The repo’s data-generation layer makes this explicit. Blender-based scene creation and physics-aware rendering give the model examples of the scene state before and after intervention, so the network can learn interactions rather than memorize masks.
Where VOID fits in the market
VOID is not trying to win the same race as every video editing product. Runway, ProPainter, DiffuEraser, and MiniMax-Remover all operate in the broad space of object removal, but VOID is aiming at a narrower and more ambitious target: repair the scene after causal intervention.
That difference matters in practice. A tool that removes a boom mic is useful. A tool that removes the boom mic, restores the occluded background, and fixes the shadow logic is a different class of instrument.
What this means for VFX workflows
The practical implication is workflow, not novelty. If the system can remove the object and automatically repair its side effects, then the editor’s job shifts from manual clean-up to review, correction, and taste.
That is a meaningful change for post-production. It does not eliminate judgment. It reduces the amount of labor spent undoing obvious physical inconsistencies after a shot has already been edited.
VOID is best understood as a proof that video editing can become causally aware. Once a model can reason about what a removed object was doing to the scene, object removal stops being a masking problem and becomes a scene reconstruction problem.