M2SVid fixes the hard part of stereo video: the holes
Google Research combines geometry, depth, and a modified Stable Video Diffusion model to turn monocular video into a right-eye view that holds together.
- M2SVid treats stereo conversion as a geometry-first repair job, not a free-form hallucination.
- Its key move is to give Stable Video Diffusion the left video, warped right video, and disocclusion masks, then let full attention focus on the missing pixels.
- The repository is laid out like a real pipeline, with preprocessing, rectification, warping, and refinement separated into distinct stages.
- The project argues that the best 3D video system is the one that constrains generative models with physics instead of asking them to invent depth from scratch.
Monocular-to-stereo conversion sounds simple until a foreground object moves. Then the new right-eye view reveals background that the original camera never saw, and those missing patches become disocclusions, the hard part of the job. M2SVid, a Google Research release accepted to 3DV 2026, is built around that failure mode.
The repo is explicit about its status. It is a research release, not an officially supported Google product, and the code reads like one, with separate preprocessing scripts, config files, and inference entry points instead of a polished end-user app. That is a strength here, because the whole system is easier to inspect stage by stage.
We extend the Stable Video Diffusion (SVD) model to utilize the input left video, the warped right video, and the disocclusion masks as conditioning input to generate a high-quality right camera view. In order to effectively exploit information from neighboring frames for inpainting, we modify the attention layers in SVD to compute full attention for discoccluded pixels.
The pipeline is the point
The pipeline starts with depth prediction, then moves to warping, then ends with inpainting and refinement. That order matters, because geometry is excellent at moving pixels where they belong, but terrible at inventing pixels that were never observed. M2SVid keeps the reliable part of the job in hard geometry and reserves the generative model for the mess geometry leaves behind.
What the model actually sees
in_channels: 13
attn_inpainting_strategy: spatial_full_attention
Those two lines tell the story. Thirteen channels mean the model is not looking at raw RGB alone. It is conditioned on a richer bundle of inputs, including the left view, the warped right view, and masks that mark the missing regions it needs to repair.
The more interesting twist is the attention strategy. Instead of treating every pixel the same, M2SVid gives disoccluded regions broader access to context, so the model can borrow clues from neighboring frames and distant parts of the scene when local evidence runs out. That is a small architectural change with a big effect on temporal consistency.
How the repository is organized
- `data_preprocess/` holds the heavy geometry work, including rectification, feature matching, and dataset preparation.
- `configs/` defines model settings, training behavior, and inference variants, including the full-attention version.
- `warping.py` and `inpaint_and_refine.py` are the main inference entry points for the stereo conversion pipeline.
- `third_party/` integrates external dependencies such as DepthCrafter and Hi3D assets.
- The codebase is overwhelmingly Python, with shell scripts doing the orchestration around it.
That shape tells you who the repo is for. It is meant for researchers and technical builders who want to reproduce a paper, swap components, or inspect where the quality comes from. The result is less like a demo and more like a carefully staged experiment.
Why the design feels disciplined
| Approach | What it trusts most | Main weakness |
|---|---|---|
| Pure warping | Geometry and depth | Leaves black holes and tearing where new background appears |
| Pure generative stereo | Model imagination | Can drift away from the original scene and shimmer over time |
| M2SVid | Geometry first, then targeted inpainting | Depends on good depth and rectification, but keeps structure grounded |
That is the real thesis of the repo. It does not ask a diffusion model to be a stereo system from scratch. It asks geometry to do the honest part, then asks the model to repair only the parts that geometry cannot recover.
According to the repository README, the method is reported as 2.6x more often the preferred result in a user study and 6x faster than earlier approaches. Read that as a research claim, not a product promise, but it explains why the pipeline is arranged so carefully.