M2SVid fixes the hard part of stereo video: the holes

Google Research combines geometry, depth, and a modified Stable Video Diffusion model to turn monocular video into a right-eye view that holds together.

11 min read • View on GitHub • More from google-research

A flat movie frame splits into a stereo pair, with the right-eye side full of jagged gaps where foreground objects moved away. Fine black ink stitches repair the missing areas, showing the handoff from geometric warping to neural refinement.
M2SVid starts with a broken stereo view on purpose, then uses a model to repair the damage geometry cannot avoid.
Key Takeaways

Monocular-to-stereo conversion sounds simple until a foreground object moves. Then the new right-eye view reveals background that the original camera never saw, and those missing patches become disocclusions, the hard part of the job. M2SVid, a Google Research release accepted to 3DV 2026, is built around that failure mode.

The repo is explicit about its status. It is a research release, not an officially supported Google product, and the code reads like one, with separate preprocessing scripts, config files, and inference entry points instead of a polished end-user app. That is a strength here, because the whole system is easier to inspect stage by stage.

We extend the Stable Video Diffusion (SVD) model to utilize the input left video, the warped right video, and the disocclusion masks as conditioning input to generate a high-quality right camera view. In order to effectively exploit information from neighboring frames for inpainting, we modify the attention layers in SVD to compute full attention for discoccluded pixels.

Nina Shvetsova, Goutam Bhat, Prune Truong, Hilde Kuehne, Federico Tombari, Authors · google-research/m2svid

The pipeline is the point

M2SVid is not one model doing everything. It is a staged system that uses geometry for alignment and diffusion for cleanup.

The pipeline starts with depth prediction, then moves to warping, then ends with inpainting and refinement. That order matters, because geometry is excellent at moving pixels where they belong, but terrible at inventing pixels that were never observed. M2SVid keeps the reliable part of the job in hard geometry and reserves the generative model for the mess geometry leaves behind.

What the model actually sees

in_channels: 13
attn_inpainting_strategy: spatial_full_attention

Those two lines tell the story. Thirteen channels mean the model is not looking at raw RGB alone. It is conditioned on a richer bundle of inputs, including the left view, the warped right view, and masks that mark the missing regions it needs to repair.

The more interesting twist is the attention strategy. Instead of treating every pixel the same, M2SVid gives disoccluded regions broader access to context, so the model can borrow clues from neighboring frames and distant parts of the scene when local evidence runs out. That is a small architectural change with a big effect on temporal consistency.

How the repository is organized

That shape tells you who the repo is for. It is meant for researchers and technical builders who want to reproduce a paper, swap components, or inspect where the quality comes from. The result is less like a demo and more like a carefully staged experiment.

Why the design feels disciplined

ApproachWhat it trusts mostMain weakness
Pure warpingGeometry and depthLeaves black holes and tearing where new background appears
Pure generative stereoModel imaginationCan drift away from the original scene and shimmer over time
M2SVidGeometry first, then targeted inpaintingDepends on good depth and rectification, but keeps structure grounded

That is the real thesis of the repo. It does not ask a diffusion model to be a stereo system from scratch. It asks geometry to do the honest part, then asks the model to repair only the parts that geometry cannot recover.

According to the repository README, the method is reported as 2.6x more often the preferred result in a user study and 6x faster than earlier approaches. Read that as a research claim, not a product promise, but it explains why the pipeline is arranged so carefully.