SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix
Peng Dai, Feitong Tan, Qiangeng Xu, David Futschik, Ruofei Du, Sean Fanello, Xiaojuan Qi, Yinda Zhang
Abstract
Video generation models have demonstrated great capabilities of producing impressive monocular videos, however, the generation of 3D stereoscopic video remains under-explored. We propose a pose-free and training-free approach for generating 3D stereoscopic videos using an off-the-shelf monocular video generation model. Our method warps a generated monocular video into camera views on stereoscopic baseline using estimated video depth, and employs a novel frame matrix video inpainting framework. The framework leverages the video generation model to inpaint frames observed from different timestamps and views. This effective approach generates consistent and semantically coherent stereoscopic videos without scene optimization or model fine-tuning. Moreover, we develop a disocclusion boundary re-injection scheme that further improves the quality of video inpainting by alleviating the negative effects propagated from disoccluded areas in the latent space. We validate the efficacy of our proposed method by conducting experiments on videos from various generative models, including Sora [4 ], Lumiere [2], WALT [8 ], and Zeroscope [ 42]. The experiments demonstrate that our method has a significant improvement over previous methods. The code will be released at https://daipengwa.github.io/SVG_ProjectPage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bf757a8-e26f-460b-b97e-6c47d5505be9Cited by top-tier papers8
- StereoWorld: Geometry-Aware Monocular-to-Stereo Video GenerationKe Xing, Longfei Li, Yuyang Yin, Hanwen Liang et al.CVPR 2026 · 3 citations
- Elastic3D: Controllable Stereo Video Conversion with Guided Latent DecodingNando Metzger, Prune Truong, Goutam Bhat, Konrad Schindler et al.CVPR 2026 · 3 citations
- Guardians of the Hair: Rescuing Soft Boundaries in Depth, Stereo, and Novel ViewsXiang Zhang, Yang Zhang, Lukas Mehl, Markus Gross et al.CVPR 2026 · 2 citations
- Stereo World Model: Camera-Guided Stereo Video GenerationYang-Tian Sun, Zehuan Huang, Yifan Niu, Lin Ma et al.CVPR 2026 · 1 citation
- DreamStereo: Towards Real-Time Stereo Inpainting for HD VideosYuan Huang, Sijie Zhao, Jing Cheng, Hao Xu et al.CVPR 2026 · 1 citation
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video GenerationJian Shi, Qian Wang, Zhenyu Li, Wenqing Cui et al.SIGGRAPH 2026
- ZeroStereo: Zero-Shot Stereo Matching from Single ImagesXianqi Wang, Hao Yang, Gangwei Xu, Junda Cheng et al.ICCV 2025 · 1 citation
- ST360D: Spatial-to-Temporal Consistency for Training-free 360 Monocular Depth EstimationZidong Cao, Jinjing Zhu, Hao Ai, Lutao Jiang et al.NeurIPS 2025 · 1 citation
- Unsupervised object-centric video generation and decomposition in 3DPaul Henderson, Christoph H. LampertNeurIPS 2020 · 41 citations
- Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video GenerationJingyi Lu, Kai HanICML 2026
