SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Input
Zhen Lv, Yangqi Long, Congzhentao Huang, Cao Li, Chengfei Lv, Hao Ren, Dian Zheng
摘要
Stereo video synthesis from monocular input is challenging in spatial computing and virtual reality due to the lack of high-quality stereo video pairs for training and the difficulty of maintaining spatio-temporal consistency between frames. Existing methods primarily address these issues by directly applying novel view synthesis (NVS) techniques to video, while facing limitations such as the inability to effectively represent dynamic scenes and the requirement for extensive training data. In this paper, we introduce a novel self-supervised stereo video synthesis paradigm via a video diffusion model, termed SpatialDreamer, which meets the challenges head-on. Firstly, to address the stereo video data insufficiency, we propose a Depth based Video Generation module DVG, which employs a forward-backward rendering mechanism to generate paired videos with geometric and temporal priors. Leveraging data generated by DVG, we propose RefinerNet along with a self-supervised synthetic framework designed to facilitate efficient and dedicated training. More importantly, we devise a consistency control module, which consists of a metric of stereo deviation strength and a Temporal Interaction Learning module TIL for geometric and temporal consistency ensurance respectively. We evaluated the proposed method against various benchmark methods, with the results showcasing its superior performance. Our project website is at: https://spatialdreamer.github.io .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular VideosKaihua Chen, Tarasha Khurana, Deva RamananNeurIPS 2025 · 被引用 17 次
- StereoWorld: Geometry-Aware Monocular-to-Stereo Video GenerationKe Xing, Longfei Li, Yuyang Yin, Hanwen Liang 等CVPR 2026 · 被引用 3 次
- DreamStereo: Towards Real-Time Stereo Inpainting for HD VideosYuan Huang, Sijie Zhao, Jing Cheng, Hao Xu 等CVPR 2026 · 被引用 1 次
- Panorama Generation From NFoV Image Done RightDian Zheng, Cheng Zhang, Xiao-Ming Wu, Cao Li 等CVPR 2025
- Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric TasksYu Zhou, Dian Zheng, Qijie Mo, Renjie Lu 等CVPR 2025
它引用的顶会 Paper32
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei 等ICCV 2023 · 被引用 1,113 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
相关 Paper
- Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video GenerationJingyi Lu, Kai HanICML 2026
- DynamicStereo: Consistent Dynamic Depth from Stereo VideosNikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova 等CVPR 2023
- V2Depth: Monocular Depth Estimation via Feature-Level Virtual-View Simulation and RefinementZizhang Wu, Zhuozheng Li, Zhi-Gang Fan, Yunzhe Wu 等ACM MM 2023 · 被引用 4 次
- MultiDiff: Consistent Novel View Synthesis from a Single ImageNorman Müller, Katja Schwarz, Barbara Rössle, Lorenzo Porzi 等CVPR 2024 · 被引用 14 次
- SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited ObservationsSongchun Zhang, Huiyao Xu, Sitong Guo, Zhongwei Xie 等ICCV 2025 · 被引用 6 次
