DynamicStereo: Consistent Dynamic Depth from Stereo Videos
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, Christian Rupprecht
Abstract
Abstract We consider the problem of reconstructing a dynamic scene observed from a stereo camera. Most existing methods for depth from stereo treat different stereo frames independently, leading to temporally inconsistent depth predictions. Temporal consistency is especially important for immersive AR or VR scenarios, where flickering greatly diminishes the user experience. We propose DynamicStereo, a novel transformer-based architecture to estimate disparity for stereo videos. The network learns to pool information from neighboring frames to improve the temporal consistency of its predictions. Our architecture is designed to process stereo videos efficiently through divided attention layers. We also introduce Dynamic Replica, a new benchmark dataset containing synthetic videos of people and animals in scanned environments, which provides complementary training and evaluation data for dynamic stereo closer to real applications than existing datasets. Training with this dataset further improves the quality of predictions of our proposed DynamicStereo as well as prior methods. Finally, it acts as a benchmark for consistent stereo methods. Project page: https://dynamic-stereo.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers70
- STream3R: Scalable Sequential 3D Reconstruction with Causal TransformerYushi Lan, Yihang Luo, Fangzhou Hong, Shangchen Zhou et al.ICLR 2026 · 84 citations
- TAPIP3D: Tracking Any Point in Persistent 3D GeometryBowei Zhang, Lei Ke, Adam W. Harley, Katerina FragkiadakiNeurIPS 2025 · 79 citations
- Efficiently Reconstructing Dynamic Scenes One D4RT at a TimeChuhan Zhang, Guillaume Le Moing, Skanda Koppula, Ignacio Rocco et al.CVPR 2026 · 52 citations
- NeoVerse: Enhancing 4D World Model with in-the-wild Monocular VideosYuxue Yang, Lue Fan, Ziqi Shi, Junran Peng et al.CVPR 2026 · 42 citations
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One SecondChenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu et al.CVPR 2026 · 38 citations
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding et al.ICCV 2021 · 380 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
Related papers
- SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular InputZhen Lv, Yangqi Long, Congzhentao Huang, Cao Li et al.CVPR 2025
- Less is More: Consistent Video Depth Estimation with Masked Frames ModelingYiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao et al.ACM MM 2022 · 23 citations
- DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video GenerationJian Shi, Qian Wang, Zhenyu Li, Wenqing Cui et al.SIGGRAPH 2026
- Deep 3D Mask Volume for View Synthesis of Dynamic ScenesKai-En Lin, Lei Xiao, Feng Liu, Guowei Yang et al.ICCV 2021 · 42 citations
- GemDepth: Geometry-Embedded Features for 3D-Consistent Video DepthYuecheng Liu, Junda Cheng, Longliang Liu, Wenjing Liao et al.ICML 2026 · 1 citation
