Stereo Any Video: Temporally Consistent Stereo Matching
Junpeng Jing, Weixun Luo, Ye Mao, Krystian Mikolajczyk
摘要
This paper introduces Stereo Any Video, a powerful framework for video stereo matching. It can estimate spatially accurate and temporally consistent disparities without relying on auxiliary information such as camera poses or optical flow. The strong capability is driven by rich priors from monocular video depth models, which are integrated with convolutional features to produce stable representations. To further enhance performance, key architectural innovations are introduced: all-to-all-pairs correlation, which constructs smooth and robust matching cost volumes, and temporal convex upsampling, which improves temporal coherence. These components collectively enhance robustness, accuracy, and temporal consistency, establishing a new standard in video stereo matching. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple datasets both qualitatively and quantitatively in zero-shot settings, as well as strong generalization to real-world indoor and outdoor scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- BANet: Bilateral Aggregation Network for Mobile Stereo MatchingGangwei Xu, Jiaxin Liu, Xianqi Wang, Junda Cheng 等ICCV 2025 · 被引用 7 次
- Lite Any Stereo: Efficient Zero-Shot Stereo MatchingJunpeng Jing, Weixun Luo, Ye Mao, Krystian MikolajczykCVPR 2026 · 被引用 4 次
- DepthFocus: Controllable Depth Estimation for See-Through Scenesjunhong min, Jimin Kim, Minwook Kim, Cheol-Hui Min 等CVPR 2026 · 被引用 4 次
- StereoWorld: Geometry-Aware Monocular-to-Stereo Video GenerationKe Xing, Longfei Li, Yuyang Yin, Hanwen Liang 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
相关 Paper
- Semantic Stereo Matching With Pyramid Cost VolumesZhenyao Wu, Xinyi Wu, Xiaoping Zhang, Song Wang 等ICCV 2019 · 被引用 125 次
- Geometry-Aware Stereo Matching via Monocular Disparity Distribution Prior and Gradient EnhancementJunze Zhang, Luoxi Jing, Yuanyuan Wang, Xueqi Li 等AAAI 2026
- DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video GenerationJian Shi, Qian Wang, Zhenyu Li, Wenqing Cui 等SIGGRAPH 2026
- Depth Any Video with Scalable Synthetic DataHonghui Yang, Di Huang, Wei Yin, Chunhua Shen 等ICLR 2025
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
