S2M2: Scalable Stereo Matching Model for Reliable Depth Estimation
Junhong Min, Youngpil Jeon, Jimin Kim, Minyong Choi
摘要
The pursuit of a generalizable stereo matching model, capable of performing well across varying resolutions and disparity ranges without dataset-specific fine-tuning, has revealed a fundamental trade-off. Iterative local search methods achieve high scores on constrained benchmarks, but their core mechanism inherently limits the global consistency required for true generalization. However, global matching architectures, while theoretically more robust, have historically been rendered infeasible by prohibitive computational and memory costs. We resolve this dilemma with S2M2: a global matching architecture that achieves state-of-the-art accuracy and high efficiency without relying on cost volume filtering or deep refinement stacks. Our design integrates a multi-resolution transformer for robust long-range correspondence, trained with a novel loss function that concentrates probability on feasible matches. This approach enables a more robust joint estimation of disparity, occlusion, and confidence. S2M2 establishes a new state of the art on Middlebury v3 and ETH3D benchmarks, significantly outperforming prior methods in most metrics while reconstructing high-quality details with competitive efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DepthFocus: Controllable Depth Estimation for See-Through Scenesjunhong min, Jimin Kim, Minwook Kim, Cheol-Hui Min 等CVPR 2026 · 被引用 4 次
- Generalized Geometry Encoding Volume for Real-time Stereo MatchingJiaxin Liu, Gangwei Xu, Xianqi Wang, Chengliang Zhang 等AAAI 2026
- DispViT: Direct Stereo Disparity Regression with a Single-Stream Vision TransformerTongfan Guan, Jiaxin Guo, Tianyu Huang, Jinhu Dong 等ICLR 2026
它引用的顶会 Paper22
- Geometric Transformer for Fast and Robust Point Cloud RegistrationZheng Qin, Hao Yu, Changjian Wang, Yulan Guo 等CVPR 2022 · 被引用 436 次
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding 等ICCV 2021 · 被引用 380 次
- Practical Stereo Matching via Cascaded Recurrent Network with Adaptive CorrelationJiankun Li, Peisen Wang, Pengfei Xiong, Tao Cai 等CVPR 2022 · 被引用 294 次
- REGTR: End-to-end Point Cloud Correspondences with TransformersZi Jian Yew, Gim Hee LeeCVPR 2022 · 被引用 242 次
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier 等NeurIPS 2022 · 被引用 189 次
相关 Paper
- ELFNet: Evidential Local-global Fusion for Stereo MatchingJieming Lou, Weide Liu, Zhuo Chen, Fayao Liu 等ICCV 2023 · 被引用 38 次
- WT-MVSNet: Window-based Transformers for Multi-view StereoJinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang 等NeurIPS 2022 · 被引用 50 次
- MonSter: Marry Monodepth to Stereo Unleashes PowerJunda Cheng, Longliang Liu, Gangwei Xu, Xianqi Wang 等CVPR 2025
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi 等ICCV 2021 · 被引用 318 次
- Eglcr: Edge Structure Guidance and Scale Adaptive Attention for Iterative Stereo MatchingZhien Dai, Zhaohui Tang, Hu Zhang, Can Tian 等ACM MM 2024 · 被引用 1 次
