Masked Representation Learning for Domain Generalized Stereo Matching
Zhibo Rao, Bangshu Xiong, Mingyi He, Yuchao Dai, Renjie He, Zhelun Shen, Xing Li
Abstract
Recently, many deep stereo matching methods have begun to focus on cross-domain performance, achieving impressive achievements. However, these methods did not deal with the significant volatility of generalization performance among different training epochs. Inspired by masked representation learning and multi-task learning, this paper designs a simple and effective masked representation for domain generalized stereo matching. First, we feed the masked left and complete right images as input into the models. Then, we add a lightweight and simple decoder following the feature extraction module to recover the original left image. Finally, we train the models with two tasks (stereo matching and image reconstruction) as a pseudo-multi-task learning framework, promoting models to learn structure information and to improve generalization performance. We implement our method on two well-known architectures (CFNet and LacGwcNet) to demonstrate its effectiveness. Experimental results on multi-datasets show that: (1) our method can be easily plugged into the current various stereo matching models to improve generalization performance; (2) our method can reduce the significant volatility of generalization performance among different training epochs; (3) we find that the current methods prefer to choose the best results among different training epochs as generalization performance, but it is impossible to select the best performance by ground truth in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f0ce556-1486-4b54-8b8e-2dc4d1086f6bCited by top-tier papers15
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingBowen Wen, Shaurya Dewan, Stan BirchfieldCVPR 2026 · 36 citations
- Robust Synthetic-to-Real Transfer for Stereo MatchingJiawei Zhang, Jiahe Li, Lei Huang, Xiaohan Yu et al.CVPR 2024 · 12 citations
- Learning Robust Stereo Matching in the Wild with Selective Mixture-of-ExpertsYun Wang, Longguang Wang, Chenghao Zhang, Yongjian Zhang et al.ICCV 2025 · 6 citations
- BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent AlignmentTongfan Guan, Jiaxin Guo, Chen Wang, Yun-Hui LiuICCV 2025 · 6 citations
- Hierarchical Object-Aware Dual-Level Contrastive Learning for Domain Generalized Stereo MatchingYikun Miao, Meiqing Wu, Siew Kei Lam, Changsheng Li et al.NeurIPS 2024 · 5 citations
Builds on15
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 690 citations
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding et al.ICCV 2021 · 380 citations
- Practical Stereo Matching via Cascaded Recurrent Network with Adaptive CorrelationJiankun Li, Peisen Wang, Pengfei Xiong, Tao Cai et al.CVPR 2022 · 294 citations
- Attention Concatenation Volume for Accurate and Efficient Stereo MatchingGangwei Xu, Junda Cheng, Peng Guo, Xin YangCVPR 2022 · 265 citations
Related papers
- CFNet: Cascade and Fused Cost Volume for Robust Stereo MatchingZhelun Shen, Yuchao Dai, Zhibo RaoCVPR 2021
- GraftNet: Towards Domain Generalized Stereo Matching with a Broad-Spectrum and Task-Oriented FeatureBiyang Liu, Huimin Yu, Guodong QiCVPR 2022 · 52 citations
- Domain Generalized Stereo Matching via Hierarchical Visual TransformationTianyu Chang, Xun Yang, Tianzhu Zhang, Meng WangCVPR 2023
- Curvature-Guided Dynamic Scale Networks for Multi-View StereoKhang Truong Giang, Soohwan Song, Sungho JoICLR 2022 · 43 citations
- Eglcr: Edge Structure Guidance and Scale Adaptive Attention for Iterative Stereo MatchingZhien Dai, Zhaohui Tang, Hu Zhang, Can Tian et al.ACM MM 2024 · 1 citation
