IINet: Implicit Intra-inter Information Fusion for Real-Time Stereo Matching
Ximeng Li, Chen Zhang, Wanjuan Su, Wenbing Tao
Abstract
Recently, there has been a growing interest in 3D CNN-based stereo matching methods due to their remarkable accuracy. However, the high complexity of 3D convolution makes it challenging to strike a balance between accuracy and speed. Notably, explicit 3D volumes contain considerable redundancy. In this study, we delve into more compact 2D implicit network to eliminate redundancy and boost real-time performance. However, simply replacing explicit 3D networks with 2D implicit networks causes issues that can lead to performance degradation, including the loss of structural information, the quality decline of inter-image information, as well as the inaccurate regression caused by low-level features. To address these issues, we first integrate intra-image information to fuse with inter-image information, facilitating propagation guided by structural cues. Subsequently, we introduce the Fast Multi-scale Score Volume (FMSV) and Confidence Based Filtering (CBF) to efficiently acquire accurate multi-scale, noise-free inter-image information. Furthermore, combined with the Residual Context-aware Upsampler (RCU), our Intra-Inter Fusing network is meticulously designed to enhance information transmission on both feature-level and disparity-level, thereby enabling accurate and robust regression. Experimental results affirm the superiority of our network in terms of both speed and accuracy compared to all other fast methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb67df5a-1b02-4de0-a254-368611d70c52Cited by top-tier papers5
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingBowen Wen, Shaurya Dewan, Stan BirchfieldCVPR 2026 · 36 citations
- BANet: Bilateral Aggregation Network for Mobile Stereo MatchingGangwei Xu, Jiaxin Liu, Xianqi Wang, Junda Cheng et al.ICCV 2025 · 7 citations
- Pip-Stereo: Progressive Iterations Pruner for Iterative Optimization based Stereo MatchingJintu Zheng, Qizhe Liu, Huangxin Xu, zhuojie ChenCVPR 2026 · 1 citation
- Enhanced Event-Based Dense Stereo via Cross-Sensor Knowledge DistillationHaihao Zhang, Yunjian Zhang, Jianing Li, Lin Zhu et al.ICCV 2025 · 1 citation
- Generalized Geometry Encoding Volume for Real-time Stereo MatchingJiaxin Liu, Gangwei Xu, Xianqi Wang, Chengliang Zhang et al.AAAI 2026
Builds on25
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch et al.ICLR 2022 · 797 citations
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding et al.ICCV 2021 · 380 citations
Related papers
- AANet: Adaptive Aggregation Network for Efficient Stereo MatchingHaofei Xu, Juyong ZhangCVPR 2020
- Semantic Stereo Matching With Pyramid Cost VolumesZhenyao Wu, Xinyi Wu, Xiaoping Zhang, Song Wang et al.ICCV 2019 · 125 citations
- Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo MatchingXiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai et al.CVPR 2020
- CFNet: Cascade and Fused Cost Volume for Robust Stereo MatchingZhelun Shen, Yuchao Dai, Zhibo RaoCVPR 2021
- Bilateral Grid Learning for Stereo Matching NetworksBin Xu, Yuhua Xu, Xiaoli Yang, Wei Jia et al.CVPR 2021
