BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment
Tongfan Guan, Jiaxin Guo, Chen Wang, Yun-Hui Liu
Abstract
Monocular and stereo depth estimation offer complementary strengths: monocular methods capture rich contextual priors but lack geometric precision, while stereo approaches leverage epipolar geometry yet struggle with ambiguities such as reflective or textureless surfaces. Despite post-hoc synergies, these paradigms remain largely disjoint in practice. We introduce a unified framework that bridges both through iterative bidirectional alignment of their latent representations. At its core, a novel cross-attentive alignment mechanism dynamically synchronizes monocular contextual cues with stereo hypothesis representations during stereo reasoning. This mutual alignment resolves stereo ambiguities (e.g., specular surfaces) by injecting monocular structure priors while refining monocular depth with stereo geometry within a single network. Extensive experiments demonstrate state-of-the-art results: it reduces zeroshot generalization error by on Middlebury and ETH3D, while addressing longstanding failures on transparent and reflective surfaces. By harmonizing multi-view geometry with monocular context, our approach enables robust 3D perception that transcends modality-specific limitations. Codes available at https://github.com/aeolusguan/BridgeDepth.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- TruckDrive: Long-Range Autonomous Highway Driving DatasetFilippo Ghilotti, Edoardo Palladin, Samuel Brucker, Adam Sigal et al.CVPR 2026 · 5 citations
- PromptStereo: Zero-Shot Stereo Matching via Structure and Motion PromptsXianqi Wang, Hao Yang, Hangtian Wang, JunDa Cheng et al.CVPR 2026 · 5 citations
- DepthFocus: Controllable Depth Estimation for See-Through Scenesjunhong min, Jimin Kim, Minwook Kim, Cheol-Hui Min et al.CVPR 2026 · 4 citations
- DispViT: Direct Stereo Disparity Regression with a Single-Stream Vision TransformerTongfan Guan, Jiaxin Guo, Tianyu Huang, Jinhu Dong et al.ICLR 2026
- Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric StereoNinghui Xu, Fabio Tosi, Lihui Wang, Jiawei Han et al.CVPR 2026
Builds on32
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
Related papers
- MonSter: Marry Monodepth to Stereo Unleashes PowerJunda Cheng, Longliang Liu, Gangwei Xu, Xianqi Wang et al.CVPR 2025
- Depth Anything with Any PriorZehan Wang, Siyu Chen, Lihe Yang, Jialei Wang et al.ICLR 2026 · 47 citations
- Zero-Shot Event-Intensity Asymmetric Stereo via Visual Prompting from Image DomainHanyue Lou, Jinxiu (Sherry) Liang, Minggui Teng, Bin Fan et al.NeurIPS 2024 · 13 citations
- MonoMVSNet: Monocular Priors Guided Multi-View Stereo NetworkJianfei Jiang, Qiankun Liu, Haochen Yu, Hongyuan Liu et al.ICCV 2025 · 3 citations
- MNSRNet: Multimodal Transformer Network for 3D Surface Super-ResolutionWuyuan Xie, Tengcong Huang, Miaohui WangCVPR 2022 · 9 citations
