ChiTransformer: Towards Reliable Stereo from Cues
Qing Su, Shihao Ji
摘要
Current stereo matching techniques are challenged by restricted searching space, occluded regions and sheer size. While single image depth estimation is spared from these challenges and can achieve satisfactory results with the extracted monocular cues, the lack of stereoscopic relationship renders the monocular prediction less reliable on its own especially in highly dynamic or cluttered environments. To address these issues in both scenarios, we present an optic-chiasm-inspired self-supervised binocular depth estimation method, wherein vision transformer (ViT) with a gated positional cross-attention (GPCA) layer is designed to enable feature-sensitive pattern retrieval between views, while retaining the extensive context information aggregated through self-attentions. Monocular cues from a single view are thereafter conditionally rectified by a blending layer with the retrieved pattern pairs. This crossover design is biologically analogous to the optic-chasma structure in human visual system and hence the name, Chi-Transformer. Our experiments show that this architecture yields substantial improvements over state-of-the-art self-supervised stereo approaches by 11%, and can be used on both rectilinear and non-rectilinear (e.g., fisheye) images. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> https://github.com/ISL-CV/ChiTransformer.git
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- IINet: Implicit Intra-inter Information Fusion for Real-Time Stereo MatchingXimeng Li, Chen Zhang, Wanjuan Su, Wenbing TaoAAAI 2024 · 被引用 22 次
- DualNet: Robust Self-Supervised Stereo Matching with Pseudo-Label SupervisionYun Wang, Jiahao Zheng, Chenghao Zhang, Zhanjie Zhang 等AAAI 2025 · 被引用 12 次
- Lite Any Stereo: Efficient Zero-Shot Stereo MatchingJunpeng Jing, Weixun Luo, Ye Mao, Krystian MikolajczykCVPR 2026 · 被引用 4 次
- Stereo Any Video: Temporally Consistent Stereo MatchingJunpeng Jing, Weixun Luo, Ye Mao, Krystian MikolajczykICCV 2025 · 被引用 1 次
- Leveraging Global Stereo Consistency for Category-Level Shape and 6D Pose Estimation from Stereo ImagesJunning Qiu, Minglei Lu, Fei Wang, Yu Guo 等CVPR 2025
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
相关 Paper
- Multi-Frame Self-Supervised Depth with TransformersVitor Guizilini, Rares Ambrus, Dian Chen, Sergey Zakharov 等CVPR 2022 · 被引用 95 次
- Deep Digging into the Generalization of Self-Supervised Monocular Depth EstimationJinwoo Bae, Sungho Moon, Sunghoon ImAAAI 2023 · 被引用 127 次
- Two-in-One Depth: Bridging the Gap Between Monocular and Binocular Self-supervised Depth EstimationZhengming Zhou, Qiulei DongICCV 2023 · 被引用 16 次
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding 等ICCV 2021 · 被引用 380 次
- GasMono: Geometry-Aided Self-Supervised Monocular Depth Estimation for Indoor ScenesChaoqiang Zhao, Matteo Poggi, Fabio Tosi, Lei Zhou 等ICCV 2023 · 被引用 27 次
