DispViT: Direct Stereo Disparity Regression with a Single-Stream Vision Transformer
Tongfan Guan, Jiaxin Guo, Tianyu Huang, Jinhu Dong, Chen Wang, Yun-Hui Liu
Abstract
Deep stereo disparity estimation has long been dominated by a matching-centric paradigm, built on constructing cost volumes and iteratively refining local correspondences. Despite its success, this paradigm exhibits an intrinsic vulnerability: visual ambiguities from occlusion or non-Lambertian surfaces invevitably induce errorneous matches that refinement cannot recover. This paper introduces DispViT, a new architecture that establishes a regression-centric paradigm. Instead of explicit matching, DispViT directly regresses disparity from tokenized binocular representations using a single-stream Vision Transformer. This is enabled by a set of lightweight yet critical designs, such as a probability-based disparity parameterization for stable training and an asymmetrically initialized stereo tokenizer for effective view distinction. To better align the two views during stereo tokenization, we introduce a novel shift-embedding mechanism that encodes different disparity shifts into channel groups, preserving geometric cues even under large view displacements. A lightweight refinement module then sharpens the regressed disparity map for fine-grained accuracy. By prioritizing holistic regression over explicit matching, DispViT streamlines the stereo pipeline while improving robustness and efficiency. Experiments on standard benchmarks show that our approach achieves state-of-the-art accuracy, with strong resilience to matching ambiguities and wide disparity ranges. Code will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding et al.ICCV 2021 · 380 citations
Related papers
- S2M2: Scalable Stereo Matching Model for Reliable Depth EstimationJunhong Min, Youngpil Jeon, Jimin Kim, Minyong ChoiICCV 2025 · 8 citations
- DynamicStereo: Consistent Dynamic Depth from Stereo VideosNikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova et al.CVPR 2023
- Multi-Frame Self-Supervised Depth with TransformersVitor Guizilini, Rares Ambrus, Dian Chen, Sergey Zakharov et al.CVPR 2022 · 95 citations
- ChiTransformer: Towards Reliable Stereo from CuesQing Su, Shihao JiCVPR 2022 · 17 citations
- Iterative Geometry Encoding Volume for Stereo MatchingGangwei Xu, Xianqi Wang, Xiaohuan Ding, Xin YangCVPR 2023
