All-in-One: Transferring Vision Foundation Models into Stereo Matching
Jingyi Zhou, Haoyu Zhang, Jiakang Yuan, Peng Ye, Tao Chen, Hao Jiang, Meiya Chen, Yangyang Zhang
摘要
As a fundamental vision task, stereo matching has made remarkable progress. While recent iterative optimization-based methods have achieved promising performance, their feature extraction capabilities still have room for improvement. Inspired by the ability of vision foundation models (VFMs) to extract general representations, in this work, we propose AIO-Stereo which can flexibly select and transfer knowledge from multiple heterogeneous VFMs to a single stereo matching model. To better reconcile features between heterogeneous VFMs and the stereo matching model and fully exploit prior knowledge from VFMs, we proposed a dual-level feature utilization mechanism that aligns heterogeneous features and transfers multi-level knowledge. Based on the mechanism, a dual-level selective knowledge transfer module is designed to selectively transfer knowledge and integrate the advantages of multiple VFMs. Experimental results show that AIO-Stereo achieves start-of-the-art performance on multiple datasets and ranks 1st on the Middlebury dataset and outperforms all the published work on the ETH3D benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingBowen Wen, Shaurya Dewan, Stan BirchfieldCVPR 2026 · 被引用 36 次
- S2M2: Scalable Stereo Matching Model for Reliable Depth EstimationJunhong Min, Youngpil Jeon, Jimin Kim, Minyong ChoiICCV 2025 · 被引用 8 次
- FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion BasesMatteo Poggi, Fabio TosiICCV 2025 · 被引用 5 次
- Lite Any Stereo: Efficient Zero-Shot Stereo MatchingJunpeng Jing, Weixun Luo, Ye Mao, Krystian MikolajczykCVPR 2026 · 被引用 4 次
- Stereo Any Video: Temporally Consistent Stereo MatchingJunpeng Jing, Weixun Luo, Ye Mao, Krystian MikolajczykICCV 2025 · 被引用 1 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
相关 Paper
- Learning Robust Stereo Matching in the Wild with Selective Mixture-of-ExpertsYun Wang, Longguang Wang, Chenghao Zhang, Yongjian Zhang 等ICCV 2025 · 被引用 6 次
- DEFOM-Stereo: Depth Foundation Model Based Stereo MatchingHualie Jiang, Zhiqiang Lou, Laiyan Ding, Rui Xu 等CVPR 2025
- PRISM: Synergizing Vision Foundation Models via Self-organized Expert SpecializationYing Tang, Dong Li, Youjia Zhang, Zikai Song 等ICML 2026
- FoundationStereo: Zero-Shot Stereo MatchingBowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz 等CVPR 2025
- Diving into the Fusion of Monocular Priors for Generalized Stereo MatchingChengtang Yao, Lidong Yu, Zhidan Liu, Jiaxi Zeng 等ICCV 2025 · 被引用 3 次
