Mind the Gap: Aligning Vision Foundation Models to Image Feature Matching
Yuhan Liu, Jingwen Fu, Yang Wu, Kangyi Wu, Pengna Li, Jiayi Wu, Sanping Zhou, Jingmin Xin
Abstract
Leveraging the vision foundation models has emerged as a mainstream paradigm that improves the performance of image feature matching. However, previous works have ignored the misalignment when introducing the foundation models into feature matching. The misalignment arises from the discrepancy between the foundation models focusing on single-image understanding and the cross-image understanding requirement of feature matching. Specifically, 1) the embeddings derived from commonly used foundation models exhibit discrepancies with the optimal embeddings required for feature matching; 2) lacking an effective mechanism to leverage the single-image understanding ability into cross-image understanding. A significant consequence of the misalignment is they struggle when addressing multi-instance feature matching problems. To address this, we introduce a simple but effective framework, called IMD (Image feature Matching with a pre-trained Diffusion model) with two parts: 1) Unlike the dominant solutions employing contrastive-learning based foundation models that emphasize global semantics, we integrate the generative-based diffusion models to effectively capture instance-level details. 2) We leverage the prompt mechanism in generative model as a natural tunnel, propose a novel cross-image interaction prompting module to facilitate bidirectional information interaction between image pairs. To more accurately measure the misalignment, we propose a new benchmark called IMIM, which focuses on multi-instance scenarios. Our proposed IMD establishes a new state-of-the-art in commonly evaluated benchmarks, and the superior improvement 12% in IMIM indicates our method efficiently mitigates the misalignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc781c46-dfe5-4314-9ee0-76c2cdbf3a23Cited by top-tier papers2
- DynamicID: Zero-Shot Multi-ID Image Personalization With Flexible Facial EditabilityXirui Hu, Jiahao Wang, Hao Chen, Weizhan Zhang et al.ICCV 2025 · 3 citations
- Think before Go: Hierarchical Reasoning for Image-goal NavigationPengna Li, Kangyi Wu, Shaoqing Xu, Fang Li et al.ACL 2026 · 2 citations
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Where's Waldo: Diffusion Features For Personalized Segmentation and RetrievalDvir Samuel, Rami Ben-Ari, Matan Levy, Nir Darshan et al.NeurIPS 2024 · 19 citations
- Explore In-Context Segmentation via Latent Diffusion ModelsChaoyang Wang, Xiangtai Li, Henghui Ding, Lu Qi et al.AAAI 2025 · 17 citations
- SynerDetect: Hierarchical Synergistic Learning for Generalizable AI-Generated Image DetectionShuaibo Li, Yijun Yang, Zhaohu Xing, Hongqiu Wang et al.AAAI 2026
- Text-Image Alignment for Diffusion-Based PerceptionNeehar Kondapaneni, Markus Marks, Manuel Knott, Rogério Guimarães et al.CVPR 2024
- SD4Match: Learning to Prompt Stable Diffusion Model for Semantic MatchingXinghui Li, Jingyi Lu, Kai Han, Victor Adrian PrisacariuCVPR 2024
