MINIMA: Modality Invariant Image Matching
Jiangwei Ren, Xingyu Jiang, Zizhuo Li, Dingkang Liang, Xin Zhou, Xiang Bai
摘要
Image matching for both cross-view and cross-modality plays a critical role in multimodal perception. In practice, the modality gap caused by different imaging systems/styles poses great challenges to the matching task. Existing works try to extract invariant features for specific modalities and train on limited datasets, showing poor generalization. In this paper, we present MINIMA, a unified image matching framework for multiple cross-modal cases. Without pursuing fancy modules, our MINIMA aims to enhance universal performance from the perspective of data scaling up. For such purpose, we propose a simple yet effective data engine that can freely produce a large dataset containing multiple modalities, rich scenarios, and accurate matching labels. Specifically, we scale up the modalities from cheap but rich RGB-only matching data, by means of generative models. Under this setting, the matching labels and rich diversity of the RGB dataset are well inherited by the generated multimodal data. Benefiting from this, we construct MD-syn, a new comprehensive dataset that fills the data gap for general multimodal image matching. With MD-syn, we can directly train any advanced matching pipeline on randomly selected modality pairs to obtain cross-modal ability. Extensive experiments on indomain and zero-shot matching tasks, including 19 crossmodal cases, demonstrate that our MINIMA can significantly outperform the baselines and even surpass modalityspecific methods. The dataset and code are available at https://github.com/LSXI7/MINIMA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- UFM: A Simple Path towards Unified Dense Correspondence with FlowYuchen Zhang, Nikhil Varma Keetha, Chenwei Lyu, Bhuvan Jhamb 等NeurIPS 2025 · 被引用 40 次
- RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim TranslationYash Jangir, Yidi Zhang, Kashu Yamazaki, Chenyu Zhang 等ICLR 2026 · 被引用 22 次
- ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image TranslationJiuhong Xiao, Roshan Nayak, Ning Zhang, Daniel Tortei 等NeurIPS 2025 · 被引用 18 次
- TherA: Thermal-Aware Visual-Language Prompting for Controllable RGB-to-Thermal Infrared TranslationDong-Guw Lee, Tai Hyoung Rhee, Hyunsoo Jang, Young-Sik Shin 等CVPR 2026 · 被引用 4 次
- CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image RegistrationXuecong Liu, Mengzhu Ding, Zixuan Sun, Zhang Li 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper20
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 被引用 936 次
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu 等CVPR 2022 · 被引用 929 次
相关 Paper
- MRGen: Segmentation Data Engine for Underrepresented MRI ModalitiesHaoning Wu, Ziheng Zhao, Ya Zhang, Yanfeng Wang 等ICCV 2025 · 被引用 3 次
- Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language ModelsXin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li 等CVPR 2025
- Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing ModalitiesJinming Zhao, Ruichen Li, Qin JinACL 2021
- MegaPairs: Massive Data Synthesis for Universal Multimodal RetrievalJunjie Zhou, Yongping Xiong, Zheng Liu, Ze Liu 等ACL 2025
- OmniDiff: A Comprehensive Benchmark for Fine-Grained Image Difference CaptioningYuan Liu, Saihui Hou, Saijie Hou, Jiabao Du 等ICCV 2025 · 被引用 1 次
