MACK: Multimodal Aligned Conceptual Knowledge for Unpaired Image-text Matching
Yan Huang, Yuming Wang, Yunan Zeng, Liang Wang
Abstract
Recently, the accuracy of image-text matching has been greatly improved by multimodal pretrained models, all of which are trained on millions or billions of paired images and texts. Different from them, this paper studies a new scenario as unpaired image-text matching, in which paired images and texts are assumed to be unavailable during model training. To deal with this, we propose a simple yet effective method namely Multimodal Aligned Conceptual Knowledge (MACK), which is inspired by the knowledge use in human brain. It can be directly used as general knowledge to correlate images and texts even without model training, or further fine-tuned based on unpaired images and texts to better generalize to certain datasets. In addition, we extend it as a re-ranking method, which can be easily combined with existing image-text matching models to substantially improve their performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5201052f-2758-408b-ac51-704fba3f1a27Cited by top-tier papers8
- Cross-modal Active Complementary Learning with Self-refining CorrespondenceYang Qin, Yuan Sun, Dezhong Peng, Joey Tianyi Zhou et al.NeurIPS 2023 · 49 citations
- RegBN: Batch Normalization of Multimodal Data with RegularizationMorteza Ghahremani, Christian WachingerNeurIPS 2023 · 15 citations
- Towards Deconfounded Image-Text Matching with Causal InferenceWenhui Li, Xinqi Su, Dan Song, Lanjun Wang et al.ACM MM 2023 · 12 citations
- Mitigating Noisy Correspondence by Geometrical Structure Consistency LearningZihua Zhao, Mengxi Chen, Tianjie Dai, Jiangchao Yao et al.CVPR 2024 · 5 citations
- Noisy Correspondence Rectification via Asymmetric Similarity LearningYunbo Wang, YuJie Wu, Zhien Dai, Can Tian et al.AAAI 2025 · 4 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image RetrievalZihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng et al.ICCV 2019 · 349 citations
Related papers
- Multimodal Aligned Semantic Knowledge for Unpaired Image-text MatchingLaiguo Yin, Yixin Zhang, YUQING SUN, Lizhen CuiICLR 2026
- Unsupervised Natural Language Inference via Decoupled Multimodal Contrastive LearningWanyun Cui, Guangyu Zheng, Wei WangEMNLP 2020 · 17 citations
- Conceptual and Syntactical Cross-modal Alignment with Cross-level Consistency for Image-Text MatchingPengpeng Zeng, Lianli Gao, Xinyu Lyu, Shuaiqi Jing et al.ACM MM 2021 · 37 citations
- Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person RetrievalDing Jiang, Mang YeCVPR 2023
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
