Token Embeddings Alignment for Cross-Modal Retrieval
Chen-Wei Xie, Jianmin Wu, Yun Zheng, Pan Pan, Xian-Sheng Hua
Abstract
Cross-modal retrieval has achieved significant progress in recent years with the help of token embeddings interaction methods. Most existing methods first extract embedding for each token of input image and text, then feed the token-level embeddings into a multi-modal transformer to learn a joint representation, this joint representation can be used to predict matching score between input image and text. However, these methods don't explicitly supervise the alignment between visual and textual tokens. In this paper, we propose a novel Token Embeddings AlignMent (TEAM) block, it first explicitly aligns visual tokens and textual tokens, then produces token-level matching scores to measure fine-grained similarity between input image and text. TEAM achieves new state-of-the-art performance on commonly used cross-modal retrieval benchmarks. Moreover, TEAM is interpretable and we provide visualization experiments to show how it works. At last, we construct a new billion-scale vision-language pre-training dataset in Chinese, which is the largest Chinese vision-language pre-training dataset so far. After pre-training on this dataset, our framework also achieves state-of-the-art performance on Chinese cross-modal retrieval benchmarks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 757af9bf-d697-432c-92f8-c9c20e673770Cited by top-tier papers3
- Causal Inference over Visual-Semantic-Aligned Graph for Image ClassificationLei Meng, Xiangxian Li, Xiaoshuo Yan, Haokai Ma et al.AAAI 2025 · 11 citations
- Explicit Modeling of Causal Factors and Confounders for Image ClassificationWei Wu, Lei Meng, Zhuang Qi, Zixuan Li et al.AAAI 2026
- Explainability and Interpretability of Multilingual Large Language Models: A SurveyLucas Resck, Isabelle Augenstein, Anna KorhonenEMNLP 2025
Related papers
- VL-Match: Enhancing Vision-Language Pretraining with Token-Level and Instance-Level MatchingJunyu Bi, Daixuan Cheng, Ping Yao, Bochen Pang et al.ICCV 2023 · 6 citations
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu et al.ICLR 2022 · 827 citations
- Seeing the Image: Prioritizing Visual Correlation by Contrastive AlignmentXin Xiao, Bohong Wu, Jiacong Wang, Chunyuan Li et al.NeurIPS 2024 · 24 citations
- CCMB: A Large-scale Chinese Cross-modal BenchmarkChunyu Xie, Heng Cai, Jincheng Li, Fanjing Kong et al.ACM MM 2023 · 10 citations
- Unsupervised Vision-and-Language Pretraining via Retrieval-based Multi-Granular AlignmentMingyang Zhou, Licheng Yu, Amanpreet Singh, Mengjiao Wang et al.CVPR 2022 · 29 citations
