Beyond General Alignment: Fine-Grained Entity-Centric Image-Text Matching with Multimodal Attentive Experts
Yaxiong Wang, Lianwei Wu, Lechao Cheng, Zhun Zhong, Yujiao Wu, Meng Wang
Abstract
Recent progress in aligning images with texts has achieved remarkable results, however, existing models tend to serve general queries and often fall short when dealing with detailed query requirements. In this paper, we work towards Entity-centric Image-Text Matching (EITM), a finer-grained image-text matching task that aligns texts and images centered around specific entities. The main challenge in EITM lies in bridging the substantial semantic gap between entity-related information in texts and images, which is more pronounced than in general image-text matching problems. To address this challenge, we adopt CLIP as our foundational model and devise a Multimodal Attentive Experts (MMAE)-based contrastive learning to adapt CLIP into an expert for EITM problem. Particularly, the core of our multimodal attentive experts learning is to generate explanation texts by Large Language Models (LLMs) as bridging clues. In specific, we first employ off-the-shelf LLMs to generate explanatory text. This text, along with the original image and text, is then fed into our Multimodal Attentive Experts module to narrow the semantic gap within a unified semantic space. Upon the enriched feature representations generated by MMAE, we have further developed an effective Gated Integrative Image-text Matching (GI-ITM) strategy. GI-ITM utilizes an adaptive gating mechanism to combine features from MMAE, followed by applying image-text matching constraints to enhance the alignment precision. Our method has been extensively evaluated on three social media news benchmarks: N24News, VisualNews, and GoodNews. The experimental results demonstrate that our approach significantly outperforms competing methods. Our code is available at: https://github.com/wangyxxjtu/ETE.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8ebbd9c7-2900-473d-bb78-d15c48119a34Cited by top-tier papers5
- Open-World 3D Scene Graph Generation for Retrieval-Augmented ReasoningFei Yu, Quan Deng, Shengeng Tang, Yuehua Li et al.AAAI 2026 · 2 citations
- Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal ManipulationsJinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu et al.ACM MM 2025 · 2 citations
- Beyond Walking: A Large-Scale Image-Text Benchmark for Text-Based Person Anomaly SearchShuyu Yang, Yaxiong Wang, Li Zhu, Zhedong ZhengICCV 2025 · 1 citation
- Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person SearchJiahao Zhang, Shaofei Huang, Yaxiong Wang, Zhedong ZhengSIGIR 2026
- Decoupled Training with Local Reinforcement Fine-Tuning in Federated LearningYuting Ma, Lechao Cheng, Xiaohua XuICML 2026
Related papers
- Entity-Level Alignment with Prompt-Guided Adapter for Remote Sensing Image-Text RetrievalShuoshuo Li, Shuli Cheng, Liejun WangACM MM 2025 · 2 citations
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 12 citations
- LLM-Enhanced Action-Aware Multi-Modal Prompt Tuning for Image-Text MatchingMengxiao Tian, Xinxiao Wu, Shuo YangICCV 2025 · 3 citations
- LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality RepresentationWeiquan Huang, Aoqi Wu, Yifan Yang, Xufang Luo et al.AAAI 2026
- Contrastive Localized Language-Image Pre-TrainingHong-You Chen, Zhengfeng Lai, Haotian Zhang, Xinze Wang et al.ICML 2025
