I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking
Ziyan Liu, Junwen Li, Kaiwen Li, Tong Ruan, Chao Wang, Xinyan He, Zongyu Wang, Xuezhi Cao, Jingping Liu
Abstract
Multimodal entity linking plays a crucial role in a wide range of applications. Recent advances in large language model-based methods have become the dominant paradigm for this task, effectively leveraging both textual and visual modalities to enhance performance. Despite their success, these methods still face two challenges, including unnecessary incorporation of image data in certain scenarios and the reliance only on a one-time extraction of visual features, which can undermine their effectiveness and accuracy. To address these challenges, we propose a novel LLM-based framework for the multimodal entity linking task, called Intra- and Inter-modal Collaborative Reflections. This framework prioritizes leveraging text information to address the task. When text alone is insufficient to link the correct entity through intra- and inter-modality evaluations, it employs a multi-round iterative strategy that integrates key visual clues from various aspects of the image to support reasoning and enhance matching accuracy. Extensive experiments on three widely used public datasets demonstrate that our framework consistently outperforms current state-of-the-art methods in the task, achieving improvements of 3.2%, 5.1%, and 1.6%, respectively. Our code is available at https://github.com/ziyan-xiaoyu/I2CR/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb23ef4d-7708-4543-8afa-8c10b2e3fa73Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Scalable Zero-shot Entity Linking with Dense Entity RetrievalLedell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel et al.EMNLP 2020 · 336 citations
Related papers
- ISR: Self-Refining Referring Expressions for Entity GroundingZhuocheng Yu, Bingchan Zhao, Yifan Song, Sujian Li et al.ACL 2025 · 1 citation
- CoMCo: Consistency-Aware Multi-Agent Coordination for Zero-Shot Cross-Modal Entity MatchingShiqi Zhang, Weixin Zeng, Ziheng Zhang, Wenzhe Hou et al.SIGIR 2026
- U-MERE: Unconstrained Multimodal Entity and Relation Extraction with Collaborative Modeling and Order-Sensitive OptimizationWei Jia, Li Jin, Kaiwen Wei, Yuying Shang et al.ACM MM 2025
- Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question AnsweringFederico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi et al.CVPR 2025
- MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity RecognitionXinkui Lin, Yuhui Zhang, Yongxiu Xu, Kun Huang et al.EMNLP 2025
