ISR: Self-Refining Referring Expressions for Entity Grounding
Zhuocheng Yu, Bingchan Zhao, Yifan Song, Sujian Li, Zhonghui He
Abstract
Entity grounding, a crucial task in constructing multimodal knowledge graphs, aims to align entities from knowledge graphs with their corresponding images. Unlike conventional visual grounding tasks that use referring expressions (REs) as inputs, entity grounding relies solely on entity names and types, presenting a significant challenge. To address this, we introduce a novel I terative S elf-R efinement ( ISR ) scheme to enhance the multimodal large language model’s capability to generate high quality REs for the given entities as explicit contextual clues. This training scheme, inspired by human learning dynamics and human annotation processes, enables the MLLM to iteratively generate and refine REs by learning from successes and failures, guided by outcome re-wards from a visual grounding model. This iterative cycle of self-refinement avoids overfitting to fixed annotations and fosters continued improvement in referring expression generation. Extensive experiments demonstrate that our methods surpasses other methods in entity grounding, highlighting its effectiveness, robustness and potential for broader applications 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0251c1b-9647-4e90-9cef-4e474df39404Cited by top-tier papers1
Ask how each one uses itBuilds on14
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity LinkingZiyan Liu, Junwen Li, Kaiwen Li, Tong Ruan et al.ACM MM 2025 · 2 citations
- Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual GroundingWenbo Chen, Zhen Xu, Ruotao Xu, Si Wu et al.CVPR 2025
- PostAlign: Multimodal Grounding as a Corrective Lens for MLLMsYixuan Wu, Yang Zhang, Jian Wu, Philip Torr et al.ICLR 2026 · 5 citations
- MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity RecognitionXinkui Lin, Yuhui Zhang, Yongxiu Xu, Kun Huang et al.EMNLP 2025
- Entity Alignment with Noisy Annotations from Large Language ModelsShengyuan Chen, Qinggang Zhang, Junnan Dong, Wen Hua et al.NeurIPS 2024 · 44 citations
