MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity Recognition
Xinkui Lin, Yuhui Zhang, Yongxiu Xu, Kun Huang, Hongzhang Mu, Yubin Wang, Gaopeng Gou, Li Qian, Li Peng, Wei Liu, Jian Luan, Hongbo Xu
Abstract
Grounded Multimodal Named Entity Recognition (GMNER), which aims to extract textual entities, their types, and corresponding visual regions from image-text data, has become a critical task in multimodal information extraction. However, existing methods face two major challenges. First, they fail to address the semantic ambiguity caused by polysemy and the long-tail distribution of datasets. Second, unlike visual grounding which provides descriptive phrases, entity grounding only offers brief entity names which carry less semantic information. Current methods lack sufficient semantic interaction between text and image, hindering accurate entity-visual region matching. To tackle these issues, we propose MAKAR, a Multi-Agent framework based Knowledge-Augmented Reasoning, comprising three agents: Knowledge Enhancement, Entity Correction, and Entity Reasoning Grounding. Specifically, in the named entity recognition phase, the Knowledge Enhancement Agent leverages a Multimodal Large Language Model (MLLM) as an implicit knowledge base to enhance ambiguous image-text content with its internal knowledge. For samples with lowconfidence entity boundaries and types, the Entity Correction Agent uses web search tools to retrieve and summarize relevant web content, thereby correcting entities using both internal and external knowledge. In the entity grounding phase, the Entity Reasoning Grounding Agent utilizes multi-step Chain-of-Thought reasoning to perform grounding for each entity. Extensive experiments show that MAKAR achieves state-of-the-art performance on two benchmark datasets. Code is available at: https://github.com/Nikol-coder/MAKAR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3812ef03-14a3-4a2d-96cd-a2977d1d3002Cited by top-tier papers1
Ask how each one uses itBuilds on13
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal TransformerJianfei Yu, Jing Jiang, Li Yang, Rui XiaACL 2020 · 260 citations
- Multi-modal Graph Fusion for Named Entity Recognition with Targeted Visual GuidanceDong Zhang, Suzhong Wei, Shoushan Li, Hanqian Wu et al.AAAI 2021 · 240 citations
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng et al.SIGIR 2022 · 227 citations
- Joint Multimodal Entity-Relation Extraction Based on Edge-Enhanced Graph Alignment Network and Word-Pair Relation TaggingLi Yuan, Yi Cai, Jin Wang, Qing LiAAAI 2023 · 89 citations
Related papers
- UnCo: Uncertainty-Driven Collaborative Framework of Large and Small Models for Grounded Multimodal NERJielong Tang, Yang Yang, Jianxing Yu, Zhen-Xing Wang et al.EMNLP 2025 · 3 citations
- MNER-QG: An End-to-End MRC Framework for Multimodal Named Entity Recognition with Query GroundingMeihuizi Jia, Lei Shen, Xin Shen, Lejian Liao et al.AAAI 2023 · 68 citations
- Grounded Multimodal Named Entity Recognition on Social MediaJianfei Yu, Ziyan Li, Jieming Wang, Rui XiaACL 2023 · 32 citations
- Ground and Reconstruct: Entity-Region Bidirectional Alignment Pre-Training for Low-Resource GMNERRunwei Situ, Yi Cai, Yong Xu, Jiexin WangACM MM 2025 · 3 citations
- Fine-Grained Multimodal Named Entity Recognition and Grounding with a Generative FrameworkJieming Wang, Ziyan Li, Jianfei Yu, Li Yang et al.ACM MM 2023 · 11 citations
