MPBoCo: Multimodal Prompt-based Boundary-enhanced Continual Framework for Joint Entity and Relation Extraction
Guanglu Sun, Xinyu Liu, Lili Liang, Yang Yu, Fei Lang, Suxia Zhu, Ming Liu
Abstract
In real-world scenarios, multimodal information continuously evolves, with new entity and relation types emerging, necessitating timely updates to multimodal knowledge graphs for supporting downstream tasks. However, existing methods struggle to balance real-time adaptability and computational efficiency in continual learning scenarios. To this end, this paper proposes the Continual Multimodal Entity and Relation Joint Extraction (CMERJE) task and a Multimodal Prompt-based Boundaryenhanced Continual (MPBoCo) framework. Specifically, MPBoCo incrementally stores task-specific knowledge via learnable multimodal prompts, dynamically matches relevant prompts for each instance, and fuses them into a frozen backbone model for task-specific reasoning. Subsequently, the boundary-enhanced dual-branch module leverages the auxiliary branch to preserve local syntactic continuity and provide boundary guidance. Experimental results demonstrate that MPBoCo achieves superior performance in real-world scenarios, significantly outperforming baseline methods by 5.5% and 7.2% in 10-task and 5-task settings, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9aecf116-537e-4b74-90e2-f46eaaa8e8cfBuilds on13
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
- Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal TransformerJianfei Yu, Jing Jiang, Li Yang, Rui XiaACL 2020 · 260 citations
- Multi-modal Graph Fusion for Named Entity Recognition with Targeted Visual GuidanceDong Zhang, Suzhong Wei, Shoushan Li, Hanqian Wu et al.AAAI 2021 · 240 citations
- RpBERT: A Text-image Relation Propagation-based BERT Model for Multimodal NERLin Sun, Jiquan Wang, Kai Zhang, Yindu Su et al.AAAI 2021 · 189 citations
- Multimodal Representation with Embedded Visual Guiding Objects for Named Entity Recognition in Social Media PostsZhiwei Wu, Changmeng Zheng, Yi Cai, Junying Chen et al.ACM MM 2020 · 139 citations
Related papers
- Joint Multimodal Entity-Relation Extraction Based on Edge-Enhanced Graph Alignment Network and Word-Pair Relation TaggingLi Yuan, Yi Cai, Jin Wang, Qing LiAAAI 2023 · 89 citations
- Towards Multimodal Continual Knowledge Embedding wth Modality Forgetting ModulationXiaowen Jiang, Jing Yang, Shundong Yang, Yuan Gao et al.AAAI 2026
- Continual Relation Extraction via Sequential Multi-Task LearningThanh-Thien Le, Manh Nguyen, Tung Thanh Nguyen, Ngo Van Linh et al.AAAI 2024 · 16 citations
- REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-ExpertsXinkui Lin, Yongxiu Xu, Minghao Tang, Shilong Zhang et al.ACM MM 2025 · 2 citations
- Adaptive Prompting for Continual Relation Extraction: A Within-Task Variance PerspectiveMinh Le, Tien Ngoc Luu, An Nguyen The, Thanh-Thien Le et al.AAAI 2025 · 12 citations
