When Large Multimodal Models Confront Evolving Knowledge: Challenges and Explorations
Kailin Jiang, Yuntao Du, Yukai Ding, Yuchen Ren, Ning Jiang, Zhi Gao, Zilong Zheng, Lei Liu, Bin Li, Qing Li
摘要
Large Multimodal Models (LMMs) store vast amounts of pretrained knowledge but struggle to remain aligned with real-world updates, making it difficult to avoid capability degradation when acquiring evolving knowledge. Furthermore, most current work focuses on exploring static textual knowledge injection, neglecting dynamic multimodal evolving knowledge injection, leaving the potential of LMMs for multimodal knowledge injection as an open question. To address this, we first propose a pipeline to construct MMEVOKE, a benchmark for evaluating LMMs' ability in multimodal evolving knowledge injection. MMEVOKE contains 9,422 samples spanning 159 subtypes. Then, based on extensive experiments with MMEVOKE, we reveal challenges such as poor injection performance and capability degradation in existing knowledge injection methods through knowledge injection tests and general capability tests. Finally, to tackle these challenges, we introduce knowledge augmentation and knowledge retention methods, finding that knowledge-aware augmentation strengthens knowledge injection performance, and that Data Replay and MoE methods effectively mitigate capability degradation. Project Page: https://evoke-lmm.github.io/ * Equal contribution. † Corresponding author.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented ControlsKailin Jiang, Hongbo Jiang, Ning Jiang, Zhi Gao 等ICML 2026
- Modality-Decoupled Online Recursive EditingSiyuan Li, Youyuan Zhang, Fangming Liu, Jing LiICML 2026
- DecomPose: Disentangling Cross-Category Optimization Contention for Category-Level 6D Object Pose EstimationYifan Gao, Lu Zou, Zhangjin Huang, Guoping WangICML 2026
它引用的顶会 Paper43
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual KnowledgeYuntao Du, Kailin Jiang, Zhi Gao, Chenrui Shi 等ICLR 2025
- MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMsHe Li, Haoang Chi, Qizhou Wang, Yunxin Mao 等ICML 2026
- MMKU-Bench: A Multimodal Update Benchmark for Diverse Visual KnowledgeBaochen Fu, Yuntao Du, Cheng Chang, Baihao Jin 等ICML 2026 · 被引用 7 次
- Benchmarking Multimodal Knowledge Conflict for Large Multimodal ModelsYifan Jia, Yuntao Du, Kailin Jiang, Yuyang Liang 等AAAI 2026
- Can Knowledge be Transferred from Unimodal to Multimodal? Investigating the Transitivity of Multimodal Knowledge EditingLingyong Fang, Xinzhong Wang, Depeng Wang, Zongru Wu 等ICCV 2025 · 被引用 4 次
