Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language Models
Zhen Zeng, Leijiang Gu, Xun Yang, Zhangling Duan, Zenglin Shi, Meng Wang
Abstract
Knowledge editing aims to efficiently and cost-effectively correct inaccuracies and update outdated information. Recently, there has been growing interest in extending knowledge editing from Large Language Models (LLMs) to Multimodal Large Language Models (MLLMs), which integrate both textual and visual information, introducing additional editing complexities. Existing multimodal knowledge editing works primarily focus on text-oriented, coarse-grained scenarios, failing to address the unique challenges posed by multimodal contexts. In this paper, we propose a visual-oriented, fine-grained multimodal knowledge editing task that targets precise editing in images with multiple interacting entities. We introduce the Fine-Grained Visual Knowledge Editing (FGVEdit) benchmark to evaluate this task. Moreover, we propose a Multimodal Scope Classifier-based Knowledge Editor (MSCKE) framework. MSCKE leverages a multimodal scope classifier that integrates both visual and textual information to accurately identify and update knowledge related to specific entities within images. This approach ensures precise editing while preserving irrelevant information, overcoming the limitations of traditional text-only editing methods. Extensive experiments on the FGVEdit benchmark demonstrate that MSCKE outperforms existing methods, showcasing its effectiveness in solving the complex challenges of multimodal knowledge editing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fdb94c84-e2ae-4423-806d-75c72fb4a192Cited by top-tier papers7
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language ModelsXinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu et al.NeurIPS 2025 · 18 citations
- MemEIC: A Step Toward Continual and Compositional Knowledge EditingJin Seong, Jiyun Park, Wencke Liermann, Hongseok Choi et al.NeurIPS 2025 · 2 citations
- Reliable Lifelong Multimodal Editing: Conflict-Aware Retrieval Meets Multi-Level GuidanceQiang Zhang, Fanrui Zhang, Jiawei Liu, Ming Hu et al.NeurIPS 2025 · 1 citation
- ReasonEdit: Editing Vision--Language Models using Human ReasoningJiaxing Qiu, Kaihua Hou, Roxana Daneshjou, Ahmed Alaa et al.ICML 2026
- Can Knowledge Editing Really Correct Hallucinations?Baixiang Huang, Canyu Chen, Xiongxiao Xu, Ali Payani et al.ICLR 2025
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual KnowledgeYuntao Du, Kailin Jiang, Zhi Gao, Chenrui Shi et al.ICLR 2025
- MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language ModelsDexuan Xu, Jieyi Wang, Zhongyan Chai, Yongzhi Cao et al.AAAI 2026 · 1 citation
- M2Edit: Locate and Edit Multi-Granularity Knowledge in Multimodal Large Language ModelYang Zhou, Pengfei Cao, Yubo Chen, Qingbin Liu et al.EMNLP 2025
- MultiMedBench: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQAShengtao Wen, Haodong Chen, Yadong Wang, Zhongying Pan et al.AAAI 2026
- Can Knowledge be Transferred from Unimodal to Multimodal? Investigating the Transitivity of Multimodal Knowledge EditingLingyong Fang, Xinzhong Wang, Depeng Wang, Zongru Wu et al.ICCV 2025 · 4 citations
