Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
Kaihang Pan, Zhaoyu Fan, Juncheng Li, Qifan Yu, Hao Fei, Siliang Tang, Richang Hong, Hanwang Zhang, Qianru Sun
Abstract
The swift advancement in Multimodal LLMs (MLLMs) also presents significant challenges for effective knowledge editing. Current methods, including intrinsic knowledge editing and external knowledge resorting, each possess strengths and weaknesses, struggling to balance the desired properties of reliability, generality, and locality when applied to MLLMs. In this paper, we propose UniKE, a novel multimodal editing method that establishes a unified perspective and paradigm for intrinsic knowledge editing and external knowledge resorting. Both types of knowledge are conceptualized as vectorized key-value memories, with the corresponding editing processes resembling the assimilation and accommodation phases of human cognition, conducted at the same semantic levels. Within such a unified framework, we further promote knowledge collaboration by disentangling the knowledge representations into the semantic and truthfulness spaces. Extensive experiments validate the effectiveness of our method, which ensures that the post-edit MLLM simultaneously maintains excellent reliability, generality, and locality. The code for UniKE is available at https://github.com/beepkh/UniKE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- GraphCLIP: Enhancing Transferability in Graph Foundation Models for Text-Attributed GraphsYun Zhu, Haizhou Shi, Xiaotang Wang, Yongchao Liu et al.WWW 2025 · 54 citations
- Unified Generative and Discriminative Training for Multi-modal Large Language ModelsWei Chow, Juncheng Li, Qifan Yu, Kaihang Pan et al.NeurIPS 2024 · 19 citations
- Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement LearningKaihang Pan, Yang Wu, Wendong Bu, Kai Shen et al.NeurIPS 2025 · 11 citations
- WiseEdit: Benchmarking Cognition- and Creativity-Informed Image EditingKaihang Pan, Weile Chen, Haiyi Qiu, Qifan Yu et al.CVPR 2026 · 9 citations
- EvolvedGRPO: Unlocking Reasoning in LVLMs via Progressive Instruction EvolutionZhebei Shen, Qifan Yu, Juncheng Li, Wei Ji et al.NeurIPS 2025 · 2 citations
Builds on29
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe et al.EMNLP 2022 · 634 citations
Related papers
- Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMsXin Gao, Cheng Yang, Chufan Shi, Taylor Berg-KirkpatrickICML 2026
- Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMsTingchao Fu, Wenkai Wang, Fanxiao Li, Huadong Zhang et al.ACL 2026
- Everything is Editable: Extend Knowledge Editing to Unstructured Data in Large Language ModelsJingcheng Deng, Zihao Wei, Liang Pang, Hanxing Ding et al.ICLR 2025
- M2Edit: Locate and Edit Multi-Granularity Knowledge in Multimodal Large Language ModelYang Zhou, Pengfei Cao, Yubo Chen, Qingbin Liu et al.EMNLP 2025
- MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual KnowledgeYuntao Du, Kailin Jiang, Zhi Gao, Chenrui Shi et al.ICLR 2025
