ReasonEdit: Editing Vision--Language Models using Human Reasoning
Jiaxing Qiu, Kaihua Hou, Roxana Daneshjou, Ahmed Alaa, Thomas Hartvigsen
Abstract
Model editing aims to correct errors in large, pretrained models without altering their unrelated behaviors. While some recent works have edited vision-language models (VLMs), no existing editors tackle reasoning-heavy tasks, which typically require humans and models to reason about images. We therefore propose ReasonEdit, the first VLM editor to let users explain their reasoning during editing, introducing a new, practical model editing setup. ReasonEdit continuously stores human reasoning in a codebook, and retrieves only relevant facts during inference using a novel topology-balanced multimodal embedding method inspired by network science. Across four VLMs on multiple rationale-based visual question answering datasets, ReasonEdit achieves stateof-the-art editing performance, ultimately showing that using human reasoning during editing greatly improves edit generalization. Our code and data are available at https://github. com/JiaxingQiu/reasonedit . ReasonEdit: Editing Vision-Language Models using Human Reasoning Topology-Balanced multi-modal embedding Codebook It's made with bamboo. Bamboo is firm and flexible. … It's a banjo. Banjo is stringed instrument. … Corona is to the left of books. Corona is a brad of beer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2568e962-0eac-4a96-8a03-a09b760659caBuilds on15
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
Related papers
- Hybrid-DMKG: A Hybrid Reasoning Framework over Dynamic Multimodal Knowledge Graphs for Multimodal Multihop QA with Knowledge EditingLi Yuan, Qingfei Huang, Bingshan Zhu, Yi Cai et al.AAAI 2026
- MemEIC: A Step Toward Continual and Compositional Knowledge EditingJin Seong, Jiyun Park, Wencke Liermann, Hongseok Choi et al.NeurIPS 2025 · 2 citations
- M2Edit: Locate and Edit Multi-Granularity Knowledge in Multimodal Large Language ModelYang Zhou, Pengfei Cao, Yubo Chen, Qingbin Liu et al.EMNLP 2025
- BalancEdit: Dynamically Balancing the Generality-Locality Trade-off in Multi-modal Model EditingDongliang Guo, Mengxuan Hu, Zihan Guan, Thomas Hartvigsen et al.ICML 2025
- MultiMedBench: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQAShengtao Wen, Haodong Chen, Yadong Wang, Zhongying Pan et al.AAAI 2026
