ICML2026

ReasonEdit: Editing Vision--Language Models using Human Reasoning

Jiaxing Qiu, Kaihua Hou, Roxana Daneshjou, Ahmed Alaa, Thomas Hartvigsen

Abstract

Model editing aims to correct errors in large, pretrained models without altering their unrelated behaviors. While some recent works have edited vision-language models (VLMs), no existing editors tackle reasoning-heavy tasks, which typically require humans and models to reason about images. We therefore propose ReasonEdit, the first VLM editor to let users explain their reasoning during editing, introducing a new, practical model editing setup. ReasonEdit continuously stores human reasoning in a codebook, and retrieves only relevant facts during inference using a novel topology-balanced multimodal embedding method inspired by network science. Across four VLMs on multiple rationale-based visual question answering datasets, ReasonEdit achieves stateof-the-art editing performance, ultimately showing that using human reasoning during editing greatly improves edit generalization. Our code and data are available at https://github. com/JiaxingQiu/reasonedit . ReasonEdit: Editing Vision-Language Models using Human Reasoning Topology-Balanced multi-modal embedding Codebook It's made with bamboo. Bamboo is firm and flexible. … It's a banjo. Banjo is stringed instrument. … Corona is to the left of books. Corona is a brad of beer.