Learning Where to Edit Vision Transformers
Yunqiao Yang, Long-Kai Huang, Shengzhuang Chen, Kede Ma, Ying Wei
Abstract
Model editing aims to data-efficiently correct predictive errors of large pre-trained models while ensuring generalization to neighboring failures and locality to minimize unintended effects on unrelated examples. While significant progress has been made in editing Transformer-based large language models, effective strategies for editing vision Transformers (ViTs) in computer vision remain largely untapped. In this paper, we take initial steps towards correcting predictive errors of ViTs, particularly those arising from subpopulation shifts. Taking a locate-then-edit approach, we first address the where-to-edit challenge by meta-learning a hypernetwork on CutMix-augmented data generated for editing reliability. This trained hypernetwork produces generalizable binary masks that identify a sparse subset of structured model parameters, responsive to real-world failure samples. Afterward, we solve the how-to-edit problem by simply fine-tuning the identified parameters using a variant of gradient descent to achieve successful edits. To validate our method, we construct an editing benchmark that introduces subpopulation shifts towards natural underrepresented images and AI-generated images, thereby revealing the limitations of pre-trained ViTs for object recognition. Our approach not only achieves superior performance on the proposed benchmark but also allows for adjustable trade-offs between generalization and locality. Our code is available at https://github.com/hustyyq/Where-to-Edit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Model Editing for Vision TransformersXinyi Huang, Kangfei Zhao, Long-Kai HuangNeurIPS 2025 · 1 citation
- Exploring and Leveraging Class Vectors for Classifier EditingJaeik Kim, Jaeyoung DoNeurIPS 2025 · 1 citation
- Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video UnderstandingKaiting Liu, Hazel DoughtyICLR 2026
- Hiding Images in Diffusion Models by Editing Learned Score FunctionsHaoyu Chen, Yunqiao Yang, Nan Zhong, Kede MaCVPR 2025
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEditQizhou Chen, Taolin Zhang, Chengyu Wang, Xiaofeng He et al.AAAI 2025 · 9 citations
- Transformer-Patcher: One Mistake Worth One NeuronZeyu Huang, Yikang Shen, Xiaofeng Zhang, Jie Zhou et al.ICLR 2023 · 10 citations
- Editing the Moving World: Model Editing for Video LLMsQian Zhang, Xinye Li, Xiaokai Wu, Junhao Xu et al.ACL 2026
- All Tokens Matter: Token Labeling for Training Better Vision TransformersZihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou et al.NeurIPS 2021 · 252 citations
- Fine-tuning Done Right in Model EditingWanli Yang, Rui Tang, Hongyu Zang, Du Su et al.ICLR 2026 · 9 citations
