Learning Where to Edit Vision Transformers
Yunqiao Yang, Long-Kai Huang, Shengzhuang Chen, Kede Ma, Ying Wei
摘要
Model editing aims to data-efficiently correct predictive errors of large pre-trained models while ensuring generalization to neighboring failures and locality to minimize unintended effects on unrelated examples. While significant progress has been made in editing Transformer-based large language models, effective strategies for editing vision Transformers (ViTs) in computer vision remain largely untapped. In this paper, we take initial steps towards correcting predictive errors of ViTs, particularly those arising from subpopulation shifts. Taking a locate-then-edit approach, we first address the where-to-edit challenge by meta-learning a hypernetwork on CutMix-augmented data generated for editing reliability. This trained hypernetwork produces generalizable binary masks that identify a sparse subset of structured model parameters, responsive to real-world failure samples. Afterward, we solve the how-to-edit problem by simply fine-tuning the identified parameters using a variant of gradient descent to achieve successful edits. To validate our method, we construct an editing benchmark that introduces subpopulation shifts towards natural underrepresented images and AI-generated images, thereby revealing the limitations of pre-trained ViTs for object recognition. Our approach not only achieves superior performance on the proposed benchmark but also allows for adjustable trade-offs between generalization and locality. Our code is available at https://github.com/hustyyq/Where-to-Edit.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Model Editing for Vision TransformersXinyi Huang, Kangfei Zhao, Long-Kai HuangNeurIPS 2025 · 被引用 1 次
- Exploring and Leveraging Class Vectors for Classifier EditingJaeik Kim, Jaeyoung DoNeurIPS 2025 · 被引用 1 次
- Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video UnderstandingKaiting Liu, Hazel DoughtyICLR 2026
- Hiding Images in Diffusion Models by Editing Learned Score FunctionsHaoyu Chen, Yunqiao Yang, Nan Zhong, Kede MaCVPR 2025
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEditQizhou Chen, Taolin Zhang, Chengyu Wang, Xiaofeng He 等AAAI 2025 · 被引用 9 次
- Transformer-Patcher: One Mistake Worth One NeuronZeyu Huang, Yikang Shen, Xiaofeng Zhang, Jie Zhou 等ICLR 2023 · 被引用 10 次
- Editing the Moving World: Model Editing for Video LLMsQian Zhang, Xinye Li, Xiaokai Wu, Junhao Xu 等ACL 2026
- All Tokens Matter: Token Labeling for Training Better Vision TransformersZihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou 等NeurIPS 2021 · 被引用 252 次
- Fine-tuning Done Right in Model EditingWanli Yang, Rui Tang, Hongyu Zang, Du Su 等ICLR 2026 · 被引用 9 次
