Improving Prototypical Visual Explanations with Reward Reweighing, Reselection, and Retraining
Aaron Jiaxun Li, Robin Netzorg, Zhihan Cheng, Zhuoqin Zhang, Bin Yu
Abstract
In recent years, work has gone into developing deep interpretable methods for image classification that clearly attributes a model's output to specific features of the data. One such of these methods is the prototypical part network (ProtoP-Net), which attempts to classify images based on meaningful parts of the input. While this architecture is able to produce visually interpretable classifications, it often learns to classify based on parts of the image that are not semantically meaningful. To address this problem, we propose the reward reweighing, reselecting, and retraining (R3) post-processing framework, which performs three additional corrective updates to a pretrained ProtoPNet in an offline and efficient manner. The first two steps involve learning a reward model based on collected human feedback and then aligning the prototypes with human preferences. The final step is retraining, which realigns the base features and the classifier layer of the original model with the updated prototypes. We find that our R3 framework consistently improves both the interpretability and the predictive accuracy of ProtoPNet and its variants.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62b9e9da-bc8c-4251-a294-8cef452ddb34Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 219 citations
- Deformable ProtoPNet: An Interpretable Image Classifier Using Deformable PrototypesJon Donnelly, Alina Jade Barnett, Chaofan ChenCVPR 2022 · 101 citations
- ProtoPShare: Prototypical Parts Sharing for Similarity Discovery in Interpretable Image ClassificationDawid Rymarczyk, Lukasz Struski, Jacek Tabor, Bartosz ZielinskiKDD 2021 · 78 citations
- This Looks Like Those: Illuminating Prototypical Concepts Using Multiple VisualizationsChiyu Ma, Brandon Zhao, Chaofan Chen, Cynthia RudinNeurIPS 2023 · 53 citations
- Evaluation and Improvement of Interpretability for Self-Explainable Part-Prototype NetworksQihan Huang, Mengqi Xue, Wenqi Huang, Haofei Zhang et al.ICCV 2023 · 47 citations
Related papers
- Interpretable Image Classification via Non-parametric Part Prototype LearningZhijie Zhu, Lei Fan, Maurice Pagnucco, Yang SongCVPR 2025
- Learning Support and Trivial Prototypes for Interpretable Image ClassificationChong Wang, Yuyuan Liu, Yuanhong Chen, Fengbei Liu et al.ICCV 2023 · 50 citations
- PIP-Net: Patch-Based Intuitive Prototypes for Interpretable Image ClassificationMeike Nauta, Jörg Schlötterer, Maurice van Keulen, Christin SeifertCVPR 2023
- ProtoArgNet: Interpretable Image Classification with Super-Prototypes and ArgumentationHamed Ayoobi, Nico Potyka, Francesca ToniAAAI 2025 · 8 citations
- Post-hoc Part-Prototype NetworksAndong Tan, Fengtao Zhou, Hao ChenICML 2024 · 7 citations
