ProtoVAE: A Trustworthy Self-Explainable Prototypical Variational Model
Srishti Gautam, Ahcène Boubekki, Stine Hansen, Suaiba Amina Salahuddin, Robert Jenssen, Marina M.-C. Höhne, Michael Kampffmeyer
Abstract
The need for interpretable models has fostered the development of self-explainable classifiers. Prior approaches are either based on multi-stage optimization schemes, impacting the predictive performance of the model, or produce explanations that are not transparent, trustworthy or do not capture the diversity of the data. To address these shortcomings, we propose ProtoVAE, a variational autoencoder-based framework that learns class-specific prototypes in an end-to-end manner and enforces trustworthiness and diversity by regularizing the representation space and introducing an orthonormality constraint. Finally, the model is designed to be transparent by directly incorporating the prototypes into the decision process. Extensive comparisons with previous self-explainable approaches demonstrate the superiority of ProtoVAE, highlighting its ability to generate trustworthy and diverse explanations, while not degrading predictive performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4134357f-badc-4cab-89f1-7f010559b18cCited by top-tier papers8
- Encoding Time-Series Explanations through Self-Supervised Model Behavior ConsistencyOwen Queen, Tom Hartvigsen, Teddy Koker, Huan He et al.NeurIPS 2023 · 55 citations
- Explaining Time Series via Contrastive and Locally Sparse PerturbationsZichuan Liu, Yingying Zhang, Tianchun Wang, Zefan Wang et al.ICLR 2024 · 26 citations
- ProtoLens: Advancing Prototype Learning for Fine-Grained Interpretability in Text ClassificationBowen Wei, Ziwei ZhuACL 2025 · 7 citations
- Pantypes: Diverse Representatives for Self-Explainable ModelsRune D. Kjærsgaard, Ahcène Boubekki, Line H. ClemmensenAAAI 2024 · 6 citations
- DISCRET: Synthesizing Faithful Explanations For Treatment Effect EstimationYinjun Wu, Mayank Keoliya, Kan Chen, Neelay Velingker et al.ICML 2024 · 3 citations
Builds on8
- Understanding Dimensional Collapse in Contrastive Self-supervised LearningLi Jing, Pascal Vincent, Yann LeCun, Yuandong TianICLR 2022 · 467 citations
- Interpretable Image Recognition by Constructing Transparent Embedding SpaceJiaqi Wang, Huafeng Liu, Xinyue Wang, Liping JingICCV 2021 · 149 citations
- ProtoPShare: Prototypical Parts Sharing for Similarity Discovery in Interpretable Image ClassificationDawid Rymarczyk, Lukasz Struski, Jacek Tabor, Bartosz ZielinskiKDD 2021 · 78 citations
- Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on ImagesRewon ChildICLR 2021 · 45 citations
- A Framework to Learn with InterpretationJayneel Parekh, Pavlo Mozharovskyi, Florence d'Alché-BucNeurIPS 2021 · 35 citations
Related papers
- Interpretable Image Classification via Non-parametric Part Prototype LearningZhijie Zhu, Lei Fan, Maurice Pagnucco, Yang SongCVPR 2025
- ProtoVAR: Efficient Dataset Distillation via Prototype-Guided Visual Autoregressive ModelingMingyu Wang, Wei JiangICML 2026
- ProtoTS: Learning Hierarchical Prototypes for Explainable Time Series ForecastingZiheng Peng, Shijie Ren, Xinyue Gu, Linxiao Yang et al.ICLR 2026 · 2 citations
- ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse AutoencodersXiangyu Liu, Haodi Lei, Yi Liu, Yang Liu et al.AAAI 2026 · 2 citations
- PEVAE: A Hierarchical VAE for Personalized Explainable RecommendationZefeng Cai, Zerui CaiSIGIR 2022 · 18 citations
