Understanding Instance-based Interpretability of Variational Auto-Encoders
Zhifeng Kong, Kamalika Chaudhuri
Abstract
Instance-based interpretation methods have been widely studied for supervised learning methods as they help explain how black box neural networks predict. However, instance-based interpretations remain ill-understood in the context of unsupervised learning. In this paper, we investigate influence functions [Koh and Liang, 2017], a popular instance-based interpretation method, for a class of deep generative models called variational auto-encoders (VAE). We formally frame the counter-factual question answered by influence functions in this setting, and through theoretical analysis, examine what they reveal about the impact of training samples on classical unsupervised learning methods. We then introduce VAE- TracIn, a computationally efficient and theoretically sound solution based on Pruthi et al. [2020], for VAEs. Finally, we evaluate VAE-TracIn on several real world datasets with extensive quantitative and qualitative analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd6f4d04-c4a1-4d3e-b169-ee15d620d069Cited by top-tier papers15
- On Memorization in Probabilistic Deep Generative ModelsGerrit J. J. van den Burg, Christopher K. I. WilliamsNeurIPS 2021 · 92 citations
- Intriguing Properties of Data Attribution on Diffusion ModelsXiaosen Zheng, Tianyu Pang, Chao Du, Jing Jiang et al.ICLR 2024 · 41 citations
- Rethinking Influence Functions of Neural Networks in the Over-Parameterized RegimeRui Zhang, Shihua ZhangAAAI 2022 · 32 citations
- An Empirical Study of Memorization in NLPXiaosen Zheng, Jing JiangACL 2022 · 30 citations
- Label-Free Explainability for Unsupervised ModelsJonathan Crabbé, Mihaela van der SchaarICML 2022 · 24 citations
Builds on11
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 236 citations
- On Memorization in Probabilistic Deep Generative ModelsGerrit J. J. van den Burg, Christopher K. I. WilliamsNeurIPS 2021 · 92 citations
- Evaluation of Similarity-based ExplanationsKazuaki Hanawa, Sho Yokoi, Satoshi Hara, Kentaro InuiICLR 2021 · 79 citations
Related papers
- CF-OPT: Counterfactual Explanations for Structured PredictionGermain Vivier-Ardisson, Alexandre Forel, Axel Parmentier, Thibaut VidalICML 2024 · 3 citations
- Distributional Training Data Attribution: What do Influence Functions Sample?Bruno Kacper Mlodozeniec, Isaac Reid, Sam Power, David Krueger et al.NeurIPS 2025
- Theoretical and Practical Perspectives on what Influence Functions DoAndrea Schioppa, Katja Filippova, Ivan Titov, Polina ZablotskaiaNeurIPS 2023 · 38 citations
- Towards Building A Group-based Unsupervised Representation Disentanglement FrameworkTao Yang, Xuanchi Ren, Yuwang Wang, Wenjun Zeng et al.ICLR 2022 · 36 citations
- PINNfluence: Interpreting PINNs through Influence FunctionsAleksander Krasowski, Jonas Naujoks, Moritz Weckbecker, Galip Yolcu et al.ICML 2026 · 1 citation
