Representer Point Selection via Local Jacobian Expansion for Post-hoc Classifier Explanation of Deep Neural Networks and Ensemble Models
Yi Sui, Ga Wu, Scott Sanner
Abstract
Explaining the influence of training data on machine learning model predictions is a critical tool for debugging models through data curation. A recent appealing and efficient approach for this task was provided via the concept of Representer Point Selection (RPS), i.e. a method the leverages the dual form of l 2 regularized optimization in the last layer of the neural network to identify the contribution of training points to the prediction. However, two key drawbacks of RPS-l 2 are that they (i) lead to disagreement between the originally trained network and the RPS-l 2 regularized network modification and (ii) often yield a static ranking of training data for test points in the same class, independent of the test point being classified. Inspired by the RPS-l 2 approach, we propose an alternative method based on a local Jacobian Taylor expansion (LJE). We empirically compared RPS-LJE with the original RPS-l 2 on image classification (with ResNet), text classification recurrent neural networks (with Bi-LSTM), and tabular classification (with XGBoost) tasks. Quantitatively, we show that RPS-LJE slightly outperforms RPS-l 2 and other state-of-the-art data explanation methods by up to 3% on a data debugging task. More critically, we qualitatively observe that RPS-LJE provides stable and individualized explanations that are more coherent to each test data point. Overall, RPS-LJE represents a novel approach to RPS-l 2 that provides a powerful tool for sample-based model explanation and debugging. * Contributions were made while the author was at the University of Toronto.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- First is Better Than Last for Language Data InfluenceChih-Kuan Yeh, Ankur Taly, Mukund Sundararajan, Frederick Liu et al.NeurIPS 2022 · 39 citations
- Rethinking Data Shapley for Data Selection Tasks: Misleads and MeritsJiachen T. Wang, Tianji Yang, James Zou, Yongchan Kwon et al.ICML 2024 · 24 citations
- Sample based Explanations via Generalized RepresentersChe-Ping Tsai, Chih-Kuan Yeh, Pradeep RavikumarNeurIPS 2023 · 13 citations
- First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence EstimationDmytro Vitel, Anshuman ChhabraICLR 2026 · 8 citations
- Representer Point Selection for Explaining Regularized High-dimensional ModelsChe-Ping Tsai, Jiong Zhang, Hsiang-Fu Yu, Eli Chien et al.ICML 2023 · 5 citations
Builds on1
Related papers
- Unsupervised Object Localization with Representer Point SelectionYeonghwan Song, Seokwoo Jang, Dina Katabi, Jeany SonICCV 2023 · 4 citations
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 209 citations
- Data Glitches Discovery using Influence-based Model ExplanationsNikolaos Myrtakis, Ioannis Tsamardinos, Vassilis ChristophidesKDD 2025
- Explaining Latent Representations with a Corpus of ExamplesJonathan Crabbé, Zhaozhi Qian, Fergus Imrie, Mihaela van der SchaarNeurIPS 2021 · 48 citations
- Effective Optimization of Root Selection Towards Improved Explanation of Deep ClassifiersXin Zhang, Shenghua Zhong, Jianmin JiangACM MM 2024
