Manifold-Aligned Guided Integrated Gradients for Reliable Feature Attribution
Soyeon Kim, Seongwoo Lim, Kyowoon Lee, Jaesik Choi
摘要
Feature attribution is central to diagnosing and trusting deep neural networks, and Integrated Gradients (IG) is widely used due to its axiomatic properties. However, IG can yield unreliable explanations when the integration path between a baseline and the input passes through regions with noisy gradients. While Guided Integrated Gradients reduces this sensitivity by adaptively updating low-gradient-magnitude features, inputspace guidance still produces intermediate inputs that deviate from the data manifold. To address this limitation, we propose Manifold-Aligned Guided Integrated Gradients (MA-GIG), which constructs attribution paths in the latent space of a pre-trained variational autoencoder. By decoding intermediate latent states, MA-GIG biases the path toward the learned generative manifold and reduces exposure to implausible inputspace regions. Through qualitative and quantitative evaluations, we demonstrate that MA-GIG produces faithful explanations by aggregating gradients on path features proximal to the input. Consequently, our method reduces offmanifold noise and outperforms prior path-based attribution methods across multiple datasets and classifiers. Our code is available at https: //github.com/leekwoon/ma-gig/ † Equal advising.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng 等NeurIPS 2024 · 被引用 758 次
- Reliable Fidelity and Diversity Metrics for Generative ModelsMuhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi 等ICML 2020 · 被引用 553 次
- Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior KnowledgeLaura Rieger, Chandan Singh, W. James Murdoch, Bin YuICML 2020 · 被引用 249 次
相关 Paper
- Manifold Integrated Gradients: Riemannian Geometry for Feature AttributionEslam Zaher, Maciej Trzaskowski, Quan Nguyen, Fred RoostaICML 2024 · 被引用 13 次
- Beyond Single Path Integrated Gradients for Reliable Input Attribution via Randomized Path SamplingGiyoung Jeon, Haedong Jeong, Jaesik ChoiICCV 2023 · 被引用 3 次
- Guided Integrated Gradients: An Adaptive Path Method for Removing NoiseAndrei Kapishnikov, Subhashini Venugopalan, Besim Avci, Ben Wedin 等CVPR 2021
- Integrated Decision Gradients: Compute Your Attributions Where the Model Makes Its DecisionChase Walker, Sumit Kumar Jha, Kenny Chen, Rickard EwetzAAAI 2024 · 被引用 25 次
- Distilled Gradient Aggregation: Purify Features for Input Attribution in the Deep Neural NetworkGiyoung Jeon, Haedong Jeong, Jaesik ChoiNeurIPS 2022 · 被引用 11 次
