Stochastic Amortization: A Unified Approach to Accelerate Feature and Data Attribution
Ian Covert, Chanwoo Kim, Su-In Lee, James Y. Zou, Tatsunori B. Hashimoto
Abstract
Many tasks in explainable machine learning, such as data valuation and feature attribution, perform expensive computation for each data point and are intractable for large datasets. These methods require efficient approximations, and although amortizing the process by learning a network to directly predict the desired output is a promising solution, training such models with exact labels is often infeasible. We therefore explore training amortized models with noisy labels, and we find that this is inexpensive and surprisingly effective. Through theoretical analysis of the label noise and experiments with various models and datasets, we show that this approach tolerates high noise levels and significantly accelerates several feature attribution and data valuation methods, often yielding an order of magnitude speedup over existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Enhancing Training Data Attribution with Representational OptimizationWeiwei Sun, Haokun Liu, Nikhil Kandpal, Colin A. Raffel et al.NeurIPS 2025 · 9 citations
- SHAP zero Explains Biological Sequence Models with Near-zero Marginal Cost for Future QueriesDarin Tsui, Aryan Musharaf, Yigit Efe Erginbas, Justin Singh Kang et al.NeurIPS 2025 · 5 citations
- On-device Content-based Recommendation with Single-shot Embedding Pruning: A Cooperative Game PerspectiveHung Vinh Tran, Tong Chen, Guanhua Ye, Quoc Viet Hung Nguyen et al.WWW 2025 · 4 citations
- Selective ExplanationsLucas Monteiro Paes, Dennis Wei, Flávio P. CalmonNeurIPS 2024 · 4 citations
- MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and TransformationHaonan Yu, Junhao Liu, Xin ZhangICML 2026 · 2 citations
Builds on27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 199 citations
- FastSHAP: Real-Time Shapley Value EstimationNeil Jethani, Mukund Sudarshan, Ian Connick Covert, Su-In Lee et al.ICLR 2022 · 186 citations
Related papers
- Efficient Shapley Values Estimation by Amortization for Text ClassificationChenghao Yang, Fan Yin, He He, Kai-Wei Chang et al.ACL 2023 · 2 citations
- Transfer and Marginalize: Explaining Away Label Noise with Privileged InformationMark Collier, Rodolphe Jenatton, Effrosyni Kokiopoulou, Jesse BerentICML 2022 · 19 citations
- SAVA: Scalable Learning-Agnostic Data ValuationSamuel Kessler, Tam Le, Vu NguyenICLR 2025
- EcoVal: An Efficient Data Valuation Framework for Machine LearningAyush K. Tarun, Vikram S. Chundawat, Murari Mandal, Hong Ming Tan et al.KDD 2024 · 3 citations
- Amortized Variational Inference for Partial-Label Learning: A Probabilistic Approach to Label DisambiguationTobias Fuchs, Nadja KleinICML 2026
