Understanding the Black Box: A Deep Empirical Dive into Shapley Value Approximations for Tabular Data
Suchit Gupte, John Paparrizos
摘要
Understanding the decisions made by machine learning models is significant for building trust and enabling the adoption of these models in real-world applications. Shapley values have emerged as a leading method for model interpretability, offering precise insights by quantifying each feature's contribution to predictions. However, computing Shapley values requires exploring all possible combinations of features, which can be computationally expensive, especially for high-dimensional data. This challenge has led to the development of various approximation techniques, often composed of estimation and replacement strategies, to compute the Shapley values efficiently. Our study focuses on the interpretability of machine learning models for tabular datasets, one of the most common and widely used data type. However, the abundance of options has created a substantial gap in determining the most appropriate technique for practical applications. Through this study, we seek to bridge this gap by comprehensively evaluating Shapley value approximations, covering 8 replacement and 17 estimation strategies across diverse regression and classification tasks. The evaluation is conducted exclusively on tabular data, leveraging 200 synthetic and real-world datasets, covering a wide range of model types, from conventional tree-based and linear models to modern neural networks. We focus on computational efficiency and the consistency of Shapley value estimates in handling high-dimensional feature spaces. Our findings reveal that traditional sampling-based approaches significantly reduce computational costs but fail to capture complex feature interactions. On the contrary, model-specific approaches that exploit the structure of the underlying model consistently outperform model-agnostic techniques, delivering higher accuracy and faster computations. Through the study, we aim to encourage further research on Shapley value approximations, advancing data-centric explainable AI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Structured Study of Multivariate Time-Series Distance MeasuresJens E. d'Hondt, Haojun Li, Fan Yang, Odysseas Papapetrou 等SIGMOD 2025 · 被引用 14 次
- SPARTAN: Data-Adaptive Symbolic Time-Series ApproximationFan Yang, John PaparrizosSIGMOD 2025 · 被引用 11 次
- The Power of Anomaly Detection in Predictive Maintenance: [Experiments & Analysis]Anastasios Papadopoulos, Apostolos Giannoulidis, Anastasios Gounaris, John PaparrizosSIGMOD 2026 · 被引用 4 次
- HYDRA: A Multi-Level Hierarchy-Driven Approach for Robust Anomaly Detection in Time SeriesMingyi Huang, Qinghua Liu, Paul Boniol, John PaparrizosSIGMOD 2026 · 被引用 4 次
它引用的顶会 Paper33
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 被引用 799 次
- On Explainability of Graph Neural Networks via Subgraph ExplorationsHao Yuan, Haiyang Yu, Jie Wang, Kang Li 等ICML 2021 · 被引用 498 次
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
- FastSHAP: Real-Time Shapley Value EstimationNeil Jethani, Mukund Sudarshan, Ian Connick Covert, Su-In Lee 等ICLR 2022 · 被引用 186 次
- Volume Under the Surface: A New Accuracy Evaluation Measure for Time-Series Anomaly DetectionJohn Paparrizos, Paul Boniol, Themis Palpanas, Ruey S. Tsay 等VLDB 2022 · 被引用 171 次
相关 Paper
- CoShap: A Scalable Coalition Growth Approach to Shapley Value ApproximationJingxuan He, Changshuo Liu, Shaofeng Cai, Xixian Han 等SIGMOD 2026
- Provably Accurate Shapley Value Estimation via Leverage Score SamplingChristopher Musco, R. Teal WitterICLR 2025
- Regression-adjusted Monte Carlo Estimators for Shapley Values and Probabilistic ValuesR. Teal Witter, Yurong Liu, Christopher MuscoNeurIPS 2025 · 被引用 22 次
- Efficient Shapley Values Estimation by Amortization for Text ClassificationChenghao Yang, Fan Yin, He He, Kai-Wei Chang 等ACL 2023 · 被引用 2 次
- Prediction via Shapley Value RegressionAmr Alkhatib, Roman Bresson, Henrik Boström, Michalis VazirgiannisICML 2025
