Leveraging Feature Bias for Scalable Misprediction Explanation of Machine Learning Models
Jiri Gesi, Xinyun Shen, Yunfan Geng, Qihong Chen, Iftekhar Ahmed
摘要
Interpreting and debugging machine learning models is necessary to ensure the robustness of the machine learning models. Explaining mispredictions can help significantly in doing so. While recent works on misprediction explanation have proven promising in generating interpretable explanations for mispredictions, the state-of-the-art techniques "blindly" deduce misprediction explanation rules from all data features, which may not be scalable depending on the number of features. To alleviate this problem, we propose an efficient misprediction explanation technique named Bias Guided Misprediction Diagnoser (BGMD), which leverages two prior knowledge about data: a) data often exhibit highly-skewed feature distributions and b) trained models in many cases perform poorly on subdataset with under-represented features. Next, we propose a technique named MAPS (Mispredicted Area UPweight Sampling). MAPS increases the weights of subdataset during model retraining that belong to the group that is prone to be mispredicted because of containing under-represented features. Thus, MAPS make retrained model pay more attention to the under-represented features. Our empirical study shows that our proposed BGMD outperformed the state-of-the-art misprediction diagnoser and reduces diagnosis time by 92%. Furthermore, MAPS outperformed two state-of-the-art techniques on fixing the machine learning model's performance on mispredicted data without compromising performance on all data. All the research artifacts (i.e., tools, scripts, and data) of this study are available in the accompanying website [1] .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in DeploymentShibbir Ahmed, Hongyang Gao, Hridesh RajanICSE 2024 · 被引用 3 次
- Efficient Understanding of Machine Learning Model MispredictionsMartin Eberlein, Jürgen Cito, Lars GrunskeASE 2025
它引用的顶会 Paper15
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan 等ICML 2021 · 被引用 683 次
- Environment Inference for Invariant LearningElliot Creager, Jörn-Henrik Jacobsen, Richard S. ZemelICML 2021 · 被引用 454 次
- Bias in machine learning software: why? how? what to do?Joymallya Chakraborty, Suvodeep Majumder, Tim MenziesFSE 2021 · 被引用 186 次
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 被引用 169 次
- Model Patching: Closing the Subgroup Performance Gap with Data AugmentationKaran Goel, Albert Gu, Sharon Li, Christopher RéICLR 2021 · 被引用 131 次
相关 Paper
- Explaining mispredictions of machine learning models using rule inductionJürgen Cito, Isil Dillig, Seohyun Kim, Vijayaraghavan Murali 等FSE 2021 · 被引用 26 次
- Explanation-based Data Augmentation for Image ClassificationSandareka Wickramanayake, Wynne Hsu, Mong-Li LeeNeurIPS 2021 · 被引用 24 次
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
- Meaningfully debugging model mistakes using conceptual counterfactual explanationsAbubakar Abid, Mert Yüksekgönül, James ZouICML 2022 · 被引用 75 次
- Explaining Hyperparameter Optimization via Partial Dependence PlotsJulia Moosbauer, Julia Herbinger, Giuseppe Casalicchio, Marius Lindauer 等NeurIPS 2021 · 被引用 109 次
