Are Data-Driven Explanations Robust Against Out-of-Distribution Data?
Tang Li, Fengchun Qiao, Mengmeng Ma, Xi Peng
Abstract
As black-box models increasingly power high-stakes applications, a variety of data-driven explanation methods have been introduced. Meanwhile, machine learning models are constantly challenged by distributional shifts. A question naturally arises: Are data-driven explanations robust against out-of-distribution data? Our empirical results show that even though predict correctly, the model might still yield unreliable explanations under distributional shifts. How to develop robust explanations against out-of-distribution data? To address this problem, we propose an end-to-end model-agnostic learning framework Distributionally Robust Explanations (DRE). The key idea is, inspired by self-supervised learning, to fully utilizes the inter-distribution information to provide supervisory signals for the learning of explanations without human annotation. Can robust explanations benefit the model's generalization capability? We conduct extensive experiments on a wide range of tasks and data types, including classification and regression on image and scientific tabular data. Our results demonstrate that the proposed method significantly improves the model's performance in terms of explanation and prediction robustness against distributional shifts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language ModelsLu Yu, Haiyang Zhang, Changsheng XuNeurIPS 2024 · 29 citations
- Beyond the Federation: Topology-aware Federated Learning for Generalization to Unseen ClientsMengmeng Ma, Tang Li, Xi PengICML 2024 · 7 citations
- Training for Stable Explanation for FreeChao Chen, Chenghua Guo, Rufeng Chen, Guixiang Ma et al.NeurIPS 2024 · 7 citations
- Beyond Accuracy: Ensuring Correct Predictions With Correct RationalesTang Li, Mengmeng Ma, Xi PengNeurIPS 2024 · 6 citations
- Ensemble Pruning for Out-of-distribution GeneralizationFengchun Qiao, Xi PengICML 2024 · 3 citations
Builds on17
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Adversarial Domain Adaptation with Domain MixupMinghao Xu, Jian Zhang, Bingbing Ni, Teng Li et al.AAAI 2020 · 499 citations
Related papers
- Model-Agnostic Random Weighting for Out-of-Distribution GeneralizationYue He, Pengfei Tian, Renzhe Xu, Xinwei Shen et al.KDD 2024 · 1 citation
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 93 citations
- Distribution-Based Feature Attribution for Explaining the Predictions of Any ClassifierXinpeng Li, Kai Ming TingAAAI 2026
- LIMEFLDL: A Local Interpretable Model-Agnostic Explanations Approach for Label Distribution LearningXiuyi Jia, Jinchi Li, Yunan Lu, Weiwei LiICML 2025
- TimeSAE: Causal Sparse Decoding for Faithful Explanations of Black-Box Time Series ModelsKhalid Oublal, Quentin Bouniot, Qi Gan, Stephan Clemencon et al.ICML 2026 · 2 citations
