Explaining Concept Shift with Interpretable Feature Attribution
Ruiqi Lyu, Alistair Turcan, Bryan Wilder
摘要
Concept shift occurs when the distribution of labels conditioned on the features changes between domains, which can make even a well-tuned ML model miscalibrated on a new domain. Identifying these shifted features provides unique insight into how feature-label relationships differ between domains, considering the difference may be across a scientifically relevant dimension, such as time, disease status, population, etc. In this paper, we propose SGShift, a method for attributing performance degradation under concept shift in tabular data to a sparse set of shifted features. We frame concept shift as a feature selection task to learn the features that can explain performance differences between models in the source and target domain. This framework enables SGShift to adapt powerful statistical tools such as generalized additive models, knockoffs, and absorption towards identifying these shifted features. We conduct extensive experiments in synthetic and real data across various ML models and find SGShift can identify shifted features much more accurately than baseline methods, requires few samples in the shifted domain, and is robust to complex cases of concept shift.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Feature Shift Detection: Localizing Which Features Have Shifted via Conditional Distribution TestsSean Kulinski, Saurabh Bagchi, David I. InouyeNeurIPS 2020 · 被引用 39 次
- Towards Explaining Distribution ShiftsSean Kulinski, David I. InouyeICML 2023 · 被引用 38 次
- "Why did the Model Fail?": Attributing Model Performance Changes to Distribution ShiftsHaoran Zhang, Harvineet Singh, Marzyeh Ghassemi, Shalmali JoshiICML 2023 · 被引用 37 次
- Estimating and Explaining Model Performance When Both Covariates and Labels ShiftLingjiao Chen, Matei Zaharia, James Y. ZouNeurIPS 2022 · 被引用 34 次
- Sequential Covariate Shift Detection Using Classifier Two-Sample TestsSooyong Jang, Sangdon Park, Insup Lee, Osbert BastaniICML 2022 · 被引用 24 次
相关 Paper
- Feature-aware Modulation for Learning from Temporal Tabular DataHaorun Cai, Han-Jia YeNeurIPS 2025 · 被引用 3 次
- TabFSBench: Tabular Benchmark for Feature Shifts in Open EnvironmentsZi-Jian Cheng, Ziyi Jia, Zhi Zhou, Yufeng Li 等ICML 2025
- "Who experiences large model decay and why?" A Hierarchical Framework for Diagnosing Heterogeneous Performance DriftHarvineet Singh, Fan Xia, Alexej Gossmann, Andrew Chuang 等ICML 2025
- Generative multitask learning mitigates target-causing confoundingTaro Makino, Krzysztof J. Geras, Kyunghyun ChoNeurIPS 2022 · 被引用 9 次
- Do causal predictors generalize better to new domains?Vivian Y. Nastl, Moritz HardtNeurIPS 2024 · 被引用 21 次
