Adversarial Learning for Feature Shift Detection and Correction
Míriam Barrabés, Daniel Mas Montserrat, Margarita Geleta, Xavier Giró-i-Nieto, Alexander G. Ioannidis
Abstract
Data shift is a phenomenon present in many real-world applications, and while there are multiple methods attempting to detect shifts, the task of localizing and correcting the features originating such shifts has not been studied in depth. Feature shifts can occur in many datasets, including in multi-sensor data, where some sensors are malfunctioning, or in tabular and structured data, including biomedical, financial, and survey data, where faulty standardization and data processing pipelines can lead to erroneous features. In this work, we explore using the principles of adversarial learning, where the information from several discriminators trained to distinguish between two distributions is used to both detect the corrupted features and fix them in order to remove the distribution shift between datasets. We show that mainstream supervised classifiers, such as random forest or gradient boosting trees, combined with simple iterative heuristics, can localize and correct feature shifts, outperforming current statistical and neural network-based techniques. The code is available at https://github.com/AI-sandbox/DataFix .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Missing Data Imputation using Optimal TransportBoris Muzellec, Julie Josse, Claire Boyer, Marco CuturiICML 2020 · 179 citations
- HyperImpute: Generalized Iterative Imputation with Automatic Model SelectionDaniel Jarrett, Bogdan Cebere, Tennison Liu, Alicia Curth et al.ICML 2022 · 129 citations
- Dataset Discovery in Data LakesAlex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos KonstantinouICDE 2020 · 118 citations
- MIRACLE: Causally-Aware Imputation via Learning Missing Data MechanismsTrent Kyono, Yao Zhang, Alexis Bellot, Mihaela van der SchaarNeurIPS 2021 · 105 citations
- Deep Weakly-supervised Anomaly DetectionGuansong Pang, Chunhua Shen, Huidong Jin, Anton van den HengelKDD 2023 · 100 citations
Related papers
- Encoding Robustness to Image Style via Adversarial Feature PerturbationsManli Shu, Zuxuan Wu, Micah Goldblum, Tom GoldsteinNeurIPS 2021 · 23 citations
- An Information-theoretic Approach to Distribution ShiftsMarco Federici, Ryota Tomioka, Patrick ForréNeurIPS 2021 · 30 citations
- Extending the WILDS Benchmark for Unsupervised AdaptationShiori Sagawa, Pang Wei Koh, Tony Lee, Irena Gao et al.ICLR 2022 · 116 citations
- Domain Adaptation with Conditional Distribution Matching and Generalized Label ShiftRemi Tachet des Combes, Han Zhao, Yu-Xiang Wang, Geoffrey J. GordonNeurIPS 2020 · 231 citations
- Gradient Distribution Alignment Certificates Better Adversarial Domain AdaptationZhiqiang Gao, Shufei Zhang, Kaizhu Huang, Qiufeng Wang et al.ICCV 2021 · 56 citations
