Distributionally Robust Feature Selection
Maitreyi Swaroop, Tamar Krishnamurti, Bryan Wilder
Abstract
We study the problem of selecting limited features to observe such that models trained on them can perform well simultaneously across multiple subpopulations. This problem has applications in settings where collecting each feature is costly, e.g. requiring adding survey questions or physical sensors, and we must be able to use the selected features to create high-quality downstream models for different populations. Our method frames the problem as a continuous relaxation of traditional variable selection using a noising mechanism, without requiring backpropagation through model training processes. By optimizing over the variance of a Bayes-optimal predictor, we develop a model-agnostic framework that balances overall performance of downstream prediction across populations. We validate our approach through experiments on both synthetic datasets and real-world data. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e96bb77-3ba3-4d70-b1cb-3f46eafd159aBuilds on3
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Locally Sparse Neural Networks for Tabular Biomedical DataJunchen Yang, Ofir Lindenbaum, Yuval KlugerICML 2022 · 45 citations
- Interpretable Deep Clustering for Tabular DataJonathan Svirsky, Ofir LindenbaumICML 2024 · 19 citations
Related papers
- Representation Matters: Assessing the Importance of Subgroup Allocations in Training DataEsther Rolf, Theodora T. Worledge, Benjamin Recht, Michael I. JordanICML 2021 · 50 citations
- Change is Hard: A Closer Look at Subpopulation ShiftYuzhe Yang, Haoran Zhang, Dina Katabi, Marzyeh GhassemiICML 2023 · 149 citations
- A Practical Upper Bound on Selection Bias Effects in Medical Prediction ModelsKara Liu, Maggie Wang, Russ B. AltmanKDD 2026
- Multi-task learning with summary statisticsParker Knight, Rui DuanNeurIPS 2023 · 17 citations
- Fairness with Overlapping Groups; a Probabilistic PerspectiveForest Yang, Mouhamadou Cisse, Oluwasanmi KoyejoNeurIPS 2020 · 71 citations
