Identifying Selection Bias from Observational Data
David Kaltenpoth, Jilles Vreeken
Abstract
Access to a representative sample from the population is an assumption that underpins all of machine learning. Unfortunately, selection effects can cause observations to instead come from a subpopulation, by which our inferences may be subject to bias. It is therefore essential to know whether or not a sample is affected by selection effects. We study under which conditions we can identify selection bias and give results for both parametric and non-parametric families of distributions. Based on these results, we develop two practical methods to determine whether or not an observed sample comes from a distribution subject to selection bias. Through extensive evaluation on synthetic and real-world data, we verify that our methods beat the state of the art both in detecting as well as characterizing selection bias.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b3862ad-43f4-4dd5-926f-0d3dec645a22Cited by top-tier papers7
- Causal Discovery from Event Sequences by Local Cause-Effect AttributionJoscha Cüppers, Sascha Xu, Ahmed Musa, Jilles VreekenNeurIPS 2024 · 13 citations
- Gene Regulatory Network Inference in the Presence of Selection Bias and Latent ConfoundersGongxu Luo, Haoyue Dai, Longkang Li, Chengqian Gao et al.NeurIPS 2025 · 9 citations
- Detecting and Identifying Selection Structure in Sequential DataYujia Zheng, Zeyu Tang, Yiwen Qiu, Bernhard Schölkopf et al.ICML 2024 · 7 citations
- Causal Modeling of Selection in EvolutionHaoyue Dai, Zeyu Tang, Peter Spirtes, Kun ZhangICML 2026
- Prompting Fairness: Integrating Causality to Debias Large Language ModelsJingling Li, Zeyu Tang, Xiaoyu Liu, Peter Spirtes et al.ICLR 2025
Builds on3
- A Group-Theoretic Framework for Data AugmentationShuxiao Chen, Edgar Dobriban, Jane H. LeeNeurIPS 2020 · 254 citations
- Learning Invariances in Neural Networks from Training DataGregory W. Benton, Marc Finzi, Pavel Izmailov, Andrew Gordon WilsonNeurIPS 2020 · 78 citations
- Invariance Learning in Deep Neural Networks with Differentiable Laplace ApproximationsAlexander Immer, Tycho F. A. van der Ouderaa, Gunnar Rätsch, Vincent Fortuin et al.NeurIPS 2022 · 56 citations
Related papers
- A Practical Upper Bound on Selection Bias Effects in Medical Prediction ModelsKara Liu, Maggie Wang, Russ B. AltmanKDD 2026
- s-ID: Causal Effect Identification in a Sub-populationAmir Mohammad Abouei, Ehsan Mokhtarian, Negar KiyavashAAAI 2024 · 4 citations
- Towards a Holistic Understanding of Selection Bias for Causal Effect IdentificationYiwen (Evie) Qiu, Filip Kovačević, Shimeng Huang, Peter Spirtes et al.ICML 2026
- Recovering the Propensity Score from Biased Positive Unlabeled DataWalter Gerych, Thomas Hartvigsen, Luke Buquicchio, Emmanuel Agu et al.AAAI 2022 · 20 citations
- Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairnessStephen Pfohl, Natalie Harris, Chirag Nagpal, David Madras et al.NeurIPS 2025 · 9 citations
