Identifying Selection Bias from Observational Data
David Kaltenpoth, Jilles Vreeken
摘要
Access to a representative sample from the population is an assumption that underpins all of machine learning. Unfortunately, selection effects can cause observations to instead come from a subpopulation, by which our inferences may be subject to bias. It is therefore essential to know whether or not a sample is affected by selection effects. We study under which conditions we can identify selection bias and give results for both parametric and non-parametric families of distributions. Based on these results, we develop two practical methods to determine whether or not an observed sample comes from a distribution subject to selection bias. Through extensive evaluation on synthetic and real-world data, we verify that our methods beat the state of the art both in detecting as well as characterizing selection bias.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Causal Discovery from Event Sequences by Local Cause-Effect AttributionJoscha Cüppers, Sascha Xu, Ahmed Musa, Jilles VreekenNeurIPS 2024 · 被引用 13 次
- Gene Regulatory Network Inference in the Presence of Selection Bias and Latent ConfoundersGongxu Luo, Haoyue Dai, Longkang Li, Chengqian Gao 等NeurIPS 2025 · 被引用 9 次
- Detecting and Identifying Selection Structure in Sequential DataYujia Zheng, Zeyu Tang, Yiwen Qiu, Bernhard Schölkopf 等ICML 2024 · 被引用 7 次
- Causal Modeling of Selection in EvolutionHaoyue Dai, Zeyu Tang, Peter Spirtes, Kun ZhangICML 2026
- Prompting Fairness: Integrating Causality to Debias Large Language ModelsJingling Li, Zeyu Tang, Xiaoyu Liu, Peter Spirtes 等ICLR 2025
它引用的顶会 Paper3
- A Group-Theoretic Framework for Data AugmentationShuxiao Chen, Edgar Dobriban, Jane H. LeeNeurIPS 2020 · 被引用 254 次
- Learning Invariances in Neural Networks from Training DataGregory W. Benton, Marc Finzi, Pavel Izmailov, Andrew Gordon WilsonNeurIPS 2020 · 被引用 78 次
- Invariance Learning in Deep Neural Networks with Differentiable Laplace ApproximationsAlexander Immer, Tycho F. A. van der Ouderaa, Gunnar Rätsch, Vincent Fortuin 等NeurIPS 2022 · 被引用 56 次
相关 Paper
- A Practical Upper Bound on Selection Bias Effects in Medical Prediction ModelsKara Liu, Maggie Wang, Russ B. AltmanKDD 2026
- s-ID: Causal Effect Identification in a Sub-populationAmir Mohammad Abouei, Ehsan Mokhtarian, Negar KiyavashAAAI 2024 · 被引用 4 次
- Towards a Holistic Understanding of Selection Bias for Causal Effect IdentificationYiwen (Evie) Qiu, Filip Kovačević, Shimeng Huang, Peter Spirtes 等ICML 2026
- Recovering the Propensity Score from Biased Positive Unlabeled DataWalter Gerych, Thomas Hartvigsen, Luke Buquicchio, Emmanuel Agu 等AAAI 2022 · 被引用 20 次
- Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairnessStephen Pfohl, Natalie Harris, Chirag Nagpal, David Madras 等NeurIPS 2025 · 被引用 9 次
