Unveiling Extraneous Sampling Bias with Data Missing-Not-At-Random
Chunyuan Zheng, Haocheng Yang, Haoxuan Li, Mengyue Yang
Abstract
Selection bias poses a widely recognized challenge for unbiased evaluation and learning in many industrial scenarios. For example, in recommender systems, it arises from the users' selective interactions with items. Recently, doubly robust and its variants have been widely studied to achieve debiased learning of prediction models, however, all of them consider a simple exact matching scenario, i.e., the units (such as user-item pairs in a recommender system) are the same between the training and test sets. In practice, there may be limited or even no overlap in units between the training and test. In this paper, we consider a more practical scenario: the joint distribution of the feature and rating is the same in the training and test sets. Theoretical analysis shows that the previous DR estimator is biased even if the imputed errors and learned propensities are correct in this scenario. In addition, we propose a novel super-population doubly robust estimator (SuperDR), which can achieve a more accurate estimation and desirable generalization error bound compared to the existing DR estimators, and extend the joint learning algorithm for training the prediction and imputation models. We conduct extensive experiments on three real-world datasets, including a large-scale industrial dataset, to show the effectiveness of our method. The code is available at https://github.com/ChunyuanZheng/neurips-25-SuperDR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Counterfactual Implicit Feedback ModelingChuan Zhou, Lina Yao, Haoxuan Li, Mingming GongNeurIPS 2025 · 8 citations
- Unbiased Reward Modeling from Implicit Feedback for LLM AlignmentHao Wang, Haocheng Yang, Licheng Pan, Zhichao Chen et al.ICML 2026 · 2 citations
- Uplift Modeling with Delayed Feedback: Identifiability and AlgorithmsChunyuan Zheng, Anpeng Wu, Chuan Zhou, Taojun Hu et al.AAAI 2026 · 2 citations
- Optimizing Marketing Subsidies via Counterfactual Learning with Asymmetric Reward FunctionXiang Li, Yanghao Xiao, Chunyuan Zheng, Qian Zou et al.SIGIR 2026 · 1 citation
- Debiased Recommendation Beyond the Positive Propensity AssumptionYanghao Xiao, Hao Wang, Xiang Li, Qian Zou et al.SIGIR 2026
Builds on31
- Causal Intervention for Leveraging Popularity Bias in RecommendationYang Zhang, Fuli Feng, Xiangnan He, Tianxin Wei et al.SIGIR 2021 · 431 citations
- A General Knowledge Distillation Framework for Counterfactual Recommendation via Uniform DataDugang Liu, Pengxiang Cheng, Zhenhua Dong, Xiuqiang He et al.SIGIR 2020 · 188 citations
- AutoDebias: Learning to Debias for RecommendationJiawei Chen, Hande Dong, Yang Qiu, Xiangnan He et al.SIGIR 2021 · 167 citations
- Information Theoretic Counterfactual Learning from Missing-Not-At-Random FeedbackZifeng Wang, Xi Chen, Rui Wen, Shao-Lun Huang et al.NeurIPS 2020 · 95 citations
- Asymmetric Tri-training for Debiasing Missing-Not-At-Random Explicit FeedbackYuta SaitoSIGIR 2020 · 90 citations
Related papers
- Doubly Calibrated Estimator for Recommendation on Data Missing Not at RandomWonbin Kweon, Hwanjo YuWWW 2024 · 23 citations
- Multiple Robust Learning for RecommendationHaoxuan Li, Quanyu Dai, Yuru Li, Yan Lyu et al.AAAI 2023 · 48 citations
- Relaxing the Accurate Imputation Assumption in Doubly Robust Learning for Debiased Collaborative FilteringHaoxuan Li, Chunyuan Zheng, Shuyi Wang, Kunhan Wu et al.ICML 2024 · 25 citations
- StableDR: Stabilized Doubly Robust Learning for Recommendation on Data Missing Not at RandomHaoxuan Li, Chunyuan Zheng, Peng WuICLR 2023 · 12 citations
- TDR-CL: Targeted Doubly Robust Collaborative Learning for Debiased RecommendationsHaoxuan Li, Yan Lyu, Chunyuan Zheng, Peng WuICLR 2023 · 14 citations
