Debiased Recommendation Beyond the Positive Propensity Assumption
Yanghao Xiao, Hao Wang, Xiang Li, Qian Zou, Cheng Bing, Wei Lin, Haoxuan Li, Zhouchen Lin
Abstract
Post-click conversion rate (CVR) prediction is a central task in recommender systems, yet selection bias creates a severe distributional gap between the clicked training samples and the entire inference space. To address selection bias, propensity-based methods such as inverse propensity scoring (IPS) and doubly robust (DR) have been adopted, which aim to estimate the unbiased learning objective from biased training samples. However, these approaches assume strictly positive propensities, implying every user-item pair has a nonzero probability of interaction. In practice, such positivity assumption maybe violated, for example, in food-delivery platforms, some restaurants located more than 10 kilometers away will be blocked for recommendation. In this study, we theoretically show that when such zero-propensity samples, termed extrapolation samples exist, both IPS and DR estimators become biased. To overcome this limitation, we propose ExtraDebias method, which enables debiased recommendation in both non-extrapolation and extrapolation samples. Specifically, we first train a propensity model to identify extrapolation samples with extremely small propensity estimates, then estimate their pseudo-label intervals, and derive an upper bound of the learning objective for extrapolation samples. By minimizing the derived upper bound, debiased learning on extrapolation samples is ensured, while unbiased learning on non-extrapolation samples is achieved by standard IPS. Experiments on four real-world offline datasets and one online A/B test show that ExtraDebias effectively minimizes prediction errors on extrapolation samples and achieves optimal performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 522a3b82-c98b-41c3-bdf5-89a594be2c19Cited by top-tier papers1
Ask how each one uses itBuilds on31
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain et al.WWW 2021 · 793 citations
- AutoDebias: Learning to Debias for RecommendationJiawei Chen, Hande Dong, Yang Qiu, Xiangnan He et al.SIGIR 2021 · 167 citations
- Asymmetric Tri-training for Debiasing Missing-Not-At-Random Explicit FeedbackYuta SaitoSIGIR 2020 · 90 citations
- ESCM2: Entire Space Counterfactual Multi-Task Model for Post-Click Conversion Rate EstimationHao Wang, Tai-Wei Chang, Tianqiao Liu, Jianmin Huang et al.SIGIR 2022 · 86 citations
- Personalized Adaptive Meta Learning for Cold-start User Preference PredictionRunsheng Yu, Yu Gong, Xu He, Yu Zhu et al.AAAI 2021 · 72 citations
Related papers
- A Generalized Doubly Robust Learning Framework for Debiasing Post-Click Conversion Rate PredictionQuanyu Dai, Haoxuan Li, Peng Wu, Zhenhua Dong et al.KDD 2022 · 45 citations
- Enhanced Doubly Robust Learning for Debiasing Post-Click Conversion Rate EstimationSiyuan Guo, Lixin Zou, Yiding Liu, Wenwen Ye et al.SIGIR 2021 · 63 citations
- Uncovering the Propensity Identification Problem in Debiased RecommendationsHonglei Zhang, Shuyi Wang, Haoxuan Li, Chunyuan Zheng et al.ICDE 2024 · 12 citations
- Adversarial-Enhanced Causal Multi-Task Framework for Debiasing Post-Click Conversion Rate EstimationXinyue Zhang, Cong Huang, Kun Zheng, Hongzu Su et al.WWW 2024 · 8 citations
- StableDR: Stabilized Doubly Robust Learning for Recommendation on Data Missing Not at RandomHaoxuan Li, Chunyuan Zheng, Peng WuICLR 2023 · 12 citations
