Policy-Gradient Training of Fair and Unbiased Ranking Functions
Himank Yadav, Zhengxiao Du, Thorsten Joachims
Abstract
While implicit feedback (e.g., clicks, dwell times, etc.) is an abundant and attractive source of data for learning to rank, it can produce unfair ranking policies for both exogenous and endogenous reasons. Exogenous reasons typically manifest themselves as biases in the training data, which then get reflected in the learned ranking policy and often lead to rich-get-richer dynamics. Moreover, even after the correction of such biases, reasons endogenous to the design of the learning algorithm can still lead to ranking policies that do not allocate exposure among items in a fair way. To address both exogenous and endogenous sources of unfairness, we present the first learning-to-rank approach that addresses both presentation bias and merit-based fairness of exposure simultaneously. Specifically, we define a class of amortized fairness-of-exposure constraints that can be chosen based on the needs of an application, and we show how these fairness criteria can be enforced despite the selection biases in implicit feedback data. The key result is an efficient and flexible policy-gradient algorithm, called FULTR, which is the first to enable the use of counterfactual estimators for both utility estimation and fairness constraints. Beyond the theoretical justification of the framework, we show empirically that the proposed algorithm can learn accurate and fair ranking policies from biased and noisy feedback. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5315d16-9839-4aa3-91d6-ae3af20a989eCited by top-tier papers12
- LLM4Rerank: LLM-based Auto-Reranking Framework for RecommendationsJingtong Gao, Bo Chen, Xiangyu Zhao, Weiwen Liu et al.WWW 2025 · 50 citations
- Joint Multisided Exposure Fairness for RecommendationHaolun Wu, Bhaskar Mitra, Chen Ma, Fernando Diaz et al.SIGIR 2022 · 48 citations
- Fairness of Exposure in Light of Incomplete Exposure EstimationMaria Heuss, Fatemeh Sarvi, Maarten de RijkeSIGIR 2022 · 21 citations
- Safe Deployment for Counterfactual Learning to Rank with Exposure-Based Risk MinimizationShashank Gupta, Harrie Oosterhuis, Maarten de RijkeSIGIR 2023 · 17 citations
- Probabilistic Permutation Graph Search: Black-Box Optimization for Fairness in RankingAli Vardasbi, Fatemeh Sarvi, Maarten de RijkeSIGIR 2022 · 9 citations
Builds on3
- Controlling Fairness and Bias in Dynamic Learning-to-RankMarco Morik, Ashudeep Singh, Jessica Hong, Thorsten JoachimsSIGIR 2020 · 205 citations
- Achieving Fairness in the Stochastic Multi-Armed Bandit ProblemVishakha Patil, Ganesh Ghalme, Vineet Nair, Y. NarahariAAAI 2020 · 131 citations
- Pairwise Fairness for Ranking and RegressionHarikrishna Narasimhan, Andrew Cotter, Maya R. Gupta, Serena Lutong WangAAAI 2020 · 125 citations
Related papers
- Optimizing Learning-to-Rank Models for Ex-Post Fair RelevanceSruthi Gorantla, Eshaan Bhansali, Amit Deshpande, Anand LouisSIGIR 2024 · 1 citation
- Individually Fair RankingsAmanda Bower, Hamid Eftekhari, Mikhail Yurochkin, Yuekai SunICLR 2021 · 4 citations
- Can Clicks Be Both Labels and Features?: Unbiased Behavior Feature Collection and Uncertainty-aware Learning to RankTao Yang, Chen Luo, Hanqing Lu, Parth Gupta et al.SIGIR 2022 · 23 citations
- Correcting for Selection Bias in Learning-to-rank SystemsZohreh Ovaisi, Ragib Ahsan, Yifan Zhang, Kathryn Vasilaky et al.WWW 2020 · 123 citations
- FairGAN: GANs-based Fairness-aware Learning for Recommendations with Implicit FeedbackJie Li, Yongli Ren, Ke DengWWW 2022 · 62 citations
