Learning from eXtreme Bandit Feedback
Romain Lopez, Inderjit S. Dhillon, Michael I. Jordan
摘要
We study the problem of batch learning from bandit feedback in the setting of extremely large action spaces. Learning from extreme bandit feedback is ubiquitous in recommendation systems, in which billions of decisions are made over sets consisting of millions of choices in a single day, yielding massive observational data. In these large-scale real-world applications, supervised learning frameworks such as eXtreme Multi-label Classification (XMC) are widely used despite the fact that they incur significant biases due to the mismatch between bandit feedback and supervised labels. Such biases can be mitigated by importance sampling techniques, but these techniques suffer from impractical variance when dealing with a large number of actions. In this paper, we introduce a selective importance sampling estimator (sIS) that operates in a significantly more favorable bias-variance regime. The sIS estimator is obtained by performing importance sampling on the conditional expectation of the reward with respect to a small subset of actions for each instance (a form of Rao-Blackwellization). We employ this estimator in a novel algorithmic procedure---named Policy Optimization for eXtreme Models (POXM)---for learning from bandit feedback on XMC tasks. In POXM, the selected actions for the sIS estimator are the top-p actions of the logging policy, where p is adjusted from the data and is significantly smaller than the size of the action space. We use a supervised-to-bandit conversion on three XMC datasets to benchmark our POXM method against three competing methods: BanditNet, a previously applied partial matching pruning strategy, and a supervised learning baseline. Whereas BanditNet sometimes improves marginally over the logging policy, our experiments show that POXM systematically and significantly improves over all baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 被引用 62 次
- On Component Interactions in Two-Stage Recommender SystemsJiri Hron, Karl Krauth, Michael I. Jordan, Niki KilbertusNeurIPS 2021 · 被引用 39 次
- Off-Policy Evaluation for Large Action Spaces via Conjunct Effect ModelingYuta Saito, Qingyang Ren, Thorsten JoachimsICML 2023 · 被引用 34 次
- Control Variates for Slate Off-Policy EvaluationNikos Vlassis, Ashok Chandrashekar, Fernando Amat Gil, Nathan KallusNeurIPS 2021 · 被引用 21 次
- Off-Policy Evaluation for Large Action Spaces via Policy ConvolutionNoveen Sachdeva, Lequn Wang, Dawen Liang, Nathan Kallus 等WWW 2024 · 被引用 17 次
它引用的顶会 Paper2
相关 Paper
- Practical Counterfactual Policy Learning for Top-K RecommendationsYaxu Liu, Jui-Nan Yen, Bo-Wen Yuan, Rundong Shi 等KDD 2022 · 被引用 11 次
- PINA: Leveraging Side Information in eXtreme Multi-label Classification via Predicted Instance Neighborhood AggregationEli Chien, Jiong Zhang, Cho-Jui Hsieh, Jyun-Yu Jiang 等ICML 2023 · 被引用 11 次
- Off-policy Bandits with Deficient SupportNoveen Sachdeva, Yi Su, Thorsten JoachimsKDD 2020 · 被引用 22 次
- ELIAS: End-to-End Learning to Index and Search in Large Output SpacesNilesh Gupta, Patrick H. Chen, Hsiang-Fu Yu, Cho-Jui Hsieh 等NeurIPS 2022 · 被引用 19 次
- Policy Optimization as Online Learning with Mediator FeedbackAlberto Maria Metelli, Matteo Papini, Pierluca D'Oro, Marcello RestelliAAAI 2021 · 被引用 11 次
