Incentivized Exploration with Stochastic Covariates: A Two-Stage Mechanism Design for Recommender System
Yuantong Li, Guang Cheng, Xiaowu Dai
Abstract
Recommender systems play a crucial role in internet economies by connecting users with relevant products. However, designing effective recommender systems faces the key challenges: the exploration-exploitation tradeoff in securing incentive to explore new products against user’s self-interested preferences. While prior work addresses Bayesian Incentive Compatibility (BIC) in fixed-design linear bandits (Sellke & Slivkins, 2023), we tackle the challenge of stochastic user covariates sampled online. Unlike standard black-box reductions (Mansour et al., 2020), our two-stage framework exploits the linear reward structure to achieve sublinear regret while satisfying incentive constraints. To address it, we propose a two-stage algorithm that integrates incentivized exploration with any efficient plug-in offline learning algorithms. In the first stage, it explores products while maintaining incentive compatibility to gather optimal samples. The second stage employs inverse proportional gap sampling strategy (IPGS) integrated with any efficient learning methods to secure sublinear regret. Theoretically, we prove that algorithm RCB achieves regret and simultaneously satisfies incentive constraints, and discovers the tradeoff between incentive budget and regret, validating in experiments. We demonstrate RCB’s strong incentive gain, sublinear regret, and robustness through a real application on personalized warfarin dosing and simulations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7f6a4c6-65b9-4ddc-bb6b-e2c4d3ac2651Builds on8
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression OraclesDylan J. Foster, Alexander RakhlinICML 2020 · 241 citations
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsJiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang et al.ICML 2024 · 200 citations
- Regret Bounds for Information-Directed Reinforcement LearningBotao Hao, Tor LattimoreNeurIPS 2022 · 31 citations
- Incentivizing Combinatorial Bandit ExplorationXinyan Hu, Dung Daniel T. Ngo, Aleksandrs Slivkins, Zhiwei Steven WuNeurIPS 2022 · 14 citations
Related papers
- Incentivizing Exploration with Linear Contexts and Combinatorial ActionsMark SellkeICML 2023 · 5 citations
- Geometry Meets Incentives: Sample-Efficient Incentivized Exploration with Linear ContextsBen Schiffer, Mark SellkeNeurIPS 2025
- Learning to Incentivize Information Acquisition: Proper Scoring Rules Meet Principal-Agent ModelSiyu Chen, Jibang Wu, Yifan Wu, Zhuoran YangICML 2023 · 9 citations
- Online Mechanism Design for Information AcquisitionFederico Cacciamani, Matteo Castiglioni, Nicola GattiICML 2023 · 3 citations
- Bandits Meet Mechanism Design to Combat Clickbait in Online RecommendationThomas Kleine Buening, Aadirupa Saha, Christos Dimitrakakis, Haifeng XuICLR 2024 · 7 citations
