Learning the Optimal Policy for Balancing Short-Term and Long-Term Rewards
Qinwei Yang, Xueqing Liu, Yan Zeng, Ruocheng Guo, Yang Liu, Peng Wu
摘要
Learning the optimal policy to balance multiple short-term and long-term rewards has extensive applications across various domains. Yet, there is a noticeable scarcity of research addressing policy learning strategies in this context. In this paper, we aim to learn the optimal policy capable of effectively balancing multiple short-term and long-term rewards, especially in scenarios where the long-term outcomes are often missing due to data collection challenges over extended periods. Towards this goal, the conventional linear weighting method, which aggregates multiple rewards into a single surrogate reward through weighted summation, can only achieve sub-optimal policies when multiple rewards are related. Motivated by this, we propose a novel decomposition-based policy learning (DPPL) method that converts the whole problem into subproblems. The DPPL method is capable of obtaining optimal policies even when multiple rewards are interrelated. Nevertheless, the DPPL method requires a set of preference vectors specified in advance, posing challenges in practical applications where selecting suitable preferences is non-trivial. To mitigate this, we further theoretically transform the optimization problem in DPPL into an ε -constraint problem, where ε represents the minimum acceptable levels of other rewards while maximizing one reward. This transformation provides intuitive into the selection of preference vectors. Extensive experiments are conducted on the proposed method and the results validate the effectiveness of the method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning Counterfactual Outcomes Under Rank PreservationPeng Wu, Haoxuan Li, Chunyuan Zheng, Yan Zeng 等NeurIPS 2025 · 被引用 7 次
- Adaptive Data-Borrowing for Improving Treatment Effect Estimation using External ControlsQinwei Yang, Jingyi Li, Peng WuNeurIPS 2025 · 被引用 4 次
- A Minimax Approach for Optimal Intervention Policy Learning with Two-Stage OutcomesChenyang Li, Hao Mei, Yue LiuICML 2026
它引用的顶会 Paper11
- Removing Hidden Confounding in Recommendation: A Unified Multi-Task Learning ApproachHaoxuan Li, Kunhan Wu, Chunyuan Zheng, Yanghao Xiao 等NeurIPS 2023 · 被引用 68 次
- Balancing Unobserved Confounding with a Few Unbiased Ratings in Debiased RecommendationsHaoxuan Li, Yanghao Xiao, Chunyuan Zheng, Peng WuWWW 2023 · 被引用 64 次
- Propensity Matters: Measuring and Enhancing Balancing for RecommendationHaoxuan Li, Yanghao Xiao, Chunyuan Zheng, Peng Wu 等ICML 2023 · 被引用 55 次
- Multiple Robust Learning for RecommendationHaoxuan Li, Quanyu Dai, Yuru Li, Yan Lyu 等AAAI 2023 · 被引用 48 次
- A Generalized Doubly Robust Learning Framework for Debiasing Post-Click Conversion Rate PredictionQuanyu Dai, Haoxuan Li, Peng Wu, Zhenhua Dong 等KDD 2022 · 被引用 45 次
相关 Paper
- Policy Learning for Balancing Short-Term and Long-Term RewardsPeng Wu, Ziyu Shen, Feng Xie, Zhongyao Wang 等ICML 2024 · 被引用 16 次
- Guided Task Planning Under Complex ConstraintsSepideh Nikookar, Paras Sakharkar, Baljinder Smagh, Sihem Amer-Yahia 等ICDE 2022 · 被引用 9 次
- Eliciting User Preferences for Personalized Multi-Objective Decision Making through Comparative FeedbackHan Shao, Lee Cohen, Avrim Blum, Yishay Mansour 等NeurIPS 2023 · 被引用 10 次
- TGRL: An Algorithm for Teacher Guided Reinforcement LearningIdan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit AgrawalICML 2023 · 被引用 22 次
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert 等ICML 2020 · 被引用 93 次
