Learning the Optimal Policy for Balancing Short-Term and Long-Term Rewards
Qinwei Yang, Xueqing Liu, Yan Zeng, Ruocheng Guo, Yang Liu, Peng Wu
Abstract
Learning the optimal policy to balance multiple short-term and long-term rewards has extensive applications across various domains. Yet, there is a noticeable scarcity of research addressing policy learning strategies in this context. In this paper, we aim to learn the optimal policy capable of effectively balancing multiple short-term and long-term rewards, especially in scenarios where the long-term outcomes are often missing due to data collection challenges over extended periods. Towards this goal, the conventional linear weighting method, which aggregates multiple rewards into a single surrogate reward through weighted summation, can only achieve sub-optimal policies when multiple rewards are related. Motivated by this, we propose a novel decomposition-based policy learning (DPPL) method that converts the whole problem into subproblems. The DPPL method is capable of obtaining optimal policies even when multiple rewards are interrelated. Nevertheless, the DPPL method requires a set of preference vectors specified in advance, posing challenges in practical applications where selecting suitable preferences is non-trivial. To mitigate this, we further theoretically transform the optimization problem in DPPL into an ε -constraint problem, where ε represents the minimum acceptable levels of other rewards while maximizing one reward. This transformation provides intuitive into the selection of preference vectors. Extensive experiments are conducted on the proposed method and the results validate the effectiveness of the method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ca4bd0d-b545-4538-af43-aacef9a50cc1Cited by top-tier papers3
- Learning Counterfactual Outcomes Under Rank PreservationPeng Wu, Haoxuan Li, Chunyuan Zheng, Yan Zeng et al.NeurIPS 2025 · 7 citations
- Adaptive Data-Borrowing for Improving Treatment Effect Estimation using External ControlsQinwei Yang, Jingyi Li, Peng WuNeurIPS 2025 · 4 citations
- A Minimax Approach for Optimal Intervention Policy Learning with Two-Stage OutcomesChenyang Li, Hao Mei, Yue LiuICML 2026
Builds on11
- Removing Hidden Confounding in Recommendation: A Unified Multi-Task Learning ApproachHaoxuan Li, Kunhan Wu, Chunyuan Zheng, Yanghao Xiao et al.NeurIPS 2023 · 68 citations
- Balancing Unobserved Confounding with a Few Unbiased Ratings in Debiased RecommendationsHaoxuan Li, Yanghao Xiao, Chunyuan Zheng, Peng WuWWW 2023 · 64 citations
- Propensity Matters: Measuring and Enhancing Balancing for RecommendationHaoxuan Li, Yanghao Xiao, Chunyuan Zheng, Peng Wu et al.ICML 2023 · 55 citations
- Multiple Robust Learning for RecommendationHaoxuan Li, Quanyu Dai, Yuru Li, Yan Lyu et al.AAAI 2023 · 48 citations
- A Generalized Doubly Robust Learning Framework for Debiasing Post-Click Conversion Rate PredictionQuanyu Dai, Haoxuan Li, Peng Wu, Zhenhua Dong et al.KDD 2022 · 45 citations
Related papers
- Policy Learning for Balancing Short-Term and Long-Term RewardsPeng Wu, Ziyu Shen, Feng Xie, Zhongyao Wang et al.ICML 2024 · 16 citations
- Guided Task Planning Under Complex ConstraintsSepideh Nikookar, Paras Sakharkar, Baljinder Smagh, Sihem Amer-Yahia et al.ICDE 2022 · 9 citations
- Eliciting User Preferences for Personalized Multi-Objective Decision Making through Comparative FeedbackHan Shao, Lee Cohen, Avrim Blum, Yishay Mansour et al.NeurIPS 2023 · 10 citations
- TGRL: An Algorithm for Teacher Guided Reinforcement LearningIdan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit AgrawalICML 2023 · 22 citations
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert et al.ICML 2020 · 93 citations
