Policy Learning for Balancing Short-Term and Long-Term Rewards
Peng Wu, Ziyu Shen, Feng Xie, Zhongyao Wang, Chunchen Liu, Yan Zeng
Abstract
Empirical researchers and decision-makers spanning various domains frequently seek profound insights into the long-term impacts of interventions. While the significance of long-term outcomes is undeniable, an overemphasis on them may inadvertently overshadow short-term gains. Motivated by this, this paper formalizes a new framework for learning the optimal policy that effectively balances both long-term and short-term rewards, where some long-term outcomes are allowed to be missing. In particular, we first present the identifiability of both rewards under mild assumptions. Next, we deduce the semiparametric efficiency bounds, along with the consistency and asymptotic normality of their estimators. We also reveal that short-term outcomes, if associated, contribute to improving the estimator of the longterm reward. Based on the proposed estimators, we develop a principled policy learning approach and further derive the convergence rates of regret and estimation errors associated with the learned policy. Extensive experiments are conducted to validate the effectiveness of the proposed method, demonstrating its practical applicability 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ea1be05-26bd-4d55-bfe9-739e1880fcf9Cited by top-tier papers4
- Learning the Optimal Policy for Balancing Short-Term and Long-Term RewardsQinwei Yang, Xueqing Liu, Yan Zeng, Ruocheng Guo et al.NeurIPS 2024 · 13 citations
- Learning Counterfactual Outcomes Under Rank PreservationPeng Wu, Haoxuan Li, Chunyuan Zheng, Yan Zeng et al.NeurIPS 2025 · 7 citations
- Adaptive Data-Borrowing for Improving Treatment Effect Estimation using External ControlsQinwei Yang, Jingyi Li, Peng WuNeurIPS 2025 · 4 citations
- A Minimax Approach for Optimal Intervention Policy Learning with Two-Stage OutcomesChenyang Li, Hao Mei, Yue LiuICML 2026
Builds on5
- Clicks can be Cheating: Counterfactual Recommendation for Mitigating Clickbait IssueWenjie Wang, Fuli Feng, Xiangnan He, Hanwang Zhang et al.SIGIR 2021 · 173 citations
- Removing Hidden Confounding in Recommendation: A Unified Multi-Task Learning ApproachHaoxuan Li, Kunhan Wu, Chunyuan Zheng, Yanghao Xiao et al.NeurIPS 2023 · 68 citations
- Addressing Unmeasured Confounder for Recommendation with Sensitivity AnalysisSihao Ding, Peng Wu, Fuli Feng, Yitong Wang et al.KDD 2022 · 43 citations
- What's the Harm? Sharp Bounds on the Fraction Negatively Affected by TreatmentNathan KallusNeurIPS 2022 · 40 citations
- Trustworthy Policy Learning under the Counterfactual No-Harm CriterionHaoxuan Li, Chunyuan Zheng, Yixiao Cao, Zhi Geng et al.ICML 2023 · 34 citations
Related papers
- Multiply Robust Off-policy Evaluation and Learning under Truncation by DeathJianing Chu, Shu Yang, Wenbin LuICML 2023 · 6 citations
- TGRL: An Algorithm for Teacher Guided Reinforcement LearningIdan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit AgrawalICML 2023 · 22 citations
- Impatient Bandits: Optimizing Recommendations for the Long-Term Without DelayThomas M. McDonald, Lucas Maystre, Mounia Lalmas, Daniel Russo et al.KDD 2023 · 12 citations
- Towards Safe Policy Learning under Partial Identifiability: A Causal ApproachShalmali Joshi, Junzhe Zhang, Elias BareinboimAAAI 2024 · 10 citations
- Robustness in the Face of Partial Identifiability in Reward LearningFilippo Lazzati, Alberto Maria MetelliICLR 2026
