Policy Learning for Balancing Short-Term and Long-Term Rewards
Peng Wu, Ziyu Shen, Feng Xie, Zhongyao Wang, Chunchen Liu, Yan Zeng
摘要
Empirical researchers and decision-makers spanning various domains frequently seek profound insights into the long-term impacts of interventions. While the significance of long-term outcomes is undeniable, an overemphasis on them may inadvertently overshadow short-term gains. Motivated by this, this paper formalizes a new framework for learning the optimal policy that effectively balances both long-term and short-term rewards, where some long-term outcomes are allowed to be missing. In particular, we first present the identifiability of both rewards under mild assumptions. Next, we deduce the semiparametric efficiency bounds, along with the consistency and asymptotic normality of their estimators. We also reveal that short-term outcomes, if associated, contribute to improving the estimator of the longterm reward. Based on the proposed estimators, we develop a principled policy learning approach and further derive the convergence rates of regret and estimation errors associated with the learned policy. Extensive experiments are conducted to validate the effectiveness of the proposed method, demonstrating its practical applicability 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning the Optimal Policy for Balancing Short-Term and Long-Term RewardsQinwei Yang, Xueqing Liu, Yan Zeng, Ruocheng Guo 等NeurIPS 2024 · 被引用 13 次
- Learning Counterfactual Outcomes Under Rank PreservationPeng Wu, Haoxuan Li, Chunyuan Zheng, Yan Zeng 等NeurIPS 2025 · 被引用 7 次
- Adaptive Data-Borrowing for Improving Treatment Effect Estimation using External ControlsQinwei Yang, Jingyi Li, Peng WuNeurIPS 2025 · 被引用 4 次
- A Minimax Approach for Optimal Intervention Policy Learning with Two-Stage OutcomesChenyang Li, Hao Mei, Yue LiuICML 2026
它引用的顶会 Paper5
- Clicks can be Cheating: Counterfactual Recommendation for Mitigating Clickbait IssueWenjie Wang, Fuli Feng, Xiangnan He, Hanwang Zhang 等SIGIR 2021 · 被引用 173 次
- Removing Hidden Confounding in Recommendation: A Unified Multi-Task Learning ApproachHaoxuan Li, Kunhan Wu, Chunyuan Zheng, Yanghao Xiao 等NeurIPS 2023 · 被引用 68 次
- Addressing Unmeasured Confounder for Recommendation with Sensitivity AnalysisSihao Ding, Peng Wu, Fuli Feng, Yitong Wang 等KDD 2022 · 被引用 43 次
- What's the Harm? Sharp Bounds on the Fraction Negatively Affected by TreatmentNathan KallusNeurIPS 2022 · 被引用 40 次
- Trustworthy Policy Learning under the Counterfactual No-Harm CriterionHaoxuan Li, Chunyuan Zheng, Yixiao Cao, Zhi Geng 等ICML 2023 · 被引用 34 次
相关 Paper
- Multiply Robust Off-policy Evaluation and Learning under Truncation by DeathJianing Chu, Shu Yang, Wenbin LuICML 2023 · 被引用 6 次
- TGRL: An Algorithm for Teacher Guided Reinforcement LearningIdan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit AgrawalICML 2023 · 被引用 22 次
- Impatient Bandits: Optimizing Recommendations for the Long-Term Without DelayThomas M. McDonald, Lucas Maystre, Mounia Lalmas, Daniel Russo 等KDD 2023 · 被引用 12 次
- Towards Safe Policy Learning under Partial Identifiability: A Causal ApproachShalmali Joshi, Junzhe Zhang, Elias BareinboimAAAI 2024 · 被引用 10 次
- Robustness in the Face of Partial Identifiability in Reward LearningFilippo Lazzati, Alberto Maria MetelliICLR 2026
