The Value of Reward Lookahead in Reinforcement Learning
Nadav Merlis, Dorian Baudry, Vianney Perchet
摘要
In reinforcement learning (RL), agents sequentially interact with changing environments while aiming to maximize the obtained rewards. Usually, rewards are observed only after acting, and so the goal is to maximize the expected cumulative reward. Yet, in many practical settings, reward information is observed in advance -- prices are observed before performing transactions; nearby traffic information is partially known; and goals are oftentimes given to agents prior to the interaction. In this work, we aim to quantifiably analyze the value of such future reward information through the lens of competitive analysis. In particular, we measure the ratio between the value of standard RL agents and that of agents with partial future-reward lookahead. We characterize the worst-case reward distribution and derive exact ratios for the worst-case reward expectations. Surprisingly, the resulting ratios relate to known quantities in offline RL and reward-free exploration. We further provide tight bounds for the ratio given the worst-case dynamics. Our results cover the full spectrum between observing the immediate rewards before acting to observing all the rewards before the interaction starts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Reinforcement Learning with Lookahead InformationNadav MerlisNeurIPS 2024 · 被引用 11 次
- Understanding LoRA as Knowledge Memory: An Empirical AnalysisSeungju Back, Dongwoo Lee, Naun Kang, Taehee Lee 等ICML 2026 · 被引用 10 次
- Fast Non-Episodic Finite-Horizon RL with K-Step Lookahead ThresholdingJiamin Xu, Kyra GanICML 2026
它引用的顶会 Paper6
- Perturbation-based Regret Analysis of Predictive Control in Linear Time Varying SystemsYiheng Lin, Yang Hu, Guanya Shi, Haoyuan Sun 等NeurIPS 2021 · 被引用 55 次
- Bounded-Regret MPC via Perturbation Analysis: Prediction Error, Constraints, and NonlinearityYiheng Lin, Yang Hu, Guannan Qu, Tongxin Li 等NeurIPS 2022 · 被引用 31 次
- Online Planning with Lookahead PoliciesYonathan Efroni, Mohammad Ghavamzadeh, Shie MannorNeurIPS 2020 · 被引用 25 次
- Near Instance-Optimal PAC Reinforcement Learning for Deterministic MDPsAndrea Tirinzoni, Aymen Al Marjani, Emilie KaufmannNeurIPS 2022 · 被引用 20 次
- Online resource allocation in Markov ChainsJianhao Jia, Hao Li, Kai Liu, Ziqi Liu 等WWW 2023 · 被引用 7 次
相关 Paper
- Policy Gradients Incorporating the FutureDavid Venuto, Elaine Lau, Doina Precup, Ofir NachumICLR 2022 · 被引用 9 次
- Reinforcement Learning with Trajectory FeedbackYonathan Efroni, Nadav Merlis, Shie MannorAAAI 2021 · 被引用 48 次
- Learning to Explore in POMDPs with Informational RewardsAnnie Xie, Logan M. Bhamidipaty, Evan Zheran Liu, Joey Hong 等ICML 2024 · 被引用 2 次
- Near-Optimal Regret for Adversarial MDP with Delayed Bandit FeedbackTiancheng Jin, Tal Lancewicki, Haipeng Luo, Yishay Mansour 等NeurIPS 2022 · 被引用 29 次
- Optimizing Traffic Control with Model-Based Learning: A Pessimistic Approach to Data-Efficient Policy InferenceMayuresh Kunjir, Sanjay Chawla, Siddarth Chandrasekar, Devika Jay 等KDD 2023 · 被引用 3 次
