The Value of Reward Lookahead in Reinforcement Learning
Nadav Merlis, Dorian Baudry, Vianney Perchet
Abstract
In reinforcement learning (RL), agents sequentially interact with changing environments while aiming to maximize the obtained rewards. Usually, rewards are observed only after acting, and so the goal is to maximize the expected cumulative reward. Yet, in many practical settings, reward information is observed in advance -- prices are observed before performing transactions; nearby traffic information is partially known; and goals are oftentimes given to agents prior to the interaction. In this work, we aim to quantifiably analyze the value of such future reward information through the lens of competitive analysis. In particular, we measure the ratio between the value of standard RL agents and that of agents with partial future-reward lookahead. We characterize the worst-case reward distribution and derive exact ratios for the worst-case reward expectations. Surprisingly, the resulting ratios relate to known quantities in offline RL and reward-free exploration. We further provide tight bounds for the ratio given the worst-case dynamics. Our results cover the full spectrum between observing the immediate rewards before acting to observing all the rewards before the interaction starts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1f39233-3825-4d6c-ba4c-c6ff81d089ffCited by top-tier papers3
- Reinforcement Learning with Lookahead InformationNadav MerlisNeurIPS 2024 · 11 citations
- Understanding LoRA as Knowledge Memory: An Empirical AnalysisSeungju Back, Dongwoo Lee, Naun Kang, Taehee Lee et al.ICML 2026 · 10 citations
- Fast Non-Episodic Finite-Horizon RL with K-Step Lookahead ThresholdingJiamin Xu, Kyra GanICML 2026
Builds on6
- Perturbation-based Regret Analysis of Predictive Control in Linear Time Varying SystemsYiheng Lin, Yang Hu, Guanya Shi, Haoyuan Sun et al.NeurIPS 2021 · 55 citations
- Bounded-Regret MPC via Perturbation Analysis: Prediction Error, Constraints, and NonlinearityYiheng Lin, Yang Hu, Guannan Qu, Tongxin Li et al.NeurIPS 2022 · 31 citations
- Online Planning with Lookahead PoliciesYonathan Efroni, Mohammad Ghavamzadeh, Shie MannorNeurIPS 2020 · 25 citations
- Near Instance-Optimal PAC Reinforcement Learning for Deterministic MDPsAndrea Tirinzoni, Aymen Al Marjani, Emilie KaufmannNeurIPS 2022 · 20 citations
- Online resource allocation in Markov ChainsJianhao Jia, Hao Li, Kai Liu, Ziqi Liu et al.WWW 2023 · 7 citations
Related papers
- Policy Gradients Incorporating the FutureDavid Venuto, Elaine Lau, Doina Precup, Ofir NachumICLR 2022 · 9 citations
- Reinforcement Learning with Trajectory FeedbackYonathan Efroni, Nadav Merlis, Shie MannorAAAI 2021 · 48 citations
- Learning to Explore in POMDPs with Informational RewardsAnnie Xie, Logan M. Bhamidipaty, Evan Zheran Liu, Joey Hong et al.ICML 2024 · 2 citations
- Near-Optimal Regret for Adversarial MDP with Delayed Bandit FeedbackTiancheng Jin, Tal Lancewicki, Haipeng Luo, Yishay Mansour et al.NeurIPS 2022 · 29 citations
- Optimizing Traffic Control with Model-Based Learning: A Pessimistic Approach to Data-Efficient Policy InferenceMayuresh Kunjir, Sanjay Chawla, Siddarth Chandrasekar, Devika Jay et al.KDD 2023 · 3 citations
