Is Pure Exploitation Sufficient in Exogenous MDPs with Linear Function Approximation?
Hao Liang, Jiayu Cheng, Sean R. Sinclair, Yali Du
Abstract
Exogenous MDPs (Exo-MDPs) capture sequential decision-making where uncertainty comes solely from exogenous inputs that evolve independently of the learner's actions. This structure is especially common in operations research applications such as inventory control, energy storage, and resource allocation, where exogenous randomness (e.g., demand, arrivals, or prices) drives system behavior. Despite decades of empirical evidence that greedy, exploitation-only methods work remarkably well in these settings, theory has lagged behind: all existing regret guarantees for Exo-MDPs rely on explicit exploration or tabular assumptions. We show that exploration is unnecessary. We propose Pure Exploitation Learning (PEL) and prove the first general finite-sample regret bounds for exploitation-only algorithms in Exo-MDPs. In the tabular case, PEL achieves O(H 2 |Ξ| √ K), where H is the horizon, Ξ the exogenous state space, and K the number of episodes. For large, continuous endogenous state spaces, we introduce LSVI-PE, a simple linear-approximation method whose regret is polynomial in the feature dimension, exogenous state space, and horizon, independent of the endogenous state and action spaces. Our analysis introduces two new tools: counterfactual trajectories and Bellman-closed feature transport, which together allow greedy policies to have accurate value estimates without optimism. Experiments on synthetic and resource-management tasks show PEL consistently outperforming baselines. Overall, our results overturn the conventional wisdom that exploration is required, demonstrating that in Exo-MDPs, pure exploitation is enough.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b3e16dc-b7b0-47fd-8548-689d843031aeRelated papers
- Provably Filtering Exogenous Distractors using Multistep Inverse DynamicsYonathan Efroni, Dipendra Misra, Akshay Krishnamurthy, Alekh Agarwal et al.ICLR 2022 · 38 citations
- Hindsight Learning for MDPs with Exogenous InputsSean R. Sinclair, Felipe Vieira Frujeri, Ching-An Cheng, Luke Marshall et al.ICML 2023 · 31 citations
- Reinforcement Learning in Linear MDPs: Constant Regret and Representation SelectionMatteo Papini, Andrea Tirinzoni, Aldo Pacchiano, Marcello Restelli et al.NeurIPS 2021 · 26 citations
- Policy Optimization as Online Learning with Mediator FeedbackAlberto Maria Metelli, Matteo Papini, Pierluca D'Oro, Marcello RestelliAAAI 2021 · 11 citations
- Learning Adversarial Low-rank Markov Decision Processes with Unknown Transition and Full-information FeedbackCanzhe Zhao, Ruofeng Yang, Baoxiang Wang, Xuezhou Zhang et al.NeurIPS 2023 · 5 citations
