Hindsight Learning for MDPs with Exogenous Inputs
Sean R. Sinclair, Felipe Vieira Frujeri, Ching-An Cheng, Luke Marshall, Hugo de Oliveira Barbalho, Jingling Li, Jennifer Neville, Ishai Menache, Adith Swaminathan
Abstract
Many resource management problems require sequential decision-making under uncertainty, where the only uncertainty affecting the decision outcomes are exogenous variables outside the control of the decision-maker. We model these problems as Exo-MDPs (Markov Decision Processes with Exogenous Inputs) and design a class of data-efficient algorithms for them termed Hindsight Learning (HL). Our HL algorithms achieve data efficiency by leveraging a key insight: having samples of the exogenous variables, past decisions can be revisited in hindsight to infer counterfactual consequences that can accelerate policy improvements. We compare HL against classic baselines in the multi-secretary and airline revenue management problems. We also scale our algorithms to a business-critical cloud resource management problem -- allocating Virtual Machines (VMs) to physical machines, and simulate their performance with real datasets from a large public cloud provider. We find that HL algorithms outperform domain-specific heuristics, as well as state-of-the-art reinforcement learning methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 709fd162-a008-48bf-b4af-4b4b07a751b8Cited by top-tier papers5
- Learning in POMDPs is Sample-Efficient with Hindsight ObservabilityJonathan Lee, Alekh Agarwal, Christoph Dann, Tong ZhangICML 2023 · 25 citations
- Sample Efficient Reinforcement Learning in Mixed Systems through Augmented Samples and Its Applications to Queueing NetworksHonghao Wei, Xin Liu, Weina Wang, Lei YingNeurIPS 2023 · 13 citations
- Lookback Prophet InequalitiesZiyad Benomar, Dorian Baudry, Vianney PerchetNeurIPS 2024 · 2 citations
- Online Reinforcement Learning in Non-Stationary Context-Driven EnvironmentsPouya Hamadanian, Arash Nasr-Esfahany, Malte Schwarzkopf, Siddhartha Sen et al.ICLR 2025
- Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight ObservationTonghe Zhang, Yu Chen, Longbo HuangICML 2024
Builds on12
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao et al.NeurIPS 2021 · 373 citations
- Reinforcement Learning for Integer Programming: Learning to CutYunhao Tang, Shipra Agrawal, Yuri FaenzaICML 2020 · 224 citations
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan et al.OSDI 2020 · 189 citations
- Adversarially Trained Actor Critic for Offline Reinforcement LearningChing-An Cheng, Tengyang Xie, Nan Jiang, Alekh AgarwalICML 2022 · 156 citations
- Provably Good Batch Off-Policy Reinforcement Learning Without Great ExplorationYao Liu, Adith Swaminathan, Alekh Agarwal, Emma BrunskillNeurIPS 2020 · 97 citations
Related papers
- Online Placement of Virtual Machines with Prior DataDavid Naori, Danny RazINFOCOM 2020 · 4 citations
- Is Pure Exploitation Sufficient in Exogenous MDPs with Linear Function Approximation?Hao Liang, Jiayu Cheng, Sean R. Sinclair, Yali DuICLR 2026
- Dual Mirror Descent for Online Allocation ProblemsSantiago R. Balseiro, Haihao Lu, Vahab S. MirrokniICML 2020 · 102 citations
- Rewriting History with Inverse RL: Hindsight Inference for Policy ImprovementBen Eysenbach, Xinyang Geng, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2020 · 96 citations
- Probabilistic Active Meta-LearningJean Kaddour, Steindór Sæmundsson, Marc Peter DeisenrothNeurIPS 2020 · 38 citations
