Learning MDPs from Features: Predict-Then-Optimize for Sequential Decision Making by Reinforcement Learning
Kai Wang, Sanket Shah, Haipeng Chen, Andrew Perrault, Finale Doshi-Velez, Milind Tambe
Abstract
In the predict-then-optimize framework, the objective is to train a predictive model, mapping from environment features to parameters of an optimization problem, which maximizes decision quality when the optimization is subsequently solved. Recent work on decision-focused learning shows that embedding the optimization problem in the training pipeline can improve decision quality and help generalize better to unseen tasks compared to relying on an intermediate loss function for evaluating prediction quality. We study the predict-then-optimize framework in the context of sequential decision problems (formulated as MDPs) that are solved via reinforcement learning. In particular, we are given environment features and a set of trajectories from training MDPs, which we use to train a predictive model that generalizes to unseen test MDPs without trajectories. Two significant computational challenges arise in applying decision-focused learning to MDPs: (i) large state and action spaces make it infeasible for existing techniques to differentiate through MDP problems, and (ii) the high-dimensional policy space, as parameterized by a neural network, makes differentiating through a policy expensive. We resolve the first challenge by sampling provably unbiased derivatives to approximate and differentiate through optimality conditions, and the second challenge by using a low-rank approximation to the high-dimensional sample-based derivatives. We implement both Bellman-based and policy gradient-based decision-focused learning on three different MDP problems with missing parameters, and show that decision-focused learning performs better in generalization to unseen tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Leaving the Nest: Going beyond Local Loss Functions for Predict-Then-OptimizeSanket Shah, Bryan Wilder, Andrew Perrault, Milind TambeAAAI 2024 · 22 citations
- First-Order Methods for Linearly Constrained Bilevel OptimizationGuy Kornowski, Swati Padmanabhan, Kai Wang, Zhe Zhang et al.NeurIPS 2024 · 21 citations
- Scalable Decision-Focused Learning in Restless Multi-Armed Bandits with Application to Maternal and Child HealthKai Wang, Shresth Verma, Aditya Mate, Sanket Shah et al.AAAI 2023 · 19 citations
- Learning for Edge-Weighted Online Bipartite Matching with Robustness GuaranteesPengfei Li, Jianyi Yang, Shaolei RenICML 2023 · 7 citations
- Robustified Learning for Online Optimization with Memory CostsPengfei Li, Jianyi Yang, Shaolei RenINFOCOM 2023 · 2 citations
Builds on6
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran et al.NeurIPS 2021 · 549 citations
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze et al.ICLR 2020 · 315 citations
- Interior Point Solving for LP-based prediction+optimisationJayanta Mandi, Tias GunsNeurIPS 2020 · 138 citations
- Automatically Learning Compact Quality-aware Surrogates for Optimization ProblemsKai Wang, Bryan Wilder, Andrew Perrault, Milind TambeNeurIPS 2020 · 37 citations
Related papers
- Solver-Free Decision-Focused Learning for Linear Optimization ProblemsSenne Berden, Ali Irfan Mahmutogullari, Dimos Tsouros, Tias GunsNeurIPS 2025 · 13 citations
- Diffusion-DFL: Decision-focused Diffusion Models for Stochastic OptimizationZihao Zhao, Christopher Yeh, Lingkai Kong, Kai WangICLR 2026 · 8 citations
- Feasibility-Aware Decision-Focused Learning for Predicting Parameters in the ConstraintsJayanta Mandi, Marianne Defresne, Senne Berden, Tias GunsNeurIPS 2025 · 9 citations
- A Solver-Free Training Method for Predict-then-OptimizeBeichen Wan, Mo LiuICML 2026 · 1 citation
- DFF: Decision-Focused Fine-Tuning for Smarter Predict-Then-Optimize with Limited DataJiaqi Yang, Enming Liang, Zicheng Su, Zhichao Zou et al.AAAI 2025 · 6 citations
