Learning MDPs from Features: Predict-Then-Optimize for Sequential Decision Making by Reinforcement Learning
Kai Wang, Sanket Shah, Haipeng Chen, Andrew Perrault, Finale Doshi-Velez, Milind Tambe
摘要
In the predict-then-optimize framework, the objective is to train a predictive model, mapping from environment features to parameters of an optimization problem, which maximizes decision quality when the optimization is subsequently solved. Recent work on decision-focused learning shows that embedding the optimization problem in the training pipeline can improve decision quality and help generalize better to unseen tasks compared to relying on an intermediate loss function for evaluating prediction quality. We study the predict-then-optimize framework in the context of sequential decision problems (formulated as MDPs) that are solved via reinforcement learning. In particular, we are given environment features and a set of trajectories from training MDPs, which we use to train a predictive model that generalizes to unseen test MDPs without trajectories. Two significant computational challenges arise in applying decision-focused learning to MDPs: (i) large state and action spaces make it infeasible for existing techniques to differentiate through MDP problems, and (ii) the high-dimensional policy space, as parameterized by a neural network, makes differentiating through a policy expensive. We resolve the first challenge by sampling provably unbiased derivatives to approximate and differentiate through optimality conditions, and the second challenge by using a low-rank approximation to the high-dimensional sample-based derivatives. We implement both Bellman-based and policy gradient-based decision-focused learning on three different MDP problems with missing parameters, and show that decision-focused learning performs better in generalization to unseen tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Leaving the Nest: Going beyond Local Loss Functions for Predict-Then-OptimizeSanket Shah, Bryan Wilder, Andrew Perrault, Milind TambeAAAI 2024 · 被引用 22 次
- First-Order Methods for Linearly Constrained Bilevel OptimizationGuy Kornowski, Swati Padmanabhan, Kai Wang, Zhe Zhang 等NeurIPS 2024 · 被引用 21 次
- Scalable Decision-Focused Learning in Restless Multi-Armed Bandits with Application to Maternal and Child HealthKai Wang, Shresth Verma, Aditya Mate, Sanket Shah 等AAAI 2023 · 被引用 19 次
- Learning for Edge-Weighted Online Bipartite Matching with Robustness GuaranteesPengfei Li, Jianyi Yang, Shaolei RenICML 2023 · 被引用 7 次
- Robustified Learning for Online Optimization with Memory CostsPengfei Li, Jianyi Yang, Shaolei RenINFOCOM 2023 · 被引用 2 次
它引用的顶会 Paper6
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran 等NeurIPS 2021 · 被引用 549 次
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- Interior Point Solving for LP-based prediction+optimisationJayanta Mandi, Tias GunsNeurIPS 2020 · 被引用 138 次
- Automatically Learning Compact Quality-aware Surrogates for Optimization ProblemsKai Wang, Bryan Wilder, Andrew Perrault, Milind TambeNeurIPS 2020 · 被引用 37 次
相关 Paper
- Solver-Free Decision-Focused Learning for Linear Optimization ProblemsSenne Berden, Ali Irfan Mahmutogullari, Dimos Tsouros, Tias GunsNeurIPS 2025 · 被引用 13 次
- Diffusion-DFL: Decision-focused Diffusion Models for Stochastic OptimizationZihao Zhao, Christopher Yeh, Lingkai Kong, Kai WangICLR 2026 · 被引用 8 次
- Feasibility-Aware Decision-Focused Learning for Predicting Parameters in the ConstraintsJayanta Mandi, Marianne Defresne, Senne Berden, Tias GunsNeurIPS 2025 · 被引用 9 次
- A Solver-Free Training Method for Predict-then-OptimizeBeichen Wan, Mo LiuICML 2026 · 被引用 1 次
- DFF: Decision-Focused Fine-Tuning for Smarter Predict-Then-Optimize with Limited DataJiaqi Yang, Enming Liang, Zicheng Su, Zhichao Zou 等AAAI 2025 · 被引用 6 次
