Speeding up Policy Simulation in Supply Chain RL
Vivek F. Farias, Joren Gijsbrechts, Aryan I. Khojandi, Tianyi Peng, Andrew Zheng
摘要
Simulating a single trajectory of a dynamical system under some state-dependent policy is a core bottleneck in policy optimization (PO) algorithms. The many inherently serial policy evaluations that must be performed in a single simulation constitute the bulk of this bottleneck. In applying PO to supply chain optimization (SCO) problems, simulating a single sample path corresponding to one month of a supply chain can take several hours. We present an iterative algorithm to accelerate policy simulation, dubbed Picard Iteration. This scheme carefully assigns policy evaluation tasks to independent processes. Within an iteration, any given process evaluates the policy only on its assigned tasks while assuming a certain 'cached' evaluation for other tasks; the cache is updated at the end of the iteration. Implemented on GPUs, this scheme admits batched evaluation of the policy across a single trajectory. We prove that the structure afforded by many SCO problems allows convergence in a small number of iterations independent of the horizon. We demonstrate practical speedups of 400x on large-scale SCO problems even with a single GPU, and also demonstrate practical efficacy in other RL environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 被引用 290 次
- Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement LearningAleksei Petrenko, Zhehui Huang, Tushar Kumar, Gaurav S. Sukhatme 等ICML 2020 · 被引用 131 次
- Dual Mirror Descent for Online Allocation ProblemsSantiago R. Balseiro, Haihao Lu, Vahab S. MirrokniICML 2020 · 被引用 102 次
相关 Paper
- From Sequential to Parallel: Reformulating Dynamic Programming as GPU Kernels for Large-Scale Stochastic Combinatorial OptimizationJingyi Zhao, Linxin Yang, Haohua ZHANG, Qile He 等ICLR 2026
- Differentiable Model Predictive Control on the GPUEmre Adabag, Marcus Greiff, John Subosits, Thomas Jonathan LewICLR 2026 · 被引用 13 次
- Penetration-free projective dynamics on the GPULei Lan, Guanqun Ma, Yin Yang, Changxi Zheng 等SIGGRAPH 2022 · 被引用 49 次
- SAPG: Split and Aggregate Policy GradientsJayesh Singla, Ananye Agarwal, Deepak PathakICML 2024 · 被引用 19 次
- Large Batch Simulation for Deep Reinforcement LearningBrennan Shacklett, Erik Wijmans, Aleksei Petrenko, Manolis Savva 等ICLR 2021 · 被引用 29 次
