RLlib Flow: Distributed Reinforcement Learning is a Dataflow Problem
Eric Liang, Zhanghao Wu, Michael Luo, Sven Mika, Joseph E. Gonzalez, Ion Stoica
摘要
Researchers and practitioners in the field of reinforcement learning (RL) frequently leverage parallel computation, which has led to a plethora of new algorithms and systems in the last few years. In this paper, we re-examine the challenges posed by distributed RL and try to view it through the lens of an old idea: distributed dataflow. We show that viewing RL as a dataflow problem leads to highly composable and performant implementations. We propose RLlib Flow, a hybrid actor-dataflow programming model for distributed RL, and validate its practicality by porting the full suite of algorithms in RLlib, a widely adopted distributed RL library. Concretely, RLlib Flow provides 2-9× code savings in real production code and enables the composition of multi-agent algorithms not possible by end users before. The open-source code is available as part of RLlib at https://github.com/ ray-project/ray/tree/master/rllib . * indicates equal contributions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu 等EuroSys 2025 · 被引用 61 次
- RAST: Reasoning Activation in LLMs via Small-model TransferSiru Ouyang, Xinyu Zhu, Zilin Xiao, Minhao Jiang 等NeurIPS 2025 · 被引用 9 次
- Scalable Option Learning in High-Throughput EnvironmentsMikael Henaff, Scott Fujimoto, Michael Matthews, Michael RabbatICML 2026 · 被引用 5 次
- Laminar: A Scalable Asynchronous RL Post-Training FrameworkGuangming Sheng, Yuxuan Tong, Borui Wan, Wang Zhang 等EuroSys 2026 · 被引用 2 次
- Zero-Shot Context Generalization in Reinforcement Learning from Few Training ContextsJames Chapman, Kedar Karhadkar, Guido F. MontúfarNeurIPS 2025
它引用的顶会 Paper3
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee 等ICLR 2020 · 被引用 608 次
- IMPACT: Importance Weighted Asynchronous Architectures with Clipped Target NetworksMichael Luo, Jiahao Yao, Richard Liaw, Eric Liang 等ICLR 2020 · 被引用 17 次
相关 Paper
- SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand CoresZhiyu Mei, Wei Fu, Jiaxuan Gao, Guangju Wang 等ICLR 2024 · 被引用 10 次
- DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Trainingzhixin wang, Jiaming Xu, Tianyi Zhou, Mingjun Zhang 等ICML 2026 · 被引用 14 次
- MSRL: Distributed Reinforcement Learning with Dataflow FragmentsHuanzhou Zhu, Bo Zhao, Gang Chen, Weifeng Chen 等USENIX ATC 2023 · 被引用 9 次
- DynaRL: Flexible and Dynamic Scheduling of Large-Scale Reinforcement Learning TrainingYuanqing Wang, Hao Lin, Junhao Hu, Chunyang Zhu 等OSDI 2026
- RLinf: Flexible and Efficient Large-Scale Reinforcement Learning via Macro-to-Micro Flow TransformationChao Yu, Yuanqing Wang, Zhen Guo, Hao Lin 等OSDI 2026 · 被引用 63 次
