RLlib Flow: Distributed Reinforcement Learning is a Dataflow Problem
Eric Liang, Zhanghao Wu, Michael Luo, Sven Mika, Joseph E. Gonzalez, Ion Stoica
Abstract
Researchers and practitioners in the field of reinforcement learning (RL) frequently leverage parallel computation, which has led to a plethora of new algorithms and systems in the last few years. In this paper, we re-examine the challenges posed by distributed RL and try to view it through the lens of an old idea: distributed dataflow. We show that viewing RL as a dataflow problem leads to highly composable and performant implementations. We propose RLlib Flow, a hybrid actor-dataflow programming model for distributed RL, and validate its practicality by porting the full suite of algorithms in RLlib, a widely adopted distributed RL library. Concretely, RLlib Flow provides 2-9× code savings in real production code and enables the composition of multi-agent algorithms not possible by end users before. The open-source code is available as part of RLlib at https://github.com/ ray-project/ray/tree/master/rllib . * indicates equal contributions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu et al.EuroSys 2025 · 61 citations
- RAST: Reasoning Activation in LLMs via Small-model TransferSiru Ouyang, Xinyu Zhu, Zilin Xiao, Minhao Jiang et al.NeurIPS 2025 · 9 citations
- Scalable Option Learning in High-Throughput EnvironmentsMikael Henaff, Scott Fujimoto, Michael Matthews, Michael RabbatICML 2026 · 5 citations
- Laminar: A Scalable Asynchronous RL Post-Training FrameworkGuangming Sheng, Yuxuan Tong, Borui Wan, Wang Zhang et al.EuroSys 2026 · 2 citations
- Zero-Shot Context Generalization in Reinforcement Learning from Few Training ContextsJames Chapman, Kedar Karhadkar, Guido F. MontúfarNeurIPS 2025
Builds on3
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
- IMPACT: Importance Weighted Asynchronous Architectures with Clipped Target NetworksMichael Luo, Jiahao Yao, Richard Liaw, Eric Liang et al.ICLR 2020 · 17 citations
Related papers
- SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand CoresZhiyu Mei, Wei Fu, Jiaxuan Gao, Guangju Wang et al.ICLR 2024 · 10 citations
- DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Trainingzhixin wang, Jiaming Xu, Tianyi Zhou, Mingjun Zhang et al.ICML 2026 · 14 citations
- MSRL: Distributed Reinforcement Learning with Dataflow FragmentsHuanzhou Zhu, Bo Zhao, Gang Chen, Weifeng Chen et al.USENIX ATC 2023 · 9 citations
- DynaRL: Flexible and Dynamic Scheduling of Large-Scale Reinforcement Learning TrainingYuanqing Wang, Hao Lin, Junhao Hu, Chunyang Zhu et al.OSDI 2026
- RLinf: Flexible and Efficient Large-Scale Reinforcement Learning via Macro-to-Micro Flow TransformationChao Yu, Yuanqing Wang, Zhen Guo, Hao Lin et al.OSDI 2026 · 63 citations
