Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
Brian R. Bartoldson, Siddarth Venkatraman, James Diffenderfer, Moksh Jain, Tal Ben-Nun, Seanie Lee, Minsu Kim, Johan S. Obando-Ceron, Yoshua Bengio, Bhavya Kailkhura
Abstract
Reinforcement learning (RL) is a critical component of large language model (LLM) post-training. However, on-policy algorithms used for post-training are not naturally robust to a diversified content of experience replay buffers, which asynchronous off-policy actors can efficiently populate in parallel to training. We propose efficiently learning on such off-policy data via Trajectory Balance with Asynchrony (TBA), an approach to asynchronous RL for LLMs that leverages the principled off-policy TB objective. On math, preference-tuning, and automated red-teaming tasks, we post-train models ranging from Pythia 410M to Qwen 2.5 7B, finding TBA offers speed and performance boosts over strong baselines like Online DPO and Dr. GRPO. Beyond TBA's performance benefits (high accuracy even as asynchrony grows) and speedups (4× or more), we show its reward-and recency-prioritizing sampling enable further gains as data generation is scaled. Our code is available at https://github.com/bbartoldson/TBA. 128 256 512 1024 2048 4096 Compute Time (Minutes) 49% 50% 51% 52% 53% 54% 55% GSM8K Test Pass@1 GRPO PPO RLOO Online DPO Online DPO VinePPO TBA (Ours) 50x faster 1.6x faster +2.0% +1.2%
Figure 1: TBA excels on the GSM8K mathematical reasoning task. All plotted points use 4xA100 GPUs (or comparable L40S GPUs). DPO and RLOO baselines taken from Noukhovitch et al. [46], PPO and VinePPO baselines taken from Kazemnejad et al. [29]. The baseline model is the SFTed RhoMath-1B [35] model, which gets 40% accuracy after SFT and before RL. Appendix B has details. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Across mathematical reasoning, preference-tuning, and automated red-teaming tasks, we find TBA broadly offers three advantages over existing LLM post-training approaches: 1) Stable off-policy RL, unlocking asynchrony for massive parallelization and reduced wall-clock times -see Figures 1 and 3.
- Sampling from a diverse replay buffer, improving exploration and preventing mode collapse. 3) Scalable search compatibility that aids sparse reward settings like automated red-teaming.
Our key contributions are summarized as follows:
• We introduce TBA, a distributed asynchronous RL framework for fast, scalable LLM post-training.
• We show that LLMs post-trained with TBA match or exceed performances from existing methods, illustrating TB's ability to use off-policy data in asynchronous RL.
• We demonstrate significant speedups (4× to 50×) for RL across mathematical reasoning, preference-tuning, and automated red-teaming.
By enabling high-quality and fast off-policy post-training, TBA contributes to scalable and effective LLM alignment, ensuring that large models can be refined more efficiently for real-world deployment.
2 Related Work RL fine-tuning of language models Reinforcement learning (RL) has been an integral component for training LLMs [14, 48]. In particular, RL has become the de facto approach for aligning language models with human preferences [8, 82, 62, 49]. Much of this work relies on Proximal Policy Optimization (PPO) [57], an on-policy RL algorithm that has become a default choice for fine-tuning LLMs due to its strong performance across different setups. Aside from PPO, other on-policy objectives such as REINFORCE [2] and variants like GRPO [59] and VinePPO [29] have also been studied in the context of language models. An alternative to PPO-based fine-tuning is rejection sampling fine-tuning, inspired by the best-of-n inference approach proposed by Nakano et al. [45]. Recent work by Dong et al. [10], Gulcehre et al. [15], and Wang et al. [72] extends this concept by generating n candidate responses for each prompt, ranking them with a learned reward function, and fine-tuning the model based on the highest-ranking responses. Separately, direct preference learning approaches [53, 3, 64] skip reward modeling entirely and train language models to directly optimize responses under a preference model. Hu et al. [23] introduced GFlowNet fine-tuning, leveraging off-policy GFlowNet algorithms for fine-tuning language models, which we build upon here.
Asynchronous distributed RL Distributed RL spreads actors/searchers, learners, and environments across a collection of computing resources. Asynchronous distributed RL does this such that searcher and trainer processes do not necessarily share the same weights, which can significantly improve training speed [44,73] and facilitate RL in complex, high-dimensional domains [19,24,21].
A foundational method in this area is Asynchronous Advantage Actor-Critic (A3C) [42]. In A3C, multiple parallel workers asynchronously interact with the environment and communicate gradients to a central node. Our approach to async distributed RL more closely resembles the Importance-Weighted Actor-Learner Architecture (IMPALA) method [11], which communicates experience trajectories (state, action, and reward tuples) to the central node. As opposed to standard RL envir
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31148afb-23e9-4e07-a23d-ae1bb27b444fCited by top-tier papers9
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language ReasoningWei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu et al.NeurIPS 2025 · 273 citations
- Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language ModelsYiran Guo, Lijie Xu, Ji Liu, Dan Ye et al.NeurIPS 2025 · 75 citations
- Steering Generative Models with Experimental Data for Protein Fitness OptimizationJason Yang, Wenda Chu, Daniel Khalil, Raul Astudillo et al.NeurIPS 2025 · 13 citations
- Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition FunctionsSimon Matrenok, Skander Moalla, Caglar GulcehreNeurIPS 2025 · 6 citations
- Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet TrainingXi Wang, Wenbo Lu, Shenji WanICML 2026 · 1 citation
Builds on33
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- D-ARL: A Distribution-Matched Asynchronous Reinforcement Learning Framework for Language Reasoning白 寅岐, Xialiang Tong, Jie Wang, Hongyu Liu et al.ICML 2026
- Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language ModelsMichael Noukhovitch, Shengyi Huang, Sophie Xhonneux, Arian Hosseini et al.ICLR 2025
- Laminar: A Scalable Asynchronous RL Post-Training FrameworkGuangming Sheng, Yuxuan Tong, Borui Wan, Wang Zhang et al.EuroSys 2026 · 2 citations
- Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMsLuke Huang, Zhuoyang Zhang, Qinghao Hu, Shang Yang et al.ICML 2026 · 3 citations
- Jackpot: Align Actor-Policy Distribution for scalable and stable RL for LLMZhuoming Chen, Hongyi Liu, Yang Zhou, Haizhong Zheng et al.ICLR 2026
