Lune

NeurIPS2025顶会

Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training

Brian R. Bartoldson, Siddarth Venkatraman, James Diffenderfer, Moksh Jain, Tal Ben-Nun, Seanie Lee, Minsu Kim, Johan S. Obando-Ceron, Yoshua Bengio, Bhavya Kailkhura

2025年份
34被引次数
9顶会引用

摘要

Reinforcement learning (RL) is a critical component of large language model (LLM) post-training. However, on-policy algorithms used for post-training are not naturally robust to a diversified content of experience replay buffers, which asynchronous off-policy actors can efficiently populate in parallel to training. We propose efficiently learning on such off-policy data via Trajectory Balance with Asynchrony (TBA), an approach to asynchronous RL for LLMs that leverages the principled off-policy TB objective. On math, preference-tuning, and automated red-teaming tasks, we post-train models ranging from Pythia 410M to Qwen 2.5 7B, finding TBA offers speed and performance boosts over strong baselines like Online DPO and Dr. GRPO. Beyond TBA's performance benefits (high accuracy even as asynchrony grows) and speedups (4× or more), we show its reward-and recency-prioritizing sampling enable further gains as data generation is scaled. Our code is available at https://github.com/bbartoldson/TBA. 128 256 512 1024 2048 4096 Compute Time (Minutes) 49% 50% 51% 52% 53% 54% 55% GSM8K Test Pass@1 GRPO PPO RLOO Online DPO Online DPO VinePPO TBA (Ours) 50x faster 1.6x faster +2.0% +1.2%

Figure 1: TBA excels on the GSM8K mathematical reasoning task. All plotted points use 4xA100 GPUs (or comparable L40S GPUs). DPO and RLOO baselines taken from Noukhovitch et al. [46], PPO and VinePPO baselines taken from Kazemnejad et al. [29]. The baseline model is the SFTed RhoMath-1B [35] model, which gets 40% accuracy after SFT and before RL. Appendix B has details. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

Across mathematical reasoning, preference-tuning, and automated red-teaming tasks, we find TBA broadly offers three advantages over existing LLM post-training approaches: 1) Stable off-policy RL, unlocking asynchrony for massive parallelization and reduced wall-clock times -see Figures 1 and 3.

  1. Sampling from a diverse replay buffer, improving exploration and preventing mode collapse. 3) Scalable search compatibility that aids sparse reward settings like automated red-teaming.

Our key contributions are summarized as follows:

• We introduce TBA, a distributed asynchronous RL framework for fast, scalable LLM post-training.

• We show that LLMs post-trained with TBA match or exceed performances from existing methods, illustrating TB's ability to use off-policy data in asynchronous RL.

• We demonstrate significant speedups (4× to 50×) for RL across mathematical reasoning, preference-tuning, and automated red-teaming.

By enabling high-quality and fast off-policy post-training, TBA contributes to scalable and effective LLM alignment, ensuring that large models can be refined more efficiently for real-world deployment.

2 Related Work RL fine-tuning of language models Reinforcement learning (RL) has been an integral component for training LLMs [14, 48]. In particular, RL has become the de facto approach for aligning language models with human preferences [8, 82, 62, 49]. Much of this work relies on Proximal Policy Optimization (PPO) [57], an on-policy RL algorithm that has become a default choice for fine-tuning LLMs due to its strong performance across different setups. Aside from PPO, other on-policy objectives such as REINFORCE [2] and variants like GRPO [59] and VinePPO [29] have also been studied in the context of language models. An alternative to PPO-based fine-tuning is rejection sampling fine-tuning, inspired by the best-of-n inference approach proposed by Nakano et al. [45]. Recent work by Dong et al. [10], Gulcehre et al. [15], and Wang et al. [72] extends this concept by generating n candidate responses for each prompt, ranking them with a learned reward function, and fine-tuning the model based on the highest-ranking responses. Separately, direct preference learning approaches [53, 3, 64] skip reward modeling entirely and train language models to directly optimize responses under a preference model. Hu et al. [23] introduced GFlowNet fine-tuning, leveraging off-policy GFlowNet algorithms for fine-tuning language models, which we build upon here.

Asynchronous distributed RL Distributed RL spreads actors/searchers, learners, and environments across a collection of computing resources. Asynchronous distributed RL does this such that searcher and trainer processes do not necessarily share the same weights, which can significantly improve training speed [44,73] and facilitate RL in complex, high-dimensional domains [19,24,21].

A foundational method in this area is Asynchronous Advantage Actor-Critic (A3C) [42]. In A3C, multiple parallel workers asynchronously interact with the environment and communicate gradients to a central node. Our approach to async distributed RL more closely resembles the Importance-Weighted Actor-Learner Architecture (IMPALA) method [11], which communicates experience trajectories (state, action, and reward tuples) to the central node. As opposed to standard RL envir

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 31148afb-23e9-4e07-a23d-ae1bb27b444f

引用它的顶会 Paper9

问问它们各自怎么用它

它引用的顶会 Paper33

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖