QuRL: Low-Precision Reinforcement Learning for Efficient Reasoning
Yuhang Li, Reena Elangovan, Xin Dong, Priyadarshini Panda, Brucek Khailany
摘要
Reinforcement learning with verifiable rewards (RLVR) has become a trending paradigm for training reasoning large language models (LLMs). However, due to the autoregressive decoding nature of LLMs, the rollout process becomes the efficiency bottleneck of RL training, consisting of up to 70% of the total training time. In this work, we propose Quantized Reinforcement Learning (QuRL) that uses a quantized actor for accelerating the rollout. We address two challenges in QuRL. First, we propose Adaptive Clipping Range (ACR) that dynamically adjusts the clipping ratio based on the policy ratio between the full-precision actor and the quantized actor, which is essential for mitigating long-term training collapse. Second, we identify the weight update problem, where weight changes between RL steps are extremely small, making it difficult for the quantization operation to capture them effectively. We mitigate this problem through the invariant scaling technique that reduces quantization noise and increases weight update. We evaluate our method with INT8 and FP8 quantization experiments on DeepScaleR and DAPO, and achieve 20% to 80% faster rollout during training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
相关 Paper
- Prune as You Generate: Online Rollout Pruning for Faster and Better RLVRHaobo Xu, Sirui Chen, Ruizhong Qiu, Yuchen Yan 等ACL 2026 · 被引用 6 次
- QeRL: Beyond Efficiency - Quantization-enhanced Reinforcement Learning for LLMsWei Huang, Yi Ge, Shuai Yang, Yicheng Xiao 等ICLR 2026 · 被引用 19 次
- Experience Augmented Policy Optimization for LLM ReasoningJinda Lu, Kexin Huang, Junkang Wu, Shuo Yang 等ICML 2026 · 被引用 2 次
- Resource-Efficient Reinforcement for Reasoning Large Language Models via Dynamic One-Shot Policy RefinementYunjian Zhang, Sudong Wang, Yang Li, Peiran Xu 等ICML 2026 · 被引用 4 次
- Contextual Rollout Bandits for Reinforcement Learning with Verifiable RewardsXiaodong Lu, Xiaohan Wang, Jiajun Chai, Guojun Yin 等ICML 2026 · 被引用 7 次
