RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
Kaiwen Zha, Zhengqi Gao, Maohao Shen, Zhang-Wei Hong, Duane S. Boning, Dina Katabi
摘要
Reinforcement learning (RL) has recently emerged as a compelling approach for enhancing the reasoning capabilities of large language models (LLMs), where an LLM generator serves as a policy guided by a verifier (reward model). However, current RL post-training methods for LLMs typically use verifiers that are fixed (rule-based or frozen pretrained) or trained discriminatively via supervised finetuning (SFT). Such designs are susceptible to reward hacking and generalize poorly beyond their training distributions. To overcome these limitations, we propose TANGO, a novel framework that uses RL to concurrently train both an LLM generator and a verifier in an interleaved manner. A central innovation of TANGO is its generative, process-level LLM verifier, which is trained via RL and co-evolves with the generator. Importantly, the verifier is trained solely based on outcome-level verification correctness rewards without requiring explicit process-level annotations. This generative RL-trained verifier exhibits improved robustness and superior generalization compared to deterministic or SFT-trained verifiers, fostering effective mutual reinforcement with the generator. Extensive experiments demonstrate that both components of TANGO achieve state-of-the-art results among 7B/8B-scale models: the generator attains best-in-class performance across five competition-level math benchmarks and four challenging out-of-domain reasoning tasks, while the verifier leads on the ProcessBench dataset. Remarkably, both components exhibit particularly substantial improvements on the most difficult mathematical reasoning problems. Code is at: https://github.com/kaiwenzha/rl-tango.
Reinforcing Generator & Verifier Together (Tango) Verifier Warmup 9.1% * Equal contribution. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Agentic Reinforcement Learning with Implicit Step RewardsXiaoqian Liu, Ke Wang, Yuchuan Wu, Fei Huang 等ICLR 2026 · 被引用 46 次
- RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL SystemYinjie Wang, Tianbao Xie, Ke Shen, Mengdi Wang 等ICML 2026 · 被引用 15 次
- Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning ModelsRunze Liu, Jiakang Wang, Yuling Shi, Zhihui Xie 等ICLR 2026 · 被引用 13 次
- Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier MathShrey Pandit, Austin Xu, Xuan-Phi Nguyen, Yifei Ming 等ACL 2026 · 被引用 13 次
- LaSeR: Reinforcement Learning with Last-Token Self-RewardingWenkai Yang, Weijie Liu, Ruobing Xie, Yiju Guo 等ICLR 2026 · 被引用 12 次
它引用的顶会 Paper26
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
相关 Paper
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language ModelsDan Shi, Zhuowen Han, Simon Ostermann, Renren Jin 等ACL 2026 · 被引用 1 次
- Generative Verifiers: Reward Modeling as Next-Token PredictionLunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi 等ICLR 2025
- General-Reasoner: Advancing LLM Reasoning Across All DomainsXueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang 等NeurIPS 2025 · 被引用 153 次
- StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language ModelsChenyu Zhou, Tianyi Xu, Jianghao Lin, Dongdong GeICLR 2026 · 被引用 29 次
- Native Reasoning Models: Training Language Models to Reason on Unverifiable DataYuanfu Wang, Zhixuan Liu, Li xiangtian, Chaochao Lu 等ICLR 2026 · 被引用 3 次
