T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Zhenyu Hou, Xin Lv, Rui Lu, Jiajie Zhang, Yujiang Li, Zijun Yao, Juanzi Li, Jie Tang, Yuxiao Dong
摘要
Large language models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks. However, existing approaches mainly rely on imitation learning and struggle to achieve effective test-time scaling. While reinforcement learning (RL) holds promise for enabling selfexploration, recent attempts yield modest improvements in complex reasoning. In this paper, we present T1 to scale RL by encouraging exploration and understand inference scaling. We first initialize the LLM using synthesized chainof-thought data that integrates trial-and-error and self-verification. To scale RL training, we promote increased sampling diversity through oversampling. We demonstrate that T1 with open LLMs as its base exhibits inference scaling behavior and achieves superior performance on challenging math reasoning benchmarks. More importantly, we present a simple strategy to examine inference scaling, where increased inference budgets directly lead to T1's better performance without any additional verification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- MobileRL: Online Agentic Reinforcement Learning for Mobile GUI AgentsYifan Xu, Xiao Liu, Xinghan Liu, Jiaqi Fu 等ICLR 2026 · 被引用 45 次
- EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-ForgetLiang Chen, Xueting Han, Qizhou Wang, Bo Han 等ICLR 2026 · 被引用 16 次
- RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-TrainingTao Ren, Jinyang Jiang, Hui Yang, Wan Tian 等ICLR 2026 · 被引用 8 次
- Learning Diverse Responses with Prefix-Conditioned Supervised Fine-TuningZhiyuan Fan, Guanqiao Chen, Yanyi Huang, Mingkuan Zhao 等ACL 2026
- : Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive EnvironmentsSangeun Park, Minhae KwonICML 2026
它引用的顶会 Paper15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
相关 Paper
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui 等NeurIPS 2025 · 被引用 20 次
- e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMsAmrith Setlur, Matthew Y. R. Yang, Charlie Victor Snell, Jeremiah Greer 等ICLR 2026 · 被引用 66 次
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 被引用 270 次
- Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo MethodsIsha Puri, Shivchander Sudalairaj, Guangxuan Xu, Abhishek Bhandwaldar 等NeurIPS 2025 · 被引用 8 次
- S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement LearningRuotian Ma, Peisong Wang, Cheng Liu, Xingyan Liu 等ACL 2025 · 被引用 13 次
