rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset
Yifei Liu, Li Lyna Zhang, Yi Zhu, Bingcheng Dong, Xudong Zhou, Ning Shang, Fan Yang, Cheng Li, Mao Yang
摘要
Advancing code reasoning in large language models (LLMs) is fundamentally limited by the scarcity of high-difficulty datasets, especially those with verifiable input-output test cases necessary for rigorous solution validation at scale. We introduce rStar-Coder, which significantly improves LLM code reasoning capabilities by constructing a large-scale, verified dataset of 418K competitionlevel code problems, 580K long-reasoning solutions along with rich test cases of varying difficulty. This is achieved through three core contributions: (1) we curate competitive programming code problems and solutions to synthesize new, solvable problems; (2) we introduce a reliable input-output test case synthesis pipeline that decouples the generation into a three-step input generation method and a mutual verification mechanism for effective output labeling; (3) we augment problems with high-quality, test-case-verified long-reasoning solutions. Extensive experiments on Qwen models (1.5B-14B) across various code reasoning benchmarks demonstrate the superiority of rStar-Coder dataset, achieving leading performance comparable to frontier reasoning LLMs with significantly smaller model sizes. On LiveCodeBench, rStar-Coder improves Qwen2.5-7B from 17.4% to an impressive 57.3%, and Qwen2.5-14B from 23.3% to 62.5%, surpassing o3-mini (low) by 3.1%. On the more challenging USA Computing Olympiad, our 7B model achieves an average pass@1 accuracy of 16.15%, outperforming the frontier-level QWQ-32B. rStar-Coder dataset is publicly available at https://huggingface.co/datasets/microsoft/rStar-Coder.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Toward Training Superintelligent Software Agents through Self-Play SWE-RLYuxiang Wei, Zhiqing Sun, Emily McMilin, Jonas Gehring 等ICML 2026 · 被引用 32 次
- LoongRL: Reinforcement Learning for Advanced Reasoning over Long ContextsSiyuan Wang, Gaokai Zhang, Li Lyna Zhang, Ning Shang 等ICLR 2026 · 被引用 28 次
- RPG: A Repository Planning Graph for Unified and Scalable Codebase GenerationJane Luo, Xin Zhang, Steven Liu, Jie Wu 等ICLR 2026 · 被引用 18 次
- AutoCode: LLMs as Problem Setters for Competitive ProgrammingShang Zhou, Zihan Zheng, Kaiyuan Liu, Zeyu Shen 等ICLR 2026 · 被引用 12 次
- Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning TrainingZheng Xin Yong, Stephen H. BachICLR 2026 · 被引用 10 次
它引用的顶会 Paper12
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- Magicoder: Empowering Code Generation with OSS-InstructYuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding 等ICML 2024 · 被引用 246 次
- Key-Point-Driven Data Synthesis with Its Enhancement on Mathematical ReasoningYiming Huang, Xiao Liu, Yeyun Gong, Zhibin Gou 等AAAI 2025 · 被引用 74 次
相关 Paper
- HARDTESTGEN: A High-Quality RL Verifier Generation Pipeline for LLM Algorithimic CodingZhongmou He, Yee Man Choi, Kexun Zhang, Ivan Bercovich 等ICLR 2026
- Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model ReasoningHonglin Lin, Qizhi Pei, Zhuoshi Pan, Yu Li 等NeurIPS 2025 · 被引用 12 次
- rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep ThinkingXinyu Guan, Li Lyna Zhang, Yifei Liu, Ning Shang 等ICML 2025
- Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMsDayu Yang, Tianyang Liu, Daoan Zhang, Antoine Simoulin 等EMNLP 2025 · 被引用 1 次
- LogicPro: Improving Complex Logical Reasoning via Program-Guided LearningJin Jiang, Yuchen Yan, Yang Liu, Jianing Wang 等ACL 2025
