SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
Xiao Liang, Zhong-Zhi Li, Yeyun Gong, Yang Wang, Hengyuan Zhang, Yelong Shen, Ying Nian Wu, Weizhu Chen
摘要
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for training large language models (LLMs) on complex reasoning tasks, such as mathematical problem solving. A prerequisite for the scalability of RLVR is a high-quality problem set with precise and verifiable answers. However, the scarcity of well-crafted human-labeled math problems and limited-verification answers in existing distillation-oriented synthetic datasets limit their effectiveness in RL. Additionally, most problem synthesis strategies indiscriminately expand the problem set without considering the model's capabilities, leading to low efficiency in generating useful questions. To mitigate this issue, we introduce a Self-aware Weakness-driven problem Synthesis framework (SwS) that systematically identifies model deficiencies and leverages them for problem augmentation. Specifically, we define weaknesses as questions that the model consistently fails to learn through its iterative sampling during RL training. We then extract the core concepts from these failure cases and synthesize new problems to strengthen the model's weak areas in subsequent augmented training, enabling it to focus on and gradually overcome its weaknesses. Without relying on external knowledge distillation, our framework enables robust generalization by empowering the model to self-identify and address its weaknesses in RL, yielding average performance gains of 10.0% and 7.7% on 7B and 32B models across eight mainstream reasoning benchmarks. Our code is available at https://github.com/MasterVito/SwS.
- Equal contribution. Work done during Xiao's and Zhongzhi's internships at Microsoft.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains RLVRXiao Liang, Zhong-Zhi Li, Yeyun Gong, Yelong Shen 等ICLR 2026 · 被引用 57 次
- Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its FriendsChaorui Yao, Yanxi Chen, Yuchang Sun, Yushuo Chen 等ICLR 2026 · 被引用 13 次
- PolySkill: Learning Generalizable Skills Through Polymorphic Abstraction For Continual LearningSimon Yu, Gang Li, Weiyan Shi, Peng QiICLR 2026 · 被引用 11 次
- Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual GuidanceXinrong Chen, Xu Chu, Yingmin Qiu, Hengyuan Zhang 等CVPR 2026 · 被引用 8 次
- Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time ScalabilityXiao Liang, Zhong-Zhi Li, Zhenghao Lin, Eric Hanchen Jiang 等ACL 2026 · 被引用 5 次
它引用的顶会 Paper31
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
相关 Paper
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
- Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive DomainsZhonghang Yuan, Zhefan Wang, Fang Hu, Zihong Chen 等ACL 2026
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language ModelsZizhuo Zhang, Jianing Zhu, Xinmu Ge, Zihua Zhao 等ICLR 2026 · 被引用 16 次
- Powering Verifiable Learning via Automated Evolutionary Data SynthesisHe Du, Bowen Li, Aijun Yang, Siyang He 等ACL 2026
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui 等NeurIPS 2025 · 被引用 20 次
