Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models
Shaoning Sun, Mingzhu Cai, Huang He, Bingjin Chen, Siqi Bao, Yujiu Yang, Hua Wu, Haifeng Wang
摘要
Language model families exhibit striking disparity in their capacity to benefit from reinforcement learning: under identical training, models like Qwen achieve substantial gains, while others like Llama yield limited improvements. Complementing data-centric approaches, we reveal that this disparity reflects a hidden structural property: distributional clarity in probability space. Through a three-stage analysis-from phenomenon to mechanism to interpretation-we uncover that RL-friendly models exhibit intra-class compactness and inter-class separation in their probability assignments to correct vs. incorrect responses. We quantify this clarity using the Silhouette Coefficient () and demonstrate that (1) high correlates strongly with RL performance; (2) low is associated with severe logic errors and reasoning instability. To confirm this property, we introduce a Silhouette-Aware Reweighting strategy that prioritizes low- samples during training. Experiments across six mathematical benchmarks show consistent improvements across all model families, with gains up to 5.9 points on AIME24. Our work establishes distributional clarity as a fundamental, trainable property underlying RL-Friendliness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang 等NeurIPS 2025 · 被引用 533 次
- Reasoning with Exploration: An Entropy PerspectiveDaixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai 等AAAI 2026 · 被引用 216 次
- The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningXinyu Zhu, Mengzhou Xia, Zhepei Wei, Wei-Lin Chen 等NeurIPS 2025 · 被引用 177 次
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu 等EuroSys 2025 · 被引用 61 次
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsCharlie Zhang, Graham Neubig, Xiang YueICML 2026 · 被引用 58 次
相关 Paper
- PRISM: Demystifying Retention and Interaction in Mid-TrainingBharat Runwal, Ashish Agrawal, Anurag Roy, Rameswar PandaICML 2026
- Spurious Rewards: Rethinking Training Signals in RLVRRulin Shao, Stella Li, Rui Xin, Scott Geng 等ICML 2026
- CLARity: Reasoning Consistency Alone Can Teach Reinforced ExpertsJiuheng Lin, Cong Jiang, Zirui Wu, Jiarui Sun 等ACL 2026
- The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical ReasoningHaolong Qian, Xianliang Yang, Ma yinuo, Lirong Che 等ICML 2026
- Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence FunctionsSiqi Kou, Qingyuan Tian, Hanwen Xu, Zihao Zeng 等NeurIPS 2025 · 被引用 9 次
