Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
Haoming Meng, Kexin Huang, Shaohang Wei, Chiyu Ma, Shuo Yang, Xue Wang, Guoyin Wang, Bolin Ding, Jingren Zhou
摘要
Reinforcement learning with verifiable rewards (RLVR) has significantly improved reasoning in large language models (LLMs), yet the token-level mechanisms underlying these improvements remain unclear. We present a systematic empirical study of RLVR's distributional effects organized around three main analyses: (1) token-level characterization of distributional shifts between base and RL models, (2) the impact of token-level distributional shifts on reasoning performance through cross-sampling interventions, and (3) fine-grained mechanics of these shifts at the token level. We find that RL fine-tuning induces highly sparse and targeted changes, with only a small fraction of token distributions exhibiting meaningful divergence. We further characterize the structure of these shifts through analyses of token entropy, positional concentration, and reallocation of probability mass. To assess the functional importance of these sparse changes, we conduct cross-sampling experiments that selectively swap token choices between the base and RL models. Inserting only a small fraction of RL-sampled tokens into base generations progressively recovers RL performance gains, while injecting a similarly small number of base token choices into RL-generated responses collapses performance to base levels, isolating a sparse set of token-level decisions directly responsible for RLVR's improvements. Finally, we explore divergenceweighted variants of the advantage signal as a diagnostic intervention, finding that they can yield improvements over baselines. Together, our results shed light on the distributional changes induced by RLVR and provide a fine-grained, token-level lens for understanding RLVR as a targeted refinement process.
Recent work has begun analyzing RL fine-tuning through token-level entropy and uncertainty perspectives (Wang et al., 2025;Cheng et al., 2025;Cui et al., 2025), highlighting the role of highentropy tokens and exploration dynamics. However, a more detailed distributional view remains missing: how such shifts are structured across positions and contexts, how probability mass is reallocated across candidate tokens, how they evolve over training, and to what extent are they responsible for RLVR's performance gains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- One-Way Policy Optimization for Self-Evolving LLMsShuo Yang, Jinda Lu, Kexin Huang, Chiyu Ma 等ICML 2026 · 被引用 2 次
- Beyond Magnitude: Leveraging Direction of RLVR Updates for LLM ReasoningKexin Huang, Haoming Meng, Junkang Wu, Jinda Lu 等ICLR 2026
- Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary SignalsShuo Yang, Jinda Lu, Chiyu Ma, Kexin Huang 等ICML 2026
- LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?Jingyuan Wang, Yankai Chen, Zhonghang Li, Chao HuangACL 2026
它引用的顶会 Paper20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
相关 Paper
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 等NeurIPS 2025 · 被引用 592 次
- Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveZhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo 等ACL 2026 · 被引用 42 次
- Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern SelectionXingwu Chen, Tianle Li, Difan ZouICLR 2026 · 被引用 8 次
- SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for ReasoningYuqian Fu, Tinghong Chen, Jiajun Chai, Xihuai Wang 等ICLR 2026 · 被引用 97 次
- Parameter-Efficient Reinforcement Learning using Prefix OptimizationItamar Rocha Filho, Rosie Zhao, Sham M. Kakade, Eran Malach 等ICLR 2026
