Mirage or Method? How Model–Task Alignment Induces Divergent RL Conclusions
Haoze Wu, Cheng Wang, Wenshuo Zhao, Junxian He
Abstract
Recent advances in applying reinforcement learning (RL) to large language models (LLMs) have led to substantial progress. In particular, a series of remarkable yet often counterintuitive phenomena have been reported in LLMs, exhibiting patterns not typically observed in traditional RL settings. For example, notable claims include that a single training example can match the performance achieved with an entire dataset, that the reward signal does not need to be very accurate, and that training solely with negative samples can match or even surpass sophisticated reward-based methods. However, the precise conditions under which these observations hold-and, critically, when they fail-remain unclear. In this work, we identify a key factor that differentiates RL observations: whether the pretrained model already exhibits strong Model-Task Alignment, as measured by pass@k accuracy on the evaluated task. Through a systematic and comprehensive examination of a series of counterintuitive claims, supported by rigorous experimental validation across different model architectures and task domains, our findings show that while standard RL training remains consistently robust across settings, many of these counterintuitive results arise only when the model and task already exhibit strong model-task alignment. In contrast, these techniques fail to drive substantial learning in more challenging regimes, where standard RL methods remain effective. Code is available at https://github.com/hkust-nlp/model-task-align-rl .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b083fec9-7a30-40e5-82fa-88fac0d25782Cited by top-tier papers4
- From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old OnesLifan Yuan, Weize Chen, Yuchen Zhang, Ganqu Cui et al.ICLR 2026 · 46 citations
- How Far Can Unsupervised RLVR Scale LLM Training?Bingxiang He, Yuxin Zuo, Zeyuan Liu, Shangziqi Zhao et al.ICLR 2026 · 31 citations
- Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMsLecheng Yan, Ruizhe Li, Guanhua CHEN, Qing Li et al.ICML 2026 · 8 citations
- ReCode: Updating Code API Knowledge with Reinforcement LearningHaoze Wu, Yunzhi Yao, Wenhao Yu, Ningyu ZhangAAAI 2026 · 7 citations
Builds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleYiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren et al.NeurIPS 2025 · 314 citations
- The Unreasonable Effectiveness of Entropy Minimization in LLM ReasoningShivam Agarwal, Zimin Zhang, Lifan Yuan, Jiawei Han et al.NeurIPS 2025 · 185 citations
Related papers
- FLAME : Factuality-Aware Alignment for Large Language ModelsSheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong et al.NeurIPS 2024 · 63 citations
- Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewardsCharles Arnal, Gaëtan Narozniak, Vivien Cabannes, Yunhao Tang et al.NeurIPS 2025 · 30 citations
- Unsupervised Large Language Model Alignment for Information Retrieval via Contrastive FeedbackQian Dong, Yiding Liu, Qingyao Ai, Zhijing Wu et al.SIGIR 2024 · 9 citations
- On the Modeling Capabilities of Large Language Models for Sequential Decision MakingMartin Klissarov, R. Devon Hjelm, Alexander T. Toshev, Bogdan MazoureICLR 2025
- Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model AlignmentYuang Cai, Yuyu Yuan, Jinsheng Shi, Qinhong LinAAAI 2025 · 5 citations
