Don't Force the Fit: Bounded Log-Likelihood Loss for Enhanced Reasoning in Large Language Models
Feng Zhao, Hong Zhang, Yu Yang, Ruilin Zhao, Guandong Xu
摘要
Supervised fine-tuning (SFT) is central to aligning large language models (LLMs) with instruction following and task-specific reasoning. Despite its success, SFT optimizes token-level likelihoods under the implicit assumption that strictly fitting all tokens in expert demonstrations induces the desired downstream behavior. However, in reasoning tasks where correctness is defined by logical validity or final outcomes rather than exact token realizations, this assumption can lead to optimization misalignment. We empirically observe that low-probability tokens in reasoning demonstrations often correspond to realization-specific or stylistic variations, and that reducing their influence during training consistently improves generalization on reasoning benchmarks. Motivated by this insight, we propose the Bounded Log-Likelihood Loss (BLL-Loss), a simple and parameter-free alternative to standard likelihood training that bounds gradient contributions from low-probability tokens while preserving conventional optimization behavior. We provide theoretical insights and extensive empirical results demonstrating that BLL-Loss improves reasoning generalization across diverse model scales and challenging benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
相关 Paper
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng 等ICLR 2026 · 被引用 130 次
- VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought SupervisionXuan Gong, Senmiao Wang, Hanbo Huang, Ruoyu Sun 等ACL 2026
- The Emperor's New Reasoning: Format Imitation Overshadows Genuine Mathematical Understanding in SFTLinyao Yang, Jian-Tao Huang, Yafei Lu, Zhenhui Jessie Li 等EMNLP 2025
- Clipping Low-Probability Tokens in SFT Yields a Generalizable Initialization for RLTian-Shuo Liu, Chengxing Jia, Haoyu Liu, Pengyuan Wang 等ICML 2026
- Navigating the Pareto Frontier of Alignment: Spectrum-Adaptive Fine-Tuning for LLMsYaoyou Fan, Chao Zhang, Xiaoyu Tan, Chenxing Sun 等ICML 2026
