Learning Dynamics of LLM Finetuning
Yi Ren, Danica J. Sutherland
Abstract
Learning dynamics, which describes how the learning of specific training examples influences the model's predictions on other examples, gives us a powerful tool for understanding the behavior of deep learning systems. We study the learning dynamics of large language models during different types of finetuning, by analyzing the step-wise decomposition of how influence accumulates among different potential responses. Our framework allows a uniform interpretation of many interesting observations about the training of popular algorithms for both instruction tuning and preference tuning. In particular, we propose a hypothetical explanation of why specific types of hallucination are strengthened after finetuning, e.g., the model might use phrases or facts in the response for question B to answer question A, or the model might keep repeating similar simple phrases when generating responses. We also extend our framework and highlight a unique "squeezing effect" to explain a previously observed phenomenon in off-policy direct preference optimization (DPO), where running DPO for too long makes even the desired outputs less likely. This framework also provides insights into where the benefits of on-policy DPO and other variants come from. The analysis not only provides a novel perspective of understanding LLM's finetuning but also inspires a simple, effective method to improve alignment performance. Code for experiments is available at https://github.com/Joshua-Ren/Learning_dynamics_LLM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 778a7000-6557-4cb7-ac96-8ea00b836fdfCited by top-tier papers57
- Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMsZhihe Yang, Xufang Luo, Zilong Wang, Dongqi Han et al.ICLR 2026 · 47 citations
- Tracing the Representation Geometry of Language Models from Pretraining to Post-trainingMelody Zixuan Li, Kumar Krishna Agrawal, Arna Ghosh, Komal Kumar Teru et al.NeurIPS 2025 · 38 citations
- On the Effect of Negative Gradient in Group Relative Deep Reinforcement OptimizationWenlong Deng, Yi Ren, Muchen Li, Danica J. Sutherland et al.NeurIPS 2025 · 36 citations
- Implicit Reward as the Bridge: A Unified View of SFT and DPO ConnectionsBo Wang, Qinyuan Cheng, Runyu Peng, Rong Bao et al.NeurIPS 2025 · 23 citations
- Bias Amplification in Language Model Evolution: An Iterated Learning PerspectiveYi Ren, Shangmin Guo, Linlu Qiu, Bailin Wang et al.NeurIPS 2024 · 23 citations
Builds on26
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma et al.ICLR 2022 · 911 citations
Related papers
- FLAME : Factuality-Aware Alignment for Large Language ModelsSheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong et al.NeurIPS 2024 · 63 citations
- Measuring memorization in RLHF for code completionJamie Hayes, Ilia Shumailov, William P. Porter, Aneesh PappuICLR 2025
- What Matters in Data for DPO?Yu Pan, Zhongze Cai, Huaiyang Zhong, Guanting Chen et al.NeurIPS 2025 · 13 citations
- Is On-Policy Data always the Best Choice for Direct Preference Optimization-Based LM Alignment?Zetian Sun, Dongfang Li, Xuhui Chen, Baotian Hu et al.ICLR 2026 · 1 citation
- Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial RegularizerZhihan Liu, Miao Lu, Shenao Zhang, Boyi Liu et al.NeurIPS 2024 · 119 citations
