Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process
Ermo Hua, Biqing Qi, Kaiyan Zhang, Kai Tian, Xingtai Lv, Ning Ding, Bowen Zhou
摘要
Supervised Fine-Tuning (SFT) and Preference Optimization (PO) are key processes for aligning Language Models (LMs) with human preferences post pre-training. While SFT excels in efficiency and PO in effectiveness, they are often combined sequentially without integrating their optimization objectives. This approach ignores the opportunities to bridge their paradigm gap and take the strengths from both. In this paper, we interpret SFT and PO with two subprocesses -Preference Estimation and Transition Optimization -defined at token level within the Markov Decision Process (MDP). This modeling shows that SFT is only a special case of PO with inferior estimation and optimization. PO estimates the model's preference by its entire generation, while SFT only scores model's subsequent predicted tokens based on prior tokens from ground truth answer. These priors deviates from model's distribution, hindering the preference estimation and transition optimization. Building on this view, we introduce Intuitive Fine-Tuning (IFT) to integrate SFT and PO into a single process. Through a temporal residual connection, IFT brings better estimation and optimization by capturing LMs' intuitive sense of its entire answers. But it solely relies on a single policy and the same volume of non-preference-labeled data as SFT. Our experiments show that IFT performs comparably or even superiorly to SFT and some typical PO methods across several tasks, particularly those requires generation, reasoning, and fact-following abilities. An explainable Frozen Lake game further validates the effectiveness of IFT for getting competitive policy. Code is available at https://github.com/ TsinghuaC3I/Intuitive-Fine-Tuning .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
相关 Paper
- Preference-Oriented Supervised Fine-Tuning: Favoring Target Model over Aligned Large Language ModelsYuchen Fan, Yuzhong Hong, Qiushi Wang, Junwei Bao 等AAAI 2025 · 被引用 7 次
- Implicit Reward as the Bridge: A Unified View of SFT and DPO ConnectionsBo Wang, Qinyuan Cheng, Runyu Peng, Rong Bao 等NeurIPS 2025 · 被引用 23 次
- Discriminative Finetuning of Generative Large Language Models without Reward Models and Human Preference DataSiqi Guo, Ilgee Hong, Vicente Balmaseda, Changlong Yu 等ICML 2025
- Expectation Preference Optimization: Reliable Preference Estimation for Improving the Reasoning Capability of Large Language ModelsZelin Li, Dawei SongEMNLP 2025
- Token-level Direct Preference OptimizationYongcheng Zeng, Guoqing Liu, Weiyu Ma, Ning Yang 等ICML 2024 · 被引用 136 次
