LIONs: An Empirically Optimized Approach to Align Language Models
Xiao Yu, Qingyang Wu, Yu Li, Zhou Yu
摘要
Alignment is a crucial step to enhance the instruction-following and conversational abilities of language models. Despite many recent work proposing new algorithms, datasets, and training pipelines, there is a lack of comprehensive studies measuring the impact of various design choices throughout the whole training process. We first conduct a rigorous analysis over a three-stage training pipeline consisting of supervised fine-tuning, offline preference learning, and online preference learning. We have found that using techniques like sequence packing, loss masking in SFT, increasing the preference dataset size in DPO, and online DPO training can significantly improve the performance of language models. We then train from Gemma-2b-base and LLama-3-8b-base, and find that our best models exceed the performance of the official instruct models tuned with closed-source data and algorithms. Our code and models can be found at https://github. com/Columbia-NLP-Lab/LionAlignment .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Aligning Large Language Models with Implicit Preferences from User-Generated ContentZhaoxuan Tan, Zheng Li, Tianyi Liu, Haodong Wang 等ACL 2025 · 被引用 10 次
- Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric ReasoningMassimiliano Pronesti, Michela Lorandi, Paul Flanagan, Oisin Redmond 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- Is On-Policy Data always the Best Choice for Direct Preference Optimization-Based LM Alignment?Zetian Sun, Dongfang Li, Xuhui Chen, Baotian Hu 等ICLR 2026 · 被引用 1 次
- Beyond Pairwise: Empowering LLM Alignment With (Ranked) Choice ModelingYuxuan Tang, Yifan FengICLR 2026 · 被引用 1 次
- PIPA: Preference Alignment as Prior-Informed Statistical EstimationJunbo Li, Zhangyang Wang, Qiang LiuICML 2025
- Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code GenerationSanjeepan Sivapiran, Gias UddinFSE 2026
- A Deep Dive into the Trade-Offs of Parameter-Efficient Preference Alignment TechniquesMegh Thakkar, Quentin Fournier, Matthew Riemer, Pin-Yu Chen 等ACL 2024
