Understanding the Learning Dynamics of Alignment with Human Feedback
Shawn Im, Yixuan Li
摘要
Aligning large language models (LLMs) with human intentions has become a critical task for safely deploying models in real-world systems. While existing alignment approaches have seen empirical success, theoretically understanding how these methods affect model behavior remains an open question. Our work provides an initial attempt to theoretically analyze the learning dynamics of human preference alignment. We formally show how the distribution of preference datasets influences the rate of model updates and provide rigorous guarantees on the training accuracy. Our theory also reveals an intricate phenomenon where the optimization is prone to prioritizing certain behaviors with higher preference distinguishability. We empirically validate our findings on contemporary LLMs and alignment tasks, reinforcing our theoretical insights and shedding light on considerations for future alignment approaches. Disclaimer: This paper contains potentially offensive text; reader discretion is advised.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Scaling Laws for Reward Model Overoptimization in Direct Alignment AlgorithmsRafael Rafailov, Yaswanth Chittepu, Ryan Park, Harshit Sikchi 等NeurIPS 2024 · 被引用 169 次
- Group Robust Preference Optimization in Reward-free RLHFShyam Sundhar Ramesh, Yifan Hu, Iason Chaimalas, Viraj Mehta 等NeurIPS 2024 · 被引用 122 次
- Can DPO Learn Diverse Human Values? A Theoretical Scaling LawShawn Im, Sharon LiNeurIPS 2025 · 被引用 8 次
- Optimizing Diversity and Quality through Base-Aligned Model CollaborationYichen Wang, Chenghao Yang, Tenghao Huang, Muhao Chen 等ICML 2026 · 被引用 7 次
- Unintentional Unalignment: Likelihood Displacement in Direct Preference OptimizationNoam Razin, Sadhika Malladi, Adithya Bhaskar, Danqi Chen 等ICLR 2025
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud 等ICLR 2024 · 被引用 762 次
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji 等ICLR 2024 · 被引用 656 次
相关 Paper
- What Matters in Data for DPO?Yu Pan, Zhongze Cai, Huaiyang Zhong, Guanting Chen 等NeurIPS 2025 · 被引用 13 次
- Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language ModelsRei Higuchi, Taiji SuzukiICML 2025
- Leveraging robust optimization for llm alignment under distribution shiftsMingye Zhu, Yi Liu, Zheren Fu, Yongdong Zhang 等NeurIPS 2025 · 被引用 2 次
- Data Selection for LLM Alignment Using Fine-Grained PreferencesJia Zhang, Yao Liu, Chen-Xi Zhang, Yi Liu 等ICLR 2026 · 被引用 1 次
- On a Connection Between Imitation Learning and RLHFTeng Xiao, Yige Yuan, Mingxiao Li, Zhengyu Chen 等ICLR 2025
