Discovering Preference Optimization Algorithms with and for Large Language Models
Chris Lu, Samuel Holt, Claudio Fanconi, Alex J. Chan, Jakob N. Foerster, Mihaela van der Schaar, Robert T. Lange
摘要
Offline preference optimization is a key method for enhancing and controlling the quality of Large Language Model (LLM) outputs. Typically, preference optimization is approached as an offline supervised learning task using manually-crafted convex loss functions. While these methods are based on theoretical insights, they are inherently constrained by human creativity, so the large search space of possible loss functions remains under explored. We address this by performing LLM-driven objective discovery to automatically discover new state-of-the-art preference optimization algorithms without (expert) human intervention. Specifically, we iteratively prompt an LLM to propose and implement new preference optimization loss functions based on previously-evaluated performance metrics. This process leads to the discovery of previously-unknown and performant preference optimization algorithms. The best performing of these we call Discovered Preference Optimization (DiscoPOP), a novel algorithm that adaptively blends logistic and exponential losses. Experiments demonstrate the state-of-the-art performance of DiscoPOP and its successful transfer to held-out tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program EvolutionRobert T. Lange, Yuki Imajuku, Edoardo CetinICLR 2026 · 被引用 162 次
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-benchEdan Toledo, Karen Hambardzumyan, Martin Josifoski, Rishi Hazra 等NeurIPS 2025 · 被引用 71 次
- Towards Execution-Grounded Automated AI ResearchChenglei Si, Zitong Yang, Yejin Choi, Emmanuel J Candes 等ICML 2026 · 被引用 12 次
- InfiGFusion: Graph-on-Logits Distillation via Efficient Gromov-Wasserstein for Model FusionYuanyi Wang, Zhaoyi Yan, Yiming Zhang, Qi Zhou 等NeurIPS 2025 · 被引用 12 次
- Meta-Learning Objectives for Preference OptimizationCarlo Alfano, Silvia Sapora, Jakob N. Foerster, Patrick Rebeschini 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper35
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
相关 Paper
- Explicit Preference Optimization: No Need for an Implicit Reward ModelXiangkun Hu, Lemin Kong, Tong He, David WipfICML 2025
- Pareto Prompt OptimizationGuang Zhao, Byung-Jun Yoon, Gilchan Park, Shantenu Jha 等ICLR 2025
- PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based SamplingYongchao Chen, Jacob Arkin, Yilun Hao, Yang Zhang 等EMNLP 2024 · 被引用 6 次
- MallowsPO: Fine-Tune Your LLM with Preference DispersionsHaoxian Chen, Hanyang Zhao, Henry Lam, David D. Yao 等ICLR 2025 · 被引用 1 次
- What Matters in Data for DPO?Yu Pan, Zhongze Cai, Huaiyang Zhong, Guanting Chen 等NeurIPS 2025 · 被引用 13 次
