Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
Yiju Guo, Ganqu Cui, Lifan Yuan, Ning Ding, Zexu Sun, Bowen Sun, Huimin Chen, Ruobing Xie, Jie Zhou, Yankai Lin, Zhiyuan Liu, Maosong Sun
摘要
Alignment in artificial intelligence pursues the consistency between model responses and human preferences as well as values. In practice, the multifaceted nature of human preferences inadvertently introduces what is known as the "alignment tax"-a compromise where enhancements in alignment within one objective (e.g., harmlessness) can diminish performance in others (e.g., helpfulness). However, existing alignment techniques are mostly unidirectional, leading to sub-optimal trade-offs and poor flexibility over various objectives. To navigate this challenge, we argue the prominence of grounding LLMs with evident preferences. We introduce controllable preference optimization (CPO), which explicitly specifies preference scores for different objectives, thereby guiding the model to generate responses that meet the requirements. Our experimental analysis reveals that the aligned models can provide responses that match various preferences among the "3H" (helpfulness, honesty, harmlessness) desiderata. Furthermore, by introducing diverse data and alignment goals, we surpass baseline methods in aligning with single objectives, hence mitigating the impact of the alignment tax and achieving improvements in multi-objective alignment. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- Decoding-Time Language Model Alignment with Multiple ObjectivesRuizhe Shi, Yifang Chen, Yushi Hu, Alisa Liu 等NeurIPS 2024 · 被引用 111 次
- Panacea: Pareto Alignment via Preference Adaptation for LLMsYifan Zhong, Chengdong Ma, Xiaoyuan Zhang, Ziran Yang 等NeurIPS 2024 · 被引用 89 次
- MetaAligner: Towards Generalizable Multi-Objective Alignment of Language ModelsKailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang 等NeurIPS 2024 · 被引用 49 次
- Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model MergingJinluan Yang, Dingnan Jin, Anke Tang, Li Shen 等NeurIPS 2025 · 被引用 23 次
- Multi-Objective Preference Optimization: Improving Human Alignment of Generative ModelsAkhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng WenICML 2026 · 被引用 15 次
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
相关 Paper
- Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language ModelsChengao Li, Hanyu Zhang, Yunkun Xu, Hongyan Xue 等ACL 2025 · 被引用 13 次
- Geometric-Averaged Preference Optimization for Soft Preference LabelsHiroki Furuta, Kuang-Huei Lee, Shixiang Shane Gu, Yutaka Matsuo 等NeurIPS 2024 · 被引用 24 次
- Dissecting Human and LLM PreferencesJunlong Li, Fan Zhou, Shichao Sun, Yikai Zhang 等ACL 2024 · 被引用 1 次
- Conflict-Aware Adaptive Alignment for LLM Hallucination MitigationRuohan Zong, Yang Zhang, WangICML 2026
- Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and FeedbackSongyang Gao, Qiming Ge, Wei Shen, Shihan Dou 等ICML 2024 · 被引用 24 次
