Multi-Value Alignment for LLMs via Value Decorrelation and Extrapolation
Hefei Xu, Le Wu, Chen Cheng, Hao Liu
Abstract
With the rapid advancement of large language models (LLMs), aligning them with human values for safety and ethics has become a critical challenge. This problem is especially challenging when multiple, potentially conflicting human values must be considered and balanced. Although several variants of existing alignment methods (such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO)) have been proposed to address multi-value alignment, they suffer from notable limitations: 1) they are often unstable and inefficient in multi-value optimization; and 2) they fail to effectively handle value conflicts. As a result, these approaches typically struggle to achieve optimal trade-offs when aligning multiple values.
To address this challenge, we propose a novel framework called Multi-Value Alignment (MVA). It mitigates alignment degradation caused by parameter interference among diverse human values by minimizing their mutual information. Furthermore, we propose a value extrapolation strategy to efficiently explore the Pareto frontier, thereby constructing a set of LLMs with diverse value preferences. Extensive experiments demonstrate that MVA consistently outperforms existing baselines in aligning LLMs with multiple human values.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2261fd5-d6ca-4ce3-b4af-3d39ce086774Cited by top-tier papers1
Ask how each one uses itBuilds on10
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri et al.NeurIPS 2023 · 516 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
- ULTRAFEEDBACK: Boosting Language Models with Scaled AI FeedbackGanqu Cui, Lifan Yuan, Ning Ding, Guanming Yao et al.ICML 2024 · 286 citations
- The HSIC Bottleneck: Deep Learning without Back-PropagationKurt Wan-Duo Ma, J. P. Lewis, W. Bastiaan KleijnAAAI 2020 · 180 citations
Related papers
- Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language ModelsChengao Li, Hanyu Zhang, Yunkun Xu, Hongyan Xue et al.ACL 2025 · 13 citations
- MACPO: Weak-to-Strong Alignment via Multi-Agent Contrastive Preference OptimizationYougang Lyu, Lingyong Yan, Zihan Wang, Dawei Yin et al.ICLR 2025
- Disentangling Consensus and Value-Specific Representations for Controllable Pluralistic Value Alignment of LLMsJianKui Zhou, Jing Yao, Xiaoyuan Yi, Peng Zhang et al.ICML 2026
- Multi-Reference Preference Optimization for Large Language ModelsHung Le, Quan Hung Tran, Dung Nguyen, Kien Do et al.AAAI 2025 · 6 citations
- The Hidden Link between RLHF and Contrastive LearningXufei Lv, Kehai Chen, Haoyuan Sun, Xuefeng Bai et al.ICML 2026
