Reward-free Alignment for Conflicting Objectives
Peter Chen, Xiaopeng Li, Xi Chen, Tianyi Lin
Abstract
Direct alignment methods are increasingly used to align large language models (LLMs) with human preferences. However, many real-world alignment problems involve multiple conflicting objectives, where naive aggregation of preferences can lead to unstable training and poor trade-offs. In particular, weighted loss methods may fail to identify update directions that simultaneously improve all objectives, and existing multi-objective approaches often rely on explicit reward models, introducing additional complexity and distorting user-specified preferences. Our first contribution is to propose a Reward-free Alignment framework for Conflicted Objectives (RACO) that directly leverages pairwise preference data and resolves gradient conflicts via a novel clipped variant of conflict-averse gradient descent. We provide convergence guarantees to Pareto-critical points that respect user-specified objective weights, and show that clipping can strictly improve convergence rate in two-objective cases. Second, we improve our method using some heuristics and conduct experiments to demonstrate the compatibility of the proposed framework for LLM alignment. Both qualitative and quantitative evaluations on multiobjective summarization and safety alignment tasks across multiple LLM families (Qwen 3, Llama 3, Gemma 3) show that our method consistently achieves better Pareto trade-offs compared to existing multi-objective alignment baselines. Warning: This paper contains examples of potentially harmful and sexually explicit content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed4d6f89-0bcb-4ddb-b691-49d420264efeBuilds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
Related papers
- Beyond Pairwise: Empowering LLM Alignment With (Ranked) Choice ModelingYuxuan Tang, Yifan FengICLR 2026 · 1 citation
- SAFT: Safety-Preserving Adaptation via Fine-Tuning Transfer for Large Language ModelsZhiwen Ruan, Yan Yang, Zhuocheng Liang, Yun Chen et al.KDD 2026
- Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language ModelsChengao Li, Hanyu Zhang, Yunkun Xu, Hongyan Xue et al.ACL 2025 · 13 citations
- Multi-Objective Preference Optimization: Improving Human Alignment of Generative ModelsAkhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng WenICML 2026 · 15 citations
- VPO: Reasoning Preferences Optimization Based on V-Usable InformationZecheng Wang, Chunshan Li, Yupeng Zhang, Han Liu et al.NeurIPS 2025 · 1 citation
