ACL2026

Understanding Conflicts in Multi-Objective Alignment through Reward Consistency

Zhihao Xu, Yongqi Tong, Xin Zhang, Jun Zhou, Xiting Wang

Abstract

Multi-objective preference alignment often faces alignment conflicts, where optimizing for one objective degrades performance on others. While prior work focuses on algorithmic solutions, the intrinsic conflict within data and its theoretical impact on training remain underexplored. To bridge this gap, we introduce the principle of REWARD CONSISTENCY (RC), a theory-grounded criterion that approximates the alignment conflicts via reward models. We prove that a sample mitigates conflicts if and only if it satisfies RC, thereby ensuring improvement across all objectives during optimization. Building on this, we propose RE-WARD CONSISTENCY SAMPLING (RCS), an automated framework for constructing pairwise data that adheres to RC, supplemented by a relaxation strategy to enhance flexibility. Extensive experiments show that RCS brings significant and consistent performance gains, achieving an average improvement of 23.07% in both harmlessness and helpfulness during simultaneous optimization compared to the vanilla dataset. Our data-centric approach is complementary to existing alignment algorithms and effective in both sequential and simultaneous optimization scenarios.