Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
Jinluan Yang, Dingnan Jin, Anke Tang, Li Shen, Didi Zhu, Zhengyu Chen, Ziyu Zhao, Daixin Wang, Qing Cui, Zhiqiang Zhang, Jun Zhou, Fei Wu, Kun Kuang
Abstract
Achieving balanced alignment of large language models (LLMs) in terms of Helpfulness, Honesty, and Harmlessness (3H optimization) constitutes a cornerstone of responsible AI. Existing methods like data mixture strategies face limitations, including heavy reliance on expert knowledge and conflicting optimization signals. While model merging offers parameter-level conflict-resolution strategies through integrating specialized models'parameters, its potential for 3H optimization remains underexplored. This paper systematically compares the effectiveness of model merging and data mixture methods in constructing 3H-aligned LLMs for the first time, revealing previously overlooked collaborative and conflict relationships among the 3H dimensions and discussing the advantages and drawbacks of data mixture (data-level) and model merging (parameter-level) methods in mitigating the conflict for balanced 3H optimization. Specially, we propose a novel Reweighting Enhanced task Singular Merging method, RESM, through outlier weighting and sparsity-aware rank selection strategies to address the challenges of preference noise accumulation and layer sparsity adaptation inherent in 3H-aligned LLM merging. Extensive evaluations can verify the effectiveness and robustness of RESM compared to previous data mixture (2%-5% gain) and model merging (1%-3% gain) methods in achieving balanced LLM alignment. We release our models through 3H_Merging for further investigations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0d49907-3345-43da-bd63-7641abe16653Cited by top-tier papers7
- OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective AlignmentLiang Lin, Zhihao Xu, Junhao Dong, Jian Zhao et al.ICLR 2026 · 7 citations
- Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem ProvingChuxue Cao, Mengze Li, Juntao Dai, Jinluan Yang et al.EMNLP 2025 · 1 citation
- Latent Score-Based Reweighting for Robust Classification on Imbalanced Tabular DataYunze Tong, Fengda Zhang, Zihao Tang, Kaifeng Gao et al.ICML 2025
- Scalable Model Merging with Progressive Layer-wise DistillationJing Xu, Jiazheng Li, Jingzhao ZhangICML 2025
- D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent SamplesZijing Hu, Fengda Zhang, Kun KuangICML 2025
Builds on43
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
Related papers
- HAF-RM: A Hybrid Alignment Framework for Reward Model TrainingShujun Liu, Xiaoyu Shen, Yuhang Lai, Siyuan Wang et al.ACL 2025 · 4 citations
- AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMsNicholas E. Corrado, Julian Katz-Samuels, Adithya M. Devraj, Hyokun Yun et al.ACL 2025
- Alignment for HonestyYuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig et al.NeurIPS 2024 · 82 citations
- Unlearners Can Lie: Evaluating and Improving Honesty in LLM UnlearningRenjie Gu, Jiazhen Du, Yihua Zhang, Sijia LiuACL 2026
- Too Helpful, Too Harmless, Too Honest or Just Right?Gautam Siddharth Kashyap, Mark Dras, Usman NaseemEMNLP 2025
