Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
Jinluan Yang, Dingnan Jin, Anke Tang, Li Shen, Didi Zhu, Zhengyu Chen, Ziyu Zhao, Daixin Wang, Qing Cui, Zhiqiang Zhang, Jun Zhou, Fei Wu, Kun Kuang
摘要
Achieving balanced alignment of large language models (LLMs) in terms of Helpfulness, Honesty, and Harmlessness (3H optimization) constitutes a cornerstone of responsible AI. Existing methods like data mixture strategies face limitations, including heavy reliance on expert knowledge and conflicting optimization signals. While model merging offers parameter-level conflict-resolution strategies through integrating specialized models'parameters, its potential for 3H optimization remains underexplored. This paper systematically compares the effectiveness of model merging and data mixture methods in constructing 3H-aligned LLMs for the first time, revealing previously overlooked collaborative and conflict relationships among the 3H dimensions and discussing the advantages and drawbacks of data mixture (data-level) and model merging (parameter-level) methods in mitigating the conflict for balanced 3H optimization. Specially, we propose a novel Reweighting Enhanced task Singular Merging method, RESM, through outlier weighting and sparsity-aware rank selection strategies to address the challenges of preference noise accumulation and layer sparsity adaptation inherent in 3H-aligned LLM merging. Extensive evaluations can verify the effectiveness and robustness of RESM compared to previous data mixture (2%-5% gain) and model merging (1%-3% gain) methods in achieving balanced LLM alignment. We release our models through 3H_Merging for further investigations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective AlignmentLiang Lin, Zhihao Xu, Junhao Dong, Jian Zhao 等ICLR 2026 · 被引用 7 次
- Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem ProvingChuxue Cao, Mengze Li, Juntao Dai, Jinluan Yang 等EMNLP 2025 · 被引用 1 次
- Latent Score-Based Reweighting for Robust Classification on Imbalanced Tabular DataYunze Tong, Fengda Zhang, Zihao Tang, Kaifeng Gao 等ICML 2025
- Scalable Model Merging with Progressive Layer-wise DistillationJing Xu, Jiazheng Li, Jingzhao ZhangICML 2025
- D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent SamplesZijing Hu, Fengda Zhang, Kun KuangICML 2025
它引用的顶会 Paper43
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
相关 Paper
- HAF-RM: A Hybrid Alignment Framework for Reward Model TrainingShujun Liu, Xiaoyu Shen, Yuhang Lai, Siyuan Wang 等ACL 2025 · 被引用 4 次
- AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMsNicholas E. Corrado, Julian Katz-Samuels, Adithya M. Devraj, Hyokun Yun 等ACL 2025
- Alignment for HonestyYuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig 等NeurIPS 2024 · 被引用 82 次
- Unlearners Can Lie: Evaluating and Improving Honesty in LLM UnlearningRenjie Gu, Jiazhen Du, Yihua Zhang, Sijia LiuACL 2026
- Too Helpful, Too Harmless, Too Honest or Just Right?Gautam Siddharth Kashyap, Mark Dras, Usman NaseemEMNLP 2025
