Towards Better Value Principles for Large Language Model Alignment: A Systematic Evaluation and Enhancement
Bingbing Xu, Jing Yao, Xiaoyuan Yi, Aishan Maoliniyazi, Xing Xie, Xiaofeng Meng
摘要
As Large Language Models (LLMs) advance, aligning them with human values is critical for their responsible development. Value principles serve as the foundation for clarifying alignment goals. Multiple sets of value principles have been proposed, such as HHH (helpful, honest, harmless) and instructions for data synthesis in reinforcement learning from AI feedback (RLAIF). However, most of them are heuristically crafted, without consideration of three primary challenges in practical LLM alignment: 1) Comprehensiveness to deal with diverse and even unforeseen scenarios in which LLMs could be applied; 2) Precision to provide LLMs with clear and actionable guidance in specific scenarios; and 3) Compatability to avoid internal contracts between principles. In this paper, we formalize quantitative metrics to evaluate value principles along the three desirable properties. Building on these metrics, we propose the Hi erarchical Va lue P rinciple framework ( HiVaP ) 1 , which constructs a hierarchical principle set and retrieves principles tailored to each scenario in a cascading way, addressing above challenges. Experimental re-sults validate that the three metrics capture the effectiveness of value principles for LLM alignment, and our HiVaP framework that enhances these metrics leads to superior alignment. Warning: This paper contains several toxic and offensive statements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual ConversationsBhada Yun, Renn Su, April Yi WangCHI 2026 · 被引用 7 次
- Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model AlignmentYucong Huang, Xiucheng Li, Kaiqi Zhao, Jing LiICML 2026
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen 等ICLR 2024 · 被引用 699 次
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji 等ICLR 2024 · 被引用 656 次
相关 Paper
- SPRI: Aligning Large Language Models with Context-Situated PrinciplesHongli Zhan, Muneeza Azmat, Raya Horesh, Junyi Jessy Li 等ICML 2025
- Alignment for HonestyYuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig 等NeurIPS 2024 · 被引用 82 次
- Towards Tool Use Alignment of Large Language ModelsZhiyuan Chen, Shiqi Shen, Guangyao Shen, Gong Zhi 等EMNLP 2024 · 被引用 5 次
- Code Red! On the Harmfulness of Applying Off-the-Shelf Large Language Models to Programming TasksAli Al-Kaswan, Sebastian Deatc, Begüm Koç, Arie van Deursen 等FSE 2025 · 被引用 1 次
- Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-SortsJingting Zheng, Yuqi Ren, Linhao Yu, Yongqi Leng 等ACL 2026
