Towards Better Value Principles for Large Language Model Alignment: A Systematic Evaluation and Enhancement
Bingbing Xu, Jing Yao, Xiaoyuan Yi, Aishan Maoliniyazi, Xing Xie, Xiaofeng Meng
Abstract
As Large Language Models (LLMs) advance, aligning them with human values is critical for their responsible development. Value principles serve as the foundation for clarifying alignment goals. Multiple sets of value principles have been proposed, such as HHH (helpful, honest, harmless) and instructions for data synthesis in reinforcement learning from AI feedback (RLAIF). However, most of them are heuristically crafted, without consideration of three primary challenges in practical LLM alignment: 1) Comprehensiveness to deal with diverse and even unforeseen scenarios in which LLMs could be applied; 2) Precision to provide LLMs with clear and actionable guidance in specific scenarios; and 3) Compatability to avoid internal contracts between principles. In this paper, we formalize quantitative metrics to evaluate value principles along the three desirable properties. Building on these metrics, we propose the Hi erarchical Va lue P rinciple framework ( HiVaP ) 1 , which constructs a hierarchical principle set and retrieves principles tailored to each scenario in a cascading way, addressing above challenges. Experimental re-sults validate that the three metrics capture the effectiveness of value principles for LLM alignment, and our HiVaP framework that enhances these metrics leads to superior alignment. Warning: This paper contains several toxic and offensive statements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8b01f61-95bf-4ef9-a86c-2f1e5a776b58Cited by top-tier papers2
- AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual ConversationsBhada Yun, Renn Su, April Yi WangCHI 2026 · 7 citations
- Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model AlignmentYucong Huang, Xiucheng Li, Kaiqi Zhao, Jing LiICML 2026
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 699 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
Related papers
- SPRI: Aligning Large Language Models with Context-Situated PrinciplesHongli Zhan, Muneeza Azmat, Raya Horesh, Junyi Jessy Li et al.ICML 2025
- Alignment for HonestyYuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig et al.NeurIPS 2024 · 82 citations
- Towards Tool Use Alignment of Large Language ModelsZhiyuan Chen, Shiqi Shen, Guangyao Shen, Gong Zhi et al.EMNLP 2024 · 5 citations
- Code Red! On the Harmfulness of Applying Off-the-Shelf Large Language Models to Programming TasksAli Al-Kaswan, Sebastian Deatc, Begüm Koç, Arie van Deursen et al.FSE 2025 · 1 citation
- Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-SortsJingting Zheng, Yuqi Ren, Linhao Yu, Yongqi Leng et al.ACL 2026
