Denevil: towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
Shitong Duan, Xiaoyuan Yi, Peng Zhang, Tun Lu, Xing Xie, Ning Gu
Abstract
Warning: this paper contains model outputs exhibiting unethical information. Large Language Models (LLMs) have made unprecedented breakthroughs, yet their increasing integration into everyday life might raise societal risks due to generated unethical content. Despite extensive study on specific issues like bias, the intrinsic values of LLMs remain largely unexplored from a moral philosophy perspective. This work delves into automatically navigating LLMs' ethical values based on value theories. Moving beyond static discriminative evaluations with poor reliability, we propose DeNEVIL, a novel prompt generation algorithm tailored to dynamically exploit LLMs' value vulnerabilities and elicit the violation of ethics in a generative manner, revealing their underlying value inclinations. On such a basis, we construct MoralPrompt, a high-quality dataset comprising 2,397 prompts covering 500+ value principles, and then benchmark the intrinsic values across a spectrum of LLMs. We discovered that most models are essentially misaligned, necessitating further ethical value alignment. In response, we develop VILMO, an in-context alignment method that enhances the value compliance of LLM outputs by learning to generate appropriate value instructions, outperforming existing competitors. Our methods are suitable for black-box and open-source models, serving as an initial step in studying LLMs' ethical values.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ec80969-392e-4e04-adc8-d48443780449Cited by top-tier papers13
- The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral DilemmasYa Wu, Qiang Sheng, Danding Wang, Guang Yang et al.EMNLP 2025 · 8 citations
- AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value DifferenceJing Yao, Shitong Duan, Xiaoyuan Yi, Dongkuan Xu et al.ICLR 2026 · 4 citations
- Mitigating Biases in Language Models via Bias UnlearningDianqing Liu, Yi Liu, Guoqing Jin, Zhendong MaoEMNLP 2025 · 4 citations
- Towards Better Value Principles for Large Language Model Alignment: A Systematic Evaluation and EnhancementBingbing Xu, Jing Yao, Xiaoyuan Yi, Aishan Maoliniyazi et al.ACL 2025 · 3 citations
- Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value CodebookJaehyeok Lee, Xiaoyuan Yi, Jing Yao, Hyunjin Hwang et al.ICML 2026 · 1 citation
Builds on22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
Related papers
- Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-SortsJingting Zheng, Yuqi Ren, Linhao Yu, Yongqi Leng et al.ACL 2026
- Unintended Harms of Value-Aligned LLMs: Psychological and Empirical InsightsSooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao et al.ACL 2025
- Code Red! On the Harmfulness of Applying Off-the-Shelf Large Language Models to Programming TasksAli Al-Kaswan, Sebastian Deatc, Begüm Koç, Arie van Deursen et al.FSE 2025 · 1 citation
- Defending Against Alignment-Breaking Attacks via Robustly Aligned LLMBochuan Cao, Yuanpu Cao, Lu Lin, Jinghui ChenACL 2024 · 34 citations
- Do LLMs have Consistent Values?Naama Rozen, Liat Bezalel, Gal Elidan, Amir Globerson et al.ICLR 2025
