The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems
Caleb Ziems, Jane A. Yu, Yi-Chia Wang, Alon Y. Halevy, Diyi Yang
Abstract
Content Warning: some examples in this paper may be offensive or upsetting. Conversational agents have come increasingly closer to human competence in open-domain dialogue settings; however, such models can reflect insensitive, hurtful, or entirely incoherent viewpoints that erode a user's trust in the moral integrity of the system. Moral deviations are difficult to mitigate because moral judgments are not universal, and there may be multiple competing judgments that apply to a situation simultaneously. In this work, we introduce a new resource, not to authoritatively resolve moral ambiguities, but instead to facilitate systematic understanding of the intuitions, values and moral judgments reflected in the utterances of dialogue systems. The MORAL INTEGRITY CORPUS, MIC , is such a resource, which captures the moral assumptions of 38k prompt-reply pairs, using 99k distinct Rules of Thumb (RoTs). Each RoT reflects a particular moral conviction that can explain why a chatbot's reply may appear acceptable or problematic. We further organize RoTs with a set of 9 moral and social attributes and benchmark performance for attribute classification. Most importantly, we show that current neural language models can automatically generate new RoTs that reasonably describe previously unseen interactions, but they still struggle with certain scenarios. Our findings suggest that MIC will be a useful resource for understanding and language models' implicit moral assumptions and flexibly benchmarking the integrity of conversational agents. To download the data, see https://github.com/GT-SALT/mic
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d48a832-c809-4cf2-9c94-f495fee9eb5eCited by top-tier papers21
- Training Socially Aligned Language Models on Simulated Social InteractionsRuibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang et al.ICLR 2024 · 97 citations
- Second Thoughts are Best: Learning to Re-Align With Human Values from Text EditsRuibo Liu, Chenyan Jia, Ge Zhang, Ziyu Zhuang et al.NeurIPS 2022 · 46 citations
- ProsocialDialog: A Prosocial Backbone for Conversational AgentsHyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu et al.EMNLP 2022 · 46 citations
- Mirages. On Anthropomorphism in Dialogue SystemsGavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser et al.EMNLP 2023 · 44 citations
- Aligning Language Models with Human Preferences via a Bayesian ApproachJiashuo Wang, Haozhao Wang, Shichao Sun, Wenjie LiNeurIPS 2023 · 42 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- Do Morals Guide How LLMs Think? The Role of Ethical Perspectives in General Problem SolvingIseo Kim, Eunjin Hong, Juae KimACL 2026
- Social Chemistry 101: Learning to Reason about Social and Moral NormsMaxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap et al.EMNLP 2020 · 11 citations
- MoVa: Towards Generalizable Classification of Human Morals and ValuesZiyu Chen, Junfei Sun, Chenxi Li, Tuan Dung Nguyen et al.EMNLP 2025
- Feeling Rules in Language Models: Mapping Norms of Emotional Appropriateness Across Roles, Institutions, and IntensityGuangrui Fan, Dandan Liu, Aznul Qalid Md Sabri, Rui Zhang et al.ACL 2026
