The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems
Caleb Ziems, Jane A. Yu, Yi-Chia Wang, Alon Y. Halevy, Diyi Yang
摘要
Content Warning: some examples in this paper may be offensive or upsetting. Conversational agents have come increasingly closer to human competence in open-domain dialogue settings; however, such models can reflect insensitive, hurtful, or entirely incoherent viewpoints that erode a user's trust in the moral integrity of the system. Moral deviations are difficult to mitigate because moral judgments are not universal, and there may be multiple competing judgments that apply to a situation simultaneously. In this work, we introduce a new resource, not to authoritatively resolve moral ambiguities, but instead to facilitate systematic understanding of the intuitions, values and moral judgments reflected in the utterances of dialogue systems. The MORAL INTEGRITY CORPUS, MIC , is such a resource, which captures the moral assumptions of 38k prompt-reply pairs, using 99k distinct Rules of Thumb (RoTs). Each RoT reflects a particular moral conviction that can explain why a chatbot's reply may appear acceptable or problematic. We further organize RoTs with a set of 9 moral and social attributes and benchmark performance for attribute classification. Most importantly, we show that current neural language models can automatically generate new RoTs that reasonably describe previously unseen interactions, but they still struggle with certain scenarios. Our findings suggest that MIC will be a useful resource for understanding and language models' implicit moral assumptions and flexibly benchmarking the integrity of conversational agents. To download the data, see https://github.com/GT-SALT/mic
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Training Socially Aligned Language Models on Simulated Social InteractionsRuibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang 等ICLR 2024 · 被引用 97 次
- Second Thoughts are Best: Learning to Re-Align With Human Values from Text EditsRuibo Liu, Chenyan Jia, Ge Zhang, Ziyu Zhuang 等NeurIPS 2022 · 被引用 46 次
- ProsocialDialog: A Prosocial Backbone for Conversational AgentsHyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu 等EMNLP 2022 · 被引用 46 次
- Mirages. On Anthropomorphism in Dialogue SystemsGavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser 等EMNLP 2023 · 被引用 44 次
- Aligning Language Models with Human Preferences via a Bayesian ApproachJiashuo Wang, Haozhao Wang, Shichao Sun, Wenjie LiNeurIPS 2023 · 被引用 42 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
- Do Morals Guide How LLMs Think? The Role of Ethical Perspectives in General Problem SolvingIseo Kim, Eunjin Hong, Juae KimACL 2026
- Social Chemistry 101: Learning to Reason about Social and Moral NormsMaxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap 等EMNLP 2020 · 被引用 11 次
- MoVa: Towards Generalizable Classification of Human Morals and ValuesZiyu Chen, Junfei Sun, Chenxi Li, Tuan Dung Nguyen 等EMNLP 2025
- Feeling Rules in Language Models: Mapping Norms of Emotional Appropriateness Across Roles, Institutions, and IntensityGuangrui Fan, Dandan Liu, Aznul Qalid Md Sabri, Rui Zhang 等ACL 2026
