ClarifyDelphi: Reinforced Clarification Questions with Defeasibility Rewards for Social and Moral Situations
Valentina Pyatkin, Jena D. Hwang, Vivek Srikumar, Ximing Lu, Liwei Jiang, Yejin Choi, Chandra Bhagavatula
摘要
Context is everything, even in commonsense moral reasoning. Changing contexts can flip the moral judgment of an action; Lying to a friend is wrong in general, but may be morally acceptable if it is intended to protect their life. We present CLARIFYDELPHI, an interactive system that learns to ask clarification questions (e.g., "why did you lie to your friend?") in order to elicit additional salient contexts of a social or moral situation. We posit that questions whose potential answers lead to diverging moral judgments are the most informative. Thus, we propose a reinforcement learning framework with a defeasibility reward that aims to maximize the divergence between moral judgments of hypothetical answers to a question. Human evaluation demonstrates that our system generates more relevant, informative and defeasible questions compared to competitive baselines. Our work is ultimately inspired by studies in cognitive science that have investigated the flexibility in moral cognition (i.e., the diverse contexts in which moral rules can be bent), and we hope that research in this direction can assist both cognitive and computational investigations of moral judgments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language ModelsArchiki Prasad, Elias Stengel-Eskin, Mohit BansalICLR 2024 · 被引用 13 次
- What Makes a Good Natural Language Prompt?Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi 等ACL 2025 · 被引用 13 次
- I Could've Asked That: Reformulating Unanswerable QuestionsWenting Zhao, Ge Gao, Claire Cardie, Alexander M. RushEMNLP 2024 · 被引用 2 次
- Social Story Frames: Contextual Reasoning about Narrative Intent and ReceptionJoel Mire, Maria Antoniak, Steven R. Wilson, Zexin Ma 等ACL 2026 · 被引用 2 次
- Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and DutiesTaylor Sorensen, Liwei Jiang, Jena D. Hwang, Sydney Levine 等AAAI 2024
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
相关 Paper
- Think about it! Improving defeasible reasoning by first modeling the question scenarioAman Madaan, Niket Tandon, Dheeraj Rajagopal, Peter Clark 等EMNLP 2021
- Delphi: A Neuro-Symbolic Framework for Individualized, Safe and Interpretable Treatment RecommendationMuchan Tao, Haonan Qin, Yuqi Fang, Caifeng Shan 等AAAI 2026
- When to Make Exceptions: Exploring Language Models as Accounts of Human Moral JudgmentZhijing Jin, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal 等NeurIPS 2022 · 被引用 146 次
- Asking What Matters: Reward-Driven Clarification for Software Engineering TasksSanidhya Vijayvargiya, Vijay Viswanathan, Graham NeubigICML 2026 · 被引用 3 次
- Epistemic Gain, Aleatoric Cost: Uncertainty Decomposition in Multi-Agent Debate for Math ReasoningDan Qiao, Binbin Chen, Fengyu Cai, Jianlong Chen 等ICML 2026 · 被引用 3 次
