Eliciting Human Preferences with Language Models
Belinda Z. Li, Alex Tamkin, Noah D. Goodman, Jacob Andreas
摘要
Language models (LMs) can be directed to perform target tasks by using labeled examples or natural language prompts. But selecting examples or writing prompts for can be challenging-especially in tasks that involve unusual edge cases, demand precise articulation of nebulous preferences, or require an accurate mental model of LM behavior. We propose to use LMs themselves to guide the task specification process. In this paper, we introduce generative active task elicitation (GATE): a learning framework in which models elicit and infer intended behavior through free-form, language-based interaction with users. We study GATE in three domains: email validation, content recommendation, and moral reasoning. In preregistered experiments, we show that LMs prompted to perform GATE (e.g., by generating open-ended questions or synthesizing informative edge cases) elicit responses that are often more informative than user-written prompts or labels. Users report that interactive task elicitation requires less effort than prompting or example labeling and surfaces novel considerations not initially anticipated by users. Our findings suggest that LM-driven elicitation can be a powerful tool for aligning models to complex human preferences and values. 1 * Equal contribution. Author order decided via coin flip.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen 等NeurIPS 2024 · 被引用 215 次
- Doing Experiments and Revising Rules with Natural Language and Probabilistic ReasoningTop Piriyakulkij, Cassidy Langenfeld, Tuan Anh Le, Kevin EllisNeurIPS 2024 · 被引用 30 次
- BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental DesignDeepro Choudhury, Sinead Williamson, Adam Golinski, Ning Miao 等ICLR 2026 · 被引用 24 次
- Doing Personal LAPS: LLM-Augmented Dialogue Construction for Personalized Multi-Session Conversational SearchHideaki Joko, Shubham Chatterjee, Andrew Ramsay, Arjen P. de Vries 等SIGIR 2024 · 被引用 24 次
- Navigating Rifts in Human-LLM Grounding: Study and BenchmarkOmar Shaikh, Hussein Mozannar, Gagan Bansal, Adam Fourney 等ACL 2025 · 被引用 21 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- MIND: A Large-scale Dataset for News RecommendationFangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu 等ACL 2020 · 被引用 454 次
- An Investigation of Why Overparameterization Exacerbates Spurious CorrelationsShiori Sagawa, Aditi Raghunathan, Pang Wei Koh, Percy LiangICML 2020 · 被引用 436 次
- Understanding the failure modes of out-of-distribution generalizationVaishnavh Nagarajan, Anders Andreassen, Behnam NeyshaburICLR 2021 · 被引用 205 次
相关 Paper
- TO-GATE: Clarifying Questions and Summarizing Responses with Trajectory Optimization for Eliciting Human PreferenceYulin Dou, Jiangming LiuAAAI 2026 · 被引用 1 次
- Mind the Gap: The Divergence Between Human and LLM-Generated TasksYi-Long Lu, Jiajun Song, Chunhui Zhang, Wei WangAAAI 2026
- Customizing Language Model Responses with Contrastive In-Context LearningXiang Gao, Kamalika DasAAAI 2024 · 被引用 23 次
- Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation FrameworkMohna Chakraborty, Lu Wang, David JurgensEMNLP 2025
- Active Task Disambiguation with LLMsKasia Kobalczyk, Nicolás Astorga, Tennison Liu, Mihaela van der SchaarICLR 2025
