Eliciting Human Preferences with Language Models
Belinda Z. Li, Alex Tamkin, Noah D. Goodman, Jacob Andreas
Abstract
Language models (LMs) can be directed to perform target tasks by using labeled examples or natural language prompts. But selecting examples or writing prompts for can be challenging-especially in tasks that involve unusual edge cases, demand precise articulation of nebulous preferences, or require an accurate mental model of LM behavior. We propose to use LMs themselves to guide the task specification process. In this paper, we introduce generative active task elicitation (GATE): a learning framework in which models elicit and infer intended behavior through free-form, language-based interaction with users. We study GATE in three domains: email validation, content recommendation, and moral reasoning. In preregistered experiments, we show that LMs prompted to perform GATE (e.g., by generating open-ended questions or synthesizing informative edge cases) elicit responses that are often more informative than user-written prompts or labels. Users report that interactive task elicitation requires less effort than prompting or example labeling and surfaces novel considerations not initially anticipated by users. Our findings suggest that LM-driven elicitation can be a powerful tool for aligning models to complex human preferences and values. 1 * Equal contribution. Author order decided via coin flip.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers29
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen et al.NeurIPS 2024 · 215 citations
- Doing Experiments and Revising Rules with Natural Language and Probabilistic ReasoningTop Piriyakulkij, Cassidy Langenfeld, Tuan Anh Le, Kevin EllisNeurIPS 2024 · 30 citations
- BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental DesignDeepro Choudhury, Sinead Williamson, Adam Golinski, Ning Miao et al.ICLR 2026 · 24 citations
- Doing Personal LAPS: LLM-Augmented Dialogue Construction for Personalized Multi-Session Conversational SearchHideaki Joko, Shubham Chatterjee, Andrew Ramsay, Arjen P. de Vries et al.SIGIR 2024 · 24 citations
- Navigating Rifts in Human-LLM Grounding: Study and BenchmarkOmar Shaikh, Hussein Mozannar, Gagan Bansal, Adam Fourney et al.ACL 2025 · 21 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- MIND: A Large-scale Dataset for News RecommendationFangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu et al.ACL 2020 · 454 citations
- An Investigation of Why Overparameterization Exacerbates Spurious CorrelationsShiori Sagawa, Aditi Raghunathan, Pang Wei Koh, Percy LiangICML 2020 · 436 citations
- Understanding the failure modes of out-of-distribution generalizationVaishnavh Nagarajan, Anders Andreassen, Behnam NeyshaburICLR 2021 · 205 citations
Related papers
- TO-GATE: Clarifying Questions and Summarizing Responses with Trajectory Optimization for Eliciting Human PreferenceYulin Dou, Jiangming LiuAAAI 2026 · 1 citation
- Mind the Gap: The Divergence Between Human and LLM-Generated TasksYi-Long Lu, Jiajun Song, Chunhui Zhang, Wei WangAAAI 2026
- Customizing Language Model Responses with Contrastive In-Context LearningXiang Gao, Kamalika DasAAAI 2024 · 23 citations
- Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation FrameworkMohna Chakraborty, Lu Wang, David JurgensEMNLP 2025
- Active Task Disambiguation with LLMsKasia Kobalczyk, Nicolás Astorga, Tennison Liu, Mihaela van der SchaarICLR 2025
