Task Ambiguity in Humans and Language Models
Alex Tamkin, Kunal Handa, Avash Shrestha, Noah D. Goodman
Abstract
Language models have recently achieved strong performance across a wide range of NLP benchmarks. However, unlike benchmarks, real world tasks are often poorly specified, and agents must deduce the user's intended behavior from a combination of context, instructions, and examples. We investigate how both humans and models behave in the face of such task ambiguity by proposing AmbiBench, a new benchmark of six ambiguously-specified classification tasks. We evaluate humans and models on AmbiBench by seeing how well they identify the intended task using 1) instructions with varying degrees of ambiguity, and 2) different numbers of labeled examples. We find that the combination of model scaling (to 175B parameters) and training with human feedback data enables models to approach or exceed the accuracy of human participants across tasks, but that either one alone is not sufficient. In addition, we show how to dramatically improve the accuracy of language models trained without large-scale human feedback training by finetuning on a small number of ambiguous in-context examples, providing a promising direction for teaching models to generalize well in the face of ambiguity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3928544b-d8aa-41eb-bb0a-df4c866473c6Cited by top-tier papers19
- Understanding Catastrophic Forgetting in Language Models via Implicit InferenceSuhas Kotha, Jacob Mitchell Springer, Aditi RaghunathanICLR 2024 · 131 citations
- Decomposing Uncertainty for Large Language Models through Input Clarification EnsemblingBairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas et al.ICML 2024 · 113 citations
- Introspective Planning: Aligning Robots' Uncertainty with Inherent Task AmbiguityKaiqu Liang, Zixu Zhang, Jaime F. FisacNeurIPS 2024 · 43 citations
- We're Afraid Language Models Aren't Modeling AmbiguityAlisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr et al.EMNLP 2023 · 35 citations
- TheoremQA: A Theorem-driven Question Answering DatasetWenhu Chen, Ming Yin, Max Ku, Pan Lu et al.EMNLP 2023 · 30 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
Related papers
- AmbiK: Dataset of Ambiguous Tasks in Kitchen EnvironmentAnastasiia Ivanova, Eva Bakaeva, Zoya Volovikova, Alexey K. Kovalev et al.ACL 2025
- AmbigNLG: Addressing Task Ambiguity in Instruction for NLGAyana Niwa, Hayate IsoEMNLP 2024 · 2 citations
- Instruction Induction: From Few Examples to Natural Language Task DescriptionsOr Honovich, Uri Shaham, Samuel R. Bowman, Omer LevyACL 2023 · 48 citations
- MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMsYunqiu Xu, Linchao Zhu, Yi YangICCV 2025 · 7 citations
- CrossCodeBench: Benchmarking Cross-Task Generalization of Source Code ModelsChangan Niu, Chuanyi Li, Vincent Ng, Bin LuoICSE 2023 · 9 citations
