CLAMBER: A Benchmark of Identifying and Clarifying Ambiguous Information Needs in Large Language Models
Tong Zhang, Peixin Qin, Yang Deng, Chen Huang, Wenqiang Lei, Junhong Liu, Dingnan Jin, Hongru Liang, Tat-Seng Chua
摘要
Large language models (LLMs) are increasingly used to meet user information needs, but their effectiveness in dealing with user queries that contain various types of ambiguity remains unknown, ultimately risking user trust and satisfaction. To this end, we introduce CLAMBER, a benchmark for evaluating LLMs using a well-organized taxonomy. Building upon the taxonomy, we construct ∼ 12K high-quality data to assess the strengths, weaknesses, and potential risks of various off-theshelf LLMs. Our findings indicate the limited practical utility of current LLMs in identifying and clarifying ambiguous user queries, even enhanced by chain-of-thought (CoT) and few-shot prompting. These techniques may result in overconfidence in LLMs and yield only marginal enhancements in identifying ambiguity. Furthermore, current LLMs fall short in generating high-quality clarifying questions due to a lack of conflict resolution and inaccurate utilization of inherent knowledge. In this paper, CLAMBER presents a guidance and promotes further research on proactive and trustworthy LLMs. Our dataset is available at https://github.com/SCUNLP/CLAMBER.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software EngineeringSanidhya Vijayvargiya, Xuhui Zhou, Akhila Yerukola, Maarten Sap 等ICLR 2026 · 被引用 35 次
- Knowledge Boundary of Large Language Models: A SurveyMoxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li 等ACL 2025 · 被引用 33 次
- Semantic Volume: Quantifying and Detecting Both External and Internal Uncertainty in LLMsXiaomin Li, Zhou Yu, Ziji Zhang, Yingying Zhuang 等AAAI 2026 · 被引用 11 次
- Do not Abstain! Identify and Solve the UncertaintyJingyu Liu, Jingquan Peng, Xiaopeng Wu, Xubin Li 等ACL 2025 · 被引用 8 次
- RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented GenerationSizhe Cheng, Jiaping Li, Huanchen Wang, Yuxin MaUIST 2025 · 被引用 6 次
它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- Generating Clarifying Questions for Information RetrievalHamed Zamani, Susan T. Dumais, Nick Craswell, Paul N. Bennett 等WWW 2020 · 被引用 238 次
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 被引用 162 次
相关 Paper
- Clarifying Ambiguities: on the Role of Ambiguity Types in Prompting Methods for Clarification GenerationAnfu Tang, Laure Soulier, Vincent GuigueSIGIR 2025 · 被引用 6 次
- CLEAR: A Parser-Independent Disambiguation Framework for NL2SQLMeng Zhang, Kexin Ma, Liyang Xu, Kedi Zhang 等ICDE 2025 · 被引用 4 次
- AmbigNLG: Addressing Task Ambiguity in Instruction for NLGAyana Niwa, Hayate IsoEMNLP 2024 · 被引用 2 次
- SQUAB: Evaluating LLM robustness to Ambiguous and Unanswerable Questions in Semantic ParsingSimone Papicchio, Luca Cagliero, Paolo PapottiEMNLP 2025
- Disambiguation in Conversational Question Answering in the Era of LLMs and Agents: A SurveyMd. Mehrab Tanjim, Yeonjun In, Xiang Chen, Victor S. Bursztyn 等EMNLP 2025 · 被引用 2 次
