SQUAB: Evaluating LLM robustness to Ambiguous and Unanswerable Questions in Semantic Parsing
Simone Papicchio, Luca Cagliero, Paolo Papotti
摘要
Large Language Models (LLMs) have demonstrated robust performance in Semantic Parsing (SP) for well-defined queries with unambiguous intent and answerable responses. However, practical user questions frequently deviate from these ideal conditions, challenging the applicability of existing benchmarks. To address this issue, we introduce SQUAB, an automatic dataset generator of Ambiguous and Unanswerable questions. SQUAB generates complex, annotated SP tests using a blend of SQL and LLM capabilities. Results show that SQUAB reduces test generation costs by up to 99% compared to human-based solutions while aligning with real-world question patterns. Furthermore, these tests challenge LLM performance while revealing disparities between public and proprietary datasets. This highlights the need for a dynamic, automatic dataset generator as SQUAB. The code is designed for user extension to accommodate new ambiguous and unanswerable patterns and is available at https: //github.com/spapicchio/squab .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- ClarifyGPT: A Framework for Enhancing LLM-Based Code Generation via Requirements ClarificationFangwen Mu, Lin Shi, Song Wang, Zhuohao Yu 等FSE 2024 · 被引用 49 次
- Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear QueriesXinyi He, Mengyu Zhou, Xinrun Xu, Xiaojun Ma 等AAAI 2024 · 被引用 48 次
- Zero and Few-shot Semantic Parsing with Ambiguous InputsElias Stengel-Eskin, Kyle Rawlins, Benjamin Van DurmeICLR 2024 · 被引用 27 次
- Benchmarking and Improving Text-to-SQL Generation under AmbiguityAdithya Bhaskar, Tushar Tomar, Ashutosh Sathe, Sunita SarawagiEMNLP 2023 · 被引用 13 次
- Data Ambiguity Profiling for the Generation of Training ExamplesEnzo Veltri, Gilbert Badaro, Mohammed Saeed, Paolo PapottiICDE 2023 · 被引用 12 次
相关 Paper
- CLAMBER: A Benchmark of Identifying and Clarifying Ambiguous Information Needs in Large Language ModelsTong Zhang, Peixin Qin, Yang Deng, Chen Huang 等ACL 2024
- STARQA: A Question Answering Dataset for Complex Analytical Reasoning over Structured DatabasesMounica Maddela, Lingjue Xie, Daniel Preotiuc-Pietro, MausamEMNLP 2025
- CLEAR: A Parser-Independent Disambiguation Framework for NL2SQLMeng Zhang, Kexin Ma, Liyang Xu, Kedi Zhang 等ICDE 2025 · 被引用 4 次
- CompTab: A Comprehensive Benchmark for Real-World TableQA with Complex Reasoning and Irregular TablesZhen Yang, Wei Du, Jie Wang, Wenze Zhou 等ACL 2026
- Disambiguation in Conversational Question Answering in the Era of LLMs and Agents: A SurveyMd. Mehrab Tanjim, Yeonjun In, Xiang Chen, Victor S. Bursztyn 等EMNLP 2025 · 被引用 2 次
