SQUAB: Evaluating LLM robustness to Ambiguous and Unanswerable Questions in Semantic Parsing
Simone Papicchio, Luca Cagliero, Paolo Papotti
Abstract
Large Language Models (LLMs) have demonstrated robust performance in Semantic Parsing (SP) for well-defined queries with unambiguous intent and answerable responses. However, practical user questions frequently deviate from these ideal conditions, challenging the applicability of existing benchmarks. To address this issue, we introduce SQUAB, an automatic dataset generator of Ambiguous and Unanswerable questions. SQUAB generates complex, annotated SP tests using a blend of SQL and LLM capabilities. Results show that SQUAB reduces test generation costs by up to 99% compared to human-based solutions while aligning with real-world question patterns. Furthermore, these tests challenge LLM performance while revealing disparities between public and proprietary datasets. This highlights the need for a dynamic, automatic dataset generator as SQUAB. The code is designed for user extension to accommodate new ambiguous and unanswerable patterns and is available at https: //github.com/spapicchio/squab .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eba35ed8-c340-40d4-bc3b-b4c849041c9aBuilds on7
- ClarifyGPT: A Framework for Enhancing LLM-Based Code Generation via Requirements ClarificationFangwen Mu, Lin Shi, Song Wang, Zhuohao Yu et al.FSE 2024 · 49 citations
- Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear QueriesXinyi He, Mengyu Zhou, Xinrun Xu, Xiaojun Ma et al.AAAI 2024 · 48 citations
- Zero and Few-shot Semantic Parsing with Ambiguous InputsElias Stengel-Eskin, Kyle Rawlins, Benjamin Van DurmeICLR 2024 · 27 citations
- Benchmarking and Improving Text-to-SQL Generation under AmbiguityAdithya Bhaskar, Tushar Tomar, Ashutosh Sathe, Sunita SarawagiEMNLP 2023 · 13 citations
- Data Ambiguity Profiling for the Generation of Training ExamplesEnzo Veltri, Gilbert Badaro, Mohammed Saeed, Paolo PapottiICDE 2023 · 12 citations
Related papers
- CLAMBER: A Benchmark of Identifying and Clarifying Ambiguous Information Needs in Large Language ModelsTong Zhang, Peixin Qin, Yang Deng, Chen Huang et al.ACL 2024
- STARQA: A Question Answering Dataset for Complex Analytical Reasoning over Structured DatabasesMounica Maddela, Lingjue Xie, Daniel Preotiuc-Pietro, MausamEMNLP 2025
- CLEAR: A Parser-Independent Disambiguation Framework for NL2SQLMeng Zhang, Kexin Ma, Liyang Xu, Kedi Zhang et al.ICDE 2025 · 4 citations
- CompTab: A Comprehensive Benchmark for Real-World TableQA with Complex Reasoning and Irregular TablesZhen Yang, Wei Du, Jie Wang, Wenze Zhou et al.ACL 2026
- Disambiguation in Conversational Question Answering in the Era of LLMs and Agents: A SurveyMd. Mehrab Tanjim, Yeonjun In, Xiang Chen, Victor S. Bursztyn et al.EMNLP 2025 · 2 citations
