STARQA: A Question Answering Dataset for Complex Analytical Reasoning over Structured Databases
Mounica Maddela, Lingjue Xie, Daniel Preotiuc-Pietro, Mausam
摘要
Semantic parsing methods for converting text to SQL queries enable question answering over structured data and can greatly benefit analysts who routinely perform complex analytics on vast data stored in specialized relational databases. Although several benchmarks measure the abilities of text to SQL, the complexity of their questions is inherently limited by the level of expressiveness in query languages and none focus explicitly on questions involving complex analytical reasoning which require operations such as calculations over aggregate analytics, time series analysis or scenario understanding. In this paper, we introduce STARQA, the first public human-created dataset of complex analytical reasoning questions and answers on three specialized-domain databases. In addition to generating SQL directly using LLMs, we evaluate a novel approach (Text2SQLCode) that decomposes the task into a combination of SQL and Python: SQL is responsible for data fetching, and Python more naturally performs reasoning. Our results demonstrate that identifying and combining the abilities of SQL and Python is beneficial compared to using SQL alone, yet the dataset still remains quite challenging for the existing state-of-the-art LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language ModelsDaman Arora, Himanshu Gaurav Singh, MausamEMNLP 2023 · 被引用 36 次
- RAR: Retrieval-augmented retrieval for code generation in low resource languagesAvik Dutta, Mukul Singh, Gust Verbruggen, Sumit Gulwani 等EMNLP 2024 · 被引用 1 次
- Dual Reader-Parser on Hybrid Textual and Tabular Evidence for Open Domain Question AnsweringAlexander Hanbo Li, Patrick Ng, Peng Xu, Henghui Zhu 等ACL 2021
- Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause FrequenciesTerrance Liu, Shuyi Wang, Daniel Preotiuc-Pietro, Yash Chandarana 等EMNLP 2025
相关 Paper
- LogicCat: A Chain-of-Thought Text-to-SQL Benchmark for Complex ReasoningLiutao, Xutao Mao, Dixuan Zhang, Yifan Li 等AAAI 2026 · 被引用 3 次
- CRT-QA: A Dataset of Complex Reasoning Question Answering over Tabular DataZhehao Zhang, Xitao Li, Yan Gao, Jian-Guang LouEMNLP 2023 · 被引用 3 次
- STaR-SQL: Self-Taught Reasoner for Text-to-SQLMingqian He, Yongliang Shen, Wenqi Zhang, Qiuying Peng 等ACL 2025
- SQUAB: Evaluating LLM robustness to Ambiguous and Unanswerable Questions in Semantic ParsingSimone Papicchio, Luca Cagliero, Paolo PapottiEMNLP 2025
- KaggleDBQA: Realistic Evaluation of Text-to-SQL ParsersChia-Hsuan Lee, Oleksandr Polozov, Matthew RichardsonACL 2021
