ATHENA++: Natural Language Querying for Complex Nested SQL Queries
Jaydeep Sen, Chuan Lei, Abdul Quamar, Fatma Özcan, Vasilis Efthymiou, Ayushi Dalmia, Greg Stager, Ashish R. Mittal, Diptikalyan Saha, Karthik Sankaranarayanan
Abstract
Natural Language Interfaces to Databases (NLIDB) systems eliminate the requirement for an end user to use complex query languages like SQL, by translating the input natural language (NL) queries to SQL automatically. Although a significant volume of research has focused on this space, most state-of-the-art systems can at best handle simple select-project-join queries. There has been little to no research on extending the capabilities of NLIDB systems to handle complex business intelligence (BI) queries that often involve nesting as well as aggregation. In this paper, we present ATHENA++, an end-to-end system that can answer such complex queries in natural language by translating them into nested SQL queries. In particular, ATHENA++ combines linguistic patterns from NL queries with deep domain reasoning using ontologies to enable nested query detection and generation. We also introduce a new benchmark data set (FIBEN), which consists of 300 NL queries, corresponding to 237 distinct complex SQL queries on a database with 152 tables, conforming to an ontology derived from standard financial ontologies (FIBO and FRO). We conducted extensive experiments comparing ATHENA++ with two state-ofthe-art NLIDB systems, using both FIBEN and the prominent Spider benchmark. ATHENA++ consistently outperforms both systems across all benchmark data sets with a wide variety of complex queries, achieving 88.33% accuracy on FIBEN benchmark, and 78.89% accuracy on Spider benchmark, beating the best reported accuracy results on the dev set by 8%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6855f8f2-3964-41cc-8e23-00a22ead5ce4Cited by top-tier papers19
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun et al.VLDB 2024 · 609 citations
- CatSQL: Towards Real World Natural Language to SQL ApplicationsHan Fu, Chang Liu, Bin Wu, Feifei Li et al.VLDB 2023 · 79 citations
- CodexDB: Synthesizing Code for Query Processing from Natural Language Instructions using GPT-3 CodexImmanuel TrummerVLDB 2022 · 77 citations
- ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL SystemsYi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten et al.VLDB 2024 · 65 citations
- PURPLE: Making a Large Language Model a Better SQL WriterTonghui Ren, Yuankai Fan, Zhenying He, Ren Huang et al.ICDE 2024 · 49 citations
Builds on1
Related papers
- Gar: A Generate-and-Rank Approach for Natural Language to SQL TranslationYuankai Fan, Zhenying He, Tonghui Ren, Dianjun Guo et al.ICDE 2023 · 12 citations
- MT-Teql: Evaluating and Augmenting Neural NLIDB on Real-world Linguistic and Schema VariationsPingchuan Ma, Shuai WangVLDB 2022 · 38 citations
- Metasql: A Generate-Then-Rank Framework for Natural Language to SQL TranslationYuankai Fan, Zhenying He, Tonghui Ren, Can Huang et al.ICDE 2024 · 23 citations
- A Natural Language Interface for Database: Achieving Transfer-learnability Using Adversarial Method for Question UnderstandingWenlu Wang, Yingtao Tian, Haixun Wang, Wei-Shinn KuICDE 2020 · 14 citations
- From Natural Language Processing to Neural DatabasesJames Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri et al.VLDB 2021 · 62 citations
