Gar: A Generate-and-Rank Approach for Natural Language to SQL Translation
Yuankai Fan, Zhenying He, Tonghui Ren, Dianjun Guo, Lin Chen, Ruisi Zhu, Guanduo Chen, Yinan Jing, Kai Zhang, X. Sean Wang
摘要
A Natural Language (NL) Interface to Databases (NLIDB) aims to help end-users access databases. State-of-the-art approaches primarily construct language translation models to convert NL queries to SQL queries. While these models exhibit good performance on NLIDB benchmarks, the translation accuracy seems to have stalled at between 70%-75%, and most erroneous translations happen with complex queries that require an understanding of the structure and semantics specific to a database. This paper proposes a Generate-And-Rank approach called Gar. Gar assumes that a set of sample SQL queries is given to represent the possible user-intended queries to the database. In order to provide a broad coverage, akin to avoiding over-fitting, Gar extracts the basic components from the sample set to form the basic building blocks to generate a set of generalized SQL queries. By leveraging a simple rule-based SQL to NL technique, a less natural NL expression called a dialect expression for each sample and generalized SQL query is obtained. Finally, a learning-to-rank method is used for a given NL query to retrieve the best dialect expression and hence the resulting SQL query. Extensive experiments are performed to study Gar in comparison with other approaches. The results show that Gar achieves better performance on the NLIDB benchmarks, including in particular a 78.5% translation accuracy on the popular Spider benchmark, outperforming the best reported accuracy in the literature. An extension to Gar, called Gar-j, is further introduced to aid the translation by annotating join semantics in the sample queries. The experimental results show that Gar-j can further improve translation accuracy on queries with joins. Code for Gar can be found at https://github.com/Kaimary/GAR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PURPLE: Making a Large Language Model a Better SQL WriterTonghui Ren, Yuankai Fan, Zhenying He, Ren Huang 等ICDE 2024 · 被引用 49 次
- Metasql: A Generate-Then-Rank Framework for Natural Language to SQL TranslationYuankai Fan, Zhenying He, Tonghui Ren, Can Huang 等ICDE 2024 · 被引用 23 次
- Grounding Natural Language to SQL Translation with Data-Based Self-ExplanationsYuankai Fan, Tonghui Ren, Can Huang, Zhenying He 等ICDE 2025 · 被引用 7 次
- The Power of Constraints in Natural Language to SQL TranslationTonghui Ren, Chen Ke, Yuankai Fan, Yinan Jing 等VLDB 2025 · 被引用 4 次
- HCT-QA: A Benchmark for Question Answering on Human-Centric TablesMohammad Shahmeer Ahmad, Zan Ahmad Naeem, Michaël Aupetit, Ahmed K. Elmagarmid 等ICDE 2026
它引用的顶会 Paper7
- Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-TrainingPeng Shi, Patrick Ng, Zhiguo Wang, Henghui Zhu 等AAAI 2021 · 被引用 124 次
- Exploring Unexplored Generalization Challenges for Cross-Database Semantic ParsingAlane Suhr, Ming-Wei Chang, Peter Shaw, Kenton LeeACL 2020 · 被引用 76 次
- GraPPa: Grammar-Augmented Pre-Training for Table Semantic ParsingTao Yu, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang 等ICLR 2021 · 被引用 59 次
- RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersBailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov 等ACL 2020 · 被引用 39 次
- MT-Teql: Evaluating and Augmenting Neural NLIDB on Real-world Linguistic and Schema VariationsPingchuan Ma, Shuai WangVLDB 2022 · 被引用 38 次
相关 Paper
- ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL SystemsYi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten 等VLDB 2024 · 被引用 65 次
- Dial: A Knowledge-Grounded Dialect-Specific NL2SQL SystemXiang Zhang, Le Zhou, Hongming Xu, Wei Zhou 等VLDB 2026 · 被引用 2 次
- A Natural Language Interface for Database: Achieving Transfer-learnability Using Adversarial Method for Question UnderstandingWenlu Wang, Yingtao Tian, Haixun Wang, Wei-Shinn KuICDE 2020 · 被引用 14 次
- DBPal: A Fully Pluggable NL2SQL Training PipelineNathaniel Weir, Prasetya Ajie Utama, Alex Galakatos, Andrew Crotty 等SIGMOD 2020 · 被引用 36 次
- ATHENA++: Natural Language Querying for Complex Nested SQL QueriesJaydeep Sen, Chuan Lei, Abdul Quamar, Fatma Özcan 等VLDB 2020
