Towards Knowledge-Intensive Text-to-SQL Semantic Parsing with Formulaic Knowledge
Longxu Dou, Yan Gao, Xuqi Liu, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Min-Yen Kan, Jian-Guang Lou
Abstract
In this paper, we study the problem of knowledge-intensive text-to-SQL, in which domain knowledge is necessary to parse expert questions into SQL queries over domainspecific tables. We formalize this scenario by building a new Chinese benchmark KNOWSQL consisting of domain-specific questions covering various domains. We then address this problem by presenting formulaic knowledge, rather than by annotating additional data examples. More concretely, we construct a formulaic knowledge bank as a domain knowledge base and propose a framework (REGROUP) to leverage this formulaic knowledge during parsing. Experiments using REGROUP demonstrate a significant 28.2% improvement overall on KNOWSQL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0765ae25-0038-41c8-8fe2-d20ee0c65794Cited by top-tier papers5
- BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation via Lens of Dynamic InteractionsNan Huo, Xiaohan Xu, Jinyang Li, Per Jacobsson et al.ICLR 2026 · 10 citations
- LogicCat: A Chain-of-Thought Text-to-SQL Benchmark for Complex ReasoningLiutao, Xutao Mao, Dixuan Zhang, Yifan Li et al.AAAI 2026 · 3 citations
- MAGNIFICo: Evaluating the In-Context Learning Ability of Large Language Models to Generalize to Novel InterpretationsArkil Patel, Satwik Bhattamishra, Siva Reddy, Dzmitry BahdanauEMNLP 2023 · 2 citations
- Structure-Guided Large Language Models for Text-to-SQL GenerationQinggang Zhang, Hao Chen, Junnan Dong, Shengyuan Chen et al.ICML 2025
- HEXGEN-FLOW: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQLYou Peng, Youhe Jiang, Wenqi Jiang, Chen Wang et al.ICDE 2026
Builds on14
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong et al.EMNLP 2022 · 222 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Constrained Language Models Yield Few-Shot Semantic ParsersRichard Shin, Christopher H. Lin, Sam Thomson, Charles Chen et al.EMNLP 2021 · 131 citations
Related papers
- DuSQL: A Large-Scale and Pragmatic Chinese Text-to-SQL DatasetLijie Wang, Ao Zhang, Kun Wu, Ke Sun et al.EMNLP 2020 · 40 citations
- Data Augmentation with Hierarchical SQL-to-Question Generation for Cross-domain Text-to-SQL ParsingKun Wu, Lijie Wang, Zhenghua Li, Ao Zhang et al.EMNLP 2021 · 22 citations
- Bridging the Generalization Gap in Text-to-SQL Parsing with Schema ExpansionChen Zhao, Yu Su, Adam Pauls, Emmanouil Antonios PlataniosACL 2022 · 19 citations
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 9 citations
- Chase: A Large-Scale and Pragmatic Chinese Dataset for Cross-Database Context-Dependent Text-to-SQLJiaqi Guo, Ziliang Si, Yu Wang, Qian Liu et al.ACL 2021
