CRUSH4SQL: Collective Retrieval Using Schema Hallucination For Text2SQL
Mayank Kothyari, Dhruva Dhingra, Sunita Sarawagi, Soumen Chakrabarti
Abstract
Existing Text-to-SQL generators require the entire schema to be encoded with the user text. This is expensive or impractical for large databases with tens of thousands of columns. Standard dense retrieval techniques are inadequate for schema subsetting of a large structured database, where the correct semantics of retrieval demands that we rank sets of schema elements rather than individual elements. In response, we propose a two-stage process for effective coverage during retrieval. First, we instruct an LLM to hallucinate a minimal DB schema deemed adequate to answer the query. We use the hallucinated schema to retrieve a subset of the actual schema, by composing the results from multiple dense retrievals. Remarkably, hallucination -generally considered a nuisance -turns out to be actually useful as a bridging mechanism. Since no existing benchmarks exist for schema subsetting on large databases, we introduce three benchmarks. Two semi-synthetic datasets are derived from the union of schemas in two wellknown datasets, SPIDER and BIRD, resulting in 4502 and 798 schema elements respectively. A real-life benchmark called SocialDB is sourced from an actual large data warehouse comprising 17844 schema elements. We show that our method 1 leads to significantly higher recall than SOTA retrieval-based augmentation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- An Audit on the Perspectives and Challenges of Hallucinations in NLPPranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs et al.EMNLP 2024 · 8 citations
- Text2sql-Flow: a Robust Sql-Aware Data Augmentation Framework for Text-To-SqlQifeng Cai, Hao Liang, Chang Xu, Tao Xie et al.ICDE 2026 · 1 citation
- Reliable Answers for Recurring Questions: Boosting Text-to-SQL Accuracy with Template Constrained DecodingSmit Jivani, Sarvam Maheshwari, Sunita SarawagiSIGMOD 2026 · 1 citation
- A Comparative Evaluation of Schema Subsetting for LLM-based NL-to-SQL over Large-Schema DatabasesKyle Luoma, Arun KumarVLDB 2026
Builds on2
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersBailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov et al.ACL 2020 · 39 citations
Related papers
- AutoLink: Autonomous Schema Exploration and Expansion for Scalable Schema Linking in Text-to-SQL at ScaleZiyang Wang, Yuanlei Zheng, Zhenbiao Cao, Xiaojin Zhang et al.AAAI 2026 · 3 citations
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 909 citations
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 9 citations
- LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQLYihan Wang, Peiyu Liu, Xin YangEMNLP 2025 · 2 citations
- Diverse Parallel Data Synthesis for Cross-Database Adaptation of Text-to-SQL ParsersAbhijeet Awasthi, Ashutosh Sathe, Sunita SarawagiEMNLP 2022 · 7 citations
