CRUSH4SQL: Collective Retrieval Using Schema Hallucination For Text2SQL
Mayank Kothyari, Dhruva Dhingra, Sunita Sarawagi, Soumen Chakrabarti
摘要
Existing Text-to-SQL generators require the entire schema to be encoded with the user text. This is expensive or impractical for large databases with tens of thousands of columns. Standard dense retrieval techniques are inadequate for schema subsetting of a large structured database, where the correct semantics of retrieval demands that we rank sets of schema elements rather than individual elements. In response, we propose a two-stage process for effective coverage during retrieval. First, we instruct an LLM to hallucinate a minimal DB schema deemed adequate to answer the query. We use the hallucinated schema to retrieve a subset of the actual schema, by composing the results from multiple dense retrievals. Remarkably, hallucination -generally considered a nuisance -turns out to be actually useful as a bridging mechanism. Since no existing benchmarks exist for schema subsetting on large databases, we introduce three benchmarks. Two semi-synthetic datasets are derived from the union of schemas in two wellknown datasets, SPIDER and BIRD, resulting in 4502 and 798 schema elements respectively. A real-life benchmark called SocialDB is sourced from an actual large data warehouse comprising 17844 schema elements. We show that our method 1 leads to significantly higher recall than SOTA retrieval-based augmentation methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- An Audit on the Perspectives and Challenges of Hallucinations in NLPPranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs 等EMNLP 2024 · 被引用 8 次
- Text2sql-Flow: a Robust Sql-Aware Data Augmentation Framework for Text-To-SqlQifeng Cai, Hao Liang, Chang Xu, Tao Xie 等ICDE 2026 · 被引用 1 次
- Reliable Answers for Recurring Questions: Boosting Text-to-SQL Accuracy with Template Constrained DecodingSmit Jivani, Sarvam Maheshwari, Sunita SarawagiSIGMOD 2026 · 被引用 1 次
- A Comparative Evaluation of Schema Subsetting for LLM-based NL-to-SQL over Large-Schema DatabasesKyle Luoma, Arun KumarVLDB 2026
它引用的顶会 Paper2
相关 Paper
- AutoLink: Autonomous Schema Exploration and Expansion for Scalable Schema Linking in Text-to-SQL at ScaleZiyang Wang, Yuanlei Zheng, Zhenbiao Cao, Xiaojin Zhang 等AAAI 2026 · 被引用 3 次
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 被引用 909 次
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 被引用 9 次
- LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQLYihan Wang, Peiyu Liu, Xin YangEMNLP 2025 · 被引用 2 次
- Diverse Parallel Data Synthesis for Cross-Database Adaptation of Text-to-SQL ParsersAbhijeet Awasthi, Ashutosh Sathe, Sunita SarawagiEMNLP 2022 · 被引用 7 次
