Knapsack Optimization-Based Schema Linking for LLM-Based Text-to-SQL Generation
Zheng Yuan, Hao Chen, Zijin Hong, Qinggang Zhang, Feiran Huang, Qing Li, Xiao Huang
摘要
Generating SQLs from user queries is a longstanding challenge, where the accuracy of initial schema linking significantly impacts subsequent SQL generation performance. However, current schema linking models still struggle with missing relevant schema elements or an excess of redundant ones. A crucial reason for this is that commonly used metrics, recall and precision, fail to capture relevant element missing and thus cannot reflect actual schema linking performance. Motivated by this, we propose enhanced schema linking metrics by introducing a restricted missing indicator. Accordingly, we introduce Knapsack optimization-based Schema Linking Approach (KaSLA), a plug-in schema linking method designed to prevent the missing of relevant schema elements while minimizing the inclusion of redundant ones. KaSLA employs a hierarchical linking strategy that first identifies the optimal table linking and subsequently links columns within the selected table to reduce linking candidate space. In each linking process, it utilizes a knapsack optimization approach to link potentially relevant elements while accounting for a limited tolerance of potentially redundant ones. With this optimization, KaSLA-1.6B achieves superior schema linking results compared to large-scale LLMs, including DeepSeek-V3 with the state-of-theart (SOTA) schema linking method. Extensive experiments on Spider and BIRD benchmarks verify that KaSLA can significantly improve the SQL generation performance of SOTA Text2SQL models by substituting their schema linking processes. The code is available at https://github.com/DEEP-PolyU/KaSLA.
Index Terms-text-to-SQL, database, large language models, natural language understanding a significant 14.91% performance gap between state-of-theart schema linking methods and ground-truth linking results. Consequently, enhancing schema linking accuracy represents a critical research frontier with significant potential to advance text-to-SQL capabilities.
Recent state-of-the-art text-to-SQL models primarily focus on the final SQL generation stage, often relying on basic or even neglecting schema linking strategies. Existing schema
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- LinearRAG: Linear Graph Retrieval Augmented Generation on Large-scale CorporaLuyao Zhuang, Shengyuan Chen, Yilin Xiao, Huachi Zhou 等ICLR 2026 · 被引用 54 次
- FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented GenerationQinggang Zhang, Zhishang Xiang, Yilin Xiao, Le Wang 等ACL 2025 · 被引用 18 次
- ReEx-SQL: Reasoning with Execution-Aware Reinforcement Learning for Text-to-SQLYaxun Dai, Wenxuan Xie, Xialie Zhuang, Tianyu Yang 等ACL 2026 · 被引用 8 次
- OpenSQL: Data-Efficient Text-to-SQL for Open-Source LLMs via Synthesized Intermediate SupervisionRuilin Hu, Yuyu Luo, Guoliang Li, Shuangqiao Wu 等VLDB 2026 · 被引用 4 次
- ErrorLLM: Modeling SQL Errors for Text-to-SQL RefinementZijin Hong, Hao Chen, Zheng Yuan, Qinggang Zhang 等KDD 2026 · 被引用 3 次
它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 被引用 909 次
相关 Paper
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 被引用 9 次
- Graph-Link: Bridging the Semantic-Structural Gap in Text-to-SQL via Constrained Subgraph InductionJianwei Zhong, Yuxi Yang, Quanxin Liu, Ruida Xu 等ICML 2026
- A Comparative Evaluation of Schema Subsetting for LLM-based NL-to-SQL over Large-Schema DatabasesKyle Luoma, Arun KumarVLDB 2026
- JOLT-SQL: Joint Loss Tuning of Text-to-SQL with Confusion-aware Noisy Schema SamplingJinwang Song, Hongying Zan, Kunli Zhang, Lingling Mu 等EMNLP 2025
- Re-examining the Role of Schema Linking in Text-to-SQLWenqiang Lei, Weixin Wang, Zhixin Ma, Tian Gan 等EMNLP 2020 · 被引用 71 次
