Generating Succinct Descriptions of Database Schemata for Cost-Efficient Prompting of Large Language Models
Immanuel Trummer
Abstract
Using large language models (LLMs) for tasks like text-to-SQL translation often requires describing the database schema as part of the model input. LLM providers typically charge as a function of the number of tokens read. Hence, reducing the length of the schema description saves money at each model invocation. This paper introduces Schemonic, a system that automatically finds concise text descriptions of relational database schemata. By introducing abbreviations or grouping schema elements with similar properties, Schemonic typically finds descriptions that use significantly fewer tokens than naive schema representations. Internally, Schemonic models schema compression as a combinatorial optimization problem and uses integer linear programming solvers to find guaranteed optimal or near-optimal solutions. It speeds up optimization by starting optimization from heuristic solutions and reducing the search space size via pre-processing. The experiments on TPC-H, SPIDER, and Public-BI demonstrate that Schemonic reduces schema description length significantly, along with fees for reading them, without reducing the accuracy in tasks such as text-to-SQL translation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9352038-9dad-4ea7-a752-772a2bbcb26dCited by top-tier papers4
- ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification QueriesKeke Huang, Yimin Shi, Dujian Ding, Yifei Li et al.VLDB 2025 · 18 citations
- Unveiling Challenges for LLMs in Enterprise Data EngineeringJan-Micha Bodensohn, Ulf Brackmann, Liane Vogel, Anupam Sanghi et al.VLDB 2026 · 13 citations
- SEFRQO: A Self-Evolving Fine-Tuned RAG-Based Query OptimizerHanwen Liu, Qihan Zhang, Ryan Marcus, Ibrahim SabekSIGMOD 2026 · 2 citations
- Accurate Table Question Answering with Accessible LLMsYangfan Jiang, Fei Wei, Ergute Bao, Yaliang Li et al.ICDE 2026 · 1 citation
Builds on14
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 682 citations
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun et al.VLDB 2024 · 609 citations
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 488 citations
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou et al.ICLR 2024 · 424 citations
Related papers
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 9 citations
- SQUiD: Synthesizing Relational Databases from Unstructured TextMushtari Sadia, Zhenning Yang, Yunming Xiao, Ang Chen et al.EMNLP 2025
- RISE: Rule-Driven SQL Dialect Translation via Query ReductionXudong Xie, Yuwei Zhang, Wensheng Dou, Yu Gao et al.ICSE 2026
- SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQLYue Gong, Chuan Lei, Xiao Qin, Kapil Vaidya et al.NeurIPS 2025 · 21 citations
- λ-Tune: Harnessing Large Language Models for Automated Database System TuningVictor Giannakouris, Immanuel TrummerSIGMOD 2025 · 20 citations
