MT-Teql: Evaluating and Augmenting Neural NLIDB on Real-world Linguistic and Schema Variations
Pingchuan Ma, Shuai Wang
Abstract
Natural Language Interface to Database (NLIDB) translates human utterances into SQL queries and enables database interactions for non-expert users. Recently, neural network models have become a major approach to implementing NLIDB. However, neural NLIDB faces challenges due to variations in natural language and database schema design. For instance, one user intent or database conceptual model can be expressed in various forms. However, existing benchmarks, using hold-out datasets, cannot provide thorough understanding of how good neural NLIDBs really are in real-world situations and its robustness against such variations. A key difficulty is to annotate SQL queries for inputs under real-world variations, requiring considerable manual effort and expert knowledge.
To systematically assess the robustness of neural NLIDBs without extensive manual effort, we propose MT-Teql, a unified framework to benchmark NLIDBs against real-world language and schema variations. Inspired by recent advances in DBMS metamorphic testing, MT-Teql implements semantics-preserving transformations on utterances and database schemas to generate their variants. NLIDBs can thus be examined for robustness utilizing utterances/schemas and their variants without requiring manual intervention. We benchmarked nine neural NLIDBs using 62,430 inputs and identified 15,433 defects. We analyzed potential root causes of defects and conducted a user study to show how MT-Teql can assist developers to systematically assess NLIDBs. We further show that the transformed (error-triggering) inputs can be used to augment popular NLIDBs and eliminate 46.5%(±5.0%) errors made by them without compromising their accuracy on standard benchmarks. We summarize lessons from this study that can provide insights to select and design NLIDBs that fit particular usage scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7fb83f6-2ec3-4669-a967-07d80433214cCited by top-tier papers13
- CatSQL: Towards Real World Natural Language to SQL ApplicationsHan Fu, Chang Liu, Bin Wu, Feifei Li et al.VLDB 2023 · 79 citations
- ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL SystemsYi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten et al.VLDB 2024 · 65 citations
- MDPFuzz: testing models solving Markov decision processesQi Pang, Yuanyuan Yuan, Shuai WangISSTA 2022 · 37 citations
- Unleashing the Power of Compiler Intermediate Representation to Enhance Neural Program EmbeddingsZongjie Li, Pingchuan Ma, Huaijin Wang, Shuai Wang et al.ICSE 2022 · 25 citations
- Testing Graph Database Systems via Graph-Aware Metamorphic RelationsZeyang Zhuang, Penghui Li, Pingchuan Ma, Wei Meng et al.VLDB 2024 · 23 citations
Builds on12
- Testing Database Engines via Pivoted Query SynthesisManuel Rigger, Zhendong SuOSDI 2020 · 150 citations
- Finding bugs in database systems via query partitioningManuel Rigger, Zhendong SuOOPSLA 2020 · 116 citations
- Generate Your Counterfactuals: Towards Controlled Counterfactual Generation for TextNishtha Madaan, Inkit Padhi, Naveen Panwar, Diptikalyan SahaAAAI 2021 · 115 citations
- Detecting optimization bugs in database engines via non-optimizing reference engine constructionManuel Rigger, Zhendong SuFSE 2020 · 104 citations
- Semantic Evaluation for Text-to-SQL with Distilled Test SuitesRuiqi Zhong, Tao Yu, Dan KleinEMNLP 2020 · 88 citations
Related papers
- Metasql: A Generate-Then-Rank Framework for Natural Language to SQL TranslationYuankai Fan, Zhenying He, Tonghui Ren, Can Huang et al.ICDE 2024 · 23 citations
- Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL RobustnessShuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan et al.ICLR 2023 · 9 citations
- EVOSCHEMA: TOWARDS TEXT-TO-SQL ROBUSTNESS AGAINST SCHEMA EVOLUTIONTianshu Zhang, Kun Qian, Siddhartha Sahai, Yuan Tian et al.VLDB 2025 · 3 citations
- Gar: A Generate-and-Rank Approach for Natural Language to SQL TranslationYuankai Fan, Zhenying He, Tonghui Ren, Dianjun Guo et al.ICDE 2023 · 12 citations
- Natural language to SQL: Where are we today?Hyeonji Kim, Byeong-Hoon So, Wook-Shin Han, Hongrae LeeVLDB 2020 · 148 citations
