Bridging the Generalization Gap in Text-to-SQL Parsing with Schema Expansion
Chen Zhao, Yu Su, Adam Pauls, Emmanouil Antonios Platanios
摘要
Text-to-SQL parsers map natural language questions to programs that are executable over tables to generate answers, and are typically evaluated on large-scale datasets like SPIDER (Yu et al., 2018) . We argue that existing benchmarks fail to capture a certain out-of-domain generalization problem that is of significant practical importance: matching domain specific phrases to composite operations over columns. To study this problem, we propose a synthetic dataset and a re-purposed train/test split of the SQUALL dataset (Shi et al., 2020) as new benchmarks to quantify domain generalization over column operations. Our results indicate that existing state-of-the-art parsers struggle in these benchmarks. We propose to address this problem by incorporating prior domain knowledge by preprocessing table schemas, and design a method that consists of two components: schema expansion and schema pruning. This method can be easily applied to multiple existing base parsers, and we show that it significantly outperforms baseline parsers on this domain generalization problem, boosting the underlying parsers' overall performance by up to 13.8% relative accuracy gain (5.1% absolute) on the new SQUALL data split.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PLOG: Table-to-Logic Pretraining for Logical Table-to-Text GenerationAo Liu, Haoyu Dong, Naoaki Okazaki, Shi Han 等EMNLP 2022 · 被引用 15 次
- Towards Knowledge-Intensive Text-to-SQL Semantic Parsing with Formulaic KnowledgeLongxu Dou, Yan Gao, Xuqi Liu, Mingyang Pan 等EMNLP 2022 · 被引用 11 次
- RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial PerturbationsYilun Zhao, Chen Zhao, Linyong Nan, Zhenting Qi 等ACL 2023 · 被引用 7 次
- MAGNIFICo: Evaluating the In-Context Learning Ability of Large Language Models to Generalize to Novel InterpretationsArkil Patel, Satwik Bhattamishra, Siva Reddy, Dzmitry BahdanauEMNLP 2023 · 被引用 2 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman 等ICLR 2020 · 被引用 401 次
- Beyond I.I.D.: Three Levels of Generalization for Question Answering on Knowledge BasesYu Gu, Sue Kase, Michelle Vanni, Brian M. Sadler 等WWW 2021 · 被引用 304 次
- Exploring Unexplored Generalization Challenges for Cross-Database Semantic ParsingAlane Suhr, Ming-Wei Chang, Peter Shaw, Kenton LeeACL 2020 · 被引用 76 次
相关 Paper
- Graphix-T5: Mixing Pre-trained Transformers with Graph-Aware Layers for Text-to-SQL ParsingJinyang Li, Binyuan Hui, Reynold Cheng, Bowen Qin 等AAAI 2023 · 被引用 164 次
- RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersBailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov 等ACL 2020 · 被引用 39 次
- KaggleDBQA: Realistic Evaluation of Text-to-SQL ParsersChia-Hsuan Lee, Oleksandr Polozov, Matthew RichardsonACL 2021
- Diverse Parallel Data Synthesis for Cross-Database Adaptation of Text-to-SQL ParsersAbhijeet Awasthi, Ashutosh Sathe, Sunita SarawagiEMNLP 2022 · 被引用 7 次
- Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL RobustnessShuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan 等ICLR 2023 · 被引用 9 次
