AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training Data
Silei Xu, Sina J. Semnani, Giovanni Campagna, Monica S. Lam
摘要
We propose AutoQA, a methodology and toolkit to generate semantic parsers that answer questions on databases, with no manual effort. Given a database schema and its data, AutoQA automatically generates a large set of high-quality questions for training that covers different database operations. It uses automatic paraphrasing combined with templatebased parsing to find alternative expressions of an attribute in different parts of speech. It also uses a novel filtered auto-paraphraser to generate correct paraphrases of entire sentences. We apply AutoQA to the Schema2QA dataset and obtain an average logical form accuracy of 62.9% when tested on natural questions, which is only 6.4% lower than a model trained with expert natural language annotations and paraphrase data collected from crowdworkers. To demonstrate the generality of AutoQA, we also apply it to the Overnight dataset. AutoQA achieves 69.8% answer accuracy, 16.4% higher than the state-of-the-art zero-shot models and only 5.2% lower than the same model trained with human data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Constrained Language Models Yield Few-Shot Semantic ParsersRichard Shin, Christopher H. Lin, Sam Thomson, Charles Chen 等EMNLP 2021 · 被引用 131 次
- GraphQ IR: Unifying the Semantic Parsing of Graph Query Languages with One Intermediate RepresentationLunyiu Nie, Shulin Cao, Jiaxin Shi, Jiuding Sun 等EMNLP 2022 · 被引用 20 次
- GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and LearningMinh Duc Vu, Han Wang, Jieshan Chen, Zhuang Li 等UIST 2024 · 被引用 17 次
- On The Ingredients of an Effective Zero-shot Semantic ParserPengcheng Yin, John Wieting, Avirup Sil, Graham NeubigACL 2022 · 被引用 15 次
- Voicify Your UI: Towards Android App Control with Voice CommandsMinh Duc Vu, Han Wang, Zhuang Li, Gholamreza Haffari 等UbiComp 2023 · 被引用 13 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta 等AAAI 2020 · 被引用 707 次
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor 等AAAI 2020 · 被引用 398 次
- Revisiting Self-Training for Neural Sequence GenerationJunxian He, Jiatao Gu, Jiajun Shen, Marc'Aurelio RanzatoICLR 2020 · 被引用 294 次
相关 Paper
- Uni-Parser: Unified Semantic Parser for Question Answering on Knowledge Base and DatabaseYe Liu, Semih Yavuz, Rui Meng, Dragomir Radev 等EMNLP 2022 · 被引用 21 次
- KaggleDBQA: Realistic Evaluation of Text-to-SQL ParsersChia-Hsuan Lee, Oleksandr Polozov, Matthew RichardsonACL 2021
- Grounded Adaptation for Zero-shot Executable Semantic ParsingVictor Zhong, Mike Lewis, Sida I. Wang, Luke ZettlemoyerEMNLP 2020 · 被引用 85 次
- Finding needles in a haystack: Sampling Structurally-diverse Training Sets from Synthetic Data for Compositional GeneralizationInbar Oren, Jonathan Herzig, Jonathan BerantEMNLP 2021
- Diverse Parallel Data Synthesis for Cross-Database Adaptation of Text-to-SQL ParsersAbhijeet Awasthi, Ashutosh Sathe, Sunita SarawagiEMNLP 2022 · 被引用 7 次
