Finding needles in a haystack: Sampling Structurally-diverse Training Sets from Synthetic Data for Compositional Generalization
Inbar Oren, Jonathan Herzig, Jonathan Berant
Abstract
Modern semantic parsers suffer from two principal limitations. First, training requires expensive collection of utterance-program pairs. Second, semantic parsers fail to generalize at test time to new compositions/structures that have not been observed during training. Recent research has shown that automatic generation of synthetic utterance-program pairs can alleviate the first problem, but its potential for the second has thus far been under-explored. In this work, we investigate automatic generation of synthetic utterance-program pairs for improving compositional generalization in semantic parsing. Given a small training set of annotated examples and an "infinite" pool of synthetic examples, we select a subset of synthetic examples that are structurally-diverse and use them to improve compositional generalization. We evaluate our approach on a new split of the schema2QA dataset, and show that it leads to dramatic improvements in compositional generalization as well as moderate improvements in the traditional i.i.d setup. Moreover, structurally-diverse sampling achieves these improvements with as few as 5K examples, compared to 1M examples when sampling uniformly at random -a 200x improvement in data efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 394acd5b-142b-4617-ac67-7126e41e1acdCited by top-tier papers9
- Diverse Demonstrations Improve In-context Compositional GeneralizationItay Levy, Ben Bogin, Jonathan BerantACL 2023 · 51 citations
- ExeDec: Execution Decomposition for Compositional Generalization in Neural Program SynthesisKensen Shi, Joey Hong, Yinlin Deng, Pengcheng Yin et al.ICLR 2024 · 21 citations
- Evaluating the Impact of Model Scale for Compositional Generalization in Semantic ParsingLinlu Qiu, Peter Shaw, Panupong Pasupat, Tianze Shi et al.EMNLP 2022 · 21 citations
- Unobserved Local Structures Make Compositional Generalization HardBen Bogin, Shivanshu Gupta, Jonathan BerantEMNLP 2022 · 15 citations
- How Do In-Context Examples Affect Compositional Generalization?Shengnan An, Zeqi Lin, Qiang Fu, Bei Chen et al.ACL 2023 · 15 citations
Builds on21
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman et al.ICLR 2020 · 401 citations
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers et al.ICML 2020 · 242 citations
Related papers
- Span-based Semantic Parsing for Compositional GeneralizationJonathan Herzig, Jonathan BerantACL 2021
- AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training DataSilei Xu, Sina J. Semnani, Giovanni Campagna, Monica S. LamEMNLP 2020 · 33 citations
- Laziness Is a Virtue When It Comes to Compositionality in Neural Semantic ParsingMaxwell Crouse, Pavan Kapanipathi, Subhajit Chaudhury, Tahira Naseem et al.ACL 2023 · 1 citation
- *-CFQ: Analyzing the Scalability of Machine Learning on a Compositional TaskDmitry Tsarkov, Tibor Tihon, Nathan Scales, Nikola Momchev et al.AAAI 2021 · 10 citations
- Compositional Generalization without Trees using Multiset Tagging and Latent PermutationsMatthias Lindemann, Alexander Koller, Ivan TitovACL 2023
