*-CFQ: Analyzing the Scalability of Machine Learning on a Compositional Task
Dmitry Tsarkov, Tibor Tihon, Nathan Scales, Nikola Momchev, Danila Sinopalnikov, Nathanael Schärli
Abstract
We present *-CFQ ("star-CFQ"): a suite of large-scale datasets of varying scope based on the CFQ semantic parsing benchmark, designed for principled investigation of the scalability of machine learning systems in a realistic compositional task setting. Using this suite, we conduct a series of experiments investigating the ability of Transformers to benefit from increased training data size under conditions of fixed computational cost. We show that compositional generalization remains a challenge at all training sizes, and we show that increasing the scope of natural language leads to consistently higher error rates, which are only partially offset by increased training data. We further show that while additional training data from a related domain improves the accuracy in data-starved situations, this improvement is limited and diminishes as the distance from the related domain to the target domain increases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74fd2397-d0b0-4dc7-b67a-086a1393710fCited by top-tier papers10
- Systematic Generalization with Edge TransformersLeon Bergen, Timothy J. O'Donnell, Dzmitry BahdanauNeurIPS 2021 · 62 citations
- Evaluating the Impact of Model Scale for Compositional Generalization in Semantic ParsingLinlu Qiu, Peter Shaw, Panupong Pasupat, Tianze Shi et al.EMNLP 2022 · 21 citations
- Understanding Robust Generalization in Learning Regular LanguagesSoham Dan, Osbert Bastani, Dan RothICML 2022 · 5 citations
- Data Factors for Better Compositional GeneralizationXiang Zhou, Yichen Jiang, Mohit BansalEMNLP 2023 · 2 citations
- On Evaluating Multilingual Compositional Generalization with Translated DatasetsZi Wang, Daniel HershcovichACL 2023 · 2 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman et al.ICLR 2020 · 401 citations
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 265 citations
- Permutation Equivariant Models for Compositional Generalization in LanguageJonathan Gordon, David Lopez-Paz, Marco Baroni, Diane BouchacourtICLR 2020 · 112 citations
- Environmental drivers of systematicity and generalization in a situated agentFelix Hill, Andrew K. Lampinen, Rosalia Schneider, Stephen Clark et al.ICLR 2020 · 109 citations
Related papers
- Compositional Generalization in Dependency ParsingEmily Goodwin, Siva Reddy, Timothy J. O'Donnell, Dzmitry BahdanauACL 2022
- Making Transformers Solve Compositional TasksSantiago Ontañón, Joshua Ainslie, Zachary Fisher, Vaclav CvicekACL 2022 · 87 citations
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
- Finding needles in a haystack: Sampling Structurally-diverse Training Sets from Synthetic Data for Compositional GeneralizationInbar Oren, Jonathan Herzig, Jonathan BerantEMNLP 2021
- Compositional Semantic Parsing with Large Language ModelsAndrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales et al.ICLR 2023 · 39 citations
