Compositional Generalization without Trees using Multiset Tagging and Latent Permutations
Matthias Lindemann, Alexander Koller, Ivan Titov
Abstract
Seq2seq models have been shown to struggle with compositional generalization in semantic parsing, i.e. generalizing to unseen compositions of phenomena that the model handles correctly in isolation. We phrase semantic parsing as a two-step process: we first tag each input token with a multiset of output tokens. Then we arrange the tokens into an output sequence using a new way of parameterizing and predicting permutations. We formulate predicting a permutation as solving a regularized linear program and we backpropagate through the solver. In contrast to prior work, our approach does not place a priori restrictions on possible permutations, making it very expressive. Our model outperforms pretrained seq2seq models and prior work on realistic semantic parsing tasks that require generalization to longer examples. We also outperform non-tree-based models on structural generalization on the COGS benchmark. For the first time, we show that a model without an inductive bias provided by trees achieves high accuracy on generalization to deeper recursion depth. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Compositional Generalization Across Distributional Shifts with Sparse Tree OperationsPaul Soulos, Henry Conklin, Mattia Opper, Paul Smolensky et al.NeurIPS 2024 · 8 citations
- SIP: Injecting a Structural Inductive Bias into a Seq2Seq Model by SimulationMatthias Lindemann, Alexander Koller, Ivan TitovACL 2024 · 2 citations
- RankGuess: Password Guessing Using Adversarial RankingTao Yang, Ding WangS&P 2025
- Compositional Generalisation for Explainable Hate Speech DetectionAgostina Calabrese, Tom Sherborne, Björn Ross, Mirella LapataEMNLP 2025
- Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic TransformationsMatthias Lindemann, Alexander Koller, Ivan TitovEMNLP 2024
Builds on13
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
- The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic GeneralizationRóbert Csordás, Kazuki Irie, Jürgen SchmidhuberICLR 2022 · 70 citations
- The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of TransformersRóbert Csordás, Kazuki Irie, Jürgen SchmidhuberEMNLP 2021 · 55 citations
- Sequence-to-Sequence Learning with Latent Neural GrammarsYoon KimNeurIPS 2021 · 44 citations
Related papers
- Span-based Semantic Parsing for Compositional GeneralizationJonathan Herzig, Jonathan BerantACL 2021
- Structural generalization in COGS: Supertagging is (almost) all you needAlban Petit, Caio F. Corro, François YvonEMNLP 2023
- Compositional Generalization and Natural Language Variation: Can a Semantic Parsing Approach Handle Both?Peter Shaw, Ming-Wei Chang, Panupong Pasupat, Kristina ToutanovaACL 2021
- Making Transformers Solve Compositional TasksSantiago Ontañón, Joshua Ainslie, Zachary Fisher, Vaclav CvicekACL 2022 · 87 citations
- Structural generalization is hard for sequence-to-sequence modelsYuekun Yao, Alexander KollerEMNLP 2022 · 8 citations
