SLOG: A Structural Generalization Benchmark for Semantic Parsing
Bingzhi Li, Lucia Donatelli, Alexander Koller, Tal Linzen, Yuekun Yao, Najoung Kim
Abstract
The goal of compositional generalization benchmarks is to evaluate how well models generalize to new complex linguistic expressions. Existing benchmarks often focus on lexical generalization, the interpretation of novel lexical items in syntactic structures familiar from training. Structural generalization tasks, where a model needs to interpret syntactic structures that are themselves unfamiliar from training, are often underrepresented, resulting in overly optimistic perceptions of how well models can generalize. We introduce SLOG, a semantic parsing dataset that extends COGS (Kim and Linzen, 2020) with 17 structural generalization cases. In our experiments, the generalization accuracy of Transformer models, including pretrained ones, only reaches 40.6%, while a structure-aware parser only achieves 70.8%. These results are far from the near-perfect accuracy existing models achieve on COGS, demonstrating the role of SLOG in foregrounding the large discrepancy between models' lexical and structural generalization capacities. * * This work was conducted during Bingzhi Li's visit to NYU. The middle authors are listed in alphabetical order. Training Generalization COGS Emma saw the dog. ; * dog(x3); see.agent(x1,Emma) ∧ see.theme(x1, x3) The cat ran. ; * cat(x1); run.agent(x2, x1) The dog ran. ; * dog(x1); run.agent(x2, x1) SLOG Emma saw the dog that Max held. ; * dog(x3); see.agent(x1,Emma) ∧ see.theme(x1, x3) ∧ dog.nmod(x3, x6) ∧ hold.agent(x6,Max) ∧ hold.theme(x6, x3) The cat ran. ; * cat(x1); run.agent(x2, x1) The dog that Max saw ran. ; * dog(x1); dog.nmod(x1, x4) ∧ see.agent(x4,Max) ∧ see.theme(x4, x1) ∧ run.agent(x5, x1)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df189d31-91c8-467c-8e78-c0cd21172494Cited by top-tier papers8
- Toward Compositional Behavior in Neural Models: A Survey of Current ViewsKate McCurdy, Paul Soulos, Paul Smolensky, Roland Fernandez et al.EMNLP 2024 · 12 citations
- Compositional Generalization Across Distributional Shifts with Sparse Tree OperationsPaul Soulos, Henry Conklin, Mattia Opper, Paul Smolensky et al.NeurIPS 2024 · 8 citations
- AMR Parsing is Far from Solved: GrAPES, the Granular AMR Parsing Evaluation SuiteJonas Groschwitz, Shay B. Cohen, Lucia Donatelli, Meaghan FowlieEMNLP 2023 · 5 citations
- Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated SurveyIvan Vegner, Sydelle de Souza, Valentin Forch, Martha Lewis et al.ACL 2025 · 3 citations
- Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic BiasesMichael Y. Hu, Jackson Petty, Chuan Shi, William Merrill et al.ACL 2025
Builds on9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz et al.NeurIPS 2022 · 267 citations
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
- The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of TransformersRóbert Csordás, Kazuki Irie, Jürgen SchmidhuberEMNLP 2021 · 55 citations
- Compositional Semantic Parsing with Large Language ModelsAndrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales et al.ICLR 2023 · 39 citations
Related papers
- Making Transformers Solve Compositional TasksSantiago Ontañón, Joshua Ainslie, Zachary Fisher, Vaclav CvicekACL 2022 · 87 citations
- Compositional Generalization without Trees using Multiset Tagging and Latent PermutationsMatthias Lindemann, Alexander Koller, Ivan TitovACL 2023
- Structural generalization in COGS: Supertagging is (almost) all you needAlban Petit, Caio F. Corro, François YvonEMNLP 2023
- On Evaluating Multilingual Compositional Generalization with Translated DatasetsZi Wang, Daniel HershcovichACL 2023 · 2 citations
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt et al.NeurIPS 2020 · 169 citations
