SLOG: A Structural Generalization Benchmark for Semantic Parsing
Bingzhi Li, Lucia Donatelli, Alexander Koller, Tal Linzen, Yuekun Yao, Najoung Kim
摘要
The goal of compositional generalization benchmarks is to evaluate how well models generalize to new complex linguistic expressions. Existing benchmarks often focus on lexical generalization, the interpretation of novel lexical items in syntactic structures familiar from training. Structural generalization tasks, where a model needs to interpret syntactic structures that are themselves unfamiliar from training, are often underrepresented, resulting in overly optimistic perceptions of how well models can generalize. We introduce SLOG, a semantic parsing dataset that extends COGS (Kim and Linzen, 2020) with 17 structural generalization cases. In our experiments, the generalization accuracy of Transformer models, including pretrained ones, only reaches 40.6%, while a structure-aware parser only achieves 70.8%. These results are far from the near-perfect accuracy existing models achieve on COGS, demonstrating the role of SLOG in foregrounding the large discrepancy between models' lexical and structural generalization capacities. * * This work was conducted during Bingzhi Li's visit to NYU. The middle authors are listed in alphabetical order. Training Generalization COGS Emma saw the dog. ; * dog(x3); see.agent(x1,Emma) ∧ see.theme(x1, x3) The cat ran. ; * cat(x1); run.agent(x2, x1) The dog ran. ; * dog(x1); run.agent(x2, x1) SLOG Emma saw the dog that Max held. ; * dog(x3); see.agent(x1,Emma) ∧ see.theme(x1, x3) ∧ dog.nmod(x3, x6) ∧ hold.agent(x6,Max) ∧ hold.theme(x6, x3) The cat ran. ; * cat(x1); run.agent(x2, x1) The dog that Max saw ran. ; * dog(x1); dog.nmod(x1, x4) ∧ see.agent(x4,Max) ∧ see.theme(x4, x1) ∧ run.agent(x5, x1)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Toward Compositional Behavior in Neural Models: A Survey of Current ViewsKate McCurdy, Paul Soulos, Paul Smolensky, Roland Fernandez 等EMNLP 2024 · 被引用 12 次
- Compositional Generalization Across Distributional Shifts with Sparse Tree OperationsPaul Soulos, Henry Conklin, Mattia Opper, Paul Smolensky 等NeurIPS 2024 · 被引用 8 次
- AMR Parsing is Far from Solved: GrAPES, the Granular AMR Parsing Evaluation SuiteJonas Groschwitz, Shay B. Cohen, Lucia Donatelli, Meaghan FowlieEMNLP 2023 · 被引用 5 次
- Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated SurveyIvan Vegner, Sydelle de Souza, Valentin Forch, Martha Lewis 等ACL 2025 · 被引用 3 次
- Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic BiasesMichael Y. Hu, Jackson Petty, Chuan Shi, William Merrill 等ACL 2025
它引用的顶会 Paper9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz 等NeurIPS 2022 · 被引用 267 次
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 被引用 149 次
- The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of TransformersRóbert Csordás, Kazuki Irie, Jürgen SchmidhuberEMNLP 2021 · 被引用 55 次
- Compositional Semantic Parsing with Large Language ModelsAndrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales 等ICLR 2023 · 被引用 39 次
相关 Paper
- Making Transformers Solve Compositional TasksSantiago Ontañón, Joshua Ainslie, Zachary Fisher, Vaclav CvicekACL 2022 · 被引用 87 次
- Compositional Generalization without Trees using Multiset Tagging and Latent PermutationsMatthias Lindemann, Alexander Koller, Ivan TitovACL 2023
- Structural generalization in COGS: Supertagging is (almost) all you needAlban Petit, Caio F. Corro, François YvonEMNLP 2023
- On Evaluating Multilingual Compositional Generalization with Translated DatasetsZi Wang, Daniel HershcovichACL 2023 · 被引用 2 次
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt 等NeurIPS 2020 · 被引用 169 次
