Learning to Substitute Spans towards Improving Compositional Generalization
Zhaoyi Li, Ying Wei, Defu Lian
摘要
Despite the rising prevalence of neural sequence models, recent empirical evidences suggest their deficiency in compositional generalization. One of the current de-facto solutions to this problem is compositional data augmentation, aiming to incur additional compositional inductive bias. Nonetheless, the improvement offered by existing handcrafted augmentation strategies is limited when successful systematic generalization of neural sequence models requires multi-grained compositional bias (i.e., not limited to either lexical or structural biases only) or differentiation of training sequences in an imbalanced difficulty distribution. To address the two challenges, we first propose a novel compositional augmentation strategy dubbed <b>Span Sub</b>stitution (SpanSub) that enables multi-grained composition of substantial substructures in the whole training set. Over and above that, we introduce the <b>L</b>earning <b>to</b> <b>S</b>ubstitute <b>S</b>pan (L2S2) framework which empowers the learning of span substitution probabilities in SpanSub in an end-to-end manner by maximizing the loss of neural sequence models, so as to outweigh those challenging compositions with elusive concepts and novel surroundings. Our empirical results on three standard compositional generalization benchmarks, including SCAN, COGS and GeoQuery (with an improvement of at most 66.5%, 10.3%, 1.2%, respectively), demonstrate the superiority of SpanSub, L2S2 and their combination. © 2023 Association for Computational Linguistics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Layer-Wise Representation Fusion for Compositional GeneralizationYafang Zheng, Lei Lin, Shuangtao Li, Yuxuan Yuan 等AAAI 2024 · 被引用 4 次
- Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language ModelsZhaoyi Li, Jiatong Li, Gangwei Jiang, Linqi Song 等ICLR 2026 · 被引用 1 次
- Benchmarking and Improving Compositional Generalization of Multi-aspect Controllable Text GenerationTianqi Zhong, Zhaoyi Li, Quan Wang, Linqi Song 等ACL 2024
- Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic TransformationsMatthias Lindemann, Alexander Koller, Ivan TitovEMNLP 2024
- CARMA: Enhanced Compositionality in LLMs via Advanced Regularisation and Mutual Information AlignmentNura Aljaafari, Danilo S. Carvalho, André FreitasEMNLP 2025
它引用的顶会 Paper13
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman 等ICLR 2020 · 被引用 401 次
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 被引用 149 次
- Permutation Equivariant Models for Compositional Generalization in LanguageJonathan Gordon, David Lopez-Paz, Marco Baroni, Diane BouchacourtICLR 2020 · 被引用 112 次
- Making Transformers Solve Compositional TasksSantiago Ontañón, Joshua Ainslie, Zachary Fisher, Vaclav CvicekACL 2022 · 被引用 87 次
相关 Paper
- Mutual Exclusivity Training and Primitive Augmentation to Induce CompositionalityYichen Jiang, Xiang Zhou, Mohit BansalEMNLP 2022 · 被引用 1 次
- Good-Enough Compositional Data AugmentationJacob AndreasACL 2020 · 被引用 15 次
- Learning to Recombine and Resample Data For Compositional GeneralizationEkin Akyürek, Afra Feyza Akyürek, Jacob AndreasICLR 2021 · 被引用 36 次
- LexSym: Compositionality as Lexical SymmetryEkin Akyürek, Jacob AndreasACL 2023 · 被引用 9 次
- Meta-Learning to Compositionally GeneralizeHenry Conklin, Bailin Wang, Kenny Smith, Ivan TitovACL 2021
