Inducing Transformer's Compositional Generalization Ability via Auxiliary Sequence Prediction Tasks
Yichen Jiang, Mohit Bansal
摘要
Systematic compositionality is an essential mechanism in human language, allowing the recombination of known parts to create novel expressions. However, existing neural models have been shown to lack this basic ability in learning symbolic structures. Motivated by the failure of a Transformer model on the SCAN compositionality challenge (Lake and Baroni, 2018), which requires parsing a command into actions, we propose two auxiliary sequence prediction tasks as additional training supervision. These automatically-generated sequences are more representative of the underlying compositional symbolic structures of the input data. During inference, the model jointly predicts the next action and the next tokens in the auxiliary sequences at each step. Experiments on the SCAN dataset show that our method encourages the Transformer to understand compositional structures of the command, improving its accuracy on multiple challenging splits from ≤ 10% to 100%. With only 418 (5%) training instances, our approach still achieves 97.8% accuracy on the MCD1 split. Therefore, we argue that compositionality can be induced in Transformers given minimal but proper guidance. We also show that a better result is achieved using less contextualized vectors as the attention's query, providing insights into architecture choices in achieving systematic compositionality. Finally, we show positive generalization results on the grounded-SCAN task (Ruis et al., 2020) . 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Evaluating the Impact of Model Scale for Compositional Generalization in Semantic ParsingLinlu Qiu, Peter Shaw, Panupong Pasupat, Tianze Shi 等EMNLP 2022 · 被引用 21 次
- Toward Compositional Behavior in Neural Models: A Survey of Current ViewsKate McCurdy, Paul Soulos, Paul Smolensky, Roland Fernandez 等EMNLP 2024 · 被引用 12 次
- NeSyCoCo: A Neuro-Symbolic Concept Composer for Compositional GeneralizationDanial Kamali, Elham J. Barezi, Parisa KordjamshidiAAAI 2025 · 被引用 4 次
- Layer-Wise Representation Fusion for Compositional GeneralizationYafang Zheng, Lei Lin, Shuangtao Li, Yuxuan Yuan 等AAAI 2024 · 被引用 4 次
- Data Factors for Better Compositional GeneralizationXiang Zhou, Yichen Jiang, Mohit BansalEMNLP 2023 · 被引用 2 次
它引用的顶会 Paper9
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman 等ICLR 2020 · 被引用 401 次
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt 等NeurIPS 2020 · 被引用 169 次
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 被引用 149 次
- Learning Compositional Rules via Neural Program SynthesisMaxwell I. Nye, Armando Solar-Lezama, Josh Tenenbaum, Brenden M. LakeNeurIPS 2020 · 被引用 120 次
- Compositional Generalization by Learning Analytical ExpressionsQian Liu, Shengnan An, Jian-Guang Lou, Bei Chen 等NeurIPS 2020 · 被引用 79 次
相关 Paper
- Learning to Recombine and Resample Data For Compositional GeneralizationEkin Akyürek, Afra Feyza Akyürek, Jacob AndreasICLR 2021 · 被引用 36 次
- Permutation Equivariant Models for Compositional Generalization in LanguageJonathan Gordon, David Lopez-Paz, Marco Baroni, Diane BouchacourtICLR 2020 · 被引用 112 次
- Mutual Exclusivity Training and Primitive Augmentation to Induce CompositionalityYichen Jiang, Xiang Zhou, Mohit BansalEMNLP 2022 · 被引用 1 次
- Towards Equipping Transformer with the Ability of Systematic CompositionalityChen Huang, Peixin Qin, Wenqiang Lei, Jiancheng LvAAAI 2024 · 被引用 3 次
- When Can Transformers Ground and Compose: Insights from Compositional Generalization BenchmarksAnkur Sikarwar, Arkil Patel, Navin GoyalEMNLP 2022 · 被引用 4 次
