Towards Equipping Transformer with the Ability of Systematic Compositionality
Chen Huang, Peixin Qin, Wenqiang Lei, Jiancheng Lv
Abstract
One of the key factors in language productivity and human cognition is the ability of Systematic Compositionality, which refers to understanding composed, unseen examples of seen primitives. However, recent evidence reveals that the Transformers have difficulty in generalizing the composed context based on the seen primitives. To this end, we take the first step to propose a compositionality-aware Transformer called CAT and two novel pre-training tasks to facilitate the systematic compositionality. We tentatively provide a successful implementation of a multi-layer CAT on the basis of the especially popular BERT. The experimental results demonstrate that CAT outperforms baselines on compositionality-aware tasks with minimal impact on effectiveness on standardized language understanding tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6cb8ef9-ee35-4f27-b999-5e595e78bf57Cited by top-tier papers2
- Can Large Language Models Understand Internet Buzzwords Through User-Generated ContentChen Huang, Junkai Luo, Xinzuo Wang, Wenqiang Lei et al.ACL 2025
- Composition-Incremental Learning for Compositional GeneralizationZhen Li, Yuwei Wu, Chenchen Jing, Che Sun et al.AAAI 2026
Builds on11
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman et al.ICLR 2020 · 401 citations
- Permutation Equivariant Models for Compositional Generalization in LanguageJonathan Gordon, David Lopez-Paz, Marco Baroni, Diane BouchacourtICLR 2020 · 112 citations
- Assessing Phrasal Representation and Composition in TransformersLang Yu, Allyson EttingerEMNLP 2020 · 60 citations
- Discrete-Valued Neural CommunicationDianbo Liu, Alex Lamb, Kenji Kawaguchi, Anirudh Goyal et al.NeurIPS 2021 · 55 citations
Related papers
- Dissecting Chain-of-Thought: Compositionality through In-Context Filtering and LearningYingcong Li, Kartik Sreenivasan, Angeliki Giannou, Dimitris Papailiopoulos et al.NeurIPS 2023 · 12 citations
- Compositional Task Representations for Large Language ModelsNan Shao, Zefan Cai, Hanwei Xu, Chonghua Liao et al.ICLR 2023
- Inducing Transformer's Compositional Generalization Ability via Auxiliary Sequence Prediction TasksYichen Jiang, Mohit BansalEMNLP 2021
- Attention as a HypernetworkSimon Schug, Seijin Kobayashi, Yassir Akram, João Sacramento et al.ICLR 2025
- How Do In-Context Examples Affect Compositional Generalization?Shengnan An, Zeqi Lin, Qiang Fu, Bei Chen et al.ACL 2023 · 15 citations
