Disentangled Sequence to Sequence Learning for Compositional Generalization
Hao Zheng, Mirella Lapata
Abstract
There is mounting evidence that existing neural network models, in particular the very popular sequence-to-sequence architecture, struggle to systematically generalize to unseen compositions of seen components. We demonstrate that one of the reasons hindering compositional generalization relates to representations being entangled. We propose an extension to sequence-to-sequence models which encourages disentanglement by adaptively reencoding (at each time step) the source input. Specifically, we condition the source representations on the newly decoded target context which makes it easier for the encoder to exploit specialized information for each prediction rather than capturing it all in a single forward pass. Experimental results on semantic parsing and machine translation empirically show that our proposal delivers more disentangled representations and better generalization. 1 1 Our code is available at https://github.com/ mswellhao/Dangle . Training Set A boy ate the cake on the table in a house. *cake(x4); *table(x7); boy(x1) AND eat.agent(x2, x1) AND eat.theme(x2, x4) AND cake.nmod.on(x4, x7) AND table.nmod.in(x7, x10) AND house(x10) Test Set (Lexical Generalization) A boy likes the cake on the table in a house. *cake(x4); *table(x7); boy(x1) AND like.agent(x2, x1) AND like.theme(x2, x4) AND cake.nmod.on(x4, x7) AND table.nmod.in(x7, x10) AND house(x10) Test Set (Structural Generalization) A boy ate the cake on the table in a house beside the tree. *cake(x4); *table(x7); *tree(x13); boy(x1) AND eat.agent(x2, x1) AND eat.theme(x2, x4) AND cake.nmod.on(x4, x7) AND table.nmod.in(x7, x10) AND house(x10) AND house.nmod.beside(x10, x13)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e15c75b7-b589-4742-9f5c-1084728ddcb7Cited by top-tier papers12
- Compositional Generalization for Multi-Label Text Classification: A Data-Augmentation ApproachYuyang Chai, Zhuang Li, Jiahui Liu, Lei Chen et al.AAAI 2024 · 18 citations
- Tripod: Three Complementary Inductive Biases for Disentangled Representation LearningKyle Hsu, Jubayer Ibn Hamid, Kaylee Burns, Chelsea Finn et al.ICML 2024 · 13 citations
- Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureLinyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu et al.EMNLP 2023 · 12 citations
- Structural generalization is hard for sequence-to-sequence modelsYuekun Yao, Alexander KollerEMNLP 2022 · 8 citations
- Sparse Universal TransformerShawn Tan, Yikang Shen, Zhenfang Chen, Aaron C. Courville et al.EMNLP 2023 · 6 citations
Builds on12
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain et al.NeurIPS 2020 · 503 citations
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman et al.ICLR 2020 · 401 citations
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
- Improving Disentangled Text Representation Learning with Information-Theoretic GuidancePengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon et al.ACL 2020 · 66 citations
Related papers
- Compositional Generalization without Trees using Multiset Tagging and Latent PermutationsMatthias Lindemann, Alexander Koller, Ivan TitovACL 2023
- Consistency Regularization Training for Compositional GeneralizationYongjing Yin, Jiali Zeng, Yafu Li, Fandong Meng et al.ACL 2023 · 5 citations
- Non-autoregressive Machine Translation with Disentangled Context TransformerJungo Kasai, James Cross, Marjan Ghazvininejad, Jiatao GuICML 2020 · 113 citations
- Layer-Wise Representation Fusion for Compositional GeneralizationYafang Zheng, Lei Lin, Shuangtao Li, Yuxuan Yuan et al.AAAI 2024 · 4 citations
- Structured Reordering for Modeling Latent Alignments in Sequence TransductionBailin Wang, Mirella Lapata, Ivan TitovNeurIPS 2021 · 20 citations
