Disentangled Sequence to Sequence Learning for Compositional Generalization
Hao Zheng, Mirella Lapata
摘要
There is mounting evidence that existing neural network models, in particular the very popular sequence-to-sequence architecture, struggle to systematically generalize to unseen compositions of seen components. We demonstrate that one of the reasons hindering compositional generalization relates to representations being entangled. We propose an extension to sequence-to-sequence models which encourages disentanglement by adaptively reencoding (at each time step) the source input. Specifically, we condition the source representations on the newly decoded target context which makes it easier for the encoder to exploit specialized information for each prediction rather than capturing it all in a single forward pass. Experimental results on semantic parsing and machine translation empirically show that our proposal delivers more disentangled representations and better generalization. 1 1 Our code is available at https://github.com/ mswellhao/Dangle . Training Set A boy ate the cake on the table in a house. *cake(x4); *table(x7); boy(x1) AND eat.agent(x2, x1) AND eat.theme(x2, x4) AND cake.nmod.on(x4, x7) AND table.nmod.in(x7, x10) AND house(x10) Test Set (Lexical Generalization) A boy likes the cake on the table in a house. *cake(x4); *table(x7); boy(x1) AND like.agent(x2, x1) AND like.theme(x2, x4) AND cake.nmod.on(x4, x7) AND table.nmod.in(x7, x10) AND house(x10) Test Set (Structural Generalization) A boy ate the cake on the table in a house beside the tree. *cake(x4); *table(x7); *tree(x13); boy(x1) AND eat.agent(x2, x1) AND eat.theme(x2, x4) AND cake.nmod.on(x4, x7) AND table.nmod.in(x7, x10) AND house(x10) AND house.nmod.beside(x10, x13)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Compositional Generalization for Multi-Label Text Classification: A Data-Augmentation ApproachYuyang Chai, Zhuang Li, Jiahui Liu, Lei Chen 等AAAI 2024 · 被引用 18 次
- Tripod: Three Complementary Inductive Biases for Disentangled Representation LearningKyle Hsu, Jubayer Ibn Hamid, Kaylee Burns, Chelsea Finn 等ICML 2024 · 被引用 13 次
- Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureLinyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu 等EMNLP 2023 · 被引用 12 次
- Structural generalization is hard for sequence-to-sequence modelsYuekun Yao, Alexander KollerEMNLP 2022 · 被引用 8 次
- Sparse Universal TransformerShawn Tan, Yikang Shen, Zhenfang Chen, Aaron C. Courville 等EMNLP 2023 · 被引用 6 次
它引用的顶会 Paper12
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain 等NeurIPS 2020 · 被引用 503 次
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman 等ICLR 2020 · 被引用 401 次
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 被引用 149 次
- Improving Disentangled Text Representation Learning with Information-Theoretic GuidancePengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon 等ACL 2020 · 被引用 66 次
相关 Paper
- Compositional Generalization without Trees using Multiset Tagging and Latent PermutationsMatthias Lindemann, Alexander Koller, Ivan TitovACL 2023
- Consistency Regularization Training for Compositional GeneralizationYongjing Yin, Jiali Zeng, Yafu Li, Fandong Meng 等ACL 2023 · 被引用 5 次
- Non-autoregressive Machine Translation with Disentangled Context TransformerJungo Kasai, James Cross, Marjan Ghazvininejad, Jiatao GuICML 2020 · 被引用 113 次
- Layer-Wise Representation Fusion for Compositional GeneralizationYafang Zheng, Lei Lin, Shuangtao Li, Yuxuan Yuan 等AAAI 2024 · 被引用 4 次
- Structured Reordering for Modeling Latent Alignments in Sequence TransductionBailin Wang, Mirella Lapata, Ivan TitovNeurIPS 2021 · 被引用 20 次
