Augmenting Transformers with Recursively Composed Multi-grained Representations
Xiang Hu, Qingyang Zhu, Kewei Tu, Wei Wu
摘要
We present ReCAT, a recursive composition augmented Transformer that is able to explicitly model hierarchical syntactic structures of raw texts without relying on gold trees during both learning and inference. Existing research along this line restricts data to follow a hierarchical tree structure and thus lacks inter-span communications. To overcome the problem, we propose novel contextual inside-outside (CIO) layers, each of which consists of a top-down pass that forms representations of high-level spans by composing low-level spans, and a bottom-up pass that combines information inside and outside a span. The bottom-up and top-down passes are performed iteratively by stacking CIO layers to fully contextualize span representations. By inserting the stacked CIO layers between the embedding layer and the attention layers in Transformer, the ReCAT model can perform both deep intra-span and deep inter-span interactions, and thus generate multi-grained representations fully contextualized with other spans. Moreover, the CIO layers can be jointly pre-trained with Transformers, making ReCAT enjoy scaling ability, strong performance, and interpretability at the same time. We conduct experiments on various sentence-level and span-level tasks. Evaluation results indicate that Re-CAT can significantly outperform vanilla Transformer models on all span-level tasks and recursive models on natural language inference tasks. More interestingly, the hierarchical structures induced by ReCAT exhibit strong consistency with human-annotated syntactic trees, indicating good interpretability brought by the CIO layers. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at ScaleXiang Hu, Pengyu Ji, Qingyang Zhu, Wei Wu 等ACL 2024 · 被引用 1 次
- Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language ModelingXiang Hu, Zhihao Teng, Jun Zhao, Wei Wu 等ICML 2025
- On Eliciting Syntax from Language Models via HashingYiran Wang, Masao UtiyamaEMNLP 2024
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive AutoencodersAndrew Drozdov, Subendhu Rongali, Yi-Pei Chen, Tim O'Gorman 等EMNLP 2020 · 被引用 27 次
- Modeling Hierarchical Structures with Continuous Recursive Neural NetworksJishnu Ray Chowdhury, Cornelia CarageaICML 2021 · 被引用 18 次
- Beam Tree Recursive CellsJishnu Ray Chowdhury, Cornelia CarageaICML 2023 · 被引用 7 次
- Fast-R2D2: A Pretrained Recursive Neural Network based on Pruned CKY for Grammar Induction and Text RepresentationXiang Hu, Haitao Mi, Liang Li, Gerard de MeloEMNLP 2022 · 被引用 7 次
相关 Paper
- R2D2: Recursive Transformer based on Differentiable Tree for Interpretable Hierarchical Language ModelingXiang Hu, Haitao Mi, Zujie Wen, Yafang Wang 等ACL 2021
- A Multi-Grained Self-Interpretable Symbolic-Neural Model For Single/Multi-Labeled Text ClassificationXiang Hu, Xinyu Kong, Kewei TuICLR 2023 · 被引用 2 次
- Compositional Generalization without Trees using Multiset Tagging and Latent PermutationsMatthias Lindemann, Alexander Koller, Ivan TitovACL 2023
- Towards Equipping Transformer with the Ability of Systematic CompositionalityChen Huang, Peixin Qin, Wenqiang Lei, Jiancheng LvAAAI 2024 · 被引用 3 次
- Pushdown Layers: Encoding Recursive Structure in Transformer Language ModelsShikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. ManningEMNLP 2023
