Tree-Structured Attention with Hierarchical Accumulation
Xuan-Phi Nguyen, Shafiq R. Joty, Steven C. H. Hoi, Richard Socher
Abstract
Incorporating hierarchical structures like constituency trees has been shown to be effective for various natural language processing (NLP) tasks. However, it is evident that state-of-the-art (SOTA) sequence-based models like the Transformer struggle to encode such structures inherently. On the other hand, dedicated models like the Tree-LSTM, while explicitly modeling hierarchical structures, do not perform as efficiently as the Transformer. In this paper, we attempt to bridge this gap with Hierarchical Accumulation to encode parse tree structures into self-attention at constant time complexity. Our approach outperforms SOTA methods in four IWSLT translation tasks and the WMT'14 English-German task. It also yields improvements over Transformer and Tree-LSTM on three text classification tasks. We further demonstrate that using hierarchical priors can compensate for data shortage, and that our model prefers phrase-level attentions over token-level attentions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01c37bb8-e9a8-47bf-90ae-5dbe3aaecb71Cited by top-tier papers16
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 179 citations
- SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationXupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng et al.ASPLOS 2024 · 105 citations
- TUTA: Tree-based Transformers for Generally Structured Table Pre-trainingZhiruo Wang, Haoyu Dong, Ran Jia, Jia Li et al.KDD 2021 · 88 citations
- Vulnerability Detection with Graph Simplification and Enhanced Graph Representation LearningXin-Cheng Wen, Yupan Chen, Cuiyun Gao, Hongyu Zhang et al.ICSE 2023 · 77 citations
Related papers
- Recursive Tree-Structured Self-Attention for Answer Sentence SelectionKhalil Mrini, Emilia Farcas, Ndapa NakasholeACL 2021
- HiTIN: Hierarchy-aware Tree Isomorphism Network for Hierarchical Text ClassificationHe Zhu, Chong Zhang, Junjie Huang, Junran Wu et al.ACL 2023 · 17 citations
- How Transformers Represent Hierarchies: A Local-to-Global MechanismZhiling Zhou, Tianhao Wang, Zhuoran YangICML 2026
- Tree Transformer's Disambiguation Ability of Prepositional Phrase Attachment and Garden Path EffectsLingling Zhou, Suzan Verberne, Gijs WijnholdsACL 2024
- StrAE: Autoencoding for Pre-Trained Embeddings using Explicit StructureMattia Opper, Victor Prokhorov, Siddharth NarayanaswamyEMNLP 2023 · 1 citation
