Syntax-Enhanced Pre-trained Model
Zenan Xu, Daya Guo, Duyu Tang, Qinliang Su, Linjun Shou, Ming Gong, Wanjun Zhong, Xiaojun Quan, Daxin Jiang, Nan Duan
摘要
We study the problem of leveraging the syntactic structure of text to enhance pre-trained models such as BERT and RoBERTa. Existing methods utilize syntax of text either in the pre-training stage or in the fine-tuning stage, so that they suffer from discrepancy between the two stages. Such a problem would lead to the necessity of having human-annotated syntactic information, which limits the application of existing methods to broader scenarios. To address this, we present a model that utilizes the syntax of text in both pre-training and fine-tuning stages. Our model is based on Transformer with a syntax-aware attention layer that considers the dependency tree of the text. We further introduce a new pre-training task of predicting the syntactic distance among tokens in the dependency tree. We evaluate the model on three downstream tasks, including relation classification, entity typing, and question answering. Results show that our model achieves state-of-the-art performance on six public benchmark datasets. We have two major findings. First, we demonstrate that infusing automatically produced syntax of text improves pre-trained models. Second, global syntactic distances among tokens bring larger performance gains compared to local head relations between contiguous tokens. 1 * Work is done during internship at Microsoft. † For questions, please contact D. Tang and Z. Xu.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- DKPLM: Decomposable Knowledge-Enhanced Pre-trained Language Model for Natural Language UnderstandingTaolin Zhang, Chengyu Wang, Nan Hu, Minghui Qiu 等AAAI 2022 · 被引用 36 次
- Rethinking Positional Encoding in Tree Transformer for Code RepresentationHan Peng, Ge Li, Yunfei Zhao, Zhi JinEMNLP 2022 · 被引用 10 次
- Do Transformers Parse while Predicting the Masked Word?Haoyu Zhao, Abhishek Panigrahi, Rong Ge, Sanjeev AroraEMNLP 2023 · 被引用 5 次
- Language Model Pre-training on True NegativesZhuosheng Zhang, Hai Zhao, Masao Utiyama, Eiichiro SumitaAAAI 2023 · 被引用 3 次
- Dependency-based Mixture Language ModelsZhixian Yang, Xiaojun WanACL 2022 · 被引用 3 次
它引用的顶会 Paper3
- Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language ModelWenhan Xiong, Jingfei Du, William Yang Wang, Veselin StoyanovICLR 2020 · 被引用 215 次
- SG-Net: Syntax-Guided Machine Reading ComprehensionZhuosheng Zhang, Yuwei Wu, Junru Zhou, Sufeng Duan 等AAAI 2020 · 被引用 192 次
- Tree-Structured Attention with Hierarchical AccumulationXuan-Phi Nguyen, Shafiq R. Joty, Steven C. H. Hoi, Richard SocherICLR 2020 · 被引用 79 次
相关 Paper
- Syntax-augmented Multilingual BERT for Cross-lingual TransferWasi Uddin Ahmad, Haoran Li, Kai-Wei Chang, Yashar MehdadACL 2021
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language UnderstandingWei Wang, Bin Bi, Ming Yan, Chen Wu 等ICLR 2020 · 被引用 297 次
- BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant SupervisionChen Liang, Yue Yu, Haoming Jiang, Siawpeng Er 等KDD 2020 · 被引用 118 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- How much pretraining data do language models need to learn syntax?Laura Pérez-Mayos, Miguel Ballesteros, Leo WannerEMNLP 2021 · 被引用 31 次
