ERNIE-Doc: A Retrospective Long-Document Modeling Transformer
Siyu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang
摘要
Transformers are not suited for processing long documents, due to their quadratically increasing memory and time consumption. Simply truncating a long document or applying the sparse attention mechanism will incur the context fragmentation problem or lead to an inferior modeling capability against comparable model sizes. In this paper, we propose ERNIE-DOC, a document-level language pretraining model based on Recurrence Transformers (Dai et al., 2019). Two welldesigned techniques, namely the retrospective feed mechanism and the enhanced recurrence mechanism, enable ERNIE-DOC 1 , which has a much longer effective context length, to capture the contextual information of a complete document. We pretrain ERNIE-DOC to explicitly learn the relationships among segments with an additional document-aware segment-reordering objective. Various experiments were conducted on both English and Chinese document-level tasks. ERNIE-DOC improved the state-of-the-art language modeling result of perplexity to 16.8 on WikiText-103. Moreover, it outperformed competitive pretraining models by a large margin on most language understanding tasks, such as text classification and question answering.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Recurrent Memory TransformerAydar Bulatov, Yuri Kuratov, Mikhail BurtsevNeurIPS 2022 · 被引用 252 次
- Museformer: Transformer with Fine- and Coarse-Grained Attention for Music GenerationBotao Yu, Peiling Lu, Rui Wang, Wei Hu 等NeurIPS 2022 · 被引用 104 次
- Beyond Attention: Breaking the Limits of Transformer Context Length with Recurrent MemoryAydar Bulatov, Yuri Kuratov, Yermek Kapushev, Mikhail BurtsevAAAI 2024 · 被引用 22 次
- PoNet: Pooling Network for Efficient Token Mixing in Long SequencesChao-Hong Tan, Qian Chen, Wen Wang, Qinglin Zhang 等ICLR 2022 · 被引用 15 次
- DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads FusionYilong Chen, Linhao Zhang, Junyuan Shang, Zhenyu Zhang 等NeurIPS 2024 · 被引用 12 次
它引用的顶会 Paper7
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingYu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng 等AAAI 2020 · 被引用 885 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
- Sparse Sinkhorn AttentionYi Tay, Dara Bahri, Liu Yang, Donald Metzler 等ICML 2020 · 被引用 391 次
相关 Paper
- Span Graph Transformer for Document-Level Named Entity RecognitionHongli Mao, Xian-Ling Mao, Hanlin Tang, Yuming Shang 等AAAI 2024 · 被引用 3 次
- In-Context Pretraining: Language Modeling Beyond Document BoundariesWeijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou 等ICLR 2024 · 被引用 87 次
- Modeling Document-Level Context for Event Detection via Important Context SelectionAmir Pouran Ben Veyseh, Minh Van Nguyen, Nghia Trung Ngo, Bonan Min 等EMNLP 2021 · 被引用 25 次
- UniDoc: Unified Pretraining Framework for Document UnderstandingJiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao 等NeurIPS 2021 · 被引用 118 次
- Recurrent Chunking Mechanisms for Long-Text Machine Reading ComprehensionHongyu Gong, Yelong Shen, Dian Yu, Jianshu Chen 等ACL 2020 · 被引用 39 次
