CogLTX: Applying BERT to Long Texts
Ming Ding, Chang Zhou, Hongxia Yang, Jie Tang
摘要
BERT is incapable of processing long texts due to its quadratically increasing memory and time consumption. The most natural ways to address this problem, such as slicing the text by a sliding window or simplifying transformers, suffer from insufficient long-range attentions or need customized CUDA kernels. The maximum length limit in BERT reminds us the limited capacity (5∼ 9 chunks) of the working memory of humans --then how do human beings Cognize Long TeXts? Founded on the cognitive theory stemming from Baddeley [2], the proposed CogLTX 1 framework identifies key sentences by training a judge model, concatenates them for reasoning, and enables multi-step reasoning via rehearsal and decay. Since relevance annotations are usually unavailable, we propose to use interventions to create supervision. As a general algorithm, CogLTX outperforms or gets comparable results to SOTA models on various downstream tasks with memory overheads independent of the length of text. BERT (reasoner) [CLS] Q yes no [SEP] z Start/End Span Q: Who is the director of the 2003 film which has scenes in it filmed at the Quality Cafe in Los Angeles? Long text x: The Quality Cafe (aka. Quality Diner) is a now-defunct diner … as a location featured in a number of Hollywood films, including "Training Day", "Old School"… Old School is a 2003 American comedy film released by DreamWorks and directed by Todd Phillips MemRecall ([Q], x) BERT (reasoner) [CLS] [SEP] z MLP Long text x: LOS ANGELES --The pilot flying Kobe Bryant and seven others to a youth basketball tournament did not have alcohol or drugs in his system, and all nine sustained immediately fatal injuries when their helicopter slammed into a hillside outside Los Angeles in January, according to autopsies released Friday. … MemRecall ([], x) 0.6 0.3 … Probabilty for each class BERT (reasoner) [CLS] x[i] [SEP] z Long text x MemRecall ([x[i]], x) NN IN DT Token-wise result of x[i] … x[0]:Confidence in the pound is widely expected to take another sharp dive if trade figures for September, due for release tomorrow, x[1]:fail to show a substantial improvement from July and August's near-record deficits" is considered, … Decompose into sub-sequences x[0]… x[n]
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Applying Contrastive Learning to Code Vulnerability Type ClassificationChen Ji, Su Yang, Hongyu Sun, Yuqing ZhangEMNLP 2024 · 被引用 6 次
- A Simple Contrastive Learning Framework for Interactive Argument Pair Identification via Argument-Context ExtractionLida Shi, Fausto Giunchiglia, Rui Song, Daqian Shi 等EMNLP 2022 · 被引用 4 次
- IRIS: Interpretable Retrieval-Augmented Classification for Long Interspersed Document SequencesFengnan Li, Elliot D. Hill, Jiang Shu, Jiaxin Gao 等ACL 2025
- Sok: Private Transformer-based Model InferenceYuntian Chen, Tianpei Lu, Zhanyong Tang, Bingsheng Zhang 等USENIX Security 2026
- R2F: A General Retrieval, Reading and Fusion Framework for Document-level Natural Language InferenceHao Wang, Yixin Cao, Yangguang Li, Zhen Huang 等EMNLP 2022
它引用的顶会 Paper4
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher 等ICLR 2020 · 被引用 322 次
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai 等EMNLP 2020 · 被引用 157 次
相关 Paper
- Recurrent Chunking Mechanisms for Long-Text Machine Reading ComprehensionHongyu Gong, Yelong Shen, Dian Yu, Jianshu Chen 等ACL 2020 · 被引用 39 次
- Let's (not) just put things in Context: Test-time Training for Long-context LLMsRachit Bansal, Aston Zhang, Rishabh Tiwari, Lovish Madaan 等ICLR 2026 · 被引用 20 次
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du 等ICLR 2026 · 被引用 225 次
- Through the Valley: Path to Effective Long CoT Training for Small Language ModelsRenjie Luo, Jiaxi Li, Chen Huang, Wei LuEMNLP 2025
- Lost in Transmission: When and Why LLMs Fail to Reason GloballyTobias Schnabel, Kiran Tomlinson, Adith Swaminathan, Jennifer NevilleNeurIPS 2025 · 被引用 9 次
