CogLTX: Applying BERT to Long Texts
Ming Ding, Chang Zhou, Hongxia Yang, Jie Tang
Abstract
BERT is incapable of processing long texts due to its quadratically increasing memory and time consumption. The most natural ways to address this problem, such as slicing the text by a sliding window or simplifying transformers, suffer from insufficient long-range attentions or need customized CUDA kernels. The maximum length limit in BERT reminds us the limited capacity (5∼ 9 chunks) of the working memory of humans --then how do human beings Cognize Long TeXts? Founded on the cognitive theory stemming from Baddeley [2], the proposed CogLTX 1 framework identifies key sentences by training a judge model, concatenates them for reasoning, and enables multi-step reasoning via rehearsal and decay. Since relevance annotations are usually unavailable, we propose to use interventions to create supervision. As a general algorithm, CogLTX outperforms or gets comparable results to SOTA models on various downstream tasks with memory overheads independent of the length of text. BERT (reasoner) [CLS] Q yes no [SEP] z Start/End Span Q: Who is the director of the 2003 film which has scenes in it filmed at the Quality Cafe in Los Angeles? Long text x: The Quality Cafe (aka. Quality Diner) is a now-defunct diner … as a location featured in a number of Hollywood films, including "Training Day", "Old School"… Old School is a 2003 American comedy film released by DreamWorks and directed by Todd Phillips MemRecall ([Q], x) BERT (reasoner) [CLS] [SEP] z MLP Long text x: LOS ANGELES --The pilot flying Kobe Bryant and seven others to a youth basketball tournament did not have alcohol or drugs in his system, and all nine sustained immediately fatal injuries when their helicopter slammed into a hillside outside Los Angeles in January, according to autopsies released Friday. … MemRecall ([], x) 0.6 0.3 … Probabilty for each class BERT (reasoner) [CLS] x[i] [SEP] z Long text x MemRecall ([x[i]], x) NN IN DT Token-wise result of x[i] … x[0]:Confidence in the pound is widely expected to take another sharp dive if trade figures for September, due for release tomorrow, x[1]:fail to show a substantial improvement from July and August's near-record deficits" is considered, … Decompose into sub-sequences x[0]… x[n]
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Applying Contrastive Learning to Code Vulnerability Type ClassificationChen Ji, Su Yang, Hongyu Sun, Yuqing ZhangEMNLP 2024 · 6 citations
- A Simple Contrastive Learning Framework for Interactive Argument Pair Identification via Argument-Context ExtractionLida Shi, Fausto Giunchiglia, Rui Song, Daqian Shi et al.EMNLP 2022 · 4 citations
- IRIS: Interpretable Retrieval-Augmented Classification for Long Interspersed Document SequencesFengnan Li, Elliot D. Hill, Jiang Shu, Jiaxin Gao et al.ACL 2025
- Sok: Private Transformer-based Model InferenceYuntian Chen, Tianpei Lu, Zhanyong Tang, Bingsheng Zhang et al.USENIX Security 2026
- R2F: A General Retrieval, Reading and Fusion Framework for Document-level Natural Language InferenceHao Wang, Yixin Cao, Yangguang Li, Zhen Huang et al.EMNLP 2022
Builds on4
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher et al.ICLR 2020 · 322 citations
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai et al.EMNLP 2020 · 157 citations
Related papers
- Recurrent Chunking Mechanisms for Long-Text Machine Reading ComprehensionHongyu Gong, Yelong Shen, Dian Yu, Jianshu Chen et al.ACL 2020 · 39 citations
- Let's (not) just put things in Context: Test-time Training for Long-context LLMsRachit Bansal, Aston Zhang, Rishabh Tiwari, Lovish Madaan et al.ICLR 2026 · 20 citations
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du et al.ICLR 2026 · 225 citations
- Through the Valley: Path to Effective Long CoT Training for Small Language ModelsRenjie Luo, Jiaxi Li, Chen Huang, Wei LuEMNLP 2025
- Lost in Transmission: When and Why LLMs Fail to Reason GloballyTobias Schnabel, Kiran Tomlinson, Adith Swaminathan, Jennifer NevilleNeurIPS 2025 · 9 citations
