Longtriever: a Pre-trained Long Text Encoder for Dense Document Retrieval
Junhan Yang, Zheng Liu, Chaozhuo Li, Guangzhong Sun, Xing Xie
Abstract
Pre-trained language models (PLMs) have achieved the preeminent position in dense retrieval due to their powerful capacity in modeling intrinsic semantics. However, most existing PLM-based retrieval models encounter substantial computational costs and are infeasible for processing long documents. In this paper, a novel retrieval model Longtriever is proposed to embrace three core challenges of long document retrieval: substantial computational cost, incomprehensive document understanding, and scarce annotations. Longtriever splits long documents into short blocks and then efficiently models the local semantics within a block and the global context semantics across blocks in a tightly-coupled manner. A pretraining phase is further proposed to empower Longtriever to achieve a better understanding of underlying semantic correlations. Experimental results on two popular benchmark datasets demonstrate the superiority of our proposal. The source code is released at https: //github.com/SamuelYang1/Longtriever .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7fc7a546-76f1-4346-b7fa-39b06cd22bb3Cited by top-tier papers2
- SEAL: Structure and Element Aware Learning Improves Long Structured Document RetrievalXinhao Huang, Zhibo Ren, Yipeng Yu, Ying Zhou et al.EMNLP 2025
- Contrastive Learning on LLM Back Generation Treebank for Cross-domain Constituency ParsingPeiming Guo, Meishan Zhang, Jianling Li, Min Zhang et al.ACL 2025
Builds on16
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
Related papers
- LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query InferenceGuangyuan Ma, Yongliang Ma, Xuanrui Gou, Zhenpeng Su et al.ICLR 2026 · 3 citations
- Revela: Dense Retriever Learning via Language ModelingFengyu Cai, Tong Chen, Xinran Zhao, Sihao Chen et al.ICLR 2026 · 3 citations
- SAILER: Structure-aware Pre-trained Language Model for Legal Case RetrievalHaitao Li, Qingyao Ai, Jia Chen, Qian Dong et al.SIGIR 2023 · 68 citations
- In-Context Pretraining: Language Modeling Beyond Document BoundariesWeijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou et al.ICLR 2024 · 87 citations
- Efficient Long Context Language Model Retrieval with CompressionMinju Seo, Jinheon Baek, Seongyun Lee, Sung Ju HwangACL 2025 · 2 citations
