Longtriever: a Pre-trained Long Text Encoder for Dense Document Retrieval
Junhan Yang, Zheng Liu, Chaozhuo Li, Guangzhong Sun, Xing Xie
摘要
Pre-trained language models (PLMs) have achieved the preeminent position in dense retrieval due to their powerful capacity in modeling intrinsic semantics. However, most existing PLM-based retrieval models encounter substantial computational costs and are infeasible for processing long documents. In this paper, a novel retrieval model Longtriever is proposed to embrace three core challenges of long document retrieval: substantial computational cost, incomprehensive document understanding, and scarce annotations. Longtriever splits long documents into short blocks and then efficiently models the local semantics within a block and the global context semantics across blocks in a tightly-coupled manner. A pretraining phase is further proposed to empower Longtriever to achieve a better understanding of underlying semantic correlations. Experimental results on two popular benchmark datasets demonstrate the superiority of our proposal. The source code is released at https: //github.com/SamuelYang1/Longtriever .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SEAL: Structure and Element Aware Learning Improves Long Structured Document RetrievalXinhao Huang, Zhibo Ren, Yipeng Yu, Ying Zhou 等EMNLP 2025
- Contrastive Learning on LLM Back Generation Treebank for Cross-domain Constituency ParsingPeiming Guo, Meishan Zhang, Jianling Li, Min Zhang 等ACL 2025
它引用的顶会 Paper16
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
相关 Paper
- LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query InferenceGuangyuan Ma, Yongliang Ma, Xuanrui Gou, Zhenpeng Su 等ICLR 2026 · 被引用 3 次
- Revela: Dense Retriever Learning via Language ModelingFengyu Cai, Tong Chen, Xinran Zhao, Sihao Chen 等ICLR 2026 · 被引用 3 次
- SAILER: Structure-aware Pre-trained Language Model for Legal Case RetrievalHaitao Li, Qingyao Ai, Jia Chen, Qian Dong 等SIGIR 2023 · 被引用 68 次
- In-Context Pretraining: Language Modeling Beyond Document BoundariesWeijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou 等ICLR 2024 · 被引用 87 次
- Efficient Long Context Language Model Retrieval with CompressionMinju Seo, Jinheon Baek, Seongyun Lee, Sung Ju HwangACL 2025 · 被引用 2 次
