Landmark Embedding: A Chunking-Free Embedding Method For Retrieval Augmented Long-Context Large Language Models
Kun Luo, Zheng Liu, Shitao Xiao, Tong Zhou, Yubo Chen, Jun Zhao, Kang Liu
Abstract
Retrieval augmentation is a promising approach to handle long-context language modeling. However, the existing retrieval methods usually work with the chunked context, which is prone to inferior quality of semantic representation and incomplete retrieval of useful information. In this work, we propose a new method for the retrieval augmentation of longcontext language modeling, called Landmark Embedding. Our method is characterized by threefold technical contributions. Firstly, we introduce a chunking-free architecture, which keeps the long context coherent such that highquality embeddings can be generated for the fine-grained units within the context. Secondly, we present a position-aware objective function, which prioritizes the ultimate boundary for a consecutive span of information. By learning to discriminate such a special position, the useful information can be comprehensively retrieved for the query. Thirdly, we design a novel multistage learning algorithm, which makes the best use of readily available data and synthetic data for cost-effective training of the landmark embedding. In our experimental study, landmark embedding is able to substantially improve the performance for both LLaMA-2 and ChatGPT in a variety of long-context tasks; meanwhile, it also outperforms the existing retrieval methods with a notable advantage. Our model and code will be made publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0588c0da-1a31-49be-b9a7-dd7ed82c640dCited by top-tier papers9
- Retrieval as Generation: A Unified Framework with Self-Triggered Information PlanningBo Li, Mingda Wang, Gexiang Fang, Shikun Zhang et al.ACL 2026 · 8 citations
- SmartCache: Context-aware Semantic Cache for Efficient Multi-turn LLM InferenceChengye Yu, Tianyu Wang, Zili Shao, Song JiangNeurIPS 2025 · 6 citations
- Modeling Uncertainty Trends for Timely Retrieval in Dynamic RAGBo Li, Tian Tian, Zhenghua Xu, Hao Cheng et al.AAAI 2026 · 5 citations
- Reinforcing Agentic Search Via Reward Density OptimizationKun Luo, Hongjin Qian, Zheng Liu, Ziyi Xia et al.ACL 2026
- Token-Free Hierarchical Indexing for RAG beyond LLM-based SummarizationYifan Wei, Dan Yuan, Xiaoyan Yu, Angsheng LiICML 2026
Builds on15
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
Related papers
- Retrieval meets Long Context Large Language ModelsPeng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee et al.ICLR 2024 · 131 citations
- Random-Access Infinite Context Length for TransformersAmirkeivan Mohtashami, Martin JaggiNeurIPS 2023 · 207 citations
- A Multi-Task Embedder For Retrieval Augmented LLMsPeitian Zhang, Zheng Liu, Shitao Xiao, Zhicheng Dou et al.ACL 2024
- Training with "Paraphrasing the Original Text" Teaches LLM to Better Retrieve in Long-Context TasksYijiong Yu, Yongfeng Huang, Zhixiao Qi, Zhe ZhouAAAI 2025 · 5 citations
- Training-Free Long-Context Scaling of Large Language ModelsChenxin An, Fei Huang, Jun Zhang, Shansan Gong et al.ICML 2024 · 68 citations
