Efficient Long Context Language Model Retrieval with Compression
Minju Seo, Jinheon Baek, Seongyun Lee, Sung Ju Hwang
Abstract
Long Context Language Models (LCLMs) have emerged as a new paradigm to perform Information Retrieval (IR), which enables the direct ingestion and retrieval of information by processing an entire corpus in their single context, showcasing the potential to surpass traditional sparse and dense retrieval methods. However, processing a large number of passages within in-context for retrieval is computationally expensive, and handling their representations during inference further exacerbates the processing time. To address this challenge, we aim to make LCLM retrieval more efficient and potentially more effective with passage compression. Specifically, we propose a new compression approach tailored for LCLM retrieval, which is trained to maximize the retrieval performance while minimizing the length of the compressed passages. To accomplish this, we generate the synthetic data, where compressed passages are automatically created and labeled as chosen or rejected according to their retrieval success for a given query, and we then train the proposed Compression model for Long context Retrieval (CoLoR) with this data via preference optimization while adding the length regularization loss on top of it to enforce brevity. We perform extensive experiments on nine datasets, and showcase that CoLoR improves the retrieval performance by 6% while compressing the in-context size by a factor of 1.91. Our code is available at: https://github.com/going-doer/CoLoR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f2c5d4b-e7c1-4375-b650-e966bfceab50Cited by top-tier papers2
- Paper2Code: Automating Code Generation from Scientific Papers in Machine LearningMinju Seo, Jinheon Baek, Seongyun Lee, Sung Ju HwangICLR 2026 · 86 citations
- ACON: Optimizing Context Compression for Long-horizon LLM AgentsMinki Kang, Wei-Ning Chen, Dongge Han, Huseyin Inan et al.ICML 2026
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Retrieval meets Long Context Large Language ModelsPeng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee et al.ICLR 2024 · 131 citations
Related papers
- RECOMP: Improving Retrieval-Augmented LMs with Context Compression and Selective AugmentationFangyuan Xu, Weijia Shi, Eunsol ChoiICLR 2024 · 260 citations
- Less Is More: Elevating RAG via Performance-Driven Context CompressionZiqiang Cui, Yunpeng Weng, Xing Tang, Peiyang Liu et al.ICML 2026
- Pretraining Context Compressor for Large Language Models with Embedding-Based MemoryYuhong Dai, Jianxun Lian, Yitian Huang, Wei Zhang et al.ACL 2025
- Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language ModelsWenhan Liu, Xinyu Ma, Yutao Zhu, Ziliang Zhao et al.ACL 2025 · 10 citations
- CompAct: Compressing Retrieved Documents Actively for Question AnsweringChanwoong Yoon, Taewhoo Lee, Hyeon Hwang, Minbyul Jeong et al.EMNLP 2024 · 10 citations
