LLoCO: Learning Long Contexts Offline
Sijun Tan, Xiuyu Li, Shishir G. Patil, Ziyang Wu, Tianjun Zhang, Kurt Keutzer, Joseph Gonzalez, Raluca A. Popa
摘要
Processing long contexts remains a challenge for large language models (LLMs) due to the quadratic computational and memory overhead of the self-attention mechanism and the substantial KV cache sizes during generation. We propose LLoCO, a novel approach to address this problem by learning contexts offline through context compression and in-domain parameter-efficient finetuning with LoRA. Our method enables an LLM to create a concise representation of the original context and efficiently retrieve relevant information to answer questions accurately. Our approach extends the effective context window of a 4k token LLaMA2-7B model to handle up to 128k tokens. We evaluate our approach on several longcontext question-answering datasets, demonstrating that LLoCO significantly outperforms in-context learning while using 30× fewer tokens during inference. LLoCO achieves up to 7.62× speed-up during inference and 11.52× higher throughput during finetuning, substantially reduces the cost of long document question answering. This makes it a promising solution for efficient long context processing. 1 * Equal contribution 1 Our code is publicly available on https://github.com/ jeffreysijuntan/lloco .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering TasksHongyuan Tao, Ying Zhang, Zhenhao Tang, Hongen Peng 等NeurIPS 2025 · 被引用 41 次
- Read As Human: Compressing Context via Parallelizable Close Reading and SkimmingJiwei Tang, Shilei Liu, Zhicheng Zhang, Qingsong Lv 等ACL 2026 · 被引用 10 次
- CoMeT: Collaborative Memory Transformer for Efficient Long Context ModelingRunsong Zhao, Shilei Liu, Jiwei Tang, Langming Liu 等ACL 2026 · 被引用 7 次
- Task Cascades for Efficient Unstructured Data ProcessingShreya Shankar, Sepanta Zeighami, Aditya G. ParameswaranSIGMOD 2026 · 被引用 7 次
- Gated Differentiable Working Memory for Long-Context Language ModelingLingrui Mei, Shenghua Liu, Yiwei Wang, Yuyao Ge 等ACL 2026 · 被引用 4 次
它引用的顶会 Paper26
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
相关 Paper
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai 等ICLR 2024 · 被引用 254 次
- LoCoCo: Dropping In Convolutions for Long Context CompressionRuisi Cai, Yuandong Tian, Zhangyang Wang, Beidi ChenICML 2024 · 被引用 17 次
- Dodo: Dynamic Contextual Compression for Decoder-only LMsGuanghui Qin, Corby Rosset, Ethan C. Chau, Nikhil Rao 等ACL 2024
- FreqKV: Key-Value Compression in Frequency Domain for Context Window ExtensionJushi Kai, Yixuan Wang, Boyi Zeng, Haoli Bai 等ICLR 2026 · 被引用 7 次
- A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial ContextsSuyu Ge, Xihui Lin, Yunan Zhang, Jiawei Han 等ICLR 2025
