LongSight: Compute-Enabled Memory to Accelerate Large-Context LLMs via Sparse Attention
Derrick Quinn, E. Ezgi Yücel, Jinkwon Kim, José F. Martínez, Mohammad Alian
Abstract
Large input context windows in transformer-based LLMs help minimize hallucinations and improve output accuracy and personalization. However, as the context window grows, the attention phase increasingly dominates execution time. Key-Value (KV) caching alleviates part of this cost by avoiding redundant computation, but the KV cache itself can quickly exceed the capacity of today's GPU high-bandwidth memory (HBM). In this work, we present LongSight, an algorithm-hardware co-design framework for accelerating attention in large-context scenarios. LongSight leverages a compute-enabled CXL memory device, originally designed for dense retrieval acceleration, to offload KV cache storage and retrieval. Therefore, LongSight effectively elevates the value of relatively low-cost LPDDR DRAM to that of high-end HBM. We demonstrate that, with just a single GPU and a single computeenabled CXL memory expander, LongSight can efficiently support context lengths of up to 1 million tokens for state-of-the-art Llama models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 93b937cd-71d4-4e42-b65a-bf7fd81b9644Cited by top-tier papers2
- Early Silicon of Raptor: The First 3D-DRAM Accelerator for Generative InferencePrashant J. Nair, Ramyad Hadidi, Subramani Ganesh, Sangamesh Kodge et al.ISCA 2026 · 4 citations
- DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory ArchitecturesPeiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati et al.ISCA 2026 · 3 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Efficient Low Rank Attention for Long-Context Inference in Large Language ModelsTenghui Li, Guoxu Zhou, Xuyang Zhao, Yuning Qiu et al.NeurIPS 2025 · 4 citations
- KVO-LLM: Boosting Long-Context Generation Throughput for Batched LLM InferenceZhenyu Li, Dongxu Lyu, Gang Wang, Yuzhou Chen et al.DAC 2025 · 1 citation
- No Buffer, No Bottleneck: Efficient Zero-Copy KV Cache Offloading for Long-Context LLMsShutian Luo, Haiying ShenOSDI 2026
- xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector ExtractionChi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Hung-Yueh Chiang et al.ICML 2026 · 3 citations
- InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM InferenceXiurui Pan, Endian Li, Qiao Li, Shengwen Liang et al.HPCA 2025 · 22 citations
