SirLLM: Streaming Infinite Retentive LLM
Yao Yao, Zuchao Li, Hai Zhao
摘要
As Large Language Models (LLMs) become increasingly prevalent in various domains, their ability to process inputs of any length and maintain a degree of memory becomes essential. However, the one-off input of overly long texts is limited, as studies have shown that when input lengths exceed the LLMs' pre-trained text length, there is a dramatic decline in text generation capabilities. Moreover, simply extending the length of pre-training texts is impractical due to the difficulty in obtaining long text data and the substantial memory consumption costs this would entail for LLMs. Recent efforts have employed streaming inputs to alleviate the pressure of excessively long text inputs, but this approach can significantly impair the model's long-term memory capabilities. Motivated by this challenge, we introduce Streaming Infinite Retentive LLM (SirLLM), which allows LLMs to maintain longer memory during infinite-length dialogues without the need for fine-tuning. SirLLM utilizes the Token Entropy metric and a memory decay mechanism to filter key phrases, endowing LLMs with both long-lasting and flexible memory. We designed three distinct tasks and constructed three datasets to measure the effectiveness of SirLLM from various angles: (1) DailyDialog; (2) Grocery Shopping; (3) Rock-Paper-Scissors. Our experimental results robustly demonstrate that SirLLM can achieve stable and significant improvements across different LLMs and tasks, compellingly proving its effectiveness. When having a coversation, "A sir could forget himself," but SirLLM never does! 043 2024). These applications, aiming to enhance user 044 interaction and conversational experience, often re-045 quire infinite input length and a certain degree of 046 memory capability. However, current LLMs are 047 usually pre-trained on texts of limited length, and 048 studies have shown that their text generation ca-049 pabilities dramatically decline when input lengths 050 exceed those of the pre-training texts (Xiao et al., 051 2023; Huang et al., 2023). Merely extending the 052 length of pre-training texts is impractical, as acquir-053 ing infinitely long text data is exceedingly challeng-054 ing, not to mention that it would result in substan-055 tial memory consumption for LLMs. Therefore, 056 researching how to enable LLMs to handle infinite 057 input lengths while maintaining memory capability 058 is an urgent issue to be addressed. 059 With the emergence of this demand, researchers 060 have gradually shifted their focus towards explor-061 ing ways to expand the input context length of 062 LLMs. A line of these studies has particularly fo-063 cused on optimizing the attention mechanism of 064 LLMs. (Beltagy et al., 2020) first proposes the 065 Sliding-window attention, as shown in Figure 1 (a). 066
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- SocialMind: LLM-based Proactive AR Social Assistive System with Human-like Perception for In-situ Live InteractionsBufang Yang, Yunqi Guo, Lilin Xu, Zhenyu Yan 等UbiComp 2025 · 被引用 26 次
- SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep LayersZicong Tang, Luohe Shi, Zuchao Li, Baoyuan Qi 等ACL 2025 · 被引用 7 次
- Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMsWanyun Cui, Mingwei XuNeurIPS 2025 · 被引用 7 次
- KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional EmbeddingLuohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi 等ACL 2025 · 被引用 5 次
- Towards Sampling Data Structures for Tensor Products in Turnstile StreamsZhao Song, Shenghao Xie, Samson ZhouICLR 2026 · 被引用 1 次
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMsSuyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang 等ICLR 2024 · 被引用 432 次
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai 等ICLR 2024 · 被引用 254 次
- CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator NeedsMajeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Zachary Henley 等CHI 2024 · 被引用 246 次
- Compressing Context to Enhance Inference Efficiency of Large Language ModelsYucheng Li, Bo Dong, Frank Guerin, Chenghua LinEMNLP 2023 · 被引用 54 次
相关 Paper
- InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context MemoryChaojun Xiao, Pengle Zhang, Xu Han, Guangxuan Xiao 等NeurIPS 2024 · 被引用 223 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- StreamingDialogue: Prolonged Dialogue Learning via Long Context Compression with Minimal LossesJianan Li, Quan Tu, Cunli Mao, Zhengtao Yu 等NeurIPS 2024 · 被引用 13 次
- ReAttention: Training-Free Infinite Context with Finite Attention ScopeXiaoran Liu, Ruixiao Li, Zhigeng Liu, Qipeng Guo 等ICLR 2025
- LLM Maybe LongLM: SelfExtend LLM Context Window Without TuningHongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang 等ICML 2024 · 被引用 167 次
