Gated Differentiable Working Memory for Long-Context Language Modeling
Lingrui Mei, Shenghua Liu, Yiwei Wang, Yuyao Ge, Baolong Bi, Jiayu Yao, Jun Wan, Ziling Yin, Jiafeng Guo, Xueqi Cheng
Abstract
Long contexts break transformers: attention scores dilute across thousands of tokens, critical information gets lost in the middle, and the model cannot adapt to novel patterns at inference time. Recent work on test-time adaptation addresses this by maintaining a form of working memory-transient parameters updated on the current context-but existing approaches employ uniform write policies that waste computation on low-value regions and suffer from high gradient variance across semantically heterogeneous contexts. In this work, we reframe test-time adaptation as a budget-constrained memory consolidation problem, asking: given limited computational budget, which parts of the context should be consolidated into working memory? We propose GDWM (Gated Differentiable Working Memory), a framework that introduces a Write Controller to gate the memory consolidation process. Our controller estimates Contextual Utility-an informationtheoretic measure quantifying how much each region depends on long-range context-and allocates gradient steps accordingly, subject to a coverage constraint that ensures global representation. Experiments on ZeroSCROLLS and LongBench v2 benchmarks demonstrate that GDWM achieves comparable or superior performance with 4× fewer gradient steps compared to uniform baselines, establishing a new efficiency-performance Pareto frontier for testtime adaptation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on49
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
Related papers
- Let's (not) just put things in Context: Test-time Training for Long-context LLMsRachit Bansal, Aston Zhang, Rishabh Tiwari, Lovish Madaan et al.ICLR 2026 · 20 citations
- GradMem: Learning to Write Context into Memory with Test-Time Gradient DescentYuri Kuratov, Matvey Kairov, Aydar Bulatov, Ivan Rodkin et al.ICML 2026 · 3 citations
- Memory efficiency and resource-rational encoding in sentence processingWeijie Xu, Brian Dillon, Richard FutrellACL 2026
- Octopus: Gated Selective Attention for Memory-Bounded Long-Context Inference in Large Language ModelsChien Van Nguyen, Ryan A. Rossi, Linh Ngo Van, Franck Dernoncourt et al.ACL 2026
- Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised ReasoningAdnan Oomerjee, Zafeirios Fountas, Haitham Bou-Ammar, Jun WangICLR 2026 · 4 citations
