Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
Jiaqi Cao, Jiarui Wang, Rubin Wei, Qipeng Guo, Kai Chen, Bowen Zhou, Zhouhan Lin
Abstract
Large Language Models (LLMs) have shown strong abilities in general language tasks, yet adapting them to specific domains remains a challenge. Current method like Domain Adaptive Pretraining (DAPT) requires costly full-parameter training and suffers from catastrophic forgetting. Meanwhile, Retrieval-Augmented Generation (RAG) introduces substantial inference latency due to expensive nearest-neighbor searches and longer context. This paper introduces Memory Decoder, a plug-and-play pretrained memory that enables efficient domain adaptation without changing the original model's parameters. Memory Decoder employs a small transformer decoder that learns to imitate the behavior of an external non-parametric retriever. Once trained, Memory Decoder can be seamlessly integrated with any pretrained language model that shares the same tokenizer, requiring no model-specific modifications. Experimental results demonstrate that Memory Decoder enables effective adaptation of various Qwen and Llama models to three distinct specialized domains: biomedicine, finance, and law, reducing perplexity by an average of 6.17 points. Overall, Memory Decoder introduces a novel paradigm centered on a specially pretrained memory component designed for domain-specific adaptation. This memory architecture can be integrated in a plug-and-play manner, consistently enhancing performance across multiple models within the target domain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e564807-6745-4dbc-b047-9b6eef7a6cebCited by top-tier papers4
- Dual Latent Memory for Visual Multi-agent SystemXinlei Yu, Chengming Xu, Zhangquan Chen, Bo Yin et al.ICML 2026 · 5 citations
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAMHaoyu Huang, Hong Ting Tsang, Jiaxin Bai, Xi Peng et al.ICLR 2026 · 4 citations
- MEMTS: Internalizing Domain Knowledge via Parameterized Memory for Retrieval-Free Domain Adaptation of Time Series Foundation ModelsXiaoyun Yu, Li Fan, Xiangfei Qiu, Nanqing Dong et al.KDD 2026 · 2 citations
- Controlled Dynamics Attractor TransformerCheng Zhang, Minnan Luo, Zesheng Yang, Ming Li et al.ICML 2026
Builds on8
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
- Adapting Large Language Models via Reading ComprehensionDaixuan Cheng, Shaohan Huang, Furu WeiICLR 2024 · 146 citations
- Adapting Language Models to Compress ContextsAlexis Chevalier, Alexander Wettig, Anirudh Ajith, Danqi ChenEMNLP 2023 · 34 citations
- Why do Nearest Neighbor Language Models Work?Frank F. Xu, Uri Alon, Graham NeubigICML 2023 · 33 citations
- Nearest Neighbor Zero-Shot InferenceWeijia Shi, Julian Michael, Suchin Gururangan, Luke ZettlemoyerEMNLP 2022 · 26 citations
Related papers
- MLP Memory: A Retriever-Pretrained Memory for Large Language ModelsRubin Wei, Jiaqi Cao, Jiarui Wang, Jushi Kai et al.ICLR 2026 · 16 citations
- G-MAP: General Memory-Augmented Pre-trained Language Model for Domain TasksZhongwei Wan, Yichun Yin, Wei Zhang, Jiaxin Shi et al.EMNLP 2022 · 2 citations
- RAPID: Long-Context Inference with Retrieval-Augmented Speculative DecodingGuanzheng Chen, Qilong Feng, Jinjie Ni, Xin Li et al.ICML 2025
- Decoupling Knowledge and Context: An Efficient and Effective Retrieval Augmented Generation Framework via Cross AttentionQian Dong, Qingyao Ai, Hongning Wang, Yiding Liu et al.WWW 2025 · 19 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
