Chunk-Distilled Language Modeling
Yanhong Li, Karen Livescu, Jiawei Zhou
Abstract
We introduce Chunk-Distilled Language Modeling (CD-LM), an approach to text generation that addresses two challenges in current large language models (LLMs): the inefficiency of token-level generation, and the difficulty of adapting to new data and knowledge. Our method combines deep network-based LLMs with a straightforward retrieval module, which allows the generation of multitoken text chunks at a single decoding step. Our retrieval framework enables flexible construction of model-or domain-specific datastores, either leveraging the internal knowledge of existing models, or incorporating expert insights from human-annotated corpora. This adaptability allows for enhanced control over the language model's distribution without necessitating additional training. We present the CD-LM formulation along with performance metrics demonstrating its ability to improve language model performance and efficiency across a diverse set of downstream applications. 1 CD-LM requires no training and can work with any off-the-shelf language model in both chunk discovery and sequence generation. We conduct a diverse set of empirical studies, including language modeling perplexity, text generation, and domain adaptation, showing the ability of CD-LM to improve inference efficiency and modeling performance. BACKGROUND While many attempts have been made to improve language modeling and generation efficiency, it remains a significant challenge to address both simultaneously. For example, non-parametric approaches like kNN-LM (Khandelwal et al., 2020) reduce LM perplexity in certain domains, but tend to require a sizable database for retrieval and adds latency during generation; specialized inference algorithms like speculative decoding (Spector & Re, 2023) speed up generation but keep the LM's distribution fixed. Unlike prior work, CD-LM can both speed up generation and adapt the LM's distribution. We include a more comprehensive overview of related work in Appendix C.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Intent-Driven Network Management with Multi-Agent LLMs: The Confucius FrameworkZhaodong Wang, Samuel Lin, Guanqing Yan, Soudeh Ghorbani et al.SIGCOMM 2025 · 22 citations
- On the Predictive Power of Representation Dispersion in Language ModelsYanhong Li, Ming Li, Karen Livescu, Jiawei ZhouICLR 2026 · 7 citations
Builds on22
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
Related papers
- Chunk-based Nearest Neighbor Machine TranslationPedro Henrique Martins, Zita Marinho, André F. T. MartinsEMNLP 2022 · 17 citations
- Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language ModelsJiaqi Cao, Jiarui Wang, Rubin Wei, Qipeng Guo et al.NeurIPS 2025 · 14 citations
- TokAlign: Efficient Vocabulary Adaptation via Token AlignmentChong Li, Jiajun Zhang, Chengqing ZongACL 2025 · 7 citations
- Drag-and-Drop LLMs: Zero-Shot Prompt-to-WeightsZhiyuan Liang, Dongwen Tang, Yuhao Zhou, Xuanlei Zhao et al.NeurIPS 2025 · 22 citations
- Language Modelling via Learning to RankArvid Frydenlund, Gagandeep Singh, Frank RudziczAAAI 2022 · 9 citations
