BTR: Binary Token Representations for Efficient Retrieval Augmented Language Models
Qingqing Cao, Sewon Min, Yizhong Wang, Hannaneh Hajishirzi
Abstract
Retrieval augmentation addresses many critical problems in large language models such as hallucination, staleness, and privacy leaks. However, running retrievalaugmented language models (LMs) is slow and difficult to scale due to processing large amounts of retrieved text. We introduce binary token representations (BTR), which use 1-bit vectors to precompute every token in passages, significantly reducing computation during inference. Despite the potential loss of accuracy, our new calibration techniques and training objectives restore performance. Combined with offline and runtime compression, this only requires 127GB of disk space for encoding 3 billion tokens in Wikipedia. Our experiments show that on five knowledge-intensive NLP tasks, BTR accelerates state-of-the-art retrievalaugmented language model inference by up to 4x and reduces storage by over 100x while maintaining over 95% task performance. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Scaling Retrieval-Based Language Models with a Trillion-Token DatastoreRulin Shao, Jacqueline He, Akari Asai, Weijia Shi et al.NeurIPS 2024 · 76 citations
- Lifting Optimized Binaries to Canonical Compiler IR via Structure-Aware Retrieval and Iterative VerificationXiaoao Zhu, Jie Ren, Zhiqiang Li, Jie Zheng et al.ACL 2026
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
Related papers
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- InstructRetro: Instruction Tuning post Retrieval-Augmented PretrainingBoxin Wang, Wei Ping, Lawrence McAfee, Peng Xu et al.ICML 2024 · 75 citations
- Beyond Attention: Breaking the Limits of Transformer Context Length with Recurrent MemoryAydar Bulatov, Yuri Kuratov, Yermek Kapushev, Mikhail BurtsevAAAI 2024 · 22 citations
- Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive StudyBoxin Wang, Wei Ping, Peng Xu, Lawrence McAfee et al.EMNLP 2023 · 19 citations
- Efficient Document Re-Ranking for Transformers by Precomputing Term RepresentationsSean MacAvaney, Franco Maria Nardini, Raffaele Perego, Nicola Tonellotto et al.SIGIR 2020 · 62 citations
