Efficient Beam Search for Large Language Models Using Trie-Based Decoding
Brian J. Chan, Mao Xun Huang, Jui-Hung Cheng, Chao-Ting Chen, Hen-Hsen Huang
摘要
This work presents a novel trie (prefix-tree)based parallel decoding method that addresses the memory inefficiency of batch-based beam search. By sharing a single KV cache across beams with common prefixes, our approach dramatically reduces memory usage and enables efficient decoding. We evaluated our method across three attention architectures, Multi-Head Attention (Phi-3.5-miniinstruct), Grouped Query Attention (Llama-3.1-8B-Instruct), and Sliding Window Attention (Mistral-Small-24B-Instruct-2501), using CN-N/DailyMail for abstractive summarization and HumanEval for code generation. Our experiments demonstrate substantial memory savings (4-8×) and up to 2.4× faster decoding, without compromising generation quality. These results highlight our method's suitability for memory-constrained environments and largescale deployments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-CompletionRahul Mehta, Kavin R. V, Indrajit Pal, Tushar Abhishek 等SIGIR 2026
- BioCG: Constrained Generative Modeling for Biochemical Interaction PredictionAmitay Sicherman, Kira RadinskyNeurIPS 2025
它引用的顶会 Paper7
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng 等ICML 2024 · 被引用 669 次
- SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationXupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng 等ASPLOS 2024 · 被引用 105 次
相关 Paper
- CoDec: Prefix-Shared Decoding Kernel for LLMsZhibin Wang, Rui Ning, Chao Fang, Zhonghui Zhang 等SIGMOD 2026 · 被引用 8 次
- DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM InferenceJinwei Yao, Kaiqi Chen, Kexun Zhang, Jiaxuan You 等ICLR 2025
- Lexico: Extreme KV Cache Compression via Sparse Coding over Universal DictionariesJunhyuck Kim, Jongho Park, Jaewoong Cho, Dimitris PapailiopoulosICML 2025
- HShare: Fast LLM Decoding by Hierarchical Key-Value SharingHuaijin Wu, Lianqiang Li, Hantao Huang, Tu Yi 等ICLR 2025
- KV Cache Transform Coding for Compact Storage in LLM InferenceKonrad Staniszewski, Adrian LancuckiICLR 2026 · 被引用 9 次
