Superposed Decoding: Multiple Generations from a Single Autoregressive Inference Pass
Ethan Shen, Alan Fan, Sarah M. Pratt, Jae Sung Park, Matthew Wallingford, Sham M. Kakade, Ari Holtzman, Ranjay Krishna, Ali Farhadi, Aditya Kusupati
摘要
Many applications today provide users with multiple auto-complete drafts as they type, including GitHub's code completion, Gmail's smart compose, and Apple's messaging auto-suggestions. Under the hood, language models support this by running an autoregressive inference pass to provide a draft. Consequently, providing drafts to the user requires running an expensive language model times. To alleviate the computation cost of running inference passes, we propose Superposed Decoding, a new decoding algorithm that generates drafts at the computation cost of one autoregressive inference pass. We achieve this by feeding a superposition of the most recent token embeddings from the drafts as input to the next decoding step of the language model. At every inference step we combine the drafts with the top- tokens to get new drafts and cache the most likely options, using an n-gram interpolation with minimal compute overhead to filter out incoherent generations. Our experiments show that drafts from Superposed Decoding are at least as coherent and factual as Nucleus Sampling and Greedy Decoding respectively, while being at least faster for . In a compute-normalized setting, user evaluations demonstrably favor text generated by Superposed Decoding over Nucleus Sampling. Superposed Decoding can also be combined with other decoding strategies, resulting in universal coverage gains when scaling inference time compute. Code and more examples open-sourced at https://github.com/RAIVNLab/SuperposedDecoding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Q-Probe: A Lightweight Approach to Reward Maximization for Language ModelsKenneth Li, Samy Jelassi, Hugh Zhang, Sham M. Kakade 等ICML 2024 · 被引用 18 次
- Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in SuperpositionZheyang Xiong, Ziyang Cai, John Cooper, Albert Ge 等ICML 2025
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
相关 Paper
- Speculative Speculative DecodingTanishq Kumar, Tri Dao, Avner MayICLR 2026 · 被引用 15 次
- Automatic Detection of Generated Text is Easiest when Humans are FooledDaphne Ippolito, Daniel Duckworth, Chris Callison-Burch, Douglas EckACL 2020 · 被引用 21 次
- Multi-Branch Self-Drafting for LLM Inference AccelerationZipeng Gao, Qingrong Xia, Tong Xu, Xinyu Duan 等AAAI 2025 · 被引用 2 次
- Cascade Speculative Drafting for Even Faster LLM InferenceZiyi Chen, Xiaocong Yang, Jiacheng Lin, Chenkai Sun 等NeurIPS 2024 · 被引用 107 次
- SAM Decoding: Speculative Decoding via Suffix AutomatonYuxuan Hu, Ke Wang, Xiaokang Zhang, Fanjin Zhang 等ACL 2025
