Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
Tian Jin, Ellie Y. Cheng, Zachary Ankner, Nikunj Saunshi, Blake M. Elias, Amir Yazdanbakhsh, Jonathan Ragan-Kelley, Suvinay Subramanian, Michael Carbin
摘要
Decoding with autoregressive large language models (LLMs) traditionally occurs sequentially, generating one token after another. An emerging line of work explored parallel decoding by identifying and simultaneously generating semantically independent chunks of LLM responses. However, they rely on hand-crafted heuristics tied to syntactic structures like lists and paragraphs, making them rigid and imprecise. We present PASTA, a learning-based system that teaches LLMs to identify semantic independence and express parallel decoding opportunities in their own responses. At its core are PASTA-LANG and its interpreter: PASTA-LANG is an annotation language that enables LLMs to express semantic independence in their own responses; the language interpreter acts on these annotations to orchestrate parallel decoding on-the-fly at inference time. Through a twostage finetuning process, we train LLMs to generate PASTA-LANG annotations that optimize both response quality and decoding speed. Evaluation on AlpacaEval, an instruction following benchmark, shows that our approach Pareto-dominates existing methods in terms of decoding speed and response quality; our results demonstrate geometric mean speedups ranging from 1.21× to 1.93× with corresponding quality changes of +2.2% to -7.1%, measured by length-controlled win rates against sequential decoding baseline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Parallel-R1: Towards Parallel Thinking via Reinforcement LearningTong Zheng, Hongming Zhang, Wenhao Yu, Xiaoyang Wang 等ICLR 2026 · 被引用 53 次
- CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot ManipulationLetian Fu, Justin Yu, Karim El-Refai, Ethan Kou 等ICML 2026 · 被引用 36 次
- Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge GenerationXinyu Yang, Yuwei An, Hongyi Liu, Tianqi Chen 等NeurIPS 2025 · 被引用 36 次
- Hogwild! Inference: Parallel LLM Generation via Concurrent AttentionGleb Rodionov, Roman Garipov, Alina Shutova, George Yakushev 等NeurIPS 2025 · 被引用 35 次
- SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning ModelsEmil Biju, Shayan Talaei, Zhemin Huang, Mohammadreza Pourreza 等NeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Chameleon: Plug-and-Play Compositional Reasoning with Large Language ModelsPan Lu, Baolin Peng, Hao Cheng, Michel Galley 等NeurIPS 2023 · 被引用 515 次
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 被引用 290 次
相关 Paper
- Planned DiffusionDaniel Mingyi Israel, Tian Jin, Ellie Y Cheng, Guy Van den Broeck 等ICLR 2026 · 被引用 8 次
- Learning to Parallel: Accelerating Diffusion Large Language Models via Learnable Parallel DecodingWenrui Bao, Zhiben Chen, Dan Xu, Yuzhang ShangICLR 2026 · 被引用 26 次
- PaDeLLM-NER: Parallel Decoding in Large Language Models for Named Entity RecognitionJinghui Lu, Yanjie Wang, Ziwei Yang, Xuejing Liu 等NeurIPS 2024 · 被引用 22 次
- Multi-Branch Self-Drafting for LLM Inference AccelerationZipeng Gao, Qingrong Xia, Tong Xu, Xinyu Duan 等AAAI 2025 · 被引用 2 次
- CLLMs: Consistency Large Language ModelsSiqi Kou, Lanxiang Hu, Zhezhi He, Zhijie Deng 等ICML 2024 · 被引用 65 次
