Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
Tian Jin, Ellie Y. Cheng, Zachary Ankner, Nikunj Saunshi, Blake M. Elias, Amir Yazdanbakhsh, Jonathan Ragan-Kelley, Suvinay Subramanian, Michael Carbin
Abstract
Decoding with autoregressive large language models (LLMs) traditionally occurs sequentially, generating one token after another. An emerging line of work explored parallel decoding by identifying and simultaneously generating semantically independent chunks of LLM responses. However, they rely on hand-crafted heuristics tied to syntactic structures like lists and paragraphs, making them rigid and imprecise. We present PASTA, a learning-based system that teaches LLMs to identify semantic independence and express parallel decoding opportunities in their own responses. At its core are PASTA-LANG and its interpreter: PASTA-LANG is an annotation language that enables LLMs to express semantic independence in their own responses; the language interpreter acts on these annotations to orchestrate parallel decoding on-the-fly at inference time. Through a twostage finetuning process, we train LLMs to generate PASTA-LANG annotations that optimize both response quality and decoding speed. Evaluation on AlpacaEval, an instruction following benchmark, shows that our approach Pareto-dominates existing methods in terms of decoding speed and response quality; our results demonstrate geometric mean speedups ranging from 1.21× to 1.93× with corresponding quality changes of +2.2% to -7.1%, measured by length-controlled win rates against sequential decoding baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1cf8ac9-6fdd-4a38-98a4-cbdf3190a539Cited by top-tier papers9
- Parallel-R1: Towards Parallel Thinking via Reinforcement LearningTong Zheng, Hongming Zhang, Wenhao Yu, Xiaoyang Wang et al.ICLR 2026 · 53 citations
- CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot ManipulationLetian Fu, Justin Yu, Karim El-Refai, Ethan Kou et al.ICML 2026 · 36 citations
- Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge GenerationXinyu Yang, Yuwei An, Hongyi Liu, Tianqi Chen et al.NeurIPS 2025 · 36 citations
- Hogwild! Inference: Parallel LLM Generation via Concurrent AttentionGleb Rodionov, Roman Garipov, Alina Shutova, George Yakushev et al.NeurIPS 2025 · 35 citations
- SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning ModelsEmil Biju, Shayan Talaei, Zhemin Huang, Mohammadreza Pourreza et al.NeurIPS 2025 · 9 citations
Builds on8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Chameleon: Plug-and-Play Compositional Reasoning with Large Language ModelsPan Lu, Baolin Peng, Hao Cheng, Michel Galley et al.NeurIPS 2023 · 515 citations
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 290 citations
Related papers
- Planned DiffusionDaniel Mingyi Israel, Tian Jin, Ellie Y Cheng, Guy Van den Broeck et al.ICLR 2026 · 8 citations
- Learning to Parallel: Accelerating Diffusion Large Language Models via Learnable Parallel DecodingWenrui Bao, Zhiben Chen, Dan Xu, Yuzhang ShangICLR 2026 · 26 citations
- PaDeLLM-NER: Parallel Decoding in Large Language Models for Named Entity RecognitionJinghui Lu, Yanjie Wang, Ziwei Yang, Xuejing Liu et al.NeurIPS 2024 · 22 citations
- Multi-Branch Self-Drafting for LLM Inference AccelerationZipeng Gao, Qingrong Xia, Tong Xu, Xinyu Duan et al.AAAI 2025 · 2 citations
- CLLMs: Consistency Large Language ModelsSiqi Kou, Lanxiang Hu, Zhezhi He, Zhijie Deng et al.ICML 2024 · 65 citations
