CLLMs: Consistency Large Language Models
Siqi Kou, Lanxiang Hu, Zhezhi He, Zhijie Deng, Hao Zhang
Abstract
Parallel decoding methods such as Jacobi decoding show promise for more efficient LLM inference as it breaks the sequential nature of the LLM decoding process and transforms it into parallelizable computation. However, in practice, it achieves little speedup compared to traditional autoregressive (AR) decoding, primarily because Jacobi decoding seldom accurately predicts more than one token in a single fixed-point iteration step. To address this, we develop a new approach aimed at realizing fast convergence from any state to the fixed point on a Jacobi trajectory. This is accomplished by refining the target LLM to consistently predict the fixed point given any state as input. Extensive experiments demonstrate the effectiveness of our method, showing 2.4× to 3.4× improvements in generation speed while preserving generation quality across both domain-specific and open-domain benchmarks. Our code is available at https://github.com/hao-ailab/Consistency LLM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers37
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Phased Consistency ModelsFu-Yun Wang, Zhaoyang Huang, Alexander William Bergman, Dazhong Shen et al.NeurIPS 2024 · 86 citations
- Attention Is All You Need for KV Cache in Diffusion LLMsQuan Nguyen-Tri, Mukul Ranjan, Zhiqiang ShenICLR 2026 · 36 citations
- Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge GenerationXinyu Yang, Yuwei An, Hongyi Liu, Tianqi Chen et al.NeurIPS 2025 · 36 citations
- d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory DistillationYu-Yang Qian, Junda Su, Lanxiang Hu, Peiyuan Zhang et al.ICML 2026 · 33 citations
Builds on17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
Related papers
- Parallel Jacobi Decoding for Fast Autoregressive Image GenerationBoya Liao, Ying Li, Siyong Jian, Huan WangCVPR 2026 · 2 citations
- Fast and Accurate Causal Parallel Decoding using Jacobi ForcingLanxiang Hu, Siqi Kou, Yichao Fu, Samyam Rajbhandari et al.ICML 2026 · 5 citations
- Accelerating Transformer Inference for Translation via Parallel DecodingAndrea Santilli, Silvio Severino, Emilian Postolache, Valentino Maiorca et al.ACL 2023 · 19 citations
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image GenerationYao Teng, Fuyun Wang, Xian Liu, Zhekai Chen et al.NeurIPS 2025 · 8 citations
- AdaDecode: Accelerating LLM Decoding with Adaptive Layer ParallelismZhepei Wei, Wei-Lin Chen, Xinyu Zhu, Yu MengICML 2025
