CLLMs: Consistency Large Language Models
Siqi Kou, Lanxiang Hu, Zhezhi He, Zhijie Deng, Hao Zhang
摘要
Parallel decoding methods such as Jacobi decoding show promise for more efficient LLM inference as it breaks the sequential nature of the LLM decoding process and transforms it into parallelizable computation. However, in practice, it achieves little speedup compared to traditional autoregressive (AR) decoding, primarily because Jacobi decoding seldom accurately predicts more than one token in a single fixed-point iteration step. To address this, we develop a new approach aimed at realizing fast convergence from any state to the fixed point on a Jacobi trajectory. This is accomplished by refining the target LLM to consistently predict the fixed point given any state as input. Extensive experiments demonstrate the effectiveness of our method, showing 2.4× to 3.4× improvements in generation speed while preserving generation quality across both domain-specific and open-domain benchmarks. Our code is available at https://github.com/hao-ailab/Consistency LLM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
- Phased Consistency ModelsFu-Yun Wang, Zhaoyang Huang, Alexander William Bergman, Dazhong Shen 等NeurIPS 2024 · 被引用 86 次
- Attention Is All You Need for KV Cache in Diffusion LLMsQuan Nguyen-Tri, Mukul Ranjan, Zhiqiang ShenICLR 2026 · 被引用 36 次
- Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge GenerationXinyu Yang, Yuwei An, Hongyi Liu, Tianqi Chen 等NeurIPS 2025 · 被引用 36 次
- d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory DistillationYu-Yang Qian, Junda Su, Lanxiang Hu, Peiyuan Zhang 等ICML 2026 · 被引用 33 次
它引用的顶会 Paper17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
相关 Paper
- Parallel Jacobi Decoding for Fast Autoregressive Image GenerationBoya Liao, Ying Li, Siyong Jian, Huan WangCVPR 2026 · 被引用 2 次
- Fast and Accurate Causal Parallel Decoding using Jacobi ForcingLanxiang Hu, Siqi Kou, Yichao Fu, Samyam Rajbhandari 等ICML 2026 · 被引用 5 次
- Accelerating Transformer Inference for Translation via Parallel DecodingAndrea Santilli, Silvio Severino, Emilian Postolache, Valentino Maiorca 等ACL 2023 · 被引用 19 次
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image GenerationYao Teng, Fuyun Wang, Xian Liu, Zhekai Chen 等NeurIPS 2025 · 被引用 8 次
- AdaDecode: Accelerating LLM Decoding with Adaptive Layer ParallelismZhepei Wei, Wei-Lin Chen, Xinyu Zhu, Yu MengICML 2025
