CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative Drafter
Yepeng Weng, Dianwen Mei, Huishi Qiu, Xujie Chen, Li Liu, Jiang Tian, Zhongchao Shi
摘要
Speculative decoding is a powerful technique that accelerates Large Language Model (LLM) inference by leveraging a lightweight speculative draft model. However, existing designs suffers in performance due to misalignment between training and inference. Recent methods have tried to solve this issue by adopting a multi-step training strategy, but the complex inputs of different training steps make it harder for the draft model to converge. To address this, we propose CORAL, a novel framework that improves both accuracy and efficiency in speculative drafting. CORAL introduces Cross-Step Representation Alignment, a method that enhances consistency across multiple training steps, significantly improving speculative drafting performance. Additionally, we identify the LM head as a major bottleneck in the inference speed of the draft model. We introduce a weight-grouping mechanism that selectively activates a subset of LM head parameters during inference, substantially reducing the latency of the draft model. We evaluate CORAL on three LLM families and three benchmark datasets, achieving speedup ratios of 2.50x-4.07x, outperforming state-of-the-art methods such as EAGLE-2 and HASS. Our results demonstrate that CORAL effectively mitigates training-inference misalignment and delivers significant speedup for modern LLMs with large vocabularies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft LearningYizhou Zhang, Ning Lv, Teng Wang, Jisheng DangICLR 2026 · 被引用 10 次
- NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context VocabulariesZhiyang Chen, Daliang Xu, Yinyuan Zhang, Chenghua Wang 等ICML 2026 · 被引用 1 次
- EDSD: Entropy-Driven Design for Faster Speculative DecodingLongkai Cheng, Ximing Wang, Jiangcai Zhu, Kailai Shao 等ACL 2026
它引用的顶会 Paper16
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- HCSpec: Two-Tier Horizontal Cascade Speculative Decoding for High-Efficiency Large Language Model InferenceYizhou Zhang, Siming Chen, Hao Ye, Erhu FengACL 2026
- PARD: Accelerating LLM Inference with Low‑Cost PARallel Draft Model AdaptationZihao An, Huajun Bai, Ziqiong Liu, Dong Li 等ICLR 2026 · 被引用 28 次
- Learning Harmonized Representations for Speculative SamplingLefan Zhang, Xiaodan Wang, Yanhua Huang, Ruiwen XuICLR 2025
- GRIFFIN: Effective Token Alignment for Faster Speculative DecodingShijing Hu, Jingyang Li, Xingyu Xie, Zhihui Lu 等NeurIPS 2025 · 被引用 14 次
- RepSpec: Structural Re-parameterized Draft Model Training for Speculative DecodingFeiye Huo, Jianchao Tan, Jiahao Liu, Zixu Jiang 等ICLR 2026
