TPP-SD: Accelerating Transformer Point Process Sampling with Speculative Decoding
Shukai Gong, Yiyang Fu, Fengyuan Ran, Quyu Kong, Feng Zhou
摘要
We propose TPP-SD, a novel approach that accelerates Transformer temporal point process (TPP) sampling by adapting speculative decoding (SD) techniques from language models. By identifying the structural similarities between thinning algorithms for TPPs and speculative decoding for language models, we develop an efficient sampling framework that leverages a smaller draft model to generate multiple candidate events, which are then verified by the larger target model in parallel. TPP-SD maintains the same output distribution as autoregressive sampling while achieving significant acceleration. Experiments on both synthetic and real datasets demonstrate that our approach produces samples from identical distributions as standard methods, but with 2-6× speedup. Our ablation studies analyze the impact of hyperparameters such as draft length and draft model size on sampling efficiency. TPP-SD bridges the gap between powerful Transformer TPP models and the practical need for rapid sequence sampling. Code is publicly available at https://github.com/GONGSHUKAI/tppsd.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 被引用 424 次
- Transformer Hawkes ProcessSimiao Zuo, Haoming Jiang, Zichong Li, Tuo Zhao 等ICML 2020 · 被引用 382 次
- Self-Attentive Hawkes ProcessQiang Zhang, Aldo Lipani, Ömer Kirnap, Emine YilmazICML 2020 · 被引用 254 次
- Intensity-Free Learning of Temporal Point ProcessesOleksandr Shchur, Marin Bilos, Stephan GünnemannICLR 2020 · 被引用 210 次
相关 Paper
- Parallel Token Prediction for Language ModelsFelix Draxler, Justus C. Will, Farrin Marouf Sofian, Theofanis Karaletsos 等ICLR 2026 · 被引用 6 次
- Dynamic-Width Speculative Beam Decoding for LLM InferenceZongyue Qin, Zifan He, Neha Prakriya, Jason Cong 等AAAI 2025 · 被引用 10 次
- Cascade Speculative Drafting for Even Faster LLM InferenceZiyi Chen, Xiaocong Yang, Jiacheng Lin, Chenkai Sun 等NeurIPS 2024 · 被引用 107 次
- ASDSV: Multimodal Generation Made Efficient with Approximate Speculative Diffusion and Speculative VerificationKaijun Zhou, Xingyu Yan, Xingda Wei, Xijun Li 等NeurIPS 2025
- A Theoretical Perspective for Speculative Decoding AlgorithmMing Yin, Minshuo Chen, Kaixuan Huang, Mengdi WangNeurIPS 2024 · 被引用 36 次
