Lune

ICML2026顶会

AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding

Shuqing Luo, Yilin Guan, Pingzhi Li, Hanrui Wang, Tianlong Chen

2026年份
2被引次数

摘要

Test-time scaling (TTS) can boost LLM reasoning through long chain-of-thought (CoT), but the linear KV-cache growth amplifies the memory-bound bottleneck of LLM decoding. Query-aware sparse decoding methods can achieve state-of-the-art performance under constrained FLOP budget, but are mainly constrained by both sequential-dependent page filtering and coarse-grained token selection, hampering the serving efficiency and model performance on TTS tasks under high concurrency and long CoT scenarios, where token selection can even occupy higher runtime than the forward pipeline itself. In this paper, we first find that the query state of the current decoding token can be approximated in a unified manner from a short sliding window of recent queries, enabling training-free query-aware sparsity without sequential dependency in the decoding loop. Based on the findings, we propose \texttt{\textbf{AsyncSpade}}, an asynchronous framework for efficient TTS, built on two core components: (1) a novel light-weight temporal-regressive module\textbf{(1) a novel light-weight temporal-regressive module} that predicts the next-token query state, and (2) an asynchronous disaggregated framework\textbf{(2) an asynchronous disaggregated framework} that decouples the KV cache selection from the auto-regressive decoding loop, overlapping the token-level KV selection with the forward inference computation through asynchronism, thereby eliminating the sequential dependency without sacrificing model performance. We validate the effectiveness of AsyncSpade\texttt{AsyncSpade} on common LLM serving setups with an A100 node, where AsyncSpade\texttt{AsyncSpade} can fully overlap KV-cache operations with the inference pipeline within a certain workload range, achieving theoretical optimal time-per-output-token (TPOT)\textbf{achieving theoretical optimal time-per-output-token~(TPOT)}. Specifically, AsyncSpade\texttt{AsyncSpade} delivers over 20% reduction on TPOT compared to SoTA baseline (i.e.\textit{i.e.} Quest) and at least 50% TPOT reduction compared to full attention on Qwen3-8B and Qwen3-32B models, while matching or surpassing their accuracy on various TTS benchmarks (AIME-24/25, GPQA-Diamond, MATH-500). Our code is available through https://github.com/UNITES-Lab/AsyncSpade.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext f0eb2875-00d8-4bdf-b1cc-cbde15c071dd

它引用的顶会 Paper15

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖