Lune

ICML2026Top-tier venue

AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding

Shuqing Luo, Yilin Guan, Pingzhi Li, Hanrui Wang, Tianlong Chen

2026Year
2Citations

Abstract

Test-time scaling (TTS) can boost LLM reasoning through long chain-of-thought (CoT), but the linear KV-cache growth amplifies the memory-bound bottleneck of LLM decoding. Query-aware sparse decoding methods can achieve state-of-the-art performance under constrained FLOP budget, but are mainly constrained by both sequential-dependent page filtering and coarse-grained token selection, hampering the serving efficiency and model performance on TTS tasks under high concurrency and long CoT scenarios, where token selection can even occupy higher runtime than the forward pipeline itself. In this paper, we first find that the query state of the current decoding token can be approximated in a unified manner from a short sliding window of recent queries, enabling training-free query-aware sparsity without sequential dependency in the decoding loop. Based on the findings, we propose \texttt{\textbf{AsyncSpade}}, an asynchronous framework for efficient TTS, built on two core components: (1) a novel light-weight temporal-regressive module\textbf{(1) a novel light-weight temporal-regressive module} that predicts the next-token query state, and (2) an asynchronous disaggregated framework\textbf{(2) an asynchronous disaggregated framework} that decouples the KV cache selection from the auto-regressive decoding loop, overlapping the token-level KV selection with the forward inference computation through asynchronism, thereby eliminating the sequential dependency without sacrificing model performance. We validate the effectiveness of AsyncSpade\texttt{AsyncSpade} on common LLM serving setups with an A100 node, where AsyncSpade\texttt{AsyncSpade} can fully overlap KV-cache operations with the inference pipeline within a certain workload range, achieving theoretical optimal time-per-output-token (TPOT)\textbf{achieving theoretical optimal time-per-output-token~(TPOT)}. Specifically, AsyncSpade\texttt{AsyncSpade} delivers over 20% reduction on TPOT compared to SoTA baseline (i.e.\textit{i.e.} Quest) and at least 50% TPOT reduction compared to full attention on Qwen3-8B and Qwen3-32B models, while matching or surpassing their accuracy on various TTS benchmarks (AIME-24/25, GPQA-Diamond, MATH-500). Our code is available through https://github.com/UNITES-Lab/AsyncSpade.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f0eb2875-00d8-4bdf-b1cc-cbde15c071dd

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines