CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
Bradley McDanel, Steven Li, Harshit Khaitan
摘要
Token-ranking heuristics accelerate the prefill bottleneck in long-context LLM inference by selectively processing semantically relevant tokens. However, current evaluation relies on end-to-end benchmarks, making it difficult to isolate the quality of the token ranking itself. To address this, we introduce an Answer-Informed Reference framework with two variants: a Model Reference (Model-Ref) that measures token importance using the model generated answer, and a Ground-Truth Reference (GT-Ref) that uses human reference answers. Using GT-Ref, we establish that as few as 10% of prompt tokens suffice to match full-context performance on LongBench across all three evaluated models. The framework also reveals that existing heuristics exhibit variance across layers, with rankings degrading sharply at specific layers. Motivated by this instability, we propose Cross-Layer Attention Aggregation (CLAA), which aggregates importance scores across consecutive layers, eliminating the layer-dependent accuracy collapse observed in single-layer methods. A meaningful gap remains between the best heuristic and GT-Ref, indicating theoretical room for improved token selection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu 等NeurIPS 2024 · 被引用 479 次
- You Only Cache Once: Decoder-Decoder Architectures for Language ModelsYutao Sun, Li Dong, Yi Zhu, Shaohan Huang 等NeurIPS 2024 · 被引用 162 次
相关 Paper
- Evolving Sparsity: Leveraging Token Importance Dynamics for Efficient LLM Decoding with Sparse AttentionRuizi Han, Miao Zhang, Ziyue Qiao, Liqiang NieACL 2026
- Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance EstimationJingyu Liu, Beidi Chen, Ce ZhangICML 2025
- SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token PruningLingkun Long, Rubing Yang, Yushi Huang, Desheng Hui 等AAAI 2026 · 被引用 8 次
- OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inferenceSeungjun Shin, Jaehoon Oh, Dokwan OhICML 2025
- Towards Long-Horizon Interpretability: Efficient and Faithful Multi-Token Attribution for Reasoning LLMsWenbo Pan, Zhichao Liu, Xianlong Wang, Yu Haining 等ICML 2026 · 被引用 3 次
