Turbo: Efficiently Serving Long-Context Large Language Models with In-Network Aggregation
Ying Wan, Yuchen Xu, Chuwen Zhang, Yingsheng Huang, Yong Feng, Wenquan Xu, Jialin Li, Mingwei Xu, Wenfei Wu, Congcong Miao
摘要
LLM supporting long contexts faces a critical memory bottleneck due to the linear growth of KV cache. Distributing the storage across multiple GPUs alleviates this burden but introduces significant communication overhead or traffic incast, especially during the decoding phase. We propose Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches. We address three key challenges to map complex attention mechanisms onto restricted switch hardware: (i) To bypass the switch's inability to buffer global states or perform complex operations, we devise online table-based aggregation, which decomposes global reduction into pairwise operations and approximates nonlinear functions via lookup tables. (ii) To circumvent the restriction on retroactive state access in RMT pipelines, we introduce a rolling forward scheme that propagates states to enable cross-stage updates. (iii) To mitigate aggregation stragglers caused by topology-induced load imbalance, we construct a load-aware aggregation tree that optimizes workload distribution. Evaluations on a Tofino2-based testbed show that Turbo reduces end-to-end inference latency by up to 37%. Large-scale simulations on NS-3 demonstrate that Turbo significantly outperforms state-of-the-art baselines in both inference latency and network traffic reduction with negligible accuracy loss.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- TurboCache: Empowering Switch-Accelerated Key-Value Caches with Accurate and Fast Cache UpdatesXiang Chen, Longlong Zhu, Linying Zheng, Lingfei Cheng 等INFOCOM 2025 · 被引用 1 次
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 被引用 67 次
- LongSight: Compute-Enabled Memory to Accelerate Large-Context LLMs via Sparse AttentionDerrick Quinn, E. Ezgi Yücel, Jinkwon Kim, José F. Martínez 等MICRO 2025 · 被引用 3 次
- Scaling Attention Beyond GPUs for LLM InferenceWeishu Deng, Yujie Yang, Peiran Du, Lingfeng Xiang 等HPDC 2026
- Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache ManagementXinjun Yang, Qingda Hu, Junru Li, Feifei Li 等SIGMOD 2026 · 被引用 24 次
