Turbo: Efficiently Serving Long-Context Large Language Models with In-Network Aggregation
Ying Wan, Yuchen Xu, Chuwen Zhang, Yingsheng Huang, Yong Feng, Wenquan Xu, Jialin Li, Mingwei Xu, Wenfei Wu, Congcong Miao
Abstract
LLM supporting long contexts faces a critical memory bottleneck due to the linear growth of KV cache. Distributing the storage across multiple GPUs alleviates this burden but introduces significant communication overhead or traffic incast, especially during the decoding phase. We propose Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches. We address three key challenges to map complex attention mechanisms onto restricted switch hardware: (i) To bypass the switch's inability to buffer global states or perform complex operations, we devise online table-based aggregation, which decomposes global reduction into pairwise operations and approximates nonlinear functions via lookup tables. (ii) To circumvent the restriction on retroactive state access in RMT pipelines, we introduce a rolling forward scheme that propagates states to enable cross-stage updates. (iii) To mitigate aggregation stragglers caused by topology-induced load imbalance, we construct a load-aware aggregation tree that optimizes workload distribution. Evaluations on a Tofino2-based testbed show that Turbo reduces end-to-end inference latency by up to 37%. Large-scale simulations on NS-3 demonstrate that Turbo significantly outperforms state-of-the-art baselines in both inference latency and network traffic reduction with negligible accuracy loss.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 04217456-87e7-401b-b319-ee4d98c91e4dRelated papers
- TurboCache: Empowering Switch-Accelerated Key-Value Caches with Accurate and Fast Cache UpdatesXiang Chen, Longlong Zhu, Linying Zheng, Lingfei Cheng et al.INFOCOM 2025 · 1 citation
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 67 citations
- LongSight: Compute-Enabled Memory to Accelerate Large-Context LLMs via Sparse AttentionDerrick Quinn, E. Ezgi Yücel, Jinkwon Kim, José F. Martínez et al.MICRO 2025 · 3 citations
- Scaling Attention Beyond GPUs for LLM InferenceWeishu Deng, Yujie Yang, Peiran Du, Lingfeng Xiang et al.HPDC 2026
- Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache ManagementXinjun Yang, Qingda Hu, Junru Li, Feifei Li et al.SIGMOD 2026 · 24 citations
