Tetris: Efficient Long-context LLM Serving with Chunkwise Dynamic Sequence Parallelism
Cong Li, Yuzhe Yang, Xuegui Zheng, Qifan Yang, Yijin Guan, Size Zheng, Li-Wen Chang, Shufan Liu, Xin Liu, Guangyu Sun
Abstract
With the advancement of large language models (LLMs), their context windows have rapidly expanded. To meet diverse demands from varying-length requests in online services, existing state-of-the-art systems adjust resource allocation by tuning the sequence parallelism (SP) allocation. However, current dynamic SP allocation lacks flexibility to (1) support stage-specific parallelism requirements in LLM inference, (2) mitigate the global latency degradation from excessive SP allocation, and (3) exploit resource fragments arising from SP size variation. To tackle this problem, we propose Chunkwise Dynamic Sequence Parallelism (CDSP), a fine-grained parallelism strategy that assigns SP sizes across intra-request token segments. Based on CDSP, we build Tetris, an LLM serving system that (1) efficiently integrates CDSP into disaggregated cluster architecture to satisfy parallelism heterogeneity, (2) dynamically regulates SP size expansion based on real-time load conditions, and (3) adaptively explores chunking plans to utilize fragmented resources while meeting per-request demands. Compared with state-ofthe-art systems, Tetris achieves up to lower time-to-firsttoken (TTFT) under max sustainable loads, reduces median timebetween-tokens (TBT) by up to 40.1%, and increases the max request capacity by up to 45%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9ef5cffe-5f6f-4e55-acf6-d89c9167ddafRelated papers
- LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence ParallelismBingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun et al.SOSP 2024 · 32 citations
- TETRIS: Optimal Draft Token Selection for Batch Speculative DecodingZhaoxuan Wu, Zijian Zhou, Arun Verma, Alok Prakash et al.ACL 2025 · 7 citations
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
- TetriServe: Efficiently Serving Mixed DiT WorkloadsRunyu Lu, Shiqi He, Wenxuan Tan, Shenggui Li et al.ASPLOS 2026 · 1 citation
- Revisiting Pipeline Parallelism for LLM ServingSoonjae Hwang, Jeongseob AhnOSDI 2026
