CONTINUUM: Restoring the Contiguous Tensor Abstraction Efficiently for Dynamic AI Workloads via Hardware Virtualization
Yangyu Zhang, shuoming zhang, Chunwei Xia, Shuaijiang Li, Zhicheng Li, Ruiyuan Xu, Zheming Yang, Lei Chen, YUAN WEN, Guangli Li, Xiaobing Feng, Huimin Cui, Jiacheng Zhao
Abstract
Emerging LLM workloads demand extreme mem- ory agility. However, state-of-the-art inference systems (e.g., vLLM) rely on software-defined paging, which sacrifices the contiguous tensor abstraction. This rigid interface exposes fragmen- tation complexity to developers, imposing a se- vere engineering burden that stifles algorithmic innovation. We introduce CONTINUUM, a tensor memory virtualization subsystem implemented as a PyTorch extension. By bypassing serialized OS bottlenecks via a lightweight GPU driver ex- tension, CONTINUUM can significantly reduce the mapping costs by orders of magnitude—from milliseconds to microseconds. Built atop this low-latency API, CONTINUUM provides Elastic Tensor, with a set of flexible tensor operations that natively supports complex memory dynamics and zero-copy topological aliasing. Evaluations demonstrate that CONTINUUM achieves signifi- cantly higher throughput across diverse dynamic scenarios, effectively democratizing the imple- mentation of next-generation LLM applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Data is all you need: Finetuning LLMs for Chip Design via an Automated design-data augmentation frameworkKaiyan Chang, Kun Wang, Nan Yang, Ying Wang et al.DAC 2024 · 59 citations
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language ModelsXin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen et al.ACL 2026 · 57 citations
Related papers
- vAttention: Dynamic Memory Management for Serving LLMs without PagedAttentionRamya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee et al.ASPLOS 2025 · 38 citations
- ConServe: Contiguity-Preserving Memory Management for Multi-Turn LLM ServingBingyao LiISCA 2026
- STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal PlanningZixiao Huang, Junhao Hu, Hao Lin, Chunyang Zhu et al.EuroSys 2026 · 1 citation
- GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory StitchingCong Guo, Rui Zhang, Jiale Xu, Jingwen Leng et al.ASPLOS 2024 · 30 citations
- T-Control: An Efficient Dynamic Tensor Rematerialization System for DNN TrainingZehua Wang, Junmin Xiao, Xiaochuan Deng, Huibing Wang et al.ASPLOS 2026 · 2 citations
