TelaMalloc: Efficient On-Chip Memory Allocation for Production Machine Learning Accelerators
Martin Maas, Ulysse Beaugnon, Arun Chauhan, Berkin Ilbeyi
2023Year
21Citations
6Top-tier citations
Abstract
Memory buffer allocation for on-chip memories is a major challenge in modern machine learning systems that target ML accelerators. In interactive systems such as mobile phones, it is on the critical path of launching ML-enabled applications. In data centers, it is part of complex optimization loops that run many times and are the limiting factor for the quality of compilation results.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bbe3f1ca-ad04-41e3-a5f1-2956c2c0e9faCited by top-tier papers6
- FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge DevicesXiangyu Li, Yuanchun Li, Yuanzhe Li, Ting Cao et al.MobiCom 2024 · 40 citations
- SoD2: Statically Optimizing Dynamic Deep Neural Network ExecutionWei Niu, Gagan Agrawal, Bin RenASPLOS 2024 · 6 citations
- STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal PlanningZixiao Huang, Junhao Hu, Hao Lin, Chunyang Zhu et al.EuroSys 2026 · 1 citation
- Automatically Generating ML Compiler Backends from Tensor Accelerator ISA DescriptionsDevansh Jain, Akash Pardeshi, Marco Frigo, Kaustubh Khulbe et al.OOPSLA 2026
- Tempo: Compiled Dynamic Deep Learning with Symbolic Dependence GraphsPedro F. Silvestre, Peter R. PietzuchSOSP 2025
Related papers
- MiniMalloc: A Lightweight Memory Allocator for Hardware-Accelerated Machine LearningMichael D. MoffittASPLOS 2023 · 5 citations
- DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory ArchitecturesPeiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati et al.ISCA 2026 · 3 citations
- Characterizing a Memory Allocator at Warehouse ScaleZhuangzhuang Zhou, Vaibhav Gogte, Nilay Vaish, Chris Kennelly et al.ASPLOS 2024 · 17 citations
- Pocket: ML Serving from the EdgeMisun Park, Ketan Bhardwaj, Ada GavrilovskaEuroSys 2023 · 10 citations
- Taming Dynamism on GPUs: Cross-SM Kernel Fusion via SM Cooperation and Just-in-Time ReductionJingkai He, Guangda Sun, TianJian Li, Dong Du et al.SOSP 2026
