MiniMalloc: A Lightweight Memory Allocator for Hardware-Accelerated Machine Learning
Michael D. Moffitt
摘要
We present a new approach to static memory allocation, a key problem that arises in the compilation of machine learning models onto the resources of a specialized hardware accelerator. Our methodology involves a recursive depth-first search that limits exploration to a special class of canonical solutions, dramatically reducing the size of the search space. We also develop a spatial inference technique that exploits this special structure by pruning unpromising partial assignments and backtracking more effectively than otherwise possible. Finally, we introduce a new mechanism capable of detecting and eliminating dominated solutions from consideration. Empirical results demonstrate orders of magnitude improvement in performance as compared to the previous state-of-the-art on many benchmarks, as well as a substantial reduction in library size.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal PlanningZixiao Huang, Junhao Hu, Hao Lin, Chunyang Zhu 等EuroSys 2026 · 被引用 1 次
- jwmalloc: A Verified Memory Allocator for Mobile DevicesJiawei Wang, Ming Fu, Ruixian Wang, Chao Xu 等OSDI 2026
- Automatically Generating ML Compiler Backends from Tensor Accelerator ISA DescriptionsDevansh Jain, Akash Pardeshi, Marco Frigo, Kaustubh Khulbe 等OOPSLA 2026
相关 Paper
- AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstractionSize Zheng, Renze Chen, Anjiang Wei, Yicheng Jin 等ISCA 2022 · 被引用 63 次
- TelaMalloc: Efficient On-Chip Memory Allocation for Production Machine Learning AcceleratorsMartin Maas, Ulysse Beaugnon, Arun Chauhan, Berkin IlbeyiASPLOS 2023 · 被引用 21 次
- Distilling Arbitration Logic from Traces using Machine Learning: A Case Study on NoCYuan Zhou, Hanyu Wang, Jieming Yin, Zhiru ZhangDAC 2021 · 被引用 8 次
- Soter: Analytical Tensor-Architecture Modeling and Automatic Tensor Program Tuning for Spatial AcceleratorsFuyu Wang, Minghua Shen, Yufei Ding, Nong XiaoISCA 2024 · 被引用 7 次
- MESA: Microarchitecture Extensions for Spatial Architecture GenerationDong Kai Wang, Jiaqi Lou, Naiyin Jin, Edwin Mascarenhas 等ISCA 2023 · 被引用 2 次
