Optimizing Memory Placement using Evolutionary Graph Reinforcement Learning
Shauharda Khadka, Estelle Aflalo, Mattias Marder, Avrech Ben-David, Santiago Miret, Shie Mannor, Tamir Hazan, Hanlin Tang, Somdeb Majumdar
Abstract
As modern neural networks have grown to billions of parameters, meeting tight latency budgets has become increasingly challenging. Approaches like compression, sparsification and network pruning have proven effective to tackle this problem - but they rely on modifications of the underlying network. In this paper, we look at a complimentary approach of optimizing how tensors are mapped to on-chip memory in an inference accelerator while leaving the network parameters untouched. Since different memory components trade off capacity for bandwidth differently, a sub-optimal mapping can result in high latency. We introduce evolutionary graph reinforcement learning (EGRL) - a method combining graph neural networks, reinforcement learning (RL) and evolutionary search - that aims to find the optimal mapping to minimize latency. Furthermore, a set of fast, stateless policies guide the evolutionary search to improve sample-efficiency. We train and validate our approach directly on the Intel NNP-I chip for inference using a batch size of 1. EGRL outperforms policy-gradient, evolutionary search and dynamic programming baselines on BERT, ResNet-101 and ResNet-50. We achieve 28-78% speed-up compared to the native NNP-I compiler on all three workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Searching for High-Value Molecules Using Reinforcement Learning and TransformersRaj Ghugare, Santiago Miret, Adriana Hugessen, Mariano Phielipp et al.ICLR 2024 · 22 citations
- Robust Scheduling with GFlowNetsDavid W. Zhang, Corrado Rainone, Markus Peschl, Roberto BondesanICLR 2023 · 2 citations
- Maintaining Sanity: Algorithm-based Comprehensive Fault Tolerance for CNNsJinhyo Jung, Hwisoo So, Woobin Ko, Sumedh Shridhar Joshi et al.DAC 2024 · 1 citation
Builds on2
Related papers
- Topology-Aware Network Pruning using Multi-stage Graph Embedding and Reinforcement LearningSixing Yu, Arya Mazaheri, Ali JannesariICML 2022 · 54 citations
- RESPECT: Reinforcement Learning based Edge Scheduling on Pipelined Coral Edge TPUsJiaqi Yin, Yingjie Li, Daniel Robinson, Cunxi YuDAC 2023 · 9 citations
- Auto Graph Encoder-Decoder for Neural Network PruningSixing Yu, Arya Mazaheri, Ali JannesariICCV 2021 · 47 citations
- OptiPIM: Optimizing Processing-in-Memory Acceleration Using Integer Linear ProgrammingJiantao Liu, Minxuan Zhou, Yue Pan, Chien-Yi Yang et al.ISCA 2025 · 6 citations
- Transferable Graph Optimizers for ML CompilersYanqi Zhou, Sudip Roy, AmirAli Abdolrashidi, Daniel Wong et al.NeurIPS 2020 · 63 citations
