Micro-Armed Bandit: Lightweight & Reusable Reinforcement Learning for Microarchitecture Decision-Making
Gerasimos Gerogiannis, Josep Torrellas
摘要
Online Reinforcement Learning (RL) has been adopted as an effective mechanism in various decision-making problems in microarchitecture. Its high adaptability and the ability to learn at runtime are attractive characteristics in microarchitecture settings. However, although hardware RL agents are effective, they suffer from two main problems. First, they have high complexity and storage overhead. This complexity stems from decomposing the environment into a large number of states and then, for each of these states, bookkeeping many action values. Second, many RL agents are engineered for a specific application and are not reusable.
In this work, we tackle both of these shortcomings by designing an RL agent that is both lightweight and reusable across different microarchitecture decision-making problems. We find that, in some of these problems, only a small fraction of the action space is useful in a given time window. We refer to this property as temporal homogeneity in the action space. Motivated by this property, we design an RL agent based on Multi-Armed Bandit algorithms, the simplest form of RL. We call our agent Micro-Armed Bandit.
We showcase our agent in two use cases: data prefetching and instruction fetch in simultaneous multithreaded (SMT) processors. For prefetching, our agent outperforms non-RL prefetchers Bingo and MLOP by 2.6% and 2.3% (geometric mean), respectively, and attains similar performance as the state-of-the-art RL prefetcher Pythia-with the dramatically lower storage requirement of only 100 bytes. For SMT instruction fetch, our agent outperforms the Hill Climbing method by 2.2% (geometric mean).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered MemoryJie Ren, Bin Ma, Shuangyan Yang, Benjamin Francis 等HPCA 2025 · 被引用 6 次
- Micro-MAMA: Multi-Agent Reinforcement Learning for Multicore PrefetchingCharles Block, Gerasimos Gerogiannis, Josep TorrellasMICRO 2025 · 被引用 6 次
- Integrating Prefetcher Selection with Dynamic Request Allocation Improves Prefetching EfficiencyMengming Li, Qijun Zhang, Yongqing Ren, Zhiyao XieHPCA 2025 · 被引用 6 次
- ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory AccessesMengming Li, Chenlu Miao, Buqing Xu, Qijun Zhang 等ISCA 2026
- DAL: A Practical Prior-Free Black-Box Framework for Piecewise Stationary BanditsArgyrios Gerogiannis, Yu-Han Huang, Subhonmesh Bose, Venugopal VeeravalliICML 2026
它引用的顶会 Paper13
- A hierarchical neural model of data prefetchingZhan Shi, Akanksha Jain, Kevin Swersky, Milad Hashemi 等ASPLOS 2021 · 被引用 100 次
- Bouquet of Instruction Pointers: Instruction Pointer Classifier-based Spatial Hardware PrefetchingSamuel Pakalapati, Biswabandan PandaISCA 2020 · 被引用 97 次
- Pythia: A Customizable Hardware Prefetching Framework Using Online Reinforcement LearningRahul Bera, Konstantinos Kanellopoulos, Anant Nori, Taha Shahroodi 等MICRO 2021 · 被引用 95 次
- Classifying Memory Access Patterns for PrefetchingGrant Ayers, Heiner Litz, Christos Kozyrakis, Parthasarathy RanganathanASPLOS 2020 · 被引用 83 次
- Designing a Cost-Effective Cache Replacement Policy using Machine LearningSubhash Sethumurugan, Jieming Yin, John SartoriHPCA 2021 · 被引用 81 次
相关 Paper
- Towards Automated RISC-V Microarchitecture Design with Reinforcement LearningChen Bai, Jianwang Zhai, Yuzhe Ma, Bei Yu 等AAAI 2024 · 被引用 20 次
- Post-Fabrication MicroarchitectureChanchal Kumar, Anirudh Seshadri, Aayush Chaudhary, Shubham Bhawalkar 等MICRO 2021 · 被引用 4 次
- ReSemble: Reinforced Ensemble Framework for Data PrefetchingPengmiao Zhang, Rajgopal Kannan, Ajitesh Srivastava, Anant V. Nori 等SC 2022 · 被引用 17 次
- I-POP: Ignite Positive PrefetchersYiquan Lin, Wenhai Lin, Yiquan Chen, Jiexiong Xu 等HPCA 2026 · 被引用 1 次
- Merging Similar Patterns for Hardware PrefetchingShizhi Jiang, Qiusong Yang, Yiwei CiMICRO 2022 · 被引用 27 次
