BLOwing Trees to the Ground: Layout Optimization of Decision Trees on Racetrack Memory
Christian Hakert, Asif Ali Khan, Kuan-Hsun Chen, Fazal Hameed, Jerónimo Castrillón, Jian-Jia Chen
摘要
Modern distributed low power systems tend to integrate machine learning algorithms, which are directly executed on the distributed devices (on the edge). In resource constrained setups (e.g. battery driven sensor nodes), the execution of the machine learning models has to be optimized for execution time and energy consumption. Racetrack memory (RTM), an emerging non-volatile memory (NVM), promises to achieve these goals by offering unprecedented integration density, smaller access-latency and reduced energy consumption. However, in order to access data in RTM, it needs to be shifted to the access port first, resulting in latency and energy penalties. In this paper, we propose B.L.O. (Bidirectional Linear Ordering), a novel domain-specific approach for placing decision trees in RTMs. We reduce the total amount of shifts during inference by exploiting the tree structure and estimated access probabilities. We further apply the state-of-the-art methods to place data structures in RTM, without exploiting any domain-specific knowledge, to the decision trees and compare them to B. L.O. We formally prove that the B.L.O. solution has an approximation ratio of 4, i.e., its number of shifts is guaranteed to be at most 4 times the optimal number of shifts for a given decision tree. Throughout the experimental evaluation, we show that for the realistic use case B.L.O. empirically outperforms the state-of-the-art data placement method on average by 54.7% in terms of shifts, 19.2% in terms of runtime and 19.2% in terms of energy consumption.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- PLIN: A Persistent Learned Index for Non-Volatile Memory with High Performance and Instant RecoveryZhou Zhang, Zhaole Chu, Peiquan Jin, Yongping Luo 等VLDB 2023 · 被引用 39 次
- SMART: on simultaneously marching racetracks to improve the performance of racetrack-based main memoryXiangjun Peng, Ming-Chang Yang, Ho Ming Tsui, Chi Ngai Leung 等DAC 2022 · 被引用 3 次
- HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI DevicesSangmin Jeon, Kangju Lee, Kyeongwon Lee, Woojoo LeeDAC 2025 · 被引用 4 次
- StreamPIM: Streaming Matrix Computation in Racetrack MemoryYuda An, Yunxiao Tang, Shushu Yi, Li Peng 等HPCA 2024 · 被引用 8 次
- A digital 3D TCAM accelerator for the inference phase of Random ForestChieh-Lin Tsai, Chun-Feng Wu, Yuan-Hao Chang, Han-Wen Hu 等DAC 2023 · 被引用 6 次
