MERCI: efficient embedding reduction on commodity hardware via sub-query memoization
Yejin Lee, Seong Hoon Seo, Hyunji Choi, Hyoung Uk Sul, Soosung Kim, Jae W. Lee, Tae Jun Ham
摘要
Deep neural networks (DNNs) with embedding layers are widely adopted to capture complex relationships among entities within a dataset. Embedding layers aggregate multiple embeddings — a dense vector used to represent the complicated nature of a data feature— into a single embedding; such operation is called embedding reduction. Embedding reduction spends a significant portion of its runtime on reading embeddings from memory and thus is known to be heavily memory-bandwidth-bound. Recent works attempt to accelerate this critical operation, but they often require either hardware modifications or emerging memory technologies, which makes it hardly deployable on commodity hardware. Thus, we propose MERCI, Memoization for Embedding Reduction with ClusterIng, a novel memoization framework for efficient embedding reduction. MERCI provides a mechanism for memoizing partial aggregation of correlated embeddings and retrieving the memoized partial result at a low cost. MERCI substantially reduces the number of memory accesses by 44% (29%), leading to 102% (74%) throughput improvement on real machines and 40.2% (28.6%) energy savings at the expense of 8×(1×) additional memory usage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Demystifying CXL Memory with Genuine CXL-Ready Systems and DevicesYan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper 等MICRO 2023 · 被引用 133 次
- Fleche: an efficient GPU embedding cache for personalized recommendationsMinhui Xie, Youyou Lu, Jiazhen Lin, Qing Wang 等EuroSys 2022 · 被引用 24 次
- Training personalized recommendation systems from (GPU) scratch: look forward not backwardsYoungeun Kwon, Minsoo RhuISCA 2022 · 被引用 24 次
- PreSto: An In-Storage Data Preprocessing System for Training Recommendation ModelsYunjae Lee, Hyeseong Kim, Minsoo RhuISCA 2024 · 被引用 8 次
- Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUsRishabh Jain, Vivek M. Bhasi, Adwait Jog, Anand Sivasubramaniam 等MICRO 2024 · 被引用 5 次
它引用的顶会 Paper2
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks 等ISCA 2020 · 被引用 235 次
- Centaur: A Chiplet-based, Hybrid Sparse-Dense Accelerator for Personalized RecommendationsRanggi Hwang, Taehun Kim, Youngeun Kwon, Minsoo RhuISCA 2020 · 被引用 94 次
相关 Paper
- GRACE: A Scalable Graph-Based Approach to Accelerating Recommendation Model InferenceHaojie Ye, Sanketh Vedula, Yuhan Chen, Yichen Yang 等ASPLOS 2023 · 被引用 16 次
- Towards Redundancy-Free Recommendation Model Training via Reusable-aware Near-Memory ProcessingHaifeng Liu, Long Zheng, Yu Huang, Haoyan Huang 等DAC 2024
- FEC: Efficient Deep Recommendation Model Training with Flexible Embedding CommunicationKaihao Ma, Xiao Yan, Zhenkun Cai, Yuzhen Huang 等SIGMOD 2023 · 被引用 8 次
- Accelerating Personalized Recommendation with Cross-level Near-Memory ProcessingHaifeng Liu, Long Zheng, Yu Huang, Chaoqiang Liu 等ISCA 2023 · 被引用 30 次
- Deep Learning Acceleration with Neuron-to-Memory TransformationMohsen Imani, Mohammad Samragh Razlighi, Yeseong Kim, Saransh Gupta 等HPCA 2020 · 被引用 31 次
