TRiM: Enhancing Processor-Memory Interfaces with Scalable Tensor Reduction in Memory
Jaehyun Park, Byeongho Kim, Sungmin Yun, Eojin Lee, Minsoo Rhu, Jung Ho Ahn
Abstract
Personalized recommendation systems are gaining significant traction due to their industrial importance. An important building block of recommendation systems consists of the embedding layers, which exhibit a highly memory-intensive characteristic. A fundamental primitive of embedding layers is the embedding vector gathers followed by vector reductions, exhibiting low arithmetic intensity and becoming bottlenecked by the memory throughput. To tackle such a challenge, recent proposals employ a near-data processing (NDP) solution at the DRAM rank-level, achieving impressive performance speedups. We observe that prior rank-level-parallelism-based NDP solutions leave significant performance potential on the table as they do not fully reap the abundant transfer throughput inherent in DRAM datapaths.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5731f2df-4caf-49b3-b34b-353ddd498954Cited by top-tier papers22
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi et al.ASPLOS 2024 · 121 citations
- Pathfinding Future PIM Architectures by Demystifying a Commercial PIM TechnologyBongjoon Hyun, Taehun Kim, Dongjae Lee, Minsoo RhuHPCA 2024 · 62 citations
- GROW: A Row-Stationary Sparse-Dense GEMM Accelerator for Memory-Efficient Graph Convolutional Neural NetworksRanggi Hwang, Minhoo Kang, Jiwon Lee, Dongyun Kam et al.HPCA 2023 · 60 citations
- SmartSAGE: training large-scale graph neural networks using in-storage processing architecturesYunjae Lee, Jinha Chung, Minsoo RhuISCA 2022 · 57 citations
- PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model InferenceYufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang et al.ASPLOS 2025 · 44 citations
Related papers
- Accelerating Personalized Recommendation with Cross-level Near-Memory ProcessingHaifeng Liu, Long Zheng, Yu Huang, Chaoqiang Liu et al.ISCA 2023 · 30 citations
- SPACE: Locality-Aware Processing in Heterogeneous Memory for Personalized RecommendationsHongju Kal, Seokmin Lee, Gun Ko, Won Woo RoISCA 2021 · 43 citations
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks et al.ISCA 2020 · 235 citations
- Towards Redundancy-Free Recommendation Model Training via Reusable-aware Near-Memory ProcessingHaifeng Liu, Long Zheng, Yu Huang, Haoyan Huang et al.DAC 2024
- Enabling Efficient Large Recommendation Model Training with Near CXL Memory ProcessingHaifeng Liu, Long Zheng, Yu Huang, Jingyi Zhou et al.ISCA 2024 · 24 citations
