Dalorex: A Data-Local Program Execution and Architecture for Memory-bound Applications
Marcelo Orenes-Vera, Esin Tureci, David Wentzlaff, Margaret Martonosi
摘要
Applications with low data reuse and frequent irregular memory accesses, such as graph or sparse linear algebra workloads, fail to scale well due to memory bottlenecks and poor core utilization. While prior work with prefetching, decoupling, or pipelining can mitigate memory latency and improve core utilization, memory bottlenecks persist due to limited off-chip bandwidth. Approaches doing processing in-memory (PIM) with Hybrid Memory Cube (HMC) overcome bandwidth limitations but fail to achieve high core utilization due to poor task scheduling and synchronization overheads. Moreover, the high memory-per-core ratio available with HMC limits strong scaling.
We introduce Dalorex, a hardware-software co-design that achieves high parallelism and energy efficiency, demonstrating strong scaling with >16,000 cores when processing graph and sparse linear algebra workloads. Over the prior work in PIM, both using 256 cores, Dalorex improves performance and energy consumption by two orders of magnitude through (1) a tilebased distributed-memory architecture where each processing tile holds an equal amount of data, and all memory operations are local; (2) a task-based parallel programming model where tasks are executed by the processing unit that is co-located with the target data; (3) a network design optimized for irregular traffic, where all communication is one-way, and messages do not contain routing metadata; (4) novel traffic-aware task scheduling hardware that maintains high core utilization; and (5) a dataplacement strategy that improves work balance.
This work proposes architectural and software innovations to provide the greatest scalability to date for running graph algorithms while still being programmable for other domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- NDPBridge: Enabling Cross-Bank Coordination in Near-DRAM-Bank Processing ArchitecturesBoyu Tian, Yiwei Li, Li Jiang, Shuangyu Cai 等ISCA 2024 · 被引用 27 次
- MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC JobsFrancesco Antici, Andrea Bartolini, Zeynep Kiziltan, Özalp Babaoglu 等SC 2024 · 被引用 10 次
- Azul: An Accelerator for Sparse Iterative Solvers Leveraging Distributed On-Chip MemoryAxel Feldmann, Courtney Golden, Yifan Yang, Joel S. Emer 等MICRO 2024 · 被引用 7 次
- PUSHtap: PIM-based In-Memory HTAP with Unified Data Storage FormatYilong Zhao, Mingyu Gao, Huanchen Zhang, Fangxin Liu 等ASPLOS 2025 · 被引用 4 次
- Evaluating Ruche Networks: Physically Scalable, Cost-Effective, Bandwidth-Flexible NoCsDai Cheol Jung, Michael B. TaylorISCA 2025 · 被引用 3 次
它引用的顶会 Paper8
- GraphPulse: An Event-Driven Hardware Accelerator for Asynchronous Graph ProcessingShafiur Rahman, Nael B. Abu-Ghazaleh, Rajiv GuptaMICRO 2020 · 被引用 67 次
- Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-DesignNishil Talati, Kyle May, Armand Behroozi, Yichen Yang 等HPCA 2021 · 被引用 62 次
- Fifer: Practical Acceleration of Irregular Applications on Reconfigurable ArchitecturesQuan M. Nguyen, Daniel SánchezMICRO 2021 · 被引用 60 次
- PolyGraph: Exposing the Value of Flexibility for Graph Processing AcceleratorsVidushi Dadu, Sihao Liu, Tony NowatzkiISCA 2021 · 被引用 60 次
- Livia: Data-Centric Computing Throughout the Memory HierarchyElliot Lockerman, Axel Feldmann, Mohammad Bakhshalipour, Alexandru Stanescu 等ASPLOS 2020 · 被引用 55 次
相关 Paper
- GraphRing: an HMC-ring based graph processing framework with optimized data movementZerun Li, Xiaoming Chen, Yinhe HanDAC 2022 · 被引用 4 次
- FALA: Locality-Aware PIM-Host Cooperation for Graph Processing with Fine-Grained Column AccessChangmin Shin, Jaeyong Song, Seongmin Na, Jun Sung 等MICRO 2025 · 被引用 5 次
- Piccolo: Large-Scale Graph Processing with Fine-Grained in-Memory Scatter-GatherChangmin Shin, Jaeyong Song, Hongsun Jang, Dogeun Kim 等HPCA 2025 · 被引用 5 次
- Tiny but mighty: designing and realizing scalable latency tolerance for manycore SoCsMarcelo Orenes-Vera, Aninda Manocha, Jonathan Balkind, Fei Gao 等ISCA 2022 · 被引用 24 次
- Large-Scale Graph Processing on FPGAs with Caches for Thousands of Simultaneous MissesMikhail Asiatici, Paolo IenneISCA 2021 · 被引用 28 次
