Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-Design
Nishil Talati, Kyle May, Armand Behroozi, Yichen Yang, Kuba Kaszyk, Christos Vasiladiotis, Tarunesh Verma, Lu Li, Brandon Nguyen, Jiawen Sun, John Magnus Morton, Agreen Ahmadi
Abstract
Irregular workloads are typically bottlenecked by the memory system. These workloads often use sparse data representations, e.g., compressed sparse row/column (CSR/CSC), to conserve space at the cost of complicated, irregular traversals. Such traversals access large volumes of data and offer little locality for caches and conventional prefetchers to exploit. This paper presents Prodigy, a low-cost hardware-software codesign solution for intelligent prefetching to improve the memory latency of several important irregular workloads. Prodigy targets irregular workloads including graph analytics, sparse linear algebra, and fluid mechanics that exhibit two specific types of data-dependent memory access patterns. Prodigy adopts a “best of both worlds” approach by using static program information from software, and dynamic run-time information from hardware. The core of the system is the Data Indirection Graph (DIG)-a proposed compact representation used to express program semantics such as the layout and memory access patterns of key data structures. The DIG representation is agnostic to a particular data structure format and is demonstrated to work with several sparse formats including CSR and CSC. Program semantics are automatically captured with a compiler pass, encoded as a DIG, and inserted into the application binary. The DIG is then used to program a low-cost hardware prefetcher to fetch data according to an irregular algorithm's data structure traversal pattern. We equip the prefetcher with a flexible prefetching algorithm that maintains timeliness by dynamically adapting its prefetch distance to an application's execution pace. We evaluate the performance, energy consumption, and transistor cost of Prodigy using a variety of algorithms from the GAP, HPCG, and NAS benchmark suites. We compare the performance of Prodigy against a non-prefetching baseline as well as state-of-the-art prefetchers. We show that by using just 0.8KB of storage, Prodigy outperforms a non-prefetching baseline by and saves energy by , on average. Prodigy also outperforms modern data prefetchers by .
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6eba6c81-29d5-422f-9b61-1f613decd2b1Cited by top-tier papers21
- Crescent: taming memory irregularities for accelerating deep point cloud analyticsYu Feng, Gunnar Hammonds, Yiming Gan, Yuhao ZhuISCA 2022 · 44 citations
- SpZip: Architectural Support for Effective Data Compression In Irregular ApplicationsYifan Yang, Joel S. Emer, Daniel SánchezISCA 2021 · 31 citations
- APT-GET: profile-guided timely software prefetchingSaba Jamilan, Tanvir Ahmed Khan, Grant Ayers, Baris Kasikci et al.EuroSys 2022 · 25 citations
- Tiny but mighty: designing and realizing scalable latency tolerance for manycore SoCsMarcelo Orenes-Vera, Aninda Manocha, Jonathan Balkind, Fei Gao et al.ISCA 2022 · 24 citations
- Near-Stream Computing: General and Transparent Near-Cache AccelerationZhengrong Wang, Jian Weng, Sihao Liu, Tony NowatzkiHPCA 2022 · 24 citations
Related papers
- Differential-Matching Prefetcher for Indirect Memory AccessGelin Fu, Tian Xia, Zhongpei Luo, Ruiyang Chen et al.HPCA 2024 · 15 citations
- RnR: A Software-Assisted Record-and-Replay Hardware PrefetcherChao Zhang, Yuan Zeng, John Shalf, Xiaochen GuoMICRO 2020 · 10 citations
- Magellan: A High-Performance Loop-Guided Prefetcher for Indirect Memory AccessGelin Fu, Tian Xia, Mingzhuo Yin, Prashant J. Nair et al.ISCA 2025 · 2 citations
- Profile-Guided Temporal PrefetchingMengming Li, Qijun Zhang, Yichuan Gao, Wenji Fang et al.ISCA 2025 · 4 citations
- ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory AccessesMengming Li, Chenlu Miao, Buqing Xu, Qijun Zhang et al.ISCA 2026
