Stream Floating: Enabling Proactive and Decentralized Cache Optimizations
Zhengrong Wang, Jian Weng, Jason Lowe-Power, Jayesh Gaur, Tony Nowatzki
摘要
As multicore systems continue to grow in scale and on-chip memory capacity, the on-chip network bandwidth and latency become problematic bottlenecks. Because of this, overheads in data transfer, the coherence protocol and replacement policies become increasingly important. Unfortunately, even in well-structured programs, many natural optimizations are difficult to implement because of the reactive and centralized nature of traditional cache hierarchies, where all requests are initiated by the core for short, cache line granularity accesses. For example, long-lasting access patterns could be streamed from shared caches without requests from the core. Indirect memory access can be performed by chaining requests made from within the cache, rather than constantly returning to the core.
Our primary insight is that if programs can embed information about long-term memory stream behavior in their ISAs, then these streams can be floated to the appropriate level of the memory hierarchy. This decentralized approach to address generation and cache requests can lead to better cache policies and lower request and data traffic by proactively sending data before the cores even request it.
To evaluate the opportunities of stream floating, we enhance a tiled multicore cache hierarchy with stream engines to process stream requests in last-level cache banks. We develop several novel optimizations that are facilitated by stream exposure in the ISA, and subsequent exposure to caches. We evaluate using a cycle-level execution-driven gem5-based simulator, using 10 data-processing workloads from Rodinia and 2 streaming kernels written in OpenMP. We find that stream floating enables 52% and 39% speedup over an inorder and OOO core with state of art prefetcher design respectively, with 64% and 49% energy efficiency advantage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- PolyGraph: Exposing the Value of Flexibility for Graph Processing AcceleratorsVidushi Dadu, Sihao Liu, Tony NowatzkiISCA 2021 · 被引用 60 次
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo 等MICRO 2022 · 被引用 37 次
- OverGen: Improving FPGA Usability through Domain-specific Overlay GenerationSihao Liu, Jian Weng, Dylan Kupsh, Atefeh Sohrabizadeh 等MICRO 2022 · 被引用 32 次
- Near-Stream Computing: General and Transparent Near-Cache AccelerationZhengrong Wang, Jian Weng, Sihao Liu, Tony NowatzkiHPCA 2022 · 被引用 24 次
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 被引用 23 次
它引用的顶会 Paper4
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang 等ISCA 2020 · 被引用 140 次
- A Hybrid Systolic-Dataflow Architecture for Inductive Matrix AlgorithmsJian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu 等HPCA 2020 · 被引用 80 次
- Livia: Data-Centric Computing Throughout the Memory HierarchyElliot Lockerman, Axel Feldmann, Mohammad Bakhshalipour, Alexandru Stanescu 等ASPLOS 2020 · 被引用 55 次
- SnackNoC: Processing in the Communication LayerKarthik Sangaiah, Michael Lui, Ragh Kuttappa, Baris Taskin 等HPCA 2020 · 被引用 13 次
相关 Paper
- DX100: Programmable Data Access Accelerator for IndirectionAlireza Khadem, Kamalavasan Kamalakkannan, Zhenyan Zhu, Akash Poptani 等ISCA 2025 · 被引用 2 次
- Tiny but mighty: designing and realizing scalable latency tolerance for manycore SoCsMarcelo Orenes-Vera, Aninda Manocha, Jonathan Balkind, Fei Gao 等ISCA 2022 · 被引用 24 次
- Decoupled Vector RunaheadAjeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M. Jones 等MICRO 2023 · 被引用 15 次
- Low-Latency Stateful Stream Processing Through Timely and Accurate PrefetchingEleni Zapridou, Anastasia AilamakiICDE 2026
- Vector RunaheadAjeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven EeckhoutISCA 2021 · 被引用 27 次
