CAESAR: Coherence-Aided Elective and Seamless Alternative Routing via on-chip FPGA
Shahin Roozkhosh, Denis Hoornaert, Renato Mancuso
摘要
Prompted by the ever-growing demand for high-performance System-on-Chip (SoC) and the plateauing of CPU frequencies, the SoC design landscape is shifting. In a quest to offer programmable specialization, the adoption of tightly-coupled FPGAs co-located with traditional compute clusters has been embraced by major vendors. Thisarchitectural paradigm opens the door to novel hardware/software co-design opportunities. The key principle is that CPU-originated memory traffic can be re-routed through the FPGA for analysis and management purposes. Albeit promising, the side-effect of this approach is that time-critical operations—such as cache-line refills—are fulfilled by moving data over slower interconnects meant for I/O traffic. In this article, we introduce a novel principle named Cache Coherence Backstabbing to precisely tackle these shortcomings. The technique leverages the ability to include the FGPA in the same coherence domain as the core processing elements. Importantly, this enables Coherence-Aided Elective and Seamless Alternative Routing (CAESAR), i.e., seamless inspection and routing of memory transactions, especially cache-line refills, through the FPGA. CAESAR allows the definition of new memory programming paradigms. We discuss the intrinsic potentials of the approach and evaluate it with a full-stack prototype implementation on a commercial platform. Our experiments show an improvement of up to 29% in read bandwidth, 23% in latency, and 13% in pragmatic workloads over the state of the art. Furthermore, we showcase the first in-coherence-domain run-time profiler design as a use-case of the CAESAR approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Livia: Data-Centric Computing Throughout the Memory HierarchyElliot Lockerman, Axel Feldmann, Mohammad Bakhshalipour, Alexandru Stanescu 等ASPLOS 2020 · 被引用 55 次
- Optimus Prime: Accelerating Data Transformation in ServersArash Pourhabibi Zarandi, Siddharth Gupta, Hussein Kassir, Mark Sutherland 等ASPLOS 2020 · 被引用 43 次
- Enzian: an open, general, CPU/FPGA platform for systems software researchDavid A. Cock, Abishek Ramdas, Daniel Schwyn, Michael Giardino 等ASPLOS 2022 · 被引用 42 次
- E-WarP: A System-wide Framework for Memory Bandwidth Profiling and ManagementParul Sohal, Rohan Tabish, Ulrich Drepper, Renato MancusoRTSS 2020 · 被引用 34 次
- Stream Floating: Enabling Proactive and Decentralized Cache OptimizationsZhengrong Wang, Jian Weng, Jason Lowe-Power, Jayesh Gaur 等HPCA 2021 · 被引用 27 次
相关 Paper
- Large-Scale Graph Processing on FPGAs with Caches for Thousands of Simultaneous MissesMikhail Asiatici, Paolo IenneISCA 2021 · 被引用 28 次
- Cohmeleon: Learning-Based Orchestration of Accelerator Coherence in Heterogeneous SoCsJoseph Zuckerman, Davide Giri, Jihye Kwon, Paolo Mantovani 等MICRO 2021 · 被引用 24 次
- Efficient Remote Memory Ordering for Non-Coherent SystemsWei Siew Liew, Md Ashfaqur Rahaman, Adarsh Patil, Ryan Stutsman 等ASPLOS 2026
- FPGA for Aggregate Processing: The Good, The Bad, and The UglyZubeyr F. Eryilmaz, Aarati Kakaraparthy, Jignesh M. Patel, Rathijit Sen 等ICDE 2021 · 被引用 12 次
- A Framework for Optimizing CPU-iGPU Communication on Embedded PlatformsFrancesco Lumpp, Hiren D. Patel, Nicola BombieriDAC 2021 · 被引用 5 次
