CAESAR: Coherence-Aided Elective and Seamless Alternative Routing via on-chip FPGA
Shahin Roozkhosh, Denis Hoornaert, Renato Mancuso
Abstract
Prompted by the ever-growing demand for high-performance System-on-Chip (SoC) and the plateauing of CPU frequencies, the SoC design landscape is shifting. In a quest to offer programmable specialization, the adoption of tightly-coupled FPGAs co-located with traditional compute clusters has been embraced by major vendors. Thisarchitectural paradigm opens the door to novel hardware/software co-design opportunities. The key principle is that CPU-originated memory traffic can be re-routed through the FPGA for analysis and management purposes. Albeit promising, the side-effect of this approach is that time-critical operations—such as cache-line refills—are fulfilled by moving data over slower interconnects meant for I/O traffic. In this article, we introduce a novel principle named Cache Coherence Backstabbing to precisely tackle these shortcomings. The technique leverages the ability to include the FGPA in the same coherence domain as the core processing elements. Importantly, this enables Coherence-Aided Elective and Seamless Alternative Routing (CAESAR), i.e., seamless inspection and routing of memory transactions, especially cache-line refills, through the FPGA. CAESAR allows the definition of new memory programming paradigms. We discuss the intrinsic potentials of the approach and evaluate it with a full-stack prototype implementation on a commercial platform. Our experiments show an improvement of up to 29% in read bandwidth, 23% in latency, and 13% in pragmatic workloads over the state of the art. Furthermore, we showcase the first in-coherence-domain run-time profiler design as a use-case of the CAESAR approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 585cc9ff-1f00-4d9e-b7b1-a3a0e12280cbCited by top-tier papers1
Ask how each one uses itBuilds on7
- Livia: Data-Centric Computing Throughout the Memory HierarchyElliot Lockerman, Axel Feldmann, Mohammad Bakhshalipour, Alexandru Stanescu et al.ASPLOS 2020 · 55 citations
- Optimus Prime: Accelerating Data Transformation in ServersArash Pourhabibi Zarandi, Siddharth Gupta, Hussein Kassir, Mark Sutherland et al.ASPLOS 2020 · 43 citations
- Enzian: an open, general, CPU/FPGA platform for systems software researchDavid A. Cock, Abishek Ramdas, Daniel Schwyn, Michael Giardino et al.ASPLOS 2022 · 42 citations
- E-WarP: A System-wide Framework for Memory Bandwidth Profiling and ManagementParul Sohal, Rohan Tabish, Ulrich Drepper, Renato MancusoRTSS 2020 · 34 citations
- Stream Floating: Enabling Proactive and Decentralized Cache OptimizationsZhengrong Wang, Jian Weng, Jason Lowe-Power, Jayesh Gaur et al.HPCA 2021 · 27 citations
Related papers
- Large-Scale Graph Processing on FPGAs with Caches for Thousands of Simultaneous MissesMikhail Asiatici, Paolo IenneISCA 2021 · 28 citations
- Cohmeleon: Learning-Based Orchestration of Accelerator Coherence in Heterogeneous SoCsJoseph Zuckerman, Davide Giri, Jihye Kwon, Paolo Mantovani et al.MICRO 2021 · 24 citations
- Efficient Remote Memory Ordering for Non-Coherent SystemsWei Siew Liew, Md Ashfaqur Rahaman, Adarsh Patil, Ryan Stutsman et al.ASPLOS 2026
- FPGA for Aggregate Processing: The Good, The Bad, and The UglyZubeyr F. Eryilmaz, Aarati Kakaraparthy, Jignesh M. Patel, Rathijit Sen et al.ICDE 2021 · 12 citations
- A Framework for Optimizing CPU-iGPU Communication on Embedded PlatformsFrancesco Lumpp, Hiren D. Patel, Nicola BombieriDAC 2021 · 5 citations
