SnackNoC: Processing in the Communication Layer
Karthik Sangaiah, Michael Lui, Ragh Kuttappa, Baris Taskin, Mark Hempstead
Abstract
In this work, we propose and evaluate a Network-on-Chip (NoC) augmented with light-weight processing elements to provide a lean dataflow-style system. We show that contemporary NoC routers can frequently experience long periods of idletime, with less than 10% link utilization in HPC applications. By repurposing the temporal and spatial slack of the NoC, the proposed platform, SnackNoC, is able to compute linear algebra kernels efficiently within the communication layer with minimal additional resource costs.
SnackNoC 'Snack' application kernels are programmed with a producer-consumer data model that uses the NoC slack to store and transmit intermediate data between processing elements. SnackNoC is demonstrated in a multi-program environment that continually executes linear algebra kernels on the NoC simultaneously with chip multiprocessor (CMP) applications on the processor cores. Linear algebra kernels are computed up to 6.15× faster on SnackNoC compared to an Intel Haswell EP x86 processing core. The cost of executing 'snack' kernels in parallel to the CMP applications is a minimal runtime impact of 0.01% to 0.83% due to higher link utilization, and an uncore area overhead of 1.1%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6df66b02-57e4-4b44-8fdf-097a0af97ee8Cited by top-tier papers7
- Stream Floating: Enabling Proactive and Decentralized Cache OptimizationsZhengrong Wang, Jian Weng, Jason Lowe-Power, Jayesh Gaur et al.HPCA 2021 · 27 citations
- NDPBridge: Enabling Cross-Bank Coordination in Near-DRAM-Bank Processing ArchitecturesBoyu Tian, Yiwei Li, Li Jiang, Shuangyu Cai et al.ISCA 2024 · 27 citations
- Near-Stream Computing: General and Transparent Near-Cache AccelerationZhengrong Wang, Jian Weng, Sihao Liu, Tony NowatzkiHPCA 2022 · 24 citations
- Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory FusionZhengrong Wang, Christopher Liu, Aman Arora, Lizy Kurian John et al.ASPLOS 2023 · 20 citations
- ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure CallsJiechen Zhao, Iris Uwizeyimana, Karthik Ganesan, Mark C. Jeffrey et al.MICRO 2022 · 11 citations
Related papers
- Flumen: Dynamic Processing in the Photonic InterconnectKyle Shiflett, Avinash Karanth, Razvan C. Bunescu, Ahmed LouriISCA 2023 · 10 citations
- A Versatile and Flexible Chiplet-based System Design for Heterogeneous Manycore ArchitecturesHao Zheng, Ke Wang, Ahmed LouriDAC 2020 · 38 citations
- FastTrackNoC: A NoC with FastTrack Router DatapathsAhsen Ejaz, Ioannis SourdisHPCA 2022 · 4 citations
- HAL: Hardware-assisted Load Balancing for Energy-efficient SNIC-Host Cooperative ComputingJinghan Huang, Jiaqi Lou, Srikar Vanavasam, Xinhao Kong et al.ISCA 2024 · 7 citations
- Ring Road: A Scalable Polar-Coordinate-based 2D Network-on-Chip ArchitectureYinxiao Feng, Wei Li, Kaisheng MaMICRO 2024 · 2 citations
