A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable Processors
Reese Kuper, Ipoom Jeong, Yifan Yuan, Ren Wang, Narayan Ranganathan, Nikhil Rao, Jiayu Hu, Sanjay Kumar, Philip Lantz, Nam Sung Kim
摘要
As semiconductor power density is no longer constant with the technology process scaling down, we need different solutions if we are to continue scaling application performance. To this end, modern CPUs are integrating capable data accelerators on the chip, aiming to improve performance and efficiency for a wide range of applications and usages. One such accelerator is the Intel® Data Streaming Accelerator (DSA) introduced since Intel® 4th Generation Xeon® Scalable CPUs (Sapphire Rapids). DSA targets data movement operations in memory that are common sources of overhead in datacenter workloads and infrastructure. In addition, it supports a wider range of operations on streaming data, such as CRC32 calculations, computation of deltas between data buffers, and data integrity field (DIF) operations. This paper aims to introduce the latest features supported by DSA, dive deep into its versatility, and analyze its throughput benefits through a comprehensive evaluation with both microbenchmarks and real use cases. Along with the analysis of its characteristics and the rich software ecosystem of DSA, we summarize several insights and guidelines for the programmer to make the most out of DSA, and use an in-depth case study of DPDK Vhost to demonstrate how these guidelines benefit a real application.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- HotTiles: Accelerating SpMM with Heterogeneous Accelerator ArchitecturesGerasimos Gerogiannis, Sriram Aananthakrishnan, Josep Torrellas, Ibrahim HurHPCA 2024 · 被引用 19 次
- PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM DevicesSi Ung Noh, Junguk Hong, Chaemin Lim, Seongyeon Park 等ISCA 2024 · 被引用 12 次
- Extended User Interrupts (xUI): Fast and Flexible Notification without PollingBerk Aydogmus, Linsong Guo, Danial Zuberi, Tal Garfinkel 等ASPLOS 2025 · 被引用 8 次
- Para-ksm: Parallelized Memory Deduplication with Data Streaming AcceleratorHouxiang Ji, Minho Kim, Seonmu Oh, Daehoon Kim 等USENIX ATC 2025 · 被引用 7 次
- RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICsJie Zhang, Hongjing Huang, Xuzheng Chen, Xiang Li 等HPCA 2025 · 被引用 6 次
它引用的顶会 Paper6
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof 等OSDI 2020 · 被引用 145 次
- HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVMAmanda Raybuck, Tim Stamler, Wei Zhang, Mattan Erez 等SOSP 2021 · 被引用 93 次
- Assise: Performance and Availability via Client-local NVM in a Distributed File SystemThomas E. Anderson, Marco Canini, Jongyul Kim, Dejan Kostic 等OSDI 2020 · 被引用 71 次
- Don't Forget the I/O When Allocating Your LLCYifan Yuan, Mohammad Alian, Yipeng Wang, Ren Wang 等ISCA 2021 · 被引用 37 次
- IDIO: Network-Driven, Inbound Network Data Orchestration on Server ProcessorsMohammad Alian, Siddharth Agarwal, Jongmin Shin, Neel Patel 等MICRO 2022 · 被引用 21 次
相关 Paper
- LightDSA: Enabling Efficient DSA Through Hardware-Aware Transparent OptimizationYuansen Wang, Teng Ma, Yuanhui Luo, Dongbiao He 等EuroSys 2026
- DSAssassin: Cross-VM Side-Channel Attacks by Exploiting Intel Data Streaming AcceleratorBen Chen, Kunlin Li, Shuwen Deng, Dongsheng Wang 等HPCA 2026 · 被引用 1 次
- DarkStream: Exploiting Internal Throughput Contention in Data Streaming Accelerator for Timing AttacksHyosang Kim, Ki-Dong Kang, Gyeongseo Park, Sungju Kim 等ISCA 2026
- Data Motion Acceleration: Chaining Cross-Domain Multi AcceleratorsShu-Ting Wang, Hanyang Xu, Amin Mamandipoor, Rohan Mahapatra 等HPCA 2024 · 被引用 10 次
- Effectively Scheduling Computational Graphs of Deep Neural Networks toward Their Domain-Specific AcceleratorsJie Zhao, Siyuan Feng, Xiaoqiang Dan, Fei Liu 等OSDI 2023 · 被引用 9 次
