A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable Processors
Reese Kuper, Ipoom Jeong, Yifan Yuan, Ren Wang, Narayan Ranganathan, Nikhil Rao, Jiayu Hu, Sanjay Kumar, Philip Lantz, Nam Sung Kim
Abstract
As semiconductor power density is no longer constant with the technology process scaling down, we need different solutions if we are to continue scaling application performance. To this end, modern CPUs are integrating capable data accelerators on the chip, aiming to improve performance and efficiency for a wide range of applications and usages. One such accelerator is the Intel® Data Streaming Accelerator (DSA) introduced since Intel® 4th Generation Xeon® Scalable CPUs (Sapphire Rapids). DSA targets data movement operations in memory that are common sources of overhead in datacenter workloads and infrastructure. In addition, it supports a wider range of operations on streaming data, such as CRC32 calculations, computation of deltas between data buffers, and data integrity field (DIF) operations. This paper aims to introduce the latest features supported by DSA, dive deep into its versatility, and analyze its throughput benefits through a comprehensive evaluation with both microbenchmarks and real use cases. Along with the analysis of its characteristics and the rich software ecosystem of DSA, we summarize several insights and guidelines for the programmer to make the most out of DSA, and use an in-depth case study of DPDK Vhost to demonstrate how these guidelines benefit a real application.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 117969e0-3856-43ba-9d0f-cd7c78afd5c2Cited by top-tier papers10
- HotTiles: Accelerating SpMM with Heterogeneous Accelerator ArchitecturesGerasimos Gerogiannis, Sriram Aananthakrishnan, Josep Torrellas, Ibrahim HurHPCA 2024 · 19 citations
- PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM DevicesSi Ung Noh, Junguk Hong, Chaemin Lim, Seongyeon Park et al.ISCA 2024 · 12 citations
- Extended User Interrupts (xUI): Fast and Flexible Notification without PollingBerk Aydogmus, Linsong Guo, Danial Zuberi, Tal Garfinkel et al.ASPLOS 2025 · 8 citations
- Para-ksm: Parallelized Memory Deduplication with Data Streaming AcceleratorHouxiang Ji, Minho Kim, Seonmu Oh, Daehoon Kim et al.USENIX ATC 2025 · 7 citations
- RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICsJie Zhang, Hongjing Huang, Xuzheng Chen, Xiang Li et al.HPCA 2025 · 6 citations
Builds on6
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof et al.OSDI 2020 · 145 citations
- HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVMAmanda Raybuck, Tim Stamler, Wei Zhang, Mattan Erez et al.SOSP 2021 · 93 citations
- Assise: Performance and Availability via Client-local NVM in a Distributed File SystemThomas E. Anderson, Marco Canini, Jongyul Kim, Dejan Kostic et al.OSDI 2020 · 71 citations
- Don't Forget the I/O When Allocating Your LLCYifan Yuan, Mohammad Alian, Yipeng Wang, Ren Wang et al.ISCA 2021 · 37 citations
- IDIO: Network-Driven, Inbound Network Data Orchestration on Server ProcessorsMohammad Alian, Siddharth Agarwal, Jongmin Shin, Neel Patel et al.MICRO 2022 · 21 citations
Related papers
- LightDSA: Enabling Efficient DSA Through Hardware-Aware Transparent OptimizationYuansen Wang, Teng Ma, Yuanhui Luo, Dongbiao He et al.EuroSys 2026
- DSAssassin: Cross-VM Side-Channel Attacks by Exploiting Intel Data Streaming AcceleratorBen Chen, Kunlin Li, Shuwen Deng, Dongsheng Wang et al.HPCA 2026 · 1 citation
- DarkStream: Exploiting Internal Throughput Contention in Data Streaming Accelerator for Timing AttacksHyosang Kim, Ki-Dong Kang, Gyeongseo Park, Sungju Kim et al.ISCA 2026
- Data Motion Acceleration: Chaining Cross-Domain Multi AcceleratorsShu-Ting Wang, Hanyang Xu, Amin Mamandipoor, Rohan Mahapatra et al.HPCA 2024 · 10 citations
- Effectively Scheduling Computational Graphs of Deep Neural Networks toward Their Domain-Specific AcceleratorsJie Zhao, Siyuan Feng, Xiaoqiang Dan, Fei Liu et al.OSDI 2023 · 9 citations
