Mitigating CPU Frontend for Complex Data Plane Applications
Yihan Dang, Hao Li, Ze Xia, Jiajun Luan, Peng Zhang
摘要
The functionality and requirement of modern networks are becoming increasingly complex, giving rise to Complex Data Plane Applications (CDPA) with rich semantics but often limited performance. However, many existing optimizations fail to improve the performance of CDPAs. This is because CDPAs usually come with excessively large code size, which is often two orders of magnitude larger than today's L1 instruction cache (I-cache) size, causing the CPU to frequently stall on accessing instructions, thus presenting a distinct performance profile that bounds on CPU frontend. This paper proposes NanoPL, an I-cache-friendly new execution model that shuffles the packet processing logic to efficiently mitigate CPU frontend for CDPAs. Stemming from the common execution pattern of CDPAs, NanoPL analyzes its code to ensure semantic consistency after shuffling. By collecting performance profile of CDPAs over underlying traffic, NanoPL partitions CDPAs into execution stages and conducts I-cachefriendly shuffling policy. Experiments show that NanoPL can achieve 17.2% 30.2% higher throughput over real world CD-PAs due to the reduction of I-cache misses by up to 86.4%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Propeller: A Profile Guided, Relinking Optimizer for Warehouse-Scale ApplicationsHan Shen, Krzysztof Pszeniczny, Rahman Lavaee, Snehasish Kumar 等ASPLOS 2023 · 被引用 37 次
- Twig: Profile-Guided BTB Prefetching for Data Center ApplicationsTanvir Ahmed Khan, Nathan Brown, Akshitha Sriraman, Niranjan K. Soundararajan 等MICRO 2021 · 被引用 33 次
- Batchy: Batch-scheduling Data Flow Graphs with Service-level ObjectivesTamás Lévai, Felicián Németh, Barath Raghavan, Gábor RétváriNSDI 2020 · 被引用 30 次
- Domain specific run time optimization for software data planesSebastiano Miano, Alireza Sanaee, Fulvio Risso, Gábor Rétvári 等ASPLOS 2022 · 被引用 20 次
- Compacting points-to sets through object clusteringMohamad Barbar, Yulei SuiOOPSLA 2021 · 被引用 15 次
相关 Paper
- hXDP: Efficient Software Packet Processing on FPGA NICsMarco Spaziani Brunella, Giacomo Belocchi, Marco Bonola, Salvatore Pontarelli 等OSDI 2020 · 被引用 25 次
- PacketMill: toward per-Core 100-Gbps networkingAlireza Farshin, Tom Barbette, Amir Roozbeh, Gerald Q. Maguire Jr. 等ASPLOS 2021 · 被引用 46 次
- NetCL: A Unified Programming Framework for In-Network ComputingGeorge Karlos, Henri E. Bal, Lin WangSC 2024 · 被引用 3 次
- P4LRU: Towards An LRU Cache Entirely in Programmable Data PlaneYikai Zhao, Wenrui Liu, Fenghao Dong, Tong Yang 等SIGCOMM 2023 · 被引用 18 次
- Hoda: a High-performance Open vSwitch Dataplane with Multiple Specialized Data PathsHeng Pan, Peng He, Zhenyu Li, Pan Zhang 等EuroSys 2024 · 被引用 6 次
