DUNE: Distributed Inference in the User Plane
Beyza Bütün, David De Andres Hernandez, Michele Gucciardo, Marco Fiore
摘要
The deployment of Machine Learning (ML) models in the user plane enables line-rate in-network inference, significantly reducing latency and improving the scalability of functions like traffic monitoring. Yet, integrating ML models into programmable network devices requires meeting stringent constraints in terms of memory resources and computing capabilities. Previous solutions have focused on implementing monolithic ML models within individual programmable network devices, which are limited by hardware constraints, especially while executing challenging classification use cases. In this paper, we propose DUNE, a novel framework that realizes for the first time a user plane inference that is distributed across the multiple devices that compose the programmable network. DUNE adopts fully automated approaches to () breaking large ML models into simpler sub-models that preserve inference accuracy while minimizing resource usage, (ii) designing the sub-models and their sequencing so as to enable an efficient distributed execution of joint packet- and flow-level inference. We implement DUNE using P4, deploy it in an experimental network with multiple industry-grade programmable switches, and run tests with real-world traffic measurements for two complex classification use cases. Our results demonstrate that DUNE not only reduces perswitch resource utilization with respect to legacy monolithic ML designs but also improves their inference accuracy by up to 7.5 %.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Re-architecting Traffic Analysis with Neural Network Interface CardsGiuseppe Siracusano, Salvator Galea, Davide Sanvito, Mohammad Malekzadeh 等NSDI 2022 · 被引用 99 次
- Flightplan: Dataplane Disaggregation and Placement for P4 ProgramsNik Sultana, John Sonchack, Hans Giesen, Isaac Pedisich 等NSDI 2021 · 被引用 95 次
- Taurus: a data plane architecture for per-packet MLTushar Swamy, Alexander Rucker, Muhammad Shahbaz, Ishan Gaur 等ASPLOS 2022 · 被引用 94 次
- Programmable Switches for in-Networking ClassificationBruno Missi Xavier, Rafael Silva Guimarães, Giovanni Comarela, Magnos MartinelloINFOCOM 2021 · 被引用 79 次
- Flowrest: Practical Flow-Level Inference in Programmable Switches with Random ForestsAristide Tanyi-Jong Akem, Michele Gucciardo, Marco FioreINFOCOM 2023 · 被引用 63 次
相关 Paper
- Jewel: Resource-Efficient Joint Packet and Flow Level Inference in Programmable SwitchesAristide Tanyi-Jong Akem, Beyza Bütün, Michele Gucciardo, Marco FioreINFOCOM 2024 · 被引用 27 次
- Monic: In-Network Mixture-of-Experts Inference on Programmable Data PlanesXiaoquan Zhang, Bowen Liang, Fung Po Tso, Yuhui Deng 等INFOCOM 2026
- FENIX: Enabling In-Network DNN Inference with FPGA-Enhanced Programmable SwitchesXiangyu Gao, Tong Li, Yinchao Zhang, Ziqiang Wang 等NSDI 2026 · 被引用 12 次
- Leo: Online ML-based Traffic Classification at Multi-Terabit Line RateSyed Usman Jafri, Sanjay G. Rao, Vishal Shrivastav, Mohit TawarmalaniNSDI 2024 · 被引用 46 次
- SPLIDT: Partitioned Decision Trees for Scalable Stateful Inference at Line RateMurayyiam Parvez, Annus Zulfiqar, Roman Beltiukov, Shir Landau Feibish 等NSDI 2026 · 被引用 1 次
