Monic: In-Network Mixture-of-Experts Inference on Programmable Data Planes
Xiaoquan Zhang, Bowen Liang, Fung Po Tso, Yuhui Deng, Zhen Zhang, Kaimin Wei, Weijia Jia, Lin Cui
摘要
In-network inference has emerged as a promising paradigm for enabling intelligent packet processing at the line rate within programmable data planes. However, it is fundamentally limited by an inherent conflict between model accuracy and the resource constraints of programmable data planes. This forces a compromise: monolithic deployments can achieve high accuracy but are constrained by the resource limits of a single switch, while distributed approaches leverage the combined resources of multiple switches but are limited in accuracy due to a lack of model coordination. We resolve this trade-off by proposing Monic, a framework that enables multiple "expert" submodels to perform collaborative inference. Inspired by the Mixture-of-Experts (MoE) paradigm, Monic uses a pipeline-compatible gating mechanism to selectively activate experts across the network. We enable this in practice through a resourceaware mapping and co-optimization strategy that automatically identifies optimal configurations under hardware constraints. We have implemented Monic using P4 hardware switches with Intel Tofino ASIC. Our evaluation shows that Monic achieves a 17.4% relative improvement over baseline methods and maintains a 29.58% Macro F1 advantage under scalability evaluation, demonstrating that a collaborative approach can simultaneously achieve resource efficiency and superior accuracy.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- DUNE: Distributed Inference in the User PlaneBeyza Bütün, David De Andres Hernandez, Michele Gucciardo, Marco FioreINFOCOM 2025 · 被引用 7 次
- Flowrest: Practical Flow-Level Inference in Programmable Switches with Random ForestsAristide Tanyi-Jong Akem, Michele Gucciardo, Marco FioreINFOCOM 2023 · 被引用 63 次
- Carlo: Cross-Plane Collaboration for Multiple In-network Computing ApplicationsXiaoquan Zhang, Lin Cui, Waiming Lau, Fung Po Tso 等INFOCOM 2024 · 被引用 1 次
- Quark: Implementing Convolutional Neural Networks Entirely on Programmable Data PlaneMai Zhang, Lin Cui, Xiaoquan Zhang, Fung Po Tso 等INFOCOM 2025 · 被引用 17 次
- Jewel: Resource-Efficient Joint Packet and Flow Level Inference in Programmable SwitchesAristide Tanyi-Jong Akem, Beyza Bütün, Michele Gucciardo, Marco FioreINFOCOM 2024 · 被引用 27 次
