NeuRO: Inference-time Profiling and Orchestration of ML Applications at the Edge
Arshad Javeed, György Dán, Viktoria Fodor
摘要
We address the problem of orchestrating machine learning (ML) workloads, such as latency-sensitive mobile applications and intelligent radio access network (RAN) control functions, in a heterogeneous edge cloud infrastructure, subject to inference time constraints. We introduce NeuRO, a modular profiling framework that learns to approximate ML task inference time summary statistics using data collected from edge deployments. NeuRO leverages two lightweight neural embeddings: one capturing the resource footprint of an individual ML task, and one capturing the aggregate server context induced by co-located workloads. We show the potential of NeuRO by formulating a performance-aware service placement problem subject to inference time summary statistic constraints. To solve this problem efficiently at scale, we develop a heuristic algorithm that incrementally places applications across servers in a sequential manner. For each server, the constrained optimization subproblem is solved using the barrier method, ensuring tractability despite the non-linear neural network (NN) constraints. We evaluate NeuRO on heterogeneous Kubernetes clusters using realistic ML/AI workloads. Experimental results show that NeuRO not only achieves substantially higher profiling accuracy compared to existing methods, but also improves operator revenue by up to 30%, all while satisfying diverse service-level agreements (SLAs).
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Tailored Learning-Based Scheduling for Kubernetes-Oriented Edge-Cloud SystemYiwen Han, Shihao Shen, Xiaofei Wang, Shiqiang Wang 等INFOCOM 2021 · 被引用 93 次
- ScalO-RAN: Energy-aware Network Intelligence Scaling in Open RANStefano Maxenti, Salvatore D'Oro, Leonardo Bonati, Michele Polese 等INFOCOM 2024 · 被引用 17 次
- A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via FaroBeomyeol Jeon, Chen Wang, Diana Arroyo, Alaa Youssef 等EuroSys 2025
- OrchestRAN: Network Automation through Orchestrated Intelligence in the Open RANSalvatore D'Oro, Leonardo Bonati, Michele Polese, Tommaso MelodiaINFOCOM 2022 · 被引用 114 次
- Kalmia: A Heterogeneous QoS-aware Scheduling Framework for DNN Tasks on Edge ServersZiyan Fu, Ju Ren, Deyu Zhang, Yuezhi Zhou 等INFOCOM 2022 · 被引用 28 次
