NeuRO: Inference-time Profiling and Orchestration of ML Applications at the Edge
Arshad Javeed, György Dán, Viktoria Fodor
Abstract
We address the problem of orchestrating machine learning (ML) workloads, such as latency-sensitive mobile applications and intelligent radio access network (RAN) control functions, in a heterogeneous edge cloud infrastructure, subject to inference time constraints. We introduce NeuRO, a modular profiling framework that learns to approximate ML task inference time summary statistics using data collected from edge deployments. NeuRO leverages two lightweight neural embeddings: one capturing the resource footprint of an individual ML task, and one capturing the aggregate server context induced by co-located workloads. We show the potential of NeuRO by formulating a performance-aware service placement problem subject to inference time summary statistic constraints. To solve this problem efficiently at scale, we develop a heuristic algorithm that incrementally places applications across servers in a sequential manner. For each server, the constrained optimization subproblem is solved using the barrier method, ensuring tractability despite the non-linear neural network (NN) constraints. We evaluate NeuRO on heterogeneous Kubernetes clusters using realistic ML/AI workloads. Experimental results show that NeuRO not only achieves substantially higher profiling accuracy compared to existing methods, but also improves operator revenue by up to 30%, all while satisfying diverse service-level agreements (SLAs).
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0682ea33-d079-4108-b38e-7d8f2465b4c9Related papers
- Tailored Learning-Based Scheduling for Kubernetes-Oriented Edge-Cloud SystemYiwen Han, Shihao Shen, Xiaofei Wang, Shiqiang Wang et al.INFOCOM 2021 · 93 citations
- ScalO-RAN: Energy-aware Network Intelligence Scaling in Open RANStefano Maxenti, Salvatore D'Oro, Leonardo Bonati, Michele Polese et al.INFOCOM 2024 · 17 citations
- A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via FaroBeomyeol Jeon, Chen Wang, Diana Arroyo, Alaa Youssef et al.EuroSys 2025
- OrchestRAN: Network Automation through Orchestrated Intelligence in the Open RANSalvatore D'Oro, Leonardo Bonati, Michele Polese, Tommaso MelodiaINFOCOM 2022 · 114 citations
- Kalmia: A Heterogeneous QoS-aware Scheduling Framework for DNN Tasks on Edge ServersZiyan Fu, Ju Ren, Deyu Zhang, Yuezhi Zhou et al.INFOCOM 2022 · 28 citations
