Lune

INFOCOM2026顶会

NeuRO: Inference-time Profiling and Orchestration of ML Applications at the Edge

Arshad Javeed, György Dán, Viktoria Fodor

2026年份
1被引次数

摘要

We address the problem of orchestrating machine learning (ML) workloads, such as latency-sensitive mobile applications and intelligent radio access network (RAN) control functions, in a heterogeneous edge cloud infrastructure, subject to inference time constraints. We introduce NeuRO, a modular profiling framework that learns to approximate ML task inference time summary statistics using data collected from edge deployments. NeuRO leverages two lightweight neural embeddings: one capturing the resource footprint of an individual ML task, and one capturing the aggregate server context induced by co-located workloads. We show the potential of NeuRO by formulating a performance-aware service placement problem subject to inference time summary statistic constraints. To solve this problem efficiently at scale, we develop a heuristic algorithm that incrementally places applications across servers in a sequential manner. For each server, the constrained optimization subproblem is solved using the barrier method, ensuring tractability despite the non-linear neural network (NN) constraints. We evaluate NeuRO on heterogeneous Kubernetes clusters using realistic ML/AI workloads. Experimental results show that NeuRO not only achieves substantially higher profiling accuracy compared to existing methods, but also improves operator revenue by up to 30%, all while satisfying diverse service-level agreements (SLAs).

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖