Pocket: ML Serving from the Edge
Misun Park, Ketan Bhardwaj, Ada Gavrilovska
2023年份
10被引次数
1顶会引用
摘要
One of the major challenges in serving ML applications is the resource pressure introduced by the underlying ML frameworks. This becomes a bigger problem at resource-constrained, multi-tenant edge server locations, where it is necessary to scale to a larger number of clients with a fixed resource envelope. Naive approaches which simply minimize the resource budget allocation of each application result in performance degradation that voids the benefits expected from operating at the edge.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- InfScaler: Enabling Efficient ML Inference Serving on Multi-Accelerator Edge Devices via Asymmetric Auto-ScalingBorui Li, Tiange Xia, Shuai Wang, Shuai WangDAC 2025 · 被引用 2 次
- NeuRO: Inference-time Profiling and Orchestration of ML Applications at the EdgeArshad Javeed, György Dán, Viktoria FodorINFOCOM 2026 · 被引用 1 次
- Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM ServingJianxiong Liao, Quanxing Dong, Yunkai Liang, Zhi Zhou 等PPoPP 2026 · 被引用 1 次
- EDDE: Container Deployment Framework Beyond the CloudHao Fan, Zhuo Huang, Shadi Ibrahim, Lin Gu 等SC 2025 · 被引用 1 次
- Optimizing Inference Serving on Serverless PlatformsAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniVLDB 2022 · 被引用 76 次
