Vulcan: Automatic Query Planning for Live ML Analytics
Yiwen Zhang, Xumiao Zhang, Ganesh Ananthanarayanan, Anand P. Iyer, Yuanchao Shu, Victor Bahl, Z. Morley Mao, Mosharaf Chowdhury
摘要
Live ML analytics have gained increasing popularity with large-scale deployments due to recent evolution of ML technologies. To serve live ML queries, experts nowadays still need to perform manual query planning, which involves pipeline construction, query configuration, and pipeline placement across multiple edge tiers in a heterogeneous infrastructure. Finding the best query plan for a live ML query requires navigating a huge search space, calling for an efficient and systematic solution.
In this paper, we propose Vulcan, a system that automatically generates query plans for live ML queries to optimize their accuracy, latency, and resource consumption. Based on the user query and performance requirements, Vulcan determines the best pipeline, placement, and query configuration for the query with low profiling cost; it also performs fast online adaptation after query deployment. Vulcan outperforms state-of-the-art ML analytics systems by 4.1×-30.1× in terms of search cost while delivering up to 3.3× better query latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Kareus: Joint Reduction of Dynamic and Static Energy in Large Model TrainingRuofan Wu, Jae-Won Chung, Mosharaf ChowdhuryOSDI 2026 · 被引用 8 次
- Compass: SLO-aware Query Planner for Compound AI Serving at ScaleBanruo Liu, Wei-Yu Lin, Minghao Fang, Yihan Jiang 等VLDB 2026 · 被引用 5 次
- AVA: Towards Agentic Video Analytics with Vision Language ModelsYuxuan Yan, Shiqi Jiang, Ting Cao, Yifan Yang 等NSDI 2026 · 被引用 4 次
- SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model InferenceHao Ma, Melis Ilayda Bal, Liang Zhang, Bingcong Li 等ICML 2026
它引用的顶会 Paper12
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- EMP: edge-assisted multi-vehicle perceptionXumiao Zhang, Anlan Zhang, Jiachen Sun, Xiao Zhu 等MobiCom 2021 · 被引用 137 次
- VIPS: real-time perception fusion for infrastructure-assisted autonomous drivingShuyao Shi, Jiahe Cui, Zhehao Jiang, Zhenyu Yan 等MobiCom 2022 · 被引用 126 次
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 被引用 103 次
- Flexible high-resolution object detection on edge devices with tunable latencyShiqi Jiang, Zhiqi Lin, Yuanchun Li, Yuanchao Shu 等MobiCom 2021 · 被引用 103 次
相关 Paper
- VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space DecompositionYang Li, Yu Shen, Wentao Zhang, Jiawei Jiang 等VLDB 2021 · 被引用 55 次
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 被引用 325 次
- A Spark Optimizer for Adaptive, Fine-Grained Parameter TuningChenghao Lyu, Qi Fan, Philippe Guyard, Yanlei DiaoVLDB 2024 · 被引用 9 次
- End-to-end Optimization of Machine Learning Prediction QueriesKwanghyun Park, Karla Saur, Dalitso Banda, Rathijit Sen 等SIGMOD 2022 · 被引用 50 次
- Optimizing Video Analytics with Declarative Model RelationshipsFrancisco Romero, Johann Hauswald, Aditi Partap, Daniel Kang 等VLDB 2023 · 被引用 37 次
