EMA: Efficient Model Adaptation for Learning-based Systems
Daiyang Yu, Xinyu Chen, Yihan Zhang, Yan Liang, Yaqi Qiao, Fan Lai
摘要
Machine learning (ML) is increasingly applied to optimize system performance in tasks such as resource management and network simulation. Unlike traditional ML tasks (e.g., image classification), networked systems often operate in heterogeneous, long-running, and dynamic environment states, where input conditions (e.g., network loads) and operational objectives can shift over time and across settings. Existing learning-based systems offer little support for adaptation, resulting in costly model training, extensive data collection, degraded system performance, and slow responsiveness.
This paper presents EMA, the first model adaptation system supporting learning-based systems to adapt to evolving environments with minimal operational overhead. EMA takes a system-driven, data-centric approach that accommodates diverse system and model designs while addressing two key deployment challenges. First, it reduces expensive model training by introducing state transformers that align the input state of a new environment with previously similar states, allowing models to warm-start adaptation. Second, it addresses the often-overlooked yet costly process of data labeling-collecting ground truth for exploring and training on various system decisions-by prioritizing labeling highutility data while balancing the tradeoff between training and labeling cost. Evaluations on eight representative learningbased systems show that EMA reduces adaptation costs (e.g., GPU training time) by 14.9-42.4% while improving system performance (e.g., network throughput) by 6.9-31.3%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Learning in situ: a randomized experiment in video streamingFrancis Y. Yan, Hudson Ayers, Chenzhi Zhu, Sadjad Fouladi 等NSDI 2020 · 被引用 360 次
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk 等OSDI 2020 · 被引用 350 次
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh 等ASPLOS 2021 · 被引用 226 次
- Active Learning on a Budget: Opposite Strategies Suit High and Low BudgetsGuy Hacohen, Avihu Dekel, Daphna WeinshallICML 2022 · 被引用 163 次
- NetLLM: Adapting Large Language Models for NetworkingDuo Wu, Xianda Wang, Yaqi Qiao, Zhi Wang 等SIGCOMM 2024 · 被引用 162 次
相关 Paper
- Caravan: Practical Online Learning of In-Network ML Models with Labeling AgentsQizheng Zhang, Ali Imran, Enkeleda Bardhi, Tushar Swamy 等OSDI 2024 · 被引用 18 次
- Maya: Optimizing Deep Learning Training Workloads using GPU Runtime EmulationSrihas Yarlagadda, Amey Agrawal, Elton Pinto, Hakesh Darapaneni 等EuroSys 2026
- Env2Vec: accelerating VNF testing with deep learningGuangyuan Piao, Patrick K. Nicholson, Diego LugonesEuroSys 2020 · 被引用 1 次
- Towards Robust and Efficient Cloud-Edge Elastic Model Adaptation via Selective Entropy DistillationYaofo Chen, Shuaicheng Niu, Yaowei Wang, Shoukai Xu 等ICLR 2024 · 被引用 18 次
- Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference EnvironmentsNikoleta Iliakopoulou, Jovan Stojkovic, Chloe Alverti, Tianyin Xu 等MICRO 2025 · 被引用 3 次
