EMA: Efficient Model Adaptation for Learning-based Systems
Daiyang Yu, Xinyu Chen, Yihan Zhang, Yan Liang, Yaqi Qiao, Fan Lai
Abstract
Machine learning (ML) is increasingly applied to optimize system performance in tasks such as resource management and network simulation. Unlike traditional ML tasks (e.g., image classification), networked systems often operate in heterogeneous, long-running, and dynamic environment states, where input conditions (e.g., network loads) and operational objectives can shift over time and across settings. Existing learning-based systems offer little support for adaptation, resulting in costly model training, extensive data collection, degraded system performance, and slow responsiveness.
This paper presents EMA, the first model adaptation system supporting learning-based systems to adapt to evolving environments with minimal operational overhead. EMA takes a system-driven, data-centric approach that accommodates diverse system and model designs while addressing two key deployment challenges. First, it reduces expensive model training by introducing state transformers that align the input state of a new environment with previously similar states, allowing models to warm-start adaptation. Second, it addresses the often-overlooked yet costly process of data labeling-collecting ground truth for exploring and training on various system decisions-by prioritizing labeling highutility data while balancing the tradeoff between training and labeling cost. Evaluations on eight representative learningbased systems show that EMA reduces adaptation costs (e.g., GPU training time) by 14.9-42.4% while improving system performance (e.g., network throughput) by 6.9-31.3%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5602ea60-343c-4e6b-8d3b-cd7394aaf5f5Builds on22
- Learning in situ: a randomized experiment in video streamingFrancis Y. Yan, Hudson Ayers, Chenzhi Zhu, Sadjad Fouladi et al.NSDI 2020 · 360 citations
- FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented MicroservicesHaoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk et al.OSDI 2020 · 350 citations
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh et al.ASPLOS 2021 · 226 citations
- Active Learning on a Budget: Opposite Strategies Suit High and Low BudgetsGuy Hacohen, Avihu Dekel, Daphna WeinshallICML 2022 · 163 citations
- NetLLM: Adapting Large Language Models for NetworkingDuo Wu, Xianda Wang, Yaqi Qiao, Zhi Wang et al.SIGCOMM 2024 · 162 citations
Related papers
- Caravan: Practical Online Learning of In-Network ML Models with Labeling AgentsQizheng Zhang, Ali Imran, Enkeleda Bardhi, Tushar Swamy et al.OSDI 2024 · 18 citations
- Maya: Optimizing Deep Learning Training Workloads using GPU Runtime EmulationSrihas Yarlagadda, Amey Agrawal, Elton Pinto, Hakesh Darapaneni et al.EuroSys 2026
- Env2Vec: accelerating VNF testing with deep learningGuangyuan Piao, Patrick K. Nicholson, Diego LugonesEuroSys 2020 · 1 citation
- Towards Robust and Efficient Cloud-Edge Elastic Model Adaptation via Selective Entropy DistillationYaofo Chen, Shuaicheng Niu, Yaowei Wang, Shoukai Xu et al.ICLR 2024 · 18 citations
- Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference EnvironmentsNikoleta Iliakopoulou, Jovan Stojkovic, Chloe Alverti, Tianyin Xu et al.MICRO 2025 · 3 citations
