Caravan: Practical Online Learning of In-Network ML Models with Labeling Agents
Qizheng Zhang, Ali Imran, Enkeleda Bardhi, Tushar Swamy, Nathan Zhang, Muhammad Shahbaz, Kunle Olukotun
Abstract
Recent work on in-network machine learning (ML) anticipates offline models to operate well in modern networking environments. However, upon deployment, these models struggle to cope with fluctuating traffic patterns and network conditions and, therefore, must be validated and updated frequently in an online fashion. This abstract presents Caravan, a practical online learning system for in-network ML models. We tackle two primary challenges in facilitating online learning for networking: (a) the automatic labeling of evolving traffic and (b) the efficient monitoring and detection of model performance degradation to trigger retraining. Caravan repurposes existing systems (e.g., heuristics, access control lists, and foundation models)---not directly suitable for such dynamic environments---into high-quality labeling sources for generating labeled data for online learning. Caravan also introduces a new metric, accuracy proxy, to track model degradation and potential drift to efficiently trigger retraining. Our evaluations show that Caravan's labeling strategy enables in-network ML models to closely follow the changes in the traffic dynamics with a 30.3% improvement in F1 score (on average), compared to offline models. Moreover, Caravan sustains comparable inference accuracy to that of a continuous-learning system while consuming 61.3% less GPU compute time (on average) via accuracy proxy and retraining triggers. This abstract summarizes a previously published work [53] at the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ae5f37b-be4c-49b9-9d03-8a424632fefcCited by top-tier papers7
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language ModelsQizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma et al.ICLR 2026 · 374 citations
- Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM AgentsQizheng Zhang, Michael Wornow, Kunle OlukotunNeurIPS 2025 · 27 citations
- METIS: Fast Quality-Aware RAG Systems with Configuration AdaptationSiddhant Ray, Rui Pan, Zhuohan Gu, Kuntai Du et al.SOSP 2025 · 3 citations
- PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model ApplicationsKuntai Du, Bowen Wang, Chen Zhang, Yiming Cheng et al.SOSP 2025 · 3 citations
- Automated Discovery of Test Oracles for Database Management Systems Using LLMsQiuyang Mang, Runyuan He, Suyang Zhong, Xiaoxuan Liu et al.SIGMOD 2026 · 1 citation
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Learning in situ: a randomized experiment in video streamingFrancis Y. Yan, Hudson Ayers, Chenzhi Zhu, Sadjad Fouladi et al.NSDI 2020 · 360 citations
- CADE: Detecting and Explaining Concept Drift Samples for Security ApplicationsLimin Yang, Wenbo Guo, Qingying Hao, Arridhana Ciptadi et al.USENIX Security 2021 · 241 citations
Related papers
- Carbon-Aware Continuous Learning for Sustainable Real-Time Machine Learning AnalyticsGwanjong Park, Osama Khan, Dongho Ha, Myeongjae Jeon et al.EuroSys 2026 · 1 citation
- EMA: Efficient Model Adaptation for Learning-based SystemsDaiyang Yu, Xinyu Chen, Yihan Zhang, Yan Liang et al.SIGCOMM 2026
- Leo: Online ML-based Traffic Classification at Multi-Terabit Line RateSyed Usman Jafri, Sanjay G. Rao, Vishal Shrivastav, Mohit TawarmalaniNSDI 2024 · 46 citations
- Autonomous Unknown-Application Filtering and Labeling for DL-based Traffic Classifier UpdateJielun Zhang, Fuhao Li, Feng Ye, Hongyu WuINFOCOM 2020 · 120 citations
- Accelerating Deep Learning Classification with Error-controlled Approximate-key CachingAlessandro Finamore, James Roberts, Massimo Gallo, Dario RossiINFOCOM 2022
