LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMs
Chansung Park, Juyong Jiang, Fan Wang, Sayak Paul, Jing Tang
Abstract
The widespread adoption of cloud-based proprietary large language models (LLMs) has introduced significant challenges, including operational dependencies, privacy concerns, and the necessity of continuous internet connectivity. In this work, we introduce an LLMOps pipeline,"LlamaDuo", for the seamless migration of knowledge and abilities from service-oriented LLMs to smaller, locally manageable models. This pipeline is crucial for ensuring service continuity in the presence of operational failures, strict privacy policies, or offline requirements. Our LlamaDuo involves fine-tuning a small language model against the service LLM using a synthetic dataset generated by the latter. If the performance of the fine-tuned model falls short of expectations, it is automatically improved through additional fine-tuning using extra similar data generated by the service LLM. This multi-turn process guarantees that the smaller model can eventually match or even surpass the service LLM's capabilities in specific downstream tasks, offering a practical and scalable solution for managing AI deployments in constrained environments. Extensive experiments with leading-edge LLMs are conducted to demonstrate the effectiveness, adaptability, and affordability of LlamaDuo across various downstream tasks. Our pipeline implementation is available at https://github.com/deep-diver/llamaduo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f175389e-b09e-4daf-bcb6-3326ecb9d44aCited by top-tier papers3
- KaSA: Knowledge-Aware Singular-Value Adaptation of Large Language ModelsFan Wang, Juyong Jiang, Chansung Park, Sunghun Kim et al.ICLR 2025
- A Strategic Coordination Framework of Small LMs Matches Large LMs in Data SynthesisXin Gao, Qizhi Pei, Zinan Tang, Yu Li et al.ACL 2025
- Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop LearningZinan Tang, Xin Gao, Qizhi Pei, Zhuoshi Pan et al.EMNLP 2025
Builds on11
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
Related papers
- Lemix: Unified Scheduling for Llm Training and Inference on Multi-Gpu SystemsYufei Li, Zexin Li, Yinglun Zhu, Cong LiuRTSS 2025 · 4 citations
- AIMS: Cost-Efficient LLM-Based Agent Deployment in Hybrid Cloud-Edge EnvironmentsShiyi Liu, Haiying Shen, Shuai Che, Mahdi Ghandi et al.EuroSys 2026 · 1 citation
- Crimson: Collaborative Parameter Updates for Efficient Pipeline Training of Large Language ModelsYapeng Jiang, Wuhui Chen, Ganhong Huang, Yuzhou Huang et al.EuroSys 2026
- A Structure-Agnostic Co-Tuning Framework for LLMs and SLMs in Cloud-Edge SystemsYuze Liu, Yunhan Wang, Tiehua Zhang, Zhishu Shen et al.WWW 2026 · 1 citation
- Crayon: Customized On-Device LLM via Instant Adapter Blending and Edge-Server Hybrid InferenceJihwan Bang, Juntae Lee, Kyuhong Shim, Seunghan Yang et al.ACL 2024 · 2 citations
