Task-Aware Cloud-End Offloading for Vision-Language Model Serving via Dynamic Modality-Specific Adapter Scheduling
Zian Wang, Ziyi Wang, Jie Xing, Yaya Wei, Ziyan Zhong, Lanshan Zhang
Abstract
Large-scale vision-language models enable powerful cross-modal understanding and generation, driving rapidly growing demand for online inference services. However, cloud-centric serving often suffers from high latency, rising costs, and network dependency, while purely on-device deployment is constrained by limited memory and reduced accuracy on complex tasks. To address this accuracy–latency–cost trilemma, we propose ShiftVL, a task-aware end–cloud serving framework that shifts suitable execution to the end device with a cloud fallback. ShiftVL serves high-frequency requests on an end-side small VLM enhanced with ViTexLoRA, a modality-disentangled parameter-efficient tuning method that preserves cross-modal alignment, while routing low-frequency or complex requests to a cloud-hosted large VLM for higher accuracy. Under tight device budgets, ShiftVL employs a predictive adapter scheduler that combines LRU-style caching with imitation learning to pre-load task-specific adapters. Experiments with InternVL models show that ShiftVL reduces cloud cost by up to 76.3% and latency by up to 42.9% while maintaining high multi-task accuracy, demonstrating its practicality for real-world vision-language model serving.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 871966ce-cc20-4e74-83f6-408fe55538d8Related papers
- Empower Vision Applications with LoRA LMMLiang Mi, Weijun Wang, Wenming Tu, Qingfeng He et al.EuroSys 2025 · 2 citations
- Eve: Efficient Multimodal Vision Language Models with Elastic Visual ExpertsMiao Rang, Zhenni Bi, Chuanjian Liu, Yehui Tang et al.AAAI 2025 · 16 citations
- TailorLLM: Collaborative End-Cloud Inference of Large and Small Language Models Based on Low-Rank AdaptationZian Wang, Ziyi Wang, Haonan Jin, Jie Xing et al.EuroSys 2026
- MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention Across Vision-Language ModelsXiaoran Fan, Zhichao Sun, Tao Ji, Lixing Shen et al.AAAI 2026
- ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM ServingJiuchen Shi, Hang Zhang, Yixiao Wang, Quan Chen et al.HPCA 2026
