C2F: Enabling Context-Aware Edge-Cloud Collaborative Inference for Foundation Models
Mingyue Zhao, Jiayi Shi, Zhengyuan Zhang, Yue Ling, Guanzhou Zhu, Dong Zhao, Huadong Ma
Abstract
Transformer-based foundation models (FMs) excel in diverse domains but struggle with high computation costs, hin-dering effective inference on resource-constrained edge devices. Edge-cloud collaborative inference offers a promising paradigm, but existing methods either neglect the contexts (computing resources and data environments) of edge devices or incur high transmission overheads when applied to FMs. In contrast, we propose C2F, a novel context-aware edge-cloud collaborative inference method for FMs, enabling open-set learning with low latency and high accuracy. In C2F, a context-aware model customization module is utilized to customize a small model (SM) from the FM based on the context, subsequently deployed on the designated edge device. During inference, an adaptive inference module is employed to determine whether to query the FM according to local results of the SM; if so, it transmits only vital data patches, effectively reducing transmission costs and end-to-end latency. Extensive experimental results demonstrate that C2F achieves efficient inference for FMs, enhancing accuracy by 3.1-30.3% and reducing end-to-end latency by 1.32-8.75× compared to various baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2aceec2c-ce60-4389-befd-0e5c8170d2f8Cited by top-tier papers1
Ask how each one uses itRelated papers
- Heterogeneous LoRA for Federated Fine-tuning of On-Device Foundation ModelsYae Jee Cho, Luyang Liu, Zheng Xu, Aldi Fahrezi et al.EMNLP 2024 · 36 citations
- Cross-Architecture Adaptation: Cloud-Edge Continual Test-Time Adaptation with Dynamic Sampling and Heterogeneous DistillationZirui Xu, Xianhang Chu, Jiahao Li, Xu Yang et al.CVPR 2026
- Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer InferenceShengyuan Ye, Jiangsu Du, Liekang Zeng, Wenzhong Ou et al.INFOCOM 2024 · 43 citations
- EdgeFormer: Latency-Aware Collaborative Multi-Head Attention of Transformer Inference in Edge NetworksYiming Yao, Jianwei Niu, Bin Dai, Tao RenACL 2026
- DeepAdapter: A Collaborative Deep Learning Framework for the Mobile Web Using Context-Aware Network PruningYakun Huang, Xiuquan Qiao, Jian Tang, Pei Ren et al.INFOCOM 2020 · 32 citations
