Outsourcing Training without Uploading Data via Efficient Collaborative Open-Source Sampling
Junyuan Hong, Lingjuan Lyu, Jiayu Zhou, Michael Spranger
Abstract
As deep learning blooms with growing demand for computation and data resources, outsourcing model training to a powerful cloud server becomes an attractive alternative to training at a low-power and cost-effective end device. Traditional outsourcing requires uploading device data to the cloud server, which can be infeasible in many real-world applications due to the often sensitive nature of the collected data and the limited communication bandwidth. To tackle these challenges, we propose to leverage widely available open-source data, which is a massive dataset collected from public and heterogeneous sources (e.g., Internet images). We develop a novel strategy called Efficient Collaborative Open-source Sampling (ECOS) to construct a proximal proxy dataset from open-source data for cloud training, in lieu of client data. ECOS probes open-source data on the cloud server to sense the distribution of client data via a communication- and computation-efficient sampling process, which only communicates a few compressed public features and client scalar responses. Extensive empirical studies show that the proposed ECOS improves the quality of automated client labeling, model compression, and label outsourcing when applied in various learning scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Differentially Private Synthetic Data via Foundation Model APIs 1: ImagesZinan Lin, Sivakanth Gopi, Janardhan Kulkarni, Harsha Nori et al.ICLR 2024 · 63 citations
- Federated Generative Model on Multi-Source Heterogeneous Data in IoTZuobin Xiong, Wei Li, Zhipeng CaiAAAI 2023 · 40 citations
- Imitation Learning from Imperfection: Theoretical Justifications and AlgorithmsZiniu Li, Tian Xu, Zeyu Qin, Yang Yu et al.NeurIPS 2023 · 26 citations
- The Eminence in Shadow: Exploiting Feature Boundary Ambiguity for Robust Backdoor AttacksZhou Feng, Jiahao Chen, Chunyi Zhou, Yuwen Pu et al.KDD 2026
Builds on15
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- FedBN: Federated Learning on Non-IID Features via Local Batch NormalizationXiaoxiao Li, Meirui Jiang, Xiaofei Zhang, Michael Kamp et al.ICLR 2021 · 1,166 citations
- Group Knowledge Transfer: Federated Learning of Large CNNs at the EdgeChaoyang He, Murali Annavaram, Salman AvestimehrNeurIPS 2020 · 605 citations
Related papers
- Efficient Augmentation for Imbalanced Deep LearningDamien A. Dablain, Colin Bellinger, Bartosz Krawczyk, Nitesh V. ChawlaICDE 2023 · 19 citations
- Collaborative Cloud-edge Generalized Category DiscoveryYingbing Liu, Fei Ma, Yanan Wu, Xinxin Zuo et al.ACM MM 2025
- Reimagining Mutual Information for Enhanced Defense against Data Leakage in Collaborative InferenceLin Duan, Jingwei Sun, Jinyuan Jia, Yiran Chen et al.NeurIPS 2024 · 5 citations
- DOS: Diverse Outlier Sampling for Out-of-Distribution DetectionWenyu Jiang, Hao Cheng, Mingcai Chen, Chongjun Wang et al.ICLR 2024 · 14 citations
- Efficient Edge Inference by Selective QueryAnil Kag, Igor Fedorov, Aditya Gangrade, Paul N. Whatmough et al.ICLR 2023
