Outsourcing Training without Uploading Data via Efficient Collaborative Open-Source Sampling
Junyuan Hong, Lingjuan Lyu, Jiayu Zhou, Michael Spranger
摘要
As deep learning blooms with growing demand for computation and data resources, outsourcing model training to a powerful cloud server becomes an attractive alternative to training at a low-power and cost-effective end device. Traditional outsourcing requires uploading device data to the cloud server, which can be infeasible in many real-world applications due to the often sensitive nature of the collected data and the limited communication bandwidth. To tackle these challenges, we propose to leverage widely available open-source data, which is a massive dataset collected from public and heterogeneous sources (e.g., Internet images). We develop a novel strategy called Efficient Collaborative Open-source Sampling (ECOS) to construct a proximal proxy dataset from open-source data for cloud training, in lieu of client data. ECOS probes open-source data on the cloud server to sense the distribution of client data via a communication- and computation-efficient sampling process, which only communicates a few compressed public features and client scalar responses. Extensive empirical studies show that the proposed ECOS improves the quality of automated client labeling, model compression, and label outsourcing when applied in various learning scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Differentially Private Synthetic Data via Foundation Model APIs 1: ImagesZinan Lin, Sivakanth Gopi, Janardhan Kulkarni, Harsha Nori 等ICLR 2024 · 被引用 63 次
- Federated Generative Model on Multi-Source Heterogeneous Data in IoTZuobin Xiong, Wei Li, Zhipeng CaiAAAI 2023 · 被引用 40 次
- Imitation Learning from Imperfection: Theoretical Justifications and AlgorithmsZiniu Li, Tian Xu, Zeyu Qin, Yang Yu 等NeurIPS 2023 · 被引用 26 次
- The Eminence in Shadow: Exploiting Feature Boundary Ambiguity for Robust Backdoor AttacksZhou Feng, Jiahao Chen, Chunyi Zhou, Yuwen Pu 等KDD 2026
它引用的顶会 Paper15
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
- FedBN: Federated Learning on Non-IID Features via Local Batch NormalizationXiaoxiao Li, Meirui Jiang, Xiaofei Zhang, Michael Kamp 等ICLR 2021 · 被引用 1,166 次
- Group Knowledge Transfer: Federated Learning of Large CNNs at the EdgeChaoyang He, Murali Annavaram, Salman AvestimehrNeurIPS 2020 · 被引用 605 次
相关 Paper
- Efficient Augmentation for Imbalanced Deep LearningDamien A. Dablain, Colin Bellinger, Bartosz Krawczyk, Nitesh V. ChawlaICDE 2023 · 被引用 19 次
- Collaborative Cloud-edge Generalized Category DiscoveryYingbing Liu, Fei Ma, Yanan Wu, Xinxin Zuo 等ACM MM 2025
- Reimagining Mutual Information for Enhanced Defense against Data Leakage in Collaborative InferenceLin Duan, Jingwei Sun, Jinyuan Jia, Yiran Chen 等NeurIPS 2024 · 被引用 5 次
- DOS: Diverse Outlier Sampling for Out-of-Distribution DetectionWenyu Jiang, Hao Cheng, Mingcai Chen, Chongjun Wang 等ICLR 2024 · 被引用 14 次
- Efficient Edge Inference by Selective QueryAnil Kag, Igor Fedorov, Aditya Gangrade, Paul N. Whatmough 等ICLR 2023
