AID: Active Distillation Machine to Leverage Pre-Trained Black-Box Models in Private Data Settings
Trong Nghia Hoang, Shenda Hong, Cao Xiao, Bryan Low, Jimeng Sun
Abstract
This paper presents an active distillation method for a local institution (e.g., hospital) to find the best queries within its given budget to distill an on-server black-box model’s predictive knowledge into a local surrogate with transparent parameterization. This allows local institutions to understand better the predictive reasoning of the black-box model in its own local context or to further customize the distilled knowledge with its private dataset that cannot be centralized and fed into the server model. The proposed method thus addresses several challenges of deploying machine learning (ML) in many industrial settings (e.g., healthcare analytics) with strong proprietary constraints. These include: (1) the opaqueness of the server model’s architecture which prevents local users from understanding its predictive reasoning in their local data contexts; (2) the increasing cost and risk of uploading local data on the cloud for analysis; and (3) the need to customize the server model with private onsite data. We evaluated the proposed method on both benchmark and real-world healthcare data where significant improvements over existing local distillation methods were observed. A theoretical analysis of the proposed method is also presented.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b0b28f0-93ed-4853-8906-bdbbae5141e6Cited by top-tier papers6
- Fault-Tolerant Federated Reinforcement Learning with Theoretical GuaranteeFlint Xiaofeng Fan, Yining Ma, Zhongxiang Dai, Wei Jing et al.NeurIPS 2021 · 102 citations
- M3Care: Learning with Missing Modalities in Multimodal Healthcare DataChaohe Zhang, Xu Chu, Liantao Ma, Yinghao Zhu et al.KDD 2022 · 78 citations
- Fair yet Asymptotically Equal Collaborative LearningXiaoqiang Lin, Xinyi Xu, See-Kiong Ng, Chuan-Sheng Foo et al.ICML 2023 · 15 citations
- Model Shapley: Equitable Model Valuation with Black-box AccessXinyi Xu, Thanh Lam, Chuan Sheng Foo, Bryan Kian Hsiang LowNeurIPS 2023 · 8 citations
- Unbiased Missing-Modality Multimodal LearningRuiting Dai, Chenxi Li, Yandong Yan, Lisi Mo et al.ICCV 2025 · 8 citations
Builds on1
Related papers
- FedED: Federated Learning via Ensemble Distillation for Medical Relation ExtractionDianbo Sui, Yubo Chen, Jun Zhao, Yantao Jia et al.EMNLP 2020 · 126 citations
- Privacy Budgeting for Growing Machine Learning DatasetsWeiting Li, Liyao Xiang, Zhou Zhou, Feng PengINFOCOM 2021 · 14 citations
- Ensemble Attention Distillation for Privacy-Preserving Federated LearningXuan Gong, Abhishek Sharma, Srikrishna Karanam, Ziyan Wu et al.ICCV 2021 · 148 citations
- Positive–Unlabeled Reinforcement Learning Distillation for On-Premise Small ModelsZhiqiang Kou, Junyang Chen, Xin-Qiang Cai, Xiaobo Xia et al.ICML 2026
- Model Distillation for Revenue Optimization: Interpretable Personalized PricingMax Biggs, Wei Sun, Markus EttlICML 2021 · 42 citations
