On the Ability of Developers' Training Data Preservation of Learnware
Hao-Yi Lei, Zhi-Hao Tan, Zhi-Hua Zhou
Abstract
The learnware paradigm aims to enable users to leverage numerous existing well-trained models instead of building machine learning models from scratch. In this paradigm, developers worldwide can submit their well-trained models spontaneously into a learnware dock system , and the system helps developers generate specification for each model to form a learnware. As the key component, a specification should characterize the capabilities of the model, enabling it to be adequately identified and reused, while preserving the developer’s original data. Recently, the RKME (Reduced Kernel Mean Embedding) specification was proposed and most commonly utilized. This paper provides a theoretical analysis of RKME specification about its preservation ability for developer’s training data. By modeling it as a geometric problem on manifolds and utilizing tools from geometric analysis, we prove that the RKME specification is able to disclose none of the developer’s original data and possesses robust defense against common inference attacks, while preserving sufficient information for effective learnware identification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0d0913f-3732-4912-8f7d-87638c4115f8Cited by top-tier papers7
- Constructive Specification for Plug-and-Play Learnware AgentsJian-Dong Liu, Zi-Chen Zhao, Hao Sun, Lin-Xing Wu et al.KDD 2026 · 3 citations
- Learnware Specification via Label-Aware Neural EmbeddingWei Chen, Junxiang Mao, Min-Ling ZhangAAAI 2025 · 1 citation
- A Study on PAVE Specification for LearnwareHao-Yu Shi, Zhi-Hao Tan, Zi-Chen Zhao, Yang Yu et al.ICLR 2026
- Tabular Learnwares Can Be Repurposed for Seemingly Irrelevant New TasksPeng Tan, Feifan Yang, Zhi-Hao Tan, Zhi-Hua ZhouAAAI 2026
- Dynamic Learnware Filtering for Efficient Learnware Identification and System SlimmingJian-Dong Liu, Zhi-Hao Tan, Zhi-Hua ZhouKDD 2025
Builds on7
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Knock Knock, Who's There? Membership Inference on Aggregate Location DataApostolos Pyrgelis, Carmela Troncoso, Emiliano De CristofaroNDSS 2018 · 293 citations
- GAN-Leaks: A Taxonomy of Membership Inference Attacks against Generative ModelsDingfan Chen, Ning Yu, Yang Zhang, Mario FritzCCS 2020 · 278 citations
- Towards Enabling Learnware to Handle Unseen JobsYu-Jie Zhang, Yu-Hu Yan, Peng Zhao, Zhi-Hua ZhouAAAI 2021 · 20 citations
- Understanding Reconstruction Attacks with the Neural Tangent Kernel and Dataset DistillationNoel Loo, Ramin M. Hasani, Mathias Lechner, Alexander Amini et al.ICLR 2024 · 14 citations
Related papers
- Identifying Useful Learnwares for Heterogeneous Label SpacesLan-Zhe Guo, Zhi Zhou, Yufeng Li, Zhi-Hua ZhouICML 2023 · 17 citations
- Identifying Learnwares via Reduced Neural Conditional Mean EmbeddingZi-Yu Mao, Ming LiICML 2026
- A Statistical Framework for Analyzing Specification Resistance to Learnware-Inversion RisksHao-Yi Lei, Zhi-Hao Tan, Zhi-Hua ZhouICML 2026
- Integrated Learnware Identification and Reuse via Reusability-Aware Metric LearningHai-Tian Liu, Peng Tan, Jian-Dong Liu, Zhi-Hao Tan et al.KDD 2026
- Handling Learnwares from Heterogeneous Feature Spaces with Explicit Label ExploitationPeng Tan, Hai-Tian Liu, Zhi-Hao Tan, Zhi-Hua ZhouNeurIPS 2024 · 8 citations
