Bucks for Buckets (B4B): Active Defenses Against Stealing Encoders
Jan Dubinski, Stanislaw Pawlak, Franziska Boenisch, Tomasz Trzcinski, Adam Dziedzic
Abstract
Machine Learning as a Service (MLaaS) APIs provide ready-to-use and high-utility encoders that generate vector representations for given inputs. Since these encoders are very costly to train, they become lucrative targets for model stealing attacks during which an adversary leverages query access to the API to replicate the encoder locally at a fraction of the original training costs. We propose Bucks for Buckets (B4B), the first active defense that prevents stealing while the attack is happening without degrading representation quality for legitimate API users. Our defense relies on the observation that the representations returned to adversaries who try to steal the encoder's functionality cover a significantly larger fraction of the embedding space than representations of legitimate users who utilize the encoder to solve a particular downstream task.vB4B leverages this to adaptively adjust the utility of the returned representations according to a user's coverage of the embedding space. To prevent adaptive adversaries from eluding our defense by simply creating multiple user accounts (sybils), B4B also individually transforms each user's representations. This prevents the adversary from directly aggregating representations over multiple accounts to create their stolen encoder copy. Our active defense opens a new path towards securely sharing and democratizing encoders over public APIs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Stealing Split Learning Bottom Models by Recovering Embedding GeometryQinbo Zhang, Yanhang Shi, Ziyi Zhang, Hao Wang et al.CVPR 2026 · 1 citation
- Efficient and Privacy-Preserving Soft Prompt Transfer for LLMsXun Wang, Jing Xu, Franziska Boenisch, Michael Backes et al.ICML 2025
- On Stealing Graph Neural Network ModelsMarcin Podhajski, Jan Dubinski, Franziska Boenisch, Adam Dziedzic et al.AAAI 2026
Builds on21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
Related papers
- On the Difficulty of Defending Self-Supervised Learning against Model ExtractionAdam Dziedzic, Nikita Dhawan, Muhammad Ahmad Kaleem, Jonas Guan et al.ICML 2022 · 34 citations
- Dataset Inference for Self-Supervised ModelsAdam Dziedzic, Haonan Duan, Muhammad Ahmad Kaleem, Nikita Dhawan et al.NeurIPS 2022 · 59 citations
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing AttacksTribhuvanesh Orekondy, Bernt Schiele, Mario FritzICLR 2020 · 194 citations
- StolenEncoder: Stealing Pre-trained Encoders in Self-supervised LearningYupei Liu, Jinyuan Jia, Hongbin Liu, Neil Zhenqiang GongCCS 2022 · 23 citations
- CloudLeak: Large-Scale Deep Learning Models Stealing Through Adversarial ExamplesHonggang Yu, Kaichen Yang, Teng Zhang, Yun-Yun Tsai et al.NDSS 2020
