Stealthy Imitation: Reward-guided Environment-free Policy Stealing
Zhixiong Zhuang, Maria-Irina Nicolae, Mario Fritz
摘要
Deep reinforcement learning policies, which are integral to modern control systems, represent valuable intellectual property. The development of these policies demands considerable resources, such as domain expertise, simulation fidelity, and real-world validation. These policies are potentially vulnerable to model stealing attacks, which aim to replicate their functionality using only black-box access. In this paper, we propose Stealthy Imitation, the first attack designed to steal policies without access to the environment or knowledge of the input range. This setup has not been considered by previous model stealing methods. Lacking access to the victim's input states distribution, Stealthy Imitation fits a reward model that allows to approximate it. We show that the victim policy is harder to imitate when the distribution of the attack queries matches that of the victim. We evaluate our approach across diverse, high-dimensional control tasks and consistently outperform prior data-free approaches adapted for policy stealing. Lastly, we propose a countermeasure that significantly diminishes the effectiveness of the attack. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Medical Multimodal Model Stealing Attacks via Adversarial Domain AlignmentYaling Shen, Zhixiong Zhuang, Kun Yuan, Maria-Irina Nicolae 等AAAI 2025 · 被引用 12 次
- Stealix: Model Stealing via Prompt EvolutionZhixiong Zhuang, Hui-Po Wang, Maria-Irina Nicolae, Mario FritzICML 2025
它引用的顶会 Paper7
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 被引用 266 次
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing AttacksTribhuvanesh Orekondy, Bernt Schiele, Mario FritzICLR 2020 · 被引用 194 次
- Error Bounds of Imitating Policies and EnvironmentsTian Xu, Ziniu Li, Yang YuNeurIPS 2020 · 被引用 141 次
- Towards Data-Free Model Stealing in a Hard Label SettingSunandini Sanyal, Sravanti Addepalli, R. Venkatesh BabuCVPR 2022 · 被引用 76 次
相关 Paper
- MAZE: Data-Free Model Stealing Attack Using Zeroth-Order Gradient EstimationSanjay Kariyappa, Atul Prakash, Moinuddin K. QureshiCVPR 2021
- Imitated Detectors: Stealing Knowledge of Black-box Object DetectorsSiyuan Liang, Aishan Liu, Jiawei Liang, Longkang Li 等ACM MM 2022 · 被引用 16 次
- Robust Deep Reinforcement Learning against Adversarial Behavior ManipulationShojiro Yamabe, Kazuto Fukuchi, Jun SakumaICLR 2026 · 被引用 1 次
- Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement LearningAmin Rakhsha, Goran Radanovic, Rati Devidze, Xiaojin Zhu 等ICML 2020 · 被引用 145 次
- Attacking Black-box Recommendations via Copying Cross-domain User ProfilesWenqi Fan, Tyler Derr, Xiangyu Zhao, Yao Ma 等ICDE 2021 · 被引用 75 次
