Experience Speaks Louder: Black-box Hard-label Adversarial Attack through Reinforcement Learning
Yilun Jin, Kun Zhu, Feng Tang, Jiaxuan Shi, Yong Chen
摘要
Black-box hard-label adversarial attack on text can be used to evaluate the robustness of text classification models (TCMs), allowing for a more thorough examination of potential security flaws before deployment. Research on this problem is still in the embryonic stage and only a few methods are available. Existing adversarial attack methods often treat each adversarial text generation as an independent task, failing to leverage past adversarial experience to facilitate subsequent generation. Consequently, these methods suffer from high query costs (i.e., the number of times TCMs need to be accessed to generate adversarial text) and pose challenges in generating adversarial texts with high semantic similarity and low perturbation rates under a query-limited setup. To achieve query-efficient adversarial text generation, we propose EGAttack, an experience-guided black-box hard-label attack method based on reinforcement learning. Specifically, we models adversarial text generation as a sequential decision process and propose a new adversarial environment to capture the intricate relationships among word perturbations, text semantic, and the target model's outputs for agent learning. Meanwhile, we extend the classical reinforcement framework DDQN into adversarial DDQN (ADDQN) through designing a new adversarial agent and a new dual-objective constraint, driving EGAttack to learn from past adversarial experience while identifying the vulnerabilities of TCMs, thus generating high-quality adversarial texts. Experimental results across various target models and datasets demonstrate that EGAttack generates adversarial texts with higher semantic similarity, lower perturbation rates, and competitive attack success rate, making it an efficient and effective tool for evaluating the robustness of TCMs.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on TextHan Liu, Zhi Xu, Xiaotong Zhang, Feng Zhang 等NeurIPS 2023 · 被引用 32 次
- Policy-Driven Attack: Learning to Query for Hard-label Black-box Adversarial ExamplesZiang Yan, Yiwen Guo, Jian Liang, Changshui ZhangICLR 2021 · 被引用 19 次
- Multi-granularity Textual Adversarial Attack with Behavior CloningYangyi Chen, Jin Su, Wei WeiEMNLP 2021 · 被引用 29 次
- Generating Natural Language Attacks in a Hard Label Black Box SettingRishabh Maheshwary, Saket Maheshwary, Vikram PudiAAAI 2021 · 被引用 128 次
- TextHoaxer: Budgeted Hard-Label Adversarial Attacks on TextMuchao Ye, Chenglin Miao, Ting Wang, Fenglong MaAAAI 2022 · 被引用 57 次
