HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on Text
Han Liu, Zhi Xu, Xiaotong Zhang, Feng Zhang, Fenglong Ma, Hongyang Chen, Hong Yu, Xianchao Zhang
摘要
Black-box hard-label adversarial attack on text is a practical and challenging task, as the text data space is inherently discrete and non-differentiable, and only the predicted label is accessible. Research on this problem is still in the embryonic stage and only a few methods are available. Nevertheless, existing methods rely on the complex heuristic algorithm or unreliable gradient estimation strategy, which probably fall into the local optimum and inevitably consume numerous queries, thus are difficult to craft satisfactory adversarial examples with high semantic similarity and low perturbation rate in a limited query budget. To alleviate above issues, we propose a simple yet effective framework to generate high quality textual adversarial examples under the black-box hard-label attack scenarios, named HQA-Attack. Specifically, after initializing an adversarial example randomly, HQA-attack first constantly substitutes original words back as many as possible, thus shrinking the perturbation rate. Then it leverages the synonym set of the remaining changed words to further optimize the adversarial example with the direction which can improve the semantic similarity and satisfy the adversarial condition simultaneously. In addition, during the optimizing procedure, it searches a transition synonym word for each changed word, thus avoiding traversing the whole synonym set and reducing the query number to some extent. Extensive experimental results on five text classification datasets, three natural language inference datasets and two real-world APIs have shown that the proposed HQA-Attack method outperforms other strong baselines significantly.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- CheatAgent: Attacking LLM-Empowered Recommender Systems via LLM AgentLiang-Bo Ning, Shijie Wang, Wenqi Fan, Qing Li 等KDD 2024 · 被引用 11 次
- TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language ModelsZhi Xu, Jiaqi Li, Xiaotong Zhang, Hong Yu 等ICLR 2026 · 被引用 2 次
- Towards a 3D Transfer-Based Black-Box Attack via Critical Feature GuidanceShuchao Pang, Zhenghan Chen, Shen Zhang, Liming Lu 等ICCV 2025 · 被引用 1 次
- Tightening Robustness Verification of MaxPool-based Neural Networks via Minimizing the Over-Approximation ZoneYuan Xiao, Yuchen Chen, Shiqing Ma, Chunrong Fang 等CVPR 2025
- HQA-VLAttack: Towards High Quality Adversarial Attack on Vision-Language Pre-Trained ModelsHan Liu, Jiaqi Li, Zhi Xu, Xiaotong Zhang 等NeurIPS 2025
它引用的顶会 Paper19
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
相关 Paper
- SSPAttack: A Simple and Sweet Paradigm for Black-Box Hard-Label Textual Adversarial AttackHan Liu, Zhi Xu, Xiaotong Zhang, Xiaoming Xu 等AAAI 2023 · 被引用 31 次
- TextHoaxer: Budgeted Hard-Label Adversarial Attacks on TextMuchao Ye, Chenglin Miao, Ting Wang, Fenglong MaAAAI 2022 · 被引用 57 次
- LeapAttack: Hard-Label Adversarial Attack on Text via Gradient-Based OptimizationMuchao Ye, Jinghui Chen, Chenglin Miao, Ting Wang 等KDD 2022 · 被引用 16 次
- Generating Natural Language Attacks in a Hard Label Black Box SettingRishabh Maheshwary, Saket Maheshwary, Vikram PudiAAAI 2021 · 被引用 128 次
- Experience Speaks Louder: Black-box Hard-label Adversarial Attack through Reinforcement LearningYilun Jin, Kun Zhu, Feng Tang, Jiaxuan Shi 等KDD 2025 · 被引用 1 次
