RTUA: Reconstruction-residual Based Targeted and Untargeted Attack Against Text-Image Person Re-Identification
Yubo Wang, Yan Lu, Bin Liu, Xulin Li, Jixiang Niu
Abstract
Text-Image Person Re-Identification (TI-ReID) is widely deployed in intelligent surveillance. Built on deep neural networks and vision-language models, TI-ReID models inherit vulnerabilities to adversarial attacks, posing security risks. Yet its security remains less explored than retrieval accuracy, and its robustness to adversarial attacks is largely unexplored. To fill this gap, we propose Reconstructionresidual based Targeted and Untargeted Attack (R 2 TUA), which generates perturbations from an image and adversarial prompt that mislead TI-ReID into matching the perturbed image with the adversarial identity. To precisely inject identity attributes into perturbations and achieve fine-grained targeted attack, R 2 TUA proposes Transformerbased Gradual Multimodal Fusion (TGMF) that fuses image and adversarial prompt progressively across layers with tunable cross-modal weight. In addition, we propose a fully differentiable Soft Clamp Function (SCF), ensuring perturbations remain inconspicuous while avoiding local gradient vanishing effects that would trap training into suboptimal local minima. To further align perturbed images with the adversarial text descriptions while leading them to mismatch their original descriptions, R 2 TUA employs Push-Pull Losses (PPLs) and matching losses during training. Extensive evaluations across multiple datasets and models demonstrate the superior untargeted attack and targeted attack performance of R 2 TUA. It also exhibits strong adaptability and transferability against black-box models, outperforming all related attacks across multiple tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbe304e8-1474-4553-8c1a-76dc4c22caddBuilds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Prompt-Driven Transferable Adversarial Attack on Person Re-identification with Attribute-Aware Textual InversionYuan Bian, Min Liu, Yunqi Yi, Xueping Wang et al.ICCV 2025 · 3 citations
- Unifying Multi-Modal Uncertainty Modeling and Semantic Alignment for Text-to-Image Person Re-identificationZhiwei Zhao, Bin Liu, Yan Lu, Qi Chu et al.AAAI 2024 · 40 citations
- Feature-Level Adversarial Attacks and Ranking Disruption for Visible-Infrared Person Re-identificationXi Yang, Huanling Liu, De Cheng, Nannan Wang et al.NeurIPS 2024 · 6 citations
- A Unified Understanding of Adversarial Vulnerability Regarding Unimodal Models and Vision-Language Pre-training ModelsHaonan Zheng, Xinyang Deng, Wen Jiang, Wenrui LiACM MM 2024 · 4 citations
- GPT-ReID: Learning Fine-grained Representation with GPT for Text-based Person RetrievalXudong Wang, Lei Tan, Pingyang Dai, Liujuan Cao et al.ACM MM 2025
