BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation
Feiran Li, Qianqian Xu, Shilong Bao, Zhiyong Yang, Xilin Zhao, Xiaochun Cao, Qingming Huang
Abstract
This paper investigates the challenging task of detecting backdoored text-to-image models under black-box settings and introduces a novel detection framework Black-Mirror. Existing approaches typically rely on analyzing image-level similarity, under the assumption that backdoor-triggered generations exhibit strong consistency across samples. However, they struggle to generalize to recently emerging backdoor attacks, where backdoored generations can appear visually diverse. BlackMirror is motivated by an observation: across backdoor attacks, only partial semantic patterns within the generated image are steadily manipulated, while the rest of the content remains diverse or benign. Accordingly, BlackMirror consists of two components: Mir-rorMatch, which aligns visual patterns with the corresponding instructions to detect semantic deviations; and MirrorVerify, which evaluates the stability of these deviations across varied prompts to distinguish true backdoor behavior from benign responses. BlackMirror is a general, training-free framework that can be deployed as a plug-and-play module in Model-as-a-Service (MaaS) applications. Comprehensive experiments demonstrate that BlackMirror achieves accurate detection across a wide range of attacks. Code is available at https://github.com/Ferry-Li/BlackMirror.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7dc077e5-1368-49f3-8ba2-3d033193e99eBuilds on48
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song et al.ICLR 2022 · 2,128 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
Related papers
- Probing Semantic Insensitivity for Inference-Time Backdoor Defense in Multimodal Large Language ModelXuankun Rong, Wenke Huang, Wenzheng Jiang, Yiming Li et al.AAAI 2026
- Towards Human-Imperceptible Backdoor Attacks on Text-to-Image Diffusion ModelsChangkun Wu, Chenghao Chen, Wu kun, Chong Fu et al.CVPR 2026
- PLA: Prompt Learning Attack Against Text-To-Image Generative ModelsXinqi Lyu, Yihao Liu, Yanjie Li, Bin XiaoICCV 2025 · 10 citations
- Black-box Detection of Backdoor Attacks with Limited Information and DataYinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang et al.ICCV 2021 · 128 citations
- Black-box Backdoor Defense via Zero-shot Image PurificationYucheng Shi, Mengnan Du, Xuansheng Wu, Zihan Guan et al.NeurIPS 2023 · 66 citations
