TIJO: Trigger Inversion with Joint Optimization for Defending Multimodal Backdoored Models
Indranil Sur, Karan Sikka, Matthew Walmer, Kaushik Koneripalli, Anirban Roy, Xiao Lin, Ajay Divakaran, Susmit Jha
Abstract
We present a Multimodal Backdoor Defense technique TIJO (Trigger Inversion using Joint Optimization). Recent work [48] has demonstrated successful backdoor attacks on multimodal models for the Visual Question Answering task. Their dual-key backdoor trigger is split across two modalities (image and text), such that the backdoor is activated if and only if the trigger is present in both modalities. We propose TIJO that defends against dual-key attacks through a joint optimization that reverse-engineers the trigger in both the image and text modalities. This joint optimization is challenging in multimodal models due to the disconnected nature of the visual pipeline which consists of an offline feature extractor, whose output is then fused with the text using a fusion module. The key insight enabling the joint optimization in TIJO is that the trigger inversion needs to be carried out in the object detection box feature space as opposed to the pixel space. We demonstrate the effectiveness of our method on the TrojVQA benchmark, where TIJO improves upon the state-of-the-art unimodal methods from an AUC of 0.6 to 0.92 on multimodal dual-key back-doors. Furthermore, our method also improves upon the unimodal baselines on unimodal backdoors. We present ablation studies and qualitative results to provide insights into our algorithm such as the critical importance of overlaying the inverted feature triggers on all visual features during trigger inversion. The prototype implementation of TIJO is available at https://github.com/SRI-CSL/TIJO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09854bca-5c3f-4da2-abea-e8c05c7109e2Cited by top-tier papers4
- BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language ModelsYi Zeng, Weiyu Sun, Tran Ngoc Huynh, Dawn Song et al.EMNLP 2024 · 10 citations
- Defending against Backdoor Attacks via Module SwitchingWeijun Li, Ansh Arora, Xuanli He, Mark Dras et al.ICLR 2026 · 2 citations
- Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and AlignmentTong Zhang, Kuofeng Gao, Jiawang Bai, Leo Yu Zhang et al.EMNLP 2025 · 1 citation
- Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal TriggersRuofei Wang, Hongzhan Lin, Ziyuan Luo, Ka Chun Cheung et al.AAAI 2025
Builds on17
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.ICLR 2021 · 548 citations
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.NeurIPS 2021 · 503 citations
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation ModelsKangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo et al.ICLR 2022 · 133 citations
Related papers
- Dual-Key Multimodal Backdoors for Visual Question AnsweringMatthew Walmer, Karan Sikka, Indranil Sur, Abhinav Shrivastava et al.CVPR 2022 · 27 citations
- Django: Detecting Trojans in Object Detection Models via Gaussian Focus CalibrationGuangyu Shen, Siyuan Cheng, Guanhong Tao, Kaiyuan Zhang et al.NeurIPS 2023 · 18 citations
- OdScan: Backdoor Scanning for Object Detection ModelsSiyuan Cheng, Guangyu Shen, Guanhong Tao, Kaiyuan Zhang et al.S&P 2024 · 14 citations
- SEER: Backdoor Detection for Vision-Language Models through Searching Target Text and Image Trigger JointlyLiuwan Zhu, Rui Ning, Jiang Li, Chunsheng Xin et al.AAAI 2024 · 11 citations
- MTAttack: Multi-Target Backdoor Attacks Against Large Vision-Language ModelsZihan Wang, Guansong Pang, Wenjun Miao, Jin Zheng et al.AAAI 2026
