Lune

CVPR2026Top-tier venue

Red-teaming Retrieval-Augmented Diffusion Models via Poisoning Knowledge Bases

Xinqi Lyu, Yihao Liu, Dong Wang, Bin Xiao

2026Year

Abstract

Retrieval-augmented diffusion models (RAG-DMs) have been increasingly deployed across applications, reflecting a broader trend of adopting retrieval-augmented pipelines in AI agent systems and the emerging OpenClaw framework. Despite the success, their trustworthiness remains underexplored. Existing backdoor attacks focus on either manipulating the generation phase or the retrieval phase under the white-box setting, which suffer from knowledge conflicts between retrieved images and user prompts. To bridge this gap, we propose a novel red-teaming approach JOB, which is the first jointly optimized backdoor attack tailored to black-box RAG-DMs. Specifically, JOB poisons the knowledge base with a small number of target class images and learns a trigger through multi-objective optimization, steering retrieval toward poisoned images and aligning the generated outputs with the target class, while preserving benign performance. Experiments show that JOB effectively attacks black-box RAG-DMs, achieving high success rates and outperforming state-of-the-art baselines.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d1399fe7-042d-4ad5-8acf-ca37159e6085

Builds on21

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines