Cross-modal Representation Learning for Diffusion-generated Image Detection
Tao Gong, Dayong Wang, Qi Chu, Bin Liu, Nenghai Yu
Abstract
The astonishing proficiency and unprecedented level of realism of diffusion models in creating and manipulating images have undoubtedly drawn concerns. Many methods have been proposed to detect generated images. Typically, they usually take RGB images as input, and use backbones like ResNet, CLIP visual encoder to extract features. Even though these backbones are capable to detect fake images, they are mainly designed to extract the high-level semantic information, rather than inherently designed for fake image detection. To this end, in this paper, we want to optimize the embedding space tailored for detecting fake images via representation learning. We notice that Neighboring Pixel Relationships (NPR) is capable to capture the intrinsic forgery clues, which means that NPR may be a good input to perform representation learning that aims at learning the embedding space tailored for detecting fake images. Therefore, we leverage features from both RGB modality and NPR modality to perform two proposed representation learning methods, Cross-Modal Contrastive Learning (CMCL) and Cross-Modal Mutual Distillation (CMMD), in order to learn the forgery-aware embedding space. The CMCL boosts the discrimination of features between real and fake images, while the CMMD simultaneously transfers the learned knowledge between two modalities, being able to learn compact features within the intra-class. CMCL and CMMD work collaboratively so that each modality learns a more comprehensive forgery-aware representation to distinguish real and fake images. Extensive experiments on GenImage, DRCT-2M, and Co-Spy-Bench datasets show that our method achieves state-of-the-art results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 179af59e-a0c5-43ab-b8f7-a2af9c4c4505Builds on36
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- RLEG: Vision-Language Representation Learning with Diffusion-based Embedding GenerationLiming Zhao, Kecheng Zheng, Yun Zheng, Deli Zhao et al.ICML 2023 · 11 citations
- Critical Forgetting-Based Multi-Scale Disentanglement for Deepfake DetectionKai Li, Wenqi Ren, Jianshu Li, Wei Wang et al.AAAI 2025 · 3 citations
- Towards Good Generalizations for Diffusion Generated Image Detection Using Multiple Reconstruction Contrastive LearningWanyi Zhuang, Qi Chu, Tao Gong, Changtao Miao et al.ACM MM 2025
- Learning Self-Consistency for Deepfake DetectionTianchen Zhao, Xiang Xu, Mingze Xu, Hui Ding et al.ICCV 2021 · 368 citations
- DRCT: Diffusion Reconstruction Contrastive Training towards Universal Detection of Diffusion Generated ImagesBaoying Chen, Jishen Zeng, Jianquan Yang, Rui YangICML 2024 · 124 citations
