CausalCLIP: Causally-Informed Feature Disentanglement and Filtering for Generalizable Detection of Generated Images
Bo Liu, Qiao Qin, Qinghui He
Abstract
The rapid advancement of generative models has increased the demand for generated image detectors capable of generalizing across diverse and evolving generation techniques. However, existing methods, including those leveraging pre-trained vision-language models, often produce highly entangled representations, mixing task-relevant forensic cues (causal features) with spurious or irrelevant patterns (non-causal features), thus limiting generalization. To address this issue, we propose CausalCLIP, a framework that explicitly disentangles causal from non-causal features and employs targeted filtering guided by causal inference principles to retain only the most transferable and discriminative forensic cues. By modeling the generation process with a structural causal model and enforcing statistical independence through Gumbel-Softmax-based feature masking and Hilbert-Schmidt Independence Criterion (HSIC) constraints, CausalCLIP isolates stable causal features robust to distribution shifts. When tested on unseen generative models from different series, CausalCLIP demonstrates strong generalization ability, achieving improvements of 6.83% in accuracy and 4.06% in average precision over state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51d1c72c-dfd7-4369-adf8-853400c3f48aCited by top-tier papers2
- DNA: Uncovering Universal Latent Forgery KnowledgeJingtong Dou, Chuancheng Shi, Anqi Yi, Shiming Guo et al.ICML 2026 · 8 citations
- Fleet: Few Shots Lead Effective AI-generated Image DetectionJiaan Wang, Sirui Liu, Yu Li, Kaiyuan Yang et al.ICML 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
Related papers
- PPM-CLIP: Probabilistic Prompt Modeling for Generalizable AI-Generated Image DetectionXinyuan Wang, Yingxin Lai, Zhiming Luo, Zhihui LiuCVPR 2026 · 1 citation
- DGS-Net: Distillation-Guided Gradient Surgery for CLIP Fine-Tuning in AI-Generated Image DetectionJiazhen Yan, Ziqiang Li, Fan Wang, Boyu Wang et al.ICML 2026 · 1 citation
- Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image DetectionChenming Zhou, Jiaan Wang, Yu Li, Lei Li et al.AAAI 2026 · 1 citation
- Diversity over Uniformity: Rethinking Representation in Generated Image DetectionQinghui He, Haifeng Zhang, Qiao Qin, Bo Liu et al.CVPR 2026
- GM-DF: Generalized Multi-Scenario Deepfake DetectionYingxin Lai, Hongyang Wang, Jing Yang, Xiangui Kang et al.ACM MM 2025 · 10 citations
