Diversity over Uniformity: Rethinking Representation in Generated Image Detection
Qinghui He, Haifeng Zhang, Qiao Qin, Bo Liu, Xiuli Bi, Bin Xiao
Abstract
With the rapid advancement of generative models, generated image detection has become an important task in visual forensics. Although existing methods have achieved remarkable progress, they often rely, after training, on only a small subset of highly salient forgery cues, which limits their ability to generalize to unseen generative mechanisms. We argue that reliably generated image detection should not depend on a single decision path but should preserve multiple judgment perspectives, enabling the model to understand the differences between real and generated images from diverse viewpoints. Based on this idea, we propose an anti-feature-collapse learning framework that filters taskirrelevant components and suppresses excessive overlap among different forgery cues in the representation space, preventing discriminative information from collapsing into a few dominant feature directions. This design maintains diverse and complementary evidence within the model, reduces reliance on a small set of salient cues, and enhances robustness under unseen generative settings. Extensive experiments on multiple public benchmarks demonstrate that the proposed method significantly outperforms the stateof-the-art approaches in cross-model scenarios, achieving an accuracy improvement of 5.02% and exhibiting superior generalization and detection reliability. The source code is available at https://github.com/Yanmou-Hui/DoU .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e5d8d64-3702-45df-b15d-01a47351fb77Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- UCF: Uncovering Common Features for Generalizable Deepfake DetectionZhiyuan Yan, Yong Zhang, Yanbo Fan, Baoyuan WuICCV 2023 · 264 citations
- DySy-Det: A Synergistic Framework with Dynamic Reconstruction-Path Consistency for AI-Generated Image DetectionFanli Jin, Feng Lin, Gaojian Wang, Tong Wu et al.AAAI 2026
- Critical Forgetting-Based Multi-Scale Disentanglement for Deepfake DetectionKai Li, Wenqi Ren, Jianshu Li, Wei Wang et al.AAAI 2025 · 3 citations
- CausalCLIP: Causally-Informed Feature Disentanglement and Filtering for Generalizable Detection of Generated ImagesBo Liu, Qiao Qin, Qinghui HeAAAI 2026 · 2 citations
- ResProto-FD: Visual-Language Residual Prototype Sets for Generalized Face Forgery DetectionJiuyao Jing, Yu Zheng, Chunlei PengAAAI 2026
