MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
Kuo Shi, Jie Lu, Shanshan Ye, Guangquan Zhang, Zhen Fang
Abstract
Recent advances in generative models have highlighted the need for robust detectors capable of distinguishing real images from AI-generated images. While existing methods perform well on known generators, their performance often declines when tested with newly emerging or unseen generative models due to overlapping feature embeddings that hinder accurate cross-generator classification. In this paper, we propose Multimodal Discriminative Representation Learning for Generalizable AI-generated Image Detection (MiraGe), a method designed to learn generator-invariant features. Motivated by theoretical insights on intra-class variation minimization and inter-class separation, MiraGe tightly aligns features within the same class while maximizing separation between classes, enhancing feature discriminability. Moreover, we apply multimodal prompt learning to further refine these principles into CLIP, leveraging text embeddings as semantic anchors for effective discriminative representation learning, thereby improving generalizability. Comprehensive experiments across multiple benchmarks show that MiraGe achieves state-of-the-art performance, maintaining robustness even against unseen generators like Sora.
• Security and privacy → Human and societal aspects of security and privacy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39d5f7bb-f21e-4fef-b5cb-25278d54b42dCited by top-tier papers1
Ask how each one uses itBuilds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
Related papers
- PPM-CLIP: Probabilistic Prompt Modeling for Generalizable AI-Generated Image DetectionXinyuan Wang, Yingxin Lai, Zhiming Luo, Zhihui LiuCVPR 2026 · 1 citation
- MIRAGE: Towards AI-Generated Image Detection in the WildOucheng Huang, Manxi Lin, Jiexiang Tan, Xiaoxiong Du et al.AAAI 2026 · 6 citations
- Breaking the Generator Barrier: Disentangled Representation for Generalizable AI-Text DetectionXiao Pu, Zepeng Cheng, Lin Yuan, Yu Wu et al.ACL 2026 · 1 citation
- Diversity over Uniformity: Rethinking Representation in Generated Image DetectionQinghui He, Haifeng Zhang, Qiao Qin, Bo Liu et al.CVPR 2026
- CausalCLIP: Causally-Informed Feature Disentanglement and Filtering for Generalizable Detection of Generated ImagesBo Liu, Qiao Qin, Qinghui HeAAAI 2026 · 2 citations
