Detecting Compressed AI-Generated Images via Phase Spectrum Robustness
Kai Li, Wenqi Ren, Wei Wang, Xiaochun Cao
Abstract
This paper aims to present a robust AI-generated image detection framework designed to address performance degradation caused by image compression in online social networks. The key challenges are twofold: 1) compression destroys fragile artifacts that are crucial to existing methods, and 2) it introduces new compression artifacts that interfere with detection. Existing methods typically enhance the compression robustness by collecting original-compression pairs and compression labels. However, the collection and annotation process is highly resource-intensive. To address these issues, we propose a Compression-Robust Phase-Harmonized Transformer, motivated by the observation that phase spectrum remains stable under compression. The framework consists of a phase-harmonized cross-modal interaction module that leverages phase spectrum information for feature fusion, enhancing compression robustness, and a multi-domain modulation adapter that further refines fused features while enabling parameter-efficient finetuning. In particular, the framework operates without requiring compression-original data pairs and compression labels. When limited compression labels are available, we introduce a difficulty-aware consistency loss to maximize their utility by prioritizing hard compressed samples during training, further boosting robustness. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches, exhibiting superior robustness against image compression.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
Related papers
- Metric Learning for Anti-Compression Facial Forgery DetectionShenhao Cao, Qin Zou, Xiuqing Mao, Dengpan Ye et al.ACM MM 2021 · 23 citations
- Decision-Driven Orthogonal Learning with Complementary Feature Mining for Robust Synthetic Image DetectionKai Li, Wei Wang, Linchao Zhang, Siying Zhu et al.AAAI 2026
- Dissect and Prune: Enhancing Robustness in AI-Generated Image DetectionDahye Kim, Jaehyun Choi, Hyun Seok Seong, Seongho Kim et al.ICML 2026
- ODDN: Addressing Unpaired Data Challenges in Open-World Deepfake Detection on Online Social NetworksRenshuai Tao, Manyi Le, Chuangchuang Tan, Huan Liu et al.AAAI 2025 · 7 citations
- Beyond Pixels: Mining Compressed Domain Artifacts for Efficient AI-Generated Video DetectionAnran Zhu, Zhengli Shi, Chende Zheng, Chenhao Lin et al.ICML 2026
