Beyond [CLS] Token: Query-Driven Token-Level Forgery Purification for Generalizable Deepfake Detection
Changshuo Wang, Jiangming Wang, Ke-Yue Zhang, Taiping Yao, Shouhong Ding, Shunli Wang, Ran Yi, Lizhuang Ma
Abstract
We investigate state-of-the-art deepfake detectors that leverage ViT-based vision foundation models and discover that the [CLS] token suffers from the Pre-trained Information Bias (PIB), i.e., it tends to mainly focus on global semantics due to the knowledge dominated by pre-trained model parameters, while struggling to emphasize subtle local forgery cues. To overcome this limitation, one potential way is incorporating the token-level features to reform a detection-specific token. To this end, we propose Query-Driven Token-Level Forgery Purification (QTFP 1 ) framework to better capture local forgery traces without losing useful pre-trained prior. Specifically, we introduce randomly initialized, learnable query tokens independent of the backbone and prior knowledge, which effectively aggregate multi-patch evidence into a global token for detection. To make query tokens focus on meaningful regions, we propose a theoretical fake-likelihood contrastive learning loss, which employs a weighting strategy to highlight significant fake regions while diminishing real-like patch impact. Using SNR theory, we verify that the designed weight is both reliable and informative. To further maintain useful authentic information, a real-attention alignment constraint is applied to query tokens. These designs go beyond relying solely on the [CLS] token by jointly reorganizing real and fake information across all tokens, which successfully enhance detector robustness. Extensive experiments on diverse datasets demonstrate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a25404c-40c6-4af3-8ac6-92d3d95ced9eBuilds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
Related papers
- Exploring Unbiased Deepfake Detection via Token-Level Shuffling and MixingXinghe Fu, Zhiyuan Yan, Taiping Yao, Shen Chen et al.AAAI 2025 · 41 citations
- FRADE: Forgery-aware Audio-distilled Multimodal Learning for Deepfake DetectionFan Nie, Jiangqun Ni, Jian Zhang, Bin Zhang et al.ACM MM 2024 · 17 citations
- Exposing the Deception: Uncovering More Forgery Clues for Deepfake DetectionZhongjie Ba, Qingyu Liu, Zhenguang Liu, Shuang Wu et al.AAAI 2024 · 101 citations
- MGQFormer: Mask-Guided Query-Based Transformer for Image Manipulation LocalizationKunlun Zeng, Ri Cheng, Weimin Tan, Bo YanAAAI 2024 · 23 citations
- Locate and Verify: A Two-Stream Network for Improved Deepfake DetectionChao Shuai, Jieming Zhong, Shuang Wu, Feng Lin et al.ACM MM 2023 · 52 citations
