Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection
Zhanhe Lei, Zhongyuan Wang, Jikang Cheng, Baojin Huang, Yuhong Yang, Zhen Han, Chao Liang, Dengpan Ye
Abstract
Standard supervised training for deepfake detection treats all samples with uniform importance, which can be suboptimal for learning robust and generalizable features. In this work, we propose a novel Tutor-Student Reinforcement Learning (TSRL) framework to dynamically optimize the training curriculum. Our method models the training process as a Markov Decision Process where a Tutor''agent learns to guide a Student''(the deepfake detector). The Tutor, implemented as a Proximal Policy Optimization (PPO) agent, observes a rich state representation for each training sample, encapsulating not only its visual features but also its historical learning dynamics, such as EMA loss and forgetting counts. Based on this state, the Tutor takes an action by assigning a continuous weight (0-1) to the sample's loss, thereby dynamically re-weighting the training batch. The Tutor is rewarded based on the Student's immediate performance change, specifically rewarding transitions from incorrect to correct predictions. This strategy encourages the Tutor to learn a curriculum that prioritizes high-value samples, such as hard-but-learnable examples, leading to a more efficient and effective training process. We demonstrate that this adaptive curriculum improves the Student's generalization capabilities against unseen manipulation techniques compared to traditional training methods. Code is available at https://github.com/wannac1/TSRL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- Learning Self-Consistency for Deepfake DetectionTianchen Zhao, Xiang Xu, Mingze Xu, Hui Ding et al.ICCV 2021 · 368 citations
- Detecting Deepfakes with Self-Blended ImagesKaede Shiohara, Toshihiko YamasakiCVPR 2022 · 366 citations
- End-to-End Reconstruction-Classification Learning for Face Forgery DetectionJunyi Cao, Chao Ma, Taiping Yao, Shen Chen et al.CVPR 2022 · 327 citations
Related papers
- Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data AugmentationYuxuan Chou, Tao Yu, Wen Huang, Yuheng Zhang et al.AAAI 2026
- Adversarial Policy Training against Deep Reinforcement LearningXian Wu, Wenbo Guo, Hua Wei, Xinyu XingUSENIX Security 2021 · 19 citations
- Proximal Supervised Fine-TuningWenhong Zhu, Ruobing Xie, Rui Wang, Xingwu Sun et al.ICLR 2026 · 13 citations
- Dynamic Mixed-Prototype Model for Incremental Deepfake DetectionJiahe Tian, Cai Yu, Xi Wang, Peng Chen et al.ACM MM 2024 · 12 citations
- Self-supervised Learning of Adversarial Example: Towards Good Generalizations for Deepfake DetectionLiang Chen, Yong Zhang, Yibing Song, Lingqiao Liu et al.CVPR 2022 · 251 citations
