Audio-Visual Asynchrony Mitigation: Cross-Modal Alignment and Feature Reconstruction for Deepfake Detection
Yan Wang, Qindong Sun, Dongzhu Rong
Abstract
The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has enabled deepfake videos to evolve from unimodal generation to audio-visual forgeries. Existing multimodal deepfake detection methods primarily rely on capturing correlations between audio-visual modalities to improve detection performance. However, in real-world scenarios, network jitter often leads to audio-visual asynchrony, disrupting inter-modal associations and limiting the effectiveness of these methods. To address this issue, we propose a deepfake detection method specifically designed for audio-visual asynchrony scenarios. First, based on the theory of open balls in metric space, we analyze the variation mechanism of joint features in both audio-visual synchrony and asynchrony scenarios, revealing the impact of audio-visual asynchrony on detection performance. Second, we design a multimodal subspace representation module to mitigate inconsistencies in feature distributions and representation heterogeneity between modalities. We then formulate audio-visual feature alignment as an integer linear programming task and employ the Hungarian algorithm to reconstruct missing inter-modal associations. Finally, we introduce a self-supervised masked reconstruction mechanism to reconstruct missing features and construct the joint correlation matrix to measure cross-modal dependencies, enhancing the robustness of detection. Extensive experiments demonstrate that our method outperforms baselines in audio-visual asynchrony scenarios and exhibits robustness against unknown disturbances.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 51708ce3-d1f3-4ef3-aed0-4f44954eafe3Cited by top-tier papers1
Ask how each one uses itRelated papers
- Joint Audio-Visual Deepfake DetectionYipin Zhou, Ser-Nam LimICCV 2021 · 232 citations
- Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and LocalizationKomal Chugh, Parul Gupta, Abhinav Dhall, Ramanathan SubramanianACM MM 2020 · 217 citations
- AVFF: Audio-Visual Feature Fusion for Video Deepfake DetectionTrevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki et al.CVPR 2024 · 51 citations
- Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt LearningHui Miao, Yuanfang Guo, Zeming Liu, Yunhong WangAAAI 2025 · 8 citations
- FRADE: Forgery-aware Audio-distilled Multimodal Learning for Deepfake DetectionFan Nie, Jiangqun Ni, Jian Zhang, Bin Zhang et al.ACM MM 2024 · 17 citations
