Query-Based Audio-Visual Temporal Forgery Localization with Register-Enhanced Representation Learning
Xiaodong Zhu, Suting Wang, Junqi Yang, Yuhong Yang, Weiping Tu, Zhongyuan Wang
摘要
Temporal forgery in multimedia-where audio or video streams are subtly manipulated-poses critical challenges for content authenticity verification. While video-level detection has advanced, Temporal Forgery Localization (TFL) remains underexplored, often limited by weak audio-visual modeling and reliance on non-learnable post-processing. To address these challenges, we propose RegQAV, a Register-enhanced Query-based Audio-Visual framework for TFL. RegQAV exploits pretrained foundation models to capture fine-grained audio-visual correspondences and learnable registers are introduced to mitigate the model's tendency to overly focus on a limited set of temporal features. A query-based localization strategy enables end-to-end optimization without post-processing. We also introduce a Modality Fusion Adapter (MFA) for effective multi-scale integration of audio-visual data, a Deepfake Queries Generation (DQG) module for efficient query initialization, and a Poisson Count-Based Approach to dynamically predict the number of forgeries. Experiments on LAV-DF and AV-Deepfake1M show that RegQAV achieves state-of-the-art performance with fewer parameters, faster inference, and stronger generalization. This work offers significant potential for real-time deepfake detection and other multimedia verification applications. The code is available at https://github.com/zxd3099/RegQAV.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video RetrievalJun Li, Xuhang Lou, Jinpeng Wang, Yuting Wang 等CVPR 2026 · 被引用 3 次
- Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake LocalizationJiayu Xiong, Jing Wang, Qi Zhang, Wanlong Wang 等CVPR 2026
- DeformTrace: A Deformable State Space Model with Relay Tokens for Temporal Forgery LocalizationXiaodong Zhu, Suting Wang, Yuanming Zheng, Junqi Yang 等AAAI 2026
- GEM-TFL: Bridging Weak and Full Supervision for Forgery Localization through EM-Guided Decomposition and Temporal RefinementXiaodong Zhu, Yuanming Zheng, Suting Wang, Junqi Yang 等CVPR 2026
相关 Paper
- Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and LocalizationKomal Chugh, Parul Gupta, Abhinav Dhall, Ramanathan SubramanianACM MM 2020 · 被引用 217 次
- Structural–Semantic Perception for Diffusion-Guided Temporal Forgery LocalizationLigong Cao, Yeting Guo, Haoang ChiCVPR 2026
- Coarse-to-Fine Proposal Refinement Framework for Audio Temporal Forgery Detection and LocalizationJunyan Wu, Wei Lu, Xiangyang Luo, Rui Yang 等ACM MM 2024 · 被引用 16 次
- UMMAFormer: A Universal Multimodal-adaptive Transformer Framework for Temporal Forgery LocalizationRui Zhang, Hongxia Wang, Mingshan Du, Hanqing Liu 等ACM MM 2023 · 被引用 42 次
- A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery LocalizationWenbo Xu, Junyan Wu, Wei Lu, Xiangyang Luo 等ACM MM 2025 · 被引用 2 次
