Consistent and Invariant Generalization Learning for Short-video Misinformation Detection
Hanghui Guo, Weijie Shi, Mengze Li, Juncheng Li, Hao Chen, Yue Cui, Jiajie Xu, Jia Zhu, Jiawei Shen, Zhangze Chen, Sirui Han
摘要
Short-video misinformation detection has attracted wide attention in the multi-modal domain, aiming to accurately identify the misinformation in the video format accompanied by the corresponding audio. Despite significant advancements, current models in this field, trained on particular domains (source domains), often exhibit unsatisfactory performance on unseen domains (target domains) due to domain gaps. To effectively realize such domain generalization on the short-video misinformation detection task, we propose deep insights into the characteristics of different domains: (1) The detection on various domains may mainly rely on different modalities (i.e., mainly focusing on videos or audios). To enhance domain generalization, it is crucial to achieve optimal model performance on all modalities simultaneously. (2) For some domains focusing on cross-modal joint fraud, a comprehensive analysis relying on cross-modal fusion is necessary. However, domain biases located in each modality (especially in each frame of videos) will be accumulated in this fusion process, which may seriously damage the final identification of misinformation. To address these issues, we propose a new DOmain generalization model via ConsisTency and invariance learning for shORt-video misinformation detection (named DOCTOR), which contains two characteristic modules: (1) We involve the cross-modal feature interpolation to map multiple modalities into a shared space and the interpolation distillation to synchronize multi-modal learning; (2) We design the diffusion model to add noise to retain core features of multi modal and enhance domain invariant features through cross-modal guided denoising. Extensive experiments demonstrate the effectiveness of our proposed DOCTOR model. Our code is publicly available at https://github.com/ghh1125/DOCTOR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Cross-modal Ambiguity Learning for Multimodal Fake News DetectionYixuan Chen, Dongsheng Li, Peng Zhang, Jie Sui 等WWW 2022 · 被引用 325 次
- Embracing Domain Differences in Fake News: Cross-domain Fake News Detection using Multi-modal DataAmila Silva, Ling Luo, Shanika Karunasekera, Christopher LeckieAAAI 2021 · 被引用 170 次
- Cross-modal Contrastive Learning for Multimodal Fake News DetectionLongzheng Wang, Chuang Zhang, Hongbo Xu, Yongxiu Xu 等ACM MM 2023 · 被引用 100 次
- SimMMDG: A Simple and Effective Framework for Multi-modal Domain GeneralizationHao Dong, Ismail Nejjar, Han Sun, Eleni N. Chatzi 等NeurIPS 2023 · 被引用 80 次
- Reinforced Adaptive Knowledge Learning for Multimodal Fake News DetectionLitian Zhang, Xiaoming Zhang, Ziyi Zhou, Feiran Huang 等AAAI 2024 · 被引用 54 次
相关 Paper
- Text-Guided Fine-grained Counterfactual Inference for Short Video Fake News DetectionLinlin Zong, Wenmin Lin, Jiahui Zhou, Xinyue Liu 等AAAI 2025 · 被引用 6 次
- Mitigating World Biases: A Multimodal Multi-View Debiasing Framework for Fake News Video DetectionZhi Zeng, Minnan Luo, Xiangzheng Kong, Huan Liu 等ACM MM 2024 · 被引用 43 次
- Event Consistency-aware Robust Fake News DetectionLiyuan Cao, Zihang Guo, Huaiwen ZhangACM MM 2025
- Navigating the Kaleidoscope of COVID-19 Misinformation Using Deep LearningYuanzhi Chen, Mohammad Rashedul HasanEMNLP 2021 · 被引用 4 次
- Detecting Fake News in Short Videos Through Multi-View AggregationNuo Li, Yuan Xiong, Chengliang Liu, Jie Wen 等AAAI 2026
