Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
Zizhao Chen, Ping Wei, Ziyang Ren, Huan Li, Xiangru Yin
摘要
As multimodal misinformation becomes more sophisticated, its detection and grounding are crucial. However, current multimodal verification methods, relying on passive holistic fusion, struggle with sophisticated misinformation. Due to 'feature dilution,' global alignments tend to average out subtle local semantic inconsistencies, effectively masking the very conflicts they are designed to find. We introduce MaLSF (Mask-aware Local Semantic Fusion), a novel framework that shifts the paradigm to active, bidirectional verification, mimicking human cognitive cross-referencing. MaLSF utilizes mask-label pairs as semantic anchors to bridge pixels and words. Its core mechanism features two innovations: 1) a Bidirectional Cross-modal Verification (BCV) module that acts as an interrogator, using parallel query streams (Text-as-Query and Image-as-Query) to explicitly pinpoint conflicts; and 2) a Hierarchical Semantic Aggregation (HSA) module that intelligently aggregates these multi-granularity conflict signals for task-specific reasoning. In addition, to extract fine-grained mask-label pairs, we introduce a set of diverse mask-label pair extraction parsers. MaLSF achieves state-of-the-art performance on both the DGM4 and multimodal fake news detection tasks. Extensive ablation studies and visualization results further verify its effectiveness and interpretability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
- Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual ConceptsYan Zeng, Xinsong Zhang, Hang LiICML 2022 · 被引用 371 次
相关 Paper
- Knowledge-Enhanced Multimodal Fake News Detection: Semantic Visual and Priority FusionQin Zhang, Jiaying Liu, Qian Tao, Zhiwei Guo 等WWW 2026
- CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation DetectionJinjie Shen, Yaxiong Wang, Yujiao Wu, Lechao Cheng 等ICML 2026
- Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media ManipulationYiheng Li, Yang Yang, Zichang Tan, Huan Liu 等CVPR 2025
- ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and GroundingZhenxing Zhang, Yaxiong Wang, Lechao Cheng, Zhun Zhong 等CVPR 2025
- RaCMC: Residual-Aware Compensation Network with Multi-Granularity Constraints for Fake News DetectionXinquan Yu, Ziqi Sheng, Wei Lu, Xiangyang Luo 等AAAI 2025 · 被引用 9 次
