ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
Zhenxing Zhang, Yaxiong Wang, Lechao Cheng, Zhun Zhong, Dan Guo, Meng Wang
Abstract
We present ASAP, a new framework for detecting and grounding multi-modal media manipulation (DGM 4 ). Upon thorough examination, we observe that accurate finegrained cross-modal semantic alignment between the image and text is vital for accurately manipulation detection and grounding. While existing DGM 4 methods pay rare attention to the cross-modal alignment, hampering the accuracy of manipulation detecting to step further. To remedy this issue, this work targets to advance the semantic alignment learning to promote this task. Particularly, we utilize the off-the-shelf large models to construct paired image-text pairs, especially for the manipulated instances. Subsequently, a cross-modal alignment learning is performed to enhance the semantic alignment. Besides the explicit auxiliary clues, we further design a Manipulation-Guided Cross Attention (MGCA) to provide implicit guidance for augmenting the manipulation perceiving. With the grounding truth available during training, MGCA encourages the model to concentrate more on manipulated components while downplaying normal ones, enhancing the model's ability to capture manipulations. Extensive experiments are conducted on the DGM 4 dataset, the results demonstrate that our model can surpass the comparison method with a clear margin. Code will be released at https://github.com/CriliasMiller/ASAP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f386360-50c6-4173-81ca-109c7fcb97a9Cited by top-tier papers8
- The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual ContextsYuchen Zhang, Yaxiong Wang, Yujiao Wu, Lianwei Wu et al.CVPR 2026 · 8 citations
- Multi-speaker Attention Alignment for Multimodal Social InteractionLiangyang Ouyang, Yifei Huang, Mingfang Zhang, Caixin Kang et al.CVPR 2026 · 8 citations
- ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake DetectionHao Gu, Jiangyan Yi, Chenglong Wang, Jianhua Tao et al.ACM MM 2025 · 5 citations
- Open-World 3D Scene Graph Generation for Retrieval-Augmented ReasoningFei Yu, Quan Deng, Shengeng Tang, Yuehua Li et al.AAAI 2026 · 2 citations
- Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal ManipulationsJinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu et al.ACM MM 2025 · 2 citations
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media ManipulationYiheng Li, Yang Yang, Zichang Tan, Huan Liu et al.CVPR 2025
- Detecting and Grounding Multi-Modal Media ManipulationRui Shao, Tianxing Wu, Ziwei LiuCVPR 2023
- IDseq: Decoupled and Sequentially Detecting and Grounding Multi-Modal Media ManipulationRunxin Liu, Tian Xie, Jiaming Li, Lingyun Yu et al.AAAI 2025 · 2 citations
- Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media VerificationZizhao Chen, Ping Wei, Ziyang Ren, Huan Li et al.CVPR 2026
- CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation DetectionJinjie Shen, Yaxiong Wang, Yujiao Wu, Lechao Cheng et al.ICML 2026
