OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL
Jinjie Shen, Jing Wu, Yaxiong Wang, Lechao Cheng, Shengeng Tang, Tianrui Hui, Nan Pu, Zhun Zhong
Abstract
Existing forgery detection methods are often limited to uni-modal or bi-modal settings, failing to handle the interleaved text, images, and videos prevalent in real-world misinformation. To bridge this gap, we propose OmniVL-Guard , a unified framework for omni vision-language forgery detection and grounding. In this unified setting, the interplay between diverse modalities and the dual requirements of simultaneous detection and localization pose significant optimization challenges. Through extensive investigations, we identify a critical difficulty bias in this multi-task optimization: the simpler veracity classification task tends to dominate the gradients, leading to suboptimal performance in fine-grained grounding. To address this imbalance, we first develop a Self-Evolving CoT Generation pipeline to synthesize high-quality reasoning paths, effectively overcoming the cold-start challenge. Building upon this, we propose A daptive R eward S caling P olicy O ptimization ( ARSPO ). By dynamically modulating reward scales and task weights, ARSPO ensures a balanced joint optimization that prioritizes challenging grounding objectives. Extensive experiments demonstrate that OmniVL-Guard significantly outperforms state-of-the-art methods and exhibits robust zero-shot generalization across out-of-domain scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efc6137d-0d26-4523-a445-c83ec23873e8Cited by top-tier papers1
Ask how each one uses itBuilds on26
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda et al.EMNLP 2020 · 562 citations
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo et al.NeurIPS 2025 · 528 citations
- DIRE for Diffusion-Generated Image DetectionZhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang et al.ICCV 2023 · 479 citations
- Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationSiwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang et al.NeurIPS 2025 · 82 citations
Related papers
- FactGuard: Agentic Video Misinformation Detection via Reinforcement LearningZehao Li, Hongwei Yu, Hao Jiang, Qiang Sheng et al.ICML 2026 · 3 citations
- OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal GroundingMinghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng et al.CVPR 2026 · 4 citations
- Forensic Prompting with Dual-Action Policy Optimization for Vision-Language Forgery Detection and LocalizationYe Zhu, Ai Zhao, Jinwei WangICML 2026
- BLM-Guard: Explainable Multimodal Ad Moderation with Chain-of-Thought and Policy-Aligned RewardsYiran Yang, Zhaowei Liu, Yuan Yuan, Yukun Song et al.AAAI 2026 · 1 citation
- Anchor-Final Self-Supervision Drives Hallucination-Aware Optimization in Large Vision-Language ModelsJiaxi Liu, Yifeng Yang, Xinbing Wang, Qinying Gu et al.ICML 2026
