Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
Jiaying Wu, Fanxiao Li, Zihang Fu, Min-Yen Kan, Bryan Hooi
Abstract
The impact of multimodal misinformation arises not only from factual inaccuracies but also from the misleading narratives that creators deliberately embed. Interpreting such creator intent is therefore essential for multimodal misinformation detection (MMD) and effective information governance. To this end, we introduce DECEPTIONDECODED, a large-scale benchmark of 12,000 image-caption pairs grounded in trustworthy reference articles, created using an intent-guided simulation framework that models both the desired influence and the execution plan of news creators. The dataset captures both misleading and non-misleading cases, spanning manipulations across visual and textual modalities, and supports three intent-centric tasks: (1) misleading intent detection, (2) misleading source attribution, and (3) creator desire inference. We evaluate 14 state-of-the-art visionlanguage models (VLMs) and find that they struggle with intent reasoning, often relying on shallow cues such as surface-level alignment, stylistic polish, or heuristic authenticity signals. To bridge this, our framework systematically synthesizes data that enables models to learn implication-level intent reasoning. Models trained on DECEPTIONDECODED demonstrate strong transferability to real-world MMD, validating our framework as both a benchmark to diagnose VLM fragility and a data synthesis engine that provides high-quality, intent-focused resources for enhancing robustness in real-world multimodal misinformation governance. 1 Content Warning: this paper contains potentially harmful text and images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d682db5-43da-478e-a746-95b3438d067eCited by top-tier papers6
- IMOL: Incomplete-Modality-Tolerant Learning for Multi-Domain Fake News Video DetectionZhi Zeng, Jiaying Wu, Minnan Luo, Herun Wan et al.ACL 2025 · 17 citations
- What's Left Unsaid? Detecting and Correcting Misleading Omissions in Multimodal News PreviewsFanxiao Li, Jiaying Wu, Tingchao Fu, Dayang Li et al.ACL 2026 · 3 citations
- From Manipulation to Mistrust: Explaining Diverse Micro-Video Misinformation for Robust Debunking in the WildZhi Zeng, Yifei Yang, Jiaying Wu, Xulang Zhang et al.WWW 2026 · 3 citations
- Reasoning About the Unsaid: Misinformation Detection with Omission-Aware Graph InferenceZhengjia Wang, Danding Wang, Qiang Sheng, Jiaying Wu et al.AAAI 2026 · 2 citations
- Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMsTingchao Fu, Wenkai Wang, Fanxiao Li, Huadong Zhang et al.ACL 2026
Builds on16
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- Can LLM-Generated Misinformation Be Detected?Canyu Chen, Kai ShuICLR 2024 · 270 citations
- Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style AttacksJiaying Wu, Jiafeng Guo, Bryan HooiKDD 2024 · 69 citations
- Visual News: Benchmark and Challenges in News Image CaptioningFuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente OrdonezEMNLP 2021 · 67 citations
Related papers
- The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual ContextsYuchen Zhang, Yaxiong Wang, Yujiao Wu, Lianwei Wu et al.CVPR 2026 · 8 citations
- From Detection to Understanding: Multi-Turn Reasoning for Video Misinformation AnalysisZhi Zeng, Jiaying Wu, Minnan Luo, Di Zhang et al.ACL 2026
- Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language ModelsSitong Fang, Shiyi Hou, Kaile Wang, Boyuan Chen et al.ICML 2026
- FakeWorld 1.0: An Omni-modal Benchmark for Fake Media and ContentYifeng Gao, Yifan Ding, Li Wang, Feida Huang et al.ICML 2026
- Probabilistic Concept Graph Reasoning for Multimodal Misinformation DetectionRuichao Yang, Wei Gao, Xiaobin Zhu, Jing Ma et al.CVPR 2026 · 1 citation
