TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection
Zehong Yan, Peng Qi, Wynne Hsu, Mong-Li Lee
Abstract
Multimodal misinformation, encompassing textual, visual, and cross-modal distortions, poses an increasing societal threat that is amplified by generative AI. Existing methods typically focus on a single type of distortion and struggle to generalize to unseen scenarios. In this work, we observe that different distortion types share common reasoning capabilities while also requiring task-specific skills. We hypothesize that joint training across distortion types facilitates knowledge sharing and enhances the model's ability to generalize. To this end, we introduce TRUST-VL, a unified and explainable vision-language model for general multimodal misinformation detection. TRUST-VL incorporates a novel Question-Aware Visual Amplifier module, designed to extract task-specific visual features. To support training, we also construct TRUST-Instruct, a large-scale instruction dataset containing 198K samples featuring structured reasoning chains aligned with human fact-checking workflows. Extensive experiments on both in-domain and zero-shot benchmarks demonstrate that TRUST-VL achieves state-of-the-art performance, while also offering strong generalization and interpretability. * Corresponding author Real Textual Distortion After the massive #HandsOff protest, Trump answered on Fox what the protesters wanted from the President. His answer is exactly why people call him #DementiaDon. Visual Distortion Donald Trump, wearing a confident smile, responded to Fox News in June 2020 regarding the protests following George Floyd's death.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc52db50-e536-4a04-a90f-fb05e01c0b9dCited by top-tier papers1
Ask how each one uses itBuilds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Can LLM-Generated Misinformation Be Detected?Canyu Chen, Kai ShuICLR 2024 · 270 citations
Related papers
- Drifting Away from Truth: GenAI-Driven News Diversity Challenges LVLM-Based Misinformation DetectionFanxiao Li, Jiaying Wu, Tingchao Fu, Yunyun Dong et al.AAAI 2026 · 4 citations
- TRUST-VLM: Thorough Red-Teaming for Uncovering Safety Threats in Vision-Language ModelsKangjie Chen, Muyang Li, Guanlin Li, Shudong Zhang et al.ICML 2025
- VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video MisinformationYongkang Zhang, Dongyu She, Baiyu Ji, Qichuan Geng et al.CVPR 2026
- Multimodal Taylor Series Network for Misinformation DetectionJiahao Sun, Chen Chen, Chunyan Hou, Yike Wu et al.WWW 2025 · 1 citation
- DOUBT: Decoupled Object-level Understanding and Bridging via vMF-based Trustworthiness for Hallucination Detection in MLLMsKaiqi Chen, Yang Qin, Changhao He, Xi Peng et al.ICML 2026
