Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual Misinformation
Jeff Da, Maxwell Forbes, Rowan Zellers, Anthony Zheng, Jena D. Hwang, Antoine Bosselut, Yejin Choi
摘要
Understanding manipulated media, from automatically generated 'deepfakes' to manually edited ones, raises novel research challenges. Because the vast majority of edited or manipulated images are benign, such as photoshopped images for visual enhancements, the key challenge is to understand the complex layers of underlying intents of media edits and their implications with respect to disinformation. In this paper, we study Edited Media Understanding Frames, a new conceptual formalism to understand visual media manipulation as structured annotations with respect to the intents, emotional reactions, effects on individuals, and the overall implications of disinformation. We introduce a dataset for our task, EMU, with 56k question-answer pairs written in rich natural language. We evaluate a wide variety of vision-and-language models for our task, and introduce a new model PELICAN, which builds upon recent progress in pretrained multimodal representations. Our model obtains promising results on our dataset, with humans rating its answers as accurate 48.2% of the time. At the same time, there is still much work to be done -and we provide analysis that highlights areas for further progress.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Combating Online Misinformation Videos: Characterization, Detection, and Future DirectionsYuyan Bu, Qiang Sheng, Juan Cao, Peng Qi 等ACM MM 2023 · 被引用 38 次
- Towards Understanding Factual Knowledge of Large Language ModelsXuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo 等ICLR 2024 · 被引用 21 次
- Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language ModelsJiaying Wu, Fanxiao Li, Zihang Fu, Min-Yen Kan 等ICLR 2026 · 被引用 9 次
- "Image, Tell me your story!" Predicting the original meta-context of visual misinformationJonathan Tonglet, Marie-Francine Moens, Iryna GurevychEMNLP 2024 · 被引用 6 次
- Harmfully Manipulated Images Matter in Multimodal Misinformation DetectionBing Wang, Shengsheng Wang, Changchun Li, Renchu Guan 等ACM MM 2024 · 被引用 5 次
它引用的顶会 Paper4
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami 等NeurIPS 2020 · 被引用 1,022 次
- Social Bias Frames: Reasoning about Social and Power Implications of LanguageMaarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky 等ACL 2020 · 被引用 16 次
- Social Chemistry 101: Learning to Reason about Social and Moral NormsMaxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap 等EMNLP 2020 · 被引用 11 次
相关 Paper
- Manipulation Intention Understanding for Zero-Shot Composed Image RetrievalYuanmin Tang, Jing Yu, Keke Gai, Gang Xiong 等AAAI 2026
- From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert UsersShahroz Tariq, Simon S. Woo, Priyanka Singh, Irena Irmalasari 等ACM MM 2025 · 被引用 13 次
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language ModelsGuo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang 等NeurIPS 2025 · 被引用 69 次
- MIPD: Exploring Manipulation and Intention In a Novel Corpus of Polish DisinformationArkadiusz Modzelewski, Giovanni Da San Martino, Pavel Savov, Magdalena Wilczynska 等EMNLP 2024 · 被引用 2 次
- Does the Emotional Understanding of LVLMs Vary Under High-Stress Environments and Across Different Demographic Attributes?Jaewook Lee, Yeajin Jang, Oh-Woog Kwon, Harksoo KimACL 2025
