Identification of Multimodal Stance Towards Frames of Communication
Maxwell A. Weinzierl, Sanda M. Harabagiu
Abstract
Frames of communication are often evoked in multimedia documents. When an author decides to add an image to a text, one or both of the modalities may evoke a communication frame. Moreover, when evoking the frame, the author also conveys her/his stance towards the frame. Until now, determining if the author is in favor of, against or has no stance towards the frame was performed automatically only when processing texts. This is due to the absence of stance annotations on multimedia documents. In this paper we introduce MMVAX-STANCE, a dataset of 11,300 multimedia documents retrieved from social media, which have stance annotations towards 113 different frames of communication. This dataset allowed us to experiment with several models of multimedia stance detection, which revealed important interactions between texts and images in the inference of stance towards communication frames. When inferring the text/image relations, a set of 46,606 synthetic examples of multimodal documents with known stance was generated. This greatly impacted the quality of identifying multimedia stance, yielding an improvement of 20% in F1-score. Component Definition Examples of Frames of Communication Confidence Trust in the security and effectiveness of 2 Pfizer COVID-19 vaccine may cause anaphylaxis vaccinations, the health authorities, and in people with polyethylene glycol (PEG) allergy. the health officials who recommend 2 The Government has provided plenty of safety and develop vaccines. information about the COVID-19 vaccines. Complacency Complacency and laziness to get vaccinated 2 Preference for getting COVID-19 and fighting due to low perceived risk of infections. it off than vaccinating. Constraints Structural or psychological hurdles that 2 It takes courage both to vaccinate against make vaccination difficult or costly. COVID-19 and to refuse the vaccine. Calculation Degree to which personal costs and benefits 2 COVID-19 vaccines protect against the emerging of vaccination are weighted. variants. Collective Willingness to protect others and to 2 Vaccination is key in protecting yourself and others Responsibility eliminate infectious diseases. against COVID-19. Compliance Support for societal monitoring and sanctioning 2 People choosing not to get the COVID-19 vaccine of people who are not vaccinated. should not lose venue access/travel to some countries. Conspiracy Conspiracy thinking and belief in 2 COVID-19 vaccines make you 5G compatible. fake news related to vaccination. 2 The COVID vaccine renders pregnancies risky.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 469a8a6c-47f9-495c-9925-0ac38325a01dCited by top-tier papers4
- Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective ModelFuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng et al.ACM MM 2024 · 13 citations
- Exploring Artificial Image Generation for Stance DetectionZhengkang Zhang, Zhongqing Wang, Guodong ZhouEMNLP 2025
- T-MAD: Target-driven Multimodal Alignment for Stance DetectionZhaoDan Zhang, Jin Zhang, Xueqi Cheng, Hui XuEMNLP 2025
- Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark ApproachXingyu Li, Chen Gong, Guohong FuACL 2025
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon et al.CVPR 2022 · 483 citations
- BridgeTower: Building Bridges between Encoders in Vision-Language Representation LearningXiao Xu, Chenfei Wu, Shachar Rosenman, Vasudev Lal et al.AAAI 2023 · 99 citations
Related papers
- Mitigating World Biases: A Multimodal Multi-View Debiasing Framework for Fake News Video DetectionZhi Zeng, Minnan Luo, Xiangzheng Kong, Huan Liu et al.ACM MM 2024 · 43 citations
- Identifying the Adoption or Rejection of Misinformation Targeting COVID-19 Vaccines in Twitter DiscourseMaxwell A. Weinzierl, Sanda M. HarabagiuWWW 2022 · 19 citations
- Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual MisinformationJeff Da, Maxwell Forbes, Rowan Zellers, Anthony Zheng et al.ACL 2021
- Generating Multimodal Metaphorical Features for Meme UnderstandingBo Xu, Junzhe Zheng, Jiayuan He, Yuxuan Sun et al.ACM MM 2024 · 6 citations
- Stance Detection in COVID-19 TweetsKyle Glandt, Sarthak Khanal, Yingjie Li, Doina Caragea et al.ACL 2021
