Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis
Changyuan Qiu, Winston Wu, Xinliang Frederick Zhang, Lu Wang
Abstract
Prior work on ideology prediction has largely focused on single modalities, i.e., text or images. In this work, we introduce the task of multimodal ideology prediction, where a model predicts binary or five-point scale ideological leanings, given a text-image pair with political content. We first collect five new large-scale datasets with English documents and images along with their ideological leanings, covering news articles from a wide range of mainstream media in US and social media posts from Reddit and Twitter. We conduct in-depth analyses on news articles and reveal differences in image content and usage across the political spectrum. Furthermore, we perform extensive experiments and ablation studies, demonstrating the effectiveness of targeted pretraining objectives on different model components. Our best-performing model, a late-fusion architecture pretrained with a triplet objective over multimodal content, outperforms the state-of-the-art text-only model by almost 4% and a strong multimodal baseline with no pretraining by over 3%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa2baaa8-c5c4-4ede-8597-236a7be09cccCited by top-tier papers1
Ask how each one uses itBuilds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
Related papers
- We Can Detect Your Bias: Predicting the Political Ideology of News ArticlesRamy Baly, Giovanni Da San Martino, James R. Glass, Preslav NakovEMNLP 2020 · 6 citations
- Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective ModelFuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng et al.ACM MM 2024 · 13 citations
- Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality PerspectiveBing Wang, Ximing Li, Yanjun Wang, Changchun Li et al.AAAI 2026 · 1 citation
- Ideology Prediction from Scarce and Biased Supervision: Learn to Disregard the "What" and Focus on the "How"!Chen Chen, Dylan Walker, Venkatesh SaligramaACL 2023 · 4 citations
- Ideology Takes Multiple Looks: A High-Quality Dataset for Multifaceted Ideology DetectionSongtao Liu, Ziling Luo, Minghua Xu, Lixiao Wei et al.EMNLP 2023 · 1 citation
