T-MAD: Target-driven Multimodal Alignment for Stance Detection
ZhaoDan Zhang, Jin Zhang, Xueqi Cheng, Hui Xu
摘要
Multimodal Stance Detection (MSD) aims to determine a user's stance -support, oppose, or neutral -toward a target by analyzing multimodal content such as texts and images from social media. Existing MSD methods struggle with generalizing to unseen targets and handling modality inconsistencies. To address these challenges, we propose the Targetdriven Multi-modal Alignment and Dynamic Weighting Model (T-MAD), which combines target-driven multi-modal alignment and dynamic weighting mechanisms to capture targetspecific relationships and balance modality contributions. The model incorporates iterative reasoning to iteratively refine predictions, achieving robust performance in both in-target and zero-shot settings. Experiments on the MMSD and MultiClimate datasets show that T-MAD outperforms state-of-the-art models, with optimal results achieved using RoBERTa, ViT, and an iterative depth of 5. Ablation studies further confirm the importance of multi-modal alignment and dynamic weighting in enhancing model effectiveness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective ModelFuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng 等ACM MM 2024 · 被引用 13 次
- MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance DetectionWeihai Lu, Zhejun Zhao, Yanshu Li, Huan HeACL 2026
- MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance DetectionBingbing Wang, Zhengda Jin, Bin Liang, Wenjie Li 等ACL 2026 · 被引用 1 次
- Incorporating Stock Market Signals for Twitter Stance DetectionCostanza Conforti, Jakob Berndt, Mohammad Taher Pilehvar, Chryssi Giannitsarou 等ACL 2022
- A New Direction in Stance Detection: Target-Stance Extraction in the WildYingjie Li, Krishna Garg, Cornelia CarageaACL 2023 · 被引用 7 次
