Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model
Fuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng, Genan Dai, Yin Chen, Hu Huang, Bowen Zhang
Abstract
Stance detection, which aims to identify public opinion towards specific targets using social media data, is an important yet challenging task. With the proliferation of diverse multimodal social media content including text, and images multimodal stance detection (MSD) has become a crucial research area. However, existing MSD studies have focused on modeling stance within individual text-image pairs, overlooking the multi-party conversational contexts that naturally occur on social media. This limitation stems from a lack of datasets that authentically capture such conversational scenarios, hindering progress in conversational MSD. To address this, we introduce a new multimodal multi-turn conversational stance detection dataset (called MmMtCSD). To derive stances from this challenging dataset, we propose a novel multimodal large language model stance detection framework (MLLM-SD), that learns joint stance representations from textual and visual modalities. Experiments on MmMtCSD show state-of-the-art performance of our proposed MLLM-SD approach for multimodal stance detection. We believe that MmMtCSD will contribute to advancing real-world applications of stance detection research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd95860c-97e1-4b70-98c7-d4c400a74db9Cited by top-tier papers4
- MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance DetectionBingbing Wang, Zhengda Jin, Bin Liang, Wenjie Li et al.ACL 2026 · 1 citation
- Exploring Artificial Image Generation for Stance DetectionZhengkang Zhang, Zhongqing Wang, Guodong ZhouEMNLP 2025
- T-MAD: Target-driven Multimodal Alignment for Stance DetectionZhaoDan Zhang, Jin Zhang, Xueqi Cheng, Hui XuEMNLP 2025
- MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance DetectionWeihai Lu, Zhejun Zhao, Yanshu Li, Huan HeACL 2026
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningZebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang et al.NeurIPS 2024 · 293 citations
- Target-adaptive Graph for Cross-target Stance DetectionBin Liang, Yonghao Fu, Lin Gui, Min Yang et al.WWW 2021 · 93 citations
Related papers
- A Chinese Multimodal Social Video Dataset for Controversy DetectionTianjiao Xu, Aoxuan Chen, Yuxi Zhao, Jinfei Gao et al.ACM MM 2024 · 4 citations
- Tree-of-Counterfactual Prompting for Zero-Shot Stance DetectionMaxwell A. Weinzierl, Sanda M. HarabagiuACL 2024
- A New Direction in Stance Detection: Target-Stance Extraction in the WildYingjie Li, Krishna Garg, Cornelia CarageaACL 2023 · 7 citations
- EZ-STANCE: A Large Dataset for English Zero-Shot Stance DetectionChenye Zhao, Cornelia CarageaACL 2024
- Bilingual Zero-Shot Stance DetectionChenye Zhao, Cornelia CarageaACL 2025 · 1 citation
