Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model
Fuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng, Genan Dai, Yin Chen, Hu Huang, Bowen Zhang
摘要
Stance detection, which aims to identify public opinion towards specific targets using social media data, is an important yet challenging task. With the proliferation of diverse multimodal social media content including text, and images multimodal stance detection (MSD) has become a crucial research area. However, existing MSD studies have focused on modeling stance within individual text-image pairs, overlooking the multi-party conversational contexts that naturally occur on social media. This limitation stems from a lack of datasets that authentically capture such conversational scenarios, hindering progress in conversational MSD. To address this, we introduce a new multimodal multi-turn conversational stance detection dataset (called MmMtCSD). To derive stances from this challenging dataset, we propose a novel multimodal large language model stance detection framework (MLLM-SD), that learns joint stance representations from textual and visual modalities. Experiments on MmMtCSD show state-of-the-art performance of our proposed MLLM-SD approach for multimodal stance detection. We believe that MmMtCSD will contribute to advancing real-world applications of stance detection research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance DetectionBingbing Wang, Zhengda Jin, Bin Liang, Wenjie Li 等ACL 2026 · 被引用 1 次
- Exploring Artificial Image Generation for Stance DetectionZhengkang Zhang, Zhongqing Wang, Guodong ZhouEMNLP 2025
- T-MAD: Target-driven Multimodal Alignment for Stance DetectionZhaoDan Zhang, Jin Zhang, Xueqi Cheng, Hui XuEMNLP 2025
- MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance DetectionWeihai Lu, Zhejun Zhao, Yanshu Li, Huan HeACL 2026
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningZebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang 等NeurIPS 2024 · 被引用 293 次
- Target-adaptive Graph for Cross-target Stance DetectionBin Liang, Yonghao Fu, Lin Gui, Min Yang 等WWW 2021 · 被引用 93 次
相关 Paper
- A Chinese Multimodal Social Video Dataset for Controversy DetectionTianjiao Xu, Aoxuan Chen, Yuxi Zhao, Jinfei Gao 等ACM MM 2024 · 被引用 4 次
- Tree-of-Counterfactual Prompting for Zero-Shot Stance DetectionMaxwell A. Weinzierl, Sanda M. HarabagiuACL 2024
- A New Direction in Stance Detection: Target-Stance Extraction in the WildYingjie Li, Krishna Garg, Cornelia CarageaACL 2023 · 被引用 7 次
- EZ-STANCE: A Large Dataset for English Zero-Shot Stance DetectionChenye Zhao, Cornelia CarageaACL 2024
- Bilingual Zero-Shot Stance DetectionChenye Zhao, Cornelia CarageaACL 2025 · 被引用 1 次
