Uncertain Multimodal Intention and Emotion Understanding in the Wild
Qu Yang, Qinghongya Shi, Tongxin Wang, Mang Ye
摘要
Understanding intention and emotion from social media poses unique challenges due to the inherent uncertainty in multimodal data, where posts often contain incomplete or missing modalities. While this uncertainty reflects realworld scenarios, it remains underexplored within the computer vision community, particularly in conjunction with the intrinsic relationship between emotion and intention. To address these challenges, we introduce the Multimodal IntentioN and Emotion Understanding in the Wild (MINE) dataset, comprising over 20,000 topic-specific social media posts with natural modality variations across text, image, video, and audio. MINE is distinctively constructed to capture both the uncertain nature of multimodal data and the implicit correlations between intentions and emotions, providing extensive annotations for both aspects. To tackle these scenarios, we propose the Bridging Emotion-Intention via Implicit Label Reasoning (BEAR) framework. BEAR consists of two key components: a BEIFormer that leverages emotion-intention correlations, and a Modality Asynchronous Prompt that handles modality uncertainty. Experiments show that BEAR outperforms existing methods in processing uncertain multimodal data while effectively mining emotion-intention relationships for social media content understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Adaptive Re-calibration Learning for Balanced Multimodal Intention RecognitionQu Yang, Xiyang Li, Fu Lin, Mang YeNeurIPS 2025 · 被引用 2 次
- Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to EmpathyJiahao Huang, Fengyan Lin, Xuechao Yang, Chen Feng 等CVPR 2026 · 被引用 2 次
- API: Adaptive Prototype Imputation for Incomplete Multimodal Sentiment AnalysisXiaotao Wang, Yiyang Fang, Wenke Huang, Bin Yang 等ICML 2026
- Scalable Medical Multimodal Fusion via Symmetric Consistency ModelingXiaowen Sun, Hui Liu, Gongguan Chen, Ning MaoICML 2026
- EMOE: Modality-Specific Enhanced Dynamic Emotion ExpertsYiyang Fang, Wenke Huang, Guancheng Wan, Kehua Su 等CVPR 2025
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 被引用 1,037 次
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh 等ACL 2020 · 被引用 584 次
相关 Paper
- Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal CuesWei Chen, Tongguan Wang, Feiyue Xue, Junkai Li 等WWW 2026
- Impact of Stickers on Multimodal Sentiment and Intent in Social Media: A New Task, Dataset and BaselineYuanchen Shi, Fang Kong, Longyin ZhangACM MM 2025 · 被引用 3 次
- Miko: Multimodal Intention Knowledge Distillation from Large Language Models for Social-Media Commonsense DiscoveryFeihong Lu, Weiqi Wang, Yangyifei Luo, Ziqin Zhu 等ACM MM 2024 · 被引用 11 次
- MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the WildYuanyuan Liu, Wei Dai, Chuanxu Feng, Wenbin Wang 等ACM MM 2022 · 被引用 83 次
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou 等ACM MM 2022 · 被引用 66 次
