MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
Hanlei Zhang, Xin Wang, Hua Xu, Qianrui Zhou, Kai Gao, Jianhua Su, Jinyue Zhao, Wenrui Li, Yanting Chen
Abstract
Multimodal intent recognition poses significant challenges, requiring the incorporation of non-verbal modalities from real-world contexts to enhance the comprehension of human intentions. However, most existing multimodal intent benchmark datasets are limited in scale and suffer from difficulties in handling outof-scope samples that arise in multi-turn conversational interactions. In this paper, we introduce MIntRec2.0, a large-scale benchmark dataset for multimodal intent recognition in multi-party conversations. It contains 1,245 high-quality dialogues with 15,040 samples, each annotated within a new intent taxonomy of 30 fine-grained classes, across text, video, and audio modalities. In addition to more than 9,300 in-scope samples, it also includes over 5,700 out-of-scope samples appearing in multi-turn contexts, which naturally occur in real-world open scenarios, enhancing its practical applicability. Furthermore, we provide comprehensive information on the speakers in each utterance, enriching its utility for multi-party conversational research. We establish a general framework supporting the organization of single-turn and multi-turn dialogue data, modality feature extraction, multimodal fusion, as well as in-scope classification and out-ofscope detection. Evaluation benchmarks are built using classic multimodal fusion methods, ChatGPT, and human evaluators. While existing methods incorporating nonverbal information yield improvements, effectively leveraging context information and detecting out-of-scope samples remains a substantial challenge. Notably, powerful large language models exhibit a significant performance gap compared to humans, highlighting the limitations of machine learning methods in the advanced cognitive intent understanding task. We believe that MIntRec2.0 will serve as a valuable resource, providing a pioneering foundation for research in human-machine conversational interactions, and significantly facilitating related applications. The full dataset and codes are available for use at https://github.com/thuiar/MIntRec2.0 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ffb7680-ac09-4095-a0e1-7b0c0aa43889Cited by top-tier papers10
- Adaptive Multimodal Fusion: Dynamic Attention Allocation for Intent RecognitionBo Hu, Kai Zhang, Yanghai Zhang, Yuyang YeAAAI 2025 · 6 citations
- Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent DetectionZhanpeng Chen, Zhihong Zhu, Xianwei Zhuang, Zhiqi Huang et al.EMNLP 2024 · 4 citations
- Impact of Stickers on Multimodal Sentiment and Intent in Social Media: A New Task, Dataset and BaselineYuanchen Shi, Fang Kong, Longyin ZhangACM MM 2025 · 3 citations
- Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent RecognitionQianrui Zhou, Hua Xu, Yunjin Gu, Yifan Wang et al.CVPR 2026 · 3 citations
- Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to EmpathyJiahao Huang, Fengyan Lin, Xuechao Yang, Chen Feng et al.CVPR 2026 · 2 citations
Builds on27
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
Related papers
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou et al.ACM MM 2022 · 66 citations
- Text Takes Over: A Study of Modality Bias in Multimodal Intent DetectionAnkan Mullick, Saransh Sharma, Abhik Jana, Pawan GoyalEMNLP 2025
- AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language ModelsZheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun et al.ICML 2025
- Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded DialoguesYoungmin Kim, Jiwan Chung, Jisoo Kim, Sunghyun Lee et al.ACL 2025
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding et al.CVPR 2026
