SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations
Satwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak Damavandi
摘要
Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment. Existing task-oriented dialog datasets aimed towards virtual assistance fall short and do not situate the dialog in the user's multimodal context. To overcome, we present a new dataset for Situated and Interactive Multimodal Conversations, SIMMC 2.0, which includes 11K taskoriented user$assistant dialogs (117K utterances) in the shopping domain, grounded in immersive and photo-realistic scenes. The dialogs are collected using a two-phase pipeline: (1) A novel multimodal dialog simulator generates simulated dialog flows, with an emphasis on diversity and richness of interactions, (2) Manual paraphrasing of the generated utterances to collect diverse referring expressions. We provide an in-depth analysis of the collected dataset, and describe in detail the four main benchmark tasks we propose. Our baseline model, powered by the state-of-theart language model, shows promising results, and highlights new challenges and directions for the community to study 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and GenerationWanwei He, Yinpei Dai, Min Yang, Jian Sun 等SIGIR 2022 · 被引用 41 次
- MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in ConversationsHanlei Zhang, Xin Wang, Hua Xu, Qianrui Zhou 等ICLR 2024 · 被引用 29 次
- End-to-end Task-oriented Dialogue: A Survey of Tasks, Methods, and Future DirectionsLibo Qin, Wenbo Pan, Qiguang Chen, Lizi Liao 等EMNLP 2023 · 被引用 12 次
- Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!Jiwan Chung, Seungwon Lim, Jaehyun Jeon, Seungbeen Lee 等EMNLP 2024 · 被引用 8 次
- PaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional ExpertsYunshui Li, Binyuan Hui, Zhichao Yin, Min Yang 等ACL 2023 · 被引用 6 次
它引用的顶会 Paper2
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta 等AAAI 2020 · 被引用 707 次
- A Simple Language Model for Task-Oriented DialogueEhsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz 等NeurIPS 2020 · 被引用 590 次
相关 Paper
- SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR StreamsTe-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab 等ACL 2023 · 被引用 4 次
- MMConv: An Environment for Multimodal Conversational Search across Multiple DomainsLizi Liao, Le Hong Long, Zheng Zhang, Minlie Huang 等SIGIR 2021 · 被引用 70 次
- Navigating Connected Memories with a Task-oriented Dialog SystemSatwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak DamavandiEMNLP 2022 · 被引用 1 次
- From Natural Alignment to Conditional Controllability in Multimodal DialogueZeyu Jin, Songtao Zhou, Haoyu Wang, Minghao Tian 等ICLR 2026 · 被引用 2 次
- SCREEN: A Benchmark for Situated Conversational RecommendationDongding Lin, Jian Wang, Chak Tou Leong, Wenjie LiACM MM 2024 · 被引用 3 次
