SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR Streams
Te-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab, Pedro Rodríguez, Babak Damavandi, Nanyun Peng, Seungwhan Moon
Abstract
Building an AI assistant that can seamlessly converse and instruct humans, in a user-centric situated scenario, requires several essential abilities: (1) spatial and temporal understanding of the situated and real-time user scenes, (2) capability of grounding the actively perceived visuals of users to conversation contexts, and (3) conversational reasoning over past utterances to perform just-in-time assistance. However, we currently lack a large-scale benchmark that captures user↔assistant interactions with all of the aforementioned features. To this end, we propose SIMMC-VR, an extension of the SIMMC 2.0 dataset to a video-grounded taskoriented dialog dataset that captures real-world AI-assisted user scenarios in VR. We propose a novel data collection paradigm that involves (1) generating object-centric multimodal dialog flows with egocentric visual streams and visually-grounded templates, and (2) manually paraphrasing the simulated dialogs for naturalness and diversity while preserving multimodal dependencies. To measure meaningful progress in the field, we propose four tasks to address the new challenges in SIMMC-VR, which require complex spatial-temporal dialog reasoning in active egocentric scenes. We benchmark the proposed tasks with strong multimodal models, and highlight the key capabilities that current models lack for future research directions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fda7330-422f-4183-a300-5d4920b1ca97Cited by top-tier papers2
- Proactive Assistant Dialogue Generation from Streaming Egocentric VideosYichi Zhang, Xin Luna Dong, Zhaojiang Lin, Andrea Madotto et al.EMNLP 2025 · 1 citation
- Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic StructuresShun Inadumi, Nobuhiro Ueda, Koichiro YoshinoACL 2025
Builds on4
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta et al.AAAI 2020 · 707 citations
- A Simple Language Model for Task-Oriented DialogueEhsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz et al.NeurIPS 2020 · 590 citations
- TEACh: Task-Driven Embodied Agents That ChatAishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange et al.AAAI 2022 · 251 citations
- SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal ConversationsSatwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak DamavandiEMNLP 2021 · 54 citations
Related papers
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human InteractionsLiang Xu, Chengqun Yang, Zili Lin, Fei Xu et al.ICCV 2025 · 2 citations
- Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video StreamsYun Wang, Junbin Xiao, Han Lyu, Yifan Wang et al.ICML 2026
- RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World UsersSuyu Ye, Haojun Shi, Darren Shih, Hyokun Yun et al.AAAI 2026 · 17 citations
- Can Vision-Language Models Answer Face to Face Questions in the Real-World?Reza Pourreza, Rishit Dagli, Apratim Bhattacharyya, Sunny Panchal et al.ICLR 2026 · 7 citations
- R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual ScenariosLu Zhu, Tiantian Geng, Yangye Chen, Teng Wang et al.AAAI 2026 · 1 citation
