MIMIC: AI and AR-enhanced Multi-Modal, Immersive, Relative Instruction Comprehension
Dhanuja Wanniarachchige, Archan Misra
摘要
We present a multimodal instruction comprehension framework, called MImIC, that utilizes visual sensing (including LIDAR and 2D RGB sensing) & AI spatial reasoning capabilities to support more seamless and immersive interaction between humans and AI-driven situated assistive agents. MImIC's key new capability is to support disambiguation of a wider set of relative spatial references that users naturally employ while issuing spatially-situated instructions. To support enhanced visual grounding via a combination of both fully-qualified and relative attribute references, MImIC uses (a) a fine-tuned transformer-based language translation DNN to accurately convert natural verbal commands into a structured set of machine understandable constraints (BLEU score=92.5), (b) a set of modules that use RGB+LIDAR sensing data to convert any relative attribute preferences to fully-qualified attribute constraints (median height/width estimation errors <=2cm), and (c) an enhanced image segmentation DNN, augmented with gesture+verbal cues, to extract target objects of interest (top-1 accuracy= 85%). To demonstrate the viability and superiority of MImIC, we implement an exemplar AR-augmented, immersive furniture shopping application, called AIRFurn. AIRFurn allows users to browse for, select and overlay furniture items of interest using natural multi-modal and relative cues. experimental studies, using 34 & 11 users over 8 different layouts in a lab setting and 11 users in 6 different real-world home setups, show that AIRFurn offers superior performance, with significantly ( 3x) lower task completion times, much higher task (17%+) accuracy and greater user satisfaction (SUS score= 78.8) compared to baselines where users perform selection using only fully-qualified verbal commands or manipulation of AR interfaces.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Towards Building Condition-Based Cross-Modality Intention-Aware Human-AI Cooperation under VR EnvironmentZiyao He, Shiyuan Li, Yunpeng Song, Zhongmin CaiCHI 2024 · 被引用 13 次
- SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR StreamsTe-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab 等ACL 2023 · 被引用 4 次
- LIRA: Reasoning Reconstruction via Multimodal Large Language ModelsZhen Zhou, Tong Wang, Yunkai Ma, Xiao Tan 等ICCV 2025 · 被引用 1 次
- Roomify: Spatially-Grounded Style Transformation for Immersive Virtual EnvironmentsXueyang Wang, Qinxuan Cen, Weitao Bi, Yunxiang Ma 等CHI 2026 · 被引用 1 次
- MIME: Human-Aware 3D Scene GenerationHongwei Yi, Chun-Hao P. Huang, Shashank Tripathi, Lea Hering 等CVPR 2023
