MIMIC: AI and AR-enhanced Multi-Modal, Immersive, Relative Instruction Comprehension
Dhanuja Wanniarachchige, Archan Misra
Abstract
We present a multimodal instruction comprehension framework, called MImIC, that utilizes visual sensing (including LIDAR and 2D RGB sensing) & AI spatial reasoning capabilities to support more seamless and immersive interaction between humans and AI-driven situated assistive agents. MImIC's key new capability is to support disambiguation of a wider set of relative spatial references that users naturally employ while issuing spatially-situated instructions. To support enhanced visual grounding via a combination of both fully-qualified and relative attribute references, MImIC uses (a) a fine-tuned transformer-based language translation DNN to accurately convert natural verbal commands into a structured set of machine understandable constraints (BLEU score=92.5), (b) a set of modules that use RGB+LIDAR sensing data to convert any relative attribute preferences to fully-qualified attribute constraints (median height/width estimation errors <=2cm), and (c) an enhanced image segmentation DNN, augmented with gesture+verbal cues, to extract target objects of interest (top-1 accuracy= 85%). To demonstrate the viability and superiority of MImIC, we implement an exemplar AR-augmented, immersive furniture shopping application, called AIRFurn. AIRFurn allows users to browse for, select and overlay furniture items of interest using natural multi-modal and relative cues. experimental studies, using 34 & 11 users over 8 different layouts in a lab setting and 11 users in 6 different real-world home setups, show that AIRFurn offers superior performance, with significantly ( 3x) lower task completion times, much higher task (17%+) accuracy and greater user satisfaction (SUS score= 78.8) compared to baselines where users perform selection using only fully-qualified verbal commands or manipulation of AR interfaces.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2c1b5213-5921-4a2f-b257-ec9813a4fc2fRelated papers
- Towards Building Condition-Based Cross-Modality Intention-Aware Human-AI Cooperation under VR EnvironmentZiyao He, Shiyuan Li, Yunpeng Song, Zhongmin CaiCHI 2024 · 13 citations
- SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR StreamsTe-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab et al.ACL 2023 · 4 citations
- LIRA: Reasoning Reconstruction via Multimodal Large Language ModelsZhen Zhou, Tong Wang, Yunkai Ma, Xiao Tan et al.ICCV 2025 · 1 citation
- Roomify: Spatially-Grounded Style Transformation for Immersive Virtual EnvironmentsXueyang Wang, Qinxuan Cen, Weitao Bi, Yunxiang Ma et al.CHI 2026 · 1 citation
- MIME: Human-Aware 3D Scene GenerationHongwei Yi, Chun-Hao P. Huang, Shashank Tripathi, Lea Hering et al.CVPR 2023
