Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System Design
Yongquan 'Owen' Hu, Jingyu Tang, Xinya Gong, Zhongyi Zhou, Shuning Zhang, Don Samitha Elvitigala, Florian 'Floyd' Mueller, Wen Hu, Aaron J. Quigley
Abstract
The recent surge in artificial intelligence, particularly in multimodal processing technology, has advanced human-computer interaction, by altering how intelligent systems perceive, understand, and respond to contextual information (i.e., context awareness). Despite such advancements, there is a significant gap in comprehensive reviews examining these advances, especially from a multimodal data perspective, which is crucial for refining system design. This paper addresses a key aspect of this gap by conducting a systematic survey of data modality-driven Vision-based Multimodal Interfaces (VMIs). VMIs are essential for integrating multimodal data, enabling more precise interpretation of user intentions and complex interactions across physical and digital environments. Unlike previous task- or scenario-driven surveys, this study highlights the critical role of the visual modality in processing contextual information and facilitating multimodal interaction. Adopting a design framework moving from the whole to the details and back, it classifies VMIs across dimensions, providing insights for developing effective, context-aware systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6dedaa42-0614-4034-a799-2e596d579cabCited by top-tier papers4
- Exploring Collaboration Patterns and Strategies in Human-AI Co-creation through the Lens of Agency: A Scoping Review of the Top-tier HCI LiteratureShuning Zhang, Hui Wang, Xin YiCSCW 2025 · 31 citations
- The Manipulative Power of Voice Characteristics: Investigating Deceptive Patterns in Mandarin Chinese Female Synthetic SpeechShuning Zhang, Han Chen, Yabo Wang, Yiqun Xu et al.UbiComp 2025 · 4 citations
- Generative Muscle Stimulation: Providing Users with Physical Assistance by Constraining Multimodal-AI with Embodied KnowledgeYun Ho, Romain Nith, Peili Jiang, Steven He et al.CHI 2026 · 1 citation
- SpeechLess: Micro-utterance with Personalized Spatial Memory-aware Assistant in Everyday Augmented RealityYoonsang Kim, Devshree Jadeja, Divyansh Pradhan, Yalong Yang et al.IEEE VR 2026 · 1 citation
Builds on56
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text DataXuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel et al.UbiComp 2024 · 281 citations
- Augmented Reality and Robotics: A Survey and Taxonomy for AR-enhanced Human-Robot Interaction and Robotic InterfacesRyo Suzuki, Adnan Karim, Tian Xia, Hooman Hedayati et al.CHI 2022 · 243 citations
- SemanticAdapt: Optimization-based Adaptation of Mixed Reality Layouts Leveraging Virtual-Physical Semantic ConnectionsYifei Cheng, Yukang Yan, Xin Yi, Yuanchun Shi et al.UIST 2021 · 143 citations
- LLMR: Real-time Prompting of Interactive Worlds using Large Language ModelsFernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski-Fahey et al.CHI 2024 · 124 citations
Related papers
- You're the One Whom I'm Talking To: The Role of Contextual External Human-Machine Interfaces in Multi-Road User Conflict ScenariosYumin Kang, Jeongju Park, Seokhyun Hwang, Minwoo Seong et al.UbiComp 2025 · 7 citations
- Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration: Seeing Eye to EyeZhuyu Teng, Pei Chen, Yichen Cai, Ruoqing Lu et al.CHI 2026 · 2 citations
- Make Interaction Situated: Designing User Acceptable Interaction for Situated Visualization in Public EnvironmentsQian Zhu, Zhuo Wang, Wei Zeng, Wai Tong et al.CHI 2024 · 15 citations
- Gaze and Speech in Multimodal Human-Computer Interaction: A Scoping ReviewAnam Ahmad Khan, Florian Weidner, Jungwoo Rhee, Yasmeen Abdrabou et al.CHI 2026 · 1 citation
- Multimodal Contextualized Semantic Parsing from SpeechJordan Voas, David Harwath, Raymond MooneyACL 2024
