Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System Design
Yongquan 'Owen' Hu, Jingyu Tang, Xinya Gong, Zhongyi Zhou, Shuning Zhang, Don Samitha Elvitigala, Florian 'Floyd' Mueller, Wen Hu, Aaron J. Quigley
摘要
The recent surge in artificial intelligence, particularly in multimodal processing technology, has advanced human-computer interaction, by altering how intelligent systems perceive, understand, and respond to contextual information (i.e., context awareness). Despite such advancements, there is a significant gap in comprehensive reviews examining these advances, especially from a multimodal data perspective, which is crucial for refining system design. This paper addresses a key aspect of this gap by conducting a systematic survey of data modality-driven Vision-based Multimodal Interfaces (VMIs). VMIs are essential for integrating multimodal data, enabling more precise interpretation of user intentions and complex interactions across physical and digital environments. Unlike previous task- or scenario-driven surveys, this study highlights the critical role of the visual modality in processing contextual information and facilitating multimodal interaction. Adopting a design framework moving from the whole to the details and back, it classifies VMIs across dimensions, providing insights for developing effective, context-aware systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Exploring Collaboration Patterns and Strategies in Human-AI Co-creation through the Lens of Agency: A Scoping Review of the Top-tier HCI LiteratureShuning Zhang, Hui Wang, Xin YiCSCW 2025 · 被引用 31 次
- The Manipulative Power of Voice Characteristics: Investigating Deceptive Patterns in Mandarin Chinese Female Synthetic SpeechShuning Zhang, Han Chen, Yabo Wang, Yiqun Xu 等UbiComp 2025 · 被引用 4 次
- Generative Muscle Stimulation: Providing Users with Physical Assistance by Constraining Multimodal-AI with Embodied KnowledgeYun Ho, Romain Nith, Peili Jiang, Steven He 等CHI 2026 · 被引用 1 次
- SpeechLess: Micro-utterance with Personalized Spatial Memory-aware Assistant in Everyday Augmented RealityYoonsang Kim, Devshree Jadeja, Divyansh Pradhan, Yalong Yang 等IEEE VR 2026 · 被引用 1 次
它引用的顶会 Paper56
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo 等NeurIPS 2024 · 被引用 1,004 次
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text DataXuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel 等UbiComp 2024 · 被引用 281 次
- Augmented Reality and Robotics: A Survey and Taxonomy for AR-enhanced Human-Robot Interaction and Robotic InterfacesRyo Suzuki, Adnan Karim, Tian Xia, Hooman Hedayati 等CHI 2022 · 被引用 243 次
- SemanticAdapt: Optimization-based Adaptation of Mixed Reality Layouts Leveraging Virtual-Physical Semantic ConnectionsYifei Cheng, Yukang Yan, Xin Yi, Yuanchun Shi 等UIST 2021 · 被引用 143 次
- LLMR: Real-time Prompting of Interactive Worlds using Large Language ModelsFernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski-Fahey 等CHI 2024 · 被引用 124 次
相关 Paper
- You're the One Whom I'm Talking To: The Role of Contextual External Human-Machine Interfaces in Multi-Road User Conflict ScenariosYumin Kang, Jeongju Park, Seokhyun Hwang, Minwoo Seong 等UbiComp 2025 · 被引用 7 次
- Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration: Seeing Eye to EyeZhuyu Teng, Pei Chen, Yichen Cai, Ruoqing Lu 等CHI 2026 · 被引用 2 次
- Make Interaction Situated: Designing User Acceptable Interaction for Situated Visualization in Public EnvironmentsQian Zhu, Zhuo Wang, Wei Zeng, Wai Tong 等CHI 2024 · 被引用 15 次
- Gaze and Speech in Multimodal Human-Computer Interaction: A Scoping ReviewAnam Ahmad Khan, Florian Weidner, Jungwoo Rhee, Yasmeen Abdrabou 等CHI 2026 · 被引用 1 次
- Multimodal Contextualized Semantic Parsing from SpeechJordan Voas, David Harwath, Raymond MooneyACL 2024
