Gesturing Toward Abstraction: Multimodal Convention Formation in Collaborative Physical Tasks
Kiyosu Maeda, William P. McCarthy, Ching-Yi Tsai, Jeffrey Mu, Haoliang Wang, Robert D. Hawkins, Judith E. Fan, Parastoo Abtahi
Abstract
A quintessential feature of human intelligence is the ability to create ad hoc conventions over time to achieve shared goals efficiently. We investigate how communication strategies evolve through repeated collaboration as people coordinate on shared procedural abstractions. To this end, we conducted an online unimodal study (n = 98) using natural language to probe abstraction hierarchies. In a follow-up lab study (n = 40), we examined how multimodal communication (speech and gestures) changed during physical collaboration. Pairs used augmented reality to isolate their partner’s hand and voice; one participant viewed a 3D virtual tower and sent instructions to the other, who built the physical tower. Participants became faster and more accurate by establishing linguistic and gestural abstractions and using cross-modal redundancy to emphasize key changes from previous interactions. Based on these findings, we extend probabilistic models of convention formation to multimodal settings, capturing shifts in modality preferences. Our findings and model provide building blocks for designing convention-aware intelligent agents situated in the physical world.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 461e6bcf-cba6-4a42-a9f0-fbf53a6d50bfBuilds on14
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real WorldXin Wang, Taein Kwon, Mahdi Rad, Bowen Pan et al.ICCV 2023 · 151 citations
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu et al.CHI 2024 · 86 citations
- YouRefIt: Embodied Reference Understanding with Language and GestureYixin Chen, Qing Li, Deqian Kong, Yik Lun Kei et al.ICCV 2021 · 57 citations
- On the Critical Role of Conventions in Adaptive Human-AI CollaborationAndy Shih, Arjun Sawhney, Jovana Kondic, Stefano Ermon et al.ICLR 2021 · 46 citations
Related papers
- Success and Cost Elicit Convention Formation for Efficient CommunicationSaujas Vaduguru, Yilun Hua, Yoav Artzi, Daniel FriedACL 2026 · 3 citations
- Exploring Communication and Collaboration in Distributed AR Escape Rooms: Design Opportunities to Support Social PlayMatthew Bradbury, Matthew Collard, Kieran Gara, Sam Gorman et al.CHI 2026 · 1 citation
- Unlocking Understanding: An Investigation of Multimodal Communication in Virtual Reality CollaborationRyan Khushan Ghamandi, Ravi Kiran Kattoju, Yahya Hmaiti, Mykola Maslych et al.CHI 2024 · 15 citations
- Talk to the Wall: The Role of Speech Interaction in Collaborative Visual AnalyticsGabriela Molina León, Anastasia Bezerianos, Olivier Gladin, Petra IsenbergIEEE VIS 2024 · 9 citations
- Using Virtual Replicas to Improve Mixed Reality Remote CollaborationHuayuan Tian, Gun A. Lee, Huidong Bai, Mark BillinghurstIEEE VR 2023 · 57 citations
